Data analysis methods based on causal relationship mapping

By using a data analysis method based on causal relationship graphs, natural language requests are parsed and causal relationship graphs are constructed, which solves the difficulty of traditional methods in identifying changes in key business indicators in dynamic business scenarios, and achieves fast and accurate data analysis and scientific decision support.

CN121858753BActive Publication Date: 2026-05-26GUANGZHOU SMART SOFTWARE CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
GUANGZHOU SMART SOFTWARE CO LTD
Filing Date
2026-03-19
Publication Date
2026-05-26

AI Technical Summary

Technical Problem

Traditional data analysis methods are difficult to adapt to dynamic business scenarios and cannot quickly and accurately identify the core driving factors of changes in key operating indicators. Furthermore, traditional attribution models suffer from lag and poor explanatory power when dealing with high-dimensional sparse data.

Method used

The data analysis method based on causal relationship graphs obtains natural language analysis requests, parses metric fields, dimension fields and analysis constraints, generates query instructions, retrieves target data from the business database, constructs causal relationship graphs, identifies key influencing factors, and calculates change indicators.

Benefits of technology

In a dynamic business environment, it can quickly and accurately identify key influencing factors, provide a scientific basis for decision-making, and improve the company's operational efficiency and market competitiveness.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121858753B_ABST
    Figure CN121858753B_ABST
Patent Text Reader

Abstract

This application relates to a data analysis method based on a causal relationship graph. It acquires and parses a user's natural language analysis request, extracting a first metric field, a first dimension field, indicator change characteristics, and analysis constraints. A query instruction is then generated to retrieve target data from a business database. Subsequently, a causal relationship graph is matched based on the first metric field. Combining the target data and the graph, key influencing factors associated with the indicator change characteristics are determined. This graph, constructed based on business data, can represent the causal relationships between metric fields. Finally, the change indicators of the key influencing factors are calculated, and the query results are obtained based on the key influencing factors and their change indicators. This application achieves accurate data analysis through natural language interaction, quickly locating the core driving factors of indicator changes, improving analysis efficiency and accuracy, and providing a scientific basis for enterprise decision-making.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of data analysis technology, and in particular to a data analysis method based on causal relationship graphs. Background Technology

[0002] With the development of the digital economy and the digital transformation of enterprises, business operations have generated massive amounts of operational data covering multiple dimensions such as sales, marketing, supply chain, customer behavior, and financial performance, providing important data support for refined management and scientific decision-making. However, given the massive scale and complex structure of this data, quickly and accurately identifying the core drivers of changes in key operational indicators such as revenue, gross profit margin, and customer retention rate has become a significant challenge for enterprise management and intelligent decision-making.

[0003] Traditional methods often rely on static reports or simple year-on-year / month-on-month comparisons, which are ill-suited to dynamically changing business environments and fail to capture the interplay of multiple factors affecting indicator changes. Furthermore, traditional attribution models such as linear regression and Shapley value decomposition are inadequate in handling high-dimensional sparse data, capturing interactions between variables, and adapting to real-time data streams. When faced with complex business data, attribution results are often lagging, lack explanatory power, and may even exhibit significant bias, failing to provide users with valuable analytical results.

[0004] Therefore, there is an urgent need for a data analysis method that can adapt to dynamic business scenarios and quickly provide accurate analysis results, so as to better serve the data analysis needs of enterprises and provide strong support for enterprise management and decision-making. Summary of the Invention

[0005] Therefore, the purpose of this application is to provide a data analysis method based on causal relationship graphs, which can quickly provide accurate analysis results in dynamic business scenarios.

[0006] The data analysis method based on causal relationship graphs in this application includes the following steps:

[0007] Obtain the natural language analysis request input by the user; parse the natural language analysis request to obtain several first metric fields, several first dimension fields, indicator change characteristics, and analysis constraints;

[0008] Based on the aforementioned first metric fields, the aforementioned first dimension fields, and the analytical constraints, a query instruction is generated; target data is then retrieved from the business database based on the query instruction.

[0009] Based on the aforementioned first metric fields, a matching causal relationship graph is obtained; based on the target data and the causal relationship graph, key influencing factors associated with the indicator change characteristics are determined; wherein, the causal relationship graph is constructed based on business data in the business database and is used to characterize the causal relationship between various metric fields;

[0010] Based on the target data and the characteristics of the indicator changes, the change indicators corresponding to the key influencing factors are calculated; based on the key influencing factors and the change indicators, the query results corresponding to the natural language analysis request are obtained.

[0011] This application embodiment, after obtaining a user's input natural language analysis request, parses key information such as measurement fields, dimension fields, indicator change characteristics, and analysis constraints in the request, and dynamically generates query instructions based on this. It then retrieves the latest business data from the business database, ensuring that the analysis is always based on the real-time state of the business. Furthermore, based on key information, it can match a causal relationship graph constructed from real-time updated business data in the business database. The causal relationship graph represents the causal relationships between various measurement fields, thus overcoming the difficulties of traditional methods in handling high-dimensional sparse data, capturing interaction effects between variables, and adapting to real-time data streams when combining target data for analysis. During the analysis process, it can quickly and accurately identify key influencing factors closely related to indicator change characteristics, no longer limited by the poor adaptability of static reports and simple year-on-year / month-on-month comparisons to dynamic business environments, and avoiding the problems of lagging attribution results, poor explanatory power, and large biases in traditional attribution models. Finally, based on the target data and indicator change characteristics, it calculates the change indicators of key influencing factors and generates query results to present to the user. The embodiments of this application enable enterprises to timely and accurately identify the core driving factors of changes in key operating indicators in dynamic business scenarios, providing a strong basis for enterprise managers to make scientific and reasonable decisions, effectively improving the enterprise's operational efficiency and market competitiveness, and helping enterprises achieve sustainable development in the digital economy era.

[0012] To better understand and implement this application, the following detailed description is provided in conjunction with the accompanying drawings. Attached Figure Description

[0013] Figure 1 This is a flowchart illustrating the data analysis method based on causal relationship graphs according to an embodiment of this application. Detailed Implementation

[0014] To make the objectives, technical solutions, and advantages of this application clearer, the embodiments of this application will be described in further detail below with reference to the accompanying drawings. Wherein, when the following description relates to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements.

[0015] It should be understood that the embodiments described below do not represent all embodiments consistent with this application. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this application as detailed in the appended claims. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without inventive effort are within the scope of protection of this application.

[0016] The terminology used in this application is for the purpose of describing particular embodiments only and is not intended to be limiting of the application. The singular forms “a,” “the,” and “the” used in this application are also intended to include the plural forms unless the context clearly indicates otherwise. Furthermore, in the description of this application, unless otherwise stated, “a plurality” means two or more. It should also be understood that the term “and / or” as used herein refers to and includes any or all possible combinations of one or more associated listed items, for example, A and / or B, which can represent: A alone, A and B together, and B alone; the character “ / ” generally indicates that the preceding and following objects are in an “or” relationship.

[0017] It should be understood that although the terms first, second, third, etc., may be used in this application to describe various information, this information should not be limited to these terms, and these terms are only used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence, nor should they be construed as indicating or implying relative importance. Those skilled in the art can understand the specific meaning of the above terms in this application according to the specific circumstances. Depending on the context, the word "if" as used in this application can be interpreted as "when," "when," or "in response to determination."

[0018] Please refer to Figure 1 This application provides a data analysis method based on causal relationship graphs, including the following steps:

[0019] S101: Obtain the natural language analysis request input by the user; parse the natural language analysis request to obtain several first metric fields, several first dimension fields, indicator change characteristics, and analysis constraints;

[0020] S102: Generate a query instruction based on the plurality of first metric fields, the plurality of first dimension fields, and the analysis constraints; obtain target data from the business database based on the query instruction;

[0021] S103: Based on the aforementioned first measurement fields, obtain a matching causal relationship graph; based on the target data and the causal relationship graph, determine the key influencing factors associated with the indicator change characteristics; wherein, the causal relationship graph is constructed based on business data in the business database and is used to characterize the causal relationship between various measurement fields;

[0022] S104: Based on the target data and the indicator change characteristics, calculate the change indicators corresponding to the key influencing factors; based on the key influencing factors and the change indicators, obtain the query results corresponding to the natural language analysis request.

[0023] This application embodiment, after obtaining a user's input natural language analysis request, parses key information such as measurement fields, dimension fields, indicator change characteristics, and analysis constraints in the request, and dynamically generates query instructions based on this. It then retrieves the latest business data from the business database, ensuring that the analysis is always based on the real-time state of the business. Furthermore, based on key information, it can match a causal relationship graph constructed from real-time updated business data in the business database. The causal relationship graph represents the causal relationships between various measurement fields, thus overcoming the difficulties of traditional methods in handling high-dimensional sparse data, capturing interaction effects between variables, and adapting to real-time data streams when combining target data for analysis. During the analysis process, it can quickly and accurately identify key influencing factors closely related to indicator change characteristics, no longer limited by the poor adaptability of static reports and simple year-on-year / month-on-month comparisons to dynamic business environments, and avoiding the problems of lagging attribution results, poor explanatory power, and large biases in traditional attribution models. Finally, based on the target data and indicator change characteristics, it calculates the change indicators of key influencing factors and generates query results to present to the user. The embodiments of this application enable enterprises to timely and accurately identify the core driving factors of changes in key operating indicators in dynamic business scenarios, providing a strong basis for enterprise managers to make scientific and reasonable decisions, effectively improving the enterprise's operational efficiency and market competitiveness, and helping enterprises achieve sustainable development in the digital economy era.

[0024] The data analysis method based on causal relationship graphs in this application uses a computer as the execution subject, and the following describes each step in detail.

[0025] For step S101, obtain the natural language analysis request input by the user; parse the natural language analysis request to obtain several first metric fields, several first dimension fields, indicator change characteristics, and analysis constraints.

[0026] Natural language processing (NLP) requests are data analysis needs submitted by users in natural language forms such as everyday spoken language and written sentences. They contain key information such as the metrics to be analyzed, changes, time, and region. These requests can be recognized and parsed by the system without the need for structured commands, triggering subsequent semantic understanding, data querying, and causal attribution analysis processes. For example, "Analyze the main reasons for the decline in sales this quarter."

[0027] Measurement fields are specific fields used in data analysis to measure business metrics, such as sales revenue, profit, and number of customers. They are key numerical indicators that reflect the business situation.

[0028] Dimension fields are used to classify or group business data, such as time dimensions (quarters, months), regional dimensions (cities, provinces), and product dimensions (product categories, product models). Different dimensions allow for more detailed analysis of measurement fields.

[0029] Indicator change characteristics are feature information describing the changes in key business indicators. They refer to the change attributes exhibited by key business indicators within a specified analysis period, mainly including the direction of change (increase, decrease, fluctuation, stagnation), the magnitude of change (numerical value, percentage), and the type of change trend (sudden increase, sharp decrease, continuous change). These characteristics are used to identify the abnormal or changing indicators that users need to analyze. In one embodiment, the indicator change characteristics include at least one of the following: indicator value increase, indicator value decrease, indicator value fluctuation, indicator year-on-year change, and indicator period-on-period change.

[0030] Analysis constraints are a set of conditions that limit the scope of data analysis. They typically consist of dimension fields and their specific values, including boundary information such as time range, region, product type, user group, and channel. These constraints are used to determine the scope of data queries and causal analysis, ensuring that the analysis results accurately correspond to user needs. In one embodiment, the analysis constraints include at least one of the following: analysis time period, analysis geographical scope, analysis object type, and analysis statistical caliber.

[0031] This step obtains the user's input natural language analysis request and uses natural language processing (NLP) techniques to perform semantic parsing on the request. Specifically, the user input text is preprocessed, removing irrelevant content and performing word segmentation and syntactic analysis. Through domain entity recognition, metric fields such as sales revenue and profit (core indicators) and dimension fields such as time, region, and channel are extracted from the statement. Rule-based and semantic understanding are used to identify the characteristics of indicator changes, including the direction and magnitude of changes such as increase, decrease, and fluctuation. Time, region, and range constraints are extracted as analytical conditions, and finally, the above content is structured and output to provide a basis for subsequent queries and attribution analysis.

[0032] In one embodiment, before step S101, which parses the natural language analysis request to obtain several first metric fields, several first dimension fields, indicator change characteristics, and analysis constraints, the method further includes:

[0033] Step S1011: Perform intent recognition on the natural language analysis request to obtain intent recognition results;

[0034] Intent recognition is performed on natural language analysis requests. Pre-trained semantic analysis models (such as BERT) or rule matching engines are used to determine whether the user's request falls under the attribution analysis scenario. For example, if a user enters "Analyze the reasons for the decline in sales in East China in Q4 2025", the attribution analysis intent can be identified through keyword extraction (such as "reason analysis" and "decline") and semantic understanding.

[0035] Step S1012: When the intent recognition result is an attribution analysis intent, the natural language analysis request is parsed to obtain several first metric fields, several first dimension fields, indicator change characteristics, and analysis constraints.

[0036] When the intent identification result is for attribution analysis, specific elements in the request are analyzed in a targeted manner, including the first metric field (such as "sales revenue"), the first dimension field (such as "East China region" and "Q4 2025"), indicator change characteristics (such as "decline for three consecutive months"), and analysis constraints (such as "only considering factors related to promotional activities"). These elements provide accurate input for subsequent causal relationship mapping and key influencing factor calculation.

[0037] This embodiment automatically determines the type of user need through intent recognition, ensuring that the system only initiates subsequent processes in attribution analysis scenarios. This avoids wasting resources on irrelevant requests and improves system processing efficiency. Once the attribution analysis intent is identified, the request elements are analyzed in a targeted manner to ensure that the extracted information is complete and meets business analysis requirements, providing an accurate data foundation for causal relationship mapping and key influencing factor calculation.

[0038] In one embodiment, after step S1011, which involves performing intent recognition on the natural language analysis request and obtaining the intent recognition result, the following steps are included:

[0039] Step S1013: When the intent recognition result is a non-attribution analysis intent, a query instruction is generated based on the natural language analysis request, and corresponding business data is queried from the business database based on the query instruction;

[0040] Step S1014: Based on the business data, obtain the query results corresponding to the natural language analysis request.

[0041] This embodiment achieves automated processing and accurate response to non-attribution analysis requests through a collaborative mechanism of intent recognition and query command generation. Upon recognizing a non-attribution analysis intent, a query command is dynamically generated and the business database is directly retrieved to obtain the corresponding business data. Combined with the technical means of this embodiment, this application can flexibly respond to different types of needs, providing accurate support for both attribution analysis and ordinary data queries, thereby improving the overall data analysis capabilities and decision-making efficiency of enterprises and building an efficient and intelligent data processing system in complex business environments.

[0042] For step S102, a query instruction is generated based on the plurality of first metric fields, the plurality of first dimension fields, and the analysis constraints; target data is obtained from the business database based on the query instruction.

[0043] A query instruction is a command generated based on the first metric field, the first dimension field, and the analysis constraints. It is used to retrieve relevant data from the business database and specifies the specific content and format of the data to be extracted.

[0044] A business database is a database that stores business-related data for an enterprise, including various types of business information such as sales data, customer data, and financial data. It serves as the data source for data analysis. In this embodiment, the business database can be a multi-source heterogeneous data system, including CRM, ERP, advertising platforms, and other multi-source data systems.

[0045] The target data is data related to the user's analysis request, which is obtained from the business database according to the query command. This data will serve as the basis for subsequent analysis.

[0046] In one embodiment, step S102, which generates a query instruction based on the plurality of first metric fields, the plurality of first dimension fields, and the analysis constraints, includes:

[0047] Step S1021: Based on the plurality of first metric fields and the plurality of first dimension fields, obtain a matching business data model; the business data model includes standard metric fields, standard dimension fields, and data source location information corresponding to each of the standard metric fields and the standard dimension fields.

[0048] A business data model is a standardized and structured data model that defines business data. It includes standard metric fields, standard dimension fields, field mapping rules, and corresponding data source location information. It is used to unify the field standards for natural language parsing and data querying, realize the accurate conversion of natural language fields to actual business database fields, and ensure the consistency between data acquisition and causal analysis.

[0049] Standard metric fields and standard dimension fields are business indicators and classification dimensions that have been standardized by the enterprise. For example, the "revenue" of each business line is unified as "operating revenue" as the standard metric field to avoid data confusion caused by differences in field names.

[0050] Data source location information records the specific path and access method of business data storage. For example, sales data for a certain region is stored in table B of database A, ensuring that query commands can accurately locate the target data.

[0051] Step S1022: Based on the business data model, determine the respective first standard metric field and the respective first standard dimension field corresponding to each of the first metric fields and each of the first dimension fields.

[0052] Through the business data model, the first metric field entered by the user is mapped to a standard metric field, and the first dimension field is mapped to a standard dimension field. For example, the user-entered "sales revenue for this quarter" is mapped to the standard metric field "quarterly operating revenue", and "East China region" is mapped to the standard dimension field "East China region", ensuring the standardization and accuracy of subsequent queries.

[0053] Step S1023: Generate a query instruction based on the data source location information corresponding to each of the first standard metric fields, each of the first standard dimension fields, each of the first standard metric fields and the first standard dimension fields, and the analysis constraints.

[0054] By combining the mapped standard metric fields, standard dimension fields, and their corresponding data source location information (such as "quarterly operating revenue" being stored in the "quarterly summary table" of the "sales module" in the database), and overlaying analysis constraints (such as "analyze only Q4 2025 data"), a query command containing details such as specific table names, field names, and time ranges is generated to ensure accurate extraction of the required target data from the business database.

[0055] This embodiment achieves end-to-end optimization from user natural language requirements to accurate data acquisition, significantly improving the accuracy and efficiency of data analysis. The application of the business data model ensures standardized mapping of fields from different business systems, avoiding data confusion caused by differences in field definitions. At the same time, the clear location information of the data source enables query commands to accurately locate target data, avoiding invalid data scanning and interference from redundant information.

[0056] In one embodiment, the business data model further includes query field mapping rules, which are used to indicate the field correspondence between the business data model and the business database.

[0057] Query field mapping rules are a set of pre-defined field correspondence rules in the business data model. They clarify the precise mapping relationship between standard fields and actual fields in the business database, and resolve query ambiguity caused by differences in field naming in different business systems.

[0058] Step S1023, which generates a query instruction based on each of the first standard metric fields, each of the first standard dimension fields, the data source location information corresponding to each of the first standard metric fields and the first standard dimension fields, and the analysis constraints, includes:

[0059] Step S10231: Based on the query field mapping rules, determine the corresponding target metric fields and target dimension fields in the business database for each of the first standard metric fields and each of the first standard dimension fields.

[0060] The target metric field / target dimension field is the field identifier that is actually stored in the business database after being transformed by the query field mapping rules and that directly corresponds to the user's needs.

[0061] This step, based on query field mapping rules, maps the first standard metric field (e.g., "quarterly operating revenue") and the first standard dimension field (e.g., "East China region") to the target metric field (e.g., "qtr_revenue") and the target dimension field (e.g., "region_code") in the business database. For example, when a user queries "2025 Q4 East China sales revenue", the mapping rules identify that "sales revenue" corresponds to the "sales_amount" field in the database, and "East China region" corresponds to the "region_east" field.

[0062] Step S10232: Based on the data source location information corresponding to each of the first standard metric fields and the first standard dimension fields, determine the data source location information corresponding to each of the target metric fields and the target dimension fields.

[0063] By combining the data source location information (such as the "sales_amount" field being stored in the "amount" column of the "order_details" table, and the "region_east" field being stored in the "code" column of the "region_mapping" table), the specific storage path of the target field in the database can be determined, avoiding performance loss or data omission caused by cross-table queries.

[0064] Step S10233: Generate a query instruction based on the data source location information corresponding to each of the target metric fields, each of the target dimension fields, each of the target metric fields and the target dimension fields, and the analysis constraints.

[0065] Integrate target fields, data source location information, and analysis constraints (such as the time range "2025-10-01 to 2025-12-31") to generate structured query commands containing specific table names, field names, and filter conditions. For example, generate the SQL command "SELECT SUM(amount) AS qtr_revenue FROM order_details JOIN region_mapping ON region=code WHERE date BETWEEN '2025-10-01' AND '2025-12-31' AND region='east'", accurately locating and aggregating target data.

[0066] This embodiment constructs a standardized channel from user natural language requirements to precise database queries by collaboratively applying query field mapping rules and data source location information. In dynamic business scenarios, this design ensures zero-error data acquisition—avoiding data confusion caused by differences in field naming and reducing invalid data scanning through clear data source location, thus significantly improving query efficiency.

[0067] For step S103, based on the plurality of first measurement fields, a matching causal relationship graph is obtained; based on the target data and the causal relationship graph, key influencing factors associated with the indicator change characteristics are determined; wherein, the causal relationship graph is constructed based on business data in the business database and is used to characterize the causal relationship between various measurement fields.

[0068] A causal relationship graph is a graphical model built upon business data in a business database to represent the causal relationships between various metric fields. Specifically, a causal relationship graph is a knowledge graph constructed with business metric fields as nodes and causal influences between fields as directed edges. It clearly shows the direct or indirect causal relationships and transmission directions between various metric fields, representing the connections and causal directions between different fields through nodes and edges, helping to understand the reasons for changes in business metrics.

[0069] Key influencing factors are business factors that, in a causal relationship graph, have a close causal relationship with the characteristics of indicator changes and significantly influence the changes in business indicators. In one embodiment, key influencing factors are those that have a significant impact on the first measurement field and whose direction of change is consistent with the characteristics of indicator changes. Influencing factors include measurement fields and / or dimension fields.

[0070] This step, based on several extracted primary metric fields, retrieves matching causal relationship graphs from a pre-built causal relationship graph library. These graphs are constructed from business data stored in the business database, having already deeply mined and analyzed the causal relationships within the business data, reflecting the inherent connections between various metric fields. Combining the target data with the causal relationship graph, key influencing factors directly related to the indicator's changing characteristics and capable of explaining these changes are selected. For example, when analyzing the reasons for a decline in sales, the causal relationship graph and target data may reveal that reduced sales volume and lower product prices are key influencing factors. Specifically, based on the target data, actual data for each metric field and dimension field is extracted. Combining this with the causal relationships, transmission directions, and correlation strengths of each node (metric field) in the causal relationship graph, feasible methods such as causal effect analysis, correlation calculation, time-series trend matching, and causal path tracing are used to select factors that significantly influence the indicator's changing characteristics. Supplementary methods such as anomaly detection and weighted ranking can also be used to ultimately determine the key influencing factors directly related to the indicator's changing characteristics and capable of explaining these changes, providing a basis for subsequent calculations of the changing indicators.

[0071] In one embodiment, step S103, which involves obtaining a matching causal relationship graph based on the plurality of first metric fields, includes:

[0072] Step S1031: Obtain a number of preset causal relationship graphs and a set of standard measurement fields corresponding to each of the causal relationship graphs.

[0073] The pre-set causal relationship graphs include causal relationship graphs for multiple scenarios pre-built based on the enterprise's historical business data. Each graph corresponds to a specific business scenario, such as quarterly sales analysis scenario and customer retention attribution scenario. The graphs include validated causal relationship paths between measurement fields, such as a complete causal chain of "promotional activity investment → customer visits → order conversion rate → quarterly operating revenue".

[0074] The standard measurement field set is a standardized indicator system associated with each causal relationship graph. For example, the set corresponding to the "Quarterly Sales Analysis" graph is {Quarterly Revenue, Promotional Activity Investment, Customer Visits, Order Conversion Rate}, ensuring a high degree of alignment between the graph and the business scenario.

[0075] Step S1032: Match the plurality of first metric fields with each of the standard metric field sets to determine the standard metric field set that matches the first metric fields.

[0076] The first metric field obtained by the user, such as "quarterly operating revenue" or "promotional activity investment," is matched with the standard metric field set of each graph. A semantic similarity algorithm (such as the BERT model) is used to calculate the field matching degree. When the matching degree exceeds a threshold (such as 90%), it is determined that the standard metric field set of that graph is highly relevant to the user's needs.

[0077] Step S1033: The causal relationship graph corresponding to the set of matching standard metric fields is determined as the matching causal relationship graph.

[0078] The corresponding causal relationship graph is selected based on the matching results. For example, if the user's first metric field has the highest matching degree with the standard metric field set of the "Market Activity Effect Attribution Graph", then this graph is selected for subsequent analysis to ensure that the analysis model accurately corresponds to the business problem.

[0079] This embodiment features a pre-defined causal relationship graph library covering all business scenarios of an enterprise. Each graph has been verified using historical data to ensure the reliability of causal relationships. Through semantic matching of a set of standard metric fields, the most suitable causal relationship graph can be automatically selected, avoiding the subjectivity and lag of manual graph selection. This intelligent graph matching mechanism enables enterprises to quickly identify the core driving factors of changes in key operating indicators in a complex and ever-changing market environment, providing managers with a scientific basis for decision-making, and ultimately achieving improved operational efficiency and enhanced market competitiveness.

[0080] In one embodiment, the causal relationship graph uses standard metric fields as nodes, the causal directions between standard metric fields as directed edges, and the standard metric fields are bound to corresponding standard dimension fields.

[0081] The preset causal relationship maps mentioned in step S1031 are obtained through the following steps:

[0082] Step S10311: Obtain each second metric field, each second dimension field, and the data source location information corresponding to each second metric field and the second dimension field of the business data in the business database.

[0083] The second metric field / second dimension field is the original business indicator and classification dimension stored in the business database, such as the "sales_amount" (sales revenue) and "region_code" (region code) fields in the database table, which have not undergone standardization processing.

[0084] This step extracts raw field information from the business database, including second measure fields (such as "sales_amount" and "discount_rate"), second dimension fields (such as "region_code" and "date"), and their data source locations (such as the "amount" column in the "sales_data" table).

[0085] Step S10312: Obtain the corresponding second standard metric field and second standard dimension field for each of the second metric fields and the second dimension fields.

[0086] The second standard metric field / second standard dimension field is a field that has been standardized by the business data model. For example, “sales_amount” is mapped to the standard metric field “quarterly operating revenue”, and “region_code” is mapped to the standard dimension field “East China region”, to ensure a consistent expression of fields from different data sources.

[0087] This step, based on the business data model, maps the second metric field and the second dimension field to the second standard metric field and the second standard dimension field, respectively. For example, "sales_amount" is mapped to "quarterly operating revenue", and "region_code" is mapped to "East China region", ensuring consistency between field naming and business semantics.

[0088] Step S10313: Obtain causal definition data, which is used to indicate the causal relationship between each of the second standard measurement fields.

[0089] Causal definition data is a predefined set of rules for causal relationships between fields, including business-verified causal chains such as "promotional activity investment → customer visits → order conversion rate → quarterly operating revenue", which clearly defines the driving relationship and direction between indicators.

[0090] This step loads the causal definition data and clarifies the causal relationships between the second standard metric fields. For example, it defines rules such as "Promotional Activity Investment" → "Customer Visits" (with a positive causal direction) and "Customer Visits" → "Order Conversion Rate" to form a causal relationship rule base.

[0091] Step S10314: Based on each of the second standard metric fields, each of the second standard dimension fields, and the causal definition data, at least one causal relationship graph is constructed.

[0092] Using the second standard metric field as nodes and the causal direction in the causal definition data as directed edges, a causal relationship graph is constructed by combining the second standard dimension fields (such as "East China Region" and "Q4 2025"). For example, in the "Quarterly Sales Analysis" graph, the node "Promotional Activity Investment (East China Region)" points to "Customer Visits (East China Region)" via directed edges, then to "Order Conversion Rate," and finally to "Quarterly Revenue," forming a complete causal chain.

[0093] This embodiment eliminates the semantic differences in multi-source heterogeneous data by mapping the second metric field and second dimension field in the business database to standardized second standard metric fields and second standard dimension fields. This ensures the cross-system consistency of the causal relationship graph, enabling indicator analysis in different business scenarios to be conducted based on a unified semantic framework. Combined with the explicit encoding of causal directions between standard fields in the causal definition data, business experience and data patterns are transformed into rules that can be directly understood by machines. This allows for the structured storage of the driving relationships of complex business logic, avoiding causal misjudgments caused by insufficient experience or subjective bias in manual analysis. Simultaneously, by binding standard dimension fields to nodes in the causal relationship graph, it supports refined causal analysis with multi-dimensional cross-cutting, accurately locating local driving factors of indicator changes in specific business scenarios, effectively solving the problem of key information loss due to data aggregation in global analysis. Furthermore, this embodiment's solution has dynamic expansion capabilities. When new business scenarios or data dimensions are added, new graphs can be quickly constructed by adding standard fields, defining causal relationships, and binding dimensions, without reconstructing the entire analysis system, thus adapting to the needs of rapid business iteration. Ultimately, this embodiment enables enterprises to achieve rapid attribution and scientific decision-making for business problems based on standardized causal graphs in complex and ever-changing market environments, thereby improving operational efficiency and market competitiveness.

[0094] In this embodiment of the application, the corresponding analysis scenario and analysis rules are determined based on the number of measurement fields in the target data and the causal relationship between the measurement fields; based on the target data and the causal relationship map, candidate key influencing factors are obtained through analysis using the corresponding analysis rules; and candidate key influencing factors whose data change trends match the change trends corresponding to the indicator change characteristics are determined as key influencing factors.

[0095] In one embodiment, step S103, which involves determining the key influencing factors associated with the indicator variation characteristics based on the target data and the causal relationship graph, includes:

[0096] Step S10300: When there is only one second metric field in the target data, obtain the associated standard dimension fields from the causal relationship graph based on the second metric field; extract the data of the second metric field and the dimension data of each associated standard dimension field from the target data.

[0097] When the target data contains only a single second metric field (such as "sales revenue"), the standard dimension fields associated with this field (such as "East China Region", "North China Region", "Q1 2025") are first located from the pre-defined causal relationship graph. Then, the specific value of this second metric field (such as "sales revenue in Q1 2025 was 5 million yuan") and the dimensional data of its associated standard dimension fields (such as "sales revenue in East China accounted for 30%) are extracted from the target data, providing a data foundation for subsequent correlation analysis.

[0098] Step S10301: Based on the data of the second metric field and the dimensional data of each of the associated standard dimension fields, calculate the correlation coefficient between the second metric field and each of the associated standard dimension fields.

[0099] Based on the extracted data from the second metric field and the standard dimension field, the correlation coefficient between them is calculated. For example, if the correlation coefficient between "sales revenue" and "sales revenue share in East China" is 0.85, it indicates a strong positive correlation; if the correlation coefficient with "promotional activity investment" is -0.72, it indicates a strong negative correlation. This step provides an objective basis for screening key influencing factors by quantifying the strength of the association between indicators and dimensions.

[0100] Step S10302: The standard dimension fields whose correlation coefficients meet the preset first correlation conditions are determined as candidate key influencing factors; when the data of any candidate key influencing factor meets the change trend corresponding to the indicator change characteristics, the candidate key influencing factor is determined as a key influencing factor.

[0101] The indicator change characteristics are the patterns of change in business indicators that users pay attention to, such as "sales have declined for three consecutive quarters" or "user activity increases significantly on weekends", which are used to screen key influencing factors that are consistent with the direction of business issues.

[0102] This step filters standard dimension fields with correlation coefficients whose absolute values ​​are greater than or equal to a first preset threshold (e.g., 0.7) as candidate key influencing factors, such as "sales share in East China" and "investment in promotional activities." Then, it checks whether the data of the candidate key influencing factors matches the changing trend of the indicator's characteristics: if the indicator change characteristic that the user is concerned about is "declining sales," then only candidate factors with a downward trend in data (e.g., "reduced investment in promotional activities") are retained, and these are ultimately determined as key influencing factors. This step ensures that the analysis results are consistent with the direction of the business problem through trend matching.

[0103] This embodiment achieves precise location and scientific attribution of the causes of changes in business indicators through dynamic interaction between target data and causal relationship graphs. First, by utilizing the correlation between standard dimension fields and second metric fields in the causal relationship graph, a multi-dimensional impact factor analysis framework is constructed, avoiding the analytical bias caused by missing dimensions in traditional methods. The correlation strength between indicators and dimensions is quantified by calculating correlation coefficients, and candidate factors are screened using preset thresholds, ensuring mathematical rigor in the identification process of key impact factors and significantly improving the credibility of the analysis results. Furthermore, by matching the data change trends of candidate factors with the change characteristics of indicators of interest to users, interfering factors unrelated to business issues can be automatically filtered out. For example, when analyzing "declining sales," dimensions with an upward trend (such as "increased proportion of new users") are excluded, ensuring that the final identified key impact factors are highly consistent with the business scenario. This solution does not require manual pre-setting of complex rules and can adapt to the dynamic analysis needs of different business scenarios. For example, switching quickly from "quarterly sales analysis" to "user retention attribution" only requires adjusting the matching logic between the target data and the causal relationship graph to achieve automatic reconstruction of the analysis model. Ultimately, based on the technical solution of this embodiment, enterprises can quickly locate the core driving factors of indicator changes in complex business environments, provide data support for adjusting operational strategies, and thus improve decision-making efficiency and business response speed.

[0104] In one embodiment, step S103, which involves determining the key influencing factors associated with the indicator variation characteristics based on the target data and the causal relationship graph, includes:

[0105] Step S10303: When there are multiple third-measure fields in the target data, and it is determined based on the causal relationship graph that there is a direct or indirect causal relationship between the multiple third-measure fields, the causal path and causal direction between the multiple third-measure fields are obtained from the causal relationship graph.

[0106] Causal paths and causal directions are causal chains in a causal relationship graph, consisting of nodes (standard metric fields) and directed edges. For example, the complete path and positive driving direction of "promotional investment → customer visits → order conversion rate → sales revenue".

[0107] This step extracts the complete causal path (e.g., A→B→C) and causal direction (positive / negative) from the graph when the target data contains multiple third-party metrics and the causal graph shows direct / indirect causal relationships. For example, in the "quarterly sales analysis" scenario, it identifies the causal chain of "promotional investment → customer visits → order conversion rate → sales revenue".

[0108] Step S10304: Based on the multiple third metric fields, extract the corresponding multiple sets of time series data from the target data.

[0109] Extract time-series data corresponding to multiple third-dimensional measurement fields from the target data (such as promotional investment, customer visits, order conversion rate, and sales data for each quarter from Q1 to Q4 of 2025) to provide time-dimensional data support for causal effect analysis.

[0110] Step S10305: Based on the causal path and causal direction, perform causal effect analysis on the multiple sets of time series data to obtain the causal effect intensity corresponding to each of the third metric fields.

[0111] Causal effect analysis is based on time-series data of multiple indicators aligned with the time dimension. It is based on the structural causal model (SCM) and combined with Do-calculus or causal discovery algorithms (PC algorithm, LiNGAM) to quantify the causal effect strength of each node measurement field in the causal path on the change of the final indicator, and identify the contribution of each causal path (e.g., quantifying the causal effect strength of "promotional investment" on "sales volume" and the contribution of the corresponding path).

[0112] The strength of causal effect measures the degree of influence of each node in the causal path on the change of the target indicator. The larger the value, the more significant the driving effect, and it is used to screen key influencing factors.

[0113] This step, based on causal paths and directions, uses causal effect analysis to calculate the causal effect strength of each third metric field. For example, analyzing the pulling effect of "a 10% increase in promotional investment" on "sales revenue" quantifies its contribution as 8%.

[0114] Step S10306: The third measurement field whose causal effect intensity meets the preset intensity condition is determined as a candidate key influence factor; when the data of any candidate key influence factor meets the change trend corresponding to the indicator change characteristics, the candidate key influence factor is determined as a key influence factor.

[0115] This step filters third-order measurement fields with causal effect strength greater than or equal to a preset strength threshold as candidate key influencing factors, and verifies whether their data trends match by combining indicator change characteristics (such as a decrease in sales), and finally determines key influencing factors that conform to the business scenario.

[0116] This embodiment achieves dynamic tracking and visual attribution of key influencing factors in complex business scenarios through causal path analysis and time-series data effect quantification in multi-indicator scenarios. First, utilizing the path structure of a causal relationship graph, the interconnected effects of multiple third-party measurement fields are transformed into an interactive causal network diagram, enabling business personnel to intuitively understand the driving relationship chain between indicators and avoiding cognitive biases caused by data fragmentation in traditional table analysis. Through causal effect analysis of time-series data, the dynamic contribution of each factor to the change of the target indicator can be quantified. Finally, combined with a trend matching mechanism based on indicator change characteristics, interfering factors irrelevant to the business problem can be automatically filtered out, retaining key influencing factors and ensuring that the analysis results are highly consistent with the business scenario.

[0117] In one embodiment, step S103, which involves determining the key influencing factors associated with the indicator variation characteristics based on the target data and the causal relationship graph, includes:

[0118] Step S10307: When there are multiple fourth metric fields in the target data, and it is determined based on the causal relationship graph that there is no direct or indirect causal relationship between the fourth metric fields, for each fourth metric field, the associated standard dimension field is obtained from the causal relationship graph, and the data of the fourth metric field and the dimension data of each associated standard dimension field are extracted from the target data.

[0119] When the target data contains multiple fourth-order measurement fields and the causal relationship graph confirms that there is no causal relationship between them, for each fourth-order measurement field (such as "product inventory"), locate its associated standard dimension field (such as "inventory percentage in East China" or "production batch number") from the graph. Then, extract the specific value of the fourth-order measurement field and the dimensional data of its associated standard dimension field from the target data to provide a data foundation for subsequent analysis.

[0120] Step S10308: Based on the data of the fourth metric field and the dimensional data of each of the associated standard dimension fields, calculate the correlation coefficient between the fourth metric field and each of the associated standard dimension fields.

[0121] Based on the extracted fourth metric field data and the standard dimension field data, the correlation coefficient between the two is calculated. For example, if the correlation coefficient between "product inventory" and "inventory ratio in East China" is 0.92, it indicates a strong positive correlation; if the correlation coefficient with "production batch number" is -0.15, it indicates a weak correlation. This step provides an objective basis for screening key influencing factors by quantifying the strength of the association between indicators and dimensions.

[0122] Step S10309: The standard dimension fields whose correlation coefficients satisfy the preset second correlation conditions are determined as candidate key influencing factors; when the data of any candidate key influencing factor satisfies the change trend corresponding to the indicator change characteristics, the candidate key influencing factor is determined as a key influencing factor.

[0123] Standard dimension fields with absolute correlation coefficients greater than or equal to a second preset threshold are selected as candidate key influencing factors, such as "East China region inventory ratio" and "supplier delivery cycle". Then, the data of the candidate factors is checked to see if they match the changing trend of the indicator's characteristics: if the indicator change characteristic that the user is concerned with is "decreasing product inventory", only candidate factors with a downward trend (such as "extended supplier delivery cycle") are retained, and these are ultimately determined as key influencing factors. This step ensures that the analysis results are consistent with the direction of the business problem through trend matching.

[0124] This embodiment addresses multi-indicator scenarios with no causal relationship. By independently analyzing the correlation between each fourth metric field and its associated dimension data, it achieves accurate positioning and scientific attribution of key influencing factors in complex business environments. First, it utilizes the dimensional correlation information of the causal relationship graph to construct an independent analysis framework covering multiple indicators, avoiding analytical biases caused by mutual interference between indicators in traditional methods. The correlation coefficient is used to quantify the strength of the association between indicators and dimensions, and candidate factors are screened using preset thresholds, ensuring mathematical rigor in the identification process of key influencing factors and significantly improving the credibility of the analysis results. Furthermore, by matching the data change trends of candidate factors with the change characteristics of indicators of interest to the user, it can automatically filter out interfering factors unrelated to the business problem. For example, when analyzing "decreasing product inventory," dimensions with an upward trend (such as "reduced return rate") are excluded, ensuring that the final identified key influencing factors are highly consistent with the business scenario. This solution does not require manual pre-setting of complex rules and can adapt to the dynamic analysis needs of different business scenarios. For example, it can quickly switch from "inventory management analysis" to "human resource efficiency assessment," requiring only adjustments to the matching logic between the target data and the causal relationship graph to automatically reconstruct the analysis model. Ultimately, based on the technical means of this embodiment, enterprises can quickly locate the core driving factors of changes in indicators without causal relationship in complex business environments, provide data support for adjusting operational strategies, and thus improve decision-making efficiency and business response speed.

[0125] For step S104, based on the target data and the indicator change characteristics, calculate the change index corresponding to the key influencing factor; and obtain the query result corresponding to the natural language analysis request based on the key influencing factor and the change index.

[0126] Change indicators are used to quantify changes in key influencing factors. Specifically, they refer to the changes in key influencing factors within a specific time window, such as absolute difference, month-on-month comparison, and year-on-year comparison. For example, when the key influencing factor is sales volume, change indicators could be the decrease or increase in sales volume, month-on-month comparison, or year-on-year comparison.

[0127] The query results are generated based on key influencing factors and their corresponding change indicators. They are the final results that can answer users' natural language analysis requests and are usually presented to users in an intuitive and easy-to-understand way.

[0128] This step calculates the change indicators corresponding to key influencing factors based on the acquired target data and the characteristics of indicator changes. For example, if the key influencing factor is sales volume, the decrease in sales volume is calculated as the change indicator based on the changes in sales volume in the target data. Finally, based on the key influencing factors and their corresponding change indicators, the query results corresponding to the natural language analysis request are compiled and presented to users in an intuitive and easy-to-understand way to help them understand the reasons for changes in business indicators.

[0129] In one embodiment, step S104, which calculates the change index corresponding to the key influencing factor based on the target data and the indicator change characteristics, includes:

[0130] Step S10401: Based on the characteristics of the indicator changes, determine the corresponding change calculation rules.

[0131] The change calculation rule is a pre-set calculation logic based on the change characteristics of the indicator. It is used to quantify the degree of change of key influencing factors, including month-on-month, year-on-year, difference, change rate, and proportion. According to the rule, quantifiable change indicators can be obtained from the target data, providing numerical basis for causal attribution analysis.

[0132] This step first analyzes the changing characteristics of the metrics that users are interested in (such as "sales have declined for three consecutive months"), and automatically matches the corresponding calculation rules based on the characteristic type. For example, for a continuous downward trend, the month-on-month growth rate formula is selected as the calculation rule; for cyclical fluctuations, volatility analysis is used. This step ensures that the analysis method is highly aligned with the business problem by dynamically adapting the calculation logic.

[0133] Step S10402: Calculate the change index corresponding to the key influence factor according to the change calculation rule and the data corresponding to the key influence factor in the target data.

[0134] Change indicators are numerical results obtained by quantifying key influencing factor data through change calculation rules, such as "promotional investment increased by 15% month-on-month" and "customer visits decreased by 8% year-on-year", which are used to intuitively reflect the specific degree of factor change.

[0135] This step, based on established change calculation rules, extracts corresponding data for key influencing factors (such as "promotional investment" and "customer visits") from the target data and performs specific calculations. For example, if the change calculation rule is a month-on-month growth rate formula, it will extract data for key influencing factors in adjacent time periods, calculate their percentage changes, and generate quantitative results such as "promotional investment increased by 12% month-on-month in Q2".

[0136] This embodiment automatically selects the optimal calculation logic based on the characteristics of indicator changes, avoiding the subjectivity and lag of manually setting calculation rules and ensuring that the analysis method always aligns with the needs of the business scenario. By calculating the changing indicators of key influencing factors, abstract business problems are transformed into concrete numerical values. For example, it clarifies the specific impact of "a 12% increase in promotional investment" on "a decrease in sales," providing quantifiable data support for decision-making. Through dynamically adaptable change calculation rules and precise quantitative analysis, it achieves a scientific measurement and intuitive presentation of changes in key influencing factors.

[0137] This application's embodiments lower the barrier to entry through natural language interaction, distinguish between correlation and causality based on causal relationship graphs, solve the data silo problem through multi-source data integration, and achieve accurate and interpretable attribution analysis by combining single-indicator attribution and multi-indicator causal reasoning, ultimately providing enterprises with a scientific basis for decision-making.

[0138] The above are merely embodiments of this application and are not intended to limit the scope of this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of the claims of this application.

Claims

1. A data analysis method based on causal relationship mapping, characterized in that, Includes the following steps: Obtain the natural language analysis request input by the user; parse the natural language analysis request to obtain several first metric fields, several first dimension fields, indicator change characteristics, and analysis constraints; Based on the aforementioned first metric fields, the aforementioned first dimension fields, and the analytical constraints, a query instruction is generated; target data is then retrieved from the business database based on the query instruction. Based on the aforementioned first metric fields, a matching causal relationship graph is obtained; based on the target data and the causal relationship graph, key influencing factors associated with the indicator change characteristics are determined; wherein, the causal relationship graph is constructed based on business data in the business database and is used to characterize the causal relationship between various metric fields; Based on the target data and the characteristics of the indicator changes, the change indicators corresponding to the key influencing factors are calculated; based on the key influencing factors and the change indicators, the query results corresponding to the natural language analysis request are obtained.

2. The data analysis method based on causal relationship graphs according to claim 1, characterized in that, The step of generating a query instruction based on the plurality of first metric fields, the plurality of first dimension fields, and the analysis constraints includes: Based on the plurality of first metric fields and the plurality of first dimension fields, a matching business data model is obtained; the business data model includes standard metric fields, standard dimension fields of business data, and data source location information corresponding to each of the standard metric fields and the standard dimension fields. Based on the business data model, determine the corresponding first standard metric field and first standard dimension field for each first metric field and each first dimension field. A query instruction is generated based on the data source location information corresponding to each of the first standard metric fields, each of the first standard dimension fields, each of the first standard metric fields and the first standard dimension fields, and the analysis constraints.

3. The data analysis method based on causal relationship graphs according to claim 2, characterized in that, The business data model also includes query field mapping rules, which are used to indicate the field correspondence between the business data model and the business database; The steps for generating a query instruction based on the data source location information corresponding to each of the first standard metric fields, each of the first standard dimension fields, each of the first standard metric fields and the first standard dimension fields, and the analysis constraints include: Based on the query field mapping rules, determine the corresponding target metric fields and target dimension fields in the business database for each of the first standard metric fields and each of the first standard dimension fields. Based on the data source location information corresponding to each of the first standard metric fields and the first standard dimension fields, the data source location information corresponding to each of the target metric fields and the target dimension fields is determined. A query instruction is generated based on the data source location information corresponding to each of the target metric fields, each of the target dimension fields, and the analysis constraints.

4. The data analysis method based on causal relationship graphs according to claim 1, characterized in that, The steps for obtaining a matching causal relationship graph based on the aforementioned first metric fields include: Obtain a set of preset causal relationship graphs and a set of standard measurement fields corresponding to each of the causal relationship graphs; The first metric fields are matched with each set of standard metric fields to determine the set of standard metric fields that match the first metric fields. The causal relationship graph corresponding to the set of standard metric fields that are matched is determined as the matched causal relationship graph.

5. The data analysis method based on causal relationship graphs according to claim 4, characterized in that, The causal relationship graph uses standard metric fields as nodes, the causal directions between standard metric fields as directed edges, and each standard metric field is bound to a corresponding standard dimension field. The preset causal relationship maps are obtained through the following steps: Obtain the second metric field, the second dimension field, and the data source location information corresponding to each second metric field and the second dimension field from the business data in the business database; Obtain the corresponding second standard metric field and second standard dimension field for each of the second metric fields and each of the second dimension fields. Obtain causal definition data, which is used to indicate the causal relationship between various second standard metric fields; Based on each of the second standard metric fields, each of the second standard dimension fields, and the causal definition data, at least one causal relationship graph is constructed.

6. The data analysis method based on causal relationship graphs according to claim 1, characterized in that, The step of determining the key influencing factors associated with the indicator variation characteristics based on the target data and the causal relationship map includes: When there is only one second metric field in the target data, the associated standard dimension fields are obtained from the causal relationship graph based on the second metric field; the data of the second metric field and the dimension data of each of the associated standard dimension fields are extracted from the target data. Based on the data of the second metric field and the dimensional data of each of the associated standard dimension fields, the correlation coefficient between the second metric field and each of the associated standard dimension fields is calculated. Standard dimension fields whose correlation coefficients meet the preset first correlation condition are identified as candidate key influencing factors; when the data of any candidate key influencing factor meets the change trend corresponding to the indicator change characteristics, the candidate key influencing factor is identified as a key influencing factor.

7. The data analysis method based on causal relationship graphs according to claim 1, characterized in that, The step of determining the key influencing factors associated with the indicator variation characteristics based on the target data and the causal relationship map includes: When multiple third-measure fields exist in the target data, and it is determined based on the causal relationship graph that there is a direct or indirect causal relationship between the multiple third-measure fields, the causal path and causal direction between the multiple third-measure fields are obtained from the causal relationship graph. Based on the multiple third-dimensional metric fields, extract multiple sets of corresponding time-series data from the target data; Based on the causal path and causal direction, causal effect analysis is performed on the multiple sets of time series data to obtain the causal effect intensity corresponding to each of the third metric fields; The third metric field whose causal effect strength meets the preset strength condition is determined as a candidate key influence factor; when the data of any candidate key influence factor meets the change trend corresponding to the indicator change characteristics, the candidate key influence factor is determined as a key influence factor.

8. The data analysis method based on causal relationship graphs according to claim 1, characterized in that, The step of determining the key influencing factors associated with the indicator variation characteristics based on the target data and the causal relationship map includes: When there are multiple fourth metric fields in the target data, and it is determined based on the causal relationship graph that there is no direct or indirect causal relationship between the fourth metric fields, for each fourth metric field, the associated standard dimension fields are obtained from the causal relationship graph, and the data of the fourth metric field and the dimension data of each associated standard dimension field are extracted from the target data. Based on the data of the fourth metric field and the dimensional data of each of the associated standard dimension fields, the correlation coefficient between the fourth metric field and each of the associated standard dimension fields is calculated. The standard dimension fields whose correlation coefficients satisfy the preset second correlation condition are identified as candidate key influencing factors; when the data of any candidate key influencing factor satisfies the change trend corresponding to the indicator change characteristics, the candidate key influencing factor is identified as a key influencing factor.

9. The data analysis method based on causal relationship graphs according to claim 1, characterized in that, The steps for calculating the change indicators corresponding to the key influencing factors based on the target data and the indicator change characteristics include: Based on the characteristics of the indicator changes, the corresponding change calculation rules are determined; Based on the change calculation rules and the data corresponding to the key influencing factors in the target data, calculate the change index corresponding to the key influencing factors.

10. The data analysis method based on causal relationship graphs according to claim 1, characterized in that, Before the steps of parsing the natural language analysis request to obtain several first metric fields, several first dimension fields, indicator change characteristics, and analysis constraints, the method further includes: The natural language analysis request is subjected to intent recognition to obtain the intent recognition result; When the intent recognition result is an attribution analysis intent, the natural language analysis request is parsed to obtain several first metric fields, several first dimension fields, indicator change characteristics, and analysis constraints.