A multi-source index matching method and system based on semantic path dynamic recall
Patent Information
- Application Number
- CN202611299961.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-08-26
- Publication Date
- 2026-09-25
AI Technical Summary
[0005]本发明的目的在于克服现有技术的不足,适应现实需要,提供一种基于语义路径动态召回的多源指标匹配方法及系统,以解决当前自然语言转结构化查询语言方案无法自动将用户的归因类查询意图转化为可执行的计算逻辑链条,导致人工操作效率低、一致性难以保证的技术问题
1、本发明通过设计多路语义召回机制与归因知识图谱的协同工作架构,在用户以口语化方式提出归因类查询时,由多路语义召回机制从同义词、知识结构、用户行为和领域知识四个维度并行检索候选指标,将模糊的业务术语映射为标准经营指标,由归因知识图谱自动编排从目标指标到各级影响因子的因子拆解路径,并从多个异构数据源中并行获取数据,在统一口径后逐级量化各影响因子对目标指标偏差的独立贡献度,将归因意图转化为可执行的多级计算逻辑链条,实现从模糊自然语言输入到归因结论输出的自动化分析,解决当前自然语言转结构化查询语言方案无法自动将用户的归因类查询意图转化为可执行的计算逻辑链条,导致人工操作效率低、一致性难以保证的问题。
Smart Images

Figure CN122817263A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of multi-source indicator matching technology, and more specifically, to a multi-source indicator matching method and system based on semantic path dynamic recall. Background Technology
[0002] In corporate financial and operational analysis, financial managers typically need to use natural language interaction to query operational indicators from multiple heterogeneous data sources such as comprehensive budgeting systems, production statistics systems, and power trading systems, and further explore the causes of indicator deviations.
[0003] When users initiate queries with attribution-based questions such as "Why did this month's profit fall short of the budget?" or "Why is the performance worse than last month?", the core task of the system is no longer simple data retrieval. Instead, it needs to complete the entire technical chain from intent understanding, fuzzy indicator matching, analysis logic arrangement, cross-source data collaborative calculation to conclusion quality assessment, and realize the leap from "passive data retrieval" to "proactive analysis".
[0004] However, existing technical solutions based on natural language to structured query language are limited to generating and executing a single query statement, only answering the "what" question—returning the numerical result of the queried indicator—and failing to automatically orchestrate the analytical intent of "why" into an executable multi-level factor decomposition path. When users need to further investigate the causes of indicator deviations, financial personnel still need to manually export raw data at different organizational granularities across systems, manually set calculation relationships according to the logic of factor analysis, and perform item-by-item decomposition. This makes the rationality of the analytical logic and the accuracy of the conclusions highly dependent on the analyst's personal experience, and the manual orchestration and calculation process is time-consuming, making it difficult to guarantee the efficiency and consistency of attribution analysis. In view of this, we propose a multi-source indicator matching method and system based on semantic path dynamic recall. Summary of the Invention
[0005] The purpose of this invention is to overcome the shortcomings of the prior art, adapt to the needs of reality, and provide a multi-source indicator matching method and system based on semantic path dynamic recall. This solves the technical problem that current natural language to structured query language schemes cannot automatically convert users' attribution query intent into an executable computational logic chain, resulting in low efficiency and difficulty in ensuring consistency of manual operations.
[0006] To address the aforementioned technical problems, this invention provides the following technical solution: a multi-source indicator matching method based on semantic path dynamic recall, comprising the following steps: It responds to natural language queries by performing semantic parsing to extract time entities, organization entities, indicator entities, and analysis action entities; When the indicator entity does not match the standard indicator name library, a multi-path semantic recall mechanism is triggered to obtain candidate indicators from the synonym retrieval path, knowledge graph retrieval path, user behavior retrieval path and domain knowledge retrieval path. Based on the number of paths hit by each candidate indicator and the corresponding preset weight, a comprehensive confidence score is calculated to determine the target matching indicator and construct the semantic path. When the analytical action entity in the semantic path represents the attribution analysis intention, the attribution knowledge graph is searched along the contribution relationship edge starting from the target matching index to generate a factor decomposition path containing factors and calculation relationships at all levels. Based on the factor decomposition path, multi-source data is requested from the corresponding heterogeneous data source and alignment processing is performed based on the organizational structure tree. On the basis of the aligned data, the independent contribution of each level of related factors to the target matching indicator is obtained by calculating the relationship level by level. A multidimensional confidence assessment based on data completeness, consistency of definition, and path matching is performed on the independent contribution to generate confidence labels, and structured attribution conclusions are generated based on the independent contribution and confidence labels.
[0007] Preferably, the formula for calculating the comprehensive confidence score is: ; in, Indicates the first The overall confidence score of each candidate indicator The index represents the semantic retrieval path and has a total number of nodes. Path, Indicates the first The candidate indicator in the first The hit status in the path, Indicates the first Preset weights for each semantic retrieval path; The preset weights are obtained by selecting sample queries with colloquial expressions as the indicator entities, calculating the recall hit rate of each path, and then normalizing them. During system operation, the preset weights of each path are dynamically updated using the query logs within the most recent time period as statistical samples.
[0008] Preferably, before constructing the semantic path, it is determined whether the current natural language query is a follow-up question to a historical query; If it is a follow-up question, at least one of the time entity, organization entity, and target matching index is inherited from the semantic path corresponding to the historical query. The inherited entity is then merged with the newly extracted entity in the current natural language query to construct the semantic path. In this way, in multi-turn dialogue scenarios, the context memory and understanding capabilities can be used to avoid the user's repeated expressions and quickly construct the semantic path of the new query.
[0009] Preferably, the attribution knowledge graph is a directed graph data structure; The nodes of the directed graph data structure are business indicators, and the directed edges between the nodes are contribution relationship edges that represent the contribution calculation relationship between the indicators. The direction of the contribution relationship edges is from the upper factor to the lower factor and encapsulates the defined source node indicators. The calculation formula for the contribution of the target node indicator is obtained, the substitution order identifier of each source node indicator in the chain substitution calculation is recorded, and the data source identifier of the heterogeneous data source corresponding to the leaf node indicator is recorded. The factor decomposition path forms a tree or network analysis path by connecting the retrieved correlation factors and indicators at all levels according to the connection order of the contribution relationship edges.
[0010] Preferably, the caliber alignment process includes: identifying the data organization granularity corresponding to each of the multi-source data returned from each heterogeneous data source; When the data organization granularity of different data sources is inconsistent, the node affiliation relationship recorded in the organizational structure tree is obtained, and the data with inconsistent organizational granularity is unified into the target statistical granularity through data aggregation or data splitting based on the node affiliation relationship, so as to obtain multi-source data with aligned caliber. The target statistical granularity is determined based on the level of the organizational entity in the semantic path and the corresponding organizational node in the organizational structure tree.
[0011] Preferably, the step-by-step calculation follows a chain substitution logic, and the correlation factors at each level include the quantity difference influence factor and the price difference influence factor under the income deviation; The formula for calculating the impact of the difference in indicator quantities is: ; The formula for calculating the impact of the indicator price spread is: ; in, This indicates the impact of volume difference caused by changes in sales volume. This indicates the impact of price differences caused by changes in selling prices. and These represent actual sales volume and budgeted sales volume, respectively. and These represent the actual selling price and the budgeted selling price, respectively. The independent influence of each factor on the change of the target indicator is isolated by sequentially replacing the base period value and the actual value of each influencing factor.
[0012] Preferably, the evaluation of the caliber consistency dimension includes: The aperture deviation value between the multi-source data after aperture alignment is calculated, and the formula for calculating the aperture deviation value is as follows: ; in, This represents the caliber deviation value, where M is the total number of child nodes participating in the aggregation. Indicates the first The actual indicator values of each child node. This represents the total theoretical budget. The affiliation coefficient is determined by the organizational structure and affiliation relationships. A deviation threshold is set to define whether the deviation between multi-source data after caliber alignment is within an acceptable range. When the caliber deviation value is less than or equal to the deviation threshold, the evaluation result of the caliber consistency dimension is determined to be consistent.
[0013] Preferably, the generation of structured attribution conclusions includes: determining whether the confidence labels of the association factors at each level are lower than the confidence threshold; For association factors with confidence labels below the confidence threshold, the evaluation results of multidimensional confidence assessment are used to trace back to the specific evaluation stage to locate abnormal data sources or controversial analysis paths and generate verification suggestions. The independent contribution values, confidence labels, and verification suggestions of association factors at all levels are then structured and assembled to generate attribution conclusions.
[0014] A multi-source index matching system based on semantic path dynamic recall includes a semantic parsing and fuzzy recall module, an attribution orchestration module, a data retrieval and calculation module, and a confidence assessment and conclusion generation module. The semantic parsing and fuzzy recall module is used to receive natural language queries input by users and extract time entities, organizational entities, indicator entities, and analysis action entities; The attribution orchestration module is used to retrieve data along the contribution relationship edge in the attribution knowledge graph when the analysis action entity in the semantic path represents the attribution analysis intention, starting from the target matching index. Based on the retrieved association factors and calculation relationships, it generates a factor decomposition path containing the target matching index, association factors at all levels, leaf node indicators, and calculation relationships between indicators, and obtains the data source identifier of each leaf node indicator. The data retrieval and calculation module is used to initiate data requests in parallel to multiple heterogeneous data sources according to the factor decomposition path and data source identifier, perform caliber alignment processing on the returned multi-source data, and calculate step by step according to the calculation relationship defined in the factor decomposition path based on the aligned data to obtain the independent contribution of each level of related factors to the target matching index. The confidence assessment and conclusion generation module is used to perform multi-dimensional confidence assessment on the independent contribution of each level of correlation factors, based at least on the dimensions of data completeness, consistency of definition, and path matching. Based on the assessment results, confidence labels are generated, and structured attribution conclusions are generated based on the independent contribution and the confidence labels.
[0015] A computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the multi-source index matching method based on semantic path dynamic recall.
[0016] Compared with the prior art, the beneficial effects of the present invention are: 1. This invention designs a collaborative architecture between a multi-path semantic recall mechanism and an attribution knowledge graph. When a user submits an attribution query in a conversational manner, the multi-path semantic recall mechanism retrieves candidate indicators in parallel from four dimensions: synonyms, knowledge structure, user behavior, and domain knowledge. This maps fuzzy business terms to standard operational indicators. The attribution knowledge graph automatically arranges factor decomposition paths from the target indicator to each level of influencing factors and acquires data in parallel from multiple heterogeneous data sources. After unifying the standards, it quantifies the independent contribution of each influencing factor to the deviation of the target indicator at each level, transforming the attribution intent into an executable multi-level computational logic chain. This achieves automated analysis from fuzzy natural language input to attribution conclusion output, solving the problem that current natural language to structured query language solutions cannot automatically transform users' attribution query intent into an executable computational logic chain, resulting in low efficiency and difficulty in ensuring consistency in manual operations.
[0017] 2. This invention also introduces a multi-dimensional confidence assessment mechanism before generating attribution conclusions. This mechanism verifies the credibility of the calculation process and the data on which the independent contribution of each factor in the factor decomposition path is calculated. The confidence labels are generated by comprehensively evaluating the results from three dimensions: data completeness, consistency of definitions, and path matching. Based on the confidence labels, decision-makers can intuitively identify which attribution conclusions are solid and credible and which conclusions require caution. This further solves the problem that automated analysis conclusions lack credibility references and that decision-makers cannot judge the reliability of the analysis results.
[0018] 3. This invention also introduces a verification suggestion generation mechanism after the confidence label is generated. If the confidence label is lower than the confidence threshold, the system traces back to the specific stage of the multidimensional assessment to locate the abnormal data source or controversial analysis path that leads to the low credibility of the conclusion. Based on this, targeted and actionable verification suggestion information is generated. Finally, the factor contribution value, confidence label and verification suggestion information are structured and assembled for output. This allows decision-makers not only to know the analysis conclusion and its credibility, but also to know what should be verified for doubtful conclusions. This further solves the problem that confidence assessment can only indicate risks but cannot guide users to locate the root cause of the problem. Attached Figure Description
[0019] Figure 1 This is a flowchart illustrating the overall method of the present invention; Figure 2 This is a schematic diagram of the multi-path semantic recall mechanism of the present invention; Figure 3 This is a schematic diagram of the factor decomposition path in the attribution knowledge graph of this invention. Detailed Implementation
[0020] Example 1: As Figures 1 to 3 As shown, the present invention relates to a multi-source indicator matching method based on semantic path dynamic recall, comprising the following steps: S100: In response to the natural language query input by the user, perform semantic parsing on the natural language query and extract the time entity, organization entity, indicator entity and analysis action entity from it; Specifically, the system receives query text input by users in natural language, performs semantic parsing on the query text, and extracts key information representing time, organization, indicators, and analytical intent, forming time entities (such as "this month" and "yesterday"), organization entities (such as "Power Plant B" and "new energy sector"), indicator entities (such as "efficiency" and "profit"), and analytical action entities (such as "query", "why", and "comparison").
[0021] For example, if a user inputs "Why is the performance of Power Plant B this month worse than the budget?", the system will parse the input and obtain the time entity "this month", the organizational entity "Power Plant B", the indicator entity "performance", and the analysis action entities "why" and "compare with the budget".
[0022] S200. When the extracted indicator entity does not match any indicator name in the standard indicator name library, a multi-path semantic recall mechanism is triggered to obtain candidate indicators from multiple semantic retrieval paths, sort each candidate indicator by confidence, determine the candidate indicator with the highest confidence in the sorting results as the target matching indicator corresponding to the natural language query, and construct a structured semantic path based on time entity, organization entity, target matching indicator and analysis action entity. Specifically, when the extracted indicator entity (such as "efficiency") fails to match precisely with the standard indicator name library (such as standard names including "daily economic profit", "net profit", "profit margin" etc.), the fuzzy indicator processing flow is triggered, that is, four semantic retrieval paths are started simultaneously to obtain candidate indicators that may be related to the user's intent from different dimensions. The standard indicator name database is a collection of searchable standard operating indicators. The database was created by collecting the official indicator names stipulated in the company's financial management documents and the corresponding standard terms from a financial terminology thesaurus. These two databases were then merged and deduplicated to form the standard indicator name database. The company's financial management documents specify the standard names for each operating indicator, such as "daily economic profit," "net profit," and "return on net assets." The financial terminology thesaurus stores the mapping relationship between colloquial expressions and standard terms; extracting the standard terms from this thesaurus can supplement indicator names not explicitly listed in the documents but actually used. Because the indicator names in the database originate from official documents and the standard terms in the thesaurus, the standard names in the database are consistent with the company's financial analysis indicator system.
[0023] The semantic retrieval paths include the synonym retrieval path, knowledge graph retrieval path, user behavior retrieval path, and domain knowledge retrieval path; The synonym retrieval path obtains candidate indicators by searching for synonyms of indicator entities in a financial terminology thesaurus. This thesaurus is formed by collecting standard terms from corporate financial management documents and extracting frequently used colloquial expressions from historical query logs, establishing a mapping relationship between standard terms and colloquial expressions, such as "efficiency" corresponding to "profit" and "revenue." A set of candidate indicators is returned after a successful search. The knowledge graph retrieval path obtains candidate indicators by searching downwards along the indicator hierarchy within a standard indicator name database. The indicator hierarchy is set according to the indicator classification logic stipulated in the company's financial management system, describing the hierarchical affiliation of various operational indicators. During the retrieval, starting from the indicator entity, the search proceeds downwards along the indicator hierarchy to obtain all subordinate indicators corresponding to that indicator entity. Indicators belonging to the standard indicator name database among the retrieved subordinate indicators are output as candidate indicators. The user behavior retrieval path obtains candidate indicators by searching for frequently occurring indicators in the current user's historical query records. These historical query records are automatically recorded after each successful indicator query and include at least the user ID, query time, and the name of the final matching standard indicator. During the retrieval, the frequency of each indicator within a preset time window is counted, and high-frequency indicators (those exceeding the average frequency) are output as candidate indicators. The domain knowledge retrieval path obtains candidate indicators by retrieving terminology association rules from the relevant business domain. These association rules, compiled from industry conventions and expert experience in the power industry's financial analysis field, store the default associations between common colloquial expressions and standard indicators in key-value pairs. During retrieval, matching is performed using the indicator entity as the key, returning the associated default indicators as candidate indicators. For example, "efficiency" in the context of a power generation group typically refers to "daily economic profit." For each of the recalled candidate metrics, a confidence ranking is performed. Specifically, this includes: counting the number of paths hit by each candidate metric in the synonym retrieval path, knowledge graph retrieval path, user behavior retrieval path, and domain knowledge retrieval path; obtaining the preset weights corresponding to each semantic retrieval path; calculating the comprehensive confidence score of each candidate metric based on the number of hit paths and the preset weights; and ranking each candidate metric from high to low according to the comprehensive confidence score, with the candidate metric having the highest score being determined as the target matching metric.
[0024] In this embodiment, multiple recall paths return a set of candidate indicators that may contain duplicates. Determining the optimal solution from this set requires a quantitative evaluation criterion. Instead of treating all paths equally, a comprehensive scoring mechanism based on preset weights is introduced. This mechanism incorporates two considerations: first, the degree of "consensus" among candidate indicators, i.e., how many paths simultaneously recognize them; the more paths that hit them, the greater the likelihood that the indicator has received cross-validation from multiple knowledge sources, and the higher its credibility; second, the authority or accuracy of each path itself, reflected through preset weights. For example, paths based on structured knowledge (such as knowledge graph paths) may be assigned higher weights than paths based on behavioral statistics. The comprehensive score is a weighted average of the weights of each hit path. Furthermore, this approach brings significant benefits: it avoids the biased errors that may arise from a single recall source, and through weighted voting from multiple evidence sources, it ensures that the final target matching indicator not only satisfies the user's colloquial expression but is also the most reasonable from the perspectives of knowledge structure and business habits.
[0025] Confidence ranking is essentially a weighted voting mechanism. The more paths a candidate metric is hit by—meaning it points to the user's true intent from multiple perspectives such as semantics (synonyms), logic (knowledge graph), experience (user behavior), and domain rules (domain knowledge)—the higher its confidence as a correct result. Specifically, the calculation logic for the comprehensive confidence score is as follows: The algorithm comprehensively considers which paths the candidate metric is hit on and the reliability weight of each path. For example, because knowledge graph paths and domain knowledge paths contain stronger logical rules and industry knowledge, they can be assigned higher weights than user behavior paths. Under this calculation logic, even if a certain metric is hit on only a few paths, if the weights of these paths are extremely high, its final score may still be very high; that is, suppose for the ... There are 10 candidate indicators, among which... The hit states in each path constitute a binary vector. ,in 1 indicates a hit, and 0 indicates a miss. The preset weight vector for each path is: This represents the confidence contribution of each path. The overall confidence score for this candidate indicator is then calculated. The weighted summation and normalization operation of the hit state vector and weight vector is performed. By calculating the fusion degree of each recall weight and hit rate, the credibility of the matching between candidate indicators and query intent is quantified. The candidate indicators are sorted according to their scores to select the optimal target matching indicator. Overall confidence score The calculation formula is: ; in, Indicates the first The overall confidence score of each candidate indicator; this score is a measure of the confidence of the candidate indicators. The quantitative evaluation of the degree of match with the user's query intent is achieved by synthesizing recall evidence from four paths and calculating a score between 0 and 1 for the candidate metric. The higher the score, the more likely the candidate metric is to be the metric that the user actually wants to query. Dimensionless. When a candidate indicator is hit by all paths with higher weights simultaneously, the score approaches 1; when a candidate indicator is not hit by any path, the score is 0. The index representing the semantic retrieval path has a total of Path; Indicates the first The candidate indicator in the first The hit status in the path is set to 1 when a hit occurs and 0 otherwise. Indicates the first The preset weights of each semantic retrieval path satisfy the normalization condition, that is, the sum of all weights is 1.
[0026] The preset weights are determined as follows: 1. Select several natural language queries as sample queries. The indicator entities in each sample query are expressed in colloquial language, and the correct standard indicator corresponding to the colloquial expression is known. For example, in the sample query "How is the performance this month?", the correct standard indicator corresponding to the indicator entity "performance" is "daily economic profit"; in the sample query "Yesterday's revenue", the correct standard indicator corresponding to the indicator entity "revenue" is "operating income". 2. For each sample query, the system executes the synonym retrieval path, knowledge graph retrieval path, user behavior retrieval path, and domain knowledge retrieval path respectively, and records the recall results for each path. If the candidate indicator set recalled by a certain path contains the correct standard indicator, it is recorded as a hit; otherwise, it is recorded as a miss. 3. Count the number of times each path is hit in all sample queries, and calculate the recall rate for each path. Let the first path be... The total number of executions for each path is Number of hits The recall hit rate of this path ; 4. Normalize the recall hit rate of each path to obtain a preset weight for each path. The calculation method is as follows: ,in For path The preset weights. After normalization, the sum of all weights is 1; The weight determination process can be executed periodically, utilizing query logs accumulated during actual operation to update hit rate statistics and achieve dynamic weight optimization. During each update, the query logs from the most recent time period are used as statistical samples, and the weights of each path are recalculated according to steps three and four above. This time period is used to limit the time frame of the statistical samples, ensuring that weight updates are based on recent system operation data, thereby promptly reflecting the current recall performance of each path.
[0027] As a specific implementation method, when the system is initially launched, initial weights can be obtained by statistically analyzing sample queries annotated by business experts. For example, through the above steps, the initial weights of the knowledge graph retrieval path are calculated to be 0.4, the synonym retrieval path to be 0.3, the user behavior retrieval path to be 0.2, and the domain knowledge retrieval path to be 0.1. If a candidate indicator is simultaneously matched by the synonym, knowledge graph, and user behavior paths, its comprehensive score is 0.3 + 0.4 + 0.2 = 0.9.
[0028] Furthermore, before constructing the structured semantic path in step S200, the process includes: determining whether the current natural language query is a follow-up to a historical query; if so, inheriting at least one of the time entity, organizational entity, and target matching index from the semantic path corresponding to the historical query, and merging the inherited entity with the newly extracted entity in the current natural language query to construct the structured semantic path. This enables contextual memory and understanding capabilities in multi-turn dialogue scenarios. For example, if a user asks "What about power plant B?" after asking "Profit of power plant A this month?", the system will inherit the contextual entities "this month" and "profit (daily economic profit)," and only replace the organizational entity with the newly extracted "power plant B" to quickly construct the semantic path for the new query, avoiding repetitive statements from the user.
[0029] S300. When the analysis action entity in the semantic path represents the attribution analysis intention, starting from the target matching index, the search is performed along the contribution relationship edge in the constructed attribution knowledge graph. Based on the related factors and the calculation relationship between the related factors obtained during the search process, a factor decomposition path containing the target matching index, related factors at all levels, leaf node indicators and the calculation relationship between the indicators is generated, and the data source identifier of each leaf node indicator is obtained. The attribution knowledge graph is a pre-constructed directed graph data structure. The nodes of the directed graph are operational indicators, and the directed edges between nodes represent the contribution calculation relationship between the indicators, i.e., contribution relationship edges. The direction of the contribution relationship edge points from the higher-level factor to its decomposed lower-level factor, indicating that the higher-level factor can be decomposed into the contribution of the lower-level factor. Each contribution relationship edge encapsulates at least the following three aspects of information: Calculation formula: Defines how to calculate the contribution to the target node indicator from the source node indicator; Substitution order identifier: Records the substitution order of each source node indicator in the chain substitution calculation when the target node indicator is affected by multiple source node indicators; Data source identifier: When the source node metric is a leaf node metric, record the heterogeneous data source corresponding to that leaf node metric.
[0030] The contribution calculation relationship is constructed based on the analytical logic of factor analysis. Factor analysis is an analytical method that uses mathematical decomposition to attribute changes in a target indicator to changes in several influencing factors. The chain substitution method is a specific implementation of factor analysis. It isolates the independent influence of each factor on the change in the target indicator by sequentially replacing the base period value and the actual value of each influencing factor. The base period value refers to the value used as a comparison benchmark. In the attribution analysis scenario, the base period value is the budgeted or planned value recorded in the benchmark data source corresponding to the leaf node indicator in the attribution knowledge graph; the actual value refers to the actual value recorded in the actual data source corresponding to the leaf node indicator. In the chain substitution calculation, the base period result of the target indicator is first calculated using the base period values of each influencing factor. Then, the base period values of each influencing factor are replaced with actual values sequentially according to the substitution order marked on the contribution relationship side. Each replacement calculates the current result of the target indicator, and the difference between the current result and the result before replacement is the independent contribution of that influencing factor to the change in the target indicator.
[0031] The attribution knowledge graph includes at least the following: a root node indicator is connected to a first-level influence factor via a first-level contribution edge (i.e., an edge connecting the root node indicator to the first-level influence factor based on the contribution relationship between the root node indicator and the first-level influence factor); the first-level influence factor is connected to a second-level influence factor via a second-level contribution edge; and the second-level influence factor is the leaf node indicator. Correspondingly, a factor decomposition path is generated based on the associated factors and the calculation relationships between them obtained during the retrieval process. This includes: connecting the retrieved associated factors at all levels and the calculation relationships between the indicators according to the connection order of the contribution relationship edges to form a tree-like or network analysis path from the root node indicator to each leaf node indicator. Specifically, the nodes of the knowledge graph represent specific, quantifiable, or calculable business indicators, while the edges encapsulate the "contribution calculation relationship" between two indicators, showing how one indicator is calculated to derive the other. This relationship is pre-constructed directly based on the chain substitution logic of factor analysis, ensuring that the decomposition logic conforms to professional standards in the field of financial analysis. When the system performs a search, it starts from the root node of the "target indicator" and traverses the edges of the graph one by one. Each time an edge is traversed, an intermediate-level correlation factor is introduced, and a calculation relationship is inherited, until the indivisible leaf node indicator is reached. Connecting such a path composed of node-edge-node naturally forms a tree-like or network-like factor decomposition path.
[0032] Furthermore, by graphically modeling the principles of factor analysis, a complex financial analysis problem is transformed into a path search problem on a directed graph. Primary influencing factors are the direct drivers of changes in target indicators, while secondary influencing factors are deeper-level factors explaining the reasons for these primary changes. For example, taking "profit deviation" as the root node, it is connected to the two primary influencing factors, "revenue deviation" and "cost deviation," via primary contribution edges (i.e., the edges connecting the root node indicator and the primary influencing factors). The calculation relationship on these edges is "revenue deviation + cost deviation." The "revenue deviation" node is further decomposed, connected to the two secondary influencing factors, "quantity difference impact" and "price difference impact," via secondary contribution edges (i.e., the edges connecting the primary and secondary influencing factors). These two are indivisible leaf node indicators, with the calculation relationships on their edges being "(actual power generation - budgeted power generation) × budgeted on-grid electricity price" and "(actual on-grid electricity price - budgeted on-grid electricity price) × actual power generation," respectively.
[0033] In one embodiment, the method for determining the attribution analysis intent represented by the action entity in step S300 is as follows: when the action entity contains attribution trigger keywords, it is determined that the action entity in the semantic path represents attribution analysis intent; the attribution trigger keywords include at least one of "why," "reason," and "influencing factors." This keyword-triggered approach can quickly and accurately identify the user's deep analytical intent, switching the process from simple query branches to complex attribution orchestration branches.
[0034] S400. Based on the factor decomposition path and the data source identifier of each leaf node indicator, initiate data requests in parallel to multiple corresponding heterogeneous data sources, perform caliber alignment processing on the multi-source data returned from multiple heterogeneous data sources, and calculate step by step according to the calculation relationship defined in the factor decomposition path based on the aligned data to obtain the independent contribution of each level of related factors to the target matching indicator. Here, multiple heterogeneous data sources refer to multiple business systems storing enterprise operational indicator data. These multiple heterogeneous data sources include at least: a production statistics system for storing actual production data, a power trading system for storing electricity trading price data, and a comprehensive budget system for storing budget preparation data. These heterogeneous data sources differ in data storage format, data organization granularity, and data update cycle. The system obtains the connection address, data organization granularity, and data update cycle of each data source through its registration information.
[0035] Perform caliber alignment processing on multi-source data returned from multiple heterogeneous data sources, including: Identify the data organization granularity corresponding to the multi-source data returned from various heterogeneous data sources; When the data organization granularity of different data sources is inconsistent, obtain the node affiliation relationship recorded in the organizational structure tree; Based on node affiliation, data with inconsistent organizational granularity are unified into the target statistical granularity through data aggregation or data splitting, resulting in multi-source data with aligned calibers. The data aggregation process is as follows: when the organizational granularity of the original data is finer than the target statistical granularity, the data of all subordinate nodes belonging to the same superior node are summed according to the node affiliation relationship in the organizational structure tree to obtain the data of that superior node level.
[0036] The data splitting process is as follows: When the granularity of the original data organization is coarser than the target statistical granularity, the data values of the corresponding upper-level nodes are obtained from the heterogeneous data source, and the actual output data of each lower-level node in the previous complete accounting cycle is obtained from the production statistics system. The proportion of the actual output of each lower-level node in the total output of the upper-level node is calculated, and the data values of the upper-level nodes are allocated according to the output proportion of each lower-level node to obtain the data at each lower-level node level. The formula for calculating the splitting ratio is: Splitting ratio = Historical actual output of each lower-level node ÷ Historical total output of the upper-level node, where the historical actual output of each lower-level node is obtained from the output records of the production statistics system, and the historical total output of the upper-level node is obtained by summing the historical actual output of each lower-level node. The target statistical granularity is determined based on the organizational entities in the semantic path. Specifically, the organizational entity extracted in step S100 corresponds to a unique organizational node in the organizational structure tree, and the level of this organizational node is the target statistical granularity. For example, when the organizational entity in the user query is "Power Plant A", "Power Plant A" corresponds to the legal entity level in the organizational structure tree, so the target statistical granularity is the legal entity; when the organizational entity in the user query is "New Energy Sector", this organizational entity corresponds to the business sector level, so the target statistical granularity is the business sector.
[0037] In this context, granular alignment is a prerequisite for ensuring the accuracy of subsequent factor calculations. If the data granularity is inconsistent—for example, production system data is at the "station" level while budget system data is at the "legal entity" level—the directly calculated contribution will be incorrect. By accessing the internal organizational structure tree, the system automatically determines that "Station A, Station B, and Station C" all belong to "Power Plant A Legal Entity," thus aggregating (summing) the power generation data of these three stations to obtain the total power generation at the "Power Plant A Legal Entity" level, which is then compared and calculated with the budget data at the same legal entity level. Conversely, if the target granularity is "station," the "legal entity" level budget data can be broken down into individual stations based on the proportion of each station's actual output in the previous complete accounting cycle to the total output of the legal entity. Historical output data for each station is obtained from the production statistics system, and the splitting ratio is dynamically calculated each time granular alignment is performed.
[0038] It's important to note that an organizational tree is a data structure used to describe the hierarchical relationships between organizational units within an enterprise. This organizational tree is stored in a tree structure, with the root node representing the highest-level organizational unit, leaf nodes representing the lowest-level organizational units, and intermediate nodes representing intermediate organizational units at various levels. The parent-child relationship between nodes indicates that a lower-level organizational unit is subordinate to its superior organizational unit in terms of administrative management and financial accounting. The data for this organizational tree originates from the organizational master data in the enterprise's human resource management system or financial accounting system.
[0039] Furthermore, the step-by-step calculation of contribution in step S400 follows the pre-defined calculation relationships in the attribution knowledge graph; that is, after the retrieval path is generated, quantitative calculations are performed accordingly. Taking a typical three-factor decomposition scenario of "profit deviation" as an example, profit deviation is defined as actual profit... With budget profit The difference is contributed by both the primary factors of revenue deviation and cost deviation. Furthermore, the revenue deviation is influenced by the quantity difference of the leaf node indicators. And the impact of price difference Composition. Influence of quantity difference. This indicates the revenue discrepancy caused by changes in sales volume, and the impact of price differences. This indicates the revenue discrepancy caused by changes in sales prices.
[0040] Specifically, when returning actual sales volume Actual selling price Budgeted sales volume and budgeted sales price Then, the calculation is performed using the following relationship: Impact of Calculation Difference At that time, the sales volume factor is replaced with the actual value from the budgeted value, while keeping the price factor at the budgeted level, to calculate the revenue difference caused by the change in sales volume; the impact of the price difference is then calculated. At that time, with the sales volume factor already updated to the actual value, the price factor is replaced from the budgeted value to the actual value, and the revenue difference caused by the price change is calculated. This achieves a clean and non-overlapping separation of the contribution of each leaf node indicator.
[0041] Impact of indicator quantity difference And the impact of price difference The calculation formulas are as follows: ; ; in, This indicates the impact of volume difference, that is, the independent revenue contribution caused by changes in sales volume, and its unit is consistent with the profit unit (e.g., ten thousand yuan). and These represent the actual sales volume and budgeted sales volume obtained from the production statistics system and the comprehensive budget system, respectively, in physical units of measurement (such as megawatt-hours, tons, etc.). This indicates the budgeted sales price obtained from the comprehensive budget system, in monetary units / physical units (e.g., yuan / megawatt-hour). This indicates the actual sales price obtained from the electricity trading system, in units of... same. This indicates the impact of price difference, that is, the independent revenue contribution caused by changes in sales price, and its unit is consistent with the profit unit.
[0042] In this way, It precisely calculated the revenue impact of changes in sales volume alone, assuming prices remain constant; and This precisely calculates the revenue impact of price changes alone at the current actual sales volume level. The sum of these two factors constitutes the total revenue deviation, achieving a complete decomposition of the factors.
[0043] For example, let's take the profit deviation attribution analysis of Power Plant B this month as an example. The system obtains the actual power generation of Power Plant B this month as 180 million kWh from the production statistics system, the budgeted power generation as 200 million kWh from the comprehensive budget system, and the budgeted on-grid electricity price as 0.5 yuan / kWh; it obtains the actual on-grid electricity price as 0.48 yuan / kWh from the power trading system; and it obtains the actual operation and maintenance cost as 8.5 million yuan and the budgeted operation and maintenance cost as 8 million yuan from the accounting system and the comprehensive budget system, respectively.
[0044] Calculate the first-level factors: revenue deviation and cost deviation. Budgeted revenue = Budgeted power generation × Budgeted on-grid electricity price = 20,000 × 0.5 = 100 million kWh = 100 million yuan; Actual revenue = actual power generation × actual grid connection price = 18000 × 0.48 = 86.4 million yuan; Revenue deviation = Actual revenue - Budgeted revenue = 86.4 million - 100 million = -13.6 million yuan; Cost deviation = Actual cost - Budgeted cost = Actual maintenance cost - Budgeted maintenance cost = 850 - 800 = 500,000 yuan; Profit deviation = Revenue deviation + Cost deviation = -1360 + 50 = -1310 million yuan.
[0045] The independent contribution of each leaf node indicator under income deviation is calculated using the chain substitution method: Quantity difference impact =(Actual power generation - Budgeted power generation) × Budgeted on-grid electricity price = (18000 - 20000) × 0.5 = (-2000) × 0.5 = -10 million yuan; Price spread impact =Actual power generation × (Actual grid connection price - Budgeted grid connection price) = 18000 × (0.48 - 0.5) = 18000 × (-0.02) = -3.6 million yuan; Verification: Quantity difference impact + Price difference impact = -1000 + (-360) = -13.6 million yuan = Revenue deviation, complete decomposition.
[0046] The final independent contributions of each factor are as follows: Quantity difference impact -10 million yuan (due to power generation being lower than the budget, a negative contribution), Price difference impact -3.6 million yuan (due to the on-grid electricity price being lower than the budget, a negative contribution), and Operation and maintenance cost deviation 500,000 yuan (due to operation and maintenance costs exceeding the budget, a negative contribution; the direction is the same as profit, but the value is positive, indicating an increase in costs). The total of the three is -1000 + (-360) + 50 = -13.1 million yuan = profit deviation, which is a complete breakdown.
[0047] S500. Perform multidimensional confidence assessment on the independent contribution of each level of correlation factors. The multidimensional confidence assessment shall be based on at least the following three dimensions: data completeness dimension, consistency of definition dimension, and path matching dimension. Generate corresponding confidence labels for each level of correlation factors based on the assessment results, and generate structured attribution conclusions based at least on the independent contribution and confidence labels.
[0048] Perform multidimensional confidence assessments on the independent contributions of each level of correlation factors, including: Check whether all heterogeneous data sources have successfully returned valid data and no missing values to obtain the evaluation result of the data completeness dimension, which is either complete or incomplete; Calculate the caliber deviation value between the multi-source data after caliber alignment, determine whether the caliber deviation value is within the deviation threshold range, and obtain the evaluation result of the caliber consistency dimension. The evaluation result is consistent or inconsistent. The factor decomposition path is compared with the pre-defined standard analysis path in the attribution knowledge graph by nodes and edges. The evaluation result of the path matching degree is determined based on the degree of overlap. This evaluation result includes at least high matching and low matching. The determination method for high matching and low matching is as follows: when all nodes and edges in the dynamically arranged factor decomposition path overlap with the nodes and edges in the standard analysis path, it is judged as a high matching; when there are non-overlapping nodes or edges, it is judged as a low matching. The confidence labels for each level of correlation factor are generated by combining the evaluation results of data completeness, consistency of definition, and path matching. Specifically, a high-confidence label is generated when the data completeness evaluation result is complete, the consistency of definition evaluation result is consistent, and the path matching evaluation result is high matching; a low-confidence label is generated when any of the dimensions' evaluation results is incomplete, inconsistent, or low matching; and a medium-confidence label is generated in all other cases.
[0049] In this embodiment, the three-dimensional evaluation constitutes a comprehensive framework for verifying the quality of the analysis conclusions. The data integrity dimension serves as a preliminary basic check to ensure that the input data itself is free from errors caused by interface or database issues; the consistency dimension re-evaluates the rationality of the processing results after data aggregation or splitting alignment processing; and the path matching dimension evaluates from a methodological perspective whether the robot's automatic orchestration analysis logic deviates from the generally accepted standard paths in the field. The evaluation of the consistency dimension involves quantifying whether the statistical error of data from different sources is within an acceptable range after consistency alignment, i.e., judging whether the deviation is within an acceptable range by using a deviation threshold. Therefore, a consistency deviation value is introduced as a quantitative indicator to measure the degree of deviation between multi-source data after consistency alignment. The formula for calculating the consistency deviation value is: ; in, This represents the diameter deviation value, which is dimensionless. An identifier representing the actual total value after aggregation (actual measured value); The identifier representing the theoretical baseline value (theoretical target value); M is the total number of child nodes participating in the aggregation, which is the next-level organizational unit that makes up the current target organizational node; Indicates the first one obtained from the data source The actual indicator value of each sub-node is in the standard unit of measurement of that indicator (such as 10,000 kilowatt-hours, 10,000 yuan, etc.). This represents the theoretical total budget or standard value stored in the benchmark database within the same time period as the target organization node, and its unit is... same; It is an affiliation coefficient determined by the organizational structure and affiliation relationships. When the first... When a child node belongs entirely to the target organization node, its value is 1; if it belongs only partially, it is its belonging ratio, which is dimensionless.
[0050] This formula calculates the normalized absolute deviation between the aggregated actual total and the theoretical baseline value. The calculated... The closer the value is to 0, the higher the consistency and the better the alignment. The system determines whether the deviation value is within the deviation threshold range, thus concluding that it is "consistent" or "there is a deviation".
[0051] The deviation threshold is used to define whether the deviation between multi-source data after alignment is within an acceptable range. The deviation threshold is determined as follows: collect multi-source data samples from multiple accounting periods, perform alignment processing on each sample, calculate the deviation value for each sample, and use the upper limit of the statistical distribution of the deviation values for each sample as the initial deviation threshold. As a specific method, for example, the mean and standard deviation of the deviation values for all samples can be calculated, and the upper limit of the statistical distribution, i.e., the initial deviation threshold, can be used as the mean plus three times the standard deviation during subsequent operation. In practice, the deviation threshold can be dynamically adjusted according to the stringency of data consistency requirements in actual business operations.
[0052] Step S500, which generates a structured attribution conclusion based at least on independent contribution and confidence labels, specifically includes: determining whether the confidence labels of each level of related factors are lower than the confidence threshold; for related factors with confidence labels lower than the confidence threshold, locating the abnormal data source or indicator node based on the evaluation results of the multidimensional confidence assessment, and generating corresponding verification suggestion information; and structuring and assembling the independent contribution values, confidence labels, and verification suggestion information of each level of related factors to generate the attribution conclusion. By comparing confidence labels with confidence thresholds, the system automatically identifies potential risks. For identified risks, the system doesn't just provide a vague warning; instead, it traces back to the specific stage of the multi-dimensional assessment, pinpointing which data source is problematic or which part of the analysis path is controversial, and generating targeted, actionable verification recommendations accordingly. The final structured assembly integrates facts (contribution values), evaluations (confidence labels), and action recommendations (verification recommendations). This allows decision-makers not only to see the analysis results but also to clearly understand which conclusions are solid, which require caution, and what to verify for questionable parts, thus enhancing the decision support value and operability of the analysis conclusions.
[0053] The confidence threshold is used to determine whether the confidence label of an attribution factor meets the acceptance criteria, and is set according to the reliability requirements of the analysis conclusions based on business decisions. As a specific implementation, the confidence threshold can be set to "medium," meaning that when the confidence label of a factor is "high" or "medium," the analysis conclusion of that factor is considered acceptable; when the confidence label of a factor is "low," the analysis conclusion of that factor needs further verification.
[0054] For example: Following the above example of profit deviation attribution analysis for power plant B, a multidimensional confidence assessment is performed on the independent contributions of the three factors: quantity difference impact, price difference impact, and operation and maintenance cost deviation. Data completeness dimension: The production statistics system, power trading system, comprehensive budget system, and accounting system all successfully returned valid data and there were no missing values. The data completeness assessment results for all three factors were "complete".
[0055] Consistency of Reference Standards: The calculation of the impact of quantity differences involves the power generation data from the production statistics system and the budgeted power generation data from the comprehensive budget system. After the two are aligned, they are compared at the legal entity level, and the reference standard deviation value is calculated. The calculation result is 0.02, which is within the deviation threshold range, and the consistency assessment result is "consistent". The price difference impact calculation involves the electricity price data of the power trading system and the budgeted electricity price data of the comprehensive budget system. The statistical methods of the two are consistent, and the consistency assessment result is "consistent". The operation and maintenance cost deviation calculation involves the actual operation and maintenance cost of the accounting system and the budgeted operation and maintenance cost of the comprehensive budget system. The consistency assessment result is "consistent".
[0056] Path matching degree dimension: The factor decomposition path arranged in this study was compared with the standard analysis path in the attribution knowledge graph. The nodes and edges of the three branches of revenue deviation to quantity difference impact, revenue deviation to price difference impact, and cost deviation to operation and maintenance cost deviation all overlap with the nodes and edges in the standard path. The path matching degree evaluation result is "high matching".
[0057] Based on the combined evaluation results of the three dimensions, the impact of quantity difference, price difference, and operation and maintenance cost deviation all meet the criteria of "complete, consistent, and highly matched", generating a "high" confidence label.
[0058] Determine whether the confidence labels of each factor are lower than the confidence threshold (the threshold is set to "medium"): If the confidence labels of all three factors are "high" and not lower than the threshold, they can be directly adopted.
[0059] The structured attribution conclusion is as follows: "Power Plant B's profit this month was 13.1 million yuan lower than budgeted. Reason breakdown:" 1. Power generation fell short of budget, resulting in a difference of -10 million yuan (Confidence level: High). 2. The on-grid electricity price was lower than the budget, resulting in a price difference impact of -3.6 million yuan (confidence level: high); 3. Maintenance costs exceeded the budget, resulting in a cost deviation of 500,000 yuan (Confidence level: High). Example 2: A multi-source index matching system based on semantic path dynamic recall, comprising: The semantic parsing and fuzzy recall module receives natural language queries from users and extracts time entities, organizational entities, indicator entities, and analysis action entities. When an indicator entity does not match any indicator name in the standard indicator name library, a multi-path semantic recall mechanism is triggered. Candidate indicators are retrieved from multiple semantic retrieval paths, ranked by confidence, and the candidate indicator with the highest confidence is determined as the target matching indicator. A structured semantic path is then constructed based on each entity and the target matching indicator. In one example, this module is deployed on an application server and integrates a natural language processing (NLP) microservice and a frequently accessed Redis cache for fast retrieval of thesaurus, user history records, etc.
[0060] The attribution orchestration module is used to retrieve, starting from the target matching index, the attribution knowledge graph along the contribution relationship edges when the analytical action entity in the semantic path represents the attribution analysis intent. Based on the retrieved association factors and computational relationships, it generates a factor decomposition path containing the target matching index, association factors at all levels, leaf node indicators, and computational relationships between indicators, and obtains the data source identifier of each leaf node indicator. The core of this module is the attribution knowledge graph stored in Neo4j or a similar graph database.
[0061] The data retrieval and calculation module initiates data requests in parallel to multiple heterogeneous data sources based on the factor decomposition path and data source identifier. It performs caliber alignment processing on the returned multi-source data and, based on the aligned data, performs step-by-step calculations according to the calculation relationships defined in the factor decomposition path to obtain the independent contribution of each level of related factors to the target matching indicator. This module sends requests in parallel to the API gateways of various business systems (such as the comprehensive budget system and production statistics system) through a thread pool or asynchronous message queue, and performs aggregation or split calculations using organizational structure tree data.
[0062] The confidence assessment and conclusion generation module is used to perform multi-dimensional confidence assessments on the independent contributions of factors at all levels, based on at least the dimensions of data completeness, consistency of definition, and path matching. Based on the assessment results, confidence labels are generated, and structured attribution conclusions are generated based on the independent contributions and the confidence labels.
[0063] Example 3: A computer-readable storage medium storing a computer program, wherein the computer program is configured to execute, at runtime, steps of a multi-source index matching method based on semantic path dynamic recall.
[0064] The embodiments disclosed in this invention are preferred embodiments, but are not limited thereto. Those skilled in the art can easily understand the spirit of this invention based on the above embodiments and make different extensions and variations, but as long as they do not depart from the spirit of this invention, they are all within the protection scope of this invention.
Claims
1. A multi-source indicator matching method based on semantic path dynamic recall, characterized in that, Includes the following steps: It responds to natural language queries by performing semantic parsing to extract time entities, organization entities, indicator entities, and analysis action entities; When the indicator entity does not match the standard indicator name library, a multi-path semantic recall mechanism is triggered to obtain candidate indicators from the synonym retrieval path, knowledge graph retrieval path, user behavior retrieval path and domain knowledge retrieval path. Based on the number of paths hit by each candidate indicator and the corresponding preset weight, a comprehensive confidence score is calculated to determine the target matching indicator and construct the semantic path. When the analytical action entity in the semantic path represents the attribution analysis intention, the attribution knowledge graph is searched along the contribution relationship edge starting from the target matching index to generate a factor decomposition path containing factors and calculation relationships at all levels. Based on the factor decomposition path, multi-source data is requested from the corresponding heterogeneous data source and alignment processing is performed based on the organizational structure tree. On the basis of the aligned data, the independent contribution of each level of related factors to the target matching indicator is obtained by calculating the relationship level by level. A multidimensional confidence assessment based on data completeness, consistency of definition, and path matching is performed on the independent contribution to generate confidence labels, and structured attribution conclusions are generated based on the independent contribution and confidence labels.
2. The multi-source index matching method based on semantic path dynamic recall according to claim 1, characterized in that, The formula for calculating the overall confidence score is as follows: ; in, Indicates the first The overall confidence score of each candidate indicator The index represents the semantic retrieval path and has a total number of nodes. Path, Indicates the first The candidate indicator in the first The hit status in the path, Indicates the first Preset weights for each semantic retrieval path; The preset weights are obtained by selecting sample queries with colloquial expressions as the indicator entities, calculating the recall hit rate of each path, and then normalizing them. During system operation, the preset weights of each path are dynamically updated using the query logs within the most recent time period as statistical samples.
3. The multi-source indicator matching method based on semantic path dynamic recall according to claim 1, characterized in that, Before constructing the semantic path, determine whether the current natural language query is a follow-up query to a historical query; If it is a follow-up question, at least one of the time entity, organization entity, and target matching index is inherited from the semantic path corresponding to the historical query. The inherited entity is then merged with the newly extracted entity in the current natural language query to construct the semantic path. In this way, in multi-turn dialogue scenarios, the context memory and understanding capabilities can be used to avoid the user's repeated expressions and quickly construct the semantic path of the new query.
4. The multi-source indicator matching method based on semantic path dynamic recall according to claim 1, characterized in that, The attribution knowledge graph is a directed graph data structure; The nodes of the directed graph data structure are business indicators, and the directed edges between the nodes are contribution relationship edges that represent the contribution calculation relationship between the indicators. The direction of the contribution relationship edges is from the upper factor to the lower factor and encapsulates the defined source node indicators. The calculation formula for the contribution of the target node indicator is obtained, the substitution order identifier of each source node indicator in the chain substitution calculation is recorded, and the data source identifier of the heterogeneous data source corresponding to the leaf node indicator is recorded. The factor decomposition path forms a tree or network analysis path by connecting the retrieved correlation factors and indicators at all levels according to the connection order of the contribution relationship edges.
5. The multi-source index matching method based on semantic path dynamic recall according to claim 1, characterized in that, The caliber alignment process includes: identifying the data organization granularity corresponding to each of the multi-source data returned from each heterogeneous data source; When the data organization granularity of different data sources is inconsistent, the node affiliation relationship recorded in the organizational structure tree is obtained, and the data with inconsistent organizational granularity is unified into the target statistical granularity through data aggregation or data splitting based on the node affiliation relationship, so as to obtain multi-source data with aligned caliber. The target statistical granularity is determined based on the level of the organizational entity in the semantic path and the corresponding organizational node in the organizational structure tree.
6. The multi-source indicator matching method based on semantic path dynamic recall according to claim 1, characterized in that, The step-by-step calculation follows a chain substitution logic, and the correlation factors at each level include the quantity difference impact factor and the price difference impact factor under income deviation; The formula for calculating the impact of the difference in indicator quantities is: ; The formula for calculating the impact of the indicator price spread is: ; in, This indicates the impact of volume difference caused by changes in sales volume. This indicates the impact of price differences caused by changes in selling prices. and These represent actual sales volume and budgeted sales volume, respectively. and These represent the actual selling price and the budgeted selling price, respectively. The independent influence of each factor on the change of the target indicator is isolated by sequentially replacing the base period value and the actual value of each influencing factor.
7. The multi-source index matching method based on semantic path dynamic recall according to claim 1, characterized in that, The evaluation of the caliber consistency dimension includes: The aperture deviation value between the multi-source data after aperture alignment is calculated, and the formula for calculating the aperture deviation value is as follows: ; in, This represents the caliber deviation value, where M is the total number of child nodes participating in the aggregation. Indicates the first The actual indicator values of each child node. This represents the total theoretical budget. The affiliation coefficient is determined by the organizational structure and affiliation relationships. A deviation threshold is set to define whether the deviation between multi-source data after caliber alignment is within an acceptable range. When the caliber deviation value is less than or equal to the deviation threshold, the evaluation result of the caliber consistency dimension is determined to be consistent.
8. The multi-source index matching method based on semantic path dynamic recall according to claim 1, characterized in that, The generation of structured attribution conclusions includes: determining whether the confidence labels of each level of association factors are lower than the confidence threshold; For association factors with confidence labels below the confidence threshold, the evaluation results of multidimensional confidence assessment are used to trace back to the specific evaluation stage to locate abnormal data sources or controversial analysis paths and generate verification suggestions. The independent contribution values, confidence labels, and verification suggestions of association factors at all levels are then structured and assembled to generate attribution conclusions.
9. A multi-source indicator matching system based on semantic path dynamic recall, characterized in that, It includes a semantic parsing and fuzzy recall module, an attribution orchestration module, a data retrieval and calculation module, and a confidence assessment and conclusion generation module; The semantic parsing and fuzzy recall module is used to receive natural language queries input by users and extract time entities, organizational entities, indicator entities, and analysis action entities; The attribution orchestration module is used to retrieve data along the contribution relationship edge in the attribution knowledge graph when the analysis action entity in the semantic path represents the attribution analysis intention, starting from the target matching index. Based on the retrieved association factors and calculation relationships, it generates a factor decomposition path containing the target matching index, association factors at all levels, leaf node indicators, and calculation relationships between indicators, and obtains the data source identifier of each leaf node indicator. The data retrieval and calculation module is used to initiate data requests in parallel to multiple heterogeneous data sources according to the factor decomposition path and data source identifier, perform caliber alignment processing on the returned multi-source data, and calculate step by step according to the calculation relationship defined in the factor decomposition path based on the aligned data to obtain the independent contribution of each level of related factors to the target matching index. The confidence assessment and conclusion generation module is used to perform multi-dimensional confidence assessment on the independent contribution of each level of correlation factors, based at least on the dimensions of data completeness, consistency of definition, and path matching. Based on the assessment results, confidence labels are generated, and structured attribution conclusions are generated based on the independent contribution and the confidence labels.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the multi-source index matching method based on semantic path dynamic recall as described in any one of claims 1 to 8.