Enterprise multi-level relation rapid retrieval method and system based on knowledge graph
By constructing an enterprise knowledge graph and combining it with graph traversal algorithms and intelligent analysis, the problem of multi-level retrieval and tracing of knowledge graphs in enterprise relationship management is solved, enabling rapid and accurate analysis of complex relationships between enterprises and improving the intelligent decision support capabilities of enterprise management.
Patent Information
- Application Number
- CN202511534697.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-27
- Publication Date
- 2025-11-25
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing knowledge graphs struggle to support rapid retrieval and accurate tracing of multi-level relationships in enterprise relationship management. They lack end-to-end integration capabilities from data collection to in-depth analysis, resulting in information lag and insufficient insights.
By collecting heterogeneous data from multiple sources both inside and outside the enterprise, performing legality verification and standardization processing, an enterprise knowledge graph is constructed. Based on graph traversal algorithms, rapid retrieval and cost-benefit analysis of multi-level relationships are performed, including associated cost accounting, benefit assessment, cost-benefit association mining, and intelligent prediction, providing anomaly early warning and source tracing.
It enables unified modeling and visualization analysis of multi-dimensional relationships between enterprises, enhancing the depth and breadth of enterprise relationship mining, supporting real-time and actionable intelligent insights, and improving the accuracy and efficiency of risk identification, investment decisions, and strategic planning.
Smart Images

Figure CN121009218A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of enterprise management technology, specifically to a method and system for rapid retrieval of multi-level relationships within an enterprise based on knowledge graphs. Background Technology
[0002] In the modern business and financial environment, inter-firm relationships are becoming increasingly complex, encompassing equity structures, supply chains, partnerships, customer relationships, legal proceedings, and many other aspects. To help businesses, financial institutions, and regulatory agencies better understand these complex relationships, knowledge graph-based enterprise relationship mining and analysis methods have emerged. Traditional relationship management methods often rely on isolated databases and static reports, making it difficult to effectively handle multi-source, heterogeneous data and reveal deep, cross-level dynamic connections in real time. This leads to challenges such as information lag and insufficient insights for enterprises in risk management, investment decisions, and operational analysis.
[0003] Knowledge graph-based enterprise relationship mining and analysis is a technical approach that reveals potential relationships and interaction patterns between enterprises by constructing and analyzing knowledge graphs of inter-enterprise relationships. As a graphical data representation, knowledge graphs can present entities such as enterprises, shareholders, and partners, and their relationships in a graph structure, making complex enterprise relationships more intuitive and understandable. However, existing knowledge graph applications often focus on displaying relationships at a macro level. When facing specific business scenarios, they often lack end-to-end integration capabilities from data collection and standardization to in-depth analysis and decision-making. This makes it difficult to support rapid retrieval and accurate tracing based on complex, multi-level relationships, limiting their ability to provide enterprises with real-time, actionable intelligent insights. Summary of the Invention
[0004] To address the aforementioned technical problems, this paper provides a method and system for rapid retrieval of multi-level enterprise relationships based on knowledge graphs. This technical solution solves at least one of the technical problems mentioned in the background section.
[0005] To achieve the above objectives, the technical solution adopted by the present invention is as follows:
[0006] A knowledge graph-based method for rapid retrieval of multi-level enterprise relationships includes:
[0007] Collect heterogeneous data from multiple sources both inside and outside the enterprise, control the collection time, and verify the legality of the collected data;
[0008] Receive the verified data, perform data standardization processing, obtain entities and relationships between entities, build an enterprise knowledge graph, and optimize the knowledge graph association rules based on the monthly new data;
[0009] Based on knowledge graphs, the system receives retrieval requests through a query interface, parses the multi-level relationship conditions in the requests, uses a graph traversal algorithm to quickly retrieve relevant entities and relationships, and returns the retrieval results.
[0010] The search results are used for cost-benefit analysis of inter-enterprise relationships, including relationship cost accounting, relationship benefit assessment, cost-benefit relationship mining, intelligent prediction, or anomaly early warning and tracing.
[0011] Preferably, the process of collecting heterogeneous data from multiple sources both inside and outside the enterprise, controlling the collection time, and verifying the legality of the collected data specifically includes:
[0012] By parsing protocols, multi-source heterogeneous data from both inside and outside the enterprise can be identified and categorized into structured data, semi-structured data, and unstructured data based on data type.
[0013] It supports switching between real-time and timed acquisition modes. When multiple acquisition tasks are triggered simultaneously, resources are scheduled based on priority order to ensure the execution of high-priority tasks, and the execution status of each acquisition task is monitored in real time.
[0014] Create a list of key fields for the key core fields in the data, check each data entry to see if it contains the key fields in the list, and generate a structured supplementary data collection list.
[0015] Preferably, the process of receiving and verifying the data, performing data standardization processing, obtaining entities and relationships between entities, constructing an enterprise knowledge graph, and optimizing the knowledge graph association rules based on monthly new data specifically includes:
[0016] Standardize and organize inter-enterprise data from multiple sources, including equity investment data, transaction data, and cooperation relationship data, using a unified standard.
[0017] With the core enterprise as the central entity and its related shareholders, partners, suppliers, and customers as surrounding entities, the relationships between entities such as enterprise-shareholders and enterprise-transactions-suppliers are clarified to form an inter-enterprise network.
[0018] Based on the monthly increase in inter-enterprise relationship data, the original entity relationship rules in the knowledge graph are automatically optimized by the relationship patterns in the data.
[0019] Data conflicts are automatically resolved through preset rules. When a rule cannot be matched, a machine learning algorithm is invoked to determine the optimal data source based on the accuracy of historical data.
[0020] Preferably, the related cost accounting specifically includes:
[0021] Based on standardized data associated with knowledge graphs, an accounting framework covering enterprise organizational structure, time dimension, and business type is built to clarify the specific components of direct investment costs and indirect related costs;
[0022] To ensure reasonable allocation of indirectly related costs, a built-in rule template library is provided for users to define their own allocation rules, which are executed according to priority settings.
[0023] Based on standardized inter-enterprise data and cost allocation rules, real-time cost calculation is performed, and cost recalculation is automatically triggered when basic data such as equity investment details and transaction records are updated.
[0024] Cost reports are generated based on real-time cost calculation results, including a detailed cost report by organization dimension, a cost comparison report by business type dimension, and a cost trend report by time dimension.
[0025] Preferably, the assessment of the associated benefits specifically includes:
[0026] An evaluation indicator framework is constructed based on three levels of indicators: cost, output, and value. The cost level focuses on the rationality of related costs, the output level emphasizes explicit results, and the value level focuses on long-term strategic value.
[0027] By using the analytic hierarchy process (AHP) to compare the importance of each level of indicators pairwise, a judgment matrix is generated. Then, by verifying the consistency of the matrix and eliminating logical contradictions, the weight values of each indicator are output.
[0028] Synchronize the latest industry data from external databases, compare enterprise indicators with industry data, and classify the associated benefit levels, including leading, good, qualified, and lagging.
[0029] Preferably, the cost-benefit correlation mining specifically includes:
[0030] The correlation strength between each associated cost item and benefit indicator is quantified using the Pearson correlation coefficient.
[0031] By constructing a multiple linear regression model, with related cost items as independent variables and benefit indicators as dependent variables, the impact coefficient of cost on benefit is quantified.
[0032] By using a propensity score matching algorithm to eliminate confounding factors, the causal relationship between associated cost items and benefit indicators can be identified.
[0033] Preferably, the intelligent prediction specifically includes:
[0034] It incorporates time series algorithms, machine learning algorithms, and deep learning algorithms, and calculates MAE and RMSE through historical data backtesting to automatically select the optimal prediction model.
[0035] It supports three forecast scenarios: custom baseline, optimistic, and pessimistic, with different parameter configurations associated with each scenario.
[0036] Based on the selected algorithm and scenario parameters, the system outputs multi-dimensional prediction results for the future time period, including predictions of total associated costs, predictions of various benefit indicators, and predictions of the cost-benefit balance point.
[0037] Preferably, the anomaly early warning and source tracing specifically includes:
[0038] Construct a comprehensive monitoring indicator system for abnormal related costs and benefits, including abnormal cost indicators, abnormal benefit indicators, and prediction deviation indicators;
[0039] It provides a visual rule configuration interface, allowing users to customize warning trigger conditions, threshold ranges, and warning levels, and the rule configuration supports combination logic;
[0040] By tracing the root cause of anomalies based on inter-enterprise knowledge graphs, the problem nodes are located layer by layer through entity relationships, generating a tracing path diagram with arrows to show the anomaly propagation chain;
[0041] Record the entire process of early warning from triggering to closing the loop, including the person in charge, the measures taken, the results and the evaluation of the effect, and form an early warning handling knowledge base.
[0042] Furthermore, a knowledge graph-based enterprise multi-level relationship fast retrieval system is proposed to implement the aforementioned knowledge graph-based enterprise multi-level relationship fast retrieval method, including:
[0043] The data acquisition module is configured to collect heterogeneous data from multiple sources both inside and outside the enterprise, control the collection time, and verify the legality of the collected data.
[0044] The knowledge graph construction module is configured to receive verified data, perform data standardization processing, obtain entities and relationships between entities, construct an enterprise knowledge graph, and optimize the knowledge graph association rules based on monthly new data.
[0045] The retrieval processing module is configured to be based on a knowledge graph. It receives retrieval requests through a query interface, parses the multi-level relationship conditions in the request, uses a graph traversal algorithm to quickly retrieve relevant entities and relationships, and returns the retrieval results.
[0046] The analysis application module is configured to use the search results for cost-benefit analysis of inter-enterprise relationships, including relationship cost accounting, relationship benefit assessment, cost-benefit relationship mining, intelligent prediction, or anomaly early warning and tracing.
[0047] Optionally, the analysis application module includes:
[0048] The associated cost accounting unit is configured to build an accounting framework covering the enterprise's organizational structure, time dimension, and business type based on standardized data associated with the knowledge graph. It performs accounting and allocation of direct investment costs and indirect associated costs, and automatically triggers recalculation based on updates to the basic data to generate multi-dimensional cost reports.
[0049] The related benefit assessment unit is configured with an assessment framework based on three levels of indicators: cost, output, and value. The weight of each indicator is calculated using the analytic hierarchy process, and the level of related benefits is classified by comparing with industry data.
[0050] The cost-benefit correlation mining unit is configured to quantify the correlation strength, influence coefficient, and causal relationship between related cost items and benefit indicators using Pearson correlation coefficient, multiple linear regression model, and propensity score matching algorithm.
[0051] The intelligent prediction unit is equipped with multiple built-in prediction algorithms. It automatically selects the optimal model through a backtesting mechanism, supports multi-scenario parameter configuration, and outputs the prediction results of the total associated costs, benefit indicators, and cost-benefit balance point for future time periods.
[0052] The anomaly early warning and source tracing unit configures and builds an anomaly monitoring indicator system, supports user-defined early warning rules, traces the root cause of anomalies based on inter-enterprise knowledge graphs and generates a source tracing path diagram, and records the entire early warning process information to form a processing knowledge base.
[0053] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0054] By constructing a knowledge graph of multi-dimensional relationships between enterprises, it achieves unified modeling and visual analysis of complex relationships such as equity, supply chain, and cooperation. It can automatically integrate multi-source heterogeneous data and dynamically optimize association rules, significantly improving the depth and breadth of enterprise relationship mining. Combined with an intelligent cost-benefit analysis model, it supports multi-level association mining from association strength quantification and impact coefficient analysis to causal identification, and provides multi-scenario prediction and anomaly tracing functions. It effectively improves the accuracy and efficiency of enterprise relationship risk identification, investment decision support, and strategic planning, providing comprehensive data support and intelligent decision support for enterprises to manage complex business relationships. Attached Figure Description
[0055] Figure 1 This is a flowchart of the knowledge graph-based fast retrieval method for enterprise multi-level relationships proposed in this solution.
[0056] Figure 2 This is a flowchart illustrating the method for collecting heterogeneous data from multiple sources both inside and outside the enterprise, as proposed in this solution.
[0057] Figure 3 This is a flowchart illustrating the method for constructing an enterprise knowledge graph proposed in this solution. Detailed Implementation
[0058] The following description is intended to disclose the invention and enable those skilled in the art to implement it. The preferred embodiments described below are merely examples, and other obvious variations will occur to those skilled in the art.
[0059] Reference Figure 1 As shown, a fast retrieval method for enterprise multi-level relationships based on knowledge graphs includes:
[0060] By collecting heterogeneous data from multiple sources both inside and outside the enterprise, controlling the collection time, and verifying the legality of the collected data, efficient and stable collection of heterogeneous data from multiple sources is achieved through protocol parsing and priority scheduling mechanisms, ensuring the comprehensiveness and real-time nature of the data source. By establishing a list of key fields for legality verification, data quality is guaranteed from the source, providing a reliable data foundation for the subsequent construction of the knowledge graph.
[0061] After receiving and verifying the data, the system performs data standardization processing, obtains entities and relationships between entities, constructs an enterprise knowledge graph, and optimizes the association rules of the knowledge graph based on monthly new data. By processing multi-source data through unified standards, it effectively eliminates data ambiguity and conflicts, laying the foundation for accurately constructing an enterprise-centric relationship network. Through automatic monthly optimization of association rules, the knowledge graph can dynamically evolve and continuously adapt to changes in business relationships, maintaining the timeliness and accuracy of the graph.
[0062] Based on knowledge graphs, the system receives retrieval requests through a query interface, parses the multi-level relationship conditions in the requests, uses graph traversal algorithms to quickly retrieve relevant entities and relationships, and returns the retrieval results. By efficiently parsing complex multi-level relationship query requests through graph traversal algorithms, the system achieves instant retrieval and visualization of cross-level and deep relationship networks from equity structure to supply chain relationships, greatly improving the query efficiency and insight capabilities of complex enterprise relationships.
[0063] The search results are used for cost-benefit analysis of inter-enterprise relationships, including relationship cost accounting, relationship benefit assessment, cost-benefit relationship mining, intelligent prediction or anomaly early warning and source tracing. Complex graph relationship data is transformed into cost-benefit analysis that can be directly used for decision-making. Through intelligent prediction and in-depth mining, the true benefits and potential risks of related transactions, investments and other activities are revealed, providing quantitative data support for corporate strategic decision-making, risk management and value assessment.
[0064] Specifically, refer to Figure 2 As shown, the process of collecting heterogeneous data from multiple sources both inside and outside the enterprise, controlling the collection time, and verifying the legality of the collected data specifically includes:
[0065] By parsing protocols, multi-source heterogeneous data from both inside and outside the enterprise can be identified and categorized into structured data, semi-structured data, and unstructured data based on data type.
[0066] It supports switching between real-time and timed acquisition modes. When multiple acquisition tasks are triggered simultaneously, resources are scheduled based on priority order to ensure the execution of high-priority tasks, and the execution status of each acquisition task is monitored in real time.
[0067] Create a list of key fields for the key core fields in the data, check each data entry to see if it contains the key fields in the list, and generate a structured supplementary data collection list.
[0068] Specifically, the above steps are implemented as follows:
[0069] Organize commonly used data transmission protocols both internally and externally to the enterprise, and establish a standardized protocol library that includes protocols such as HTTP / HTTPS, FTP / SFTP, JDBC, REST API, SOAP API, and JSON-RPC;
[0070] Check whether the parsed data meets the characteristics of fixed fields, standardized format, and direct storage in a two-dimensional table. For example, the employee basic information table exported from the HRM system is directly classified as structured data. If the data has expandable fields and a format with certain rules but not strictly a two-dimensional table, extract the core fixed fields from the data, mark the expanded fields, and classify it as semi-structured data. For binary or text data without fixed field formats, identify them by file extension or data encoding format and classify them as unstructured data.
[0071] A three-tier priority system is established, and the importance of tasks is quantified through priority scoring. A priority-based preemptive scheduling algorithm is adopted. When multiple tasks are triggered simultaneously: high-priority tasks are given priority in allocating core resources such as CPU, memory, and network bandwidth, with a resource utilization rate of no less than 70%; if a high-priority task is completed, the released resources are allocated to medium-priority tasks; low-priority tasks only occupy the remaining resources when there are no high- or medium-priority tasks. If a high-priority task is triggered, the low-priority task is immediately paused and the resources are released.
[0072] The data validity verification process is as follows: Step 1: Non-empty verification: Traverse each piece of collected data and check if it contains any "non-empty fields" from the key field list. If a non-empty field is missing, mark the data as "invalid" and record the name of the missing field. Step 2: Format and type verification: For existing key fields, verify whether they conform to the "field rule definition": Data type verification: For example, if the pre-tax salary field is a string, it does not meet the "numeric type" requirement and is marked as invalid; Format verification: For example, if the ID number field is 17 or 19 digits, it does not meet the "18-character" requirement and is marked as invalid. Step 3: Uniqueness verification: For fields with a "unique identifier" attribute, query historical collected data or the database. If duplicate values exist, mark the data as "invalid" and record the duplicate fields and duplicate values.
[0073] Reference Figure 3 As shown, the process involves receiving verified data, performing data standardization, obtaining entities and relationships between entities, constructing an enterprise knowledge graph, and optimizing the knowledge graph association rules based on monthly new data. Specifically, this includes:
[0074] Standardize and organize inter-enterprise data from multiple sources, including equity investment data, transaction data, and cooperation relationship data, using a unified standard.
[0075] With the core enterprise as the central entity and its related shareholders, partners, suppliers, and customers as surrounding entities, the relationships between entities such as enterprise-shareholders and enterprise-transactions-suppliers are clarified to form an inter-enterprise network.
[0076] Based on the monthly increase in inter-enterprise relationship data, the original entity relationship rules in the knowledge graph are automatically optimized by the relationship patterns in the data.
[0077] Data conflicts are automatically resolved through preset rules. When a rule cannot be matched, a machine learning algorithm is invoked to determine the optimal data source based on the accuracy of historical data.
[0078] Specifically, the newly added inter-enterprise relationship data each month, such as new investment relationships, transaction data, and cooperative relationship updates, are standardized to extract new entities and attributes; "frequent itemset mining" is used to identify entity association patterns in the new data, calculate the support and confidence of association patterns, and screen high-confidence patterns.
[0079] Use new data to verify the effectiveness of existing association rules in the graph. If the rule matching rate is ≥90%, the existing rule is effective, and the matching status is retained and recorded. If the rule matching rate is 70% ≤ rule matching rate <90%, analyze the unmatched data and supplement the rule conditions. If the rule matching rate is <70%, the existing rule is invalid, and the original rule is replaced based on the high-confidence association patterns mined.
[0080] The conflict types are divided into two core conflict categories: Attribute conflict: The same attribute of the same entity has different values, such as the shareholding ratio of company "A" being "60%" in the industrial and commercial system and "55%" in its own disclosure report; Relationship conflict: The related relationship of the same entity is contradictory, such as "B company" being shown as "a subsidiary of company C" in the equity system, but as "a competitor of company C" in the bidding system.
[0081] Extract the feature dimensions of conflicting data, including data source credibility score, data update time difference, historical data accuracy, and data format integrity; use historical data as samples to train a data credibility assessment model, calculate the comprehensive credibility score of each conflicting data, select the data with the highest comprehensive credibility score as the final value, and if there are multiple data with the highest score, trigger the manual review process.
[0082] Specifically, in some preferred embodiments, rapid retrieval of multi-level enterprise relationships based on knowledge graphs to support enterprise association cost accounting includes the following steps:
[0083] Based on standardized data associated with knowledge graphs, an accounting framework covering enterprise organizational structure, time dimension, and business type is built to clarify the specific components of direct investment costs and indirect related costs;
[0084] To ensure reasonable allocation of indirectly related costs, a built-in rule template library is provided for users to define their own allocation rules, which are executed according to priority settings.
[0085] Based on standardized inter-enterprise data and cost allocation rules, real-time cost calculation is performed, and cost recalculation is automatically triggered when basic data such as equity investment details and transaction records are updated.
[0086] Cost reports are generated based on real-time cost calculation results, including a detailed cost report by organization dimension, a cost comparison report by business type dimension, and a cost trend report by time dimension.
[0087] Specifically, based on business needs, three major accounting dimensions are defined:
[0088] Organizational dimension: divided by corporate group, subsidiary, and business unit; Time dimension: divided by accounting period (monthly, quarterly, annual); Business relationship dimension: divided by related business type (equity investment, supply chain, strategic cooperation, etc.);
[0089] Dismantling cost components:
[0090] Definition of direct cost components: These are costs that can be directly attributed to specific related parties, including: direct investment costs (equity acquisition consideration, capital increase); transaction-related costs (raw material procurement costs, product distribution costs); cooperation performance costs (project cooperation deposits, technology licensing fees); and related taxes and fees (equity transaction taxes and fees, cross-border related transaction taxes and fees).
[0091] Definition of indirect cost components: Identify related costs that cannot be directly attributed and need to be allocated, including: relationship maintenance costs (allocation of due diligence fees, legal fees, and audit and valuation fees); management support costs (allocation of group management fees and cross-border transaction management fees); risk reserves (allocation of related-party transaction bad debt reserves and investment impairment reserves); and other indirect costs (allocation of strategic cooperation forum expenses and inter-company cultural exchange activity expenses).
[0092] A hierarchical model is established using dimensions and cost components as the core framework:
[0093] First level: Total related costs;
[0094] The second layer: Sub-total costs broken down by organizational dimension, time dimension, and business relationship dimension;
[0095] The third level: direct and indirect related costs under each sub-total cost;
[0096] The fourth level: Specific components of directly related costs and indirectly related costs.
[0097] In some preferred embodiments, rapid retrieval of multi-level enterprise relationships based on knowledge graphs is used to support the assessment of enterprise association benefits, including the following steps:
[0098] An evaluation indicator framework is constructed based on three levels of indicators: cost, output, and value. The cost level focuses on the rationality of related costs, the output level emphasizes explicit results, and the value level focuses on long-term strategic value.
[0099] By using the analytic hierarchy process (AHP) to compare the importance of each level of indicators pairwise, a judgment matrix is generated. Then, by verifying the consistency of the matrix and eliminating logical contradictions, the weight values of each indicator are output.
[0100] Synchronize the latest industry data from external databases, compare enterprise indicators with industry data, and classify the associated benefit levels, including leading, good, qualified, and lagging.
[0101] Specifically, based on the indicator system framework, a three-tiered structure of "target layer - criterion layer - indicator layer" is constructed:
[0102] Target level: Comprehensive evaluation of enterprise benefits;
[0103] Criterion layer: cost layer, output layer, value layer;
[0104] Indicator layer: Specific indicators under each criterion layer;
[0105] Invite 5-10 experts to conduct pairwise comparisons of the importance of indicators within the same level, assign values using a 1-9 scale, and form a judgment matrix;
[0106] Normalize each column of the judgment matrix, sum the rows of the normalized matrix, and calculate the product of the judgment matrix and the summation result. Calculate the largest eigenvalue. Obtain the consistency ratio by calculating the consistency index and searching for the random consistency index. If the consistency ratio is less than 0.1, the judgment matrix meets the consistency requirements and can proceed to weight calculation; otherwise, the judgment matrix has logical contradictions.
[0107] The formula for calculating the consistency index is:
[0108]
[0109] in, As a consistency indicator, It is the largest eigenvalue. To determine the order of a matrix, The smaller the value, the better the consistency of the judgment matrix. Calculating the consistency index provides basic data for subsequent consistency checks. If the value is too large, it indicates a significant contradiction in the judgment matrix, which needs to be adjusted first.
[0110] The formula for the consistency ratio is:
[0111]
[0112] in, As a consistency indicator, To find a random consistency index, The consensus ratio, whose core function is to provide a decision-making criterion, is when... When <0.1, the inconsistency of the judgment matrix is considered to be within an acceptable range; when When the value is ≥0.1, it indicates that the contradiction in the judgment matrix has exceeded the reasonable range and the judgment matrix needs to be adjusted.
[0113] In some preferred embodiments, rapid retrieval of multi-level enterprise relationships based on knowledge graphs is used to support cost-effective association mining for enterprises, including the following steps:
[0114] The correlation strength between each associated cost item and benefit indicator is quantified using the Pearson correlation coefficient.
[0115] By constructing a multiple linear regression model, with related cost items as independent variables and benefit indicators as dependent variables, the impact coefficient of cost on benefit is quantified.
[0116] By using a propensity score matching algorithm to eliminate confounding factors, the causal relationship between associated cost items and benefit indicators can be identified.
[0117] Specifically, each related cost item is paired with each benefit indicator to form multiple analytical pairs. The direction and strength of the correlation are calculated: for each analytical pair, the Pearson correlation coefficient is used to determine the direction of the correlation between the two. Positive correlation: when the related cost increases, the benefit increases simultaneously; negative correlation: when the related cost increases, the benefit decreases. The correlation strength is quantified: the closer the absolute value of the coefficient is to 1, the stronger the correlation; the closer it is to 0, the weaker the correlation. Analytical pairs with an absolute correlation coefficient greater than 0.3 are selected and sorted from high to low correlation strength, prioritizing cost-benefit combinations with strong correlation.
[0118] A multiple linear regression model is constructed using historical data. The influence coefficients of each associated cost item on the benefit index are determined through the training set. The model's predictive performance is tested using validation set data. If the error between the predicted value and the actual benefit value is within an acceptable range, the model is effective. If the error is large, the data needs to be cleaned again or the variables adjusted during the data preparation stage. The influence coefficients of each associated cost item are analyzed. A positive coefficient with a large value indicates that the associated cost item has a strong positive driving effect on the benefit.
[0119] Using a propensity score matching algorithm, a control group of companies (such as those that have made specific equity investments) that are similar in terms of confounding factors such as size, industry, and financial status but have not made such investments are selected and processed. The benefit indicators of the two matched groups are compared. If the benefit difference between the two groups is significant, it indicates that there is a causal relationship between the related cost item and the benefit indicator. By adjusting the calculation parameters of the propensity score, the matching and comparison process is repeated. If multiple results show that there is a significant causal relationship between the related cost item and the benefit, the causal relationship is finally confirmed. If the results are unstable, further data on confounding factors need to be added.
[0120] In some preferred embodiments, rapid retrieval of multi-level enterprise relationships based on knowledge graphs to support intelligent prediction for enterprises includes the following steps:
[0121] It incorporates time series algorithms, machine learning algorithms, and deep learning algorithms, and calculates MAE and RMSE through historical data backtesting to automatically select the optimal prediction model.
[0122] It supports three forecast scenarios: custom baseline, optimistic, and pessimistic, with different parameter configurations associated with each scenario.
[0123] Based on the selected algorithm and scenario parameters, the system outputs multi-dimensional prediction results for the future time period, including predictions of total associated costs, predictions of various benefit indicators, and predictions of the cost-benefit balance point.
[0124] Specifically, for each algorithm and its specific model, backtesting is performed sequentially. The model is trained using the training dataset, and then the model outputs prediction results using the backtesting dataset. The prediction results are compared with the actual values in the backtesting dataset. For each model's prediction results, two error metrics are calculated: MAE (Mean Absolute Error, reflecting the average deviation between the predicted and actual values) and RMSE (Root Mean Square Error, more sensitive to larger deviations and highlighting extreme errors). The values of these two error metrics for each model are recorded. All models are sorted by their MAE and RMSE values from smallest to largest, and the model with the smallest values for both errors is selected first. If a model has the smallest MAE but a large RMSE, a decision is made based on business requirements.
[0125] Import the three scenario parameters (baseline, optimistic, and pessimistic) determined by the scenario setting unit into the optimal model selected by the algorithm selection unit, start the model calculation, and generate prediction data for the corresponding scenario. Repeat the prediction process for each scenario 2-3 times. If the deviation of the multiple prediction results is less than 3%, it indicates that the prediction is stable. If the deviation is large, return to the scenario setting unit to check the rationality of the parameters, or return to the algorithm selection unit to re-verify the model.
[0126] In some preferred embodiments, rapid retrieval of multi-level enterprise relationships based on knowledge graphs is used to support enterprise anomaly early warning and tracing, including the following steps:
[0127] Construct a comprehensive monitoring indicator system for abnormal related costs and benefits, including abnormal cost indicators, abnormal benefit indicators, and prediction deviation indicators;
[0128] It provides a visual rule configuration interface, allowing users to customize warning trigger conditions, threshold ranges, and warning levels, and the rule configuration supports combination logic;
[0129] By tracing the root cause of anomalies based on inter-enterprise knowledge graphs, the problem nodes are located layer by layer through entity relationships, generating a tracing path diagram with arrows to show the anomaly propagation chain;
[0130] Record the entire process of early warning from triggering to closing the loop, including the person in charge, the measures taken, the results and the evaluation of the effect, and form an early warning handling knowledge base.
[0131] Specifically, the core entities in the related fields between enterprises are sorted out, the relationships between entities are defined, a structured knowledge network is formed, and abnormal data of early warning indicators are connected to the knowledge graph to automatically match the corresponding entity nodes;
[0132] First-level source tracing (locating direct abnormal nodes): For abnormal indicators, extract directly related entities from the knowledge graph, compare the actual data of each node with the warning threshold, and lock in the primary abnormal node, such as identifying abnormal fluctuations in specific equity investment relationships or transactions.
[0133] Secondary source tracing (mining indirect influencing factors): Based on the knowledge graph relationship, trace down from the primary anomaly node, analyze the data of related entities such as shareholders, suppliers, and customers, and locate the indirect root cause of the anomaly, such as the transmission of risk from upstream suppliers to the core enterprise.
[0134] Once an alert is triggered, the system automatically records the alert number, trigger time, abnormal indicators, and alert level. Based on the responsible department configured according to the rules (such as the investment management department or risk control department), the system notifies the person in charge through a system message, recording the assignment time and the name of the person in charge.
[0135] Furthermore, a rapid retrieval system for enterprise multi-level relationships based on knowledge graphs is proposed, including:
[0136] The data acquisition module is configured to collect heterogeneous data from multiple sources both inside and outside the enterprise, control the collection time, and verify the legality of the collected data.
[0137] The knowledge graph construction module is configured to receive verified data, perform data standardization processing, obtain entities and relationships between entities, construct an enterprise knowledge graph, and optimize the knowledge graph association rules based on monthly new data.
[0138] The retrieval processing module is configured to be based on a knowledge graph. It receives retrieval requests through a query interface, parses the multi-level relationship conditions in the request, uses a graph traversal algorithm to quickly retrieve relevant entities and relationships, and returns the retrieval results.
[0139] The analysis application module is configured to use the search results for cost-benefit analysis of inter-enterprise relationships, including relationship cost accounting, relationship benefit assessment, cost-benefit relationship mining, intelligent prediction, or anomaly early warning and tracing.
[0140] The analysis application module includes:
[0141] The associated cost accounting unit is configured to build an accounting framework covering the enterprise's organizational structure, time dimension, and business type based on standardized data associated with the knowledge graph. It performs accounting and allocation of direct investment costs and indirect associated costs, and automatically triggers recalculation based on updates to the basic data to generate multi-dimensional cost reports.
[0142] The related benefit assessment unit is configured with an assessment framework based on three levels of indicators: cost, output, and value. The weight of each indicator is calculated using the analytic hierarchy process, and the level of related benefits is classified by comparing with industry data.
[0143] The cost-benefit correlation mining unit is configured to quantify the correlation strength, influence coefficient, and causal relationship between related cost items and benefit indicators using Pearson correlation coefficient, multiple linear regression model, and propensity score matching algorithm.
[0144] The intelligent prediction unit is equipped with multiple built-in prediction algorithms. It automatically selects the optimal model through a backtesting mechanism, supports multi-scenario parameter configuration, and outputs the prediction results of the total associated costs, benefit indicators, and cost-benefit balance point for future time periods.
[0145] The anomaly early warning and source tracing unit configures and builds an anomaly monitoring indicator system, supports user-defined early warning rules, traces the root cause of anomalies based on inter-enterprise knowledge graphs and generates a source tracing path diagram, and records the entire early warning process information to form a processing knowledge base.
[0146] In summary, the advantages of this invention are as follows: By constructing a knowledge graph of multi-dimensional relationships between enterprises, it achieves unified modeling and visual analysis of complex relationships such as equity, supply chain, and cooperation. It can automatically integrate multi-source heterogeneous data and dynamically optimize association rules, significantly improving the depth and breadth of enterprise relationship mining. Combined with an intelligent cost-benefit analysis model, it supports multi-level association mining from association strength quantification and influence coefficient analysis to causal identification, and provides multi-scenario prediction and anomaly tracing functions. This effectively improves the accuracy and efficiency of enterprise relationship risk identification, investment decision support, and strategic planning, providing comprehensive data support and intelligent decision support for enterprises to manage complex business relationships.
[0147] The foregoing has shown and described the basic principles, main features, and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited to the above embodiments. The embodiments and descriptions in the specification are merely principles of the invention. Various changes and modifications can be made to the invention without departing from its spirit and scope, and all such changes and modifications fall within the scope of the claimed invention. The scope of protection claimed by the appended claims and their equivalents is defined.
Claims
1. A method for quickly searching multi-level relations of enterprises based on a knowledge graph, characterized in that, The method comprises the following steps: Collecting multi-source heterogeneous data inside and outside the enterprise, controlling the time of collection, and verifying the legality of the collected data; Receiving the verified data, performing data standardization processing, obtaining entities and relationships between entities, constructing an enterprise knowledge graph, and optimizing the knowledge graph association rules based on the monthly new data; Based on the knowledge graph, receive the search request through the query interface, analyze the multi-level relationship conditions in the request, use the graph traversal algorithm to quickly search for related entities and relationships, and return the search results; The search results are used for cost-benefit analysis of inter-enterprise association, including association cost accounting, association benefit evaluation, cost-benefit association mining, intelligent prediction or abnormal early warning traceability. 2.The method of claim 1, wherein, The collection of multi-source heterogeneous data inside and outside the enterprise, the control of the collection time, and the verification of the legality of the collected data specifically includes: Identify multi-source heterogeneous data inside and outside the enterprise by analyzing the protocol, and divide the data based on data type into structured data, semi-structured data, and unstructured data; Support real-time collection and timed collection switching, when multiple collection tasks are triggered at the same time, schedule resources based on priority order, prioritize high-priority task execution, and monitor the execution status of each collection task in real time; Establish a list of key fields for key core fields in the data, check each piece of data for the presence of key fields in the list one by one, and generate a structured supplementary collection list. 3.The method of claim 2, wherein, The receiving of the verified data, the execution of the data standardization processing, the obtaining of the entities and the relationships between the entities, the construction of the enterprise knowledge graph, and the optimization of the knowledge graph association rules based on the monthly new data specifically includes: Standardize multi-source inter-enterprise association data based on a unified standard, including equity investment data, transaction data, and cooperation relationship data; Take the core enterprise as the central entity, and take the associated shareholders, partners, suppliers and customers of the core enterprise as the peripheral entities, clearly define the entity-ownership-shareholder, enterprise-transaction-supplier and other entity association relationships, and form an inter-enterprise association network; Based on the monthly new inter-enterprise association data, automatically optimize the original entity association rules in the knowledge graph through the association mode in the data; Automatically solve data conflicts through pre-set rules, and when the rules cannot be matched, call the machine learning algorithm to determine the optimal data source based on the historical data accuracy. 4.The method of claim 1, wherein, The association cost accounting specifically includes: Based on the standardized data of the knowledge graph association, build a calculation framework covering the enterprise organizational structure, time dimension, and business type, and clearly define the specific composition of direct investment cost and indirect association cost; In order to reasonably allocate indirect association costs, a rule template library is built to allow users to customize allocation rules, which are executed according to priority; Based on the standardized inter-enterprise association data and cost allocation rules, perform real-time cost calculation, and automatically trigger cost recalculation when basic data such as equity investment details and transaction records are updated; Based on the real-time cost calculation results, generate a cost report, including an organization dimension cost detail table, a business type dimension cost comparison table, and a time dimension cost trend table. 5.The method of claim 1, wherein, The association benefit evaluation specifically includes: An evaluation index framework is constructed based on three layers of indexes, i.e. cost, output and value. The cost layer focuses on the rationality of associated costs, the output layer emphasizes on explicit achievements, and the value layer pays attention to long-term strategic value. The importance of each level index is compared with each other by using the analytic hierarchy process to generate a judgment matrix, and the consistency of the matrix is checked to exclude logical contradictions, and the weight value of each index is output. The latest industry data is synchronized from an external database, the enterprise indexes are compared with the industry data, and the associated benefit level is divided, including leading, good, qualified and lagging. 6.The method of claim 1, wherein, The cost-benefit association mining specifically includes: The correlation between each associated cost item and the benefit index is quantified by Pearson correlation. A multiple linear regression model is constructed with the associated cost item as the independent variable and the benefit index as the dependent variable to quantify the influence coefficient of cost on benefit. The propensity score matching algorithm is used to exclude the interference of confounding factors to identify the causal relationship between the associated cost item and the benefit index. 7.The method of claim 1, wherein, The intelligent prediction specifically includes: Built-in time series algorithm, machine learning algorithm and deep learning algorithm are used to calculate MAE and RMSE through historical data backtesting mechanism to automatically select the optimal prediction model. Support for customizing three prediction scenarios, i.e. benchmark, optimistic and pessimistic, each scenario is associated with different parameter configurations. Based on the selected algorithm and scenario parameters, multi-dimensional prediction results in the future time period are output, including associated cost total amount prediction, each benefit index prediction and cost-benefit balance point prediction. 8.The method of claim 1, wherein, The abnormal early warning tracing specifically includes: A comprehensive associated cost-benefit abnormal monitoring index system is constructed, including cost abnormal index, benefit abnormal index and prediction deviation index. A visual rule configuration interface is provided to support user-defined early warning trigger conditions, threshold range and early warning level, and the rule configuration supports combination logic. Based on the associated knowledge graph between enterprises, the abnormal root cause is traced, the problem node is located layer by layer through entity association relationship, an arrowed traceability path graph is generated, and the abnormal transmission chain is displayed. The whole process information from early warning triggering to closing loop is recorded, including the handler, handling measures, handling results and effect evaluation, forming an early warning handling knowledge base.
9. A knowledge graph-based fast search system for multi-level relationships of enterprises, characterized in that, The method for quickly retrieving multi-level relationships based on a knowledge graph of an enterprise according to any one of claims 1-8 comprises: A data acquisition module configured to acquire multi-source heterogeneous data inside and outside the enterprise, control the acquisition time, and perform legality verification on the acquired data; A knowledge graph construction module configured to receive the verified data, perform data standardization processing, obtain entities and relationships between entities, construct an enterprise knowledge graph, and optimize the knowledge graph association rules based on monthly new data; A retrieval processing module configured to receive a retrieval request through a query interface based on the knowledge graph, parse the multi-level relationship conditions in the request, quickly retrieve related entities and relationships using a graph traversal algorithm, and return the retrieval results; An analysis and application module configured to use the retrieval results for cost-benefit analysis of the association between enterprises, including associated cost accounting, associated benefit evaluation, cost-benefit association mining, intelligent prediction or abnormal early warning tracing. 10.The knowledge graph based enterprise multi-level relationship quick search system according to claim 9, characterized in that, The analysis and application module includes: The associated cost accounting unit is configured to build an accounting framework covering enterprise organizational structure, time dimension and business type based on standardized data associated with the knowledge graph, to account and allocate direct investment costs and indirect associated costs, and to automatically trigger re-accounting based on basic data updates to generate multi-dimensional cost reports; The associated benefit evaluation unit is configured to build an evaluation framework based on three layers of cost, output and value indicators, to calculate the weight of each indicator using the analytic hierarchy process, and to divide the associated benefits into grades by comparing with industry data; The cost-benefit association mining unit is configured to quantify the correlation strength, influence coefficient and causal relationship between associated cost items and benefit indicators through Pearson correlation coefficient, multiple linear regression model and propensity score matching algorithm; The intelligent prediction unit is configured with multiple built-in prediction algorithms, automatically selects the optimal model through backtesting mechanism, supports multi-scenario parameter configuration, and outputs the prediction results of associated cost total, benefit indicators and cost-benefit balance point in the future time period; The abnormal early warning tracing unit is configured to build an abnormal monitoring indicator system, support user-defined early warning rules, trace the root cause of abnormalities based on the associated knowledge graph between enterprises and generate a trace path diagram, and record the whole process information of early warning to form a processing knowledge base.
Citation Information
Patent Citations
Knowledge graph construction method and device based on structural data
CN104462501A
Graph-database-based real-time police analysis application platform and construction method therefor
CN104850601A
Method for apportioning costs of real estate enterprises
CN108305191A
Low-altitude maneuvering target tracking method based on multiple interactive models
CN114966667A
International oil price prediction method combining scene analysis and neural network algorithm
CN119205168A