Enterprise global data analysis method based on knowledge graph and large language model
By combining the two-way inference fusion of knowledge graphs and large language models, natural language intention analysis and adaptive knowledge completion technologies, multiple limitations in enterprise global multi-source heterogeneous data analysis are solved, and data analysis and decision support with high accuracy and transparency are achieved.
Patent Information
- Application Number
- CN202510694911.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-28
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2045-05-28
AI Technical Summary
The existing technology has limitations in the integration, reasoning and intelligent analysis of multi-source heterogeneous data in the entire enterprise domain, including single-path reasoning, insufficient natural language intention analysis, lack of real-time conflict detection and knowledge completion, weak semantic alignment capabilities of multi-source data, poor domain adaptability and weak interpretability, and lack of inference process traceability mechanism.
A method of enterprise whole-domain data analysis based on knowledge graph and large language model is proposed. Through two-way inference fusion, natural language intention analysis, adaptive knowledge completion, cross-origin data semantic alignment and inference chain traceability mechanism, the intelligent integration and interpretability analysis of enterprise whole-domain data are realized.
It improves the accuracy of reasoning, data fusion intelligence and decision-making transparency, enhances the data integration and intelligent reasoning capabilities of enterprises in a multi-source heterogeneous data environment, and supports the reliability and transparency of enterprise intelligent decision-making.
Smart Images

Figure CN120218256A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data analysis, and particularly to an enterprise-wide data analysis method based on a knowledge graph and a large language model. Background Art
[0002] With the rapid development of information technology and data intelligence technology, enterprises' dependence on multi-source data in the process of operation and decision-making has been continuously deepening. Especially in the context of the increasing demand for large-scale data integration and intelligent analysis, how to effectively integrate heterogeneous data from enterprise internal systems, external platforms, and third-party resources has become an important issue for improving the intelligent level of enterprise decision-making. As a means of structuring and expressing knowledge, the knowledge graph, and as a tool with powerful natural language understanding and reasoning capabilities, the large language model, the combination of the two provides new possibilities for the intelligent processing of enterprise data.
[0003] However, in the prior art, the methods for fusing, reasoning, and intelligent analysis of enterprise-wide multi-source heterogeneous data still have obvious limitations, mainly reflected in the following aspects: 1. Single reasoning chain: Existing methods mostly rely on single-path reasoning in the knowledge graph or language model, lacking the bidirectional integration of symbolic reasoning and semantic reasoning, resulting in insufficient accuracy and robustness of the reasoning results.
[0004] 2. Insufficient parsing of natural language intentions: Traditional methods based on structured query language cannot effectively parse complex natural language query intentions, resulting in limited generation of reasoning paths, poor flexibility, and inability to adapt to the diverse query needs of users.
[0005] 3. Lack of real-time conflict detection and knowledge completion mechanisms during the reasoning process: Existing reasoning processes usually perform unified verification after the reasoning ends, and cannot dynamically identify and complete semantic conflicts and logical omissions during the reasoning process, affecting the integrity and reliability of the reasoning chain.
[0006] 4. Weak semantic alignment ability for multi-source data: When facing multi-source heterogeneous data inside and outside the enterprise, existing methods often have difficulty effectively completing entity alignment and semantic fusion, resulting in fragmented and poorly correlated data analysis results.
[0007] 5. Poor domain adaptability and weak interpretability: When applied to specific enterprise domains, the reasoning results of general language models or graph reasoning engines often lack pertinence and are poorly interpretable, making it difficult to meet the requirements of enterprise audit compliance and decision-making transparency.
[0008] 6. Lack of a reasoning process traceability mechanism: Traditional reasoning systems generally ignore the complete records of the reasoning chain path, the evolution of reasoning nodes, and the changes in reasoning relationships, making it difficult to support subsequent decision verification and problem tracing requirements.
[0009] Therefore, how to provide an enterprise-wide data analysis method based on a knowledge graph and a large language model is an urgent problem to be solved by those skilled in the art. Summary of the Invention
[0010] An object of the present invention is to propose an enterprise-wide data analysis method based on a knowledge graph and a large language model. The present invention combines structured knowledge representation and natural language understanding and reasoning capabilities, and details how to achieve intelligent integration and interpretable analysis of enterprise-wide data through bidirectional reasoning fusion, natural language intent parsing, adaptive knowledge completion, cross-source data semantic alignment, and reasoning chain tracing mechanisms, with the advantages of high reasoning accuracy, strong data fusion intelligence, and good decision transparency.
[0011] According to an embodiment of the present invention, the enterprise-wide data analysis method based on a knowledge graph and a large language model includes the following steps: S1. Collect enterprise internal data sources and enterprise external data sources, perform data cleaning, normalization, and entity mapping processing to generate an enterprise initial knowledge graph, where the enterprise initial knowledge graph includes knowledge graph entity nodes, knowledge graph attribute nodes, and knowledge graph relationship edges; S2. Receive a natural language query request, use the large language model to perform semantic parsing processing on the natural language query request, extract query intent, query target entity, and query associated relationship elements, and generate a query intent graph and reasoning task description information; S3. Invoke the enterprise initial knowledge graph, perform symbolic reasoning processing based on the query intent graph and reasoning task description information to generate a first reasoning chain, and perform semantic reasoning processing in combination with the large language model to generate a second reasoning chain; S4. Perform reasoning chain fusion processing on the first reasoning chain and the second reasoning chain, verify according to the reasoning path structure similarity index and the reasoning conclusion consistency index, and adopt a dynamic iterative optimization mechanism to generate a fused and optimized reasoning chain; S5. Based on the fused and optimized reasoning chain, perform local conflict detection processing and adaptive knowledge completion processing, and output a corrected reasoning chain; S6. Extract reasoning entity nodes and reasoning relationship edges according to the corrected reasoning chain, perform semantic alignment processing and intelligent association processing on enterprise internal data sources and enterprise external data sources, and construct an enterprise unified dynamic optimization data analysis view; S7. Based on the enterprise unified dynamic optimization data analysis view, output enterprise-wide data analysis results and generate a reasoning chain traceability file.
[0012] Optionally, S2 includes the following steps: S21. Receive the natural language query request input by the enterprise user, perform clause splitting and word segmentation processing on the natural language query request, and generate a set of natural language parsing fragments; S22. Use the large language model to perform semantic classification processing on the set of natural language parsing fragments, classify the set of natural language parsing fragments into query intention fragments, query target entity fragments, and query association relationship fragments, and label the query intention category label, query target entity category label, and query association relationship category label respectively; S23. For the query target entity fragment, perform entity mapping retrieval in the enterprise initial knowledge graph to obtain a set of entity matching results of the knowledge graph entity nodes, and assign entity matching confidence to each knowledge graph entity node; S24. Based on the query intention fragment, query target entity fragment, and query association relationship fragment, construct a query intention graph. The query intention graph includes query intention layer nodes, query target entity layer nodes, and query association relationship layer edges, and the association relationships between them are labeled according to semantic relevance and context consistency; S25. Set any node in the query intention graph and the comprehensive association weight between nodes . The comprehensive association weight is calculated according to the following formula: Among them, represents the semantic correlation function between node and node , represents the context consistency score between node and node , represents the shortest path length between node and node in the enterprise initial knowledge graph, and are comprehensive association weight adjustment coefficients; S26. Generate inference task description information based on the node hierarchy of the query intention graph, the set of query association relationship layer edges, and the comprehensive association weight matrix. The inference task description information includes a set of query intention layer nodes, a set of query target entity layer nodes, a set of query association relationship layer edges, a comprehensive association weight matrix, and inference priority sorting information.
[0013] Optionally, the S24 includes the following steps: S241. Based on the query intention fragment, query target entity fragment, and query association relationship fragment, respectively extract the query intention element set, query target entity element set, and query association relationship element set, and generate a query intention graph; S242. Map each query intention element in the query intention element set, query target entity element set, and query association relationship element set to a query intention layer node, a query target entity layer node, and a query association relationship layer edge respectively; S243. Based on the nodes and edges in the query intention graph, establish semantic associations between each query intention layer node in the query intention graph and one or more corresponding query target entity layer nodes, and connect each query target entity layer node in the query intention graph to other query target entity layer nodes according to the query association relationship layer edges; S244. In the query intention graph, set initial attributes for each query intention layer node, query target entity layer node, and query association relationship layer edge respectively. The initial attributes include node category identification, node semantic vector representation, node context vector representation, and edge association relationship category identification; S245. For any query intention layer node and query target entity layer node in the query intention graph and node , based on the node semantic vector representation of node and the node semantic vector representation of node , calculate the semantic relevance function between node and node ; S246. Based on the node context vector representation of node and the node context vector representation of node , calculate the context consistency score between node and node .
[0014] Optionally, the S3 includes the following steps: S31. Based on the inference task description information, call the enterprise initial knowledge graph, and locate the corresponding knowledge graph entity nodes in the enterprise initial knowledge graph according to each query target entity node in the query target entity layer node set, and generate a knowledge graph entity node set; S32. Based on the knowledge graph entity node set and the query association relationship layer edge set, perform symbolic reasoning processing along the knowledge graph relationship edges in the enterprise initial knowledge graph to generate a symbolic reasoning path set, where each symbolic reasoning path starts and ends with a knowledge graph entity node and meets the semantic requirements of the query association relationship layer edges; S33. Perform screening processing on the symbolic reasoning path set according to the comprehensive association weight matrix, retain the symbolic reasoning paths with comprehensive association weight values greater than the preset weight threshold, and sort the screened symbolic reasoning paths according to the reasoning priority sorting information to generate a first inference chain; S34. Input the inference task description information into the large language model, perform semantic inference processing based on the query intent layer node set, the query target entity layer node set, and the query association relationship layer edge set, and derive and generate a semantic inference path set; S35. For each semantic inference path in the semantic inference path set, adjust the inference association order of the inference nodes in the semantic inference path according to the node hierarchy structure of the query intent graph, and calculate the semantic inference path score; S36. According to the comprehensive association weight matrix and the semantic inference path score, perform a scoring process on the semantic inference path set, filter out the semantic inference paths with a scoring value greater than the preset scoring threshold, and sort them according to the inference priority sorting information to generate a second inference chain.
[0015] Optionally, S34 includes the following steps: S331. Traverse each symbolic inference path in the symbolic inference path set, and for each adjacent node pair in the symbolic inference path, extract the corresponding comprehensive association weight value from the comprehensive association weight matrix; S332. For each symbolic inference path, calculate the comprehensive weight cumulative value of the symbolic inference path based on the comprehensive association weight value , and the comprehensive weight cumulative value of each symbolic inference path is compared with the preset weight threshold , and the symbolic inference paths with a comprehensive weight cumulative value greater than the preset weight threshold are filtered out to form a symbolic inference path screening set; S333. For each symbolic inference path in the symbolic inference path screening set, extract the inference priority sorting information in the inference task description information, and perform a priority sorting process on the symbolic inference paths in the symbolic inference path screening set according to the inference priority sorting information; S334. According to the symbolic inference path set sorted by inference priority, connect the symbolic inference paths in sequence according to the priority order to generate a first inference chain.
[0016] Optionally, S4 includes the following steps: S41. Extract the inference path sets in the first inference chain and the second inference chain, and mark the start nodes and end nodes of each inference path in the inference path sets respectively; S42. For any two inference paths in the first inference chain and the second inference chain, calculate the inference path structure similarity index , and the inference path structure similarity index is calculated based on the node sequence matching degree of the two inference paths, and is calculated using the following formula: ; Among them, represents the total number of nodes in two inference paths, and respectively represent the number of nodes in two inference paths; S43. For any two inference paths in the first inference chain and the second inference chain, calculate the inference conclusion consistency index , and the inference conclusion consistency index is determined according to the consistency scores of the target entities and associated relationships derived from two inference paths; S44. Based on the inference path structure similarity index and the inference conclusion consistency index , filter out the inference path pairs that meet the conditions that the similarity index is greater than the preset structure similarity threshold and the consistency index is greater than the preset consistency threshold, and form an inference path matching pair set; S45. Based on the inference path matching pair set, adopt a dynamic iterative optimization mechanism, and according to the inference priority and path importance index, adjust the connection order of inference paths and the node merging strategy round by round to generate a fused and optimized inference chain.
[0017] Optionally, S44 includes the following steps: S441. Set a preset structure similarity threshold , and the preset structure similarity threshold is set according to the statistical data of the historical inference path node sequence matching degree and is used as the benchmark condition for filtering the inference path structure similarity; S442. Set a preset consistency threshold , and the preset consistency threshold is set according to the weighted calculation result of the inference target entity consistency score and the inference associated relationship consistency score and is used as the benchmark condition for filtering the inference conclusion consistency; S443. For each pair of inference paths in the first inference chain and the second inference chain, if the inference path structure similarity index is greater than or equal to the preset structure similarity threshold , and the inference conclusion consistency index is greater than or equal to the preset consistency threshold , then add this inference path pair to the inference path matching pair set; S444. Perform a deduplication process on the inference path matching pair set to eliminate the duplicate inference path pairs with exactly the same start nodes, end nodes, and intermediate nodes, and retain the unique inference path matching pairs.
[0018] Optionally, S5 includes the following steps: S51. Receive the inference chain after fusion optimization, and extract the set of inference paths and the set of inference nodes in the inference chain after fusion optimization; S52. For the set of inference paths in the inference chain after fusion optimization, detect the logical conflicts existing within each inference path and between different inference paths in the set of inference paths. The logical conflicts include inference direction conflicts, entity attribute conflicts, and association relationship conflicts; S53. For each detected logical conflict, extract the set of conflict nodes and the set of conflict relationships corresponding to the logical conflict, and calculate the conflict severity according to the preset conflict severity scoring rule The conflict severity is calculated according to the following formula: ; where represents the entity attribute conflict score, represents the association relationship conflict score, represents the inference direction conflict score, , , are the conflict score weighting coefficients, and satisfy ; S54. Compare the conflict severity with the preset conflict severity threshold, and screen out the logical conflicts with the conflict severity greater than the preset conflict severity threshold. For the screened logical conflicts, call the enterprise's initial knowledge graph to retrieve and complete the knowledge fragments. The completed knowledge fragments include entity nodes, attribute nodes, and relationship edge information associated with the set of conflict nodes; S55. Based on the completed knowledge fragments, perform local inference expansion processing, and insert the entity nodes, attribute nodes, and relationship edges in the completed knowledge fragments into the corresponding conflict positions in the inference chain after fusion optimization to form a set of corrected inference paths; S56. Integrate the set of corrected inference paths and the inference paths in the inference chain after fusion optimization that have not had conflicts to generate a corrected inference chain.
[0019] Optionally, S54 includes the following steps: S541. For each detected logical conflict in the inference chain after fusion optimization, call the preset conflict severity scoring rule to calculate the corresponding conflict severity ; S542. Set the preset conflict severity threshold . The preset conflict severity threshold is set according to the statistics of historical inference chain conflict correction samples and is used as the benchmark condition for screening whether logical conflicts need to be completed; S543. Compare the severity of conflicts with a preset conflict severity threshold If the conflict severity is greater than the preset conflict severity threshold then add the logical conflict to the set of logical conflicts to be completed S544. For each logical conflict in the set of logical conflicts to be completed, extract the corresponding conflict node set and conflict relationship set to form a conflict element set S545. Based on the conflict element set, perform a complementary knowledge retrieval process in the enterprise's initial knowledge graph, and the retrieval conditions include the entity category, attribute category, and association relationship category of the nodes in the conflict node set S546. According to the retrieval conditions, retrieve matching complementary knowledge fragments in the enterprise's initial knowledge graph. The complementary knowledge fragments include entity nodes, attribute nodes, and relationship edge information, and screen the retrieved complementary knowledge fragments according to the semantic matching degree score with the conflict node set S547. Output the screened complementary knowledge fragments as the input data for the local reasoning and extension process
[0020] Optionally, the S6 includes the following steps S61. Extract the set of reasoning entity nodes and the set of reasoning relationship edges in the corrected reasoning chain, and record the entity identification information of each reasoning entity node in the set of reasoning entity nodes, and the relationship type information of each reasoning relationship edge in the set of reasoning relationship edges S62. For the set of reasoning entity nodes in the corrected reasoning chain, perform entity node matching retrieval in the enterprise's internal data source and the enterprise's external data source respectively, and establish a preliminary entity matching set based on the entity identification information, entity attribute description information, and entity context semantic information of the reasoning entity nodes S63. Calculate the semantic alignment similarity for each pair of candidate matching entity nodes in the preliminary entity matching set. The semantic alignment similarity is calculated based on the entity attribute similarity and the context relationship consistency score using the following formula ; where is the weighting coefficient of the attribute similarity and the context relationship consistency score, and satisfies ; S64. Compare the semantic alignment similarity with the preset semantic alignment threshold, and screen out the semantic alignment similarity Matching entity node pairs greater than the preset semantic alignment threshold, and establishing an intelligent association relationship for reasoning entity nodes; S65. For the screened matching entity node pairs, extract the corresponding inference relation types in the inference relation edge set, and establish an associated relationship mapping of the inference relation edges between the internal data source and the external data source of the enterprise according to the inference relation types; S66. Based on the set of inference entity nodes and the set of inference relation edges that have undergone semantic alignment processing and intelligent association processing, construct a unified dynamic optimization data analysis view for the enterprise, and the unified dynamic optimization data analysis view of the enterprise is used as the basic data input for generating the enterprise-wide data analysis results.
[0021] The beneficial effects of the present invention are as follows: (1) By combining the structured reasoning of the knowledge graph and the semantic reasoning of the large language model, the present invention proposes a two-way reasoning fusion mechanism, realizes the mutual calibration and dynamic optimization of the reasoning paths, breaks through the limitations of the traditional single reasoning method in terms of data analysis accuracy and reasoning chain integrity, and effectively improves the enterprise's data integration and intelligent reasoning capabilities in a multi-source heterogeneous data environment.
[0022] (2) Through natural language intention parsing and inference task dynamic generation technology, and combining the construction of a dual-channel reasoning chain of symbolic reasoning and semantic reasoning, the present invention breaks through the limitations of traditional structured query-based methods, enabling enterprise users to directly drive the reasoning and data analysis processes through natural language, and improving the query flexibility and reasoning depth of the system.
[0023] (3) Through the local logical conflict detection and adaptive knowledge completion mechanism, the present invention realizes the real-time identification and dynamic correction of logical conflicts and semantic deficiencies during the reasoning process, ensures the coherence and integrity of the reasoning chain, and enhances the stability of the reasoning process and the reliability of the reasoning results.
[0024] (4) Through semantic alignment and intelligent association processing technology, the present invention establishes a dynamic semantic association system for multi-source heterogeneous data inside and outside the enterprise, significantly improves the intelligence level of data fusion, and promotes the unified management and in-depth analysis and application of the enterprise-wide data resources.
[0025] (5) By combining the domain-fine-tuned large language model with knowledge reasoning optimization, the present invention forms a highly adaptable and highly interpretable reasoning and analysis ability for enterprise-specific domain data, which is different from the traditional general model reasoning method, and enhances the business relevance and audit compliance of the reasoning results.
[0026] (6) Through the reasoning chain traceability mechanism, the present invention records the evolution process of the reasoning chain path, reasoning nodes and reasoning relations, supports the full traceability and transparent audit of the reasoning process, and improves the reliability, transparency and interpretability of the system in the decision-making support scenario. Brief Description of the Drawings
[0027] The accompanying drawings are used to provide a further understanding of the present invention and form a part of the specification. They are used together with the embodiments of the present invention to explain the present invention and do not constitute a limitation to the present invention. In the drawings: Figure 1 is a flowchart of the enterprise-wide data analysis method based on knowledge graph and large language model proposed by the present invention; Figure 2 is a flowchart of the inference chain generation and fusion optimization of the enterprise-wide data analysis method based on knowledge graph and large language model proposed by the present invention; Figure 3 is a flowchart of the local conflict detection and adaptive knowledge completion of the enterprise-wide data analysis method based on knowledge graph and large language model proposed by the present invention. Detailed Description of the Embodiments
[0028] The present invention will now be described in further detail with reference to the accompanying drawings. These drawings are all simplified schematic diagrams, only showing the basic structure of the present invention in a schematic way, so they only show the components related to the present invention.
[0029] Referring to Figures 1 - 3 , the enterprise-wide data analysis method based on knowledge graph and large language model includes the following steps: S1. Collect enterprise internal data sources and enterprise external data sources, perform data cleaning, normalization, and entity mapping processing, and generate an enterprise initial knowledge graph, where the enterprise initial knowledge graph includes knowledge graph entity nodes, knowledge graph attribute nodes, and knowledge graph relationship edges; In this embodiment, enterprise internal data sources (such as ERP systems, CRM systems, production monitoring platforms, etc.) and enterprise external data sources (such as industry databases, third-party intelligence platforms, open-source data sets, etc.) are collected through a unified access interface. The collected data is first subjected to data cleaning processing to remove redundant, abnormal, and duplicate information, and then normalized to unify data in different formats and standards into a standard structure. Based on entity recognition and mapping rules, entity mapping processing is performed to map data items to uniformly defined entity categories and attribute relationships. Finally, an enterprise initial knowledge graph is generated. This enterprise initial knowledge graph is composed of entity nodes, attribute nodes, and relationship edges, forming a unified structured semantic representation, realizing the standardized integration of multi-source heterogeneous data, breaking through enterprise data islands, laying a foundation for subsequent reasoning and intelligent analysis, and effectively improving data consistency, semantic clarity, and analysis reliability.
[0030] S2. Receive a natural language query request, perform semantic parsing on the natural language query request using a large language model, extract query intent, query target entity, and query associated relationship elements, and generate a query intent graph and inference task description information; S3. Invoke the enterprise's initial knowledge graph, perform symbolic reasoning processing based on the query intent graph and inference task description information to generate a first inference chain, and perform semantic reasoning processing in combination with the large language model to generate a second inference chain; S4. Perform inference chain fusion processing on the first inference chain and the second inference chain, verify according to the inference path structure similarity index and the inference conclusion consistency index, and adopt a dynamic iterative optimization mechanism to generate a fused and optimized inference chain; S5. Based on the fused and optimized inference chain, perform local conflict detection processing and adaptive knowledge completion processing, and output a corrected inference chain; S6. Extract inference entity nodes and inference relationship edges according to the corrected inference chain, perform semantic alignment processing and intelligent association processing on the enterprise's internal data source and the enterprise's external data source, and construct an enterprise unified dynamic optimization data analysis view; S7. Based on the enterprise unified dynamic optimization data analysis view, output the enterprise's global data analysis results and generate an inference chain traceability file.
[0031] Based on the enterprise unified dynamic optimization data analysis view, this embodiment extracts the inference entity nodes and inference relationship edges after semantic alignment and intelligent association processing, combines the evolution process of each node in the inference process, the relationship change record, and the inference chain dynamic adjustment information to generate a complete enterprise global data analysis result. At the same time, an inference chain traceability file is generated in real time. The traceability file includes the inference starting point, the sequence of inference path nodes, the relationship evolution trajectory, and the formation process of the inference conclusion, realizing the full-link visualization and traceability of the inference process. It not only improves the transparency and audit compliance of the data analysis results, but also enhances the enterprise's intelligent level in data governance, decision support, and problem tracing, providing a solid data support for subsequent business optimization and risk prevention and control.
[0032] In this embodiment, S2 includes the following steps: S21. Receive the natural language query request input by the enterprise user, perform clause splitting processing and word segmentation processing on the natural language query request, and generate a set of natural language parsing fragments; S22. Use the large language model to perform semantic classification processing on the set of natural language parsing fragments, classify the set of natural language parsing fragments according to query intent fragments, query target entity fragments, and query associated relationship fragments, and respectively label query intent category labels, query target entity category labels, and query associated relationship category labels; S23. For the query target entity segment, perform entity mapping retrieval in the enterprise initial knowledge graph to obtain a set of entity matching results of the knowledge graph entity nodes, and assign entity matching confidence levels to each knowledge graph entity node. S24. Based on the query intent segment, query target entity segment, and query association relationship segment, construct a query intent graph. The query intent graph includes query intent layer nodes, query target entity layer nodes, and query association relationship layer edges, and the association relationships between them are labeled according to semantic relevance and context consistency. S25. Set the comprehensive association weight between any node and node in the query intent graph. The comprehensive association weight is calculated according to the following formula: ; where represents the semantic correlation function between node and node , represents the context consistency score between node and node , represents the shortest path length between node and node in the enterprise initial knowledge graph, and are comprehensive association weight adjustment coefficients. By combining the semantic correlation score and context consistency score between node pairs in the inference path, this formula comprehensively considers the closeness of the semantic relationship between nodes and the rationality of the inference path in the context. At the same time, it introduces the shortest path distance between nodes in the knowledge graph as a regulatory factor to construct a weight index reflecting the overall inference association strength of node pairs. The principle is to ensure the logical coherence between nodes in the inference chain and take into account the reasonable distribution of the inference path in the knowledge network through weighted fusion of the static semantic relationship and dynamic inference context between nodes, so as to screen out a path structure that more conforms to the inference task intention during the inference process and improve the accuracy and robustness of the inference chain generation.
[0033] S26. Generate inference task description information based on the node hierarchy structure of the query intent graph, the set of query association relationship layer edges, and the comprehensive association weight matrix. The inference task description information includes the query intent layer node set, query target entity layer node set, query association relationship layer edge set, comprehensive association weight matrix, and inference priority sorting information.
[0034] In this embodiment, by receiving a natural language query request input by an enterprise user, sentence splitting processing and word segmentation processing are performed to generate a set of natural language parsing fragments. A large language model is used to perform semantic classification on the set of parsing fragments, which are classified into query intention picture fragments, query target entity fragments, and query association relationship fragments, and category labels are respectively marked. Subsequently, entity mapping retrieval is performed on the query target entity fragments in the enterprise initial knowledge graph to generate an entity matching result set and allocate matching confidence. Further, a query intention graph is constructed based on each fragment. By setting a comprehensive association weight function between nodes, the weights are calculated by combining the node semantic relevance, context consistency, and graph structure distance. Finally, based on the node hierarchy structure of the query intention graph, the edge set of the association relationship layer, and the comprehensive association weight matrix, inference task description information is generated. The present invention effectively improves the understanding and inference conversion ability of natural language queries, enables the inference task to fully reflect the user's query intention and data semantic association characteristics, and enhances the accuracy, flexibility, and intelligence of the inference path generation.
[0035] In this embodiment, the specific steps of S24 include: S241. Respectively extract a query intention element set, a query target entity element set, and a query association relationship element set from the query intention picture fragment, the query target entity fragment, and the query association relationship fragment to generate a query intention graph; S242. Map each query intention element in the query intention element set, the query target entity element set, and the query association relationship element set to a query intention layer node, a query target entity layer node, and a query association relationship layer edge respectively; S243. According to the nodes and edges in the query intention graph, establish a semantic association between each query intention layer node in the query intention graph and one or more corresponding query target entity layer nodes, and each query target entity layer node in the query intention graph is connected to other query target entity layer nodes according to the query association relationship layer edge; S244. In the query intention graph, set initial attributes for each query intention layer node, query target entity layer node, and query association relationship layer edge respectively. The initial attributes include node category identification, node semantic vector representation, node context vector representation, and edge association relationship category identification; S245. For any query intention layer node and query target entity layer node in the query intention graph and the node , based on the node semantic vector representation of node and the node semantic vector representation of node , calculate the semantic relevance function between node and node ; S246. Based on node The node context vector representation of and the node The node context vector representation of, calculate the node and the node The context consistency score between .
[0036] In this embodiment, by parsing the query intent fragment, query target entity fragment, and query association relationship fragment in the natural language query, the corresponding query intent element set, query target entity element set, and query association relationship element set are respectively extracted, a query intent graph is constructed, and each element set is mapped to nodes in the query intent layer, query target entity layer nodes, and edges in the query association relationship layer. The initial attributes of each node and edge are set in the query intent graph, including node category identifiers, semantic vector representations, and context vector representations. Subsequently, the semantic relevance between nodes is calculated based on the semantic vectors of the nodes, and the context consistency score between nodes is calculated based on the context vectors, thereby realizing structured semantic association and context semantic fusion when constructing the query intent graph, improving the semantic accuracy and logical coherence of the inference path generation, providing high-quality inference task inputs for subsequent symbolic reasoning and semantic reasoning, and significantly enhancing the intelligent level and interpretability of natural language understanding and inference intent modeling in the enterprise data analysis process.
[0037] In this embodiment, S3 includes the following steps: S31. Based on the inference task description information, call the enterprise initial knowledge graph, and locate the corresponding knowledge graph entity nodes in the enterprise initial knowledge graph according to each query target entity node in the query target entity layer node set, and generate a knowledge graph entity node set; S32. Based on the knowledge graph entity node set and the query association relationship layer edge set, perform symbolic reasoning processing along the knowledge graph relationship edges in the enterprise initial knowledge graph to generate a symbolic reasoning path set, where each symbolic reasoning path starts and ends with a knowledge graph entity node and satisfies the semantic requirements of the query association relationship layer edge; S33. Perform screening processing on the symbolic reasoning path set according to the comprehensive association weight matrix, retain the symbolic reasoning paths with comprehensive association weight values greater than the preset weight threshold, and sort the screened symbolic reasoning paths according to the inference priority sorting information to generate a first inference chain; S34. Input the inference task description information into the large language model, and perform semantic reasoning processing based on the query intent layer node set, query target entity layer node set, and query association relationship layer edge set to derive and generate a semantic reasoning path set; S35. For each semantic inference path in the semantic inference path set, according to the node hierarchy of the query intent graph, adjust the inference association order of the inference nodes in the semantic inference path, and calculate the semantic inference path score; S36. According to the comprehensive association weight matrix and the semantic inference path score, perform a scoring process on the semantic inference path set, filter out the semantic inference paths with scoring values greater than the preset scoring threshold, and sort them according to the inference priority sorting information to generate a second inference chain.
[0038] In this embodiment, based on the inference task description information, first locate the query target entity node in the enterprise initial knowledge graph, perform symbolic reasoning processing along the knowledge graph relationship edges to generate a set of symbolic reasoning paths that meet the query intent, and filter out the high-correlation paths according to the comprehensive association weight matrix to construct the first inference chain; at the same time, input the inference task description information into the large language model, perform semantic inference processing based on the query intent graph to generate a set of semantic inference paths, adjust the inference order and score according to the inference node hierarchical relationship, filter out high-quality semantic inference paths to construct the second inference chain. Through the parallel generation and optimization of symbolic reasoning and semantic reasoning, the dual-channel construction of the inference chain is realized, the accuracy, rationality and flexibility of the inference path generation are improved, the rigor of structural reasoning and the flexibility of semantic reasoning are taken into account, a solid foundation is laid for the subsequent inference chain fusion and optimization, and the robustness and intelligence of enterprise data analysis and reasoning are enhanced.
[0039] In this embodiment, S34 includes the following steps: S331. Traverse each symbolic reasoning path in the symbolic reasoning path set, and for each adjacent node pair in the symbolic reasoning path, extract the corresponding comprehensive association weight value from the comprehensive association weight matrix; S332. For each symbolic reasoning path, calculate the comprehensive weight cumulative value of the symbolic reasoning path based on the comprehensive association weight value , and compare the comprehensive weight cumulative value of each symbolic reasoning path with the preset weight threshold , filter out the symbolic reasoning paths with comprehensive weight cumulative values greater than the preset weight threshold to form a symbolic reasoning path screening set; S333. For each symbolic reasoning path in the symbolic reasoning path screening set, extract the inference priority sorting information in the inference task description information, and perform priority sorting processing on the symbolic reasoning paths in the symbolic reasoning path screening set according to the inference priority sorting information; S334. According to the set of symbolic reasoning paths sorted by inference priority, connect the symbolic reasoning paths in sequence according to the priority order to generate the first inference chain.
[0040] In this embodiment, by traversing the set of symbolic inference paths, extracting the weight values of adjacent node pairs in the comprehensive association weight matrix, calculating the cumulative comprehensive weight value of the symbolic inference paths, and comparing it with a preset weight threshold, symbolic inference paths with high association degrees are screened out to form a screening set. Subsequently, the symbolic inference paths in the screening set are sorted according to the inference priorities in the inference task description information. Finally, the first inference chain is generated by connecting them in the order of priorities, which can effectively improve the relevance and rationality of the inference path screening, ensure that the inference process is more in line with the user's intention and the logic of the knowledge structure, enhance the accuracy of the inference chain generation and the overall reliability of the inference, and provide high-quality input support for the subsequent inference chain fusion and optimization.
[0041] In this embodiment, S4 includes the following steps: S41. Extract the set of inference paths in the first inference chain and the set of inference paths in the second inference chain, and respectively mark the start node and the end node of each inference path in the set of inference paths. S42. For any two inference paths in the first inference chain and the second inference chain, calculate the inference path structure similarity index , and the inference path structure similarity index is calculated based on the node sequence matching degree of the two inference paths, and the following formula is used: ; Among them, represents the number of common nodes in the two inference paths, and respectively represent the number of nodes in the two inference paths; This formula evaluates the structural matching degree between inference paths through the inference path structure similarity index, and plays a role in measuring the similarity degree of the node structure arrangements of the two inference paths. Specifically, the calculation principle of the inference path structure similarity is based on the proportional relationship between the number of common nodes in the two inference paths and the minimum value of their respective node numbers, reflecting the coincidence degree of the node composition and order of the two inference paths. The more the number of common nodes, and the closer it is to the total number of nodes in the shorter path, the higher the inference path structure similarity score, indicating that the two inference paths are closer in structure, which helps to screen out inference path pairs with stronger structural consistency in the subsequent inference chain fusion process, and improve the accuracy and reliability of the inference chain fusion and optimization.
[0042] S43. For any two inference paths in the first inference chain and the second inference chain, calculate the inference conclusion consistency index , and the inference conclusion consistency index is determined based on the consistency score of the target entity and the association relationship derived from the two inference paths; S44. According to the inference path structure similarity index and the inference conclusion consistency index , filter out the inference path pairs that meet the condition that the similarity index is greater than the preset structure similarity threshold and the consistency index is greater than the preset consistency threshold, and form a set of inference path matching pairs; S45. Based on the set of inference path matching pairs, adopt a dynamic iterative optimization mechanism, and according to the inference priority and path importance index, adjust the connection order of the inference paths and the node merging strategy round by round to generate a fused and optimized inference chain.
[0043] In this embodiment, by extracting the inference path sets in the first inference chain and the second inference chain, respectively marking the starting nodes and ending nodes of each inference path, calculating the inference path structure similarity index based on the node sequence matching, and calculating the inference conclusion consistency index according to the derived target entity and associated relationship scores, further filter out the inference path pairs that meet the dual criteria of structure similarity and conclusion consistency, construct a set of inference path matching pairs, and on this basis, adopt a dynamic iterative optimization mechanism to gradually adjust the inference path connection order and node merging strategy according to the inference priority and path importance index, and finally generate a fused and optimized inference chain, realizing the two-way interactive verification and dynamic adjustment of the inference chain, effectively improving the accuracy, robustness and consistency of the inference process in the enterprise's global data analysis, and enhancing the stability and interpretability of the inference results.
[0044] In this embodiment, S44 includes the following steps: S441. Set a preset structure similarity threshold , and the preset structure similarity threshold is set according to the historical inference path node sequence matching degree statistical data and is used as the benchmark condition for screening the inference path structure similarity; S442. Set a preset consistency threshold , and the preset consistency threshold is set according to the weighted calculation result of the inference target entity consistency score and the inference associated relationship consistency score and is used as the benchmark condition for screening the inference conclusion consistency; S443. For each pair of inference paths in the first inference chain and the second inference chain, if the inference path structure similarity index is greater than or equal to the preset structure similarity threshold , and the inference conclusion consistency index is greater than or equal to the preset consistency threshold , then add this inference path pair to the set of inference path matching pairs; S444. Perform deduplication on the set of inference path matching pairs, eliminate duplicate inference path pairs with exactly the same starting node, ending node, and intermediate nodes, and retain the unique inference path matching pairs.
[0045] In this embodiment, by setting a preset structural similarity threshold based on the statistical matching degree of historical inference path node sequences and a preset consistency threshold calculated by weighting the inference target entity consistency score and the inference association relationship consistency score, double-index screening is performed on the inference path pairs in the first inference chain and the second inference chain. The inference path pairs that simultaneously meet the structural similarity and inference conclusion consistency are screened out, and node and path deduplication processing is performed on the matched inference path pairs, thereby retaining the unique and effective inference path matching relationship, achieving high-precision path screening and structural optimization before inference chain fusion, effectively improving the accuracy, reliability, and consistency of the inference chain in the inference fusion process, and laying a solid foundation for subsequent dynamic optimization of the inference chain and the credibility of the inference results.
[0046] In this embodiment, S5 includes the following steps: S51. Receive the inference chain after fusion optimization, and extract the inference path set and inference node set in the inference chain after fusion optimization; S52. For the inference path set in the inference chain after fusion optimization, detect the logical conflicts existing within each inference path and between different inference paths in the inference path set. The logical conflicts include inference direction conflicts, entity attribute conflicts, and association relationship conflicts; S53. For each detected logical conflict, extract the conflict node set and conflict relationship set corresponding to the logical conflict, and calculate the conflict severity according to the preset conflict severity scoring rule The conflict severity is calculated according to the following formula: ; where represents the entity attribute conflict score, represents the association relationship conflict score, represents the inference direction conflict score, . . are conflict score weighting coefficients and satisfy ; This formula performs weighted summation respectively according to the entity attribute conflict score, association relationship conflict score, and reasoning direction conflict score involved in each logical conflict. The weight coefficients used are for adjusting the contribution ratio of different types of conflicts to the overall severity. Each type of conflict score reflects the deviation degree of the conflict in terms of attribute consistency, relationship rationality, and reasoning path directionality. The larger the finally synthesized conflict severity value, the greater the impact of the logical conflict on the coherence and accuracy of the reasoning chain, thereby guiding the system to prioritize the processing of high-severity conflicts and enhancing the pertinence and efficiency of the reasoning chain correction process.
[0047] S54. According to the conflict severity Compare with the preset conflict severity threshold, and screen out the logical conflicts with conflict severity greater than the preset conflict severity threshold. For the screened logical conflicts, call the enterprise's initial knowledge graph to retrieve and complete the knowledge fragments, where the completed knowledge fragments include entity nodes, attribute nodes, and relationship edge information associated with the conflict node set; S55. Based on the completed knowledge fragments, perform local reasoning expansion processing, and insert the entity nodes, attribute nodes, and relationship edges in the completed knowledge fragments into the corresponding conflict positions of the reasoning chain after fusion and optimization to form a set of corrected reasoning paths; S56. Integrate the set of corrected reasoning paths and the reasoning paths in the reasoning chain after fusion and optimization that have not had conflicts to generate a corrected reasoning chain.
[0048] In this embodiment, by receiving the reasoning chain after fusion and optimization, extracting the set of reasoning paths and the set of reasoning nodes, detecting the reasoning direction conflicts, entity attribute conflicts, and association relationship conflicts inside and between the reasoning paths, and calculating the conflict severity according to the preset scoring rules. When the conflict severity exceeds the preset threshold, call the enterprise's initial knowledge graph to retrieve and complete the knowledge fragments, and perform local reasoning expansion based on the completed knowledge fragments, effectively correcting the logical defects of the reasoning chain. Finally, integrate the corrected reasoning paths to generate a complete reasoning chain, realizing the dynamic identification and repair of logical conflicts and semantic deficiencies during the reasoning process, enhancing the integrity of the reasoning chain, the accuracy of reasoning inference, and the stability of the data analysis process, and providing reliable and coherent data support for enterprise intelligent decision-making.
[0049] In this embodiment, S54 includes the following steps: S541. For each logical conflict detected in the reasoning chain after fusion and optimization, call the preset conflict severity scoring rule to calculate the corresponding conflict severity ; S542. Set the preset conflict severity threshold , and the preset conflict severity threshold Revise the sample statistics setting according to the historical reasoning chain conflict, which is used as the benchmark condition for screening whether the logical conflict needs to be complemented; S543. Compare the conflict severity With the preset conflict severity threshold If the conflict severity Is greater than the preset conflict severity threshold Then add the logical conflict to the set of logical conflicts to be complemented; S544. For each logical conflict in the set of logical conflicts to be complemented, extract the corresponding conflict node set and conflict relationship set to form a conflict element set; S545. Based on the conflict element set, perform complementary knowledge retrieval processing in the enterprise's initial knowledge graph, and the retrieval conditions include the entity category, attribute category, and association relationship category of the nodes in the conflict node set; S546. According to the retrieval conditions, retrieve the matching complementary knowledge fragments in the enterprise's initial knowledge graph. The complementary knowledge fragments include entity nodes, attribute nodes, and relationship edge information, and screen the retrieved complementary knowledge fragments according to the semantic matching degree score with the conflict node set; S547. Output the screened complementary knowledge fragments as the input data for local reasoning expansion processing.
[0050] In this embodiment, for the logical conflicts detected in the reasoning chain after fusion and optimization, the conflict severity scoring rule is called to calculate the conflict severity, and screening is performed according to the preset conflict severity threshold set by the historical reasoning chain conflict samples. The logical conflicts with a conflict severity exceeding the threshold are included in the process of being complemented. Further, the conflict node set and conflict relationship set are extracted to form a conflict element set, and complementary knowledge retrieval is performed in the enterprise's initial knowledge graph with this as the retrieval condition. Complementary knowledge fragments with a high semantic matching degree with the conflict node set are screened. Finally, the screened complementary knowledge fragments are used as the input data for local reasoning expansion, so as to realize the dynamic detection of key conflicts, knowledge complementation, and self-adaptive optimization of the reasoning chain during the reasoning process, effectively improving the integrity and logical coherence of the reasoning chain structure, and enhancing the accuracy and stability of reasoning in enterprise-wide data analysis.
[0051] In this embodiment, S6 includes the following steps: S61. Extract the set of reasoning entity nodes and the set of reasoning relationship edges in the corrected reasoning chain, and record the entity identification information of each reasoning entity node in the set of reasoning entity nodes, and the relationship type information of each reasoning relationship edge in the set of reasoning relationship edges; S62. For the set of inference entity nodes in the corrected inference chain, perform entity node matching and retrieval in the internal enterprise data source and the external enterprise data source respectively, and establish a preliminary entity matching set based on the entity identification information, entity attribute description information, and entity context semantic information of the inference entity nodes; S63. For each pair of candidate matching entity nodes in the preliminary entity matching set, calculate the semantic alignment similarity , where the semantic alignment similarity is calculated based on the entity attribute similarity and the context relationship consistency score using the following formula: ; where is the weighting coefficient of the attribute similarity and the context relationship consistency score, and satisfies ; This formula is used to measure the comprehensive matching degree between the inference entity node and the data source entity node in terms of attribute description and context relationship, and plays a role in screening entity node pairs with a higher semantic correlation degree. The calculation principle of the semantic alignment similarity is to perform weighted fusion on the similarity score between entity attributes and the consistency score of the context relationship. The weight coefficient can be set according to actual needs to reflect the importance ratio of attribute features and context environment. By comprehensively considering the similarity of entity attributes themselves and the consistency of their context environment, the semantic correspondence relationship between entity nodes can be evaluated more comprehensively and accurately, providing a reliable basis for subsequent data semantic alignment and intelligent association processing.
[0052] S64. Compare the semantic alignment similarity with the preset semantic alignment threshold, screen out the matching entity node pairs with the semantic alignment similarity greater than the preset semantic alignment threshold, and establish an intelligent association relationship for the inference entity nodes. The preset semantic alignment threshold is used as a judgment criterion for screening the semantic matching relationship between the inference entity node and the data source entity node. When the semantic alignment similarity is greater than this threshold, it is determined that the matching is successful and an intelligent association relationship is established; S65. For the screened matching entity node pairs, extract the corresponding inference relationship types in the inference relationship edge set, and establish an associated relationship mapping of the inference relationship edges between the internal enterprise data source and the external enterprise data source; S66. Based on the set of inference entity nodes and the set of inference relationship edges that have undergone semantic alignment processing and intelligent association processing, construct a unified dynamic optimization data analysis view for the enterprise. The unified dynamic optimization data analysis view for the enterprise serves as the basic data input for generating the enterprise-wide data analysis results.
[0053] In this embodiment, by extracting the set of inference entity nodes and the set of inference relationship edges in the corrected inference chain, and performing matching retrieval in the internal enterprise data source and the external enterprise data source based on entity identification information, attribute description information, and context semantic information, calculating the semantic alignment similarity, screening out the matching entity node pairs that meet the preset semantic alignment threshold, establishing the intelligent association relationship of the inference entity nodes, and constructing the associated mapping of the inference relationship edges between different data sources according to the inference relationship type. Finally, based on the set of inference entity nodes and the set of inference relationship edges after semantic alignment and intelligent association processing, a unified enterprise data analysis view is dynamically generated, realizing the deep integration and intelligent fusion of multi-source heterogeneous data, improving the coherence and accuracy of data analysis, and providing a solid data foundation for subsequent global data intelligent analysis and decision support.
[0054] Example 1: To verify the feasibility of the present invention in implementation, the present invention is applied to the data intelligent analysis platform of a well-known large manufacturing group across the country. The group has accumulated a large amount of data sets from multiple fields such as production, supply chain, customer service, and external market environment during its daily operation, involving multi-source heterogeneous data such as ERP system data, MES system data, CRM system data, as well as external industry reports and regulatory information. Due to the diverse data sources, complex formats, and inconsistent semantics, traditional data integration and analysis methods are difficult to meet the group's requirements for real-time, intelligent, and interpretable data analysis and decision support, and there are problems such as data islands, broken inference chains, and frequent semantic conflicts, resulting in scattered and incoherent decision-making bases, which restricts the development of the enterprise's intelligent transformation. Therefore, the group decides to introduce the enterprise global data analysis method based on the fusion inference of knowledge graph and large language model proposed by the present invention, aiming to break through the data barriers and improve the depth of data integration and decision support capabilities.
[0055] In specific applications, the group first performs standardized preprocessing on the data of internal enterprise systems such as ERP, MES, and CRM, including data cleaning, format normalization, and entity mapping, and uniformly maps them into the initial enterprise knowledge graph. The entity nodes, attribute nodes, and relationship edges in the knowledge graph are constructed according to the unified enterprise ontology specification. At the same time, the industry data, supplier reports, and documents purchased externally by the group also go through the same standardization process and are added to the external data knowledge graph of the enterprise, finally forming a dual-graph fusion structure of the enterprise's internal and external.
[0056] In the actual business analysis process, business personnel put forward complex data analysis requests through natural language, such as "Under the influence of the new environment, the most severely affected raw material categories in the supply chain this quarter and the situation of their upstream and downstream enterprises". The method of the present invention first analyzes the natural language request through a large language model, extracts the query intention, target entity and associated relationship elements, and generates a query intention graph and inference task description information. In the inference stage, on the one hand, the system performs symbolic reasoning based on the knowledge graph to generate a symbolic reasoning chain, and on the other hand, combines the large language model for semantic reasoning to generate a semantic reasoning chain, and performs two-way fusion verification through the inference path structure similarity and inference conclusion consistency indicators, and adopts a dynamic iteration mechanism to generate a fused and optimized inference chain.
[0057] During the inference process, the system automatically detects logical conflicts in the inference chain, such as broken supply chain entity relationships, contradictory attribute descriptions, etc., and in real time calls relevant domain knowledge in the enterprise's initial knowledge graph to perform local knowledge completion and conflict resolution to ensure the integrity and consistency of the inference chain. For the fused and optimized inference chain, extract the inference entity nodes and inference relationship edges, perform entity matching in the enterprise internal data source and the enterprise external data source respectively, and use the semantic alignment similarity index based on weighted calculation of entity attribute similarity and context relationship consistency to screen out the best matching entity pairs, complete semantic alignment, and establish an intelligent entity association relationship.
[0058] Based on the inference chain after completing semantic alignment and intelligent association processing, the system dynamically constructs a unified and optimized data analysis view, connects the raw material category nodes affected by the new environment and their upstream and downstream enterprise nodes into a dynamic relationship graph with clear semantics and logical coherence, and completely records the changes of each path, node and relationship evolution process in the inference chain into the inference chain traceability file, providing a transparent and traceable basis for subsequent decision-making audit and compliance.
[0059] In order to verify the actual effect of the method of the present invention, the group selected data for nearly one year for a comparative experiment. The experiment selected three typical business scenarios, including supply chain risk analysis, customer complaint traceability analysis, and financial anomaly early warning analysis, and analyzed and compared them using the traditional data warehouse query method and the method of the present invention respectively. The experimental period was three months, and the specific data is shown in the following table.
[0060] Table 1 Comparison of data analysis effects of traditional method and the method of the present invention in different business scenarios ; It can be seen from the experimental results that the method of the present invention has shortened the supply chain risk analysis response time by about 78%, increased the accuracy rate of customer complaint traceability by nearly 20 percentage points, and extended the early warning period of financial anomalies from 2 days to 7 days, effectively enhancing the group's sensitivity to and response ability to business risks. In terms of the integrity of data fusion, the present invention has increased the integration rate of multi-source data to more than 95% through semantic alignment and intelligent association mechanisms, and both the coherence score of the inference chain and the complete traceability rate of the inference chain have been greatly improved, providing high-confidence and high-transparency data analysis support for enterprises. At the same time, the accuracy rate of understanding the natural language query intent of users has been increased to more than 93%, greatly improving the data interaction experience of business personnel and reducing the technical threshold.
[0061] In practical applications, the present invention also demonstrates strong dynamic adaptation capabilities. Taking supply chain analysis as an example, when there are external updates or changes in the industry environment, the system can update the inference chain in real time based on the supplementary knowledge fragments and dynamically reconstruct the data analysis view, avoiding the cumbersome process of manually reconstructing the data warehouse and logical rules under the traditional static model. For example, in March 2025, after a certain raw material was exported, the system completed the adjustment of the inference chain and the update of the data view of the affected supply chain nodes within only 6 hours and automatically pushed the early warning report to the group's decision-making level, greatly improving the enterprise's emergency response speed.
[0062] In summary, this embodiment fully verifies the significant advantages of the method of the present invention in aspects such as intelligent fusion of enterprise-wide multi-source heterogeneous data, dynamic construction of inference chains, understanding of data analysis requests driven by natural language, local knowledge supplementation, cross-source semantic alignment, intelligent entity association, and full traceability of inference chains. Through the application of the present invention, enterprises have not only achieved synchronous improvement in data integration efficiency and analysis depth, but also significantly enhanced the transparency, traceability, and adaptability of the decision-making process, providing strong technical support for intelligent enterprise operation in complex environments.
[0063] The above is only a preferred specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention, according to the technical solution and inventive concept of the present invention, makes equivalent substitutions or changes, and all should be covered within the protection scope of the present invention.
Claims
1. An enterprise-wide data analysis method based on a knowledge graph and a large language model, characterized in that, It includes the following steps: S1. Collect the internal data sources and external data sources of the enterprise, perform data cleaning, normalization, and entity mapping processing, and generate an initial enterprise knowledge graph, where the initial enterprise knowledge graph includes knowledge graph entity nodes, knowledge graph attribute nodes, and knowledge graph relationship edges; S2. Receive a natural language query request, use a large language model to perform semantic parsing on the natural language query request, extract query intent, query target entity, and query associated relationship elements, and generate a query intent graph and inference task description information; S3. Invoke the initial enterprise knowledge graph, perform symbolic reasoning processing based on the query intent graph and inference task description information to generate a first inference chain, and perform semantic reasoning processing in combination with the large language model to generate a second inference chain; S4. Perform inference chain fusion processing on the first inference chain and the second inference chain, verify according to the inference path structure similarity index and inference conclusion consistency index, and use a dynamic iterative optimization mechanism to generate a fused and optimized inference chain; S5. Based on the fused and optimized inference chain, perform local conflict detection processing and adaptive knowledge completion processing, and output a corrected inference chain; S6. Extract inference entity nodes and inference relationship edges according to the corrected inference chain, perform semantic alignment processing and intelligent association processing on the internal data sources and external data sources of the enterprise, and construct a unified dynamic optimization data analysis view of the enterprise; S7. Based on the unified dynamic optimization data analysis view of the enterprise, output the enterprise-wide data analysis results and generate an inference chain traceability file.
2. The enterprise-wide data analysis method based on a knowledge graph and a large language model according to claim 1, wherein The S2 includes the following steps: S21. Receive the natural language query request input by the enterprise user, perform sentence splitting and word segmentation on the natural language query request, and generate a set of natural language parsing fragments; S22. Use a large language model to perform semantic classification on the set of natural language parsing fragments, classify the set of natural language parsing fragments into query intent fragments, query target entity fragments, and query associated relationship fragments, and label the query intent category label, query target entity category label, and query associated relationship category label respectively; S23. For the query target entity fragment, perform entity mapping retrieval in the initial enterprise knowledge graph to obtain a set of entity matching results of the knowledge graph entity nodes, and assign entity matching confidence to each knowledge graph entity node; S24. Based on the query intent fragment, query target entity fragment, and query associated relationship fragment, construct a query intent graph, where the query intent graph includes query intent layer nodes, query target entity layer nodes, and query associated relationship layer edges, and the association relationships between them are labeled according to semantic relevance and context consistency; S25. Set any node in the query intention graph and the node The comprehensive association weight between , and the comprehensive association weight is calculated according to the following formula: ; Among them, represents the semantic relevance function between node and node ; represents the context consistency score between node and node ; represents the shortest path length between node and node in the enterprise's initial knowledge graph; and is the comprehensive correlation weight adjustment coefficient; S26. Generate inference task description information based on the node hierarchy of the query intent graph, the set of query associated relationship layer edges, and the comprehensive association weight matrix, where the inference task description information includes the set of query intent layer nodes, the set of query target entity layer nodes, the set of query associated relationship layer edges, the comprehensive association weight matrix, and inference priority sorting information.
3. The enterprise-wide data analysis method based on a knowledge graph and a large language model according to claim 2, wherein The S24 specifically includes: S241. Extract the query intent element set, query target entity element set, and query association relationship element set based on the query intent picture fragment, query target entity fragment, and query association relationship fragment respectively, and generate a query intent graph; S242. Map each query intent element in the query intent element set, query target entity element set, and query association relationship element set to a query intent layer node, a query target entity layer node, and a query association relationship layer edge respectively; S243. Based on the nodes and edges in the query intent graph, establish a semantic association between each query intent layer node in the query intent graph and one or more corresponding query target entity layer nodes, and each query target entity layer node in the query intent graph is connected to other query target entity layer nodes according to the query association relationship layer edge; S244. In the query intent graph, set initial attributes for each query intent layer node, query target entity layer node, and query association relationship layer edge respectively. The initial attributes include node category identification, node semantic vector representation, node context vector representation, and edge association relationship category identification; S245. For any query intent layer node and query target entity layer node in the query intent graph and node , based on the node semantic vector representation of node and the node semantic vector representation of node , calculate the semantic relevance function between node ; ; S246. Calculate the context consistency score between node and node based on the node context vector representation of node and the node context vector representation of node . 4. The enterprise-wide data analysis method based on a knowledge graph and a large language model according to claim 1, wherein, The S3 includes the following steps: S31. Based on the inference task description information, call the enterprise initial knowledge graph, and locate the corresponding knowledge graph entity node in the enterprise initial knowledge graph according to each query target entity node in the query target entity layer node set, and generate a knowledge graph entity node set; S32. Based on the knowledge graph entity node set and the query association relationship layer edge set, perform symbolic reasoning processing along the knowledge graph relationship edge in the enterprise initial knowledge graph to generate a symbolic reasoning path set, where each symbolic reasoning path takes the knowledge graph entity node as the starting point and the ending point, and meets the semantic requirements of the query association relationship layer edge; S33. Perform a screening process on the symbolic reasoning path set according to the comprehensive association weight matrix, retain the symbolic reasoning paths with the comprehensive association weight value greater than the preset weight threshold, and sort the screened symbolic reasoning paths according to the inference priority sorting information to generate the first inference chain; S34. Input the inference task description information into the large language model, and perform semantic reasoning processing based on the query intent layer node set, query target entity layer node set, and query association relationship layer edge set to derive and generate a semantic reasoning path set; S35. For each semantic reasoning path in the semantic reasoning path set, adjust the inference association order of the inference nodes in the semantic reasoning path according to the node hierarchy structure of the query intent graph, and calculate the semantic reasoning path score; S36. Perform a scoring process on the semantic reasoning path set according to the comprehensive association weight matrix and the semantic reasoning path score, screen the semantic reasoning paths with the scoring value greater than the preset scoring threshold, and sort them according to the inference priority sorting information to generate the second inference chain.
5. The enterprise-wide data analysis method based on the knowledge graph and the large language model according to claim 4, wherein The S34 includes the following steps: S331. Traverse each symbolic reasoning path in the symbolic reasoning path set, and extract the corresponding comprehensive association weight value from the comprehensive association weight matrix for each adjacent node pair in the symbolic reasoning path; S332. For each symbol inference path, calculate the cumulative comprehensive weight value of the symbol inference path based on the comprehensive correlation weight value , and use the cumulative comprehensive weight value of each symbol inference path to compare with the preset weight threshold , and filter out the symbol inference paths with cumulative comprehensive weight values greater than the preset weight threshold to form a symbol inference path screening set; S333. For each symbolic reasoning path in the symbolic reasoning path screening set, extract the reasoning priority ranking information in the reasoning task description information, and perform priority ranking processing on the symbolic reasoning paths in the symbolic reasoning path screening set according to the reasoning priority ranking information; S334. According to the symbolic reasoning path set after priority ranking, connect the symbolic reasoning paths in sequence according to the priority order to generate the first reasoning chain.
6. The method for enterprise-wide data analysis based on a knowledge graph and a large language model according to claim 5, wherein, The S4 includes the following steps: S41. Extract the reasoning path set in the first reasoning chain and the reasoning path set in the second reasoning chain, and respectively mark the start node and end node of each reasoning path in the reasoning path set; S42. Calculate the inference path structure similarity index for any two inference paths in the first inference chain and the second inference chain , where the inference path structure similarity index is calculated based on the matching degree of the node sequences of the two inference paths and is calculated using the following formula: ; Among them, represents the number of common nodes in two inference paths, and respectively represent the number of nodes in two inference paths; S43. Calculate the inference conclusion consistency index for any two inference paths in the first inference chain and the second inference chain , where the inference conclusion consistency index is determined based on the consistency score of the target entity and the associated relationship derived from the two inference paths; S44. According to the inference path structure similarity index and the inference conclusion consistency index , filter out the inference path pairs that meet the condition that the similarity index is greater than the preset structure similarity threshold and the consistency index is greater than the preset consistency threshold, and form a set of inference path matching pairs; S45. Based on the reasoning path matching pair set, adopt a dynamic iterative optimization mechanism, and according to the reasoning priority and path importance index, adjust the connection order of the reasoning paths and the node merging strategy round by round to generate a fused and optimized reasoning chain.
7. The enterprise-wide data analysis method based on a knowledge graph and a large language model according to claim 6, wherein The S44 includes the following steps: S441. Set a preset structural similarity threshold , and the preset structural similarity threshold is set according to the statistical data of the matching degree of the historical inference path node sequence and is used as the benchmark condition for screening the structural similarity of the inference path; S442. Set a preset consistency threshold , and the preset consistency threshold is set according to the weighted calculation result of the inference target entity consistency score and the inference association relationship consistency score, and is used as the benchmark condition for screening the inference conclusion consistency; S443. For each pair of inference paths in the first inference chain and the second inference chain, if the inference path structure similarity index is greater than or equal to the preset structure similarity threshold , and the inference conclusion consistency index is greater than or equal to the preset consistency threshold , then add this pair of inference paths to the set of inference path matching pairs; S444. Perform duplicate removal processing on the reasoning path matching pair set, eliminate duplicate reasoning path pairs with exactly the same start node, end node, and intermediate nodes, and retain the unique reasoning path matching pairs.
8. The enterprise-wide data analysis method based on a knowledge graph and a large language model according to claim 1, wherein The S5 includes the following steps: S51. Receive the fused and optimized reasoning chain, and extract the reasoning path set and reasoning node set in the fused and optimized reasoning chain; S52. For the reasoning path set in the fused and optimized reasoning chain, detect the logical conflicts existing within each reasoning path and between different reasoning paths in the reasoning path set. The logical conflicts include reasoning direction conflicts, entity attribute conflicts, and association relationship conflicts; S53. For each detected logical conflict, extract the conflict node set and conflict relationship set corresponding to the logical conflict, and calculate the conflict severity according to the preset conflict severity scoring rule , the conflict severity is calculated according to the following formula: ; Among them, represents the entity attribute conflict score, represents the association relationship conflict score, represents the inference direction conflict score, , , are the conflict score weighting coefficients, and satisfy ; S54. According to the conflict severity Compare with the preset conflict severity threshold to screen out the logical conflicts with a conflict severity greater than the preset conflict severity threshold. For the screened logical conflicts, call the enterprise's initial knowledge graph to retrieve and complete the knowledge fragments, and the completed knowledge fragments include entity nodes, attribute nodes, and relationship edge information associated with the conflict node set; S55. Based on the completed knowledge fragments, perform local reasoning extension processing, and insert the entity nodes, attribute nodes, and relationship edges in the completed knowledge fragments into the corresponding conflict positions in the fused and optimized reasoning chain to form a corrected reasoning path set; S56. Integrate the corrected reasoning path set and the reasoning paths in the fused and optimized reasoning chain that have not had conflicts to generate a corrected reasoning chain.
9. The enterprise-wide data analysis method based on a knowledge graph and a large language model according to claim 8, wherein The S54 includes the following steps: S541. For each logical conflict detected in the optimized inference chain after fusion, call the preset conflict severity scoring rule to calculate the corresponding conflict severity ; S542. Set the preset conflict severity threshold , and the preset conflict severity threshold is set according to the historical inference chain conflict correction samples, and is used as a benchmark condition for screening whether logical conflicts need to be supplemented and processed; S543. Compare conflict severity with a preset conflict severity threshold . If the conflict severity is greater than the preset conflict severity threshold , add the logical conflict to the set of logical conflicts to be completed S544. For each logical conflict in the logical conflict set to be completed, extract the corresponding conflict node set and conflict relationship set to form a conflict element set; S545. Based on the conflict element set, perform completed knowledge retrieval processing in the enterprise initial knowledge graph, and the retrieval conditions include the entity category, attribute category, and association relationship category of the nodes in the conflict node set; S546. According to the retrieval conditions, retrieve the matching completed knowledge fragments in the enterprise initial knowledge graph. The completed knowledge fragments include entity nodes, attribute nodes, and relationship edge information, and screen the retrieved completed knowledge fragments according to the semantic matching degree score with the conflict node set; S547. Output the screened completed knowledge fragments as the input data for local reasoning extension processing.
10. The enterprise-wide data analysis method based on a knowledge graph and a large language model according to claim 1, wherein The S6 includes the following steps: S61. Extract the reasoning entity node set and reasoning relationship edge set in the corrected reasoning chain, and respectively record the entity identification information of each reasoning entity node in the reasoning entity node set, and the relationship type information of each reasoning relationship edge in the reasoning relationship edge set; S62. For the set of inference entity nodes in the corrected inference chain, perform entity node matching retrieval in the enterprise internal data source and the enterprise external data source respectively, and establish a preliminary entity matching set based on the entity identification information, entity attribute description information, and entity context semantic information of the inference entity nodes; S63. For each pair of candidate matching entity nodes in the preliminary entity matching set, calculate the semantic alignment similarity , where the semantic alignment similarity is calculated based on the entity attribute similarity and the context relationship consistency score using the following formula: ; Among them, is the weighted coefficient of the attribute similarity and the context relationship consistency score, and satisfies ; S64. Compare the semantic alignment similarity with a preset semantic alignment threshold, and filter out the matching entity node pairs whose semantic alignment similarity is greater than the preset semantic alignment threshold, and establish an intelligent association relationship of the inference entity nodes; S65. For the selected matching entity node pairs, extract the corresponding inference relation types in the inference relation edge set, and establish an associated relationship mapping of the inference relation edges between the enterprise internal data source and the enterprise external data source according to the inference relation types; S66. Based on the set of inference entity nodes and the set of inference relation edges that have undergone semantic alignment processing and intelligent association processing, construct a unified dynamic optimization data analysis view for the enterprise, and the unified dynamic optimization data analysis view for the enterprise is used as the basic data input for generating the enterprise-wide data analysis results.
Citation Information
Patent Citations
Fusion reasoning method, system and equipment based on large language model and medium
CN117709468A
Data asset identification and risk early warning system and method based on financial knowledge graph and large language model
CN118247057A
Enterprise knowledge graph construction method and device
CN118410181A
Large language model knowledge graph construction method in financial field
CN119378660A
Multi-source heterogeneous big data fusion and reasoning method and system based on knowledge graph
CN119442151A
Cited By
Contract text key payment index extraction method based on natural language processing
CN120523929A
Enterprise intelligent decision-making method and system driven by causal atlas
CN120542981A
Management method of multi-mode enterprise knowledge base system
CN120561342A
Multi-model collaborative enterprise credit risk analysis method and system
CN120579829A
Large language model reasoning enhancement method based on structure perception knowledge graph representation
CN120633869A