Enterprise global data analysis method based on knowledge graph and large language model

By combining knowledge graphs and large language models, the fusion and intelligent analysis of enterprise multi-source heterogeneous data across the entire domain are solved, and data integration and reasoning with high accuracy and robustness are achieved, and intelligent integration and interpretability analysis of enterprise whole domain data is supported, which improves the flexibility and transparency of data analysis.

CN120218256BActive Publication Date: 2025-08-08SHANGSHANG (SUZHOU) DIGITAL TECHNOLOGY CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510694911.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-28
Publication Date
2025-08-08
Estimated Expiration
2045-05-28

AI Technical Summary

Technical Problem

The existing technology has problems such as single reasoning chain, insufficient natural language intention analysis, lack of real-time conflict detection and knowledge completion mechanism, weak semantic alignment ability of multi-source heterogeneous data, poor domain adaptability and interpretability, and lack of inference process traceability mechanism in the integration, inference and intelligent analysis of enterprise whole-domain multi-source heterogeneous data, resulting in insufficient accuracy and robustness of inference results.

Method used

Combining knowledge graphs and large language models, through two-way inference fusion, natural language intent analysis, adaptive knowledge completion, cross-original data semantic alignment and inference chain traceability mechanisms, we realize intelligent integration and interpretability analysis of enterprise whole-domain data, including data cleaning, entity mapping, symbol and semantic inference, logical conflict detection and completion, semantic alignment and inference chain traceability steps.

Benefits of technology

It improves inference accuracy, data fusion intelligence and decision-making transparency, enhances the data integration capabilities of enterprises in a multi-source heterogeneous data environment, improves the system's query flexibility and inference depth, ensures the consistency and integrity of the inference chain, and supports the full traceability and transparent audit of the inference process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120218256B_ABST
    Figure CN120218256B_ABST
Patent Text Reader

Abstract

The present invention discloses an enterprise global data analysis method based on a knowledge graph and a large language model, comprising the following steps: by collecting multi-source data from inside and outside the enterprise, and after cleaning, standardization and mapping, a unified data graph is constructed. After the user inputs a natural language query, the reasoning goal and task description are generated in combination with semantic understanding, and two reasoning paths are formed through structural reasoning and semantic reasoning respectively, and the two reasoning paths are fused for interactive verification and optimization. During the reasoning process, logical conflicts and data missing are detected in real time, and supplementary knowledge is automatically called to perform local reasoning corrections. Based on the corrected reasoning results, the system performs semantic alignment and association analysis on multi-source data, dynamically generates a unified enterprise data analysis view, and records the reasoning path and evolution information during the reasoning and analysis process to support subsequent traceable queries. The present invention improves the data integration, intelligent reasoning and decision support capabilities of enterprises in complex data environments.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of data analysis technology, and in particular to an enterprise global data analysis method based on knowledge graphs and large language models. Background Art

[0002] With the rapid development of information technology and data intelligence, enterprises are increasingly relying on multi-source data in their operations and decision-making. Especially given the growing demand for large-scale data integration and intelligent analysis, effectively integrating heterogeneous data from internal enterprise systems, external platforms, and third-party resources has become a crucial issue for improving intelligent enterprise decision-making. The combination of knowledge graphs, a means of expressing structured knowledge, and large language models, tools with powerful natural language understanding and reasoning capabilities, offers new possibilities for intelligent enterprise data processing.

[0003] However, existing technologies for integrating, reasoning, and intelligently analyzing enterprise-wide, multi-source, and heterogeneous data still have significant limitations, primarily in the following areas:

[0004] 1. Single reasoning chain: Existing methods mostly rely on single-path reasoning in knowledge graphs or language models, lacking the two-way integration of symbolic reasoning and semantic reasoning, resulting in insufficient accuracy and robustness of reasoning results.

[0005] 2. Insufficient natural language intent analysis: Traditional methods based on structured query language cannot effectively analyze complex natural language query intent, the generation of reasoning paths is limited, the flexibility is poor, and it cannot adapt to users' diverse query needs.

[0006] 3. Lack of real-time conflict detection and knowledge completion mechanisms during the reasoning process: The existing reasoning process usually performs unified verification after the reasoning is completed. It is unable to dynamically identify and complete semantic conflicts and logical gaps during the reasoning process, affecting the integrity and reliability of the reasoning chain.

[0007] 4. Weak semantic alignment capabilities for multi-source data: When faced with multi-source heterogeneous data from within and outside the enterprise, existing methods often find it difficult to effectively complete entity alignment and semantic fusion, resulting in fragmented and poorly correlated data analysis results.

[0008] 5. Poor domain adaptability and weak interpretability: When general language models or graph reasoning engines are applied in specific enterprise fields, the reasoning results are often not targeted and have poor interpretability, making it difficult to meet the enterprise's audit compliance and decision-making transparency requirements.

[0009] 6. Lack of reasoning process tracing mechanism: Traditional reasoning systems generally ignore the complete record of reasoning chain paths, reasoning node evolution, and reasoning relationship changes, making it difficult to support subsequent decision verification and problem tracing needs.

[0010] Therefore, how to provide an enterprise-wide data analysis method based on knowledge graphs and large language models is an urgent problem that technicians in this field need to solve. Summary of the Invention

[0011] One purpose of the present invention is to propose an enterprise global data analysis method based on knowledge graph and large language model. The present invention combines structured knowledge representation with natural language understanding and reasoning capabilities, and describes in detail the intelligent integration and explainable analysis of enterprise global data through two-way reasoning fusion, natural language intent analysis, adaptive knowledge completion, cross-source data semantic alignment and reasoning chain tracing mechanism in a multi-source heterogeneous data environment. It has the advantages of high reasoning accuracy, strong data fusion intelligence and good decision-making transparency.

[0012] According to an embodiment of the present invention, a method for analyzing enterprise global data based on a knowledge graph and a large language model includes the following steps:

[0013] S1. Collect internal and external enterprise data sources, perform data cleaning, normalization, and entity mapping, and generate an initial enterprise knowledge graph, which includes knowledge graph entity nodes, knowledge graph attribute nodes, and knowledge graph relationship edges.

[0014] S2. Receive a natural language query request, perform semantic parsing on the natural language query request using a large language model, extract the query intent, query target entity, and query association relationship elements, and generate a query intent graph and reasoning task description information;

[0015] S3. Call the enterprise's initial knowledge graph, perform symbolic reasoning based on the query intent graph and reasoning task description information to generate the first reasoning chain, and perform semantic reasoning in combination with the large language model to generate the second reasoning chain;

[0016] S4. Performing reasoning chain fusion processing on the first reasoning chain and the second reasoning chain, verifying them based on the reasoning path structure similarity index and the reasoning conclusion consistency index, and using a dynamic iterative optimization mechanism to generate a fused and optimized reasoning chain;

[0017] S5. Based on the fused and optimized reasoning chain, perform local conflict detection and adaptive knowledge completion processing, and output the corrected reasoning chain;

[0018] S6. Extract inference entity nodes and inference relationship edges based on the corrected inference chain, perform semantic alignment and intelligent association processing between the enterprise's internal and external data sources, and build a unified, dynamically optimized data analysis view for the enterprise.

[0019] S7. Based on the enterprise's unified dynamic optimization data analysis view, output the enterprise's global data analysis results and generate a reasoning chain traceability file.

[0020] Optionally, S2 includes the following steps:

[0021] S21. Receive a natural language query request input by an enterprise user, perform sentence segmentation and word segmentation processing on the natural language query request, and generate a set of natural language parsed segments;

[0022] S22. Using the large language model, perform semantic classification processing on the natural language parsed segment set, classify the natural language parsed segment set according to query intent segments, query target entity segments, and query association segment sets, and label the query intent category label, query target entity category label, and query association category label, respectively;

[0023] S23. Execute entity mapping search in the initial knowledge graph of the enterprise for the query target entity fragment, obtain a set of entity matching results of the knowledge graph entity nodes, and assign entity matching confidence to each knowledge graph entity node;

[0024] S24. Construct a query intent graph based on the query intent fragment, the query target entity fragment, and the query association relationship fragment. The query intent graph includes query intent layer nodes, query target entity layer nodes, and query association relationship layer edges, and the association relationships between them are annotated based on semantic relevance and context consistency.

[0025] S25. Set any node in the query intent graph With node The comprehensive correlation weight between , the comprehensive association weight is calculated according to the following formula:

[0026] ;

[0027] in, Representation node With node The semantic relevance function between Representation node With node The contextual consistency score between Representation node With node The shortest path length in the enterprise's initial knowledge graph, and is the comprehensive correlation weight adjustment coefficient;

[0028] S26. Generate reasoning task description information based on the node hierarchy of the query intent graph, the query association relationship layer edge set and the comprehensive association weight matrix. The reasoning task description information includes the query intent layer node set, the query target entity layer node set, the query association relationship layer edge set, the comprehensive association weight matrix and the reasoning priority sorting information.

[0029] Optionally, the S24 includes the following steps:

[0030] S241, extracting a query intent element set, a query target entity element set, and a query association relationship element set based on the query intent fragment, the query target entity fragment, and the query association relationship fragment, respectively, to generate a query intent graph;

[0031] S242, mapping each query intent element in the query intent element set, the query target entity element set, and the query association relationship element set to a query intent layer node, a query target entity layer node, and a query association relationship layer edge respectively;

[0032] S243. Based on the nodes and edges in the query intent graph, establish a semantic association between each query intent layer node in the query intent graph and one or more corresponding query target entity layer nodes, and connect each query target entity layer node in the query intent graph to other query target entity layer nodes according to the query association relationship layer edge;

[0033] S244. In the query intent graph, initial attributes are set for each query intent layer node, query target entity layer node, and query association layer edge, respectively. The initial attributes include a node category identifier, a node semantic vector representation, a node context vector representation, and an edge association category identifier.

[0034] S245: For any query intent layer node and query target entity layer node in the query intent graph With node , based on the node The node semantic vector representation and node The node semantic vector representation of the calculation node With node Semantic relevance function between ;

[0035] S246, Node-based The node context vector representation of the node The node context vector representation of the calculation node With node The contextual consistency score between .

[0036] Optionally, S3 includes the following steps:

[0037] S31. Based on the reasoning task description information, call the enterprise's initial knowledge graph, locate the corresponding knowledge graph entity node in the enterprise's initial knowledge graph according to each query target entity node in the query target entity layer node set, and generate a knowledge graph entity node set;

[0038] S32. Based on the set of knowledge graph entity nodes and the set of query association layer edges, perform symbolic reasoning along the knowledge graph relationship edges in the initial enterprise knowledge graph to generate a set of symbolic reasoning paths, where each symbolic reasoning path starts and ends at a knowledge graph entity node and meets the semantic requirements of the query association layer edges;

[0039] S33: Screening the symbolic reasoning path set based on the comprehensive association weight matrix, retaining symbolic reasoning paths whose comprehensive association weight values are greater than a preset weight threshold, and sorting the screened symbolic reasoning paths based on the reasoning priority sorting information to generate a first reasoning chain;

[0040] S34, inputting the reasoning task description information into the large language model, performing semantic reasoning processing based on the query intent layer node set, the query target entity layer node set, and the query association layer edge set, and deriving and generating a semantic reasoning path set;

[0041] S35. For each semantic reasoning path in the semantic reasoning path set, adjust the reasoning association order of the reasoning nodes in the semantic reasoning path according to the node hierarchy of the query intent graph, and calculate the semantic reasoning path score;

[0042] S36. Based on the comprehensive association weight matrix and the semantic reasoning path score, perform scoring processing on the semantic reasoning path set, filter the semantic reasoning paths with score values greater than the preset score threshold, and sort them according to the reasoning priority sorting information to generate a second reasoning chain.

[0043] Optionally, the S34 includes the following steps:

[0044] S331, traversing each symbolic reasoning path in the symbolic reasoning path set, and for each adjacent node pair in the symbolic reasoning path, extracting the corresponding comprehensive association weight value from the comprehensive association weight matrix;

[0045] S332: For each symbolic reasoning path, calculate the cumulative comprehensive weight of the symbolic reasoning path based on the comprehensive association weight value. , and the comprehensive weight of each symbolic reasoning path is accumulated With preset weight threshold Compare and filter out the cumulative value of comprehensive weight Greater than the preset weight threshold The symbolic reasoning paths constitute the symbolic reasoning path screening set;

[0046] S333: For each symbolic reasoning path in the symbolic reasoning path screening set, extract the reasoning priority sorting information in the reasoning task description information, and perform priority sorting processing on the symbolic reasoning paths in the symbolic reasoning path screening set according to the reasoning priority sorting information;

[0047] S334 , based on the set of symbolic reasoning paths sorted by reasoning priority, connect the symbolic reasoning paths in order of priority to generate a first reasoning chain.

[0048] Optionally, the S4 includes the following steps:

[0049] S41, extracting a set of reasoning paths in the first reasoning chain and a set of reasoning paths in the second reasoning chain, and marking the starting node and the ending node of each reasoning path in the set of reasoning paths respectively;

[0050] S42. Calculate the reasoning path structure similarity index for any two reasoning paths in the first reasoning chain and the second reasoning chain. , the reasoning path structure similarity index The node sequence matching degree calculation based on two reasoning paths uses the following formula:

[0051] ;

[0052] in, Indicates the number of nodes in the two reasoning paths. and Represents the number of nodes of the two reasoning paths respectively;

[0053] S43. Calculate the consistency index of the reasoning conclusion for any two reasoning paths in the first reasoning chain and the second reasoning chain. , the consistency index of the reasoning conclusion The consistency score between the target entity and the associated relationship derived from the two reasoning paths is determined;

[0054] S44, based on the inference path structure similarity index Consistency index with reasoning conclusion , filter out the reasoning path pairs that satisfy the similarity index greater than the preset structural similarity threshold and the consistency index greater than the preset consistency threshold, and form a reasoning path matching pair set;

[0055] S45. Based on the set of inference path matching pairs, a dynamic iterative optimization mechanism is adopted to adjust the connection order and node merging strategy of the inference path in rounds according to the inference priority and path importance index to generate a fused and optimized inference chain.

[0056] Optionally, the S44 includes the following steps:

[0057] S441. Setting a preset structural similarity threshold , the preset structural similarity threshold The statistical data of the matching degree of the historical reasoning path node sequence is set as the benchmark condition for screening the structural similarity of the reasoning path;

[0058] S442. Setting a preset consistency threshold , the preset consistency threshold The result of the weighted calculation of the consistency score of the inference target entity and the consistency score of the inference association relationship is set as the benchmark condition for the consistency screening of the inference conclusion;

[0059] S443. For each pair of reasoning paths in the first reasoning chain and the second reasoning chain, if the reasoning path structure similarity index Greater than or equal to the preset structural similarity threshold , and the consistency index of the inference conclusion Greater than or equal to the preset consistency threshold , then add the reasoning path pair to the reasoning path matching pair set;

[0060] S444: De-duplication processing is performed on the set of reasoning path matching pairs to eliminate duplicate reasoning path pairs with identical starting nodes, end nodes, and intermediate nodes, and retain unique reasoning path matching pairs.

[0061] Optionally, S5 includes the following steps:

[0062] S51, receiving the fused and optimized reasoning chain, and extracting the reasoning path set and the reasoning node set in the fused and optimized reasoning chain;

[0063] S52: Detecting logical conflicts within each reasoning path and between different reasoning paths in the reasoning path set after fusion optimization, wherein the logical conflicts include reasoning direction conflicts, entity attribute conflicts, and association relationship conflicts;

[0064] S53: For each logical conflict detected, extract the conflict node set and conflict relationship set corresponding to the logical conflict, and calculate the conflict severity according to the preset conflict severity scoring rules. , the severity of the conflict Calculated according to the following formula:

[0065] ;

[0066] in, represents the entity attribute conflict score, represents the association conflict score, represents the reasoning direction conflict score, 、 、 is the conflict score weighting coefficient, and satisfies ;

[0067] S54. Based on the severity of the conflict Compare with the preset conflict severity threshold to filter out the conflict severity For logical conflicts with a severity greater than a preset threshold, the enterprise's initial knowledge graph is used to retrieve and retrieve supplementary knowledge fragments for the selected logical conflicts. The supplementary knowledge fragments contain entity nodes, attribute nodes, and relationship edge information associated with the conflicting node set.

[0068] S55. Based on the completed knowledge fragment, perform local reasoning expansion processing, insert the entity nodes, attribute nodes, and relationship edges in the completed knowledge fragment into the corresponding conflict positions of the fused and optimized reasoning chain to form a corrected reasoning path set;

[0069] S56. Integrate the corrected reasoning path set with the non-conflicting reasoning paths in the fused and optimized reasoning chain to generate a corrected reasoning chain.

[0070] Optionally, the S54 includes the following steps:

[0071] S541: For each logical conflict detected in the optimized reasoning chain, the preset conflict severity scoring rule is called to calculate the corresponding conflict severity. ;

[0072] S542. Setting a preset conflict severity threshold , the preset conflict severity threshold Correct sample statistics based on historical reasoning chain conflicts, used as a benchmark for screening whether logical conflicts need to be completed;

[0073] S543. Comparison of conflict severity Conflict severity threshold If the severity of the conflict Greater than the preset conflict severity threshold , then add the logical conflict to the set of logical conflicts to be completed;

[0074] S544. For each logical conflict in the set of logical conflicts to be completed, extract the corresponding conflict node set and conflict relationship set to form a conflict element set;

[0075] S545. Based on the conflicting element set, a complete knowledge search process is performed in the initial knowledge graph of the enterprise, where the search conditions include the entity category, attribute category, and association relationship category of the nodes in the conflicting node set;

[0076] S546. Based on the search conditions, matching complementary knowledge fragments are retrieved in the enterprise initial knowledge graph, wherein the complementary knowledge fragments include entity nodes, attribute nodes, and relationship edge information, and the retrieved complementary knowledge fragments are screened according to the semantic matching scores with the conflicting node set;

[0077] S547: Output the filtered completed knowledge fragments as input data for local reasoning expansion processing.

[0078] Optionally, S6 includes the following steps:

[0079] S61, extracting the inference entity node set and the inference relationship edge set in the corrected inference chain, and recording the entity identification information of each inference entity node in the inference entity node set and the relationship type information of each inference relationship edge in the inference relationship edge set;

[0080] S62: For the set of inference entity nodes in the corrected inference chain, perform entity node matching retrieval in the enterprise internal data source and the enterprise external data source, respectively, and establish a preliminary entity matching set based on the entity identification information, entity attribute description information, and entity context semantic information of the inference entity node;

[0081] S63. Calculate the semantic alignment similarity for each pair of candidate matching entity nodes in the preliminary entity matching set. , the semantic alignment similarity Based on entity attribute similarity Consistency score with context Calculation, using the following formula:

[0082] ;

[0083] in, is the weighted coefficient of attribute similarity and contextual relationship consistency score, and satisfies ;

[0084] S64. Semantic alignment similarity Compare with the preset semantic alignment threshold to filter out the semantic alignment similarity Matching entity node pairs greater than a preset semantic alignment threshold and establishing intelligent association relationships between inferred entity nodes;

[0085] S65. For the selected matching entity node pairs, extract the corresponding reasoning relationship type in the reasoning relationship edge set, and establish an association relationship mapping of the reasoning relationship edge between the internal enterprise data source and the external enterprise data source based on the reasoning relationship type;

[0086] S66. Based on the set of inference entity nodes and the set of inference relationship edges that have undergone semantic alignment and intelligent association processing, a unified dynamic optimization data analysis view for the enterprise is constructed, and the unified dynamic optimization data analysis view for the enterprise serves as the basic data input for generating the enterprise's global data analysis results.

[0087] The beneficial effects of the present invention are:

[0088] (1) This invention proposes a bidirectional reasoning fusion mechanism by combining knowledge graph structured reasoning with large language model semantic reasoning, realizes mutual verification and dynamic optimization of reasoning paths, breaks the limitations of traditional single reasoning methods in data analysis accuracy and reasoning chain integrity, and effectively improves the data integration and intelligent reasoning capabilities of enterprises in multi-source heterogeneous data environments.

[0089] (2) This invention breaks through the limitations of traditional structured query-based systems by using natural language intent parsing and dynamic reasoning task generation technology, combined with symbolic reasoning and semantic reasoning dual-channel reasoning chain construction, enabling enterprise users to directly drive the reasoning and data analysis process through natural language, thereby improving the query flexibility and reasoning depth of the system.

[0090] (3) The present invention realizes the real-time identification and dynamic correction of logical conflicts and semantic gaps in the reasoning process through local logical conflict detection and adaptive knowledge completion mechanism, ensures the coherence and integrity of the reasoning chain, and enhances the stability of the reasoning process and the reliability of the reasoning results.

[0091] (4) The present invention establishes a dynamic semantic association system for internal and external multi-source heterogeneous data through semantic alignment and intelligent association processing technology, significantly improving the intelligence level of data fusion and promoting the unified management and in-depth analysis application of the enterprise's global data resources.

[0092] (5) The present invention combines a domain-fine-tuned large language model with knowledge reasoning optimization to form a highly adaptable and highly interpretable reasoning and analysis capability for enterprise-specific domain data. This is different from the traditional general model reasoning method and enhances the business relevance and audit compliance of the reasoning results.

[0093] (6) The present invention records the evolution of the reasoning chain path, reasoning nodes and reasoning relationships through the reasoning chain tracing mechanism, supports the full traceability and transparent audit of the reasoning process, and improves the reliability, transparency and explainability of the system in decision support scenarios. BRIEF DESCRIPTION OF THE DRAWINGS

[0094] The accompanying drawings are used to provide a further understanding of the present invention and constitute a part of the specification. Together with the embodiments of the present invention, they are used to explain the present invention and do not constitute a limitation of the present invention. In the accompanying drawings:

[0095] Figure 1 This is a flowchart of the enterprise global data analysis method based on knowledge graph and large language model proposed by the present invention;

[0096] Figure 2 This is a flowchart of the reasoning chain generation and fusion optimization of the enterprise global data analysis method based on knowledge graph and large language model proposed in this invention;

[0097] Figure 3 This is a flowchart of local conflict detection and adaptive knowledge completion for the enterprise global data analysis method based on knowledge graph and large language model proposed in the present invention. DETAILED DESCRIPTION

[0098] The present invention will now be described in further detail with reference to the accompanying drawings, which are simplified schematic diagrams that illustrate the basic structure of the present invention in a schematic manner.

[0099] refer to Figure 1-Figure 3 The enterprise global data analysis method based on knowledge graph and large language model includes the following steps:

[0100] S1. Collect internal and external enterprise data sources, perform data cleaning, normalization, and entity mapping, and generate an initial enterprise knowledge graph, which includes knowledge graph entity nodes, knowledge graph attribute nodes, and knowledge graph relationship edges.

[0101] This implementation method collects enterprise internal data sources (such as ERP systems, CRM systems, production monitoring platforms, etc.) and enterprise external data sources (such as industry databases, third-party intelligence platforms, open source data sets, etc.) through a unified access interface. The collected data is first cleaned to remove redundant, abnormal and duplicate information, and then normalized to unify data of different formats and standards into a standard structure. Entity mapping is performed based on entity recognition and mapping rules to map data items to uniformly defined entity categories and attribute relationships, and finally generate the enterprise's initial knowledge graph. The enterprise's initial knowledge graph consists of entity nodes, attribute nodes and relationship edges to form a unified structured semantic representation, realizing the standardized integration of multi-source heterogeneous data, breaking through enterprise data silos, laying the foundation for subsequent reasoning and intelligent analysis, and effectively improving data consistency, semantic clarity and analysis reliability.

[0102] S2. Receive a natural language query request, perform semantic parsing on the natural language query request using a large language model, extract the query intent, query target entity, and query association relationship elements, and generate a query intent graph and reasoning task description information;

[0103] S3. Call the enterprise's initial knowledge graph, perform symbolic reasoning based on the query intent graph and reasoning task description information to generate the first reasoning chain, and perform semantic reasoning in combination with the large language model to generate the second reasoning chain;

[0104] S4. Performing reasoning chain fusion processing on the first reasoning chain and the second reasoning chain, verifying them based on the reasoning path structure similarity index and the reasoning conclusion consistency index, and using a dynamic iterative optimization mechanism to generate a fused and optimized reasoning chain;

[0105] S5. Based on the fused and optimized reasoning chain, perform local conflict detection and adaptive knowledge completion processing, and output the corrected reasoning chain;

[0106] S6. Extract inference entity nodes and inference relationship edges based on the corrected inference chain, perform semantic alignment and intelligent association processing between the enterprise's internal and external data sources, and build a unified, dynamically optimized data analysis view for the enterprise.

[0107] S7. Based on the enterprise's unified dynamic optimization data analysis view, output the enterprise's global data analysis results and generate a reasoning chain traceability file.

[0108] This implementation method is based on the enterprise's unified dynamic optimization data analysis view, extracts reasoning entity nodes and reasoning relationship edges after semantic alignment and intelligent association processing, and combines the evolution process of each node in the reasoning process, relationship change records and dynamic adjustment information of the reasoning chain to generate complete enterprise-wide data analysis results. At the same time, it generates a reasoning chain traceability file in real time. The traceability file contains the reasoning starting point, reasoning path node sequence, relationship evolution trajectory and the formation process of the reasoning conclusion, realizing the full-link visualization and traceability of the reasoning process, which not only improves the transparency and audit compliance of data analysis results, but also enhances the enterprise's intelligence level in data governance, decision support and problem tracing, and provides solid data support for subsequent business optimization and risk prevention and control.

[0109] In this embodiment, S2 includes the following steps:

[0110] S21. Receive a natural language query request input by an enterprise user, perform sentence segmentation and word segmentation processing on the natural language query request, and generate a set of natural language parsed segments;

[0111] S22. Using the large language model, perform semantic classification processing on the natural language parsed segment set, classify the natural language parsed segment set according to query intent segments, query target entity segments, and query association segment sets, and label the query intent category label, query target entity category label, and query association category label, respectively;

[0112] S23. Execute entity mapping search in the initial knowledge graph of the enterprise for the query target entity fragment, obtain a set of entity matching results of the knowledge graph entity nodes, and assign entity matching confidence to each knowledge graph entity node;

[0113] S24. Construct a query intent graph based on the query intent fragment, the query target entity fragment, and the query association relationship fragment. The query intent graph includes query intent layer nodes, query target entity layer nodes, and query association relationship layer edges, and the association relationships between them are annotated based on semantic relevance and context consistency.

[0114] S25. Set any node in the query intent graph With node The comprehensive correlation weight between , the comprehensive association weight is calculated according to the following formula:

[0115] ;

[0116] in, Representation node With node The semantic relevance function between Representation node With node The contextual consistency score between Representation node With node The shortest path length in the enterprise's initial knowledge graph, and is the comprehensive correlation weight adjustment coefficient;

[0117] This formula combines the semantic relevance score and context consistency score between node pairs in the reasoning path, comprehensively considers the closeness of the node semantic relationship and the rationality of the reasoning path in the context, and introduces the shortest path distance between nodes in the knowledge graph as a regulating factor to construct a weight index that reflects the overall reasoning association strength of the node pair. The principle is that by weightedly fusing the static semantic relationship between nodes and the dynamic reasoning context environment, it not only ensures the logical coherence between nodes in the reasoning chain, but also takes into account the rational distribution of the reasoning path in the knowledge network, thereby screening out path structures that are more in line with the intention of the reasoning task during the reasoning process, and improving the accuracy and robustness of the reasoning chain generation.

[0118] S26. Generate reasoning task description information based on the node hierarchy of the query intent graph, the query association relationship layer edge set and the comprehensive association weight matrix. The reasoning task description information includes the query intent layer node set, the query target entity layer node set, the query association relationship layer edge set, the comprehensive association weight matrix and the reasoning priority sorting information.

[0119] This embodiment receives the natural language query request input by the enterprise user, performs sentence processing and word segmentation processing, generates a set of natural language parsed fragments, uses a large language model to semantically classify the parsed fragment set, and distinguishes them into query intent fragments, query target entity fragments and query association relationship fragments, and annotates them with category labels respectively. Then, for the query target entity fragments, an entity mapping search is performed in the enterprise's initial knowledge graph to generate an entity matching result set and assign matching confidence. A query intent graph is further constructed based on each fragment. By setting a comprehensive association weight function between nodes, the weight is calculated by combining the node semantic relevance, context consistency and graph structure distance. Finally, according to the query intent graph node hierarchy, association relationship layer edge set and comprehensive association weight matrix, the reasoning task description information is generated. The present invention effectively improves the understanding and reasoning conversion capabilities of natural language queries, enables the reasoning task to fully reflect the user's query intent and data semantic association characteristics, and enhances the accuracy, flexibility and intelligence of the reasoning path generation.

[0120] In this embodiment, the S24 specifically includes:

[0121] S241, extracting a query intent element set, a query target entity element set, and a query association relationship element set based on the query intent fragment, the query target entity fragment, and the query association relationship fragment, respectively, to generate a query intent graph;

[0122] S242, mapping each query intent element in the query intent element set, the query target entity element set, and the query association relationship element set to a query intent layer node, a query target entity layer node, and a query association relationship layer edge respectively;

[0123] S243. Based on the nodes and edges in the query intent graph, establish a semantic association between each query intent layer node in the query intent graph and one or more corresponding query target entity layer nodes, and connect each query target entity layer node in the query intent graph to other query target entity layer nodes according to the query association relationship layer edge;

[0124] S244. In the query intent graph, initial attributes are set for each query intent layer node, query target entity layer node, and query association layer edge, respectively. The initial attributes include a node category identifier, a node semantic vector representation, a node context vector representation, and an edge association category identifier.

[0125] S245: For any query intent layer node and query target entity layer node in the query intent graph With node , based on the node The node semantic vector representation and node The node semantic vector representation of the calculation node With node Semantic relevance function between ;

[0126] S246, Node-based The node context vector representation of the node The node context vector representation of the calculation node With node The contextual consistency score between .

[0127] This embodiment parses the query intent fragments, query target entity fragments and query association relationship fragments in natural language queries, extracts the corresponding query intent element sets, query target entity element sets and query association relationship element sets respectively, constructs a query intent graph, maps each element set into query intent layer nodes, query target entity layer nodes and query association relationship layer edges, and sets the initial attributes of each node and edge in the query intent graph, including node category identification, semantic vector representation and context vector representation, and then calculates the semantic relevance between nodes based on the semantic vectors of the nodes, and calculates the context consistency score between nodes based on the context vector, thereby realizing the fusion of structured semantic association and contextual semantics when constructing the query intent graph, improving the semantic accuracy and logical coherence of the reasoning path generation, providing high-quality reasoning task input for subsequent symbolic reasoning and semantic reasoning, and significantly enhancing the intelligence level and interpretability of natural language understanding and reasoning intent modeling in the process of enterprise data analysis.

[0128] In this embodiment, S3 includes the following steps:

[0129] S31. Based on the reasoning task description information, call the enterprise's initial knowledge graph, locate the corresponding knowledge graph entity node in the enterprise's initial knowledge graph according to each query target entity node in the query target entity layer node set, and generate a knowledge graph entity node set;

[0130] S32. Based on the set of knowledge graph entity nodes and the set of query association layer edges, perform symbolic reasoning along the knowledge graph relationship edges in the initial enterprise knowledge graph to generate a set of symbolic reasoning paths, where each symbolic reasoning path starts and ends at a knowledge graph entity node and meets the semantic requirements of the query association layer edges;

[0131] S33: Screening the symbolic reasoning path set based on the comprehensive association weight matrix, retaining symbolic reasoning paths whose comprehensive association weight values are greater than a preset weight threshold, and sorting the screened symbolic reasoning paths based on the reasoning priority sorting information to generate a first reasoning chain;

[0132] S34, inputting the reasoning task description information into the large language model, performing semantic reasoning processing based on the query intent layer node set, the query target entity layer node set, and the query association layer edge set, and deriving and generating a semantic reasoning path set;

[0133] S35. For each semantic reasoning path in the semantic reasoning path set, adjust the reasoning association order of the reasoning nodes in the semantic reasoning path according to the node hierarchy of the query intent graph, and calculate the semantic reasoning path score;

[0134] S36. Based on the comprehensive association weight matrix and the semantic reasoning path score, perform scoring processing on the semantic reasoning path set, filter the semantic reasoning paths with score values greater than the preset score threshold, and sort them according to the reasoning priority sorting information to generate a second reasoning chain.

[0135] This implementation method is based on the reasoning task description information. First, the query target entity node is located in the enterprise's initial knowledge graph, and symbolic reasoning processing is performed along the relationship edge of the knowledge graph to generate a set of symbolic reasoning paths that meet the query intent. The high-correlation paths are screened out based on the comprehensive correlation weight matrix to construct the first reasoning chain. At the same time, the reasoning task description information is input into the large language model, and semantic reasoning processing based on the query intent graph is performed to generate a set of semantic reasoning paths. The reasoning order and score are adjusted according to the hierarchical relationship of the reasoning nodes, and high-quality semantic reasoning paths are screened out to construct the second reasoning chain. Through the parallel generation and optimization of symbolic reasoning and semantic reasoning, the dual-channel construction of the reasoning chain is realized, which improves the accuracy, rationality and flexibility of the reasoning path generation, takes into account the rigor of structural reasoning and the flexibility of semantic reasoning, lays a solid foundation for the subsequent reasoning chain fusion and optimization, and enhances the robustness and intelligence of enterprise data analysis reasoning.

[0136] In this embodiment, the step S34 includes the following steps:

[0137] S331, traversing each symbolic reasoning path in the symbolic reasoning path set, and for each adjacent node pair in the symbolic reasoning path, extracting the corresponding comprehensive association weight value from the comprehensive association weight matrix;

[0138] S332: For each symbolic reasoning path, calculate the cumulative comprehensive weight of the symbolic reasoning path based on the comprehensive association weight value. , and the comprehensive weight of each symbolic reasoning path is accumulated With preset weight threshold Compare and filter out the cumulative value of comprehensive weight Greater than the preset weight threshold The symbolic reasoning paths constitute the symbolic reasoning path screening set;

[0139] S333: For each symbolic reasoning path in the symbolic reasoning path screening set, extract the reasoning priority sorting information in the reasoning task description information, and perform priority sorting processing on the symbolic reasoning paths in the symbolic reasoning path screening set according to the reasoning priority sorting information;

[0140] S334 , based on the set of symbolic reasoning paths sorted by reasoning priority, connect the symbolic reasoning paths in order of priority to generate a first reasoning chain.

[0141] This implementation traverses the symbolic reasoning path set, extracts the weight values of adjacent node pairs in the comprehensive association weight matrix, calculates the cumulative comprehensive weight value of the symbolic reasoning path, and compares it with the preset weight threshold, screens out symbolic reasoning paths with high correlation to form a screening set, then sorts the symbolic reasoning paths in the screening set according to the reasoning priority in the reasoning task description information, and finally connects them in priority order to generate the first reasoning chain. This can effectively improve the relevance and rationality of reasoning path screening, ensure that the reasoning process is more in line with user intentions and knowledge structure logic, enhance the accuracy of reasoning chain generation and the reliability of the reasoning as a whole, and provide high-quality input support for subsequent reasoning chain fusion and optimization.

[0142] In this embodiment, the step S4 includes the following steps:

[0143] S41, extracting a set of reasoning paths in the first reasoning chain and a set of reasoning paths in the second reasoning chain, and marking the starting node and the ending node of each reasoning path in the set of reasoning paths respectively;

[0144] S42. Calculate the reasoning path structure similarity index for any two reasoning paths in the first reasoning chain and the second reasoning chain. , the reasoning path structure similarity index The node sequence matching degree calculation based on two reasoning paths uses the following formula:

[0145] ;

[0146] in, Indicates the number of nodes in the two reasoning paths. and Represents the number of nodes of the two reasoning paths respectively;

[0147] This formula uses the inference path structural similarity index to evaluate the structural matching between inference paths, which serves to measure the degree of similarity between the node structure arrangements of two inference paths. Specifically, the calculation principle of inference path structural similarity is based on the proportional relationship between the number of shared nodes in two inference paths and the minimum number of nodes in each path, reflecting the degree of overlap in the node composition and order of the two inference paths. The more shared nodes there are and the closer they are to the total number of nodes in the shorter path, the higher the inference path structural similarity score, indicating that the two inference paths are structurally closer. This helps to screen out inference path pairs with strong structural consistency in the subsequent inference chain fusion process, thereby improving the accuracy and reliability of inference chain fusion optimization.

[0148] S43. Calculate the consistency index of the reasoning conclusion for any two reasoning paths in the first reasoning chain and the second reasoning chain. , the consistency index of the reasoning conclusion The consistency score between the target entity and the associated relationship derived from the two reasoning paths is determined;

[0149] S44, based on the inference path structure similarity index Consistency index with reasoning conclusion , filter out the reasoning path pairs that satisfy the similarity index greater than the preset structural similarity threshold and the consistency index greater than the preset consistency threshold, and form a reasoning path matching pair set;

[0150] S45. Based on the set of inference path matching pairs, a dynamic iterative optimization mechanism is adopted to adjust the connection order and node merging strategy of the inference path in rounds according to the inference priority and path importance index to generate a fused and optimized inference chain.

[0151] This implementation method extracts the reasoning path sets in the first reasoning chain and the second reasoning chain, marks the starting node and the end node of each reasoning path respectively, calculates the reasoning path structure similarity index based on node sequence matching, and calculates the reasoning conclusion consistency index based on the derived target entity and association relationship score, further screens out the reasoning path pairs that meet the dual standards of structural similarity and conclusion consistency, constructs a reasoning path matching pair set, and adopts a dynamic iterative optimization mechanism on this basis to gradually adjust the reasoning path connection order and node merging strategy according to the reasoning priority and path importance index, and finally generates a fused and optimized reasoning chain, realizing two-way interactive verification and dynamic adjustment of the reasoning chain, effectively improving the accuracy, robustness and consistency of the reasoning process of the reasoning chain fusion in the enterprise's global data analysis, and enhancing the stability and interpretability of the reasoning results.

[0152] In this embodiment, the S44 includes the following steps:

[0153] S441. Setting a preset structural similarity threshold , the preset structural similarity threshold The statistical data of the matching degree of the historical reasoning path node sequence is set as the benchmark condition for screening the structural similarity of the reasoning path;

[0154] S442. Setting a preset consistency threshold , the preset consistency threshold The result of the weighted calculation of the consistency score of the inference target entity and the consistency score of the inference association relationship is set as the benchmark condition for the consistency screening of the inference conclusion;

[0155] S443. For each pair of reasoning paths in the first reasoning chain and the second reasoning chain, if the reasoning path structure similarity index Greater than or equal to the preset structural similarity threshold , and the consistency index of the inference conclusion Greater than or equal to the preset consistency threshold , then add the reasoning path pair to the reasoning path matching pair set;

[0156] S444: De-duplication processing is performed on the set of reasoning path matching pairs to eliminate duplicate reasoning path pairs with identical starting nodes, end nodes, and intermediate nodes, and retain unique reasoning path matching pairs.

[0157] This embodiment performs dual-index screening on the reasoning path pairs in the first reasoning chain and the second reasoning chain by setting a preset structural similarity threshold based on the statistics of the matching degree of the historical reasoning path node sequence, and a preset consistency threshold based on the weighted calculation of the inference target entity consistency score and the inference association relationship consistency score, and screens out the reasoning path pairs that meet both the structural similarity and the consistency of the reasoning conclusion, and performs node and path deduplication processing on the matched reasoning path pairs, so as to retain the only valid reasoning path matching relationship, realizing high-precision path screening and structural optimization before the reasoning chain fusion, effectively improving the accuracy, reliability and consistency of the reasoning chain in the reasoning fusion process, and laying a solid foundation for the subsequent dynamic optimization of the reasoning chain and the credibility of the reasoning results.

[0158] In this embodiment, S5 includes the following steps:

[0159] S51, receiving the fused and optimized reasoning chain, and extracting the reasoning path set and the reasoning node set in the fused and optimized reasoning chain;

[0160] S52: Detecting logical conflicts within each reasoning path and between different reasoning paths in the reasoning path set after fusion optimization, wherein the logical conflicts include reasoning direction conflicts, entity attribute conflicts, and association relationship conflicts;

[0161] S53: For each logical conflict detected, extract the conflict node set and conflict relationship set corresponding to the logical conflict, and calculate the conflict severity according to the preset conflict severity scoring rules. , the severity of the conflict Calculated according to the following formula:

[0162] ;

[0163] in, represents the entity attribute conflict score, represents the association conflict score, represents the reasoning direction conflict score, 、 、 is the conflict score weighting coefficient, and satisfies ;

[0164] This formula takes a weighted sum of the entity attribute conflict scores, relationship conflict scores, and reasoning direction conflict scores involved in each logical conflict. The weight coefficients used adjust the contribution of different conflict types to the overall severity. Each conflict score reflects the degree of deviation in attribute consistency, relationship rationality, and reasoning path direction. A higher overall conflict severity value indicates a greater impact on the coherence and accuracy of the reasoning chain. This guides the system to prioritize high-severity conflicts, improving the targeted and efficient correction of reasoning chains.

[0165] S54. Based on the severity of the conflict Compare with the preset conflict severity threshold to filter out the conflict severity For logical conflicts with a severity greater than a preset threshold, the enterprise's initial knowledge graph is used to retrieve and retrieve supplementary knowledge fragments for the selected logical conflicts. The supplementary knowledge fragments contain entity nodes, attribute nodes, and relationship edge information associated with the conflicting node set.

[0166] S55. Based on the completed knowledge fragment, perform local reasoning expansion processing, insert the entity nodes, attribute nodes, and relationship edges in the completed knowledge fragment into the corresponding conflict positions of the fused and optimized reasoning chain to form a corrected reasoning path set;

[0167] S56. Integrate the corrected reasoning path set with the non-conflicting reasoning paths in the fused and optimized reasoning chain to generate a corrected reasoning chain.

[0168] This implementation receives the fused and optimized reasoning chain, extracts the reasoning path set and the reasoning node set, detects reasoning direction conflicts, entity attribute conflicts and association relationship conflicts within and between reasoning paths, and calculates the conflict severity based on preset scoring rules. When the conflict severity exceeds the preset threshold, the enterprise's initial knowledge graph is called to retrieve and complete the knowledge fragments, and local reasoning expansion is performed based on the completed knowledge fragments to effectively correct the logical defects of the reasoning chain. Finally, the corrected reasoning path is integrated to generate a complete reasoning chain, which realizes the dynamic identification and repair of logical conflicts and semantic deficiencies during the reasoning process, improves the integrity of the reasoning chain, the accuracy of reasoning inferences and the stability of the data analysis process, and provides reliable and coherent data support for enterprise intelligent decision-making.

[0169] In this embodiment, the S54 includes the following steps:

[0170] S541: For each logical conflict detected in the optimized reasoning chain, the preset conflict severity scoring rule is called to calculate the corresponding conflict severity. ;

[0171] S542. Setting a preset conflict severity threshold , the preset conflict severity threshold Correct sample statistics based on historical reasoning chain conflicts, used as a benchmark for screening whether logical conflicts need to be completed;

[0172] S543. Comparison of conflict severity Conflict severity threshold If the severity of the conflict Greater than the preset conflict severity threshold , then add the logical conflict to the set of logical conflicts to be completed;

[0173] S544. For each logical conflict in the set of logical conflicts to be completed, extract the corresponding conflict node set and conflict relationship set to form a conflict element set;

[0174] S545. Based on the conflicting element set, a complete knowledge search process is performed in the initial knowledge graph of the enterprise, where the search conditions include the entity category, attribute category, and association relationship category of the nodes in the conflicting node set;

[0175] S546. Based on the search conditions, matching complementary knowledge fragments are retrieved in the enterprise initial knowledge graph, wherein the complementary knowledge fragments include entity nodes, attribute nodes, and relationship edge information, and the retrieved complementary knowledge fragments are screened according to the semantic matching scores with the conflicting node set;

[0176] S547: Output the filtered completed knowledge fragments as input data for local reasoning expansion processing.

[0177] This embodiment calls the conflict severity scoring rule to calculate the conflict severity of the logical conflicts detected in the integrated and optimized reasoning chain, screens them according to the preset conflict severity threshold set by the historical reasoning chain conflict samples, and incorporates the logical conflicts with conflict severity exceeding the threshold into the process to be completed. It further extracts the conflict node set and the conflict relationship set to form a conflict element set, and uses this as the retrieval condition to perform completed knowledge retrieval in the enterprise's initial knowledge graph, screens the completed knowledge fragments with a high semantic match with the conflict node set, and finally uses the screened completed knowledge fragments as input data for local reasoning expansion, thereby realizing dynamic detection of key conflicts, knowledge completion and adaptive optimization of the reasoning chain during the reasoning process, effectively improving the integrity and logical coherence of the reasoning chain structure, and enhancing the accuracy and stability of reasoning and inference in the enterprise's global data analysis.

[0178] In this embodiment, S6 includes the following steps:

[0179] S61, extracting the inference entity node set and the inference relationship edge set in the corrected inference chain, and recording the entity identification information of each inference entity node in the inference entity node set and the relationship type information of each inference relationship edge in the inference relationship edge set;

[0180] S62: For the set of inference entity nodes in the corrected inference chain, perform entity node matching retrieval in the enterprise internal data source and the enterprise external data source, respectively, and establish a preliminary entity matching set based on the entity identification information, entity attribute description information, and entity context semantic information of the inference entity node;

[0181] S63. Calculate the semantic alignment similarity for each pair of candidate matching entity nodes in the preliminary entity matching set. , the semantic alignment similarity Based on entity attribute similarity Consistency score with context Calculation, using the following formula:

[0182] ;

[0183] in, is the weighted coefficient of attribute similarity and contextual relationship consistency score, and satisfies ;

[0184] This formula measures the comprehensive match between the attribute descriptions and contextual relationships of the inferred entity node and the data source entity node, helping to screen pairs of entity nodes with high semantic relevance. The semantic alignment similarity calculation principle is a weighted fusion of the similarity scores between entity attributes and the consistency scores of contextual relationships. The weight coefficient can be set according to actual needs to reflect the relative importance of attribute features and context. By comprehensively considering the similarity of the entity attributes themselves and the consistency of their contextual environment, the semantic correspondence between entity nodes can be more comprehensively and accurately assessed, providing a reliable basis for subsequent data semantic alignment and intelligent association processing.

[0185] S64. Semantic alignment similarity Compare with the preset semantic alignment threshold to filter out the semantic alignment similarity Matching entity node pairs greater than a preset semantic alignment threshold and establishing an intelligent association relationship between the inference entity nodes. The preset semantic alignment threshold is used as a judgment standard for screening the semantic matching relationship between the inference entity node and the data source entity node. When the semantic alignment similarity is greater than the threshold, it is determined that the match is successful and an intelligent association relationship is established;

[0186] S65. For the selected matching entity node pairs, extract the corresponding reasoning relationship type in the reasoning relationship edge set, and establish an association relationship mapping of the reasoning relationship edge between the internal enterprise data source and the external enterprise data source based on the reasoning relationship type;

[0187] S66. Based on the set of inference entity nodes and the set of inference relationship edges that have undergone semantic alignment and intelligent association processing, a unified dynamic optimization data analysis view for the enterprise is constructed, and the unified dynamic optimization data analysis view for the enterprise serves as the basic data input for generating the enterprise's global data analysis results.

[0188] This embodiment extracts the inference entity node set and the inference relationship edge set in the corrected inference chain, performs matching retrieval in the enterprise's internal data source and the enterprise's external data source based on entity identification information, attribute description information and contextual semantic information, calculates semantic alignment similarity and screens out matching entity node pairs that meet the preset semantic alignment threshold, establishes intelligent association relationships of inference entity nodes, and constructs association mappings of inference relationship edges between different data sources according to the inference relationship type. Finally, based on the inference entity node set and the inference relationship edge set after semantic alignment and intelligent association processing, a unified enterprise data analysis view is dynamically generated, realizing deep integration and intelligent fusion of multi-source heterogeneous data, improving the consistency and accuracy of data analysis, and providing a solid data foundation for subsequent global data intelligent analysis and decision support.

[0189] Example 1:

[0190] In order to verify the feasibility of the present invention in implementation, the present invention was applied to the data intelligent analysis platform of a nationally renowned large-scale manufacturing group. In the course of daily operations, the group has accumulated huge data sets from multiple fields such as production, supply chain, customer service, external market environment, etc., involving ERP system data, MES system data, CRM system data, as well as multi-source heterogeneous data such as external industry reports and regulatory information. Due to the diverse sources of data, complex formats and inconsistent semantics, traditional data integration and analysis methods are difficult to meet the group's needs for real-time, intelligent, and explainable data analysis and decision support. There are problems such as data islands, broken reasoning chains, and frequent semantic conflicts, which lead to fragmented and incoherent decision-making basis, restricting the development of enterprise intelligent transformation. Therefore, the group decided to introduce the enterprise-wide data analysis method based on the fusion reasoning of knowledge graph and large language model proposed in the present invention, aiming to break through data barriers and enhance data integration depth and decision-making support capabilities.

[0191] In specific applications, the group first standardized data preprocessing for internal ERP, MES, CRM, and other systems, including data cleaning, format normalization, and entity mapping, before uniformly mapping it to the company's initial knowledge graph. The entity nodes, attribute nodes, and relationship edges in the knowledge graph were constructed according to a unified enterprise ontology specification. Simultaneously, externally purchased industry data, supplier reports, and documents were also standardized and incorporated into the company's external data knowledge graph, ultimately forming a dual-graph fusion structure that integrates both internal and external data.

[0192] In the actual business analysis process, business personnel make complex data analysis requests through natural language, such as "under the influence of the new environment, the most affected raw material categories and their upstream and downstream enterprises in the supply chain this quarter". The method of the present invention first parses the natural language request through a large language model, extracts the query intent, target entity and association relationship elements, and generates a query intent graph and reasoning task description information. In the reasoning stage, on the one hand, the system performs symbolic reasoning based on the knowledge graph to generate a symbolic reasoning chain. On the other hand, it combines the large language model for semantic reasoning to generate a semantic reasoning chain, and performs two-way fusion verification through the reasoning path structure similarity and the reasoning conclusion consistency index, and uses a dynamic iteration mechanism to generate a fused and optimized reasoning chain.

[0193] During the inference process, the system automatically detects logical conflicts within the inference chain, such as broken supply chain entity relationships and contradictory attribute descriptions. It then calls upon relevant domain knowledge from the enterprise's initial knowledge graph in real time to complete local knowledge and resolve conflicts, ensuring the integrity and consistency of the inference chain. For the fused and optimized inference chain, the system extracts inference entity nodes and inference relationship edges, performs entity matching in both internal and external enterprise data sources, and uses a semantic alignment similarity metric, calculated based on a weighted calculation of entity attribute similarity and contextual relationship consistency, to select the best matching entity pairs, complete semantic alignment, and establish intelligent entity association relationships.

[0194] Based on the reasoning chain after semantic alignment and intelligent association processing, the system dynamically constructs a unified and optimized data analysis view, connecting the raw material category nodes affected by the new environment and their upstream and downstream enterprise nodes into a dynamic relationship diagram with clear semantics and logical coherence, and fully records each path, node change and relationship evolution process in the reasoning process in the reasoning chain traceability file, providing a transparent and traceable basis for subsequent decision-making audits and compliance.

[0195] To verify the effectiveness of the proposed method, the Group conducted a comparative experiment using data from the past year. The experiment selected three typical business scenarios: supply chain risk analysis, customer complaint tracing analysis, and financial anomaly warning analysis. Traditional data warehouse query methods were used for analysis and comparison, respectively. The experimental period lasted three months, and the specific data is shown in the table below.

[0196] Table 1 Comparison of data analysis effects between traditional methods and the method of the present invention in different business scenarios

[0197] ;

[0198] From the experimental results, it can be seen that the method of the present invention has shortened the response time of supply chain risk analysis by about 78%, increased the accuracy of customer complaint tracing by nearly 20 percentage points, and extended the lead time of financial anomaly warning from 2 days to 7 days, effectively improving the group's sensitivity and response capabilities to operating risks. In terms of data fusion integrity, the present invention has increased the multi-source data integration rate to more than 95% through semantic alignment and intelligent association mechanisms, and significantly improved the reasoning chain coherence score and the reasoning chain traceability completeness rate, providing enterprises with highly reliable and transparent data analysis support. At the same time, the accuracy of understanding users' natural language query intentions has increased to more than 93%, greatly improving the data interaction experience of business personnel and lowering the technical threshold.

[0199] In practical applications, the present invention also demonstrates strong dynamic adaptability. Taking supply chain analysis as an example, when external updates or changes in the industry environment occur, the system can update the reasoning chain in real time based on supplementary knowledge fragments and dynamically rebuild the data analysis view, avoiding the tedious process of manually rebuilding the data warehouse and logical rules under the traditional static model. For example, in March 2025, after a certain raw material was exported, the system completed the reasoning chain adjustment and data view update of the affected supply chain nodes in just 6 hours, and automatically pushed the early warning report to the group decision-making level, greatly improving the company's emergency response speed.

[0200] In summary, this example fully demonstrates the significant advantages of the present invention's method in intelligently integrating enterprise-wide, multi-source, heterogeneous data, dynamically building reasoning chains, understanding data analysis requests driven by natural language, completing local knowledge, aligning cross-source semantics, intelligently associating entities, and tracing reasoning chains throughout the entire process. Through the application of this invention, enterprises not only achieve simultaneous improvements in data integration efficiency and analytical depth, but also significantly enhance the transparency, traceability, and adaptability of the decision-making process, providing strong technical support for intelligent enterprise operations in complex environments.

[0201] The above description is only a preferred specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any technician familiar with the technical field, within the technical scope disclosed by the present invention, who makes equivalent replacements or changes based on the technical solution and inventive concept of the present invention, should be covered by the scope of protection of the present invention.

Claims

1. A method for analyzing enterprise global data based on knowledge graph and large language model, characterized by: The steps include: S1. Collect internal and external enterprise data sources, perform data cleaning, normalization, and entity mapping, and generate an initial enterprise knowledge graph, which includes knowledge graph entity nodes, knowledge graph attribute nodes, and knowledge graph relationship edges. S2. Receive a natural language query request, perform semantic parsing on the natural language query request using a large language model, extract the query intent, query target entity, and query association relationship elements, and generate a query intent graph and reasoning task description information; S3. Call the enterprise's initial knowledge graph, perform symbolic reasoning based on the query intent graph and reasoning task description information to generate the first reasoning chain, and perform semantic reasoning in combination with the large language model to generate the second reasoning chain; S4. Performing reasoning chain fusion processing on the first reasoning chain and the second reasoning chain, verifying them based on the reasoning path structure similarity index and the reasoning conclusion consistency index, and using a dynamic iterative optimization mechanism to generate a fused and optimized reasoning chain; S5. Based on the fused and optimized reasoning chain, perform local conflict detection and adaptive knowledge completion processing, and output the corrected reasoning chain; S6. Extract inference entity nodes and inference relationship edges based on the corrected inference chain, perform semantic alignment and intelligent association processing between the enterprise's internal and external data sources, and build a unified, dynamically optimized data analysis view for the enterprise. S7. Based on the enterprise's unified dynamic optimization data analysis view, output the enterprise's global data analysis results and generate a reasoning chain traceability file; The S4 comprises the following steps: S41, extracting a set of reasoning paths in the first reasoning chain and a set of reasoning paths in the second reasoning chain, and marking the starting node and the ending node of each reasoning path in the set of reasoning paths respectively; S42. Calculate the reasoning path structure similarity index for any two reasoning paths in the first reasoning chain and the second reasoning chain. , the reasoning path structure similarity index The node sequence matching degree calculation based on two reasoning paths uses the following formula: ; in, Indicates the number of nodes in the two reasoning paths. and Represents the number of nodes of the two reasoning paths respectively; S43. Calculate the consistency index of the reasoning conclusion for any two reasoning paths in the first reasoning chain and the second reasoning chain. , the consistency index of the reasoning conclusion The consistency score between the target entity and the associated relationship derived from the two reasoning paths is determined; S44, based on the inference path structure similarity index Consistency index with reasoning conclusion , filter out the reasoning path pairs that satisfy the similarity index greater than the preset structural similarity threshold and the consistency index greater than the preset consistency threshold, and form a reasoning path matching pair set; S45. Based on the set of inference path matching pairs, a dynamic iterative optimization mechanism is used to adjust the connection order and node merging strategy of the inference paths in rounds according to the inference priority and path importance index to generate a fused and optimized inference chain. The S44 includes the following steps: S441. Setting a preset structural similarity threshold , the preset structural similarity threshold The statistical data of the matching degree of the historical reasoning path node sequence is set as the benchmark condition for screening the structural similarity of the reasoning path; S442. Setting a preset consistency threshold , the preset consistency threshold The result of the weighted calculation of the consistency score of the inference target entity and the consistency score of the inference association relationship is set as the benchmark condition for the consistency screening of the inference conclusion; S443. For each pair of reasoning paths in the first reasoning chain and the second reasoning chain, if the reasoning path structure similarity index Greater than or equal to the preset structural similarity threshold , and the consistency index of the inference conclusion Greater than or equal to the preset consistency threshold , then add the reasoning path pair to the reasoning path matching pair set; S444: performing deduplication processing on the set of reasoning path matching pairs, eliminating duplicate reasoning path pairs with identical starting nodes, end nodes, and intermediate nodes, and retaining unique reasoning path matching pairs; The S5 comprises the following steps: S51, receiving the fused and optimized reasoning chain, and extracting the reasoning path set and the reasoning node set in the fused and optimized reasoning chain; S52: Detecting logical conflicts within each reasoning path and between different reasoning paths in the reasoning path set after fusion optimization, wherein the logical conflicts include reasoning direction conflicts, entity attribute conflicts, and association relationship conflicts; S53: For each logical conflict detected, extract the conflict node set and conflict relationship set corresponding to the logical conflict, and calculate the conflict severity according to the preset conflict severity scoring rules. , the severity of the conflict Calculated according to the following formula: ; in, represents the entity attribute conflict score, represents the association conflict score, represents the reasoning direction conflict score, 、 、 is the conflict score weighting coefficient, and satisfies ; S54. Based on the severity of the conflict Compare with the preset conflict severity threshold to filter out the conflict severity For logical conflicts with a severity greater than a preset threshold, the enterprise's initial knowledge graph is used to retrieve and retrieve supplementary knowledge fragments for the selected logical conflicts. The supplementary knowledge fragments contain entity nodes, attribute nodes, and relationship edge information associated with the conflicting node set. S55. Based on the completed knowledge fragment, perform local reasoning expansion processing, insert the entity nodes, attribute nodes, and relationship edges in the completed knowledge fragment into the corresponding conflict positions of the fused and optimized reasoning chain to form a corrected reasoning path set; S56. Integrate the corrected reasoning path set with the non-conflicting reasoning paths in the fused and optimized reasoning chain to generate a corrected reasoning chain.

2. The enterprise global data analysis method based on knowledge graph and large language model according to claim 1 is characterized in that: The S2 comprises the following steps: S21. Receive a natural language query request input by an enterprise user, perform sentence segmentation and word segmentation processing on the natural language query request, and generate a set of natural language parsed segments; S22. Using the large language model, perform semantic classification processing on the natural language parsed segment set, classify the natural language parsed segment set according to query intent segments, query target entity segments, and query association segment sets, and label the query intent category label, query target entity category label, and query association category label, respectively; S23. Execute entity mapping search in the initial knowledge graph of the enterprise for the query target entity fragment, obtain a set of entity matching results of the knowledge graph entity nodes, and assign entity matching confidence to each knowledge graph entity node; S24. Construct a query intent graph based on the query intent fragment, the query target entity fragment, and the query association relationship fragment. The query intent graph includes query intent layer nodes, query target entity layer nodes, and query association relationship layer edges, and the association relationships between them are annotated based on semantic relevance and context consistency. S25. Set the comprehensive association weight between any node i and node j in the query intent graph , the comprehensive association weight is calculated according to the following formula: ; in, Representation node With node The semantic relevance function between Representation node With node The contextual consistency score between Representation node With node The shortest path length in the enterprise's initial knowledge graph, and is the comprehensive correlation weight adjustment coefficient; S26. Generate reasoning task description information based on the node hierarchy of the query intent graph, the query association relationship layer edge set and the comprehensive association weight matrix. The reasoning task description information includes the query intent layer node set, the query target entity layer node set, the query association relationship layer edge set, the comprehensive association weight matrix and the reasoning priority sorting information.

3. The enterprise global data analysis method based on knowledge graph and large language model according to claim 2 is characterized in that: The S24 specifically includes: S241, extracting a query intent element set, a query target entity element set, and a query association relationship element set based on the query intent fragment, the query target entity fragment, and the query association relationship fragment, respectively, to generate a query intent graph; S242, mapping each query intent element in the query intent element set, the query target entity element set, and the query association relationship element set to a query intent layer node, a query target entity layer node, and a query association relationship layer edge respectively; S243. Based on the nodes and edges in the query intent graph, establish a semantic association between each query intent layer node in the query intent graph and one or more corresponding query target entity layer nodes, and connect each query target entity layer node in the query intent graph to other query target entity layer nodes according to the query association relationship layer edge; S244. In the query intent graph, initial attributes are set for each query intent layer node, query target entity layer node, and query association layer edge, respectively. The initial attributes include a node category identifier, a node semantic vector representation, a node context vector representation, and an edge association category identifier. S245: For any query intent layer node and query target entity layer node in the query intent graph With node , based on the node The node semantic vector representation and node The node semantic vector representation of the calculation node With node Semantic relevance function between ; S246, Node-based The node context vector representation of the node The node context vector representation of the calculation node With node The contextual consistency score between .

4. The enterprise global data analysis method based on knowledge graph and large language model according to claim 1 is characterized in that: The S3 comprises the following steps: S31. Based on the reasoning task description information, call the enterprise's initial knowledge graph, locate the corresponding knowledge graph entity node in the enterprise's initial knowledge graph according to each query target entity node in the query target entity layer node set, and generate a knowledge graph entity node set; S32. Based on the set of knowledge graph entity nodes and the set of query association layer edges, perform symbolic reasoning along the knowledge graph relationship edges in the initial enterprise knowledge graph to generate a set of symbolic reasoning paths, where each symbolic reasoning path starts and ends at a knowledge graph entity node and meets the semantic requirements of the query association layer edges; S33: Screening the symbolic reasoning path set based on the comprehensive association weight matrix, retaining symbolic reasoning paths whose comprehensive association weight values are greater than a preset weight threshold, and sorting the screened symbolic reasoning paths based on the reasoning priority sorting information to generate a first reasoning chain; S34, inputting the reasoning task description information into the large language model, performing semantic reasoning processing based on the query intent layer node set, the query target entity layer node set, and the query association layer edge set, and deriving and generating a semantic reasoning path set; S35. For each semantic reasoning path in the semantic reasoning path set, adjust the reasoning association order of the reasoning nodes in the semantic reasoning path according to the node hierarchy of the query intent graph, and calculate the semantic reasoning path score; S36. Based on the comprehensive association weight matrix and the semantic reasoning path score, perform scoring processing on the semantic reasoning path set, filter the semantic reasoning paths with score values greater than the preset score threshold, and sort them according to the reasoning priority sorting information to generate a second reasoning chain.

5. The enterprise global data analysis method based on knowledge graph and large language model according to claim 4 is characterized in that: The S33 includes the following steps: S331, traversing each symbolic reasoning path in the symbolic reasoning path set, and for each adjacent node pair in the symbolic reasoning path, extracting the corresponding comprehensive association weight value from the comprehensive association weight matrix; S332: For each symbolic reasoning path, calculate the cumulative comprehensive weight of the symbolic reasoning path based on the comprehensive association weight value. , and the comprehensive weight of each symbolic reasoning path is accumulated With preset weight threshold Compare and filter out the cumulative value of comprehensive weight Greater than the preset weight threshold The symbolic reasoning paths constitute the symbolic reasoning path screening set; S333: For each symbolic reasoning path in the symbolic reasoning path screening set, extract the reasoning priority sorting information in the reasoning task description information, and perform priority sorting processing on the symbolic reasoning paths in the symbolic reasoning path screening set according to the reasoning priority sorting information; S334 , based on the set of symbolic reasoning paths sorted by reasoning priority, connect the symbolic reasoning paths in order of priority to generate a first reasoning chain.

6. The enterprise global data analysis method based on knowledge graph and large language model according to claim 1 is characterized in that: The S54 includes the following steps: S541: For each logical conflict detected in the optimized reasoning chain, the preset conflict severity scoring rule is called to calculate the corresponding conflict severity. ; S542. Setting a preset conflict severity threshold , the preset conflict severity threshold Correct sample statistics based on historical reasoning chain conflicts, used as a benchmark for screening whether logical conflicts need to be completed; S543. Comparison of conflict severity Conflict severity threshold If the severity of the conflict Greater than the preset conflict severity threshold , then add the logical conflict to the set of logical conflicts to be completed; S544. For each logical conflict in the set of logical conflicts to be completed, extract the corresponding conflict node set and conflict relationship set to form a conflict element set; S545. Based on the conflicting element set, a complete knowledge search process is performed in the initial knowledge graph of the enterprise, where the search conditions include the entity category, attribute category, and association relationship category of the nodes in the conflicting node set; S546. Based on the search conditions, matching complementary knowledge fragments are retrieved in the enterprise initial knowledge graph, wherein the complementary knowledge fragments include entity nodes, attribute nodes, and relationship edge information, and the retrieved complementary knowledge fragments are screened according to the semantic matching scores with the conflicting node set; S547: Output the filtered completed knowledge fragments as input data for local reasoning expansion processing.

7. The enterprise global data analysis method based on knowledge graph and large language model according to claim 1 is characterized in that: The S6 comprises the following steps: S61, extracting the inference entity node set and the inference relationship edge set in the corrected inference chain, and recording the entity identification information of each inference entity node in the inference entity node set and the relationship type information of each inference relationship edge in the inference relationship edge set; S62: For the set of inference entity nodes in the corrected inference chain, perform entity node matching retrieval in the enterprise internal data source and the enterprise external data source, respectively, and establish a preliminary entity matching set based on the entity identification information, entity attribute description information, and entity context semantic information of the inference entity node; S63. Calculate the semantic alignment similarity for each pair of candidate matching entity nodes in the preliminary entity matching set. , the semantic alignment similarity Based on entity attribute similarity Consistency score with context Calculation, using the following formula: ; in, is the weighted coefficient of attribute similarity and contextual relationship consistency score, and satisfies ; S64. Semantic alignment similarity Compare with the preset semantic alignment threshold to filter out the semantic alignment similarity Matching entity node pairs greater than a preset semantic alignment threshold and establishing intelligent association relationships between inferred entity nodes; S65. For the selected matching entity node pairs, extract the corresponding reasoning relationship type in the reasoning relationship edge set, and establish an association relationship mapping of the reasoning relationship edge between the internal enterprise data source and the external enterprise data source based on the reasoning relationship type; S66. Based on the set of inference entity nodes and the set of inference relationship edges that have undergone semantic alignment and intelligent association processing, a unified dynamic optimization data analysis view for the enterprise is constructed, and the unified dynamic optimization data analysis view for the enterprise serves as the basic data input for generating the enterprise's global data analysis results.

Citation Information

Patent Citations

  • Data asset identification and risk early warning system and method based on financial knowledge graph and large language model

    CN118247057A