Multi-source data integration management system based on knowledge graph and large model

By establishing a multi-source data integration and management system based on knowledge graphs and large models, the problem of merging structured and unstructured data has been solved, achieving efficient data integration and intelligent business decision support, and improving the accuracy and efficiency of data processing and knowledge extraction.

CN121901432APending Publication Date: 2026-04-21HAIZHI INFORMATION TECH (NANJING) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
HAIZHI INFORMATION TECH (NANJING) CO LTD
Filing Date
2025-12-31
Publication Date
2026-04-21

AI Technical Summary

Technical Problem

Existing technologies cannot efficiently integrate structured and unstructured data. Data formats are inconsistent, relationships are unclear, and the coordination between knowledge storage and retrieval is insufficient, making it difficult to support intelligent business applications. Knowledge extraction from unstructured data is inefficient and lacks accuracy, and structured data and knowledge graphs are difficult to map efficiently.

Method used

A multi-source data integration and management system based on knowledge graphs and large models is adopted, including an infrastructure layer, a data resource layer, a data access layer, a data processing layer, a knowledge graph layer, and an application management layer. Through distributed deployment, multi-type connection mechanism, D2R mapping technology, and dual-database collaborative retrieval mechanism, it realizes intelligent processing and efficient integration of multi-source data throughout the entire process.

Benefits of technology

It enables intelligent processing of multi-source data throughout the entire process, improves the accuracy of knowledge extraction from unstructured data and the mapping efficiency between structured data and knowledge graphs, enhances the accuracy and efficiency of business decision-making, and supports multi-dimensional knowledge fusion and rapid response.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121901432A_ABST
    Figure CN121901432A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-source data integration management system based on a knowledge graph and a large model, and the system comprises an infrastructure layer which is used for providing hardware support and an operation environment; the data resource layer is used for integrating the structured data and the unstructured data and constructing a knowledge base; the data access layer is used for designing a connection mechanism and completing automatic acquisition and circulation of multi-source data; the data processing layer is used for executing data cleaning, mode mapping, knowledge extraction, knowledge normalization fusion and vector database generation; the knowledge graph layer is used for constructing a knowledge graph including mode definition, graph reasoning, entity management and graph retrieval; and the application management layer is used for designing a double-database collaborative retrieval mechanism based on the vector database and the knowledge graph and realizing business checking and intelligent response. According to the invention, multi-source heterogeneous data can be integrated, knowledge values can be deeply mined, and business intelligent decisions can be endowed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of data integration and management technology, specifically to a multi-source data integration and management system based on knowledge graphs and large models. Background Technology

[0002] In the long-term development of core business scenarios such as industrial project management, quality control, and troubleshooting, the effective utilization of data and information has gradually become a key factor driving industry progress. With the continuous expansion and increasing complexity of business operations, the scale and diversity of data have grown dramatically, triggering a series of pressing issues. Therefore, how to properly handle and utilize this multi-source, heterogeneous data has become a focus of attention within the industry.

[0003] In the past, various methods have been adopted in the industry to address these issues. Regarding business data, structured data is typically stored in a dispersed manner across QMS / MRO systems, Excel spreadsheets, and other media, while unstructured data exists in the form of source quality remediation reports, zero-point fault analysis reports, technical standard documents, and fault tree diagrams. In terms of business processing models, during new project planning, staff rely on manually retrieving historical cases, verifying parameter rationality, and identifying potential risks; after a fault occurs, similar fault cases are located manually, the root cause of the fault is deduced, and corresponding countermeasures are determined. Regarding data processing technology, traditional rule-driven text extraction methods are used to process text data, lacking effective correlation and fusion mechanisms for multimodal data (text + images), resulting in a relatively simplistic knowledge storage and retrieval model.

[0004] Clearly, existing technologies have significant shortcomings. They cannot efficiently integrate structured and unstructured data (text, images), resulting in inconsistent data formats and ambiguous relationships, leading to fragmented knowledge and hindering support for intelligent business applications. Furthermore, unstructured data knowledge extraction is inefficient and lacks precision; structured data and knowledge graphs are difficult to map efficiently; and there are issues with insufficient coordination between knowledge storage and retrieval. Against this backdrop, there is an urgent need to build an integrated system capable of integrating multi-source heterogeneous data, deeply mining knowledge value, and empowering intelligent business decision-making, thereby addressing the current pain points and challenges in data utilization and business processing. Summary of the Invention

[0005] In order to provide an integrated system that can integrate multi-source heterogeneous data, deeply mine knowledge value, and empower intelligent business decision-making, this application provides a multi-source data integration and management system based on knowledge graphs and large models.

[0006] In the first aspect, this application provides a multi-source data integration and management system based on knowledge graphs and large models, including: an infrastructure layer, a data resource layer, a data access layer, a data processing layer, a knowledge graph layer, and an application management layer; The infrastructure layer provides hardware support and operating environment; the data resource layer integrates structured and unstructured data to build a knowledge base; and the data access layer designs access mechanisms to automate the collection and flow of multi-source data. The data processing layer is used to perform data cleaning, pattern mapping, knowledge extraction, knowledge normalization fusion, and vector database generation on the multi-source data. Knowledge extraction includes: for unstructured text data, constructing prompt words based on a preset knowledge graph pattern definition and corresponding knowledge base text blocks, calling a large model to perform triple extraction, and iteratively optimizing the prompt word template based on the verification results; for unstructured image data, acquiring document image and text entity information in parallel, constructing prompt words containing image context and entity lists, calling a large model to parse image semantics, establishing the association between images and text entities through a cross-modal alignment mechanism, and generating structured indexing results; knowledge normalization fusion includes: fusing the extracted triples and structured indexing results; and mapping structured data to the knowledge graph through D2R mapping technology; vector database generation includes: generating multimodal vectors based on the knowledge base and completing index construction. The knowledge graph layer is used to construct a knowledge graph that includes pattern definition, graph reasoning, entity management, and graph retrieval; the application management layer is used to design a dual-database collaborative retrieval mechanism based on vector database and knowledge graph to achieve business verification and intelligent response.

[0007] By adopting the above scheme, intelligent processing of multi-source data from access to application is realized. Structured and unstructured data are efficiently integrated and standardized. Hint engineering, large models and D2R mapping technology are used to improve the accuracy of knowledge extraction. Triple data extraction of unstructured data, structured indexing and flexible mapping of structured data to knowledge graphs are realized. Combined with the synergistic linkage of semantic information and structured knowledge, the retrieval accuracy and efficiency are improved.

[0008] Preferably, the infrastructure layer comprises a cluster of four distributed servers; wherein, the first server carries core components and runs containerized services; the second server deploys a message queue and file storage system; the third server stores relational data and graph databases; and the fourth server stores cached data and index data; each server achieves separation of static and dynamic resources and traffic distribution through a load balancing layer.

[0009] By adopting the above solution and utilizing distributed deployment, the system components are decoupled and flexibly expanded, ensuring the stability of data transmission and file storage, enabling the classified storage and efficient access of multiple types of data, and improving the system's concurrent processing capabilities and availability.

[0010] Preferably, the data access layer design includes an interface connection mechanism, structured connection, unstructured connection, and connection task scheduling, which coordinates the various connection mechanisms to complete the automated collection and transfer of different types of data.

[0011] By adopting the above scheme and designing a multi-type connection mechanism, the automated collection and transfer of different data types can be realized, improving the flexibility and efficiency of data access and further promoting the integration of multi-source heterogeneous data.

[0012] Preferably, the data processing layer is further configured to classify the current unstructured text data into types based on data scale, sample richness, and real-time performance, including high real-time and small sample scenario types, low real-time and large sample scenario types, high real-time and large sample scenario types, and low real-time and small sample scenario types; and to adapt knowledge extraction path decisions according to different data types, including a first knowledge extraction path decision matching the high real-time and small sample scenario type, a second knowledge extraction path decision matching the low real-time and large sample scenario type, and a third knowledge path decision matching the high real-time and large sample scenario type and the low real-time and small sample scenario type. The first knowledge extraction path decision includes prompting engineering and large model iteration path; knowledge extraction using the first knowledge extraction path decision includes: for unstructured text data, constructing prompt words based on the knowledge base text blocks corresponding to the unstructured text data defined and annotated according to a preset knowledge graph pattern, calling the large model to perform triple extraction, and iteratively optimizing the prompt word template based on the verification results; the second knowledge extraction path decision includes pre-annotation and large model fine-tuning path; knowledge extraction using the second knowledge extraction path decision includes: annotating some unstructured text data with triples, directly inputting the knowledge base text blocks corresponding to the unstructured text data annotated with triples into the large model for model parameter fine-tuning, and using the fine-tuned large model data to perform triple extraction; the third knowledge path decision is a hybrid application of the first and second knowledge extraction path decisions.

[0013] By adopting the above scheme, different knowledge extraction path decisions are adapted according to different unstructured text data types. In high real-time small sample scenarios, the knowledge extraction efficiency is improved through prompting engineering and large model iteration. In low real-time large sample scenarios, the knowledge extraction effect is optimized through pre-labeling and large model fine-tuning. In other scenarios, a hybrid strategy is used to improve the knowledge extraction efficiency, thereby improving the pertinence and accuracy of knowledge extraction from unstructured text data.

[0014] Preferably, the infrastructure layer is further used to provide streaming computing resources; the data resource layer is further used to divide the real-time data pool and the batch data pool to achieve hot and cold data separation; and the data access layer is further used to add streaming access interfaces for multimodal data. The data processing layer is further divided into a streaming processing sublayer and a batch processing sublayer; the streaming processing sublayer is used to perform data cleaning, pattern mapping, knowledge extraction and knowledge normalization fusion, and semantic vector generation in real time; the batch processing sublayer is used to perform data cleaning, pattern mapping, knowledge extraction and knowledge normalization fusion, and semantic vector generation in batches. The knowledge graph layer is further used to divide the graph into layers, including an ontology layer storing domain ontology definitions and cross-modal ontology mapping rules, an instance layer storing triples that map core entities and relations and structured data, and a multimodal attribute layer storing attribute triples extracted from unstructured data and associations with multimodal feature vectors. The ontology layer for constructing the knowledge graph adopts a strategy of static generation and batch updating; the instance layer adopts a strategy of real-time generation and incremental updating; and the multimodal attribute layer adopts a strategy of real-time generation and streaming updating.

[0015] By adopting the above scheme, the provision of streaming computing resources and the division of real-time data pools and batch data pools have achieved the separation of hot and cold data, improving the efficiency and targeting of data processing; the addition of multimodal data streaming access interfaces has enabled better processing of multimodal data; the data processing layer is designed to be split into streaming processing sub-layers and batch processing sub-layers to perform real-time and batch data processing operations respectively, making data processing more flexible and efficient; the knowledge graph layer is designed to divide the graph into levels and adopt different generation and update strategies, making the construction and maintenance of the knowledge graph more reasonable.

[0016] Preferably, the data processing layer integrates a confidence-driven automatic verification module. This module is used to construct a rule base based on the graph pattern definition, combine a lightweight decision tree to detect logical contradictions, and obtain the confidence of the extracted triples. It also uses multimodal contrastive learning to verify the consistency between the image description and the associated triples, and obtain the confidence of the structured indexing results. When the confidence of the extracted triples or the confidence of the structured indexing results is lower than the corresponding preset confidence, manual review is triggered to verify the completeness and accuracy of the triple extraction results and structured indexing results output by the large model. The review results obtained from the verification are then fed back to the prompt word template optimization process.

[0017] By adopting the above scheme, the confidence level of triple extraction and structured indexing results is automatically obtained. When the confidence level is lower than the preset value, manual review is triggered to verify the completeness and accuracy of the large model output results. The review results can also be fed back to optimize the prompt word template, thereby improving the accuracy and reliability of knowledge extraction.

[0018] Preferably, the data processing layer is further used to classify the types of structured data based on table features, data characteristics, and business characteristics; in the process of mapping structured data to knowledge graphs through D2R mapping technology, a type-customized differentiated mapping strategy is adopted according to the structured data type to complete the mapping from data tables in structured data to knowledge graph concepts, the mapping from data fields in structured data to knowledge graph attributes, and the mapping from foreign keys between tables or business logic in structured data to relationships in knowledge graphs.

[0019] By adopting the above scheme, customized mapping is performed according to different types of structured data, making the mapping between structured data and knowledge graphs more accurate and flexible, and improving the matching degree and mapping efficiency of structured data and knowledge graphs.

[0020] Preferably, the data processing layer is further configured to support parallel mapping of different data types, full mapping of batch processing, and incremental mapping of different data types based on triggering mechanisms and processing strategies during the mapping process between structured data and knowledge graphs via D2R mapping technology, and to complete the mapping accuracy verification according to the automated verification rules of different data types.

[0021] By adopting the above scheme, in the process of mapping structured data to knowledge graphs, it supports full mapping in parallel and batch by type and incremental mapping triggered and processed by type. Combined with automated rules to verify the accuracy of mapping, it improves mapping efficiency and quality, and realizes real-time data synchronization and dynamic knowledge updates.

[0022] Preferably, the knowledge graph layer includes a normalized dictionary and a thesaurus, used for continuously learning and updating term mapping rules.

[0023] By adopting the above scheme, the knowledge graph layer sets up a normalized dictionary and a thesaurus and continuously learns and updates the term mapping rules, making cross-document entity associations more standardized and unified, and facilitating the tracing of associated information.

[0024] Preferably, the application management layer is further configured to design a dual-database collaborative retrieval mechanism based on a vector database and a knowledge graph. During the business verification and intelligent response process, it receives input application business parameters and determines the type of the current application business, including precise query, fuzzy query, complex reasoning, and hybrid query types. Based on the determined application business type, it matches a suitable collaborative strategy, including: a collaborative strategy adapted to precise queries that first performs query retrieval based on the knowledge graph, followed by query retrieval based on the vector database; a collaborative strategy adapted to fuzzy queries that first performs query retrieval based on the vector database, followed by query retrieval based on the knowledge graph; a collaborative strategy adapted to complex reasoning queries that performs query retrieval based on the knowledge graph; and a collaborative strategy adapted to hybrid queries that performs parallel retrieval based on both the vector database and the knowledge graph.

[0025] By adopting the above solution, corresponding dual-database collaborative retrieval strategies can be matched according to different application business types, improving the pertinence and accuracy of business verification and intelligent response, and further enhancing the efficiency and quality of business decision-making.

[0026] In summary, this application has the following beneficial effects: 1. By designing a six-layer architecture and a distributed deployment model, intelligent processing of multi-source heterogeneous data is achieved throughout the entire process from access to application. Each layer operates collaboratively, enabling automated collection, integration, and standardized processing of structured and unstructured data. A large model is used for knowledge extraction, improving the accuracy of knowledge extraction from unstructured text and image data, uncovering key information in images, and achieving multi-dimensional knowledge fusion. D2D mapping technology enables precise mapping from structured data to knowledge graphs. A dual-database collaborative retrieval mechanism is employed to achieve synergistic linkage between semantic information and structured knowledge, improving the accuracy and efficiency of similar case retrieval and providing fast and reliable support for business decisions. 2. Considering the characteristics of multi-source heterogeneous data, we will improve the efficiency and accuracy of knowledge extraction from unstructured text data by matching knowledge types; improve the mapping adaptability between structured data and knowledge graphs by matching knowledge types; and improve the retrieval accuracy and efficiency of dual-database collaborative retrieval by matching types, providing fast and reliable support for business decisions. 3. Considering the processing timeliness of multi-source heterogeneous data, in order to efficiently process data with different timeliness, realize the separation of hot and cold data and the streaming access of multimodal data, and at the same time, flexibly construct and update the knowledge graph according to the characteristics of different levels, improve the overall efficiency of the system for data processing and knowledge management, and better meet the needs of different business scenarios for data processing and knowledge application. Attached Figure Description

[0027] Figure 1 This is an architecture diagram of the multi-source data integration and management system based on knowledge graphs and large models described in a specific embodiment; Figure 2 This is a flowchart illustrating the process of multi-source data integration in the multi-source data integration management system based on knowledge graphs and large models described in a specific embodiment. Figure 3 This is a flowchart illustrating the overall process framework for multi-source data integration using the knowledge graph and large model-based multi-source data integration management system described in a specific application embodiment. Figure 4 The flowchart for triple extraction in multi-source data integration during the application of the multi-source data integration management system based on knowledge graphs and large models described in the specific embodiment is as follows: Figure 5This is a flowchart illustrating image indexing in multi-source data integration using the multi-source data integration management system based on knowledge graphs and large models described in a specific embodiment. Figure 6 This is a flowchart illustrating the structured data and knowledge graph in the multi-source data integration process of the multi-source data integration management system based on knowledge graphs and large models described in a specific embodiment. Detailed Implementation

[0028] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.

[0029] like Figure 1 As shown in the figure, this application discloses a multi-source data integration and management system based on knowledge graphs and large models, including: an infrastructure layer 100, a data resource layer 200, a data access layer 300, a data processing layer 400, a knowledge graph layer 500, and an application management layer 600. Each layer cooperates with each other to achieve intelligent processing of multi-source data from access to application, and solves problems such as multi-source heterogeneous data integration, knowledge extraction, data mapping, business decision-making, and knowledge storage and retrieval. The functional design and application of each layer are described in detail below.

[0030] like Figure 1 As shown, the infrastructure layer 100 includes servers, storage devices, virtualization software, and a compilation environment. The servers provide computing power, the storage devices provide data storage, the virtualization software optimizes resource utilization, and the compilation environment supports system development. It is used to provide stable hardware support and operating environment, and to provide highly available basic resources for the upper layers.

[0031] Specifically, the server is the core hardware device for system operation, employing high-performance servers with powerful computing capabilities and stability. Storage devices are used to store various data during system operation, such as databases and files, and can be disk arrays, tape libraries, etc. Virtualization software can virtualize the physical server into multiple virtual machines, improving server utilization; examples include VMware and Hyper-V. The compilation environment provides the necessary tools and environment for system development and operation, such as Java development environments and Python environments. The server and storage devices are connected via a high-speed network to ensure rapid data transmission; the virtualization software runs on the server, managing and allocating server resources; and the compilation environment provides support for system development and maintenance.

[0032] To achieve component decoupling and flexible expansion, the infrastructure layer adopts distributed deployment and storage technologies, setting up a cluster of four distributed servers. The first server hosts core components and runs containerized services; specifically, it acts as an application server, integrating over 20 core components such as data access, knowledge representation, and gateway services, and is responsible for business logic processing. The second server deploys a message queue and file storage system; specifically, it acts as a middleware server, ensuring the stability of data transmission and file storage. The third server stores relational data and graph databases. The fourth server stores cached data and index data, enabling categorized storage and efficient access to multiple data types. Each server utilizes a load balancing layer to achieve separation of static and dynamic resources and traffic distribution. For example, load balancing, reverse proxying, and separation of static and dynamic resources are achieved through LVS, Nginx, and Keepalived, improving system concurrency and availability.

[0033] like Figure 1 As shown, the data resource layer 200 is used to integrate structured and unstructured data, build a multi-source heterogeneous original knowledge base, and break down data silos.

[0034] Specifically, structured data includes QMS / MRO system data and Excel files. QMS / MRO system data typically contains information related to industrial project management and quality control, and is stored in a relational database format. Excel files can be tabular records of various business data. Unstructured data includes quality reports, technical standards, and fault tree diagrams. Quality reports and technical standards are usually text files, while fault tree diagrams are image files. The data resource layer integrates these different types of data to build a multi-source heterogeneous original knowledge base. Through the data access layer, data from different sources is collected into the system, then stored and managed, breaking down data silos and enabling different types of data to be interconnected and utilized.

[0035] like Figure 1 As shown, the data access layer 300 is used to uniformly access data and complete the automated collection and flow of multi-source data through the design of the connection mechanism.

[0036] Specifically, the data access layer's connection mechanism includes: interface connection, structured connection, unstructured connection, and connection task scheduling. Interface connection interacts with other systems through standardized interfaces, such as connecting to QMS / MRO systems via RESTful APIs. Structured connection handles structured data, using specific algorithms and tools to extract data from the data source and perform format conversion and cleaning. Unstructured connection specifically processes unstructured data, such as using OCR technology to recognize text information in images. Connection task scheduling manages and schedules data acquisition tasks to ensure timely data collection and flow. Interface connection works in conjunction with structured and unstructured connection, selecting appropriate connection methods based on different data types; task scheduling coordinates and manages the entire data access process. The specific combination logic is: based on the data type and characteristics, select an appropriate connection method for data acquisition, and then ensure the efficiency and orderliness of data acquisition through task scheduling.

[0037] As described Figure 1 As shown, the data processing layer 400 is used to perform data cleaning, pattern mapping, knowledge extraction, knowledge normalization fusion, and vector database generation.

[0038] First, perform data cleaning. Data cleaning is used to remove noise, duplicate data, and erroneous data from the original data, thereby improving the data quality.

[0039] Second, perform model mapping. Schema mapping matches and transforms the schema of the original data with the schema of the knowledge graph, achieving data standardization. For example, if the current knowledge graph schema is an e-commerce knowledge graph schema, the schema of the original data will be matched to the corresponding e-commerce data schema.

[0040] Third, knowledge extraction is performed. This mainly involves extracting triples from unstructured text data and extracting and indexing multimodal information from other data, such as semantic parsing and structured indexing of images.

[0041] Specifically, for unstructured text data, in order to extract entities, relations, and objects from the text data to form knowledge triples, a closed-loop approach of schema definition, prompting engineering, and large-scale model iteration is adopted. First, the knowledge graph schema definition and the corresponding knowledge base text blocks of the unstructured text data are queried in parallel, and prompt words are constructed based on both. Then, for the constructed prompt words and the unstructured text data to be processed, a large model (such as Qwen3) is called to perform triple (subject, relation, object) extraction. After the extraction results are verified for completeness and accuracy, if they do not meet the standards, the prompt word template is optimized and re-extracted until structured knowledge that meets the standards is output.

[0042] The entire knowledge extraction logic can be understood as follows: Based on pattern definitions and the characteristics of unstructured data, examples with few sample annotations are generated. These examples are used to optimize the prompt word template of the large model and stored as baseline knowledge in the graph database. The few-sample learning capability of the large model is then utilized to extract triples from other unstructured data. The following example illustrates this, analyzing the characteristics of the knowledge base text block (medical record text) corresponding to unstructured text data: long text, many technical terms, and a fixed structure. Based on the pattern definition of the knowledge graph, prompt words are designed, such as: "Please extract information from the following medical record text and output it in JSON format." The entity types to be extracted include: disease, symptoms, medicine, and examination items; relationships include: the relationship between disease and symptoms, the relationship between disease and medicine, and the relationship between disease and examination items; text: {input text}. A small number of already annotated medical records are used as examples to form prompt words for few-sample learning. The large model is then called for extraction, the effect is evaluated, and the prompt words are adjusted until the effect is stable. New medical record text is input into the large model through the prompt words to obtain the extracted triples.

[0043] For unstructured image data, after extracting entities from the text side (such as subjects or objects in triples for fault, fault cause, and fault phenomenon), document image and text-side entity information are acquired in parallel. Hint words containing image context and entity lists are constructed, and a large model is invoked to parse image semantics (e.g., using a multimodal large model (Qwen3-VL) for deep semantic understanding of embedded images in the document). A cross-modal alignment mechanism is used to establish the association between image and text entities, generating structured indexing results and completing entity linking. Similarly, after the indexing results are verified for completeness and accuracy, if they fail, the hint word template can be optimized, the hint strategy adjusted or refined, and the hint words reconstructed. The model is then invoked again for extraction until the results meet the requirements.

[0044] In addition, in order to achieve completeness and accuracy verification of extraction and indexing results from unstructured data, manual review can be directly selected. However, in order to reduce manual costs, a confidence-driven automatic verification module is integrated into the data processing layer. Based on the automatic verification results, it is determined whether to trigger further manual review, thereby improving the efficiency of result verification.

[0045] Specifically, the confidence-driven automatic verification module is used to construct a rule base based on the graph pattern definition, and combine it with a lightweight decision tree to detect logical contradictions and obtain the confidence of the extracted triples. The lightweight decision tree is trained by inputting the extracted triples, valid and invalid labels, and confidence (multi-dimensional evaluation criteria can be pre-set). Correspondingly, the confidence-driven automatic verification module is also used to use multimodal contrastive learning to verify the consistency between image descriptions and associated triples, and obtain the confidence of the structured indexing results. Verifying the consistency between image descriptions and associated triples includes using a CLIP model to generate joint embedding vectors of images and text, calculating similarity and comparing whether it is greater than a similarity threshold. If it is greater than the similarity threshold, the verification passes, and the confidence is output simultaneously. The confidence-driven automatic verification module is also used to trigger manual review when the confidence of the triplet extraction or the confidence of the structured indexing result is lower than the corresponding preset confidence. This is to verify the completeness and accuracy of the triplet extraction result and the structured indexing result output by the large model, and to feed the verification result back to the prompt word template optimization process.

[0046] Fourth, knowledge normalization and fusion. Based on the triple extraction, multimodal extraction and entity linking mentioned above, the original data has been transformed into standardized knowledge units. The standardized knowledge is then further fused to assist in the generation of knowledge graphs.

[0047] Specifically, after knowledge extraction from unstructured data, the resulting structured knowledge is fused (e.g., fused for the same entity) and stored in a knowledge graph database. This includes: supplementing the knowledge graph based on triples to assist in subsequent knowledge graph construction and updates; and supplementing the knowledge graph based on other multimodal information associated with the entity, such as image context information associated with the entity, to assist in subsequent knowledge graph construction and updates.

[0048] Specifically, for structured data, D2R (Database-to-RDF) mapping technology is used to map structured data to knowledge graphs. This involves converting existing structured data (such as relational databases) into RDF format conforming to knowledge graph standards, achieving cross-structure data adaptation and semantic alignment, and further realizing knowledge normalization and fusion. This includes: parallel acquisition of database table field information (table name, field name, data type) and knowledge graph ontology schema (concept class, attribute, relationship); sequential completion of concept mapping (data table mapping to knowledge graph concepts), attribute mapping (data field mapping to knowledge graph attributes), and relationship mapping (inter-table foreign keys / business logic mapping to knowledge graph object attributes). After defining all mapping rules, the system supports two data import methods: full and incremental dual-mode import. In full import mode, all eligible structured data is converted into RDF triples and written to the graph database at once; in incremental import mode, only newly added or modified structured data is processed, achieving real-time synchronization of structured data and dynamic knowledge updates, improving data processing efficiency.

[0049] Fifth, vector database generation. Specifically, vector database generation includes: generating multimodal vectors based on a knowledge base, including: text vectorization (using pre-trained language models (such as BERT, Sentence-BERT, RoBERTa, etc.) to generate text vectors and store them in the vector database), image vectorization (using pre-trained convolutional neural networks (such as ResNet, VGG, EfficientNet, etc.) or visual transformers (ViT) to extract image feature vectors), and structured data vectorization (converting structured data into text descriptions and then using text vectorization methods); and completing index construction, which can use various indexing methods to complete vector indexing, and then storing index files and vector data.

[0050] like Figure 1 As shown, the knowledge graph layer 500 is used to construct a knowledge graph that includes pattern definition, graph reasoning, entity management, and graph retrieval.

[0051] Specifically, the knowledge graph layer includes pattern / terminology definition, graph reasoning, entity management, graph retrieval, a normalized dictionary, and a thesaurus. The pattern / terminology definition defines and describes the concepts, attributes, and relationships in the knowledge graph, providing the foundation for its construction. Graph reasoning, based on rules and facts in the knowledge graph, discovers new knowledge. Entity management performs operations such as creating, updating, and deleting entities in the knowledge graph. Graph retrieval provides query functionality for the knowledge graph, facilitating users' access to the knowledge they need. The normalized dictionary and the thesaurus are used to normalize terms, eliminating discrepancies. Graph reasoning and entity management maintain and expand the knowledge graph, graph retrieval facilitates user access to the knowledge graph, and the normalized dictionary and thesaurus improve the accuracy and consistency of the knowledge graph. Together, they constitute a complete knowledge graph system, supporting structured organization of knowledge, intelligent reasoning, and efficient querying.

[0052] Specifically, the knowledge graph layer facilitates predefined patterns and terminology definitions, assisting the data processing layer in completing data processing. Based on the structured knowledge output from the historical data processing layer and stored in the graph database, a historical knowledge graph is generated in the historical graph database, and entity management is performed. Subsequently, graph retrieval and graph reasoning can be performed based on the structured knowledge output from the real-time data processing layer.

[0053] like Figure 1 As shown, the application management layer 600 is used to design a dual-database collaborative retrieval mechanism based on vector database and knowledge graph to achieve business verification and intelligent response.

[0054] Specifically, the application management layer includes three core functional modules: business verification, intelligent response, and backend management. In this embodiment, business verification takes "new project planning auxiliary verification" as an example. When planning a new project, a dual-database collaborative retrieval mechanism based on a vector database and a knowledge graph can be used to search for similar cases in both databases. This allows for parameter verification and risk identification of the planning proposal, improving the accuracy and reliability of project planning. For example, after the user inputs the planning proposal parameters, the parameters are first vectorized to generate a query vector. Semantically similar text fragments are then retrieved from the vector database using an approximate nearest neighbor algorithm. Simultaneously, the entity retrieval function of the knowledge graph is invoked to find relevant entities based on triplet association information. Finally, the search scope is expanded through graph exploration, integrating the results from both databases to output accurate similar cases to assist in decision-making.

[0055] Taking "Intelligent Rapid Response to Fault Analysis" as an example, when a fault occurs, a dual-database collaborative retrieval mechanism based on a vector database and a knowledge graph quickly locates similar fault cases, infers the root cause of the fault, and recommends the optimal countermeasures, thereby improving the efficiency and accuracy of fault analysis. For example, it parses the user's fault question through an intent recognition interface and extracts an entity list; after determining that the parameters are met, it recalls graph data and document data, infers the fault chain based on the knowledge graph, and recommends similar cases and countermeasures.

[0056] Backend management involves managing and maintaining system users, permissions, and data to ensure the system's normal operation. These three functional modules cover the needs of the entire business process, collaborate with each other, and support business decision-making.

[0057] In a specific embodiment, selecting appropriate knowledge extraction paths based on different unstructured text data types can more effectively utilize data resources and improve the efficiency and accuracy of knowledge extraction. For high real-time small sample data, accurate knowledge can be quickly obtained through prompting engineering and iterative optimization; for low real-time large sample data, information in the data can be fully mined through pre-labeling and model fine-tuning, improving the system's performance and adaptability; the system includes: To address the low efficiency and insufficient accuracy of knowledge extraction from unstructured data, two differentiated knowledge extraction paths are designed based on the characteristics of unstructured data: large-model iterative extraction and supervised fine-tuning extraction, thereby improving knowledge extraction efficiency. Specifically, the data processing layer 400 is also used to classify the current unstructured text data into types based on data scale, sample richness, and real-time performance, including: high real-time and small sample (less than 1000 records) scenario type, low real-time and large sample (not less than 1000 records) scenario type, high real-time and large sample scenario type, and low real-time and small sample scenario type; and to adapt knowledge extraction path decisions according to different data types, including: a first knowledge extraction path decision matching the high real-time and small sample scenario type, a second knowledge extraction path decision matching the low real-time and large sample scenario type, and a third knowledge path decision matching the high real-time and large sample scenario type and the low real-time and small sample scenario type.

[0058] The first knowledge extraction path decision includes prompting engineering and a large model iteration path. This design is based on the fact that in small sample cases, insufficient labeled data makes supervised fine-tuning prone to overfitting. Prompt word iteration, however, leverages the zero-shot or small-sample learning capabilities of the large model, allowing for rapid adaptation to new domains through appropriate prompt word design and iterative optimization. Furthermore, it offers high real-time performance, requiring rapid deployment; prompt word iteration does not require lengthy training and can be adjusted in real-time. Knowledge extraction using the first knowledge extraction path decision includes: for unstructured text data, constructing prompt words based on a predefined knowledge graph pattern definition and the corresponding knowledge base text blocks of the labeled unstructured text data; calling the large model to perform triple extraction; and iteratively optimizing the prompt word template based on the validation results. Note that the labeled unstructured text data here is not manually labeled data, but rather prompt words generated based on pattern definitions with a small number of labeled examples.

[0059] The second knowledge extraction path decision includes pre-labeling and large model fine-tuning. This design is based on the following principles: large samples provide sufficient training data to train a high-quality domain-specific model; supervised fine-tuning fully utilizes labeled data, resulting in higher extraction accuracy. Low real-time requirements allow for model training and optimization, providing stable and efficient services during the inference phase. Knowledge extraction using the second knowledge extraction path decision includes: triple labeling of some unstructured text data; directly inputting the knowledge base text blocks corresponding to the labeled unstructured text data into the large model for model parameter fine-tuning; and using the fine-tuned large model data to perform triple extraction. Compared to the first knowledge extraction route, because the data is large-sample, the domain is fixed, and the ontology can be solidified and expanded in detail, the direct choice is to label unstructured text data with triples, and then use the labeled data for large model fine-tuning. For example, similarly, analyzing medical record text and defining patterns; labeling medical record text, annotating each text with entities and relationships to form training data; selecting a large model (such as BERT, GPT, etc.) for fine-tuning, and modeling the knowledge extraction task as a sequence labeling or text generation task. Train the model, evaluate its performance, and fine-tune it. Deploy the finely tuned model to extract data from new medical record texts in real time.

[0060] The third knowledge path decision is a hybrid application of the first and second knowledge extraction path decisions. In high real-time and large-sample scenarios, a large sample size can train a high-quality model, but the model's inference speed needs to be optimized to meet high real-time requirements, such as through model distillation, quantization, and optimized deployment. Alternatively, a hybrid strategy can be adopted, which involves training a base model with supervised fine-tuning and then using prompt words for iterative adjustments and adaptation to new changes. Therefore, the hybrid strategy includes: using the second knowledge extraction path decision to directly input the knowledge base text blocks corresponding to the unstructured text data labeled with triples into the large model for model parameter fine-tuning, training and generating a base model; then using the first knowledge extraction path decision to construct prompt words based on a preset knowledge graph pattern definition and the knowledge base text blocks corresponding to the labeled unstructured text data; calling the base model to perform triple extraction; and iteratively optimizing the prompt word template based on the verification results.

[0061] In scenarios with low real-time requirements and small sample sizes, supervised fine-tuning is prone to overfitting in small sample situations, but the low real-time requirements allow for data augmentation and more refined model adjustments. A lightweight supervised fine-tuning approach can be adopted. Therefore, a hybrid strategy is used, including utilizing a second knowledge extraction path decision to directly input the knowledge base text blocks corresponding to the unstructured text data labeled with triples into a large model for lightweight fine-tuning of model parameters (a preset number of fine-tuning attempts). After training to generate a base model, the first knowledge extraction path decision is then used to construct prompt words based on a preset knowledge graph pattern definition and the knowledge base text blocks corresponding to the labeled unstructured text data. The base model is then called to perform triple extraction, and the prompt word template is iteratively optimized based on the validation results.

[0062] One specific embodiment, by introducing streaming computing resources, separating hot and cold data, providing multimodal data streaming access, and implementing layered processing and knowledge graph updates, better handles real-time and multimodal data, improving the system's real-time performance and data processing capabilities. The system also includes: The infrastructure layer is also used to provide streaming computing resources, such as GPU clusters and streaming processing nodes; The data resource layer is also used to divide the real-time data pool and the batch data pool to achieve the separation of hot and cold data; The data access layer is also used to add streaming access interfaces for multimodal data, supporting real-time image uploading and video frame extraction streaming input; for example, using Kafka Connect for image data access. The data processing layer is further divided into a streaming processing sublayer and a batch processing sublayer. The streaming processing sublayer is used to perform data cleaning, pattern mapping, knowledge extraction and knowledge normalization fusion, and semantic vector generation in real time, such as using FLINK for streaming computation. The batch processing sublayer is used to perform data cleaning, pattern mapping, knowledge extraction and knowledge normalization fusion, and semantic vector generation in batches, such as using SPARK for batch computation. When streaming conflicts exist, a timestamp-first and confidence-first strategy is adopted to process conflicting knowledge in real time.

[0063] Considering the characteristics of multimodal knowledge (stable ontology layer and high-frequency change in multimodal attribute layer), a layered architecture is adopted to design the knowledge graph. Specifically, the knowledge graph layer is also used to divide the graph into layers, including: ontology layer, instance layer and multimodal attribute layer.

[0064] The ontology layer stores domain ontology definitions (entity types, relation types, attribute types) and cross-modal ontology mapping rules (such as association rules between text entities and image contexts). This layer has high stability, and the ontology layer used to build knowledge graphs generally adopts a strategy of static generation and batch updates. For example, batch analysis of uncovered entity / relation types (such as material attributes that frequently appear in e-commerce images) is performed, and domain experts expand the ontology in batches; batch analysis of mapping conflicts between text and image knowledge (such as product A-material-pure cotton in text and product A-material-polyester extracted from the image) is performed, and mapping rules are optimized (such as adding material feature vector matching rules).

[0065] The instance layer stores core entities and relationships (such as product entities, user entities, and "purchase" relationships) as triples mapped from structured data. The stability of this layer is moderate. The instance layer is constructed using a real-time generation and incremental update strategy. For example, Flink streaming computation is used to disambiguate newly added entities in the instance layer in real time; similar entities added on the same day are merged in batches; Flink CDC is used to monitor changes in the relational database (such as user table updates), and real-time mapping is used to map them into triples in the instance layer and write them into the graph.

[0066] The multimodal attribute layer stores attribute triples (such as product color, size, and appearance) extracted from unstructured data (images / audio / video) and associated with multimodal feature vectors. This layer has low content stability. The construction of the multimodal attribute layer adopts a strategy of real-time generation and streaming update, such as streaming deduplication, lightweight streaming inference, and automatic cleaning of expired attributes.

[0067] In a specific embodiment, based on structured data types, a type-customized differentiated mapping strategy is adopted when using D2R mapping technology to complete the mapping of structured data to concepts, attributes, and relationships in a knowledge graph. This improves the mapping adaptability between structured data and knowledge graphs, achieves flexible adaptation between relational database data and knowledge graph schemas, and meets the needs of real-time data synchronization and dynamic knowledge updates. The system also includes: The data processing layer is also used to classify the types of structured data based on table features, data characteristics, and business characteristics; the specific classifications include: traditional relational databases (such as MySQL / Oracle), time-series databases (such as InfluxDB / TDengine), data warehouses (such as Hive / BigQuery), wide-table databases (such as ClickHouse / HBase), and semi-structured data (such as JSON / CSV / Parquet), etc.

[0068] In the process of mapping structured data to knowledge graphs using D2R mapping technology, a type-customized differentiated mapping strategy is adopted according to the structured data type to complete the mapping from data tables in structured data to knowledge graph concepts, the mapping from data fields in structured data to knowledge graph attributes, and the mapping from foreign keys between tables or business logic in structured data to relationships in the knowledge graph.

[0069] Specifically, structured data tables / datasets are mapped to entity concepts in a knowledge graph. A dynamic mapping principle prioritizing data structure adaptation and business meaning is adopted. The core business meaning of the table / dataset determines the concept type, and the data structure (such as table partitioning and partitioning) determines the concept aggregation and splitting method. Taking a traditional relational database as an example, for table partitioning scenarios, a unified concept mapping is performed across multiple tables. Tables with pre-defined close relationships, such as the `user_detail` table and the `user` table, are merged into a unified concept. Taking a time-series database as an example, concepts are split according to static entities and time-series events. Measurement point tables (such as the `device` table, storing device IDs and models) are mapped to the static entity concept (Device), and time-series data tables (such as the `temperature` table, storing timestamps and temperature values) are mapped to time-series event concepts. Taking a data warehouse as an example, dimension tables are mapped to static entity concepts, preserving the dimension hierarchy, and fact tables are mapped to event concepts, associated with the corresponding dimension concepts. Taking a wide-table database as an example, the concept is split according to the business domain of the field; taking semi-structured data as an example, the root node of nested JSON is mapped to the main concept, and the nested nodes are mapped to the sub-concepts, associating the main / sub-concepts; for CSV data, the concept is mapped by grouping the business domain of the table header.

[0070] Specifically, data fields are mapped to attributes of knowledge graph entities, following the principles of field characteristic adaptation and business value selection. The attribute storage method is determined by the field type, and the business value of the field determines whether to retain the mapping. Taking traditional relational databases as an example, a one-to-one mapping from fields to attributes is used, with field type standardization conversion as the foundation. For example, enumeration fields are mapped to enumeration attributes, and the meaning of enumeration values ​​is annotated. Foreign key fields are not directly mapped to attributes but are used for subsequent relation mapping. Taking time-series databases as an example, through classification mapping, static fields are mapped to static attributes of static entity concepts, and time-series fields are mapped to indicator attributes of time-series event concepts. Aggregation mapping is performed on high-frequency time-series fields. Taking data warehouses as an example, dimension table fields are mapped to hierarchical / descriptive attributes of static entity concepts, and fact table fields are mapped to metric attributes of event concepts, with units annotated. Taking wide-table databases as an example, redundant fields are filtered, fields are grouped according to business domains, and mapped to attributes of corresponding concepts. Taking semi-structured data as an example, nested fields in JSON correspond to the attributes of sub-concepts, and dynamic key fields are mapped to dynamic attributes; null fields in CSV data are mapped to attributes and set with default values, and the field types are automatically identified when they are not fixed.

[0071] Specifically, table relationships (foreign keys / business logic) are mapped to entity relationships in a knowledge graph, following the principles of foreign key priority and business logic completion. Explicit foreign keys are prioritized for establishing relationships; when no foreign keys exist, implicit relationships are mined through table / field name semantics and business logic to ensure that relationship types match business scenarios. Taking a traditional relational database as an example, when there are foreign key relationships, relationships are directly established based on the foreign key, and intermediate relationships are established between many-to-many related tables; when there are no foreign key relationships, relationships are mined through field semantic matching. Taking a time-series database as an example, a `has_timeseries_data` relationship is established between measurement point tables and time-series tables; a `same_device` relationship is established between time-series tables for relationships related to the same device; and a `time_sequence` relationship is established for time-series data relationships. Taking a data warehouse as an example, single participation or inclusion relationships exist between dimension tables and fact tables; hierarchical dimensions establish attribution relationships; and implicit relationships between facts are mined through business logic. Taking a wide-table database as an example, relationships between entities are mined through field semantics for relationships within the wide table; and cross-table relationships are established between the wide table and other tables through core fields. Taking semi-structured data as an example, in JSON, a has_sub relationship is established between the main concept and the sub-concept, and in CSV data, entities from different CSV / tables are associated through the semantics of the table header.

[0072] In addition, the data processing layer is also used to support parallel mapping of different data types and full mapping of batch processing, as well as incremental mapping of different data types based on triggering mechanisms and processing strategies, in the process of mapping structured data and knowledge graphs through D2R mapping technology, and to complete the mapping accuracy verification according to the automatic verification rules of different data types.

[0073] Specifically, taking relational databases as an example, the full mapping strategy employs multi-threaded parallel processing and foreign key association grouping algorithms; the incremental mapping strategy uses row-level triggering of INSERT / UPDATE / DELETE events in Flink CDC to achieve row-level incremental mapping and incremental linkage of related tables; the automated verification rules are: integrity of foreign key associations, correctness of field type mapping, and entity uniqueness. Taking time-series databases as an example, the full mapping strategy uses sharding mapping by device ID and time granularity, and batch import; the incremental mapping strategy uses time window triggering and data volume threshold triggering to achieve incremental mapping within the window and aggregation of time-series attributes; the automated verification rules are: integrity of timestamp attributes, rationality of time-series metrics, and correctness of device associations. Taking data warehouses as an example, the full mapping strategy prioritizes dimension table mapping and batch maps fact tables by partition; the incremental mapping strategy uses partition completion triggering and batch data arrival triggering to achieve partition incremental mapping; the automated verification rules are: consistency of dimension levels, integrity of the association between fact tables and dimension tables, and standardization of metric units. Taking wide tables as an example, the full mapping strategy uses block mapping based on field business domains for parallel processing; the incremental mapping strategy uses batch data arrival triggers and field change triggers to achieve batch incremental mapping; the automated verification rules are: the rationality of field classification, the completeness of filtering redundant fields, and the non-nullability of core attributes. Taking semi-structured data as an example, the full mapping strategy uses data volume-based shard mapping and asynchronous parsing of nested structures; the incremental mapping strategy uses file upload triggers or directory file addition triggers to achieve file-level mapping; the automated verification rules are: the completeness of nested structure parsing, the correctness of dynamic field mapping, and the rationality of null value handling.

[0074] In one specific embodiment, selecting appropriate collaborative retrieval strategies based on different application business types can improve retrieval efficiency and accuracy. For different types of query needs, the system leverages the characteristics of vector databases and knowledge graphs to achieve synergistic interaction between semantic information and structured knowledge; the system also includes: The application management layer is also used to design a dual-database collaborative retrieval mechanism based on vector databases and knowledge graphs. In the process of business verification and intelligent response, it receives input application business parameters and determines the type of the current application business. Based on the determined type of application business, it matches a collaborative strategy that is suitable for it.

[0075] Specifically, the application service types include: precise query, fuzzy query, complex reasoning query, and hybrid query. The corresponding service type can be obtained based on user input by training a neural network model.

[0076] For precise queries, where the user query has clearly defined entities and relationships, it's suitable to first use a knowledge graph for precise matching. If the answer cannot be found in the knowledge graph, then a vector database is used for similarity matching. Therefore, the appropriate collaborative strategy is to first perform a query based on the knowledge graph, followed by a query based on the vector database. For fuzzy queries, where the user query is more ambiguous, it's suitable to first use a vector database for similarity matching to find relevant content, then extract entities and relationships from these results, and finally use a knowledge graph for supplementation and verification. Therefore, the appropriate collaborative strategy is also to first perform a query based on the vector database, followed by a query based on the knowledge graph.

[0077] In this embodiment, taking the new project planning auxiliary verification process as an example, the new project planning document uploaded by the user is mainly intended to obtain decision support information such as similar cases, risk warnings, and optimization suggestions, which belongs to the fuzzy query type. Starting from user input, the system receives planning parameters (such as dimensions, tolerances, material types, and performance requirements). First, these parameters are vectorized to generate a query vector. Next, the system uses this vector to perform similarity matching in the vector database, retrieving the text fragments with the most semantically similar meaning. Simultaneously, the system further calls the knowledge graph entity retrieval function, searching for graph entities related to the query based on the associated triplet information. Then, the system enters the graph exploration stage, combining semantic matching results and graph relationship paths to expand the search scope and uncover potential related cases. Finally, the system integrates the results of vector retrieval and graph reasoning, outputting structured similar case retrieval result graph data.

[0078] Among them, complex reasoning queries require multiple steps of reasoning from the user. These queries need the graph traversal capabilities of a knowledge graph, and the knowledge graph can be used directly. Therefore, the appropriate collaborative strategy is a collaborative strategy based on knowledge graph query retrieval. Hybrid queries refer to user queries that contain both precise structured information and fuzzy semantic information. These queries require collaborative retrieval, which can execute knowledge graph and vector database retrieval in parallel, and then fuse the results. Therefore, the appropriate collaborative strategy is a collaborative strategy based on parallel retrieval of vector databases and knowledge graphs.

[0079] In this embodiment, taking business process failure analysis response as an example, this type of data is complex reasoning type. Therefore, when a user uploads a file, the user's failure problem is parsed through the intent recognition interface, the entity list is extracted, and it is determined whether the problem is ambiguous. If it is ambiguous, clarification options are generated to supplement the text problem. If it is not ambiguous, the knowledge graph data and document data are recalled, and the failure link is reasoned based on the knowledge graph to recommend similar cases and countermeasures.

[0080] like Figure 2As shown in the embodiments, this application discloses a multi-source data integration and management method based on knowledge graphs and large models. Using the multi-source data integration and management system described in the above embodiments, the following steps are performed: S1. Automated collection of multi-source heterogeneous data is achieved through the data access layer.

[0081] Specifically, it involves collecting structured and unstructured data to build a multi-source heterogeneous original knowledge base.

[0082] S2. Data cleaning, pattern mapping, and multimodal knowledge extraction and fusion are completed through the data processing layer.

[0083] Specifically, the process of extracting multimodal knowledge from collected data is divided into the extraction of structured and unstructured data, and the fusion of extracted knowledge. For example... Figure 4 As shown, the knowledge extraction process for knowledge base text blocks in unstructured data in multimodal datasets includes: parallel querying of preset knowledge graph pattern definitions and knowledge base text blocks; constructing accurate prompt words based on both; calling a large model to perform triple extraction; after the extraction results are verified for completeness and accuracy, if they do not meet the standards, the prompt word template is optimized and re-extracted until structured knowledge that meets the standards is output. Figure 5 As shown, the knowledge extraction process for image data in unstructured data within multimodal knowledge includes: achieving semantic association between images and text based on a large model (such as Qwen3-VL); after completing text-side entity extraction, acquiring document image and text entity information in parallel, constructing prompt words containing image context and entity lists, calling the large model to parse image semantics, establishing associations between image and text entities through a cross-modal alignment mechanism, and generating structured indexing results to assist in achieving knowledge fusion of multimodal data. Figure 6 As shown, the knowledge graph mapping process for structured data includes: parallel acquisition of database table field information and knowledge graph ontology schema; sequential completion of concept mapping, attribute mapping, and relation mapping; in full import mode, all data that meets the conditions is converted into RDF triples and written to the graph database at once; in incremental import mode, only newly added or changed data is processed to achieve real-time data synchronization and dynamic knowledge updates, thereby improving data processing efficiency.

[0084] S3. Construct a graph knowledge network that integrates structured and unstructured data.

[0085] S4. Accurate knowledge matching and intelligent decision support in business scenarios are achieved through dual-database collaborative retrieval technology. For example... Figure 3 As shown, based on actual business scenarios, searches are performed by business scenario to achieve knowledge matching for different business scenarios, obtain similar case recommendations or fault analysis, and assist in making intelligent decisions.

[0086] This application also discloses a computer-readable storage medium.

[0087] Specifically, the computer-readable storage medium stores a computer program that can be loaded by a processor and executed as described above in the multi-source data integration and management method based on knowledge graphs and large models. The computer-readable storage medium includes, for example, various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0088] This application also discloses a computer device.

[0089] Specifically, the computer device includes a memory and a processor. The memory stores computer programs that can be loaded by the processor and executed using the aforementioned multi-source data integration and management method based on knowledge graphs and large models.

[0090] The above are all preferred embodiments of this application and are not intended to limit the scope of protection of this application. Any feature disclosed in this specification (including the abstract and drawings) may be replaced by other equivalent or similar features unless specifically stated otherwise. That is, unless specifically stated otherwise, each feature is only one example of a series of equivalent or similar features.

Claims

1. A multi-source data integration and management system based on knowledge graphs and large models, characterized in that, include: Infrastructure layer, data resource layer, data access layer, data processing layer, knowledge graph layer, and application management layer; The infrastructure layer provides hardware support and a runtime environment; the data resource layer integrates structured and unstructured data to build a knowledge base. The data access layer is used to design the connection mechanism to complete the automated collection and flow of multi-source data; The data processing layer is used to perform data cleaning, pattern mapping, knowledge extraction, knowledge normalization fusion, and vector database generation on the multi-source data. The knowledge extraction includes: for unstructured text data, constructing prompt words based on a preset knowledge graph pattern definition and the corresponding knowledge base text blocks of the unstructured text data, calling a large model to perform triple extraction, and iteratively optimizing the prompt word template based on the verification results; for unstructured image data, acquiring document image and text entity information in parallel, constructing prompt words containing image context and entity list, calling a large model to parse image semantics, establishing the association between image and text entities through a cross-modal alignment mechanism, and generating structured indexing results; the knowledge normalization fusion includes: fusing the extracted triples and structured indexing results; and realizing the mapping between structured data and knowledge graph through D2R mapping technology; the vector database generation includes: generating multimodal vectors based on the knowledge base and completing index construction; The knowledge graph layer is used to construct a knowledge graph that includes pattern definition, graph reasoning, entity management, and graph retrieval; the application management layer is used to design a dual-database collaborative retrieval mechanism based on vector database and knowledge graph to achieve business verification and intelligent response.

2. The multi-source data integration and management system based on knowledge graphs and large models according to claim 1, characterized in that, The infrastructure layer comprises a cluster of four distributed servers; the first server carries core components and runs containerized services; the second server deploys a message queue and file storage system; the third server stores relational data and graph databases; and the fourth server stores cached data and index data. The distributed server clusters all achieve separation of static and dynamic resources and traffic distribution through a load balancing layer.

3. The multi-source data integration and management system based on knowledge graphs and large models according to claim 1, characterized in that, The data access layer design includes the following connection mechanisms: interface connection, structured connection, unstructured connection, and connection task scheduling, which coordinate the various connection mechanisms to complete the automated collection and flow of different types of data.

4. The multi-source data integration and management system based on knowledge graphs and large models according to claim 1, characterized in that, The data processing layer is also used to classify the current unstructured text data into types based on data scale, sample richness, and real-time performance, including: high real-time and small sample scenario type, low real-time and large sample scenario type, high real-time and large sample scenario type, and low real-time and small sample scenario type; and to adapt knowledge extraction path decisions according to different data types, including: a first knowledge extraction path decision matching the high real-time and small sample scenario type, a second knowledge extraction path decision matching the low real-time and large sample scenario type, and a third knowledge path decision matching the high real-time and large sample scenario type and the low real-time and small sample scenario type. The first knowledge extraction path decision includes prompting engineering and large model iteration path; knowledge extraction using the first knowledge extraction path decision includes: for unstructured text data, constructing prompt words based on the knowledge base text blocks corresponding to the unstructured text data defined and annotated according to a preset knowledge graph pattern, calling the large model to perform triple extraction, and iteratively optimizing the prompt word template based on the verification results; the second knowledge extraction path decision includes pre-annotation and large model fine-tuning path; knowledge extraction using the second knowledge extraction path decision includes: annotating some unstructured text data with triples, directly inputting the knowledge base text blocks corresponding to the unstructured text data annotated with triples into the large model for model parameter fine-tuning, and using the fine-tuned large model data to perform triple extraction; the third knowledge path decision is a hybrid application of the first and second knowledge extraction path decisions.

5. The multi-source data integration and management system based on knowledge graphs and large models according to claim 1, characterized in that, The infrastructure layer is also used to provide streaming computing resources; the data resource layer is also used to divide the real-time data pool and the batch data pool to achieve hot and cold data separation; the data access layer is also used to add streaming access interfaces for multimodal data. The data processing layer is further divided into a streaming processing sublayer and a batch processing sublayer; the streaming processing sublayer is used to perform data cleaning, pattern mapping, knowledge extraction and knowledge normalization fusion, and semantic vector generation in real time; the batch processing sublayer is used to perform data cleaning, pattern mapping, knowledge extraction and knowledge normalization fusion, and semantic vector generation in batches. The knowledge graph layer is further used to divide the graph into layers, including an ontology layer storing domain ontology definitions and cross-modal ontology mapping rules, an instance layer storing triples that map core entities and relations and structured data, and a multimodal attribute layer storing attribute triples extracted from unstructured data and associations with multimodal feature vectors. The ontology layer for constructing the knowledge graph adopts a strategy of static generation and batch updating; the instance layer adopts a strategy of real-time generation and incremental updating; and the multimodal attribute layer adopts a strategy of real-time generation and streaming updating.

6. The multi-source data integration and management system based on knowledge graphs and large models according to claim 1, characterized in that, The data processing layer integrates a confidence-driven automatic verification module, which is used to build a rule base based on the graph pattern definition, combine a lightweight decision tree to detect logical contradictions, and obtain the confidence of the triplet extraction. Multimodal contrastive learning is used to verify the consistency between image descriptions and associated triples, and to obtain the confidence level of structured indexing results. When the confidence level of triple extraction or the confidence level of structured indexing results is lower than the corresponding preset confidence level, manual review is triggered to verify the completeness and accuracy of triple extraction results and structured indexing results output by the large model. The review results obtained from the verification are then fed back to the prompt word template optimization process.

7. The multi-source data integration and management system based on knowledge graphs and large models according to claim 1, characterized in that, The data processing layer is also used to classify the types of structured data based on table features, data characteristics, and business characteristics. In the process of mapping structured data to knowledge graphs through D2R mapping technology, a type-customized differentiated mapping strategy is adopted according to the structured data type to complete the mapping from data tables in structured data to knowledge graph concepts, the mapping from data fields in structured data to knowledge graph attributes, and the mapping from foreign keys between tables or business logic in structured data to relationships in knowledge graphs.

8. The multi-source data integration and management system based on knowledge graphs and large models according to claim 7, characterized in that, The data processing layer is also used to support parallel mapping of different data types, full mapping of batch processing, and incremental mapping of different data types with triggering mechanisms and processing strategies during the mapping of structured data and knowledge graphs through D2R mapping technology, and to complete the mapping accuracy verification according to the automatic verification rules of different data types.

9. The multi-source data integration and management system based on knowledge graphs and large models according to claim 1, characterized in that, The knowledge graph layer includes a normalized dictionary and a thesaurus, which are used to continuously learn and update term mapping rules.

10. The multi-source data integration and management system based on knowledge graphs and large models according to claim 1, characterized in that, The application management layer is also used to design a dual-database collaborative retrieval mechanism based on vector database and knowledge graph. In the process of realizing business verification and intelligent response, it receives input application business parameters and determines the type of current application business, including precise query type, fuzzy query type, complex reasoning type and hybrid query type. Based on the defined type of application business, a matching collaborative strategy is matched, including: a collaborative strategy that pre-queries based on knowledge graphs and then queries based on vector databases to match precise query types; A collaborative strategy adapted to fuzzy queries, which involves pre-querying and retrieving from a vector database followed by querying and retrieving from a knowledge graph; a collaborative strategy adapted to complex reasoning queries, which involves querying and retrieving from a knowledge graph; and a collaborative strategy adapted to hybrid queries, which involves parallel retrieval from a vector database and a knowledge graph.