Metadata-based data governance knowledge graph construction method

By building a unified meta-model structure and dynamically adjusting the ontology model, combined with the real-time update capability of Graph RAG model, the flexibility and timeliness of metadata management in the existing technology are solved, and efficient and accurate knowledge graph construction and update are achieved.

CN120181206AActive Publication Date: 2025-06-20BEIJING INST OF TECH

Patent Information

Application Number
CN202510653509.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-21
Publication Date
2025-06-20
Estimated Expiration
2045-05-21

AI Technical Summary

Technical Problem

The existing methods of building knowledge graphs based on metadata are not flexible and universal, and it is difficult to adapt to the data characteristics of different enterprises and fields, and lack the ability to adjust dynamically in real time, resulting in limited accuracy and timeliness of knowledge graphs.

Method used

By formulating multiple acquisition task areas, building a unified meta-model structure, dynamically adjusting the ontology model, and using the Graph RAG model to update real-time knowledge graphs, automatic identification, extraction and fusion of metadata are realized.

Benefits of technology

It realizes flexible, accurate and intelligent management of metadata, can be customized and adjusted according to the data characteristics of different enterprises and fields, has the ability to update dynamically in real time, and improves the timeliness and accuracy of the knowledge graph.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120181206A_ABST
    Figure CN120181206A_ABST
Patent Text Reader

Abstract

The invention provides a data governance knowledge graph construction method based on metadata, and the method comprises the steps: defining a meta-model structure of each class of metadata and a relationship between meta-models, summarizing and abstracting all meta-model structures, and forming a unified meta-model structure after constructing a corresponding relationship; dynamically adjusting the uniform meta-model structure; constructing a dynamic ontology model, mapping the unified meta-model structure to the dynamic ontology model, and adjusting the dynamic ontology model according to the dynamic change of the unified meta-model structure; performing data processing on the metadata in the metadata lake; according to a corresponding mapping relation between the unified meta-model structure and a dynamic ontology model, enabling the structured data after data processing to correspond to the dynamic ontology model, and completing the construction of a primary knowledge graph; and detecting and preprocessing the real-time data stream to form updated content of the knowledge graph. The method aims at achieving effective fusion and dynamic updating of multi-source heterogeneous metadata knowledge in the data governance process.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of metadata management technology, and in particular to a method for constructing a data governance knowledge graph based on metadata. Background Art

[0002] With the rapid development of information technology, the digital transformation of enterprises has become an irreversible trend. In this process, enterprises continue to accumulate massive, diverse, and heterogeneous data resources, which contain rich business value and innovation potential. However, the explosive growth of data has also brought unprecedented management challenges. How to effectively manage this data and tap its intrinsic value has become an urgent issue facing enterprises.

[0003] Metadata is the cornerstone of data governance. Data governance must first manage metadata. When metadata is clearly defined and reasonably designed, data quality will inevitably improve. The most common definition of metadata is "data about data" or "data describing data." "GB / T 18391.1-2009 Information Technology Metadata Registration System (MDR). Part 1: Framework" defines metadata as data that defines and describes other data. It describes the structure, content, context and management rules of data.

[0004] Metadata is generally divided into business metadata, technical metadata, and operational metadata. Business metadata describes the business meaning, business rules, and relationships of data. It includes business definitions, business terms, business rules (business engine rules, data quality detection rules), business models (conceptual models, logical models), data security, sensitivity levels, etc. Technical metadata describes relevant conceptual information in the technical field of the system. It includes data structure (name, length, type, constraints, relationships, etc.), data processing (ETL information, scheduling information, update frequency, etc.), data storage (type, location, file format, compression type, etc.). Operational metadata describes the operational attributes of data and clarifies data management responsibilities. It includes data owners, data custodians, data access control information (organizational roles, access methods, access cycles, access scopes), data backup (archiving location, archiving date, archiving cycle), etc.

[0005] Microscopically, data governance refers to the data management of individuals, covering the entire data life cycle, that is, the overall management of the practicability, availability, integrity, and security of data. The core purpose is to achieve the security and availability of data. Currently, it includes data integration management, data exchange management, data modeling management, metadata management, data standard management, data quality management, master data management, data asset management, data security management, data life cycle management, etc. Traditional data governance methods often rely on manual intervention, which is not only inefficient but also error-prone and difficult to meet the needs of modern enterprises for efficient and accurate data management. Therefore, it is particularly important to explore an automated and intelligent data governance method, especially a metadata management solution.

[0006] Currently, as a powerful knowledge representation and reasoning tool, the knowledge graph has shown great potential in the field of data governance. Essentially, the knowledge graph is a knowledge base of semantic networks. It stores entities related to the real world and the relationships between entities in the form of a graph. From the perspective of practical applications, the knowledge graph can actually be simply understood as a multi-relational graph. Entities are the basic units in the knowledge graph, representing objects or concepts in the real world. Relationships are the relationships between entities, describing how entities are related to each other. Attributes are the characteristics or properties of entities, providing detailed information about entities. The graph structure means that the knowledge graph stores data in the form of a graph, where nodes represent entities and edges represent relationships. Semantics means that the information in the knowledge graph has clear semantics, enabling machines to understand the meaning of data. Reasoning means that the knowledge graph can be used for complex reasoning to discover implicit relationships between entities. Data integration means that the knowledge graph can integrate data from different sources and provide a unified view.

[0007] Knowledge Fusion is based on multi-source heterogeneous data. With the support of ontology libraries and rule libraries, it obtains knowledge factors and their associated relationships hidden in data resources through knowledge extraction and transformation, fuses the description information of the same entity or concept from multiple sources, and combines, reasons, and creates new knowledge at the semantic level. And this process needs to be adjusted in real time dynamically according to changes in data sources and user feedback.

[0008] However, the existing methods for constructing knowledge graphs based on metadata still have many deficiencies. On the one hand, these methods often rely on preset features or rules for classifying metadata categories and extracting relationships, which limits the flexibility and universality of the methods and makes it difficult to adapt to the data characteristics of different enterprises and fields. On the other hand, some methods overly rely on data quality and are insufficient in dealing with noise and outliers in the data, resulting in limited accuracy of the constructed knowledge graph. In addition, when implementing knowledge fusion, the existing methods often lack the ability to adjust in real-time and dynamically, and it is difficult to update in a timely manner according to changes in data sources and user feedback.

[0009] In summary, there is an urgent need for a more flexible, accurate, and intelligent metadata-based knowledge graph construction method to address the data governance challenges in enterprise digital transformation. This method should be able to automatically identify and extract key information in metadata, construct a unified and comprehensive metadata knowledge graph; at the same time, it should have good adaptability and scalability, and be able to perform customized adjustments according to the data characteristics of different enterprises and fields; in addition, it should also have the ability to update in real-time and dynamically to ensure the timeliness and accuracy of the knowledge graph. Summary of the Invention

[0010] Aiming at the problems existing in the prior art, the purpose of the present invention is to provide a method for constructing a data governance knowledge graph based on metadata, aiming to solve the problem of how to effectively fuse and dynamically update the multi-source heterogeneous metadata knowledge generated in multiple links during the data governance process.

[0011] The present invention realizes the above purpose through the following technical solutions: A method for constructing a data governance knowledge graph based on metadata, comprising the following steps: Formulate multiple collection task areas, and divide the collection task areas into real-time metadata collection tasks and real-time business data collection tasks; Define the meta-model structure of each type of metadata and the relationships between meta-models according to the collected metadata, summarize and abstract all the meta-model structures, and form a unified meta-model structure after constructing the corresponding relationships, which is used to accommodate information of all meta-model types; Perform real-time detection through the metadata collection task, and dynamically adjust the unified meta-model structure; Construct a dynamic ontology model, map the unified meta-model structure to the dynamic ontology model, including mapping metadata types to classes, metadata to entities, metadata attribute sets to attributes, and metadata relationship sets to relationships, and adjust the dynamic ontology model according to the dynamic changes of the unified meta-model structure; Preprocess, extract knowledge, and fuse knowledge for the metadata in the metadata lake; According to the corresponding mapping relationship between the unified meta-model structure and the dynamic ontology model, the structured data after knowledge fusion is mapped to the dynamic ontology model, and the transformed data is stored in the graph database to complete the construction of the initial knowledge graph; After detecting and preprocessing the real-time data stream, connect it to the pre-trained Graph RAG model, automatically extract entities, relationships and their attributes from the preprocessed data, and represent the processing in the form of a graph structure to form the updated content of the knowledge graph; add the newly extracted entities and relationships to the existing initial knowledge graph, and at the same time update the existing entity and relationship information.

[0012] According to a method for constructing a data governance knowledge graph based on metadata provided by the present invention, for the structured database data source, the change data capture CDC framework is used to monitor the binlog log of the database to capture the change events in the database in real time, and the captured change events are parsed to obtain structured metadata or business data; For the log file data source, use the Logstash log collection tool for real-time collection. By configuring the input, filter, and output plugins of Logstash, realize the real-time reading, processing, and forwarding of the log file, and then collect the data in the log file to the specified storage location; For the text file data source, use a file monitoring tool to monitor the update of the text file in real time. When the file changes, trigger the collection task, read and parse the updated file content to obtain the metadata or business data in the text file.

[0013] According to a method for constructing a data governance knowledge graph based on metadata provided by the present invention, the unified meta-model structure further defines the specific composition and hierarchical relationship of metadata types, specifically including: Entities related to data modeling, including database connection entities, database instance entities, database entities, data table entities, field entities, view entities, function entities, and data modeling rule entities. Among them, the data modeling rule entity is used to specify the standards and specifications for database design, data table structure, field definition, view creation, and function writing; Entities related to data integration, including task entities and scheduling entities. The task entity is used to describe the specific jobs or operations in the data integration process, and the scheduling entity is used to define the order, time, and frequency of task execution; Entities related to data services, including system entities, page entities, API entities, and dataset entities. The system entity represents the software system that provides data services, the page entity refers to the user interface within the system, the API entity represents the interface provided by the system to the outside, and the dataset entity is the specific data collection of the data service; Data management-related entities, including account entities, log entities, message entities, and document entities. Account entities are used to manage user access permissions, log entities are used to record the system operation history, message entities are used for information transmission within the system or across systems, and document entities contain various types of document materials in the data governance process.

[0014] According to a method for constructing a data governance knowledge graph based on metadata provided by the present invention, the defined unified meta-model structure is classified and labeled according to three classification methods: entity information, relationship information, and attribute information. Among them, entity information treats each metadata item as a metadata entity, which includes data modeling-related entities, data integration-related entities, data service-related entities, and data management-related entities; relationship information is used to define the associations, dependencies, or hierarchical relationships existing between various entities of the meta-model; and attribute information defines various attributes possessed by the meta-model entities themselves. The classified and labeled meta-model structure, together with its corresponding label information, is stored in the meta-model structure library, and a unified meta-model structure framework is summarized and abstracted based on the classification and labeling results. This framework at least includes meta-model type, meta-model entity information, relationship information, and attribute information.

[0015] According to a method for constructing a data governance knowledge graph based on metadata provided by the present invention, the unified meta-model structure is dynamically adjusted, including: For structured databases, set metadata collection tasks, real-time detect the log information in the database, parse the DDL statement information related to data structure changes in the log, and match the type of DDL operation object with the meta-model types pre-stored in the meta-model structure library. If the type match is successful, further parse the attributes of the operation object in the DDL statement, and match these attributes with the attributes of the corresponding meta-model in the meta-model structure library one by one. If the attribute match is successful, it indicates that the data structure change caused by this DDL statement has been reflected in the meta-model structure library, and no processing is required; if the attribute match is not successful, it indicates that the DDL statement has introduced new attributes or changed existing attributes, then new attributes are added to the meta-model in the meta-model structure library, and this change information is synchronously updated to the unified meta-model structure.

[0016] According to a method for constructing a data governance knowledge graph based on metadata provided by the present invention, for newly generated data structures, natural language processing technology and machine learning algorithms are used for active identification and classification to determine the meta-model type to which they belong; the newly identified meta-model is automatically stored in the meta-model structure library and classified and labeled. For the newly recognized meta-model, according to its meta-model type and predefined standardized conversion rules, conversion operations are automatically performed to convert it into a unified meta-model structure; Among them, the standardized conversion rules are used to describe how to convert different types of meta-models into a unified meta-data model structure.

[0017] According to a method for constructing a data governance knowledge graph based on metadata provided by the present invention, the metadata in the metadata lake is preprocessed according to data standards and preset data quality rules, specifically including: Data integrity verification: According to the data standards, perform integrity checks on the metadata in the metadata lake; Data uniqueness verification: When the metadata has a business unique key, perform uniqueness verification, and its specific indicators include the unique key conflict rate; Data accuracy verification: According to the data standards, verify the generation trigger mechanism and attribute value filling of the metadata, and its specific indicators include the accuracy rate; Data consistency verification: Check the logical consistency between different attributes of business objects and the consistency of business process magnitude fluctuations, which is carried out by means of cross-table and value comparison, and fluctuation consistency comparison; Data timeliness verification: According to the data standards, verify the generation time and arrival delay of the metadata; Data standardization / normalization processing: Perform standardization or normalization processing on the metadata to eliminate the influence of dimensions or numerical values on the calculation results, specifically including Z-Score standardization and Max-Min normalization; Data quality report generation: According to the definition of data quality rules, regularly generate a data quality report to record the quality status of the metadata in terms of integrity, uniqueness, accuracy, consistency, and timeliness; Data quality problem handling: According to the records in the data quality report, respond to, locate, analyze, and solve the discovered metadata quality problems to ensure that the metadata in the metadata lake continuously meets the data standards and preset data quality rules.

[0018] According to a method for constructing a data governance knowledge graph based on metadata provided by the present invention, the knowledge extraction process includes entity extraction, attribute extraction, and relationship extraction, specifically including: Meta-model structure matching: Use the meta-model structure library to match the preprocessed metadata with the meta-model structure; directly identify and locate the entity and attribute information in the metadata according to the classification labels in the meta-model structure; Entity extraction: Apply natural language processing techniques and machine learning algorithms to preprocess and extract features from the text content in the metadata, and perform entity extraction, attribute extraction, and relationship extraction; Attribute extraction: Construct an attribute list for each recognized entity, apply an attribute extraction algorithm to extract entity-related attribute information from the metadata text, and assign attribute values to these attributes; Relationship extraction: On the basis of entity extraction and attribute extraction, further identify and extract the relationships between entities; apply a relationship extraction algorithm to analyze the context and semantics in the metadata text to determine the association types between entities; construct an entity relationship graph to intuitively display the complex relationship network between entities; Knowledge integration and storage: Integrate the results of entity extraction, attribute extraction, and relationship extraction to form a structured knowledge representation, and store the structured knowledge in a knowledge base.

[0019] According to a method for constructing a data governance knowledge graph based on metadata provided by the present invention, the knowledge extraction and fusion process specifically includes: Select the entities to be aligned, and for each entity, select its key attributes. The entity itself, its attributes and attribute values, and the relationships with other entities respectively form three groups of features; use vectorization technology to convert the three groups of features into vector representations respectively, and fuse these vectors into a unified representation in a low-dimensional vector space through a concatenation operation; In the constructed low-dimensional vector space, use a semantic distance calculation method to measure the similarity between entities; set a threshold for semantic distance. For entity pairs with a semantic distance less than this threshold, initially judge that they may be aligned or equivalent entities; For entity pairs initially judged to be possibly aligned, perform detection and processing of attribute conflicts and relationship conflicts. By comparing whether the attribute values of entities are consistent or compatible, and whether the relationships between entities match each other, determine whether these entities can actually be merged into a unified entity; among them, for entities confirmed to be mergeable, perform a merge operation and integrate them into a new entity representation.

[0020] According to a method for constructing a data governance knowledge graph based on metadata provided by the present invention, embed the preprocessed new information into the graph structure of the constructed knowledge graph to form a graph structure fragment containing new entities, new relationships and their attributes, and use the embedding layer of the GraphRAG model to convert the new information and the existing information into low-dimensional vector representations; Based on the reasoning ability of the GraphRAG model, check the logical relationship and consistency between the new information and the existing information in the knowledge graph; Using the attention mechanism of the GraphRAG model, focus on entities and relationships directly related to new information, as well as the connections and constraints between them; by comparing the distance and similarity between new information and existing information in the vector space, evaluate the rationality and credibility of new information; according to the results of logical relationship and consistency checks, the GraphRAG model automatically infers whether the new information conforms to the existing information in the knowledge graph and whether there are conflicts or contradictions. The GraphRAG model outputs a verification decision, including accepting the new information, rejecting the new information, or marking it as information to be further verified; for the accepted new information, it is formally incorporated into the knowledge graph to update the structure and content of the graph; for the rejected or to-be-verified information, corresponding feedback is output.

[0021] It can be seen that the present invention proposes a method for constructing a data governance knowledge graph based on metadata. By integrating multi-source heterogeneous metadata knowledge generated during the data governance process, this method can form a unified metadata governance view. Therefore, the present invention brings the following remarkable beneficial effects: 1. The present invention constructs a unified metadata governance view by integrating metadata from different sources and in different formats, enabling data governance personnel to quickly and accurately locate the required metadata and its corresponding data information, greatly improving the efficiency and convenience of data governance.

[0022] 2. Under the unified metadata governance view, data governance personnel can quickly find relevant metadata and data information without having to switch back and forth between multiple heterogeneous data sources, greatly saving time costs and improving the response speed of data governance.

[0023] 3. The knowledge graph constructed by the present invention contains rich metadata information and the relationships between them, providing a solid data foundation for the automatic generation of data quality rules. By analyzing and mining the metadata relationships in the knowledge graph, data quality rules that meet the requirements of data governance can be automatically generated to ensure the quality and accuracy of data.

[0024] 4. The metadata and their relationships in the knowledge graph also provide strong support for the automatic generation of data standards. The present invention can automatically generate data standards that meet business requirements and industry standards by classifying, summarizing, and organizing the metadata in the knowledge graph, thereby standardizing the use and management of data.

[0025] 5. In the process of data governance, data security auditing is an important task. The knowledge graph constructed by the present invention can clearly display the association relationships between metadata, enabling data security auditing personnel to quickly locate potential security risk points and conduct targeted auditing and inspections, thereby ensuring the security of data.

[0026] 6. By monitoring changes in data sources and using natural language processing technology to automatically identify the characteristics of newly introduced datasets, the present invention can recommend appropriate classifications and tags, thereby reducing the degree of manual intervention. This not only improves the automation level of data governance but also reduces the risk of human errors.

[0027] 7. For the complex relationships involved in the data governance process, the present invention designs a unified knowledge representation model that can clearly express the association relationships among data, business, and management, enabling data governance personnel to better understand and analyze various problems and challenges in the data governance process, thereby enhancing the ability and level of data governance.

[0028] 8. Based on the Graph RAG technology, the present invention realizes incremental knowledge graph construction, which can avoid reconstructing the entire knowledge graph, thereby maintaining the timeliness and accuracy of information, enabling data governance personnel to obtain the latest metadata and its relationship information in a timely manner, and providing strong support for data governance work.

[0029] In summary, the method for constructing a data governance knowledge graph based on metadata proposed by the present invention brings many beneficial effects, can significantly improve the efficiency and accuracy of data governance, and provides strong support for the digital transformation and business development of enterprises.

[0030] The following further elaborates on the present invention in detail in conjunction with the accompanying drawings and specific implementation manners. Description of the Drawings

[0031] Figure 1 is a flowchart of an embodiment of a method for constructing a data governance knowledge graph based on metadata of the present invention.

[0032] Figure 2 is a schematic diagram of an embodiment of a method for constructing a data governance knowledge graph based on metadata of the present invention. Detailed Implementation Manner

[0033] To make the objectives, technical solutions, and advantages of the present invention clearer, the technical solutions in the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art without creative efforts based on the embodiments in the present invention belong to the scope of protection of the present invention.

[0034] References herein to "embodiments" mean that the particular features, structures, or characteristics described in connection with the embodiments can be included in at least one embodiment of the present application. The phrase appears in various places in the specification and does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment mutually exclusive with other embodiments. Those skilled in the art will explicitly and implicitly understand that the embodiments described herein can be combined with other embodiments.

[0035] See Figure 1 With Figure 2 , this embodiment provides a method for constructing a data governance knowledge graph based on metadata, and the method includes the following steps: Step S1, formulate multiple collection task areas, and divide the collection task areas into real-time metadata collection tasks and real-time business data collection tasks; Step S2, define the meta-model structure of each type of metadata and the relationships between meta-models according to the collected metadata, summarize and abstract all the meta-model structures, and form a unified meta-model structure after constructing the corresponding relationships, which is used to accommodate information of all meta-model types; Step S3, perform real-time detection through the metadata collection task, and dynamically adjust the unified meta-model structure; Step S4, construct a dynamic ontology model, map the unified meta-model structure to the dynamic ontology model, including mapping metadata types to classes, metadata to entities, metadata attribute sets to attributes, and metadata relationship sets to relationships, and adjust the dynamic ontology model according to the dynamic changes of the unified meta-model structure; Step S5, preprocess, extract knowledge, and fuse knowledge for the metadata in the metadata lake; Step S6, according to the corresponding mapping relationship between the unified meta-model structure and the dynamic ontology model, correspond the structured data after knowledge fusion to the dynamic ontology model, and store the converted data in the graph database to complete the construction of the initial knowledge graph; Step S7, after detecting and preprocessing the real-time data stream, connect it to a pre-trained Graph RAG model, automatically extract entities, relationships, and their attributes from the preprocessed data, and represent the processing in the form of a graph structure to form the updated content of the knowledge graph; add the newly extracted entities and relationships to the existing initial knowledge graph, and at the same time update the existing entity and relationship information.

[0036] In the above step S1, when collecting data in real time, two different types of real-time collection tasks are distinguished: real-time metadata collection tasks and real-time business data collection tasks. The data collected in real time for metadata is stored in the metadata lake, and the data collected by the real-time business data collection task is stored in the data lake.

[0037] Among them, the real-time metadata collection task and the real-time business data collection task are separated logically and / or physically to ensure the independent collection, processing, and storage of metadata and business data, thereby improving the efficiency of data collection and the flexibility of data management.

[0038] In this embodiment, the data source types involved in the data governance process are mainly divided into structured databases and text files. For structured database data sources, the Change Data Capture (CDC) framework is used to monitor the binlog logs of the database to capture change events in the database in real time, and the captured change events are parsed to obtain structured metadata or business data. For log file data sources, the Logstash log collection tool is used for real-time collection. By configuring the input, filter, and output plugins of Logstash, real-time reading, processing, and forwarding of log files are achieved, and then the data in the log files is collected to the specified storage location. For text file data sources, a file monitoring tool is used to monitor the update status of text files in real time. When a file changes, a collection task is triggered to read and parse the content of the updated file to obtain metadata or business data in the text file.

[0039] Among them, the above specific collection methods for different data source types can ensure the real-time and accuracy of data collection, while adapting to the characteristics and formats of different data sources, improving the flexibility and generality of data collection.

[0040] Specifically, for structured databases, Flink CDC is used to monitor changes in the binlog logs of the database, trigger the capture of change events, and parse the changes. Since the Flink CDC connector does not directly support DDL statement parsing, a custom parser is built and integrated into the Flink job. After synchronizing the binlog logs to Kafka based on Flink CDC, the Kafka Source of Flink is used to read the Kafka topic containing DDL statements, then the DDL statements are parsed, and the parsed results are passed to the downstream. For text files, since they are uniformly stored in the file management system, the changed files can be identified through the operation logs of the file system, and then all the files involved in the changes are extracted.

[0041] In the above step S2, the unified meta-model structure further defines the specific composition and hierarchical relationship of metadata types, specifically including: Data modeling-related entities, including database link entities, database instance entities, database entities, data table entities, field entities, view entities, function entities, and data modeling rule entities. Among them, the data modeling rule entities are used to specify the standards and specifications for database design, data table structure, field definition, view creation, and function writing.

[0042] Data integration-related entities, including task entities and scheduling entities. Task entities are used to describe specific jobs or operations in the data integration process, and scheduling entities are used to define the order, time, and frequency of task execution.

[0043] Data service-related entities, including system entities, page entities, API entities, and dataset entities. System entities represent software systems that provide data services, page entities refer to user interfaces within the system, API entities represent interfaces provided by the system to the outside, and dataset entities are specific data collections of data services.

[0044] Data management-related entities, including account entities, log entities, message entities, and document entities. Account entities are used to manage user access permissions, log entities are used to record system operation history, message entities are used for information transmission within the system or across systems, and document entities contain various types of document materials in the data governance process.

[0045] Moreover, the unified meta-model structure also defines the association relationships and hierarchical structures among the above various entities, as well as the relationships between internal attributes and attributes within the entities, thus forming a complete and standardized metadata model system, providing a solid foundation for data governance and knowledge graph construction.

[0046] Specifically, define a unified metadata model: Through the previous investigation and identification of all metadata in data governance, the metadata types include: four major categories of data modeling-related, data integration-related, data service-related entities, and data management-related entities. Among them, data modeling-related entities include database links, database instances, databases, data tables, fields, views, functions, etc.; data integration-related entities include tasks, scheduling; data service-related entities include systems, pages, APIs, datasets, etc.; data management-related entities mainly include accounts, logs, messages, documents, etc. A meta-model is a model about models, which defines the specifications of a certain model, that is, the elements that make up the model and the relationships between the elements. The specific meta-model structure is shown in Table 1.

[0047] [Table 1]

[0048] In this embodiment, the defined unified meta-model structure is classified and labeled according to three classification methods: entity information, relationship information, and attribute information. Among them, entity information treats each metadata item as a metadata entity, which includes data modeling-related entities, data integration-related entities, data service-related entities, and data management-related entities. Relationship information is used to define the associations, dependencies, or hierarchical relationships existing among various entities in the meta-model. Attribute information, on the other hand, defines various attributes possessed by the meta-model entities themselves, such as name, description, type, status, etc.

[0049] The classified and labeled meta-model structure, together with its corresponding tag information, is stored in the meta-model structure library for subsequent query, invocation, and management. Based on the classification and labeling results, a unified meta-model structure framework is summarized and abstracted. This framework at least includes meta-model type, meta-model entity information, relationship information, and attribute information, providing a standardized template and guidelines for constructing metadata models in specific domains, as shown in Table 2.

[0050] [Table 2]

[0051] In step S3 above, the unified meta-model structure is dynamically adjusted, including: For a structured database, a metadata collection task is set up to detect the log information in the database in real time, parse the DDL (Data Definition Language) statement information related to data structure changes in the log, and match the type of the DDL operation object (such as table, database, etc.) with the meta-model types pre-stored in the meta-model structure library. If the type match is successful, further parse the attributes of the operation object in the DDL statement and match these attributes one by one with the attributes of the corresponding meta-model in the meta-model structure library. If the attribute match is successful, it indicates that the data structure change caused by this DDL statement has been reflected in the meta-model structure library and no processing is required. If the attribute match is not successful, it indicates that the DDL statement has introduced new attributes or changed existing attributes. In this case, new attributes are added to the meta-model in the meta-model structure library, and this change information is synchronously updated to the unified meta-model structure.

[0052] Through the above steps, the automatic recognition, dynamic update, and unified management of the meta-model structure in the structured database are realized, greatly improving the efficiency and accuracy of metadata management.

[0053] In this embodiment, if a new data structure is generated, text parsing and semantic understanding are performed on the name, description, field information, etc. of the newly generated data structure, and active recognition and classification are carried out using natural language processing techniques and machine learning algorithms (such as classifiers, clustering algorithms, etc.) to determine the type of meta-model to which it belongs, such as a table model, a view model, or a function model, etc. The newly recognized meta-model is automatically stored in the meta-model structure library and classified and labeled; for the newly recognized meta-model, according to its meta-model type and predefined standardized conversion rules, a conversion operation is automatically performed to convert it into a unified meta-model structure. Among them, the standardized conversion rules are used to describe how to convert different types of meta-models into a unified meta-data model structure.

[0054] It can be seen that through steps such as active recognition and classification, automatic storage and classification labeling, and standardized conversion and consistency maintenance, this embodiment realizes the automatic recognition, classification, and standardized conversion of the meta-model of the newly generated data structure in the database, providing strong support for the meta-data management, data governance, and value mining of data assets in the database.

[0055] In the above step S4, regarding the construction of the dynamic ontology model, an ontology is a formal, canonical, and explicit description of a shared conceptual model, including elements such as classes (concepts), relationships, functions, axioms, and instances. Through a one-to-one mapping between the ontology model and the unified meta-model, where the "type" of the meta-model is mapped to the "class" of the ontology model, the "entity" of the meta-model is mapped to the "entity" of the ontology model, the "attribute" of the meta-model is mapped to the "attribute" of the ontology model, and the "relationship" of the meta-model is mapped to the "relationship" of the ontology model. When the unified meta-model structure is dynamically adjusted by detecting the structural changes of the meta-model, the ontology model will also dynamically adjust its model structure accordingly.

[0056] In the above step S5, according to data standards and preset data quality rules, the meta-data in the meta-data lake is preprocessed to convert the data into a unified format and unit, eliminate data inconsistencies; complete missing values; remove duplicate records for duplicate values, and supplement standard and rule information. It specifically includes: Data integrity verification: According to data standards, the meta-data in the meta-data lake is checked for integrity to ensure that the meta-data is not lost or duplicated during upload, transmission, and storage. Specific indicators include, but are not limited to, the transmission loss rate and the transmission duplication rate.

[0057] Data uniqueness verification: When the meta-data has a business unique key, uniqueness verification is performed to ensure the uniqueness of the meta-data in terms of business logic. Specific indicators include the unique key conflict rate.

[0058] Data accuracy verification: According to the data standard, verify the generation trigger mechanism and attribute value filling of the metadata to ensure that the metadata meets the business logic and accuracy requirements. The specific indicators include accuracy rate.

[0059] Data consistency verification: Check the logical consistency between different attributes of the business object and the consistency of the magnitude fluctuations in the business process. This is carried out through cross-table and value comparison, fluctuation consistency comparison, etc., to ensure the consistency of the metadata in terms of logic and magnitude.

[0060] Data timeliness verification: According to the data standard, verify the generation time and arrival delay of the metadata to ensure that the metadata meets the preset requirements in terms of timeliness. The specific indicators include generation delay, date drift, etc.

[0061] Data standardization / normalization processing: For machine learning models or data analysis requirements, perform standardization or normalization processing on the metadata to eliminate the influence of dimension or numerical values on the calculation results. Specifically, it includes Z-Score standardization and Max-Min normalization to ensure the effectiveness and accuracy of the metadata in model calculations.

[0062] Data quality report generation: According to the definition of data quality rules, regularly generate data quality reports to record the quality status of the metadata in terms of integrity, uniqueness, accuracy, consistency, and timeliness, providing decision-making support for metadata management.

[0063] Data quality problem handling: According to the records in the data quality report, respond to, locate, analyze, and solve the discovered metadata quality problems to ensure that the metadata in the metadata lake continuously meets the data standard and the preset data quality rules.

[0064] Through the above preprocessing steps, ensure that the metadata in the metadata lake meets the preset data quality rules in terms of integrity, uniqueness, accuracy, consistency, timeliness, and standardization / normalization, providing a solid guarantee for the effective management and efficient utilization of the metadata.

[0065] In the above step S5, the knowledge extraction process includes entity extraction, attribute extraction, and relationship extraction, specifically including: Meta-model structure matching: Use the meta-model structure library to match the preprocessed metadata with the meta-model structure; according to the classification labels in the meta-model structure, directly identify and locate the entity and attribute information in the metadata, providing a basis for subsequent entity extraction and attribute extraction tasks.

[0066] Entity extraction: Apply natural language processing techniques and machine learning algorithms to the text content in the metadata for preprocessing and feature extraction, and perform entity extraction, attribute extraction, and relationship extraction; Through a trained named entity recognition model, identify entities with specific meanings from the text, such as person names, place names, organization names, etc., and assign corresponding entity labels.

[0067] Attribute extraction: Construct an attribute list for each identified entity, apply an attribute extraction algorithm to extract attribute information related to the entity from the metadata text, and assign attribute values to these attributes.

[0068] Relationship extraction: On the basis of entity extraction and attribute extraction, further identify and extract the relationships between entities; Apply a relationship extraction algorithm to analyze the context and semantics in the metadata text, determine the association types between entities, such as parent-child relationship, subordination relationship, cooperation relationship, etc.; Construct an entity relationship graph to intuitively display the complex relationship network between entities; Knowledge integration and storage: Integrate the results of entity extraction, attribute extraction, and relationship extraction to form a structured knowledge representation, and store the structured knowledge in a knowledge base for subsequent knowledge query, analysis, and application.

[0069] Specifically, knowledge extraction includes three core tasks: entity extraction, attribute extraction, and relationship extraction. Entity extraction is also called Named Entity Recognition (NER), which identifies entities with specific meanings from the text; Attribute extraction constructs an attribute list for each identified entity and attaches attribute values; Relationship extraction identifies and extracts the relationships between entities, such as parent-child relationship, subordination relationship, etc. Through the above meta-model structure library, match the meta-model structure with the metadata, and directly find entity and attribute information according to the classification labels in the meta-model structure. Take the meta-model structure in Table 3 as an example: [Table 3]

[0070] For the text in the metadata, preprocess and extract features from the data through natural language processing techniques and machine learning algorithms, and then perform entity extraction, attribute extraction, and relationship extraction respectively.

[0071] In step S5 above, the knowledge extraction fusion process specifically includes: Select the entities to be aligned, and for each entity, select its key attributes. The entity itself, its attributes and attribute values, and the relationships with other entities respectively form three groups of features; Use vectorization technology to convert the three groups of features into vector representations respectively, and fuse these vectors into a unified representation in a low-dimensional vector space through a concatenation operation.

[0072] In the constructed low-dimensional vector space, semantic distance calculation methods, such as cosine similarity, Euclidean distance, etc., are used to measure the similarity degree between entities; a threshold of semantic distance is set, and for entity pairs with a semantic distance less than this threshold, it is preliminarily judged that they may be aligned or equivalent entities.

[0073] For entity pairs preliminarily judged to be possibly aligned, detection and handling of attribute conflicts and relationship conflicts are carried out. By comparing whether the attribute values of entities are consistent or compatible, and whether the relationships between entities match each other, it is determined whether these entities can actually be merged into a unified entity; among them, for entities confirmed to be mergeable, a merge operation is performed and integrated into a new entity representation.

[0074] Specifically, entity alignment is an important task for solving the same entities in the knowledge graph and is the key foundation for realizing the integration of the knowledge graph. This embodiment is completed through four steps: entity calculation, entity matching, entity alignment, and result evaluation and optimization. Entity calculation includes selecting the key attributes of the entity, then dividing the entity, attributes and attribute values, and relationships into three groups, representing them with vectors respectively, and then forming a unified low-dimensional vector space through splicing. Entity matching includes in this vector space, by calculating the semantic distance between entities and comparing with the threshold, to judge whether the entities are aligned. Entities with a smaller semantic distance can be considered equivalent. Entity alignment includes merging entities with a smaller semantic distance into one entity through the handling of attribute conflicts and relationship conflicts.

[0075] Finally, result evaluation and optimization are carried out: the evaluation indicators precision and recall are defined, and according to the evaluation results, the entity calculation method and the threshold are optimized and adjusted.

[0076] Through the above steps, this embodiment can achieve efficient and accurate entity alignment in the metadata lake, thereby promoting the integration of the knowledge graph and knowledge fusion, and providing strong support for knowledge management, intelligent analysis, and decision-making support in the metadata lake.

[0077] In the above step S6, the unified metadata model in data governance includes metadata models such as databases, tables, fields, tasks, schedules, APIs, etc., and the relationships between them. According to the one-to-one correspondence between the metadata model and the ontology model in the dynamic ontology model construction step, the ontology model also includes classes such as databases, tables, fields, tasks, schedules, APIs, etc., and the association relationships between them. In this embodiment, it is only necessary to import the metadata of each type of metadata model into the entities of the corresponding classes in the ontology model. Suppose there are two tables, "Sales Order" and "Product Information", and a data integration task, "Sales Order Data Synchronization". Then, import the metadata of the "Sales Order" and "Product Information" tables into the table entities of the ontology model, import the metadata of the "Sales Order Data Synchronization" task into the task entities of the ontology model, and construct the association relationship between the task entities and the table entities. Through the above steps, the final data is stored in the graph database, and the construction of the initial knowledge graph is completed.

[0078] In the above step S7, load the pre-trained and optimized GraphRAG (Relational Attention Graph) model, which combines graph neural networks and attention mechanisms, is good at processing graph-structured data, and has powerful reasoning and verification capabilities. Embed the preprocessed new information into the graph structure of the constructed knowledge graph to form a graph structure fragment containing new entities, new relationships, and their attributes. Use the embedding layer of the GraphRAG model to convert the new information and the existing information into low-dimensional vector representations; based on the reasoning ability of the GraphRAG model, check the logical relationship and consistency between the new information and the existing information in the knowledge graph; use the attention mechanism of the GraphRAG model to focus on the entities and relationships directly related to the new information, as well as the connections and constraints between them; by comparing the distance and similarity between the new information and the existing information in the vector space, evaluate the rationality and credibility of the new information; according to the results of the logical relationship and consistency check, the GraphRAG model automatically infers whether the new information conforms to the existing information in the knowledge graph, and whether there are conflicts or contradictions; the GraphRAG model outputs a verification decision, including accepting the new information, rejecting the new information, or marking it as information to be further verified; for the accepted new information, integrate it formally into the knowledge graph and update the structure and content of the graph; for the rejected or to-be-verified information, output the corresponding feedback, and the system provides the corresponding feedback or suggestions for the user or the system to further process. Record the results of each reasoning and verification, as well as the user's feedback and correction information, as the data source for the model's continuous learning. Regularly retrain and optimize the GraphRAG model to improve the accuracy and efficiency of its reasoning and verification, and ensure the quality of the update and maintenance of the knowledge graph.

[0079] In summary, this embodiment proposes a method for constructing a data governance knowledge graph based on metadata. By integrating multi-source heterogeneous metadata knowledge generated during the data governance process, this method can form a unified metadata governance view.

[0080] Furthermore, this embodiment constructs a unified metadata governance view by integrating metadata from different sources and in different formats, enabling data governance personnel to quickly and accurately locate the required metadata and its corresponding data information, greatly improving the efficiency and convenience of data governance.

[0081] Furthermore, under the unified metadata governance view, data governance personnel can quickly find relevant metadata and data information without having to switch back and forth between multiple heterogeneous data sources, saving a great deal of time and enhancing the response speed of data governance.

[0082] Furthermore, the knowledge graph constructed in this embodiment contains rich metadata information and the relationships between them, providing a solid data foundation for the automatic generation of data quality rules. By analyzing and mining the metadata relationships in the knowledge graph, data quality rules that meet the requirements of data governance can be automatically generated to ensure the quality and accuracy of data.

[0083] Furthermore, the metadata and their relationships in the knowledge graph also provide strong support for the automatic generation of data standards. The method of this embodiment can automatically generate data standards that meet business requirements and industry standards by classifying, summarizing, and organizing the metadata in the knowledge graph, thus standardizing the use and management of data.

[0084] Furthermore, in the process of data governance, data security auditing is an important task. The knowledge graph constructed in this embodiment can clearly display the association relationships between metadata, enabling data security auditors to quickly locate potential security risk points and conduct targeted audits and inspections, thereby ensuring data security.

[0085] Furthermore, by monitoring changes in data sources and using natural language processing technology to automatically identify the characteristics of newly introduced data sets, the method of this embodiment can recommend appropriate classifications and labels, thereby reducing the degree of manual intervention. This not only improves the automation level of data governance but also reduces the risk of human errors.

[0086] Furthermore, for the complex relationships involved in the data governance process, this embodiment designs a unified knowledge representation model that can clearly express the association relationships between data, business, and management, enabling data governance personnel to better understand and analyze various problems and challenges in the data governance process, thereby enhancing the ability and level of data governance.

[0087] Furthermore, based on the Graph RAG technology, this embodiment realizes incremental knowledge graph construction, which can avoid reconstructing the entire knowledge graph, thereby maintaining the timeliness and accuracy of information, enabling data governance personnel to obtain the latest metadata and its relationship information in a timely manner, and providing strong support for data governance work.

[0088] Therefore, the method for constructing a data governance knowledge graph based on metadata proposed in this embodiment brings many beneficial effects, can significantly improve the efficiency and accuracy of data governance, and provides strong support for the digital transformation and business development of enterprises.

[0089] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered to be within the scope described in this specification.

[0090] The above embodiments are only the preferred embodiments of the present invention and cannot be used to limit the scope of protection of the present invention. Any non-substantial changes and substitutions made by those skilled in the art based on the present invention belong to the scope of protection required by the present invention.

Claims

1. A method for constructing a data governance knowledge graph based on metadata, characterized in that: The following steps are involved: Establish multiple collection task areas, and divide the collection tasks into metadata real-time collection tasks and business data real-time collection tasks; Define the metamodel structure of each type of metadata and the relationship between metamodels based on the collected metadata, summarize and abstract all metamodel structures, and form a unified metamodel structure after building corresponding relationships. This structure is used to accommodate information of all metamodel types. Real-time detection is performed through metadata collection tasks to dynamically adjust the unified metamodel structure; Construct a dynamic ontology model, map the unified metamodel structure to the dynamic ontology model, including mapping metadata types to classes, metadata to entities, metadata attribute sets to attributes, and metadata relationship sets to relationships, and adjust the dynamic ontology model according to the dynamic changes of the unified metamodel structure; Preprocess, extract knowledge, and fuse knowledge from the metadata lake; According to the corresponding mapping relationship between the unified metamodel structure and the dynamic ontology model, the structured data after knowledge fusion is mapped to the dynamic ontology model, and the converted data is stored in the graph database to complete the construction of the first generation of knowledge graph; After detecting and preprocessing the real-time data stream, the pre-trained Graph RAG model is connected to automatically extract entities, relationships and their attributes from the preprocessed data, and represent the processing in the form of a graph structure to form the updated content of the knowledge graph; the newly extracted entities and relationships are added to the existing first-generation knowledge graph, and the existing entity and relationship information is updated at the same time.

2. The method according to claim 1, characterized in that: For structured database data sources, the change data capture CDC framework is used to monitor the binlog logs of the database to capture change events in the database in real time, and the captured change events are parsed to obtain structured metadata or business data; For log file data sources, use the Logstash log collection tool for real-time collection. By configuring the Logstash input, filter, and output plug-ins, the real-time reading, processing, and forwarding of log files can be realized, and the data in the log files can be collected to the specified storage location. For text file data sources, use file monitoring tools to monitor the update status of text files in real time. When the file changes, the collection task is triggered to read and parse the updated file content to obtain the metadata or business data in the text file.

3. The method according to claim 1, characterized in that: The unified metamodel structure further defines the specific composition and hierarchical relationship of metadata types, including: Data modeling related entities, including database link entities, database instance entities, database entities, data table entities, field entities, view entities, function entities and data modeling rule entities, where the data modeling rule entities are used to specify the standards and specifications for database design, data table structure, field definition, view creation and function writing; Data integration related entities, including task entities and scheduling entities. Task entities are used to describe specific jobs or operations in the data integration process, while scheduling entities are used to define the order, time, and frequency of task execution. Data service related entities include system entities, page entities, API entities and data set entities. System entities represent the software system that provides data services, page entities refer to the user interface within the system, API entities represent the interface provided by the system to the outside world, and data set entities are the specific data sets of data services; Data management related entities include account entities, log entities, message entities and document entities. Account entities are used to manage user access rights, log entities are used to record system operation history, message entities are used for information transmission within the system or across systems, and document entities contain various types of documents in the data governance process.

4. The method according to claim 3, characterized in that: The defined unified metamodel structure is classified and labeled according to three classification methods: entity information, relationship information, and attribute information. Among them, entity information is used to regard each metadata item as a metadata entity, which includes data modeling related entities, data integration related entities, data service related entities, and data management related entities. Relationship information is used to define the association, dependency, or hierarchical relationship between various entities in the metamodel. Attribute information defines the various attributes of the metamodel entity itself. The classified and labeled metamodel structures, together with their corresponding label information, are stored in the metamodel structure library, and a unified metamodel structure framework is abstracted based on the classification and labeling results. The framework contains at least metamodel type, metamodel entity information, relationship information, and attribute information.

5. The method according to claim 4, characterized in that Dynamically adjust the unified metamodel structure, including: For structured databases, set up metadata collection tasks, detect log information in the database in real time, parse DDL statement information related to data structure changes in the log, and match the type of DDL operation objects with the metamodel type pre-stored in the metamodel structure library by identifying them; If the type matching is successful, the attributes of the operation object in the DDL statement are further parsed, and these attributes are matched one by one with the attributes of the corresponding metamodel in the metamodel structure library; If the attribute matching is successful, it indicates that the data structure changes caused by the DDL statement have been reflected in the metamodel structure library, and no processing will be performed; if the attribute matching is not successful, it indicates that the DDL statement introduces new attributes or changes existing attributes, then new attributes are added to the metamodel in the metamodel structure library, and this change information is synchronously updated to the unified metamodel structure.

6. The method according to claim 5, characterized in that: For newly generated data structures, natural language processing technology and machine learning algorithms are used to actively identify and classify them to determine the type of metamodel they belong to; the newly identified metamodels are automatically stored in the metamodel structure library and classified and labeled; For the newly identified metamodel, the conversion operation is automatically performed to convert it into a unified metamodel structure according to the metamodel type to which it belongs and the pre-defined standardized conversion rules; Among them, standardized transformation rules are used to describe how to transform different types of metamodels into a unified metadata model structure.

7. The method according to claim 1, characterized in that Preprocess the metadata in the metadata lake according to data standards and preset data quality rules, including: Data integrity verification: Perform integrity checks on metadata in the metadata lake according to data standards; Data uniqueness verification: When metadata has a business-unique key, uniqueness verification is performed. Specific indicators include the unique key conflict rate. Data accuracy verification: Verify the metadata generation trigger mechanism and attribute filling according to data standards, and its specific indicators include accuracy; Data consistency verification: Check the logical consistency between different attributes of business objects and the consistency of business process magnitude fluctuations, which is done through cross-table and value comparisons and fluctuation consistency comparisons; Data timeliness verification: Verify the generation time and arrival delay of metadata according to data standards; Data standardization / normalization: Standardize or normalize metadata to eliminate the impact of dimensions or values ​​on calculation results, including Z-Score standardization and Max-Min normalization; Data quality report generation: According to the definition of data quality rules, data quality reports are generated regularly to record the quality status of metadata in terms of completeness, uniqueness, accuracy, consistency and timeliness; Data quality problem handling: Based on the records in the data quality report, respond to, locate, analyze, and resolve metadata quality issues found to ensure that the metadata in the metadata lake continues to meet data standards and preset data quality rules.

8. The method according to claim 4, characterized in that The knowledge extraction process includes entity extraction, attribute extraction and relationship extraction, specifically including: Metamodel structure matching: Use the metamodel structure library to match the preprocessed metadata with the metamodel structure; directly identify and locate the entity and attribute information in the metadata based on the classification labels in the metamodel structure; Entity extraction: Apply natural language processing technology and machine learning algorithms to preprocess and extract features from the text content in metadata, and perform entity extraction, attribute extraction, and relationship extraction; Attribute extraction: construct an attribute list for each identified entity, apply attribute extraction algorithms to extract attribute information related to the entity from the metadata text, and assign attribute values ​​to these attributes; Relationship extraction: Based on entity extraction and attribute extraction, further identify and extract the relationship between entities; apply relationship extraction algorithms to analyze the context and semantics in metadata texts to determine the types of associations between entities; construct entity relationship graphs to intuitively display the complex relationship network between entities; Knowledge integration and storage: Integrate the results of entity extraction, attribute extraction and relationship extraction to form a structured knowledge representation, and store the structured knowledge in the knowledge base.

9. The method according to any one of claims 1 to 8, characterized in that: The knowledge extraction and fusion process specifically includes: Select the entities to be aligned and select the key attributes of each entity. The entity itself, its attributes and attribute values, and its relationship with other entities are respectively composed of three sets of features. Use vectorization technology to convert the three sets of features into vector representations, and merge these vectors into a unified representation in a low-dimensional vector space through splicing operations. In the constructed low-dimensional vector space, the semantic distance calculation method is used to measure the similarity between entities. A semantic distance threshold is set, and for entity pairs with a semantic distance less than the threshold, it is preliminarily determined that they may be aligned or equivalent entities. For entity pairs that are initially judged to be potentially aligned, attribute conflicts and relationship conflicts are detected and processed. By comparing whether the attribute values ​​of the entities are consistent or compatible, and whether the relationships between the entities match each other, it is determined whether these entities can indeed be merged into a unified entity. Among them, for entities that are confirmed to be mergeable, a merge operation is performed and integrated into a new entity representation.

10. The method according to any one of claims 1 to 8, characterized in that: The preprocessed new information is embedded into the graph structure of the constructed knowledge graph to form a graph structure fragment containing new entities, new relationships and their attributes. The embedding layer of the GraphRAG model is used to convert the new information and existing information into low-dimensional vector representations. Based on the reasoning ability of the GraphRAG model, the logical relationship and consistency between the new information and the existing information in the knowledge graph are checked; The attention mechanism of the GraphRAG model is used to focus on entities and relationships directly related to the new information, as well as the connections and constraints between them. The rationality and credibility of the new information are evaluated by comparing the distance and similarity between the new information and the existing information in the vector space. Based on the results of the logical relationship and consistency check, the GraphRAG model automatically infers whether the new information is consistent with the existing information in the knowledge graph, and whether there is any conflict or contradiction. The GraphRAG model outputs verification decisions, including accepting new information, rejecting new information, or marking it as information for further verification. For accepted new information, it is formally integrated into the knowledge graph and the structure and content of the graph are updated. For rejected or pending information, corresponding feedback is output.

Citation Information

Patent Citations

  • Fusion platform and fusion method of heterogeneous multi-source data

    CN107633075A

  • Autonomous data lake construction system and method based on associated data

    CN110941612A

  • Semantic-based data lake query system and method

    CN114218400A

  • Methods and systems for controlled modeling and optimization of a natural language database interface

    US20240184829A1

Cited By

  • Method for improving metadata storage and query efficiency

    CN120578797A

  • Data processing method and device, computer readable storage medium and electronic equipment

    CN120780749A

  • Multi-modal data retrieval method and device, storage medium and computer equipment

    CN121051286A

  • Zero-configuration knowledge graph construction method, equipment and medium

    CN121257685A

  • Zero-configuration knowledge graph construction method, device and medium

    CN121257685B