Standard knowledge graph database construction method for retrieval

By building a standard knowledge graph library, the problems of incomplete data and difficulty in updating are solved, the comprehensive collection and accuracy of data are achieved, and the timeliness and reliability of the knowledge graph library is ensured.

CN120235224APending Publication Date: 2025-07-01CHINA AEROSPACE STANDARDIZATION INST
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510188842.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-20
Publication Date
2025-07-01

AI Technical Summary

Technical Problem

The data in the existing standard knowledge graph library is not comprehensive enough, information updates are difficult to be timely, and there are problems of errors or inconsistencies.

Method used

By clarifying construction goals, collecting structured, semi-structured and unstructured data, pre-processing and modeling, detecting errors using automation tools, regularly updating knowledge graphs to ensure timeliness and accuracy.

Benefits of technology

It realizes the comprehensiveness and accuracy of data collection, ensures the timeliness and reliability of the knowledge graph library, and improves the speed and efficiency of information updates.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120235224A_ABST
    Figure CN120235224A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of retrieval, and particularly relates to a standard knowledge graph library construction method for retrieval, which comprises the following specific steps of: 1, firstly, determining a construction target of a knowledge graph, determining the field of a knowledge graph library, and determining an application scene and purpose of the knowledge graph; selecting a specific field or theme according to application requirements; secondly, the structured data, the semi-structured data, the unstructured data and the public data are collected, data collection can be carried out through the structured data, the semi-structured data, the unstructured data and the public data, the comprehensiveness of data collection is improved, quality control is carried out on the data after data collection, and the data collection efficiency is improved. According to the method, the accuracy of the data is ensured, the data is regularly and comprehensively reviewed, whether outdated information, wrong data or a new entity relationship needing to be added exists or not is checked, and the timeliness of the standard knowledge graph database is ensured.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of retrieval, and specifically provides a method for constructing a standard knowledge graph library for retrieval. Background Art

[0002] A standard knowledge graph library is a structured knowledge representation and storage system that organizes and manages data in the form of a graph, where nodes represent entities (such as people, places, events, etc.) and edges represent the relationships between these entities. Standard knowledge graph libraries generally follow some common specifications and standards to facilitate data sharing, interoperability, and scalability;

[0003] However, the following technical problems still exist when using existing standard knowledge graph libraries:

[0004] The data in existing standard knowledge graph libraries may not be comprehensive enough, and the information on certain fields or entities may be very limited. Even after strict quality control, there may still be incorrect data or inconsistent information. The information in the knowledge graph needs to be updated regularly to maintain timeliness, but it is often difficult to achieve timely updates in actual operations.

[0005] Therefore, a method for constructing a standard knowledge graph library for retrieval is proposed to solve the above-mentioned problems. Summary of the Invention

[0006] The purpose of the present invention is to provide a method for constructing a standard knowledge graph library for retrieval to solve the problems raised in the above background art.

[0007] To achieve the above purpose, the present invention provides the following technical solution: A method for constructing a standard knowledge graph library for retrieval, and the specific steps of this method for constructing a standard knowledge graph library for retrieval are as follows:

[0008] Step 1: First, clarify the construction goal of the knowledge graph, determine the field of the knowledge graph library, clarify the application scenario and purpose of the knowledge graph, and select a specific field or topic according to the application requirements;

[0009] Step 2: Collect structured data, semi-structured data, unstructured data, and public data respectively;

[0010] Step 3: Preprocess the collected data, remove duplicate items, handle outliers, and unify the format of these data; convert data from different sources into a consistent form, automatically extract entities in the text through named entity recognition technology, and link them to the existing knowledge base;

[0011] Step 4: Model the knowledge, define concepts and relationships within the field to form an ontology model, set relevant attributes for each entity, and create relationships for the logical connections between entities;

[0012] Step Five: Select a suitable storage solution to store this knowledge;

[0013] Step Six: Integrate the stored knowledge with each other, integrate information from different channels, solve the synonym problem, eliminate ambiguity, and ensure the consistency and accuracy of the information;

[0014] Step Seven: Conduct a quality assessment of the stored knowledge, regularly check and update the content of the knowledge graph to ensure its timeliness and reliability, and use automated tools to detect errors or incomplete information;

[0015] Step Eight: During the later retrieval process, maintain and regularly update the stored knowledge, continuously monitor the performance of the knowledge graph, promptly adjust strategies to adapt to changing requirements, regularly add new data, and maintain the timeliness of the knowledge graph.

[0016] Preferably, in Step Two, the structured data is sourced from databases and API interfaces. The databases include relational databases and NoSQL databases; the API interfaces include publicly available network data and enterprise data; the semi-structured data includes XML / JSON files and CSV / TSV files; the unstructured data includes text files, web content, social media, and multimedia content. The text files include news articles and research reports. The web content includes information crawled from web pages by web crawlers. The social media includes posts or comments on various platforms. The multimedia content includes pictures, videos, and audio.

[0017] Preferably, in Step Two, the tools for collecting data include web crawlers, API calls, database exports, text processing, and multimedia processing.

[0018] Preferably, in Step Three, when preprocessing the data, ensure that each entity and relationship appears only once to avoid redundancy. For missing data, methods such as deletion, filling, or interpolation can be used. Identify and handle outliers, for example, by statistical methods or machine learning methods to detect outliers and decide whether to delete or correct them; solve the problem of entities with the same name, ensure that each entity is unique and correctly linked to the corresponding node in the knowledge graph, and link the identified entities to the existing knowledge base to utilize existing resources; ensure the consistency of data between different sources and ensure that the data conforms to logical relationships.

[0019] Preferably, in the fourth step, it is necessary to clarify the field covered by the knowledge graph, define the core concepts in this field, establish the hierarchical relationships between the concepts, define the relevant attributes for each concept, and define the relationships between the concepts; use the Resource Description Framework to represent the data, each data item consists of a triple, use the Web Ontology Language to represent more complex semantic relationships, support reasoning and logical expressions, use a graph database to store and query the data, which is suitable for processing complex relationship networks, and ensure that each entity has a unique identifier to avoid duplication and confusion.

[0020] Preferably, in the fifth step, the storage scheme includes storing in the form of RDF triples and a graph database.

[0021] Preferably, in the eighth step, set periodic time points to comprehensively review the data in the knowledge graph, check whether there is outdated information, incorrect data, or new entity relationships that need to be added. This is assisted by automated tools, but ultimately may still require manual intervention to ensure accuracy. Automatically extract structured data from text materials through NLP technology and integrate it into the existing knowledge graph, greatly improving the speed and efficiency of information update. If there are other professional databases with a good reputation already established, establish an interface with them to achieve data sharing and synchronization.

[0022] Compared with the prior art, the beneficial effects of the present invention are:

[0023] This application can collect data through structured data, semi-structured data, unstructured data, and public data, increasing the comprehensiveness of data collection. After data collection, it conducts quality control on the data to ensure data accuracy, and regularly conducts a comprehensive review of the data to check for outdated information, incorrect data, or new entity relationships that need to be added, ensuring the timeliness of the standard knowledge graph library. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] Figure 1 It is a flowchart of the steps of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0025] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0026] In the description of the present invention, it should be understood that the orientation or positional relationship indicated by the terms "upper", "lower", "front", "rear", "left", "right", "top", "bottom", "inner", "outer", etc. is based on the orientation or positional relationship shown in the drawings. It is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation. Therefore, it should not be construed as a limitation to the present invention.

[0027] Embodiment:

[0028] A method for constructing a standard knowledge graph library for retrieval, and the specific steps of the method for constructing the standard knowledge graph library for retrieval are as follows:

[0029] Step 1: First, clarify the construction objective of the knowledge graph, determine the field of the knowledge graph library, clarify the application scenario and purpose of the knowledge graph, and select a specific field or topic according to the application requirements;

[0030] Step 2: Collect structured data, semi-structured data, unstructured data, and public data respectively;

[0031] Step 3: Preprocess the collected data, remove duplicate items, handle outliers, and unify the format of these data; convert data from different sources into a consistent form, automatically extract entities in the text through named entity recognition technology, and link them to the existing knowledge base;

[0032] Step 4: Model the knowledge, define concepts and relationships within the field to form an ontology model, set relevant attributes for each entity, and create relationships for the logical connections between entities;

[0033] Step 5: Select a suitable storage scheme to store this knowledge;

[0034] Step 6: Integrate the stored knowledge with each other, integrate information from different channels, solve the problem of synonyms, eliminate ambiguity, and ensure the consistency and accuracy of the information;

[0035] Step 7: Evaluate the quality of the stored knowledge, regularly check and update the content of the knowledge graph to ensure its timeliness and reliability, and use automated tools to detect errors or incomplete information;

[0036] When conducting a quality assessment of stored knowledge, it is necessary to ensure that all necessary fields have values without omission, ensure that each entity and relationship is complete and reasonable, check the consistency of data for the same entity in different places, compare with external authoritative data sources to ensure data consistency; use known correct datasets for verification to ensure data accuracy, check whether the logical relationships between data are correct, check the timestamps of data to ensure that the data is up-to-date, evaluate the credibility of data sources, give priority to authoritative and reliable sources, and record the data sources, processing procedures, and change histories for easy traceability and auditing.

[0037] Step Eight: During the later retrieval and use process, maintain and update the stored knowledge regularly, continuously monitor the performance of the knowledge graph, adjust strategies in a timely manner to adapt to changing requirements, add new data regularly to keep the knowledge graph up-to-date.

[0038] In Step Two, the structured data comes from databases and API interfaces. The databases include relational databases and NoSQL databases; the API interfaces include publicly available network data and enterprise data; the semi-structured data includes XML / JSON files and CSV / TSV files; the unstructured data includes text files, web page content, social media, and multimedia content. The text files include news articles and research reports. The web page content includes information crawled from web pages by web crawlers. The social media includes posts or comments on various platforms. The multimedia content includes pictures, videos, and audio.

[0039] In Step Two, the tools for data collection include web crawlers, API calls, database exports, text processing, and multimedia processing.

[0040] In Step Three, when preprocessing the data, ensure that each entity and relationship appears only once to avoid redundancy. For missing data, methods such as deletion, filling, or interpolation can be used. Identify and handle outliers, for example, detect outliers through statistical methods or machine learning methods and decide whether to delete or correct them; solve the problem of entities with the same name, ensure that each entity is unique and correctly linked to the corresponding node in the knowledge graph, and link the identified entities to the existing knowledge base to utilize existing resources; ensure the consistency of data between different sources and ensure that the data conforms to logical relationships.

[0041] In step 4, it is necessary to clarify the fields covered by the knowledge graph, define the core concepts in this field, establish the hierarchical relationships between concepts, define the relevant attributes for each concept, and define the relationships between concepts; use the Resource Description Framework to represent data, with each data item consisting of triples, use the Web Ontology Language to represent more complex semantic relationships, support reasoning and logical expressions, use a graph database to store and query data, which is suitable for processing complex relationship networks, and ensure that each entity has a unique identifier to avoid duplication and confusion.

[0042] In step 5, the storage scheme includes storing in the form of RDF triples and a graph database. The advantages of storing in the form of RDF triples follow the W3C standard, facilitating interoperability with other RDF data sources, supporting the SPARQL query language, being powerful and flexible, and suitable for large-scale data storage and processing; the advantages of the graph database are that it optimizes query performance for graph structures, especially performing excellently in complex association queries, is easy to represent and query complex many-to-many relationships, and many graph databases support horizontal expansion and can handle large-scale data sets.

[0043] In step 8, set periodic time points to comprehensively review the data in the knowledge graph, check for outdated information, incorrect data, or new entity relationships that need to be added. This is assisted by automated tools but may ultimately still require manual intervention to ensure accuracy. Automatically extract structured data from text materials through NLP technology and integrate it into the existing knowledge graph, greatly improving the speed and efficiency of information update. If there are other well-established professional databases, establish an interface with them to achieve data sharing and synchronization.

[0044] The above shows and describes the basic principles, main features, and advantages of the present invention. For those skilled in the art, it is obvious that the present invention is not limited to the details of the above exemplary embodiments, and can be implemented in other specific forms without departing from the spirit or basic features of the present invention; therefore, from any point of view, the embodiments should be regarded as exemplary and non-restrictive. The scope of the present invention is defined by the appended claims rather than the above description. Therefore, all changes falling within the meaning and scope of the equivalent elements of the claims are intended to be encompassed within the present invention, and any reference signs in the claims should not be regarded as limiting the claims involved.

[0045] Although the embodiments of the present invention have been shown and described, for those of ordinary skill in the art, it can be understood that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.

Claims

1. A method for constructing a standard knowledge graph library for retrieval, characterized in that: The specific steps of the standard knowledge graph library construction method for retrieval are as follows: Step 1: First, clarify the construction goal of the knowledge graph, determine the field of the knowledge graph library, clarify the application scenarios and purposes of the knowledge graph, and select specific fields or topics based on application requirements; Step 2: Collect structured data, semi-structured data, unstructured data and public data respectively; Step 3: Preprocess the collected data to remove duplicates, process outliers, and unify the format; convert data from different sources into a consistent format, automatically extract entities from the text through named entity recognition technology, and link them to the existing knowledge base; Step 4: Model the knowledge, define the concepts and relationships within the domain, form an ontology model, set relevant attributes for each entity, and create relationships for the logical connections between entities; Step 5: Select a suitable storage solution to store this knowledge; Step 6: Integrate these stored knowledge with each other, integrate information from different channels, solve synonym problems, eliminate ambiguity, and ensure the consistency and accuracy of information; Step 7: Evaluate the quality of stored knowledge, regularly check and update the content of the knowledge graph to ensure its timeliness and reliability, and use automated tools to detect errors or incomplete information; Step 8: During later retrieval, the stored knowledge should be maintained and regularly updated, the performance of the knowledge graph should be continuously monitored, strategies should be adjusted in a timely manner to adapt to changing needs, and new data should be added regularly to maintain the timeliness of the knowledge graph.

2. A method for constructing a standard knowledge graph library for retrieval according to claim 1, characterized in that: In the step 2, the structured data comes from a database and an API interface, the database includes a relational database and a NoSQL database; the API interface includes online public data and enterprise data; the semi-structured data includes XML / JSON files and CSV / TSV files; the unstructured data includes text files, web page content, social media and multimedia content, the text files include news articles and research reports, the web page content includes information captured on web pages by web crawlers, the social media includes posts or comments on various platforms, and the multimedia content includes pictures, videos, and audio.

3. The method for constructing a standard knowledge graph library for retrieval according to claim 1, characterized in that: In step 2, the tools for collecting data include web crawlers, API calls, database exports, text processing, and multimedia processing.

4. The method for constructing a standard knowledge graph library for retrieval according to claim 1, characterized in that: In step 3, when preprocessing the data, ensure that each entity and relationship appears only once to avoid redundancy. For missing data, methods such as deletion, filling or interpolation can be used to identify and process outliers. For example, statistical methods or machine learning methods are used to detect outliers and decide whether to delete or correct them. Solve the problem of entities with the same name, ensure that each entity is uniquely and correctly linked to the corresponding node in the knowledge graph, and link the identified entities to the existing knowledge base to utilize existing resources; Ensure data consistency across different sources and ensure data is logically related.

5. The method for constructing a standard knowledge graph library for retrieval according to claim 1, characterized in that: In step 4, it is necessary to clarify the fields covered by the knowledge graph, define the core concepts of the field, establish a hierarchical relationship between concepts, define relevant attributes for each concept, and define the relationship between concepts; use the resource description framework to represent data, each data item consists of a triple, use the Web ontology language to represent more complex semantic relationships, support reasoning and logical expression, use a graph database to store and query data, which is suitable for processing complex relational networks, and ensure that each entity has a unique identifier to avoid duplication and confusion.

6. The method for constructing a standard knowledge graph library for retrieval according to claim 1, characterized in that: In the step 5, the storage solution includes RDF triple form storage and graph database.

7. The method for constructing a standard knowledge graph library for retrieval according to claim 1, characterized in that: In step eight, periodic time points are set to conduct a comprehensive review of the data in the knowledge graph to check whether there is outdated information, erroneous data, or new entity relationships that need to be added. This is completed with the assistance of automated tools, but human intervention may still be required in the end to ensure accuracy. Structured data is automatically extracted from text materials through NLP technology and integrated into the existing knowledge graph, greatly improving the speed and efficiency of information updates. If there are other professional databases that have established a good reputation, interfaces can be established with them to achieve data sharing and synchronization.