A fully automatic knowledge graph construction method, system, electronic device and storage medium based on a large model
By automatically identifying entity information and generating query statements through large-scale language models, the problems of traditional knowledge graph construction being time-consuming, labor-intensive and user-unfriendly are solved, efficient and automated knowledge graph construction and query are achieved, and the user-friendliness and flexibility of the system are improved.
Patent Information
- Application Number
- CN202410209268.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-02-26
- Publication Date
- 2025-09-12
- Estimated Expiration
- 2044-02-26
AI Technical Summary
Traditional knowledge graph construction methods rely on manual annotation and rule definition, which is time-consuming and labor-intensive, difficult to adapt to diverse needs, has high update costs and poor user-friendliness, and the query language is complex, which limits the rapid construction and expansion of knowledge graphs.
Use large language models to automatically identify entity information, build knowledge graphs, understand user query intent through large language models, generate query statements, reduce manual intervention, and improve automation and user-friendliness.
It realizes efficient and automated knowledge graph construction and query, reduces dependence on manual annotation, improves construction efficiency and user interaction experience, and enhances the flexibility and real-time performance of the knowledge graph.
Smart Images

Figure CN118278508B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of knowledge graph technology, and in particular to a method, system, electronic device and storage medium for fully automatic knowledge graph construction based on a large model. Background Art
[0002] Traditional graph construction suffers from numerous drawbacks, primarily: Traditional knowledge graph construction methods often rely on manual annotation and rule definition. Specifically, current graph schemas rely on the development of business experts, who struggle to fully understand all relationships and entity types in large, complex domains. Graph instance construction relies on extensive manual effort and the involvement of domain experts to manually annotate and define rules. This process is not only time-consuming and labor-intensive, but also difficult to adapt to the diverse needs of different domains, limiting the rapid construction and expansion of knowledge graphs. Traditional methods also suffer from high costs and low efficiency when it comes to updating knowledge graphs. Because information in knowledge graphs needs to be dynamically updated, traditional methods often require frequent manual data updates, including the addition of new events, relationships, or entities, as well as revisions to existing information. This traditional manual data maintenance and update process is time-consuming and error-prone, impacting the real-time, flexibility, and scalability of knowledge graphs. Furthermore, traditional graph database query languages (such as Cypher) often have complex syntax and structure, creating a high learning curve for novices. Users must spend time learning the query language's syntax rules and features, limiting the ability of ordinary users and non-experts to use graph databases and the flexibility of query logic. Summary of the Invention
[0003] In response to the shortcomings of the above problems, the present invention provides a method, system, electronic device and storage medium for constructing a knowledge graph based on a large model and fully automatic.
[0004] To achieve the above objectives, the present invention provides a fully automatic knowledge graph construction method based on a large model, comprising:
[0005] Obtain corresponding data samples based on business type;
[0006] Preprocessing the data sample;
[0007] Inputting the pre-processed data sample into a large language model to identify entity information, wherein the entity information includes entities, entity relationships, and entity attributes;
[0008] Based on the entity information, a knowledge graph is constructed;
[0009] The knowledge graph is automatically recalled based on the large language model.
[0010] Preferably, the data sample includes structured data and unstructured data.
[0011] Preferably, preprocessing the data sample includes:
[0012] Removing HTML tags and special formatting from the data sample and removing duplicate content;
[0013] The data samples are standardized.
[0014] Preferably, the large language model is selected based on the business type.
[0015] Preferably, the entities include nouns, the entity relationships include terms indicating the belonging relationships between the entities, and the entity attributes include terms describing the characteristics or status of the entities.
[0016] Preferably, the entity information is obtained, similar entity information is merged, and entities with an occurrence probability lower than a threshold are removed.
[0017] The present invention also provides a fully automatic knowledge graph construction system based on a large model, including:
[0018] An acquisition module is used to obtain corresponding data samples based on business types;
[0019] A processing module, configured to pre-process the data sample;
[0020] A recognition module, configured to input the pre-processed data sample into a large language model to recognize entity information, wherein the entity information includes entities, entity relationships, and entity attributes;
[0021] A construction module, used to construct a knowledge graph based on the entity information;
[0022] An application module is used to automatically recall the knowledge graph based on the large language model.
[0023] The present invention also provides an electronic device, comprising at least one processing unit and at least one storage unit, wherein the storage unit stores a computer program, and when the program is executed by the processing unit, the processing unit executes the above method.
[0024] The present invention also provides a storage medium storing a computer program executable by a device, wherein when the program runs on the device, the device executes the above method.
[0025] Compared with the prior art, the present invention has the following beneficial effects:
[0026] The present invention uses a large language model to infer entity information from text data and construct a highly intelligent knowledge graph without relying too much on manual annotation and rule definition, thereby improving the efficiency and automation of knowledge graph construction. The large language model learns the semantic information in the context to more accurately identify entity information, that is, it automatically extracts entities, entity relationships and entity attributes from the text, reducing dependence on manual annotation; in addition, by understanding the natural language of the user's query intention, it generates corresponding graph query statements, thereby realizing automatic recall of relevant graph data. This function greatly simplifies the process of user interaction with the knowledge graph and improves the user-friendliness of the system. BRIEF DESCRIPTION OF THE DRAWINGS
[0027] Figure 1 This is a flow chart of the fully automatic knowledge graph construction method based on a large model of the present invention. DETAILED DESCRIPTION
[0028] To make the objectives, technical solutions, and advantages of the embodiments of the present invention more clear, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts shall fall within the scope of protection of the present invention.
[0029] Reference Figure 1 The present invention provides a method for constructing a knowledge graph based on a large model and fully automatic, comprising:
[0030] Obtain corresponding data samples based on business type;
[0031] Specifically, data samples include structured data and unstructured data. Structured data is data organized according to a predetermined model, with elements having clearly defined relationships. This data is typically stored in a table format, consisting of rows and columns, with each column containing a specific type of data. Unstructured data is data without a fixed data model or predetermined structure. This type of data, such as text data, is not easily stored in tables or databases.
[0032] Preprocess the data samples;
[0033] Specifically, remove content containing HTML tags and special formats from the data sample, as well as remove duplicate content. Delete duplicate data to ensure that there is no repeated counting or duplicate information during the analysis process. If the text data contains HTML tags or other special formats, these contents need to be removed to retain the plain text information. Standardize the data to ensure data consistency. This may include operations such as unit conversion and date format unification to make the data more comparable, delete special characters, punctuation marks and other non-alphanumeric symbols in the text to retain the substantive information in the text, and remove common stop words, which appear frequently in the text but usually do not carry useful information. Stop words usually include "ah", "um", etc. During the data conversion and processing process, keep a backup of the original data to prevent information loss. This helps to trace back to the original data in subsequent analysis.
[0034] Input the preprocessed data samples into a large language model to identify entity information, which includes entities, entity relationships, and entity attributes;
[0035] Specifically, large language models are selected based on the business type, and users may choose large language models based on the performance requirements of their applications. Different models may differ in processing speed, accuracy, and resource consumption. Users may choose the model that best meets their performance standards based on their specific needs. Different models may focus on processing different types of tasks, such as natural language processing (NLP) tasks, computer vision tasks, or recommendation system tasks. Users will choose the most appropriate model based on the task type of their application. The use of the model may be limited by computing resources. Users may choose an appropriate model based on their available hardware resources (such as CPU, GPU) to ensure the best performance under given hardware conditions.
[0036] In this example, a visual configuration and result preview interface is provided, allowing users to easily configure and debug the prompts and parameters of large models to meet the needs of specific scenarios. Prompt setting: allows users to set the prompt of the model to standardize the model's extraction specifications for knowledge graph entities, relationships and events, and the granularity of knowledge graph construction. Model parameter setting: the temperature coefficient of the model, that is, the degree of divergence of the schema or triples extracted by the model. Result preview: provides a preview interface for model debugging, showing the model's response to the input prompt and the extraction results. Users can intuitively view the output of the model to check whether the extracted entities and relationships meet expectations. Allow users to switch the model operation mode according to actual needs, supporting both online real-time application and offline batch application modes.
[0037] Online Application: For scenarios with high real-time requirements, models must be launched online before model-level applications can be initiated. This feature provides a quick model launch capability. Users can trigger model launch with a single click, and the system automatically calls underlying computing resources to launch the corresponding online service. After the service is successfully launched, users can directly enter a text message to test the online service. The online service interface can also be directly imported into business systems for application.
[0038] Offline Applications: Allows users to set offline batch generation tasks to meet business needs that require regular updates or large-scale data processing. Task scheduling allows users to set parameters such as the time and frequency of generation tasks. This optimizes the task execution process to ensure that offline generation tasks are executed efficiently and as planned.
[0039] When model performance or operational status anomalies occur, administrators or relevant personnel are promptly notified. Detailed error logs and diagnostic information are provided to help users quickly locate and resolve issues. The system also monitors model performance in real time, including response time, memory usage, and CPU utilization. Performance reports are generated, displaying model performance metrics in charts or tables to help users understand the model's operational status.
[0040] Specifically, large-scale language models are used to identify entity information in the data. Entity information includes entities, entity relationships, and entity attributes. Entities include nouns. For example, entities include persons and locations. Persons identify entities containing names, titles, and identities within text. Locations identify geographic locations within text, including cities and countries. That is, people such as politicians, scientists, artists, etc., and places such as cities, countries, landmarks, etc.; use large language models for named entity recognition (NER) to identify entities in the text, and use the model to infer the identified entities to determine their specific types; entity relationships include words that represent the relationship between entities; for example, employment relationship, location relationship, and founding relationship; employment relationship is to identify the employment relationship between a person and an organization, location relationship is to identify the location relationship between a person and a place, and founding relationship is to identify the founding relationship between an organization and a person, that is, use the semantic understanding ability of the model to extract relationships, learn the relationship between entity pairs through training data, and use contextual information and keywords to infer the specific type of relationship; entity attributes include words that describe the characteristics or status of the entity, that is, the model deeply analyzes the text, understands the semantic information of the entity, extracts attributes and understands the contextual meaning of the attributes, and uses contextual context and common patterns to infer the specific type of attributes.
[0041] Build a knowledge graph based on entity information;
[0042] Automatically recall the knowledge graph based on a large language model.
[0043] In this embodiment, after obtaining entity information, similar entity information is merged and entity types with an occurrence probability below a threshold are removed. This setting ensures the simplicity and effectiveness of the graph.
[0044] Specifically, the generated graph schema is reviewed and managed through a visual interface to discover possible errors, misjudgments or inconsistencies, thereby ensuring the quality of the knowledge graph. At the same time, users are allowed to customize based on business needs. Utilizing the entity, attribute and relationship information in the generated graph schema, combined with the language understanding capabilities of the large model, specific instances of entities, attribute values, and relationships between entities are extracted from the corresponding data source. Through prompt engineering, appropriate natural language descriptions are designed to enable the large model to understand the context of the entities, attributes, and relationships that need to be extracted, and generate corresponding triples. Using the generated graph, semantic reasoning is performed through the large model to discover potential associations between entities and relationships. Utilizing the contextual understanding capabilities of the large model, logical associations between entities are inferred, and more comprehensive triple information is generated. Based on the extraction results of entity, relationship, and attribute information, knowledge graph instances are automatically generated.
[0045] Suppose there is a movie graph that contains nodes such as actors, movies, and directors, as well as the relationships between them. Use Cypher query language to query movies that Liu xx has starred in. The query process is as follows:
[0046] User natural language query: "What TV series has Liu xx acted in?"
[0047] Semantic understanding: The large model performs semantic understanding on user queries and identifies that the user's main intention is to query TV series in which Liu xx has participated, as well as other actors and directors related to her.
[0048] Query statement generation
[0049] By leveraging the semantic understanding capabilities of the large model, the natural language queries provided by users are converted into graph query statements, and corresponding Cypher query statements are generated.
[0050] Graph data recall
[0051] The generated Cypher query statement will be sent to the graph database to execute the query operation. The database will recall relevant data from the graph based on the query statement, including entity, relationship and attribute information.
[0052] Graph Data Traceability
[0053] Provide original information traceability for recalled entities, relationships and other data to help users understand the credibility of the data.
[0054] Related knowledge recommendation
[0055] By discovering related entities and relationships through graph algorithms, the system can recommend information related to the user's query topic but not explicitly requested based on the user's question, thereby providing more comprehensive and in-depth knowledge.
[0056] User feedback mechanism
[0057] It supports users to provide feedback and evaluate whether they are "satisfied" or "unsatisfied" on the graph information retrieved from the application system, forming a feedback and reflow mechanism for graph data, and continuously improving and adjusting the query and return results of the knowledge graph.
[0058] The present invention also provides a fully automatic knowledge graph construction system based on a large model, including:
[0059] An acquisition module is used to obtain corresponding data samples based on business types;
[0060] A processing module, used for preprocessing data samples;
[0061] The recognition module is used to input the preprocessed data samples into the large language model to identify entity information, which includes entities, entity relationships and entity attributes;
[0062] Construction module, used to build knowledge graph based on entity information;
[0063] Application module for automatically recalling knowledge graphs based on large language models.
[0064] The present invention also provides an electronic device, comprising at least one processing unit and at least one storage unit, wherein the storage unit stores a computer program, and when the program is executed by the processing unit, the processing unit executes the above method.
[0065] The present invention also provides a storage medium storing a computer program executable by a device, which, when executed on the device, causes the device to execute the above method.
[0066] The above are merely preferred embodiments of the present invention and are not intended to limit the present invention. Those skilled in the art will readily appreciate that various modifications and variations of the present invention are possible. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of the present invention shall be included within the scope of protection of the present invention.
Claims
1. A fully automatic knowledge graph construction method based on a large model, characterized by: include: Obtain corresponding data samples based on business type; Preprocessing the data sample; Inputting the pre-processed data sample into a large language model to identify entity information, wherein the entity information includes entities, entity relationships, and entity attributes; Based on the entity information, a knowledge graph is constructed; Automatically recalling the knowledge graph based on the large language model; The large language model is selected based on the business type; the entities include nouns, and the model is used to infer the identified entities to determine their specific types; the entity relationships include the terms of the belonging relationships between the entities, and the model's semantic understanding ability is used to extract the relationships, learn the relationships between entity pairs through training data, and use contextual information and keywords to infer the specific types of the relationships; the entity attributes include words that describe the characteristics or status of the entity, and the model deeply analyzes the text to understand the semantic information of the entity, extracts the attributes and understands the contextual meaning of the attributes, and uses contextual context and common patterns to infer the specific types of the attributes; Utilizing the entity, attribute, and relationship information in the generated graph schema, combined with the language understanding capabilities of the large model, specific instances of entities, attribute values, and relationships between entities are extracted from the corresponding data source. Through prompt engineering, appropriate natural language descriptions are designed to enable the large model to understand the context of the entities, attributes, and relationships to be extracted and generate corresponding triples. Using the generated graph, semantic reasoning is performed through the large model to discover potential associations between entities and relationships. The large model's contextual understanding capabilities are used to infer logical associations between entities and generate more comprehensive triple information. Based on the extracted entity, relationship, and attribute information, knowledge graph instances are automatically generated. Acquire the entity information, merge similar entity information, and remove entities whose occurrence probability is lower than a threshold; The data sample includes structured data and unstructured data; Preprocessing the data sample includes: Removing HTML tags and special formatting from the data sample and removing duplicate content; performing standardization processing on the data sample; The database recalls relevant data from the graph based on the query statement, including entity, relationship and attribute information; the graph data traceability provides original information traceability for the recalled data.
2. A fully automatic knowledge graph construction system based on a large model, characterized by: include: An acquisition module is used to obtain corresponding data samples based on business types; A processing module, configured to pre-process the data sample; A recognition module, configured to input the pre-processed data sample into a large language model to recognize entity information, wherein the entity information includes entities, entity relationships, and entity attributes; A construction module, used to construct a knowledge graph based on the entity information; An application module, configured to automatically recall the knowledge graph based on the large language model; Among them, the large language model is selected based on the business type; the entities include nouns, and the model is used to infer the identified entities to determine their specific types; the entity relationships include the terms of the belonging relationships between the entities, and the semantic understanding ability of the model is used to extract the relationships, and the relationship between entity pairs is learned through training data, and the specific type of the relationship is inferred using context information and keywords; the entity attributes include words that describe the characteristics or status of the entity, and the model deeply analyzes the text to understand the semantic information of the entity, extracts attributes and understands the contextual meaning of the attributes, and uses context context and common patterns to infer the specific type of the attributes; the entity, attribute and relationship information in the generated graph schema are used, combined with the language understanding ability of the large model, to extract the specific instances of entities, attribute values, and relationships between entities from the corresponding data sources; through Through prompt engineering, appropriate natural language descriptions are designed to enable the big model to understand the context of entities, attributes and relationships that need to be extracted, and generate corresponding triples; using the generated graph, semantic reasoning is performed through the big model to discover potential associations between entities and relationships; using the context understanding ability of the big model, the logical associations between entities are inferred to generate more comprehensive triple information, and based on the extraction results of entity, relationship and attribute information, knowledge graph instances are automatically generated; the entity information is obtained, similar entity information is merged, and entities with a probability of occurrence below a threshold are removed; the data sample includes structured data and unstructured data; preprocessing the data sample includes: removing content containing HTML tags and special formats in the data sample and removing duplicate content; and standardizing the data sample; The database recalls relevant data from the graph based on the query statement, including entity, relationship and attribute information; the graph data traceability provides original information traceability for the recalled data.
3. An electronic device, characterized in that: The computer comprises at least one processing unit and at least one storage unit, wherein the storage unit stores a computer program, and when the program is executed by the processing unit, the processing unit executes the method according to claim 1.
4. A storage medium, characterized in that The computer program that can be executed by a device is stored therein, and when the program is run on the device, the device is caused to execute the method according to claim 1 .
Citation Information
Patent Citations
Knowledge graph retrieval method and system fusing pre-training language model
CN117555985A