AI-Based Intelligent Knowledge Graph Construction Method

The method addresses polysemous word ambiguity in knowledge graphs by determining their meanings through a multi-scoring mechanism and comprehensive evaluation, enhancing accuracy and reliability.

CN119129727BActive Publication Date: 2025-07-15SHANDONG UNIV OF FINANCE & ECONOMICS
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202411595065.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-11
Publication Date
2025-07-15
Estimated Expiration
2044-11-11

AI Technical Summary

Technical Problem

In the prior art, entity recognition and relationship extraction are inaccurate due to uncertain meaning of polysemes, which affects the quality and usability of the knowledge graph.

Method used

Multiple scoring mechanism is used to determine the meaning of polysimilars, and a high-quality knowledge graph is constructed through context consistency, commonness and emotional consistency scoring, combined with data cleaning and integrity evaluation.

Benefits of technology

It effectively reduces the ambiguity caused by polysynonyms, improves the accuracy and usability of the knowledge graph, enhances the query and reasoning capabilities of the knowledge graph, and improves maintenance efficiency and update speed.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119129727B_ABST
    Figure CN119129727B_ABST
Patent Text Reader

Abstract

The present invention belongs to the technical field of knowledge graph construction, and particularly relates to an intelligent knowledge graph construction method based on AI. The steps include obtaining a target subject domain, collecting knowledge data from a data source according to the target subject domain; performing data cleaning on the collected knowledge data to obtain the cleaned knowledge data; performing data extraction on the cleaned knowledge data to obtain knowledge extraction data; extracting polysemous words in the knowledge extraction data, adopting a multiple scoring mechanism for the polysemous words to obtain the meanings of the polysemous words; performing knowledge representation and storage according to the knowledge extraction data and the meanings of the polysemous words; performing integrity evaluation on the constructed knowledge graph to obtain an overall integrity index, and the overall integrity index is used to measure the integrity of the knowledge graph; verifying and optimizing the constructed knowledge graph according to the overall integrity index. The present invention effectively prevents the situation that paragraphs are ambiguous due to uncertain meanings of polysemous words, and enhances the accuracy of the knowledge graph.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of knowledge graph construction, and particularly relates to an intelligent knowledge graph construction method based on AI. Background Art

[0002] Knowledge graph construction refers to the process of integrating scattered and heterogeneous data sources, and automatically identifying and extracting entities, relationships, and attributes through technologies such as natural language processing and machine learning to form a structured semantic knowledge base. This knowledge base is represented in the form of a graph, consisting of nodes (representing entities) and edges (representing relationships between entities), and can provide rich background knowledge for machine learning models, enhancing their understanding and reasoning abilities. Knowledge graphs are widely used in fields such as search engines, recommendation systems, intelligent question answering, and the education industry. Nowadays, knowledge graphs have become a key technology for organizing and understanding massive amounts of information, and how to obtain more accurate knowledge graphs has attracted widespread attention.

[0003] In the prior art, CN114328951B discloses a knowledge graph construction method that integrates information acquisition and triple extraction, including the following steps: S1: Regularly use web crawler technology to crawl text content related to the ocean, including news, from specified web pages; S2: Use natural language processing tools to perform entity extraction and relationship extraction on the text content to obtain the triples of the news, and then store the triples of the news in a database; S3: Construct a knowledge graph based on the triples in the database and implement the visualization of the knowledge graph in a data browser; S4: Obtain the associations of knowledge based on the visualized knowledge graph.

[0004] CN115017335A discloses a knowledge graph construction method and system, including the following steps: Set a new word discovery algorithm, and organize proper nouns and open knowledge graphs as databases for word segmentation recognition; Obtain the input text, and extract triples including subjects, predicates, and objects in the text as knowledge extraction results based on the database and a word segmentation extractor; Query triples related to the corresponding nodes of the extracted triples in the open knowledge graph, and form a new knowledge graph with all the triples.

[0005] However, the prior art and the above-mentioned existing patents still have some problems, mainly manifested in that many words and phrases have multiple meanings, and the ambiguity caused by this multiple meaning will lead to inaccurate entity recognition and relationship extraction, thereby affecting the quality and usability of the knowledge graph, the user's trust in the data, and the application performance based on the knowledge graph. Summary of the Invention

[0006] In view of the deficiencies of the above prior art, the purpose of the present invention is to provide an AI-based intelligent knowledge graph construction method, which effectively prevents the ambiguity of paragraphs caused by the uncertain meaning of polysemous words and enhances the accuracy of the knowledge graph.

[0007] To achieve the above purpose, the technical solution adopted by the present invention is an AI-based intelligent knowledge graph construction method, including the following steps:

[0008] Obtain the target subject domain, and collect knowledge data from the data source according to the target subject domain;

[0009] Clean the collected knowledge data to obtain the cleaned knowledge data;

[0010] Extract the data from the cleaned knowledge data to obtain the knowledge extraction data;

[0011] Extract the polysemous words in the knowledge extraction data, adopt a multiple scoring mechanism for the polysemous words, determine the meaning of the polysemous words through the multiple scoring mechanism, and obtain the meaning of the polysemous words;

[0012] According to the knowledge extraction data and the meaning of the polysemous words, perform knowledge representation and storage;

[0013] Evaluate the integrity of the constructed knowledge graph to obtain an overall integrity index, which is used to measure the integrity of the knowledge graph;

[0014] Verify and optimize the constructed knowledge graph according to the overall integrity index.

[0015] Preferably, the collection of knowledge data from the data source according to the target subject domain includes the following steps:

[0016] Obtain the target subject domain covered by the knowledge graph;

[0017] Connect to the data source and identify the data source type;

[0018] Determine the data collection tool according to the data source type, and collect and extract knowledge data based on the target subject domain using data mining algorithms and text analysis techniques.

[0019] Preferably, the cleaning of the collected knowledge data to obtain the cleaned knowledge data and the extraction of the data from the cleaned knowledge data to obtain the knowledge extraction data include the following steps:

[0020] Remove duplicate records through a unique identifier, and unify data in different formats into a standard format;

[0021] Through the collation of different data sources, merge relevant data, remove redundant data, and obtain the cleaned knowledge data;

[0022] Based on the cleaned knowledge data, entity extraction, relationship extraction, and attribute extraction are performed through data extraction rules to obtain knowledge extraction data.

[0023] Preferably, the polysemous words in the extracted knowledge extraction data are selected, and a multiple scoring mechanism is adopted for the polysemous words. The meaning of the polysemous words is determined through the multiple scoring mechanism, and the meaning of the polysemous words is obtained, including the following steps:

[0024] Based on the polysemous words in the knowledge extraction data, the meanings 1, 2,..., N of the polysemous words are extracted, where N is the meaning number;

[0025] The meanings 1, 2,..., N of the polysemous words are scored through the multiple scoring mechanism, and based on the scoring results, the meaning of the polysemous words is obtained.

[0026] Preferably, the multiple scoring mechanism includes the following steps:

[0027] Obtain the meanings 1, 2,..., N of the polysemous words;

[0028] Substitute the meanings 1, 2,..., N into the original paragraph where the polysemous word is located one by one to obtain paragraphs 1, 2,..., G, where G is the paragraph number. The original paragraph is scored for context consistency with paragraphs 1, 2,..., G in turn to obtain a context consistency scoring sequence, and the context consistency scoring sequence includes context consistency scores 1, 2,..., F, where F is the context consistency scoring number;

[0029] Count the number of the same highest scores in the context consistency scoring sequence:

[0030] If the number of the same highest scores in the context consistency scoring sequence is unique, the meaning of the polysemous word corresponding to the highest score is the appropriate meaning of the polysemous word in its original paragraph;

[0031] If the number of the same highest scores in the context consistency scoring sequence is not unique, then the paragraphs with the highest context consistency scores are scored for commonness to obtain a commonness scoring sequence, and the commonness scores include commonness scores 1, 2,..., M, where M is the number of the same highest scores in the context consistency scoring sequence;

[0032] If the number of the same highest scores in the commonness scoring sequence is unique, the meaning of the polysemous word corresponding to the highest score is the appropriate meaning of the polysemous word in its original paragraph;

[0033] If the number of the highest scores in the commonality scoring sequence is not unique, then perform sentiment consistency scoring on the paragraphs with the highest commonality scores to obtain a sentiment consistency scoring sequence, which includes sentiment consistency score 1, sentiment consistency score 2, ……, sentiment consistency score P, where P is the number of the highest scores that are the same in the commonality scoring sequence;

[0034] If the number of the highest scores in the sentiment consistency scoring sequence is unique, then the meaning of the polysemous word corresponding to this highest score is the adapted meaning of the polysemous word in its original paragraph; otherwise, notify the administrator to perform manual adaptation on the meaning of the polysemous word.

[0035] Preferably, constructing a knowledge graph based on the knowledge extraction data and the meaning of the polysemous word includes the following steps:

[0036] Identify similar entities through string matching and semantic similarity calculation, perform entity alignment, and use the weighted average method to fill in the attribute values of the knowledge extraction data;

[0037] Structure the knowledge extraction data and the meaning of the polysemous word, convert the structured knowledge extraction data and the meaning of the polysemous word into triple form to obtain the nodes and edges of the knowledge graph;

[0038] Store the obtained nodes and edges of the knowledge graph into a graph database to construct the knowledge graph.

[0039] Preferably, evaluating the integrity of the constructed knowledge graph to obtain an overall integrity index includes the following steps:

[0040] Obtain integrity evaluation indicators, and based on the obtained integrity evaluation indicators, perform quantitative analysis on the nodes, edges, and attributes in the knowledge graph to calculate the actual values of each indicator;

[0041] Based on the actual values of the integrity evaluation indicators, obtain the overall integrity index of the knowledge graph through the overall integrity index calculation formula.

[0042] Preferably, verifying and optimizing the constructed knowledge graph according to the overall integrity index includes the following steps:

[0043] Based on the obtained overall integrity index of the knowledge graph, compare the overall integrity index of the knowledge graph with the lowest expected threshold of integrity;

[0044] If the overall integrity index of the knowledge graph is not less than the lowest expected threshold of integrity, then it is determined that the integrity of the knowledge graph has reached the expected level;

[0045] If the overall integrity index of the knowledge graph is less than the minimum expected threshold of integrity, it is determined that the integrity of the knowledge graph does not reach the expected level, and it is output that the knowledge graph does not meet the integrity standard, and the administrator is notified to verify and optimize the knowledge graph.

[0046] Preferably, the steps for obtaining the overall integrity index are as follows:

[0047] The integrity evaluation indicators include the difference in the number of entities, the difference in the number of relationships, the difference in the number of attributes, the total number of retrieval times, and the number of unsuccessful comparisons;

[0048] Obtain the dictionary of word meanings;

[0049] Through a deep learning model different from the data extraction rules, entity extraction, relationship extraction, and attribute extraction are performed again from the knowledge extraction data;

[0050] Combined with the dictionary of word meanings, the newly extracted entities, relationships, or attributes are constructed into a comparison thesaurus;

[0051] Respectively count the number of entities, relationships, and attributes in the comparison thesaurus and the knowledge graph, and compare them with each other to obtain the difference in the number of entities, the difference in the number of relationships, and the difference in the number of attributes;

[0052] Randomly select an entity, relationship, or attribute from the comparison thesaurus, place it in the knowledge graph for comparison and retrieval, and count the total number of retrieval times and the number of unsuccessful comparisons;

[0053] Combined with the difference in the number of entities, the difference in the number of relationships, the difference in the number of attributes, the total number of retrieval times, and the number of unsuccessful comparisons, the overall integrity index is calculated through the overall integrity index calculation formula.

[0054] Preferably, the overall integrity index calculation formula is as follows:

[0055] ;

[0056] ;

[0057] In the formula, β is the overall integrity index, is the difference in the number of entities, is the difference in the number of relationships, is the difference in the number of attributes, Y is the total number of retrieval times, X is the number of unsuccessful comparisons, e is the natural constant, α1 is the weight of the difference in the number of entities for the overall integrity index, α2 is the weight of the difference in the number of relationships for the overall integrity index, α3 is the weight of the difference in the number of attributes for the overall integrity index, α4 is the weight of the ratio of the number of unsuccessful comparisons to the total number of retrieval times for the overall integrity index, and f is the transfer function.

[0058] The present invention has the following advantages:

[0059] In the present invention, several meanings included in a polysemous word are respectively substituted into the original paragraph, and then the new paragraph after substituting the meanings is scored for context consistency with the original paragraph. By evaluating the coherence and logical consistency of the paragraph, the ambiguity caused by the polysemous word is effectively reduced. By substituting the meanings of the polysemous word into the original paragraph for commonness scoring, the meanings with the same context consistency score can be scored again, effectively improving the disambiguation efficiency. By substituting the meanings of the polysemous word into the original paragraph for sentiment consistency scoring, the meanings with the same commonness score can be scored three times, improving the accuracy of the adaptation of the meanings of the polysemous word and enhancing the accuracy of the knowledge graph.

[0060] In the present invention, the knowledge extraction data and the meanings of polysemous words are structured and transformed into a triple form to obtain the nodes and edges of the knowledge graph. Thereby, the knowledge graph can be easily stored and queried, facilitating subsequent knowledge reasoning. At the same time, the knowledge graph can support complex queries and intelligent applications, having the effect of enhancing the queryability and inferability of the knowledge graph, and improving the practicality and flexibility of the knowledge graph.

[0061] In the present invention, the constructed knowledge graph is evaluated for integrity to obtain an overall integrity index, and the knowledge graph is verified and optimized according to the overall integrity index, ensuring that the knowledge graph can continuously reflect the latest domain knowledge. And through an automated evaluation and optimization mechanism, the maintenance efficiency and update speed of the knowledge graph are improved, ensuring the timeliness and reliability of the knowledge graph.

[0062] The knowledge graph constructed by the present invention can be applied in an educational and teaching dedicated large model. By combining technologies such as big data, machine learning, and knowledge graph, an intelligent science and education environment is built to create a teaching and research service platform with the characteristics of professionalism, personalization, systematicness, and one-stop. BRIEF DESCRIPTION OF THE DRAWINGS

[0063] Figure 1 It is a flowchart of an AI-based intelligent knowledge graph construction method provided by the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0064] The following further describes the present invention with reference to the drawings.

[0065] The general idea in the present invention is as follows:

[0066] By substituting several meanings contained in a polysemous word into the original paragraph respectively, and then performing a context consistency score on the new paragraph after substituting the meaning with the original paragraph, by evaluating the coherence and logical consistency of the paragraph, by substituting the meaning of the polysemous word into the original paragraph for a commonness score, thus the meanings with the same context consistency score can be scored again, by substituting the meaning of the polysemous word into the original paragraph for an emotional consistency score, thus the meanings with the same commonness score can be scored three times, achieving the effect of quickly reducing the ambiguity caused by polysemous words and improving the accuracy of polysemous word meaning adaptation.

[0067] As Figure 1 shown, the AI-based intelligent knowledge graph construction method includes the following steps:

[0068] Obtain the target topic domain, and collect knowledge data from the data source according to the target topic domain;

[0069] Clean the collected knowledge data to obtain the cleaned knowledge data;

[0070] Extract the data from the cleaned knowledge data to obtain the knowledge extraction data;

[0071] Extract the polysemous words in the knowledge extraction data, adopt a multiple scoring mechanism for the polysemous words, determine the meaning of the polysemous words through the multiple scoring mechanism, and obtain the polysemous word meaning;

[0072] According to the knowledge extraction data and the polysemous word meaning, perform knowledge representation and storage;

[0073] Evaluate the integrity of the constructed knowledge graph to obtain an overall integrity index, and the overall integrity index is used to measure the integrity of the knowledge graph;

[0074] Verify and optimize the constructed knowledge graph according to the overall integrity index.

[0075] In this embodiment, in the data collection stage, a large amount of information is obtained from various data sources, including structured basic information such as colleges, majors, courses, and textbooks extracted from school databases, educational administration systems, etc., and structured or semi-structured information such as table of contents, basic theories, and main viewpoints, as well as unstructured information such as innovation trends and social practices extracted from text resources such as professional textbooks, curriculum outlines, academic papers, research reports, scientific research data, and internship and training cases through large models (natural language processing technology).

[0076] Outliers are removed from the collected data through the IQR (Interquartile Range) technique, and noise is removed from the collected data through the principal component analysis (PCA) technique. Standardized naming is performed on keywords such as disciplines, departments, and majors, which can ensure data consistency, improve data accuracy, and facilitate subsequent knowledge extraction. Machine learning algorithms, such as conditional random fields (CRF), long short-term memory networks (LSTM), etc., are used to automatically annotate and classify knowledge extraction information, extract the required information, and convert the extracted knowledge elements into triple form to initially form a knowledge graph.

[0077] By converting unstructured data into structured knowledge representation, the readability of the data is enhanced. When encountering polysemous words, a multiple scoring mechanism is adopted to determine their meanings in specific contexts, effectively solving the ambiguity problem in natural language processing, improving the accuracy and usability of the knowledge graph. After extraction, the extracted knowledge elements are stored in a structured form for subsequent querying and reasoning, improving the application value and practicality of knowledge.

[0078] Integrity assessment helps identify missing parts in the knowledge graph by measuring the completeness of the knowledge graph, while verification and optimization further improve the accuracy and reliability of the knowledge graph, not only improving the quality of the knowledge graph but also ensuring that the knowledge graph can continuously support various intelligent applications, enhancing the practicality and effectiveness of the knowledge graph.

[0079] Collect knowledge data from the data source according to the target subject domain, including the following steps: obtain the target subject domain covered by the knowledge graph; connect to the data source and identify the data source type; determine the data collection tool according to the data source type, and collect and extract knowledge data based on the target subject domain using data mining algorithms and text analysis techniques.

[0080] In this embodiment, combined with the expected application scenario, the target subject domain is clarified to collect knowledge data from the data source to ensure that the constructed knowledge graph can meet the actual needs. Connect to the data source through NLP technology, identify the type of the connected data source through a deep learning model, extract structured data in the data source by using an ETL tool, extract semi-structured data in the data source by using a JSON parsing tool, and extract unstructured data in the data source through API calls. Selecting appropriate tools can improve the efficiency and accuracy of data collection. Finally, collect and extract knowledge data based on the target subject domain using data mining algorithms and text analysis techniques, which can ensure the extraction of high-quality knowledge data related to the target subject domain from the data source.

[0081] Data cleaning is performed on the collected knowledge data to obtain the cleaned knowledge data. Data extraction is then performed on the cleaned knowledge data to obtain the knowledge extraction data, including the following steps: duplicate records are removed through unique identifiers, and data in different formats are unified into a standard format; by organizing different data sources, relevant data are merged, and redundant data are removed to obtain the cleaned knowledge data; based on the cleaned knowledge data, entity extraction, relationship extraction, and attribute extraction are performed through data extraction rules to obtain the knowledge extraction data.

[0082] In this embodiment, the data extraction rules include rule-based extraction and statistic-based extraction. By removing duplicate records through unique identifiers and unifying data in different formats into a standard format, duplicate data in the data can be avoided, improving the quality and accuracy of the data. Standardized naming is performed on keywords such as disciplines, departments, and majors, and data from different sources are integrated into a unified data set to ensure data consistency and improve the quality of data collection.

[0083] Among them, the statistic-based extraction method: using machine learning algorithms such as conditional random fields (CRF), long short-term memory networks (LSTM), etc., to extract the required information from the cleaned knowledge data. The extraction process includes entity extraction, relationship extraction, and attribute extraction. The entity extraction process is to automatically discover entities such as disciplines, departments, majors, authors, and publishers from the cleaned knowledge data through named entity recognition (NER) technology. By training the model through machine learning methods, it is possible to learn how to identify and classify the relationships between entities from a large amount of labeled data, so as to extract the relationships between entities; for knowledge points, specific attributes such as table of contents structure, basic theory, and main viewpoints are extracted from the cleaned knowledge data through in-depth semantic understanding and analysis, so as to complete the attribute extraction and obtain the knowledge extraction data.

[0084] Extract the polysemous words in the knowledge extraction data, and adopt a multiple scoring mechanism for the polysemous words. The meaning of the polysemous words is determined through the multiple scoring mechanism to obtain the meaning of the polysemous words, including the following steps: based on the polysemous words in the knowledge extraction data, extract the meanings 1, 2,..., N of the polysemous words, where N is the meaning number; score the meanings 1, 2,..., N of the polysemous words through the multiple scoring mechanism, and based on the scoring results, obtain the meaning of the polysemous words.

[0085] The multiple scoring mechanism includes the following steps: Obtain the meanings 1, 2, …, N of the polysemous word; Substitute meanings 1, 2, …, N into the original paragraph where the polysemous word is located one by one to obtain paragraphs 1, 2, …, G, where G is the paragraph number, and perform context consistency scoring on the original paragraph with paragraphs 1, 2, …, G in turn. The context consistency scoring method is as follows: Evaluate the consistency of the paragraph structure through dependency syntactic analysis, and then calculate the correctness, integrity, and logic of the dependency relationship for scoring to obtain a context consistency scoring sequence. The context consistency scoring sequence includes context consistency scores 1, 2, …, F, where F is the context consistency score number; Count the number of the same highest scores in the context consistency scoring sequence: If the number of the same highest scores in the context consistency scoring sequence is unique, the meaning of the polysemous word corresponding to the highest score is the appropriate meaning of the polysemous word in its original paragraph, which has the effect of being able to select the appropriate meaning of the polysemous word through one scoring, saving working time.

[0086] If the number of the same highest scores in the context consistency scoring sequence is not unique, then perform commonness scoring on the paragraphs with the highest context consistency scores. The commonness scoring method is as follows: Based on the statistical data of the N-gram model, determine the commonness of a meaning, and perform scoring by calculating the frequency of a specific meaning appearing in the corpus. The higher the frequency, the higher the commonness score, to obtain a commonness scoring sequence. The commonness scores include commonness scores 1, 2, …, M, where M is the number of the same highest scores in the context consistency scoring sequence; If the number of the same highest scores in the commonness scoring sequence is unique, the meaning of the polysemous word corresponding to the highest score is the appropriate meaning of the polysemous word in its original paragraph. By performing secondary scoring on the paragraphs, it prevents the situation that a single context consistency score may not be sufficient to distinguish polysemous words with similar meanings, enhancing the robustness of the system and improving the user experience.

[0087] If the number of the same highest scores in the commonality score sequence is not unique, then the emotional consistency of the paragraphs with the highest commonality scores is scored again. The emotional consistency scoring method is as follows: By using an emotion analysis tool, such as the Snownlp library, to perform emotional polarity analysis on the paragraphs and output the probabilities of positive or negative emotions, and then calculating the emotional polarity probabilities output by the Snownlp library for scoring. The more consistent the emotional tendency is with the original paragraph, the higher the emotional consistency score. The emotional consistency score sequence is obtained, including emotional consistency score 1, emotional consistency score 2,..., emotional consistency score P, where P is the number of the same highest scores in the commonality score sequence. If the number of the same highest scores in the emotional consistency score sequence is unique, then the meaning of the polysemous word corresponding to the highest score is the appropriate meaning of the polysemous word in its original paragraph. Otherwise, the administrator is notified to perform manual adaptation on the meaning of the polysemous word. The multiple scoring mechanism helps to provide more accurate results and reduce users' doubts about the results and the situations that need to be manually corrected.

[0088] According to the knowledge extraction data and the meaning of the polysemous word, construct a knowledge graph, including the following steps: Identify similar entities through string matching and semantic similarity calculation. Machine learning algorithms such as cosine similarity or Jaccard similarity can be used to quantify the similarity between entities to ensure the accuracy of entity recognition, and then perform entity alignment and use the weighted average method to fill in the attribute values of the knowledge extraction data.

[0089] Structurize the knowledge extraction data and the meaning of the polysemous word through natural language processing (NLP) tools, transform the structured knowledge extraction data and the meaning of the polysemous word into triple form to obtain the nodes and edges of the knowledge graph, and thus the relationships and attributes between entities can be clearly represented. Store the obtained nodes and edges of the knowledge graph into a graph database to construct the knowledge graph.

[0090] Conduct integrity evaluation on the constructed knowledge graph to obtain the overall integrity index, including the following steps: Obtain integrity evaluation indicators, and based on the obtained integrity evaluation indicators, perform quantitative analysis on the nodes, edges and attributes in the knowledge graph and calculate the actual values of each indicator. Based on the actual values of the integrity evaluation indicators, obtain the overall integrity index of the knowledge graph through the overall integrity index calculation formula, which helps to improve the reliability and usability of the knowledge graph and ensure the accuracy and reliability of the information in the knowledge graph.

[0091] Check and optimize the constructed knowledge graph according to the overall integrity index, including the following steps: Based on the obtained overall integrity index of the knowledge graph, compare the overall integrity index of the knowledge graph with the lowest expected threshold of integrity. If the overall integrity index of the knowledge graph is not less than the lowest expected threshold of integrity, it is determined that the integrity of the knowledge graph has reached the expected level.

[0092] If the overall integrity index of the knowledge graph is less than the minimum expected threshold of integrity, it is determined that the integrity of the knowledge graph does not reach the expected level. Output that the knowledge graph does not meet the integrity standard, and notify the administrator to verify and optimize the knowledge graph. When new knowledge is added, it can ensure seamless integration of the newly added knowledge with the existing knowledge graph, maintaining overall consistency and integrity.

[0093] The steps to obtain the overall integrity index are as follows: The integrity evaluation indicators include the difference in the number of entities, the difference in the number of relationships, the difference in the number of attributes, the total number of retrieval times, and the number of unsuccessful comparisons; Obtain the dictionary of word meanings; Through a deep learning model different from the data extraction rules, perform entity extraction, relationship extraction, and attribute extraction from the knowledge extraction data again; Combine the dictionary of word meanings to construct a comparison vocabulary with the newly extracted entities, relationships, or attributes; Count the number of entities, relationships, and attributes in the comparison vocabulary and the knowledge graph respectively, and compare them with each other to obtain the difference in the number of entities, the difference in the number of relationships, and the difference in the number of attributes; Randomly select an entity, relationship, or attribute from the comparison vocabulary, place it in the knowledge graph for comparison retrieval, and count the total number of retrieval times and the number of unsuccessful comparisons; Combine the difference in the number of entities, the difference in the number of relationships, the difference in the number of attributes, the total number of retrieval times, and the number of unsuccessful comparisons, and calculate the overall integrity index through the overall integrity index calculation formula.

[0094] In this embodiment, the deep learning model specifically refers to: performing pre-training on a large amount of unlabeled text data and unsupervised pre-training tasks on the BERT pre-trained language representation model, so that the BERT pre-trained language representation model has richer language representation capabilities and better context understanding capabilities. Through the BERT pre-trained language representation model, entity extraction, relationship extraction, and attribute extraction are performed from the knowledge extraction data, and a comparison vocabulary is constructed by combining the dictionary of word meanings with the newly extracted entities, relationships, or attributes, which can more accurately identify key information in the text and improve the extraction coverage rate.

[0095] The overall integrity index calculation formula is:

[0096] ;

[0097] ;

[0098] In the formula, β is the overall integrity index, is the difference in the number of entities, is the difference in the number of relationships, Let Δ be the difference in the number of attributes, Y be the total number of retrievals, X be the number of unsuccessful comparisons, e be the natural constant, α1 be the weight of the difference in the number of entities for the overall integrity index, α2 be the weight of the difference in the number of relationships for the overall integrity index, α3 be the weight of the difference in the number of attributes for the overall integrity index, α4 be the weight of the ratio of the number of unsuccessful comparisons to the total number of retrievals for the overall integrity index, and f be the transfer function.

[0099] An example calculation is that when = 1, α1 = 0.2, = 1, α2 = 0.2, = 1, α3 = 0.2, X = 0, Y = 28, α4 = 0.4, the above formula can be used to calculate that β = f = 0.8019. The closer the overall integrity is to 1, the better the integrity. The administrator sets the minimum expected threshold for integrity. When the overall integrity is lower than the minimum expected threshold for integrity, it is output that the knowledge graph does not meet the integrity standard, and the administrator conducts the next verification and optimization.

[0100] In summary, the present invention effectively prevents the paragraph from being ambiguous due to the uncertain meaning of the polysemous word by substituting the meanings of the polysemous word into the original paragraph respectively for scoring through a multiple scoring mechanism, and enhances the accuracy of the knowledge graph.

Claims

1. AI-based intelligent knowledge graph construction method, characterized in that It includes the following steps: Obtain the target subject domain and collect knowledge data from the data source according to the target subject domain; Clean the collected knowledge data to obtain the cleaned knowledge data; Extract data from the cleaned knowledge data to obtain knowledge extraction data; Extract the polysemous words in the knowledge extraction data, adopt a multiple scoring mechanism for the polysemous words, and determine the meanings of the polysemous words through the multiple scoring mechanism to obtain the meanings of the polysemous words; The step of cleaning the collected knowledge data to obtain the cleaned knowledge data and extracting data from the cleaned knowledge data to obtain the knowledge extraction data includes the following steps: Remove duplicate records through a unique identifier and unify data in different formats into a standard format; By sorting out different data sources, merging relevant data, and removing redundant data, obtain the cleaned knowledge data; Based on the cleaned knowledge data, perform entity extraction, relationship extraction, and attribute extraction through data extraction rules to obtain knowledge extraction data; The multiple scoring mechanism includes the following steps: Obtain the meanings 1, 2,..., N of the polysemous word; Substitute the meanings 1, 2,..., N into the original paragraph where the polysemous word is located in turn to obtain paragraphs 1, 2,..., G (G is the paragraph number). Perform context consistency scoring on the original paragraph with paragraphs 1, 2,..., G in turn to obtain a context consistency scoring sequence. The context consistency scoring sequence includes context consistency scores 1, 2,..., F (F is the context consistency scoring number); Count the number of the same highest scores in the context consistency scoring sequence: If the number of the same highest scores in the context consistency scoring sequence is unique, the meaning of the polysemous word corresponding to the highest score is the appropriate meaning of the polysemous word in its original paragraph; If the number of the same highest scores in the context consistency scoring sequence is not unique, then perform commonness scoring on the paragraphs with the highest context consistency scores to obtain a commonness scoring sequence. The commonness scoring includes commonness scores 1, 2,..., M (M is the number of the same highest scores in the context consistency scoring sequence); If the number of the same highest scores in the commonness scoring sequence is unique, the meaning of the polysemous word corresponding to the highest score is the appropriate meaning of the polysemous word in its original paragraph; If the number of the same highest scores in the commonness scoring sequence is not unique, then perform emotional consistency scoring on the paragraphs with the highest commonness scores to obtain an emotional consistency scoring sequence. The emotional consistency scoring sequence includes emotional consistency scores 1, 2,..., P (P is the number of the same highest scores in the commonness scoring sequence); If the number of the same highest scores in the emotional consistency scoring sequence is unique, the meaning of the polysemous word corresponding to the highest score is the appropriate meaning of the polysemous word in its original paragraph. Otherwise, notify the administrator to perform manual adaptation on the meaning of the polysemous word; Perform knowledge representation and storage according to the knowledge extraction data and the meanings of the polysemous words; The construction of a knowledge graph according to the knowledge extraction data and the meanings of the polysemous words includes the following steps: By string matching and semantic similarity calculation, similar entities are identified for entity alignment, and the weighted average method is used to fill in the attribute values of the knowledge extraction data; The knowledge extraction data and the meanings of polysemous words are structured, and the structured knowledge extraction data and the meanings of polysemous words are transformed into triple forms to obtain the nodes and edges of the knowledge graph; The obtained nodes and edges of the knowledge graph are stored in a graph database to construct a knowledge graph; The constructed knowledge graph is evaluated for integrity to obtain an overall integrity index, which is used to measure the completeness of the knowledge graph; The evaluation of the integrity of the constructed knowledge graph to obtain an overall integrity index includes the following steps: Obtain integrity evaluation indicators, and based on the obtained integrity evaluation indicators, conduct quantitative analysis on the nodes, edges, and attributes in the knowledge graph to calculate the actual values of each indicator; Based on the actual values of the integrity evaluation indicators, the overall integrity index of the knowledge graph is obtained through the overall integrity index calculation formula; The constructed knowledge graph is verified and optimized according to the overall integrity index; The steps for obtaining the overall integrity index are as follows: The integrity evaluation indicators include the difference in the number of entities, the difference in the number of relationships, the difference in the number of attributes, the total number of retrieval times, and the number of unsuccessful comparisons; Obtain a dictionary of word meanings; Through a deep learning model different from the data extraction rules, entity extraction, relationship extraction, and attribute extraction are performed again from the knowledge extraction data; Combined with the dictionary of word meanings, the newly extracted entities, relationships, or attributes are constructed into a comparison thesaurus; Count the number of entities, relationships, and attributes in the comparison thesaurus and the knowledge graph respectively, and compare them with each other to obtain the difference in the number of entities, the difference in the number of relationships, and the difference in the number of attributes; Randomly select an entity, relationship, or attribute from the comparison thesaurus, place it in the knowledge graph for comparison and retrieval, and count the total number of retrieval times and the number of unsuccessful comparisons; Combined with the difference in the number of entities, the difference in the number of relationships, the difference in the number of attributes, the total number of retrieval times, and the number of unsuccessful comparisons, the overall integrity index is calculated through the overall integrity index calculation formula; The overall integrity index calculation formula is as follows: ; ; In the formula, β is the overall integrity index, is the difference in the number of entities, is the difference in the number of relationships, is the difference in the number of attributes, Y is the total number of retrievals, X is the number of unsuccessful comparisons, e is the natural constant, α1 is the weight of the difference in the number of entities for the overall integrity index, α2 is the weight of the difference in the number of relationships for the overall integrity index, α3 is the weight of the difference in the number of attributes for the overall integrity index, α4 is the weight of the ratio of the number of unsuccessful comparisons to the total number of retrievals for the overall integrity index, and f is the transfer function; The verification and optimization of the constructed knowledge graph according to the overall integrity index includes the following steps: Based on the overall integrity index of the obtained knowledge graph, compare the overall integrity index of the knowledge graph with the lowest expected threshold of completeness; If the overall integrity index of the knowledge graph is not less than the lowest expected threshold of completeness, it is determined that the completeness of the knowledge graph has reached the expected level; If the overall integrity index of the knowledge graph is less than the lowest expected threshold of completeness, it is determined that the completeness of the knowledge graph has not reached the expected level, and it is output that the knowledge graph does not meet the integrity standard, and the administrator is notified to verify and optimize the knowledge graph.

2. The AI-based intelligent knowledge graph construction method according to claim 1, wherein: The collection of knowledge data from the data source according to the target subject domain includes the following steps: Obtain the target subject domain covered by the knowledge graph; Connect to the data source and identify the data source type; Determine the data collection tool according to the data source type, and collect and extract knowledge data based on the target subject domain using data mining algorithms and text analysis techniques.

3. The AI-based intelligent knowledge graph construction method according to claim 1, wherein: Extract the polysemous words in the knowledge extraction data, adopt a multiple scoring mechanism for the polysemous words, and determine the meanings of the polysemous words through the multiple scoring mechanism to obtain the meanings of the polysemous words, including the following steps: Based on the polysemous words in the knowledge extraction data, extract the meanings 1, 2,..., N of the polysemous words, where N is the meaning number; Score the meanings 1, 2,..., N of the polysemous words through the multiple scoring mechanism, and obtain the meaning of the polysemous word based on the scoring results.

Citation Information

Patent Citations

  • Word sense disambiguation method and system

    CN101840397A

  • Method and system for disambiguating polysemy in sentences

    CN111858952A

  • Energy industry knowledge graph construction method and device based on multi-source heterogeneous data fusion technology

    CN117313849A