Method and system for constructing knowledge graph by using artificial intelligence
Through artificial intelligence systems, data on educational knowledge graphs are collected and optimized, and the accuracy and readability of knowledge graphs in the existing technology are solved, and efficient and accurate knowledge graph construction and optimization are achieved.
Patent Information
- Application Number
- CN202411935815.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-26
- Publication Date
- 2025-05-30
AI Technical Summary
The existing educational knowledge graph generation methods rely on artificial intelligence, which can easily lead to reduced accuracy of knowledge graphs and complex content, affecting readability.
Create a database through an artificial intelligence system, enter keywords, collect relevant data, and perform quality detection and preprocessing. Then, entities and relationships are extracted, fusion and optimization of entities and relationships are carried out, and the knowledge graph is finally constructed and optimized.
Improve the accuracy and readability of the knowledge graph, reduce construction errors through manual intervention and data optimization, and provide simplified and detailed knowledge graphs for different users to choose.
Smart Images

Figure CN120069020A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of knowledge graph construction, and particularly to a method and system for constructing a knowledge graph using artificial intelligence. Background Art
[0002] With the rapid development and popularization of information technology, the knowledge graph, as an efficient and intuitive way of knowledge representation and organization, has gradually gained favor in various industries. Based on a graph structure, the knowledge graph presents concepts, entities in the real world and the relationships between them in a networked form, thus providing a new mode of knowledge acquisition, management and application. Especially in key fields such as education, healthcare, and finance, the application of the knowledge graph has had a profound impact.
[0003] However, although the knowledge graph has great application potential in various fields, there are some problems with existing methods for generating educational knowledge graphs: First, due to the rise of big data and artificial intelligence, many current methods for constructing knowledge graphs are completely based on artificial intelligence to complete the construction of the knowledge graph. Although this way simplifies the workload, artificial intelligence can only construct the knowledge graph based on the data in the network system, which may lead to incorrect judgments of certain knowledge, thereby affecting the accuracy of the knowledge graph; Second, currently, the knowledge graphs constructed by artificial intelligence basically extract all entities and relationships, and then establish a knowledge graph supported by multiple data. This will result in too many entities in the knowledge graph constructed in some fields, the entire knowledge graph framework is too large, and the content is too redundant, which is not convenient for different people to view and reduces the readability of the knowledge graph. Summary of the Invention
[0004] The present invention aims to provide a method and system for constructing a knowledge graph using artificial intelligence to solve the technical problems raised in the above background art.
[0005] To achieve the above object, the present invention provides the following technical solutions:
[0006] A method for constructing a knowledge graph using artificial intelligence, comprising the following steps:
[0007] S1. Establish a database, input keywords, and collect all materials related to the keywords based on artificial intelligence and big data to obtain preliminary data, which includes structured data, semi-structured data, and unstructured data. Classify different data types and add label remarks to different types of data;
[0008] Data Detection: Based on the artificial intelligence system, quality inspection is carried out on all retrieved materials. Quality standards are set, a standard database and a differential database are established. The data that meet the quality standards are output to the standard database, and the data that do not meet the quality standards are output to the differential database;
[0009] S2. Data Preprocessing: Preprocess the text data in the standard database.
[0010] S3. Entity Extraction: Based on the artificial intelligence system, key entities are identified and extracted from the preprocessed data. After the key entities are extracted, corresponding labels are attached to each key entity for subsequent extraction use;
[0011] S4. Relationship Extraction: Based on the key entities extracted in step S3, taking two different key entities as a group combination, the relationship between each group combination is extracted from the preprocessed data based on artificial intelligence, and their relationship is labeled.
[0012] S5. Knowledge Fusion
[0013] Constructing a Comparison Group: All the above-extracted key entities are named with individual labels, and the key entities are attached with the named labels again, that is, the key entities are named as Entity 1, Entity 2... Entity n. Different key entities are grouped in pairs to form a comparison group, and at the same time, an ambiguity analysis database is established.
[0014] Entity Fusion: Based on artificial intelligence, entity fusion determination is carried out on the key entities in the comparison group. When it is judged that the comparison group is the same key entity, it is determined that they are mutually transformable entities, and they are fused into the same key entity. All comparison groups are judged, and then all mutually transformable entities are fused into the same key entity.
[0015] Second Relationship Fusion: After the mutually transformable entities are fused into the same key entity, the relationships between the key entities before fusion are also merged, and the merged relationships are assisted with named labels, that is, the relationships are named as Relationship 1, Relationship 2... Relationship n. Different relationships are grouped in pairs to form a comparison group. Based on artificial intelligence, relationship fusion judgment between the comparison groups is realized. When it is judged that the relationships between the relationships are the same relationship, they are fused into the same relationship to further eliminate referential ambiguity;
[0016] S6. Manual Intervention
[0017] Manual Judgment: The key entities and relationships for which the artificial intelligence cannot judge whether they are mutually transformable entities or relationships enter the ambiguity analysis database. At this time, the artificial is reminded to intervene and judge whether the key entities or relationships that the artificial intelligence cannot judge are mutually transformable entities or relationships; thus, the fusion of all key entities and the disambiguation of all relationships are completed.
[0018] Ambiguous data intelligent analysis: Analyze the data in the differential database through an artificial intelligence system, compare all the information in the differential database, set the duplication rate standard, extract the data whose data duplication rate reaches the set duplication rate standard, and output this information to the ambiguity analysis database;
[0019] S7. Knowledge processing, the artificial intelligence system further reasons based on all the key entities and all the relationships that have been fused, further expands the relationships between key entities based on the database information, and then completes the supplementation of key entities or relationships.
[0020] S8. Construction of the first knowledge graph
[0021] Design the knowledge graph structure: Define the schema of the knowledge graph, including entity types, relationship types, and attribute types; Determine the association relationships between entities, as well as the attributes of each entity and relationship;
[0022] Select and configure the graph database: Select the graph database algorithm for storing and managing the knowledge graph; Configure the graph database environment and set storage parameters and query parameters;
[0023] Construct the first knowledge graph: Import all the processed key entities and all relationship data into the graph database to form the preliminary structure of the first knowledge graph; Create entity nodes and relationship edges in the graph database, and add appropriate labels and attributes to the entity nodes and relationship edges, and save the first knowledge graph.
[0024] S9. Optimization of the first knowledge graph
[0025] Determine the connection number: Take the number of the remaining key entities connected to a group of key entities in the first knowledge graph as the connection number, and name the two parts of key entities as the upper-level key entity and the lower-level key entity. The connection number is the number of lower-level key entities connected to the upper-level key entity.
[0026] Determine the minimum number of lower-level key entities: Through extensive analysis of the original data by the artificial intelligence system, determine the minimum number of lower-level key entities required by the upper-level key entity.
[0027] Determine the importance ratio of all lower-level key entities: Extract and analyze all the lower-level key entities through the artificial intelligence system, obtain the importance degree of each lower-level key entity in the upper-level key entity through comprehensive analysis of the original data by the artificial intelligence system, export the importance ratio, sort all the lower-level key entities according to the importance ratio, and obtain the duplication rate of each lower-level key entity through comprehensive analysis of the original data by the artificial intelligence system to determine the usage rate of the lower-level key entities, and sort all the lower-level key entities according to the usage rate;
[0028] Extract the main entities: Calculate the quantity of importance proposed based on the artificial intelligence model, determine the quantity of importance proposed based on the minimum number of lower-level key entities, input the minimum number of lower-level key entities into the artificial intelligence model, calculate the quantity of importance proposed through the artificial intelligence model, and based on the above-mentioned lower-level key entities sorted by importance rate, propose the lower-level key entities whose importance rate is before the quantity of importance proposed as the first main entity; Compare the first main entity with the lower-level key entities based on the repetition rate, delete the lower-level key entities that repeat with the first main entity in the sorting of the lower-level key entities based on the usage rate, determine the remaining main entities based on the artificial intelligence model, input the usage rate and importance rate of the remaining lower-level key entities that have not been extracted as the first main entity into the artificial intelligence model, calculate the median proportion through the artificial intelligence model, sort the remaining lower-level key entities based on the median proportion, and select the second main entity based on the median proportion. The number of the second main entities is the minimum number of lower-level entities minus the number of the first main entity;
[0029] S10. Construction of the second knowledge graph: Delete the remaining key entities in the first knowledge graph except the first main entity and the second main entity to obtain the second knowledge graph; Establish a retrieval engine: Create a retrieval engine for the knowledge graph, and both the first knowledge graph and the second knowledge graph will pop up during retrieval.
[0030] Preferably, the data preprocessing includes the following steps:
[0031] Data conversion: Extract and convert the information in the unstructured data and semi-structured data in the standard database into the same text data,
[0032] Data supplementation: Delete the noise data and redundant information in the data based on the artificial intelligence system; Supplement the data integrity, fill in the missing information in the data based on the artificial intelligence system, and at the same time identify basic problems such as misspelling, speech errors, and word order disorders in the data.
[0033] Preferably, the step S5 further includes the following steps,
[0034] The first relationship fusion: Based on the artificial intelligence, realize the relationship fusion judgment between the key entities of the comparison group. When it is judged that the relationship between the key entities of the comparison group is the same relationship, it is determined as a mutual transformation relationship, and they are fused into the same relationship to eliminate the reference ambiguity; The first relationship fusion is carried out after constructing the comparison group.
[0035] Preferably, the step S5 further includes the following steps,
[0036] Ambiguous data output: When it is impossible to determine whether two sets of relationships or two sets of entities are mutually convertible entities or relationships during the above-mentioned first relationship fusion, entity fusion, and second relationship fusion processes, the comparison group or the comparison set is output to the ambiguous analysis database; the ambiguous data output is performed after the second relationship fusion step.
[0037] Preferably, step S6 further includes the following steps
[0038] Artificial analysis of ambiguous data: Manual intervention is carried out to manually judge the quality of the data and whether the content needs to be added to the construction of the knowledge graph. If necessary, entity extraction, relationship extraction, and knowledge fusion steps are performed on the data, and all the fused data are merged; the artificial analysis of ambiguous data is carried out after the intelligent analysis of ambiguous data.
[0039] Preferably, the method further includes
[0040] S11. Data update and coverage. By updating or covering the data in the knowledge graph, the knowledge graph can quickly and timely reflect new knowledge, new viewpoints, and new discoveries.
[0041] Preferably, the data update and coverage include:
[0042] Set the update interval: Adaptively set the data update time according to the industry development situation;
[0043] Retrieve new data: After reaching the set update time, the artificial intelligence system automatically retrieves relevant data to obtain information related to the keywords;
[0044] New data screening and application: Perform quality analysis on the obtained new data. The data that meets the set quality standards enters the standard database. The artificial intelligence system compares the new data with the old data. When the new data and the old data are mutually convertible data, no subsequent update is performed. When the new data and the old data are not mutually convertible data, a discrimination is made to determine whether the new data is newly added data or corrected data. When the artificial intelligence system determines that the new data is corrected data, the new data is processed according to the above steps and covers the original key entity or relationship to achieve the purpose of data update. When the artificial intelligence system determines that the new data is corrected data, the new data is processed according to the above steps and added to the first knowledge graph to achieve the purpose of expanding and improving the knowledge graph; and the second knowledge graph is modified according to steps S9 and S10 above.
[0045] A system for constructing a knowledge graph using artificial intelligence, the system includes:
[0046] A data collection module, which is used to collect preliminary data, store the corresponding data in the corresponding database after preliminary processing of the data;
[0047] A data preprocessing module for preprocessing the data in the corresponding database;
[0048] A knowledge extraction module for extracting entities and relationships and making preliminary processing of the extracted entities and relationships;
[0049] A knowledge fusion module for judging, qualitative determination, disambiguation, coreference resolution and fusion of the extracted entities and relationships;
[0050] An artificial processing module for human-computer interaction to facilitate manual judgment and screening of data, entities and relationships;
[0051] A knowledge processing module for reasoning, supplementing data, entities and relationships to complete the construction of a knowledge system;
[0052] A knowledge graph construction module for constructing an education knowledge graph according to the identified key entities and the extracted relationships;
[0053] An optimization and export module for optimizing the first knowledge graph, selecting important key entities, and generating a relatively concise knowledge graph.
[0054] Beneficial effects of this technical solution:
[0055] (1) The technical solution provided by the present invention extracts entities and relationships through an artificial intelligence system, and detects, identifies and selects entities and relationships, which can more accurately identify the entities and relationships corresponding to keywords, improve the accuracy of knowledge graph construction, and sets up an artificial intervention section to complete the construction of the entire knowledge graph through the method of human-machine interaction. When there is information that cannot be judged by the artificial intelligence system, it is analyzed manually, and manual can also choose whether to add this information to the knowledge graph according to the information, which greatly improves the accuracy of the knowledge graph while ensuring the efficiency of constructing the knowledge graph through the artificial intelligence system; at the same time, it increases the repeated screening process of unqualified data and reduces the construction error of the knowledge graph.
[0056] (2) The technical solution provided by the present invention can provide two groups of knowledge graphs at the same time, one is a comprehensive knowledge graph and the other is a simplified knowledge graph. Readers can choose a suitable knowledge graph for viewing according to their actual needs, which greatly improves the readability of the knowledge graph. By retrieving and calculating all data through the artificial intelligence system, the minimum number of lower-layer key entities of the upper-layer key entities can be accurately obtained. Through the artificial intelligence system, the importance rate and usage rate of the lower-layer key entities can also be judged, and then the lower-layer key entities that the upper-layer key entities in the simplified knowledge graph should most connect can be accurately calculated, ensuring the accuracy of the simplified knowledge graph.
[0057] (3) The technical solution provided by the present invention can automatically update the knowledge graph at regular intervals, thereby avoiding errors in the knowledge graph information caused by outdated data or the overthrow of the original data by new data. The automatic update and supplementation of the knowledge graph are completed based on the artificial intelligence system, enabling the knowledge graph to quickly and timely reflect new knowledge, new viewpoints, and new discoveries, thereby improving the actual application effect of the knowledge graph. BRIEF DESCRIPTION OF THE DRAWINGS
[0058] Figure 1 FIG. is a functional block diagram of a system for constructing a knowledge graph using artificial intelligence provided by the present invention; DETAILED DESCRIPTION OF THE EMBODIMENTS
[0059] The present invention will be further described in detail below in conjunction with the drawings and embodiments:
[0060] A method for constructing a knowledge graph using artificial intelligence includes the following steps:
[0061] S1. Establish a database, input keywords, and collect all materials related to the keywords based on artificial intelligence and big data to obtain preliminary data. This data includes structured data, semi-structured data, and unstructured data, and the data includes information such as audio, video, and pictures. Classify different data types and add label remarks to different types of data;
[0062] Data detection: Based on the artificial intelligence system, perform quality detection on all retrieved materials, set quality standards, establish a standard database and a differential database, output the data that meets the quality standards to the standard database, and output the data that does not meet the quality standards to the differential database;
[0063] S2. Data preprocessing, preprocess the text data in the standard database,
[0064] Data conversion: Extract and convert the information in the unstructured data and semi-structured data in the standard database into the same text data,
[0065] Data supplementation: Based on the artificial intelligence system, delete the noise data and redundant information in the data; supplement the data integrity, fill in the missing information in the data based on the artificial intelligence system, and at the same time identify basic problems such as misspelling, speech errors, and incorrect word order in the data;
[0066] S3. Entity extraction, based on the artificial intelligence system, identify and extract key entities from the preprocessed data, and attach corresponding labels to each key entity after extraction for subsequent extraction and use;
[0067] S4. Relationship extraction. Based on the key entities extracted in step S3, taking two different groups of key entities as a combination, extract the relationships between each combination from the preprocessed data based on the artificial intelligence system, and label the relationship types between them and attach corresponding labels.
[0068] S5. Knowledge fusion
[0069] Construct a comparison group: Name each of the above-extracted key entities separately, attach the naming labels to the key entities again, that is, name the key entities as Entity 1, Entity 2... Entity n, form a comparison group with different key entities in pairs, and establish an ambiguity analysis database at the same time.
[0070] First relationship fusion: Based on artificial intelligence, implement the relationship fusion determination between the key entities in the comparison group. When it is determined that the relationship between the key entities in the comparison group is the same relationship, it is determined as a mutual transformation relationship, and they are fused into the same relationship to eliminate referential ambiguity.
[0071] Entity fusion: Based on artificial intelligence, implement the entity fusion determination for the key entities in the comparison group. When it is determined that the comparison group consists of the same key entities, they are determined as mutual transformation entities, and they are fused into the same key entity. Determine all comparison groups and then fuse all mutual transformation entities into the same key entity.
[0072] For example, if Entity 1 and Entity 2 currently form a comparison group, when the system determines that Entity 1 and Entity 2 are mutual transformation entities, it means that Entity 1 and Entity 2 are the same concept. Fuse Entity 1 and Entity 2 into the same key entity. At this time, if there is another Entity 3, Entity 1 and Entity 3 will form a comparison group, and Entity 2 and Entity 3 will also form a comparison group. At this time, based on the artificial intelligence system, determine whether Entity 1 and Entity 3 are mutual transformation entities, and at the same time determine whether the relationship between Entity 2 and Entity 3 is a mutual transformation relationship. If it is determined that all three are mutual transformation relationships, then the three are fused into the same key entity. This can also review the determination situation between Entity 1 and Entity 2. Since Entity 1 and Entity 2 are mutual transformation entities, therefore, if the determination situation between Entity 1 and Entity 3 is definitely the same as that between Entity 2 and Entity 3, so the determination review can be carried out to further improve the accuracy.
[0073] Second relationship fusion: After the mutual transformation entities are fused into the same key entity, merge the relationships between the key entities before fusion, and assist in naming the merged relationships with labels, that is, name the relationships as Relationship 1, Relationship 2... Relationship n, form a comparison group with different relationships in pairs, and based on artificial intelligence, implement the relationship fusion judgment between the comparison groups. When it is determined that the relationship between the relationships is the same relationship, fuse them into the same relationship to further eliminate referential ambiguity.
[0074] Ambiguous data output: When it is impossible to determine whether two sets of relationships or two sets of entities are mutually convertible entities or relationships during the above-mentioned first relationship fusion, entity fusion, and second relationship fusion processes, the comparison group or comparison set is output to the ambiguous analysis database.
[0075] S6. Manual intervention
[0076] Manual judgment: The key entities and relationships for which artificial intelligence cannot determine whether they are mutually convertible entities or relationships enter the ambiguous analysis database. At this time, it reminds the human to intervene, and the human judges whether the entities or relationships that artificial intelligence cannot determine are mutually convertible; and then completes the fusion of all key entities and the disambiguation of all relationships.
[0077] Intelligent analysis of ambiguous data: Analyze the data in the differential database through the artificial intelligence system, compare all the information in the differential database, set the duplication rate standard, extract the data whose data duplication rate reaches the set duplication rate standard, and output this information to the ambiguous analysis database.
[0078] Manual analysis of ambiguous data: Manual intervention, manually judge the quality of the data and whether the content needs to be added to the construction of the knowledge graph. If so, perform entity extraction, relationship extraction, and knowledge fusion steps on the data, and merge all the key entities for fusion.
[0079] S7. Knowledge processing: The artificial intelligence system makes further inferences based on all the key entities and all the relationships that have been fused, and further expands the relationships between key entities based on the database information, thereby completing the supplementation of key entities or relationships.
[0080] S8. Construction of the first knowledge graph
[0081] Design the knowledge graph structure: Define the schema of the knowledge graph, including entity types, relationship types, and attribute types; determine the association relationships between entities, as well as the attributes of each entity and relationship.
[0082] Select and configure the graph database: Select the graph database algorithm for storing and managing the knowledge graph; configure the graph database environment and set the storage parameters and query parameters.
[0083] Construct the first knowledge graph: Import all the processed key entities and all the relationship data into the graph database to form the preliminary structure of the first knowledge graph; create entity nodes and relationship edges in the graph database, and add appropriate labels and attributes to the entity nodes and relationship edges, and save the first knowledge graph.
[0084] S9. Optimization of the first knowledge graph
[0085] Determine the number of connections: Take the number of the remaining key entities connected to a group of key entities in the first knowledge graph as the number of connections, and name the two parts of key entities as upper-layer key entities and lower-layer key entities. The number of connections is the number of lower-layer key entities connected to the upper-layer key entity;
[0086] For example, if there are 5 groups of lower-layer key entities connected to the upper-layer key entity, then the number of connections is 5;
[0087] Determine the minimum number of lower-layer key entities: Through extensive analysis of the original data by the artificial intelligence system, determine the minimum number of lower-layer key entities required by the upper-layer key entity;
[0088] Here, the minimum number of lower-layer key entities is calculated based on the retrieval of the overall data by the artificial intelligence model. The determined minimum number of lower-layer key entities is greater than or equal to 1, and the degree of simplification of the first knowledge graph is determined by the number of lower-layer key entities; for example, the determined minimum number of lower-layer key entities is 4 groups;
[0089] Determine the importance ratio of all lower-layer key entities: Extract and analyze all lower-layer key entities through the artificial intelligence system. Obtain the importance degree of each lower-layer key entity in the upper-layer key entity through the comprehensive analysis of the original data by the artificial intelligence system, and derive the importance ratio. Sort all lower-layer key entities based on the importance ratio. Obtain the repetition rate of each lower-layer key entity through the comprehensive analysis of the original data by the artificial intelligence system to determine the usage rate of the lower-layer key entity, and sort all lower-layer key entities based on the usage rate;
[0090] According to the above description, there are 5 groups of lower-layer key entities. Sort these 5 groups of lower-layer key entities based on the importance ratio, which are in turn: lower-layer key entity 1, lower-layer key entity 2, lower-layer key entity 3, lower-layer key entity 4, lower-layer key entity 5. Sort them based on the usage rate, which are in turn: lower-layer key entity 1, lower-layer key entity 3, lower-layer key entity 5, lower-layer key entity 2, lower-layer key entity 4;
[0091] Extract the main entities: Based on the importance calculated by the artificial intelligence model, propose a quantity. Determine the importance-based quantity according to the minimum number of lower-level key entities. Input the minimum number of lower-level key entities into the artificial intelligence model, and calculate the importance-based quantity through the artificial intelligence model. According to the above-mentioned lower-level key entities sorted by importance rate, propose the lower-level key entities whose importance rate is before the importance-based quantity as the first main entity; compare the first main entity with the lower-level key entities based on the repetition rate, and delete the lower-level key entities that repeat with the first main entity in the sorting of the lower-level key entities based on the usage rate. Determine the remaining main entities based on the artificial intelligence model. Input the usage rate and importance rate of the remaining lower-level key entities that have not been extracted as the first main entity into the artificial intelligence model, calculate the median proportion through the artificial intelligence model, sort the remaining lower-level key entities based on the median proportion, and select the second main entity according to the median proportion. The number of the second main entities is the minimum number of lower-level entities minus the number of the first main entities;
[0092] For example, if the importance-based quantity calculated by the artificial intelligence model is 3, then propose the lower-level key entity 1, the lower-level key entity 2, and the lower-level key entity 3 as the first main entities; then calculate the median proportion of the lower-level key entity 4 and the lower-level key entity 5, and sort them based on the median proportion. The sequence is: the lower-level key entity 5, the lower-level key entity 4; at this time, the number of the first main entities is 3 groups, and the minimum number of lower-level key entities is 4 groups. Therefore, select one group of second main entities according to the median proportion, and use the lower-level key entity 5 as the second main entity;
[0093] S10. Construction of the second knowledge graph: Delete the remaining key entities in the first knowledge graph except the first main entity and the second main entity to obtain the second knowledge graph; Establish a retrieval engine: Create a retrieval engine for the knowledge graph, and both the first knowledge graph and the second knowledge graph will pop up during retrieval;
[0094] For example, at this time, the lower-level key entity 1, the lower-level key entity 2, the lower-level key entity 3, and the lower-level key entity 5 are used as the first main entity and the second main entity respectively. Then the lower-level key entity 4 can be deleted from the first knowledge graph to obtain the second knowledge graph;
[0095] S11. Data update and coverage,
[0096] Set the update interval and adaptively set the data update time according to the industry development situation;
[0097] Retrieve new data. After reaching the set update time, the artificial intelligence system automatically retrieves relevant data to obtain information related to the keywords;
[0098] New data screening and application: The obtained new data is subjected to quality analysis. The data that meets the set quality standards enters the standard database. The new data is compared with the old data through the artificial intelligence system. When the new data and the old data are convertible data to each other, no subsequent update is performed. When the new data and the old data are not convertible data to each other, discrimination is carried out to determine whether the new data is newly added data or corrected data. When the artificial intelligence system determines that the new data is corrected data, the new data is processed according to the above steps and covers the original key entity or relationship to achieve the purpose of data update. When the artificial intelligence system determines that the new data is corrected data, the new data is processed according to the above steps and added to the first knowledge graph to achieve the purpose of expanding and improving the knowledge graph; and the second knowledge graph is modified according to the above steps S9 and S10.
[0099] 10. A system for constructing a knowledge graph using artificial intelligence, characterized in that the system includes:
[0100] A data collection module, which is used to collect preliminary data, store the corresponding data in the corresponding database after preliminary processing of the data;
[0101] A data preprocessing module, which is used to preprocess the data in the corresponding database;
[0102] A knowledge extraction module, which is used to extract entities and extract relationships, and make preliminary processing on the extracted entities and relationships;
[0103] A knowledge fusion module, which is used to judge, determine the quality, disambiguate, resolve anaphora and fuse the extracted entities and relationships;
[0104] An artificial processing module, which is used for human-computer interaction to facilitate manual judgment and screening of data, entities and relationships;
[0105] A knowledge processing module, which is used to reason, supplement and improve the construction of the knowledge system for data, entities and relationships;
[0106] A knowledge graph construction module, which is used to construct an education knowledge graph according to the identified key entities and the extracted relationships;
[0107] An optimization and export module, which is used to optimize the first knowledge graph, select important key entities, and generate a relatively concise knowledge graph.
[0108] The above are only embodiments of the present invention, and common general technical solutions and / or characteristics in the solution are not described in detail herein. It should be noted that for those skilled in the art, without departing from the technical solution of the present invention, several modifications and improvements can be made, which should also be regarded as the protection scope of the present invention, and these will not affect the implementation effect of the present invention and the practicability of the patent. The protection scope claimed in this application shall be subject to the content of its claims, and the specific implementation manners described in the specification can be used to interpret the content of the claims.
Claims
1. A method for constructing a knowledge graph using artificial intelligence, characterized in that: The following steps are involved: S1. Establish a database, input keywords, and collect all information related to the keywords based on artificial intelligence and big data to obtain preliminary data, which includes structured data, semi-structured data, and unstructured data. Different data types are classified and labeled; Data detection: Based on the artificial intelligence system, all retrieved data are quality-checked, quality standards are set, and a standard database and a differential database are established. Data that meets the quality standards are output to the standard database, and data that does not meet the quality standards are output to the differential database; S2, data preprocessing, preprocessing the text data in the standard database, S3, entity extraction, based on the artificial intelligence system to identify and extract key entities from the preprocessed data, and attach corresponding labels to each key entity after extracting the key entities for subsequent extraction and use; S4, relationship extraction, based on the key entities extracted in step S3, takes two different groups of key entities as a group of combinations, extracts the relationship between each group of combinations from the preprocessed data based on artificial intelligence, and annotates the relationship between them. S5. Knowledge fusion, Constructing comparison groups: All the key entities extracted above are individually labeled and named, and the key entities are labeled again, that is, the key entities are named entity 1, entity 2, ... entity n, and different key entities are grouped into comparison groups in pairs, and an ambiguous analysis database is established at the same time. Entity fusion: Based on artificial intelligence, the key entities of the comparison group are judged for entity fusion. When the comparison groups are judged to be the same key entities, they are judged as mutually transformed entities and merged into the same key entity. All comparison groups are judged and all mutually transformed entities are merged into the same key entity. Second relationship fusion: After mutually transformed entities are merged into the same key entity, the relationships between the key entities before fusion are also merged, and the merged relationships are given auxiliary naming labels, that is, the relationships are named relationship 1, relationship 2...relationship n, and different relationships are grouped into comparison groups in pairs. The relationship fusion judgment between the comparison groups is realized based on artificial intelligence. When the relationship between the relationships is judged to be the same, they are merged into the same relationship to further eliminate reference ambiguity; S6. Human intervention Manual judgment: The key entities and relationships that AI cannot judge whether they are mutually transformable entities or relationships enter the ambiguity analysis database. At this time, manual intervention is prompted to judge whether the entities or relationships that AI cannot judge are mutually transformable entities or relationships. This completes the fusion of all key entities and the fusion disambiguation of all relationships. Intelligent analysis of ambiguous data: Analyze the data in the difference database through the artificial intelligence system, compare the information in all the difference databases, set the repetition rate standard, extract the data whose repetition rate reaches the set repetition rate standard, and output the information to the ambiguous analysis database; S7. Knowledge processing: The artificial intelligence system performs further reasoning based on all key entities and all relationships that have been integrated, and further expands the relationships between key entities based on database information, thereby completing the supplementation of key entities or relationships. S8. Construction of the first knowledge graph, Design the knowledge graph structure: define the schema of the knowledge graph, including entity types, relationship types, and attribute types; Determine the relationships between entities and the attributes of each entity and relationship; Select and configure the graph database: select a graph database algorithm for storing and managing the knowledge graph; configure the graph database environment, and set storage parameters and query parameters; Construct the first knowledge graph: import all processed key entities and all relationship data into the graph database to form the preliminary structure of the first knowledge graph; create entity nodes and relationship edges in the graph database, add appropriate labels and attributes to the entity nodes and relationship edges, and save the first knowledge graph. S9. Optimization of the first knowledge graph, Determine the number of connections: The number of remaining key entities connected to a group of key entities in the first knowledge graph is taken as the number of connections, and the two parts of key entities are named upper key entities and lower key entities. The number of connections is the number of lower key entities connected to the upper key entity. Determine the minimum number of lower-level key entities: Extensive analysis of raw data by an AI system to determine the minimum number of lower-level key entities required for upper-level key entities; Determine the importance ratio of all lower-level key entities: extract and analyze all lower-level key entities through the artificial intelligence system, analyze all original data through the artificial intelligence system to obtain the importance of each lower-level key entity in the upper-level key entity, derive the importance rate, and sort all lower-level key entities based on the importance rate. Analyze all original data through the artificial intelligence system to obtain the repetition rate of each lower-level key entity to determine the utilization rate of the lower-level key entity, and sort all lower-level key entities based on the utilization rate. Extract backbone entities: calculate the number of proposed importance based on the artificial intelligence model, determine the number of proposed importance based on the minimum number of lower-level key entities, input the minimum number of lower-level key entities into the artificial intelligence model, calculate the number of proposed importance through the artificial intelligence model, and based on the lower-level key entities sorted according to the importance rate, propose the lower-level key entities with an importance rate before the number of proposed importance as the first backbone entity; compare the first backbone entity with the lower-level key entities based on the repetition rate, delete the lower-level key entities that are repeated with the first backbone entity in the sorting of lower-level key entities based on the usage rate, determine the remaining backbone entities based on the artificial intelligence model, input the usage rate and importance rate of the remaining lower-level key entities that have not been extracted as the first backbone entity into the artificial intelligence model, calculate the median proportion through the artificial intelligence model, sort the remaining lower-level key entities based on the median proportion, select the second backbone entity based on the median proportion, and the number of the second backbone entities is the minimum number of lower-level entities minus the number of the first backbone entities; S10. Constructing the second knowledge graph, deleting the remaining key entities in the first knowledge graph except the first backbone entity and the second backbone entity to obtain the second knowledge graph; establishing a search engine: creating a search engine for the knowledge graph, and the first knowledge graph and the second knowledge graph pop up simultaneously during the search.
2. A method for constructing a knowledge graph using artificial intelligence as claimed in claim 1, characterized in that: The data preprocessing includes the following steps: Data conversion: extract information from unstructured data and semi-structured data in standard databases and convert them into the same text data. Data supplementation: Delete noise data and redundant information in the data based on artificial intelligence system; Supplement data integrity by filling in missing information in the data based on artificial intelligence systems, while also identifying basic problems in the data such as misspellings, phonetic errors, and disordered text order.
3. The method for constructing a knowledge graph using artificial intelligence according to claim 1, characterized in that: The step S5 further comprises the following steps: First relationship fusion: Based on artificial intelligence, the relationship fusion judgment between the key entities of the comparison group is realized. When the relationship between the key entities of the comparison group is judged to be the same, it is determined to be a mutual conversion relationship and merged into the same relationship to eliminate the ambiguity of reference; The first relational integration was conducted after the construction of the comparison group.
4. The method for constructing a knowledge graph using artificial intelligence as claimed in claim 1, characterized in that: The step S5 further comprises the following steps: Ambiguous data output: When it is impossible to determine whether two groups of relationships or two groups of entities are mutually transformed entities or relationships during the first relationship fusion, entity fusion and second relationship fusion processes, the comparison group or comparison group will be output to the ambiguity analysis database; ambiguous data output is performed after the second relationship fusion step.
5. The method for constructing a knowledge graph using artificial intelligence as claimed in claim 1, characterized in that: The step S6 also includes the following steps: Manual analysis of ambiguous data: manual intervention is required to manually judge the quality of the data and whether the content needs to be included in the construction of the knowledge graph. If necessary, the data will be subjected to entity extraction, relationship extraction and knowledge fusion steps, and all fused data will be merged. Manual analysis of ambiguous data is performed after intelligent analysis of ambiguous data.
6. The method for constructing a knowledge graph using artificial intelligence as claimed in claim 1, characterized in that: The method further comprises S11. Data update coverage: By updating or covering the data in the knowledge graph, the knowledge graph can quickly and timely reflect new knowledge, new ideas and new discoveries.
7. A method for constructing a knowledge graph using artificial intelligence as claimed in claim 6, characterized in that: The data update covers: Set update interval: Set data update time according to industry development; Retrieve new data: After the set update time is reached, the artificial intelligence system automatically retrieves relevant data and obtains information related to the keyword; New data screening and application: Perform quality analysis on the new data obtained, and the data that meets the set quality standards will enter the standard database. The new data will be compared with the old data through the artificial intelligence system. When the new data and the old data are mutually converted data, no subsequent update will be performed. When the new data and the old data are not mutually converted data, it will be judged whether the new data is newly added data or corrected data. When the artificial intelligence system determines that the new data is corrected data, the new data will be processed according to the above steps and overwrite the original key entities or relationships to achieve the purpose of data updating. When the artificial intelligence system determines that the new data is corrected data, the new data will be processed according to the above steps and added to the first knowledge graph to achieve the purpose of expanding and improving the knowledge graph; and the second knowledge graph will be modified according to the above steps S9 and S10.
8. A system for constructing a knowledge graph using artificial intelligence, characterized in that: The system comprises: A data acquisition module, wherein the data acquisition module is used to collect preliminary data, perform preliminary processing on the data and store the corresponding data in a corresponding database; A data preprocessing module, used for preprocessing the data in the corresponding database; The knowledge extraction module is used to extract entities and relations, and perform preliminary processing on the extracted entities and relations; The knowledge fusion module is used to judge, characterize, disambiguate, resolve references, and fuse entities and relationships. The manual processing module is used for human-computer interaction to facilitate manual judgment and screening of data, entities and relationships; The knowledge processing module is used to reason and supplement data, entities and relationships to improve the construction of the knowledge system; A knowledge graph construction module, used to construct an education knowledge graph based on the identified key entities and the extracted relationships; The optimization and export module is used to optimize the first knowledge graph, delete important key entities, and generate a simpler knowledge graph.