Automatic construction method for mapping knowledge domain in geophysical field

Through multi-source data integration, knowledge extraction, fusion and optimization and iterative update methods, a geophysical knowledge graph is constructed, which solves the problems of low efficiency and poor accuracy of knowledge graph processing in the geophysical field in the existing technology, and achieves efficient, real-time and accurate knowledge sharing.

CN120338074APending Publication Date: 2025-07-18CHONGQING UNIVERSITY OF SCIENCE AND TECHNOLOGY
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510479215.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-16
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

The existing automated knowledge graph method is not suitable for geophysical fields with many relationships and complexity, resulting in low data processing efficiency and poor accuracy.

Method used

The methods of multi-source data integration, knowledge extraction, fusion and optimization, storage and iterative update are adopted to build a geophysical knowledge graph through the word vector model, syntax analysis, semantic role annotation, rule engine and machine learning model to ensure the comprehensiveness, timeliness and accuracy of the data.

Benefits of technology

It significantly improves the processing efficiency and accuracy of geophysical knowledge graphs, supports multilingual translation and cross-language information retrieval, ensures the real-time and accuracy of knowledge graphs, and promotes knowledge sharing among different languages.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120338074A_ABST
    Figure CN120338074A_ABST
Patent Text Reader

Abstract

The invention discloses an automatic construction method for a knowledge graph in the field of geophysics, and relates to the technical field of geophysics. According to the method, the comprehensiveness, timeliness and accuracy of the data can be ensured through data preparation and preprocessing, the quality of the data can be ensured, the unification of the data can also be ensured, and subsequent processing is facilitated; key information such as entity names and subject classification can be automatically extracted from massive texts through knowledge extraction, the processing efficiency is remarkably improved, multi-language translation and cross-language information retrieval are supported, communication and knowledge sharing among different languages are promoted, and the accuracy of data can also be guaranteed. Through knowledge fusion and optimization, similar entities can be merged, repeated information is eliminated, the accuracy of the knowledge graph is ensured, a rule engine or a machine learning model is utilized to verify the rationality of a relationship, and an implicit relationship is expanded according to domain knowledge, so that the accuracy of the knowledge graph is further improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of geophysics, and particularly relates to an automated construction method for a knowledge graph in the field of geophysics. Background Art

[0002] Geophysics is an interdisciplinary subject between physics, geology, atmospheric science, ocean science and astronomy. Its main research object is the earth where humans live and the space around it. It uses the principles and methods of physics, and through the use of advanced electronic and information technologies, aerospace technologies and space exploration technologies to observe various geophysical fields, so as to explore the medium structure, material composition, formation and evolution of the earth's interior and the space around it, the near-earth space, and study various natural phenomena related to it and their changing laws.

[0003] At present, there are many methods for automatically constructing knowledge graphs, but most of them are for extracting triple data of specified relationships. This method is not applicable to professional fields with more and more complex relationships. Therefore, we propose an automated construction method for a knowledge graph in the field of geophysics. Summary of the Invention

[0004] The purpose of the present invention is to provide an automated construction method for a knowledge graph in the field of geophysics to solve the problems mentioned in the above background art.

[0005] The present invention specifically adopts the following technical solutions to achieve the above purpose:

[0006] An automated construction method for a knowledge graph in the field of geophysics includes the following steps:

[0007] Step 1, data preparation and preprocessing: Obtain a dataset of geophysical knowledge and process it;

[0008] Step 2, knowledge extraction: Extract information on entities, attributes and relationships from the inside of the dataset. The extracted content includes named entities, relationships between entities and attributes of entities;

[0009] Step 3, knowledge fusion and optimization: Integrate knowledge from different sources, eliminate conflicting and duplicate information, and generate unified entities, attributes and relationships;

[0010] Step 4, knowledge storage: Store the fused knowledge in a knowledge base;

[0011] Step 5, generate a knowledge graph: Generate a knowledge graph based on the knowledge base that has completed storage;

[0012] Step 6, Iterative Update: Continuously update and maintain the entities, attributes, and relationships in the knowledge base, and continuously update and iterate the model based on user feedback, and the increase and update of the corpus.

[0013] Further, the data preparation and preprocessing include the following steps:

[0014] Step 11, Multi-source Data Integration: Collect geophysics-related data from channels such as scientific literature, databases, academic journals, and encyclopedias to establish an initial dataset;

[0015] Step 12, Text Cleaning and Annotation: Denoise, tokenize, part-of-speech tag, named entity recognition, and relationship preprocessing on the original data;

[0016] Step 13, Format Conversion and Standardization: Convert data from different sources into a unified format and perform standardized annotation using ontology language.

[0017] Further, the database in the multi-source data integration is based on geological survey data, and the data includes text descriptions, experimental data, and observation records.

[0018] Further, the named entity recognition includes mineral names and geographical locations, and the relationships include formation relationships and distribution relationships.

[0019] Further, the knowledge extraction includes the following steps:

[0020] Step 21, Language Processing: Extract entities and their relationships from the text through techniques such as word vector models, syntactic analysis, and semantic role annotation;

[0021] Step 22, Relationship Extraction and Indicator Matching: Use a predefined relationship indicator library or machine learning-based similarity calculation to identify the relationships between entity pairs and filter out noise data;

[0022] Step 23, Graph Model Construction: Convert the extracted entities and relationships into a graph structure, with nodes representing concepts and edges representing associations, initially forming the skeleton of the knowledge graph.

[0023] Further, the knowledge fusion and optimization include the following steps:

[0024] Step 31, Entity Alignment and Duplicate Removal: Merge similar entities through techniques such as ontology matching and clustering analysis to eliminate duplicate information;

[0025] Step 32, Relationship Verification and Extension: Use a rule engine or machine learning model to verify the rationality of the relationships and extend implicit relationships based on domain knowledge.

[0026] Further, the knowledge storage uses one of a relational database or a graph database.

[0027] Furthermore, the iterative update includes the following steps:

[0028] Step 61, data collection and incremental update: By monitoring the timestamps of data sources, identify newly added or modified entities and relationships, and only update the changed parts. Compare the old and new graph structures, and identify different nodes and edges through graph traversal algorithms for local update;

[0029] Step 62, entity alignment and fusion: Align entities in different knowledge graphs to the same semantic space through predefined rules;

[0030] Step 63, conflict resolution and knowledge integration: Learn the association weights between entities through the graph attention mechanism of GNN to achieve the collaborative representation of multi-source knowledge;

[0031] Step 64, online reasoning and optimization: Combine real-time data streams with offline batch processing to dynamically adjust the knowledge base structure to ensure the real-time performance of the system;

[0032] Step 65, quality evaluation and monitoring: Judge the integrity and accuracy of the knowledge graph through statistical analysis, use machine learning models to discover abnormal data, and automatically trigger the repair mechanism.

[0033] Furthermore, it also includes: a quality evaluation and feedback mechanism, which evaluates the quality of the knowledge graph through methods such as expert review and cross-validation, and establishes a feedback mechanism for continuous optimization.

[0034] The beneficial effects of the present invention are as follows:

[0035] 1. Through multi-source data integration, the present invention can ensure the comprehensiveness, timeliness, and accuracy of data; through text cleaning and annotation, denoising, word segmentation, part-of-speech tagging, named entity recognition, and relationships of the original data, the quality of data can be ensured; through format conversion and standardization, the unification of data can be ensured, which is conducive to subsequent processing.

[0036] 2. Through knowledge extraction, the present invention can automatically extract key information from a large amount of text, such as entity names, topic classifications, etc., significantly improve the processing efficiency, support multi-language translation and cross-language information retrieval, promote communication and knowledge sharing between different languages, and can also ensure the accuracy of data.

[0037] 3. Through knowledge fusion and optimization, the present invention can merge similar entities, eliminate duplicate information, ensure the accuracy of the knowledge graph, verify the rationality of relationships using a rule engine or machine learning model, and expand implicit relationships according to domain knowledge, further improving the accuracy of the knowledge graph.

[0038] 4. Through iterative updates, the present invention can ensure the timeliness of the entire knowledge graph. Meanwhile, when performing iterative updates, the graph attention mechanism of GNN is also used to learn the association weights between entities, realizing the collaborative representation of multi-source knowledge. Combining real-time data streams with offline batch processing, the knowledge base structure is dynamically adjusted to ensure the real-time performance of the system. By statistical analysis, the integrity and accuracy of the knowledge graph are judged, and machine learning models are used to discover abnormal data and automatically trigger the repair mechanism, thereby ensuring the practicality and real-time performance of the entire knowledge graph. By means of expert review, cross-validation, etc., the quality of the knowledge graph is evaluated, and a feedback mechanism is established for continuous optimization, which can ensure the timeliness and accuracy of the entire knowledge graph and guarantee the quality of the knowledge graph. BRIEF DESCRIPTION OF THE DRAWINGS

[0039] Figure 1 is the workflow diagram of the present invention;

[0040] Figure 2 is the workflow diagram of data preparation and preprocessing in the present invention;

[0041] Figure 3 is the workflow diagram of knowledge extraction in the present invention;

[0042] Figure 4 is the workflow diagram of knowledge fusion and optimization in the present invention;

[0043] Figure 5 is the workflow diagram of iterative update in the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0044] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention.

[0045] Please refer to Figure 1 - Figure 5 , the present invention provides an automated construction method for a knowledge graph in the geophysical field, including the following steps:

[0046] Step 1, data preparation and preprocessing: Obtain a dataset of geophysical knowledge and process it;

[0047] Step 2, knowledge extraction: Extract information on entities, attributes, and relationships from the inside of the dataset, and the extracted content includes named entities, relationships between entities, and attributes of entities;

[0048] Step 3, knowledge fusion and optimization: Integrate knowledge from different sources, eliminate conflicting and duplicate information, and generate unified entities, attributes, and relationships;

[0049] Step 4, Knowledge Storage: Store the fused knowledge in the knowledge base;

[0050] Step 5, Generate Knowledge Graph: Generate a knowledge graph based on the knowledge base that has completed storage;

[0051] Step 6, Iterative Update: Continuously update and maintain the entities, attributes, and relationships in the knowledge base, and continuously update and iterate the model based on user feedback, the increase and update of corpora.

[0052] In this embodiment, preferably, data preparation and preprocessing include the following steps:

[0053] Step 11, Multi-source Data Integration: Collect geophysics-related data from channels such as scientific literature, databases, academic journals, and encyclopedias to establish an initial dataset;

[0054] Step 12, Text Cleaning and Annotation: Denoise, tokenize, perform part-of-speech tagging, named entity recognition, and relationship preprocessing on the original data;

[0055] Step 13, Format Conversion and Standardization: Convert data from different sources into a unified format and perform standardized annotation using ontology language.

[0056] Through multi-source data integration, the comprehensiveness, timeliness, and accuracy of the data can be ensured; through text cleaning and annotation, denoising, tokenizing, part-of-speech tagging, named entity recognition, and relationships on the original data, the quality of the data can be ensured; through format conversion and standardization, the uniformity of the data can be ensured, which is conducive to subsequent processing.

[0057] In this embodiment, preferably, the database in multi-source data integration is based on geological survey data, and the data includes text descriptions, experimental data, and observation records.

[0058] In this embodiment, preferably, named entity recognition includes mineral names and geographical locations, and relationships include formation relationships and distribution relationships.

[0059] In this embodiment, preferably, knowledge extraction includes the following steps:

[0060] Step 21, Language Processing: Extract entities and their relationships from the text through techniques such as word vector models, syntactic analysis, and semantic role annotation;

[0061] Step 22, Relationship Extraction and Indicator Word Matching: Use a predefined relationship indicator word library or machine learning-based similarity calculation to identify the relationships between entity pairs and filter out noisy data;

[0062] Step 23, Graph Model Construction: Convert the extracted entities and relationships into a graph structure, where nodes represent concepts and edges represent associations, initially forming the skeleton of the knowledge graph.

[0063] Knowledge extraction can automatically extract key information from a vast amount of text, such as entity names, topic classifications, etc., significantly improving processing efficiency, supporting multilingual translation and cross - language information retrieval, promoting communication and knowledge sharing between different languages, and also ensuring the accuracy of data.

[0064] In this embodiment, preferably, knowledge fusion and optimization include the following steps:

[0065] Step 31, entity alignment and deduplication: By techniques such as ontology matching and clustering analysis, similar entities are merged, duplicate information is eliminated, and the accuracy of the knowledge graph is ensured;

[0066] Step 32, relationship verification and extension: Use a rule engine or a machine learning model to verify the rationality of relationships and extend implicit relationships based on domain knowledge.

[0067] Through knowledge fusion and optimization, similar entities can be merged, duplicate information can be eliminated, the accuracy of the knowledge graph can be ensured, and a rule engine or a machine learning model can be used to verify the rationality of relationships and extend implicit relationships based on domain knowledge, further improving the accuracy of the knowledge graph.

[0068] In this embodiment, preferably, the knowledge storage uses one of a relational database or a graph database.

[0069] In this embodiment, preferably, iterative update includes the following steps:

[0070] Step 61, data collection and incremental update: By monitoring the timestamps of data sources, newly added or modified entities and relationships are identified, and only the changed parts are updated. By comparing the old and new graph structures, different nodes and edges are identified through graph traversal algorithms for local update;

[0071] Step 62, entity alignment and fusion: Align entities in different knowledge graphs to the same semantic space through predefined rules;

[0072] Step 63, conflict resolution and knowledge fusion: Learn the association weights between entities through the graph attention mechanism of GNN to achieve the collaborative representation of multi - source knowledge;

[0073] Step 64, online reasoning and optimization: Combine real - time data streams with offline batch processing to dynamically adjust the knowledge base structure and ensure the real - time performance of the system;

[0074] Step 65, quality evaluation and monitoring: Judge the integrity and accuracy of the knowledge graph through statistical analysis, use a machine learning model to discover abnormal data, and automatically trigger a repair mechanism, thereby ensuring the practicality and real - time performance of the entire knowledge graph.

[0075] In this embodiment, preferably, it further includes: a quality assessment and feedback mechanism, which evaluates the quality of the knowledge graph through methods such as expert review and cross-validation, and establishes a feedback mechanism for continuous optimization, which can ensure the timeliness and accuracy of the entire knowledge graph and can ensure the quality of the knowledge graph.

[0076] The working principle and usage process of the present invention:

[0077] Step 1. Data preparation and preprocessing: Obtain a dataset of geophysical knowledge and process it; it includes the following steps:

[0078] Step 11. Multi-source data integration: Collect geophysics-related data from channels such as scientific literature, databases, academic journals, and encyclopedias to establish an initial dataset; the database is based on geological survey data, and the data includes text descriptions, experimental data, and observation records.

[0079] Step 12. Text cleaning and annotation: Denoise, tokenize, part-of-speech tag, named entity recognition, and relationship preprocessing are performed on the original data; named entity recognition includes mineral names and geographical locations, and relationships include formation relationships and distribution relationships.

[0080] Step 13. Format conversion and standardization: Convert data from different sources into a unified format and perform standardized annotation using ontology language.

[0081] Through multi-source data integration, the comprehensiveness, timeliness, and accuracy of the data can be ensured; through text cleaning and annotation, denoising, tokenization, part-of-speech tagging, named entity recognition, and relationships of the original data can be performed, which can ensure the quality of the data; through format conversion and standardization, the uniformity of the data can be ensured, which is conducive to subsequent processing.

[0082] Step 2. Knowledge extraction: Extract information on entities, attributes, and relationships from the inside of the dataset, and the extracted content includes named entities, relationships between entities, and attributes of entities;

[0083] It includes the following steps:

[0084] Step 21. Language processing: Extract entities and their relationships from the text through techniques such as word vector models, syntactic analysis, and semantic role annotation; key information such as entity names and topic classifications can be automatically extracted from a large amount of text, significantly improving the processing efficiency, supporting multilingual translation and cross-language information retrieval, and promoting communication and knowledge sharing between different languages.

[0085] Step 22. Relationship extraction and indicator word matching: Use a predefined relationship indicator word library or similarity calculation based on machine learning to identify the relationships between entity pairs and filter out noise data; it can ensure the accuracy of the data.

[0086] Step 23, Graph Model Construction: Convert the extracted entities and relationships into a graph structure, where nodes represent concepts and edges represent associations, initially forming the skeleton of the knowledge graph.

[0087] Knowledge extraction can automatically extract key information from a large amount of text, such as entity names, topic classifications, etc., significantly improving processing efficiency, supporting multilingual translation and cross-lingual information retrieval, promoting communication and knowledge sharing between different languages, and also ensuring the accuracy of data.

[0088] Step 3, Knowledge Fusion and Optimization: Integrate knowledge from different sources, eliminate conflicting and duplicate information, and generate unified entities, attributes, and relationships; it includes the following steps:

[0089] Step 31, Entity Alignment and Duplicate Removal: Merge similar entities through techniques such as ontology matching and clustering analysis, eliminate duplicate information, and ensure the accuracy of the knowledge graph;

[0090] Step 32, Relationship Verification and Extension: Use a rule engine or machine learning model to verify the rationality of relationships and extend implicit relationships based on domain knowledge.

[0091] Through knowledge fusion and optimization, similar entities can be merged, duplicate information can be eliminated, the accuracy of the knowledge graph can be ensured, the rationality of relationships can be verified using a rule engine or machine learning model, and implicit relationships can be extended based on domain knowledge, further improving the accuracy of the knowledge graph.

[0092] Step 4, Knowledge Storage: Store the fused knowledge in a knowledge base; use either a relational database or a graph database.

[0093] Step 5, Generate Knowledge Graph: Generate a knowledge graph based on the knowledge base that has been completed for storage; it also includes: a quality assessment and feedback mechanism, which evaluates the quality of the knowledge graph through methods such as expert review and cross-validation, and establishes a feedback mechanism for continuous optimization, which can ensure the timeliness and accuracy of the entire knowledge graph and can ensure the quality of the knowledge graph.

[0094] Step 6, Iterative Update: Continuously update and maintain the entities, attributes, and relationships in the knowledge base, and continuously update and iterate the model based on user feedback, the increase and update of corpus; it includes the following steps:

[0095] Step 61, Data Collection and Incremental Update: By monitoring the timestamps of data sources, identify newly added or modified entities and relationships, and only update the changed parts. Compare the old and new graph structures, identify different nodes and edges through graph traversal algorithms, and perform local updates;

[0096] Step 62, Entity Alignment and Fusion: Align entities in different graphs to the same semantic space through predefined rules;

[0097] Step 63, Conflict Resolution and Knowledge Fusion: Learn the association weights between entities through the graph attention mechanism of GNN to achieve the collaborative representation of multi-source knowledge;

[0098] Step 64, Online Inference and Optimization: Combine real-time data streams with offline batch processing to dynamically adjust the knowledge base structure and ensure the real-time performance of the system;

[0099] Step 65, Quality Evaluation and Monitoring: Judge the integrity and accuracy of the knowledge graph through statistical analysis, use machine learning models to discover abnormal data, and automatically trigger the repair mechanism, so as to ensure the practicality and real-time performance of the entire knowledge graph.

[0100] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present invention. Various modifications to these embodiments will be apparent to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to these embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. An automated construction method for a knowledge graph in the geophysical field, characterized in that It includes the following steps: Step 1, Data Preparation and Preprocessing: Obtain a dataset of geophysical knowledge and process it; Step 2, Knowledge Extraction: Extract information on entities, attributes, and relationships from within the dataset. The extracted content includes named entities, relationships between entities, and attributes of entities; Step 3, Knowledge Fusion and Optimization: Integrate knowledge from different sources, eliminate conflicting and duplicate information, and generate unified entities, attributes, and relationships; Step 4, Knowledge Storage: Store the fused knowledge in a knowledge base; Step 5, Generate Knowledge Graph: Generate a knowledge graph based on the knowledge base that has been completed for storage; Step 6, Iterative Update: Continuously update and maintain the entities, attributes, and relationships in the knowledge base. Based on user feedback, and the increase and update of corpus, continuously update and iterate the model.

2. The automated construction method for a knowledge graph in the geophysical field according to claim 1, wherein The data preparation and preprocessing include the following steps: Step 11, Multi-source Data Integration: Collect geophysical-related data from channels such as scientific literature, databases, academic journals, encyclopedias, etc., and establish an initial dataset; Step 12, Text Cleaning and Annotation: Denoise, tokenize, part-of-speech tag, named entity recognition, and relationship preprocessing on the original data; Step 13, Format Conversion and Standardization: Convert data from different sources into a unified format and perform standardized annotation using ontology language.

3. The automated construction method for a knowledge graph in the geophysical field according to claim 2, characterized in that: The database in the multi-source data integration is based on geological survey data, and the data includes text descriptions, experimental data, and observation records.

4. The automated construction method for a knowledge graph in the geophysical field according to claim 2, wherein: The named entity recognition includes mineral names and geographical locations, and the relationships include formation relationships and distribution relationships.

5. The automated construction method for a knowledge graph in the geophysical field according to claim 1, wherein The knowledge extraction includes the following steps: Step 21, Language Processing: Extract entities and their relationships from the text through techniques such as word vector models, syntactic analysis, and semantic role annotation; Step 22, Relationship Extraction and Indicator Matching: Use a predefined relationship indicator library or machine learning-based similarity calculation to identify relationships between entity pairs and filter out noisy data; Step 23, Graph Model Construction: Convert the extracted entities and relationships into a graph structure, where nodes represent concepts and edges represent associations, initially forming the skeleton of the knowledge graph.

6. The automated construction method for a knowledge graph in the geophysical field according to claim 1, characterized in that The knowledge fusion and optimization include the following steps: Step 31, Entity Alignment and Duplicate Removal: Merge similar entities through techniques such as ontology matching and clustering analysis to eliminate duplicate information; Step 32, Relationship Verification and Extension: Use a rule engine or machine learning model to verify the rationality of relationships and expand implicit relationships based on domain knowledge.

7. The automated construction method for a knowledge graph in the geophysical field according to claim 1, wherein: The knowledge storage uses one of a relational database or a graph database.

8. The automated construction method for a knowledge graph in the geophysical field according to claim 1, characterized in that The iterative update includes the following steps: Step 61, Data Collection and Incremental Update: By monitoring the timestamps of data sources, identify newly added or modified entities and relationships, and only update the changed parts. Compare the old and new graph structures, and identify different nodes and edges through graph traversal algorithms for local update; Step 62, Entity Alignment and Fusion: Align entities in different knowledge graphs to the same semantic space through predefined rules; Step 63, Conflict Resolution and Knowledge Fusion: Learn the association weights between entities through the graph attention mechanism of GNN to achieve the collaborative representation of multi-source knowledge. Step 64, Online Inference and Optimization: Combine real-time data streams with offline batch processing to dynamically adjust the knowledge base structure and ensure the real-time performance of the system; Step 65, Quality Evaluation and Monitoring: Judge the integrity and accuracy of the knowledge graph through statistical analysis, use machine learning models to detect abnormal data, and automatically trigger the repair mechanism.

9. The automated construction method for a knowledge graph in the geophysical field according to claim 1, wherein It also includes: Quality Assessment and Feedback Mechanism: Evaluate the quality of the knowledge graph through methods such as expert review and cross-validation, and establish a feedback mechanism for continuous optimization.

Citation Information

Cited By

  • Method and system for automatically extracting defense knowledge in unstructured text

    CN121351951A