A knowledge graph automatic construction method embedded in a scientific field knowledge system architecture
Patent Information
- Application Number
- CN202511593634.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-03
- Publication Date
- 2026-09-18
- Estimated Expiration
- 2045-11-03
AI Technical Summary
[0003]本申请解决了现有构建科学领域知识图谱的方法缺乏对科学领域知识体系指导、难以捕获深层认知关系的技术问题
本申请通过构建含知识架构编码器、层级关系建模器和语义推理引擎的知识体系底座模型,预处理科学文献后用该模型做多层级深度知识处理,再以知识图谱构建机制抽取实体、识别逻辑链关联,构建科学领域知识图谱,提升图谱的科学性与复杂科学推理支撑能力,达到了实现深层语义与关联关系处理,构建出符合科学认知规律、支持精准推理的科学领域知识图谱的技术效果。
Smart Images

Figure CN121478984B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of natural language processing technology, and in particular to a method for automatically constructing a knowledge graph embedded in a scientific knowledge system architecture. Background Technology
[0002] Knowledge graphs in the scientific domain are crucial for scientific research innovation and knowledge management, and their accurate construction is key to supporting complex scientific reasoning and knowledge application. Current technologies for knowledge graph construction mostly employ traditional data-driven methods, utilizing entity extraction and relation recognition techniques. While these methods have some effectiveness in general domains, they have limitations when applied to the scientific domain: they lack the embedded knowledge system architecture of the scientific domain and the deep semantic processing supported by AI computing models, making it impossible to accurately capture the hierarchical structure and deep relationships of scientific knowledge. This results in insufficient logical consistency and incomplete data in the constructed graphs, failing to meet the needs of accurate knowledge modeling and efficient reasoning in the scientific domain, and thus unable to provide reliable knowledge support for scientific research. Summary of the Invention
[0003] This application addresses the technical problems of existing methods for constructing knowledge graphs in scientific fields, which lack guidance on the knowledge system of scientific fields and struggle to capture deep cognitive relationships.
[0004] To address the aforementioned technical problems, this application proposes an automatic knowledge graph construction method embedding a scientific domain knowledge system architecture. The method includes: constructing a knowledge system foundation model based on a pre-trained large language model and a knowledge system fusion mechanism; the knowledge system foundation model includes a knowledge architecture encoder, a hierarchical relationship modeler, and a semantic reasoning engine; preprocessing embedded scientific literature; and using the knowledge system foundation model to perform multi-level deep knowledge processing on the preprocessed embedded scientific literature to obtain deep semantic analysis results. The multi-level deep knowledge processing includes semantic analysis, knowledge hierarchy identification and classification, and deep relational mining. Finally, using a knowledge graph construction mechanism, entity extraction and logical chain relational identification are performed on the deep semantic analysis results to construct a scientific domain knowledge graph.
[0005] This application proposes one or more technical solutions, which have at least the following technical effects: This application constructs a knowledge system foundation model containing a knowledge architecture encoder, a hierarchical relationship modeler, and a semantic reasoning engine. After preprocessing scientific literature, this model is used to perform multi-level deep knowledge processing. Then, a knowledge graph construction mechanism is used to extract entities and identify logical chain relationships to construct a scientific domain knowledge graph. This enhances the scientific rigor and ability to support complex scientific reasoning, achieving the technical effect of realizing deep semantic and relational processing and constructing a scientific domain knowledge graph that conforms to the laws of scientific cognition and supports accurate reasoning. Attached Figure Description
[0006] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0007] Figure 1 This is a flowchart illustrating an automatic knowledge graph construction method embedded in a scientific domain knowledge system architecture, as provided in an embodiment of this application.
[0008] Figure 2 This is a flowchart illustrating the process of multi-level deep knowledge processing in an automatic knowledge graph construction method embedded in a scientific domain knowledge system architecture, as provided in an embodiment of this application. Detailed Implementation
[0009] This application provides an automatic knowledge graph construction method that embeds a knowledge system architecture in a scientific field, solving the technical problems of existing methods for constructing knowledge graphs in a scientific field lacking guidance on the knowledge system of the scientific field and having difficulty capturing deep cognitive relationships.
[0010] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of them. All other embodiments obtained by those skilled in the art based on the embodiments of this application without creative effort are within the scope of protection of this application.
[0011] It should be noted that any variation of the terms "comprising" and "having" is intended to cover non-exclusive inclusion, for example, a process, method, system, product, or server that includes a series of steps or units is not necessarily limited to those steps or units that are explicitly listed, but may include other steps or modules that are not explicitly listed or that are inherent to such processes, methods, products, or devices.
[0012] like Figure 1 As shown, an automatic knowledge graph construction method embedding a scientific domain knowledge system architecture is provided, wherein the method includes: A knowledge system foundation model is constructed based on a pre-trained large language model and a knowledge system fusion mechanism. The knowledge system foundation model includes a knowledge architecture encoder, a hierarchical relationship modeler, and a semantic reasoning engine.
[0013] In this embodiment of the application, the knowledge system is a structured knowledge framework formed by the organic organization of concepts, theories, methods and laws in a specific scientific field according to logical relationships.
[0014] Specifically, the knowledge system foundation model first includes a basic knowledge system foundation model, which is built upon a knowledge architecture encoder, a hierarchical relationship modeler, and a semantic reasoning engine. The functions of each component are defined: the knowledge architecture encoder is responsible for learningable vector encoding; the hierarchical relationship modeler is used to identify hierarchical relationships, causal relationships, and logical dependencies; and the semantic reasoning engine performs semantic consistency checks and logical reasoning under domain knowledge constraints.
[0015] Next, the specific construction method of the knowledge system foundation model is clarified: the pre-trained large language model adopts 4-bit quantization and low-rank adaptive matrix update mechanism to complete the update of its own model parameters. Then, using this pre-trained large language model, a three-stage QLoRA domain knowledge fine-tuning is carried out on the basic knowledge system foundation model, and this three-stage fine-tuning specifically includes a classification system learning stage, a concept relationship modeling stage, and a reasoning ability enhancement stage.
[0016] Finally, when defining the foundational model of the knowledge system, the knowledge system fusion mechanism adopted is the RAG retrieval enhancement mechanism, which includes a knowledge retrieval module, a relevance evaluation module, and a knowledge fusion module. The knowledge retrieval module retrieves relevant knowledge fragments from a multi-source heterogeneous knowledge retrieval system. The relevance evaluation module scores and filters the retrieved relevant knowledge fragments based on semantic relevance, authority, and timeliness. The knowledge fusion module then enhances and fuses the filtered retrieved relevant knowledge fragments with the internal knowledge of the foundational knowledge system model.
[0017] Preprocessing embedded scientific literature involves using the knowledge system foundation model to perform multi-level deep knowledge processing on the preprocessed embedded scientific literature, resulting in deep semantic analysis results. The multi-level deep knowledge processing includes semantic analysis, knowledge level identification and classification, and deep relationship mining.
[0018] Optionally, the embedded scientific literature is first preprocessed. Then, using the attention mechanism of the knowledge system foundation model, semantic analysis of the preprocessed embedded scientific literature is performed on its conceptual hierarchy, logical relationships, and reasoning patterns to obtain semantically analyzed knowledge fragments. Subsequently, a graph neural network classification model is used to perform hierarchical recognition and classification of the semantically analyzed knowledge fragments, outputting the hierarchical recognition and classification results. Finally, high-order semantic relationship modeling is performed on the hierarchical recognition and classification results to obtain the deep semantic analysis results.
[0019] By utilizing knowledge graph construction mechanisms, entity extraction and logical chain relationship identification are performed on the deep semantic analysis results to construct a knowledge graph for the scientific domain.
[0020] In this embodiment, a knowledge graph is a data structure used to represent and organize knowledge. It describes the relationships between entities in the real world in the form of a graph. Its basic components include three core elements: entities (nodes), relationships (edges), and attributes. Entities can be any concrete or abstract thing such as people, places, items, or concepts. Relationships describe the connections between entities, while attributes provide detailed information about the entities.
[0021] In one embodiment of this application, it is first clarified that the knowledge graph construction mechanism includes three components: an entity extractor, a relation builder, and a graph structure verification optimizer. The entity extractor is used to extract entities from the deep semantic analysis results, the relation builder is used to establish the association between the entities extracted by the entity extractor, and the graph structure verification optimizer optimizes the knowledge graph in the scientific field by verifying the integrity of the topological structure and the semantic consistency.
[0022] Furthermore, the method provided in this application embodiment includes: The knowledge system foundation model includes a basic knowledge system foundation model, which is built based on the knowledge architecture encoder, hierarchical relationship modeler, and semantic reasoning engine. The knowledge architecture encoder is used to perform learnable vector encoding, the hierarchical relationship modeler is used to identify hierarchical relationships, causal relationships, and logical dependencies, and the semantic reasoning engine is used to perform semantic consistency checks and logical reasoning under domain knowledge constraints.
[0023] Specifically, the foundational knowledge system model includes a knowledge architecture encoder, a hierarchical relationship modeler, and a semantic reasoning engine. The specific construction process is as follows: When constructing a knowledge architecture encoder, the first step is to select a publicly available scientific classification system as the input data source, such as the Chinese Library Classification. The hierarchical structure data of this classification system is then organized to clarify the subject areas corresponding to the first to fifth level categories. Simultaneously, domain-specific terms under each category are collected to form a structured input dataset. For example, "fundamental theory of quantum mechanics" is associated with terms such as "quantum state" and "measurement theory" under this category. Next, a pre-trained language model embedding layer fine-tuning method, well-known to those skilled in the art, is employed. Using the publicly available SciBERT pre-trained model as a foundation, the organized structured input data is converted into a text pair format that the model can process. Using "category + term" as input samples, the model's embedding layer parameters are fine-tuned using gradient descent. This allows the model to learn the association between category hierarchy relationships and term semantics, ultimately enabling the knowledge architecture encoder to output learnable vectors containing category hierarchy attributes and term semantic information, thus completing the construction of the knowledge architecture encoder.
[0024] When constructing a hierarchical relationship modeler, the first step is to acquire publicly available domain ontology annotation datasets, such as the GO ontology dataset for biomedicine or the ChEBI ontology dataset for chemistry. Annotated samples in the format "Concept A - Relationship Type - Concept B" are extracted from these datasets, such as "DNA - Contains - Base," "Catalyst - Causality - Reaction Rate Increase," and "Animal - Hypothesis - Human," forming the training dataset for the hierarchical relationship modeler. Then, using the approach of text classification models, a multi-classification head is added to the model's output layer based on the publicly available RoBERTa pre-trained model. The categories of the classification heads correspond to three target relationship types: hypothesis, causal, and logical dependency. The labeled concept pairs from the training dataset are input into the model. The error between the model's prediction and the labeled relationship type is calculated using the cross-entropy loss function. The Adam optimizer is used to iteratively update the model parameters until the model's relationship recognition accuracy on the validation set stabilizes. At this point, the hierarchical relationship modeler can identify relationships between input concept pairs and output the confidence score for each relationship type, thus completing the construction of the hierarchical relationship modeler.
[0025] When building a semantic reasoning engine, the first step is to compile a domain constraint library based on publicly available scientific axioms and industry standards. The content of this library is presented in an IF-THEN rule format, such as "IF concept A is a subordinate concept of concept B, THEN concept A possesses the general attributes of concept B." Next, a reasoning module is built, employing a combination of rule matching and semantic similarity verification. First, the knowledge fragment to be reasoned is matched against the rules in the domain constraint library to check for logical conflicts. For example, if the fragment "enzymes do not lower the activation energy of chemical reactions" is matched against the rules "catalysts lower the activation energy of chemical reactions" and "enzymes are biological catalysts," a preliminary conflict is determined. Then, the publicly available Sentence-BERT model is used to calculate the semantic similarity between the knowledge fragment to be reasoned and related knowledge in the domain constraint library. If the similarity is below a preset threshold of 0.8, semantic inconsistency is further confirmed. Finally, the semantic reasoning engine outputs a report containing the reasoning conclusion and semantic consistency verification results, completing the construction of the semantic reasoning engine.
[0026] Finally, the completed knowledge architecture encoder, hierarchical relationship modeler, and semantic reasoning engine are integrated to build a basic knowledge system foundation model according to the functional logic of each component. The above steps, through the use of pre-trained model fine-tuning, classification model construction, and the combination of rule base and semantic verification, realize learnable vector encoding of scientific knowledge, hierarchical relationship recognition, and semantic reasoning under domain constraints, providing a knowledge system support that conforms to the laws of scientific cognition for subsequent knowledge graph construction.
[0027] Furthermore, the method provided in this application embodiment includes: The pre-trained large language model updates its model parameters based on a 4-bit quantization and low-rank adaptive matrix update mechanism. The pre-trained large language model is used to fine-tune the foundational knowledge system model in three stages using QLoRA domain knowledge. These three stages include a classification system learning stage, a concept relationship modeling stage, and a reasoning ability enhancement stage.
[0028] Optionally, the raw weight parameters of the publicly available Qwen3-8B pre-trained large language model are first loaded, and the model weights are quantized to 4 bits using the GPTQ quantization algorithm. Specifically, the 32-bit floating-point weight data in the model is mapped to a 4-bit integer range according to the GPTQ algorithm's quantization rules. Simultaneously, error correction parameters generated during quantization are calculated and stored to ensure that the model retains its core semantic understanding capabilities after quantization. This step yields the 4-bit quantized Qwen3-8B model, significantly reducing model memory usage.
[0029] Next, in the Transformer layer of the quantized Qwen3-8B model, low-rank adaptation matrices A and B are inserted into the weight matrices of the attention module and fully connected layer, respectively. The dimension of matrix A is set to the product of the hidden layer dimension and the low-rank dimension, with 8 chosen as the low-rank dimension. The dimension of matrix B is set to the product of the low-rank dimension and the hidden layer dimension. The parameters of matrices A and B are initialized using a random normal distribution, forming a Qwen3-8B model structure containing the low-rank adaptation matrices. The correction logic and optimization objective of the above low-rank adaptation matrices for the model's forward propagation can be clarified by the following formula: First, the corrected model forward propagation formula is: In the formula The corrected output representation of the i-th layer model. This represents the operation of dequantizing the 4-bit weights of the i-th layer into floating-point numbers. This represents the input of the i-th layer. This represents the output of the original (uncorrected) model at layer i. This represents the low-rank correction term introduced by QLoRA. and These are two defined projection matrices with dimensions respectively. and Where d is the hidden layer dimension of the model, r is the projection dimension, and Secondly, the optimization objective of model parameter updates is to minimize the loss of the new task, and the loss function formula is: Where L represents the loss function, used to measure the prediction error of the corrected model on the new task. This represents the original model parameters after quantization. This represents all projection matrices introduced by QLoRA, i.e., all low-rank adaptation matrix pairs introduced by QLoRA from layer 1 to layer L. D is the training dataset for the new task, containing the input x and the corresponding label y. l For task-related loss functions, This indicates the model's response to the input after adjusting for quantized original parameters and low-rank adaptation matrix. The predicted output, Indicates input The corresponding real labels. During the above optimization process, only the following are updated: and maintain constant.
[0030] Then, those skilled in the art need to prepare training data in the scientific field, including data from the Chinese Library Classification system, domain concept terminology data, and simple scientific knowledge relationship data. This data is then processed into a text sequence format acceptable to the pre-trained large language model. The processed training data is input into the Qwen3-8B model containing a low-rank adaptation matrix. The cross-entropy loss function is used to calculate the loss value between the model's prediction results and the labeled results. The parameters of the low-rank adaptation matrices A and B in the model are updated using the backpropagation algorithm, without updating the quantized original weight parameters. Iterative training continues until the loss value stabilizes and converges, completing the model parameter update.
[0031] Finally, the foundational knowledge system model is fine-tuned in three stages using a pre-trained large language model. The specific training content is as follows: the classification system learning stage is trained using graph classification data; the concept relationship modeling stage is trained around equivalence relations, inclusion relations, causal relations, and conditional relations; and the reasoning ability enhancement stage is trained by setting up scientific reasoning task samples. This makes the execution goal of the three-stage QLoRA domain knowledge fine-tuning clearer and more feasible. This step will be explained in detail in the following content.
[0032] By using the publicly available Qwen3-8B pre-trained large language model as a foundation, and combining 4-bit quantization processing and low-rank adaptation matrix parameter update methods, the goal is to reduce model memory consumption and computational costs while enabling the pre-trained large language model to initially adapt to scientific domain knowledge, thus laying the foundation for subsequent fine-tuning of the basic knowledge system base model.
[0033] Furthermore, the method provided in this application embodiment includes: The three-stage QLoRA domain knowledge fine-tuning includes a classification system learning stage, a concept relationship modeling stage, and a reasoning ability enhancement stage. The classification system learning stage includes training on graph classification data. The concept relationship modeling stage includes training on equivalence relations, inclusion relations, causal relations, and conditional relations. The reasoning ability enhancement stage includes training on scientific reasoning task samples.
[0034] Specifically, the hierarchical data of graph classification is first organized. The first to fifth level categories of graph classification in the scientific field are processed into text sequences according to the structure of "superior category - subordinate category - core category definition," such as "O Mathematical Sciences - O4 Physics - O413 Quantum Mechanics - Study of the laws of motion and interaction of microscopic particles," etc. These text sequences are then input into a pre-trained Qwen3-8B large language model that has undergone 4-bit quantization. Using the QLoRA fine-tuning method, the original model weights are fixed, and only the low-rank adaptation matrix parameters are updated. "Category-level semantic matching" is used as the training task. The cross-entropy loss function is used to calculate the error between the model's predicted category-level associations and the actual category-level relationships. The Adam optimizer iteratively adjusts the low-rank matrix parameters until the model's category-level recognition accuracy on the validation set stabilizes above a preset threshold. At the same time, the category semantic encoding parameters learned by the model after this stage of fine-tuning are synchronized to the knowledge architecture encoder of the basic knowledge system base model, so that when the knowledge architecture encoder generates learnable vectors, it can directly incorporate the hierarchical features of graph classification and output vector representations that are more in line with the classification logic of the scientific field.
[0035] Next, we collected concept relation samples from the scientific domain and extracted "concept pair-relation type" labeled data covering equivalence relations, inclusion relations, causal relations, and conditional relations from authoritative domain ontology databases. For example, "water-equivalence relation-" Examples such as "mammals - inclusion relationship - humans," "increased temperature - causal relationship - accelerated molecular motion," and "presence of oxygen - conditional relationship - combustion" are processed into a text format of "concept A - relationship description - concept B." These text samples are then input into the Qwen3-8B model, which has been fine-tuned in the previous stage. Further fine-tuning using QLoRA is performed, with "concept relationship classification prediction" as the training task. The model's output relationship type probability distribution and the loss value of the labeled results are calculated, and the low-rank adaptation matrix is updated. After training, the learned concept relationship recognition logic parameters are passed to the hierarchical relationship modeler of the foundational knowledge system model. This allows the hierarchical relationship modeler to more accurately determine the four target relationship types based on the fine-tuned parameters when receiving concept pair data, improving the accuracy of relationship modeling.
[0036] Subsequently, scientific reasoning task samples were designed, constructing a sample set according to the format of "scientific premise - reasoning question - correct conclusion," such as "Premise: The types of elements remain unchanged before and after a chemical reaction; Question: Hydrogen and oxygen react to produce water. Before the reaction, there are hydrogen and oxygen elements. After the reaction, what elements are contained in the water? Conclusion: Water contains hydrogen and oxygen elements." The samples cover multiple scientific fields such as physics, chemistry, and biology. Next, the reasoning task samples were input into the Qwen3-8B model, continuing the QLoRA fine-tuning mode, with "reasoning conclusion generation and correctness verification" as the training task. The semantic difference between the model's generated conclusion and the correct conclusion was calculated using the cosine similarity loss function, iteratively optimizing the low-rank matrix parameters. After fine-tuning, the reasoning logic parameters learned by the model were synchronized to the semantic reasoning engine of the basic knowledge system foundation model. This enables the semantic reasoning engine to efficiently complete semantic consistency checks and scientific question reasoning based on the fine-tuned parameters when performing logical reasoning under domain knowledge constraints, enhancing its reasoning ability.
[0037] By implementing QLoRA domain knowledge fine-tuning in three stages based on the Qwen3-8B pre-trained large language model, and linking and integrating the model parameters trained in each stage with the knowledge architecture encoder, hierarchical relationship modeler, and semantic reasoning engine of the basic knowledge system base model, the basic knowledge system base model was able to deeply adapt to the scientific domain classification rules, concept relationship features, and reasoning logic, and thus have a more accurate scientific knowledge processing capability.
[0038] Furthermore, the method provided in this application embodiment includes: The knowledge system fusion mechanism is a RAG retrieval enhancement mechanism, which includes a knowledge retrieval module, a relevance evaluation module, and a knowledge fusion module. The knowledge retrieval module is used to retrieve relevant knowledge fragments from a multi-source heterogeneous knowledge retrieval system. The relevance evaluation module is used to score and filter the retrieved relevant knowledge fragments according to semantic relevance, authority, and timeliness. The knowledge fusion module is used to enhance the fusion of the filtered retrieved relevant knowledge fragments with the internal knowledge of the basic knowledge system foundation model.
[0039] In this embodiment, the RAG retrieval enhancement mechanism is a technical mechanism that improves the accuracy, coverage, and reliability of inference by retrieving relevant knowledge fragments from multi-source heterogeneous knowledge resources, combining the model's filtering and fusion of retrieved knowledge, and supplementing the basic model with external knowledge.
[0040] Specifically, the RAG retrieval enhancement mechanism is first selected as the knowledge system fusion mechanism. The RAG retrieval enhancement mechanism includes a knowledge retrieval module, a relevance assessment module, and a knowledge fusion module. The specific construction process is as follows: First, the data from the multi-source heterogeneous knowledge retrieval system is processed. Different types of knowledge resources are collected, including academic journal articles, domain ontology libraries, and industry standard documents. Text parsing tools are used to convert documents in formats such as PDF and XML into a unified text format, while extracting core concepts and key sentences. Then, the vector encoding function of the knowledge architecture encoder in the foundational knowledge system model is invoked to convert the processed text knowledge into learnable vectors consistent with the encoder's output format, ensuring that the representation dimensions of external knowledge vectors and internal knowledge vectors in the foundational model are unified. Next, the FAISS vector retrieval tool is used to build a retrieval index based on the converted knowledge vectors. The vector dimension of the index is consistent with the output vector dimension of the knowledge architecture encoder. When retrieving relevant knowledge fragments, the query requirement is converted into a query vector by the knowledge architecture encoder, and then the FAISS tool performs an approximate nearest neighbor search in the index to match knowledge fragments semantically similar to the query vector, completing the construction of the knowledge retrieval module.
[0041] Next, a relevance assessment module was constructed. First, the calculation methods for the three scoring dimensions—semantic relevance, authority, and timeliness—were determined. For semantic relevance scoring, relevant knowledge fragment vectors output by the knowledge retrieval module were obtained and input along with the query vector into the knowledge architecture encoder. The encoder calculated the cosine similarity between the two, and the similarity value became the semantic relevance score, with a score range of 0 to 1. For authority scoring, a domain authority rating standard was established. Academic literature was assigned authority scores of 0 to 1 based on journal impact factor, ontology databases on domain recognition, and standard documents on the level of the publishing institution. The authority attributes of the corresponding knowledge resources were directly extracted and matched with the scores. For timeliness scoring, the publication or update time of the knowledge resource was obtained, and the interval between the current time and that time was calculated. The shorter the interval, the higher the timeliness score. This score was also normalized to a range of 0 to 1. Next, weights are set for three dimensions, such as semantic relevance 0.5, authority 0.3, and timeliness 0.2. These weights can be adjusted by those skilled in the art according to actual needs. The comprehensive score of each knowledge fragment is obtained by weighted summation, and knowledge fragments with a comprehensive score higher than the preset threshold of 0.6 are selected. At the same time, referring to the conceptual relationships identified by the hierarchical relationship modeler in the foundation knowledge system base model, if the selected knowledge fragment contains hierarchical or causal relationships with concepts within the base model, its comprehensive score weight can be appropriately increased to further optimize the selection results.
[0042] In the subsequent construction of the knowledge fusion module, the selected knowledge fragments are first subjected to entity and relation extraction. Entity links are used to match and align the entities in the fragments with the entities within the foundational knowledge system's base model. During matching, the relationship recognition function of the hierarchical relationship modeler is invoked to confirm the association type between the fragment entities and internal entities, such as inclusion, equivalence, and causality, ensuring the accuracy of entity alignment. Then, the aligned knowledge fragment vectors are fused with the corresponding entity vectors within the foundational knowledge system's base model using a weighted average method. The fusion weight is set based on the comprehensive score of the knowledge fragments; the higher the score, the greater the weight of the fragment vector. The fused vector serves as a supplementary update value for the entity vectors within the foundational knowledge system's base model. After fusion, the updated knowledge data is input into the semantic reasoning engine. The semantic reasoning engine performs a semantic consistency check under domain knowledge constraints to check whether the fused knowledge has logical conflicts with existing knowledge within the base model, such as contradictory conceptual relationships. If a conflict exists, the fusion result of that knowledge fragment is discarded; otherwise, the fused knowledge is updated to the internal knowledge base of the foundational knowledge system's base model, completing the construction of the knowledge fusion module.
[0043] By constructing a knowledge retrieval module, a relevance evaluation module, and a knowledge fusion module that are deeply associated with the knowledge architecture encoder, hierarchical relationship modeler, and semantic reasoning engine components of the basic knowledge system base model, and combining methods such as FAISS retrieval, cosine similarity calculation, and entity linking, the system achieves the effect of accurately supplementing the knowledge system base model with external multi-source knowledge and improving the knowledge coverage and accuracy of the base model.
[0044] Furthermore, such as Figure 2 As shown, the method provided in this application embodiment includes: The attention mechanism of the knowledge system foundation model is used to perform semantic analysis on the preprocessed embedded scientific literature in terms of concept hierarchy, logical relationship and reasoning pattern to obtain semantic analysis knowledge fragments; the semantic analysis knowledge fragments are classified hierarchically according to the graph neural network classification model to output the hierarchical identification and classification results; high-order semantic relationship modeling is performed on the hierarchical identification and classification results to obtain deep semantic analysis results.
[0045] Specifically, the scientific literature text is first preprocessed by using regular expressions to remove special symbols, redundant spaces, and meaningless characters. Then, terminology standardization is performed against a standard scientific terminology database, unifying synonyms and abbreviations into standardized expressions, such as standardizing DNA as deoxyribonucleic acid. Next, the knowledge architecture encoder in the knowledge system foundation model is invoked to split the standardized text into sentences. Each sentence is then converted into a fixed-dimensional vector, with the vector dimension matching the output dimension of the knowledge architecture encoder. This completes the preprocessing embedding, resulting in the embedded scientific literature text data.
[0046] Then, the embedded scientific literature text data is input into the knowledge system foundation model. Next, the knowledge architecture encoder is invoked to align the text vectors with the scientific domain concept hierarchy vectors stored in the knowledge architecture encoder, such as the category vectors of the Chinese Library Classification, ensuring that the text vectors match the concept representation logic of the foundation model. The attention weight of each word in the text is calculated based on the model's self-attention mechanism. During the calculation, the priority of the concept hierarchy output by the knowledge architecture encoder is considered, with words under core categories having higher base weights. Furthermore, the logical relationships identified by the hierarchical relationship modeler are combined, such as "cause" words in causal relationships and "superior concept" words in hierarchical relationships, adding a correlation coefficient during weight calculation. Simultaneously, based on the inference pattern rules built into the semantic reasoning engine, such as "premise" and "conclusion" keywords in the reasoning chain, higher weights are assigned. The specific calculation of attention weights requires first determining the attention score and then normalizing it using the softmax function. The core formula is as follows: First, the attention score calculation needs to incorporate knowledge system constraints, the formula is: Secondly, the normalized attention weights are calculated based on the scores, using the following formula: .in, This represents the attention score between position i and position j, used to quantify the strength of the semantic association between them. Indicates position Position Attention weights This represents the attention score between position i and position k, where n is the total length of the text sequence, i.e., the total number of positions involved in the attention calculation, and k represents the summation index, which iterates through all positions from 1 to n. This represents the attention computation function that incorporates the constraints of a knowledge system. This represents the text feature vector corresponding to position i, obtained by the model encoding the text at that position. This represents the text feature vector corresponding to position j. Indicates position The corresponding knowledge system constraint information.
[0047] The final attention weights are then obtained through weighted summation; higher weights indicate greater value of the words for semantic analysis. Based on these weights, core words and related phrases are selected from the text. During this selection process, a hierarchical relationship modeler is invoked to verify whether the relationships between words conform to scientific logic. For example, isolated words without clear relationships are excluded, and redundant modifying or supplementary content is filtered out. Then, based on semantic relevance and the relationship types identified by the hierarchical relationship modeler (e.g., causal, hyponymous, etc.), the selected content is divided into multiple fragments. Each fragment focuses on a set of related conceptual levels, logical relationships, or reasoning patterns, ultimately yielding semantic analysis knowledge fragments.
[0048] A graph neural network classification model was then constructed, using a graph convolutional network as the basic architecture. The model consists of an input layer, a graph convolutional layer, and an output layer. The input layer receives semantic fragment vectors, the graph convolutional layer aggregates features from neighboring nodes, and the output layer uses a softmax activation function to output the category level probability distribution. To train the model, a dataset was prepared by collecting data on the first to fifth levels of categories in the Chinese Graph Classification method and the hierarchical relationships between categories. Semantic fragments labeled with corresponding category levels were then selected from publicly available scientific literature databases. Each training sample consists of a semantic fragment vector and a category level label. During training, the semantic fragment vectors were input into the model's input layer, and the category hierarchical relationships were used as edge features of the graph structure. The graph convolutional layer calculated the feature interactions between nodes and their neighbors, and the output layer output the predicted category level probability. The cross-entropy loss function was used to calculate the error between the predicted value and the true label. The Adam optimizer was used to iteratively adjust the model parameters. After each training round, the accuracy was validated on a validation set until the accuracy stabilized above a preset value, completing the model training. When performing hierarchical recognition and classification, semantic analysis knowledge fragments are converted into vectors and input into the trained model. The model outputs the probability of the fragment corresponding to different category levels, and the category level with the highest probability is selected as the hierarchical recognition and classification result.
[0049] Finally, the concept and category information in the hierarchical identification and classification results are organized to construct triples in the form of "concept-relationship-concept", such as "quantum mechanics-belongs to-physics" and "Schrödinger equation-belongs to-quantum mechanics". A relation path mining method is used to extract multi-step association paths from the triples, such as "Schrödinger equation-belongs to-quantum mechanics-belongs to-physics". Simultaneously, cosine similarity calculation is used to compare the similarity of concept vectors at different levels to uncover implicit associations, such as the semantic association between "quantum entanglement" and "quantum state". The multi-step association paths and implicit associations are integrated to complete high-order semantic relationship modeling, resulting in deep semantic analysis results that include concept hierarchical affiliation, multi-step association paths, and implicit semantic associations.
[0050] Through the above steps, combined with the functions of the knowledge system foundation model, the goal is to achieve precise multi-level deep knowledge processing of pre-processed embedded scientific literature and output structured deep semantic analysis results.
[0051] Furthermore, the method provided in this application embodiment includes: The deep semantic analysis results are subjected to quality inspection and scoring, and a comprehensive quality score is output. The comprehensive quality score is a weighted calculation result of consistency score, integrity score, and reliability score. When the comprehensive quality score is less than a preset threshold, the deep semantic analysis results are reconstructed and optimized.
[0052] In one embodiment, the semantic reasoning engine of the knowledge system foundation model is first invoked to calculate the knowledge units in the deep semantic analysis results. The knowledge units in the scientific domain knowledge set KS stored internally within the knowledge system foundation model. similarity Through the formula: A consistency score is calculated; the higher the average similarity, the higher the consistency score. This represents the total number of knowledge units in the knowledge system KS, i.e., the total number of samples participating in the consistency comparison. This represents the i-th knowledge unit in the knowledge system KS, and... The objects to be compared for similarity. Representing knowledge units and The similarity.
[0053] Next, compare the standard attribute set of this type of knowledge unit in the scientific knowledge system. Statistical knowledge unit The actual set of attributes Through formula The completeness score is calculated based on attribute coverage; a higher attribute coverage results in a higher completeness score. Representing knowledge units The actual number of attributes, i.e., the number of attributes that have been extracted. Representing knowledge units The total number of attributes that should be present in a scientific knowledge system, i.e., the expected number of attributes.
[0054] Then check the knowledge units in the deep semantic analysis results. Source Through the formula: A reliability score is calculated, where Authority represents the source's authority, such as the journal's impact factor; Citation represents the number of citations received from the source. A higher average of both scores indicates a higher reliability score. Representing knowledge units Sources include authoritative scientific literature, industry standards, and domain ontology libraries.
[0055] The overall quality score is quantitatively calculated through a weighted summation, using the following formula: ,in, , , These are the weighting coefficients, and Then, those skilled in the art set weights for the different quality dimensions based on their knowledge of the scientific field. For example, the consistency score weight is set to 0.4, the integrity score weight is set to 0.3, and the reliability score weight is set to 0.3. The three scores are multiplied by their respective weights and then summed to obtain the overall quality score.
[0056] Next, the acquisition of the preset threshold needs to be based on historical qualified data. First, collect the deep semantic analysis results that have been manually reviewed and confirmed to meet the requirements of scientific knowledge over a period of time. Calculate the comprehensive quality score of these qualified results, take the average of the comprehensive quality scores of all qualified results, and then adjust it in combination with the minimum acceptable quality error range for scientific knowledge processing. For example, if the average comprehensive quality score of qualified results is 80 points, and considering that a certain error is allowed in actual processing, the threshold is lowered by 5 points, and the preset threshold is finally determined to be 75 points.
[0057] Subsequently, when the overall quality score is lower than a preset threshold, the reasons for the deduction of each individual score are analyzed. If the consistency score is low, the semantic reasoning engine locates the specific logical conflict point and corrects the conflicting content based on the scientific knowledge within the knowledge system's foundation model. If the completeness score is low, the knowledge retrieval module in the knowledge system fusion mechanism is invoked to supplement the retrieval of knowledge fragments corresponding to the uncovered categories, and the supplemented knowledge fragments are integrated into the results. If the reliability score is low, knowledge fragments from low-reliability sources are replaced with knowledge content from authoritative sources. After the correction is completed, the consistency score, completeness score, and reliability score are recalculated, and a new overall quality score is calculated according to the weights. This process is repeated until the overall quality score is not lower than the preset threshold, thereby completing the reconstruction and optimization of the deep semantic analysis results.
[0058] Furthermore, the method provided in this application embodiment includes: A knowledge graph construction mechanism is used to extract entities and identify logical chain relationships from the deep semantic analysis results. The knowledge graph construction mechanism includes an entity extractor, a relation builder, and a graph structure verification optimizer. The entity extractor is used to extract entities from the deep semantic analysis results, the relation builder is used to establish the relationships between the entities extracted by the entity extractor, and the graph structure verification optimizer is used to verify the integrity of the topological structure and semantic consistency to optimize the scientific domain knowledge graph.
[0059] Optionally, an entity extractor is first constructed by collecting annotated data from authoritative scientific literature, extracting entities such as physical laws, chemical substances, and biological species from the literature, and labeling the entity types to form an entity extraction training and validation set. A BERT pre-trained model is selected as the basic architecture, and an entity type classification layer is added to the model's output layer. The training set is input into the model, and the error between the predicted entity type and the labeled type is calculated using the cross-entropy loss function. The model parameters are iteratively adjusted using the Adam optimizer, and the entity recognition F1 score is calculated on the validation set after each training round. When the F1 score stabilizes, the entity prediction confidence score that results in the highest F1 score on the validation set is selected as the preset threshold, thus completing the entity extractor construction. This entity extractor can process text fragments from deep semantic analysis results and output entities with confidence scores higher than the preset threshold and their corresponding types.
[0060] Then, a relation builder is constructed, collecting entity pair association samples from the scientific field and labeling the relation types between entity pairs, such as inference, inclusion, and reaction, to form a relation building training set and a validation set. A Siamese network architecture is used, inputting the vectors obtained from the entity extractor for each entity pair into the network to calculate entity vector similarity. A relation type classification layer is added to the network output layer. The training set is input into the model, and the model parameters are optimized using the cross-entropy loss function. During training, the relation recognition accuracy is calculated on the validation set. The similarity value that achieves the highest accuracy on the validation set is selected as a preset threshold, thus completing the relation builder construction. The relation builder can receive entities output by the entity extractor, calculate the similarity between entities, and if the similarity is higher than the preset threshold, establish the corresponding association and output the relation type.
[0061] Next, a graph structure verification optimizer is constructed. This involves reviewing the topological rules of scientific knowledge graphs and clarifying the requirements for category hierarchy completeness. For example, under the Chinese Library Classification, a category must contain specified core sub-nodes. Scientific axioms and industry standards are compiled to form a semantic constraint library. A topological integrity scoring index is designed, which measures the percentage of nodes in the knowledge graph that conform to the topological rules. A semantic consistency scoring index is also designed, which measures the percentage of relations in the graph that conform to the rules of the semantic constraint library. Historical qualified scientific knowledge graph data is collected, and the topological integrity and semantic consistency scores of these graphs are calculated. The weighted sum of these two scores is taken as the average, and combined with the allowable error range in actual processing, a 5% reduction is used as the preset threshold for graph structure verification. This completes the construction of the graph structure verification optimizer. The graph structure verification optimizer can receive the initial graph output by the relation builder, calculate the topological integrity and semantic consistency scores, and if the combined score is lower than the preset threshold, it locates missing nodes or conflicting relations and optimizes them.
[0062] Finally, the completed entity extractor, relation builder, and graph structure verification optimizer are integrated to form a knowledge graph construction mechanism. The entity extractor extracts entities and types from the deep semantic analysis results, the relation builder establishes relationships between entities to form an initial scientific domain knowledge graph, and the graph structure verification optimizer verifies the topological integrity and semantic consistency of the scientific domain knowledge graph. Knowledge graphs that do not meet the requirements are optimized, and the final scientific domain knowledge graph is output.
[0063] Furthermore, the method provided in this application embodiment includes: Obtain a candidate entity set, calculate the annotation probability for each candidate entity in the candidate entity set, determine the type of the candidate entity set according to the annotation probability calculation result, perform attribute feature completion for each entity according to the type of the candidate entity set, and output the entity extraction result.
[0064] Optionally, the text content from the deep semantic analysis results is first split into sentences to obtain several text segments, and then a pre-trained BERT-NER model is used to process these text segments. The construction and training process of the BERT-NER model is as follows: The BERT-base pre-trained model, a general pre-trained model, was chosen as the basic architecture. This model already possesses general text semantic understanding capabilities, thereby reducing the initial cost of model training in the scientific field. Next, training data was prepared by selecting literature texts covering disciplines such as physics, chemistry, and biology from publicly available authoritative scientific literature databases. Target entities in the texts were manually annotated, specifically including the names of physical laws, molecular formulas and names of chemical substances, and scientific names of biological species. Annotation files were generated according to the format of "entity start position - entity end position - entity type," and the annotated data was divided into training, validation, and test sets in an 8:1:1 ratio.
[0065] During the training phase, the pre-defined training set is input into the BERT-base pre-trained model. A linear classification layer is added to the model's output layer to predict whether each token (smallest semantic unit) belongs to the target entity and its specific type. The cross-entropy loss function is used to calculate the error between the model's predictions and the manually labeled results. The Adam optimizer is used to adjust the model parameters, with a learning rate of 2e-5 and a batch size of 32. After each training round, the F1 score for entity recognition is calculated on the validation set. If the F1 score on the validation set does not improve for three consecutive rounds, training is stopped, the model parameters are saved, and the construction and training of the BERT-NER model are complete. The trained model needs to be validated on the test set to ensure that its F1 score is not lower than 0.85, to meet the accuracy requirements for entity extraction in scientific fields.
[0066] Next, each text segment is input into the model. The model outputs the probability of each word belonging to different entity types, filtering out consecutive text segments that the model classifies as entities to form a candidate entity set. Then, for each candidate entity in the candidate entity set, its overall annotation probability is calculated, and the average of the entity type probabilities corresponding to all words contained within that entity is taken as the annotation probability of that candidate entity. Subsequently, an annotation probability threshold of 0.7 is set, and the number of candidate entities in the candidate entity set with annotation probabilities higher than this threshold is counted. The type distribution of these high-probability entities is analyzed: if the proportion of high-probability entities of a certain entity type exceeds 60% of the total number of entities in the set, then the type of the candidate entity set is determined to be that entity type.
[0067] Next, based on the determined candidate entity set type, a pre-built scientific domain entity attribute library is invoked. This attribute library is categorized by entity type and stores the core attributes of each type of entity. For example, the attributes corresponding to the "physical laws" type are the proposer, formula expression, and scope of application, while the attributes corresponding to the "chemical substances" type are chemical formula, melting point, boiling point, etc. For each entity in the candidate entity set, the text content of the deep semantic analysis results is first checked to extract the entity attribute information already contained therein. If a certain attribute information is missing, for example, if a "chemical substance" entity does not mention a chemical formula, then a publicly available scientific domain database, such as the ChemSpider chemical database or the Physics Database, is invoked. Through precise matching of entity names, the missing attribute values corresponding to the entity are queried, and the queried attribute values are added to the entity information to complete the attribute feature completion for each entity. Finally, the entity extraction results containing the entity name, type, and complete attributes are output.
[0068] The BERT-NER model is used to obtain a set of candidate entities and calculate the labeled probabilities to determine the types. By combining the entity attribute library in the scientific field with public databases to complete the entity attribute features, the system can accurately identify entities in the scientific field, improve entity information, and output accurate and complete entity extraction results.
[0069] Furthermore, the method provided in this application embodiment includes: The entity extraction results are subjected to local relation identification, and a local relation network is output. The local relation network is subjected to inference path mining and the relationship strength between adjacent entities is calculated, and an indirect relation network is output. The local relation network and the indirect relation network are fused to construct a knowledge graph of the scientific domain.
[0070] Optionally, a rule base for common entity relationships in the scientific field is first compiled. This rule base contains explicit association patterns such as "physical law-derivation-formula," "chemical substance-reaction-product," and "biological species-inclusion-subspecies." Simultaneously, the vector representation of each entity in the entity extraction results from the previous steps is used. All entities in the entity extraction results are paired to form entity pairs. Each entity pair is then matched against the association patterns in the rule base. If an entity pair matches a certain association pattern, their local relationship is directly determined. If no rule is matched, the cosine similarity between the two entity vectors is calculated. A similarity threshold of 0.6 is set. If the similarity is higher than the threshold, a "relationship" type local relationship is determined. Finally, all determined entity pairs and their corresponding relationships are organized into a local relationship network with entities as nodes and relationships as edges.
[0071] Next, a breadth-first search algorithm is used to mine inference paths in the local relation network. Starting from each entity in the network, it sequentially searches for entities directly related to that entity (a 1-step path) and entities further related to the directly related entities (a 2-step path), mining indirect paths within a maximum of 3 steps. For example, starting from "Newton's Second Law," the path "derive-acceleration formula" and "association-force calculation" is used to find the entity representing the "unit of force" indirectly related to "Newton's Second Law." For each mined indirect path, the strength of the relationship between adjacent entities is calculated. Specifically, the confidence of each direct relationship in the local relation network is first obtained, which can be determined by the score during rule matching or similarity calculation. The confidence of all adjacent entity relationships in an indirect path is multiplied and then divided by the path length, i.e., the number of edges in the path, to obtain the indirect relationship strength of that path. A strength threshold of 0.4 is set, and indirect entity pairs and their corresponding relationships with strengths higher than the threshold are selected to form an indirect relation network.
[0072] Finally, the entity nodes in the local relation network and the indirect relation network are aligned to ensure that the same entity has the same identifier in both networks. The relation edges in the two networks are then integrated. If the same entity pair has the same relation in both networks, the edge with the higher confidence or strength is retained; if the same entity pair has different relations, both edges are retained; if an entity or relation exists only in one network, it is directly incorporated into the integrated network. After integration, the integrated entity and relation data are stored using the Neo4j graph database. Each entity is stored as a node in the database, containing information such as entity name, type, and attributes; each relation is stored as an edge, containing information such as relation type, confidence, or strength. This process ultimately constructs a knowledge graph for the scientific domain.
[0073] By combining rule matching with cosine similarity to identify local relationships, breadth-first search to mine indirect paths and calculate relationship strength, and integrating the network and storing it in a graph database, the goal of constructing a complete scientific domain knowledge graph containing both local direct and indirect relationships was achieved.
[0074] In summary, the method for automatically constructing a knowledge graph embedded in a scientific domain knowledge system architecture provided in this application has the following technical effects: This application utilizes a knowledge graph construction mechanism to process the results of deep semantic analysis. Through entity extraction, relational identification, and graph structure verification and optimization, entity and relational data are obtained. Adjustments are made based on the verification and optimization results, thereby accurately constructing a scientific domain knowledge graph. This makes the construction results of the scientific domain knowledge graph more accurate and reliable, achieving the technical effect of realizing deep semantic and relational processing and constructing a scientific domain knowledge graph that conforms to the laws of scientific cognition and supports accurate reasoning.
[0075] The above description of the disclosed embodiments enables those skilled in the art to make or use this application. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of this application. Therefore, this application is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.
[0076] Obviously, those skilled in the art can make various modifications and variations to this application without departing from the spirit and scope of this application. Therefore, if such modifications and variations fall within the scope of this application and its equivalents, this application also intends to include such modifications and variations.
Claims
1. A method for automatically constructing a knowledge graph embedded in a scientific domain knowledge system architecture, characterized in that, The method includes: A knowledge system foundation model is constructed based on a pre-trained large language model and a knowledge system fusion mechanism. The knowledge system foundation model includes a basic knowledge system foundation model, which includes a knowledge architecture encoder, a hierarchical relationship modeler, and a semantic reasoning engine. The knowledge architecture encoder uses a publicly available scientific domain classification system as its input data source. Based on the SciBERT pre-trained model, it uses categories and terms as input samples. The model embedding layer parameters are fine-tuned using gradient descent, allowing the model to learn the association between category hierarchy and term semantics. This results in a learnable vector for outputting category hierarchy attributes and term semantic information. The hierarchical relationship modeler uses concept A-relationship type-concept B format labeled samples extracted from the domain ontology labeled dataset as training samples. It is based on the RoBERTa pre-trained model and adds multi-class heads corresponding to the relationship types to the model output layer. The error is calculated by the cross-entropy loss function and the model parameters are iteratively updated by the Adam optimizer. The training results in a hierarchical relationship modeler that can identify hierarchical relationships, causal relationships and logical dependencies, and outputs the confidence level corresponding to each relationship type. The semantic reasoning engine is based on axioms and industry standards in the scientific field to form a domain constraint library in the IF-THEN rule format. The reasoning module adopts a combination of rule matching and Sentence-BERT semantic similarity verification to obtain a semantic consistency test and logical reasoning under domain knowledge constraints, and outputs a semantic reasoning engine containing reasoning conclusions and semantic consistency test result reports. The foundational knowledge system model is fine-tuned in three stages using a pre-trained large language model. The three stages of QLoRA domain knowledge fine-tuning include a classification system learning stage, a concept relationship modeling stage, and a reasoning ability enhancement stage. The knowledge system fusion mechanism is the RAG retrieval enhancement mechanism, which retrieves relevant knowledge fragments from a multi-source heterogeneous knowledge retrieval system and enhances the knowledge fusion with the internal knowledge of the basic knowledge system base model. Preprocessing embedded scientific literature, and using the knowledge system foundation model to perform multi-level deep knowledge processing on the preprocessed embedded scientific literature to obtain deep semantic analysis results, wherein the multi-level deep knowledge processing includes semantic analysis, knowledge level identification and classification, and deep relationship mining; By utilizing knowledge graph construction mechanisms, entity extraction and logical chain relationship identification are performed on the deep semantic analysis results to construct a knowledge graph for the scientific domain.
2. The method as described in claim 1, characterized in that, The knowledge architecture encoder is used to perform learnable vector encoding, the hierarchical relationship modeler is used to identify hierarchical relationships, causal relationships and logical dependencies, and the semantic reasoning engine is used to perform semantic consistency checks and logical reasoning under domain knowledge constraints.
3. The method as described in claim 2, characterized in that, Methods for constructing a foundational model for a knowledge system include: The pre-trained large language model updates its parameters based on a 4-bit quantization and low-rank adaptive matrix update mechanism.
4. The method as described in claim 2, characterized in that, Methods for constructing a foundational model for a knowledge system also include: The RAG retrieval enhancement mechanism includes a knowledge retrieval module, a relevance assessment module, and a knowledge fusion module; The knowledge retrieval module is used to retrieve relevant knowledge fragments from a multi-source heterogeneous knowledge retrieval system. The relevance evaluation module is used to score and filter the retrieved relevant knowledge fragments according to semantic relevance, authority, and timeliness. The knowledge fusion module is used to enhance the knowledge fusion of the filtered retrieved relevant knowledge fragments with the internal knowledge of the basic knowledge system base model.
5. The method as described in claim 1, characterized in that, The classification system learning phase includes training on graph classification data; the concept relationship modeling phase includes training on equivalence relations, inclusion relations, causal relations, and conditional relations; and the reasoning ability enhancement phase includes training on scientific reasoning task samples.
6. The method as described in claim 1, characterized in that, The method involves using the aforementioned knowledge system foundation model to perform multi-level deep knowledge processing on preprocessed embedded scientific literature, including: The attention mechanism of the knowledge system foundation model is used to perform semantic analysis on the preprocessed embedded scientific literature in terms of conceptual hierarchy, logical relationship and reasoning pattern to obtain semantic analysis knowledge fragments; The semantic analysis knowledge fragments are classified hierarchically according to the graph neural network classification model, and the hierarchical classification results are output. The hierarchical identification and classification results are modeled with higher-order semantic relationships to obtain deep semantic analysis results.
7. The method as described in claim 6, characterized in that, After obtaining the deep semantic analysis results, the method also includes: The deep semantic analysis results are subjected to quality inspection and scoring, and a comprehensive quality score is output. The comprehensive quality score is a weighted calculation result among consistency score, integrity score, and reliability score. When the overall quality score is less than a preset threshold, the deep semantic analysis results are reconstructed and optimized.
8. The method as described in claim 1, characterized in that, The deep semantic analysis results are used to extract entities and identify logical chain relationships using a knowledge graph construction mechanism, which includes an entity extractor, a relationship builder, and a graph structure verification optimizer. The entity extractor is used to extract entities from the deep semantic analysis results, the relation builder is used to establish the association relationships between the entities extracted by the entity extractor, and the graph structure verification optimizer is used to verify the topological integrity and semantic consistency of the scientific domain knowledge graph.
9. The method as described in claim 8, characterized in that, The method for entity extraction from the deep semantic analysis results includes: Obtain a candidate entity set, calculate the labeling probability for each candidate entity in the candidate entity set, and determine the type of the candidate entity set according to the labeling probability calculation results; The attribute features of each entity are completed according to the type of the candidate entity set, and the entity extraction results are output.
10. The method as described in claim 9, characterized in that, After the entity extraction results are output, logical chain relationships are identified. Methods include: Local relation identification is performed on the entity extraction results, and a local relation network is output. The local relationship network is subjected to inference path mining and the relationship strength between adjacent entities is calculated to output the indirect relationship network; By integrating the local relationship network and the indirect relationship network, a knowledge graph for the scientific domain is constructed.
Citation Information
Patent Citations
Medical knowledge relation extraction method and system based on large language model fine tuning and retrieval enhancement generation
CN118569263A
Knowledge graph construction method and system based on large language model technology
CN120523966A