Drug research and development knowledge base construction method and system
By combining the BERT model and RE model, data integration and timeliness in the drug development knowledge base are solved, efficient and accurate knowledge base construction and dynamic updates are achieved, and the latest and reliable knowledge support is provided for drug development.
Patent Information
- Application Number
- CN202510305653.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-14
- Publication Date
- 2025-06-27
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
It is difficult for existing drug R&D knowledge base construction technology to effectively integrate multi-source data, deal with data consistency and accuracy issues, and it is difficult to quickly respond to the update of new research results and clinical trial data, resulting in insufficient timeliness of the knowledge base.
The BERT model and RE model are used to combine the BERT model and RE model to perform entity annotation and relationship extraction on multi-source data, and calculate entity similarity through knowledge graph construction and embedding models, perform entity alignment and contradiction correction, and dynamically update and expand the knowledge base.
It realizes efficient, accurate construction and dynamic update of the drug research and development knowledge base, ensures the timeliness and completeness of the knowledge base, and provides the latest and reliable knowledge support for drug research and development.
Smart Images

Figure CN120218208A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of drug research and development, and specifically relates to a method and system for constructing a knowledge base for drug research and development. Background Art
[0002] With the rapid development of information technology and the continuous deepening of biomedical research, a vast amount of data related to drug research and development has emerged, such as scientific research literature, clinical trial reports, regulatory documents, and enterprise internal materials. These data contain rich knowledge and are of crucial value to all aspects of drug research and development, including drug target discovery, drug design, clinical trial design, and drug safety assessment.
[0003] However, in the data processing stage, the integration of multi-source data faces huge challenges. Due to the large differences in data formats, term expressions, and data quality among different data sources, there is a lack of effective preprocessing means to unify terms and eliminate the problem of ambiguous entity references. For example, in different literatures, the same drug may have multiple names, some using generic names and some using trade names, making it difficult to determine whether they refer to the same drug when integrating data; for some disease names, there are also different expression methods, which bring great difficulties to subsequent data processing and knowledge extraction, resulting in the inability to guarantee the consistency and accuracy of the data, seriously affecting the quality of knowledge base construction.
[0004] In terms of knowledge extraction, traditional methods mostly rely on entity annotation and relationship extraction techniques based on rules or simple machine learning algorithms. These methods often cannot fully understand the context information of the text and have limited capabilities in dealing with complex sentence structures and semantic relationships. For example, in the text describing the interaction between a drug and a target, the sentence structure may be very complex, containing multiple modifiers and nested relationships, and traditional methods are difficult to accurately identify the drug entity, target entity, and their interaction relationships, resulting in incomplete and inaccurate knowledge extraction, and there are a large number of information omissions and errors in the constructed knowledge graph, unable to provide comprehensive and reliable knowledge support for scientific research personnel.
[0005] New research results are emerging continuously, and clinical trial data is also being updated continuously. If the knowledge base cannot be updated in a timely manner, it will lead to knowledge lag, unable to provide the latest knowledge reference for drug research and development work, hindering the progress of drug research and development. At present, the knowledge base construction technology often stays at static data storage and is difficult to quickly respond to the release of new research results and the update of clinical trial data, resulting in insufficient timeliness of the knowledge base.
[0006] To this end, those skilled in the art have proposed a method and system for constructing a drug R & D knowledge base, aiming to provide a more efficient, accurate and dynamically updated way to construct a drug R & D knowledge base, and provide more reliable and comprehensive knowledge support for drug R & D. Summary of the Invention
[0007] To solve the above technical problems, the present invention provides a method and system for constructing a drug R & D knowledge base to solve the problems raised in the background technology.
[0008] According to the first aspect of the present disclosure, a method for constructing a drug R & D knowledge base is proposed, including the following steps:
[0009] S1. Identify and collect multi-source data including literature, clinical trial results, regulatory announcements and company annual reports, preprocess the data, and synchronously perform synonym mapping and entity disambiguation to obtain the preprocessed original text;
[0010] S2. Use the BERT model to perform entity annotation on the preprocessed original text, identify the key entities in the text to obtain the entity annotation result, input the entity annotation result into the RE model to extract the relationships between entities, and integrate the extracted entities and relationships into structured knowledge data to obtain the integrated knowledge data;
[0011] S3. According to the defined entities and relationships, and by the architecture and node connection method of the knowledge graph, fill the integrated knowledge data into the knowledge graph to form a drug R & D knowledge graph, reason about the entities and relationships of the drug R & D knowledge graph to obtain the reasoning result;
[0012] S4. Use the embedding model to calculate the similarity between entities for entity alignment, compare the reasoning result with the clinical trial data, correct the contradictions through a voting mechanism to obtain the verified graph, and store the verified graph in the main database to obtain the drug R & D knowledge base;
[0013] S5. Predict potential entity relationships through SWRL rules and GNN reasoning to dynamically update and expand the drug R & D knowledge base.
[0014] Preferably, the use of the BERT model to perform entity annotation on the original text includes:
[0015] By splitting the original text into a sequence of tokens, adding corresponding labels to each token, and inputting the processed token sequence into the pre-trained BERT model to obtain the context representation of each token, the input token sequence is X = [x1, x2,..., x n , and after BERT encoding, the hidden state representation of each token is obtained as H = [h1, h2,..., h n , where hi is the hidden state vector of the i-th token;
[0016] A fully connected layer is added to the output of the BERT model for classification, mapping the hidden state of each token to different entity categories, expressed as:
[0017] s i = Wh i + b
[0018] where W is the weight matrix of the BERT model, b is the bias vector, and s i is the classification score of the i-th token; perform a softmax operation on the classification score to obtain the probability distribution of each token belonging to different entity categories, and select the category with the highest probability as the predicted label of the token, expressed as:
[0019]
[0020] where C is the number of entity categories, is the predicted label of the i-th token.
[0021] Preferably, inputting the entity annotation result into the RE model to extract the relationship between entities includes:
[0022] According to the entity annotation result, extract all possible entity pairs from the text. If there are m entities in the entity annotation result, the number of entity pairs is For each entity pair, use its corresponding token sequence and entity information as input, and obtain the feature representation through encoding, that is, the feature vector corresponding to the entity pair (e1, e2) is f;
[0023] Input the feature vector f into the RE model, and obtain the probability distribution of the relationship categories between entity pairs through a fully connected layer and a softmax operation, expressed as:
[0024] r = W re f + b re
[0025] where W re is the weight matrix of the RE model, b re is the bias vector, and s i is the relationship score; the predicted relationship category is expressed as:
[0026]
[0027] where R is the number of relationship categories; according to the entity annotation result and the relationship extraction result, generate a triple set, and for each predicted entity pair (e1, e2) and its relationship generate a triple Store all the generated triples in a data structure to obtain the integrated knowledge data.
[0028] Preferably, the embedding model adopts the DistMult model. Calculating the similarity between entities includes:
[0029] Map both entities and relationships to vectors, and calculate the triples through the dot product of the vectors. Confidence, expressed as:
[0030]
[0031] Among them, the larger the dot product, the higher the similarity between entities e1 and e2 under the relationship ;
[0032] Through the set similarity threshold θ, for two entities e1 and e2, calculate their maximum similarity sim max (e1, e2) under all possible relationships. Then the entity alignment judgment is expressed as:
[0033]
[0034] Among them, align(e1, e2) = 1 indicates that e1 and e2 are aligned, and align(e1, e2) = 0 indicates that they are not aligned.
[0035] Preferably, comparing the inference result with the clinical trial data includes:
[0036] The inference result about entity e and relationship r obtained from the knowledge graph is R inf (e, r), and the actual result about entity e and relationship r obtained from the clinical trial data is R exp (e, r);
[0037] When the difference metric function Q(R inf (e, r), R exp (e, r)) exceeds the preset tolerance ε, it is considered that there is a contradiction. The difference metric function is expressed as: Q(R inf (e, r), R exp (e, r)) = |R inf (e, r) - R exp (e, r)|. The contradiction judgment is expressed as:
[0038]
[0039] Among them, conflict(e, r) = 1 indicates that there is a contradiction, and conflict(e, r) = 0 indicates that there is no contradiction;
[0040] For the entity-relationship pairs (e, r) with contradictions, collect n pieces of evidence E1, E2,..., E n , and each piece of evidence votes on the results O1, O2,..., O m to calculate the total number of votes A j for each result, expressed as:
[0041]
[0042] where a ij is the voting situation of the i-th piece of evidence for the j-th result, and a ij = 1 indicates voting, and a ij = 0 indicates not voting; select the result with the most total votes as the corrected result O corr , that is According to the entity alignment result, merge the aligned entities; according to the contradiction correction result, update the corresponding entity-relationship pairs in the knowledge graph to obtain the verified drug R & D knowledge graph.
[0043] According to the second aspect of the present disclosure, a drug R & D knowledge base construction system is also proposed, which adopts the construction method proposed in the first aspect, including:
[0044] A data acquisition and processing module for identifying and collecting multi-source data and preprocessing the collected data;
[0045] An entity and relationship extraction module for using the BERT model to perform entity annotation on the original text, identifying key entities in the text, and inputting the annotation results into the RE model to extract the relationships between entities, and integrating the entities and relationships into structured knowledge data;
[0046] A knowledge graph construction module for defining the architecture and node connection method of the knowledge graph according to the extracted entities and relationships, and filling the integrated knowledge data into the knowledge graph to form a complete drug R & D knowledge graph;
[0047] A graph verification and correction module for calculating the similarity between entities using an embedding model, performing entity alignment, comparing the inference results with clinical trial data, and correcting contradictions through a voting mechanism to obtain a verified knowledge graph;
[0048] A data storage module for storing the verified graph into the main database, marking the confidence level for each entity and relationship in the graph to obtain a drug R & D knowledge base;
[0049] A dynamic update module for predicting potential entity relationships through SWRL rules and GNN inference, and dynamically updating and expanding the drug R & D knowledge graph according to the inference results.
[0050] Compared with the prior art, the present invention has the following beneficial effects:
[0051] 1. By combining the use of the BERT model and the RE model, the present invention can accurately label entities and extract the relationships between entities, thereby constructing a more complete and reliable knowledge graph for drug research and development, which helps researchers more comprehensively understand the relevant knowledge of drug research and development and the relationships between them. During preprocessing, synonym mapping and entity disambiguation are performed simultaneously, which can effectively solve the problems of term inconsistency and entity reference ambiguity in multi-source data.
[0052] 2. By calculating similarities using the embedding model for entity alignment and comparing the inference results with clinical trial data to correct contradictions, the present invention ensures the quality of the knowledge graph and provides a reliable basis for drug research and development.
[0053] 3. By using SWRL rules and GNN inference to predict potential entity relationships and dynamically updating and expanding the knowledge graph for drug research and development according to the inference results, the present invention can quickly integrate new research results and data into the knowledge base, keeping the knowledge base timely and complete, providing the latest knowledge support for drug research and development, and helping to promote the continuous progress of drug research and development work. BRIEF DESCRIPTION OF THE DRAWINGS
[0054] Figure 1 is a flowchart of the method for constructing a knowledge base for drug research and development according to the present invention;
[0055] Figure 2 is a framework diagram of the system for constructing a knowledge base for drug research and development according to the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0056] The following further describes in detail the embodiments of the present invention in conjunction with the drawings. The following embodiments are used to illustrate the present invention, but cannot be used to limit the scope of the present invention.
[0057] As shown in the attached Figure 1 figure:
[0058] Embodiment 1: The present invention provides a method for constructing a knowledge base for drug research and development, including the following steps:
[0059] S1. Identify and collect multi-source data including literature, clinical trial results, regulatory announcements, and company annual reports, preprocess the data, and perform synonym mapping and entity disambiguation simultaneously;
[0060] Among them, synonym mapping constructs a synonym dictionary D through medical terms and their synonyms in the collected multi-source data. Each entry in the dictionary is a key-value pair, where the key is the standard term and the value is a set containing synonyms;
[0061] For each term t in the input text data, look it up in the thesaurus D. When t is a key in the thesaurus, uniformly replace all occurrences of the term and its corresponding value in the text with the standard term, i.e., the input text is T = {t1, t2,..., t n}, and the text after synonym mapping is T′ = {t′1, t′2,..., t′ n}; for each i ∈ [1, n], when t i ∈ the key set of D, then t′ i = t i ; when there exists a key k ∈ D such that t i ∈ D[k], i.e., t i is a synonym of a certain standard term, then t′ i = k; otherwise t′ i = t i ;
[0062] Among them, entity disambiguation is to represent each entity e in the text as a feature vector Determine the true entity set C of each entity according to domain knowledge and medical databases e , calculate the similarity between the feature vector of each candidate entity c ∈ C e and the entity e. When the feature vector of the candidate entity c then the cosine similarity between the entity e and the candidate entity c is expressed as:
[0063]
[0064] Select the candidate entity with the highest similarity as the true entity of the entity candidate word, that is, determine the true entity of the entity e as
[0065] S2. Use the BERT model to perform entity annotation on the original text. By splitting the original text into a sequence of tokens and adding corresponding labels to each token, input the processed token sequence into the pre-trained BERT model to obtain the context representation of each token. The input token sequence is X = [x1, x2,..., x n , and after BERT encoding, the hidden state representation of each token is obtained as H = [h1, h2,..., h n , where h i is the hidden state vector of the i-th token;
[0066] Add a fully connected layer on the output of the BERT model for classification, and map the hidden state of each token to different entity categories, expressed as:
[0067] s i = Wh i + b
[0068] Among them, W is the weight matrix of the BERT model, b is the bias vector, and s i is the classification score of the i-th token; perform a softmax operation on the classification score to obtain the probability distribution of each token belonging to different entity categories, and select the category with the highest probability as the predicted label of the token, which is expressed as:
[0069]
[0070] Among them, C is the number of entity categories, is the predicted label of the i-th token.
[0071] The entity annotation results are also input into the RE model to extract the relationships between entities. According to the entity annotation results, all possible entity pairs are extracted from the text. If there are m entities in the entity annotation results, the number of entity pairs is For each entity pair, its corresponding token sequence and entity information are used as inputs, and after encoding, a feature representation is obtained, that is, the feature vector corresponding to the entity pair (e1, e2) is f;
[0072] The feature vector f is input into the RE model, and through the fully connected layer and softmax operation, the probability distribution of the relationship categories between entity pairs is obtained, which is expressed as:
[0073] r = W re f + b re
[0074] Among them, W re is the weight matrix of the RE model, b re is the bias vector, and s i is the relationship score; the predicted relationship category is expressed as:
[0075]
[0076] Among them, R is the number of relationship categories; according to the entity annotation results and relationship extraction results, a triple set is generated. For each predicted entity pair (e1, e2) and its relationship generate a triple Store all the generated triples in a data structure to obtain the integrated knowledge data.
[0077] S3. Define according to the extracted entities and relationships, where the entity types include drug entities, target entities, disease entities, clinical trial entities, and research institution entities, and the relationship types between entities include drug-target relationships, drug-disease relationships, clinical trial-drug relationships, and research institution-drug relationships;
[0078] Based on the defined entity types and relationship types, determine the connection rules between nodes, and fill the integrated knowledge data into the knowledge graph. That is, for each entity in the triple, create a corresponding node in the knowledge graph, set the correct type label for the node according to the previously defined entity type, and create an edge between the corresponding nodes according to the relationship in the triple, and set the relationship type and attributes for the edge;
[0079] After filling the data, verify the knowledge graph. After verification, a complete knowledge graph for drug R & D is formed.
[0080] S4. Use the DistMult embedding model to calculate the similarity between entities, map both entities and relationships to vectors, and calculate the triple through the dot product of vectors Confidence, expressed as:
[0081]
[0082] The larger the dot product, the higher the similarity between entities e1 and e2 under the relationship ;
[0083] Through the set similarity threshold θ, for two entities e1 and e2, calculate their maximum similarity sim max (e1, e2) under all possible relationships, then the entity alignment judgment is expressed as:
[0084]
[0085] Among them, align(e1, e2) = 1 indicates that e1 and e2 are aligned, and align(e1, e2) = 0 indicates non-alignment.
[0086] The inference result is also compared with the clinical trial data. The inference result about entity e and relationship r obtained from the knowledge graph is R inf (e, r), and the actual result about entity e and relationship r obtained from the clinical trial data is R exp (e, r);
[0087] When the difference metric function Q(R inf (e, r), R exp (e, r)) exceeds the preset tolerance ε, it is considered that there is a contradiction. The difference metric function is expressed as: Q(R inf (e, r), R exp (e, r)) = |R inf (e, r) - R exp (e, r)|, and the contradiction judgment is expressed as:
[0088]
[0089] Among them, conflict(e, r) = 1 indicates the existence of a contradiction, and conflict(e, r) = 0 indicates the absence of a contradiction;
[0090] For the entity-relationship pair (e, r) with a contradiction, collect n evidences E1, E2,..., E n , and each evidence votes on the results O1, O2,..., O m to calculate the total number of votes A for each result j , expressed as:
[0091]
[0092] Among them, a ij is the voting situation of the i-th evidence for the j-th result, and a ij = 1 indicates voting, and a ij = 0 indicates not voting; select the result with the most total votes as the corrected result O corr , that is According to the entity alignment result, merge the aligned entities; according to the contradiction correction result, update the corresponding entity-relationship pairs in the knowledge graph to obtain the verified drug R & D knowledge graph.
[0093] S5. Predict potential entity relationships through SWRL rules and GNN reasoning. Among them, the SWRL rules are defined based on the professional knowledge and experience in the field of drug R & D. Apply the SWRL rules to the knowledge graph, and by matching the prerequisite conditions of the rules, find out the combinations of entities and relationships that meet the conditions, so as to deduce new potential relationships; among them, GNN represents the knowledge graph as a graph G = (V, E), where V is the set of nodes and E is the set of edges, and assign initial feature vectors to each node and edge;
[0094] GNN updates the feature representation of nodes through a message passing mechanism. In each layer, a node updates its own features according to the features of its neighbor nodes and the features of the edges, expressed as:
[0095]
[0096] Among them, is the feature vector of node v at the l-th layer, N(v) is the set of neighbor nodes of node v, W (l) and b (l) are the learnable parameters of the l-th layer, and σ is the activation function;
[0097] After several layers of message passing, use the final feature representation of the nodes for relationship prediction. Measure the possibility that there is a relationship r between nodes u and v through the scoring function g(u, v, r), and use the dot product as the scoring function, expressed as where h u and h v are the final feature vectors of nodes u and v, and W r is the learnable parameter matrix for relation r;
[0098] Add the potential relations obtained through SWRL rules and GNN reasoning to the knowledge graph. For each new relation (u, v, r), create a corresponding edge in the knowledge graph and update the attributes of the nodes and edges. As new data is continuously added and the inference results are updated, the knowledge graph is dynamically updated.
[0099] As can be seen from the above, by combining the BERT model and the RE model, entities can be accurately annotated and the relationships between entities can be extracted, thereby constructing a more complete and reliable knowledge graph for drug R & D, which helps researchers to more comprehensively understand the relevant knowledge of drug R & D and the relationships between them. And dynamically update and expand the knowledge graph of drug R & D according to the inference results, so as to quickly integrate new research results and data into the knowledge base, keep the knowledge base timely and complete, and provide the latest knowledge support for drug R & D.
[0100] As shown in the appendix Figure 2 as follows:
[0101] Embodiment 2: The present invention also provides a drug R & D knowledge base construction system, which is applied in Embodiment 1 and includes: a data collection and processing module for identifying and collecting multi-source data and preprocessing the collected data; the multi-source data includes data such as literature, clinical trial results, regulatory announcements, and company annual reports. Through data preprocessing, the quality and consistency of the data can be improved, and synonym mapping and entity disambiguation work are carried out synchronously during preprocessing;
[0102] An entity and relation extraction module for using the BERT model to perform entity annotation on the original text, identifying key entities in the text, and inputting the annotation results into the RE model to extract the relationships between entities, and integrating the entities and relationships into structured knowledge data;
[0103] A knowledge graph construction module for defining the architecture of the knowledge graph and the node connection method according to the extracted entities and relationships, and filling the integrated knowledge data into the knowledge graph to form a complete knowledge graph for drug R & D; through the graph, the relevant knowledge in the field of drug R & D and the relationships between them can be clearly displayed;
[0104] The atlas verification and correction module is used to calculate the similarity between entities using the embedding model, perform entity alignment to ensure that the entities in the atlas are accurate and consistent, compare the inference results with clinical trial data to verify the accuracy and reliability of the information in the atlas, and correct contradictions through a voting mechanism for information with contradictions or inconsistencies to obtain a verified knowledge graph;
[0105] The data storage module is used to store the verified atlas in the main database and mark the confidence level for each entity and relationship in the atlas;
[0106] The dynamic update module is used to predict potential entity relationships through SWRL rules and GNN inference, dynamically update and expand the drug R & D knowledge graph according to the inference results, so that the knowledge graph can timely reflect the latest information and research results in the field of drug R & D and maintain the timeliness and integrity of knowledge.
[0107] Importantly, it should be noted that the construction and arrangement of the present application shown in multiple different exemplary embodiments are only illustrative. Although only a few embodiments are described in detail in this disclosure, those who refer to this disclosure should easily understand that many modifications are possible without substantially departing from the novel teachings and advantages of the subject matter described in this application. Other substitutions, modifications, changes, and omissions can be made in the design, operating conditions, and arrangement of the exemplary embodiments without departing from the scope of the present invention. Therefore, the present invention is not limited to a specific embodiment, but extends to various modifications that still fall within the scope of the appended claims.
[0108] In addition, in order to provide a concise description of the exemplary embodiments, all features of the actual embodiments may not be described (i.e., those features that are not relevant to the currently considered best mode of implementing the present invention or those features that are not relevant to implementing the present invention).
[0109] It should be understood that in the development process of any actual implementation, a large number of specific implementation decisions can be made in any engineering or design project. Such development efforts may be complex and time-consuming, but for those ordinary technical personnel who benefit from this disclosure, without excessive experimentation, the development efforts will be a routine work of design, manufacturing, and production.
[0110] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the technical solutions of the present invention, and they should all be covered by the scope of the claims of the present invention.
Claims
1. A method for constructing a drug development knowledge base, characterized in that: The following steps are involved: S1. Identify and collect multi-source data including literature, clinical trial results, regulatory agency announcements and company annual reports, pre-process the data, and simultaneously perform synonym mapping and entity disambiguation to obtain the pre-processed original text; S2. Use the BERT model to perform entity annotation on the preprocessed original text, identify the key entities in the text, obtain the entity annotation results, and input the entity annotation results into the RE model to extract the relationship between entities, integrate the extracted entities and relationships into structured knowledge data, and obtain the integrated knowledge data; S3. According to the extracted entity and relationship definitions, the knowledge graph is populated with the integrated knowledge data based on the knowledge graph architecture and node connection method to form a drug development knowledge graph, and the entities and relationships of the drug development knowledge graph are inferred to obtain the inference results; S4. Use the embedding model to calculate the similarity between entities, perform entity alignment, compare the inference results with the clinical trial data, correct the contradictions through the voting mechanism, obtain the verified map, and store the verified map in the main database to obtain the drug development knowledge base; S5. Predict potential entity relationships through SWRL rules and GNN reasoning, and dynamically update and expand the drug development knowledge base.
2. A method for constructing a drug development knowledge base according to claim 1, characterized in that: The use of the BERT model to perform entity annotation on the original text includes: By dividing the original text into word-unit sequences and adding corresponding labels to each word-unit, the processed word-unit sequence is input into the pre-trained BERT model to obtain the contextual representation of each word-unit. The input word-unit sequence is X = [x1, x2, ..., x n ], after BERT encoding, the hidden state representation of each word is obtained as H = [h1,h2,...,h n ], where h i is the hidden state vector of the i-th word; Add a fully connected layer to the output of the BERT model for classification, mapping the hidden state of each word to different entity categories, expressed as: s i =Wh i +b Among them, W is the weight matrix of the BERT model, b is the bias vector, and s i is the classification score of the i-th word; perform softmax operation on the classification score to obtain the probability distribution of each word belonging to different entity categories, and select the category with the largest probability as the predicted label of the word, which is expressed as: Where C is the number of entity categories, is the predicted label of the i-th word.
3. A method for constructing a drug development knowledge base according to claim 1, characterized in that: The entity annotation results are input into the RE model to extract the relationship between entities, including: According to the entity annotation results, all possible entity pairs are extracted from the text. If the entity annotation results contain m entities, the number of entity pairs is For each entity pair, the corresponding word sequence and entity information are used as input, and the feature representation is obtained after encoding, that is, the feature vector corresponding to the entity pair (e1, e2) is f; The feature vector f is input into the RE model, and the probability distribution of the relationship category between entity pairs is obtained through the fully connected layer and softmax operation, which is expressed as: r=W re f+b re Among them, W re is the weight matrix of the RE model, b re is the bias vector, s i Score the relationship; predict the relationship category It is expressed as: Where R is the number of relationship categories; based on the entity labeling results and relationship extraction results, a set of triples is generated. For each predicted entity pair (e1, e2) and its relationship Generate a triple All generated triples are stored in a data structure to obtain integrated knowledge data.
4. A method for constructing a drug development knowledge base according to claim 1, characterized in that: The embedding model adopts the DistMult model, and the similarity between entities is calculated, including: Map both entities and relations into vectors and calculate triples by dot product of vectors Credibility, expressed as: The larger the dot product, the more entities e1 and e2 are in the relationship. The higher the similarity, By setting the similarity threshold θ, for two entities e1 and e2, the maximum similarity sim under all possible relationships is calculated. max (e1,e2), then the entity alignment judgment is expressed as: Among them, align(e1, e2) = 1 means that e1 and e2 are aligned, and align(e1, e2) = 0 means that they are not aligned.
5. A method for constructing a drug development knowledge base according to claim 1, characterized in that: The inference results are compared with clinical trial data, including: The inference result about entity e and relationship r obtained from the knowledge graph is R inf (e,r), the actual result obtained from the clinical trial data about entity e and relationship r is R exp (e,r); When the difference metric function Q(R inf (e,r),R exp (e,r)) exceeds the preset tolerance ε, it is considered that there is a contradiction, and the difference measurement function is expressed as: Q(R inf (e,r),R exp (e,r))=|R inf (e,r)-R exp (e,r)|, the contradictory judgment is expressed as: Where conflict(e,r)=1 means there is a conflict, conflict(e,r)=0 means there is no conflict; For contradictory entity-relation pairs (e,r), collect n pieces of evidence E1, E2, ..., E n , each piece of evidence has a significant effect on the results O1, O2, ..., O m Conduct a vote and calculate the total number of votes for each result A j , expressed as: Among them, a ij is the vote of the i-th evidence on the j-th result, and a ij =1 means voting, a ij =0 means no vote; the result with the largest total votes is selected as the revised result. corr ,Right now According to the entity alignment results, the aligned entities are merged; according to the contradiction correction results, the corresponding entity-relationship pairs in the knowledge graph are updated to obtain the verified drug development knowledge graph.
6. A drug development knowledge base construction system, using the drug development knowledge base construction method according to claims 1 to 5, characterized in that: include: Data acquisition and processing module, used to identify and collect multi-source data and pre-process the collected data; The entity and relationship extraction module is used to use the BERT model to perform entity annotation on the original text, identify key entities in the text, and input the annotation results into the RE model to extract the relationship between entities and integrate entities and relationships into structured knowledge data; The knowledge graph construction module is used to define the architecture and node connection mode of the knowledge graph based on the extracted entities and relationships, and fill the integrated knowledge data into the knowledge graph to form a complete drug development knowledge graph; The graph verification and correction module is used to use the embedding model to calculate the similarity between entities, perform entity alignment, compare the inference results with the clinical trial data, correct the contradictions through the voting mechanism, and obtain the verified knowledge graph; A data storage module is used to store the verified graph into the main database, mark the confidence level for each entity and relationship in the graph, and obtain a drug development knowledge base; The dynamic update module is used to predict potential entity relationships through SWRL rules and GNN reasoning, and dynamically update and expand the drug development knowledge graph based on the reasoning results.
Citation Information
Cited By
Inference method, system and equipment of wide constraint large language model and medium
CN121146067A
Relay protection defect diagnosis method and device based on dynamic knowledge graph
CN122064971A