A document-level biomedical relation extraction method based on prompt optimization model
Patent Information
- Application Number
- CN202411283594.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-13
- Publication Date
- 2026-09-29
- Estimated Expiration
- 2044-09-13
AI Technical Summary
但是,现有的基于提示学习的方法大多忽略了对实体对本身所蕴含事实知识的融合,此外也往往忽略了生物医学文献中部分冗余噪声对模型抽取效果的负面影响
[0030]与现有技术相比,本发明一方面设计了选区约束模块,通过新颖的选区约束方法可以有效地剔除冗余信息,简化实体间的复杂交互,从而使抽取过程聚焦于目标实体,提升关系抽取的质量;另一方面设计了知识优化提示模板构建模块,可利用实体类型知识来优化提示模板,使提示模板学习充足的特定实体对的事实知识,从而提高了提示模板的语义丰富性和文档级生物医学关系抽取的总体性能。总之,本发明能使文本关系信息集中,剔除冗余的负面信息,充分挖掘实体对本身包含的关系事实知识,实现复杂文本情况下文档级生物医学关系的精确抽取。实验结果表明,本发明在CDR和GDA两个数据集上表现出了卓越的竞争性能,有效提高了文档级生物医学关系抽取的总体精确度。
Smart Images

Figure CN119226510B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of biomedical natural language processing technology, and in particular to a document-level biomedical relation extraction method based on a prompting optimization model. Background Technology
[0002] In recent years, with the continuous deepening of biomedical research, the number of biomedical literatures has grown exponentially. These literatures contain a wealth of biomedical knowledge and are of great significance for advancing research and applications in the medical field. However, for researchers, reading such a large volume of literature presents enormous challenges. Therefore, automatically extracting structured knowledge from biomedical literature has become a crucial task, and biomedical relation extraction occupies a very important position in the field of biomedical natural language processing.
[0003] The goal of biomedical relation extraction is to identify the relationships between biomedical entity pairs from unstructured biomedical text. Currently, biomedical relation extraction techniques primarily focus on sentence-level or fragment-level relation extraction. While these techniques perform well in certain specific tasks, they often fall short when dealing with complex relationships across sentences or even paragraphs. Document-level biomedical relation extraction aims to extract intra-sentence and cross-sentence entity pair relationships from the perspective of the entire document, helping to improve the efficiency of information acquisition and processing. However, compared to sentence-level relation extraction, documents involve a larger number of entities, and the processing requires comprehensive consideration of contextual information and document structure to extract the complex logical relationships between entity pairs. This makes document-level biomedical relation extraction a more challenging and complex task.
[0004] Previous methods for document-level biomedical relation extraction have mainly fallen into two categories: sequential structure-based methods and graph structure-based methods. Sequential structure-based methods utilize deep neural networks (such as CNN, RNN, LSTM, and GRU) and attention mechanisms to capture long-range document information, thereby achieving document-level biomedical relation extraction. Graph structure-based methods construct document graphs by finding connections between keywords in documents to obtain global structural information, thus enabling document-level biomedical relation extraction. In recent years, many Transformer-based pre-trained language models have achieved state-of-the-art performance in capturing long-range text contextual information and have been widely applied in document-level biomedical relation extraction research. Cue-based learning methods have become a new paradigm for many natural language processing tasks, formalizing classification tasks as a masked language modeling problem.
[0005] Currently, in document-level biomedical relation extraction tasks, the mainstream methods are based on Transformer and inference techniques. However, these methods often rely excessively on complex downstream network architectures and contextual representations provided by pre-trained language models, thus affecting inference quality. In contrast, cue-based learning methods are more efficient in training and inference, require fewer parameters and computational resources, and offer greater flexibility and interpretability. However, most existing cue-based learning methods neglect the integration of factual knowledge inherent in the entity pairs themselves, and also often overlook the negative impact of redundant noise in biomedical literature on model extraction performance. Summary of the Invention
[0006] The present invention aims to solve the aforementioned technical problems existing in the prior art by providing a document-level biomedical relation extraction method based on a prompting optimization model.
[0007] The technical solution of this invention is: a document-level biomedical relation extraction method based on a prompt optimization model, comprising a selection constraint module, a knowledge optimization prompt template construction module, and a prompt relation prediction module, which are performed in the following steps:
[0008] Step 1. Obtain a relation instance text from the document to be extracted using the selection constraint module. Process the text with selection constraints to obtain the document sequence after selection constraints. The specific steps are as follows:
[0009] Step 1.1 Obtain a relation instance text from the document to be extracted, and determine the set of sentence positions p containing all mentions of the head entity. h The set of sentence positions p containing all mentions of the last entity. t The head entity and tail entity constitute the target entity pair that needs to be predicted in the document, i.e., each relation instance;
[0010] Step 1.2 Obtain the set p of all sentences containing information related to the head and tail entities according to formula (1);
[0011] p = p h ∪p t (1)
[0012] Step 1.3 uses set p as the selection constraint and retains only sentences in the document that exist in set p through a judgment mechanism, thereby obtaining the document sequence after selection constraint;
[0013] Step 2. Construct a knowledge optimization prompt template using the knowledge optimization prompt template construction module. The specific steps are as follows:
[0014] Step 2.1 Use ENTITY / TYPE MARKER to index the beginning and end entity positions of the document sequence after selection constraint. That is, insert special markers around the target entities in the document and add markers at the beginning of the text to obtain the text sequence after adding markers, denoted as x.
[0015] Step 2.2 Define the prompt template as shown in (2):
[0016] T(·)=“·The relations between ET .h and ET ·t is[MASK].” (2)
[0017] Among them, ET ·h and ET ·t These represent the types of the head entity and the tail entity, respectively.
[0018] Map x to T(·) to obtain the knowledge optimization prompt template as shown in (3);
[0019] x prompt =T(x)=“x The relations between ET xh and ET xt is[MASK].” (3)
[0020] The [MASK] is a mask marker;
[0021] Step 3. Complete the relationship instance classification using the prompt relationship prediction module. The specific steps are as follows:
[0022] Step 3.1 Build the completed knowledge optimization hint template x prompt Input a pre-trained language model T5 based on the Transformer architecture, and compute the hidden layer vector representation h at the mask marker [MASK] position. [MASK] ;
[0023] Step 3.2 Generate the probability distribution of each possible tag word at the mask position, that is, perform relation prediction through linear transformation and softmax function, thereby completing the relation instance classification, that is, obtaining the relation type of the document relation instance;
[0024] The linear transformation is shown in (4):
[0025] z = h [MASK] ·W+b (4)
[0026] Where W is the weight matrix in the output layer, and b is the bias vector in the output layer;
[0027] The softmax function, as shown in formula (5), transforms the unnormalized score vector z into a probability distribution, thereby generating the probability distribution P(y|x) of each possible tag word at the mask position.
[0028] P(y|x)=P([MASK]=v|x prompt = softmax(z) (5)
[0029] in, It is a collection of all tag words.
[0030] Compared with existing technologies, this invention, on the one hand, designs a selection constraint module. Through a novel selection constraint method, redundant information is effectively eliminated, and complex interactions between entities are simplified, thus focusing the extraction process on the target entities and improving the quality of relation extraction. On the other hand, it designs a knowledge-optimized prompt template construction module. This module utilizes entity type knowledge to optimize the prompt template, enabling it to learn sufficient factual knowledge about specific entity pairs, thereby improving the semantic richness of the prompt template and the overall performance of document-level biomedical relation extraction. In summary, this invention enables the concentration of textual relational information, eliminates redundant negative information, fully mines the relational factual knowledge contained within entity pairs, and achieves accurate extraction of document-level biomedical relations in complex text scenarios. Experimental results show that this invention exhibits superior competitive performance on both the CDR and GDA datasets, effectively improving the overall accuracy of document-level biomedical relation extraction. Attached Figure Description
[0031] Figure 1 This is a schematic diagram of the overall architecture of an embodiment of the present invention. Detailed Implementation
[0032] This invention provides a document-level biomedical relation extraction method based on a prompting optimization model, as follows: Figure 1 As shown, it consists of a selection constraint module, a knowledge optimization prompt template construction module, and a prompt relationship prediction module, and is performed in the following steps:
[0033] Step 1. Obtain a relation instance text from the document to be extracted using the selection constraint module. Process the text with selection constraints to obtain the document sequence after selection constraints. The specific steps are as follows:
[0034] Step 1.1 Obtain a relation instance text from the document to be extracted, and determine the set of sentence positions p containing all mentions of the head entity. h The set of sentence positions p containing all mentions of the last entity. t The head entity and tail entity constitute the target entity pair that needs to be predicted in the document, i.e., each relation instance; the mentions are multiple occurrences of each entity in the document;
[0035] Step 1.2 Obtain the set p of all sentences containing information related to the head and tail entities according to formula (1);
[0036] p = p h ∪p t (1)
[0037] Step 1.3 uses set p as the selection constraint and retains only sentences in the document that exist in set p through a judgment mechanism, thereby obtaining the document sequence after selection constraint;
[0038] Step 2. Construct a knowledge optimization prompt template using the knowledge optimization prompt template construction module. The specific steps are as follows:
[0039] Step 2.1 Use ENTITY / TYPE MARKER to index the beginning and end entity positions of the document sequence after selection constraints. That is, insert special markers [E] and [ / E] around the target entities in the document and add the marker [CLS] at the beginning of the text to obtain the text sequence after adding the markers, denoted as x, such as [CLS]...[E2]vocal fold palsy[ / E2]...[E1]disulfiram[ / E1]..., [E1][ / E1] marks the beginning entity, and [E2][ / E2] marks the end entity;
[0040] Step 2.2 Define the prompt template as shown in (2):
[0041] T(·)=“·The relations between ET ·h and ET ·t is[MASK].” (2)
[0042] Among them, ET ·h and ET ·t These represent the types of the head entity and the tail entity, respectively.
[0043] Map x to T(·) to obtain the knowledge optimization prompt template as shown in (3);
[0044] x prompt =T(x)=“x The relations between ET xh and ET xt is[MASK].” (3)
[0045] The [MASK] is a mask marker; in addition to the original marker, the prompt template also contains at least one mask marker [MASK], which is used to guide the pre-trained language model to predict the label relationship at the mask position. The constructed optimized template introduces additional entity type knowledge into the natural language text prompt to improve performance.
[0046] Step 3. Complete the relationship classification using the prompt relationship prediction module. The specific steps are as follows:
[0047] Step 3.1 Build the completed knowledge optimization hint template x prompt Input a T5 (Text-to-Text Transfer Transformer) language model pre-trained based on the Transformer architecture, and compute the hidden layer vector representation h at the mask marker [MASK] position. [MASK] ;
[0048] The T5 is a pre-trained language model (Text-to-Text Transfer Transformer) based on the Transformer architecture proposed by Google.
[0049] Step 3.2 Generate the probability distribution of each possible tag word at the mask position, that is, perform relation prediction through linear transformation and softmax function, thereby completing relation classification and obtaining the relation type of the document relation instance;
[0050] The linear transformation is shown in (4):
[0051] z = h [MASK] ·W+b (4)
[0052] Where W is the weight matrix in the output layer, and b is the bias vector in the output layer;
[0053] The softmax function, as shown in formula (5), transforms the unnormalized score vector z into a probability distribution, thereby generating the probability distribution P(y|x) of each possible tag word at the mask position.
[0054] P(y|x)=P([MASK]=v|x prompt = softmax(z) (5)
[0055] in, It is a collection of all tag words.
[0056] To verify the effectiveness of this invention, experiments were conducted on two widely used document-level biomedical relation extraction datasets, CDR and GDA, based on the method provided herein. The Biocreative V Chemical Disease Relation benchmark (CDR) is a human-annotated relation extraction dataset in the biomedical field, composed of PubMed summaries, containing relationships between the concepts of "chemical" and "disease": chemical-induced disease (CID). The Gene Disease Associations (GDA) is a large-scale relation extraction dataset in the biomedical field, whose task is to predict the interaction between the concepts of "gene" and "disease." Statistical data for both datasets are shown in Table 1.
[0057] Table 1. Statistics of CDR and GDA datasets
[0058]
[0059] The evaluation criteria used in the experiment were P, R, and F1 scores, where P is precision, R is recall, and F1 score is a comprehensive evaluation metric for general classification problems. In addition, Intra-F1 and Inter-F1 metrics were used to evaluate the model's performance on intra-sentence and inter-sentence relationships.
[0060] The performance comparison results between the method of this invention and existing methods are shown in Tables 2 and 3. The results show that the method of this invention comprehensively outperforms existing methods. On the CDR dataset, compared with the best-performing Transformer-based DocuNet model, the method of this invention improves the overall F1 score by 1.30%; compared with the best-performing cue-based CPT-RI model, the method of this invention improves the overall F1 score by 6.00%. Furthermore, compared with the CGM2IR Transformer-based model, the method of this invention improves the Intra-F1 and Inter-F1 scores by 0.87% and 14.76%, respectively.
[0061] Table 2 shows the experimental results of this invention on the CDR dataset.
[0062]
[0063] Note: "-" in this table refers to data not given in the original paper.
[0064] The method presented in this invention demonstrates superior overall performance on the GDA dataset, achieving P, R, and F1 scores of 84.76%, 87.20%, and 85.80%, respectively, outperforming most existing models and reaching a state-of-the-art level. It is noteworthy that the GDA dataset is nearly 50 times larger than the CDR dataset, thus exhibiting lower sensitivity to model performance, resulting in less significant improvements across all evaluation metrics. Nevertheless, compared to the best-performing Transformer-based model, DocuNet, the method presented in this invention still improves the overall F1 score by 0.50%. Furthermore, compared to the best-performing prompt-based model, CPT-RI, the method presented in this invention improves the overall F1 score by 0.20%. This significant performance improvement is attributed to the effective utilization of selection constraint strategies and knowledge-optimized prompting methods in this invention. Selection constraint strategies can eliminate the influence of redundant noise, allowing the model to focus on key information related to the target entity and simplifying complex interactions between entities. Simultaneously, injecting knowledge into the prompt template allows it to learn rich factual knowledge about specific entity pairs, thereby improving the quality of relational reasoning. Furthermore, the prompt prediction module can fully utilize the powerful contextual understanding and generation capabilities of pre-trained language models, enabling it to capture relational information in the text more accurately.
[0065] Table 3 Experimental results of this invention on the GDA dataset
[0066]
[0067] All test results validated the effectiveness and advancement of the method of this invention in document-level biomedical relation extraction tasks.
Claims
1. A document-level biomedical relation extraction method based on a prompt optimization model, characterized by comprising a selection constraint module, a knowledge optimization prompt template construction module, and a prompt relation prediction module, which proceed in the following steps: Step 1. Obtain a relation instance text from the document to be extracted using the selection constraint module. Process the text with selection constraints to obtain the document sequence after selection constraints. The specific steps are as follows: Step 1.1 Obtain a relation instance text from the document to be extracted, and determine the set of sentence positions p containing all mentions of the head entity. h The set of sentence positions p containing all mentions of the last entity. t The head entity and tail entity constitute the target entity pair that needs to be predicted in the document, i.e., each relation instance; Step 1.2 Obtain the set p of all sentences containing information related to the head and tail entities according to formula (1); p=p h ∪p t (1) Step 1.3 uses set p as the selection constraint and retains only sentences in the document that exist in set p through a judgment mechanism, thereby obtaining the document sequence after selection constraint; Step 2. Construct a knowledge optimization prompt template using the knowledge optimization prompt template construction module. The specific steps are as follows: Step 2.1 Use ENTITY / TYPE MARKER to index the beginning and end entity positions of the document sequence after selection constraint. That is, insert special markers around the target entities in the document and add markers at the beginning of the text to obtain the text sequence after adding markers, denoted as x. Step 2.2 Define the prompt template as shown in (2): T(·)="·The relations between ET .h and ET .t is[MASK].” (2) in, ET .h and ET ·t These represent the types of the head entity and the tail entity, respectively. Map x to T(·) to obtain the knowledge optimization prompt template as shown in (3); x prompt =T(x)="x The relations between ET xh and ET xt is[MASK].” (3) The [MASK] is a mask marker; Step 3. Complete the relationship instance classification using the prompt relationship prediction module. The specific steps are as follows: Step 3.1 Build the completed knowledge optimization hint template x prompt Input a pre-trained language model T5 based on the Transformer architecture, and compute the hidden layer vector representation h at the mask marker [MASK] position. [MASK] ; Step 3.2 Generate the probability distribution of each possible tag word at the mask position, that is, perform relation prediction through linear transformation and softmax function, thereby completing the relation instance classification, that is, obtaining the relation type of the document relation instance; The linear transformation is shown in (4): z=h [MASK] ·W+b (4) Where W is the weight matrix in the output layer, and b is the bias vector in the output layer; The softmax function, as shown in formula (5), transforms the unnormalized score vector z into a probability distribution, thereby generating the probability distribution P(y|x) of each possible tag word at the mask position. P(y|x)=P([MASK]=v|x prompt )=softmax(z) (5) in, It is a collection of all tag words.
Citation Information
Patent Citations
Chinese medical entity relation joint extraction method and system
CN114036934A
Traditional Chinese medicine document level relation extraction method and system based on local path enhancement, electronic equipment and medium
CN117151102A