Self-adaptive data correction method for document relation extraction

Through the adaptive data correction method, in response to the noise labeling problem in the document-level relationship extraction model, denoising training and confidence-based data revision are used to gradually optimize the data set, solving the problem of unsatisfactory performance of the document-level relationship extraction model, and achieving high-quality data revision and model performance improvement.

CN120031034APending Publication Date: 2025-05-23SCHOOL OF MILITARY MANAGEMENT NAT DEFENSE UNIV OF THE CHINESE PEOPLES LIBERATION ARMY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411923426.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-25
Publication Date
2025-05-23

AI Technical Summary

Technical Problem

The experimental performance of the existing document-level relationship extraction model on mainstream data sets is not ideal, mainly due to false negatives and false positives caused by noise annotation problems, which hinders the training and evaluation of the model.

Method used

A method of adaptive data correction for document relationship extraction is proposed. Through denoising training, confidence-based data revision and iterative training, noisy data sets are gradually optimized to improve the recognition ability and data quality of the model.

Benefits of technology

High-quality data revisions on data sets of any size are achieved, significantly improving the performance of document-level relationship extraction models, reducing the cost of data revisions, and automated processing, avoiding the high cost of manual intervention.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120031034A_ABST
    Figure CN120031034A_ABST
Patent Text Reader

Abstract

The invention discloses a document relationship extraction-oriented adaptive data correction method. The method comprises the following steps of: obtaining a drug database; training a classifier for identifying potential relation facts, including positive training and negative training; a self-adaptive threshold is introduced for each relation fact, false positive is filtered in a self-adaptive mode, and false negative is marked again; iteratively executing de-noising training and data revision based on confidence coefficient; and the interaction between the medicines is identified and output, so that side effects or adverse reactions caused by improper medication are avoided. According to the invention, the distinguishing capability of relation facts is enhanced, and the influence of noisy annotations is reduced; self-adaptive data revision is introduced to revise the relation fact of long-tail distribution; iterative training mines valuable relational facts from a noisy dataset.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of medical information technology, and in particular relates to an adaptive data correction method for document relationship extraction. Background Art

[0002] Relation Extraction (RE), as an important foundation for knowledge acquisition, is part of natural language processing technology and has made important contributions to the advancement of many tasks such as medical artificial intelligence. For example, through relation extraction technology, patients' symptoms can be identified from electronic medical records and associated with known diseases to help doctors make diagnoses faster; or interactions between drugs can be identified to avoid side effects or adverse reactions caused by improper medication; or personalized treatment plans can be recommended based on the patient's specific situation and the latest medical research results; or key information such as drug efficacy and side effects can be extracted from a large number of clinical trial reports to support new drug research and development. For example: Use named entity recognition algorithms to mark drug names. For example, in the sentence "Taking warfarin while taking aspirin may increase the risk of bleeding", "warfarin" and "aspirin" should be marked as drug entities. Then a relation extraction model needs to be trained to identify the relationship between drug entities. In this example, this application focuses on the "interaction" relationship. The model needs to learn to recognize which vocabulary patterns indicate interactions, such as "increase risk", "reduce efficacy", etc.

[0003] However, relation extraction in practical applications is far more complex than the above examples. To adapt to practical applications, the research focus of relation extraction is gradually shifting from sentence-level data to document-level data. Document-level relation extraction (DocRE) aims to extract the relationship between entities in the entire document. Previous studies have built a powerful model structure from multiple aspects, starting from document-level graph construction, enhancing logical rule-based reasoning, and improving context understanding, and have conducted extremely detailed analysis and modeling of the task from multiple angles. Although existing studies have made considerable efforts in designing model structures, from the perspective of practical applications, the experimental performance of DocRE models on mainstream datasets is still unsatisfactory. The frustrating performance is mainly due to the large amount of noisy annotations introduced in the complex document-level data annotation process, which seriously hinders the training and evaluation of the model. However, previous studies rarely consider the noisy annotation problem of DocRE datasets.

[0004] The noisy annotation problem refers to the presence of a large number of false negatives (FNs) and false positives (FPs) in the dataset, which is a common problem that seriously affects model training and evaluation. FN represents relational facts that are ignored in the annotation process due to the complex text structure and contextual semantics of document-level data. For example, a fact is observed from a sentence, thereby inferring a relational fact, which is not annotated in the previous annotation process, which introduces the FN problem. In addition, most popular RE datasets use external knowledge bases as supporting resources for annotations, which inevitably leads to the FP problem.

[0005] To address this problem, early researchers manually re-annotated a subset of previously manually annotated documents in the dataset to generate a higher quality dataset. Although these manual methods improve the performance of the model, they mainly target FNs in the dataset and only perform on manually annotated documents, while ignoring the remaining large amount of distantly supervised documents. Due to the low efficiency and high labor cost, the proposed re-labeling strategy is not suitable for large-scale datasets, and in the strategy, FPs are not explicitly focused. Summary of the invention

[0006] In view of this, this application introduces the idea of ​​automatic relabeling based on existing medical / pharmaceutical datasets, and comprehensively analyzes the challenges that will be encountered: 1) A document consists of multiple sentences, each of which contains different entities and expresses different relationships. In addition, there are different dependencies between entities in sentences. Due to the complex structure and contextual semantics, it becomes challenging to effectively identify the potential relationships between FNs / FPs. 2) Data statistics show that automatic relabeling is severely affected by the long-tail distribution of relationship facts: a few relationships are common, while many relationships are rare. During model training, this situation leads to different degrees of convergence, which confuses the model when distinguishing whether the relationship facts are correct or noise. 3) Due to the limited number of high-quality samples in the noisy dataset, it is challenging to train a model that effectively handles noisy annotations in a single round of data revision. The model may only be able to identify false negatives based on existing posterior knowledge from the noisy dataset. After revising the noisy dataset using the identified false negatives, the model can be trained with richer knowledge and deep reasoning can be performed to identify false negative facts.

[0007] In order to address the challenges and achieve high-quality automatic data re-annotation, the present invention proposes a data revision framework (this application) for document-level relationship extraction, which aims to iteratively optimize noisy datasets. Compared with previous studies, this application can be easily applied to datasets of any size and achieve high-quality data revision. Specifically, this application optimizes the recognition ability of long-tail distribution relationships by dynamically adjusting the threshold; enhances the modeling ability of multi-entity relationships in complex documents through local context pooling; uses knowledge distillation to reduce the noise problem caused by remote supervision annotation through knowledge transfer; uses iterative optimization: gradually improves the quality of data revision and realizes the coordinated optimization of classifiers and data.

[0008] Given a predefined set of relationship types and a document-level dataset Each document (d i ) contains a predefined collection of entities The goal of document-level relation extraction (DocRE) is to extract i ) to extract the relationship between entities ({(e_{im},e_{in})|e_{im},e_{in}∈E_i}). In practical scenarios, the complex structure and semantics of document-level data inevitably affect the data annotation process, resulting in noise in the dataset. This application represents this noisy dataset as in is a document i ) is a possibly noisy label, coming from a noise distribution (P * After data annotation, this application defines the following two types of relationship facts: The first type is document (d i ) but not labeled ({(e_{im},r_k,e_{in})|e_{im},e_{in}∈E_i,r_k∈R}), which is referred to as false negative (FN) in this application; the second category is the document (d i ) but are annotated, are referred to as false positives (FP) in this application.

[0009] The noisy annotations mentioned above (i.e., false negative and false positive annotations distributed in documents) have long been an obstacle to the development of document-level relation extraction (DocRE). Previous studies have explicitly explored this problem by manually re-annotating. However, this approach is inefficient, costly, and difficult to apply to large-scale datasets. To address this problem, the present invention constructs a revision method for annotated data. Assume that there is a noise-free true data distribution (P), and a clean dataset (D = {(d 1 ,y 1),…,(d n ,y n )}), the goal of this application is to use a noisy dataset (D * ), find the optimal estimated parameters (θ) of the true mapping (f:x→y), where ((x,y)∈D). Specifically, this application consists of three customized modules: 1. Denoising training module: performs robust classifier training to accurately distinguish between false negatives (FNs) and false positives (FPs) distributed in documents. 2. Confidence-based data revision module: revise noisy datasets with long-tail distributions in an adaptive manner. 3. Iterative training module: converts the revised dataset into useful training data to further enhance the classifier, thereby building a virtuous data revision cycle. These key modules together automatically solve the problem of noisy annotations. Through this application, reliable data revision can be achieved in a cost-effective manner, and reliable training datasets can be continuously generated for document-level relationship extraction (DocRE).

[0010] To achieve the above purpose, the adaptive data correction method for document relationship extraction disclosed in the present application includes the following steps:

[0011] Data acquisition steps: obtain drug database;

[0012] Denoising training step: training a classifier for identifying potential relational facts, including positive training and negative training; the positive training attempts to force the classifier to learn features from given noisy data; the negative training separates positive and negative instances, thereby reducing the confidence score of false positives;

[0013] Confidence-based data revision step: For the long-tailed relationship facts, an adaptive threshold is introduced for each relationship fact to adaptively filter out false positives and relabel false negatives, thereby revising the noisy dataset D. * Get the revised dataset

[0014] Iterative training step: using the revised dataset Enhance the classifier, conduct a new round of data revision, iteratively perform denoising training and confidence-based data revision, and * The relationship facts between drugs are gradually mined;

[0015] Data output step: Based on the mined facts about the relationship between drugs, identify and output the interactions between drugs to avoid side effects or adverse reactions caused by improper use of drugs.

[0016] Furthermore, given the original noisy dataset D * , a set of predefined relationship types and the entity set of the i-th document l represents the lth relationship type, M is the total number of relationship types, j represents the jth entity in the ith document, and T represents the number of entities contained in the ith document. The denoising training first uses the DocRE model as the relationship classifier f to obtain the mth pair of entities (e ij ,e ik ) m :

[0017] p m =f(d i ,e ij ,e ik ),

[0018] Among them, e ik is the kth entity in the i-th document, p m ∈R |R| , p m ∈[0,1] represents document d i Relationship facts ij ,r n ,e ik ) exists; after obtaining the predicted probability of the relationship facts in the dataset, positive training and negative training are performed to optimize the relationship classifier.

[0019] Furthermore, a relational fact discrimination model is constructed using cross entropy loss to improve the recognition ability of the classifier on positive instances;

[0020] The loss function for forward training is defined as:

[0021]

[0022] Among them, p ij represents the predicted probability of the jth entity in the i-th document, y ij is the label of the j-th pair of entities in the ith document, 1 indicates that the relationship exists, and 0 indicates that the relationship does not exist.

[0023] Furthermore, complementary labels are generated to reduce the confidence of mislabeling. The false negative problem is solved by constructing complementary labels; the loss of the negative training is defined as:

[0024]

[0025] represents complementary labels, N is the total number of documents, and M is the number of entity pairs in each document.

[0026] Furthermore, the final loss function combines positive and negative training:

[0027]

[0028] α is the weighting factor.

[0029] Furthermore, the structure of the multi-level classifier includes an input layer, a local context pooling layer, a relation classification layer, and a confidence calculation module;

[0030] The input layer encodes the entity pairs and their context information in the input document into vector representations through the embedding layer. Mathematical description:

[0031] X ij =f embed (E i ,E j ,C ij ), where E i ,E j Represents an entity, C ij Indicates context;

[0032] The local context pooling layer uses the axial attention mechanism to extract the local context features of each entity pair, mathematically described as: H ij =Attention(X ij ,Context);

[0033] The relationship classification layer calculates the relationship probability score based on the classification head; mathematical description: p ij (k)=softmax(W k H ij +b k ), where k represents the relationship type;

[0034] The confidence calculation module performs confidence aggregation on the relationship probability scores.

[0035] Furthermore, the confidence is calculated as follows:

[0036] After denoising training, for each entity pair, the relationship prediction probability output by the classifier is used to calculate the maximum value of all possible relationships as the confidence score of the entity pair:

[0037]

[0038] Among them, p ij (k) represents the predicted probability of the real number pair (i, j) under relationship type k.

[0039] Furthermore, in the confidence-based data revision step, an adaptive threshold T is calculated for each relationship type k. k :

[0040] T k =λ 1 ·max(s ij)+λ 2 ·mean(s ij )

[0041] λ 1 and λ 2 is a hyperparameter used to balance the contribution of the maximum value and the average value, max(s ij ) is the maximum confidence score of relation type k, mean(s ij ) is the average confidence score of relation type k.

[0042] Furthermore, the iterative training establishes a virtuous cycle, thereby gradually improving the quality of data revision; in each iteration, the classifier is retrained using the revised dataset, allowing it to accumulate more knowledge, thereby gradually improving the data revision process;

[0043] Specifically include:

[0044] If s ij <T k , then remove the annotation relationship of entity pair (i,j),

[0045] If s ij ≥T k , then re-label k as an existence relationship;

[0046] Finally, the revised dataset is sent to the next iteration, and the dataset that makes the relation classifier perform best on the development set throughout the iterative training and data revision process is retained.

[0047] This application designs a cost-effective data revision method that can automatically solve the noisy annotation problem without manual intervention; builds a solid training framework to train a robust classifier based on an existing noisy dataset to accurately distinguish potential relational facts; 3. Designs an adaptive threshold to distinguish false negatives (FNs) and false positives (FPs) in long-tailed relational facts instead of relying on a uniform threshold; uses iterative training in combination with newly mined relational facts to enhance classifier performance and improve the quality of data revision.

[0048] The method proposed in this application has the following beneficial effects:

[0049] 1. Cost-effective and applicable to datasets of various sizes: This application can not only significantly reduce the cost of data revision, but also be economically and efficiently applied to datasets of any size.

[0050] 2. Robust data revision capability: In the data revision process, this application achieves robust data revision through denoising training and confidence-based revision strategies, especially for noisy data sets with long-tail distributions. Through the denoising training module, this application can further utilize the revised data to gradually mine knowledge in noisy data sets, forming a virtuous circle.

[0051] 3. Competitive performance compared to manual data revision: Even compared to manual revision, the present application achieves competitive performance. BRIEF DESCRIPTION OF THE DRAWINGS

[0052] Figure 1 The framework diagram of the present invention;

[0053] Figure 2 The iterative training algorithm process of the present invention. DETAILED DESCRIPTION

[0054] The present invention is further described below in conjunction with the accompanying drawings, but the present invention is not limited in any way. Any changes or substitutions made based on the teachings of the present invention belong to the protection scope of the present invention.

[0055] Artificial intelligence is the theory, method, technology and application system that uses digital computers or machines controlled by digital computers to simulate, extend and expand human intelligence, perceive the environment, acquire knowledge and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science that attempts to understand the essence of intelligence and produce a new intelligent machine that can respond in a similar way to human intelligence. Artificial intelligence is to study the design principles and implementation methods of various intelligent machines so that machines have the functions of perception, reasoning and decision-making.

[0056] Natural Language Processing (NLP) is an important direction in the fields of computer science and artificial intelligence. It studies various theories and methods that can achieve effective communication between people and computers using natural language. Natural language processing is a science that integrates linguistics, computer science, and mathematics. Therefore, research in this field will involve natural language, that is, the language people use in daily life, so it is closely related to the study of linguistics. Natural language processing technology usually includes text processing, semantic understanding, machine translation, robot question answering, knowledge graph and other technologies.

[0057] The technical solutions provided in the embodiments of the present application involve technologies such as machine learning and natural language processing of artificial intelligence, which are specifically introduced and explained through the following embodiments.

[0058] Before introducing the embodiments of the present application, some terms involved in the present application are first explained.

[0059] 1. Document-level relation extraction: Document-level relation extraction (DocRE) aims to extract relations between entities distributed across multiple sentences. Early DocRE research focused on improving the relation extraction capability of DocRE models, and the experimental performance of DocRE models was far from satisfactory. The ubiquitous noisy annotations hinder the training of DocRE models. As a common problem that hinders the development of various NLP tasks, the noisy annotation problem has attracted a lot of attention from researchers. However, previous studies either focused on sentence-level tasks or incurred excessive computational overhead and memory usage, making it challenging to apply to DocRE. In order to further explore the potential of the proposed DocRE model, researchers attempted to solve the problem of DocRE by manual re-annotation. However, manually annotating document-level data requires high labor costs and can hardly be applied to large-scale datasets.

[0060] 2. Machine-assisted data generation: The task of machine-assisted data generation aims to provide reliable data sets using artificial intelligence technology, thereby alleviating the overall expensive manual labor. Due to the rapid development of deep learning, the contradiction between the exponentially growing data demand and the difficulty of data annotation continues to escalate. Therefore, remote supervision is adopted in various tasks to generate annotated training data without manual operation. In addition to data generation, remote supervision is also used in data denoising and training processes.

[0061] One of the goals of this application is to construct a data relabeling framework to establish a reliable DocRE training data generation pipeline in a cost-effective manner.

[0062] Notation: Given a neural network (f) and a set of learnable parameters and input set (X = {(d_i,e_{ij},e_{ik})|d_i∈D,e_{ij},e_{ik}∈E_i}), where (d i ) represents a document, (e ij ) and (e ik ) represents the document (d i ). The structure of the neural network (f) is initialized by a typical document-level relation extraction (DocRE) model structure and used as a relation classifier. The goal of the relation classifier (f) is to find the optimal estimated parameters (θ) to minimize the loss between the predicted relation probability distribution of the entity pair and its corresponding true label: min θ (f(d i ,e ij ,e ik ),y).

[0063] Figure 1The revision method framework of the present invention is presented. The framework is divided into three modules: denoising training, confidence-based data revision and iterative training. Denoising training aims to minimize the noise data (D * ) in the noisy data (D_n^*={(d_i,y_i^n)|y_i^n∈y_i^*,y_i^n≠y_i}) to * ) to train a robust DocRE classifier. Through denoising training, the classifier is able to generate reliable confidence scores. To deal with the long-tail distribution problem, confidence-based data revision adaptively filters false positives (FPs) and re-labels false negatives (FNs), thereby revising the dataset (D * ). Iterative training uses the revised dataset Further enhancing the classifier, denoising training and confidence-based data revision are performed iteratively, from (D * ) to gradually mine valuable relational facts. Therefore, ensuring that the revised dataset Gradually align with the real dataset (D).

[0064] The structure of the multi-level classifier includes an input layer, a local context pooling layer, a relation classification layer, and a confidence calculation module;

[0065] The input layer encodes the entity pairs and their context information in the input document into vector representations through the embedding layer. Mathematical description:

[0066] X ij =f embed (E i ,E j ,C ij ), where E i ,E j Represents an entity, C ij Indicates context;

[0067] The local context pooling layer uses the axial attention mechanism to extract the local context features of each entity pair, mathematically described as: H ij =Attention(X ij ,Context);

[0068] The relationship classification layer calculates the relationship probability score based on the classification head; mathematical description: p ij (k)=softmax(W k H ij +b k ), where k represents the relationship type;

[0069] The confidence calculation module performs confidence aggregation on the relationship probability scores.

[0070] S1 denoising training

[0071] Denoising training includes positive training and negative training, aiming to build a DocRE model with strong relational pattern recognition capabilities. Given the original noisy dataset (D * ), a set of predefined relationship types and the entity set of the (i)th document The framework first uses the typical DocRE model as a relation classifier (f) to obtain the (m)th pair of entities ((e ij ,e ik ))The probability score (p m ):

[0072] p m =f(d i ,e ij ,e ik ),

[0073] Among them (p m ∈R |R| ), (p mn ∈[0,1]) represents the document (d i ) in the relationship facts ((e ij ,r n ,e ik )) existence. Once the predicted probability of the relational facts in the dataset is obtained, the present application performs positive training and negative training to optimize the relational classifier. Specifically, the positive training aims to enhance the recognition ability of relational patterns by improving the prediction scores of instances marked as positive examples. In addition, through negative training, the influence of noisy annotations in the model training process can be effectively reduced.

[0074] S11 positive training

[0075] Forward training uses cross entropy loss to build a robust relational fact discrimination model. The loss function of forward training is defined as:

[0076]

[0077] Among them, p ij represents the predicted probability of the jth entity in the i-th document, y ij is the label of the j-th pair of entities in the ith document, 1 indicates that the relationship exists, and 0 indicates that the relationship does not exist.

[0078] S12 negative training

[0079] Negative training aims to minimize the impact of noisy annotations. For each label The input document (d i), this application generates its complementary tag Considering the common false negative problem in DocRE dataset, complementary labels are constructed by randomly extracting from label space, eliminating Right now The loss for negative training is defined as:

[0080]

[0081] represents complementary labels, N is the total number of documents, and M is the number of entity pairs in each document.

[0082] The final loss function combines positive and negative training:

[0083]

[0084] α is the weighting factor.

[0085] S2 Confidence-Based Data Revision

[0086] Confidence-based data revision aims to filter false positives (FPs) from the original dataset and re-annotate false negatives (FNs) that were overlooked by previous annotation methods. With proper revision, the revised data will provide more useful knowledge, thereby further improving model performance.

[0087] S21 confidence calculation

[0088] After denoising training, find a set of well-estimated parameters (φ) for the relation classifier (f) and use these parameters to compute the document (d i ) in the (m)th pair of entities ((e ij ,e ik There is a relationship (r k )’s confidence score:

[0089] Using the relationship prediction probability output by the classifier, calculate the maximum value of all possible relationships as the confidence score of the entity pair:

[0090]

[0091] Among them, p ij (k) represents the predicted probability of the real number pair (i, j) under relationship type k.

[0092] S22 Adaptive Data Revision

[0093] In obtaining (D *), adaptive data revision aims to filter out false positives (FPs) and re-label false negatives (FNs). In order to accurately revise the data of long-tail distributed relational facts, this application constructs an adaptive threshold for each relational fact based on the confidence score of the classifier. Specifically, it includes:

[0094] If s ij <T k , then remove the annotation relationship of entity pair (i,j),

[0095] If s ij ≥T k , then re-label k as an existence relationship;

[0096] Therefore, the adaptive threshold is constructed based on the convergence degree of each relation. In this way, data revision not only depends on the convergence degree of each relation, but also is dynamically adjusted throughout the iterative training process. In addition, the adaptive threshold is also more effective in dealing with the multi-label long-tail distribution problem that is widely present in the DocRE dataset.

[0097] Based on the confidence scores of relational facts, this application empirically assumes that these confidence scores follow a certain distribution, the vast majority of which do not exist in (D * ) is in the low value area, while the relational facts in (D * ) are usually in the medium or high value area. Based on this assumption, this application designs a revised noisy DocRE dataset strategy.

[0098] S3 Iterative Training

[0099] Considering the complexity of document-level relation extraction (DocRE), simple denoising training and adaptive data revision process are not enough to achieve effective data revision. Therefore, this application introduces iterative training to establish a virtuous cycle, thereby gradually improving the quality of data revision. In each iteration, this application retrains the classifier with the revised dataset, allowing it to accumulate more useful knowledge, thereby gradually improving the data revision process. Figure 2 As shown in the algorithm, in order to ensure that the relation classifier adapts to the distribution of the current dataset and avoids the accumulation of errors introduced by noisy annotations, the relation classifier is first reinitialized in each iteration. In addition, reinitialization also introduces randomness, which further enhances the effect of data revision. Through denoising training and the guidance of a small amount of artificial knowledge, the current optimal relation classifier is obtained. Subsequently, according to Figure 2The method described in the algorithm performs confidence-based data revision. Finally, the revised dataset will be sent to the next iteration, and the dataset that makes the relation classifier perform best on the development set (Dev) throughout the iterative training and data revision process is retained.

[0100] Dataset

[0101] This application conducted experiments on three different versions of drug datasets:

[0102] DrugBank: It is a comprehensive drug database, including basic drug information, targets, metabolic pathways, etc. It is suitable for drug development, drug interaction research, etc.

[0103] ChEMBL: is a large-scale small molecule biological activity database, including compound structure, biological activity data, target information, etc., suitable for drug discovery, chemical genomics research, etc.

[0104] BindingDB: is a database about the interaction between small molecules and their protein targets, including the binding affinity data between compounds and proteins; it is suitable for drug design, protein-ligand interaction research, etc.

[0105] The documents of this application consist of 5,053 manually annotated documents (HA-Doc) and 101,873 remotely supervised annotated documents (DS-Doc) from the above three databases.

[0106] The dataset of this application introduces 57.1% of relationship facts and effectively reduces NA (i.e., no relationship), which shows that this application can effectively solve the false negative problem. For the convenience of comparison, this application adjusts the dataset to different sizes by random sampling, including a mini dataset of 3,053 documents, a large dataset of 10,000 documents, and a full dataset of 101,873 documents.

[0107] In order to verify the versatility of the proposed framework and the quality of data revision, this application introduces five typical document-level relation extraction (DocRE) models: DocRE-Rec, SSAN, ATLOP, DocuNet and AA.

[0108] -DocRE-Rec proposes an encoder-classifier reconstructor to force the model to pay more attention to entity pairs with relations, thereby effectively identifying relation categories.

[0109] -SSAN explicitly models the entity structure information in DocRE. With structure prior, SSAN can combine contextual reasoning and structural reasoning to improve performance.

[0110] -ATLOP uses adaptive threshold loss for flexible entity pair relationship classification. In order to effectively capture relevant contextual information, the model further introduces local context pooling.

[0111] -DocuNet introduces a U-shaped network for DocRE, which processes the DocRE task in parallel with the semantic segmentation task in computer vision.

[0112] -AA developed a framework to enhance the DocRE task by introducing anaphora and coreference information to provide a detailed understanding of entity interactions.

[0113] The framework proposed in this application is highly versatile and supports seamless integration with any DocRE model. In one embodiment, this application initializes the classifier of this application using the DocuNet architecture for data revision, and uses RoBERTa-large as the encoder of the classifier. The hyperparameters of the classifier are fixed according to their original research. In the process of revising the DS-Doc application (full) data set, 12 iterations of denoising training and adaptive data revision were performed. In each iteration, this application trains the classifier for 30 cycles. For negative training, negative samples are randomly selected for optimization with a probability of 35%. In the adaptive data revision stage, the values ​​of (μ) and (α) depend on the degree of convergence of each target relational pattern. The convergence threshold is set to 0.99. For the relational pattern with the highest confidence score greater than the convergence value, the values ​​of (μ) and (α) are set to 0.55 and 0.3, respectively. Otherwise, the values ​​of (μ) and (α) are set to 0.385 and 0.27.

[0114] Table 1 Experimental results

[0115]

[0116] This application compares the manually annotated training data of HA-DocRED with the data of this application. In order to eliminate the impact of data volume, this application compares data sets of the same size. As can be seen from Table 1, DocuNet trained using the mini data set shows almost the same competitiveness as HA-RED in F1 scores (50.95 vs. 50.78 on the development set, and 49.93 vs. 49.97 on the test set). In addition, the remaining models trained using this application even surpass the performance of the models trained using HA-Doc this application. Specifically, except for DocuNet, the models trained using this application achieved an improvement of 0.03 to 2.4 in F1 scores. In addition to the F1 score, the models trained using this application also showed significant performance improvements in recall. This transcendence of model performance indicates that this application can replace manual annotation to a certain extent.

[0117] The beneficial effects of this application are as follows:

[0118] The confidence-based adaptive data revision framework proposed in this application is used for automated, high-quality and cost-effective data revision. First, denoising training aims to enhance the distinguishing ability of relational facts and minimize the impact of noisy annotations. Subsequently, this application introduces adaptive data revision to revise relational facts with long-tail distributions. Finally, this application constructs iterative training to gradually mine valuable relational facts from noisy datasets by utilizing revised data. Experimental results show that this application can effectively improve data quality and replace manual work to a certain extent, continuously providing reliable training data for document-level relation extraction.

[0119] As used herein, the word "preferred" is intended to be used as an example, instance, or illustration. Any aspect or design described herein as "preferred" is not necessarily to be construed as being more advantageous than other aspects or designs. On the contrary, the use of the word "preferred" is intended to present concepts in a specific way. The term "or" as used in this application is intended to mean an inclusive "or" rather than an exclusive "or". That is, unless otherwise specified or clear from the context, "X uses A or B" means any one of the naturally included permutations. That is, if X uses A; X uses B; or X uses both A and B, then "X uses A or B" is satisfied in any of the foregoing examples.

[0120] Moreover, although the present disclosure has been shown and described with respect to one or implementations, those skilled in the art will think of equivalent variations and modifications based on the reading and understanding of this specification and the accompanying drawings. The present disclosure includes all such modifications and variations, and is limited only by the scope of the appended claims. In particular, with respect to the various functions performed by the above-mentioned components (such as elements, etc.), the terms used to describe such components are intended to correspond to any component (unless otherwise indicated) that performs the specified function of the component (such as it is functionally equivalent), even if the structure is not equivalent to the disclosed structure of the function in the exemplary implementation of the present disclosure shown herein. In addition, although the specific features of the present disclosure have been disclosed with respect to only one of several implementations, such features can be combined with one or other features of other implementations that may be desired and advantageous for a given or specific application. Moreover, insofar as the terms "including", "having", "containing" or their variations are used in specific embodiments or claims, such terms are intended to be included in a manner similar to the term "comprising".

[0121] The functional units in the embodiments of the present invention may be integrated into a processing module, or each unit may exist physically separately, or multiple or more units may be integrated into one module. The above-mentioned integrated module may be implemented in the form of hardware or in the form of a software functional module. If the integrated module is implemented in the form of a software functional module and sold or used as an independent product, it may also be stored in a computer-readable storage medium. The above-mentioned storage medium may be a read-only memory, a disk or an optical disk, etc. The above-mentioned devices or systems may execute the storage method in the corresponding method embodiment.

[0122] To sum up, the above embodiment is an implementation mode of the present invention, but the implementation mode of the present invention is not limited by the embodiment. Any other changes, modifications, substitutions, combinations, and simplifications that deviate from the spirit and principles of the present invention should be equivalent replacement methods and are included in the protection scope of the present invention.

Claims

1. An adaptive data correction method for document relationship extraction, characterized in that: The following steps are involved: Data acquisition steps: obtain drug database; Denoising training step: training a classifier to identify potential relational facts, including positive training and negative training; the positive training attempts to force the classifier to learn features from the given noisy data; The negative training separates positive and negative instances, thereby reducing the confidence scores of false positives; Confidence-based data revision step: For the long-tailed relationship facts, an adaptive threshold is introduced for each relationship fact to adaptively filter out false positives and relabel false negatives, thereby revising the noisy dataset D. * Get the revised dataset Iterative training step: using the revised dataset Enhance the classifier, conduct a new round of data revision, iteratively perform denoising training and confidence-based data revision, and * The relationship facts between drugs are gradually mined; Data output step: Based on the mined facts about the relationship between drugs, identify and output the interactions between drugs to avoid side effects or adverse reactions caused by improper use of drugs.

2. The adaptive data correction method for document relationship extraction according to claim 1, characterized in that: Given an original noisy dataset D * , a set of predefined relationship types and the entity set of the i-th document l represents the lth relationship type, M is the total number of relationship types, j represents the jth entity in the ith document, and T represents the number of entities contained in the ith document. The denoising training first uses the DocRE model as the relationship classifier f to obtain the mth pair of entities (e ij ,e ik ) m : p m =f(d i ,e ij ,e ik ), Among them, e ik is the kth entity in the i-th document, p m ∈R |R| , p m ∈[0,1] represents document d i Relationship facts ij ,r n ,e ik ) exists; after obtaining the predicted probability of the relationship facts in the dataset, positive training and negative training are performed to optimize the relationship classifier.

3. The adaptive data correction method for document relationship extraction according to claim 2 is characterized in that: Using cross entropy loss, a relational fact discrimination model is constructed to improve the recognition ability of the classifier on positive instances; The loss function for forward training is defined as: Among them, p ij represents the predicted probability of the jth entity in the i-th document, y ij is the label of the j-th pair of entities in the ith document, 1 indicates that the relationship exists, and 0 indicates that the relationship does not exist.

4. The adaptive data correction method for document relationship extraction according to claim 3 is characterized in that: Generate complementary labels to reduce the confidence of mislabeling. Solve the false negative problem by constructing complementary labels; the loss of negative training is defined as: represents complementary labels, N is the total number of documents, and M is the number of entity pairs in each document.

5. The method for adaptive data correction for document relationship extraction according to claim 4, characterized in that: The final loss function combines positive and negative training: α is the weighting factor.

6. The method for adaptive data correction for document relationship extraction according to claim 5, characterized in that: The structure of the multi-level classifier includes an input layer, a local context pooling layer, a relation classification layer, and a confidence calculation module; The input layer encodes the entity pairs and their context information in the input document into vector representations through the embedding layer. Mathematical description: X ij =f embed (E i ,E j ,C ij ), where E i ,E j Represents an entity, C ij Indicates context; The local context pooling layer uses the axial attention mechanism to extract the local context features of each entity pair, mathematically described as: H ij =Attention(X ij ,Context); The relation classification layer calculates relation probability scores based on the classification head; Mathematical description: p ij (k)=softmax(W k H ij +b k ), where k represents the relationship type, W k is the weight, b k is the threshold value; The confidence calculation module performs confidence aggregation on the relationship probability scores.

7. The adaptive data correction method for document relationship extraction according to claim 6, characterized in that: The confidence level is calculated as follows: After denoising training, for each entity pair, the relationship prediction probability output by the classifier is used to calculate the maximum value of all possible relationships as the confidence score of the entity pair: Among them, p ij (k) represents the predicted probability of the real number pair (i, j) under relationship type k.

8. The method for adaptive data correction for document relationship extraction according to claim 7, characterized in that: In the confidence-based data revision step, an adaptive threshold T is calculated for each relationship type k. k : T k =λ1·max(s ij )+λ2·mean(s ij ) λ1 and λ2 are hyperparameters used to balance the contribution of the maximum value and the average value. ij ) is the maximum confidence score of relation type k, mean(s ij ) is the average confidence score of relation type k.

9. The method for adaptive data correction for document relationship extraction according to claim 8, characterized in that: The iterative training establishes a virtuous cycle, thereby gradually improving the quality of data revision; In each iteration, the classifier is retrained using the revised dataset, allowing it to accumulate more knowledge, thereby gradually improving the data revision process; Specifically include: If s ij <T k , then remove the annotation relationship of entity pair (i,j), If s ij ≥T k , then re-label k as an existence relationship; Finally, the revised dataset is sent to the next iteration, and the dataset that makes the relation classifier perform best on the development set throughout the iterative training and data revision process is retained.