Document-Level Relation Extraction Method, Apparatus, Device, and Medium Based on Duplicate Sampling Removal

By de-resampling and multi-grained text encoding, combined with the relationship extraction method of graph convolution neural network and hybrid expert system, the problem of unbalanced distribution of relationship categories is solved, and the accuracy and accuracy of document-level relationship extraction is improved.

CN119807329BActive Publication Date: 2025-07-29深圳市规划和自然资源数据管理中心(深圳市空间地理信息中心) +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510309131.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-17
Publication Date
2025-07-29
Estimated Expiration
2045-03-17

AI Technical Summary

Technical Problem

When traditional document-level relationship extraction methods face unbalanced distribution of relationship categories, they are prone to ignore low-frequency relationship categories, resulting in inaccurate extraction results.

Method used

The document-level relationship extraction method based on deresampling is adopted, and documents are processed through preset marks, and document relationship extraction model is used for deresampling and multi-grained text encoding. The relationship extraction process is performed in combination with graph convolution neural network and hybrid expert system.

Benefits of technology

It effectively reduces the impact of unbalanced distribution of relationship categories on the extraction results, improves the accuracy and accuracy of relationship extraction, and improves the accuracy of target relationship modeling.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119807329B_ABST
    Figure CN119807329B_ABST
Patent Text Reader

Abstract

This application relates to the field of natural language processing technology. This application discloses a document-level relation extraction method, device, equipment, and medium based on duplicate removal sampling, which can reduce the impact of the unbalanced distribution of relation categories on the accuracy of relation extraction results, thereby improving the precision of relation extraction results. The document-level relation extraction method based on duplicate removal sampling includes obtaining a text document; performing a marking process on the text document using a preset marker to obtain a marked document, where the marked document includes at least one set of entity pairs; inputting the marked document into a document relation extraction model, and the document relation extraction model performs duplicate removal sampling and relation extraction processing on the marked document to obtain a relation extraction result.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of natural language processing technology. More specifically, this application relates to a document-level relation extraction method, device, equipment, and medium based on duplicate sampling removal. Background Art

[0002] Traditional document-level relation extraction methods obtain document text, convert the document text to obtain target entity pair vectors and non-target entity pair vectors; combine the target entity pair vectors and non-target entity pair vectors to obtain entity pair combined vectors; input the entity pair combined vectors into a classifier to obtain relation extraction results. However, most document texts have an unbalanced distribution of relation categories. In the process of converting such document texts with unbalanced relation category distributions into target entity pair vectors and non-target entity pair vectors, this method is prone to ignoring low-frequency relation categories, resulting in inaccurate relation extraction results. Summary of the Invention

[0003] The purpose of the embodiments of this application is to provide a document-level relation extraction method, device, equipment, and medium based on duplicate sampling removal, which can reduce the impact of unbalanced relation category distribution on the accuracy of relation extraction results, thereby improving the precision of relation extraction results. The embodiments of this application are mainly implemented through the following technical solutions:

[0004] In the first aspect of the embodiments of this application, a document-level relation extraction method based on duplicate sampling removal is provided, including:

[0005] Obtain a text document;

[0006] Perform a marking process on the text document using a preset marker to obtain a marked document, where the marked document includes at least one set of entity pairs;

[0007] Input the marked document into a document relation extraction model, and the document relation extraction model performs duplicate sampling removal and relation extraction processing on the marked document to obtain a relation extraction result.

[0008] According to an embodiment of this application, the step of inputting the marked document into a document relation extraction model, and the document relation extraction model performing duplicate sampling removal and relation extraction processing on the marked document to obtain a relation extraction result includes:

[0009] Input the marked document into the duplicate sampling removal module of the document relation extraction model for duplicate sampling removal processing to obtain a set of entity pairs to be processed;

[0010] Input the marked document into the multi-granularity text encoding module of the document relation extraction model for encoding processing to obtain context embedding vectors;

[0011] Input the set of entity pairs to be processed into the graph convolutional neural network of the document relation extraction model for calculation and processing, to obtain a first global embedding representation of the subject and a second global embedding representation of the object in the target entity pair, where the target entity pair is any entity pair in the set of entity pairs to be processed;

[0012] Input the target entity pair and the context embedding vector into the context pooling module of the document relation extraction model for local context pooling processing to obtain relation features;

[0013] Input the first global embedding representation, the second global embedding representation, and the relation features into the mixture of experts system of the document relation extraction model for scoring calculation processing to obtain a relation score corresponding to the target entity pair;

[0014] Obtain the relation extraction result based on all relation scores.

[0015] According to an embodiment of the present application, the step of inputting the labeled document into the deduplication module of the document relation extraction model for deduplication sampling processing to obtain the set of entity pairs to be processed includes:

[0016] Set a first set to be processed and initialize the first set to be processed as an empty set;

[0017] Obtain the set of related entity pairs and the set of unrelated entity pairs in the labeled document;

[0018] Randomly shuffle all entity pairs in the set of unrelated entity pairs to obtain a second set to be processed;

[0019] In the second set to be processed, count the first occurrence frequency of the subject and the second occurrence frequency of the object in each unrelated entity pair;

[0020] In the case where both the first occurrence frequency of the subject and the second occurrence frequency of the object in the target unrelated entity pair do not exceed a preset threshold, add the target unrelated entity pair to the first set to be processed, where the target unrelated entity pair is any unrelated entity pair in the second set to be processed;

[0021] After counting the first occurrence frequency of the subject and the second occurrence frequency of the object in each unrelated entity pair, merge the set of related entity pairs and the first set to be processed to obtain the set of entity pairs to be processed.

[0022] According to an embodiment of the present application, the step of inputting the labeled document into the multi-granularity text encoding module of the document relation extraction model for encoding processing to obtain the context embedding vector includes:

[0023] Use the pre-trained language model with the multi-granularity text encoding module to encode the labeled document to obtain a first embedding vector at the word level;

[0024] Use the phrase detection model of the multi-granularity text encoding module to encode the labeled document to obtain a second embedding vector at the phrase level;

[0025] Use the sentence encoder of the multi-granularity text encoding module to encode the labeled document to obtain a third embedding vector at the sentence level;

[0026] Use the fusion module of the multi-granularity text encoding module to fuse the first embedding vector, the second embedding vector, and the third embedding vector to obtain the context embedding vector.

[0027] According to an embodiment of the present application, the step of inputting the set of entity pairs to be processed into the graph convolutional neural network of the document relation extraction model for calculation to obtain a first global embedding representation of the subject and a second global embedding representation of the object in the target entity pair, where the target entity pair is any entity pair in the set of entity pairs to be processed includes:

[0028] Construct an entity graph based on the set of entity pairs to be processed;

[0029] Based on the entity graph, use the graph convolutional neural network to calculate multiple-layer first embedding representations of the subject in the target entity pair, and use the first embedding representation of the last layer as the first global embedding representation of the subject;

[0030] Based on the entity graph, use the graph convolutional neural network to calculate multiple-layer second embedding representations of the object in the target entity pair, and use the second embedding representation of the last layer as the second global embedding representation of the object.

[0031] According to an embodiment of the present application, the step of inputting the first global embedding representation, the second global embedding representation, and the relationship feature into the mixture-of-experts system of the document relation extraction model for scoring calculation to obtain a relationship score corresponding to the target entity pair includes:

[0032] Use the mixture-of-experts system to calculate the first global embedding representation and the relationship feature to obtain a first joint feature;

[0033] Use the mixture-of-experts system to calculate the second global embedding representation and the relationship feature to obtain a second joint feature;

[0034] Obtain a relationship score corresponding to the target entity pair based on the first joint feature and the second joint feature.

[0035] According to an embodiment of the present application, the method for document-level relation extraction based on de-duplication sampling further includes a training step of the document relation extraction model, and the training step of the document relation extraction model includes:

[0036] Obtain a training data set and a true label set, and there is a one-to-one correspondence between each training data in the training data set and one of the true labels in the true label set;

[0037] Use the preset marking to mark the target training data to obtain a training document, and the training document contains at least one set of entity pairs;

[0038] Input the training document into the original document relation extraction model, and the original document relation extraction model performs de-duplication sampling and relation extraction processing on the training document to obtain a prediction result;

[0039] Calculate a loss function based on the prediction result and the true label corresponding to the target training data;

[0040] Adjust the model parameters of the original document relation extraction model based on the loss function to form the document relation extraction model.

[0041] In the second aspect of the embodiments of the present application, a device for document-level relation extraction based on de-duplication sampling is provided, including:

[0042] A text document acquisition module, configured to acquire a text document;

[0043] A marked document acquisition module, configured to use a preset marking to mark the text document to obtain a marked document, and the marked document contains at least one set of entity pairs;

[0044] A model processing module, configured to input the marked document into a document relation extraction model, and the document relation extraction model performs de-duplication sampling and relation extraction processing on the marked document to obtain a relation extraction result.

[0045] In the third aspect of the embodiments of the present application, a terminal device is provided, including: a processor and a memory, where the memory is used to store a computer program, and the processor is used to call and run the computer program stored in the memory to execute the steps of the method for document-level relation extraction based on de-duplication sampling provided in the first aspect of the embodiments of the present application.

[0046] In the fourth aspect of the embodiments of the present application, a computer-readable storage medium is provided. The computer-readable storage medium is used to store a computer program, and the computer program causes a computer to execute the steps of the method for document-level relation extraction based on de-duplication sampling provided in the first aspect of the embodiments of the present application.

[0047] The beneficial effects of the embodiments of the present application include:

[0048] In the embodiments of the present application, a de-duplication sampling method is adopted to balance the distribution of relation categories in the document. Specifically, in the embodiments of the present application, a document relation extraction model with a de-duplication sampling function is set, and the document relation extraction model is used to perform de-duplication sampling and relation extraction processing on the labeled documents that have been marked, so as to obtain the relation extraction result. Compared with the prior art, the embodiments of the present application can reduce the influence of the imbalance of the relation category distribution on the accuracy of the relation extraction result, thereby improving the precision of the relation extraction result. Description of the Drawings

[0049] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required to be used in the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0050] Figure 1 It is a flowchart of the method for document-level relation extraction based on de-duplication sampling of the present application in some embodiments;

[0051] Figure 2 It is a training flowchart of the original document relation extraction model of the present application in some embodiments;

[0052] Figure 3 It is a principle block diagram of the device for document-level relation extraction based on de-duplication sampling of the present application in some embodiments;

[0053] Figure 4 It is a principle block diagram of the terminal device of the present application in some embodiments. Detailed Embodiments

[0054] In order to make the above objects, features, and advantages of the present application more obvious and understandable, the following will describe the detailed embodiments of the present application in conjunction with the drawings. Many specific details are set forth in the following description to facilitate a full understanding of the present application. However, the present application can be implemented in many other ways different from those described herein. Those skilled in the art can make similar improvements without departing from the connotation of the present application. Therefore, the present application is not limited by the specific embodiments disclosed below.

[0055] It should be noted that the terms "first" and "second" are only used for descriptive purposes and should not be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, features defined with "first" and "second" may explicitly or implicitly include at least one of such features. In the description of this application, "a plurality of" means at least two, such as two, three, etc., unless otherwise specifically defined.

[0056] The term "exemplary" or "for example" and the like are used to indicate examples, illustrations or explanations. Any embodiment or design solution described as "exemplary" or "for example" in the embodiments of this application should not be construed as being more preferred or having more advantages than other embodiments or design solutions. Rather, the use of the term "exemplary" or "for example" and the like is intended to present relevant concepts in a specific manner.

[0057] The term "comprise", "include" or any other variant thereof is intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0058] Unless otherwise defined, all technical and scientific terms used in the specification of this application have the same meaning as commonly understood by those skilled in the technical field to which this application belongs. The terms used in the specification of this application are only for the purpose of describing specific embodiments and are not intended to limit this application. The term "and / or" used in the specification of this application includes any and all combinations of one or more of the related listed items.

[0059] The following further describes the specific embodiments of this application in conjunction with the accompanying drawings.

[0060] Refer to Figure 1 As shown, it is a flowchart of a document-level relation extraction method provided in the first aspect of the embodiments of this application. In Figure 1 it, the document-level relation extraction method based on duplicate sampling includes:

[0061] S1. Obtain a text document.

[0062] The text document may be a word document, and the text document records text content.

[0063] S2. Perform a marking process on the text document using a preset marker to obtain a marked document, and the marked document includes at least one set of entity pairs.

[0064] Specifically, the preset mark is "*". In the embodiments of the present application, the preset mark is used to emphasize the starting position and the ending position where the entity appears.

[0065] S3. Input the marked document into a document relationship extraction model, and the document relationship extraction model performs duplicate sampling and relationship extraction processing on the marked document to obtain a relationship extraction result.

[0066] Further, the step S3 includes:

[0067] S31. Input the marked document into the duplicate removal module of the document relationship extraction model for duplicate sampling processing to obtain a set of entity pairs to be processed.

[0068] Further, the step S31 includes:

[0069] S311. Set a first set to be processed and initialize the first set to be an empty set.

[0070] The first set to be processed can be expressed as .

[0071] S312. Obtain a set of related entity pairs and a set of unrelated entity pairs in the marked document.

[0072] Each entity pair in the set of related entity pairs corresponds to at least one non-empty relationship category. The non-empty relationship category can be "founder relationship" or "partner relationship". Of course, the non-empty relationship category is not limited to the above examples and can be set by those skilled in the art according to actual needs.

[0073] The relationship category of each entity pair in the set of unrelated entity pairs is only "unrelated".

[0074] The set of related entity pairs can be expressed as . The set of unrelated entity pairs can be expressed as .

[0075] S313. Randomly shuffle all entity pairs in the set of unrelated entity pairs to obtain a second set to be processed.

[0076] S314. In the second set to be processed, count the first occurrence frequency of the subject and the second occurrence frequency of the object in each unrelated entity pair.

[0077] S315. In the case where both the first occurrence frequency of the subject and the second occurrence frequency of the object in the target non-related entity pair do not exceed a preset threshold, add the target non-related entity pair to the first set to be processed, where the target non-related entity pair is any non-related entity pair in the second set to be processed.

[0078] The target non-related entity pair consists of two entities, one of which is the subject and the other is the object.

[0079] The preset threshold is a specific value, which can be specifically set by those skilled in the art according to actual needs.

[0080] S316. After counting the first occurrence frequency of the subject and the second occurrence frequency of the object in each non-related entity pair, merge the set of related entity pairs and the first set to be processed to obtain the set of entity pairs to be processed.

[0081] The setting of step S31 can limit the occurrence frequency of each entity in the set of non-related entity pairs not to exceed the preset threshold, thereby reducing redundant data and maintaining the integrity of the set of related entity pairs at the same time.

[0082] S32. Input the labeled document into the multi-granularity text encoding module of the document relation extraction model for encoding processing to obtain context embedding vectors.

[0083] The multi-granularity text encoding module is used to perform encoding at the word level, phrase level, and sentence level to capture semantic information at different levels.

[0084] Further, step S32 includes:

[0085] S321. Use the pre-trained language model of the multi-granularity text encoding module to encode the labeled document to obtain the first embedding vector at the word level.

[0086] The pre-trained language model is the BERT (Bidirectional Encoder Representations from Transformers) pre-trained language model. In other embodiments, the pre-trained language model can also be the RoBerta (Robustly Optimized BERT Pretraining Approach) pre-trained language model or the DeBerta (Decoding-enhanced BERT with Disentangled Attention) pre-trained language model, which can be specifically set by those skilled in the art according to actual needs.

[0087] S322. Use the phrase detection model of the multi-granularity text encoding module to encode the labeled document to obtain a second embedding vector at the phrase level.

[0088] The phrase detection model is the SpanBERT model, and the full name of the SpanBERT model is Improving Pre-training by Representing and Predicting Spans (that is, improving pre-training by representing and predicting spans). The SpanBERT model is an extension of BERT.

[0089] S323. Use the sentence encoder of the multi-granularity text encoding module to encode the labeled document to obtain a third embedding vector at the sentence level.

[0090] The sentence encoder is the Sentence-BERT (Sentence Bidirectional Encoder Representations from Transformers) model.

[0091] S324. Use the fusion module of the multi-granularity text encoding module to fuse the first embedding vector, the second embedding vector, and the third embedding vector to obtain the context embedding vector.

[0092] Further, the calculation formula in step S324 is:

[0093] ;

[0094] where is the context embedding vector, , is a real number space of dimension is the length of the labeled document, is the encoding dimension; is the first embedding vector; is the second embedding vector; is the third embedding vector.

[0095] Further, step S32 also includes:

[0096] S325. Use the pre-trained language model to perform a calculation process on the labeled document to obtain word-level attention.

[0097] S326. Use the phrase detection model to perform a calculation process on the labeled document to obtain phrase-level attention.

[0098] S327. Use the sentence encoder to perform computational processing on the labeled document to obtain sentence-level attention.

[0099] S328. Use the fusion module to fuse the word-level attention, the phrase-level attention, and the sentence-level attention to obtain the context attention.

[0100] Further, the calculation formula in step S328 is:

[0101] ;

[0102] where is the context attention, , is a real number space of dimension is the number of attention heads, is the length of the labeled document; is the word-level attention; is the phrase-level attention; is the sentence-level attention; is a fusion function, and the fusion function has a merging function.

[0103] S33. Input the set of entity pairs to be processed into the graph convolutional neural network of the document relation extraction model for computational processing to obtain a first global embedding representation of the subject and a second global embedding representation of the object in the target entity pair, where the target entity pair is any entity pair in the set of entity pairs to be processed.

[0104] Further, step S33 includes:

[0105] S331. Construct an entity graph based on the set of entity pairs to be processed.

[0106] Each entity pair in the set of entity pairs to be processed contains two entities, one of which is the subject and the other is the object.

[0107] The entity graph is expressed as ; is the entity graph; is all the nodes of the entity graph, that is, the set of all entities in the set of entity pairs to be processed, , is the first node of the entity graph, the first entity in the set of entity pairs to be processed, is the second entity in the set of entity pairs to be processed, and so on, is the th entity in the set of entities to be processed, and is the total number of all entities in the set of entities to be processed; entity and have a semantic relationship ; is the th node of the entity graph, that is, the th entity in the set of entities to be processed ; is the th node of the entity graph, that is, the th entity in the set of entities to be processed .

[0108] S332. Based on the entity graph, use the graph convolutional neural network to calculate the multi-layer first embedding representation of the subject in the target entity pair, and use the first embedding representation of the last layer as the first global embedding representation of the subject.

[0109] Furthermore, the calculation formula for the first embedding representation of the th layer in the multi-layer first embedding representation is:

[0110] ;

[0111] where is the first embedding representation of the th layer of the subject in the target entity pair; is the non-linear activation function; is the set of neighbor nodes of node , that is, the set of neighbor nodes of the subject in the target entity pair; is the set of neighbor nodes of node , that is, the set of neighbor nodes of entity ; is the learnable weight matrix of the first embedding representation of the th layer of the subject in the target entity pair, , , is dimensional real space, is the embedding dimension; is the first embedding representation of the th layer of the subject ; is the first embedding representation of the th layer of the subject The bias term of the first-layer first embedding representation is a normalization coefficient used to balance the influence of neighbor nodes.

[0112] It should be understood that the first-layer first embedding representation of each entity in the set of entity pairs to be processed belongs to , where is a real number space of dimension; is the embedding dimension. Exemplarily, the first-layer first embedding representation of the subject in the target entity pair is .

[0113] Furthermore, the calculation formula for the first global embedding representation is:

[0114] ;

[0115] where is the first global embedding representation, is the subject in the target entity pair; is the first embedding representation of the last layer; is the last layer in the multi-layer first embedding representation.

[0116] S333. Based on the entity graph, use the graph convolutional neural network to calculate the multi-layer second embedding representation of the object in the target entity pair, and use the second embedding representation of the last layer as the second global embedding representation of the object.

[0117] Furthermore, the calculation formula for the second embedding representation of the th layer in the multi-layer second embedding representation is:

[0118] ;

[0119] where is the second embedding representation of the th layer of the object in the target entity pair ; is a non-linear activation function; is the set of neighbor nodes of node , that is, the set of neighbor nodes of the object in the target entity pair ; is the set of neighbor nodes of node , that is, the set of neighbor nodes of entity ; is the learnable weight matrix of the second embedding representation of the th layer of the object in the target entity pair , ,​ is a real number space of dimension ; is the th second embedding representation of the layer of the object is the bias term of the th second embedding representation of the layer of the object in the target entity pair is a normalization coefficient used to balance the influence of neighbor nodes.

[0120] Furthermore, the calculation formula for the second global embedding representation is:

[0121] ;

[0122] where is the second global embedding representation is the object in the target entity pair; is the second embedding representation of the last layer; is the last layer in the multi-layer second embedding representation.

[0123] S34. Input the target entity pair and the context embedding vector into the context pooling module of the document relation extraction model for local context pooling processing to obtain relation features.

[0124] Furthermore, step S34 includes:

[0125] S341. Calculate the first attention weight of the subject in the target entity pair.

[0126] Furthermore, the calculation formula for step S341 is:

[0127] ;

[0128] where is the first attention weight; is the number of mentions of the subject in the target entity pair is the self-attention weight of the th appearance of the subject in the target entity pair at the is a real number space of dimension is the number of attention heads; is the length of the labeled document; is the subscript corresponding to the subject in the target entity pair.

[0129] S342. Calculate the second attention weight of the object in the target entity pair.

[0130] Further, the calculation formula for step S342 is:

[0131] ;

[0132] Wherein, is the second attention weight; is the number of mentions of the object in the target entity pair ; is the self-attention weight of the position where the th appearance of the object in the target entity pair is located ; is a real number space of dimension ; is the number of attention heads; is the length of the labeled document;

[0133] S343. Aggregate the first attention weight and the second attention weight in a way of summing and averaging to obtain the importance distribution of the target entity pair.

[0134] Further, the calculation formula for step S343 is:

[0135] ;

[0136] Wherein, is the importance distribution of the target entity pair; is the first attention weight; is the Hadamard product; is the second attention weight; is the transpose of the first attention weight, is the subscript corresponding to the subject in the target entity pair, is the subscript corresponding to the object in the target entity pair.

[0137] S344. Calculate the relationship feature based on the importance distribution and the context embedding vector.

[0138] Further, the calculation formula for step S344 is:

[0139] ;

[0140] Wherein, is the relationship feature, and it is also the weighted average of all words in the labeled document; is the transpose of the context embedding vector ; is the importance distribution of the target entity pair; is the subscript corresponding to the subject in the target entity pair, is the subscript corresponding to the object in the target entity pair.

[0141] S35. Input the first global embedding representation, the second global embedding representation, and the relation feature into the mixture-of-experts system of the document relation extraction model for scoring calculation processing to obtain a relation score corresponding to the target entity pair.

[0142] The mixture-of-experts system MoE (Mixture of Experts) combines an expert network and a gating mechanism to improve the flexibility and performance of the model. The expert network focuses on the fine-grained analysis of input features, while the gating mechanism is responsible for dynamically selecting the top weights of the experts and weighted generating the final feature representation.

[0143] MoE contains 16 experts, and the calculation formula for each expert is:

[0144] ;

[0145] where, is the layer normalization operation; is the Gaussian error linear unit activation function; is the learnable weight matrix of the th expert, is a real number space of dimension ; is the input feature; is the learnable bias term of the th expert, is a real number space of dimension

[0146] The gating mechanism dynamically assigns weights according to the input feature and selects the top weights of the experts.

[0147] The calculation formula of the gating mechanism is:

[0148] ;

[0149] where, is the function for normalizing the weights; Before selection the top learnable weight matrix for the gating mechanism , is a real number space of dimension coding dimension; input feature; learnable bias term for the gating mechanism , is a real number space of dimension 16.

[0150] The final output of the Mixture of Experts (MoE) system is the weighted sum of the top experts. The calculation formula for the final output of the MoE system is:

[0151] ;

[0152] where is the final output; is the weight assigned by the gating mechanism to the -th expert; is the -th expert's output for the input feature .

[0153] Furthermore, step S35 includes:

[0154] S351. Use the MoE system to perform calculation processing on the first global embedding representation and the relation feature to obtain a first joint feature.

[0155] Furthermore, the calculation formula for step S351 is:

[0156] ;

[0157] where is the first joint feature; is the MoE system; is the first global embedding representation; is the relation feature.

[0158] S352. Use the MoE system to perform calculation processing on the second global embedding representation and the relation feature to obtain a second joint feature.

[0159] Furthermore, the calculation formula for step S352 is:

[0160] ;

[0161] where is the second combined feature; is the hybrid expert system; is the second global embedding representation; is the relational feature.

[0162] S353. Obtain a relational score corresponding to the target entity pair based on the first combined feature and the second combined feature.

[0163] Further, the calculation formula for step S353 is:

[0164] ;

[0165] where is the relational score corresponding to the target entity pair; is a learnable parameter, , is a real number space of dimension is the encoding dimension; is the first combined feature; is a learnable parameter, ; is the second combined feature.

[0166] S36. Obtain the relation extraction result based on all relational scores.

[0167] In the embodiments of the present application, it may be to perform screening processing on all relational scores to obtain the relation extraction result.

[0168] In the embodiments of the present application, a resampling method is adopted to balance the distribution of relation categories in the document. Specifically, in the embodiments of the present application, a document relation extraction model with a resampling function is set, and the marked documents after being marked are subjected to resampling and relation extraction processing by using the document relation extraction model, so as to obtain the relation extraction result. Compared with the prior art, the embodiments of the present application can reduce the influence of the imbalance of the relation category distribution on the accuracy of the relation extraction result, thereby improving the precision of the relation extraction result.

[0169] The use of the hybrid expert system can improve the accuracy of target relation modeling and overcome the problem of information focus in current document-level relation extraction that is vulnerable to noise interference.

[0170] In some embodiments, the document-level relation extraction method based on resampling further includes a training step of the document relation extraction model, and the training step of the document relation extraction model includes:

[0171] S4. Obtain a training data set and a true label set, where each training data in the training data set has a one-to-one correspondence with one of the true labels in the true label set. The steps of S4 can refer to Figure 2 the step of "obtaining a training data set and a true label set" in

[0172] Each training data in the training data set is a kind of text document.

[0173] S5. Use the preset tags to mark the target training data to obtain a training document, and the training document contains at least one set of entity pairs.

[0174] S6. Input the training document into the original document relation extraction model, and the original document relation extraction model performs duplicate removal sampling and relation extraction processing on the training document to obtain a prediction result.

[0175] The original document relation extraction model includes an original duplicate removal module, an original multi-granularity text encoding module, an original graph convolutional neural network, an original context pooling module, and an original mixture of experts system.

[0176] Further, the steps of S6 include:

[0177] S61. Input the training document into the original duplicate removal module of the original document relation extraction model for duplicate removal sampling processing to obtain a training set of entity pairs to be processed. The steps of S61 can refer to Figure 2 the step of "duplicate removal sampling" in

[0178] S62. Input the training document into the original multi-granularity text encoding module of the original document relation extraction model for encoding processing to obtain a training context embedding vector. The steps of S62 can refer to Figure 2 the step of "multi-granularity text encoding" in

[0179] S63. Input the training set of entity pairs to be processed into the original graph convolutional neural network of the original document relation extraction model for calculation processing to obtain a first training global embedding representation of the subject and a second training global embedding representation of the object in the target training entity pair, where the target training entity pair is any entity pair in the training set of entity pairs to be processed. The steps of S63 can refer to Figure 2 the step of "entity embedding based on graph neural network" in

[0180] S64. Input the target training entity pair and the training context embedding vector into the original context pooling module of the original document relation extraction model for local context pooling processing to obtain training relation features. The steps of S64 can refer to Figure 2 the step of "local context pooling" in

[0181] S65. Input the first training global embedding representation, the second training global embedding representation, and the training relationship features into the original mixture-of-experts system of the original document relationship extraction model for scoring calculation to obtain a training relationship score corresponding to the target training entity pair. Step S65 can refer to Figure 2 the step of "calculating the relationship score by the original mixture-of-experts system" in

[0182] S66. Obtain the prediction result based on all the training relationship scores.

[0183] The difference between the original document relationship extraction model and the document relationship extraction model lies in the different model parameters.

[0184] S7. Calculate the loss function based on the prediction result and the true label corresponding to the target training data. Step S7 can refer to Figure 2 the step of "calculating the loss function" in

[0185] Furthermore, the calculation formula of the loss function is:

[0186] ;

[0187] ;

[0188] where is the loss function; is the relationship of the th entity pair in the target training data; is the true label, that is, the true label corresponding to the th entity pair in the target training data; is the relationship prediction probability; is the relationship score corresponding to the th entity pair in the target training data; is the relationship score corresponding to the th entity pair in the target training data; is the set of all relationships in the target training data; is the exponential function. When is the correct relationship, is equal to 1, otherwise is equal to 0.

[0189] S8. Adjust the model parameters of the original document relationship extraction model based on the loss function to form the document relationship extraction model. Step S8 can refer to Figure 2 the step of "correcting the model parameters" in

[0190] The document relationship extraction model is a trained model.

[0191] Reference Figure 3 As shown, it is a schematic block diagram of a document-level relationship extraction device based on duplicate removal sampling provided in the second aspect of the embodiments of the present application. In Figure 3 it, the document-level relationship extraction device 100 based on duplicate removal sampling includes:

[0192] A text document acquisition module 101, configured to acquire a text document;

[0193] A labeled document acquisition module 102, configured to perform a labeling process on the text document using a preset label to obtain a labeled document, where the labeled document includes at least one set of entity pairs;

[0194] A model processing module 103, configured to input the labeled document into a document relationship extraction model, and the document relationship extraction model performs duplicate removal sampling and relationship extraction processing on the labeled document to obtain a relationship extraction result.

[0195] The third aspect of the embodiments of the present application provides a terminal device, and the schematic block diagram of the terminal device can be as Figure 4 shown. The terminal device includes a processor, a memory, a network interface, a display screen, and a temperature sensor connected through a system bus. Among them, the processor is used to provide computing and control capabilities. The memory includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the terminal device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, it implements a document-level relationship extraction method based on duplicate removal sampling. The display screen can be a liquid crystal display screen or an electronic ink display screen, and the temperature sensor is pre-set inside the terminal device to detect the operating temperature of the internal device.

[0196] Those skilled in the art can understand that Figure 4 the schematic block diagram shown in it is only a block diagram of some structures related to the solution of the present invention, and does not constitute a limitation on the terminal device to which the solution of the present invention is applied. The specific terminal device may include more or fewer components than those shown in the figure, or combine some components, or have a different component layout.

[0197] In some embodiments, the embodiments of the present application provide a terminal device, which includes a processor and a memory. The memory is used to store a computer program, and the processor is used to call and run the computer program stored in the memory to execute the steps of the method for document-level relation extraction based on resampling provided in the first aspect of the embodiments of the present application. In the fourth aspect of the embodiments of the present application, a computer-readable storage medium is provided. The computer-readable storage medium is used to store a computer program, and the computer program causes a computer to execute the steps of the method for document-level relation extraction based on resampling provided in the first aspect of the embodiments of the present application.

[0198] Those of ordinary skill in the art can understand that all or part of the processes of implementing the methods in the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium used in the various embodiments provided by the present invention can include non-volatile and / or volatile memories. Non-volatile memories can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memories can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and Rambus dynamic RAM (RDRAM), etc.

[0199] Without changing the basic principles of the present application, the technical features of the above embodiments can be combined. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.

[0200] The above embodiments only illustrate several implementation manners of the present application. The description thereof is relatively specific and detailed, but it should not be construed as a limitation to the scope of the patented application. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all fall within the protection scope of the present application. Therefore, the patent protection scope of the present application shall be subject to the appended claims.

Claims

1. A document-level relation extraction method based on deduplication sampling, characterized in that Including: Obtain a text document; Perform a marking process on the text document using a preset marker to obtain a marked document, where the marked document contains at least one set of entity pairs; Input the marked document into a document relation extraction model, and the document relation extraction model performs duplicate removal sampling and relation extraction processing on the marked document to obtain a relation extraction result; The steps of inputting the marked document into a document relation extraction model, where the document relation extraction model performs duplicate removal sampling and relation extraction processing on the marked document to obtain a relation extraction result include: inputting the marked document into the duplicate removal module of the document relation extraction model for duplicate removal sampling processing to obtain a set of entity pairs to be processed; inputting the marked document into the multi-granularity text encoding module of the document relation extraction model for encoding processing to obtain context embedding vectors; the multi-granularity text encoding module is used to perform encoding at the word level, phrase level, and sentence level to capture semantic information at different levels; inputting the set of entity pairs to be processed into the graph convolutional neural network of the document relation extraction model for calculation processing to obtain a first global embedding representation of the subject and a second global embedding representation of the object in the target entity pair, where the target entity pair is any entity pair in the set of entity pairs to be processed; inputting the target entity pair and the context embedding vectors into the context pooling module of the document relation extraction model for local context pooling processing to obtain relation features; inputting the first global embedding representation, the second global embedding representation, and the relation features into the mixture of experts system of the document relation extraction model for scoring calculation processing to obtain a relation score corresponding to the target entity pair; obtaining the relation extraction result based on all relation scores; The steps of inputting the marked document into the duplicate removal module of the document relation extraction model for duplicate removal sampling processing to obtain a set of entity pairs to be processed include: setting a first set to be processed and initializing the first set to be processed as an empty set; obtaining a set of related entity pairs and a set of unrelated entity pairs in the marked document; randomly shuffling all entity pairs in the set of unrelated entity pairs to obtain a second set to be processed; counting the first occurrence frequency of the subject and the second occurrence frequency of the object in each unrelated entity pair in the second set to be processed; in the case where the first occurrence frequency of the subject and the second occurrence frequency of the object in the target unrelated entity pair do not exceed a preset threshold, adding the target unrelated entity pair to the first set to be processed, where the target unrelated entity pair is any unrelated entity pair in the second set to be processed; after counting the first occurrence frequency of the subject and the second occurrence frequency of the object in each unrelated entity pair, merging the set of related entity pairs and the first set to be processed to obtain the set of entity pairs to be processed; Input the set of entity pairs to be processed into the graph convolutional neural network of the document relation extraction model for computational processing to obtain a first global embedding representation of the subject and a second global embedding representation of the object in the target entity pair. The steps for the target entity pair, which is any entity pair in the set of entity pairs to be processed, include: constructing an entity graph based on the set of entity pairs to be processed; using the graph convolutional neural network to calculate multiple-layer first embedding representations of the subject in the target entity pair based on the entity graph, and taking the first embedding representation of the last layer as the first global embedding representation of the subject; using the graph convolutional neural network to calculate multiple-layer second embedding representations of the object in the target entity pair based on the entity graph, and taking the second embedding representation of the last layer as the second global embedding representation of the object; In the multi-layer first embedding representation, the -th layer first embedding representation is calculated as follows: ; where is the -th layer first embedding representation of the subject in the target entity pair; is a non-linear activation function; is the set of neighbor nodes of node ; is the set of neighbor nodes of node ; is the learnable weight matrix of the -th layer first embedding representation of the subject in the target entity pair, , is a real number space of dimension ; is the -th layer first embedding representation of the subject; is the bias term of the -th layer first embedding representation of the subject in the target entity pair, is the normalization coefficient.​​​​ 2. The method for document-level relation extraction based on de-resampling according to claim 1, wherein The steps for inputting the labeled document into the multi-granularity text encoding module of the document relation extraction model for encoding processing to obtain a context embedding vector include: Using the pre-trained language model of the multi-granularity text encoding module to encode the labeled document to obtain a first embedding vector at the word level; Using the phrase detection model of the multi-granularity text encoding module to encode the labeled document to obtain a second embedding vector at the phrase level; Using the sentence encoder of the multi-granularity text encoding module to encode the labeled document to obtain a third embedding vector at the sentence level; Using the fusion module of the multi-granularity text encoding module to fuse the first embedding vector, the second embedding vector, and the third embedding vector to obtain the context embedding vector.

3. The method for document-level relation extraction based on deduplication sampling according to claim 1, wherein The steps for inputting the first global embedding representation, the second global embedding representation, and the relation feature into the mixture of experts system of the document relation extraction model for scoring calculation processing to obtain a relation score corresponding to the target entity pair include: Using the mixture of experts system to calculate and process the first global embedding representation and the relation feature to obtain a first joint feature; Using the mixture of experts system to calculate and process the second global embedding representation and the relation feature to obtain a second joint feature; Obtaining a relation score corresponding to the target entity pair based on the first joint feature and the second joint feature.

4. The method for document-level relation extraction based on de-duplication sampling according to claim 1, wherein The document-level relation extraction method based on de-duplication sampling further includes the training step of the document relation extraction model. The training step of the document relation extraction model includes: Obtaining a training data set and a true label set, where each training data in the training data set has a one-to-one correspondence with one of the true labels in the true label set; Using the preset marking to mark the target training data to obtain a training document, and the training document contains at least one group of entity pairs; Inputting the training document into the original document relation extraction model, and the original document relation extraction model performs de-duplication sampling and relation extraction processing on the training document to obtain a prediction result; Calculating a loss function based on the prediction result and the true label corresponding to the target training data; Adjust the model parameters of the original document relation extraction model based on the loss function to form the document relation extraction model.

5. A document-level relation extraction device based on de-duplication sampling, characterized in that, It includes: A text document acquisition module for acquiring text documents; A labeled document acquisition module for performing labeling processing on the text document using a preset label to obtain a labeled document, where the labeled document contains at least one set of entity pairs; A model processing module for inputting the labeled document into a document relation extraction model, and the document relation extraction model performs duplicate sampling and relation extraction processing on the labeled document to obtain a relation extraction result; The document-level relation extraction device based on duplicate sampling is further configured to input the labeled document into the duplicate sampling module of the document relation extraction model for duplicate sampling processing to obtain a set of entity pairs to be processed; Input the labeled document into the multi-granularity text encoding module of the document relation extraction model for encoding processing to obtain a context embedding vector; the multi-granularity text encoding module is used to perform encoding at the word level, phrase level, and sentence level to capture semantic information at different levels; input the set of entity pairs to be processed into the graph convolutional neural network of the document relation extraction model for calculation processing to obtain a first global embedding representation of the subject and a second global embedding representation of the object in the target entity pair, where the target entity pair is any entity pair in the set of entity pairs to be processed; input the target entity pair and the context embedding vector into the context pooling module of the document relation extraction model for local context pooling processing to obtain a relation feature; input the first global embedding representation, the second global embedding representation, and the relation feature into the mixture of experts system of the document relation extraction model for scoring calculation processing to obtain a relation score corresponding to the target entity pair; obtain the relation extraction result based on all relation scores; The document-level relation extraction device based on duplicate sampling is further configured to set a first set to be processed and initialize the first set to be processed as an empty set; obtain a set of related entity pairs and a set of unrelated entity pairs in the labeled document; randomly shuffle all entity pairs in the set of unrelated entity pairs to obtain a second set to be processed; count the first occurrence frequency of the subject and the second occurrence frequency of the object in each unrelated entity pair in the second set to be processed; in the case where the first occurrence frequency of the subject and the second occurrence frequency of the object in the target unrelated entity pair do not exceed a preset threshold, add the target unrelated entity pair to the first set to be processed, where the target unrelated entity pair is any unrelated entity pair in the second set to be processed; after counting the first occurrence frequency of the subject and the second occurrence frequency of the object in each unrelated entity pair, merge the set of related entity pairs and the first set to be processed to obtain the set of entity pairs to be processed; The document-level relation extraction device based on de-duplication sampling is further configured to construct an entity graph based on the set of entity pairs to be processed; based on the entity graph, use the graph convolutional neural network to calculate multiple-layer first embedding representations of the subject in the target entity pair, and use the first embedding representation of the last layer as the first global embedding representation of the subject; based on the entity graph, use the graph convolutional neural network to calculate multiple-layer second embedding representations of the object in the target entity pair, and use the second embedding representation of the last layer as the second global embedding representation of the object. The calculation formula for the first embedding representation of the -th layer in the multi-layer first embedding representation is: ; where is the first embedding representation of the subject in the target entity pair at the -th layer; is a non-linear activation function; is the set of neighbor nodes of node ; is the set of neighbor nodes of node ; is the learnable weight matrix of the first embedding representation of the subject in the target entity pair at the -th layer, , is a real number space of dimension ; is the first embedding representation of the subject at the -th layer; is the bias term of the first embedding representation of the subject in the target entity pair at the -th layer, is the normalization coefficient.

6. A terminal device, characterized in that, It includes: a processor and a memory, the memory is used to store a computer program, and the processor is used to call and run the computer program stored in the memory to execute the steps of the document-level relation extraction method based on de-duplication sampling according to any one of claims 1 to 4 above.

7. A computer-readable storage medium, characterized in that, For storing a computer program, the computer program causes a computer to execute the steps of the document-level relation extraction method based on de-duplication sampling according to any one of claims 1 to 4 above.

Citation Information

Patent Citations

  • Improved relation extraction method combining entity relation and local information

    CN114791953A

  • Reference document-oriented document-level legal relationship extraction model product and method

    CN118396103A