Entity Relationship Extraction Model Training Method and Entity Relationship Joint Extraction Method

Through the knowledge distillation framework and active learning method of teacher and student models, the training of entity relationship extraction models in vertical fields is optimized, the problem of insufficient accuracy and generalization ability is solved, and high-quality triple data set construction is achieved.

CN119227742BActive Publication Date: 2025-06-24INNER MONGOLIA UNIVERSITY
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202411292512.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-09-14
Publication Date
2025-06-24
Estimated Expiration
2044-09-14

AI Technical Summary

Technical Problem

In the vertical field, the prior art has low accuracy, poor generalization ability and low recall rate when extracting entity relationships.

Method used

A knowledge distillation framework composed of teacher-student models is adopted, and the teacher model and student model is iteratively optimized, combined with manual annotation and active learning methods, the training data set is expanded to improve the training accuracy and generalization ability of the model.

Benefits of technology

In vertical fields, the accuracy and generalization ability of entity relationship extraction have been significantly improved, and a high-quality triple data set has been constructed, suitable for multiple vertical fields.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119227742B_ABST
    Figure CN119227742B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for training an entity relationship extraction model and a method for jointly extracting entity relationships. The method for training the entity relationship extraction model includes: Step S1, combining a teacher model and a student model to form an entity relationship extraction model with a knowledge distillation framework, inputting a training data set, and outputting triple information of text data by the student model; Step S2, determining whether the accuracy of the triple information output by the current model is the maximum. If not, screening out the data that needs to be manually labeled, manually labeling it, and fusing it with the initial training data set to obtain an optimized training data set and continuing to input it into the model to train the model until the accuracy reaches the maximum and then completing the model training. By combining the active learning method with the entity relationship joint extraction model, the present invention adopts the active learning method to expand the corpus, and can continuously form new training sets to refine the parameter values, having good generalization ability and accuracy, and being applicable to vertical fields with small data volumes.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of natural language processing, and particularly to a method for training an entity relationship extraction model and a method for jointly extracting entity relationships. Background Art

[0002] Natural language processing (NPL) is an important direction in the fields of computer science and artificial intelligence. Information extraction is a basic research in natural language processing. Entity relationship extraction is one of the key tasks in information extraction, and its purpose is to extract entities and relationships from unstructured text data to form structured knowledge. Relational triples are usually represented in the triple form of subject, relational predicate, and object, where the subject and object are meaningful named entities, and the relational predicate is usually one of several predefined relationship types.

[0003] Traditional joint extraction models generally fall into two categories: joint extraction models with shared parameters and joint decoding joint extraction models. The former realizes joint extraction by sharing parameters (sharing input features or internal hidden layer states). This model usually uses a shared neural network to encode the input text and performs entity recognition and relationship extraction on the basis of the encoding. The latter is to strengthen the interaction between the entity model and the relationship model, proposes a joint decoding algorithm, and uses two independent decoders to perform entity recognition and relationship extraction on the input text. The difficulty of joint extraction lies in how to strengthen the interaction between the entity model and the relationship model and further utilize the potential information between the two models.

[0004] With the continuous development of deep learning, deep learning methods are mostly used for entity extraction, but the training of its models requires a large amount of labeled data and is costly. Therefore, active learning is applied to the deep learning models for entity and relationship extraction to minimize the manually labeled data required for model training while maintaining the performance of the extraction model. At the same time, with the rapid development of information technology, the amount of data in each field shows an explosive growth trend, generating a large number of domain-specific terms. The information contained in these redundant and complex domain data is difficult to utilize. Therefore, the application of entity extraction technology in each vertical field has important value. However, due to the small amount of data in each vertical field and the small amount of corpus labeled data, when using the above traditional joint extraction model for entity relationship extraction in the vertical field, there are defects such as low accuracy, poor generalization ability, and low recall rate. Summary of the Invention

[0005] The present application proposes a method for training an entity relationship extraction model and a method for jointly extracting entity relationships, aiming to solve the problems of low accuracy and poor generalization ability existing in the entity relationship extraction in the vertical field in the prior art.

[0006] The technical solution adopted by the invention is: a method for training an entity relationship extraction model, and the method for training an entity relationship extraction model includes: Step S1, combining a teacher model and a student model to form an entity relationship extraction model with a knowledge distillation framework, inputting a training data set, and outputting triple information of text data by the student model;

[0007] Step S2, determining whether the accuracy of the triple information output by the current model is the maximum. If not, screening out the data that needs to be manually labeled, manually labeling it, and then fusing it with the initial training data set to obtain an optimized training data set, and continuing to input the model to train the model until the accuracy reaches the maximum, and then completing the model training.

[0008] In an optional implementation manner, further, it further includes an iterative step of the teacher model and the student model: training the teacher and student models on the latest optimized training data set until the test accuracy reaches the maximum, and using knowledge distillation to synchronize the knowledge between the teacher model and the student model in the current iteration.

[0009] Further, through cross-training of the teacher model and the student model, optimizing the student model with the minimum loss function as the goal, the steps include: S3.1: Calculating the cross-entropy between the Soft-target obtained by the teacher model through softmax at temperature T and the Soft-prediction obtained by the student model through softmax at the same temperature as the first loss: where, refers to the value of the softmax output of the teacher model on the i-th class under the condition that the temperature is equal to T, refers to the value of the softmax output of the student model on the i-th class under the condition that the temperature is equal to T;

[0010] S3.2: Calculating the cross-entropy between the Hard-target obtained by the student model through softmax at temperature 1 and the actual value as the second loss: where, c i refers to the true label value on the i-th class, c i ∈{0,1}, where 1 represents a positive label and 0 represents a negative label, refers to the value of the softmax output of the student model on the i-th class under the condition that the temperature is equal to 1;

[0011] S3.3: Combining the first loss and the second loss, the total loss function: L = αL soft +βL hard , optimizing the parameters of the student model with the total loss function L as the minimum goal, where both α and β are hyperparameters.

[0012] Further, step S1 further includes the step of expanding the training data set, and the specific method includes:

[0013] S1.1, divide the input training data set into a training set and a test set;

[0014] S1.2, input the training set of the training data set into the model for training, and test the model on the test set of the training data set, and calculate the test accuracy of the training data set;

[0015] S1.3, determine whether the test accuracy of the current training data set is greater than the test accuracy of the previous training data set. If so, go to step S1.4; if not, go to step S1.5;

[0016] S1.4, use the current teacher model as the new model and synchronously perform knowledge distillation of the student model, and determine whether the uncertain data is exhausted. If exhausted, stop training; if not, send the training set into the example selection algorithm and perform manual annotation and then fuse it into the current training data set again to obtain an optimized training data set, and return the optimized training data set to step S1.1 to continue iterative training until the data is exhausted;

[0017] S1.5, determine whether the uncertain data is exhausted. If exhausted, stop training; if not, send the training set into the example selection algorithm and perform manual annotation and then fuse it into the current training data set again to obtain an optimized training data set, and return the optimized training data set to step S1.1 to continue iterative training. The iterative training continuously feeds back to the teacher model to refine the parameter values until the performance of the teacher model and the student model reaches the maximum.

[0018] Further, an entity relationship joint extraction method is proposed. The entity relationship joint extraction method is an entity relationship extraction model composed of a teacher model and a student model with a knowledge distillation framework obtained based on the entity relationship extraction model training method described in any one of claims 1 to 4.

[0019] In an alternative embodiment, further, in the student model, the input mapping relationship of the entity is extracted, and in the teacher model, the entity is mapped by the input of the relationship. An entity relationship extraction model with a knowledge distillation framework is composed of the teacher model and the student model.

[0020] Further, the teacher model is a pre-trained BERT model. The method for the teacher model to obtain triple information includes:

[0021] Step S4.1: Data preprocessing, perform word segmentation on the text data, perform vector conversion on the word sequence to obtain the text feature h of the input text, and perform sub-word alignment to obtain h avg, in the set h of vectorized representations of sentences; h = BERT(s), h avg = Avgpool(h), where Avgpool is the average pooling operation, aiming to align the embedding transformation length with the original sentence length;

[0022] Step S4.2: Use a label classifier to perform multi-label binary classification on the sentence text features and the relationship information, and obtain a subset P of potential relationships that may exist in the sentence according to the input privileged relationship information rel : P rel = σ(W r h avg + b r ), where σ represents the sigmoid function, and and respectively represent the weight and bias parameters when calculating the subset of potential relationships, reflecting the existence probability of different relationships;

[0023] Step S4.3: Perform two sequence tagging operations to extract the corresponding subject and object respectively, so as to extract the complete triple information: where and respectively represent the probability distributions that the i-th tag is the subject or object of the j-th relationship, u j is the j-th relationship representation in the embeddable matrix U, is the encoding representation of the i-th token after subword alignment, W sub 、b sub and W obj are the weight and bias parameters when calculating the subject and object corresponding to the relationship.

[0024] Furthermore, the method for the student model to obtain triple information includes:

[0025] Step S5.1: Combine the GloVe embedding X g and the trainable position embedding X p , and use a convolutional encoder containing L stacked blocks to encode the text: X = [X g ; X p ; H = Block(…(Block(X))); where, [;] represents the concatenation operation, and each Block contains two dilated convolutions with a dilation rate of p i , a gated unit, and a residual link; padding is used for each Block to ensure that the output dimension matches the input dimension:

[0026] Y a = DilatedConv a (X);

[0027] Y b= DilatedConv b (X);

[0028]

[0029] where DilatedConv represents the dilated convolution module, represents element-wise multiplication, and Y i refers to the output of the i-th Block and the input of the (i + 1)-th Block. The sentence means that H is equivalent to the output Y of the last Block; L equivalent;

[0030] Step S5.2: Generate the subject auxiliary feature H and the object auxiliary feature H from H through two different self-attention modules; h and the object auxiliary feature H t ;

[0031] Step S5.3: The two lines of S → O and O → S are performed in parallel. The sentence is concatenated with the corresponding auxiliary feature and fed into the feed-forward network to extract the subject information S and the object information O respectively. Then, the obtained subject information S is used to guide the extraction of the object, and the obtained object information O is used to guide the extraction of the subject.

[0032] A computer-readable storage medium stores computer program instructions, which, when executed by a processor, implement the above-mentioned entity relationship extraction model training method or the above-mentioned entity relationship joint extraction method.

[0033] Compared with the prior art, the present invention relies on a teacher-student model to form an entity relationship extraction model with a knowledge distillation framework. Through continuous optimization of the teacher model and the student model, in the case of less data, artificial annotation is combined to actively expand the corpus, so as to continuously form a new training set, which is then continuously fed back to the teacher model to refine the parameter values until the performance of the teacher model and the student model saturates. The active learning method is used to expand the corpus, improve the model effect and construct a high-quality vertical domain triple dataset at the same time, making the present invention good at solving the entity relationship joint extraction task in the low-resource background of the vertical domain, and having good generalization ability in vertical domains with small data volume and less labeled data, and can be flexibly applied in multiple vertical domains such as cultural tourism, medical care, and finance.

[0034] In addition, the present invention also aims at the previous framework model for extracting the subject entity and the object entity in sequence. A bidirectional entity extraction framework is adopted to avoid the failure of the entire triple extraction caused by the failure of subject entity extraction. The two-step entity relationship extraction model is combined. In the student model, the input mapping relationship of the entity is extracted, and in the teacher model, the entity of the relationship input mapping is used. The entity relationship triple is extracted by the joint training method of the teacher-student model; by combining the active learning method with the entity relationship joint extraction model, the situation of less labeled data in the vertical domain corpus is effectively dealt with, thereby further improving the accuracy of triple extraction. BRIEF DESCRIPTION OF THE DRAWINGS

[0035] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the embodiments or the description of the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0036] Figure 1 It is the overall structure diagram of the entity relationship joint extraction model in the present invention;

[0037] Figure 2 It is the overall framework diagram of the entity relationship joint extraction model in the present invention;

[0038] Figure 3 It is the framework diagram of the teacher model in the present invention;

[0039] Figure 4 It is the framework diagram of the student model in the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0040] In order to make the technical problems, technical solutions and beneficial effects to be solved by the present invention more clearly understood, the present invention will be further described in detail below with reference to the drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0041] The present invention proposes a training method for an entity relationship extraction model, as Figure 1As shown, the entity relation extraction model in the present invention is mainly a relation extraction model with a knowledge distillation framework structure composed of a teacher model and a student model. Input text data from a corpus dataset into the teacher model and the student model. Combine the text data with manually annotated examples to pre-train the teacher model to obtain a teacher model with relatively high accuracy. The teacher model synchronously trains a student model through knowledge distillation. The student model outputs the triple information of the corresponding text data. Judge the accuracy of the triple information output by the student model. Send the text data with unqualified accuracy to the example selection module. The example selection module uses the example selection algorithm and combines the active learning method to iteratively update the teacher model. Each updated teacher model synchronously trains the student model. Continue to output triple information by the updated student model until the relation extraction capabilities of the teacher model and the student model for the text data reach the best. The following will be described in detail.

[0042] First, in combination with the attached Figure 1 、 2 As shown, the entity relation extraction model in the present application is a relation extraction model with a knowledge distillation framework composed of a teacher model and a student model. Among them, knowledge distillation is a model compression technology, aiming to transfer the knowledge of a complex model (that is, its parameters and learned representations) to a simplified model, so as to reduce the computational and storage requirements of the model while maintaining performance. The relatively complex model is called the teacher model, and the relatively simple model is called the student model. The teacher model is usually a complex or large model, which performs well on a given task and has high performance. The student model is a simplified model, usually with fewer parameter and computational resource requirements. The main implementation of knowledge distillation is to use the output of the teacher model as soft labels (output probability distribution) to train the student model. The student model learns by minimizing the cross-entropy or other loss functions with the output of the teacher model. The goal of the student model is to fit the predictions of the teacher model as much as possible, so as to obtain similar generalization ability and performance.

[0043] Specifically, in an optional embodiment, the teacher model uses the BERT model, and the student model uses the convolutional neural network model; in combination with the attached Figure 3 、 4It can be seen that the teacher model combines the input privileged relationship information (including but not limited to manual annotation, known relational databases) to perform relationship marking on the input text data, and uses the entity extraction component to mark the subject and object. For example, for the input "Mark Zuckerberg is the founder of Facebook", "founder" is the relationship mark, and the entity extraction component marks the subject "Mark Zuckerberg" and the object "Facebook" respectively.

[0044] The specific implementation steps of the teacher model are as follows:

[0045] Step S4.1: Data preprocessing. Use the BERT pre-trained model to obtain the text feature h, and perform sub-word alignment in the following way to obtain h avg ,

[0046] h = BERT(s);

[0047] h avg = Avgpool(h),

[0048] where Avgpool is the average pooling operation, the purpose of which is to align the embedding transformation length with the original sentence length. In this process, first, the input text data needs to be tokenized by the tokenizer to obtain the tokenized words (tokens), and then the model outputs the predicted part-of-speech for each word. After that, the tagged set {B, I, O} is also used to label the entity boundaries of the words in the sentence. B represents the start, I represents the inside, and O represents the outside. After the annotation, it is input into the model for training. By performing BIO annotation before the relation extraction task, the words starting with B and followed by I are concatenated until the position of the next B label, which is equivalent to separating a word phrase, thus completing the tokenization of the input text. Then, through encoding processing, a word sequence W = {w1, w2... w n} of the input text is obtained, and then the word sequence is vectorized to obtain the text feature h = {h1, h2... h n} of the input text. Then, [CLS] and [SEP] tags are added. Similarly, all sentences in the training set can be made of the same length, but the length of each sentence is not fixed. It is necessary to ensure that the length of all word lists is the same. Suppose the length of the sentence is maintained at 7, and the length of the input sentence after tokenization is 5. Then, padding tokens [PAD] need to be added to the sentence to make the sentence length become 7. The aggregation information of the entire complete sentence is ensured by the [CLS] tag at the beginning of the sentence. In the set h of vectorized representations of the sentence, each word vector combines three vectors: token embedding, segment embedding, and position embedding into the embedding layer as the input of the BERT model.

[0049] Step S4.2: Use a label classifier to perform multi-label binary classification on the sentence text features and relationship information, perform relationship marking according to the input privilege relationship information, and the probability P of potential relationships that may exist in the sentence can also be obtained by combining the potential relationship prediction module rel ,

[0050] P rel = σ(W r g avg + b r ),

[0051] where σ represents the sigmoid function, W r and b r respectively represent the weight and bias parameters when calculating the subset of potential relationships, reflecting the existence probability of different relationships. When P rel is greater than the set threshold, the relationship is marked as 1 and placed in the set of potential relationships R; calculate the probability of each relationship through the sigmoid function to obtain the subset of potential relationships R; for example: Suppose there are three potential relationship labels: "founder", "located in", and "cooperate", the label classifier will output the probability of the existence of each relationship, such as founder, 0.85; located in, 0.12; cooperate, 0.03, which means that according to the content of the sentence, the model believes that the "is the founder" relationship has the highest possibility of existence and is higher than the preset value, and the relationship "is the founder" is put into the relationship set R;

[0052] Step S4.3: Perform two sequence marking operations, use Softmax with a fully connected layer and the set of potential relationships R to mark the subject and object, and extract triple information: The fully connected layer traverses the relationships in the set of relationships R for each subject, calculates whether there is a relevant object, and if so, outputs the triple, and outputs the triple in combination with the following formula:

[0053]

[0054]

[0055] where and respectively represent the probability distributions that the i-th mark is the subject and object of the j-th relationship, u j is the representation of the j-th relationship in the embeddable matrix U, is the encoding representation of the i-th token after subword alignment, W sub , b sub , W obj and b objThey are the weight and bias parameters when calculating the weights of the subject and object corresponding to the relationship. If, for the Jth relationship, the output probabilities of the subject and object are greater than the preset value, then the subject and object are marked as 1. At this time, the triple {subject, relationship, object} is output for successful matching. Through the above method, the subject and object labels can be successfully extracted according to the input privilege relationship, and the triple information can be obtained.

[0056] The training of the student model specifically includes the following steps:

[0057] Step S5.1: Encode the text using a convolutional encoder; in this embodiment, we combine the GloVe embedding X g and the trainable position embedding X p to concatenate and form the vector representation X = [X g ; X p of the input statement. In this way, we can combine the word embedding with the position information to obtain a richer and more contextually informative representation of the input statement. For example: Suppose we use a pre-trained GloVe word embedding model to convert each word in the sentence into a 300-dimensional word vector representation. For each word, we calculate its position information relative to the subject "Mark Zuckerberg" and the object "Facebook". We can define the distance relative to the subject and the distance relative to the object. These position information can be represented by a small trainable embedding matrix. For example, each position feature is 50-dimensional. For each word, we concatenate its GloVe word embedding vector and the corresponding position embedding vector to form a vector representation with a total length of 350 dimensions, so as to obtain a word-level vector representation sequence.

[0058] In this way, we can combine the word embedding and the position information to obtain a richer and more contextually informative vector representation of the input statement. Such a vector representation can be used in a relationship extraction model or other natural language processing tasks, which helps to improve the performance and generalization ability of the model.

[0059] Step S5.2: Use self-attention to calculate the subject auxiliary feature and the object auxiliary feature, connect the sentence representation with the auxiliary features, and send them into the FFN to extract the subject information and the object information respectively;

[0060] This step specifically includes the following sub-steps:

[0061] S5.2.1: Combine the vector obtained from the GloVe embedding X g and the trainable position embedding X p to encode the text using a convolutional encoder with L stacked blocks to obtain the text representation: H = Block(…(Block(X))), where [;] represents the concatenation operation, and each Block contains two dilation rates of p iA dilated convolution, a gating unit, and a residual link. For each Block, padding is used to ensure that the output dimension matches the input dimension:

[0062] The final text representation is obtained as follows:

[0063] Y a = DilatedConv a (X)

[0064] Y b = DilatedConv b (X)

[0065]

[0066] where DilatedConv represents the dilated convolution module, represents element-wise multiplication, and Y i refers to the output of the i-th Block and the input of the (i + 1)-th Block. Obviously, the text representation H is equivalent to the output YL of the last Block.

[0067] S5.2.2: Two different self-attention modules are used to calculate and generate the subject auxiliary feature H h and the object auxiliary feature H t from the text representation H. The text features are connected to the corresponding auxiliary features in a bidirectional parallel manner to extract the corresponding subject information S and object information 0, which are then used to guide the extraction of object and subject information respectively;

[0068] Taking the generation of the subject auxiliary feature H h as an example for illustration:

[0069]

[0070] Q = W q ·H + b q

[0071] K = W k ·H + b k

[0072] V = W v .H + b v

[0073] where d k is the dimension of the matrix K, and W q , W k , W v and b q , b k , b vThey are the weight and bias parameters for obtaining the Q, K, and V matrices. The object auxiliary feature H t is obtained in the same way as the subject auxiliary feature H h , which will not be elaborated here.

[0074] S5.2.3: Connect the sentence with the corresponding auxiliary feature and send them into the feed-forward network respectively to extract the subject information S and the object information O, and then use the obtained subject information S to guide the extraction of the object, and use the obtained object information O to guide the extraction of the subject.

[0075] Here, two lines of S→O and O→S are carried out in parallel. Connect the sentence with the corresponding auxiliary feature and send them into the feed-forward network respectively to extract the subject information S and the object information 0, and then use the obtained subject information S to guide the extraction of the object, and use the obtained object information 0 to guide the extraction of the subject. Here, the S→O direction is taken as an example:

[0076]

[0077]

[0078]

[0079]

[0080] Among them, [;] represents the connection operation, and respectively represent the start and end of the i-th token being the subject with the j-th relationship, and respectively represent the start and end of the i-th token being the object with the j-th relationship.

[0081] Step S5.3: Use the object and entity information to form triples through relationship mapping. Two lines of S→O and O→S are carried out in parallel. Connect the sentence with the corresponding auxiliary feature and send them into the feed-forward network respectively to extract the subject information S and the object information O, and then use the obtained subject information S to guide the extraction of the object, and use the obtained object information O to guide the extraction of the subject.

[0082] Here, the teacher model extracts relationship features in advance, and the student model extracts the subject and object in parallel. For the previous framework model that extracts the subject entity and the object entity in sequence, a bidirectional entity extraction framework is adopted to avoid the failure of the entire triple extraction caused by the failure of the subject entity extraction. The two-step entity relationship extraction models are combined. In the student model, the input mapping relationship of the entity is extracted, and in the teacher model, the entity is mapped by the input of the relationship. The teacher-student model is jointly trained to extract entity relationship triples. By combining the active learning method with the entity relationship joint extraction model, the accuracy of the model is improved, and the triples can be extracted from two directions, improving the recognition ability of triples in unstructured corpora. Moreover, the system can obtain information from multiple perspectives and multiple levels. This diverse information acquisition method helps the system to more comprehensively understand the text content and enhances the ability to grasp the complex relationships between the subject, object, and relationship.

[0083] The above is an example of text relationship extraction for the teacher model and the student model. For the situation where there is less data in the vertical domain, the key is how to increase the amount of corpus so that the student model and the teacher model can perform deep learning training, thereby relying on less corpus data to improve the accuracy and generalization ability of the model in relationship extraction in this domain. The present invention effectively addresses the situation of less labeled data in the vertical domain corpus by combining the active learning method with the entity relationship joint extraction model, thereby improving the accuracy of triple extraction. It has good generalization ability for most vertical domains and can be flexibly applied in multiple fields such as cultural tourism, healthcare, and finance.

[0084] The method for training the entity relationship extraction model in the present invention includes:

[0085] Step S1: An entity relationship extraction model with a knowledge distillation framework is composed of a teacher model and a student model. The training data set is input, and the student model outputs the triple information of the text data. The teacher model and the student model input the training data simultaneously. The teacher model is only used for training to guide the student model with the learned knowledge, and the student model is used for prediction. By combining the teacher model extracting relationship features in advance and the student model extracting the subject and object in parallel, the system can obtain information from multiple perspectives and multiple levels. This diverse information acquisition method helps the system to more comprehensively understand the text content and enhances the ability to grasp the complex relationships between the subject, object, and relationship.

[0086] Step S2: Determine whether the accuracy of the output triple information is the maximum. If not, the text data corresponding to the triple information with unqualified accuracy is manually marked and fused with the initial training data set to obtain an optimized training data set until the output accuracy of the model reaches the maximum. The optimized training data set is used as the input training data set of the entity relationship extraction model, and steps S1 and S2 are repeated to train and update the teacher model and the student model.

[0087] Specifically, the iterative steps of the teacher model and the student model are further included in step S2: the text data corresponding to the triple information with unqualified accuracy is sent into the example selection algorithm and combined with the active learning method to train the teacher model on the latest optimized training data set until the test accuracy of the teacher model reaches the maximum, and knowledge distillation is used to synchronize the knowledge between the teacher model and the student model in the current iteration.

[0088] Among them, the example selection algorithm is used to select the most representative or most informative samples from the unlabeled data. The example selection algorithm includes uncertainty sampling, maximum margin sampling, etc. The selected text data is sent to manual annotation through the example selection algorithm to obtain their true labels. Active learning effectively utilizes the limited labeled resources by selecting the most informative samples for labeling, thereby improving the model performance.

[0089] The training steps for the teacher model include: dividing the training data set into a training set and a test set, continuously training the model through the training set, and testing whether the effect of the current teacher model is better than the previous teacher model through the test set. If so, the current teacher model is used as the new teacher model to continue the iterative training of the teacher model until the test accuracy reaches the maximum, and knowledge distillation is used to synchronize the knowledge between the teacher model and the student model in the current iteration.

[0090] Through the cross-training of the teacher model and the student model, the student model is optimized with the goal of minimizing the loss function. The steps include:

[0091] S3.1: Calculate the cross-entropy between the Soft-target obtained by the teacher model through softmax at temperature T and the Soft-prediction obtained by the student model through softmax at the same temperature as the first loss:

[0092] Among them, refers to the value of the softmax output of the teacher model at the i-th class under the condition that the temperature is equal to T, refers to the value of the softmax output of the student model at the i-th class under the condition that the temperature is equal to T;

[0093] S3.2: Calculate the cross-entropy between the Hard-target obtained by the student model through softmax at temperature 1 and the actual value as the second loss:

[0094] Among them, c i refers to the true label value at the i-th class, c i ∈{0, 1}, where 1 represents the positive label and 0 represents the negative label, It refers to the value of the softmax output of the student model on the i-th class when the temperature is equal to 1;

[0095] S3.3: Combine the first loss and the second loss. The total loss function: L = αL soft + βL hard , and optimize the parameters of the student model with the minimum total loss function L as the goal, where both α and β are hyperparameters. During the training process of the student model, through the new loss function for backpropagation, it attempts to gradually adjust its prediction ability by simulating the prediction results of the teacher model, improving the generalization ability of the model. Finally, the performance of the large model is achieved on the lightweight model.

[0096] Furthermore, step S1 also includes the step of expanding the training data set. The specific method includes:

[0097] S1.1, divide the input training data set into a training set and a test set;

[0098] S1.2, input the training set of the training data set into the model for training, and test the model on the test set of the training data set, and calculate the test accuracy of this training data set;

[0099] S1.3, determine whether the test accuracy of the current training data set is greater than that of the previous training data set. If so, go to step S1.4; if not, go to step S1.5;

[0100] S1.4, use the current teacher model as the new model and synchronously perform knowledge distillation of the student model, and determine whether the uncertain data is exhausted. If exhausted, stop training; if not, send the training set into the example selection algorithm and perform manual annotation, and then fuse it into the current training data set again to obtain an optimized training data set. Return the optimized training data set to step S1.1 to continue iterative training until the data is exhausted;

[0101] S1.5, determine whether the uncertain data is exhausted. If exhausted, stop training; if not, send the training set into the example selection algorithm and perform manual annotation, and then fuse it into the current training data set again to obtain an optimized training data set. Return the optimized training data set to step S1.1 to continue iterative training. The iterative training continuously feeds back to the teacher model to refine the parameter values, and continues until the performance of the teacher model and the student model reaches the maximum.

[0102] Among them, in the above steps, after the training set is input into the model, if the precision of the triple information output by the student model is lower than that of the previous model, then the text data below the threshold is mixed with the training set of the previous round and the iterative training of the teacher model is carried out again; at the same time, during the training of the teacher model, some or all of the text data below the threshold can be manually annotated (mainly in the form of relation annotation here), and the text data after manual annotation is mixed with the training set of the previous round to obtain a larger data set for iterative training; during the iterative process, the training set data is mainly sent into the teacher model for iterative training, and the teacher model tests the corresponding effects (such as accuracy rate, F1 value, etc.) through the test set after each iteration. The effect of the teacher model is evaluated by testing the F1 value output by the current teacher model, and a new student model will be synchronously updated after each iteration of the teacher model for outputting triples. Here, the iterative training of the teacher model is described in multiple cases:

[0103] Case 1: After the current teacher model after iteration is tested on the test set, if the effect of the teacher model is better than that of the previous generation, then the current version of the teacher model is used. The teacher model obtains an optimized student model through knowledge distillation. Output by the student model, continue to use the training set as the input data set to train the model. If there are still triple data with unqualified precision in the output of this version of the student model, then at this time, the methods of step S1 and step S2 need to be used to continue training the model, that is, these new unqualified triple data are continuously manually annotated and mixed into the previous training set to form a new training set for training until the student model no longer outputs unqualified triple information and the output precision of the student model reaches the maximum.

[0104] Case 2: If the current teacher model after iteration is tested on the test set and the effect of the teacher model is worse than that of the previous generation, then it is necessary to judge whether there is still unqualified triple information at this time. If not, these unqualified text data have been exhausted, then stop training and use the teacher model with the best effect of the previous version as the model to be used. If there is still unqualified triple information in the output, then re-mix these unqualified data into the training set of the previous round and continue training in combination with the method of manual annotation until the student model no longer outputs unqualified triple information.

[0105] It should be noted that the process of selecting pseudo-labels through the student model is (that is, using the text data corresponding to the triples with unqualified precision as new marked text data for training with the training set):

[0106]

[0107] Among them, Pnew is the probability score obtained by predicting the samples that need to be relabeled through the teacher-student model, P old refers to the probability score of the most original labels of these samples. γ is the boundary threshold, (d(u new ) ≤ d(u old ) - γ indicates that the classification boundary between the new non-compliant triple information and the training set in the previous round is in the low-density region. None refers to the identifier of no category. If the output probability of the text data is lower than 0.1, the labeled text data will not be used as pseudo-labels to be mixed with the training set, but mainly use the method of manual annotation to form text data with relation annotation; if the output probability score of the text data in the current version model is greater than the output probability score in the previous version model, and the classification boundary between the new non-compliant triple information and the training set in the previous round is in the low-density region, the newly labeled text data will be used as pseudo-labels to be mixed with the training set. It is also possible to manually annotate some or all of these text data below the threshold during the training of the teacher model (mainly using the method of relation annotation here), and mix the text data after manual annotation with the training set in the previous round, so as to obtain a larger dataset for iterative training; when the output probability score of the text data in the current version model is less than the output probability score in the previous version model, the text data labeled last time will be used as pseudo-labels to be mixed with the training set. It is also possible to manually annotate some or all of these text data below the threshold during the training of the teacher model (mainly using the method of relation annotation here), and mix the text data after manual annotation with the training set in the previous round, so as to obtain a larger dataset for iterative training.

[0108] Through this process, new training sets can be continuously formed, which are then continuously fed back to the teacher model to refine the parameter values until the performance of the teacher model and the student model saturates. The active learning method is used to expand the corpus, improve the model effect, and construct a high-quality vertical domain triple dataset. It has good generalization ability for most vertical domains and can be flexibly applied in multiple fields such as cultural tourism, medical care, and finance.

[0109] Verification: The present invention uses the introduction materials of tourist attractions in Inner Mongolia Autonomous Region collected from the network to construct an unannotated corpus. The initial unlabeled data includes tourist attractions in 13 leagues and cities in Inner Mongolia Autonomous Region. 12 entity types and 11 relation types are designed. The first batch of 600 sentence data is marked as the training set, and 200 sentences are used as the validation set.

[0110] The basic data information used for verification in the present invention is shown in Table 1.

[0111] Table 1 Statistical Table of Inner Mongolia Autonomous Region Tourist Attraction Dataset (IMTAD)

[0112]

[0113] The baseline models in the experiment are: CasRel, one of the new tag-based methods, which models relationships as functions that map subjects to objects. PRGC decomposes the relationship triple extraction task into three subtasks: relationship judgment, entity extraction, and subject-object alignment, greatly reducing the judgment of relationship redundancy. The evaluation index used in the experiment of the present invention is mainly the F1 value. In order to evaluate the advantages and disadvantages of different algorithms, the concept of the F1 value is proposed based on Precision and Recall to comprehensively evaluate Precision (accuracy rate) and Recall (recall rate). The definition of F1 is as follows: F1 value = accuracy rate * recall rate * 2 / (accuracy rate + recall rate).

[0114] First, use the initial training data to evaluate and compare the model proposed in this patent application on different baseline models. For the sake of clear and convenient description, the entity relationship joint extraction model proposed in this patent application is abbreviated as the ActiveLRel model in Table 2. The main experimental results are shown in Table 2. It can be seen that all the indicators of the ActiveLRel model proposed in this patent application are significantly higher than those of the other baseline models in the Inner Mongolia tourist attraction dataset.

[0115] Table 2 Training and Evaluation at the Initial Stage of Different Baselines

[0116]

[0117] In addition, after the initial stage, the F1 value of the model reaches 86.97. Predict the unlabeled corpus through the model, connect the results predicted by the model to LabelStudio for visual display, and use manual intervention to modify to ensure the correctness of the predicted labels. Reduce the workload of manual annotation through the model prediction results, and finally obtain 1,780 labeled data in the field of Inner Mongolia tourist attractions.

[0118] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, and improvements made within the spirit and principles of the present invention should be included within the protection scope of the present invention.

Claims

1. A method for training an entity relationship extraction model, characterized in that: The entity relationship extraction model training method includes: Step S1: In the student model, the input mapping relationship of the extracted entity is extracted, and in the teacher model, the entity is mapped by the input of the relationship. The teacher model and the student model are combined into an entity relationship extraction model with a knowledge distillation framework, and the training data set is input, and the student model outputs the triple information of the text data; Step S2: determine whether the accuracy of the output triple information reaches a threshold. If not, select the data that needs to be manually labeled, perform manual labeling, and fuse it with the initial training data set to obtain an optimized training data set, and continue to input the model to train the model until the accuracy reaches the threshold and the model training is completed; wherein, The teacher model is a pre-trained BERT model, and the method for the teacher model to obtain triple information includes: Step S4.1: Data preprocessing: segment the text data, convert the word sequence into a vector to obtain the text feature h of the input text, and perform subword alignment to obtain , in the vectorized representation set h of the sentence; , , where Avgpool is an average pooling operation, the purpose of which is to align the embedding transformation length with the original length of the sentence; Step S4.2: Use the label classifier to perform multi-label binary classification on the sentence text features and relation information, and obtain a subset of potential relations that may exist in the sentence based on the input privileged relation information. : ,in, represents the sigmoid function, and They represent the weight and bias parameters when calculating the potential relationship subset, reflecting the existence probability of different relationships; Step S4.3: Perform two sequence labeling operations to extract the corresponding subject and object respectively, thereby extracting the complete triple information: , ,in and denote the probability distribution that the i-th tag is the subject or object of the j-th relation, respectively. is the jth relation representation that can be embedded in the matrix U, is the encoding representation of the i-th token after subword alignment, , and are the weights and bias parameters when calculating the relationship between subject and object. The method for obtaining triple information of the student model includes: Step S5.1: Combine GloVe embedding X g and a trainable position embedding X p , using the L A convolutional encoder with stacked blocks encodes text: ; ; where [;] represents a connection operation, and each Block contains two expansion rates p i The expanded convolution, a gated unit and a residual connection are used; each block is padded to ensure that the output dimension matches the input dimension: ; ; ; in, represents the dilated convolution module, represents element-by-element multiplication, Y i Refers to i The output of the first block and the i +1) Block input, sentence representation H The output Y of the last Block L equivalence; Step S5.2: Two different self-attention modules are used to H Calculate and generate the main auxiliary feature H h and object auxiliary feature H t ; Step S5.3: The two lines S→O and O→S are carried out in parallel. The sentences are connected with the corresponding auxiliary features and sent to the feedforward network to extract the subject information S and the object information O respectively. The subject information S is then used to guide the extraction of the object, and the object information O is used to guide the extraction of the subject.

2. The entity relationship extraction model training method according to claim 1, characterized in that: It also includes the iterative steps of the teacher model and the student model: the teacher and student models are trained on the latest optimized training data set until the test accuracy reaches the maximum, and knowledge distillation is used to synchronize the knowledge between the teacher model and the student model in the current iteration.

3. The entity relationship extraction model training method according to claim 2 is characterized in that: Through cross-training between the teacher model and the student model, the student model is optimized with the goal of minimizing the loss function. The steps include: S3.1: Calculate the cross entropy of the Soft-target obtained by the teacher model at temperature T and the Soft-prediction obtained by the student model at the same temperature as the first loss: , in, Refers to the value of the softmax output of the teacher model on the i-th category under the condition that the temperature is equal to T, Refers to the value of the softmax output of the student model on the i-th category under the condition that the temperature is equal to T; S3.2: Calculate the cross entropy between the Hard-target and the actual value obtained by the student model at a temperature of 1 and softmax as the second loss: , in, refers to the true label value of the i-th category, ∈{0,1}, where 1 represents a positive label and 0 represents a negative label, Refers to the value of the softmax output of the student model on the i-th category when the temperature is equal to 1; S3.3: Combining the first loss and the second loss, the total loss function is: , with the total loss function Optimize the student model parameters for the minimum objective, where and are all hyperparameters.

4. The entity relationship extraction model training method according to claim 3 is characterized in that: Step S1 also includes the step of expanding the training data set, and the specific method includes: S1.1, divide the input training data set into training set and test set; S1.2, input the training set of the training data set into the model for training, and test the model on the test set of the training data set, and calculate the test accuracy of the training data set; S1.3, determine whether the test accuracy of the current training data set is greater than the test accuracy of the previous training data set, if so, proceed to step S1.4, if not, proceed to step S1.5; S1.4, take the current teacher model as the new model and perform knowledge distillation of the student model simultaneously, and determine whether the uncertain data is exhausted. If so, stop training; if not, send the training set to the sample selection algorithm and perform manual annotation before integrating it into the current training data set to obtain an optimized training data set. Return the optimized training data set to step S1.1 to continue iterative training until the data is exhausted; S1.5, determine whether the uncertain data is exhausted. If so, stop training. If not, send the training set to the sample selection algorithm and perform manual annotation before merging it into the current training data set to obtain an optimized training data set. Return the optimized training data set to step S1.1 to continue iterative training. Iterative training continuously feeds back to the teacher model to refine the parameter values ​​until the performance of the teacher model and the student model reaches the maximum.

5. A method for joint extraction of entity relations, characterized in that: The entity relationship joint extraction method is an entity relationship extraction model with a knowledge distillation framework obtained by combining a teacher model and a student model based on the entity relationship extraction model training method described in any one of claims 1 to 4.

6. A computer-readable storage medium having computer program instructions stored thereon, which, when executed by a processor, implement the entity relationship extraction model training method according to any one of claims 1 to 4 or the entity relationship joint extraction method according to claim 5.

Citation Information

Patent Citations

  • Incremental relation extraction method based on knowledge distillation

    CN115203404A

  • Biomedical relationship extraction method and system fusing iterative active learning

    CN116070700A

  • Relationship extraction method for electronic medical record analysis

    CN117116408A