Supply chain management field entity relationship joint extraction method based on active learning

By adopting an adversarial training method based on active learning in the field of supply chain management, combining entity and relational feature extraction of BERT and BiLSTM models, the problem of weak semantic association caused by insufficient labeling samples is solved, and efficient joint extraction of entity relationships is achieved.

CN120068869APending Publication Date: 2025-05-30GUANGZHOU INST OF RAILWAY TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510039891.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-10
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

In the field of supply chain management, the existing joint extraction method of entity relationships is insufficient in the number of labeled samples, resulting in weak semantic correlations, and the cumulative propagation of sub-task interaction errors, affecting the extraction effect.

Method used

Using an active learning-based approach, through adversarial training of GAN generator and discriminator, the pre-trained BERT model and BiLSTM are used to extract entity and relational feature, combined with the context entity label spatial attention mechanism and relational decoding attention mechanism, reduce the propagation of error information, and add labeled samples through a multi-stage sample selection strategy.

Benefits of technology

It effectively reduces the cost of manual labeling samples, solves the problem of underfitting caused by insufficient labeling samples, and improves the accuracy and efficiency of joint extraction of entity relationships.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120068869A_ABST
    Figure CN120068869A_ABST
Patent Text Reader

Abstract

The invention relates to an active learning-based supply chain management field entity relationship joint extraction method, which comprises the following steps of: pre-training a GAN generator for a training set in manual annotation data in a to-be-processed supply chain management text; processing test samples in the test set by using a pre-trained GAN generator, and training a GAN discriminator; and after the training of the GAN discriminator is completed, performing adversarial training on a pre-trained GAN generator for fine tuning, predicting unlabeled data in a to-be-processed supply chain management text of the fine-tuned GAN generator, adding the predicted unlabeled data to a training set, and repeating the training process until the number of iterations is met, thereby obtaining the trained GAN generator. According to the method, manually labeled samples are effectively reduced, and the problem that the method falls into under-fitting due to the lack of large-scale labeled training samples in the field of supply chain management is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer technology, and particularly to a method for jointly extracting entity relationships in the field of supply chain management based on active learning. Background Art

[0002] In the supply chain management text corpus, sentences contain multiple entity relationship triples. Some sentences are relatively long, entities involve multiple entity relationships, some head entities and tail entities are far apart in the sentence, and the semantic association is weak. Moreover, there is a lack of a large number of labeled samples and less entity semantic co-occurrence information. In view of the fact that the number of relationship categories involved by some entities varies, researchers have transformed the joint extraction problem into a sequence labeling problem. For example, in the prior art, entity boundaries, relationship categories, and entity position information are synchronously labeled, and then a BERT-CRF method that embeds lexical information is constructed to identify entity relationship labels, and finally entity relationship matching is designed to obtain triples. Or, entity relationship labels of entity position-entity type-relationship category-entity role are defined to solve the limitation of not being able to extract overlapping relationships. Secondly, a domain-specific pre-training PowerRoberta method for the power field is constructed to synchronously complete entity annotation and relationship extraction.

[0003] In view of the problem of weak semantic association of sentence entity relationships, researchers have proposed a method for joint extraction by sharing method parameters. In the prior art, based on the input of character features of Bidirectional Encoder Representations from Transformers (BERT), a partition filtering network is proposed to simulate the bidirectional interaction between tasks. Experimental results show that the performance of this method is better than previous methods. Also, in the prior art, a joint extraction method with position-aware attention and relationship embedding is proposed. This method first identifies entities, then enhances sentence representation through entity- and relationship-guided attention mechanisms, and uses position-aware attention to extract the subject and object of the sentence. Experimental results show that the problem of extracting overlapping triples can be solved.

[0004] However, the above-mentioned methods for jointly extracting entity relationships mainly rely on a certain scale of labeled samples for training and ignore unlabeled samples. In fact, in a specific domain, a large number of samples lack labels.

[0005] In the face of insufficient labeled samples, some research works aim to maximize the extraction effect of the method with the least labeled samples by using unlabeled samples to participate in training or remote supervision. For example, a unified label space is proposed, and entity relation joint extraction is realized through matrix annotation. On this basis, a sample sampling strategy based on entropy-based active learning is adopted, which can achieve the optimal method performance with the least training data. Or, combined with the sampling method of active learning to reduce the text annotation cost, data augmentation strategies such as word replacement and generation are proposed to solve problems such as unbalanced relationship classification between samples and insufficient annotation materials.

[0006] In the case of a limited number of labeled samples, due to the weak semantic association and the varying number of entity relationships in sentences, the existing supervised learning methods for entity relation joint extraction mainly achieve joint extraction by studying annotation strategies or sharing method parameters. In these methods, the error accumulation and propagation may occur in the interaction mode of sub-tasks, making it difficult for the method to effectively obtain useful information between tasks during the joint extraction process, thus affecting the extraction effect. In addition, there is a lack of large-scale labeled samples in the corpus of the supply chain management field. The existing active learning methods need to pre-label all samples or manually label the selected samples, and the manual annotation cost has not been reduced. Although samples can be automatically annotated by generating pseudo-labels, there may be noisy texts with inference errors, which will also affect the entity relation joint extraction effect.

[0007] In view of this, to solve the problem of the underfitting effect caused by weak semantic association samples in the entity relation joint extraction method in the supply chain management field, the present invention provides an entity relation joint extraction method based on active learning. Summary of the Invention

[0008] (1) Technical Problems to be Solved

[0009] In view of the above-mentioned disadvantages and deficiencies of the prior art, the present invention provides an entity relation joint extraction method in the supply chain management field based on active learning, which solves the problem of the underfitting effect caused by weak semantic association samples in the entity relation joint extraction method in the supply chain management field.

[0010] (2) Technical Solutions

[0011] To achieve the above object, the main technical solutions adopted by the present invention include:

[0012] In the first aspect, an embodiment of the present invention provides an entity relation joint extraction method in the supply chain management field based on active learning, including:

[0013] A100. Pre-train the GAN generator for the training set in the manually annotated data of the supply chain management text to be processed;

[0014] A200. Process the test samples in the test set using a pre-trained GAN generator, and output the relationship label prediction results and entity label prediction results of the test samples;

[0015] A300. Obtain the character mixed features of the test samples according to the entity label prediction results of the test samples and the intermediate character related variables for obtaining the entity label prediction results;

[0016] A400. Train the GAN discriminator based on the character mixed features, relationship label prediction results and entity label prediction results of the test samples, and the manually annotated relationship labels and entity label results;

[0017] A500. After the training of the GAN discriminator is completed, perform adversarial training on the pre-trained GAN generator for fine-tuning, and use the fine-tuned GAN generator to predict the unlabeled sample entities in the supply chain management text to be processed. Add the selected samples to the training set, and repeat the above training process until the iteration times are met to obtain the trained GAN generator.

[0018] Optionally, the A100 includes:

[0019] A101. Use the BERT pre-training method to obtain the character feature vector sequence x=(x 1 ,x 2 ,x i ...,x n ) of the sentence samples; n is the number of characters in the sentence length, and x i represents the character feature vector of the i-th character;

[0020] A102. Use stacked BiLSTM to extract the context features of the character feature vector sequence and output the text deep context feature sequence L is the stacking layer number, represents the text deep context feature of the i-th character;

[0021] A103. Use a Conditional Random Field (CRF) to identify the sentence entities in the text deep context feature sequence to obtain the entity labels; each identified entity label has a corresponding entity label embedding e; the entity label embedding sequence of all sentence entities is E=(e 1 ,e 2 ,...,e n ). Input E=(e 1 ,e 2 ,...,e n ) into BiLSTM to obtain the context features with the entity label embedding sequence

[0022] A104. Incorporate into through the multi-head self-attention mechanism to , obtaining a new context feature sequence Z; and performing relationship label prediction on the new context feature sequence Z, outputting the relationship label prediction result, and using the relationship label prediction result and the entity label as the entity label prediction and relationship label prediction results of the first stage;

[0023] A105. Calculate the character-weighted relationship feature for the relationship label prediction result of the first stage through the relationship decoding attention mechanism, obtaining a weighted character relationship feature sequence;

[0024] Add the weighted character relationship feature sequence to and input it into the CRF to obtain the entity label, and obtain the relationship label prediction result through the multi-head attention mechanism, obtaining the entity label prediction and relationship label prediction results output in the second stage, and using the result output in the second stage as the output result of the GAN generator.

[0025] Optionally, the A104 includes:

[0026] A1041. The text deep context feature sequence H is transformed into the query (Q) and key (K) matrix spaces through the linear matrix W i H , and the entity label embedding context feature sequence H E is transformed into the value (V) matrix space through the linear matrix . The attention function calculates the correlation parameters between the text deep context feature sequences themselves, and assigns attention weights to the entity label vector sequence according to the correlation parameters, obtaining a weighted entity label vector sequence;

[0027] The weighted entity label vector sequence and the text deep context feature sequence are added to obtain the text deep context feature sequence integrating entity label information, that is, the new context feature sequence Z;

[0028] The score i calculation process is repeated N times, and the results of the text deep context feature sequences integrating entity label information for N times are concatenated, and z = concat(score 1 , score 2 ,..., score N ) is output, where concat is the concatenation operation;

[0029]

[0030] where i is the calculation times;

[0031] A1042. Predict the relationship of the text deep context feature Z that integrates entity label information for the character c i and the character c j The corresponding text deep context feature is denoted as and

[0032] The character c i and the character c j The relationship probability calculation formula between them is as follows:

[0033]

[0034] Among them, p(c i , r, c j ) represents the predicted probability of each relationship r between the character c i and the character c j . is the scoring function for calculating the relationship r between the character c i and the character c j , and σ is the sigmoid function.

[0035] Optionally, A105 includes: A1051. The entity relationship mainly predicts the relationship category probability between entity pairs. The entity involves multiple fonts, and the relationship annotation is mainly marked on the last character of the head and tail entities. Therefore, the predicted relationship is converted into predicting the relationship category probability between character pairs. If there is a relationship between entity pairs, the entity relationship is connected to the last character of the head and tail entities. If there is no relationship, it is predicted as N. Each relationship has a corresponding character combination pair relationship category probability matrix.

[0036] According to the positions of the characters in the head and tail entities, the character pair relationship category probability matrix is divided into a forward character pair relationship category probability matrix and a backward character pair relationship category probability matrix, which are used as the input of the relationship decoding attention mechanism.

[0037] A1052. For the forward character pair relationship category probability matrix, the relationship decoding attention mechanism first calculates the scoring coefficient between the text deep context feature sequence that integrates entity label information and the forward character pair relationship category probability matrix. The scoring coefficient calculation formula is as follows:

[0038] Among them, is the tail entity character probability vector for the i-th character predicted as the head entity pair relationship category r, and w a , v, and w b are training parameters.

[0039] A1053. After obtaining , calculate the weight The calculation formula is as follows:

[0040] Among them, r is the number of relationship categories. Multiply the weights with the corresponding and sum them up to obtain the weighted probability matrix of the forward character pair for relationship category r;

[0041] A1054. Obtain the weighted probability matrix of the backward character pair for relationship category r, add up the weighted probability matrices corresponding to all relationship categories, and the two matrices update and fuse the deep context features of the text with entity label information through matrix mapping. The update formula is as follows:

[0042] Among them, and are the weighted probability matrices of the forward character pair relationship category and the backward character pair relationship category respectively, and w bb and w cc are training parameters.

[0043] Optionally, the A200 includes:

[0044] For each test sample, generate the position feature of the predicted entity in the sentence, that is, the entity label prediction result. The calculation formula is as follows: Among them, N represents the number of entity predictions in the sentence, p i,j represents the number of characters separated between the i-th character and the j-th entity character. The position encoding of all characters predicted as entities in the sentence is 0, the position encoding of the characters on the left of the predicted entity is negative, and the position encoding of the characters on the right is positive. Calculate the average distance of each character from all predicted entities in the sentence to obtain the character position feature;

[0045] The A300 includes:

[0046] A301. For each test sample, splice the entity label prediction result of the test sample, the character feature vectors in the character feature vector sequence, and the entity label embeddings in the entity label embedding sequence to form character mixed features;

[0047] The character feature vectors in the character feature vector sequence and the entity label embeddings in the entity label embedding sequence are intermediate character-related variables during the process of using the pre-trained GAN generator to process the test sample.

[0048] Optionally, the A400 includes: A401. For each test sample, input the character mixed features of each test sample into the hybrid neural network to obtain the sentence features output by the hybrid neural network;

[0049] A402. Calculate the support degree of sentence features for each type of entity relationship label based on a multi-stage sample selection strategy. The support degree is used as the weight of each test sample. Sort the test samples according to the support degree weight, and filter out a preset number of samples based on the sorting as the sample selection in the first stage.

[0050] A403. For each test sample in the sample selection of the first stage, use the characters in the test sample as nodes, and based on the relationship label prediction result of the test sample, perform graph convolution operations on the character mixed features of the test sample using a graph convolutional neural network to output the convolved character features. Also, add the convolved character features and the character mixed features to obtain a new character feature sequence.

[0051] A404. Based on the new character feature sequence, use CNN and max pooling to calculate the first prominent feature of the test sample, and use this first prominent feature as the latent variable of the test sample.

[0052] Based on the manually labeled entity label result and entity relationship label result of the test sample, repeat the processes of A300 and A400 to generate the second prominent feature corresponding to the test sample, and use this second prominent feature as the true variable of the test sample. The latent variable and the true variable are input into the Softmax function for training, and after training, they are used for the second-stage selection of unlabeled samples.

[0053] The purpose of training with the Softmax function is to perform the second-stage selection among unlabeled samples with many entity relationships after training to select as many correct unlabeled samples as possible. Step A402 selects test samples during the training of the GAN discriminator and unlabeled samples after training, mainly selecting sentence samples with many entity relationships.

[0054] A500 includes: After the adversarial training is completed, the GAN generator predicts unlabeled samples, and the GAN discriminator sorts and selects the unlabeled sample set. The GAN discriminator selects samples with many entity relationship triples in the first stage according to the support degree of sentence features for relationship r, and then selects correct samples through the Softmax function in the second stage.

[0055] Regarding the entity label result and relationship label result predicted by the GAN generator corresponding to the selected unlabeled samples as the annotation information of the unlabeled samples, and adding them to the manually labeled training set. Repeat the training processes of the training set and the test set for the GAN generator and the GAN discriminator until the number of iterations is satisfied.

[0056] Optionally, A401 includes: The hybrid neural network includes: BiLSTM and a segmented convolutional neural network; A4011. Capture the context information of the character hybrid features through BiLSTM to obtain the character hybrid context feature sequence h = {h 1 , h 2 ,..., h n};

[0057] A4012. Split the character hybrid context features in the character hybrid context feature sequence into character context features, entity prediction label embedding context features, and character position context features, and perform segmented convolution operations on each segment of feature vectors respectively. The convolution formula for each segment is as follows:

[0058] c i = f(F·h i:i+l-1 + b); where l is the height of the convolution kernel, i is the i-th convolution window, F is the convolution kernel, f represents the RELU non-linear function, and the segmented feature matrix output by the segmented convolutional neural network at the end is C j = (c 1,j , c 2,j ,..., c m,j ), where m is the number of convolution kernels and j is the j-th feature segment;

[0059] A4013. Use segmented max pooling to extract segmented prominent features from the segmented feature matrix;

[0060] Perform max pooling operations on each feature segment vector respectively. The max pooling operation is:

[0061] p ij = max(c ij ), 1 ≤ i ≤ m, 1 ≤ j ≤ 3

[0062] Concatenate the three max pooling vectors together to obtain the vector p i = {p i1 , p i2 , p i3}; Connect all the p i And perform non-linear function operations to obtain the sentence feature vector output by the max pooling layer as s = tanh(p 1:m ).

[0063] Optionally, A402 includes: Calculate the support degree of the sentence features for the relationship r through the Sigmoid function, and screen the samples containing more triples based on the support degree;

[0064] A403 includes: The graph convolution operation introduces relationship prediction information into the character node features. The graph convolution calculation formula is as follows:

[0065]

[0066] Among them, is the node feature of node u in the l-th layer graph convolutional neural network, where V and R are the total number of characters and the total number of relationship categories respectively.

[0067] Update the character node features according to the graph convolution iterative features, and the update formula is as follows:

[0068] Among them, h u represents the mixed feature of character u. Finally, set the updated feature sequence as H'=(h 1 ', h' 2 ,..., h' n ).

[0069] Optionally, A404 includes:

[0070] Perform convolution operation on the updated character node feature sequence using convolution kernels of different sizes, splice the l-gram features of different convolution kernel sizes, perform max pooling operation on the spliced l-gram features to obtain the first prominent feature, and define this first prominent feature as the latent variable of the test sample;

[0071] Replace the manually labeled entities and entity relationship labels of the test sample with the predicted entities and predicted entity relationship labels to obtain the second prominent feature corresponding to the test sample as the true variable of the test sample;

[0072] Fix the generator parameter variables, input the latent variable and the true variable into the Softmax function for classification training, and the loss function is as follows:

[0073] Among them, v P and v R are the latent variable and the true variable of the sample respectively, and are the predicted probabilities of the sample variables, D represents the discriminator, and E represents the mathematical expectation; the discriminator training is used to distinguish the feature distributions of the predicted labels and the true labels.

[0074] In a second aspect, an embodiment of the present invention further provides an electronic device, including a memory and a processor, where the memory stores a computer program, and the processor executes the computer program in the memory and performs the steps of any of the methods for jointly extracting entity relationships in the field of supply chain management based on active learning described in the first aspect above.

[0075] (III) Beneficial effects

[0076] In the method of the present invention for the generator of the GAN, aiming at the problem of cumulative propagation of sub-task interaction errors that may exist in the existing entity relationship joint extraction method, the present invention proposes a context entity label space attention mechanism and a relationship decoding attention mechanism. Through the self-training of the generator, it selectively fuses entity and entity relationship features, reduces the propagation of error information, and improves the knowledge extraction effect. Aiming at the problem of limited number of labeled management corpus, a multi-stage sample selection strategy is proposed in the discriminator of the GAN. Through a two-stage sample selection process, samples containing more triples can be selected as much as possible. After the discriminator of the GAN is trained, adversarial training is introduced to train the generator, so that the generator predicts as correct entity and relationship labels as possible for unlabeled samples. Through the screening of the discriminator, the selected samples are automatically labeled with predicted labels and added to the labeled training sample set for the next round of iterative training of the generator. This method reduces the cost of manually labeled samples and solves the problem that the method falls into underfitting due to the lack of large-scale labeled training samples in the supply chain management field. BRIEF DESCRIPTION OF THE DRAWINGS

[0077] Figure 1 It is a schematic structural diagram of the generator in the embodiment of the present invention;

[0078] Figure 2 It is a schematic diagram of the multi-stage sample selection strategy method in the embodiment of the present invention;

[0079] Figure 3 It is a partial flow chart of the GA-BASBN method in the embodiment of the present invention;

[0080] Figure 4 It is a structural diagram of the fusion method of the entity label embedding context feature sequence and the text deep context feature sequence in the embodiment of the present invention;

[0081] Figure 5 It is an example diagram of entity relationship annotation shown in the embodiment of the present invention;

[0082] Figure 6 It is the overall flow chart of the GA-BASBN method shown in the embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0083] In order to better explain the present invention for easy understanding, the exemplary embodiments of the present invention will be described in more detail below with reference to the accompanying drawings. Although the exemplary embodiments of the present invention are shown in the drawings, it should be understood that the present invention can be implemented in various forms and should not be limited by the embodiments set forth herein. On the contrary, these embodiments are provided so that the present invention can be understood more clearly and thoroughly, and the scope of the present invention can be completely conveyed to those skilled in the art.

[0084] Embodiment 1

[0085] This embodiment provides a method for jointly extracting entity relationships in the field of supply chain management based on active learning. The method may include:

[0086] A100. Pre-train the GAN generator for the training set in the manually annotated data of the supply chain management text to be processed;

[0087] A200. Process the test samples in the test set using the pre-trained GAN generator, and output the relationship label prediction results and entity label prediction results of the test samples.

[0088] For example, for a test sample, generate the position feature of the predicted entity in the sentence, that is, the entity label prediction result. The calculation formula is as follows:

[0089] where N represents the number of entity predictions in the sentence, p i,j represents the number of characters between the i-th character and the j-th entity. The position encoding of all characters of the predicted entity in the sentence is 0, the position encoding of the characters on the left of the predicted entity is negative, and the position encoding of the characters on the right is positive. Calculate the average distance of each character from all predicted entities in the sentence to obtain the character position feature.

[0090] A300. Obtain the character mixed feature of the test sample according to the entity label prediction result of the test sample and the intermediate character-related variables for obtaining the entity label prediction result.

[0091] For example, for a test sample, splice the entity label prediction result of the test sample, the character feature vector in the character feature vector sequence, and the entity label embedding in the entity label embedding sequence to form a character mixed feature;

[0092] The character feature vector in the character feature vector sequence and the entity label embedding in the entity label embedding sequence are intermediate character-related variables in the process of using the pre-trained GAN generator to process the test sample.

[0093] A400. Train the GAN discriminator based on the character mixed feature, relationship label prediction result, entity label prediction result, manually annotated relationship label, and entity label result of the test sample;

[0094] A500. After the training of the GAN discriminator is completed, perform adversarial training on the pre-trained GAN generator for fine-tuning, and use the fine-tuned GAN generator to predict the entities of the unlabeled samples in the supply chain management text to be processed. Then, use the GAN discriminator to perform sorting and selection, add the selected samples to the training set, and repeat the above training process until the iteration times are met to obtain the trained GAN generator.

[0095] In this embodiment, after the adversarial training is completed, the GAN generator predicts unlabeled samples, and sorts and selects the unlabeled sample set;

[0096] For the entity label result and the relationship label result predicted by the GAN generator corresponding to the selected unlabeled samples, they are used as the annotation information of the unlabeled samples and added to the manually labeled training set; repeat the training process of the GAN generator and the GAN discriminator for the training set and the test set until the number of iterations is satisfied.

[0097] The method of this embodiment can effectively reduce the cost of manually annotating samples and solve the problem that the method falls into underfitting due to the lack of a large number of labeled training samples in the field of supply chain management.

[0098] To better understand the above steps, the following will be elaborated in combination with each sub-step.

[0099] In a possible implementation manner, step A100 of this embodiment may include the following sub-steps:

[0100] A101. Obtain the character feature vector sequence x = (x 1 , x 2 , x i ..., x n ) of the sentence sample by using the BERT pre-training method; n is the number of characters in the sentence length, and x i represents the character feature vector of the i-th character;

[0101] A102. Use a stacked BiLSTM to extract the context features of the character feature vector sequence and output the deep context feature sequence of the text L is the number of stacked layers, represents the deep context feature of the i-th character of the text;

[0102] A103. Use a conditional random field (CRF) to identify the sentence entities in the deep context feature sequence of the text to obtain entity labels; each identified entity label has a corresponding entity label embedding e; the entity label embedding sequence of all sentence entities is E = (e 1 , e 2 ,..., e n ), input E = (e 1 , e 2 ,..., e n ) into the BiLSTM to obtain the context features with the entity label embedding sequence

[0103] A104. Through the multi-head self-attention mechanism, is fused into In this process, a new context feature sequence Z is obtained; and the new context feature sequence Z is used for relationship label prediction, and the relationship label prediction result is output. The relationship label prediction result and the entity label are used as the entity label prediction and relationship label prediction results in the first stage;

[0104] A105. Calculate the character-weighted relationship feature for the relationship label prediction result in the first stage through the relationship decoding attention mechanism to obtain the weighted character relationship feature sequence;

[0105] The weighted character relationship feature sequence is added to and input into the CRF to obtain the entity label, and the relationship label prediction result is obtained through the multi-head attention mechanism, and the entity label prediction and relationship label prediction results output in the second stage are used as the output results of the GAN generator.

[0106] For example, sub-step A104 may include:

[0107] A1041. The text deep context feature sequence H is transformed into the query (Q) and key (K) matrix spaces through the linear matrix W i H The entity label embedding context feature sequence H E is transformed into the value (V) matrix space through the linear matrix The attention function calculates the correlation parameters between the text deep context feature sequences themselves, and the attention weight distribution for the entity label vector sequence is realized according to the correlation parameters to obtain the weighted entity label vector sequence;

[0108] The weighted entity label vector sequence and the text deep context feature sequence are added to obtain the text deep context feature sequence integrating entity label information, that is, the new context feature sequence Z;

[0109] The score i calculation process is repeated N times, and the results of the text deep context feature sequences integrating entity label information for N times are concatenated, and z = concat(score 1 , score 2 ,..., score N ) is output, where concat is the concatenation operation;

[0110]

[0111] where i is the calculation times;

[0112] A1042. Perform relationship prediction on the text deep context feature Z integrating entity label information, and the character ci The text deep context feature corresponding to the character cj is denoted as and

[0113] character c i The formula for calculating the relationship probability between the character c and the character cj is as follows:

[0114]

[0115] where p(c i , r, c j ) represents the predicted probability of each relationship r between the character c i and the character c j , and σ is the sigmoid function. is the scoring function for calculating the relationship r between the character c i and the character c j .

[0116] In the second possible implementation, the sub-step A105 may include:

[0117] A1051. The entity relationship mainly predicts the relationship category probability between entity pairs. The entity involves multiple fonts, and the relationship annotation is mainly marked on the last character of the head and tail entities. Therefore, predicting the relationship is transformed into predicting the relationship category probability between character pairs. If there is a relationship between entity pairs, the entity relationship is connected to the last character of the head and tail entities. If there is no relationship, it is predicted as N. Each relationship has a corresponding character combination pair relationship category probability matrix.

[0118] According to the positions of the characters in the head and tail entities, the character pair relationship category probability matrix is divided into a forward character pair relationship category probability matrix and a backward character pair relationship category probability matrix, which are used as the input of the relationship decoding attention mechanism.

[0119] A1052. For the forward character pair relationship category probability matrix, the relationship decoding attention mechanism first calculates the scoring coefficient between the text deep context feature sequence that fuses entity label information and the forward character pair relationship category probability matrix. The formula for calculating the scoring coefficient is as follows:

[0120] where is the probability vector of the tail entity character predicted for the relationship category r of the head entity pair by the i-th character, and w a , v, and w b are training parameters.

[0121] A1053. When is obtained, calculate the weight The calculation formula is as follows: where r is the number of relationship categories, and the weight Multiply with the corresponding and sum them up to obtain the weighted probability matrix of the forward character pair relationship category r;

[0122] A1054. Obtain the weighted probability matrix of the backward character pair relationship category r, add up the weighted probability matrices corresponding to all relationship categories, and update and fuse the deep context features of the text with entity label information through matrix mapping for the two matrices. The update formula is as follows:

[0123] where, and are the weighted probability matrices of the forward character pair relationship category and the backward character pair relationship category respectively, and w bb and w cc are training parameters.

[0124] In the fourth possible implementation manner, step A400 of this embodiment may include the following sub-steps:

[0125] A401. For each test sample, input the character mixed feature of each test sample into the hybrid neural network to obtain the sentence feature output by the hybrid neural network;

[0126] A402. Calculate the support degree of the sentence feature for each type of entity relationship label based on the multi-stage sample selection strategy. The support degree is used as the weight of each test sample. Sort the test samples according to the support degree weight, and filter out a preset number of samples based on the sorting as the sample selection in the first stage.

[0127] For example, the support degree of the sentence feature for relationship r is calculated through the Sigmoid function, and samples containing more triples are filtered based on the support degree.

[0128] A403. Based on each test sample in the sample selection in the first stage, use the characters in the test sample as nodes, and based on the relationship label prediction result of the test sample, perform graph convolution operation on the character mixed feature of the test sample using the graph convolutional neural network, output the convolved character feature, and add the convolved character feature and the character mixed feature to obtain a new character feature sequence.

[0129] For example, the graph convolution operation introduces relationship prediction information into the character node feature. The graph convolution calculation formula is as follows:

[0130]

[0131] where, is the node feature of node u of the l-th layer graph convolutional neural network, and V and R are the total number of characters and the total number of relationship categories respectively.

[0132] Update the character node features according to the graph convolution iterative features, and the update formula is as follows:

[0133] Among them, h u represents the mixed feature of character u. Finally, set the updated feature sequence as H'=(h 1 ', h' 2 ,..., h' n ).

[0134] A404. Based on the new character feature sequence, use CNN and max pooling to calculate the first prominent feature of the test sample, and use this first prominent feature as the latent variable of the test sample;

[0135] Based on the manually labeled entity label results and entity relationship label results of the test sample, repeat the processes of A300 and A400 to generate the second prominent feature corresponding to the test sample, and use this second prominent feature as the true variable of the test sample; The latent variable and the true variable are input into the Softmax function for training, and after training, they are used for the second-stage selection of unlabeled samples. It can be understood that the purpose of training the softmax function is to perform the second-stage selection among unlabeled samples with many entity relationships after training, and select as many correct unlabeled samples as possible. Step A402 is to select test samples when the GAN discriminator is trained, and select unlabeled samples after training, mainly to select sentence samples with many entity relationships.

[0136] For example, in this embodiment, convolution operation can be performed on the updated character node feature sequence using convolution kernels of different sizes, the l-gram features of different convolution kernel sizes are concatenated, and max pooling operation is performed on the concatenated l-gram features to obtain the first prominent feature, which is defined as the latent variable of the test sample;

[0137] Replace the predicted entity and predicted entity relationship labels with the manually labeled entities and entity relationships of the test sample to obtain the second prominent feature corresponding to the test sample as the true variable of the test sample;

[0138] Fix the generator parameter variables, input the latent variable and the true variable into the Softmax function for classification training, and the loss function is as follows:

[0139] Among them, v P and v R are the latent variable and the true variable of the sample respectively, and are the predicted probabilities of the sample variables, D represents the discriminator, and E represents the mathematical expectation; The discriminator training is used to distinguish the feature distributions of the predicted labels and the true labels.

[0140] For example, sub-step A401 may include:

[0141] The hybrid neural network includes: BiLSTM and a segmented convolutional neural network;

[0142] A4011. Capture the context information of the character hybrid features through BiLSTM to obtain the character hybrid context feature sequence h = {h 1 , h 2 ,..., h n};

[0143] A4012. Segment the character hybrid context features in the character hybrid context feature sequence into character context features, entity prediction label embedding context features, and character position context features, and perform segmented convolution operations on each segment of feature vectors respectively.

[0144] The convolution formula for each segment is as follows: c i = f(F·h i:i+l-1 + b)

[0145] where l is the height of the convolution kernel, i is the i-th convolution window, F is the convolution kernel, f represents the RELU non-linear function, and the segmented feature matrix output by the segmented convolutional neural network is C j = (c 1,j , c 2,j ,..., c m,j ), where m is the number of convolution kernels and j is the j-th feature segment.

[0146] The segmentation in this step is mainly performed according to the character position feature, entity label embedding feature, and dimension of the character feature. The previous step is spliced according to the three feature positions and dimensions, and the segmentation is also performed according to the position and dimension.

[0147] A4013. Use segmented max pooling to extract segmented prominent features from the segmented feature matrix;

[0148] Perform max pooling operations on each feature segment vector respectively. The max pooling operation is:

[0149] p ij = max(c ij ), 1 ≤ i ≤ m, 1 ≤ j ≤ 3;

[0150] Concatenate the three max pooling vectors together to obtain the vector p i = {p i1 , p i2 , p i3}; Connect all p iAnd perform nonlinear function operation to obtain the sentence feature vector output by the maximum pooling layer as s = tanh (p 1:m ).

[0151] Embodiment 2

[0152] The solution of this embodiment aims to solve the problems that the number of corpus annotations is limited, the number of entity relationships contained in sentences is different, some sentences are long, entities involve multiple entity relationships, some head entities and tail entities are far apart in sentences, and the semantic association is weak. A method for joint extraction of entity relationships using active learning of a generative adversarial network (GAN) is proposed. In the GAN generator, the generator structure is as follows: Figure 1 shown.

[0153] The GAN generator method uses stacked BiLSTM to obtain deep context features of the text, which is very important for understanding the semantic structure of samples with weak semantic associations. On this basis, in order to solve the problem of error accumulation and propagation of interaction between subtasks (i.e., entity extraction and relationship collection) that may exist in the existing entity-relationship joint extraction method, a context entity label space attention mechanism and a relationship decoding attention mechanism are proposed to fuse entity label information and entity relationship information between subtasks in a weighted distribution manner, thereby improving the mutual connection between the two subtasks and reducing error accumulation and propagation. In the GAN discriminator, a multi-stage sample selection strategy is proposed. The structure diagram of this strategy method is shown in the figure below. Figure 2 As shown in the figure. First, by calculating the support of the sample for the relationship, the samples with a large number of entity relations are screened out. Then, the sample prediction variables and true variables are calculated in the samples with a large number of entity relations. The predicted labeled samples with semantic similarity to the labeled samples and the ability to retain positive instances to the maximum extent are selected through adversarial training, and added to the training samples. While reducing the cost of manually labeled samples, the problem of underfitting of entity relationship joint extraction methods due to the lack of large-scale entity relationship labeled training samples in specific fields is solved by increasing the number of training samples.

[0154] The overall structure of the active learning logic knowledge entity relationship joint extraction method of the generative adversarial network of this embodiment (Binocular-attention-based Stacked BiLSTM with two-stage GAN, GA-BASBN) is as follows: Figure 1 As shown in the flow chart of the method Figure 3As shown in the figure. The method of this embodiment designs an entity relationship joint extraction method (Binocular-attention-based Stacked BiLSTM Network, BASBN) through a GAN generator, and increases the mutual connection between the two subtasks of entity recognition and relationship extraction by means of attention weights; a multi-stage sample selection strategy is proposed in the GAN discriminator, mainly selecting unlabeled samples for automatic annotation and adding them to the labeled dataset for the training of the GAN generator method, reducing the cost of manually labeled samples and solving the problem of the lack of a large-scale labeled dataset in the field.

[0155] To achieve the above object, the technical solution of the following steps is adopted in this embodiment:

[0156] A generative adversarial network-based active learning entity relationship joint extraction method, the goal of which is to predict the entity labels in a sentence, and then predict the relationship labels for the relationship between two characters. When there is an entity relationship between entities, the tail character pair corresponding to the head entity and the tail entity is predicted as the entity relationship label, and the entity and the entity relationship form an entity relationship triple for output; at the same time, during the training process of the joint extraction method, through active learning, the joint extraction method automatically labels samples, expands the manually labeled sample training set, and reduces the human and time costs of manually labeled data. Its main steps include:

[0157] Before the computing device executes, the supply chain management text is preprocessed first. Mainly, the corpus is segmented into sentences, and sentences are used as samples. Some sentences are selected for entity label and relationship label marking to form an entity relationship joint extraction manually labeled dataset. The manually labeled dataset of this embodiment includes a training set and a test set. The supply chain management text includes manually labeled samples and unlabeled samples.

[0158] Preprocess the supply chain management text. Mainly, segment the corpus into sentences, use sentences as samples, select some sentences for entity label and relationship label marking, and the marking examples are as follows Figure 5 As shown, to form an entity relationship joint extraction manually labeled dataset.

[0159] For a sequence of sentences, when there is a relationship between entity pairs, the relationship category is labeled on the tail character of the head entity, and then the position of the tail character of the tail entity is marked. The same applies to other relationship categories. When the same entity involves multiple relationship category labels, the corresponding relationship labels and the positions of the tail characters of the tail entities are marked separately according to the above method. Entity categories include Industry, Company, Field, Problem, Trigger_word, Object, Method, Model, Index, Index_system, Attribute, Attribute_value, etc. Entity relationship categories include Belong_to, composed_of, contrapose, used_for, has_attribute, appear, event, followed, and lead_to.

[0160] Step 1: In the adversarial network generator part, for a sentence sample in the training set, the pre-determined composition of the sentence character vector includes three parts: the character embedding vector, the sentence embedding vector, and the character position embedding vector. The sum of the three embedding vectors is used as the input of the BERT pre-training method. Through BERT fine-tuning training, a sequence of character feature vectors x = (x 1 , x 2 ,..., x n ) is obtained.

[0161] In this embodiment, the input of BERT is three parts. The character embedding vector is the vocabulary vector provided by BERT. The character position embedding vector is calculated by sine and cosine functions. The sentence embedding vectors are all 0 (single sentence). BERT fine-tunes and trains these three parts of features and outputs a sequence of character feature vectors.

[0162] Step 2: Input the sequence of character feature vectors into a stacked BiLSTM for context feature extraction, and output a sequence of deep text context features

[0163] For a sequence of sentences, the composition of the character vector includes three parts: the character embedding vector, the sentence embedding vector, and the character position embedding vector. The sum of the three embedding vectors is used as the input of the BERT method. Through BERT fine-tuning training, a sequence of character feature vectors is obtained and represented as x = (x 1 , x 2 ,..., x n )).

[0164] Specifically, the stacked BiLSTM unit consists of a forward LSTM unit and a backward LSTM unit. Given a character feature sequence x, the output calculation formula of the hidden layer of each layer of BiLSTM is as follows, where L is the number of stacked layers.

[0165]

[0166] Step 3: Use a Conditional Random Field (CRF) to identify sentence entities based on the deep context features of the text, and obtain the entity labels of the sentence entities.

[0167] Each identified entity label has a corresponding entity label embedding e, and the entity label embedding sequence E=(e 1 ,e 2 ,...,e n ) is input into the BiLSTM to obtain the context features of the entity label embedding sequence.

[0168] The BiLSTM here belongs to a new BiLSTM, which is different from the stacked BiLSTM in Step 2.

[0169] Specifically, this step can be elaborated as follows:

[0170] Sub-step 31: Use a Conditional Random Field (CRF) to train the influence between event argument entity labels. Given a deep context feature sequence V, let the output label sequence be Y, and the joint probability formula of the label sequence and the feature sequence is as follows, where

[0171] represents the character label transition matrix. represents the entity label probability predicted by character i.

[0172]

[0173] Sub-step 32: In the CRF, the maximum likelihood estimation function is used for the loss function, and the calculation formula is as follows, where f(V) represents the set of all possible predicted label sequences. Y' represents the label sequence that the method may predict. During the entity category prediction process, the viterbi algorithm is used to predict the optimal label sequence.

[0174] Sub-step 33: Define an entity label embedding vector matrix and update the parameters during the training process. Given the predicted label sequence Y' in sub-step 32), each entity predicted label in the sequence has a corresponding entity label vector. In this way, the predicted label sequence generates an entity label embedding vector sequence E=(e 1 , e 2 ,..., e n ). Use BiLSTM to obtain the context features of the entity label embedding vector sequence The calculation formula for the context features of the entity label embedding vector is as follows:

[0175] Step 4: Incorporate the entity label embedding context feature sequence into the text deep context feature sequence through the multi-head self-attention mechanism, and the incorporation method is as Figure 4 shown

[0176] Then, perform relation label prediction on the text deep context feature sequence Z incorporating entity label information, and output the relation label prediction result. The relation label prediction result in this embodiment is presented in the form of a relation category probability matrix, which mainly indicates whether there is an entity relation in the character pair relation category probability matrix prediction

[0177] Thus, the entity label prediction result and the relation label prediction result in the first stage are output through the GAN generator. The entity label prediction result here is obtained in step 3

[0178] In step 4, in order to reduce the training parameters of the traditional multi-head self-attention mechanism, this step simplifies the multi-head self-attention mechanism, reduces the number of linear transformation matrices, and the score i and the calculation formula of the attention function are as follows. Here, i is the number of calculations. The text deep context feature sequence H is transformed into the query (Q) and key (K) matrix spaces through the linear matrix W i H . The entity label embedding context feature sequence H E is transformed into the value (V) matrix space through the linear matrix . The attention function calculates the correlation parameters between the text deep context feature sequences themselves, and assigns attention weights to the entity label vector sequence according to the correlation parameters. The weighted entity label vector sequence and the text deep context feature sequence are added together to obtain the text deep context feature sequence incorporating entity label information. Repeat the score i calculation process N times, and splice the results of the N text deep context feature sequences incorporating entity label information, and output z = concat(score 1 , score 2,...,score N ), concat is the concatenation operation;

[0179]

[0180]

[0181] Secondly, for the text deep context feature Z that incorporates entity label information, perform relationship prediction on the character c i and the character c j The corresponding text deep context features are denoted as and For the character c i and the character c j The calculation formula for the relationship probability between them is as follows. Among them, p(c i , r, c j ) represents the predicted probability of each relationship r between the character c i and the character c j . is the scoring function for calculating the relationship between the character c i and the character c j . σ is the sigmoid function. Using the sigmoid function can independently consider each relationship between character pairs.

[0182] Step 5: Based on the relationship label prediction results predicted in the first stage, each type of relationship prediction label has a corresponding character pair relationship category probability matrix, which serves as the input to the relationship decoding attention mechanism. The relationship decoding attention mechanism calculates the character-weighted relationship features based on the relationship label prediction results in the first stage. The weighted character relationship feature sequence is added to the text deep context feature sequence in Step 2, introducing the relationship prediction result information. Repeat the processes of Step 3 and Step 4 to perform entity relationship joint extraction and relationship label prediction in the second stage.

[0183] The entity prediction results and relationship label prediction results output in the second stage are the final outputs of the GAN generator part.

[0184] The above Steps 1 to 5 can be the pre-training process of the GAN generator.

[0185] To better understand Step 5, the following is an expanded explanation:

[0186] a) The entity relationship mainly predicts the relationship category probability between character pairs. If there is a relationship between characters, the entity relationship connection is at the last character of the head and tail entities. If there is no relationship, it is predicted as N. Since the relationship connects the head and tail entities, there is a corresponding character pair - to - relationship category probability matrix for each relationship category. Given a relationship category r, in the corresponding character pair - to - relationship category probability matrix, p r p(u, v) represents the probability that characters u and v predict relationship r, where u is the head character and v is the tail character. At the same time, there is also the relationship p r p(v, u), which represents the probability that characters v and u predict relationship r. Therefore, according to the positions of the characters in the head and tail entities, the character pair - to - relationship category probability matrix is divided into a forward character pair relationship category probability matrix and a backward character pair relationship category probability matrix, which are used as the input of the relationship decoding attention mechanism.

[0187] b) Taking the forward character pair relationship category probability matrix as an example, the relationship decoding attention mechanism first calculates the scoring coefficient between the text deep context feature sequence that fuses entity label information and the forward character pair relationship category probability matrix. The formula for the scoring coefficient is as follows:

[0188] Among them, is the probability vector of the tail entity character for the i - th character predicting the head entity pair relationship category r, and w a , v, and w b are training parameters.

[0189] c) When obtaining , calculate the weight The calculation formula is as follows:

[0190] Among them, r is the number of relationship categories. Multiply the weight by the corresponding and sum them to obtain the weighted probability matrix of the forward character pair relationship category r.

[0191] d) Similarly, calculate the weighted probability matrix of the backward character pair relationship category r according to steps b) and c). Add the weighted probability matrices corresponding to all relationship categories. The two matrices update the text deep context feature that fuses entity label information through matrix mapping. The update formula is as follows:

[0192] Among them, and are the weighted probability matrices of the forward character pair relationship category and the backward character pair relationship category respectively, and w bb and w ccAs training parameters, the updated deep context features of the text have incorporated the entities and entity relationships predicted in the first stage. On this basis, joint extraction of entity relationships in the second stage is carried out, and the entities and entity relationship labels predicted in the second stage are the output results of the generator method.

[0193] Step 6: For the training set and test set of the manually annotated sample set, the training set is used to train the generator part. After training, the GAN generator predicts entity and relationship labels for the test samples in the test set. The position features of the entity characters in the sentence are obtained through the predicted entity labels, and are concatenated with the character features obtained in Step 1 and the entity label embeddings obtained in Step 3 to form character hybrid features. That is, for the test set, the GAN generator outputs the entity prediction results and relationship label prediction results of each test sample, as well as the intermediate process information based on the GAN generator, to obtain the character hybrid features of the test samples.

[0194] Step 7: Use the character hybrid as the input of the hybrid neural network structure to obtain the output sentence features.

[0195] The hybrid neural network includes BiLSTM and a segmented convolutional neural network, mainly for obtaining sentence features; the BiLSTM in this hybrid neural network belongs to an independent structure and does not repeat with the foregoing steps.

[0196] Specifically, this step may include the following sub-steps:

[0197] a) After the pre-training of the GAN generator method is completed, the generator method predicts the manually annotated test samples and generates the position features of the predicted entities in the sentence. The calculation formula for the position features is as follows:

[0198] where N represents the number of entity predictions in the sentence, p i,j represents the number of characters between the i-th character and the j-th entity, and the position encodings of all characters of the predicted entity in the sentence are 0, the position encodings of the characters to the left of the predicted entity are negative, and the position encodings of the characters on the right are positive. Calculate the average distance of each character from all predicted entities in the sentence to obtain the character position features.

[0199] b) The entity label embedding features are mainly mapped to the predicted entity embedding features according to the entity categories predicted by CRF in Step 3, and are concatenated with the character features output in Step 1 to form character hybrid features, which are used as the input of the hybrid neural network structure.

[0200] c) Use BiLSTM to capture the context information of the character hybrid features to obtain the character hybrid context feature sequence h = {h 1 ,h 2 ,...,h n}. Secondly, the character-mixed context features are segmented into character context features, entity prediction label embedding context features, and character position context features, and a segmented convolution operation is performed on each segment of feature vectors. The convolution formula for each segment is as follows:

[0201] c i = f(F·h i:i+l-1 + b);

[0202] where l is the height of the convolution kernel, i is the i-th convolution window, F is the convolution kernel, f represents the RELU non-linear function, and the segmented feature matrix output by the segmented convolution neural network is C j = (c 1,j , c 2,j ,..., c m,j ), where m is the number of convolution kernels and j is the j-th feature segment.

[0203] d) The segmented feature matrix is further processed by segmented max pooling to extract segmented prominent features. The max pooling operation is performed on each feature segment vector respectively, and the max pooling operation is as shown:

[0204] p ij = max(c ij ), 1 ≤ i ≤ m, 1 ≤ j ≤ 3;

[0205] After that, the three max pooling vectors are concatenated together to obtain the vector p i = {p i1 , p i2 , p i3}. Connect all the p i and perform a non-linear function operation, then the sentence feature vector output by the max pooling layer can be obtained as s = tanh(p 1:m ).

[0206] Step 8: Process the sentence features using a multi-stage sample selection strategy.

[0207] First, calculate the support of the sentence features for each type of entity relationship label. The support is used as the weight of each test sample. The greater the support, the more entity relationship labels and triples there are in the sample. Sort the test samples according to the support weights and select some test samples. The purpose of this step is to select samples with more triples, which is the first stage of sample selection.

[0208] Specifically, the output sentence features are used to calculate the support of the sentence features for the relationship r through the Sigmoid function. Since each sample sentence can contain multiple relationship categories, the Sigmoid function can independently assign a larger prediction probability to multiple relationships, and better screen samples with more triples into the next step. This is the first stage of sample selection.

[0209] Step 9: Based on the test samples selected in Step 8, using characters as nodes, perform graph convolution operations on the character mixed features in Step 6 with a graph convolutional neural network to output the convolved character features, and add the convolved character features and the character mixed features to obtain a new character feature sequence.

[0210] The adjacency matrix of the sentence samples used in the graph convolutional neural network is the relationship category probability matrix between the characters predicted in Step 5. The purpose of this step is to add the predicted entity relationship information to the character mixed features to form a new character feature sequence.

[0211] Specifically, this step may include the following sub-steps:

[0212] a) For each sample in the selected test samples, using characters as nodes, taking the character mixed features of the sample as the character node features, the character adjacency matrix required by the graph convolutional neural network uses the forward character pair relationship category probability matrix and the backward character pair relationship category probability matrix obtained in Step 5. The graph convolution operation introduces the relationship prediction information into the character node features. The graph convolution calculation formula is as follows:

[0213]

[0214] Among them, is the node feature of node u in the l-th layer of the graph convolutional neural network. V and R are the total number of characters and the total number of relationship categories respectively.

[0215] b) Update the character node features according to the graph convolution iterative features. The update formula is as follows:

[0216] Among them, h u represents the mixed feature of character u. Finally, set the updated feature sequence as H'=(h 1 ', h' 2 ,..., h' n ).

[0217] Step 10: The new character feature sequence is used to calculate the first prominent feature of the test sample through CNN and max pooling, and use this first prominent feature as the latent variable of the test sample;

[0218] Furthermore, based on the manually labeled entities and entity relationship labels of each test sample in the test set, generate the corresponding second prominent feature in the manner of Steps 6 - 10, and use the second prominent feature as the true variable of the test sample.

[0219] Also, the latent variables and real variables are input into the Softmax function for classification, which is the sample selection in the second stage. After the latent variables and real variables are input into the Softmax for training and the training is completed, the GAN generator is fine-tuned. Finally, in step 12, the unlabeled samples with a large number of predicted entity relationships are sorted and classified in the second stage.

[0220] Specifically, a) perform convolution operation on the updated character node feature sequence using convolution kernels of different sizes, splice the l-gram features of different convolution kernel sizes, and perform max pooling operation on the spliced l-gram features to obtain prominent features, which are defined as the latent variables of the test samples;

[0221] b) Replace the predicted entities and predicted entity relationship labels with the manually labeled entities and entity relationship labels in the test set, and output the corresponding prominent features in the manner of steps 6 - 10, which are defined as the real variables of the test samples;

[0222] c) Fix the generator parameter variables, and input the latent variables and real variables into the Softmax function for classification training. The loss function is as follows:

[0223]

[0224] where, v P and v R are the latent variable and real variable of the sample respectively, and are the predicted probabilities of the sample variables. D represents the discriminator, and E represents the mathematical expectation. The discriminator training is used to distinguish the feature distributions of the predicted labels and real labels.

[0225] Step 11: When the discriminator training is completed, fix the discriminator parameters and perform adversarial training on the parameters of the generator method.

[0226] Step 12: After the adversarial training is completed, use the trained GAN generator to predict the unlabeled samples. Using the aforementioned process, sort and select the unlabeled sample set. For example, label the selected unlabeled samples with the entity and relationship labels predicted by the GAN generator, add them to the manually labeled training set, and delete them from the unlabeled sample set. Repeat steps 1 - 5 to train the generator method, repeat steps 6 - 11 to train the discriminator method, and step 12 to select unlabeled samples until the number of iterations is met.

[0227] In the above steps 8 to 10, the trained GAN generator is used to predict unlabeled samples, and the entity relationship prediction labels for two stages are output. In step 6, the sentence character mixed features of unlabeled samples are generated, and in step 7, the sentence features of unlabeled samples are obtained. In step 8, unlabeled samples with a large number of predicted entity relationships are selected. Based on the unlabeled samples with a large number of predicted entity relationships, in steps 9 to 10, the softmax function is used to select unlabeled samples for which the entity relationships can be predicted correctly as much as possible. The selected samples are automatically labeled with the entity and relationship labels predicted by the generator and added to the manually labeled training set.

[0228] For example, a) Fix the discriminator parameters and perform adversarial training on the generator. The adversarial training makes the generator attempt to deceive the discriminator into believing that the entity and entity relationship prediction labels generated by the generator are as close as possible to the true labels, thereby minimizing the feature distribution gap between the prediction labels and the true labels for subsequent step 12 to select unlabeled samples.

[0229] In the generator of the GAN in the method of this embodiment, in view of the problem of cumulative propagation of sub-task interaction errors that may exist in the existing entity relationship joint extraction method, the present invention proposes a context entity label space attention mechanism and a relationship decoding attention mechanism. Through the self-training of the generator, entity and entity relationship features are selectively fused to reduce the propagation of error information and improve the knowledge extraction effect. In view of the problem of limited management corpus annotation quantity, a multi-stage sample selection strategy is proposed in the discriminator of the GAN. Through a two-stage sample selection process, samples containing more triples can be selected as much as possible. After the GAN discriminator training is completed, adversarial training is introduced to train the generator to make the generator predict entity and relationship labels as correct as possible for unlabeled samples, and through the screening of the discriminator, the selected samples are automatically labeled with prediction labels and added to the labeled training sample set for the next round of generator iterative training. The method reduces the cost of manually labeled samples and solves the problem that the method falls into underfitting due to the lack of large-scale labeled training samples in the supply chain management field.

[0230] According to another aspect of the embodiments of the present invention, the embodiments of the present invention also provide an electronic device, including a memory and a processor. The memory stores a computer program, and the processor executes the computer program in the memory and executes the steps of any of the above-mentioned entity relationship joint extraction methods based on active learning in the supply chain management field in Embodiment 1 or Embodiment 2.

[0231] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a system or a computer program product. Therefore, the present invention can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memory, CD-ROM, optical memory, etc.) that contain computer-usable program code.

[0232] In the description of this specification, the description of terms such as "one embodiment", "some embodiments", "embodiment", "example", "specific example" or "some examples", etc., means that the specific features, structures, materials or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. Although the preferred embodiments of the present invention have been described, those skilled in the art can make additional changes and modifications after learning the basic creative concepts. Therefore, the claims should be construed to include the preferred embodiments as well as all changes and modifications falling within the scope of the present invention.

[0233] Obviously, those skilled in the art can make various modifications and variations to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the claims of the present invention and their equivalent technologies, the present invention should also include these modifications and variations.

Claims

1. A joint extraction method of entity relationships in the field of supply chain management based on active learning, characterized in that: include: A100, pre-training the GAN generator based on the training set of manually annotated data in the supply chain management text to be processed; A200, use the pre-trained GAN generator to process the test samples in the test set, and output the relationship label prediction results and entity label prediction results of the test samples; A300, obtaining a character mixed feature of the test sample according to the entity label prediction result of the test sample and the intermediate character related variable of the entity label prediction result; A400, train the GAN discriminator based on the character mixed features of the test samples, the relationship label prediction results and the entity label prediction results, and the manually annotated relationship labels and entity label results; A500. After the training of the GAN discriminator is completed, the pre-trained GAN generator is subjected to adversarial training for fine-tuning, and the fine-tuned GAN generator is used to predict unlabeled sample entities in the supply chain management text to be processed, and the GAN discriminator is used to sort and select them. The selected samples are added to the training set, and the above training process is repeated until the number of iterations is met to obtain the trained GAN generator.

2. The method according to claim 1, characterized in that The A100, includes: A101. Use the BERT pre-training method to obtain the character feature vector sequence x=(x1,x2,x i ...,x n ); n is the number of characters in the sentence length, x i Represents the character feature vector of the i-th character; A102. Use stacked BiLSTM to extract contextual features from character feature vector sequences and output text deep context feature sequences L is the number of stacking layers, Represents the deep contextual features of the text of the i-th character; A103. Use CRF to identify sentence entities in the deep context feature sequence of the text and obtain entity labels. Each identified entity label has a corresponding entity label embedding e. The entity label embedding sequence of all sentence entities is E = (e1, e2, ..., e n ), E=(e1,e2,...,e n ) is input to BiLSTM to obtain contextual features with entity tag embedding sequence A104, through the multi-head self-attention mechanism Fusion to , obtaining a new context feature sequence Z; and performing relationship label prediction on the new context feature sequence Z, outputting the relationship label prediction result, and using the relationship label prediction result and the entity label as the entity label prediction and relationship label prediction result of the first stage; A105. Calculate the character weighted relationship features for the relationship label prediction results of the first stage through the relationship decoding attention mechanism to obtain a weighted character relationship feature sequence; The weighted character relationship feature sequence is combined with The entity labels are obtained by adding the inputs to the CRF, and the relationship label prediction results are obtained through the multi-head attention mechanism to obtain the entity label prediction and relationship label prediction results output by the second stage. The output results of the second stage are used as the output results of the GAN generator.

3. The method according to claim 2, characterized in that The A104 includes: A1041, text deep context feature sequence H through the linear matrix W i H Transformed into the query (query, Q) and key (key, K) matrix space, the entity label is embedded in the context feature sequence H E Through the linear matrix Transformed into the value (value, V) matrix space, the attention function calculates the correlation parameters between the deep context feature sequences of the text itself, and implements the attention weight allocation to the entity label vector sequence according to the correlation parameters to obtain the weighted entity label vector sequence; The weighted entity label vector sequence and the text deep context feature sequence are added together to obtain the text deep context feature sequence that integrates the entity label information, i.e., the new context feature sequence Z; Score i The calculation process is repeated N times, and the N-times fusion entity label information text deep context feature sequence results are spliced, and the output z = concat (score1, score2, ..., score N ), concat is a concatenation operation; Where i is the number of calculations; A1042, perform relationship prediction on the deep context feature Z of the text that incorporates the entity label information, character c i and the character c j The corresponding text deep context feature is denoted as and Character c i The relationship probability calculation formula between and character cj is: Among them, p(c i ,r,c j ) represents the character c i and the character c j The predicted probability of each relationship r between is the calculation character c i and the character c j The relationship between them belongs to the score function, and σ is the sigmoid function.

4. The method according to claim 2, characterized in that: A105 includes: A1051, Entity Relationship mainly predicts the probability of the relationship category between entity pairs. If there is a relationship between entity pairs, the entity relationship is connected to the last character of the head and tail entities. If there is no relationship, it is predicted to be N. Each type of relationship has a corresponding character combination relationship category probability matrix. According to the head and tail entity positions of the characters, the character pair relationship category probability matrix is ​​divided into a forward character pair relationship category probability matrix and a backward character pair relationship category probability matrix as the input of the relationship decoding attention mechanism; A1052. For the forward character pair relationship category probability matrix, the relationship decoding attention mechanism first calculates the score coefficient between the text deep context feature sequence fused with entity label information and the forward character pair relationship category probability matrix. The score coefficient calculation formula is as follows: in, The probability vector of the tail entity character predicted as the head entity pair relationship category r for the i-th character, w a , v and w b is the training parameter; A1053, when you get Then, calculate the weight The calculation formula is as follows: Among them, r is the number of relationship categories, and the weight With the corresponding Multiply and sum to obtain the weighted probability matrix of the forward character pair relationship category r; A1054. Obtain the weighted probability matrix of the backward character pair relationship category r, add the weighted probability matrices corresponding to all relationship categories, and update the deep context features of the text that incorporates the entity label information through matrix mapping. The update formula is as follows: in, and are the weighted probability matrix of the forward character pair relationship category and the weighted probability matrix of the backward character pair relationship category, w bb and w cc is the training parameter.

5. The method according to claim 1, characterized in that The A200 includes: For each test sample, the position feature of the predicted entity in the sentence, i.e., the entity label prediction result, is generated. The calculation formula is as follows: Where N represents the number of entity predictions in the sentence, and p i,j It represents the length of the number of characters between the ith character and the jth entity character, and the position of all characters predicted as entities in the sentence is encoded as 0, the character position to the left of the predicted entity is encoded as a negative number, and the character position to the right is encoded as a positive number. The average distance of each character from all predicted entities in the sentence is calculated to obtain the character position feature; The A300 includes: A301. For each test sample, concatenate the entity label prediction result of the test sample, the character feature vector in the character feature vector sequence, and the entity label embedding in the entity label embedding sequence to form a character mixed feature; The character feature vectors in the character feature vector sequence and the entity label embeddings in the entity label embedding sequence are intermediate character-related variables in the process of processing test samples using a pre-trained GAN generator.

6. The method according to claim 1, characterized in that The A400 includes: A401, for each test sample, inputting the character mixed features of each test sample into the hybrid neural network to obtain the sentence features output by the hybrid neural network; A402. Based on a multi-stage sample selection strategy, the support of sentence features for each type of entity relationship label is calculated. The support is used as the weight of each test sample. The test samples are sorted according to the support weight, and a preset number of samples are screened out based on the sorting as the first stage of sample selection. A403, based on each test sample in the sample selection of the first stage, taking the characters in the test sample as nodes and based on the relationship label prediction result of the test sample, using a graph convolution neural network to perform a graph convolution operation on the character mixed feature of the test sample, outputting the convolved character feature, and adding the convolved character feature and the character mixed feature to obtain a new character feature sequence; A404, based on the new character feature sequence, using CNN and maximum pooling to calculate the first prominent feature of the test sample, and using the first prominent feature as a latent variable of the test sample; Based on the manually annotated entity label results and entity relationship label results of the test sample, the process of A300 and A400 is repeated to generate the second salient feature corresponding to the test sample, and the second salient feature is used as the real variable of the test sample; the latent variable and the real variable are input into the Softmax function for training, and after the training is completed, they are used for the second stage selection of unlabeled samples; A500 includes: After adversarial training is completed, the GAN generator predicts unlabeled samples, and the GAN discriminator sorts and selects the unlabeled sample set; the GAN discriminator selects samples with more entity relationship triplets in the first stage according to the support of the sentence features for the relation r, and then selects the correct samples through the softmax function in the second stage; The entity label results and relationship label results predicted by the GAN generator corresponding to the selected unlabeled samples are used as the annotation information of the unlabeled samples and added to the manually labeled training set; the training process of the GAN generator and the GAN discriminator using the training set and the test set is repeated until the number of iterations is met.

7. The method according to claim 6, characterized in that A401 includes: The hybrid neural network includes: BiLSTM and segmented convolutional neural network; A4011. Capture the context information of character mixed features through BiLSTM to obtain the character mixed context feature sequence h = {h1,h2,...,h n }; A4012. The character mixed context features in the character mixed context feature sequence are divided into character context features, entity prediction label embedding context features and character position context features, and each segment of the feature vector is subjected to a segmented convolution operation. The convolution formula for each segment is as follows: c i =f(F·h i:i+l-1 +b) Among them, l is the height of the convolution kernel, i is the i-th convolution window, F is the convolution kernel, f represents the RELU nonlinear function, and the segmented feature matrix output by the segmented convolutional neural network is C j =(c 1,j ,c 2,j ,...,c m,j ), where m is the number of convolution kernels and j is the jth feature segment; A4013, segmented feature matrix is ​​subjected to segmented maximum pooling to extract segmented salient features; The maximum pooling operation is performed on each feature segment vector respectively. The maximum pooling operation is: p ij =max(c ij ),1≤i≤m,1≤j≤3 Concatenate the three maximum pooling vectors together to get vector p i ={p i1 ,p i2 ,p i3 }; Connect all p i And perform nonlinear function operation to obtain the sentence feature vector output by the maximum pooling layer as s = tanh (p 1:m ).

8. The method according to claim 6, characterized in that A402 includes: The sentence features are used to calculate the support of the sentence features for the relation r through the Sigmoid function, and samples containing more triples are selected based on the support; A403 includes: The graph convolution operation introduces relationship prediction information into the character node features. The graph convolution calculation formula is as follows: in, is the node feature of the node u of the l-th layer graph convolutional neural network, V and R are the total number of characters and the total number of relationship categories, respectively. Update the character node features according to the graph convolution iterative features. The update formula is as follows: Among them, h u represents the mixed features of character u. Finally, the updated feature sequence is set to H'=(h1',h'2,...,h' n ).

9. The method according to claim 8, characterized in that A404 includes: Using convolution kernels of different sizes to perform convolution operations on the updated character node feature sequence, concatenating l-gram features of different convolution kernel sizes, and performing a maximum pooling operation on the concatenated l-gram features to obtain a first salient feature, which is defined as a latent variable of the test sample; Replace the predicted entity and predicted entity relationship labels with the manually annotated entity and entity relationship labels of the test sample, and obtain the second salient feature corresponding to the test sample as the true variable of the test sample; Fixed generator parameter variables, latent variables and real variables are input into the Softmax function for classification training, and the loss function is as follows: Among them, v P and v R are the latent variables and true variables of the samples, and is the predicted probability of the sample variable, D represents the discriminator, and E represents the mathematical expectation; the discriminator training is used to distinguish the feature distribution of the predicted label and the true label.

10. An electronic device, characterized in that: It includes a memory and a processor, the memory stores a computer program, the processor executes the computer program in the memory, and executes the steps of the method for joint extraction of entity relationships in the supply chain management field based on active learning as described in any one of claims 1 to 9.