An entity relationship extraction method, device, equipment and storage medium
By introducing pre-trained relationship coding modules and clustering modules into the entity relationship extraction model, using the clustering module to generate pseudo-labels and optimize model parameters, the problem of high cost of manual labeling in the existing technology is solved, and efficient entity relationship extraction without manual labeling is achieved.
Patent Information
- Application Number
- CN202211324149.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-10-27
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2042-10-27
AI Technical Summary
The existing entity relationship extraction method requires a large amount of manual labeling data, which leads to high labor costs and makes it difficult to efficiently extract entity relationships in unlabeled sentences.
By entering the sentence sample set into the initial entity relationship extraction model, the model includes a pre-trained relationship encoding module and a clustering module, using the clustering module to cluster to obtain pseudo-labels, and by iteratively adjusting the model parameters, the model is gradually optimized to achieve entity relationship extraction without manual labeling.
It realizes entity relationship extraction without manual marking, reduces labor costs, and improves the efficiency and accuracy of entity relationship extraction.
Smart Images

Figure CN115544273B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer processing technologies, and in particular, to a method, apparatus, device, and storage medium for entity relationship extraction. Background Art
[0002] As an important research topic in the field of information extraction, entity relationship extraction mainly aims to extract the semantic relationships between the marked entities in a sentence, that is, to determine the relationship categories between entity pairs in unstructured text on the basis of entity recognition, and form structured data for storage and access, providing data support for downstream tasks such as knowledge graphs.
[0003] Existing relationship extraction methods can effectively learn the semantic relationships between entities in a sentence, but most are based on supervised models and require a large amount of labeled data, and high-quality labeled data often requires a lot of manpower. Summary of the Invention
[0004] The present invention provides a method, apparatus, device, and storage medium for entity relationship extraction to solve the problem of high labor costs for training an entity relationship extraction model by manually labeling tags for sentence samples, and to implement an entity relationship extraction method without manual labeling.
[0005] According to one aspect of the present invention, there is provided a method for training an entity relationship extraction model, including:
[0006] Inputting a sentence sample set into an initial entity relationship extraction model, where the relationship extraction model includes: a pre-trained relationship encoding module and a clustering module; and the sentence sample set is composed of unlabeled sentence samples;
[0007] Performing entity relationship prediction on the sentence samples in the sentence sample set through the relationship encoding module to obtain an entity relationship graph, and inputting the entity relationship graph into the clustering module;
[0008] Performing clustering on the basis of the similarity between entities in the entity relationship graph through the clustering module to obtain at least one first entity relationship cluster and the pseudo-labels of the sentence samples included in the first entity relationship cluster; where the sentence samples included in the same first entity relationship cluster are labeled with the same pseudo-label;
[0009] Updating the sentence sample set according to the sentence samples with pseudo-labels, inputting the updated sentence sample set into the initial entity relationship extraction model to obtain at least one second entity relationship cluster and the prediction labels of the sentence samples included in the second entity relationship cluster;
[0010] Calculate the loss function value based on the pseudo-labels and predicted labels corresponding to the sentence samples, and iteratively adjust the network parameters in the initial entity relationship extraction model based on the loss function value to obtain a target entity relationship extraction model.
[0011] According to another aspect of the present invention, there is provided an entity relationship extraction method, including:
[0012] Obtain a sentence to be extracted;
[0013] Input the sentence to be extracted into a target entity relationship extraction model trained by using the training method of the entity relationship extraction model described in any embodiment;
[0014] Obtain the entity relationship category of the sentence to be extracted output by the target entity relationship extraction model.
[0015] According to another aspect of the present invention, there is provided a training device for an entity relationship extraction model, including:
[0016] An input module, configured to input a sentence sample set into an initial entity relationship extraction model, where the relationship extraction model includes: a pre-trained relationship encoding module and a clustering module; the sentence sample set is composed of unlabeled sentence samples;
[0017] The relationship encoding module is configured to perform entity relationship prediction on the sentence samples in the sentence sample set to obtain an entity relationship graph, and input the entity relationship graph into the clustering module;
[0018] The clustering module is configured to perform clustering based on the similarity between entities in the entity relationship graph to obtain at least one first entity relationship cluster and the pseudo-labels of the sentence samples included in the first entity relationship cluster; where the sentence samples included in the same entity relationship cluster are marked with the same pseudo-labels;
[0019] The update module is configured to update the sentence sample set according to the sentence samples with pseudo-labels, input the updated sentence sample set into the initial entity relationship extraction model to obtain at least one second entity relationship cluster and the predicted labels of the sentence samples included in the second entity relationship cluster;
[0020] The adjustment module is configured to calculate the loss function value based on the pseudo-labels and predicted labels corresponding to the sentence samples, and iteratively adjust the network parameters in the initial entity relationship extraction model based on the loss function value to obtain a target entity relationship extraction model.
[0021] According to another aspect of the present invention, there is provided an entity relationship extraction device, including:
[0022] A sentence acquisition module, configured to obtain a sentence to be extracted;
[0023] An input module, configured to input the sentence to be extracted into a target entity relationship extraction model trained by using the training method of the entity relationship extraction model according to any one of the embodiments.
[0024] A result acquisition module, configured to acquire an entity relationship extraction result of the sentence to be extracted output by the target entity relationship extraction model.
[0025] According to another aspect of the present invention, there is provided an electronic device, including:
[0026] At least one processor; and
[0027] A memory communicatively connected to the at least one processor; wherein,
[0028] The memory stores a computer program executable by the at least one processor, and when the computer program is executed by the at least one processor, the at least one processor is enabled to execute the training method of the entity relationship extraction model or the entity relationship extraction method according to any one of the embodiments of the present invention.
[0029] According to another aspect of the present invention, there is provided a computer-readable storage medium storing computer instructions for causing a processor to implement the training method of the entity relationship extraction model or the entity relationship extraction method according to any one of the embodiments of the present invention when executed.
[0030] The technical solution of the embodiment of the present invention is to input a sentence sample set into an initial entity relationship extraction model, where the relationship extraction model includes: a pre-trained relationship encoding module and a clustering module; the sentence sample set is composed of unlabeled sentence samples; the relationship encoding module is used to perform entity relationship prediction on the sentence samples in the sentence sample set to obtain an entity relationship graph, and input the entity relationship graph into the clustering module; the clustering module is used to perform clustering based on the similarity between entities in the entity relationship graph to obtain at least one first entity relationship cluster and the pseudo-labels of the sentence samples included in the first entity relationship cluster; where the sentence samples included in the same entity relationship cluster are marked with the same pseudo-labels; update the sentence sample set according to the sentence samples with pseudo-labels, and input the updated sentence sample set into the initial entity relationship extraction model to obtain at least one second entity relationship cluster and the predicted labels of the sentence samples included in the second entity relationship cluster; calculate the loss function value according to the pseudo-labels and predicted labels corresponding to the sentence samples, and iteratively adjust the network parameters in the initial entity relationship extraction model based on the loss function value to obtain a target entity relationship extraction model, which can solve the problem of high labor cost in training an entity relationship extraction model by manually labeling sentence samples, and realize an entity relationship extraction method without manual labeling.
[0031] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present invention, nor is it used to limit the scope of the present invention. Other features of the present invention will become easily understood through the following description. Brief Description of the Drawings
[0032] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0033] Figure 1A It is a flowchart of a method for training an entity relationship extraction model provided in Embodiment 1 of the present invention;
[0034] Figure 1B It is a structural schematic diagram of an entity relationship extraction model;
[0035] Figure 1C It is a structural schematic diagram of the relationship encoding module of an entity relationship extraction model;
[0036] Figure 2 It is a flowchart of an entity relationship extraction method provided in Embodiment 2 of the present invention;
[0037] Figure 3 It is a schematic structural diagram of a training device for an entity relationship extraction model provided in Embodiment 3 of the present invention;
[0038] Figure 4 It is a schematic structural diagram of an entity relationship extraction device provided in Embodiment 4 of the present invention;
[0039] Figure 5 It is a schematic structural diagram of an electronic device for implementing the training method or the entity relationship extraction method of the entity relationship extraction model in the embodiments of the present invention. Detailed implementation manners
[0040] In order to enable those skilled in the art to better understand the solution of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0041] It should be noted that the terms "first", "second", etc. in the specification and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily need to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present invention described here can be implemented in an order other than those illustrated or described here. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device including a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0042] Embodiment 1
[0043] Figure 1A For Embodiment 1 of the present invention, a flowchart of a training method for an entity relationship extraction model is provided. This embodiment is applicable to the situation where an entity relationship extraction model is trained without manually labeling a sentence sample set. This method can be executed by a training device for an entity relationship extraction model. The training device for an entity relationship extraction model can be implemented in the form of hardware and / or software, and the training device for an entity relationship extraction model can be configured in an electronic device. As Figure 1A shown, the method includes:
[0044] S110. Input the sentence sample set into the initial entity relation extraction model. Among them, the relation extraction model includes: a pre-trained relation encoding module and a clustering module; the sentence sample set consists of unlabeled sentence samples.
[0045] Among them, the initial entity relation extraction model refers to the initially established entity relation extraction model that has not been trained or is not completely trained. The entity relation extraction model is a model used to extract entity relations from the input sentences. The initial entity relation extraction model needs to be iteratively trained with the input sentence sample set to obtain a completely trained target entity relation extraction model.
[0046] In the embodiment of the present invention, the relation extraction model includes: a pre-trained relation encoding module and a clustering module. The pre-trained relation encoding module can be understood as pre-training the initially established relation encoding module with a certain number of sentence samples, so that the pre-trained relation encoding module can predict entity relations by extracting the context features between entities in the input sentence samples. The clustering module aims to cluster the input sentence samples into multiple semantically meaningful relation clusters, thereby determining the potential relation classification of the sentence samples.
[0047] Specifically, input the sentence sample set into the initial relation extraction model, and process the sentence samples in the input sentence sample set sequentially through the pre-trained relation encoding module and the clustering module.
[0048] S120. Perform entity relation prediction on the sentence samples in the sentence sample set through the relation encoding module to obtain an entity relation graph, and input the entity relation graph into the clustering module.
[0049] Among them, the entity relation graph can be understood as a structure graph used to describe the relations between entities in the sentence samples.
[0050] Specifically, input the sentence sample set into the pre-trained relation encoding module, and perform feature extraction and relation prediction on the sentence samples in the sentence sample set sequentially through the relation encoding module to obtain an entity relation graph.
[0051] Exemplarily, the relation encoding module may include: an embedding layer, an encoding layer, a decoding layer, and a prediction layer. Through the embedding layer, the entity feature extraction of the input sentence samples can be continued to obtain entity vectors in embedded representation. The encoding layer and the decoding layer are used to perform semantic feature learning on the entity vectors of the input sentence samples to obtain sentence vectors. The prediction layer is used to perform relation prediction on the input sentence vectors to obtain an entity relation graph.
[0052] S130. Based on the similarity between entities in the entity relationship graph, the clustering module performs clustering to obtain at least one first entity relationship cluster and the pseudo-labels of the sentence samples included in the first entity relationship cluster; wherein, the sentence samples included in the same first entity relationship cluster are marked with the same pseudo-label.
[0053] Among them, the first entity relationship cluster can be understood as an entity cluster obtained by clustering the entities in the entity relationship graph through the clustering module, and each entity cluster is a first entity relationship cluster.
[0054] The pseudo-label can be regarded as the label marked for the sentence samples included in the first entity relationship cluster. Among them, the sentence samples included in the same first entity relationship cluster are marked with the same pseudo-label; the sentence samples included in different first entity relationship clusters are marked with different pseudo-labels.
[0055] It should be noted that this pseudo-label is different from the label marked manually, and this pseudo-label needs to be continuously optimized and updated through the training process of the model. And this pseudo-label can be used to represent the first entity relationship cluster to which the relationship between entities in the sentence sample belongs, that is, to distinguish sentence samples with different entity relationships, but the relationship between entities in this sentence sample cannot be determined. For example, after clustering, two first entity relationship clusters are obtained, then the pseudo-label of the sentence samples included in the first first entity relationship cluster can be class A; the pseudo-label of the sentence samples included in the first first entity relationship cluster can be class B. Class A and class B are different, but the actual entity relationship types represented by class A and class B cannot be determined.
[0056] Specifically, the entity relationship graph is input into the clustering module, and the clustering model fuses and clusters the entity relationship graph to obtain at least one first entity relationship cluster, and marks the pseudo-labels for the sentence samples included in each first entity relationship cluster according to the standard that the sentence samples included in the same first entity relationship cluster are marked with the same pseudo-label, and the sentence samples included in different first entity relationship clusters are marked with different pseudo-labels.
[0057] S140. Update the sentence sample set according to the sentence samples with pseudo-labels, and input the updated sentence sample set into the initial entity relationship extraction model to obtain at least one second entity relationship cluster and the prediction labels of the sentence samples included in the second entity relationship cluster.
[0058] Among them, the second entity relationship cluster can be understood as an entity cluster obtained by clustering the entities in the entity relationship graph of the updated sentence samples through the clustering module, and each entity cluster is a second entity relationship cluster.
[0059] The predicted label of a sentence sample refers to the result predicted by the initial entity relationship extraction model for relationship extraction of the sentence sample.
[0060] Specifically, update the sentence sample set according to the sentence samples marked with pseudo-labels. After the update, some sentence samples in the sentence sample set may be marked with pseudo-labels, while some sentence samples may not be marked with pseudo-labels. Input the updated sentence sample set into the initial entity relationship extraction model. Process the input updated sentence sample set through the relationship encoding module and the clustering module of the initial entity relationship extraction model to obtain at least one second entity relationship cluster.
[0061] It should be noted that inputting the updated sentence sample set into the initial entity relationship extraction model to obtain at least one second entity relationship cluster is basically the same as the steps from step S110 to step S130 of inputting the sentence sample set into the initial entity relationship extraction model to obtain at least one first entity relationship cluster, and this step will not be elaborated here.
[0062] S150. Calculate the loss function value according to the pseudo-label and the predicted label corresponding to the sentence sample, and iteratively adjust the network parameters in the initial entity relationship extraction model based on the loss function value to obtain the target entity relationship extraction model.
[0063] Specifically, for a sentence sample that has both a pseudo-label and a predicted label, calculate the loss function value according to the pseudo-label and the predicted label, and continue to adjust the network parameters of each module in the initial entity relationship extraction model according to this loss function value until the loss function value reaches the minimum, and determine the initial entity relationship extraction module at this time as the target entity relationship extraction model.
[0064] The technical solution of the embodiment of the present invention is to input a sentence sample set into an initial entity relationship extraction model, where the relationship extraction model includes: a pre-trained relationship encoding module and a clustering module; the sentence sample set is composed of unlabeled sentence samples; the relationship encoding module is used to perform entity relationship prediction on the sentence samples in the sentence sample set to obtain an entity relationship graph, and input the entity relationship graph into the clustering module; the clustering module is used to perform clustering based on the similarity between entities in the entity relationship graph to obtain at least one first entity relationship cluster and the pseudo-labels of the sentence samples included in the first entity relationship cluster; where the sentence samples included in the same entity relationship cluster are marked with the same pseudo-labels; update the sentence sample set according to the sentence samples with pseudo-labels, input the updated sentence sample set into the initial entity relationship extraction model to obtain at least one second entity relationship cluster and the prediction labels of the sentence samples included in the second entity relationship cluster; calculate the loss function value according to the pseudo-labels and prediction labels corresponding to the sentence samples, and iteratively adjust the network parameters in the initial entity relationship extraction model based on the loss function value to obtain a target entity relationship extraction model, which can solve the problem of high labor cost in training an entity relationship extraction model by manually labeling tags for sentence samples, and realize an entity relationship extraction method without manual labeling.
[0065] Optionally, the relationship encoding module includes: a preset number of parallel convolutional trapezoidal networks, and each convolutional trapezoidal network includes: an embedding layer, an encoding layer, a decoding layer, and a prediction layer; the encoding layer includes: a noise-adding encoder and a denoising encoder; the decoding layer includes a first decoder and a second decoder.
[0066] Specifically, as Figure 1B shown, the relationship encoding module includes: a preset number of parallel convolutional trapezoidal networks, and the convolutional trapezoidal network is a network proposed based on the ladder network (Ladder Networks, Ladder Nets). As Figure 1C shown, the convolutional trapezoidal network includes: an embedding layer, an encoding layer, a decoding layer, and a prediction layer.
[0067] The embedding layer is used to convert text information into vector representations, enabling the encoding layer to better understand the semantics of sentences. The relationship between a pair of entities can be reflected by their context. Exemplarily, the embedding layer can use a pre-trained BERT model to extract features from sentences, obtaining context-based word embedding representations. Or a method of obtaining features for text classification based on Bi-LSTM can be used. A typical end-to-end Bi-LSTM model is an improvement based on the recurrent neural network (RNN), which can give priority to temporal features and solve the problem that RNN cannot handle word features with long-distance spans. At the same time, the positions of entity pairs in sentence samples are converted into randomly initialized position embeddings, which are used to characterize the relative distance between the two entities in the sentence. In addition, entity relationships can be associated with certain types of entities. For example, the birthplace relationship links a person and a location. Therefore, in order to better extract the deep semantic relationships between entities, in addition to their own position information, entity type information can also provide inductive biases for relationship discovery.
[0068] The encoding layer includes a noisy encoder and a denoising encoder. Among them, random Gaussian noise is applied to each layer of the noisy encoder. By learning to reconstruct sentence samples with superimposed noise, it can prevent the encoder from simply retaining the information of the original input, so that the sentence features learned by the encoding layer are more robust and the generalization ability of the network is improved. Since all layers of the noisy encoder are damaged by noise, another denoising encoder with shared parameters is needed to provide a clean reconstruction target and assist the decoding layer in unsupervised training to achieve the best mapping effect for noisy data. Using the hidden layer representation of each layer of the denoising encoder as the target value, the mean square error between the noisy encoder and the denoising encoder is minimized to obtain the unsupervised reconstruction loss of each layer of the convolutional trapezoidal network.
[0069] Exemplarily, the structures of the noisy encoder and the denoising encoder are the same, both consisting of a convolutional layer, a pooling layer, a fully connected layer, and a classification layer. A reasonable convolutional kernel height is set for the convolutional layer, and the width is the dimension of the sentence embedding vector. At the same time, in order to make the obtained features diversified, certain convolutional kernels are used to extract feature information. After convolutional, pooling, and fully connected operations, sentence vectors are obtained.
[0070] The decoding layer includes a first decoder and a second decoder. The first decoder is used to decode the vector output by the noisy encoder; the second decoder is used to decode the vector output by the denoising encoder.
[0071] The convolutional trapezoidal network adds mutual information loss on the basis of the reconstruction loss of the conventional trapezoidal network, that is, the prediction layer maximizes the mutual information between the input and output of the encoder, enabling the convolutional trapezoidal network to give similar relationship category predictions for sentences expressing similar relationships, and overall maximizing the diversity of relationship predictions, so that sentences with different relationships are predicted to have different entity relationships.
[0072] Optionally, the entity relationship prediction is performed on the sentence samples in the sentence sample set by the relationship encoding module to obtain an entity relationship graph, including:
[0073] For each of the convolutional trapezoidal networks, the embedding layer is used to extract features from the input sentence sample to obtain an embedding vector, and the embedding vector is input into the noise-adding encoder;
[0074] The noise-adding encoder is used to perform noise-adding encoding on the embedding vector to obtain a noise-adding encoded feature vector; the noise-adding encoded feature vector is input into the first decoder in the decoding layer to perform noise-adding decoding to obtain a noise-adding decoded feature vector;
[0075] The noise-adding decoded feature vector is input into the denoising encoder to perform denoising encoding to obtain a denoising encoded feature vector, and the denoising encoded feature vector is input into the second decoder in the decoding layer to perform denoising decoding to obtain a denoising decoded feature vector;
[0076] The denoising decoded feature vector is input into the prediction layer to obtain an entity relationship graph.
[0077] Specifically, the sentence sample is input into each convolutional trapezoidal network in the relationship encoding module. In the convolutional trapezoidal network, the embedding layer is used to extract features from the sentence sample, so that the sentence sample is represented as an embedding vector of the entity relationship, and the embedding vector is input into the noise-adding encoder; the noise-adding encoder is used to perform noise-adding encoding on the embedding vector to obtain a noise-adding encoded feature vector, and the noise-adding encoded feature vector is input into the first decoder in the decoding layer to perform noise-adding decoding to obtain a noise-adding decoded feature vector; since all layers of the noise-adding encoder are damaged by noise, the noise-adding decoded feature vector needs to be input into the denoising encoder to perform denoising encoding to obtain a denoising encoded feature vector, and the denoising encoded feature vector is input into the second decoder in the decoding layer to perform denoising decoding to obtain a denoising decoded feature vector, assisting the decoding layer in unsupervised training to achieve the best mapping effect for noisy data.
[0078] Optionally, the clustering module includes: a fusion network and a clustering network; correspondingly, at least one first entity relationship cluster is obtained by clustering based on the similarity between entities in the entity relationship graph through the clustering module, including:
[0079] The fusion network is used to fuse the entity relationship graphs output by each convolutional trapezoidal network in the relationship encoding module to obtain an entity relationship similarity graph, and the entity relationship similarity graph is input into the clustering network;
[0080] The clustering network is used to cluster the sentence samples in the entity relationship similarity graph whose entity relationship similarity is greater than the preset confidence level to obtain at least one first entity relationship cluster.
[0081] Among them, the entity relationship similarity graph is a graph obtained by fusing based on the similarity between the entity relationships of multiple sentence samples.
[0082] Specifically, the clustering module includes: a fusion network and a clustering network. The fusion network traverses the entire sentence sample set and fuses the entity relationship graphs corresponding to the sentence samples output by each convolutional trapezoidal network in the relationship encoding module to obtain an entity relationship similarity graph. The fusion method can be to perform node fusion according to the relationship similarity between entities in the entity relationship graph corresponding to the sentence sample to form a relationship similarity graph with multiple nodes. Input the entity relationship similarity graph into the clustering network, and cluster the sentence samples with the similarity of entity relationships in the entity relationship similarity graph greater than the preset confidence level through the clustering network to obtain at least one first entity relationship cluster.
[0083] It should be noted that the clustering algorithm adopted by the clustering module in the embodiments of the present invention is different from clustering algorithms such as k-means. In the clustering module, it does not directly cluster the entire sentence sample set, but only clusters a small part of the sample set that can obtain high accuracy, that is, not all sentence samples in the sample set participate in the clustering. Only the sentence samples with the similarity of entity relationships greater than the preset confidence level are clustered. In this way, high-confidence samples belonging to each class are extracted, which is very important for improving the accuracy of the next semi-supervised training.
[0084] In this step, by determining high-confidence samples, the unsupervised clustering performance is further improved, and unknown relationships in the open domain are more effectively extracted.
[0085] Optionally, iteratively adjusting the network parameters in the initial entity relationship extraction model based on the loss function value to obtain a target entity relationship extraction model includes:
[0086] Adjusting the network parameters in the initial entity relationship extraction model based on the loss function value;
[0087] Determining the predicted label of the sentence sample as the pseudo-label of the sentence sample;
[0088] Return to execute the step of updating the sentence sample set according to the sentence sample with the pseudo-label and inputting the updated sentence sample set into the clustering module;
[0089] Until the loss function value is the minimum value, determine the initial entity relationship extraction model corresponding to the loss function value as the target entity relationship extraction model.
[0090] Specifically, the loss function value calculated based on the pseudo-labels and predicted labels corresponding to the sentence samples is used to adjust the network parameters in the initial entity relationship extraction model; the predicted label of the sentence sample is determined as the pseudo-label of the sentence sample, and the step of updating the sentence sample set according to the sentence sample with the pseudo-label is returned. The updated sentence sample set is input into the initial entity relationship extraction model. The relationship encoding module in the initial entity relationship extraction model after parameter adjustment is used to perform entity relationship prediction on the sentence samples in the sentence sample set to obtain an entity relationship graph, and the entity relationship graph is input into the clustering module; the clustering module performs clustering based on the similarity between entities in the entity relationship graph to obtain at least one second entity relationship cluster, and the predicted labels of the sentence samples included in the second entity relationship cluster; the loss function value calculated based on the pseudo-labels and predicted labels corresponding to the sentence samples is used to adjust the network parameters in the initial entity relationship extraction model until the loss function value reaches the minimum value and the iterative training process stops. The initial entity relationship extraction model corresponding to the loss function value at this time is determined as the target entity relationship extraction model, and the training of the entity relationship extraction model is completed.
[0091] The training method of the entity relationship extraction model provided by the embodiment of the present invention does not require any labeled data and obtains supervision from the data itself in a completely unsupervised manner, which not only maintains the advantages of unsupervised learning but also has strong feature discrimination ability of supervised learning.
[0092] Embodiment 2
[0093] Figure 2 FIG. 10 is a flowchart of an entity relationship extraction method provided by Embodiment 1 of the present invention. This embodiment is applicable to the case of extracting the entity relationship of a sentence using the target entity relationship extraction model trained by the model training method described in the embodiment. This method can be executed by an entity relationship extraction device, and the entity relationship extraction device can be implemented in the form of hardware and / or software. The entity relationship extraction device can be configured in an electronic device. As Figure 2 shown, the method includes:
[0094] S210. Obtain a sentence to be extracted.
[0095] Among them, the sentence to be extracted can be understood as a sentence that requires entity relationship extraction.
[0096] S220. Input the sentence to be extracted into the target entity relationship extraction model trained by using the training method of the entity relationship extraction model.
[0097] Specifically, first, the training method of the entity relationship extraction model provided in the first embodiment is adopted. By inputting the sentence sample set into the initial entity relationship extraction model, where the relationship extraction model includes: a pre-trained relationship encoding module and a clustering module; the sentence sample set is composed of unlabeled sentence samples; the relationship encoding module performs entity relationship prediction on the sentence samples in the sentence sample set to obtain an entity relationship graph, and inputs the entity relationship graph into the clustering module; the clustering module performs clustering based on the similarity between entities in the entity relationship graph to obtain at least one first entity relationship cluster and the pseudo-labels of the sentence samples included in the first entity relationship cluster; update the sentence sample set according to the sentence samples with pseudo-labels, and input the updated sentence sample set into the initial entity relationship extraction model to obtain at least one second entity relationship cluster and the predicted labels of the sentence samples included in the second entity relationship cluster; calculate the loss function value according to the pseudo-labels and predicted labels corresponding to the sentence samples, and iteratively adjust the network parameters in the initial entity relationship extraction model based on the loss function value to obtain the target entity relationship extraction model. Then, input the sentence to be extracted into the target entity relationship extraction model.
[0098] S230. Obtain the entity relationship category of the sentence to be extracted output by the target entity relationship extraction model.
[0099] Specifically, obtain the entity relationship category obtained by the target entity relationship extraction model for continuing to perform entity relationship extraction on the input sentence to be extracted.
[0100] The technical solution of the embodiment of the present invention includes obtaining the sentence to be extracted; inputting the sentence to be extracted into the target entity relationship extraction model trained by using the training method of the entity relationship extraction model; and obtaining the entity relationship category of the sentence to be extracted output by the target entity relationship extraction model.
[0101] Embodiment III
[0102] Figure 3 It is a schematic structural diagram of a training device for an entity relationship extraction model provided in Embodiment III of the present invention. As Figure 3 shown, the device includes: an input module 310, a relationship encoding module 320, a clustering module 330, an update module 340, and an adjustment module 350;
[0103] The input module 310 is configured to input a sentence sample set into an initial entity relationship extraction model, where the relationship extraction model includes: a pre-trained relationship encoding module and a clustering module; the sentence sample set is composed of unlabeled sentence samples;
[0104] The relationship encoding module 320 is configured to perform entity relationship prediction on the sentence samples in the sentence sample set to obtain an entity relationship graph, and input the entity relationship graph into the clustering module;
[0105] A clustering module 330, configured to perform clustering based on the similarity between entities in the entity relationship graph to obtain at least one first entity relationship cluster and pseudo-labels of sentence samples included in the first entity relationship cluster; wherein, sentence samples included in the same entity relationship cluster are marked with the same pseudo-labels.
[0106] An updating module 340, configured to update the sentence sample set according to the sentence samples with pseudo-labels, input the updated sentence sample set into the initial entity relationship extraction model to obtain at least one second entity relationship cluster and predicted labels of sentence samples included in the second entity relationship cluster.
[0107] An adjustment module 350, configured to calculate a loss function value according to the pseudo-labels and predicted labels corresponding to the sentence samples, and iteratively adjust network parameters in the initial entity relationship extraction model based on the loss function value to obtain a target entity relationship extraction model.
[0108] Optionally, the relationship encoding module includes: a preset number of parallel convolutional trapezoidal networks, and each convolutional trapezoidal network includes: an embedding layer, an encoding layer, a decoding layer, and a prediction layer; the encoding layer includes: a noise-adding encoder and a denoising encoder; the decoding layer includes a first decoder and a second decoder.
[0109] Optionally, the relationship encoding module is specifically configured to:
[0110] For each convolutional trapezoidal network, perform feature extraction on the input sentence sample through the embedding layer to obtain an embedding vector, and input the embedding vector into the noise-adding encoder.
[0111] Perform noise-added encoding on the embedding vector through the noise-adding encoder to obtain a noise-added encoding feature vector; input the noise-added encoding feature vector into the first decoder in the decoding layer to perform noise-added decoding to obtain a noise-added decoding feature vector.
[0112] Input the noise-added decoding feature vector into the denoising encoder to perform noise reduction encoding to obtain a denoising encoding feature vector, and input the denoising encoding feature vector into the second decoder in the decoding layer to perform noise reduction decoding to obtain a denoising encoding feature vector.
[0113] Input the denoising encoding feature vector into the prediction layer to obtain an entity relationship graph.
[0114] Optionally, the clustering module includes: a fusion network and a clustering network; correspondingly, the clustering module is specifically configured to:
[0115] Fuse the entity relationship graphs output by each convolutional trapezoidal network in the relationship encoding module through the fusion network to obtain an entity relationship similarity graph, and input the entity relationship similarity graph into the clustering network;
[0116] Cluster the sentence samples in the entity relationship similarity graph whose entity relationship similarity is greater than a preset confidence level through the clustering network to obtain at least one first entity relationship cluster.
[0117] Traceable, the adjustment module is specifically used for:
[0118] Adjust the network parameters in the initial entity relationship extraction model based on the loss function value;
[0119] Determine the predicted label of the sentence sample as the pseudo-label of the sentence sample;
[0120] Return to execute the step of updating the sentence sample set according to the sentence sample with the pseudo-label, and input the updated sentence sample set into the clustering module;
[0121] Until the loss function value is the minimum value, determine the initial entity relationship extraction model corresponding to the loss function value as the target entity relationship extraction model.
[0122] The training device of the entity relationship extraction model provided by the embodiments of the present invention can execute the training method of the entity relationship extraction model provided by any embodiment of the present invention, and has the corresponding functional modules and beneficial effects for executing the method.
[0123] Embodiment 4
[0124] Figure 4 It is a schematic structural diagram of an entity relationship extraction device provided by Embodiment 4 of the present invention. As Figure 4 shown, the device includes: a sentence acquisition module 410, an input module 420, and a result acquisition module 430;
[0125] The sentence acquisition module 410 is used to acquire sentences to be extracted;
[0126] The input module 420 is used to input the sentences to be extracted into the target entity relationship extraction model trained by using the training method of the entity relationship extraction model according to any one of claims 1-6;
[0127] The result acquisition module 430 is used to acquire the entity relationship extraction result of the sentences to be extracted output by the target entity relationship extraction model.
[0128] The entity relationship extraction training device provided by the embodiments of the present invention can execute the entity relationship extraction method provided by any embodiment of the present invention, and has the corresponding functional modules and beneficial effects for executing the method.
[0129] Example 5
[0130] Figure 5 FIG. shows a schematic structural diagram of an electronic device 10 that can be used to implement the embodiments of the present invention. The electronic device is intended to represent various forms of digital computers, such as, for example, laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, for example, personal digital assistants, cellular phones, smart phones, wearable devices (such as helmets, glasses, watches, etc.) and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely exemplary and are not intended to limit the implementation of the present invention described and / or claimed herein.
[0131] As Figure 5 shown, the electronic device 10 includes at least one processor 11, and a memory communicatively connected to the at least one processor 11, such as a read-only memory (ROM) 12, a random access memory (RAM) 13, etc., wherein the memory stores a computer program executable by the at least one processor. The processor 11 can perform various appropriate actions and processes according to the computer program stored in the read-only memory (ROM) 12 or the computer program loaded from the storage unit 18 into the random access memory (RAM) 13. In the RAM 13, various programs and data required for the operation of the electronic device 10 can also be stored. The processor 11, the ROM 12, and the RAM 13 are connected to each other via a bus 14. The input / output (I / O) interface 15 is also connected to the bus 14.
[0132] A plurality of components in the electronic device 10 are connected to the I / O interface 15, including: an input unit 16, such as a keyboard, a mouse, etc.; an output unit 17, such as various types of displays, speakers, etc.; a storage unit 18, such as a magnetic disk, an optical disk, etc.; and a communication unit 19, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 19 allows the electronic device 10 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0133] The processor 11 can be various general-purpose and / or special-purpose processing components having processing and computing capabilities. Some examples of the processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The processor 11 executes the various methods and processes described above, such as the training method or the entity relationship extraction method of the entity relationship extraction model.
[0134] In some embodiments, the training method or the entity relationship extraction method of the entity relationship extraction model can be implemented as a computer program, which is tangibly included in a computer-readable storage medium, such as storage unit 18. In some embodiments, part or all of the computer program can be loaded and / or installed onto the electronic device 10 via the ROM 12 and / or the communication unit 19. When the computer program is loaded into the RAM 13 and executed by the processor 11, one or more steps of the training method or the entity relationship extraction method described above can be executed. Alternatively, in other embodiments, the processor 11 can be configured to execute the training method or the entity relationship extraction method by any other suitable means (e.g., by means of firmware).
[0135] The various embodiments of the systems and techniques described above in this document can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a dedicated or general-purpose programmable processor that receives data and instructions from a storage system, at least one input device, and at least one output device, and transmits the data and instructions to the storage system, the at least one input device, and the at least one output device.
[0136] The computer programs for implementing the methods of the present invention can be written in any combination of one or more programming languages. These computer programs can be provided to the processor of a general-purpose computer, a dedicated computer, or other programmable data processing devices, such that when the computer programs are executed by the processor, the functions / operations specified in the flowcharts and / or block diagrams are implemented. The computer programs can be executed entirely on the machine, partially on the machine, as an independent software package partially on the machine and partially on a remote machine, or entirely on a remote machine or server.
[0137] In the context of the present invention, a computer-readable storage medium can be a tangible medium that can contain or store a computer program for use by or in connection with an instruction execution system, apparatus, or device. The computer-readable storage medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. Alternatively, the computer-readable storage medium can be a machine-readable signal medium. More specific examples of the machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0138] To provide for interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the electronic device. Other kinds of devices can also be used to provide for interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, speech input, or tactile input).
[0139] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: local area network (LAN), wide area network (WAN), blockchain network, and the Internet.
[0140] A computing system may include a client and a server. The client and the server are generally far from each other and usually interact via a communication network. The client-server relationship is created by computer programs running on respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or a cloud host, which is a host product in the cloud computing service system, and solves the defects of difficult management and weak business scalability existing in traditional physical hosts and VPS services.
[0141] It should be understood that various forms of the processes shown above can be used, steps can be reordered, added, or deleted. For example, the steps recited in the present invention can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solution of the present invention can be achieved, and no limitation is made herein.
[0142] The above specific embodiments do not constitute a limitation to the protection scope of the present invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A training method for an entity relationship extraction model, characterized in that, it includes: Inputting a sentence sample set into an initial entity relationship extraction model, where the relationship extraction model includes: a pre-trained relationship encoding module and a clustering module; the sentence sample set is composed of unlabeled sentence samples; Performing entity relationship prediction on the sentence samples in the sentence sample set through the relationship encoding module to obtain an entity relationship graph, and inputting the entity relationship graph into the clustering module; Clustering through the clustering module based on the similarity between entities in the entity relationship graph to obtain at least one first entity relationship cluster and the pseudo-labels of the sentence samples included in the first entity relationship cluster; where the sentence samples included in the same first entity relationship cluster are marked with the same pseudo-labels; Updating the sentence sample set according to the sentence samples with pseudo-labels, inputting the updated sentence sample set into the initial entity relationship extraction model to obtain at least one second entity relationship cluster and the predicted labels of the sentence samples included in the second entity relationship cluster; Calculating a loss function value according to the pseudo-labels and predicted labels corresponding to the sentence samples, and iteratively adjusting the network parameters in the initial entity relationship extraction model based on the loss function value to obtain a target entity relationship extraction model; Iteratively adjusting the network parameters in the initial entity relationship extraction model based on the loss function value to obtain a target entity relationship extraction model, including: Adjusting the network parameters in the initial entity relationship extraction model based on the loss function value; Determining the predicted labels of the sentence samples as the pseudo-labels of the sentence samples; Returning to execute the step of updating the sentence sample set according to the sentence samples with pseudo-labels and inputting the updated sentence sample set into the initial entity relationship extraction model to obtain at least one second entity relationship cluster; Until the loss function value is the minimum value, determining the initial entity relationship extraction model corresponding to the loss function value as the target entity relationship extraction model.
2. The method according to claim 1, characterized in that, The relationship encoding module includes: a preset number of parallel convolutional trapezoidal networks, and each convolutional trapezoidal network includes: an embedding layer, an encoding layer, a decoding layer and a prediction layer; the encoding layer includes: a noise-adding encoder and a denoising encoder; the decoding layer includes a first decoder and a second decoder.
3. The method according to claim 2, characterized in that, Performing entity relationship prediction on the sentence samples in the sentence sample set through the relationship encoding module to obtain an entity relationship graph, including: For each convolutional trapezoidal network, extracting features from the input sentence sample through the embedding layer to obtain an embedding vector, and inputting the embedding vector into the noise-adding encoder; Performing noise-added encoding on the embedding vector through the noise-adding encoder to obtain a noise-added encoded feature vector; inputting the noise-added encoded feature vector into the first decoder in the decoding layer for noise-added decoding to obtain a noise-added decoded feature vector; Input the noisy decoded feature vector into the denoising encoder for denoising encoding to obtain a denoised encoded feature vector, and input the denoised encoded feature vector into the second decoder in the decoding layer for denoising decoding to obtain a denoised encoded feature vector; Input the denoised encoded feature vector into the prediction layer to obtain an entity relationship graph.
4. The method according to claim 2, wherein, the clustering module includes: a fusion network and a clustering network; correspondingly, at least one first entity relationship cluster is obtained by clustering based on the similarity between entities in the entity relationship graph through the clustering module, including: fusing the entity relationship graphs output by each convolutional trapezoidal network in the relationship encoding module through the fusion network to obtain an entity relationship similarity graph, and inputting the entity relationship similarity graph into the clustering network; clustering the sentence samples with the similarity of entity relationships in the entity relationship similarity graph greater than a preset confidence level through the clustering network to obtain at least one first entity relationship cluster.
5. An entity relationship extraction method, wherein, includes: obtaining a sentence to be extracted; inputting the sentence to be extracted into a target entity relationship extraction model trained by using the training method of the entity relationship extraction model according to any one of claims 1-4; obtaining the entity relationship category of the sentence to be extracted output by the target entity relationship extraction model.
6. A training device for an entity relationship extraction model, wherein, includes: an input module, configured to input a sentence sample set into an initial entity relationship extraction model, wherein the relationship extraction model includes: a pre-trained relationship encoding module and a clustering module; the sentence sample set is composed of unlabeled sentence samples; a relationship encoding module, configured to perform entity relationship prediction on the sentence samples in the sentence sample set to obtain an entity relationship graph, and input the entity relationship graph into the clustering module; a clustering module, configured to cluster based on the similarity between entities in the entity relationship graph to obtain at least one first entity relationship cluster, and the pseudo-labels of the sentence samples included in the first entity relationship cluster; wherein, the sentence samples included in the same entity relationship cluster are labeled with the same pseudo-label; an update module, configured to update the sentence sample set according to the sentence samples with pseudo-labels, input the updated sentence sample set into the initial entity relationship extraction model to obtain at least one second entity relationship cluster, and the prediction labels of the sentence samples included in the second entity relationship cluster; an adjustment module, configured to calculate a loss function value according to the pseudo-labels and prediction labels corresponding to the sentence samples, and iteratively adjust the network parameters in the initial entity relationship extraction model based on the loss function value to obtain a target entity relationship extraction model; the adjustment module is specifically configured to: adjust the network parameters in the initial entity relationship extraction model based on the loss function value; determine the prediction label of the sentence sample as the pseudo-label of the sentence sample; return to execute the step of updating the sentence sample set according to the sentence samples with pseudo-labels, and inputting the updated sentence sample set into the clustering module; When the loss function value reaches the minimum value, the initial entity relation extraction model corresponding to the loss function value is determined as the target entity relation extraction model.
7. An entity relation extraction device Characterized in that Comprising: A sentence acquisition module, configured to acquire a sentence to be extracted; An input module, configured to input the sentence to be extracted into a target entity relation extraction model trained by using the training method of the entity relation extraction model according to any one of claims 1-4; A result acquisition module, configured to acquire the entity relation extraction result of the sentence to be extracted output by the target entity relation extraction model.
8. An electronic device Characterized in that The electronic device includes: At least one processor; and A memory communicatively connected to the at least one processor; wherein, The memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute the training method of the entity relation extraction model according to any one of claims 1-4, or the entity relation extraction method according to claim 5.
9. A computer-readable storage medium Characterized in that The computer-readable storage medium stores computer instructions, and the computer instructions are used to implement the training method of the entity relation extraction model according to any one of claims 1-4, or the entity relation extraction method according to claim 5 when executed by a processor.
Citation Information
Patent Citations
False news identification method based on heterogeneous graph contrast learning
CN114020928A
Open relationship extraction method and device and storage medium
CN115204142A