Relation Extraction Method, Device and Storage Medium

By integrating the current representative target text and prior knowledge relationships in the relationship extraction model and using the memory network for memory reproduction, the problem of the relationship extraction model forgetting the old relationship in continuous learning is solved, and the continuous stability and accuracy of the relationship between entities is achieved.

CN115510852BActive Publication Date: 2025-07-29HARBIN INST OF TECH SHENZHEN GRADUATE SCHOOL
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202210989377.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-08-17
Publication Date
2025-07-29
Estimated Expiration
2042-08-17

AI Technical Summary

Technical Problem

In the field of continuous learning, the relationship extraction model is prone to forgetting the old relationship or changing the old relationship when extracting new relationships, resulting in inaccurate prediction of the old relationship.

Method used

By obtaining the target text of the current task, using the relationship extraction model to perform relationship prediction, determining the current representative target text, and integrating it with the prior knowledge relationship, using the memory network to memory reproduce and adjust the network parameters of the relationship extraction model to maintain the continuous stability of the relationship extraction.

Benefits of technology

The accuracy of the relationship between entities in the continuous learning process of the relationship extraction model is improved, and the stability and accuracy of relationship extraction is avoided.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115510852B_ABST
    Figure CN115510852B_ABST
Patent Text Reader

Abstract

The present application discloses a relation extraction method, device, and storage medium. The relation extraction method includes: using a relation extraction model to perform relation prediction on the target text of the current task to obtain a first target relation corresponding to each target text, and determining the current representative target text based on the first target relation; fusing the current representative target text, the corresponding first target relation, and the prior knowledge relation to obtain a current relation prototype corresponding to the prior knowledge relation; using the relation extraction model to perform memory reproduction on the current representative target text and the historical representative target text to obtain a second target relation corresponding to the current representative target text and the historical representative target text; and finally, adjusting the network parameters of the relation extraction model based on the second target relation. Through the above method, the accuracy of the relation between entities obtained by using the relation extraction model for relation extraction can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of natural language processing, and particularly to a relationship extraction method, device, and storage medium. Background Art

[0002] Relationship extraction is an important technical field in natural language processing tasks, and its purpose is to identify the relationships between entities from text. However, in the current field of continuous learning, relationship extraction is carried out continuously, and there will be problems of forgetting old relationships when extracting new relationships, or changing old relationships when extracting new relationships, resulting in inaccurate prediction of old relationships. Summary of the Invention

[0003] This application provides a relationship extraction method, device, and storage medium, which can improve the accuracy of the relationships between entities obtained by using a relationship extraction model for relationship extraction.

[0004] To solve the above technical problems, a technical solution adopted by this application is: obtain a current task, where the current task includes multiple target texts, and each target text includes a corresponding first entity and a second entity; use a relationship extraction model to perform relationship prediction on each target text to obtain a first target relationship corresponding to each target text; use the first target relationship corresponding to each target text to determine a current representative target text from the multiple target texts; use the current representative target text, the target relationship corresponding to the current representative target text, and prior knowledge relationships for fusion to obtain a current relationship prototype corresponding to each prior knowledge relationship; based on the current relationship prototype and the historical relationship prototype, use the memory network in the relationship extraction model to perform memory reproduction on the current representative target text and the historical representative target text to obtain a second target relationship corresponding to the current representative target text and the historical representative target text, and the historical relationship prototype and the historical representative target text are obtained based on a historical task; adjust the network parameters of the relationship extraction model based on the second target relationship.

[0005] Among them, using the first target relationship corresponding to each target text to determine a current representative target text from the multiple target texts includes: classifying the multiple target texts according to the first target relationship to obtain a target text set corresponding to each first target relationship, and the target text set includes at least one target text; using a clustering algorithm to process the target texts in each target text set to obtain a current representative target text.

[0006] Among them, fusing the current representative target text, the target relationship corresponding to the current representative target text, and the prior knowledge relationship to obtain the current relationship prototype corresponding to each prior knowledge relationship includes: encoding the current representative target text and the target relationship corresponding to the current representative target text to obtain a first encoded feature; encoding the prior knowledge relationship to obtain a second encoded feature; fusing the first encoded feature and the second encoded feature to obtain the current relationship prototype feature corresponding to each prior knowledge relationship.

[0007] Among them, after obtaining the current relationship prototype feature corresponding to each prior knowledge relationship, the gradient of the relationship extraction model is not updated.

[0008] Among them, based on the current relationship prototype and the historical relationship prototype, using the memory network in the relationship extraction model to perform memory reproduction on the current representative target text and the historical representative target text to obtain the second target relationship corresponding to the current representative target text and the historical representative target text, including: encoding the current representative target text and the first target relationship, and the historical representative target text and the historical target relationship to obtain a third encoded feature; based on the current relationship prototype and the historical relationship prototype, using the memory network to perform memory reproduction on the third encoded feature to refine the third encoded feature and obtain a fourth encoded feature; using the fourth encoded feature to obtain the second target relationship corresponding to the current representative target text and the historical representative target text.

[0009] Among them, before using the fourth encoded feature to obtain the second target relationship corresponding to the current representative target text and the historical representative target text, it includes: fusing the fourth encoded feature and the third encoded feature to obtain a target encoded feature, and using the target encoded feature as the fourth encoded feature.

[0010] Among them, after using the memory network to perform memory reproduction on the third encoded feature based on the current relationship prototype and the historical relationship prototype to refine the third encoded feature and obtain the fourth encoded feature, it includes: calculating the gating value of the fourth encoded feature using a gating module to obtain the corresponding gating value.

[0011] Among them, adjusting the network parameters of the relationship extraction model based on the second target relationship includes: adjusting the network parameters of the relationship extraction model based on the second target relationship and the gating value.

[0012] Among them, after obtaining the current task, it includes: constructing a prompt format for each target text to generate the corresponding target text pair.

[0013] Among them, a relationship extraction model is used to predict the relationship of each target text, and a first target relationship corresponding to each target text is obtained, including: using the relationship extraction model to predict the relationship of each target text pair, and obtaining a first target relationship corresponding to each target text pair; writing the first target relationship, the first entity and the second entity corresponding to the first target relationship into the target text pair according to the prompt format.

[0014] To solve the above technical problems, another technical solution adopted by this application is: to provide a relationship extraction device, which includes a memory and a processor. Among them, the memory is used to store program data, and the processor is used to execute the program data to implement the above relationship extraction method.

[0015] To solve the above technical problems, another technical solution adopted by this application is: to provide a computer-readable storage medium, in which program data is stored, and when the program data is executed by a processor, it is used to execute the above relationship extraction method.

[0016] The beneficial effect of this application is: different from the prior art, the relationship extraction method provided by this application obtains the current task, where the current task includes multiple target texts, and each target text includes a corresponding first entity and a second entity; uses the relationship extraction model to predict the relationship of each target text, and obtains a first target relationship corresponding to each target text; determines the current representative target text from multiple target texts using the first target relationship corresponding to each target text; fuses the current representative target text, the target relationship corresponding to the current representative target text, and the prior knowledge relationship to obtain a current relationship prototype corresponding to each prior knowledge relationship; based on the current relationship prototype and the historical relationship prototype, uses the memory network in the relationship extraction model to perform memory reproduction on the current representative target text and the historical representative target text, and obtains a second target relationship corresponding to the current representative target text and the historical representative target text, where the historical relationship prototype and the historical representative target text are obtained based on the historical task; finally, adjusts the network parameters of the relationship extraction model based on the second target relationship. In the above manner, the relationship extraction model is continuously adjusted through memory reproduction, enabling the relationship extraction model to continuously learn the relationships between entities in previous tasks, so that when performing relationship extraction on subsequent tasks, it will not forget the relationships between entities obtained from historical extraction tasks, maintaining the continuous stability of relationship extraction, and being able to improve the accuracy of the relationships between entities obtained by using the relationship extraction model for relationship extraction. Description of the Drawings

[0017] To more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts. Among them:

[0018] Figure 1 is a schematic flowchart of the first embodiment of the relationship extraction method provided by the present application;

[0019] Figure 2 is a schematic flowchart of an embodiment of step 13 provided by the present application;

[0020] Figure 3 is a schematic flowchart of an embodiment of step 14 provided by the present application;

[0021] Figure 4 is a schematic flowchart of an embodiment of step 15 provided by the present application;

[0022] Figure 5 is a schematic flowchart of the second embodiment of the relationship extraction method provided by the present application;

[0023] Figure 6 is a schematic flowchart of an embodiment of step 53 provided by the present application;

[0024] Figure 7 is a schematic structural diagram of an embodiment of the relationship extraction device provided by the present application;

[0025] Figure 8 is a schematic structural diagram of an embodiment of the computer-readable storage medium provided by the present application. Detailed implementation manners

[0026] The following will clearly and completely describe the technical solutions in the embodiments of the present application with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only some embodiments of the present application, rather than all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present application.

[0027] Refer to Figure 1 , Figure 1 which is a schematic flowchart of the first embodiment of the relationship extraction method provided by the present application. The method includes:

[0028] Step 11: Obtain the current task, where the current task includes multiple target texts, and each target text includes a corresponding first entity and a second entity.

[0029] Specifically, each task contains multiple target texts.

[0030] In some embodiments, the target text is unstructured text.

[0031] Optionally, the target text includes at least one or more of txt text, word text, and pdf text, which are unstructured texts.

[0032] Optionally, the first entity and the second entity can be a place name, a person name, an item name, etc.

[0033] For example, if the target text is "Henry was born in London", then the first entity can be "Henry" or "London". After determining the first entity, the remaining noun corresponds to the second entity.

[0034] In some embodiments, according to the order of language habits, by default, the first noun is the first entity and the following noun is the second entity, that is, "Henry" is the first entity and "London" is the second entity.

[0035] Step 12: Use the relation extraction model to perform relation prediction on each target text to obtain the first target relation corresponding to each target text.

[0036] In some embodiments, the relation extraction model is the BERT model. The BERT (Bidirectional Encoder Representations from Transformer) model is a bidirectional encoder representation based on Transformer and is a pre-trained language representation model. It adopts a new masked language model (MLM) to generate deep bidirectional language representations. The goal of the BERT model is to train with a large-scale unlabeled corpus to obtain the Representation of the text containing rich semantic information, that is, the semantic representation of the text, and then fine-tune the semantic representation of the text in a specific NLP (Natural Language Processing) task, and finally apply it to this NLP task.

[0037] In some embodiments, to perform relation prediction on the target text, it is first necessary to determine the first entity and the second entity in the target text, and then use the relation extraction model to extract / predict the target relation between the first entity and the second entity.

[0038] In some embodiments, the relation between the first entity and the second entity is represented by a relation character.

[0039] Specifically, the target relation of each target text is the relation between the first entity and the second entity. Depending on the differences between the first entity and the second entity, the target relation between the two is also different.

[0040] For example, if the target text is "Henry was born in London", where the first entity is "Henry" and the second entity is "London", then the relationship between the first entity and the second entity is "was born in", and the target relationship "was born in" can be classified into the category of the relationship character "place of birth relationship" (the first target relationship).

[0041] For example, if the target text is "Li Hua is Xiao An's mother", where the first entity is "Li Hua" and the second entity is "Xiao An", then the relationship between the first entity and the second entity is "is...'s mother" or "is...'s daughter", and the relationships "is...'s mother" and "is...'s daughter" can be classified into the category of the relationship characters "parent-child relationship" or "blood relationship" (the first target relationship).

[0042] For example, if the target text is "Li Hua founded the English League", where the first entity is "Li Hua" and the second entity is "English League", then the relationship between the first entity and the second entity is "founded", and the relationship "founded" can be classified into the category of the relationship character "founding relationship" (the first target relationship).

[0043] Step 13: Determine the current representative target text from multiple target texts using the first target relationship corresponding to each target text.

[0044] Specifically, each task contains multiple target texts, and a target text has at least one first target relationship. In other words, a task contains multiple first target relationships.

[0045] Therefore, it is necessary to determine the representative target text from multiple first target relationships to reduce the number of texts participating in subsequent operations.

[0046] In some embodiments, referring to Figure 2 , Step 13 may be the following process:

[0047] Step 21: Classify multiple target texts according to the first target relationship to obtain a set of target texts corresponding to each first target relationship, where the set of target texts includes at least one target text.

[0048] In some embodiments, different relationship characters can be used as the division criteria to classify and group different first target relationships, so as to achieve the classification and grouping of target texts.

[0049] For example, target texts such as "Henry was born in London" and "Xiao An was born in Chengdu" can both be classified into the relationship category of "place of birth relationship". That is to say, the set of target texts with the first target relationship of "place of birth relationship" includes "Henry was born in London" and "Xiao An was born in Chengdu".

[0050] For example, the target texts "Li Hua founded the English League", "Xiao An established the Chinese learning group", and "Henry organized the Math team" can all be classified into the relationship category of "founding relationship". That is to say, the set of target texts with the first target relationship being "founding relationship" includes "Li Hua founded the English League", "Xiao An established the Chinese learning group", and "Henry organized the Math team".

[0051] It should be noted that the relationship between the first entity and the second entity is represented by relationship characters. The same target relationship can include multiple different relationship characters. For example, "founded", "established", and "organized" can all represent the first target relationship of "founding relationship" within a specific range.

[0052] Step 22: Use the clustering algorithm to process the target texts in each target text set to obtain the current representative target text.

[0053] Optionally, the clustering algorithm includes prototype clustering algorithm, density clustering algorithm, hierarchical clustering algorithm, grid model clustering algorithm, model clustering algorithm, and spectral clustering algorithm. Among them, the prototype clustering algorithm includes K-means algorithm, Learning Vector Quantization (LVQ) algorithm, and K-Nearest Neighbor (KNN) algorithm; the density clustering algorithm is DBSCAN algorithm; the hierarchical clustering algorithm is AGNES algorithm; the model clustering algorithm is Gaussian Mixture Clustering (EM) algorithm.

[0054] Specifically, in one embodiment, the clustering algorithm used is the K-means algorithm. The K-means algorithm involves a clustering principle, which means clustering with k points in space as the centers, classifying the objects closest to them, and successively calculating the values of the cluster centers as the new center values, and iteratively updating until the cluster center positions no longer change or reach the maximum number of iterations.

[0055] In some embodiments, use the K-means algorithm to process the target texts in the target text set to obtain the representative target text corresponding to the target text set.

[0056] Step 14: Integrate the current representative target text, the target relationship corresponding to the current representative target text, and the prior knowledge relationship to obtain the current relationship prototype corresponding to each prior knowledge relationship.

[0057] In some embodiments, use the multi-head attention mechanism to perform feature integration on the current representative target text, the target relationship corresponding to the current representative target text, and the prior knowledge relationship.

[0058] Specifically, the multi-head attention mechanism is a weighting mechanism for input features. The input features Query (query), Key (key), and Value (value) (in the Self-Attention mechanism, Query = Key = Value) first undergo a linear mapping and are then divided into N segments (N heads) for combining features to parallelly calculate and select multiple pieces of information from the input information. Each attention focuses on different parts of the input information and then they are concatenated. The formula for the multi-head attention mechanism is as follows:

[0059]

[0060] Among them, Attention(Q, K, V) is the obtained attention value, Q, K, and V are the query volume (Query), key (Key), and value (Value) respectively, softmax is the normalization function, and d k is the dimension of K.

[0061] Specifically, the prior knowledge relationship refers to the relationship between the type corresponding to the first entity and the type corresponding to the second entity that is known in advance. For example, if the first entity is Henry and the second entity is London, Henry corresponds to a person and London corresponds to a location, then the prior knowledge relationship is the relationship between a person and a location. The prior knowledge relationship can be regarded as a prompt for the relationship extraction model to improve the accuracy of relationship extraction.

[0062] In some embodiments, referring to Figure 3 , step 14 may be the following process:

[0063] Step 31: Encode the current representative target text and the target relationship corresponding to the current representative target text to obtain a first encoded feature.

[0064] In some embodiments, by using a relationship extraction model (such as the BERT model) to encode the current representative target text and the target relationship corresponding to the current representative target text, a first encoded feature can be obtained.

[0065] In some embodiments, the first encoded feature can be used as the input in the multi-head attention mechanism of the relationship extraction model. Specifically, the first encoded feature is used as the key (Key) and value (Value) in the multi-head attention mechanism, and Key and Value are the same.

[0066] Step 32: Encode the prior knowledge relationship to obtain a second encoded feature.

[0067] Specifically, the prior knowledge relationship is the literal definition between the type corresponding to the first entity and the type corresponding to the second entity.

[0068] In some embodiments, by using a relation extraction model (such as a BERT model) to encode prior knowledge relations, second encoded features can be obtained.

[0069] In some embodiments, the second encoded features can be used as the input in the multi-head attention mechanism of the relation extraction model. Specifically, the second encoded features are used as the query in the multi-head attention mechanism.

[0070] Step 33: Fuse the first encoded features and the second encoded features to obtain the current relation prototype features corresponding to each prior knowledge relation.

[0071] In some embodiments, use the multi-head attention mechanism to perform feature fusion on the first encoded features and the second encoded features to obtain the current relation prototype features corresponding to the prior knowledge relations.

[0072] It can be represented by the following formula:

[0073] P r = MultiHeadATT(Query, Key, Value)

[0074] where MultiHeadATT(*) represents the attention function, Query represents the first encoded features, Key and Value represent the second encoded features, and P r represents the current relation prototype features.

[0075] It should be noted that after obtaining the current relation prototype features corresponding to each prior knowledge relation, the relation extraction model is not updated with gradients.

[0076] Step 15: Based on the current relation prototype and the historical relation prototype, use the memory network in the relation extraction model to perform memory reproduction on the current representative target text and the historical representative target text to obtain the second target relations corresponding to the current representative target text and the historical representative target text, where the historical relation prototype and the historical representative target text are obtained based on historical tasks.

[0077] It can be understood that the memory network (Memory Networks, abbreviated as MemNN) is a network structure based on the multi-head attention mechanism, used to combine the encoded replay samples and the relation prototype features. In some embodiments, the memory network can be a long short-term memory network or a bidirectional long short-term memory network.

[0078] In some embodiments, referring to Figure 4 , Step 15 can be the following process:

[0079] Step 41: Encode the current representative target text and the first target relationship, as well as the historical representative target text and the historical target relationship, to obtain a third encoded feature.

[0080] In some embodiments, a multi-head attention mechanism is used to encode the current representative target text and the first target relationship, as well as the historical representative target text and the historical target relationship, to obtain a third encoded feature.

[0081] Step 42: Based on the current relationship prototype and the historical relationship prototype, use a memory network to perform memory reproduction on the third encoded feature to refine the third encoded feature and obtain a fourth encoded feature.

[0082] In some embodiments, the memory network for performing memory reproduction on the third encoded feature is a one-layer multi-head attention network.

[0083] Since the fourth encoded feature is refined from the third encoded feature based on the current relationship prototype and the historical relationship prototype, the fourth encoded feature has more relationship information between entities possessed by the current relationship prototype and the historical relationship prototype.

[0084] Step 43: Use a gating module to calculate the gating value for the fourth encoded feature to obtain the corresponding gating value.

[0085] In some embodiments, there is a problem of feature distribution shift in refining the third encoded feature. To avoid the imbalance between data plasticity and stability caused by feature distribution shift, a gating module is used to calculate the gating value for the fourth encoded feature to obtain the corresponding gating value, and the loss function of the data is controlled based on the corresponding gating value. Among them, the gating value can be used as the weight of the corresponding feature in the loss function.

[0086] Specifically, the calculation formula for the gating value is:

[0087]

[0088]

[0089] where \(b(x,r)\) represents the weight value of the loss function, \(R\) k represents the prototype set of the \(k\)th task, that is, the relationship prototype set of the current task, represents the prototype set of all tasks including the \(k\)th task, that is, the prototype sets of the current task and the historical tasks, \(\gamma\) is a parameter greater than 1, represents the encoded feature corresponding to the \(x\)th text, that is, the above-mentioned fourth encoded feature, represents the relationship prototype corresponding to the \(x\)th text, represents the remaining relationship prototypes.

[0090] Step 44: Fuse the fourth encoded feature and the third encoded feature to obtain a target encoded feature, and use the target encoded feature as the fourth encoded feature.

[0091] In some embodiments, a residual connection is performed on the fourth encoded feature and the third encoded feature, that is, a residual network is introduced, so as to realize the feature fusion of the third encoded feature and the fourth encoded feature, and obtain a target encoded feature.

[0092] It should be noted that the residual network can be easily implemented with mainstream automatic differentiation deep learning frameworks, and the residual network can directly update parameters using the BP algorithm (backpropagation). That is to say, the network parameters of the relation extraction model can be updated through residual connections.

[0093] Step 45: Use the fourth encoded feature to obtain the second target relation corresponding to the current representative target text and the historical representative target text.

[0094] It can be understood that in continual learning, continual relation extraction (CRE) needs to be performed. Continual relation extraction means training multiple regular relation extraction tasks in sequence according to time order. Among them, each relation extraction task has an independent training set and test set. Generally, it is assumed that the relation types between tasks are different from each other. When all tasks are trained, the model needs to be tested on the test sets and the test set union of each task. An excellent continual learning model should have the ability to recognize accurate relations on the test set of any one task.

[0095] It can be understood that the continual relation extraction method is used to continuously learn newly emerging relations and at the same time overcome the catastrophic forgetting problem faced in continual learning, and is usually completed in an incremental training framework.

[0096] In some embodiments, the determination of the second target relation does not change the first target relation. That is to say, in continual relation extraction, the latter target relation does not affect the accuracy of the former target relation, and the former target relation is not forgotten when obtaining the latter target relation.

[0097] Step 16: Adjust the network parameters of the relation extraction model based on the second target relation.

[0098] In some embodiments, the network parameters of the relation extraction model are adjusted based on the second target relation and the gating value.

[0099] Different from the prior art, the relation extraction method provided by this application continuously adjusts the relation extraction model by means of memory reproduction, enabling the relation extraction model to continuously learn the relations between entities in the past. Consequently, when performing relation extraction on subsequent tasks, it will not forget the relations between entities obtained from historical extraction tasks, maintaining the continuous stability of relation extraction and enhancing the accuracy of the relations between entities obtained by using the relation extraction model for relation extraction.

[0100] Referring to Figure 5 , Figure 5 which is a schematic flowchart of the second embodiment of the relation extraction method provided by this application. The method includes:

[0101] Step 51: Obtain the current task, where the current task includes multiple target texts, and each target text includes a corresponding first entity and a second entity.

[0102] Step 52: Construct a prompt format for each target text to generate a corresponding target text pair.

[0103] In some embodiments, the prompt format is "First entity? [MASK], Second entity", where "[MASK]" is the relation character to be predicted. In other words, the prompt format is "First entity - relation character - Second entity".

[0104] For example, if the target text is "Henry was born in London", the corresponding prompt format is "Henry? [MASK], London", then the target text pair is "Henry? [MASK], London - Henry was born in London", that is, the target text pair consists of the target text pair and the corresponding prompt format.

[0105] Step 53: Use the relation extraction model to perform relation prediction on each target text to obtain a first target relation corresponding to each target text.

[0106] In some embodiments, referring to Figure 6 , Step 53 can be the following process:

[0107] Step 61: Use the relation extraction model to perform relation prediction on each target text pair to obtain a first target relation corresponding to each target text pair.

[0108] Step 62: Write the first target relation, as well as the first entity and the second entity corresponding to the first target relation, into the target text pair according to the prompt format.

[0109] In some embodiments, based on the relation extraction model, the first target relation, the first entity, and the second entity are written into the target text pair in the form of a cloze test.

[0110] For example, if the first target relationship is the "birthplace relationship", the first entity is "Henry", and the second entity is "London", then a cloze test is performed in the form of the prompt format "First entity? [MASK], second entity" to obtain "Henry was born in London".

[0111] Step 54: Determine the current representative target text from multiple target texts using the first target relationship corresponding to each target text.

[0112] Step 55: Integrate the current representative target text, the target relationship corresponding to the current representative target text, and the prior knowledge relationship to obtain the current relationship prototype corresponding to each prior knowledge relationship.

[0113] Step 56: Based on the current relationship prototype and the historical relationship prototype, use the memory network in the relationship extraction model to perform memory reproduction on the current representative target text and the historical representative target text to obtain the second target relationship corresponding to the current representative target text and the historical representative target text, where the historical relationship prototype and the historical representative target text are obtained based on historical tasks.

[0114] Step 57: Adjust the network parameters of the relationship extraction model based on the second target relationship.

[0115] Steps 54 to 57 may have the same or similar technical solutions as the above embodiments, and will not be elaborated here.

[0116] Different from the prior art, the relationship extraction method provided by this application obtains the current task, where the current task includes multiple target texts, and each target text includes a corresponding first entity and a second entity; uses the relationship extraction model to perform relationship prediction on each target text to obtain the first target relationship corresponding to each target text; determines the current representative target text from multiple target texts using the first target relationship corresponding to each target text; integrates the current representative target text, the target relationship corresponding to the current representative target text, and the prior knowledge relationship to obtain the current relationship prototype corresponding to each prior knowledge relationship; based on the current relationship prototype and the historical relationship prototype, uses the memory network in the relationship extraction model to perform memory reproduction on the current representative target text and the historical representative target text to obtain the second target relationship corresponding to the current representative target text and the historical representative target text, where the historical relationship prototype and the historical representative target text are obtained based on historical tasks; finally, adjusts the network parameters of the relationship extraction model based on the second target relationship.

[0117] In the above - mentioned manner, the relation extraction model is continuously adjusted through the way of memory reproduction, enabling the relation extraction model to continuously learn the relations between entities in the previous tasks. Consequently, when performing relation extraction on subsequent tasks, it will not forget the relations between entities obtained from historical extraction tasks, maintaining the continuous stability of relation extraction and improving the accuracy of the relations between entities obtained by using the relation extraction model for relation extraction.

[0118] Refer to Figure 7 , Figure 7 FIG. Figure 7 is a schematic structural diagram of an embodiment of the relation extraction device provided by the present application. The relation extraction device 70 includes a memory 701 and a processor 702. The memory 701 is used to store program data, and the processor is used to execute the program data to implement the relation extraction method in any of the above - mentioned embodiments, which will not be elaborated here.

[0119] Refer to Figure 8 , Figure 8 FIG. Figure 8 is a schematic structural diagram of an embodiment of the computer - readable storage medium provided by the present application. The computer - readable storage medium 80 stores program data 801. When the program data 801 is executed by the processor, it is used to implement the relation extraction method in any of the above - mentioned embodiments, which will not be elaborated here.

[0120] In summary, the present application can effectively handle the relation extraction problem with continuously increasing relation categories, providing a solution for text structured information extraction. Moreover, the present application can significantly reduce the forgetting speed during the continuous learning process and can provide reference value for other continuous learning - type problems. Additionally, the present application introduces prior - knowledge relations into the relation extraction task, enabling more accurate recognition ability in the relation recognition task.

[0121] The processor involved in the present application may be referred to as a CPU (Central Processing Unit). It may be an integrated circuit chip, and may also be a general - purpose processor, a digital signal processor (DSP), an application - specific integrated circuit (ASIC), a field - programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.

[0122] The storage media used in the present application include various media such as USB flash drives, mobile hard disks, read - only memories (ROM, Read - Only Memory), random - access memories (RAM, Random Access Memory), or optical discs that can store program codes.

[0123] The above are only the embodiments of the present application, and do not thus limit the patent scope of the present application. Any equivalent structure or equivalent process transformation made by using the content of the specification and drawings of the present application, or directly or indirectly applied in other related technical fields, shall similarly be included within the patent protection scope of the present application.

Claims

1. A relation extraction method, characterized in that, The method includes: Obtaining a current task, where the current task includes multiple target texts, and each target text includes a corresponding first entity and a second entity; Using a relation extraction model to perform relation prediction on each target text to obtain a first target relation corresponding to each target text; Determining a current representative target text from the multiple target texts by using the first target relation corresponding to each target text; Encoding the current representative target text and the target relation corresponding to the current representative target text to obtain a first encoded feature; Encoding prior knowledge relations to obtain a second encoded feature; Fusing the first encoded feature and the second encoded feature to obtain a current relation prototype feature corresponding to each prior knowledge relation; Encoding the current representative target text and the first target relation, as well as a historical representative target text and a historical target relation, to obtain a third encoded feature; Based on the current relation prototype and the historical relation prototype, using a memory network to perform memory reproduction on the third encoded feature to refine the third encoded feature and obtain a fourth encoded feature; Using the fourth encoded feature to obtain a second target relation corresponding to the current representative target text and the historical representative target text; the historical relation prototype and the historical representative target text are obtained based on a historical task; Adjusting the network parameters of the relation extraction model based on the second target relation.

2. The method according to claim 1, characterized in that, The determining a current representative target text from the multiple target texts by using the first target relation corresponding to each target text includes: Classifying the multiple target texts according to the first target relation to obtain a set of target texts corresponding to each first target relation; the set of target texts includes at least one of the target texts; Processing the target texts in each set of target texts by using a clustering algorithm to obtain the current representative target text.

3. The method according to claim 1, wherein After obtaining the current relation prototype feature corresponding to each prior knowledge relation, no gradient update is performed on the relation extraction model.

4. The method according to claim 1, characterized in that Before using the fourth encoded feature to obtain a second target relation corresponding to the current representative target text and the historical representative target text, it includes: Fusing the fourth encoded feature and the third encoded feature to obtain a target encoded feature, and using the target encoded feature as the fourth encoded feature.

5. The method according to claim 1, wherein After using the memory network to perform memory reproduction on the third encoded feature based on the current relation prototype and the historical relation prototype to refine the third encoded feature and obtain a fourth encoded feature, it includes: Calculating a gating value for the fourth encoded feature by using a gating module to obtain a corresponding gating value; The adjusting the network parameters of the relation extraction model based on the second target relation includes: Adjusting the network parameters of the relation extraction model based on the second target relation and the gating value.

6. The method according to claim 1, wherein After obtaining the current task, it includes: Constructing a prompt format for each target text to generate a corresponding target text pair; Performing relationship prediction on each target text by using the relationship extraction model to obtain a first target relationship corresponding to each target text, including: Performing relationship prediction on each of the target text pairs by using the relationship extraction model to obtain a first target relationship corresponding to each of the target text pairs; Writing the first target relationship, and the first entity and the second entity corresponding to the first target relationship into the target text pair according to the prompt format.

7. A relationship extraction device, characterized in that The relationship extraction device includes a memory and a processor, the memory stores program data, and the processor is configured to execute the program data to implement the method according to any one of claims 1-6.

8. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores program data, and the program data is used to implement the method according to any one of claims 1-6 when being executed by a processor.

Citation Information

Patent Citations

  • Text relation extraction method and device, equipment and storage medium

    CN114610903A