Small-sample relation extraction method under continuous learning based on prompt contrastive learning

Through the continuous learning method based on prompt comparison learning, the problems of knowledge forgetting and insufficient samples in the continuous small sample relationship extraction model are solved, and the accuracy of entity relationship extraction and the robustness of the model are improved.

CN116719934BActive Publication Date: 2025-05-13NAT UNIV OF DEFENSE TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310583107.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-05-23
Publication Date
2025-05-13
Estimated Expiration
2043-05-23

AI Technical Summary

Technical Problem

The existing continuous small sample relationship extraction model faces catastrophic forgetting problems and insufficient samples for new relationship training, resulting in low accuracy of entity relationship extraction.

Method used

The continuous learning method based on prompt comparison learning is adopted, and the enhanced sentence representation is obtained through information enhancement and template mapping functions. The sentence vector representation is mapped to the low-dimensional embedding space using pre-trained BERT model and projection head, and the total loss function is constructed and optimized. Combined with contrast learning and knowledge distillation methods, the model is optimized to reduce knowledge forgetting.

Benefits of technology

The accuracy of entity relationship extraction is improved, and by enhancing sample representation and optimizing model parameters, the limited training samples are effectively utilized, reducing the error of relationship extraction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116719934B_ABST
    Figure CN116719934B_ABST
Patent Text Reader

Abstract

The present application relates to a method for extracting small sample relations under continuous learning based on prompt contrastive learning. The method comprises: inputting prompt-based data into a pre-trained BERT model to obtain a vector representation of a sentence and mapping the vector representation of the sentence to a low-dimensional embedding space to obtain a hidden representation; normalizing the hidden representation to obtain a normalized embedding representation; constructing a total loss function of an initial CFRE model according to the probability distribution, hidden representation, normalized embedding representation and data enhancement strategy on a label set, training the initial CFRE model using a training data set and a total loss function, minimizing the loss function of the initial CFRE model to optimize the parameters in the model, optimizing the initial continuous small sample relation extraction model according to contrastive learning and knowledge distillation methods, and extracting relations using the continuous small sample relation extraction model. The use of this method can improve the accuracy of entity relation extraction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of data processing technology, and in particular to a method for extracting small sample relationships under continuous learning based on prompt contrastive learning. Background Art

[0002] Relation extraction (RE) plays a vital role in knowledge extraction and management, and can be used for a range of tasks in natural language processing, such as question answering, text understanding, etc. Given a sentence x with annotated entity pairs e1 and e2, the task of RE aims to classify it into a predefined relation R k The traditional RE model requires a predefined fixed relation set, which can only classify the relationship between entity pairs in a sentence into a category in the fixed relation set. However, since the real world is open, the emergence of new data leads to an increasing number of relations. In response to this situation, a continuous relation extraction paradigm is proposed.

[0003] In existing research, most CRE models assume that there is enough labeled data available for training in subsequent tasks. However, in the real world, most relations do not have enough labeled data, especially for newly emerged relations. Therefore, considering the high cost of obtaining enough labeled data, Qin et al. proposed continuous small-sample relation extraction, which requires the model to learn new relations from very few training examples while not forgetting old relations. Current CFRE models face two challenges: catastrophic forgetting of existing knowledge and insufficient training samples for newly emerged relations, resulting in low accuracy in entity relation extraction. Summary of the invention

[0004] Based on this, it is necessary to provide a small sample relationship extraction method under continuous learning based on prompt contrastive learning, which can improve the accuracy of entity relationship extraction in response to the above technical problems.

[0005] A method for extracting small sample relations under continuous learning based on prompt contrastive learning, the method comprising:

[0006] Obtain a task set and an initial CFRE model; the task set contains multiple tasks; each task contains its own training data set; the data set contains multiple training samples; the training samples include multiple sentences;

[0007] Performing information enhancement on the sentences in the task to obtain enhanced sentences; mapping the enhanced sentences according to a pre-built template mapping function to obtain prompt-based data;

[0008] Input the prompt-based data into the pre-trained BERT model to obtain the vector representation of the sentence, and calculate the probability distribution on the label set based on the vector representation of the sentence;

[0009] The projection head is used to map the vector representation of the sentence to a low-dimensional embedding space to obtain a hidden representation; the hidden representation is normalized to obtain a normalized embedding representation;

[0010] In each task, the total loss function of the initial CFRE model is constructed according to the probability distribution, hidden representation, normalized embedding representation and data augmentation strategy on the label set. The initial CFRE model is trained using the training dataset and the total loss function. The loss function of the initial CFRE model is minimized to optimize the parameters in the model and obtain the initial continuous small sample relationship extraction model.

[0011] The initial continuous small sample relationship extraction model is optimized according to contrastive learning and knowledge distillation methods to obtain a small sample relationship extraction model under continuous learning.

[0012] Relation extraction is performed using a small sample relation extraction model under continuous learning.

[0013] In one embodiment, information enhancement is performed on a sentence in a task to obtain an enhanced sentence, including:

[0014] The information of the sentence in the task is enhanced, and the enhanced sentence is

[0015] x aug ={w 1 ,v,[E 11 ],e 1 ,[E 12 ],...,[E 21 ],e 2 ,[E 22 ],...,w x}

[0016] Among them, w x Indicates a word in a sentence. [E11], [E12], [E21], and [E22] are all special tags. 1 and e 2 Represents the head entity and the tail entity in a sentence.

[0017] In one embodiment, the enhanced sentence is mapped according to a pre-built template mapping function to obtain prompt-based data, including:

[0018] The enhanced sentences are mapped according to the pre-built template mapping function, and the prompt-based data is obtained as

[0019] x prompt =T(x aug )=I think 1 is[MASK]of e2 [SEP]x aug

[0020] Where T(·) represents the template mapping function, [MASK] represents the task location, and [SEP] represents the sentence separator.

[0021] In one embodiment, the total loss function of the initial CFRE model includes a cross entropy loss function, a sentence feature mixed loss function, and a contrast loss function; in each task, the total loss function of the initial CFRE model is constructed according to the probability distribution, hidden representation, normalized embedding representation, and data enhancement strategy on the label set, including:

[0022] Construct a cross entropy loss function based on the probability distribution on the label set;

[0023] According to the data enhancement strategy, linear interpolation calculation is performed on the hidden representation and the class label corresponding to the hidden representation to obtain the interpolation result; the sentence feature mixed loss function is constructed using the interpolation result;

[0024] Construct a contrastive loss function based on the normalized embedding representation;

[0025] The total loss function of the initial CFRE model constructed using the cross entropy loss function, sentence feature mixed loss function and contrast loss function is:

[0026] L All =α 1 ·L CE +α 2 ·L SM +α 3 ·L CL

[0027] Among them, L CE represents the cross entropy loss function, L SM represents the sentence feature mixed loss function, L CL represents the contrast loss function, α 1 , α 2 , α 3 Both represent tuning parameters.

[0028] In one embodiment, a cross entropy loss function is constructed according to the probability distribution on the label set, including:

[0029] According to the probability distribution on the label set, the cross entropy loss function is constructed as follows:

[0030]

[0031] Among them, Θ represents the initial CFRE model, represents the training data set, k represents the task number, P(y=y i ∣zi ,y i ) represents the probability distribution on the label set, y represents the target label word, and y i represents the true label of sample i in the training data set, z i represents the normalized embedding representation, and i represents the sample number.

[0032] In one embodiment, the sentence feature hybrid loss function is constructed using the interpolation result, including:

[0033] The sentence feature mixed loss function is constructed using the interpolation results:

[0034]

[0035] Among them, v i Indicates the representation of different words obtained by BERT. Represents hidden representation, represents the mixed sentence features obtained by the prompt-based encoder, Represents the mixed label of mixed sentence features and the correct word representation.

[0036] In one embodiment, constructing a contrastive loss function based on the normalized embedding representation includes:

[0037] The contrast loss function is constructed based on the normalized embedding representation:

[0038]

[0039] Where n represents The total number of samples in the training data set, P(i) represents the number of samples corresponding to the training sample x i The sample set with the same relationship, |P(i)| represents the cardinality of P(i), and p represents the cardinality of the training sample x. i Samples with the same relationship, S I Represents the set of some samples in the memory, z i represents the normalized embedding representation, z p Represents the training sample x i The samples with the same relationship are represented by , and τ represents the temperature coefficient.

[0040] In one embodiment, before optimizing the initial continuous small sample relationship extraction model according to the contrastive learning and knowledge distillation method, the method further includes:

[0041] The memory of all tasks is initialized according to the normalized embedding representation. The cluster embedding of each relation training sample in the training data set is obtained in the memory using the k-means method. The typical samples of the memory are selected and saved according to the cluster embedding.

[0042] In one embodiment, the initial continuous small sample relationship extraction model is optimized according to contrastive learning and knowledge distillation methods to obtain a small sample relationship extraction model under continuous learning, including:

[0043] According to the knowledge distillation method, the similarity between the typical samples and the corresponding relationship prototypes in the previous task, and the dissimilarity between the typical samples and other relationship prototypes are retained. The consistency between the similarity and dissimilarity between each sample and all relationship prototypes in the previous task is calculated. The initial continuous small sample relationship extraction model is optimized using the calculation results and the pre-set contrastive learning loss function to obtain the continuous small sample relationship extraction model.

[0044] In one embodiment, the consistency between the similarity and dissimilarity between each sample and all relation prototypes in the previous task is calculated, including:

[0045] The consistency between the similarity and dissimilarity between each sample and all relation prototypes in the previous task is calculated, and the calculation result is

[0046]

[0047] in, is the metric distribution of similarity between samples and other relation prototypes before training new tasks, represents the set of all relations that appear in the first k tasks, represents the calculation of the similarity between prototype i and relation prototype j before training a new task, is the metric distribution of similarity between samples and other relation prototypes after training a new task, represents the calculation of the similarity between prototype i and relation prototype j after training a new task, Sample z calculated by dynamic embedding i and the relational prototype p j is the cosine similarity between them, and τ represents the temperature coefficient.

[0048] The above-mentioned small-sample relationship extraction method under continuous learning based on prompt contrastive learning, this application obtains enhanced sentence representation by marking the entities in the task, and then designs a template and obtains prompt-based data through a template mapping function, and uses the pre-trained language model BERT as an encoder to obtain a better relationship representation, which can better utilize the knowledge in a limited number of training samples, and then uses the projection head of the pre-trained BERT model to map the vector representation of the sentence to a low-dimensional embedding space to obtain a hidden representation; the hidden representation is normalized to obtain a normalized embedding representation, and the total loss function of the initial CFRE model is constructed according to the probability distribution on the label set, the hidden representation, the normalized embedding representation and the data enhancement strategy, and the initial CFRE model is trained using the training data set and the total loss function, and the loss function of the initial CFRE model is minimized to optimize the parameters in the model to obtain an initial continuous small-sample relationship extraction model, and the robustness of the model is improved by an enhancement strategy of mixed data features, and overfitting of small-sample tasks is prevented, thereby improving the accuracy of model relationship extraction. Secondly, in order to alleviate the forgetting of previous knowledge, in addition to replaying samples in memory, the initial continuous small sample relation extraction model is optimized through contrastive learning and knowledge distillation methods, maintaining the stability of instance embedding, and retaining the knowledge of learned relations through instance-level distillation loss, reducing relation extraction errors. BRIEF DESCRIPTION OF THE DRAWINGS

[0049] Figure 1 It is a flowchart of a method for extracting small sample relations under continuous learning based on prompt contrastive learning in one embodiment;

[0050] Figure 2 FIG. 4 is a diagram showing the internal structure of a computer device in one embodiment. DETAILED DESCRIPTION

[0051] In order to make the purpose, technical solution and advantages of the present application more clearly understood, the present application is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0052] In one embodiment, Figure 1 As shown, a small sample relationship extraction method under continuous learning based on prompt contrastive learning is provided, including the following steps:

[0053] Step 102, obtaining a task set and an initial CFRE model; the task set includes multiple tasks; each task includes its own training data set; the data set includes multiple training samples; the training samples include multiple sentences.

[0054] Step 104, perform information enhancement on the sentences in the task to obtain enhanced sentences; map the enhanced sentences according to a pre-built template mapping function to obtain prompt-based data; input the prompt-based data into the pre-trained BERT model to obtain a vector representation of the sentence, and calculate the probability distribution on the label set according to the vector representation of the sentence.

[0055] This application uses some special tags [E11], [E12], [E21] and [E22] at the beginning and end of the sentence entity to increase sentence information for better representation. The enhanced sentence is

[0056] x aug ={w 1 ,...,[E 11 ],e 1 ,[E 12 ],...,[E 21 ],e 2 ,[E 22 ],...,w x} (1)

[0057] A template mapping function T(·) is constructed to obtain the prompt input, which can enable PLM to generate better representations for instances.

[0058] x aug As the input of T(·) to obtain the prompt-based data x prompt .

[0059] x prompt =T(x aug )=I think 1 is[MASK]of e 2 [SEP]x aug (2)

[0060] Then, x prompt Provided to BERT, generate the representation of the [MASK] position of the sentence, and calculate the label set based on the vector representation of the sentence The probability distribution on is:

[0061]

[0062] in is the target label word, and v is the word obtained by BERT. Representation

[0063] Step 106, using a projection head to map the vector representation of the sentence to a low-dimensional embedding space to obtain a hidden representation; and normalizing the hidden representation to obtain a normalized embedded representation.

[0064] To obtain a denser representation, we use the projection head Pro of the pre-trained BERT model j Features are further extracted and mapped into a low-dimensional embedding space.

[0065]

[0066] In this model, the hidden representation is used for relationship classification, and the hidden representation is normalized to obtain the normalized embedding representation Do comparative learning.

[0067] Step 108, in each task, a total loss function of the initial CFRE model is constructed according to the probability distribution, hidden representation, normalized embedding representation and data augmentation strategy on the label set, the initial CFRE model is trained using the training data set and the total loss function, and the loss function of the initial CFRE model is minimized to optimize the parameters in the model to obtain an initial continuous small sample relationship extraction model.

[0068] According to the general assumption of CFRE, all relations expressed in the historical tasks are k Therefore, in the kth task, we first initialize T with the normalized embedding z. k Memory in M all .

[0069]

[0070] Among them, M all is the memory of all tasks, R k It is T k The relationship set, is the set of memory samples stored in relation r, O is the number of samples stored in the memory for each relation. Assume O=1.

[0071] In the training data set Train the model Θ to minimize the total loss L all To optimize the parameters in the model, where L all Including cross entropy loss, sentence feature mixed loss and contrast loss.

[0072] The cross entropy loss describes the similarity between the actual output probability and the expected output probability. The cross entropy loss function formula is as follows:

[0073]

[0074] where y i is the sample i in The real label in .

[0075] Mix-up is a data augmentation method that uses linear interpolation to generate more training samples, which can improve the generalization ability and robustness of the model on few-sample tasks.

[0076] In this model, sentence mixing is used. In the senMix-up loss, a pair of hidden representations And its corresponding class label y is linearly interpolated, the formula is as follows:

[0077]

[0078]

[0079] Hidden Representation Vector This is used to obtain the distribution of possible target labels in the softmax layer. We then use a multi-class cross entropy loss for training:

[0080]

[0081] The contrast loss attempts to reduce the distance between samples with the same relational labels and expand the distance between samples with different relational labels. The contrast loss function is as follows:

[0082]

[0083] Where n is a batch The number of samples, P(i) is the number of samples with the training sample x i The sample set with the same relationship, |P(i)| is its cardinality, z i is the training sample x i The expression S I It is M all A collection of some samples.

[0084] The total training loss can be written as:

[0085] L All =α 1 ·L CE +α 2 ·L SM +α 3 ·L CL (11)

[0086] Minimize the overall training loss to quickly adapt to new tasks.

[0087] Step 110, optimizing the initial continuous small sample relationship extraction model according to contrastive learning and knowledge distillation methods to obtain a small sample relationship extraction model under continuous learning; and performing relationship extraction using the small sample relationship extraction model under continuous learning.

[0088] The initial continuous small sample relation extraction model is optimized by contrastive learning and knowledge distillation methods, maintaining the stability of instance embedding, and retaining the knowledge of learned relations through instance-level distillation loss, reducing relation extraction errors. When the continuous small sample relation extraction model is used for relation extraction, given an input sentence Input into the continuous small sample relationship extraction model, and use formula (3) to calculate the relationship label: The probability of , select the relationship with the largest probability as the predicted relationship.

[0089] In the above-mentioned small-sample relationship extraction method under continuous learning based on prompt contrastive learning, the present application obtains enhanced sentence representation by marking the entities in the task, then designs a template and obtains prompt-based data through a template mapping function, and uses the pre-trained language model BERT as an encoder to obtain better relationship representation, which can better utilize the knowledge in a limited number of training samples, and then uses the projection head of the pre-trained BERT model to map the vector representation of the sentence to a low-dimensional embedding space to obtain a hidden representation; the hidden representation is normalized to obtain a normalized embedding representation, and the total loss function of the initial CFRE model is constructed according to the probability distribution on the label set, the hidden representation, the normalized embedding representation and the data enhancement strategy, and the initial CFRE model is trained using the training data set and the total loss function, and the loss function of the initial CFRE model is minimized to optimize the parameters in the model to obtain an initial continuous small-sample relationship extraction model, and the robustness of the model is improved by an enhancement strategy of mixed data features, and overfitting of few-sample tasks is prevented, thereby improving the accuracy of model relationship extraction. Secondly, in order to alleviate the forgetting of previous knowledge, in addition to replaying samples in memory, the initial continuous small sample relation extraction model is optimized through contrastive learning and knowledge distillation methods, maintaining the stability of instance embedding, and retaining the knowledge of learned relations through instance-level distillation loss, reducing relation extraction errors.

[0090] In one embodiment, information enhancement is performed on a sentence in a task to obtain an enhanced sentence, including:

[0091] The information of the sentence in the task is enhanced, and the enhanced sentence is

[0092] x aug ={w 1 ,...,[E 11 ],e 1 ,[E 12 ],...,[E 21 ],e 2 ,[E 22 ],...,w x}

[0093] Among them, wx Indicates a word in a sentence. [E11], [E12], [E21], and [E22] are all special tags. 1 and e 2 Represents the head entity and the tail entity in a sentence.

[0094] In one embodiment, the enhanced sentence is mapped according to a pre-built template mapping function to obtain prompt-based data, including:

[0095] The enhanced sentences are mapped according to the pre-built template mapping function, and the prompt-based data is obtained as

[0096] x prompt =T(x aug )=I think 1 is[MASK]of e 2 [SEP]x aug

[0097] Where T(·) represents the template mapping function, [MASK] represents the task location, and [SEP] represents the sentence separator.

[0098] In one embodiment, the total loss function of the initial CFRE model includes a cross entropy loss function, a sentence feature mixed loss function, and a contrast loss function; in each task, the total loss function of the initial CFRE model is constructed according to the probability distribution, hidden representation, normalized embedding representation, and data enhancement strategy on the label set, including:

[0099] Construct a cross entropy loss function based on the probability distribution on the label set;

[0100] According to the data enhancement strategy, linear interpolation calculation is performed on the hidden representation and the class label corresponding to the hidden representation to obtain the interpolation result; the sentence feature mixed loss function is constructed using the interpolation result;

[0101] Construct a contrastive loss function based on the normalized embedding representation;

[0102] The total loss function of the initial CFRE model constructed using the cross entropy loss function, sentence feature mixed loss function and contrast loss function is:

[0103] L All =α 1 ·L CE +α 2 ·L SM +α 3 ·L CL

[0104] Among them, L CE represents the cross entropy loss function, L SMrepresents the sentence feature mixed loss function, L CL represents the contrast loss function, α 1 , α 2 , α 3 Both represent tuning parameters.

[0105] In one embodiment, a cross entropy loss function is constructed according to the probability distribution on the label set, including:

[0106] According to the probability distribution on the label set, the cross entropy loss function is constructed as follows:

[0107]

[0108] Among them, Θ represents the initial CFRE model, represents the training data set, k represents the task number, P(y=y i ∣z i ,y i ) represents the probability distribution on the label set, y represents the target label word, and y i represents the true label of sample i in the training data set, z i represents the normalized embedding representation, and i represents the sample number.

[0109] In one embodiment, the sentence feature hybrid loss function is constructed using the interpolation result, including:

[0110] The sentence feature mixed loss function is constructed using the interpolation results:

[0111]

[0112] Among them, v i Indicates the representation of different words obtained by BERT. Represents hidden representation, represents the mixed sentence features obtained by the prompt-based encoder, Represents the mixed label of mixed sentence features, The correct word indicates that.

[0113] In one embodiment, constructing a contrastive loss function based on the normalized embedding representation includes:

[0114] The contrast loss function is constructed based on the normalized embedding representation:

[0115]

[0116] Where n represents The total number of samples in the training data set, P(i) represents the number of samples corresponding to the training sample x iThe sample set with the same relationship, |P(i)| represents the cardinality of P(i), and p represents the cardinality of the training sample x. i Samples with the same relationship, S I Represents the set of some samples in the memory, z i represents the normalized embedding representation, z p Represents the training sample x i The samples with the same relationship are represented by , and τ represents the temperature coefficient.

[0117] In one embodiment, before optimizing the initial continuous small sample relationship extraction model according to the contrastive learning and knowledge distillation method, the method further includes:

[0118] The memory of all tasks is initialized according to the normalized embedding representation. The cluster embedding of each relation training sample in the training data set is obtained in the memory using the k-means method. The typical samples of the memory are selected and saved according to the cluster embedding.

[0119] In a specific embodiment, after the model is trained using formula (11), in each new task, in order to enable the model to remember the knowledge of the relationship in the previous task, a typical sample is saved in the memory for each relationship. i , using the k-means method to obtain The cluster embedding of each relation training sample in M ​​is stored in the closest sample. all middle.

[0120] In one embodiment, the initial continuous small sample relationship extraction model is optimized according to contrastive learning and knowledge distillation methods to obtain a small sample relationship extraction model under continuous learning, including:

[0121] According to the knowledge distillation method, the similarity between the typical samples and the corresponding relationship prototypes in the previous task, and the dissimilarity between the typical samples and other relationship prototypes are retained. The consistency between the similarity and dissimilarity between each sample and all relationship prototypes in the previous task is calculated. The initial continuous small sample relationship extraction model is optimized using the calculation results and the pre-set contrastive learning loss function to obtain the continuous small sample relationship extraction model.

[0122] In a specific embodiment, two replay strategies are used to prevent the encoder from forgetting the knowledge of learned relations when learning new relations, namely contrastive learning and knowledge distillation. When completing the training of a new task, the samples are replayed in the memory using formula (10) to review the learned knowledge. The SI here is not a part of the samples in the memory, but all the samples in the memory. By replaying all the samples in the memory, the model regains a stable understanding of the learned relations and consolidates the relations learned in the current task.

[0123] However, since only one memory sample is stored for each relation in the new task, it is easy to cause overfitting and disrupt the distribution of relations. Therefore, this application uses the knowledge distillation method to retain the knowledge between relations and samples in the previous task. The model needs to maintain the similarity between the sample and its corresponding relation prototype as much as possible while maintaining the stability of the embedding space, as well as the dissimilarity between the sample and other relation prototypes.

[0124] The prototype for each relation is obtained by averaging the samples of each relation stored in memory:

[0125]

[0126] Sample z i With relation prototype p j The similarity is calculated by cosine similarity:

[0127]

[0128] where α ij For sample z i With relation prototype p j The cosine similarity between .

[0129] After training a new task, the parameters of the model will change, and the parameters of the encoder will also change. Therefore, the sample embedding stored in the memory is dynamically changed. When learning a new task, KL divergence is used to make the similarity and dissimilarity between each sample and all relation prototypes in the previous task as consistent as possible to preserve the knowledge of those learned relations between this step and the previous step.

[0130]

[0131] in is the metric distribution of the similarity between the sample and other relation prototypes before training the new task, where Similarly, is the metric distribution of the similarity between the sample and other relation prototypes after training the new task, where Sample z calculated by dynamic embedding i and prototype p j The cosine similarity between .

[0132] In one embodiment, the consistency between the similarity and dissimilarity between each sample and all relation prototypes in the previous task is calculated, including:

[0133] The consistency between the similarity and dissimilarity between each sample and all relation prototypes in the previous task is calculated, and the calculation result is

[0134]

[0135] in, is the metric distribution of similarity between samples and other relation prototypes before training new tasks, represents the set of all relations that appear in the first k tasks, represents the calculation of the similarity between prototype i and relation prototype j before training a new task, is the metric distribution of similarity between samples and other relation prototypes after training a new task, represents the calculation of the similarity between prototype i and relation prototype j after training a new task, Sample z calculated by dynamic embedding i and the relational prototype p j is the cosine similarity between them, and τ represents the temperature coefficient.

[0136] It should be understood that although Figure 1 The steps in the flowchart are shown in sequence as indicated by the arrows, but these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise specified in this document, there is no strict order restriction for the execution of these steps, and these steps can be executed in other orders. Moreover, Figure 1 At least part of the steps may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily executed at the same time, but can be executed at different times. The execution order of these sub-steps or stages is not necessarily sequential, but can be executed in turn or alternately with other steps or at least part of the sub-steps or stages of other steps.

[0137] In one embodiment, a computer device is provided. The computer device may be a terminal, and its internal structure diagram may be as follows: Figure 2 As shown. The computer device includes a processor, a memory, a network interface, a display screen and an input device connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, a small sample relationship extraction method under continuous learning based on prompt contrast learning is implemented. The display screen of the computer device can be a liquid crystal display screen or an electronic ink display screen, and the input device of the computer device can be a touch layer covered on the display screen, or a button, trackball or touchpad set on the computer device shell, or an external keyboard, touchpad or mouse, etc.

[0138] Those skilled in the art will understand that Figure 2 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine certain components, or have a different arrangement of components.

[0139] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. As an illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM).

[0140] The technical features of the above embodiments may be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0141] The above-mentioned embodiments only express several implementation methods of the present application, and the descriptions thereof are relatively specific and detailed, but they cannot be understood as limiting the scope of the invention patent. It should be pointed out that, for a person of ordinary skill in the art, several variations and improvements can be made without departing from the concept of the present application, and these all belong to the protection scope of the present application. Therefore, the protection scope of the patent of the present application shall be subject to the attached claims.

Claims

1. A small sample relationship extraction method under continuous learning based on prompt contrastive learning, characterized in that: The method comprises: Obtain a task set and an initial CFRE model; the task set includes multiple tasks; each task includes its own training data set; the data set includes multiple training samples; the training samples include multiple sentences; Performing information enhancement on sentences in the task to obtain enhanced sentences; mapping the enhanced sentences according to a pre-built template mapping function to obtain prompt-based data; Inputting the prompt-based data into a pre-trained BERT model to obtain a vector representation of the sentence, and calculating a probability distribution on the label set based on the vector representation of the sentence; Mapping the vector representation of the sentence to a low-dimensional embedding space using a projection head to obtain a hidden representation; normalizing the hidden representation to obtain a normalized embedding representation; In each task, a total loss function of the initial CFRE model is constructed according to the probability distribution, hidden representation, normalized embedding representation and data augmentation strategy on the label set, the initial CFRE model is trained using the training data set and the total loss function, and the loss function of the initial CFRE model is minimized to optimize the parameters in the model, thereby obtaining an initial continuous small sample relationship extraction model; The initial continuous small sample relationship extraction model is optimized according to contrastive learning and knowledge distillation methods to obtain a small sample relationship extraction model under continuous learning; Relation extraction is performed using the small sample relationship extraction model under continuous learning.

2. The method according to claim 1, characterized in that Perform information enhancement on the sentences in the task to obtain enhanced sentences, including: The information of the sentence in the task is enhanced, and the enhanced sentence is x aug ={w1,...,[E 11 ],e1,[E 12 ],...,[AND 21 ],e2,[E 22 ],...,w x } Among them, w x represents a word in a sentence, [E11], [E12], [E21] and [E22] all represent special tags, and e1 and e2 represent the head entity and tail entity in a sentence.

3. The method according to claim 2, characterized in that The enhanced sentence is mapped according to a pre-built template mapping function to obtain prompt-based data, including: The enhanced sentence is mapped according to the pre-built template mapping function to obtain the prompt-based data: x prompi =T(x aug )=I think e1is[MASK]of e2[SEP]x aug Where T(·) represents the template mapping function, [MASK] represents the task location, and [SEP] represents the sentence separator.

4. The method according to claim 3, characterized in that The total loss function of the initial CFRE model includes a cross entropy loss function, a sentence feature mixed loss function and a contrast loss function; in each task, the total loss function of the initial CFRE model is constructed according to the probability distribution, hidden representation, normalized embedding representation and data enhancement strategy on the label set, including: Constructing a cross entropy loss function according to the probability distribution on the label set; According to the data enhancement strategy, linear interpolation calculation is performed on the hidden representation and the class label corresponding to the hidden representation to obtain an interpolation result; and a sentence feature mixed loss function is constructed using the interpolation result; Construct a contrastive loss function based on the normalized embedding representation; The total loss function of the initial CFRE model constructed using the cross entropy loss function, sentence feature mixed loss function and contrast loss function is: L All =α1·L CE +α2·L SM +α3·L CL Among them, L CE represents the cross entropy loss function, L SM represents the sentence feature mixed loss function, L CL represents the contrast loss function, and α1, α2, and α3 represent tuning parameters.

5. The method according to claim 4, characterized in that A cross entropy loss function is constructed according to the probability distribution on the label set, including: The cross entropy loss function is constructed based on the probability distribution on the label set: Among them, Θ represents the initial CFRE model, represents the training data set, k represents the task number, P(y=y i ∣z i ,y i ) represents the probability distribution on the label set, y represents the target label word, yi represents the true label of sample i in the training data set, and z i represents the normalized embedding representation, and i represents the sample number.

6. The method according to claim 5, characterized in that The sentence feature mixed loss function is constructed using the interpolation result, including: The sentence feature mixed loss function is constructed using the interpolation result: Among them, v i Indicates the representation of different words obtained by BERT, Represents hidden representation, represents the mixed sentence features obtained by the prompt-based encoder, Represents the mixed label of mixed sentence features and the correct word representation.

7. The method according to claim 6, characterized in that Construct a contrastive loss function based on the normalized embedding representation, including: The contrast loss function is constructed based on the normalized embedding representation: Where n represents The total number of samples in the training data set, P(i) represents the number of samples corresponding to the training sample x i The sample set with the same relationship, |P(i)| represents the cardinality of P(i), and p represents the cardinality of the training sample x. i Samples with the same relationship, S I Represents the set of some samples in the memory, z i represents the normalized embedding representation, z p Represents the training sample x i The samples with the same relationship are represented by , and τ represents the temperature coefficient.

8. The method according to claim 1, characterized in that Before optimizing the initial continuous small sample relationship extraction model according to the contrastive learning and knowledge distillation method, it also includes: The memory of all tasks is initialized according to the normalized embedding representation, the cluster embedding of each relation training sample in the training data set is obtained in the memory using the k-means method, and typical samples of the memory are selected and saved according to the cluster embedding.

9. The method according to claim 8, characterized in that The initial continuous small sample relationship extraction model is optimized according to contrastive learning and knowledge distillation methods to obtain a small sample relationship extraction model under continuous learning, including: According to the knowledge distillation method, the similarity between the typical sample and the corresponding relationship prototype in the previous task, and the dissimilarity between the typical sample and other relationship prototypes are retained, and the consistency between the similarity and dissimilarity between each sample and all relationship prototypes in the previous task is calculated. The initial continuous small sample relationship extraction model is optimized using the calculation results and a pre-set contrastive learning loss function to obtain a continuous small sample relationship extraction model.

10. The method according to claim 9, characterized in that The consistency between the similarity and dissimilarity between each sample and all relation prototypes in the previous task is calculated, including: The consistency between the similarity and dissimilarity between each sample and all relation prototypes in the previous task is calculated, and the calculation result is in, is the metric distribution of similarity between samples and other relation prototypes before training new tasks, represents the set of all relations that appear in the first k tasks, represents the calculation of the similarity between prototype i and relation prototype j before training a new task, is the metric distribution of similarity between samples and other relation prototypes after training a new task, represents the calculation of the similarity between prototype i and relation prototype j after training a new task, Sample z calculated by dynamic embedding i and the relational prototype p j is the cosine similarity between them, and τ represents the temperature coefficient.

Citation Information

Patent Citations

  • Zero sample relation extraction method and system based on dual contrast learning

    CN114548325A

  • Small sample relation classification method and system based on prompt learning, medium and electronic equipment

    CN115982363A