Knowledge distillation and parameter efficient fine tuning fused few-sample relationship classification method

By integrating knowledge distillation and efficient parameter fine-tuning in the small sample relationship classification, the problem of overfitting and not using pre-trained model knowledge in the prior art is solved, and a more accurate and stable relationship representation and classification effect is achieved.

CN120086374AActive Publication Date: 2025-06-03NAT UNIV OF DEFENSE TECH
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
CN202510567812.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-30
Publication Date
2025-06-03
Estimated Expiration
2045-04-30

AI Technical Summary

Technical Problem

Existing low-sample relationship classification methods are prone to overfitting when fine-tuning parameters, and do not make full use of the factual knowledge in the pre-trained language model, resulting in inaccurate prototype representation.

Method used

Using a method of fusion knowledge distillation and efficient parameter fine-tuning, factual knowledge in the pre-trained language model is captured through parameter fine-tuning of the mixing prompts, and the prototype representation of the relationship category is optimized through knowledge distillation.

Benefits of technology

The accuracy of the classification of relationships with few samples is significantly improved, and through efficient parameter fine-tuning and knowledge distillation technology, more stable relationship representation and higher classification accuracy are obtained.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120086374A_ABST
    Figure CN120086374A_ABST
Patent Text Reader

Abstract

The invention relates to a few-sample relation classification method integrating knowledge distillation and efficient parameter fine tuning. According to the method, a few-sample relation classification model comprising a mixed prompt input module, a coding module, a relation classifier and a consistency discriminator is constructed. According to the model, a parameter efficient fine tuning strategy is adopted, only a small number of trainable virtual marks in mixed prompt input need to be optimized, meanwhile, all other model parameters are frozen, and the training efficiency is remarkably improved. Moreover, the model constructs a teacher-student framework with consistency constraint through a knowledge distillation technology, can effectively extract fact knowledge in the pre-training model, and obtains more stable relation representation from the perspective of memory enhancement. Besides, the model can further refine and extract more accurate prototype representation and reduce noise interference by performing weighted average on the relationship representation of the support set instances, and finally can improve the classification prediction accuracy of the query set instances in a few-sample relationship classification task.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of natural language processing, and particularly to a few-shot relation classification method that integrates knowledge distillation and parameter-efficient fine-tuning. Background Art

[0002] Few-shot relation classification is a core task in natural language processing, aiming to identify the semantic relations between entity pairs in text through extremely few labeled instances. Existing research usually relies on prototype networks to learn dense representations of relation categories, mainly by integrating text descriptions of entities and relations or pre-trained language models to enhance prototype representations. However, these methods usually fine-tune all parameters, which may lead to overfitting problems in few-shot learning. At the same time, they fail to fully utilize the factual knowledge in pre-trained language models, resulting in inaccurate prototype representations and affecting the accuracy of few-shot relation classification. Summary of the Invention

[0003] Based on this, in view of the above technical problems, it is necessary to provide a few-shot relation classification method that integrates knowledge distillation and parameter-efficient fine-tuning. The parameter fine-tuning based on hybrid prompts can efficiently capture the factual knowledge in pre-trained language models, and based on knowledge distillation, it can optimize the prototype representations of relation categories through memory enhancement, thereby obtaining more stable and effective relation representations and improving the accuracy of few-shot relation classification.

[0004] A few-shot relation classification method that integrates knowledge distillation and parameter-efficient fine-tuning, the method comprising: Dividing a text data set into a training set, a validation set, and a test set, and randomly extracting multiple meta-tasks composed of a support set and a query set from each of the divided data sets; the support set and the query set contain multiple relation categories and multiple instances corresponding to each relation category, and each instance is composed of a sentence and the semantic relation of the entity pair in the sentence; Constructing a few-shot relation classification model including a hybrid prompt input module, an encoding module, a relation classifier, and a consistency discriminator; wherein, the hybrid prompt input module is used to splice the virtual tokens in the constructed teacher model and student model with the sentences and entity pairs in the input instances respectively to obtain a hybrid prompt input; the encoding module is used to encode the hybrid prompt input according to a pre-trained language model to obtain relation representations of support set instances and query set instances; the relation classifier is used to perform weighted averaging on the relation representations of support set instances to obtain prototype representations of each relation category, and predict the relation category to which the query set instance belongs by calculating the similarity between the prototype representation and the relation representation of the query set instance; the consistency discriminator is used to construct a consistency constraint between the teacher model and the student model; Construct a training objective that includes a relationship prediction loss and a consistency constraint loss, and perform model training and evaluation on the meta-tasks extracted from the training set and the validation set until a few-shot relationship classification model that meets the training objective is obtained and the few-shot relationship classification task in the test set is executed; among them, only the virtual labels are optimized during model training, and all other model parameters are kept frozen.

[0005] In one embodiment, the text dataset is divided into a training set, a validation set, and a test set, and multiple meta-tasks composed of a support set and a query set are randomly selected from each of the divided datasets, including: The text dataset containing multiple relationship categories and corresponding instances is divided into a training set, a validation set, and a test set according to categories; among them, the relationship categories included in the training set, the validation set, and the test set are non-overlapping; Randomly select multiple meta-training tasks, meta-validation tasks, and meta-test tasks from the training set, the validation set, and the test set respectively; each meta-task consists of a support set and a query set ; among them, is the number of relationship categories in the support set or the query set, K is the number of support set instances included in each relationship category in the support set, is the number of query set instances included in each relationship category in the query set, and the support set instances and the query set instances are non-overlapping; and respectively represent the sentence and the semantic relationship of the entity pair in the sentence in the th support set instance; and respectively represent the sentence and the semantic relationship of the entity pair in the sentence in the jth query set instance; Support set and query set in each instance in, represents a sentence containing a pair of entities, represents the semantic relationship between the entity pair in the sentence; the objective of the meta-task is to use a small number of labeled instances in the support set to predict the relationship category of the entity pair in the sentence in any query set instance ; among them represents a sentence containing words, is the ith word in the sentence and , and respectively represent the head entity and the tail entity in the sentence, represents and the semantic relationship between them, It is a predefined set of relationships.

[0006] In one embodiment, the virtual tokens in the constructed teacher model and student model are respectively concatenated with the sentence and entity pairs in the input instance to obtain a mixed prompt input, including: The teacher model and student model are constructed using the mean teacher algorithm, and the continuous dense vectors of the learnable virtual tokens in the teacher model and student model are respectively obtained ; where represents the i-th vector, and ; The continuous dense vectors in the teacher model and student model are respectively concatenated with the sentence in the input instance and the triple composed of the head entity , the mask token and the tail entity to obtain the mixed prompt input of the teacher model and the mixed prompt input of the student model with consistent structures, both represented as: ; where is the mixed prompt input; the input instance includes a support set instance and a query set instance.

[0007] In one embodiment, the mixed prompt input is encoded according to the pre-trained language model to obtain the relationship representation of the support set instance and the query set instance, including: The pre-trained language model is represented as a function that maps the mixed prompt input to the feature representation of the mask token , represented as: ; where is the hidden representation of the mask token , and is used as the relationship representation of the support set instance and the query set instance obtained from the teacher model or the student model; are the fixed parameters in the backbone module of the pre-trained language model, are the trainable parameters, specifically representing the parameters to be trained in the teacher model or the student model, and .

[0008] In one embodiment, the parameters of the teacher model in the training step are the exponential moving average weights of the parameters of the student model, represented as: ; Among them, is the hyperparameter of the smoothing coefficient, and the initial parameters of the student model are randomly initialized, and the initial parameters of the teacher model are equal to .

[0009] In one embodiment, a weighted average is performed on the relational representations of the support set instances to obtain the prototype representation of each relational category, including: A weighted average is performed on the relational representations of the support set instances obtained from the teacher model and the student model to obtain the prototype representation of each relational category. The expression is: ; Among them, represents the prototype representation of the i-th relational category; is the teacher model t or the student model s; represents the relational representation of the j-th support set instance of the i-th relational category in the support set S obtained from the teacher model t or the student model s, is the support set The number of support set instances included in each relational category in is an adjustable parameter, defined as: .

[0010] In one embodiment, by calculating the similarity between the prototype representation and the relational representation of the query set instance, the relational category to which the query set instance belongs is predicted, including: According to the Euclidean distance function Calculate the prototype representation of the i-th relational category and the relational representation of the query set instance The similarity between them is predicted to obtain the sentence of the query set instance The probability that the entity pair in belongs to the i-th relational category , expressed as: ; Among them, represents the i-th relational category, is the semantic relationship between entity pairs in the sentence, and N is the number of relational categories in the support set or the query set.

[0011] In one embodiment, a training objective including a relational prediction loss and a consistency constraint loss is constructed, including: According to the relational category prediction probability of the entity pair in the sentence of the query set instance Construct a relational prediction loss , expressed as: ; Construct the consistency constraint loss between the teacher model and the student model according to the normalized relationship representations of each instance in the support set and the query set , which is expressed as: ; where K is the number of support set instances included in each relationship category in the support set , is the query set the number of query set instances included in each relationship category in; and respectively represent the normalized relationship representations of the i-th instance obtained from the student model s and the teacher model t , and , is the relationship representation of the i-th instance; The comprehensive relationship prediction loss and the consistency constraint loss , the final training objective is obtained as: ; where are the fixed parameters in the backbone module of the pre-trained language model, and respectively represent the parameters of the teacher model and the student model; is the dynamic balance coefficient.

[0012] The above few-shot relation classification method that combines knowledge distillation and parameter-efficient fine-tuning constructs a few-shot relation classification model including a hybrid prompt input module, an encoding module, a relation classifier, and a consistency discriminator. This model adopts a parameter-efficient fine-tuning strategy, only needs to optimize a small number of trainable virtual tokens in the hybrid prompt input, and freezes all other model parameters at the same time, significantly improving the training efficiency. Moreover, this model constructs a teacher-student framework with consistency constraints through knowledge distillation technology, can effectively extract factual knowledge from the pre-trained model, and obtain a more stable relation representation from the perspective of memory enhancement. In addition, this model can further refine and extract more accurate prototype representations by weighted averaging the relation representations of the support set instances, reduce noise interference, and finally improve the classification prediction accuracy of the query set instances in the few-shot relation classification task. BRIEF DESCRIPTION OF THE DRAWINGS

[0013] Figure 1 is a schematic flowchart of a few-shot relation classification method that combines knowledge distillation and parameter-efficient fine-tuning in an embodiment; Figure 2 is a schematic diagram of the overall architecture of a few-shot relation classification model in an embodiment. DETAILED DESCRIPTION OF THE INVENTION

[0014] To make the objectives, technical solutions and advantages of this application clearer and more understandable, the following further details this application in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not used to limit this application.

[0015] In one embodiment, as Figure 1 shown, a few-shot relation classification method integrating knowledge distillation and parameter-efficient fine-tuning is provided, including the following steps: Step S1, divide the text dataset into a training set, a validation set and a test set, and randomly extract multiple meta-tasks composed of a support set and a query set from each of the divided datasets.

[0016] Specifically, step S1 includes: First, divide the text dataset containing multiple relation categories and corresponding instances into a training set, a validation set and a test set according to the categories; among them, the relation categories included in the training set, the validation set and the test set are non-overlapping. Such a setting can ensure that the model makes accurate predictions and evaluations when facing unseen relation categories.

[0017] Then, randomly extract multiple meta-training tasks, meta-validation tasks and meta-test tasks from the training set, the validation set and the test set respectively; in a typical N-way-K-shot (N categories and K samples) setting, each meta-task consists of a support set and a query set ; among them, is the number of relation categories in the support set or the query set, K is the number of support set instances included in each relation category in the support set, is the number of query set instances included in each relation category in the query set, and the support set instances and the query set instances are non-overlapping; and respectively represent the sentence and the semantic relation of the entity pair in the sentence in the th support set instance; and respectively represent the sentence and the semantic relation of the entity pair in the sentence in the jth query set instance.

[0018] In the support set and the query set for each instance , represents a sentence containing a pair of entities, represents the semantic relation between the entity pair in the sentence; the objective of the meta-task is to use a small number of labeled instances in the support set to predict the relation category of the entity pair in the sentence in any query set instance ; where denotes a sentence containing words, is the i-th word in the sentence and , and respectively denote the head entity and the tail entity in the sentence, denotes and the semantic relationship between them, is a predefined set of relationships.

[0019] Step S2, construct a few-shot relation classification model including a hybrid prompt input module, an encoding module, a relation classifier, and a consistency discriminator.

[0020] The overall architecture of the few-shot relation classification model is as Figure 2 shown, Figure 2 where the yellow color block and the orange color block respectively represent the learnable virtual tokens from the student model and the teacher model, the blue color block represents the frozen parameters of the model, and the green color block represents the relation representation of the instances in the model. The specific structure and data processing logic of the few-shot relation classification model are as follows: (1) Hybrid prompt input module: For the relation classification task, this application proposes a new hybrid prompt input method, which consists of some learnable task-specific virtual tokens and the structural patterns of triples. Specifically, the process of constructing the hybrid prompt input includes: Adopt the mean teacher algorithm to construct the teacher model and the student model, and respectively obtain the continuous dense vectors of the learnable virtual tokens in the teacher model and the student model ; where represents the i-th vector, and ; Respectively concatenate the continuous dense vectors in the teacher model and the student model with the sentence in the input instance and the triple composed of the head entity , the mask token and the tail entity to obtain the hybrid prompt input of the teacher model and the hybrid prompt input of the student model with consistent structures, both represented as: ; where is the hybrid prompt input; the input instance includes the support set instance and the query set instance.

[0021] (2) Encoding module: Use the pre-trained language model as the encoder, and use the pre-trained language model Expressed as a function that maps a mixed prompt input to a masked token and is represented as: ; where is the hidden representation of the masked token and is used as the relationship representation between the support set instance and the query set instance obtained from the teacher model or the student model; is a fixed parameter in the backbone module of the pre-trained language model, is a trainable parameter, specifically representing the parameters to be trained in the teacher model or the student model, and . Compared with the traditional model training method that optimizes all parameters , during the training of this application, only the continuous embedding vectors of some virtual tokens in the training mixed prompt input need to be optimized, while keeping all other model parameters frozen, so as to achieve efficient and concise parameter fine-tuning, avoid overfitting problems, and improve the model training efficiency.

[0022] Constructing the teacher model and the student model based on the mean teacher algorithm as described above is a knowledge distillation method. It forms new semantic memories by using the exponential moving average of the student model parameters as the parameters of the teacher model during the training process. Different from directly sharing the parameters of the student model, the parameters of the teacher model use the exponential moving average weights of the student model parameters, which can be regarded as enhancing the model through memory without directly optimizing. For few-shot relation classification, the mean teacher is a self-integrated intermediate model state that can obtain better relation representations. Specifically, the training steps The parameters of the teacher model in are the exponential moving average weights of the parameters of the student model and are represented as: where is the smoothing coefficient hyperparameter, the initial parameters of the student model are randomly initialized, and the initial parameters of the teacher model are equal to

[0023] (3) Relation classifier: The main idea of the traditional prototype network is to use prototype representations to represent each relation. The traditional method for calculating the prototype representation is to average the relation representations of all support set instances in the support set. Thus, the traditional prototype representation of the th relation is: ; where is from the support set the relationship representation of the th support set instance in the th relationship category. Compared with the traditional prototype representation calculation method, in this application, by adopting the mean teacher algorithm in the few-shot relationship classification model, the relationship representation can be obtained from the student model respectively, and the relationship representation can be obtained from the teacher model ; where, represents the prototype representation of the i-th relationship category; is the teacher model t or the student model s; represents the relationship representation of the j-th support set instance of the i-th relationship category in the support set S obtained from the teacher model t or the student model s, is the support set the number of support set instances included in each relationship category; is an adjustable parameter, defined as: .

[0024] Then, according to the Euclidean distance function calculate the similarity between the prototype representation of the i-th relationship category and the relationship representation of the query set instance, and predict the probability that the entity pair in the query set instance sentence belongs to the i-th relationship category, expressed as: ; where, represents the i-th relationship category, is the semantic relationship between the entity pairs in the sentence, and N is the number of relationship categories in the support set or the query set. Figure 2 in and respectively represent the relationship representations of each support set instance in the support set S obtained from the student model s and the teacher model t.

[0025] Further, according to the predicted probability of the relationship category of the entity pair in the query set instance sentence Construct the relational prediction loss in the form of cross - entropy objective function , expressed as: .

[0026] (4) Consistency discriminator: Used to construct the consistency constraint between the teacher model and the student model. Specifically, construct the consistency constraint loss between the teacher model and the student model according to the normalized relational representations of each instance in the support set and the query set , expressed as: ; where K is the number of support set instances included in each relation category in the support set , is the query set the number of query set instances included in each relation category in; and respectively represent the normalized relational representations of the i - th instance obtained from the student model s and the teacher model t , and , is the relational representation of the i - th instance. Figure 2 In and respectively represent the relational representations of each query set instance in the query set Q obtained from the student model s and the teacher model t.

[0027] Step S3, construct a training objective that includes relational prediction loss and consistency constraint loss, and perform model training and evaluation on the meta - tasks extracted from the training set and the validation set until a few - shot relational classification model that meets the training objective is obtained and perform the few - shot relational classification task in the test set; among them, only the virtual labels are optimized during model training, and all other model parameters are kept frozen.

[0028] Specifically, combining the relational prediction loss and the consistency constraint loss , the final training objective is: ; where are the fixed parameters in the backbone module of the pre - trained language model, and respectively represent the parameters of the teacher model and the student model; is the dynamic balance coefficient.

[0029] Furthermore, experiments are conducted on two widely used public datasets, FewRel 1.0 and FewRel 2.0, to evaluate the performance of the few - shot relational classification model constructed in this application.

[0030] The FewRel 1.0 dataset contains 100 relation categories, with 700 instances for each relation category. These instances are all extracted from Wikipedia articles. According to the official evaluation setting, Fewrel 1.0 is divided into a training set, a validation set, and a test set based on 100 relation categories. Among them, the training set contains 64 relation categories, the validation set contains 16 relation categories, and the test set contains 20 relation categories, with 700 instances for each relation category.

[0031] The FewRel 2.0 dataset mainly focuses on the problem of domain adaptability. Its training set is the same as that of FewRel 1.0; the validation set is the SemEval-2010 task 8 dataset, which contains 17 relation categories, with 520 instances for each relation category; the test set is the PubMed dataset, which comes from biomedical literature and contains 25 relation categories, with 100 instances for each relation category.

[0032] It should be noted that all the data of FewRel 1.0 belongs to the Wikipedia domain, while the validation set and test set of FewRel 2.0 come from the biomedical domain. The domain difference brings more challenges to the model's rapid learning and generalization. Therefore, FewRel 2.0 is more challenging than FewRel 1.0.

[0033] Four N-way-K-shot few-shot learning settings, namely 5-way-1-shot (hereinafter referred to as 5-w-1-s, and the same for others), 5-way-5-shot, 10-way-1-shot, and 10-way-5-shot, are used to evaluate the performance of the few-shot relation classification model constructed in this application, and the average accuracy is used as the evaluation metric. During training, 30,000 tasks are randomly sampled from the training data for training, 10,000 tasks are sampled from the validation data for evaluation, and 20,000 tasks are sampled from the test data for testing. Specifically, a 12-layer Transformer is used and initialized with BERT_BASE (the BERT base model) as the pre-trained language model. During the training process, a batch size of 4 is used. The optimizer is AdamW, and the learning rate is 2e-3, which is used to optimize the learning of continuous embeddings.

[0034] Table 1 Comparison results of different model performances on the FewRel 1.0 test set

[0035] Table 2 Comparison results of different model performances on the FewRel 2.0 test set

[0036] Table 1 and Table 2 respectively list the comparison results of different model performances on the FewRel 1.0 and FewRel 2.0 test sets. The few-shot relation classification model constructed in this application is labeled as ETKD in the table. Other comparison models include: MAML-BERT (BERT model adapted by meta-learning), REGRAB (relation-enhanced graph attention BERT), BERT-PAIR (BERT sentence pair model), Proto-BERT (BERT model enhanced by prototype network), MTB (pre-trained BERT for matching tasks), Big ProtoBERT (large-scale prototype BERT), Proto-HP (hybrid prompt prototype BERT), Proto-SP (soft prompt prototype BERT), Proto-HbP (hybrid block prompt prototype BERT), and Proto-HbP (Efficient) (efficient hybrid block prompt prototype BERT).

[0037] As can be seen from Table 1 and Table 2, the few-shot classification model constructed in this application not only trains model parameters efficiently but also achieves the best classification performance in all N-way-K-shot few-shot learning settings, which verifies the effectiveness of the few-shot relation classification method that integrates knowledge distillation and parameter-efficient fine-tuning in capturing key information in few-shot learning.

[0038] The technical features of the above embodiments can be combined arbitrarily. For the sake of concise description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope recorded in this specification.

[0039] The above-described embodiments only represent several implementation manners of this application, and their descriptions are relatively specific and detailed. However, it should not be construed as a limitation to the scope of this application. It should be noted that for those of ordinary skill in the art, without departing from the concept of this application, several modifications and improvements can still be made, and these all belong to the protection scope of this application. Therefore, the protection scope of this application should be subject to the appended claims.

Claims

1. A few-sample relation classification method integrating knowledge distillation and efficient parameter fine-tuning, characterized in that: The method comprises: The text data set is divided into a training set, a validation set and a test set, and a plurality of meta-tasks consisting of a support set and a query set are randomly extracted from each divided data set; the support set and the query set contain a plurality of relation categories and a plurality of instances corresponding to each relation category, and each instance consists of a sentence and a semantic relationship between entity pairs in the sentence; Construct a few-sample relation classification model including a hybrid prompt input module, an encoding module, a relation classifier and a consistency discriminator; wherein the hybrid prompt input module is used to respectively concatenate the virtual tags in the constructed teacher model and the student model with the sentence and entity pairs in the input instance to obtain a hybrid prompt input; the encoding module is used to encode the hybrid prompt input according to the pre-trained language model to obtain the relational representation of the support set instance and the query set instance; the relation classifier is used to perform weighted averaging on the relational representation of the support set instance to obtain the prototype representation of each relational category, and predict the relational category to which the query set instance belongs by calculating the similarity between the prototype representation and the relational representation of the query set instance; the consistency discriminator is used to construct a consistency constraint between the teacher model and the student model; A training objective including a relationship prediction loss and a consistency constraint loss is constructed, and model training and evaluation are performed on the meta-tasks extracted from the training set and the validation set until a few-sample relationship classification model that meets the training objective is obtained and the few-sample relationship classification task in the test set is performed; wherein, during model training, only the virtual label is optimized, and all other model parameters remain frozen.

2. The method according to claim 1, characterized in that The text dataset is divided into training set, validation set and test set, and multiple meta-tasks consisting of support set and query set are randomly selected from each divided dataset, including: Dividing a text data set containing multiple relationship categories and corresponding instances into a training set, a validation set, and a test set according to the categories; wherein the relationship categories contained in the training set, the validation set, and the test set are disjoint; A plurality of meta-training tasks, meta-verification tasks and meta-testing tasks are randomly selected from the training set, the validation set and the test set respectively; each meta-task consists of a support set and a queryset Composition; among them, is the number of relation categories in the support set or query set, K is the number of support set instances contained in each relation category in the support set, is the number of query set instances contained in each relation category in the query set, and the support set instances are disjoint with the query set instances; and Respectively represent The semantic relationship between the sentences in the support set instances and the entity pairs in the sentences; and Respectively represent the sentence in the j-th query set instance and the semantic relationship between the entity pairs in the sentence; Support set With queryset Each instance in middle, represents a sentence containing a pair of entities, Represents the semantic relationship between entity pairs in a sentence; the goal of the meta-task is to use the support set A small number of labeled instances in The relationship category of entity pairs in ;in Indicates that it contains A sentence of words, is the i-th word in the sentence and , and Represent the head entity and tail entity in the sentence respectively, express and The semantic relationship between Is a predefined set of relations.

3. The method according to claim 2, characterized in that The virtual tags in the constructed teacher model and student model are concatenated with the sentences and entity pairs in the input instance to obtain a mixed prompt input, including: The mean teacher algorithm is used to construct the teacher model and the student model, and the teacher model and the student model are obtained respectively. A continuous dense vector of learnable virtual labels ;in, represents the i-th vector, and ; The continuous dense vectors in the teacher model and the student model are respectively combined with the sentences in the input instance and the source entity , mask mark and tail entity The composed triplets are concatenated to obtain the mixed prompt input of the teacher model and the mixed prompt input of the student model with the same structure, both of which are expressed as: ; in, It is a mixed prompt input; the input instances include support set instances and query set instances.

4. The method according to claim 3, characterized in that Encoding the mixed prompt input according to the pre-trained language model to obtain a relationship representation between the support set instance and the query set instance, including: The pre-trained language model Represents a mixed prompt input Mapping to mask tags The function of the characteristic representation is expressed as: ; in, Is a mask mark The hidden representation of As a representation of the relationship between the support set instances and the query set instances obtained from the teacher model or the student model; is a fixed parameter in the backbone module of the pre-trained language model, are trainable parameters, including parameters that need to be trained in the teacher model or the student model, and .

5. The method according to claim 4, characterized in that Training steps The parameters of the teacher model in are the parameters of the student model The exponential moving average weight is expressed as: ; in, is the smoothing coefficient hyperparameter, the initial parameter of the student model is randomly initialized, the initial parameters of the teacher model equal .

6. The method according to claim 4, characterized in that The relation representations of the support set instances are weighted averaged to obtain the prototype representation of each relation category, including: The relation representations of the support set instances obtained from the teacher model and the student model are weighted averaged to obtain the prototype representation of each relation category, expressed as: ; in, represents the prototype representation of the i-th relationship category; is the teacher model t or the student model s; represents the relation representation of the jth support instance of the i-th relation category in the support S obtained from the teacher model t or the student model s, is the support set The number of support set instances contained in each relation category; is a tunable parameter defined as: 。 7. The method according to claim 6, characterized in that By calculating the similarity between the prototype representation and the relation representation of the query set instance, predicting the relation category to which the query set instance belongs, including: According to the Euclidean distance function Calculate the prototype representation of the i-th relationship category Relationship representation with queryset instances The similarity between them is used to predict the sentence of the query set instance. The probability that the entity pair in belongs to the i-th relation category , expressed as: ; in, represents the i-th relationship category, is the semantic relationship between entity pairs in the sentence, and N is the number of relationship categories in the support set or query set.

8. The method according to claim 7, characterized in that Construct a training objective that includes relationship prediction loss and consistency constraint loss, including: Sentences based on query set instances The predicted probability of the relationship category of the entity pair in Constructing Relationship Prediction Loss , expressed as: ; The consistency constraint loss between the teacher model and the student model is constructed based on the normalized relational representation of each instance in the support set and the query set. , expressed as: ; Among them, K is the support set The number of support set instances contained in each relation category in , For query set The number of query set instances contained in each relation category; and Respectively represent the normalized relational representation of the i-th instance obtained from the student model s and the teacher model t ,and , is the relation representation of the i-th instance; Comprehensive relationship prediction loss and consistency constraint loss , the final training goal is: ; in, are fixed parameters in the backbone module of the pre-trained language model, and Represent the parameters of the teacher model and the student model respectively; is the dynamic balance coefficient.

Citation Information

Patent Citations

  • Multi-cross-domain few-sample classification method based on knowledge distillation

    CN113610173A

  • Student model training method based on pre-training language model and text classification system

    CN115526332A

  • Cross-domain small sample relation extraction method and system based on enhanced contrast learning fine tuning

    CN116561308A

  • Small sample text classification method based on semi-supervised teacher-student model

    CN117150021A

  • Deep fusion multi-cross-domain few-sample classification method based on knowledge distillation

    CN118799645A