General information extraction method based on contrast supervision and cross-stage distillation

Through the method of comparative supervision and cross-stage distillation, the catastrophic forgetting problem of the general information extraction model in continuous learning is solved, and the model is efficiently adapted and precisely extracted between different tasks, which is suitable for edge device deployment.

CN120338086AActive Publication Date: 2025-07-18TIANJIN JIZHI TECH CO LTD +1

Patent Information

Application Number
CN202510813075.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-18
Publication Date
2025-07-18
Estimated Expiration
2045-06-18

AI Technical Summary

Technical Problem

The existing general information extraction model has catastrophic forgetting problems during the continuous learning process, resulting in a significant decline in the performance of the model when learning new tasks, especially when the overlap of new and old tasks is high or the differences are significant.

Method used

Using a method based on contrast supervision and cross-stage distillation, the LoRA matrix is trained in stages and compounded parameters, retaining auxiliary task knowledge, using natural language instructions to drive the model to process different tasks, freezing the weight of the base model, independently training auxiliary tasks, adding playback samples and comparison losses to force recalling old knowledge, and distillation loss constrains the output of the new model to approximate the teacher model.

Benefits of technology

It significantly reduces the loss of knowledge transfer across stages, improves the model's adaptability and accuracy to different tasks, reduces computing resource consumption, is suitable for edge device deployment, and enhances the model's fault tolerance and learning speed of instructions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120338086A_ABST
    Figure CN120338086A_ABST
Patent Text Reader

Abstract

The invention provides a general information extraction method based on contrast supervision and cross-stage distillation. The general information extraction method comprises the following steps: selecting a pre-trained model as a base model and performing initialization; initializing a first low-rank decomposition matrix and a second low-rank decomposition matrix for the weight matrix of the base model; training the first low-rank decomposition matrix and the second low-rank decomposition matrix; combining the weight of the base model, the first low-rank decomposition matrix and the second low-rank decomposition matrix to obtain a reasoning model; inputting natural language text information into the inference model; and assigning a task type and an output format through a natural language instruction, and outputting a structured information extraction result. The method has the beneficial effects that the LoRA matrix is trained in stages and the parameters are compounded, so that the target task is learned while the auxiliary task knowledge is reserved, and the loss of cross-stage knowledge migration is remarkably reduced; the natural language instruction drives the same model to process different tasks, and model architectures do not need to be switched.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of information processing, and in particular, relates to a general information extraction method based on contrastive supervision and cross-stage distillation. Background Art

[0002] General information extraction aims to implement a general model that can handle multiple information extraction tasks simultaneously, including core subtasks such as named entity recognition, relation extraction, and event extraction. Existing research designs auxiliary tasks closely related to the target task and constructs a training paradigm based on multi-stage continual learning to enhance the extraction performance of the target task. This learning paradigm based on multi-stage training gradually improves the extraction ability by sequentially learning extraction targets of different difficulties and importance levels, and has made breakthrough progress in the field of general information extraction. In particular, this progressive learning strategy can be called continual learning.

[0003] Continual Learning, also known as lifelong learning or incremental learning, aims to learn different tasks in a certain order to cope with dynamically changing knowledge and environment. Although existing research has achieved better extraction effects than traditional multi-task learning architectures based on the continual learning framework, they often also face the problem of catastrophic forgetting (CF). Catastrophic forgetting refers to the phenomenon that when the model learns a new task, the inference performance on past tasks drops significantly. In particular, when the new and old tasks have a high degree of overlap, the model may be able to better retain the memory of previous knowledge. When the new and old tasks are significantly different, the model tends to update and overwrite the previously learned knowledge, making the problem of catastrophic forgetting more prominent.

[0004] Auxiliary tasks usually provide underlying language representations for the target task and may serve as a key sub-link in the target task. For example, if the auxiliary task is to extract entity fragments, and the target task is to extract entity fragments and assign them to a predefined entity type label, then forgetting the auxiliary task will directly break the logical link between tasks, thus having a negative impact on the performance of the target task. In this case, the model tends to overfit the data distribution of the target task, resulting in poor performance in cross-task adaptive learning and task feature discrimination, ultimately affecting the overall extraction performance. However, existing general information extraction works based on multi-stage continual learning often ignore the impact of the catastrophic forgetting problem on the effectiveness of cross-stage knowledge transfer, resulting in suboptimal performance of the model during the continual learning process. Therefore, there is an urgent need for a general information extraction method to solve the above problems. Summary of the Invention

[0005] To solve the above technical problems, the present invention provides a general information extraction method based on contrastive supervision and cross-stage distillation, which is particularly suitable for coping with dynamically changing knowledge and environment through continuous learning in general information extraction.

[0006] The technical solution adopted by the present invention is as follows: In the first aspect, a general information extraction method based on contrastive supervision and cross-stage distillation is provided, including the following steps:

[0007] Select a pre-trained model as the base model and initialize it;

[0008] Initialize a first low-rank decomposition matrix and a second low-rank decomposition matrix for the weight matrix of the base model;

[0009] Train the first low-rank decomposition matrix and the second low-rank decomposition matrix;

[0010] Merge the weights of the base model, the first low-rank decomposition matrix and the second low-rank decomposition matrix to obtain an inference model;

[0011] Input the natural language text information into the inference model;

[0012] Specify the task type and output format through natural language instructions, and output the structured information extraction result.

[0013] Further, training the first low-rank decomposition matrix includes the following steps:

[0014] Freeze the weights of the base model;

[0015] Collect the first input text according to the task purpose of the first low-rank decomposition matrix;

[0016] The first input text is preprocessed to generate a first input vector;

[0017] Input the first input vector into the base model;

[0018] According to the equation Calculate the first output vector, where is the first output vector, is the first input vector, is the original weight matrix of the base model, is the first low-rank decomposition matrix;

[0019] Construct a first loss function to optimize the parameters of the first low-rank decomposition matrix.

[0020] Further, training the second low-rank decomposition matrix includes the following steps:

[0021] Freeze the weights of the base model and the first low-rank decomposition matrix;

[0022] Sample the first input text according to the sampling ratio by the random sampling method to form a replay sample;

[0023] Collect the second input text according to the task purpose of the second low-rank decomposition matrix;

[0024] The second input text and the replay sample are preprocessed to generate a second input vector;

[0025] According to the equation Calculate the second output vector, where is the second output vector, is the second input vector, is the original weight matrix of the base model, is the first low-rank decomposition matrix, is the second low-rank decomposition matrix;

[0026] Construct a second loss function to optimize the parameters of the first low-rank decomposition matrix.

[0027] Furthermore, the first loss function has the formula , where is the minimized negative log-likelihood loss, is the supervised contrastive loss, is the weight of the supervised contrastive loss term.

[0028] Furthermore, the second loss function has the formula , where is the minimized negative log-likelihood loss, is the supervised contrastive loss, is the weight of the supervised contrastive loss term, is the replay sample loss, is the distillation loss, is the weight of the distillation loss term.

[0029] Furthermore, the supervised contrastive loss has the equation , where is the triplet loss, is the sample anchor, represents one of the positive samples, represents one of the negative samples, is a margin value greater than zero, For updating the model parameters; the supervised contrastive loss introduces the Euclidean distance to quantify the distance between two input example instructions , where is the embedding vector of the instruction, denotes the L2 norm.

[0030] Furthermore, the equation of the replay sample loss is , where are the weights of the base model and the first low-rank decomposition matrix, is the parameter updated on the second low-rank decomposition matrix, is the size of the replay sample set, is the i-th sample input vector; the equation of the distillation loss is , where is the probability distribution of the weights of the base model and the first low-rank decomposition matrix, is the probability distribution of the weights of the second low-rank decomposition matrix, is the output sequence length.

[0031] In a second aspect, there is provided an electronic device, including: at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the general information extraction method based on contrastive supervision and cross-stage distillation provided by the present disclosure.

[0032] In a third aspect, there is provided a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause a computer to execute the general information extraction method based on contrastive supervision and cross-stage distillation provided by the present disclosure.

[0033] In a fourth aspect, there is provided a computer program product, including computer programs / instructions, and the computer programs / instructions are used to execute the general information extraction method based on contrastive supervision and cross-stage distillation provided by the present disclosure when being executed by a processor.

[0034] The advantages and positive effects of the present invention are as follows: By adopting the above technical solution, the LoRA matrix is trained in stages and the parameters are combined, so that while retaining the knowledge of the auxiliary task, the target task is learned, and the loss of cross-stage knowledge transfer is significantly reduced; natural language instructions drive the same model to process different tasks without switching the model architecture; freezing the weights of the base model ensures that the general language ability does not degenerate, and focuses on optimizing the special features of the auxiliary task; independent training avoids the interference of target task noise and improves the accuracy of basic extraction tasks such as entities / relationships; by adding replay samples, the knowledge of the auxiliary task is forced to be recalled at the data level; the contrast loss enables the model to map instructions with the same semantics to a similar representation space; the fault tolerance ability of the model to interference instructions is improved, and the extraction errors caused by differences in user instruction expressions are reduced; the replay loss retains the performance of the auxiliary task, the distillation loss constrains the output of the new model to approximate the teacher model (old knowledge), and the contrast loss enhances the discrimination of task instructions; the quadruple loss acts synergistically to improve the learning speed of the target task. Description of the Drawings

[0035] Figure 1 It is a schematic flowchart of a general information extraction method according to an embodiment of the present invention Detailed Embodiments

[0036] The present disclosure will be described more fully hereinafter with reference to the accompanying drawings, in which exemplary embodiments of the present disclosure are shown. The technical solutions in the embodiments of the present disclosure will be clearly and completely described below with reference to the drawings in the embodiments of the present disclosure. It is obvious that the described embodiments are only a part of the embodiments of the present disclosure, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present disclosure without creative efforts shall fall within the protection scope of the present disclosure.

[0037] As Figure 1 shown, the present invention provides a general information extraction method based on contrastive supervision and cross-stage distillation, including the following steps:

[0038] S100. Select a pre-trained model as the base model and initialize it;

[0039] The target tasks are clearly defined as Named Entity Recognition (NER), Relation Extraction (RE), Event Trigger Word Extraction (EET), and Event Argument Extraction (EEA). Meanwhile, an auxiliary task set is constructed, including entity recognition, entity classification, entity pair recognition, entity pair classification, and event element extraction. It should be noted that the target tasks are not limited to the above names and are determined by the settings of model training. Named Entity Recognition (NER) is used to identify entities of specific categories such as person names, place names, organization names, time, diseases, drugs, etc. in the text; Relation Extraction (RE) is used to identify the semantic relationships (such as "treatment", "causing", "belonging to") between two entities in the text; Event Trigger Word Extraction (EET) is used to identify the words that mark the occurrence of an event in the text (for example, in the sentence "The company announced a new merger and acquisition plan yesterday", the word "announced" can be regarded as a trigger word because it marks the occurrence of an announcement event); Event Argument Extraction (EEA) is used to identify information such as entities or times related to the event.

[0040] S200. Initialize the first low-rank decomposition matrix and the second low-rank decomposition matrix for the weight matrix of the base model;

[0041] S300. Train the first low-rank decomposition matrix and the second low-rank decomposition matrix;

[0042] S400. Combine the weights of the base model, the first low-rank decomposition matrix, and the second low-rank decomposition matrix to obtain an inference model;

[0043] S500. Input the natural language text information into the inference model;

[0044] S600. Specify the task type and output format through natural language instructions, and output the structured information extraction result.

[0045] Using the above method, by training the LoRA matrix in stages and compounding the parameters, while retaining the knowledge of the auxiliary task (i.e., the first low-rank decomposition matrix), learning the target task (i.e., the second low-rank decomposition matrix), significantly reducing the loss of cross-stage knowledge transfer; natural language instructions drive the same model to process different tasks (such as entity recognition → entity classification), without switching the model architecture; the merged single inference model reduces the consumption of computing resources and is suitable for deployment on edge devices.

[0046] To solve the problem of avoiding feature confusion caused by direct multi-task training, an implementation method is provided in this embodiment.

[0047] In one embodiment, training the first low-rank decomposition matrix includes the following steps:

[0048] Freeze the weights of the base model;

[0049] Collect the first input text according to the task objective of the first low-rank decomposition matrix;

[0050] Generate the first input vector after preprocessing the first input text;

[0051] Input the first input vector into the base model;

[0052] According to the equation Calculate the first output vector, where is the first output vector, is the first input vector, is the original weight matrix of the base model, is the first low-rank decomposition matrix;

[0053] Construct the first loss function to optimize the parameters of the first low-rank decomposition matrix.

[0054] Adopt the above method, freeze the weights of the base model to ensure that the general language ability does not degenerate, and focus on optimizing the features specific to the auxiliary task; independent training avoids the interference of target task noise and improves the accuracy of basic extraction tasks such as entities / relationships.

[0055] To solve the catastrophic forgetting problem in multi-stage continual learning, an implementation method is provided in this embodiment.

[0056] In one embodiment, training the second low-rank decomposition matrix includes the following steps:

[0057] Freeze the weights of the base model and the first low-rank decomposition matrix;

[0058] Use the random sampling method to sample the first input text according to the sampling ratio to form a replay sample, and the sampling ratio range is 10%-30%, and the sampling ratio can be adjusted according to specific situations;

[0059] Collect the second input text according to the task objective of the second low-rank decomposition matrix;

[0060] Generate the second input vector after preprocessing the second input text and the replay sample;

[0061] According to the equation Calculate the second output vector, where is the second output vector, is the second input vector, is the original weight matrix of the base model, is the first low-rank decomposition matrix, is the second low-rank decomposition matrix;

[0062] Construct the second loss function to optimize the parameters of the first low-rank decomposition matrix.

[0063] Using the above method, the knowledge of the auxiliary task is forcibly recalled from the data level by adding replay samples.

[0064] To solve the problem of poor stability of instruction diversity and output fluctuations caused by differences in instruction expressions, an implementation method is provided in this embodiment.

[0065] In one embodiment, the first loss function has the formula , where is the minimized negative log-likelihood loss, is the supervised contrastive loss, is the weight of the supervised contrastive loss term.

[0066] By minimizing the negative log-likelihood loss a basic loss function is constructed. For the auxiliary task learning process, in the training loss function, represents the pre-trained model weights, represents the change amount of the first low-rank decomposition matrix during the parameter update process. For the target task learning process, in the training loss function, represents the pre-trained model weights and the weights of the first low-rank decomposition matrix, represents the change amount of the second low-rank decomposition matrix during the parameter update process.

[0067] Using the above method, the contrastive loss enables the model to map instructions with the same semantics to a similar representation space; improves the fault tolerance of the model to interfering instructions, and reduces extraction errors caused by differences in user instruction expressions.

[0068] To solve the problems of having to face forgetting old tasks, aligning feature spaces, and distinguishing new task instructions, an implementation method is provided in this embodiment.

[0069] In one embodiment, the second loss function has the formula , where is the minimized negative log-likelihood loss, is the supervised contrastive loss, is the weight of the supervised contrastive loss term, is the replay sample loss, is the distillation loss, is the weight of the distillation loss term.

[0070] Apply the knowledge distillation technique in the learning stage of the target task (i.e., the second low-rank decomposition matrix) to alleviate the limitation of the reduced effectiveness of cross-stage knowledge transfer. Consider the fixed weights of the auxiliary task model (i.e., the first low-rank decomposition matrix) as the teacher model, and the continuously updated weights of the target task model (i.e., the second low-rank decomposition matrix) as the student model. Quantify the difference in probability distributions between the new and old models through a distillation regularization term based on the Kullback-Leibler (KL) divergence. And introduce contrastive losses for the two training stages respectively to alleviate the catastrophic forgetting problem in the continuous learning process. By constructing a positive sample set and a negative sample set , and define the current input as the anchor sample. Based on the Triplet Loss, make the positive sample pairs close to each other in the feature space and the negative sample pairs far from each other in the feature space. The contrastive loss designed in the present invention enhances the model's sensitivity and compliance ability to instructions from the perspective of embedded representation encoding, thereby alleviating the confusion of past input instructions and ensuring that the model accurately outputs the extraction results according to the instruction requirements.

[0071] Using the above method, the replay loss preserves the performance of the auxiliary task, the distillation loss constrains the output of the new model to approximate the teacher model (old knowledge), and the contrastive loss enhances the distinguishability of task instructions; the quadruple losses work together to improve the learning speed of the target task.

[0072] To solve the problem that traditional contrastive learning does not consider the logical relevance of cross-task instructions, resulting in instruction confusion during task switching, an implementation method is provided in this embodiment.

[0073] In one embodiment, the supervised contrastive loss has the equation , where is the triplet loss, is the sample anchor, represents one of the positive samples, represents one of the negative samples, is a margin value greater than zero, is the model update parameter; the supervised contrastive loss introduces the Euclidean distance to quantify the distance between the instructions of two input samples , where is the embedded vector of the instruction, represents the L2 norm.

[0074] Using the above method, the triplet loss explicitly pulls closer the cross-task instructions with consistent outputs and pushes away the irrelevant instructions; the model can seamlessly switch task instructions, and the error rate caused by instruction confusion is reduced.

[0075] To solve the problem that it is difficult to solve the feature space shift with pure data replay, and distillation lacks constraints on the original samples, an implementation method is provided in this embodiment.

[0076] In one embodiment, the replay sample loss has the equation , where is the weight of the base model and the first low-rank decomposition matrix, is the parameter updated on the second low-rank decomposition matrix, is the size of the replay sample set, is the input vector of the i-th sample; the distillation loss has the equation , where is the probability distribution of the weights of the base model and the first low-rank decomposition matrix, is the probability distribution of the weights of the second low-rank decomposition matrix, is the output sequence length.

[0077] Using the above method, the average distance between the old and new tasks in the representation space is effectively reduced, proving that feature degradation is effectively suppressed.

[0078] Next, in combination with a preferred embodiment, the content involved in the above embodiment will be described.

[0079] To verify the effectiveness of the present invention, an example verification was carried out on the standard general information extraction benchmark. The standard general information extraction benchmark covers 25 named entity recognition datasets, 10 relation extraction datasets, and 3 event extraction datasets, totaling 200 entity categories, 81 relation categories, and 40 event categories. Among them, each dataset is divided into a training set, a validation set, and a test set. To ensure the balance of samples in the corpus, a random sampling strategy is adopted, so that each dataset contains at most 10,000 instances.

[0080] The model used is the flan-t5-xxl model, the metric used for evaluation is Micro-F1 based on span shift, the contrast loss weight is 0.2, and the distillation loss weight is 0.7.

[0081] On the standard general information extraction benchmark, the flan-t5-xxl model achieved accuracies of 87.89%, 69.04%, and 72.97% in named entity recognition, relation extraction, and event extraction tasks respectively.

[0082] Furthermore, to explore the effectiveness of the learning strategy of the present invention, the present invention additionally conducted ablation experiments based on the flan-t5-large model. On the standard general information extraction benchmark, the continuous learning model with contrast loss and distillation loss constraints outperformed the traditional continuous learning method in named entity recognition, relation extraction, event trigger word extraction, and event argument extraction tasks by 3.22%, 4.27%, 5.94%, and 6.79% in terms of F1 score, respectively, fully verifying the effectiveness of the present invention.

[0083] Based on the embodiments of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product.

[0084] An electronic device includes at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the general information extraction method based on contrastive supervision and cross-stage distillation provided by the present disclosure.

[0085] The electronic device is intended to represent various forms of digital computers, such as, laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as, personal digital processors, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0086] A non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause a computer to execute the general information extraction method based on contrastive supervision and cross-stage distillation provided by the present disclosure.

[0087] The various embodiments in the present disclosure may be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), system-on-chip systems (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include: implemented in one or more computer programs that may be executed and / or interpreted on a programmable system including at least one programmable processor, the programmable processor may be a dedicated or general-purpose programmable processor, and may receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit the data and instructions to the storage system, the at least one input device, and the at least one output device.

[0088] A computer program product includes computer programs / instructions which, when executed by a processor, implement the general information extraction method based on contrastive supervision and cross-stage distillation provided by the present disclosure.

[0089] The program code for implementing the method of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing devices, such that when the program codes are executed by the processor or controller, the functions / operations specified in the flowchart and / or block diagram are implemented. The program code can be executed entirely on the machine, partially on the machine, executed partially on the machine as an independent software package and partially on a remote machine, or executed entirely on a remote machine or server.

[0090] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media would include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disc read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0091] The above has described the embodiments of the present invention in detail, but the above content is only the preferred embodiments of the present invention and cannot be considered as defining the scope of implementation of the present invention. All equivalent changes and improvements made in accordance with the scope of the application of the present invention should still fall within the scope covered by the patent of the present invention.

Claims

1. A general information extraction method based on contrastive supervision and cross-stage distillation, characterized in that It includes the following steps: Select a pre-trained model as the base model and initialize it; Initialize the first low-rank decomposition matrix and the second low-rank decomposition matrix for the weight matrix of the base model; Train the first low-rank decomposition matrix and the second low-rank decomposition matrix; Merge the weights of the base model, the first low-rank decomposition matrix and the second low-rank decomposition matrix to obtain an inference model; Input natural language text information into the inference model; Specify the task type and output format through natural language instructions, and output the structured information extraction result.

2. The general information extraction method according to claim 1, wherein Training the first low-rank decomposition matrix includes the following steps: Freeze the weights of the base model; Collect the first input text according to the task purpose of the first low-rank decomposition matrix; The first input text generates a first input vector after preprocessing; Input the first input vector into the base model; According to the equation calculate the first output vector, where is the first output vector, is the first input vector, is the original weight matrix of the base model, is the first low-rank decomposition matrix; Construct a first loss function to optimize the parameters of the first low-rank decomposition matrix.

3. The general information extraction method according to claim 2, wherein Training the second low-rank decomposition matrix includes the following steps: Freeze the weights of the base model and the first low-rank decomposition matrix; Use the random sampling method to sample the first input text according to the extraction ratio to form a replay sample; Collect the second input text according to the task purpose of the second low-rank decomposition matrix; The second input text and the replay sample generate a second input vector after preprocessing; According to the equation calculate the second output vector, where is the second output vector, is the second input vector, is the original weight matrix of the base model, is the first low-rank decomposition matrix, is the second low-rank decomposition matrix; Construct a second loss function to optimize the parameters of the first low-rank decomposition matrix.

4. The general information extraction method according to claim 2 or 3, characterized in that: The first loss function has the formula , where is the minimized negative log-likelihood loss, is the supervised contrastive loss, is the weight of the supervised contrastive loss term.

5. The general information extraction method according to claim 3, wherein: The second loss function has the formula , where is the minimized negative log-likelihood loss, is the supervised contrastive loss, is the weight of the supervised contrastive loss term, is the replay sample loss, is the distillation loss, is the weight of the distillation loss term.

6. The general information extraction method according to claim 5, characterized in that: The supervised contrastive loss has the following equation , where is the triplet loss, is the sample anchor, represents one of the positive samples, represents one of the negative samples, is a margin value greater than zero, is the model update parameter; the supervised contrastive loss introduces the Euclidean distance to quantify the distance between two input example instructions , where is the embedding vector of the instruction, represents the L2 norm.

7. The general information extraction method according to claim 5, characterized in that: The replay sample loss has the equation , where are the weights of the base model and the first low-rank decomposition matrix, are the parameters updated on the second low-rank decomposition matrix, is the size of the replay sample set, is the i-th sample input vector; the distillation loss has the equation , where is the probability distribution of the weights of the base model and the first low-rank decomposition matrix, is the probability distribution of the weights of the second low-rank decomposition matrix, is the output sequence length.

8. An electronic device, comprising: At least one processor; And A memory communicatively connected to at least one processor; wherein, The memory stores instructions executable by at least one processor, and the instructions are executed by at least one processor so that at least one processor can execute the method according to any one of claims 1 to 7.

9. A non-transitory computer-readable storage medium storing computer instructions, wherein, Computer instructions are used to cause a computer to execute the method according to any one of claims 1 to 7.

10. A computer program product, comprising computer programs / instructions which, when executed by a processor, implement the steps of the method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Face recognition method and device based on Insight Face and LIS file management

    CN118230395A

  • Low-Rank Adaptation of Neural Network Models

    US20220383126A1

  • Multitask prompt tuning for parameter-efficient transfer learning

    US20250005370A1

Cited By

  • General information extraction method based on task specific mixed low-rank adaptation

    CN121168579A