Model training method and device, electronic equipment, storage medium and program product

By using the knowledge retention mechanism for model training in text classification tasks, the problem that knowledge retention technology in the existing technology is not applied to text classification is solved, and the effect of the model retaining old knowledge while learning new knowledge is achieved.

CN120045704APending Publication Date: 2025-05-27TENCENT TECHNOLOGY (SHENZHEN) CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311536585.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-11-16
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

In the prior art, knowledge retention technology is mainly applied to the field of image classification and has not yet been extended to the field of text classification in natural language processing. The linear layer classifiers used in related art ignore commonalities and differences between categories.

Method used

A model training method is proposed, which obtains text samples for pre-training, and obtains language model and prototype vectors; then transfer training is carried out based on the knowledge retention mechanism, and the language model and prototype vectors are updated to ensure that the model retains old knowledge while learning new knowledge.

Benefits of technology

It realizes the effective preservation of knowledge in text classification tasks, avoids the model quickly forgetting old knowledge, and optimizes the learning effect of the language model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120045704A_ABST
    Figure CN120045704A_ABST
Patent Text Reader

Abstract

The invention provides a model training method and device, electronic equipment, a computer readable storage medium and a computer program product, and the method comprises the steps: obtaining a first text sample, and carrying out the pre-training processing of a first language model and a first prototype vector corresponding to an original type based on the first text sample, obtaining a second language model and a second prototype vector corresponding to the original type; obtaining a third prototype vector and a second text sample corresponding to the new type; and based on the second text sample, carrying out migration training processing based on a knowledge retention mechanism on a second language model, a second prototype vector and a third prototype vector to obtain a third language model, a fourth prototype vector corresponding to an original type and a fifth prototype vector corresponding to a new type. Through the method and the device, knowledge retention can be realized in the text classification task.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technologies, and in particular, to a model training method, apparatus, electronic device, computer-readable storage medium, and computer program product. Background Art

[0002] Resisting catastrophic forgetting of old data while a model continuously learns new data is the focus of research on the incremental learning problem. In related technologies, maintaining a virtual memory and retaining old knowledge about image classification through knowledge distillation are proposed. In related technologies, a three-stage learning framework is also proposed to solve the generalized few-shot image classification task and the incremental learning few-shot classification task: in the first stage, the model is trained on data of visible classes; in the second stage, a new classification head is added, and then the data of new classes is used to continue training on this base model; in the third stage, both old data and new data are used for training simultaneously, and parameter constraints are supplemented to help the model not forget the knowledge learned in the first stage.

[0003] The knowledge retention technology used in related technologies is applied to the field of image classification and has not been extended to the field of text classification in natural language processing. Summary of the Invention

[0004] Embodiments of this application provide a model training method, apparatus, electronic device, computer-readable storage medium, and computer program product, which can achieve knowledge retention in text classification tasks.

[0005] The technical solution of the embodiments of this application is implemented as follows:

[0006] Embodiments of this application provide a model training method, including:

[0007] Obtain a first text sample, and perform pre-training processing on a first language model and a first prototype vector of the corresponding original type based on the first text sample to obtain a second language model and a second prototype vector of the corresponding original type;

[0008] Obtain a third prototype vector of the corresponding new type and a second text sample;

[0009] Based on the second text sample, perform transfer training processing on the second language model, the second prototype vector, and the third prototype vector based on a knowledge retention mechanism to obtain a third language model, a fourth prototype vector of the corresponding original type, and a fifth prototype vector of the corresponding new type;

[0010] Wherein, the knowledge retention mechanism includes at least one of a model parameter constraint mechanism and an original type constraint mechanism.

[0011] Embodiments of this application provide a text processing method, including:

[0012] Obtain the input text;

[0013] Perform text representation processing on the input text to obtain an input text representation, and add a classification label to the input text representation to obtain a text representation to be processed;

[0014] Perform mapping processing on the text representation to be processed through a third language model to obtain a fifth feature vector at the classification label, where the third language model is trained through the model training method provided by the embodiments of the present application;

[0015] Determine the fifth similarity between the fifth feature vector and the fourth prototype vector of each original type and the fifth prototype vector of each new type, where the fourth prototype vector and the fifth prototype vector are trained through the model training method provided by the embodiments of the present application;

[0016] Take the original type or new type corresponding to the maximum fifth similarity as the classification result of the input text.

[0017] An embodiment of the present application provides a model training device, including:

[0018] A pre-training module, configured to obtain a first text sample, and perform pre-training processing on a first language model and a first prototype vector corresponding to an original type based on the first text sample to obtain a second language model and a second prototype vector corresponding to the original type;

[0019] An acquisition module, configured to acquire a third prototype vector corresponding to a new type and a second text sample;

[0020] A knowledge transfer module, configured to perform transfer training processing based on the knowledge retention mechanism on the second language model, the second prototype vector, and the third prototype vector based on the second text sample to obtain a third language model, a fourth prototype vector corresponding to the original type, and a fifth prototype vector corresponding to the new type, where the knowledge retention mechanism includes at least one of a model parameter constraint mechanism and an original type constraint mechanism.

[0021] In the above solution, the pre-training module is further configured to obtain an original text sample, perform text representation processing on the original text sample to obtain a text representation, and add a classification label to the text representation to obtain the first text sample.

[0022] In the above solution, the pre-training module is further configured to classify the first text sample through the first language model to obtain a first predicted classification result, calculate a first loss function based on the first labeled classification label of the first text sample and the first predicted classification result, and update the first language model and the first prototype vector based on the first loss function to obtain the second language model and the second prototype vector.

[0023] In the above solution, the pre-training module is further configured to map the first text sample through the first language model to obtain a first feature vector at the classification label, determine a first similarity between the first feature vector and the first prototype vector of each of the original types, and use the first similarity corresponding to each of the original types as the first predicted classification result.

[0024] In the above solution, the pre-training module is further configured to obtain a third text sample belonging to the same original type as the first text sample, obtain a second feature vector of the third text sample through the first language model, calculate a first sample contrast loss function based on the error between the first feature vector and the second feature vector, calculate a first prototype contrast loss function based on the error between the first feature vector and the first prototype vector of the original type to which the first text sample belongs, calculate a first classification loss function based on the error between the first labeled classification label of the first text sample and the first predicted classification result, and fuse the first sample contrast loss function, the first prototype contrast loss function, and the first classification loss function to obtain the first loss function.

[0025] In the above solution, the transfer training module is further configured to classify the second text sample through the second language model to obtain a second predicted classification result, calculate a second loss function based on the second labeled classification label of the second text sample and the second predicted classification result, obtain a knowledge retention loss function corresponding to the knowledge retention mechanism based on the second language model before processing the second text sample, fuse the second loss function and the knowledge retention loss function to obtain a transfer learning loss function, and update the second language model, the second prototype vector, and the third prototype vector based on the transfer learning loss function to obtain the third language model, the fourth prototype vector, and the fifth prototype vector.

[0026] In the above solution, the transfer training module is further configured to perform mapping processing on the second text sample through the second language model to obtain a third feature vector at the classification tag, determine a second similarity between the third feature vector and the first prototype vector of each of the original types and the third prototype vector of each of the new types, and use the second similarity corresponding to each of the original types and each of the new types as the second predicted classification result.

[0027] In the above solution, the transfer training module is further configured to obtain a fourth text sample that belongs to the same original type as the second text sample, obtain a fourth feature vector of the fourth text sample through the second language model, calculate a second sample contrast loss function based on the error between the third feature vector and the fourth feature vector, calculate a second prototype contrast loss function based on the error between the third feature vector and the third prototype vector of the original type to which the second text sample belongs, calculate a second classification loss function based on the error between the second labeled classification label of the second text sample and the second predicted classification result, and perform fusion processing on the second sample contrast loss function, the second prototype contrast loss function, and the second classification loss function to obtain the second loss function.

[0028] In the above solution, the transfer training module is further configured to obtain the first model parameters of the second language model before processing the second text sample and the second model parameters of the second language model at the current moment, and determine a knowledge retention loss function corresponding to the model parameter constraint mechanism based on the error between the first model parameters and the second model parameters.

[0029] In the above solution, the transfer training module is further configured to classify the first text sample through the second language model before processing the second text sample to obtain a third similarity corresponding to each of the original types and each of the new types, classify the first text sample through the second language model to obtain a fourth similarity corresponding to each of the original types and each of the new types, and determine a knowledge retention loss function corresponding to the original type constraint mechanism based on the third similarity and the fourth similarity.

[0030] In the above solution, the transfer training module is further configured to obtain an input text, perform text representation processing on the input text to obtain an input text representation, add a classification label to the input text representation to obtain a text representation to be processed, perform mapping processing on the text representation to be processed through the third language model to obtain a fifth feature vector at the classification label, determine a fifth similarity between the fifth feature vector and the fourth prototype vector of each original type and the fifth prototype vector of each new type, and use the original type or new type corresponding to the maximum fifth similarity as the classification result of the input text.

[0031] An embodiment of the present application provides a text processing device, including:

[0032] A text acquisition module, configured to acquire an input text;

[0033] A text representation processing module, configured to perform text representation processing on the input text to obtain an input text representation, and add a classification label to the input text representation to obtain a text representation to be processed;

[0034] A mapping processing module, configured to perform mapping processing on the text representation to be processed through a third language model to obtain a fifth feature vector at the classification label, where the third language model is trained through the model training method provided by the embodiment of the present application;

[0035] A similarity determination module, configured to determine a fifth similarity between the fifth feature vector and the fourth prototype vector of each original type and the fifth prototype vector of each new type, where the fourth prototype vector and the fifth prototype vector are trained through the model training method provided by the embodiment of the present application;

[0036] A result output module, configured to use the original type or new type corresponding to the maximum fifth similarity as the classification result of the input text.

[0037] An embodiment of the present application provides an electronic device, including:

[0038] A memory, configured to store computer-executable instructions;

[0039] A processor, configured to implement the model training method and the text processing method provided by the embodiment of the present application when executing the computer-executable instructions stored in the memory.

[0040] An embodiment of the present application provides a computer-readable storage medium, storing computer-executable instructions, which are used to cause a processor to implement the model training method and the text processing method provided by the embodiment of the present application when executed.

[0041] An embodiment of the present application provides a computer program product, including computer-executable instructions, which, when executed by a processor, implement the model training method and text processing method provided by the embodiment of the present application.

[0042] The embodiment of the present application has the following beneficial effects:

[0043] Through the pre-training of classifying the language model, the language model learns the classification knowledge related to the original type, and after adding new types, the language model is trained to learn the classification knowledge related to the new types. At the same time, the language model is subjected to transfer training processing based on the knowledge retention mechanism of at least one of the model parameter constraint mechanism and the original type constraint mechanism, so that the language model retains the original type classification knowledge learned in the pre-training stage while learning the new type knowledge, avoiding the language model quickly forgetting old knowledge and optimizing the learning effect of the language model. BRIEF DESCRIPTION OF THE DRAWINGS

[0044] Figure 1 It is a schematic structural diagram of the model training system architecture provided by the embodiment of the present application;

[0045] Figure 2 It is a schematic structural diagram of the server provided by the embodiment of the present application;

[0046] Figure 3A It is a schematic flowchart of the model training method provided by the embodiment of the present application;

[0047] Figure 3B It is an alternative flowchart of the model training method provided by the embodiment of the present application;

[0048] Figure 3C It is an alternative flowchart of the model training method provided by the embodiment of the present application;

[0049] Figure 3D It is an alternative flowchart of the model training method provided by the embodiment of the present application;

[0050] Figure 3E It is an alternative flowchart of the model training method provided by the embodiment of the present application;

[0051] Figure 3F It is an alternative flowchart of the model training method provided by the embodiment of the present application;

[0052] Figure 3G It is an alternative flowchart of the model training method provided by the embodiment of the present application;

[0053] Figure 3H It is an alternative flowchart of the model training method provided by the embodiment of the present application;

[0054] Figure 4 It is a schematic flowchart of the text processing method provided by the embodiments of the present application;

[0055] Figure 5 It is a schematic framework diagram of the model training scheme provided by the embodiments of the present application;

[0056] Figure 6 It is a data graph of the experimental results provided by the embodiments of the present application;

[0057] Figure 7 It is a line graph of the test performance data provided by the embodiments of the present application. Detailed implementation manners

[0058] In order to make the objectives, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings. The described embodiments should not be construed as limiting the present application. All other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the scope of protection of the present application.

[0059] In the following description, reference is made to "some embodiments", which describe a subset of all possible embodiments. However, it can be understood that "some embodiments" can be the same subset or different subsets of all possible embodiments, and can be combined with each other without conflict.

[0060] In the following description, the terms "first / second / third" are only used to distinguish similar objects and do not represent a specific order for the objects. It can be understood that "first / second / third" can be interchanged with a specific order or sequence when allowed, so that the embodiments of the present application described herein can be implemented in an order other than that illustrated or described herein.

[0061] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which this application belongs. The terms used herein are only for the purpose of describing the embodiments of the present application and are not intended to limit the present application.

[0062] Before further elaborating on the embodiments of the present application, the nouns and terms involved in the embodiments of the present application are described. The nouns and terms involved in the embodiments of the present application are subject to the following explanations.

[0063] 1) One-shot learning: It means that during the training process, the model can only use 1 sample to learn a new task. This sample contains the original text and the corresponding label.

[0064] 2) Five-shot learning: It means that during the training process, the model uses 5 samples to learn a new task. This sample contains the original text and the corresponding label.

[0065] 3) Data-Dependent Knowledge Preservation (DDKP): A method we proposed, which is used to preserve old knowledge when learning new types of knowledge.

[0066] 4) Data-Agnostic Knowledge Prservation (DAKP): A method we proposed, which is used to preserve old knowledge when learning new types of knowledge.

[0067] The following solutions have been proposed for knowledge preservation techniques in related technologies:

[0068] 1) Maintain a virtual memory to prevent the model from quickly forgetting the image classification knowledge learned in the past through knowledge distillation. In addition, they also proposed to update the data in the virtual memory: create a prototype set (i.e., the virtual memory is used to store real old samples), which calculates the center of each type, selects a certain number of samples closest to the center and adds them to this prototype set. The storage space of the prototype set is limited. Therefore, when the number of samples that can be stored in the storage space reaches the maximum value, the old samples in the prototype set will be reduced, that is, the samples relatively far from the class center will be removed.

[0069] 2) Use a three-stage learning framework to solve the generalized few-shot image classification task and the incremental learning few-shot classification task: In the first stage, the model is trained on the data of visible classes. In the second stage, a new classification head is added, and then the data of the new type is used to continue training on this basic model. In the third stage, both the old data and the new data are used for training, and parameter constraints are supplemented to help the model not forget the knowledge learned in the first stage.

[0070] The applicant found the following defects in the related technologies when implementing the embodiments of the present application:

[0071] Currently, the knowledge preservation techniques in related technologies are applied to the field of image classification and have not been extended to the text classification field of natural language processing. In addition, using a linear layer as a classifier in related technology 2) will ignore the commonalities and differences between classes.

[0072] Embodiments of the present application provide a model training method, apparatus, electronic device, computer-readable storage medium, and computer program product, which can achieve knowledge retention in text classification tasks. The following describes an exemplary application of the electronic device provided by the embodiments of the present application. The electronic device provided by the embodiments of the present application can be implemented as various types of user terminals such as laptop computers, tablet computers, desktop computers, set-top boxes, mobile devices (e.g., mobile phones, portable music players, personal digital assistants, dedicated messaging devices, portable game devices), etc., or can also be implemented as a server. The following describes an exemplary application of the electronic device provided by the embodiments of the present application. The electronic device provided by the embodiments of the present application can be implemented as a terminal or a server.

[0073] See Figure 1 , Figure 1 FIG. is a schematic architecture diagram of the model training system 100 provided by the embodiments of the present application. To support a model training application, the terminal 400 is connected to the server 200 through the network 300. The network 300 can be a wide area network, a local area network, or a combination of the two.

[0074] The terminal 400 (running a client that can perform model training) is used to obtain a model training request. For example, the user generates a model training request through the input interface of the terminal 400. The server 200, according to the model training request, obtains a first text sample, and performs pre-training processing on the first language model and the first prototype vector of the corresponding original type based on the first text sample to obtain a second language model and the second prototype vector of the corresponding original type, obtains the third prototype vector of the corresponding new type and the second text sample, and performs transfer training processing based on the knowledge retention mechanism on the second language model, the second prototype vector, and the third prototype vector based on the second text sample to obtain a third language model, the fourth prototype vector of the corresponding original type, and the fifth prototype vector of the corresponding new type.

[0075] In some embodiments, the server 200 can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or can also be a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms. The terminal 400 can be a smart phone, a tablet computer, a laptop computer, a desktop computer, a smart speaker, a smart watch, etc., but is not limited thereto. The terminal and the server can be directly or indirectly connected through wired or wireless communication methods, which are not limited in the embodiments of the present invention.

[0076] See Figure 2 , Figure 2 FIG. is a schematic structure diagram of the server 200 provided by the embodiments of the present application.Figure 2 The terminal 400 shown includes: at least one processor 210, a memory 250, at least one network interface 220 and a user interface 230. The various components in the server 200 are coupled together via a bus system 240. It is understood that the bus system 240 is used to achieve connection and communication between these components. In addition to the data bus, the bus system 240 also includes a power bus, a control bus and a status signal bus. However, for the sake of clarity, the bus system 240 is not described in detail. Figure 2 Various buses are labeled as bus system 240 .

[0077] The processor 210 can be an integrated circuit chip with signal processing capabilities, such as a general-purpose processor, a digital signal processor (DSP), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc., where the general-purpose processor can be a microprocessor or any conventional processor, etc.

[0078] The user interface 230 includes one or more output devices 231 that enable presentation of media content, including one or more speakers and / or one or more visual display screens. The user interface 230 also includes one or more input devices 232, including user interface components that facilitate user input, such as a keyboard, mouse, microphone, touch screen display, camera, other input buttons and controls.

[0079] The memory 250 may be removable, non-removable, or a combination thereof. Exemplary hardware devices include solid-state memory, hard disk drives, optical disk drives, etc. The memory 250 may optionally include one or more storage devices that are physically remote from the processor 210.

[0080] The memory 250 includes a volatile memory or a non-volatile memory, and may also include both volatile and non-volatile memories. The non-volatile memory may be a read-only memory (ROM), and the volatile memory may be a random access memory (RAM). The memory 250 described in the embodiments of the present application is intended to include any suitable type of memory.

[0081] In some embodiments, memory 250 can store data to support various operations, examples of which include programs, modules, and data structures, or a subset or superset thereof, as exemplarily described below.

[0082] Operating system 251, including system programs for processing various basic system services and performing hardware-related tasks, such as a framework layer, a core library layer, a driver layer, etc., for implementing various basic services and processing hardware-based tasks;

[0083] A network communication module 252 for reaching other computing devices via one or more (wired or wireless) network interfaces 220. Exemplary network interfaces 220 include: Bluetooth, Wireless Fidelity (WiFi), and Universal Serial Bus (USB), etc.

[0084] In some embodiments, the model training device provided by the embodiments of the present application can be implemented in software. Figure 2 Shown is a model training device 253 stored in the memory 250, which can be software in the form of a program and a plug-in, etc., including the following software modules: a pre-training module 2531, an acquisition module 2532, and a transfer training module 2533. These modules are logical, so they can be combined arbitrarily or further split according to the functions to be implemented. The functions of each module will be described below.

[0085] In other embodiments, the model training device provided by the embodiments of the present application can be implemented in hardware. As an example, the model training device provided by the embodiments of the present application can be a processor in the form of a hardware decoding processor, which is programmed to execute the model training method provided by the embodiments of the present application. For example, a processor in the form of a hardware decoding processor can adopt one or more Application Specific Integrated Circuits (ASICs), DSPs, Programmable Logic Devices (PLDs), Complex Programmable Logic Devices (CPLDs), Field-Programmable Gate Arrays (FPGAs), or other electronic components.

[0086] In some embodiments, a terminal or a server can implement the model training method provided by the embodiments of the present application by running a computer program. For example, the computer program can be a native program or a software module in an operating system; it can be a native application (APP), that is, a program that needs to be installed in the operating system to run, such as a model training APP; it can also be a small program, that is, a program that only needs to be downloaded to a browser environment to run; it can also be a small program that can be embedded in any APP. In short, the above computer program can be any form of application program, module, or plug-in.

[0087] The model training method provided by the embodiments of the present application will be described in combination with the exemplary applications and implementations of the server provided by the embodiments of the present application.

[0088] See Figure 3A , Figure 3A which is a schematic flowchart of the model training method provided by an embodiment of the present application, and will be described in conjunction with Figure 3A the steps 101 to 103 shown.

[0089] In step 101, a first text sample is obtained, and based on the first text sample, pre-training processing is performed on a first language model and a first prototype vector of the corresponding original type to obtain a second language model and a second prototype vector of the corresponding original type.

[0090] Refer to Figure 3B , Figure 3B which is an alternative schematic flowchart of the model training method provided by an embodiment of the present application. In some embodiments,[[]] Figure 3A step 101 in Figure 3B can be implemented through

[0091] the steps 1011 and 1012 shown, and the details will be described below.

[0092] In step 1011, a first text sample is obtained.

[0093] As an example, an original text sample x = w 1 ,..., w n is obtained. The original text sample x is subjected to text representation processing to obtain a text representation. By introducing the addition template function T(·), a classification label is added to the text representation through x prompt = T(x) to obtain a first text sample x prompt , and its specific form is x. The intent is to[MASK], where [MASK] is the classification label.

[0094] Through the embodiments of the present application, based on prompt learning, the classification task of the original text sample is transformed into a fill-in-the-blank masked language modeling problem, which helps to effectively use the old knowledge related to the original type learned by the first language model in the pre-training stage.

[0095] In step 1012, based on the first text sample, pre-training processing is performed on the first language model and the first prototype vector of the corresponding original type to obtain a second language model and a second prototype vector of the corresponding original type.

[0096] Refer to Figure 3C , Figure 3CIt is an optional process schematic diagram of the model training method provided by the embodiments of the present application. In some embodiments, Figure 3B Step 1012 in Figure 3C can be implemented through

[0097] Steps 10121 to 10123 in

[0098] which will be described in detail below.

[0099] In step 10121, the first text sample is classified by the first language model to obtain the first predicted classification result.

[0100] As an example, the first language model is a language model that has not been pre-trained. For example, an initialized general language model. [MASK] :

[0101] h [MASK] = M(T(x)) (1)

[0102] where M is the first language model. The first feature vector h [MASK] is mapped to the vector space where the first prototype vectors of the original types are located through the linear transformation matrix W to obtain v [MASK] , v [MASK] is the first feature vector of h [MASK] mapped to the vector space where the first prototype vectors of the original types are located. Each original type corresponds to a standard classification label in the label definition set y of the original types. The first prototype vector of the original type is also a feature vector. Classification is achieved by calculating the first similarity between v [MASK] and the first prototype vector, and the probability distribution of the first similarity corresponding to each original type is used as the first predicted classification result.

[0103] Through the embodiments of the present application, the similarity between the classification mark of the first text sample and each original prototype is evaluated according to the similarity between the first feature vector and the first prototype vector of each original prototype, and the probability distribution of the similarity between the classification mark of the first text sample and each original prototype is used as the first prediction result, enabling the first language model to effectively learn the old knowledge related to the original types during the pre-training stage.

[0104] In step 10122, based on the first labeled classification label and the first predicted classification result of the first text sample, calculate the first loss function.

[0105] In some embodiments, step 10122 can be implemented as follows: Obtain a third text sample that belongs to the same original type as the first text sample, and obtain the second feature vector of the third text sample through the first language model; Calculate the first sample contrast loss function based on the error between the first feature vector and the second feature vector; Calculate the first prototype contrast loss function based on the error between the first feature vector and the first prototype vector of the original type to which the first text sample belongs; Calculate the first classification loss function based on the error between the first labeled classification label and the first predicted classification result of the first text sample; Perform a fusion process on the first sample contrast loss function, the first prototype contrast loss function, and the first classification loss function to obtain the first loss function.

[0106] As an example, the first feature vector of the first text sample is v j , and the second feature vector of the third text sample is v i , and calculate the first sample contrast loss function through the following formula:

[0107]

[0108] Calculate the first sample contrast loss function through the following formula:

[0109]

[0110] where c k is the first prototype vector corresponding to the kth original type, c i is the first prototype vector corresponding to the ith original type, T is the total number of samples, and C is the total number of types. Calculate the first classification loss function L cls based on the cross-entropy loss function between the first labeled classification label and the first predicted classification result of the first text sample, and perform a fusion process on the first sample contrast loss function, the first prototype contrast loss function, and the first classification loss function through the following formula:

[0111] L = L cls + L ii + L is (4)

[0112] where L is the first loss function.

[0113] Through the embodiments of the present application, the first classification loss function, the first sample contrast loss function, and the first prototype contrast loss function are calculated. Using the method of supervised learning, not only does the first language model learn the original type knowledge, but also the differences between different original types are considered. The functions of the first sample contrast loss function and the first prototype contrast loss function are to update the first prototype vectors of the original types, increasing the differences between the updated first prototype vectors and other first prototype vectors to improve the accuracy of the classification results.

[0114] In step 10123, based on the first loss function, update processing is performed on the first language model and the first prototype vectors to obtain a second language model and second prototype vectors.

[0115] As an example, based on the first loss function L, update processing is performed on the model parameters of the first language model and the first prototype vectors corresponding to each original type. Among them, through the update processing of the first language model, a second language model is obtained, and through the update processing of the first prototype vectors of each original type, second prototype vectors corresponding to each original type are obtained.

[0116] Through the embodiments of the present application, the first language model is used to classify the first text sample, and the first loss function is calculated based on the first predicted classification result obtained from the classification processing. The first speech model is updated with the first loss function, so that the updated second language model has the text classification ability, and the initialized first prototype vectors are updated, making the parameters of the updated second prototype vectors close to the vector parameters required for the classification task.

[0117] Continue to refer to Figure 3A , in step 102, a third prototype vector corresponding to the new type and a second text sample are obtained.

[0118] As an example, the new type is a classification prototype different from the original type, and the second text sample is a text sample different from the first text sample.

[0119] In step 103, based on the second text sample, transfer training processing based on the knowledge retention mechanism is performed on the second language model, the second prototype vectors, and the third prototype vectors to obtain a third language model, fourth prototype vectors corresponding to the original types, and fifth prototype vectors corresponding to the new type.

[0120] As an example, the knowledge retention mechanism includes at least one of a model parameter constraint mechanism and an original type constraint mechanism.

[0121] Refer to Figure 3D , Figure 3D is an optional process schematic diagram of the model training method provided by the embodiments of the present application. In some embodiments,Figure 3A Step 103 in Figure 3D can be implemented by steps 1031 to 1035 in

[0122] In step 1031, the second text sample is classified by the second language model to obtain a second predicted classification result.

[0123] Referring to Figure 3E , Figure 3E is an optional process schematic diagram of the model training method provided by the embodiments of the present application. In some embodiments, Figure 3D step 1031 in Figure 3E can be implemented by steps 10311 and 10312 in

[0124] In step 10311, the second text sample is mapped by the second language model to obtain a third feature vector at the classification label.

[0125] In step 10312, the second similarity between the third feature vector and the first prototype vector of each original type and the third prototype vector of each new type is determined, and the second similarity corresponding to each original type and each new type is used as the second predicted classification result.

[0126] As an example, the mapping process, the second similarity calculation, and the second predicted classification result determination process in this step are the same as those in step 10121, and will not be elaborated here.

[0127] Through the embodiments of the present application, according to the similarity between the third feature vector of the second text sample and the first prototype vector of each original prototype and the third prototype vector of each new type, the similarity between the classification label of the second text sample and each original prototype and each new type is evaluated, and the probability distribution of the similarity between the classification label of the second text sample and each original prototype and each new type is used as the second predicted classification result, so that the second language model can effectively learn the old knowledge related to the original type and the new type during the new knowledge learning stage.

[0128] Continuing to refer to Figure 3D , in step 1032, a second loss function is calculated based on the second labeled classification label and the second predicted classification result of the second text sample.

[0129] Referring to Figure 3F , Figure 3F is an optional process schematic diagram of the model training method provided by the embodiments of the present application. In some embodiments, Figure 3D step 1032 in Figure 3F can be implemented by steps 10321 to 10325 in

[0130] In step 10321, obtain a fourth text sample that belongs to the same original type as the second text sample, and obtain a fourth feature vector of the fourth text sample through a second language model.

[0131] As an example, through the above mapping processing method, perform mapping processing on the fourth text sample to obtain a fourth feature vector.

[0132] In step 10322, calculate a second sample comparison loss function based on the error between the third feature vector and the fourth feature vector.

[0133] As an example, the method for calculating the second sample loss comparison function in this step is the same as that in step 10122, and will not be elaborated here.

[0134] In step 10323, calculate a second prototype comparison loss function based on the error between the third feature vector and the third prototype vector of the original type to which the second text sample belongs.

[0135] As an example, the method for calculating the second sample loss comparison function in this step is the same as that in step 10122, and will not be elaborated here.

[0136] In step 10324, calculate a second classification loss function based on the error between the second labeled classification label and the second predicted classification result of the second text sample.

[0137] As an example, the method for calculating the second sample loss comparison function in this step is the same as that in step 10122, and will not be elaborated here.

[0138] In step 10325, perform fusion processing on the second sample comparison loss function, the second prototype comparison loss function, and the second classification loss function to obtain a second loss function.

[0139] As an example, the method for calculating the second sample loss comparison function in this step is the same as that in step 10122, and will not be elaborated here.

[0140] Through the embodiments of the present application, calculate a second loss function based on the second predicted classification result obtained through classification processing. The second loss function corresponds to the first loss function in the pre-training stage. The calculated second loss function is used for subsequent fusion with the knowledge retention loss function and jointly serves as a transfer learning loss function for updating the second speech model, to update the second language model, so that the updated second language model has the text classification ability for new type knowledge, and update the second prototype vector and the third prototype vector, so that the parameters of the updated fourth prototype vector and fifth prototype vector are close to the vector parameters required for the classification task.

[0141] Continue to refer to Figure 3D In step 1033, based on the second language model before processing the second text sample, obtain the knowledge retention loss function corresponding to the knowledge retention mechanism.

[0142] Refer to Figure 3G , Figure 3G is an alternative flowchart of the model training method provided by the embodiments of the present application. In some embodiments, Figure 3D Step 1033 in Figure 3G can be implemented by step 10331A and step 10332A in

[0143] In step 10331A, the knowledge retention mechanism can be a model parameter constraint mechanism, and obtain the first model parameters of the second language model before processing the second text sample and the second model parameters of the second language model at the current moment.

[0144] As an example, for an application scenario where old data cannot be accessed repeatedly, since old data cannot be obtained, a data-agnostic knowledge preservation method (DAKP, Data-Agnostic Knowledge Prservation) is adopted, and the constraint loss function is calculated through the following formula to constrain the model parameters of the second language model and the second prototype vector of the corresponding original type:

[0145]

[0146] where p joint are the model parameters of the second language model at the current moment (i.e., the model parameters of the second language model when learning new type knowledge), and p seen are the model parameters of the second language model at the previous moment (i.e., the model parameters of the second language model when new type knowledge has not been learned).

[0147] In step 10332A, based on the error between the first model parameters and the second model parameters, determine the knowledge retention loss function corresponding to the model parameter constraint mechanism.

[0148] Through the embodiments of the present application, in an application scenario where old data cannot be obtained, an explicit weight constraint method is used to enable the second language model to retain the old knowledge of the original type learned during the pre-training process during the process of learning new type knowledge.

[0149] Refer to Figure 3H , Figure 3H is an alternative flowchart of the model training method provided by the embodiments of the present application. In some embodiments, Figure 3D Step 1033 inFigure 3H are implemented by steps 10331B to 10333B in

[0150] In step 10331B, the first text sample is classified by the second language model before processing the second text sample to obtain the third similarity corresponding to each original type and each new type.

[0151] As an example, the knowledge retention mechanism is the original type constraint mechanism. Before processing the second text sample, the second language model is used to classify the first text sample to obtain the third similarity corresponding to each original type and each new type of the first text sample. The classification process is the same as step 10121 and will not be elaborated here.

[0152] In step 10332B, the first text sample is classified by the second language model to obtain the fourth similarity corresponding to each original type and each new type.

[0153] As an example, after processing the second text sample using the second language model, the second language model is used again to classify the first text sample to obtain the fourth similarity corresponding to each original type and each new type of the first text sample. The classification process is the same as step 10121 and will not be elaborated here.

[0154] In step 10333B, based on the third similarity and the fourth similarity, a knowledge retention loss function corresponding to the original type constraint mechanism is determined.

[0155] As an example, since the model parameters of the second language model before processing the second text sample and the second language model after processing the second text sample are different for the same processing object, that is, the first text sample, due to the addition of new types, there is a knowledge retention loss caused by learning new types between the third similarity and the fourth similarity. The following formula is used to calculate the knowledge retention loss function:

[0156]

[0157] where p i is the label probability distribution in the third similarity, q iIt is the label probability distribution in the fourth similarity obtained by processing the first text sample with the second language model after processing the second text sample. N is the number of labels. Specifically, the second language model before processing the second text sample (i.e., the second language model that has not learned new type of knowledge) is used to predict the third similarity of the first text sample, and the label probability distribution in the third similarity is used to replace the true label probability distribution. The true label probability distribution refers to the label probability distribution in the form of 0, 1, 0, 0, 0, while the label probability distribution in the third similarity refers to the label probability distribution in the form of 0.02, 0.95, 0.01, 0.01, 0.01.

[0158] Through the embodiments of the present application, using the knowledge distillation method, in the application scenarios where old data can be obtained, old knowledge is migrated, so that the second language model can retain the old knowledge of the original type learned during the pre-training process while learning new type of knowledge, thereby improving the learning efficiency of the language model.

[0159] It should be noted that in the application scenarios where old data can be obtained, since the model parameters and old data at different times can be obtained, the knowledge retention can be achieved by using the model parameter constraint mechanism and the original type constraint mechanism according to actual needs.

[0160] Continue to refer to Figure 3D , in step 1034, the second loss function and the knowledge retention loss function are fused to obtain the transfer learning loss function.

[0161] As an example, for the application scenarios where the old data cannot be accessed repeatedly, the knowledge retention loss function corresponding to the model parameter constraint mechanism is obtained by using the data-independent knowledge retention method, and the second loss function and the knowledge retention loss function are fused, that is, the transfer learning loss function corresponding to the model parameter constraint mechanism is calculated by the following formula:

[0162]

[0163] Among them, L cls is the second classification loss function, L ii is the second sample loss function, L is is the second prototype loss function, is the knowledge retention loss function corresponding to the model parameter constraint mechanism, and λ is a manually adjustable hyperparameter. For the application scenarios where the old data can be accessed, the knowledge retention loss function corresponding to the original type constraint mechanism is obtained by using the data-dependent knowledge retention method, and the second loss function and the knowledge retention loss function are fused, that is, the transfer learning loss function corresponding to the original type constraint mechanism is calculated by the following formula:

[0164] L = L cls + Lii +L is +L KD (8)

[0165] Among them, L cls is the second classification loss function, L ii is the second sample loss function, L is is the second prototype loss function, L KD is the knowledge retention loss function corresponding to the original type constraint mechanism.

[0166] For scenarios that allow access to old data, during training, the original type knowledge retention is achieved by using old data and knowledge distillation. For application scenarios that do not allow access to old data, by constraining the model parameters of the second language model and the second prototype vector of the original type, the second language model will not forget old knowledge too quickly during training.

[0167] In step 1035, based on the transfer learning loss function, the second language model, the second prototype vector, and the third prototype vector are updated to obtain the third language model, the fourth prototype vector, and the fifth prototype vector.

[0168] As an example, the third language model is a speech model for classification processing that has completed knowledge transfer processing. The fourth prototype vector is the prototype vector obtained by updating the second prototype vector of the original type. The fifth prototype vector is the prototype vector obtained by updating the third prototype vector of the new type.

[0169] As an example, for application scenarios where old data cannot be accessed repeatedly, the second language model is updated using the transfer learning loss function corresponding to the model parameter constraint mechanism to obtain the third language model that realizes knowledge retention. The second prototype vector is updated to obtain the fourth prototype vector, and the third prototype vector is updated to obtain the fifth prototype vector. For application scenarios that can access old data, the second language model is updated using the transfer learning loss function corresponding to the original type constraint mechanism to obtain the third language model that realizes knowledge retention. The second prototype vector is updated to obtain the fourth prototype vector, and the third prototype vector is updated to obtain the fifth prototype vector.

[0170] Through the embodiments of the present application, two knowledge retention methods are proposed according to different application scenarios. For scenarios that allow access to old data, during training, old data and knowledge distillation are used to achieve prototype knowledge retention in different learning stages. For application scenarios that do not allow access to old data, by constraining the model parameters of the second language model and the second prototype vector of the original type, the second language model will not quickly forget old knowledge during training. At the same time, calculate the second loss function corresponding to the first loss function in the pre-training stage, fuse the second loss function with the knowledge retention function, and update the second language model, the second prototype vector, and the third prototype vector, so that the updated third language model retains the old knowledge of the original type while learning new type knowledge, and the vector forms of the updated fourth prototype vector and the fifth prototype vector are adapted to the third language model, improving the learning efficiency and classification processing efficiency of the third language model.

[0171] It should be noted that for application scenarios that allow access to old data, the data-dependent knowledge retention method and the data-independent knowledge retention method can be used in combination.

[0172] The text processing method provided by the embodiments of the present application will be described in detail below.

[0173] Referring to Figure 4 , Figure 4 is a schematic flowchart of the text processing method provided by the embodiments of the present application. In some embodiments, after step 103, the steps 201 to 205 in Figure 4 can also be executed, which will be described in detail below.

[0174] In step 201, obtain the input text.

[0175] As an example, apply the third language model for classification processing to obtain the input text x′, and input the input text into the third language model for processing.

[0176] In step 202, perform text representation processing on the input text to obtain the input text representation, and add a classification mark to the input text representation to obtain the text representation to be processed.

[0177] As an example, perform text representation processing on the input text x′ to obtain the input text representation, and add a classification mark to the input text representation to obtain the text representation to be processed. Introduce the addition template function T(·), and add a classification mark to the text representation through x′ prompt = T(x′) to obtain the first text sample x′ prompt , and its specific form is x′.The intent isto[MASK′], where [MASK′] is the classification mark.

[0178] In step 203, the third language model is used to perform a mapping process on the text representation to be processed, and a fifth feature vector at the classification token is obtained.

[0179] As an example, the third language model obtains the fifth feature vector h corresponding to the classification label at the [MASK′] position. [MASK′] , and maps the fifth feature vector h [MASK′] to the vector space where the fourth prototype vector of the original type and the fifth prototype vector of the new type are located through the linear transformation matrix W, and v is obtained. [MASK′] , v [MASK′] is the fifth feature vector obtained by mapping h [MASK′] to the vector space where the fourth prototype vector of the original type and the fifth prototype vector of the new type are located.

[0180] In step 204, the fifth similarity between the fifth feature vector and each fourth prototype vector of the original type and each fifth prototype vector of the new type is determined.

[0181] As an example, the classification is achieved by calculating the fifth similarity between v [MASK] and the fourth prototype vector and the fifth prototype vector, and the fifth similarity corresponding to each fourth prototype vector and the fifth prototype vector is used as the predicted classification result.

[0182] In step 205, the original type or the new type corresponding to the maximum fifth similarity is used as the classification result of the input text.

[0183] As an example, the original type or the new type corresponding to the maximum fifth similarity in the predicted classification result is used as the classification result of the input text.

[0184] Through the embodiments of the present application, by using the third language model that has been pre-trained and knowledge-preserved, based on prompt learning, the input text is classified, and the classification task of the input text is transformed into a fill-in-the-blank masked language modeling problem. The third language model is effectively used to learn knowledge, and the classification of the input text can be achieved according to the original type in the old knowledge and the new type in the new knowledge, improving the accuracy of the classification result.

[0185] Next, an exemplary application of the embodiments of the present application in an actual text classification task application scenario will be described.

[0186] Since the overall framework is divided into multiple parts, each module's function will be introduced separately first, and then the overall training process and inference process will be introduced.

[0187] (1) Prototype Classification Method Based on Prompt Learning

[0188] Prompt Learning is an emerging technology that incorporates additional supplementary information (referred to as "Prompt") into the original text, transforming the downstream task into a fill-in-the-blank masked language modeling problem. This additional Prompt consists of characters, and after being transformed into a fill-in-the-blank masked language modeling problem, it can help us effectively utilize the knowledge learned by the pre-trained language model during the pre-training stage. Additionally, in Prompt Learning, there is an important component called "Verbalizer", through which the original labels can be projected onto a set of texts or continuous vectors.

[0189] First, represent the input original text as x = w 1 ,..., w n , and define the corresponding label as y. Then introduce a function T(·) for adding templates, which takes the input original text x. Through x prompt = T(x), the text template x prompt with additional information can be obtained. Its specific form is x. The intent is to [MASK]. [MASK] is a special label whose role is to predict the classification label of the original text x. Like other words, [MASK] has a corresponding hidden feature vector h [MASK] . During classification, the cosine similarity between h [MASK] and the first prototype vector of all original types will be calculated to classify [MASK]. Next, given a pre-trained language model M, the hidden feature vector h [MASK] corresponding to the classification label at the [MASK] position is obtained through the following formula:

[0190] h [MASK] = M(T(x)) (9)

[0191] where h [MASK] ∈ R h is the output of the last layer of the pre-trained language model M at the [MASK] position. Next, the hidden feature vector h [MASK] is mapped into the vector space where the first prototype vector of the original type is located through a linear transformation matrix W:

[0192] v [MASK] = W × h [MASK] (10)

[0193] where v [MASK] is the feature vector obtained by mapping h [MASK] into the vector space where the first prototype vector of the original type is located. Each first prototype vector of the original type corresponds to a standard classification label in the label definition set y. By calculating h [MASK]Implement classification using the cosine similarity between the sample and the first prototype vector of the original type.

[0194] Finally, calculate the similarity between the sample and the original type using cosine similarity to implement classification:

[0195]

[0196] where v i corresponds to v [MASK] and v j corresponds to the first prototype vector of the original type.

[0197] (2) Knowledge retention method

[0198] (2.1) Data-Agnostic Knowledge Prservation

[0199] It is crucial to retain the knowledge of previously seen prototypes. However, in many application scenarios, due to certain reasons (such as protecting user privacy), it is not possible to repeatedly access past data. In this case, we propose to use explicit weight constraints of the model, specifically, to constrain all parameters of the pre-trained language model and all first prototype vectors of the original types, that is, the original types in the old knowledge. The specific constraint loss function is as follows:

[0200]

[0201] where p joint is the model parameter of the language model at the current moment (i.e., the model parameter of the second language model when learning new type knowledge), and p seen is the model parameter of the language model at the previous moment (the model parameter of the second language model before learning new type knowledge). By means of explicit weight constraints, the language model can be forced not to quickly forget old knowledge during the process of learning new type knowledge. The constraint loss function in formula (12) is used after calculating the second classification loss function.

[0202] (2.2) Data-Dependent Knowledge Preservation: In some cases, the system or model is still allowed to access past data. In such a case, reserve a fixed-size virtual memory for storing the observed data (i.e., the first text samples), which are randomly selected from the original dataset at a fixed ratio, and use knowledge distillation to help the model remember old knowledge. To transfer old knowledge, use the following knowledge distillation loss function formula:

[0203]

[0204] Among them, p i is a soft label, and q i is the probability predicted by the model at the current moment. Specifically, use the language model of the previous moment (i.e., the second language model that has not learned the knowledge of the new type) to predict the label probability distribution of each piece of data in the virtual memory (i.e., the third similarity, called the soft label) (i.e., use the second language model that has not learned the knowledge of the new type to predict the label probability distribution of the first text sample stored in the virtual memory), and use these soft labels to replace the original hard labels (the hard labels refer to the label distribution in the form of 0, 1, 0, 0, 0, and the soft labels are in the form of 0.02, 0.95, 0.01, 0.01, 0.01). Among them, the hard label is the probability distribution of the true label, that is, the label corresponding to the sample is 1, and the rest are 0. For example, there are three prototype categories A, B, and C, and the label corresponding to x is A, then the probability distribution of the true label (hard label) is 1, 0, 0 (corresponding to A, B, C). The soft label is the probability distribution obtained by x through the model output, which may be 0.98, 0.01, 0.01 (corresponding to A, B, C). Using soft labels instead of hard labels can enable the model to learn more information during training. For example, if the probability distribution obtained by x through the model output is 0.5, 0.4, 0.1 (corresponding to A, B, C), then it means that from the perspective of the model, although x belongs to the A prototype category, it also has the characteristics of the B prototype category.

[0205] (3) Contrastive learning technique

[0206] In order to better classify the labels of the input text of the input language model, a supervised contrastive learning method is adopted, so that a sample has a high similarity with other samples of the same category and a high similarity with its prototype at the same time. Specifically, two samples belonging to the same category are regarded as a positive sample pair, and two samples with different intentions are regarded as a negative sample pair, and the following two loss functions are optimized:

[0207]

[0208]

[0209] Among them, L ii is the sample contrast loss function between a sample i and other samples j of the same category, and L is is the prototype contrast loss function between a sample i and the prototype c. T is the total number of samples, C is the total number of categories (corresponding to the total number of prototypes), and c i refers to the i-th prototype.

[0210] (4) Overall framework

[0211] Reference Figure 5 , Figure 5 is a schematic diagram of the model training scheme framework provided by the embodiments of this application. As Figure 5 shown, the entire framework is divided into two stages. In the first stage, a large number of pre-trained labeled examples (i.e., the first text samples) are used to learn the knowledge of the original types. Specifically, the pre-trained labeled examples are processed by adding templates to obtain text representations and templates including classification labels, and the text representations and templates are input into a pre-trained language model (i.e., the first language model) to output hidden feature vectors (i.e., the first feature vectors) corresponding to the classification labels. According to the original prototype vectors in the original prototype space (i.e., the first prototype vectors of each original type in the vector space), the hidden feature vectors are classified.

[0212] As Figure 5 shown, in the second stage, the prototype vectors of the newly added types are initialized in the original prototype space, that is, the corresponding newly added prototype vectors (i.e., the third prototype vectors) are randomly initialized for the newly added types, and the newly added prototype vectors are added to the original prototype space. For example, if there are 3 newly added types, then 3 corresponding newly added prototype vectors will be randomly initialized. These newly added prototype vectors have the same form as other original prototype vectors (i.e., the first prototype vectors or the second prototype vectors of the original types). The original prototype space with the newly added prototype vectors is the mixed prototype space. Then, a small number of fine-tuning training labeled examples (i.e., the second text samples) are used to fine-tune the model parameters of the language model. At the same time, the knowledge retention method is used to transfer the old knowledge to the new vector space. Specifically, the fine-tuning training labeled examples are processed in the same way as in the first stage to obtain hidden feature vectors (i.e., the third feature vectors). When classifying the hidden feature vectors, the knowledge retention method needs to be used to transfer the old knowledge to the mixed prototype space. Specifically, for application scenarios where old knowledge cannot be obtained, a data-independent knowledge retention method is used to train the mixed prototype vectors in the mixed prototype space to obtain the final prototype space. For application scenarios where old knowledge can be obtained, the old knowledge is stored using virtual memory, and a data-dependent knowledge retention method is used to train the mixed prototype vectors in the mixed prototype space according to the old knowledge to obtain the final prototype space. The final prototype space includes the final prototype vectors of the original types (i.e., the fourth prototype vectors) and the final prototype vectors of the newly added types (i.e., the fifth prototype vectors).

[0213] In addition to training the prototype vectors, the language model also needs to be trained.

[0214] In the first stage, the language model (i.e., the first language model) takes the pre-trained labeled examples as input, outputs hidden feature vectors, and determines the original prototype vectors corresponding to the hidden feature vectors. At this time, the total loss of the first stage can be calculated. Here, let θ seenExpressed as the parameters of the pre-trained language model and the known prototypes (i.e., the original prototype vectors). The total loss in the first stage is defined as follows:

[0215] L = L cls + L ii + L is (16)

[0216] where L cls is the first classification loss function, L ii is the first sample contrast loss function between a sample i and other samples j of the same class in the first stage, and L is is the first prototype contrast loss function between a sample i and the prototype c in the first stage.

[0217] In the second stage, the pre-trained language model (i.e., the second language model) is trained using fine-tuning training labeled examples, where the fine-tuning training labeled examples include labeled examples corresponding to the original prototype vectors and the newly added prototype vectors. The fine-tuning training labeled examples are the labeled examples not used in the first stage, that is, there is no intersection between the pre-training labeled examples and the fine-tuning training labeled examples.

[0218] Let θ joint represent the parameters excluding the prototype vectors of the newly added types (i.e., all parameters except the prototype vectors of the newly added types, including the model parameters of the pre-trained language model and the prototype vectors of the original types).

[0219] In the process of old knowledge transfer (i.e., knowledge retention) in the second stage, different knowledge retention methods need to be used for different application scenarios.

[0220] For application scenarios where old knowledge cannot be accessed, such as application scenarios where access to old data is not allowed due to privacy policies or other issues, a data-independent knowledge retention method is adopted. At this time, the optimized loss function is:[[]]

[0221]

[0222] where λ is a manually adjustable hyperparameter.

[0223] For application scenarios where access to old data is allowed, a virtual memory can be designed to store a small amount of old data, and knowledge distillation, that is, a data-dependent knowledge retention method, can be used to consolidate the old knowledge. Therefore, the optimized loss function is as follows:

[0224] L = L cls + L ii + L is + L KD (18)

[0225] As described above, the training process of the language model is divided into two stages. In each stage, the prototype vectors of all seen class prototypes (i.e., original types) are randomly initialized.

[0226] In the first stage, the original types are preset as known class prototypes in the vector space of class prototypes. A large number of labeled examples are used as pre-training labeled examples (i.e., the first text samples), which are input into the language model. Through the classification process of the language model, the first predicted classification result is obtained. According to the first classification result, the first classification loss function L cls , the first sample contrast loss function L ii (i.e., formula (14)) and the first prototype contrast loss function L is (i.e., formula (15)) are calculated. According to formula (16), the total loss function of the first stage (i.e., the first loss function) is calculated. Based on the total loss function of the first stage, the language model (i.e., the first language) and the prototype vectors of the original types are updated.

[0227] In the second stage, different knowledge retention methods are adopted according to different application scenarios (formula (13) corresponds to the scenario where access to old data is allowed, while formula (12) corresponds to the scenario where access to old data is restricted due to privacy policies).

[0228] That is, in the second stage, the prototype vectors of new types are added to the vector space of the original types. A small number of labeled examples are used as the labeled examples for fine-tuning. Among them, the labeled examples for fine-tuning are different from the labeled examples for pre-training. Using the labeled examples for fine-tuning (i.e., the second text samples) as the input of the language model (i.e., the second language model) pre-trained in the first stage, a classification process is carried out to obtain the second predicted classification result. According to the second predicted classification result, the second classification loss function Lcls, the second sample contrast loss function L ii (i.e., formula (14)) and the second prototype contrast loss function L is (i.e., formula (15)) are calculated. According to different application scenarios, formula (12) or formula (13) is used to calculate the knowledge retention loss function. Among them, formula (13) corresponds to the application scenario where access to old data is allowed, while formula (12) corresponds to the application scenario where access to old data is restricted due to privacy policies. Finally, according to formula (17) or formula (18), the second classification loss function L cls , the second sample contrast loss function L ii and the second prototype contrast loss function L isIt is fused with the knowledge retention loss function to obtain the total loss function in the second stage (i.e., the transfer learning loss function). The total loss function in the second stage is used to update all the parameters of the language model (i.e., the second language model) pre-trained in the first stage, as well as the prototype vectors of the original types and the prototype vectors of the new types, to obtain the language model for text classification (i.e., the third language model) and the final prototype space.

[0229] The inference process of the language model for text classification is also very simple. For a given input text x, it passes through the functions mentioned above in sequence, and finally obtains v. [MASK] . When calculating the similarity between v [MASK] and each class prototype using cosine similarity, the class with the highest similarity is used as the class predicted by the model. For example, for a scenario with C prototypes, the algorithm flow for inferring unlabeled input text is as follows:

[0230] 1: Input: An unlabeled utterance x / / Input the unlabeled statement x

[0231] 2: Output: The corresponding intent y of unlabeled utterance x / / Output: The intent classification y corresponding to the unlabeled statement x

[0232] 3: X prompt ← T(x) / / Obtain the text template x through x prompt = T(x) to get the text template x prompt

[0233] 4: h[MASK] ← M(T(x)) with Eq.9 / / Use formula (9) h [MASK] = M(T(x)) to process x prompt to obtain h [MASK]

[0234] 5: υ [MASK] ← W × h [MASK] with Eq.10 / / Use formula (10) v [MASK] = W × h [MASK] to process h [MASK] to obtain v [MASK]

[0235] 6: for i = 1 to C do / / Perform operations in sequence from the first prototype to the Cth prototype

[0236] 7: Eq.11 / / Use formula (11) to calculate v[MASK] with each prototype vector v i cosine similarity

[0237] 8: end for / / End

[0238] 9: Get intent y whose corresponding prototype and utterance x have the highest similarity / / Obtain the intent y whose corresponding prototype and statement x have the highest similarity

[0239] 10: return y / / Output the classification label y.

[0240] Refer to Figure 6 , Figure 6 is the experimental result data graph provided by the embodiments of the present application. Figure 6 The full-sample test and batch-sample test in [] are two different experimental settings. The full-sample test refers to directly testing all unlabeled samples, and the batch-sample test refers to conducting small-sample learning tests in a certain number of batches. For example, it is set to conduct 1000 small-sample learning scenario experiments. For the setting of five times of learning, for the SNIPS dataset, 20 unlabeled samples are selected for each test. Among them, the 20 unlabeled samples are obtained by extracting 5 samples from each of the 4 prototype categories. For the NLUE dataset, 50 unlabeled samples are selected for each test. Among them, the 50 unlabeled samples are obtained by extracting 5 samples from each of the 10 prototype categories. Figure 6 The data marked in bold in [] are the data with the highest values under the same experimental conditions. According to Figure 6 the experimental results in [], it can be found that in the full-sample test and batch-sample test experiments of one-time learning and five-time learning respectively with the SNIPS dataset and NLU dataset as samples, compared with the meta-learning adversarial domain adaptation network (MLADA, Meta-Learning Adversarial Domain Adaptation Network) and diversity feature enhanced prototype network (DFEPN, Diversity Feature Enhanced Prototypical Network) in the related technology, DDKP and DAKP in the model training method provided by the embodiments of the present application can achieve optimal or sub-optimal performance in all cases, significantly better than the previous methods (especially in the scenario of one-time learning). Among them, DAKP does not rely on any old data during training and can adapt to various scenarios.

[0241] Refer to Figure 7 ,Figure 7 This is the line graph of the test performance data provided by the embodiments of the present application. To further explore the effectiveness of the model training method provided by the embodiments of the present application, two types of tests were conducted respectively: (1) After learning new category knowledge (i.e., new added types), test the performance on the test data of the old categories using DDKP and DAKP; (2) After learning new category knowledge, test the performance on the test data of the new categories using DDKP and DAKP. Through the above method, the forgetting speed of the language model for old knowledge can be better observed. As Figure 7 shown, conduct a one-time learning test using the SNIPS-NLU public dataset. Figure 7 In, the line corresponding to DDKP-old category represents the performance of the language model using the DDKP model training method provided by the embodiments of the present application on the test data of the old categories, the line corresponding to DAKP-old category represents the performance of the language model using the DAKP model training method provided by the embodiments of the present application on the test data of the old categories, the line corresponding to the control group-old category represents the performance of the language model not trained using the DAKP and DDKP methods on the test data of the old categories, the line corresponding to DDKP-new category represents the performance of the language model using the DDKP model training method provided by the embodiments of the present application on the test data of the new categories, the line corresponding to DAKP-new category represents the performance of the language model using the DAKP model training method provided by the embodiments of the present application on the test data of the new categories, and the line corresponding to the control group-new category represents the performance of the language model not trained using the DAKP and DDKP methods on the test data of the new categories. As Figure 7 shown, the experimental results show that the speed at which the language models trained using DDKP and DAKP forget old knowledge is significantly slower than that of the language models not trained using any knowledge retention techniques.

[0242] It can be understood that in the embodiments of the present application, data related to user information, etc. is involved. When the embodiments of the present application are applied to specific products or technologies, user permission or consent needs to be obtained, and the collection, use, and processing of relevant data need to comply with the relevant laws, regulations, and standards of the relevant countries and regions.

[0243] Next, continue to illustrate the exemplary structure of the software module implementation of the model training device 253 provided by the embodiments of the present application. In some embodiments, as Figure 2 shown, the software module stored in the model training device 253 in the memory 250 may include:

[0244] A pre-training module 2531, configured to obtain a first text sample, and perform pre-training processing on a first language model and a first prototype vector corresponding to the original type based on the first text sample, so as to obtain a second language model and a second prototype vector corresponding to the original type;

[0245] An acquisition module 2532, configured to acquire a third prototype vector and a second text sample corresponding to a new added type.

[0246] A transfer training module 2533, configured to perform transfer training processing based on a knowledge retention mechanism on a second language model, a second prototype vector, and a third prototype vector based on the second text sample, to obtain a third language model, a fourth prototype vector corresponding to an original type, and a fifth prototype vector corresponding to the new added type, where the knowledge retention mechanism includes at least one of a model parameter constraint mechanism and an original type constraint mechanism.

[0247] In some embodiments, the pre-training module 2531 is further configured to acquire an original text sample; perform text representation processing on the original text sample to obtain a text representation; and add a classification label to the text representation to obtain a first text sample.

[0248] In some embodiments, the pre-training module 2531 is further configured to perform classification processing on the first text sample through a first language model to obtain a first predicted classification result, calculate a first loss function based on a first labeled classification label and the first predicted classification result of the first text sample, and update the first language model and the first prototype vector based on the first loss function to obtain a second language model and a second prototype vector.

[0249] In some embodiments, the pre-training module 2531 is further configured to perform mapping processing on the first text sample through the first language model to obtain a first feature vector at the classification label, determine a first similarity between the first feature vector and a first prototype vector of each original type, and use the first similarity corresponding to each original type as the first predicted classification result.

[0250] In some embodiments, the pre-training module 2531 is further configured to acquire a third text sample belonging to the same original type as the first text sample, and obtain a second feature vector of the third text sample through the first language model; calculate a first sample contrast loss function based on an error between the first feature vector and the second feature vector; calculate a first prototype contrast loss function based on an error between the first feature vector and a first prototype vector of the original type to which the first text sample belongs; calculate a first classification loss function based on an error between the first labeled classification label and the first predicted classification result of the first text sample; and perform fusion processing on the first sample contrast loss function, the first prototype contrast loss function, and the first classification loss function to obtain a first loss function.

[0251] In some embodiments, the transfer training module 2533 is further configured to classify the second text sample through the second language model to obtain a second predicted classification result; calculate a second loss function based on the second labeled classification label and the second predicted classification result of the second text sample; obtain a knowledge retention loss function of the corresponding knowledge retention mechanism based on the second language model before processing the second text sample; perform a fusion process on the second loss function and the knowledge retention loss function to obtain a transfer learning loss function; and update the second language model, the second prototype vector, and the third prototype vector based on the transfer learning loss function to obtain a third language model, a fourth prototype vector, and a fifth prototype vector.

[0252] In some embodiments, the transfer training module 2533 is further configured to perform a mapping process on the second text sample through the second language model to obtain a third feature vector at the classification label; determine a second similarity between the third feature vector and the first prototype vector of each original type and the third prototype vector of each new type, and use the second similarities corresponding to each original type and each new type as the second predicted classification result.

[0253] In some embodiments, the transfer training module 2533 is further configured to obtain a fourth text sample that belongs to the same original type as the second text sample, and obtain a fourth feature vector of the fourth text sample through the second language model; calculate a second sample contrast loss function based on the error between the third feature vector and the fourth feature vector; calculate a second prototype contrast loss function based on the error between the third feature vector and the third prototype vector of the original type to which the second text sample belongs; calculate a second classification loss function based on the error between the second labeled classification label and the second predicted classification result of the second text sample; and perform a fusion process on the second sample contrast loss function, the second prototype contrast loss function, and the second classification loss function to obtain a second loss function.

[0254] In some embodiments, the transfer training module 2533 is further configured to obtain the first model parameters of the second language model before processing the second text sample and the second model parameters of the second language model at the current moment, and determine a knowledge retention loss function of the corresponding model parameter constraint mechanism based on the error between the first model parameters and the second model parameters.

[0255] In some embodiments, the transfer training module 2533 is further configured to classify the first text sample by using the second language model before processing the second text sample, to obtain a third similarity corresponding to each original type and each new type; classify the first text sample by using the second language model, to obtain a fourth similarity corresponding to each original type and each new type; and determine a knowledge retention loss function for the original type constraint mechanism based on the third similarity and the fourth similarity.

[0256] In some embodiments, the transfer training module 2533 is further configured to obtain an input text, perform text representation processing on the input text to obtain an input text representation, add a classification mark to the input text representation to obtain a text representation to be processed, perform mapping processing on the text representation to be processed by using a third language model to obtain a fifth feature vector at the classification mark, determine a fifth similarity between the fifth feature vector and a fourth prototype vector of each original type and a fifth prototype vector of each new type, and use the original type or the new type corresponding to the maximum fifth similarity as the classification result of the input text.

[0257] Next, the exemplary structure of the text processing device provided in the embodiments of the present application implemented as a software module will be further described. In some embodiments, the text processing device may include:

[0258] A text acquisition module, configured to acquire an input text;

[0259] A text representation processing module, configured to perform text representation processing on the input text to obtain an input text representation, and add a classification mark to the input text representation to obtain a text representation to be processed;

[0260] A mapping processing module, configured to perform mapping processing on the text representation to be processed by using a third language model to obtain a fifth feature vector at the classification mark, where the third language model is trained by using the model training method provided in the embodiments of the present application;

[0261] A similarity determination module, configured to determine a fifth similarity between the fifth feature vector and a fourth prototype vector of each original type and a fifth prototype vector of each new type, where the fourth prototype vector and the fifth prototype vector are trained by using the model training method provided in the embodiments of the present application;

[0262] A result output module, configured to use the original type or the new type corresponding to the maximum fifth similarity as the classification result of the input text.

[0263] An embodiment of the present application provides a computer program product, which includes computer-executable instructions stored in a computer-readable storage medium. A processor of an electronic device reads the computer-executable instructions from the computer-readable storage medium, and the processor executes the computer-executable instructions, so that the electronic device executes the model training method and the text processing method described above in the embodiments of the present application.

[0264] An embodiment of the present application provides a computer-readable storage medium storing computer-executable instructions, wherein the computer-executable instructions are stored. When the computer-executable instructions are executed by a processor, the processor will be caused to execute the model training method and the text processing method provided by the embodiments of the present application. For example, as Figure 3A the model training method shown Figure 4 and the text processing method shown.

[0265] In some embodiments, the computer-readable storage medium may be a memory such as FRAM, ROM, PROM, EPROM, EEPROM, flash memory, magnetic surface memory, optical disc, or CD-ROM; or may be various devices including one or any combination of the above memories.

[0266] In some embodiments, the computer-executable instructions may be in the form of a program, software, software module, script, or code, written in any form of programming language (including compiled or interpreted language, or declarative or procedural language), and may be deployed in any form, including being deployed as an independent program or being deployed as a module, component, subroutine, or other unit suitable for use in a computing environment.

[0267] As an example, the computer-executable instructions may or may not correspond to a file in the file system, and may be stored as a part of a file storing other programs or data. For example, they may be stored in one or more scripts in a HyperText Markup Language (HTML) document, stored in a single file dedicated to the program being discussed, or stored in multiple cooperating files (for example, files storing one or more modules, subroutines, or code portions).

[0268] As an example, the computer-executable instructions may be deployed to execute on one electronic device, or on multiple electronic devices located at one location, or on multiple electronic devices distributed at multiple locations and interconnected by a communication network.

[0269] In summary, by implementing classification using a prototype based on prompt learning and combining supervised contrastive learning in the embodiments of the present application, the knowledge learned during the pre-training of the language model is utilized, and at the same time, the differences between categories are more considered. In addition, two knowledge retention methods are proposed according to different real-life scenarios: (1) for scenarios where access to old data is allowed, a virtual memory is maintained to store old data, and old data and knowledge distillation are used during training to achieve prototype knowledge retention in different learning stages; (2) for scenarios where access to old data is not allowed, by constraining the parameters of the pre-trained language model and the parameters of the category prototypes, the language model can avoid quickly forgetting old knowledge during training. The embodiments of the present application are intended to be applied to the scenario of text classification, and more specifically, the scenario of intent recognition. In theory, it can be directly transferred to other text classification scenarios.

[0270] As described above, the above are only embodiments of the present application and are not intended to limit the protection scope of the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and scope of the present application are included in the protection scope of the present application.

Claims

1. A model training method, characterized in that, the method includes: obtaining a first text sample, and performing pre-training processing on a first language model and a first prototype vector of a corresponding original type based on the first text sample to obtain a second language model and a second prototype vector of the corresponding original type; obtaining a third prototype vector of a corresponding new type and a second text sample; performing transfer training processing based on a knowledge retention mechanism on the second language model, the second prototype vector, and the third prototype vector based on the second text sample to obtain a third language model, a fourth prototype vector of the corresponding original type, and a fifth prototype vector of the corresponding new type; wherein, the knowledge retention mechanism includes at least one of a model parameter constraint mechanism and an original type constraint mechanism.

2. The method according to claim 1, characterized in that, the obtaining of the first text sample includes: obtaining an original text sample; performing text representation processing on the original text sample to obtain a text representation; adding a classification label to the text representation to obtain the first text sample.

3. The method according to claim 1, characterized in that, the performing pre-training processing on the first language model and the first prototype vector of the corresponding original type based on the first text sample to obtain the second language model and the second prototype vector of the corresponding original type includes: performing classification processing on the first text sample through the first language model to obtain a first predicted classification result; calculating a first loss function based on a first labeled classification label of the first text sample and the first predicted classification result; performing update processing on the first language model and the first prototype vector based on the first loss function to obtain the second language model and the second prototype vector.

4. The method according to claim 3, characterized in that, the first text sample includes a classification label; the performing classification processing on the first text sample through the first language model to obtain a first predicted classification result includes: performing mapping processing on the first text sample through the first language model to obtain a first feature vector at the classification label; determining a first similarity between the first feature vector and the first prototype vector of each of the original types, and taking the first similarity corresponding to each of the original types as the first predicted classification result.

5. The method according to claim 4, characterized in that, the calculating a first loss function based on the first labeled classification label of the first text sample and the first predicted classification result includes: obtaining a third text sample belonging to the same original type as the first text sample, and obtaining a second feature vector of the third text sample through the first language model; calculating a first sample contrast loss function based on an error between the first feature vector and the second feature vector; calculating a first prototype contrast loss function based on an error between the first feature vector and the first prototype vector of the original type to which the first text sample belongs. Calculate a first classification loss function based on the error between the first labeled classification label of the first text sample and the first predicted classification result; Perform a fusion process on the first sample contrast loss function, the first prototype contrast loss function, and the first classification loss function to obtain the first loss function.

6. The method according to claim 1, wherein, the performing, based on the second text sample, a transfer training process on the second language model, the second prototype vector, and the third prototype vector based on a knowledge retention mechanism to obtain a third language model, a fourth prototype vector corresponding to the original type, and a fifth prototype vector corresponding to the new type includes: Performing a classification process on the second text sample through the second language model to obtain a second predicted classification result; Calculating a second loss function based on the second labeled classification label of the second text sample and the second predicted classification result; Obtaining a knowledge retention loss function corresponding to the knowledge retention mechanism based on the second language model before processing the second text sample; Performing a fusion process on the second loss function and the knowledge retention loss function to obtain a transfer learning loss function; Updating the second language model, the second prototype vector, and the third prototype vector based on the transfer learning loss function to obtain the third language model, the fourth prototype vector, and the fifth prototype vector.

7. The method according to claim 6, wherein, the performing a classification process on the second text sample through the second language model to obtain a second predicted classification result includes: Performing a mapping process on the second text sample through the second language model to obtain a third feature vector at the classification label; Determining a second similarity between the third feature vector and the first prototype vector of each original type and the third prototype vector of each new type, and taking the second similarity corresponding to each original type and each new type as the second predicted classification result.

8. The method according to claim 7, wherein, the calculating a second loss function based on the second labeled classification label of the second text sample and the second predicted classification result includes: Obtaining a fourth text sample belonging to the same original type as the second text sample, and obtaining a fourth feature vector of the fourth text sample through the second language model; Calculating a second sample contrast loss function based on the error between the third feature vector and the fourth feature vector; Calculating a second prototype contrast loss function based on the error between the third feature vector and the third prototype vector of the original type to which the second text sample belongs; Calculating a second classification loss function based on the error between the second labeled classification label of the second text sample and the second predicted classification result; Performing a fusion process on the second sample contrast loss function, the second prototype contrast loss function, and the second classification loss function to obtain the second loss function.

9. The method according to claim 6, It is characterized in that the knowledge retention mechanism is the model parameter constraint mechanism; obtaining a knowledge retention loss function corresponding to the knowledge retention mechanism based on the second language model before processing the second text sample includes: obtaining first model parameters of the second language model before processing the second text sample and second model parameters of the second language model at the current moment; determining a knowledge retention loss function corresponding to the model parameter constraint mechanism based on the error between the first model parameters and the second model parameters.

10. The method according to claim 6, It is characterized in that the knowledge retention mechanism is the original type constraint mechanism; obtaining a knowledge retention loss function corresponding to the knowledge retention mechanism based on the second language model before processing the second text sample includes: performing classification processing on the first text sample by the second language model before processing the second text sample to obtain a third similarity corresponding to each of the original types and each of the new types; performing classification processing on the first text sample by the second language model to obtain a fourth similarity corresponding to each of the original types and each of the new types; determining a knowledge retention loss function corresponding to the original type constraint mechanism based on the third similarity and the fourth similarity.

11. A text processing method, It is characterized in that the method includes: obtaining an input text; performing text representation processing on the input text to obtain an input text representation, and adding classification tags to the input text representation to obtain a text representation to be processed; performing mapping processing on the text representation to be processed by a third language model to obtain a fifth feature vector at the classification tag, where the third language model is trained by the method according to any one of claims 1 to 10; determining a fifth similarity between the fifth feature vector and a fourth prototype vector of each original type and a fifth prototype vector of each new type, where the fourth prototype vector and the fifth prototype vector are trained by the method according to any one of claims 1 to 10; taking the original type or new type corresponding to the maximum fifth similarity as the classification result of the input text.

12. A model training device, It is characterized in that the device includes: a pre-training module, configured to obtain a first text sample, and perform pre-training processing on a first language model and a first prototype vector corresponding to an original type based on the first text sample to obtain a second language model and a second prototype vector corresponding to the original type; an obtaining module, configured to obtain a third prototype vector corresponding to a new type and a second text sample; A transfer training module, configured to perform transfer training processing based on a knowledge retention mechanism on the second language model, the second prototype vector, and the third prototype vector based on the second text sample, to obtain a third language model, a fourth prototype vector corresponding to the original type, and a fifth prototype vector corresponding to the new type, where the knowledge retention mechanism includes at least one of a model parameter constraint mechanism and an original type constraint mechanism.

13. A text processing device, characterized in that the device includes: a text acquisition module, configured to acquire an input text; a text representation processing module, configured to perform text representation processing on the input text to obtain an input text representation, and add a classification tag to the input text representation to obtain a text representation to be processed; a mapping processing module, configured to perform mapping processing on the text representation to be processed through a third language model to obtain a fifth feature vector at the classification tag, where the third language model is trained by the method according to any one of claims 1 to 10; a similarity determination module, configured to determine a fifth similarity between the fifth feature vector and the fourth prototype vector of each original type and the fifth prototype vector of each new type, where the fourth prototype vector and the fifth prototype vector are trained by the method according to any one of claims 1 to 10; a result output module, configured to use the original type or the new type corresponding to the maximum fifth similarity as the classification result of the input text.

14. An electronic device, characterized in that the electronic device includes: a memory, configured to store computer-executable instructions; a processor, configured to implement the method according to any one of claims 1 to 10 or claim 11 when executing the computer-executable instructions stored in the memory.

15. A computer-readable storage medium storing computer-executable instructions, characterized in that the computer-executable instructions, when executed by a processor, implement the method according to any one of claims 1 to 10 or claim 11.

16. A computer program product including computer-executable instructions, characterized in that the computer-executable instructions, when executed by a processor, implement the method according to any one of claims 1 to 10 or claim 11.