A method, device and storage medium for training text matching model
Through multiple rounds of knowledge distillation and data enhancement methods, the problem of insufficient data volume and diversity in text matching model training is solved, the generalization and matching degree of the model are improved, and the efficient knowledge distillation effect is achieved.
Patent Information
- Application Number
- CN202210231382.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-03-09
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2042-03-09
AI Technical Summary
During the training process of existing text matching models, the training effect is poor due to insufficient data volume and diversity of the training data, and the existing data enhancement methods are poor or difficult to control.
By obtaining labeled and unlabeled data sets in different fields, the pre-trained language model is subjected to multiple rounds of knowledge distillation training, combined with search engines and approximate neighbor search libraries, the data is targeted and enhanced, and finally high-quality distillation data is generated for training student models.
It improves the generalization of the student model and the degree of matching with the teacher model, enhances the effect of knowledge distillation, and solves the problem of insufficient data volume and diversity.
Smart Images

Figure CN114781477B_ABST
Abstract
Description
Background Art
[0002] In the training process of the text matching model, a large amount of targeted training data is required to train the model in order to improve the training effect of the model. However, in most cases, there is a problem of insufficient data volume and diversity of the training data. Since the deep learning used for text matching model training generally has a high data dependence, the training effect is often severely limited. In order to increase the quality and quantity of training data as much as possible, data enhancement has begun to attract attention. There are currently two main methods for text data enhancement. One is the random addition, deletion, and back translation method, and the other is the adversarial enhancement method based on autoencoders.
[0003] Methods such as random addition, deletion, and back translation are very mechanical and cannot be used for targeted enhancement based on the training of the neural network, and have poor flexibility. Although adversarial enhancement methods based on autoencoders can dynamically adjust the enhanced data according to the neural network, most of them are complex and difficult to control, and cannot be achieved in one step. Summary of the invention
[0004] In order to overcome the above technical problems, the present invention proposes a method for training a text matching model, and the technical solution of the method is as follows:
[0005] S1, obtaining a first training set, a second training set and a third training set, wherein the first training set comprises labeled data for text matching in a first domain, the second training set comprises unlabeled data for text matching in the first domain, and the third training set comprises labeled data for text matching in the second domain;
[0006] S2, inputting the third training set into a pre-trained language model, training the pre-trained language model based on a text matching task to obtain a candidate language model, using the candidate language model as a first teacher model in knowledge distillation, and using the first teacher model to perform knowledge distillation on a pre-established model to be distilled to obtain a first student model;
[0007] S3, using the first training set to train the first teacher model to obtain a trained second teacher model, storing the unlabeled data in the second training set in a search engine, using the second teacher model to extract feature vectors corresponding to the unlabeled data in the second training set, storing the feature vectors in an approximate nearest neighbor search library, the search engine segmenting the unlabeled data in the second training set to establish an inverted index, and the approximate nearest neighbor search library using a predetermined approximate nearest neighbor search algorithm to index the feature vectors;
[0008] S4, the annotated data in the first training set includes a first text and a second text, a first result set is obtained from the search engine and the approximate neighbor search library according to the first text, a second result set is obtained from the search engine and the approximate neighbor search library according to the second text, and distilled data for knowledge distillation is determined by concatenating the first text with the second result set and concatenating the second text with the first result set;
[0009] S5. Based on the distilled data and the second teacher model, the first student model is trained using a knowledge distillation method to generate a trained target model.
[0010] Furthermore, in step S3, the first teacher model is trained using the first training set to obtain a trained second teacher model, including:
[0011] The labeled data contained in the first training set corresponds to multiple categories. The first training set is divided into a sub-training set, a sub-validation set and a sub-test set. The second teacher model is determined based on the sub-training set, the sub-validation set and the sub-test set, so that the prediction accuracy of the second teacher model in each category in the sub-test set is greater than a preset threshold, and the value range of the preset threshold is (0,1).
[0012] Furthermore, the preset threshold is 0.95.
[0013] Furthermore, the search engine is an Elasticsearch engine, and the approximate neighbor search library is a Faiss framework.
[0014] Furthermore, in step S4, obtaining a first result set from the search engine and the approximate neighbor search library according to the first text includes:
[0015] Recalling n texts from the search engine according to the first text, where n is a natural number;
[0016] Recalling m texts from the approximate nearest neighbor search library according to the first text, where m is a natural number;
[0017] The second teacher model is used to determine the matching degree between the n texts and the m texts and the first text, and the texts whose matching degree is not less than the preset threshold are taken as the first result set.
[0018] Furthermore, in step S4, obtaining a second result set from the search engine and the approximate neighbor search library according to the second text includes:
[0019] Recalling p pieces of text from the search engine according to the second text, where p is a natural number;
[0020] Recalling q texts from the approximate neighbor search library according to the second text, where q is a natural number;
[0021] The second teacher model is used to determine the matching degree between the p pieces of text and the q pieces of text and the second text, and the texts whose matching degree is not less than the preset threshold are taken as the second result set.
[0022] Furthermore, in step S4, determining the distilled data for knowledge distillation by concatenating the first text with the second result set and concatenating the second text with the first result set includes:
[0023] Concatenate the first text with each text in the second result set to obtain a first concatenated result set;
[0024] Concatenate each text in the first result set with the second text to obtain a second concatenated result set;
[0025] The first training set, the first splicing result set and the second splicing result set are combined to obtain the distilled data.
[0026] The present invention also proposes a device for training a text matching model, the device for training a text matching model comprising a memory and a processor, the memory storing at least one program, the at least one program being executed by the processor to implement the method for training a text matching model as described above.
[0027] The present invention also proposes a computer-readable storage medium, in which at least one program is stored, and when the at least one program is run, the method for training a text matching model as described above is executed.
[0028] The beneficial effects brought by the technical solution provided by the present invention are:
[0029] A method and device for training a text matching model in an embodiment of the present invention can generate a teacher model that meets the text matching task objectives in a field by training a pre-trained language model using a training set containing labeled data in different fields. In conjunction with a search engine and an approximate neighbor search library, targeted enhancements can be made to the training set containing labeled data and the training set containing unlabeled data, and high-quality distilled data for knowledge distillation can be generated. A student model can be trained based on the teacher model and the distilled data, thereby improving the generalization of the student model and the degree of matching between the student model and the teacher model, thereby improving the effect of knowledge distillation. BRIEF DESCRIPTION OF THE DRAWINGS
[0030] Figure 1A flowchart of a method for training a text matching model according to an embodiment of the present invention;
[0031] Figure 2 The present invention is a schematic diagram of the structure of a device for training a text matching model according to an embodiment of the present invention. DETAILED DESCRIPTION
[0032] In order to make the purpose, technical solution and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0033] Embodiment 1:
[0034] like Figure 1 FIG. 1 is a flow chart of a method for training a text matching model according to an embodiment of the present invention, showing specific implementation steps of the method, including:
[0035] S1, obtaining a first training set, a second training set and a third training set, wherein the first training set comprises labeled data for text matching in a first domain, the second training set comprises unlabeled data for text matching in the first domain, and the third training set comprises labeled data for text matching in the second domain;
[0036] In practical applications, the first field is a pre-selected specific field, such as news, medical or financial fields. The first training set belonging to the first field contains a small amount of labeled data, and the second training set belonging to the first field contains a large amount of unlabeled data. The second field is a public domain, and public data sets for text matching can be obtained from the public domain, including but not limited to the following data sets: Chinese-MNLI data set, Chinese-SNLI data set, OCNLI data set, Chinese-STS-B data set and PKU data set.
[0037] S2, inputting the third training set into a pre-trained language model, i.e., a pre-trained language model, training the pre-trained language model based on a text matching task to obtain a candidate language model, using the candidate language model as a first teacher model in knowledge distillation, and using the first teacher model to perform knowledge distillation on a pre-established model to be distilled to obtain a first student model;
[0038] S3, using the first training set to train the first teacher model to obtain a trained second teacher model, storing the unlabeled data in the second training set in a search engine, using the second teacher model to extract feature vectors corresponding to the unlabeled data in the second training set, storing the feature vectors in an approximate nearest neighbor search library, the search engine segmenting the unlabeled data in the second training set to establish an inverted index, and the approximate nearest neighbor search library using a predetermined approximate nearest neighbor search algorithm to index the feature vectors;
[0039] S4, the annotated data in the first training set includes a first text and a second text, a first result set is obtained from the search engine and the approximate neighbor search library according to the first text, a second result set is obtained from the search engine and the approximate neighbor search library according to the second text, and distilled data for knowledge distillation is determined by concatenating the first text with the second result set and concatenating the second text with the first result set;
[0040] S5. Based on the distilled data and the second teacher model, the first student model is trained using a knowledge distillation method to generate a trained target model.
[0041] Specifically, in step S3, the first teacher model is trained using the first training set to obtain a trained second teacher model, including:
[0042] The labeled data contained in the first training set corresponds to multiple categories. The first training set is divided into a sub-training set, a sub-validation set and a sub-test set. The second teacher model is determined based on the sub-training set, the sub-validation set and the sub-test set, so that the prediction accuracy of the second teacher model in each category in the sub-test set is greater than a preset threshold, and the value range of the preset threshold is (0,1).
[0043] Specifically, the preset threshold is 0.95.
[0044] Specifically, the search engine is the Elasticsearch engine, and the approximate neighbor search library is the Faiss framework.
[0045] Specifically, obtaining a first result set from the search engine and the approximate neighbor search library according to the first text in step S4 includes:
[0046] Recalling n texts from the search engine according to the first text, where n is a natural number;
[0047] Recalling m texts from the approximate nearest neighbor search library according to the first text, where m is a natural number;
[0048] The second teacher model is used to determine the matching degree between the n texts and the m texts and the first text, and the texts whose matching degree is not less than the preset threshold are taken as the first result set.
[0049] Specifically, obtaining a second result set from the search engine and the approximate neighbor search library according to the second text in step S4 includes:
[0050] Recalling p pieces of text from the search engine according to the second text, where p is a natural number;
[0051] Recalling q texts from the approximate neighbor search library according to the second text, where q is a natural number;
[0052] The second teacher model is used to determine the matching degree between the p pieces of text and the q pieces of text and the second text, and the texts whose matching degree is not less than the preset threshold are taken as the second result set.
[0053] Specifically, in step S4, determining the distilled data for knowledge distillation by concatenating the first text with the second result set and concatenating the second text with the first result set includes:
[0054] Concatenate the first text with each text in the second result set to obtain a first concatenated result set;
[0055] Concatenate each text in the first result set with the second text to obtain a second concatenated result set;
[0056] The first training set, the first splicing result set and the second splicing result set are combined to obtain the distilled data.
[0057] Embodiment 2:
[0058] A method for training a text matching model proposed in this embodiment mainly includes the following steps:
[0059] S1. Obtain a small amount of labeled data D1 in the field, i.e., the first training set, a large amount of unlabeled data D2 in the field, i.e., the second training set, and a public data set D3 of text matching in the public domain, i.e., the third training set.
[0060] The above-mentioned field corresponds to the first field mentioned above, that is, it refers to a specific field, such as news, medical or financial fields. The training set in a specific field has the problem of high acquisition cost, so the training set in the field is divided into a small amount of labeled data D1 and a large amount of unlabeled data D2. Data sets for text matching can be easily obtained in the public domain, for example, relevant data sets can be obtained through the Internet, including but not limited to the following data sets: Chinese-MNLI data set, Chinese-SNLI data set, OCNLI data set, Chinese-STS-B data set and PKU data set.
[0061] S2. Generate the first teacher model T1 and the first student model S1.
[0062] The public dataset D3 for text matching in the public domain has the advantages of being large and easily accessible, and the public dataset D3 can be used to perform the first step of training the pre-trained language model T0. Using data from the public domain to train the model allows the model to have a certain degree of prior knowledge, thereby reducing dependence on labeled data in a specific field. Therefore, the first step of this embodiment uses the public dataset D3 to train the pre-trained language model T0 to obtain the first teacher model T1. The first teacher model T1 is used to perform knowledge distillation on the pre-established small model S0 to be distilled to obtain the first student model S1.
[0063] Exemplarily, the pre-trained language model T0 may be a BERT model, and the small model S0 to be distilled may be a one-layer bidirectional LSTM model, which is not specifically limited in this embodiment.
[0064] S3. Generate a second teacher model T2, and establish a search engine and an approximate neighbor search library.
[0065] After step S2, the first teacher model T1 has the ability to judge the matching degree of the public domain text. Now what this embodiment needs to do is to let the first teacher model T1 learn more matching features in the field. This step can greatly enhance the accuracy of the student model after knowledge distillation. The steps are as follows:
[0066] S3.0 divides a small amount of in-domain annotated data D1 into a training set TRAIN, a validation set DEV and a test set TEST; the annotated data contained in the in-domain annotated data D1 corresponds to multiple categories; in one application scenario, for example, in a text matching task, the categories can be divided into 1 and 0, where category 1 indicates that the meanings of the two texts of the annotated data are close, and category 0 indicates that the meanings of the two texts of the annotated data are distant; in another application scenario, the categories can be divided into contradiction, entailment and neutral; the specific categories contained in the in-domain annotated data D1 are not specifically limited in this embodiment and can be set according to the specific application scenario;
[0067] S3.1 Fine-tune the teacher model T1 using the training set TRAIN, take the one with the highest prediction accuracy on the validation set DEV as the second teacher model T2, and determine on the test set TEST that the prediction accuracy of the second teacher model T2 for each category in the test set TEST is greater than the preset threshold of 0.95;
[0068] S3.2 stores a large amount of unlabeled data D2 in the field into the Elasticsearch database and creates an index;
[0069] S3.3 uses the second teacher model T2 to convert the unlabeled data D2 in the domain into vectors, stores the converted vectors into the Faiss framework, and creates an index.
[0070] S4. Generate target distillation data.
[0071] S4.0 traverses the labeled data D1 in the domain, each piece of data consists of two texts to be matched, namely the first text and the second text, named text_a and text_b respectively;
[0072] S4.1 For text_a, recall n texts through the Elasticsearch database, recall m texts through the Faiss framework, and then use the second teacher model T2 to infer the matching degree between the n+m texts and text_a to obtain the inference score; select the text whose inference score is not less than the preset threshold from the n+m texts to obtain the first result set;
[0073] S4.2 For text_b, recall p texts through the Elasticsearch database, recall q texts through the Faiss framework, and then use the second teacher model T2 to infer the matching degree between the p+q texts and text_b to obtain the inference score; select the texts whose inference scores are not less than the preset threshold from the p+q texts to obtain the second result set;
[0074] S4.3 concatenates text_a with the text in the first result set, and concatenates the text in the second result set with text_b, to obtain an enhanced text set ENHANCE;
[0075] S4.4 merges the enhanced text set ENHANCE with the first training set to obtain distilled data for knowledge distillation.
[0076] S5. Generate a target model.
[0077] Based on the distilled data and the second teacher model T2, the first student model S1 is trained using the knowledge distillation method to generate a trained target model.
[0078] A method for training a text matching model in this embodiment can enhance the data of the training set in a targeted manner during the training process to improve the effect of knowledge distillation, including:
[0079] 1) Compared with random addition, deletion, modification and back-translation, which can be enhanced before model training, this method can enhance the vulnerable features of the neural network in a targeted manner during its training;
[0080] 2) Compared with random addition, deletion, modification and back-translation, which can be enhanced before model training, this method can enhance any data, greatly improving the generalization of the student model and its matching degree with the teacher model;
[0081] 3) Compared with adversarial enhancement methods such as autoencoders, this method is very simple and does not require additional training of a complex autoencoder.
[0082] Embodiment three:
[0083] The present invention also provides a device for training a text matching model, such as Figure 2 As shown, the device includes a processor 201, a memory 202, a bus 203, and a computer program stored in the memory 202 and executable on the processor 201. The processor 201 includes one or more processing cores. The memory 202 is connected to the processor 201 via the bus 203. The memory 202 is used to store program instructions. When the processor executes the computer program, the steps in the above method embodiment of the present invention are implemented.
[0084] Further, as an executable solution, the device for training the text matching model can be a computing device such as a desktop computer, a notebook, a PDA, and a cloud server. The system / electronic device may include, but is not limited to, a processor and a memory. Those skilled in the art will understand that the composition structure of the above-mentioned system / electronic device is only an example of the system / electronic device and does not constitute a limitation on the system / electronic device. It may include more or less components than the above, or a combination of certain components, or different components. For example, the system / electronic device may also include input and output devices, network access devices, buses, etc., which are not limited in the embodiments of the present invention.
[0085] Further, as an executable solution, the processor may be a central processing unit (CPU), other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or any conventional processor, etc. The processor is the control center of the system / electronic device, and uses various interfaces and lines to connect various parts of the entire system / electronic device.
[0086] The memory can be used to store computer programs and / or modules. The processor realizes various functions of the system / electronic device by running or executing the computer programs and / or modules stored in the memory and calling the data stored in the memory. The memory can mainly include a program storage area and a data storage area, wherein the program storage area can store an operating system and at least one application required for a function; the data storage area can store data created according to the use of the mobile phone, etc. In addition, the memory can include a high-speed random access memory, and can also include a non-volatile memory, such as a hard disk, a memory, a plug-in hard disk, a smart memory card (Smart Media Card, SMC), a secure digital (Secure Digital, SD) card, a flash card (Flash Card), at least one disk storage device, a flash memory device, or other volatile solid-state storage devices.
[0087] Embodiment 4:
[0088] The present invention also provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, the steps of the above method in the embodiment of the present invention are implemented.
[0089] If the module / unit integrated in the system / electronic device is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the present invention implements all or part of the processes in the above-mentioned embodiment method, and can also be completed by instructing the relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium, and the computer program can implement the steps of the above-mentioned various method embodiments when executed by the processor. Among them, the computer program includes computer program code, and the computer program code can be in source code form, object code form, executable file or some intermediate form. The computer-readable medium may include: any entity or device capable of carrying computer program code, recording medium, U disk, mobile hard disk, disk, optical disk, computer memory, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory) and software distribution medium, etc. It should be noted that the content contained in the computer-readable medium can be appropriately increased or decreased according to the requirements of legislation and patent practice in the jurisdiction.
[0090] Although the present invention has been specifically shown and described in conjunction with the preferred embodiments, it should be understood by those skilled in the art that various changes may be made to the present invention in form and details without departing from the spirit and scope of the present invention as defined by the appended claims, all of which are within the scope of protection of the present invention.
Claims
1. A method for training a text matching model, It is characterized in that Includes steps: S1, obtaining a first training set, a second training set and a third training set, wherein the first training set comprises labeled data for text matching in a first domain, the second training set comprises unlabeled data for text matching in the first domain, and the third training set comprises labeled data for text matching in the second domain; S2, inputting the third training set into a pre-trained language model, training the pre-trained language model based on a text matching task to obtain a candidate language model, using the candidate language model as a first teacher model in knowledge distillation, and using the first teacher model to perform knowledge distillation on a pre-established model to be distilled to obtain a first student model; S3, using the first training set to train the first teacher model to obtain a trained second teacher model, storing the unlabeled data in the second training set in a search engine, using the second teacher model to extract feature vectors corresponding to the unlabeled data in the second training set, storing the feature vectors in an approximate nearest neighbor search library, the search engine segmenting the unlabeled data in the second training set to establish an inverted index, and the approximate nearest neighbor search library using a predetermined approximate nearest neighbor search algorithm to index the feature vectors; S4, the annotated data in the first training set includes a first text and a second text, a first result set is obtained from the search engine and the approximate neighbor search library according to the first text, a second result set is obtained from the search engine and the approximate neighbor search library according to the second text, and distilled data for knowledge distillation is determined by concatenating the first text with the second result set and concatenating the second text with the first result set; S5. Based on the distilled data and the second teacher model, the first student model is trained using a knowledge distillation method to generate a trained target model.
2. The method according to claim 1, It is characterized in that In step S3, the first teacher model is trained using the first training set to obtain a trained second teacher model, which includes: The labeled data contained in the first training set corresponds to multiple categories. The first training set is divided into a sub-training set, a sub-validation set and a sub-test set. The second teacher model is determined based on the sub-training set, the sub-validation set and the sub-test set, so that the prediction accuracy of the second teacher model in each category in the sub-test set is greater than a preset threshold, and the value range of the preset threshold is (0,1).
3. The method according to claim 2, It is characterized in that The preset threshold is 0.
95.
4. The method according to claim 1, It is characterized in that The search engine is the Elasticsearch engine, and the approximate neighbor search library is the Faiss framework.
5. The method according to claim 2, It is characterized in that The step S4 of acquiring a first result set from the search engine and the approximate neighbor search library according to the first text includes: Recalling n texts from the search engine according to the first text, where n is a natural number; Recalling m texts from the approximate nearest neighbor search library according to the first text, where m is a natural number; The second teacher model is used to determine the matching degree between the n texts and the m texts and the first text, and the texts whose matching degree is not less than the preset threshold are taken as the first result set.
6. The method according to claim 2, It is characterized in that The step S4 of acquiring a second result set from the search engine and the approximate neighbor search library according to the second text includes: Recalling p pieces of text from the search engine according to the second text, where p is a natural number; Recalling q texts from the approximate neighbor search library according to the second text, where q is a natural number; The second teacher model is used to determine the matching degree between the p pieces of text and the q pieces of text and the second text, and the texts whose matching degree is not less than the preset threshold are taken as the second result set.
7. The method according to claim 1, It is characterized in that Determining the distilled data for knowledge distillation by concatenating the first text with the second result set and concatenating the second text with the first result set in step S4 includes: Concatenate the first text with each text in the second result set to obtain a first concatenated result set; Concatenate each text in the first result set with the second text to obtain a second concatenated result set; The first training set, the first splicing result set and the second splicing result set are combined to obtain the distilled data.
8. A device for training a text matching model, It is characterized in that The method comprises a memory and a processor, wherein the memory stores at least one program, and the at least one program is executed by the processor to implement the method for training a text matching model as described in any one of claims 1 to 7.
9. A computer-readable storage medium, It is characterized in that The storage medium stores at least one program, and when the at least one program is run, the method for training a text matching model according to any one of claims 1 to 7 is executed.
Citation Information
Patent Citations
Cross-domain text classification model training method, classification method and device
CN111831826A
Multi-field adaptive end-to-end speech recognition method and system, and electronic device
CN113436616A