Model training method and device, electronic equipment and storage medium

Through repeated training of text matching models and updating error sample labels, the problem of improving text matching accuracy in the prior art is solved, and the controllability and accuracy of model optimization are achieved.

CN120030340APending Publication Date: 2025-05-23TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311584949.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-11-23
Publication Date
2025-05-23

AI Technical Summary

Technical Problem

When training text matching models in the prior art, it is difficult to effectively improve the accuracy of text matching, especially in the repair of error samples and model optimization, there are problems of link length and uncontrollability.

Method used

By obtaining the target error sample set and the original training data, the updated text matching model is repeatedly trained, and the label of the error sample is updated during the training process until the training end condition is met to obtain improved text matching accuracy.

Benefits of technology

This method can effectively repair error samples, shorten the link to optimize the model, reduce uncontrollability, and improve the accuracy and response speed of the text matching model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120030340A_ABST
    Figure CN120030340A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a model training method and device, electronic equipment and a storage medium, which are applied to the technical field of computers and can relate to the fields of artificial intelligence, machine learning and the like, in the embodiment of the invention, a target error sample set is acquired, the target error sample set comprises a plurality of error samples with first tags, and the first tags are used as first tags; each error sample comprises an unmatched text pair, then a to-be-updated text matching model and original training data are obtained, the to-be-updated text matching model is obtained by training an initial text matching model through the original training data, and the original training data comprise a plurality of training text pairs with second labels; the second label of one training text pair represents the second similarity between the training text pairs, and then based on the target error sample set and the original training data, training operation is repeatedly executed on the to-be-updated text matching model until the training ending condition is met, and a trained target text matching model is obtained.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application belongs to the field of computer technology. Specifically, the present application relates to a method, device, electronic device and storage medium for model training. Background Art

[0002] Text matching is a core issue in natural language processing. Many natural language processing tasks can be abstracted into text matching problems. For example, information retrieval can be reduced to the matching of search terms and document resources, question-answering systems can be reduced to the matching of questions and candidate answers, paraphrase questions can be reduced to the matching of two synonymous sentences, dialogue systems can be reduced to the matching of dialogues and replies, and machine translation can be reduced to the matching of two languages.

[0003] Currently, most companies in the industry train text matching models when implementing natural language processing tasks, and perform text matching based on the trained text matching models. For example, in a question-answering system, the target object inputs a query, and the question-answering system matches the target query in the database based on the query input by the target object, and then obtains the answer based on the target query matching. However, if the wrong target query is matched, the output answer may be wrong. Therefore, how to train the text matching model to improve the accuracy of text matching has become a key issue. Summary of the invention

[0004] The purpose of the embodiments of the present application is to provide a method, device, electronic device and storage medium for model training. In order to achieve the above purpose, the technical solutions provided by the embodiments of the present application are as follows:

[0005] In a first aspect, a model training method is provided, comprising:

[0006] Acquire a target error sample set, wherein the target error sample set includes a plurality of error samples with a first label, each of the error samples includes an unmatched text pair, and the first label represents a first similarity between the text pairs;

[0007] Obtaining a text matching model to be updated and original training data, wherein the text matching model to be updated is obtained by training an initial text matching model using the original training data, and the original training data includes a plurality of training text pairs with second labels, and the second label of a training text pair represents a second similarity between the training text pairs;

[0008] Based on the target error sample set and the original training data, the training operation is repeatedly performed on the text matching model to be updated until the training end condition is met to obtain a trained target text matching model, wherein the training operation includes the following steps:

[0009] Based on the target error example set and the original training data, perform at least one training on the text matching model obtained from the previous training operation to obtain a trained first model;

[0010] Perform a label update operation on each of the error examples, where the label update operation includes:

[0011] Through the first model, perform similarity recognition on each of the error examples respectively to obtain a third similarity corresponding to each of the error examples, and update the first label of the corresponding error example based on the third similarity corresponding to each of the error examples, and use the updated label as the first label of the error example during the next training operation.

[0012] In a possible implementation manner, the performing a label update operation on each of the error examples includes:

[0013] If the first condition is not satisfied, perform a label update operation on each of the error examples;

[0014] The training operation further includes:

[0015] If the first condition is satisfied, do not perform a label update operation on each of the error examples during the current training operation and each subsequent training operation after the current training operation;

[0016] Wherein, the satisfaction of the first condition includes satisfying at least one of the following:

[0017] The number of executed training operations reaches a preset number;

[0018] The label modification degree satisfies a preset condition, where the label modification degree is determined based on the first difference corresponding to each of the error examples, and the first difference corresponding to an error example is the absolute value of the difference between the third similarity of the error example obtained from the current training operation and the third similarity of the error example obtained from the previous training operation.

[0019] In another possible implementation manner, the satisfaction of the preset condition by the label modification degree includes at least one of the following:

[0020] The first difference corresponding to each of the error examples is less than a first threshold;

[0021] The sum of the first differences corresponding to each of the error examples is less than a second threshold;

[0022] The first ratio is greater than a set value, where the first ratio is the ratio of the error examples corresponding to the first difference less than the first threshold in the target error example set.

[0023] In another possible implementation, for each of the error samples, the training operation further includes:

[0024] Determine a second difference corresponding to the error sample, where the second difference is an absolute value of a difference between a third similarity corresponding to the error sample and the first similarity;

[0025] If the second difference corresponding to the error sample is greater than a second preset threshold, and there is a target training text pair matching the error sample in the original training data, the target training text pair is deleted from the original training data, and the original training data with the target training text pair deleted is used as the original training data for the next training operation.

[0026] In another possible implementation, for each of the error samples, updating the first label of the corresponding error sample based on the third similarity corresponding to each of the error samples includes:

[0027] If the second difference corresponding to the error sample is less than a third preset threshold, determining the third similarity corresponding to the error sample as the updated first label of the error sample;

[0028] If the second difference corresponding to the error sample is not less than the third preset threshold and not greater than the second preset threshold, the third similarity corresponding to the error sample or the first label corresponding to the sample is determined as the updated first label of the error sample.

[0029] In another possible implementation, obtaining a target error sample set includes:

[0030] Obtaining the multiple error samples;

[0031] Based on the text matching model to be updated, the texts corresponding to the respective error samples are matched respectively to obtain the fourth similarities corresponding to the respective error samples;

[0032] Based on the fourth similarities corresponding to the respective error samples, the first labels corresponding to the respective error samples are determined.

[0033] In another possible implementation, each of the error samples includes: a first text and a second text;

[0034] Before determining the first labels corresponding to the respective error samples based on the fourth similarities corresponding to the respective error samples, the method further includes:

[0035] For each error sample, extract a first intention representation and a second intention representation, wherein the first intention representation is the intention representation of the first text in a target scenario, and the second intention representation is the intention representation of the second text in the target scenario;

[0036] Determine the intent similarity between the first intent representation and the second intent representation as the intent similarity corresponding to the error sample;

[0037] The determining of the first labels corresponding to the respective error samples based on the fourth similarities corresponding to the respective error samples includes:

[0038] Based on the intention similarities corresponding to the error samples, the fourth similarities corresponding to the error samples are updated to obtain the first labels corresponding to the error samples.

[0039] In another possible implementation, the updating of the fourth similarities corresponding to the respective error samples based on the intention similarities corresponding to the respective error samples to obtain the first labels corresponding to the respective error samples includes:

[0040] For each of the error samples, the fourth similarity corresponding to the error sample is updated in the following manner to obtain the corresponding first similarity:

[0041] If the intention similarity corresponding to the error sample is less than the fourth preset threshold and greater than the fifth preset threshold, adjusting the fourth similarity corresponding to the error sample to the first target interval to obtain the corresponding first similarity;

[0042] If the intention similarity corresponding to each of the error samples is not less than the fourth preset threshold, the fourth similarity corresponding to the error sample is adjusted to the second target interval to obtain the corresponding first similarity, the maximum value of the second target interval is greater than the maximum value of the first target interval, the minimum value of the second target area is greater than the minimum value of the first target interval, and the difference between the minimum value of the second target interval and the maximum value of the first target interval is less than a preset value;

[0043] If the intention similarity corresponding to the error sample is less than the fifth preset threshold, the character similarity of the error sample is determined, and based on the character similarity, the corresponding fourth similarity is updated within the third target interval to obtain the corresponding first similarity, and the maximum value in the third target interval is less than the minimum value in the first target interval;

[0044] If the fourth similarity corresponding to the error sample does not belong to the fourth target interval, and the intended similarity of the error sample belongs to the first target interval, the fourth similarity corresponding to the error sample is adjusted to the first target interval to obtain the corresponding first similarity, and the first similarities corresponding to each error sample whose fourth similarity is adjusted to the first target interval are different.

[0045] In another possible implementation, the method further includes:

[0046] Acquire an initial sample set, wherein the initial sample set includes a plurality of initial error samples;

[0047] Filtering a prompt command from the instruction database, where the prompt command is used to prompt the data enhancement model to perform data enhancement processing on the multiple initial error samples;

[0048] Based on the initial sample set and the prompt command, data enhancement is performed using a trained data enhancement model to obtain enhanced error samples, wherein the relationship between text pairs in each enhanced error sample is obtained from the relationship between text pairs in each initial sample;

[0049] The target error sample set is determined based on the initial sample set and the error samples after the enhancement process.

[0050] In a second aspect, a model training device is provided, the device comprising:

[0051] A first acquisition module is used to acquire a target error sample set, wherein the target error sample set includes a plurality of error samples with a first label, each of the error samples includes an unmatched text pair, and the first label represents a first similarity between the text pairs;

[0052] A second acquisition module is used to acquire a text matching model to be updated and original training data, wherein the text matching model to be updated is obtained by training an initial text matching model with the original training data, and the original training data includes a plurality of training text pairs with second labels, and the second label of a training text pair represents a second similarity between the training text pairs;

[0053] A repeated training module is used to repeatedly perform a training operation on the text matching model to be updated based on the target error sample set and the original training data until a training end condition is met to obtain a trained target text matching model, wherein the training operation includes the following steps:

[0054] Based on the target error sample set and the original training data, the text matching model obtained in the last training operation is trained at least once to obtain a trained first model;

[0055] Performing a label update operation on each of the error samples, the label update operation comprising:

[0056] The first model is used to perform similarity identification on each of the error samples to obtain a third similarity corresponding to each of the error samples, and the first label of the corresponding error sample is updated based on the third similarity corresponding to each of the error samples, and the updated label is used as the first label of the error sample in the next training operation.

[0057] In a possible implementation, when performing a label update operation on each of the error samples, the repeated training module is specifically used to:

[0058] If the first condition is not met, performing a label update operation on each of the error samples;

[0059] The training operation further includes:

[0060] If the first condition is met, then in the current training operation and each training operation after the current training operation, the label update operation for each of the erroneous samples is not performed;

[0061] Wherein, satisfying the first condition includes satisfying at least one of the following:

[0062] The number of training operations executed reaches the preset number;

[0063] The degree of label modification meets a preset condition, wherein the degree of label modification is determined based on the first difference corresponding to each of the error samples, and the first difference corresponding to an error sample is the absolute value of the difference between the third similarity of the error sample obtained in the current training operation and the third similarity of the error sample obtained in the previous training operation.

[0064] In another possible implementation, the label modification degree satisfies a preset condition including at least one of the following:

[0065] The first difference values ​​corresponding to the error samples are all smaller than the first threshold value;

[0066] The sum of the first difference values ​​corresponding to the error samples is less than a second threshold;

[0067] The first proportion is greater than a set value, and the first proportion is the proportion of error samples whose corresponding first difference values ​​are less than a first threshold in the target error sample set.

[0068] In another possible implementation, for each of the error samples, the training operation further includes:

[0069] Determine a second difference corresponding to the error sample, where the second difference is an absolute value of a difference between a third similarity corresponding to the error sample and the first similarity;

[0070] If the second difference corresponding to the error sample is greater than a second preset threshold, and there is a target training text pair matching the error sample in the original training data, the target training text pair is deleted from the original training data, and the original training data with the target training text pair deleted is used as the original training data for the next training operation.

[0071] In another possible implementation, for each of the error samples, when the repeated training module updates the first label of the corresponding error sample based on the third similarity corresponding to each of the error samples, it is specifically configured to:

[0072] If the second difference corresponding to the error sample is less than a third preset threshold, determining the third similarity corresponding to the error sample as the updated first label of the error sample;

[0073] If the second difference corresponding to the error sample is not less than the third preset threshold and not greater than the second preset threshold, the third similarity corresponding to the error sample or the first label corresponding to the sample is determined as the updated first label of the error sample.

[0074] In another possible implementation, when acquiring the target error sample set, the first acquisition module is specifically configured to:

[0075] Obtaining the multiple error samples;

[0076] Based on the text matching model to be updated, the texts corresponding to the respective error samples are matched respectively to obtain the fourth similarities corresponding to the respective error samples;

[0077] Based on the fourth similarities corresponding to the respective error samples, the first labels corresponding to the respective error samples are determined.

[0078] In another possible implementation, each of the error samples includes: a first text and a second text;

[0079] The device further comprises: an extraction module and a similarity determination module, wherein:

[0080] The extraction module is used to extract a first intention representation and a second intention representation for each error sample, wherein the first intention representation is an intention representation of the first text in a target scenario, and the second intention representation is an intention representation of the second text in the target scenario;

[0081] The similarity determination module is used to determine the intent similarity between the first intent representation and the second intent representation as the intent similarity corresponding to the error sample;

[0082] Wherein, when the first acquisition module determines the first labels corresponding to the respective error samples based on the fourth similarities corresponding to the respective error samples, it is specifically used to:

[0083] Based on the intention similarities corresponding to the error samples, the fourth similarities corresponding to the error samples are updated to obtain the first labels corresponding to the error samples.

[0084] In another possible implementation, when the first acquisition module updates the fourth similarities corresponding to the respective error samples based on the intention similarities corresponding to the respective error samples to obtain the first labels corresponding to the respective error samples, it is specifically configured to:

[0085] For each of the error samples, the fourth similarity corresponding to the error sample is updated in the following manner to obtain the corresponding first similarity:

[0086] If the intention similarity corresponding to the error sample is less than the fourth preset threshold and greater than the fifth preset threshold, adjusting the fourth similarity corresponding to the error sample to the first target interval to obtain the corresponding first similarity;

[0087] If the intention similarity corresponding to each of the error samples is not less than the fourth preset threshold, the fourth similarity corresponding to the error sample is adjusted to the second target interval to obtain the corresponding first similarity, the maximum value of the second target interval is greater than the maximum value of the first target interval, the minimum value of the second target area is greater than the minimum value of the first target interval, and the difference between the minimum value of the second target interval and the maximum value of the first target interval is less than a preset value;

[0088] If the intention similarity corresponding to the error sample is less than the fifth preset threshold, the character similarity of the error sample is determined, and based on the character similarity, the corresponding fourth similarity is updated within the third target interval to obtain the corresponding first similarity, and the maximum value in the third target interval is less than the minimum value in the first target interval;

[0089] If the fourth similarity corresponding to the error sample does not belong to the fourth target interval, and the intended similarity of the error sample belongs to the first target interval, the fourth similarity corresponding to the error sample is adjusted to the first target interval to obtain the corresponding first similarity, and the first similarities corresponding to each error sample whose fourth similarity is adjusted to the first target interval are different.

[0090] In another possible implementation, the device further includes: a third acquisition module, a screening module, a data enhancement processing module, and a target sample set determination module, wherein:

[0091] The third acquisition module is used to acquire an initial sample set, wherein the initial sample set includes a plurality of initial error samples;

[0092] The screening module is used to screen out prompt commands from the instruction database, wherein the prompt commands are used to prompt the data enhancement model to perform data enhancement processing on the multiple initial error samples;

[0093] The data enhancement processing module is used to perform data enhancement based on the initial sample set and the prompt command using a trained data enhancement model to obtain enhanced error samples, wherein the relationship between text pairs in each enhanced error sample is obtained from the relationship between text pairs in each initial sample;

[0094] The target sample set determination module is used to determine the target error sample set based on the initial sample set and the error samples after the enhancement process.

[0095] In a third aspect, an embodiment of the present application further provides an electronic device, comprising a memory and a processor, wherein a computer program is stored in the memory, and the processor executes the computer program to implement the model training method provided by any possible implementation method of the first aspect.

[0096] In a fourth aspect, an embodiment of the present application further provides a computer-readable storage medium, which stores a computer program, and when the computer program is executed by a processor, the model training method provided by any possible implementation method of the first aspect is implemented.

[0097] In the fifth aspect, an embodiment of the present application further provides a computer program product, which includes a computer program, and when the computer program is executed by a processor, it implements the model training method provided by any possible implementation method of the fifth aspect.

[0098] The beneficial effects of the technical solution provided by the embodiment of the present application are as follows:

[0099] The embodiments of the present application provide a method, device, electronic device and storage medium for model training. In the embodiments of the present application, a plurality of error samples with a first label, a text matching model to be updated and original training data used for training the text matching model to be updated are obtained, and then based on the original training data of the target error sample set, the training operation is repeatedly performed on the text matching model to be updated until the training end condition is met to obtain a trained text matching model, and in the process of repeatedly performing the training, after the text matching model obtained in the previous operation is trained at least once by the target error sample set and the original training data, a label update operation is also performed on the error sample using the training model to update the label corresponding to the error sample, and then the model is updated based on the target error sample set after the label is updated and the original training data until the training end condition is met to obtain a trained target text matching model, because when the model is cyclically trained and updated, the label in the target error sample set also needs to be updated to improve the accuracy of the training samples, thereby improving the accuracy of text matching of the trained model. BRIEF DESCRIPTION OF THE DRAWINGS

[0100] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings required for use in describing the embodiments of the present application are briefly introduced below.

[0101] Figure 1 A flowchart of model training in related technologies;

[0102] Figure 2 This is a schematic diagram of the model structure of the Teacher Model;

[0103] Figure 3 This is a schematic diagram of the network structure of the Student Model;

[0104] Figure 4 This is a flow chart of a method for model training in an embodiment of the present application;

[0105] Figure 5 A schematic diagram of a sub-process of a model training method in an embodiment of the present application;

[0106] Figure 6 A schematic diagram of applying the model for matching in a social security scenario in an embodiment of the present application;

[0107] Figure 7 A schematic diagram of performing model training in a specific scenario in an embodiment of the present application;

[0108] Figure 8 A schematic diagram of a device structure for model training in an embodiment of the present application;

[0109] Fig. 9 It is a schematic diagram of the device structure of an electronic device in an embodiment of the present application. DETAILED DESCRIPTION

[0110] The embodiments of the present application are described below in conjunction with the drawings in the present application. It should be understood that the implementation methods described below in conjunction with the drawings are exemplary descriptions for explaining the technical solutions of the embodiments of the present application and do not constitute a limitation on the technical solutions of the embodiments of the present application.

[0111] It will be understood by those skilled in the art that, unless specifically stated, the singular forms "one", "said", and "the" used herein may also include plural forms. It should be further understood that the terms "including" and "comprising" used in the embodiments of the present application refer to that the corresponding features can be implemented as the presented features, information, data, steps, operations, elements and / or components, but do not exclude the implementation as other features, information, data, steps, operations, elements, components and / or combinations thereof supported by the technical field. It should be understood that when we say that an element is "connected" or "coupled" to another element, the one element may be directly connected or coupled to the other element, or it may refer to that the one element and the other element establish a connection relationship through an intermediate element. In addition, the "connection" or "coupling" used herein may include wireless connection or wireless coupling. The term "and / or" used herein indicates at least one of the items defined by the term, for example, "A and / or B" may be implemented as "A", or as "B", or as "A and B". When describing multiple (two or more) items, if the relationship between the multiple items is not clearly defined, the multiple items may refer to one, multiple or all of the multiple items. For example, the description of "parameter A includes A1, A2, A3" can be implemented as parameter A including A1 or A2 or A3, and can also be implemented as parameter A including at least two of the three items A1, A2, A3.

[0112] Artificial Intelligence (AI) is the theory, method, technology and application system that uses digital computers or machines controlled by digital computers to simulate, extend and expand human intelligence, perceive the environment, acquire knowledge and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science that attempts to understand the essence of intelligence and produce a new intelligent machine that can respond in a similar way to human intelligence. Artificial intelligence is to study the design principles and implementation methods of various intelligent machines so that machines have the functions of perception, reasoning and decision-making.

[0113] Artificial intelligence technology is a comprehensive discipline that covers a wide range of fields, including both hardware-level and software-level technologies. Basic artificial intelligence technologies generally include sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, pre-trained model technology, operation / interaction systems, mechatronics, etc. Among them, pre-trained models are also called large models and basic models. After fine-tuning, they can be widely used in downstream tasks in various major directions of artificial intelligence. Artificial intelligence software technology mainly includes computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning.

[0114] Machine Learning (ML) is a multi-disciplinary interdisciplinary subject involving probability theory, statistics, approximation theory, convex analysis, algorithm complexity theory and other disciplines. It specializes in studying how computers simulate or implement human learning behavior to acquire new knowledge or skills and reorganize existing knowledge structures to continuously improve their performance. Machine learning is the core of artificial intelligence and the fundamental way to make computers intelligent. Its applications are spread across all areas of artificial intelligence. Machine learning and deep learning usually include artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and learning by teaching. The pre-trained model is the latest development of deep learning, which integrates the above technologies.

[0115] Before further describing the embodiments of the present application in detail, the nouns and terms involved in the embodiments of the present application are explained. The nouns and terms involved in the embodiments of the present application are subject to the following interpretations.

[0116] Text matching model: used to calculate the similarity between the user query and the query in the preset knowledge base. Example: Applying for a social security card vs. how to apply for a social security card;

[0117] Bi-encoder: For example, in Dense Passage Retrieval (DPR), two texts are respectively represented by the model, and finally the similarity between the two text representations is calculated through a relevance discriminant function. Because the interaction occurs only in the final relevance discriminant function, the text representation of the candidate set can be calculated offline, which is very suitable for industrial scenarios and has extremely high matching efficiency. However, the disadvantage is that the fine-grained interaction features between the two texts are ignored, resulting in poor matching quality.

[0118] Cross-encoder: For example, BERT, which concatenates two texts and directly outputs the similarity through BERT. The advantage is that it is simple and can achieve deep text interaction. The disadvantage is that it cannot be used in the recall stage due to the large amount of calculation.

[0119] Late-interaction encoder: For example, ColBERT, based on the Bi-encoder, performs fine-grained representation interaction at the final stage to obtain the final similarity; since the two towers are still separated, the amount of calculation is small and a more accurate similarity can be obtained, but due to the use of the post-interaction sum operation, it is not possible to directly use ANN for retrieval;

[0120] Flexible interaction encoder, which represents the combination between Encoder and InteractionMode, and there can be a rich combination relationship between the two;

[0121] In the embodiments of the present application, text matching is a core issue in natural language processing. Many natural language processing tasks can be abstracted into text matching problems. For example, information retrieval can be reduced to the matching of search terms and document resources, the question-answering system can be reduced to the matching of questions and candidate answers, the paraphrase question can be reduced to the matching of two synonymous sentences, the dialogue system can be reduced to the matching of dialogue and reply, and machine translation can be reduced to the matching of two languages.

[0122] There are four general computational paradigms for text matching: (a) Bi-encoder, (b) Cross-encoder, (c) Late-interaction encoder, and (d) Flexible interaction encoder. The performance is: b>d>c>a.

[0123] Currently, most companies in the industry use fine-grained annotation-based text matching models to implement intelligent customer service knowledge question answering. In addition, in order to improve the effect of the text matching model, they use multi-view training of the Teachers-Student model. This training process, whether incremental learning, full learning, or continuous learning, is bound to be more uncontrollable and cumbersome when optimizing error samples than optimizing a single model.

[0124] Furthermore, the relevant technical processes are as follows Figure 1 As shown, it can be divided into four stages:

[0125] Phase 1 includes data stream 1 and data stream 2, which are mainly used to obtain training data for the Teacher Model. Specifically, the error samples (in the form of sentence pairs) are rewritten into triples and mixed with the basic training data (Common Data, CD);

[0126] Phase 2 includes data stream 3, which is mainly used to train the Multi-View Teachers Model (composed of Teacher1Model, Teacher2 Model, ..., and TeacherN Model). The model structure of the Teacher Model is as follows: Figure 2 As shown, it is a ranking model based on Sentence Bert; in the embodiment of the present application, the positive sample (Doc+), the query input by the target object, and the negative sample (Doc-) are respectively encoded by their respective encoders (Encoder) to obtain their respective corresponding encoding representations (R doc+ , R query , R doc- ), then based on R doc+ and R query , calculate the similarity between the positive sample and the query (S θ (Q, DoC+)), and, based on R doc- and R query , calculate the similarity between the negative sample and the query (S θ (Q, DoC-)), then based on (S θ (Q, DoC+)) and (S θ (Q, DoC-)) calculate the loss Loss.

[0127] Phase 3 includes data streams 4 and 5, which are mainly used to prepare the training data for the Student Model. Specifically, the mixed triples in Phase 1 are split into sentence pairs again, and the sentence pairs are scored using the N Teacher Models trained in Phase 2. The N scores are weighted averaged as the training score for the Phase 4 model.

[0128] Phase 4 includes data stream 6, which is mainly used to train the Student Model. The Student Model is a regression model based on the Cross-encoder. The model network structure is as follows: Figure 3 Specifically, the training samples and queries in the Doc library are encoded through the BERT model to obtain the corresponding vector representation R, and then passed through the full connection layer (FullConnection, FC), and then the FC output results are subjected to logistic regression (Logistic regression) to obtain the logistic regression results.

[0129] Then, the related technical solutions have the following problems: 1. The link for repairing error samples is too long, which makes the effect of model retraining uncontrollable and the training cost high; 2. There is a requirement for the number of error samples, and the model can only be retrained when the number of error samples accumulates to 500-1000. In this way, when a key customer (KeyAccount, KA) reports a small number of error samples, it is impossible to complete the optimization of the model within two weeks.

[0130] Based on this, the embodiment of the present application proposes a method for directional repair of error samples that can improve the above-mentioned problems at the same time. In the embodiment of the present application, error samples are first collected, and data enhancement is performed based on a large model or manually; secondly, the error samples are pseudo-labeled in combination with data labeling rules and business requirements (these data are recorded as new data); then, the new data is mixed with the training data of the original model and the regression model is retrained; specifically, the stages one to three of the above-mentioned related technical solutions are to obtain the "similarity score" of the error samples, and it is expected to repair the error samples by participating in the training of the Student Model. In the embodiment of the present application, stages one to three can be removed, and pseudo labels can be assigned to erroneous samples with the help of data annotation rules and business requirements; in another possible implementation method, a self-learning method is introduced to enable the Student Model to learn appropriate label values ​​for erroneous samples; specifically, the erroneous samples can be first pseudo-labeled manually with the help of data annotation rules and business requirements, or pseudo-labeled by other means, and then mixed with the basic training data CD to retrain the Student Model, and then the Student Model is allowed to score the erroneous samples and update their pseudo-labels according to the update rules; repeating 1 to 2 cycles can enable the original model to repair the erroneous samples without affecting the original test set indicators.

[0131] Furthermore, based on the above embodiments, a model training method is introduced in detail through a specific embodiment, which is as follows:

[0132] The present application embodiment provides a model training method, which is performed by an electronic device, such as Figure 4 As shown, the method may include:

[0133] Step S401: Obtain a target error sample set.

[0134] The target error sample set includes multiple error samples with a first label, each error sample includes an unmatched text pair. In the embodiment of the present application, an unmatched text pair generally refers to a text pair whose recognition result does not match the intention of the text pair itself when the model is used for text matching. The first label represents the first similarity between the text pairs. That is, the target error sample set is composed of the error sample sentence pair set D = (d1 , d 2 , …, d n ) is obtained after annotating each text pair, d 1 , d 2 , …, d n Represent n pairs of texts respectively.

[0135] Specifically, in an embodiment of the present application, a text pair includes a first text and a second text. The first text and the second text themselves have a high similarity in intent, but when the model performs text matching, they are often considered to be an unmatched text pair, that is, the matching degree is low. At this time, the first text and the second text can be called an error sample; or, the first text and the second text themselves have a low degree of intent matching, but in model matching, they are considered to be a text pair with a high degree of matching. At this time, the text pair can also be called an error sample. For example, in the social security scenario, text pair 1 includes: text 1 and text 2, where text 1 is "receive social security card" and text 2 is "apply for social security card". The intent matching degree itself is low, but it is considered to be a high degree of intent matching in model matching. At this time, text 1 and text 2 belong to an error sample.

[0136] It should be noted that the text pairs in the target error sample set may be the text pairs obtained in the text matching process through the initial matching model, and of course may also be the text pairs obtained by other means, which is not limited in the embodiments of the present application.

[0137] Furthermore, each text pair in the target error sample set may be obtained for a certain scenario or for multiple scenarios, which is not limited in the embodiments of the present application.

[0138] For example, each text pair in the target error sample set may be obtained for a social security scenario, or may be obtained for a social security scenario and a provident fund scenario.

[0139] Step S402: Obtain the text matching model to be updated and the original training data.

[0140] The text matching model to be updated is obtained by training the initial text matching model with original training data, wherein the original training data includes a plurality of training text pairs with second labels, and the second label of a training text pair represents the second similarity between the training text pairs.

[0141] In the embodiment of the present application, the original training data may be all the training data used to obtain the text matching model to be updated, or may be part of the training data used to obtain the text matching model to be updated, which is not limited in the embodiment of the present application. For example, 2000 training data are required to train the text matching model to be updated from the initial text matching model, wherein the original training data used subsequently may be all the training data, that is, 2000 training data, or may be part of the 2000 training data, such as 1000 training data.

[0142] Furthermore, in an embodiment of the present application, the text matching model to be updated may also be an initial text matching model, the original training data is part of the training data to be used to train the initial text matching model, and the other part of the training data may be the target error sample set shown in the above embodiment.

[0143] It should be noted that in the embodiment of the present application, the target error sample set can be obtained first and then the text matching model to be updated and the original training data can be obtained. The text matching model to be updated and the original training data can also be obtained first, and then the target error sample set can be obtained. Of course, the target error sample set, the text matching model to be updated and the original training data can also be obtained at the same time.

[0144] Furthermore, for the text matching model to be updated and the original training data, the text matching model to be updated can be obtained first, and then the original training data can be obtained, or the original training data can be obtained first, and then the text matching model to be updated can be obtained, or the original training data and the text matching model to be updated can be obtained at the same time, which is not limited in the embodiments of the present application.

[0145] Step S403: Based on the target error sample set and the original training data, the training operation is repeatedly performed on the updated text matching model until the training end condition is met, thereby obtaining a trained target text matching model.

[0146] That is, after obtaining the target error sample set and the original training data through the above-mentioned embodiment, the text matching model to be updated is repeatedly trained through the target error sample set and the original training data until the training end condition is met. In the embodiment of the present application, the training end condition may include: the number of training times reaches a preset number threshold, the loss function reaches fitting, or the result obtained by the test training set through the trained target text matching model meets at least one of the requirements. For example, the test training set shown here can be the test training set used in the process of obtaining the text matching model to be updated by training the initial text matching model, or it can be a newly acquired test training set.

[0147] Furthermore, in an embodiment of the present application, in the process of updating when repairing erroneous samples to obtain a trained target text matching model, 1-2 repeated trainings are generally required, so that a certain stability can be achieved, that is, the training conditions are met, so as to ensure that the indicators of the test training set do not decrease, while also achieving the effect of repairing erroneous samples.

[0148] The training operation includes the following steps, namely step Sa to step Sb. Figure 5 As shown,

[0149] Step Sa: Based on the target error sample set and the original training data, the text matching model obtained by the previous training operation is trained at least once to obtain a trained first model.

[0150] For the embodiment of the present application, after obtaining the target error sample set and the original training data, the target error sample set and the original training data are mixed, and then the text matching model obtained by the previous training operation is trained at least once based on the mixed data to obtain a trained first model.

[0151] For example, the target error sample set is represented by D+A′, where D is used to represent each text pair in the target error sample set, and A′=(a′ 1 ,a′ 2 ,…,a′ n ), and A′ is used to represent the target error sample set and the first label corresponding to each error sample. The original training data is represented by CD′, and the text training model can be a student model, that is, D+A′ and CD′ are mixed, and the mixed data is used to train the student model to obtain a trained student model.

[0152] Specifically, if this training is the first training, the text matching model obtained in the previous training operation is the text matching model to be updated; if this training is the Nth training, the text matching model obtained in the previous training operation is the text matching model obtained in the N-1th training.

[0153] Step Sb: perform label update operation on each error sample.

[0154] Specifically, in the embodiments of the present application, a label update operation may be performed after a training operation is performed, or a label update operation may be performed after at least two training operations are performed, which is not limited in the embodiments of the present application.

[0155] Specifically, in an embodiment of the present application, the label update operation includes: performing similarity identification on each error sample through the first model to obtain a third similarity corresponding to each error sample, and updating the first label of the corresponding error sample based on the third similarity corresponding to each error sample, and using the updated label as the first label of the error sample in the next training operation.

[0156] For the embodiment of the present application, the first model can be characterized by M1. Since the first label of any error sample is also used to characterize the first similarity between the text pairs of the error sample, after similarity recognition is performed by M1, the third similarity of the error sample is obtained.

[0157] Among them, for a certain error sample, its corresponding first similarity may be the same or different. For example, the target error sample set includes: error sample 1 and error sample 2, wherein the first similarity corresponding to error sample 1 is the same as its corresponding first similarity, and the first similarity corresponding to error sample 2 is different from its corresponding first similarity, or, the first similarity corresponding to error sample 1 is different from its corresponding first similarity, and the first similarity corresponding to error sample 2 is different from its corresponding first similarity. Of course, it is also possible that the first similarity corresponding to error sample 1 is the same as its corresponding first similarity, and the first similarity corresponding to error sample 2 is the same as its corresponding first similarity.

[0158] Furthermore, in an embodiment of the present application, for an error sample, the first label of the corresponding error sample is updated based on the third similarity corresponding to the error sample, and the updated label is used as the first label of the error sample in the next training operation. Specifically, it may include: the updated label of the error sample may be the third similarity corresponding to the error sample, and the third similarity corresponding to the error sample is used as the first label of the error sample in the next training operation.

[0159] For example, for error sample 1, the updated label of error sample 1, that is, the third similarity corresponding to error sample 1, that is, the third similarity corresponding to error sample 1 is used as the first label corresponding to error sample 1 in the next training.

[0160] Of course, the first label of the corresponding error sample may also be updated based on the third similarity corresponding to the error sample in other ways. Please refer to the following embodiment for details, which will not be described in detail here.

[0161] Another possible implementation of the embodiment of the present application, the method also includes: obtaining an initial sample set, the initial sample set includes multiple initial error samples; filtering out prompt commands from the instruction database, the prompt commands are used to prompt the data enhancement model to perform data enhancement processing on multiple initial error samples; based on the initial sample set and the prompt commands, and using the trained data enhancement model to perform data enhancement to obtain enhanced error samples, the relationship between the text pairs in each enhanced error sample and the relationship between the text pairs in each initial sample are obtained; based on the initial sample set and the enhanced error samples, determine the target error sample set. That is to say, after obtaining multiple initial error samples, the training samples obtained may be small. If the text matching model to be updated is trained only based on these initial error samples, it may lead to poor training effect. In order to improve the training effect, after obtaining these initial training samples, data enhancement can also be performed through a large model to obtain more error samples, so as to update the text matching model to be updated with these error samples and the initial training set obtained before. Furthermore, through the above implementation method, automatic mining can be carried out through algorithms, and large models can be introduced for corpus production to enrich the diversity of error samples. In this way, problems can be solved before KA customers find the problem, thereby improving the KA customer experience.

[0162] Furthermore, the method may also include: after obtaining the enhanced error sample sets, a target error sample set may be obtained only through these enhanced error sample sets to train the text matching model to be updated, or a partial error sample set may be obtained from these enhanced error sample sets, and the target error sample set may be determined based on these partial error sample sets to train the text matching model to be updated, or after obtaining these partial error sample sets, a target error sample set may be obtained based on these error sample sets and the above-mentioned initial training samples to train the text matching model to be updated.

[0163] Furthermore, in addition to performing data enhancement on the initial error samples through a large model to obtain more error samples, more error samples can also be obtained through manual labeling.

[0164] Furthermore, after obtaining multiple error samples in the target error sample set in the above manner, the labels of each error sample in these error samples can also be obtained, and then the target error sample set can be obtained. Specifically, obtaining the target error sample set in step S101 can specifically include: obtaining multiple error samples; matching the corresponding texts of each error sample based on the text matching model to be updated to obtain the fourth similarity corresponding to each error sample; determining the first label corresponding to each error sample based on the fourth similarity corresponding to each error sample. In other words, after these initial sample sets pass through the text matching model to be updated, the fourth similarity corresponding to each error sample is obtained, and the first label corresponding to each error sample is determined based on the fourth similarity corresponding to each error sample.

[0165] Specifically, in the embodiment of the present application, determining the first labels corresponding to the error samples based on the fourth similarities corresponding to the error samples may include: directly determining the fourth similarities corresponding to the error samples as the first labels corresponding to the error samples, or updating the fourth similarities corresponding to the error samples in other ways to determine the first labels corresponding to the error samples. For example, the fourth similarities corresponding to the error samples are determined by A=(a 1 、a 2 , …, a n ) is characterized, that is, based on A=(a 1 、a 2 , …, a n ) to update A′=(a′ 1 ,a′ 2 ,…,a′ n ).

[0166] Furthermore, in another possible implementation of the embodiment of the present application, the first labels corresponding to these error samples in the target error sample set can partially directly determine the corresponding fourth similarity as the corresponding first label, and another part obtains the corresponding first label after updating the fourth similarity in other ways.

[0167] For example, the target error sample set includes error sample 1 and error sample 2, wherein the first label of error sample 1 is the fourth similarity corresponding thereto, and the first label of error sample 2 is obtained by updating the fourth similarity in other ways.

[0168] Specifically, that is to say, the fourth similarity is updated based on other methods to obtain the first label, that is, based on the fourth similarity corresponding to each error sample, the first label corresponding to each error sample is determined. Specifically, it can also include: based on the intention similarity corresponding to each error sample, the fourth similarity corresponding to each error sample is updated to obtain the first label corresponding to each error sample. In an embodiment of the present application, the intention representation corresponding to the error sample can characterize the intention similarity of the text pair in the error sample. The fourth similarity of the error sample is updated by the intention similarity of the text pair in the error sample, so that the updated similarity is more in line with the actual intention of the text pair, that is, the accuracy of the first label corresponding to each error sample can be improved, and then the accuracy of the trained first model in text matching can be improved.

[0169] Specifically, each error sample includes: a first text and a second text; based on the fourth similarities corresponding to each error sample, determining the first label corresponding to each error sample, and also including: for each error sample, extracting the first intention representation and the second intention representation, determining the intention similarity between the first intention representation and the second intention representation, as the intention similarity corresponding to the error sample.

[0170] Among them, the first intention representation is the intention representation of the first text in the target scenario, and the second intention representation is the intention representation of the second text in the target scenario; in an embodiment of the present application, the intention similarity model can be used to extract each error sample (the first text and the second text) to obtain the first intention representation corresponding to the first text and the second intention representation corresponding to the second text, and after obtaining the first intention representation corresponding to the first text and the second intention representation corresponding to the second text, the similarity between the first intention representation and the second intention representation is determined, and then, after obtaining the similarity between the first intention representation and the second intention representation, the corresponding fourth similarity is updated based on the intention similarity.

[0171] Specifically, based on the intention similarities corresponding to the respective error samples, the fourth similarities corresponding to the respective error samples are updated to obtain the first labels corresponding to the respective error samples, specifically including: for each error sample, the fourth similarity corresponding to the error sample is updated by the following methods (method 1, method 2, method 3 and method 4) to obtain the corresponding first similarity, wherein,

[0172] Method 1: If the intention similarity corresponding to the error sample is less than the fourth preset threshold and greater than the fifth preset threshold, the fourth similarity corresponding to the error sample is adjusted to the first target interval to obtain the corresponding first similarity.

[0173] For the embodiment of the present application, if the intention similarity is not less than the fourth preset threshold, it indicates that the intention matching degree between the first text and the second text in the error sample in the target scenario is high, which is a direct hit situation. For example, in an intelligent question-and-answer scenario, if the target object inputs the first query, the system can directly match the second query in the system based on the first query, and then give an accurate reply, that is, the similarity between the first query and the second query is high, which is a direct hit situation.

[0174] If the intent similarity is less than the fourth preset threshold but greater than the fifth preset threshold, it means that the intent matching degree between the first text and the second text in the error sample in the target scenario is medium, which is a recommendation situation. Continuing with the previous example, in the intelligent question and answer scenario, if the target object inputs the first query, the system may not be able to accurately match the second query directly in the system based on the first query, but instead matches multiple candidate queries. At this time, multiple candidate queries are returned to the target object, and the corresponding reply is returned only after the target object selects the target query from the multiple candidate queries. This is a recommendation situation.

[0175] That is to say, through the intent similarity recognition, the intent similarity corresponding to the error sample is less than the fourth preset threshold and greater than the fifth preset threshold, that is, at this time it may belong to the recommendation situation, then the fourth similarity corresponding to the error sample is adjusted to the first target interval to obtain the corresponding first similarity. In the embodiment of the present application, the first target interval is the similarity area corresponding to the error sample (text pair) belonging to the recommendation situation.

[0176] Furthermore, for an error sample whose intention similarity is less than the fourth preset threshold and greater than the fifth preset threshold, if the corresponding fourth similarity belongs to the first target interval, there is no need to adjust the fourth similarity corresponding to the error sample.

[0177] Method 2: If the intention similarity corresponding to each error sample is not less than the fourth preset threshold, the fourth similarity corresponding to the error sample is adjusted to the second target interval to obtain the corresponding first similarity.

[0178] Specifically, it can be seen from the above embodiments that if the intention similarity corresponding to a certain error sample is not less than the fourth preset threshold, it indicates that the intention similarity between the first text and the second text in the error sample in the target scene is high, which is a direct hit. At this time, the fourth similarity corresponding to the error sample is adjusted to the second target interval to obtain the corresponding first similarity.

[0179] Furthermore, for an error sample whose intention similarity is not less than a fourth preset threshold, if its corresponding fourth similarity belongs to the second target interval, it is not necessary to adjust the fourth similarity corresponding to the error sample.

[0180] The maximum value of the second target interval is greater than the maximum value of the first target interval, the minimum value of the second target interval is greater than the minimum value of the first target interval, and the difference between the minimum value of the second target interval and the maximum value of the first target interval is less than a preset value. For example, the first target interval is [a, b], and the second target interval can be [b-a1, b+a2], where a1 and a2 can be the same or different.

[0181] Method 3: If the intention similarity corresponding to the error sample is less than the fifth preset threshold, determine the character similarity of the error sample, and based on the character similarity, update the corresponding fourth similarity within the third target interval to obtain the corresponding first similarity.

[0182] For the embodiment of the present application, if the intent similarity is less than the fifth preset threshold, it indicates that the intent matching degree between the first text and the second text in the error sample in the target scenario is low, and it is a recommendation failure. In the intelligent question-answering system, if the target object inputs the first query, the system may not be able to accurately match the second query directly in the system based on the first query, nor can it match the candidate query. In this case, it is a recommendation failure.

[0183] Specifically, if the intention similarity corresponding to the error sample is less than the fifth preset threshold, the character similarity of the error sample is determined, and based on the character similarity, the corresponding fourth similarity is updated within the third target interval to obtain the corresponding first similarity. Among them, the higher the character similarity, the lower its fourth similarity needs to be adjusted. For example, the intention similarities corresponding to error sample 1 and error sample 2 are both less than the fifth preset threshold, but the similarity between the two texts of error sample 1 is greater than the similarity between the two texts of error sample 2. At this time, the fourth similarity of error sample 1 needs to be adjusted to be lower than the adjusted fourth similarity of error sample 2.

[0184] The maximum value in the third target interval is smaller than the minimum value in the first target interval. For example, the third target interval is [c, d], where d is smaller than a.

[0185] Method 4: If the fourth similarity corresponding to the error sample does not belong to the fourth target interval, and the intended similarity of the error sample belongs to the first target interval, the fourth similarity corresponding to the error sample is adjusted to the first target interval to obtain the corresponding first similarity, and the first similarities corresponding to each error sample whose fourth similarity is adjusted to the first target interval are different.

[0186] Specifically, the minimum value of the fourth target interval is greater than the maximum value of the first target interval. In the embodiment of the present application, the fourth target interval belongs to the target interval corresponding to the similarity of the error sample that can be directly hit. For example, the fourth target interval is [e, f], where e is greater than b.

[0187] Furthermore, if a certain error sample is identified by similarity through the text matching model to be updated, its corresponding fourth similarity does not belong to the fourth target interval, but may belong to the first target interval, or the third target interval. At this time, if the intention similarity of the error sample belongs to the first target interval, or the intention similarity belongs to the fourth target interval, the fourth similarity corresponding to the error sample can be adjusted to the fourth target interval or the first target interval; in an embodiment of the present application, if there are multiple error samples that all belong to this situation, the similarities adjusted to the fourth target interval or the first target interval are different.

[0188] Furthermore, in another possible implementation of the embodiment of the present application, the fourth similarity corresponding to each error sample can be updated with reference to the labeling rules and business requirements to obtain the first label corresponding to each error sample. The labeling rules and business requirements can be as follows:

[0189] a) The hit is unreasonable but the recommendation is acceptable. This is adjusted to the recommendation range based on the existing scores, outside of some more similar corpus scores; for example, taking the intelligent question-answering scenario as an example, if the target object enters the first query, the system can directly match the second query in the system based on the first query, and then give an accurate reply, that is, the similarity between the first query and the second query is high, which is a direct hit situation, and the hit is not a hit situation; in the intelligent question-answering scenario, if the target object enters the first query, the system can be based on the first query, but cannot accurately match the second query directly in the system, but matches multiple candidate queries. At this time, multiple candidate queries are returned to the target object, and the target object selects the target query from the multiple candidate queries before returning the corresponding reply, which is a recommendation situation;

[0190] b) generate reasonable sentence pairs and adjust them to the range of direct hit threshold according to the existing scores;

[0191] c) No recommendation can be made, which is beyond the recommendation threshold according to the literal similarity. Continuing with the above example, in the intelligent question-answering system, if the target object inputs the first query, the system cannot accurately match the second query directly in the system based on the first query, nor can it match the candidate query. In this case, no recommendation can be made;

[0192] d) If misses need to be adjusted to recommendations or hits, batch shifts should be performed based on existing scores. It is best not to give the same score to sentence pairs with different similarities.

[0193] e) If there are more similar corpora, the scores will be lower than others, and the scores need to be adjusted according to the similarity of the corpora;

[0194] Furthermore, after the fourth similarity corresponding to the error sample is updated through the above embodiment to obtain the first similarity of the error sample, that is, after obtaining the target error sample set, the target error sample set and the original training data are mixed to repeat the training operation on the text matching model to be updated.

[0195] Specifically, the training operation includes: based on the target error sample set and the original training data, training the text matching model obtained in the previous training operation at least once to obtain a trained first model; and performing a label update operation on each of the error samples.

[0196] Furthermore, based on the target error sample set and the original training data, the text matching model obtained in the last training operation is trained at least once, and after obtaining the trained first model, the original training data may also need to be updated. In an embodiment of the present application, for each error sample, the training operation may also include: determining a second difference corresponding to the error sample; if the second difference corresponding to the error sample is greater than a second preset threshold, and there is a target training text pair matching the error sample in the original training data, then deleting the target training text pair from the original training data, and using the original training data from which the target training text pair has been deleted as the original training data for the next training operation.

[0197] The second difference is the absolute value of the difference between the third similarity corresponding to the error sample and the first similarity. In the embodiment of the present application, the third similarity corresponding to each error sample is represented by B = (b1, b2, ..., bn), and the second difference corresponding to the i-th error sample can be represented by |a i ′-b i | is represented by, where i can be 1, 2, ..., n, θ 1 Used to characterize the second preset threshold. In the embodiment of the present application, when |a i ′-b i |>θ 1, then it represents that the original training data contains the error samples in the target error samples, then it is necessary to find the training text pairs that overlap with the error samples in the target error samples from the original training data, that is, the target training pairs, and delete the target training pairs and their corresponding second labels from the original training data, and the original training data with the target training pairs and the corresponding second labels deleted is used as the original training data for the next training operation.

[0198] Specifically, in an embodiment of the present application, searching for training text pairs that overlap with error samples in target error samples from original training data may specifically include: extracting at least one keyword from each error sample, and determining a first repeated training sample from multiple training text pairs, wherein the first repeated training sample is at least one training text pair that contains the at least one keyword in the multiple training text pairs; or, performing character matching on each error sample and multiple training text pairs to obtain their respective corresponding matching degrees, and deleting training text pairs whose matching degrees are greater than a preset matching degree and the corresponding second similarity from the first original training data.

[0199] It should be noted that, in addition to deleting the original training data of the target training pair and the corresponding second label during the next training, the first similarities corresponding to each error sample can also be used as the target sample corresponding to the next training, or the third similarities corresponding to each training sample can be used as the first labels corresponding to each error sample during the next training.

[0200] Specifically, in the embodiment of the present application, performing a label update operation on each error sample may include: if the first condition is not met, performing a label update operation on each error sample. That is, if the number of training operations that have been executed does not reach a preset number of times, or the degree of label modification does not meet the preset condition, performing a label update operation on each error sample.

[0201] In an embodiment of the present application, for each error sample, the first label of the corresponding error sample is updated based on the third similarity corresponding to each error sample, which may specifically include: if the second difference corresponding to the error sample is less than a third preset threshold, the third similarity corresponding to the error sample is determined as the updated first label of the error sample.

[0202] Among them, in the embodiment of the present application, when |a i ′-b i |<θ 2 ,θ 2 It is used to represent the third preset threshold. At this time, the third similarity corresponding to the error sample is determined as the updated first label of the error sample, that is, the third similarity corresponding to the i-th error sample is used as the first label corresponding to the i-th error sample.

[0203] In another possible implementation of the embodiment of the present application, if the second difference corresponding to the error sample is not less than the third preset threshold and not greater than the second preset threshold, the third similarity corresponding to the error sample or the first label corresponding to the sample is determined as the updated first label of the error sample. i ′-b i |≥θ 2 , and, |a i ′-b i |≤θ 1 , then at this time a i ′=b i , that is, the third similarity of the i-th error sample is the same as the first similarity, that is, the third similarity or the first label can be used as the updated first label.

[0204] Furthermore, the training operation may also include: if the first condition is met, then in the current training operation and each training operation after the current training operation, the label update operation for each erroneous sample is not performed.

[0205] Wherein, satisfying the first condition includes satisfying that the number of training operations executed reaches the preset number, and the degree of label modification satisfies at least one of the preset conditions. That is, when the number of training operations executed reaches the preset number, the label update operation for each error sample is not performed in the current training operation and each training operation after the current training operation; or, when the degree of label modification satisfies the preset condition, the label update operation for each error sample is not performed in the current training operation and each training operation after the current training operation; or, when the number of training operations executed reaches the preset number, and the degree of label modification satisfies the preset condition, the label update operation for each error sample is not performed in the current training operation and each training operation after the current training operation.

[0206] Specifically, the degree of label modification is determined based on the first difference corresponding to each error sample, and the first difference corresponding to an error sample is the absolute value of the difference between the third similarity of the error sample obtained in the current training operation and the third similarity of the error sample obtained in the previous training operation. In the embodiment of the present application, the degree of label modification satisfies the preset conditions including: at least one of condition 1, condition 2 or condition 3, wherein condition 1 is that the first difference corresponding to each error sample is less than the first threshold; condition 2 is that the sum of the first difference corresponding to each error sample is less than the second threshold; condition 3 is that the first proportion is greater than the set value, and the first proportion is the proportion of error samples whose corresponding first difference is less than the first threshold in the target error sample set.

[0207] For example, the target error sample set includes error sample 1 and error sample 2. If the absolute value of the difference between the third similarity obtained in the current training and the third similarity obtained in the previous training corresponding to error sample 1 is less than the first threshold, and the absolute value of the difference between the third similarity obtained in the current training and the third similarity obtained in the previous training corresponding to error sample 2 is less than the first threshold, then in each training operation after the current training operation, the label update operation for each error sample is not performed, that is, only the text matching model is trained subsequently, and the label update for the error sample is no longer performed; if the sum of the first difference corresponding to error sample 1 and the first difference corresponding to error sample 2 is less than the second threshold, then in each training operation after the current training operation, the label update operation for each error sample is not performed. In the training operation, the label update operation for each error sample is not performed, that is, only the text matching model is trained subsequently, and the label update operation for the error sample is no longer performed; if the set value is 50%, such as the absolute value of the difference between the third similarity obtained in the current training corresponding to error sample 1 and the third similarity obtained in the previous training is less than the first threshold, and the absolute value of the difference between the third similarity obtained in the current training corresponding to error sample 2 and the third similarity obtained in the previous training is less than the first threshold, at this time, the proportion of error samples whose first difference is less than the first threshold in the target error sample set is 100%, that is, greater than 50%, and at this time, the label update operation for each error sample is not performed in each training operation after the current training operation.

[0208] Furthermore, by training and updating in the above-mentioned manner, a trained text matching model can be obtained, and it can be applied to various scenarios to repair the erroneous examples obtained by the previous model, and it can improve the accuracy of calculating the similarity between the query input by the target object and the query in the preset database. For example, in the social security scenario, if the method described in the embodiment of the present application is not applied to train and update the model, at this time, if the query input by the target object is "get a social security card", the query matched by the model is "apply for a social security card", and the specific method of "applying for a social security card" is returned; if the method described in the present application is applied to train and update the model, at this time, if the query input by the target object is "get a social security card", the query matched by the model is "I want to get a social security card", and the relevant information of "getting a social security card" is returned, as follows: Figure 6 shown.

[0209] The following is an introduction to the specific method of training the model in the embodiment of the present application through a specific application scenario. The sentence pair set of error samples is D = (d 1 , d 2 , …, d n ) as an example. Figure 7 As shown:

[0210] The first step is to assign pseudo labels: use the previous version of the model (the text matching model to be updated) to score the sentence pairs in set D and obtain a score set A = (a 1 、a 2 , …, a n ), that is, we get the scores a corresponding to each sentence pair in D;

[0211] In the second step, update the pseudo-labels of the first step, refer to the annotation rules and business requirements to adjust the score set A to obtain A′=(a′ 1 ,a′ 2 ,…,a′ n ), that is, at this time, we get the scores a′ corresponding to each sentence pair in D;

[0212] The third step is to mix the error sample D+A′ with the basic data CD′ and train the Student Model from scratch, denoted as: M1;

[0213] The fourth step is to score the new model M1 for D in the error sample set and obtain B = (b 1 , b 2 , …, b n );

[0214] Step 5: Update the pseudo-label in step 2: By comparing A′ and B, we get B′:

[0215] (1)|a′ i -b i |>θ 1 , then it means that the basic data contains the error samples in the target error samples, then it is necessary to find the training text pairs that overlap with the error samples in the target error samples from the basic data, that is, the target training pairs, and delete the target training pairs from the original training data;

[0216] (2)|a′ i -b i |<θ 2 , then use b directly i As b′ i ;

[0217] (3) In other cases, b i =a′ i ;

[0218] Step 6: Mix D+B′ and basic data CD′, train M1 to obtain a trained text matching model, denoted as M2; further, through the above model training method, the regression model, that is, the link of retraining the above Student Model, is shortened, and the uncontrollability of the model when repairing error samples is reduced; and error samples can also be repaired quickly, and negative feedback from KA customers can be responded to quickly, providing effect guarantee for KA customers.

[0219] Based on the same principle as the model training method provided in the embodiment of the present application, the embodiment of the present application also provides a model training device, such as Figure 8 As shown, the device 80 may include: a first acquisition module 81, a second acquisition module 82 and a repeated training module 83, wherein:

[0220] A first acquisition module 81 is used to acquire a target error sample set, the target error sample set includes multiple error samples with a first label, each error sample includes an unmatched text pair, and the first label represents a first similarity between the text pairs;

[0221] A second acquisition module 82 is used to acquire a text matching model to be updated and original training data, wherein the text matching model to be updated is obtained by training the initial text matching model with the original training data, and the original training data includes a plurality of training text pairs with second labels, and the second label of a training text pair represents a second similarity between the training text pairs;

[0222] The repeated training module 83 is used to repeatedly perform the training operation on the updated text matching model based on the target error sample set and the original training data until the training end condition is met to obtain a trained target text matching model. The training operation includes the following steps:

[0223] Based on the target error sample set and the original training data, the text matching model obtained by the last training operation is trained at least once to obtain a first model after training;

[0224] Perform label update operations on each error sample. The label update operations include:

[0225] The first model is used to perform similarity identification on each error sample, and the third similarity corresponding to each error sample is obtained. The first label of the corresponding error sample is updated based on the third similarity corresponding to each error sample, and the updated label is used as the first label of the error sample in the next training operation.

[0226] In another possible implementation of the embodiment of the present application, when the repeated training module 83 performs a label update operation on each error sample, it is specifically used to:

[0227] If the first condition is not met, the label update operation is performed on each error sample;

[0228] The training operation also includes:

[0229] If the first condition is met, the label update operation for each erroneous sample is not performed in the current training operation and each training operation after the current training operation;

[0230] Wherein, satisfying the first condition includes satisfying at least one of the following:

[0231] The number of training operations executed reaches the preset number;

[0232] The degree of label modification meets a preset condition, wherein the degree of label modification is determined based on the first difference corresponding to each error sample, and the first difference corresponding to an error sample is the absolute value of the difference between the third similarity of the error sample obtained in the current training operation and the third similarity of the error sample obtained in the previous training operation.

[0233] In another possible implementation of the embodiment of the present application, the label modification degree satisfies the preset condition including at least one of the following:

[0234] The first difference corresponding to each error sample is smaller than the first threshold;

[0235] The sum of the first difference values ​​corresponding to each error sample is less than the second threshold;

[0236] The first proportion is greater than a set value, and the first proportion is the proportion of error samples whose corresponding first difference values ​​are less than a first threshold in the target error sample set.

[0237] In another possible implementation of the embodiment of the present application, for each error sample, the training operation further includes:

[0238] Determine a second difference corresponding to the error sample, where the second difference is an absolute value of a difference between a third similarity corresponding to the error sample and the first similarity;

[0239] If the second difference corresponding to the error sample is greater than a second preset threshold, and there is a target training text pair matching the error sample in the original training data, the target training text pair is deleted from the original training data, and the original training data with the target training text pair deleted is used as the original training data for the next training operation.

[0240] In another possible implementation of the embodiment of the present application, for each error sample, when the repeated training module 83 updates the first label of the corresponding error sample based on the third similarity corresponding to each error sample, it is specifically used to:

[0241] If the second difference corresponding to the error sample is less than the third preset threshold, the third similarity corresponding to the error sample is determined as the updated first label of the error sample;

[0242] If the second difference corresponding to the error sample is not less than the third preset threshold and not greater than the second preset threshold, the third similarity corresponding to the error sample or the first label corresponding to the sample is determined as the updated first label of the error sample.

[0243] In another possible implementation of the embodiment of the present application, when the first acquisition module 81 acquires the target error sample set, it is specifically used to:

[0244] Get multiple error samples;

[0245] Based on the text matching model to be updated, the texts of the error samples are matched respectively to obtain the fourth similarities corresponding to the error samples;

[0246] Based on the fourth similarities corresponding to the respective error samples, the first labels corresponding to the respective error samples are determined.

[0247] In another possible implementation manner of the embodiment of the present application, each error sample includes: a first text and a second text;

[0248] The device 80 further includes: an extraction module and a similarity determination module, wherein:

[0249] An extraction module, configured to extract a first intention representation and a second intention representation for each error sample, wherein the first intention representation is an intention representation of the first text in a target scenario, and the second intention representation is an intention representation of the second text in the target scenario;

[0250] A similarity determination module, used to determine the intent similarity between the first intent representation and the second intent representation as the intent similarity corresponding to the error sample;

[0251] The first acquisition module 81 is specifically used to determine the first labels corresponding to the respective error samples based on the fourth similarities corresponding to the respective error samples:

[0252] Based on the intention similarities corresponding to the error samples, the fourth similarities corresponding to the error samples are updated to obtain the first labels corresponding to the error samples.

[0253] In another possible implementation of the embodiment of the present application, the first acquisition module 81 updates the fourth similarities corresponding to the respective error samples based on the intention similarities corresponding to the respective error samples to obtain the first labels corresponding to the respective error samples, specifically for:

[0254] For each error sample, the fourth similarity corresponding to the error sample is updated in the following manner to obtain the corresponding first similarity:

[0255] If the intention similarity corresponding to the error sample is less than the fourth preset threshold and greater than the fifth preset threshold, the fourth similarity corresponding to the error sample is adjusted to the first target interval to obtain the corresponding first similarity;

[0256] If the intention similarity corresponding to each error sample is not less than a fourth preset threshold, the fourth similarity corresponding to the error sample is adjusted to the second target interval to obtain the corresponding first similarity, the maximum value of the second target interval is greater than the maximum value of the first target interval, the minimum value of the second target area is greater than the minimum value of the first target interval, and the difference between the minimum value of the second target interval and the maximum value of the first target interval is less than a preset value;

[0257] If the intent similarity corresponding to the error sample is less than the fifth preset threshold, the character similarity of the error sample is determined, and based on the character similarity, the corresponding fourth similarity is updated within the third target interval to obtain the corresponding first similarity, and the maximum value in the third target interval is less than the minimum value in the first target interval;

[0258] If the fourth similarity corresponding to the error sample does not belong to the fourth target interval, and the intended similarity of the error sample belongs to the first target interval, the fourth similarity corresponding to the error sample is adjusted to the first target interval to obtain the corresponding first similarity, and the first similarities corresponding to each error sample whose fourth similarity is adjusted to the first target interval are different.

[0259] In another possible implementation of the embodiment of the present application, the device 80 further includes: a third acquisition module, a screening module, a data enhancement processing module, and a target sample set determination module, wherein:

[0260] A third acquisition module is used to acquire an initial sample set, where the initial sample set includes a plurality of initial error samples;

[0261] A screening module is used to screen out prompt commands from the instruction database, where the prompt commands are used to prompt the data enhancement model to perform data enhancement processing on multiple initial error samples;

[0262] A data enhancement processing module is used to perform data enhancement based on the initial sample set and the prompt command and using the trained data enhancement model to obtain enhanced error samples, where the relationship between text pairs in each enhanced error sample is obtained from the relationship between text pairs in each initial sample;

[0263] The target sample set determination module is used to determine the target error sample set based on the initial sample set and the error samples after enhancement processing.

[0264] It should be noted that the first acquisition module 81, the second acquisition module 82 and the third acquisition module may all be the same acquisition module, may not all be the same acquisition module, and may also partially be the same acquisition module;

[0265] An embodiment of the present application provides a device for model training. In the embodiment of the present application, a plurality of error samples with a first label, a text matching model to be updated, and original training data used to train the text matching model to be updated are obtained, and then based on the target error sample set original training data, the training operation is repeatedly performed on the text matching model to be updated until the training end condition is met to obtain a trained text matching model, and in the process of repeatedly performing the training, after the text matching model obtained in the previous operation is trained at least once by the target error sample set and the original training data, a label update operation is also performed on the error sample using the training model to update the label corresponding to the error sample, and then the model is updated based on the target error sample set after the label is updated and the original training data until the training end condition is met to obtain a trained target text matching model, because when the model is cyclically trained and updated, the label in the target error sample set also needs to be updated to improve the accuracy of the training samples, thereby improving the accuracy of text matching of the trained model.

[0266] The device of the embodiments of the present application can execute the method provided by the embodiments of the present application, and the implementation principles are similar. The actions performed by each module in the device of each embodiment of the present application correspond to the steps in the method of each embodiment of the present application. For the detailed functional description of each module of the device, please refer to the description in the corresponding method shown in the previous text, which will not be repeated here.

[0267] Fig. 9 A schematic diagram of the structure of an electronic device applicable to the embodiment of the present application is shown. Fig. 9 As shown, the electronic device can be used to implement the method provided in any embodiment of the present application.

[0268] like Fig. 9 As shown in FIG. 1 , the electronic device 900 may mainly include at least one processor 901 ( Fig. 9 ), memory 902, communication module 903 and input / output interface 904 and other components. Optionally, the components can be connected and communicated through bus 905. It should be noted that Fig. 9 The structure of the electronic device 900 shown in the figure is merely illustrative and does not constitute a limitation on the electronic device to which the method provided in the embodiment of the present application is applicable.

[0269] Among them, the memory 902 can be used to store operating systems and applications, etc. The application may include a computer program that implements the method shown in the embodiment of the present application when called by the processor 901, and may also include a program for implementing other functions or services. The memory 902 may be a ROM (Read Only Memory) or other types of static storage devices that can store static information and instructions, a RAM (Random Access Memory) or other types of dynamic storage devices that can store information and computer programs, or an EEPROM (Electrically Erasable Programmable Read Only Memory), a CD-ROM (Compact Disc Read Only Memory) or other optical disk storage, optical disk storage (including compressed optical disk, laser disc, optical disk, digital versatile disc, Blu-ray disc, etc.), a magnetic disk storage medium or other magnetic storage device, or any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto.

[0270] The processor 901 is connected to the memory 902 via a bus 905, and implements corresponding functions by calling the application program stored in the memory 902. Among them, the processor 901 can be a CPU (Central Processing Unit), a general processor, a DSP (Digital Signal Processor), an ASIC (Application Specific Integrated Circuit), an FPGA (Field Programmable Gate Array) or other programmable logic devices, transistor logic devices, hardware components or any combination thereof, which can implement or execute various exemplary logic blocks, modules and circuits described in conjunction with the disclosure of this application. The processor 901 can also be a combination that implements computing functions, such as a combination of one or more microprocessors, a combination of a DSP and a microprocessor, etc.

[0271] The electronic device 900 can be connected to the network through the communication module 903 (which may include but is not limited to components such as a network interface) to communicate with other devices (such as a user terminal or a server, etc.) through the network to achieve data interaction, such as sending data to other devices or receiving data from other devices. Among them, the communication module 903 may include a wired network interface and / or a wireless network interface, etc., that is, the communication module may include at least one of a wired communication module or a wireless communication module.

[0272] The electronic device 900 can be connected to the required input / output devices, such as a keyboard, a display device, etc., through the input / output interface 904. The electronic device 900 itself can have a display device, and can also be connected to other display devices through the interface 904. Optionally, a storage device, such as a hard disk, can also be connected through the interface 904, so that data in the electronic device 900 can be stored in the storage device, or data in the storage device can be read, and data in the storage device can also be stored in the memory 902. It can be understood that the input / output interface 904 can be a wired interface or a wireless interface. Depending on the actual application scenario, the device connected to the input / output interface 904 can be a component of the electronic device 900, or it can be an external device connected to the electronic device 900 when needed.

[0273] The bus 905 for connecting the components may include a path to transmit information between the above components. The bus 905 may be a PCI (Peripheral Component Interconnect) bus or an EISA (Extended Industry Standard Architecture) bus, etc. According to different functions, the bus 905 may be divided into an address bus, a data bus, a control bus, etc.

[0274] Optionally, for the solution provided in the embodiments of the present application, the memory 902 can be used to store a computer program for executing the solution of the present application, and run by the processor 901. When the processor 901 runs the computer program, the actions of the method or device provided in the embodiments of the present application are implemented.

[0275] An embodiment of the present application provides an electronic device. In the embodiment of the present application, a plurality of error samples with a first label, a text matching model to be updated, and original training data used to train the text matching model to be updated are obtained, and then based on the target error sample set original training data, the training operation is repeatedly performed on the text matching model to be updated until the training end condition is met to obtain a trained text matching model, and in the process of repeatedly performing the training, after the text matching model obtained in the previous operation is trained at least once by the target error sample set and the original training data, a label update operation is also performed on the error sample using the training model to update the label corresponding to the error sample, and then the model is updated based on the target error sample set after the label is updated and the original training data until the training end condition is met to obtain a trained target text matching model, because when the model is cyclically trained and updated, the label in the target error sample set also needs to be updated to improve the accuracy of the training samples, thereby improving the accuracy of the text matching of the trained model.

[0276] Based on the same principle as the method provided in the embodiment of the present application, the embodiment of the present application provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the corresponding content of the aforementioned method embodiment can be implemented. In the embodiment of the present application, a plurality of error samples with a first label, a text matching model to be updated, and original training data used for training the text matching model to be updated are obtained, and then based on the original training data of the target error sample set, the training operation is repeatedly performed on the text matching model to be updated until the training end condition is met to obtain a trained text matching model, and in the process of repeatedly performing the training, after the text matching model obtained in the previous operation is trained at least once by the target error sample set and the original training data, a label update operation is also performed on the error sample using the training model to update the label corresponding to the error sample, and then the model is updated based on the target error sample set after the label is updated and the original training data until the training end condition is met to obtain a trained target text matching model, because when the model is cyclically trained and updated, the label in the target error sample set also needs to be updated to improve the accuracy of the training sample, thereby improving the accuracy of the text matching of the trained model.

[0277] The embodiments of the present application also provide a computer program product, which includes a computer program. When the computer program is executed by a processor, it can implement the corresponding content of the foregoing method embodiments. In the embodiments of the present application, a plurality of error examples with a first label, a text matching model to be updated, and the original training data used to train the text matching model to be updated are obtained. Then, based on the target error example set and the original training data, the training operation is repeatedly performed on the text matching model to be updated until the training end condition is satisfied, so as to obtain a trained text matching model. And during the repeated training process, after the text matching model obtained from the previous operation is trained at least once by the target error example set and the original training data, a label update operation will also be performed on the error examples using the training model to update the labels corresponding to the error examples. Then, based on the target error example set with updated labels and the original training data, the model is updated until the training end condition is satisfied, and a trained target text matching model is obtained. Since when the model is cyclically trained and updated, the labels in the target error example set also need to be updated to improve the accuracy of the training samples, the accuracy of text matching of the trained model can be improved accordingly.

[0278] It should be noted that the terms "first", "second", "third", "fourth", "1", "2", etc. (if any) in the description and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily need to be used to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present application described herein can be implemented in an order other than that shown in the drawings or described in words.

[0279] It should be understood that although the flowcharts in the embodiments of the present application indicate various operation steps by arrows, the execution order of these steps is not limited to the order indicated by the arrows. Unless there is a clear description in this application, in some implementation scenarios of the embodiments of the present application, the implementation steps in each flowchart can be executed in other orders according to requirements. In addition, some or all of the steps in each flowchart may include multiple sub-steps or multiple stages based on the actual implementation scenario. Some or all of these sub-steps or stages can be executed at the same time, and each sub-step or stage of these sub-steps or stages can also be executed at different times respectively. In the scenario where the execution times are different, the execution order of these sub-steps or stages can be flexibly configured according to requirements, and the embodiments of the present application do not limit this.

[0280] In the embodiments of the present application, the term "module" or "unit" refers to a computer program or a part of a computer program with a predetermined function, and works together with other related parts to achieve a predetermined goal, and can be implemented in whole or in part by using software, hardware (such as processing circuits or memories) or a combination thereof. Similarly, a processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be part of an overall module or unit that includes the function of the module or unit.

[0281] The above is only an optional implementation method for some implementation scenarios of the present application. It should be pointed out that for ordinary technicians in this technical field, without departing from the technical concept of the solution of the present application, other similar implementation methods based on the technical ideas of the present application are also within the protection scope of the embodiments of the present application.

Claims

1. A model training method, It is characterized in that The method comprises: Acquire a target error sample set, wherein the target error sample set includes a plurality of error samples with a first label, each of the error samples includes an unmatched text pair, and the first label represents a first similarity between the text pairs; Obtaining a text matching model to be updated and original training data, wherein the text matching model to be updated is obtained by training an initial text matching model using the original training data, and the original training data includes a plurality of training text pairs with second labels, and the second label of a training text pair represents a second similarity between the training text pairs; Based on the target error sample set and the original training data, the training operation is repeatedly performed on the text matching model to be updated until the training end condition is met to obtain a trained target text matching model, wherein the training operation includes the following steps: Based on the target error sample set and the original training data, the text matching model obtained in the last training operation is trained at least once to obtain a trained first model; Performing a label update operation on each of the error samples, the label update operation comprising: The first model is used to perform similarity identification on each of the error samples to obtain a third similarity corresponding to each of the error samples, and the first label of the corresponding error sample is updated based on the third similarity corresponding to each of the error samples, and the updated label is used as the first label of the error sample in the next training operation.

2. The method according to claim 1, It is characterized in that The step of performing a label updating operation on each of the error samples includes: If the first condition is not met, performing a label update operation on each of the error samples; The training operation further includes: If the first condition is met, then in the current training operation and each training operation after the current training operation, the label update operation for each of the erroneous samples is not performed; Wherein, satisfying the first condition includes satisfying at least one of the following: The number of training operations executed reaches the preset number; The degree of label modification meets a preset condition, wherein the degree of label modification is determined based on the first difference corresponding to each of the error samples, and the first difference corresponding to an error sample is the absolute value of the difference between the third similarity of the error sample obtained in the current training operation and the third similarity of the error sample obtained in the previous training operation.

3. The method according to claim 2, It is characterized in that The label modification degree meets the preset condition including at least one of the following: The first difference values ​​corresponding to the error samples are all smaller than the first threshold value; The sum of the first difference values ​​corresponding to the error samples is less than a second threshold; The first proportion is greater than a set value, and the first proportion is the proportion of error samples whose corresponding first difference values ​​are less than a first threshold in the target error sample set.

4. The method according to claim 1, It is characterized in that For each of the error samples, the training operation further includes: Determine a second difference corresponding to the error sample, where the second difference is an absolute value of a difference between a third similarity corresponding to the error sample and the first similarity; If the second difference corresponding to the error sample is greater than a second preset threshold, and there is a target training text pair matching the error sample in the original training data, the target training text pair is deleted from the original training data, and the original training data with the target training text pair deleted is used as the original training data for the next training operation.

5. The method according to claim 4, It is characterized in that For each of the error samples, updating the first label of the corresponding error sample based on the third similarity corresponding to each of the error samples includes: If the second difference corresponding to the error sample is less than a third preset threshold, determining the third similarity corresponding to the error sample as the updated first label of the error sample; If the second difference corresponding to the error sample is not less than the third preset threshold and not greater than the second preset threshold, the third similarity corresponding to the error sample or the first label corresponding to the sample is determined as the updated first label of the error sample.

6. The method according to claim 1, It is characterized in that The step of obtaining a target error sample set includes: Obtaining the multiple error samples; Based on the text matching model to be updated, the texts corresponding to the respective error samples are matched respectively to obtain the fourth similarities corresponding to the respective error samples; Based on the fourth similarities corresponding to the respective error samples, the first labels corresponding to the respective error samples are determined.

7. The method according to claim 6, It is characterized in that Each of the error samples includes: a first text and a second text; Before determining the first labels corresponding to the respective error samples based on the fourth similarities corresponding to the respective error samples, the method further includes: For each error sample, extract a first intention representation and a second intention representation, wherein the first intention representation is the intention representation of the first text in a target scenario, and the second intention representation is the intention representation of the second text in the target scenario; Determine the intent similarity between the first intent representation and the second intent representation as the intent similarity corresponding to the error sample; The determining of the first labels corresponding to the respective error samples based on the fourth similarities corresponding to the respective error samples includes: Based on the intention similarities corresponding to the error samples, the fourth similarities corresponding to the error samples are updated to obtain the first labels corresponding to the error samples.

8. The method according to claim 7, It is characterized in that The updating of the fourth similarities corresponding to the respective error samples based on the intention similarities corresponding to the respective error samples to obtain the first labels corresponding to the respective error samples includes: For each of the error samples, the fourth similarity corresponding to the error sample is updated in the following manner to obtain the corresponding first similarity: If the intention similarity corresponding to the error sample is less than the fourth preset threshold and greater than the fifth preset threshold, adjusting the fourth similarity corresponding to the error sample to the first target interval to obtain the corresponding first similarity; If the intention similarity corresponding to each of the error samples is not less than the fourth preset threshold, the fourth similarity corresponding to the error sample is adjusted to the second target interval to obtain the corresponding first similarity, the maximum value of the second target interval is greater than the maximum value of the first target interval, the minimum value of the second target area is greater than the minimum value of the first target interval, and the difference between the minimum value of the second target interval and the maximum value of the first target interval is less than a preset value; If the intention similarity corresponding to the error sample is less than the fifth preset threshold, the character similarity of the error sample is determined, and based on the character similarity, the corresponding fourth similarity is updated within the third target interval to obtain the corresponding first similarity, and the maximum value in the third target interval is less than the minimum value in the first target interval; If the fourth similarity corresponding to the error sample does not belong to the fourth target interval, and the intended similarity of the error sample belongs to the first target interval, the fourth similarity corresponding to the error sample is adjusted to the first target interval to obtain the corresponding first similarity, and the first similarities corresponding to each error sample whose fourth similarity is adjusted to the first target interval are different.

9. The method according to claim 1, It is characterized in that The method further comprises: Acquire an initial sample set, wherein the initial sample set includes a plurality of initial error samples; Filtering a prompt command from the instruction database, where the prompt command is used to prompt the data enhancement model to perform data enhancement processing on the multiple initial error samples; Based on the initial sample set and the prompt command, data enhancement is performed using a trained data enhancement model to obtain enhanced error samples, wherein the relationship between text pairs in each enhanced error sample is obtained from the relationship between text pairs in each initial sample; The target error sample set is determined based on the initial sample set and the error samples after the enhancement process.

10. A device for model training, It is characterized in that The device comprises: A first acquisition module is used to acquire a target error sample set, wherein the target error sample set includes a plurality of error samples with a first label, each of the error samples includes an unmatched text pair, and the first label represents a first similarity between the text pairs; A second acquisition module is used to acquire a text matching model to be updated and original training data, wherein the text matching model to be updated is obtained by training an initial text matching model with the original training data, and the original training data includes a plurality of training text pairs with second labels, and the second label of a training text pair represents a second similarity between the training text pairs; A repeated training module is used to repeatedly perform a training operation on the text matching model to be updated based on the target error sample set and the original training data until a training end condition is met to obtain a trained target text matching model, wherein the training operation includes the following steps: Based on the target error sample set and the original training data, the text matching model obtained in the last training operation is trained at least once to obtain a trained first model; Performing a label update operation on each of the error samples, the label update operation comprising: The first model is used to perform similarity identification on each of the error samples to obtain a third similarity corresponding to each of the error samples, and the first label of the corresponding error sample is updated based on the third similarity corresponding to each of the error samples, and the updated label is used as the first label of the error sample in the next training operation.

11. An electronic device, It is characterized in that The electronic device includes a memory and a processor, the memory stores a computer program, and the processor executes the model training method described in any one of claims 1 to 9 when running the computer program.

12. A computer-readable storage medium, It is characterized in that The storage medium stores a computer program, which, when executed by a processor, implements the model training method described in any one of claims 1 to 9.