A training method and device for a text retrieval model
Through global analysis of sample types and optimization of loss function weights during text retrieval model training, the problem of weak overfitting and generalization capabilities of the model is solved, and the training efficiency and accuracy of the text retrieval model are improved.
Patent Information
- Application Number
- CN202211699919.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-12-28
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2042-12-28
AI Technical Summary
In the prior art, during the training process of text retrieval model based on neural networks, due to the large number of easily divided negative samples obtained by negative sampling, the model is overfitting and generalizing, resulting in inaccurate search results and poor user experience.
By conducting global analysis of training text samples, different sample types are identified, and different loss function weight values are set according to the sample type, the text retrieval model is trained to improve the model training efficiency and effect.
It improves the training efficiency and performance of the text retrieval model, improves the accuracy of text retrieval results, and improves the user experience.
Smart Images

Figure CN115841144B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of artificial intelligence technology, and particularly relates to a method and device for training a text retrieval model. Background Art
[0002] In the search service, during the training process of a commonly used neural network-based text retrieval model, a large number of easy-to-separate negative samples that have relatively little effect on the neural network training are obtained through negative sampling. These easily obtainable and relatively easy-to-distinguish negative samples have limited training gain for the model, which may lead to problems such as model overfitting and weak generalization ability due to insufficient information in the representation features of the negative samples during the training process of the text retrieval model. In this way, in the scenario of using the text retrieval model for text retrieval, the retrieved text retrieval results are not the ones that the user really wants, resulting in poor user experience. Summary of the Invention
[0003] In view of this, embodiments of the present disclosure provide a method, device, computer device, and computer-readable storage medium for training a text retrieval model to solve the problem in the prior art that a large number of easy-to-separate negative samples obtained through negative sampling have relatively little effect on the neural network training, resulting in limited training gain for the model, which may lead to problems such as model overfitting and weak generalization ability due to insufficient information in the representation features of the negative samples during the training process of the text retrieval model. In this way, in the scenario of using the text retrieval model for text retrieval, the retrieved text retrieval results are not the ones that the user really wants, resulting in poor user experience.
[0004] In a first aspect of the embodiments of the present disclosure, a method for training a text retrieval model is provided. The method includes:
[0005] Obtain a training text sample set; wherein, the training text sample set includes a number of training text samples, and each group of training text samples includes a sample query text and a true article title corresponding to the sample query text; the true article title is an article title in the preset text database;
[0006] For each round of training of each group of training text samples, input the sample query text in the training text sample into a preset text retrieval model to obtain a text statement feature corresponding to the sample query text; query in the preset text database according to the text statement feature corresponding to the sample query text to obtain a predicted article title corresponding to the sample query text, wherein the predicted article title is an article title in the preset text database; determine the loss function value of this round of the training text sample according to the predicted article title and the true article title;
[0007] Determine the sample type of each group of training text samples according to the loss function values of each group of training text samples in each round of N rounds of training.
[0008] Use the training text sample set, the sample type of each training text sample in the training text sample set, and the loss function weight values corresponding to the preset sample types to train the text retrieval model to obtain a trained text retrieval model.
[0009] In a second aspect of the embodiments of the present disclosure, there is provided a training device for a text retrieval model, the device including:
[0010] A set acquisition unit for acquiring a training text sample set; wherein, the training text sample set includes a number of training text samples, each group of training text samples includes a sample query text and the true article title corresponding to the sample query text; the true article title is an article title in the preset text database.
[0011] A numerical value determination unit for, for each round of training of each group of training text samples, inputting the sample query text in the training text sample into a preset text retrieval model to obtain the text statement feature corresponding to the sample query text; querying in the preset text database according to the text statement feature corresponding to the sample query text to obtain the predicted article title corresponding to the sample query text, wherein the predicted article title is an article title in the preset text database; determining the loss function value of this round of the training text sample according to the predicted article title and the true article title.
[0012] A type determination unit for determining the sample type of each group of training text samples according to the loss function values of each group of training text samples in each round of N rounds of training.
[0013] A model training unit for using the training text sample set, the sample type of each training text sample in the training text sample set, and the loss function weight values corresponding to the preset sample types to train the text retrieval model to obtain a trained text retrieval model.
[0014] In a third aspect of the embodiments of the present disclosure, there is provided a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, and when the processor executes the computer program, the steps of the above method are implemented.
[0015] In a fourth aspect of the embodiments of the present disclosure, there is provided a computer-readable storage medium storing a computer program, and when the computer program is executed by a processor, the steps of the above method are implemented.
[0016] The beneficial effects of the embodiments of the present disclosure compared with the prior art are as follows: In the embodiments of the present disclosure, a training text sample set can be obtained first; wherein, the training text sample set includes a number of training text samples, and each group of training text samples includes a sample query text and the true article title corresponding to the sample query text; the true article title is an article title in the preset text database. Then, for each round of training of each group of training text samples, the sample query text in the training text sample is input into a preset text retrieval model to obtain the text statement features corresponding to the sample query text; according to the text statement features corresponding to the sample query text, a query is performed in the preset text database to obtain the predicted article title corresponding to the sample query text, wherein the predicted article title is an article title in the preset text database; according to the predicted article title and the true article title, the loss function value of this round of the training text sample is determined. Next, according to the loss function values of each group of training text samples in each round of N rounds of training in the training text sample set, the sample type of each group of training text samples can be determined. Finally, the text retrieval model can be trained by using the training text sample set, the sample type of each training text sample in the training text sample set, and the loss function weight value corresponding to the preset sample type to obtain a trained text retrieval model.It can be seen that in this embodiment, first, a sample query text in each group of training text samples and the true article title corresponding to the sample query text are used to determine the loss function value of each group of training text samples in each round of training. Then, according to the loss function value of each group of training text samples in each round of training, each group of training text samples is classified to determine the sample type of the group of training text samples. In this way, a full global analysis and mining of the training text samples is carried out. By classifying the loss function values of the training text samples in the first-stage training process (i.e., N rounds of training), different sample types (such as simple samples, ordinary samples, and difficult samples) are identified. Since the influence degrees of training text samples of different sample types on the training effect of the text retrieval model (such as affecting the generalization ability of the text retrieval model) are different, different loss function weight values can be set according to different sample types. So, in the process of training the text retrieval model by using the training text sample set, the sample type of each group of training text samples in the training text sample set, and the loss function weight value corresponding to the preset sample type, the loss function weight value of the training text samples of the sample type that does not affect the model training effect can be reduced, and the loss function weight value of the training text samples of the sample type that has a greater impact on the model training effect can be increased, so as to give full play to the potential of the training text samples of the sample type that has a greater impact on the model training effect and reduce the impact of the training text samples of the sample type that has no impact or a poor impact on the model training effect on the training of the text retrieval model. It can be seen that this embodiment can effectively improve the distribution of the training weights (i.e., loss function weight values) of the training text samples of different sample types, so that the training process of the text retrieval model is more sufficient, the training efficiency and effect of the text retrieval model can be improved, and then the performance of the text retrieval model can be improved, and further the text retrieval effect of the text retrieval model in the actual business scenario can be improved (such as improving the accuracy of the text retrieval results of the text retrieval model). BRIEF DESCRIPTION OF THE DRAWINGS
[0017] In order to more clearly illustrate the technical solutions in the embodiments of the present disclosure, the accompanying drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the accompanying drawings in the following description are only some embodiments of the present disclosure. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0018] Figure 1 is a schematic diagram of the application scenario of the embodiment of the present disclosure;
[0019] Figure 2 is a flowchart of the training method of the text retrieval model provided by the embodiment of the present disclosure;
[0020] Figure 3It is a block diagram of a training device for a text retrieval model provided by an embodiment of the present disclosure;
[0021] Figure 4 It is a schematic diagram of a computer device provided by an embodiment of the present disclosure. Specific embodiments
[0022] In the following description, specific details such as specific system architectures and technologies are presented for the purpose of illustration rather than limitation, so as to thoroughly understand the embodiments of the present disclosure. However, those skilled in the art should clearly understand that the present disclosure can also be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid unnecessary details from interfering with the description of the present disclosure.
[0023] A training method and device for a text retrieval model according to an embodiment of the present disclosure will be described in detail below with reference to the accompanying drawings.
[0024] In the prior art, in the traditional search service, during the training process of a commonly used neural network-based text retrieval model, a large number of easy-to-separate negative samples that have relatively little effect on the neural network training are obtained through negative sampling. These easily obtainable and relatively easy-to-distinguish negative samples have limited training gain for the model, which may lead to problems such as model overfitting and weak generalization ability due to insufficient information in the characterization features of the negative samples during the training process of the text retrieval model. In this way, in the scenario of text retrieval using the text retrieval model, the retrieved text retrieval results are not the ones that the user really wants, resulting in a poor user experience.
[0025] To solve the above problems, the present invention provides a method for training a text retrieval model. In this method, since in this embodiment, a sample query text in each group of training text samples and the true article title corresponding to the sample query text can be first used to determine the loss function value of each group of training text samples in each round of training. Then, according to the loss function value of each group of training text samples in each round of training, each group of training text samples is classified to determine the sample type of the group of training text samples. In this way, a full global analysis and mining of the training text samples are carried out. By classifying the loss function values of the training text samples in the first-stage training process (i.e., N rounds of training), different sample types (such as simple samples, ordinary samples, and difficult samples) are identified. Since the influence degrees of training text samples of different sample types on the training effect of the text retrieval model (such as affecting the generalization ability of the text retrieval model) are different, different loss function weight values can be set according to different sample types. So, in the process of training the text retrieval model by using the training text sample set, the sample type of each group of training text samples in the training text sample set, and the loss function weight value corresponding to the preset sample type, the loss function weight value of the training text samples of the sample type that does not affect the model training effect can be reduced, and the loss function weight value of the training text samples of the sample type that has a greater impact on the model training effect can be increased, so as to give full play to the potential of the training text samples of the sample type that has a greater impact on the model training effect and reduce the influence of the training text samples of the sample type that has no impact or a poor impact on the model training effect on the training of the text retrieval model. It can be seen that this embodiment can effectively improve the distribution of the training weights (i.e., loss function weight values) of the training text samples of different sample types, so that the training process of the text retrieval model is more sufficient, the training efficiency and effect of the text retrieval model can be improved, and then the performance of the text retrieval model can be improved, and further the text retrieval effect of the text retrieval model in the actual business scenario can be improved (such as improving the accuracy of the text retrieval results of the text retrieval model).
[0026] For example, the embodiment of the present invention can be applied to an application scenario as Figure 1 shown. In this scenario, it may include a terminal device 1 and a server 2.
[0027] The terminal device 1 can be either hardware or software. When the terminal device 1 is hardware, it can be various electronic devices with a display screen and supporting communication with the server 2, including but not limited to smartphones, tablet computers, laptop portable computers, desktop computers, etc.; when the terminal device 1 is software, it can be installed in the above-mentioned electronic devices. The terminal device 1 can be implemented as multiple software or software modules, or can also be implemented as a single software or software module, and the embodiments of the present disclosure do not limit this. Further, various applications can be installed on the terminal device 1, such as data processing applications, instant messaging tools, social platform software, search applications, shopping applications, etc.
[0028] The server 2 can be a server providing various services. For example, it can be a background server that receives requests sent by a terminal device establishing a communication connection with it. This background server can receive and analyze requests sent by the terminal device and generate a processing result. The server 2 can be a single server, or can also be a server cluster composed of several servers, or can also be a cloud computing service center, and the embodiments of the present disclosure do not limit this.
[0029] It should be noted that the server 2 can be either hardware or software. When the server 2 is hardware, it can be various electronic devices providing various services for the terminal device 1. When the server 2 is software, it can be multiple software or software modules providing various services for the terminal device 1, or can also be a single software or software module providing various services for the terminal device 1, and the embodiments of the present disclosure do not limit this.
[0030] The terminal device 1 and the server 2 can be communicatively connected through a network. The network can be a wired network connected by coaxial cables, twisted pairs, and optical fibers, or can also be a wireless network that can interconnect various communication devices without cabling, such as Bluetooth, Near Field Communication (NFC), Infrared, etc., and the embodiments of the present disclosure do not limit this.
[0031] Specifically, the user can input to obtain a training text sample set through the terminal device 1; the terminal device 1 sends the obtained training text sample set to the server 2. The server 2 stores a text retrieval model to be trained; for each round of training of each group of training text samples, the server 2 can first input the sample query text in the training text sample into the preset text retrieval model to obtain the text statement feature corresponding to the sample query text; according to the text statement feature corresponding to the sample query text, query in the preset text database to obtain the predicted article title corresponding to the sample query text, where the predicted article title is an article title in the preset text database; determine the loss function value of this round of the training text sample according to the predicted article title and the true article title; then, the server 2 can determine the sample type of each group of training text samples according to the loss function values of each group of training text samples in each round of the N rounds of training in the training text sample set; then, the server 2 can use the training text sample set, the sample type of each training text sample in the training text sample set, and the loss function weight value corresponding to the preset sample type to train the text retrieval model to obtain a trained text retrieval model.In this way, since the present application can first use a sample query text in each group of training text samples and the true article title corresponding to the sample query text to determine the loss function value of each group of training text samples in each round of training, and then classify each group of training text samples according to the loss function value of each group of training text samples in each round of training to determine the sample type of the group of training text samples; in this way, a full global analysis and mining of the training text samples is carried out. By classifying the loss function values of the training text samples in the first-stage training process (i.e., N rounds of training), different sample types (such as simple samples, ordinary samples, and difficult samples) are identified. Since the influence degrees of training text samples of different sample types on the training effect of the text retrieval model (such as affecting the generalization ability of the text retrieval model) are different, therefore, different loss function weight values can be set according to different sample types. So that in the process of training the text retrieval model by using the training text sample set, the sample type of each group of training text samples in the training text sample set, and the loss function weight value corresponding to the preset sample type, the loss function weight value of the training text samples of the sample type that does not affect the model training effect can be reduced, and the loss function weight value of the training text samples of the sample type that has a greater impact on the model training effect can be increased, so as to give full play to the potential of the training text samples of the sample type that has a greater impact on the model training effect and reduce the impact of the training text samples of the sample type that has no impact or a poor impact on the model training effect on the training of the text retrieval model; it can be seen that this embodiment can effectively improve the distribution of the training weights (i.e., loss function weight values) of the training text samples of different sample types, so that the training process of the text retrieval model is more sufficient, the training efficiency and effect of the text retrieval model can be improved, and further the performance of the text retrieval model can be improved, and the text retrieval effect of the text retrieval model in the actual business scenario can be further improved (such as improving the accuracy of the text retrieval results of the text retrieval model).
[0032] It should be noted that the specific types, quantities, and combinations of the terminal device 1, the server 2, and the network can be adjusted according to the actual needs of the application scenario, and the embodiments of the present disclosure do not limit this.
[0033] It should be noted that the above application scenarios are only shown for the convenience of understanding the present disclosure, and the embodiments of the present disclosure are not limited in this regard. On the contrary, the embodiments of the present disclosure can be applied to any applicable scenario.
[0034] Figure 2 is a flowchart of a training method for a text retrieval model provided by an embodiment of the present disclosure. Figure 2 A training method for a text retrieval model can be performed by Figure 1 the terminal device or the server. As Figure 2As shown, the training method of the text retrieval model includes:
[0035] S201: Obtain a training text sample set.
[0036] Among them, the training text sample set includes several training text samples. Each group of training text samples includes a sample query text and the true article title corresponding to the sample query text.
[0037] In this embodiment, the training text sample set may include several training text samples. Among them, each group of training recommendation samples includes a sample query text and the true article title corresponding to the sample query text. Among them, the sample query text can be understood as the query statement text that needs to be queried; in one implementation, the sample query text can be the historical query text input by the user. The true article title corresponding to the sample query text can be understood as the true article title corresponding to the sample query text.
[0038] It should be noted that the true article title is an article title in the preset text database. It can be understood that the true article title corresponding to the sample query text is the article title in the preset text database that is most similar to the text content of the sample query text; among them, multiple article titles and the text statement features corresponding to each article title are pre-stored in the preset text database. Among them, the text statement features corresponding to the article title can be understood as the features that can reflect the text content meaning of the article title. For example, if the sample query text is "World Cup start time", the true article title corresponding to the sample query text can be "World Cup opening match announced".
[0039] S202: For each round of training of each group of training text samples, input the sample query text in the training text samples into a preset text retrieval model to obtain the text statement features corresponding to the sample query text; query in the preset text database according to the text statement features corresponding to the sample query text to obtain the predicted article title corresponding to the sample query text; determine the loss function value of this round of the training text samples according to the predicted article title and the true article title.
[0040] After obtaining the training samples, each group of training text samples in the training text sample set can be used to train a preset text retrieval model for N rounds respectively, and the loss function values of each group of training text samples in the training text sample set in each round of training can be obtained. It can be understood that since a group of training text samples is used to train the preset text retrieval model for N rounds, N loss function values corresponding to this group of training text samples can be obtained. If there are Z groups of training text samples in the training text sample set, then Z*N loss function values can be obtained. It should be noted that N can be a positive integer greater than 1, and N rounds can be the number of times preset according to actual needs, or the number of training rounds that can cause the text retrieval model to show a slight overfitting situation (that is, the loss function value of the training text sample set decreases, but the loss function value of the validation set begins to increase) can be used as N rounds. For example, assuming that a training text sample set is used to train the text retrieval model, when the training reaches the 100th round, the text retrieval model shows a slight overfitting situation, then N can be set to 100. It should be noted that in one implementation, the text retrieval model can be a recurrent neural network or a self-attention network, such as neural network models like transformer, BERT, BERTa, RNN, LSTM, GRU, ElMo, CNN, fastText, Albert, etc.
[0041] Specifically, for each round of training of each group of training text samples, the sample query text in the training text sample is input into the preset text retrieval model to obtain the text statement feature corresponding to the sample query text. That is to say, the sample query text in the training text sample is input into the preset text retrieval model so that the text retrieval model can extract the latent representation vector of the sample query text in order to obtain the text statement feature corresponding to the sample query text. It can be understood that the text statement feature corresponding to the sample query text can be understood as a feature vector that can reflect the text content of the sample query text.
[0042] Then, according to the text statement feature corresponding to the sample query text, a query can be made in the preset text database to obtain the predicted article title corresponding to the sample query text; that is to say, using the text statement feature corresponding to the sample query text, query the text database for an article title that is closest to the text statement feature, and this article title can be called the predicted article title; it can be understood that the predicted article title is an article title in the preset text database.
[0043] Specifically, in one implementation, for each article title in the preset text database, the matching value between the article title and the sample query text can be determined first according to the text statement features corresponding to the article title and the text statement features corresponding to the sample query text. For example, the inner product of the text statement features corresponding to the article title and the text statement features corresponding to the sample query text can be calculated, and this inner product can be used as the matching value between the article title and the sample query text. It can be understood that the higher the matching value between the article title and the sample query text, the more similar the text content meanings of the article title and the sample query text are. On the contrary, the lower the matching value between the article title and the sample query text, the less similar the text content meanings of the article title and the sample query text are. And the article title with the highest matching value with the sample query text in the preset text database can be used as the predicted article title corresponding to the sample query text.
[0044] Next, according to the predicted article title and the true article title, the loss function value of this round of the training text sample is determined. For example, in one implementation, the loss function of the text retrieval model is binary cross-entropy, multi-class cross-entropy loss function, logistic regression loss function, triplet loss function. The loss function of the text retrieval model can be used to calculate the loss function value of the predicted article title and the true article title, and this loss function value can be used as the loss function value of this round of training of this group of training image samples.
[0045] S203: Determine the sample type of each group of training text samples in the training text sample set according to the loss function values of each group of training text samples in each round of N rounds of training.
[0046] Since the training text samples of different sample types have different degrees of influence on the training effect of the text retrieval model (such as affecting the generalization ability of the text retrieval model), it is necessary to classify the training text samples in the training text sample set so as to perform different training on the text retrieval model for the training text samples of different sample types. In this embodiment, by using the loss function value as an index, the classification of the training text samples can be completed adaptively and efficiently, avoiding cumbersome data analysis and business knowledge intervention, being efficient and having a relatively wide adaptability, and being able to be well adapted in other scenarios.
[0047] Specifically, in this embodiment, for each group of training text samples in the training text sample set, the sample type of this group of training text samples can be determined according to the loss function values of this group of training text samples in each round of training. That is to say, in this embodiment, after a global and sufficient analysis and mining of this group of training text samples, the loss function values of this group of training text samples in each round of training can be obtained, and, according to the loss function values of this group of training text samples in each round of training, classification is performed to determine the sample type of this group of training text samples.
[0048] It should be noted that the sample types of training text samples can be divided into a first sample type (such as a normal sample), a second sample type (such as a simple sample), and a third sample type (such as a difficult sample).
[0049] S204: Use the training text sample set, the sample type of each training text sample in the training text sample set, and the loss function weight values corresponding to the preset sample types to train the text retrieval model to obtain a trained text retrieval model.
[0050] In this embodiment, different loss function weight values can be set in advance for training text samples of different sample types during the training process, so as to give full play to the potential of the training text samples, enable the text retrieval model to be trained more fully during the training process, improve the performance of the model, and further improve the text retrieval effect of the text retrieval model in the actual business scenario; and, a large number of useless simple samples can be effectively identified, so that the training weights of some simple samples that do not affect the model training effect can be reduced, that is, during the process of training the text retrieval model, the loss function weight values of the training text samples of the sample types that do not affect the model training effect can be reduced, and the loss function weight values of the training text samples of the sample types that have a greater impact on the model training effect can be increased, so as to give full play to the potential of the training text samples of the sample types that have a greater impact on the model training effect, reduce the impact of the training text samples of the sample types that have no impact or a poor impact on the model training effect on the training of the text retrieval model, so as to effectively improve the training weight distribution of simple samples and difficult samples, and further improve the efficiency and effect of model training.
[0051] In this embodiment, since the third sample type (such as difficult samples) has a greater impact on the model training effect, the weight value of the loss function of the training text samples of the third sample type can be increased. That is to say, the weight value of the loss function corresponding to the training text samples of the third sample type is greater than the weight value of the loss function corresponding to the training text samples of the first sample type. For example, the weight value of the loss function of the training text samples of the third sample type is between 2 and 5 times the weight value of the loss function of the training text samples of the first sample type. Since the second sample type (such as easy samples) has a smaller impact on the model training effect, for example, it has no gain in the training of the text retrieval model and will cause overfitting of the model and reduce the generalization ability online. Therefore, the weight value of the loss function of the training text samples of the second sample type can be reduced. That is to say, the weight value of the loss function corresponding to the training text samples of the second sample type is less than the weight value of the loss function corresponding to the training text samples of the first sample type. For example, the weight value of the loss function of the training text samples of the first sample type is between 2 and 5 times the weight value of the loss function of the training text samples of the second sample type.
[0052] In this implementation, after determining the sample type of each group of training text samples in the training text sample set, the text retrieval model can be trained using the training text sample set, the sample type of each group of training text samples in the training text sample set, and the weight value of the loss function corresponding to the preset sample type to obtain a trained text retrieval model. That is to say, the text retrieval model is trained using the training text samples in the training text sample set. Among them, in the process of adjusting the model parameters of the text retrieval model according to the loss function value, when calculating the loss function value, it is necessary to calculate according to the sample type of each group of training text samples in the training text sample set, the weight value of the loss function corresponding to the preset sample type, and the preset loss function; train until the loss function value of the text retrieval model meets the preset conditions, or the number of training times reaches the preset number of times, and a trained text retrieval model can be obtained.
[0053] Next, for example, illustrate how to calculate the loss function value according to the sample type of each group of training text samples in the training text sample set, the preset loss function weight value corresponding to the sample type, and the preset loss function in S204. For example: the training text sample set includes 4 training text samples: the first training text sample (including the sample query text s1), the second training text sample (including the sample query text s2), the third training text sample (including the sample query text s3), and the fourth training text sample (including the sample query text s4); assume that the sample types of the first training text sample and the second training text sample are difficult samples (i.e., the third sample type), and the sample types of the third training text sample and the fourth training text sample are ordinary samples (i.e., the first sample type). The true article titles corresponding to the sample query text s1, the sample query text s2, the sample query text s3, and the sample query text s4 are l1, l2, l3, and l4 respectively. The loss function weight of the third sample type is w times that of the first sample type. Among them, the neural network (i.e., the text retrieval model) is f(), and the loss function is loss(). Then, during this gradient calculation, the loss function value Loss of the text retrieval model after reassigning the weights 平均 is calculated as follows: Loss 平均 = w * (loss(l1, f(s1)) + loss(l2, f(s2)) + loss(l3, f(s3)) + loss(l4, f(s4))). Then, the model training can be carried out according to the normal process.
[0054] It can be seen that in the embodiments of the present disclosure, a training text sample set can be obtained first; wherein, the training text sample set includes a number of training text samples, and each group of training text samples includes a sample query text and the true article title corresponding to the sample query text; the true article title is an article title in the preset text database. Then, for each round of training of each group of training text samples, the sample query text in the training text samples is input into a preset text retrieval model to obtain the text statement features corresponding to the sample query text; according to the text statement features corresponding to the sample query text, a query is performed in the preset text database to obtain the predicted article title corresponding to the sample query text, wherein the predicted article title is an article title in the preset text database; according to the predicted article title and the true article title, the loss function value of this round of the training text samples is determined. Next, according to the loss function values of each group of training text samples in each round of the N rounds of training in the training text sample set, the sample type of each group of training text samples can be determined. Finally, the training text sample set, the sample type of each training text sample in the training text sample set, and the loss function weight value corresponding to the preset sample type can be used to train the text retrieval model to obtain a trained text retrieval model.It can be seen that in this embodiment, first, a sample query text in each group of training text samples and the true article title corresponding to the sample query text are used to determine the loss function value of each group of training text samples in each round of training. Then, according to the loss function value of each group of training text samples in each round of training, each group of training text samples is classified to determine the sample type of the group of training text samples. In this way, a full global analysis and mining of the training text samples are carried out. By classifying the loss function values of the training text samples in the first-stage training process (i.e., N rounds of training), different sample types (such as simple samples, ordinary samples, and difficult samples) are identified. Since the influence degrees of the training text samples of different sample types on the training effect of the text retrieval model (such as affecting the generalization ability of the text retrieval model) are different, different loss function weight values can be set according to different sample types. So, in the process of training the text retrieval model by using the training text sample set, the sample type, and the loss function weight value corresponding to the preset sample type of each group of training text samples in the training text sample set, the loss function weight value of the training text samples of the sample type that does not affect the model training effect can be reduced, and the loss function weight value of the training text samples of the sample type that has a greater impact on the model training effect can be increased to give full play to the potential of the training text samples of the sample type that has a greater impact on the model training effect and reduce the influence of the training text samples of the sample type that has no impact or a poor impact on the model training effect on the training of the text retrieval model. It can be seen that this embodiment can effectively improve the distribution of the training weights (i.e., loss function weight values) of the training text samples of different sample types, making the training process of the text retrieval model more sufficient, improving the training efficiency and effect of the text retrieval model, further enhancing the performance of the text retrieval model, and further improving the text retrieval effect of the text retrieval model in the actual business scenario (such as improving the accuracy of the text retrieval results of the text retrieval model).
[0055] In some embodiments, the step of S203 "determine the sample type of each group of training text samples according to the loss function value of each group of training text samples in each round of the N rounds of training in the training text sample set" may include the following steps:
[0056] S203a: For each group of training text samples in the training text sample set, determine the overall training loss reduction degree of the training text sample according to the loss function value of the training text sample in each round of training.
[0057] In an implementation manner of this embodiment, the text retrieval model can be trained for M rounds by using the training text sample set, that is, in each round of training, all the training text samples in the training text sample set are used to train the text retrieval model once.
[0058] First, based on the loss function values of the training text samples in each round of training, the average value of the loss function values in the first M rounds of training and the loss function values in the subsequent X rounds of training can be determined, where M and X are both positive integers, and both M and X are less than N. It can be understood that, first, based on the loss function values of the training text samples in each of the first M rounds of training, the average value of the loss function values of the training text samples in the first M rounds of training is determined, and, based on the loss function values of the training text samples in each of the subsequent X rounds of training, the average value of the loss function values of the training text samples in the subsequent X rounds of training is determined. For example, based on all the loss function values of the training text samples in the first 10% of the rounds, the average value of the loss function values of the training text samples in the first 10% of the rounds can be calculated, and, based on all the loss function values of the training text samples in the last 10% of the rounds, the average value of the loss function values of the training text samples in the last 10% of the rounds can be calculated.
[0059] Then, based on the average value of the loss function values in the first M rounds of training and the loss function values in the subsequent X rounds of training, the overall training loss reduction degree of the training text samples can be determined. That is, based on the average value of the loss function values of the training text samples in the first M rounds of training and the average value of the loss function values of the training text samples in the subsequent X rounds of training, the overall training loss reduction degree of the training text samples can be determined. In one implementation, the following formula can be used to calculate the overall training loss reduction degree of the training text samples: c = (a - b) / a; where c is the overall training loss reduction degree of the training text samples, a is the average value of the loss function values of the training text samples in the first M rounds of training, and b is the average value of the loss function values of the training text samples in the subsequent X rounds of training.
[0060] S203b: Determine the sample type of each group of training text samples in the training text sample set according to the overall training loss reduction degree of each group of training text samples in the training text sample set.
[0061] After determining the overall training loss reduction degree of each training text sample in the training text sample set, the overall training loss reduction degrees of the training text samples in the training text sample set can be sorted from high to low first to obtain a sorting result. Then, the sample type of the training text samples ranked in the top n (such as the top 30%) in the sorting result can be determined as the first sample type (such as a normal sample); the sample type of the training text samples ranked in the bottom n (such as the bottom 30%) in the sorting result can be determined as the second sample type (such as an easy sample); and the sample type of the training text samples not ranked in the top n and the bottom n (such as after the top 30% and before the bottom 30%) in the sorting result can be determined as the third sample type (such as a difficult sample).
[0062] For example, assume that the training text sample set includes 100 training text samples. Sort the training process loss reduction degrees of each training text sample in the training text sample set from high to low to obtain a sorting result. The sample types of the training text samples ranked first to thirtieth can be determined as ordinary samples, the sample types of the training text samples ranked thirty-first to seventieth can be determined as difficult samples, and the sample types of the training text samples ranked seventy-first to one hundredth can be determined as simple samples.
[0063] Any combination of the above all optional technical solutions can form an optional embodiment of the present disclosure, which will not be elaborated herein one by one.
[0064] The following is an embodiment of the device of the present disclosure, which can be used to execute the method embodiment of the present disclosure. For the details not disclosed in the embodiment of the device of the present disclosure, please refer to the method embodiment of the present disclosure.
[0065] Figure 3 is a schematic diagram of a training device for a text retrieval model provided by an embodiment of the present disclosure. As Figure 3 shown, the training device for the text retrieval model includes:
[0066] A set acquisition unit 301, configured to acquire a training text sample set; wherein, the training text sample set includes a plurality of training text samples, and each group of training text samples includes a sample query text and a true article title corresponding to the sample query text; the true article title is an article title in the preset text database;
[0067] A numerical value determination unit 302, configured to, for each round of training of each group of training text samples, input the sample query text in the training text sample into a preset text retrieval model to obtain a text statement feature corresponding to the sample query text; query in the preset text database according to the text statement feature corresponding to the sample query text to obtain a predicted article title corresponding to the sample query text, wherein the predicted article title is an article title in the preset text database; determine the loss function value of this round of the training text sample according to the predicted article title and the true article title;
[0068] A type determination unit 303, configured to determine the sample type of each group of training text samples in the training text sample set according to the loss function values of each round of training of each group of training text samples in N rounds of training;
[0069] A model training unit 304, configured to train the text retrieval model by using the training text sample set, the sample types of each training text sample in the training text sample set, and the loss function weight values corresponding to the preset sample types, so as to obtain a trained text retrieval model.
[0070] Optionally, the preset text database stores multiple article titles and the text statement features respectively corresponding to each article title; the numerical value determination unit 302 is configured to:
[0071] For each article title in the preset text database, determine a matching value between the article title and the sample query text according to the text statement features corresponding to the article title and the text statement features corresponding to the sample query text;
[0072] Use the article title with the highest matching value with the sample query text in the preset text database as the predicted article title corresponding to the sample query text.
[0073] Optionally, the text retrieval model is a recurrent neural network or a self-attention network.
[0074] Optionally, the type determination unit 303 is configured to:
[0075] For each group of training text samples in the training text sample set, determine the overall training loss reduction degree of the training text sample according to the loss function values of the training text sample in each round of training;
[0076] Determine the sample type of each group of training text samples in the training text sample set according to the overall training loss reduction degrees of the groups of training text samples in the training text sample set.
[0077] Optionally, the type determination unit 303 is configured to:
[0078] Determine the average value of the loss function values in the first M rounds of training and the loss function values in the last X rounds of training according to the loss function values of the training text sample in each round of training, where both M and X are less than N;
[0079] Determine the overall training loss reduction degree of the training text sample according to the average value of the loss function values in the first M rounds of training and the loss function values in the last X rounds of training.
[0080] Optionally, the type determination unit 303 is configured to:
[0081] Sort the overall training loss reduction degrees of the groups of training text samples in the training text sample set from high to low to obtain a sorting result;
[0082] Determine the sample type of the training text samples ranked in the top n in the sorting result as the first sample type;
[0083] Determine the sample type of the training text samples ranked in the last n in the sorting result as the second sample type;
[0084] Determine the sample type of the training text samples that are not ranked in the top n and the last n in the sorting result as the third sample type.
[0085] Optionally, the loss function weight value corresponding to the training text sample with the sample type of the third sample type is greater than the loss function weight value corresponding to the training text sample with the sample type of the first sample type; the loss function weight value corresponding to the training text sample with the sample type of the second sample type is less than the loss function weight value corresponding to the training text sample with the sample type of the first sample type.
[0086] Optionally, the loss function of the text retrieval model is binary cross-entropy, multi-class cross-entropy loss function, logistic regression loss function, triplet loss function.
[0087] The beneficial effects of the embodiments of the present disclosure compared with the prior art are as follows: The embodiments of the present disclosure provide a training device for a text retrieval model, and the device includes: a set acquisition unit, configured to acquire a training text sample set; wherein, the training text sample set includes a plurality of training text samples, and each group of training text samples includes a sample query text and a true article title corresponding to the sample query text; the true article title is an article title in the preset text database; a numerical determination unit, configured to, for each round of training of each group of training text samples, input the sample query text in the training text sample into the preset text retrieval model to obtain a text statement feature corresponding to the sample query text; query in the preset text database according to the text statement feature corresponding to the sample query text to obtain a predicted article title corresponding to the sample query text, wherein the predicted article title is an article title in the preset text database; determine the loss function value of this round of the training text sample according to the predicted article title and the true article title; a type determination unit, configured to determine the sample type of each group of training text samples according to the loss function values of each group of training text samples in the training text sample set in each round of N rounds of training; a model training unit, configured to use the training text sample set, the sample type of each training text sample in the training text sample set, and the loss function weight value corresponding to the preset sample type to train the text retrieval model to obtain a trained text retrieval model.It can be seen that in this embodiment, first, a sample query text in each group of training text samples and the true article title corresponding to the sample query text are used to determine the loss function value of each group of training text samples in each round of training. Then, according to the loss function value of each group of training text samples in each round of training, each group of training text samples is classified to determine the sample type of the group of training text samples. In this way, a full global analysis and mining of the training text samples is carried out. By classifying the loss function values of the training text samples during the first-stage training process (i.e., N rounds of training), different sample types (such as simple samples, ordinary samples, and difficult samples) are identified. Since the influence degrees of training text samples of different sample types on the training effect of the text retrieval model (such as affecting the generalization ability of the text retrieval model) are different, different loss function weight values can be set according to different sample types. So, during the process of training the text retrieval model by using the training text sample set, the sample type of each group of training text samples in the training text sample set, and the loss function weight value corresponding to the preset sample type, the loss function weight value of the training text samples of the sample type that does not affect the model training effect can be reduced, and the loss function weight value of the training text samples of the sample type that has a greater impact on the model training effect can be increased, so as to give full play to the potential of the training text samples of the sample type that has a greater impact on the model training effect and reduce the influence of the training text samples of the sample type that has no impact or a poor impact on the model training effect on the training of the text retrieval model. It can be seen that this embodiment can effectively improve the distribution of the training weights (i.e., loss function weight values) of the training text samples of different sample types, so that the training process of the text retrieval model is more sufficient, the training efficiency and effect of the text retrieval model can be improved, and then the performance of the text retrieval model can be improved, and further the text retrieval effect of the text retrieval model in the actual business scenario can be improved (such as improving the accuracy of the text retrieval results of the text retrieval model).
[0088] It should be understood that the magnitudes of the sequence numbers of the steps in the above embodiments do not mean the order of execution is prior or posterior. The execution order of each process should be determined according to its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present disclosure.
[0089] Figure 4 is a schematic diagram of the computer device 4 provided by the embodiment of the present disclosure. As Figure 4 shown, the computer device 4 of this embodiment includes: a processor 401, a memory 402, and a computer program 403 stored in the memory 402 and capable of running on the processor 401. When the processor 401 executes the computer program 403, the steps in the above various method embodiments are implemented. Or, when the processor 401 executes the computer program 403, the functions of each module / module in the above various device embodiments are implemented.
[0090] Exemplarily, the computer program 403 can be divided into one or more modules, which are stored in the memory 402 and executed by the processor 401 to implement the present disclosure. One or more modules can be a series of computer program instruction segments capable of performing specific functions, and these instruction segments are used to describe the execution process of the computer program 403 in the computer device 4.
[0091] The computer device 4 can be a desktop computer, a notebook, a palm computer, a cloud server, or other computer devices. The computer device 4 can include, but is not limited to, the processor 401 and the memory 402. Those skilled in the art can understand that Figure 4 merely examples of the computer device 4, which do not constitute a limitation on the computer device 4, and may include more or fewer components than shown in the figure, or combine certain components, or different components. For example, the computer device may further include input / output devices, network access devices, a bus, etc.
[0092] The processor 401 can be a central processing module (Central Processing Unit, CPU), or other general-purpose processors, digital signal processors (Digital Signal Processor, DSP), application specific integrated circuits (Application Specific Integrated Circuit, ASIC), field programmable gate arrays (Field-Programmable Gate Array, FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc.
[0093] The memory 402 can be an internal storage module of the computer device 4. For example, the hard disk or memory of the computer device 4. The memory 402 can also be an external storage device of the computer device 4. For example, a plug-in hard disk, a smart media card (Smart Media Card, SMC), a secure digital (Secure Digital, SD) card, a flash card (Flash Card), etc. equipped on the computer device 4. Further, the memory 402 can also include both the internal storage module and the external storage device of the computer device 4. The memory 402 is used to store computer programs and other programs and data required by the computer device. The memory 402 can also be used to temporarily store data that has been output or will be output.
[0094] Those skilled in the art can clearly understand that, for the convenience and simplicity of description, only the above-mentioned functional modules and module divisions are used as examples. In actual applications, the above functions can be allocated to different functional modules or modules as needed, that is, the internal structure of the device can be divided into different functional modules or modules to complete all or part of the functions described above. Each functional module and module in the embodiments can be integrated into a processing module, or each module can exist physically alone, or two or more modules can be integrated into one module. The above integrated modules can be implemented in the form of hardware or in the form of software functional modules. In addition, the specific names of the functional modules and modules are only for the convenience of distinguishing from each other and do not limit the protection scope of the present disclosure. The specific working processes of the modules and modules in the above system can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.
[0095] In the above embodiments, the descriptions of the respective embodiments have their own emphases. For parts not detailed or recorded in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0096] Those of ordinary skill in the art can realize that the modules and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or by a combination of computer software and electronic hardware. Whether these functions are executed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present disclosure.
[0097] In the embodiments provided by the present disclosure, it should be understood that the disclosed device / computer device and method can be implemented in other ways. For example, the device / computer device embodiments described above are only illustrative. For example, the division of modules or modules is only a logical function division, and there can be other division methods in actual implementation. Multiple modules or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces, and the indirect couplings or communication connections of the devices or modules can be in electrical, mechanical or other forms.
[0098] The modules described as separate components may or may not be physically separated, and the components shown as modules may or may not be physical modules, that is, they can be located in one place or distributed to multiple network modules. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0099] In addition, in each of the embodiments of the present disclosure, each functional module may be integrated into one processing module, or each module may exist physically alone, or two or more modules may be integrated into one module. The above-mentioned integrated module may be implemented in the form of hardware or in the form of a software functional module.
[0100] When the integrated module / module is implemented in the form of a software functional module and sold or used as an independent product, it may be stored in a computer-readable storage medium. Based on such an understanding, all or part of the processes in the above-mentioned method embodiments of the present disclosure may also be completed by instructing relevant hardware through a computer program. The computer program may be stored in the computer-readable storage medium. When the computer program is executed by a processor, the steps of the above-mentioned method embodiments may be implemented. The computer program may include computer program code, and the computer program code may be in the form of source code, object code, executable file or some intermediate form, etc. The computer-readable medium may include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal, and software distribution medium, etc. It should be noted that the content included in the computer-readable medium may be appropriately increased or decreased according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, the computer-readable medium does not include electrical carrier signals and telecommunication signals.
[0101] The above embodiments are only used to illustrate the technical solutions of the present disclosure, rather than to limit them; although the present disclosure has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present disclosure, and should all be included in the protection scope of the present disclosure.
Claims
1. A training method for a text retrieval model, characterized in that, The method includes: Obtain a training text sample set; wherein, the training text sample set includes a number of training text samples, and each group of training text samples includes a sample query text and the true article title corresponding to the sample query text; the true article title is an article title in a preset text database; For each round of training of each group of training text samples, input the sample query text in the training text sample into a preset text retrieval model to obtain the text statement features corresponding to the sample query text; according to the text statement features corresponding to the sample query text, query in the preset text database to obtain the predicted article title corresponding to the sample query text, wherein the predicted article title is an article title in the preset text database; determine the loss function value of this round of the training text sample according to the predicted article title and the true article title; Determine the sample type of each group of training text samples according to the loss function values of each group of training text samples in the training text sample set in each round of N rounds of training; Use the training text sample set, the sample type of each training text sample in the training text sample set, and the loss function weight values corresponding to the preset sample types to train the text retrieval model to obtain a trained text retrieval model; The determining the sample type of each group of training text samples according to the loss function values of each group of training text samples in the training text sample set in each round of N rounds of training includes: For each group of training text samples in the training text sample set, determine the overall training loss reduction degree of the training text sample according to the loss function values of the training text sample in each round of training; Determine the sample type of each group of training text samples according to the overall training loss reduction degrees of each group of training text samples in the training text sample set; The determining the overall training loss reduction degree of the training text sample according to the loss function values of the training text sample in each round of training includes: Determine the average value of the loss function values of the first M rounds of training and the loss function values of the last X rounds of training according to the loss function values of the training text sample in each round of training, wherein both M and X are less than N; Determine the overall training loss reduction degree of the training text sample according to the average value of the loss function values of the first M rounds of training and the loss function values of the last X rounds of training.
2. The method according to claim 1, wherein The preset text database stores multiple article titles and the text statement features respectively corresponding to each article title; The querying in the preset text database according to the text statement features corresponding to the sample query text to obtain the predicted article title corresponding to the sample query text includes: For each article title in the preset text database, determine the matching value between the article title and the sample query text according to the text statement features corresponding to the article title and the text statement features corresponding to the sample query text; Use the article title with the highest matching value between the preset text database and the sample query text as the predicted article title corresponding to the sample query text.
3. The method according to claim 1, wherein The text retrieval model is a recurrent neural network or a self-attention network.
4. The method according to claim 1, wherein Determining the sample type of each group of training text samples in the training text sample set according to the training overall loss reduction degree of each group of training text samples in the training text sample set includes: Sort the training overall loss reduction degrees of each group of training text samples in the training text sample set from high to low to obtain a sorting result; Determine the sample type of the training text samples in the top n positions in the sorting result as the first sample type; Determine the sample type of the training text samples in the last n positions in the sorting result as the second sample type; Determine the sample type of the training text samples that are not in the top n positions and the last n positions in the sorting result as the third sample type.
5. The method according to claim 4, characterized in that, The loss function weight value corresponding to the training text sample with the sample type of the third sample type is greater than the loss function weight value corresponding to the training text sample with the sample type of the first sample type; the loss function weight value corresponding to the training text sample with the sample type of the second sample type is less than the loss function weight value corresponding to the training text sample with the sample type of the first sample type.
6. The method according to claim 1, wherein The loss function of the text retrieval model is binary cross-entropy, multi-class cross-entropy loss function, logistic regression loss function, or triplet loss function.
7. A training device for a text retrieval model, which is used to implement the method according to any one of claims 1 to 6, characterized in that, The device includes: A set acquisition unit for acquiring a training text sample set; wherein, the training text sample set includes a plurality of training text samples, and each group of training text samples includes a sample query text and the true article title corresponding to the sample query text; the true article title is an article title in a preset text database; A value determination unit for, for each round of training of each group of training text samples, inputting the sample query text in the training text sample into a preset text retrieval model to obtain the text statement feature corresponding to the sample query text; querying in the preset text database according to the text statement feature corresponding to the sample query text to obtain the predicted article title corresponding to the sample query text, wherein the predicted article title is an article title in the preset text database; determining the loss function value of this round of the training text sample according to the predicted article title and the true article title; A type determination unit for determining the sample type of each group of training text samples in the training text sample set according to the loss function value of each round of training of each group of training text samples in the N rounds of training; A model training unit for training the text retrieval model by using the training text sample set, the sample type of each training text sample in the training text sample set, and the loss function weight value corresponding to the preset sample type to obtain a trained text retrieval model.
8. A computer device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 6.
9. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the steps of the method according to any one of claims 1 to 6.
Citation Information
Patent Citations
Machine reading comprehension method based on multi-task joint training, and computer storage medium
CN110309305A
Text generation model training method and device and electronic equipment
CN111625645A