Question and answer system performance improvement method based on multi-task balance and generalization enhancement

CN119988969APending Publication Date: 2025-05-13北京粉笔上岸科技有限公司
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510064846.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-15
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

The existing technology is difficult to coordinate the capabilities of multiple professional fields, resulting in data equalization, difficulty equalization and convergence equalization problems, which in turn affects the generalization of the model and the performance of the question-and-answer system.

Method used

Using Loose Teacher Forcing technology, Focal Loss and data balance, difficulty balance, and convergence balance strategies are introduced to coordinate the training and reasoning process of the model through these means to ensure that the model performs balancedly in multiple professional fields and has strong generalization capabilities.

Benefits of technology

The generalization problems caused by data balance, difficulty balance, convergence balance and gap between training and reasoning were solved, and a question-and-answer system with excellent abilities in multiple professional fields and strong generalization was trained.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119988969A_ABST
    Figure CN119988969A_ABST
Patent Text Reader

Abstract

The invention provides a question and answer system performance improvement method based on multi-task balance and generalization enhancement, and relates to the technical field of artificial intelligence, comprising the following steps: randomly selecting a window w of a training sample sequence in the middle and later periods of training by adopting a Loose Teacher Forching technology, and replacing an existing token of an original sample with a token of a model in the window w; the total number of tokens corresponding to the field is used when the loss of the sample without the field is calculated; according to the method, the Loose Teacher Forcing is used for replacing an existing Hard Teacher Forcing training scheme, gap during model training and reasoning is reduced, and then data balance, difficulty balance and convergence balance are freely combined and used according to the characteristics of a training set, so that stable training is carried out, the ability of learning each specialty is coordinated, and a trained model is obtained.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and in particular to a method for improving the performance of a question-answering system based on multi-task balance and generalization enhancement. Background Art

[0002] With the rapid development of artificial intelligence technology, the application of intelligent question-answering systems in various fields is becoming more and more extensive. At present, in the field of building professional question-answering systems, there are mainly two technical routes: (1) Post-training scheme based on the Instruct model: This scheme is based on the existing instruction model and optimizes the model through the post-training process, including: fine-tuning, using domain professional data to supervise the model, adopting the teacherforcing training paradigm, optimizing the model parameters through maximum likelihood estimation, while maintaining the original capabilities of the model, while injecting professional domain knowledge; human preference alignment, using reinforcement learning human feedback (RLHF) technology to make the model output more in line with human expectations, using new direct alignment methods such as direct preference optimization (DPO) to improve the reliability and security of the model's answers, and optimizing the output quality and applicability of the model through human feedback data; (2) Continuing pre-training scheme based on the Base model: Starting from the basic language model, this scheme achieves specialization through multi-stage training, mainly including: continuing pre-training stage (Continue Pretraining) uses large-scale professional domain corpus for self-supervised learning, and enhances the model's domain knowledge understanding through pre-training tasks such as masked language modeling, maintaining the model's basic language ability while injecting professional knowledge; in the question-answering ability optimization stage, supervised fine-tuning is performed based on professional question-answering datasets, and the teacher forcing mechanism is also used to predict the next token to optimize the model's question-answering performance and professional ability. The common features of these two technical solutions are: both use the teacher forcing method for training, predict the next token through autoregression, and use the cross-entropy loss function to optimize model parameters;

[0003] However, the above technical solutions cannot coordinate the capabilities of multiple professional fields well. On the one hand, different professional fields can collect different amounts of high-quality data, and existing research shows that in large model training, especially in the post-training stage, the quality of data is far greater than the amount of data. On the other hand, the learning difficulty of different professional fields is also different. The existing technical solutions use the average calculation method for the data samples with the above data volume and learning difficulty, which will cause the following problems:

[0004] Data balance problem: The amount of data available for training in different professional fields is unbalanced. Existing technologies generally do not distinguish or simply upsample data from low-data-volume data sources. The former will cause low-data-volume fields to contribute less to the loss, which in turn has little impact on model updates. Ultimately, the trained model performs poorly in these fields. The latter simply upsampling will cause the model to quickly remember samples that are repeated many times, which in turn affects the generalization of the model, so it is also suboptimal.

[0005] Difficulty balance problem: The learning difficulty of different samples in different majors or even in the same major field is also very different. The existing technical solutions do not distinguish these samples of different difficulty levels, which will cause simple samples to be learned early or learned many times while difficult samples are not learned enough, that is, there is a problem of difficulty balance;

[0006] Convergence balance problem: From the perspective of the degree of model learning and fitting the training set, the convergence speed of different specialties is different. The existing technical solutions ignore this difference. Even the current mainstream large-model training frameworks such as HuggingfaceTRL, Llama Factory, and ms-swift do not support observing the loss of different sub-datasets. The neglect of the progress of training in different fields conceals a very important fact: the convergence speed of different specialties is different. If we observe the respective validation set losses and other indicators, we will find that in the later stages of training, some professional fields are still converging, while others have begun to overfit. The existing technical solutions lack the understanding and response methods for convergence balance, which will result in the final model failing to perform well in multiple downstream application fields.

[0007] Generalization problems caused by the gap between training and reasoning: Existing technical solutions generally use the teacher forcing method for training, and predict the next token through autoregression. That is to say, the prefix used to predict the next token during training is a fixed prefix in the given data set. The reason for adopting this method is to reduce the difficulty of training and improve the efficiency of training. However, there is no existing answer for the model during reasoning. The model can only continue to predict the next token based on the token it has inferred each time, until the EOS token is inferred. Therefore, there is an inherent gap between the training and reasoning of the model. In theory, if the training set is infinitely large to include all token paths that may be encountered during reasoning, this gap can be offset. However, in the domain fine-tuning scenario, the opposite is true, which is limited by limited high-quality professional knowledge. Therefore, this gap between training and reasoning will greatly affect the generalization ability of the model.

[0008] Therefore, the present invention proposes a question answering system performance improvement method based on multi-task balance and generalization enhancement to solve the problems existing in the prior art. Summary of the invention

[0009] In response to the above problems, the present invention proposes a method for improving the performance of a question-answering system based on multi-task balance and generalization enhancement. The method for improving the performance of a question-answering system based on multi-task balance and generalization enhancement solves the data balance problem, the difficulty balance problem, the convergence balance problem, and the generalization problem caused by the gap between training and reasoning. It trains a question-answering system with excellent capabilities in multiple professional fields and that can effectively deal with problems that are not in the training field, thereby providing users with more comprehensive and helpful assistance.

[0010] To achieve the purpose of the present invention, the present invention is implemented by the following technical solution: a method for improving the performance of a question-answering system based on multi-task balance and generalization enhancement, comprising the following steps:

[0011] S1: Loose Teacher Forching technology is used to randomly select a window w of the training sample sequence in the middle and late stages of training, and the model's own token is used in the window w to replace the existing token of the original sample;

[0012] S2: When calculating the loss of samples in different fields, the total number of tokens corresponding to the field is used to achieve data balance;

[0013] S3: Introduce Focal Loss into the model Post Training, and focus on key tokens through Focal Loss;

[0014] S4: During the training process, the indicators of the validation sets of different professional fields are recorded, and the slope of the linear fit of the indicator set in the most recent observation interval k is used as an indicator of the convergence status of the current professional field to coordinate different convergence speeds.

[0015] A further improvement is that in S1, the behavior in window w is consistent with that during inference, thereby alleviating the gap between model training and inference. At the same time, the model's own token is used to stimulate a larger exploration space.

[0016] A further improvement is that: S2 specifically includes the following steps:

[0017] When calculating the loss of samples in different fields, the total number of tokens corresponding to the field is used instead of the tokens of all samples;

[0018] Shield the impact of the size of field data on data in other fields, and shield the impact of the amount of data in other fields;

[0019] Finally, data balance is achieved.

[0020] A further improvement is that in S3, while focusing on key tokens, less attention is paid to excessive tokens.

[0021] A further improvement is that in S3, through the focus adjustment of Focal Loss, during training, the overall perplexity of a sample is counted, and a higher weight is assigned to samples with high perplexity.

[0022] A further improvement is to combine S3 with S1 and select the subsequence interval with high difficulty as the window in the Loss Teaching Forcing to explore the more difficult interval.

[0023] A further improvement is that in S4, when the convergence speed of a certain field slows down, its weight on the loss is dynamically increased; when the convergence speed of a certain field is too fast or starts to overfit, its weight on the loss is dynamically lowered to achieve the effect of coordinating different convergence speeds.

[0024] The beneficial effects of the present invention are:

[0025] 1. The present invention uses Loose Teacher Forcing to replace the existing Hard Teacher Forcing training scheme to reduce the gap between model training and reasoning, and then uses a balanced training method to freely combine data balance, difficulty balance, and convergence balance according to the characteristics of the training set to perform stable training and coordinate the learning capabilities of various professions to obtain the final trained model. In summary, the data balance problem, difficulty balance problem, convergence balance problem, and generalization problem caused by the gap between training and reasoning are solved, and a question-answering system with excellent capabilities in multiple professional fields and strong generalization that can effectively deal with problems that are not in the training field is trained, providing users with more comprehensive and helpful assistance. BRIEF DESCRIPTION OF THE DRAWINGS

[0026] Figure 1 Schematic diagram of the joint training framework of the present invention. DETAILED DESCRIPTION

[0027] In order to deepen the understanding of the present invention, the present invention will be further described in detail below in conjunction with examples. The examples are only used to explain the present invention and do not constitute a limitation on the protection scope of the present invention.

[0028] Embodiment 1

[0029] according to Figure 1 As shown, this embodiment proposes a method for improving the performance of a question-answering system based on multi-task balance and generalization enhancement, comprising the following steps:

[0030] S1: Loose Teacher Forching technology is used to randomly select a window w of the training sample sequence in the middle and late stages of training. The model's own token is used in window w to replace the existing token of the original sample. The behavior in window w is consistent with that during inference, thereby reducing the gap between model training and inference. At the same time, the model's own token is used to stimulate a larger exploration space.

[0031] S2: When calculating the loss of samples in different fields, the total number of tokens corresponding to the field is used to achieve data balance; specifically, the following steps are included: when calculating the loss of samples in different fields, the total number of tokens corresponding to the field is used to replace the tokens of all samples; the influence of the size of the field data on the data in other fields is shielded, and the influence of the amount of data in other fields is shielded; finally, data balance is achieved;

[0032] S3: Introduce Focal Loss into the model Post Training, and use Focal Loss to focus on key tokens; while focusing on key tokens, reduce the focus on excessive tokens, and adjust the focus of Focal Loss. During training, count the overall perplexity of a sample and assign higher weights to samples with high perplexity; when selecting Loss Teaching Forcing, select the subsequence interval with high difficulty as the window to explore the more difficult interval;

[0033] S4: During the training process, the indicators of the validation sets of different professional fields are recorded, and the slope of the linear fit of the indicator set in the most recent observation interval k is used as an indicator of the convergence of the current professional field to coordinate different convergence speeds. When the convergence speed of a certain field slows down, its weight on the loss is dynamically increased. When the convergence speed of a certain field is too fast or starts to overfit, its weight on the loss is dynamically lowered to achieve the effect of coordinating different convergence speeds.

[0034] The present invention enables the question-answering system to perform in a balanced manner in multiple professional fields and possess strong generalization capabilities, which has important practical significance for improving user experience and meeting market demand.

[0035] Embodiment 2

[0036] according to Figure 1 As shown, this embodiment proposes a method for improving the performance of a question-answering system based on multi-task balance and generalization enhancement, such as Figure 1The left half of the figure represents the training data of different professional fields, and the size of the rectangle represents the size of different professional data sets. During model training, the present invention proposes to use Loose Teacher Forcing to replace the existing Hard Teacher Forcing training scheme to reduce the gap between model training and reasoning. Finally, through the balanced training module, according to the characteristics of the training set, the data balance, difficulty balance, and convergence balance modules are freely combined for stable training, and the ability of each major is coordinated to obtain the final trained model.

[0037] Loose Teacher Forching

[0038] In order to solve the generalization problem caused by the gap between training and reasoning in the prior art, the present invention proposes Losse Teacher Forching to replace the existing Teacher Forcing technology. Loose Teacher Forching randomly selects a window w of the training sample sequence in the middle and late stages of training, and uses the model's own token in this window to replace the existing token of the original sample. The behavior in this window w is consistent with that during reasoning. This approach reduces the gap between model training and reasoning. In addition, using the model's own token can stimulate a larger exploration space, which can play a role in regularization and enhance the generalization ability of the model.

[0039] Data Balance

[0040] The present invention proposes data balancing to solve the problem that the existing technology cannot cope with unbalanced data. Specifically, the existing technology treats tokens in different professional fields in the same way, that is, the total number of tokens in all professional fields is used as the normalized denominator of the loss function of all field samples. This will roughly make professional field data with a large number contribute more to the training loss, and the model will be more inclined to the field with a large amount of data. The data balancing of the present invention will use the total number of tokens corresponding to the field when calculating the loss of samples in different fields instead of the tokens of all samples. In this way, no matter the size of the field data, it will not affect the data in other fields, nor will it be affected by the amount of data in other fields. The effect of data balancing is achieved.

[0041] Difficulty Balance

[0042] The present invention proposes difficulty balancing to solve the problem that the existing technology cannot cope with the different learning difficulties of different samples. Specifically, Focal Loss is introduced into the model Post Training. The essence of a large model predicting the next token is to do a classification task of the vocabulary size, so Focal Loss can focus on difficult-to-predict tokens, which are generally key tokens, and reduce easier-to-predict tokens, which are generally transitional tokens. Further based on similar ideas, it is chosen to count the overall perplexity of a sample during training, and assign higher weights to samples with high perplexity.

[0043] In addition, when selecting Loss Teaching Forcing, we select the subsequence interval with high difficulty as the window to achieve more exploration of the more difficult interval.

[0044] Convergent Equilibrium

[0045] The present invention proposes convergence balance to solve the problem that the existing technology cannot cope with the different actual convergence speeds of different professions; specifically, during the training process, the indicators of the validation sets of different professional fields are recorded, and the slope of the linear fit of the indicator set in the most recent observation interval k is used as an indicator of the convergence status of the current professional field: when the convergence speed of a certain field slows down, its weight on the loss is dynamically increased; when the convergence speed of a certain field is very fast or starts to overfit, its weight on the loss is dynamically lowered, so as to achieve the effect of coordinating different convergence speeds.

[0046] This method for improving the performance of a question-answering system based on multi-task balance and generalization enhancement uses Loose Teacher Forcing to replace the existing Hard Teacher Forcing training scheme to reduce the gap between model training and reasoning. Then, through a balanced training method, data balance, difficulty balance, and convergence balance are freely combined according to the characteristics of the training set to perform stable training and coordinate the learning capabilities of various professions to obtain the final trained model. In summary, the data balance problem, difficulty balance problem, convergence balance problem, and generalization problem caused by the gap between training and reasoning are solved, and a question-answering system with excellent capabilities in multiple professional fields and strong generalization that can effectively cope with fields that are not in the training field is trained, providing users with more comprehensive and helpful assistance.

[0047] The above shows and describes the basic principles, main features and advantages of the present invention. It should be understood by those skilled in the art that the present invention is not limited to the above embodiments. The above embodiments and descriptions are only for explaining the principles of the present invention. Without departing from the spirit and scope of the present invention, the present invention may have various changes and improvements, which fall within the scope of the present invention. The scope of protection of the present invention is defined by the attached claims and their equivalents.

Claims

1. A method for improving the performance of a question-answering system based on multi-task balance and generalization enhancement, characterized in that: The following steps are involved: S1: Loose Teacher Forching technology is used to randomly select a window w of the training sample sequence in the middle and late stages of training, and the model's own token is used in the window w to replace the existing token of the original sample; S2: When calculating the loss of samples in different fields, the total number of tokens corresponding to the field is used to achieve data balance; S3: Introduce Focal Loss into the model Post Training, and focus on key tokens through Focal Loss; S4: During the training process, the indicators of the validation sets of different professional fields are recorded, and the slope of the linear fit of the indicator set in the most recent observation interval k is used as an indicator of the convergence status of the current professional field to coordinate different convergence speeds.

2. The method for improving the performance of a question-answering system based on multi-task balance and generalization enhancement according to claim 1, characterized in that: In S1, the behavior in window w is consistent with that during inference, thereby alleviating the gap between model training and inference. At the same time, the model's own token is used to stimulate a larger exploration space.

3. The method for improving the performance of a question-answering system based on multi-task balance and generalization enhancement according to claim 1, characterized in that: The S2 specifically includes the following steps: When calculating the loss of samples in different fields, the total number of tokens corresponding to the field is used instead of the tokens of all samples; Shield the impact of the size of field data on data in other fields, and shield the impact of the amount of data in other fields; Finally, data balance is achieved.

4. The method for improving the performance of a question-answering system based on multi-task balance and generalization enhancement according to claim 1, characterized in that: In the above S3, while focusing on key tokens, less attention is paid to excessive tokens.

5. The method for improving the performance of a question-answering system based on multi-task balance and generalization enhancement according to claim 4, characterized in that: In S3, through the focus adjustment of Focal Loss, during training, the overall perplexity of a sample is counted, and a higher weight is assigned to samples with high perplexity.

6. The method for improving the performance of a question-answering system based on multi-task balance and generalization enhancement according to claim 5, characterized in that: Combine S3 with S1, and when selecting Loss Teaching Forcing, select the subsequence interval with high difficulty as the window to explore the more difficult interval.

7. The method for improving the performance of a question-answering system based on multi-task balance and generalization enhancement according to claim 1, characterized in that: In S4, when the convergence speed of a certain field slows down, its weight on the loss is dynamically increased. When the convergence speed of a certain field is too fast or starts to overfit, its weight on the loss is dynamically lowered to achieve the effect of coordinating different convergence speeds.