Model training and model question and answer method and device, electronic equipment and storage medium

By generating the same type of question-and-answer sample data and adjusting the data proportional weight and teachability scores, and optimizing model training with the negative feedback mechanism, the problem of poor quality of self-learning model training is solved, and the generalization ability and application effect of the model are improved.

CN120523912APending Publication Date: 2025-08-22CHINA TELECOM ARTIFICIAL INTELLIGENCE TECHNOLOGY (BEIJING) CO LTD
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202510646942.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-19
Publication Date
2025-08-22

AI Technical Summary

Technical Problem

The quality of the training data generated by the self-learning model in the prior art is uneven, resulting in poor generalization capabilities of the model and poor application effect.

Method used

By generating second question-and-answer sample data of the same type as the original training data, adjusting the data proportional weight and teachability scores, optimizing the model training process, and optimizing the learning of the model on self-generated data in combination with the negative feedback mechanism.

Benefits of technology

The model's performance and generalization ability in real-world tasks is improved, and the poor application effect caused by poor data generated by self-learning models is solved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120523912A_ABST
    Figure CN120523912A_ABST
Patent Text Reader

Abstract

The invention discloses a model training method and device, a model question-answering method and device, electronic equipment and a storage medium. The method comprises the steps that a first model is adopted, second question and answer sample data of the same type as first question and answer sample data is generated, and an extended training data set is obtained; according to a data matching weight corresponding to each type of data in the extended training data set, obtaining various types of second question and answer sample data to train the basic model to obtain a second model; determining a corresponding teaching score when the second model processes the question text in the original training data set, and adjusting the data matching weight according to the teaching score to obtain a target data matching weight; and training the basic model according to the target data matching weight to obtain a target model. The technical problem of poor application effect of the model obtained by final training due to poor quality of training data generated by a self-learning model in related technologies is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of artificial intelligence technology, and more specifically, to a model training and model question-answering method, device, electronic device, and storage medium. Background Art

[0002] In the field of artificial intelligence, particularly in deep learning and large-scale model applications, self-learning mechanisms have become a key means of improving model performance. Traditional machine learning methods rely on large amounts of labeled data to train models, a process that is both expensive and time-consuming. To overcome these challenges, various self-learning strategies have emerged in related technologies, enabling models to learn and fine-tune using self-generated data.

[0003] However, the self-learning methods in related technologies have problems such as uneven quality of training data generated by the model and poor generalization ability of the model, which leads to poor application effect of the final trained model.

[0004] To address the above-mentioned problems, no effective solutions have been proposed so far. Summary of the Invention

[0005] The embodiments of the present application provide a model training and model question-answering method, device, electronic device and storage medium to at least solve the technical problem of poor application effect of the final trained model caused by the poor quality of training data generated by the self-learning model in the related technology.

[0006] According to one aspect of an embodiment of the present application, a model training method is provided, comprising: using a first model to analyze first question-answer sample data, generating second question-answer sample data of the same type as the first question-answer sample data, and obtaining an extended training data set, wherein the first model is obtained by training a basic model using an original training data set, and the original training data set contains a plurality of manually annotated first question-answer sample data of different types, and each type corresponds to a plurality of second question-answer sample data; according to the data ratio weight corresponding to each type of data in the extended training data set, obtaining various types of second question-answer sample data in the extended training data set, and obtaining the second question-answer sample data according to the obtained data ratio weight. The second question and answer sample data obtained are used to train the basic model to obtain the second model, wherein the data matching weight is used to indicate the collection ratio of each type of second question and answer sample data when training the model; the teachability score corresponding to the second model when processing the question text in the original training data set is determined, and the data matching weight is adjusted according to the teachability score to obtain the target data matching weight, wherein the teachability score is used to characterize the degree of learning value of each type of data for the basic model; the basic model is trained according to the original training data set, the expanded training data set, and the target data matching weight to obtain the target model.

[0007] Optionally, using the first model to generate second question and answer sample data of the same type as the first question and answer sample data includes: clustering the first question and answer sample data in the original training data set to obtain multiple clusters, wherein each cluster corresponds to a type; extracting a first preset number of first question and answer sample data from each cluster to form a first training data set, wherein the total number of first question and answer sample data in the first training data set is less than the total number of first question and answer sample data in the original training data set; using the first question and answer sample data of each type in the first training data set as seed data; using the first model to generate diversity based on the seed data to obtain second question and answer sample data of the same type as the seed data, and using the second question and answer sample data to form an extended training data set, wherein, in the process of diversity generation, a second preset number of second question and answer sample data are generated based on each type of seed data.

[0008] Optionally, the basic model is trained based on the acquired second question and answer sample data to obtain the second model, including: determining the data ratio weight corresponding to each type of data in the extended training data set, wherein each type corresponds to a data ratio weight; for each type of second question and answer sample data, according to the collection ratio corresponding to the data ratio weight of the type, the second question and answer sample data of the collection ratio is randomly obtained from all the second question and answer sample data of the type in the extended training data set to form a second training data set; based on the second training data set, the basic model is trained to obtain the second model.

[0009] Optionally, determining the teachability score corresponding to the question text in the original training data set when the second model processes the question text includes: using the second model to analyze and process the question text in the first question and answer sample data of each type in the original training data set to obtain the answer text output by the second model; determining a loss value parameter and a confidence parameter based on the answer text, wherein the loss value parameter is used to characterize the degree of error between the answer text output by the second model and the answer text corresponding to the question text in the first question and answer sample data, and the confidence parameter is used to characterize the average confidence of each word in the answer text output by the second model; determining the teachability score corresponding to the type based on the loss value parameter and the confidence parameter, wherein each type of data in the original training data set corresponds to a teachability score.

[0010] Optionally, the adjustment of the data matching weight includes multiple training batches; the data matching weight is adjusted according to the teachability score to obtain the target data matching weight, including: in each training batch, obtaining the first teachability score corresponding to the second model of the first training batch, and the second teachability score corresponding to the second model of the second training batch, wherein the second training batch is the training batch immediately before the first training batch; comparing the size changes of each type of teachability score between the first teachability score and the second teachability score, and adjusting the data matching weight corresponding to each type according to the size changes; according to the adjusted data matching weight, re-obtaining the second question and answer sample data of various types in the expanded training data set, and re-obtaining the second question and answer sample data according to the obtained second question and answer sample data. Data, train the basic model to obtain the second model corresponding to the next batch, and iteratively repeat the above steps of adjusting the data matching weight according to the teachability score until the preset iteration termination condition is met, wherein the preset iteration termination condition includes at least one of the following: the change amplitude between the teachability scores corresponding to two consecutive training batches is less than the preset amplitude threshold, the number of times the sum of the teachable scores continues to increase exceeds the preset number threshold, and the loss value parameter corresponding to the second model when processing the question text in the original training data set is less than the loss value parameter corresponding to the first model when processing the question text in the original training data set; when the preset iteration termination condition is met, the data matching weight obtained after the last adjustment is determined as the target data matching weight.

[0011] Optionally, based on the size change, the data allocation weight corresponding to each type is adjusted, including: when the first teachability score corresponding to the type is greater than the second teachability score, the data allocation weight corresponding to the type is increased by a preset proportional step; when the first teachability score corresponding to the type is less than the second teachability score, the data allocation weight corresponding to the type is reduced by a preset proportional step.

[0012] Optionally, the basic model is trained based on the original training data set, the extended training data set, and the target data ratio weight to obtain the target model, including: obtaining various types of second question and answer sample data in the extended training data set according to the collection ratio corresponding to the target data ratio weight; the obtained second question and answer sample data and the first question and answer sample data in the original training data set are combined to form a target training data set; based on the target training data set, the basic model is trained to obtain the target model.

[0013] According to another aspect of the embodiment of the present application, a model question-answering method is also provided, including: obtaining a question text, wherein the question text is a text containing a question to be solved by the user; using a target model to analyze the question text, and obtaining an answer text corresponding to the question text generated by the target model; wherein the training step of the target model includes: using a first model to analyze the first question-answering sample data, generating second question-answering sample data of the same type as the first question-answering sample data, and obtaining an extended training data set, wherein the first model is obtained by training the basic model using the original training data set, and the original training data set contains multiple manually annotated first question-answering sample data of different types, and each type corresponds to multiple second question-answering sample data; according to each type in the extended training data set, The data matching weight corresponding to the type of data is obtained, and the second question and answer sample data of various types in the extended training data set are obtained. The basic model is trained based on the obtained second question and answer sample data to obtain a second model, wherein the data matching weight is used to indicate the collection ratio of each type of second question and answer sample data when training the model; the teachability score corresponding to the second model when processing the question text in the original training data set is determined, and the data matching weight is adjusted based on the teachability score to obtain the target data matching weight, wherein the teachability score is used to characterize the degree of learning value of each type of data for the basic model; the basic model is trained based on the original training data set, the extended training data set, and the target data matching weight to obtain the target model.

[0014] According to another aspect of the embodiment of the present application, a model training device is also provided, including: a data generation module, for using a first model to analyze the first question and answer sample data, generate second question and answer sample data of the same type as the first question and answer sample data, and obtain an extended training data set, wherein the first model is obtained by training the basic model using the original training data set, and the original training data set contains multiple manually labeled first question and answer sample data of different types, and each type corresponds to multiple second question and answer sample data; a first fine-tuning module, for obtaining various types of second question and answer sample data in the extended training data set according to the data ratio weight corresponding to each type of data in the extended training data set, and The second question and answer sample data obtained is used to train the basic model to obtain a second model, wherein the data ratio weight is used to indicate the collection ratio of each type of second question and answer sample data when training the model; the ratio adjustment module is used to determine the teachability score corresponding to the second model when processing the question text in the original training data set, and adjust the data ratio weight according to the teachability score to obtain the target data ratio weight, wherein the teachability score is used to characterize the degree of learning value of each type of data for the basic model; the second fine-tuning module is used to train the basic model according to the original training data set, the expanded training data set, and the target data ratio weight to obtain the target model.

[0015] According to another aspect of an embodiment of the present application, an electronic device is provided, including: a memory and a processor, the processor being configured to run a program stored in the memory, wherein the program executes a model training method or a model question-answering method when the program is running.

[0016] According to another aspect of an embodiment of the present application, a non-volatile storage medium is further provided, wherein the non-volatile storage medium includes a stored computer program, wherein the device where the non-volatile storage medium is located executes a model training method or a model question-answering method by running the computer program.

[0017] According to another aspect of the embodiments of the present application, a computer program product is also provided, including a computer program, which implements the steps of the model training method or the model question-answering method when the computer program is executed by a processor.

[0018] In an embodiment of the present application, the first question and answer sample data is analyzed by adopting the first model to generate second question and answer sample data of the same type as the first question and answer sample data, and an extended training data set is obtained, wherein the first model is obtained by training the basic model using the original training data set, and the original training data set contains multiple manually labeled first question and answer sample data of different types, and each type corresponds to multiple second question and answer sample data; according to the data ratio weight corresponding to each type of data in the extended training data set, the various types of second question and answer sample data in the extended training data set are obtained, and based on the obtained second question and answer sample data, the basic model is trained to obtain the second model, wherein the data ratio weight is used to indicate the collection of each type of second question and answer sample data when performing model training. proportion; determine the teachability score corresponding to the second model when processing the question text in the original training data set, and adjust the data matching weight according to the teachability score to obtain the target data matching weight, wherein the teachability score is used to characterize the degree of learning value of each type of data to the basic model; train the basic model according to the original training data set, the expanded training data set, and the target data matching weight to obtain the target model. By using negative feedback indicators to adjust the learning process of the optimization model on self-generated data, the purpose of improving the performance and generalization ability of the model in real-world tasks is achieved, thereby solving the technical problem of poor application effect of the final trained model caused by the poor quality of training data generated by the self-learning model in the related technology. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings:

[0020] Figure 1 This is a hardware structure block diagram of a computer terminal (or electronic device) for implementing a model training method or a model question-answering method provided in an embodiment of the present application;

[0021] Figure 2 This is a schematic diagram of a model training method flow provided according to an embodiment of the present application;

[0022] Figure 3 This is a schematic diagram of a method flow for fine-tuning a self-learning large model based on data negative feedback provided in an embodiment of the present application;

[0023] Figure 4 This is a schematic diagram of a model question-answering method flow according to an embodiment of the present application;

[0024] Figure 5It is a structural diagram of a model training device provided according to an embodiment of the present application. DETAILED DESCRIPTION

[0025] In order to enable those skilled in the art to better understand the present invention, the following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments in the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of this application.

[0026] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequential order. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the present application described herein can be implemented in a sequence other than those illustrated or described herein. In addition, the terms "including" and "having" and any of their variations are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device comprising a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0027] To facilitate those skilled in the art to better understand the embodiments of the present application, some technical terms or nouns involved in the embodiments of the present application are explained as follows:

[0028] Fine-tuning is an optimization technique that aims to improve the performance of a large pre-trained model on a specific task by fine-tuning it using a small amount of sample data from the target domain. The goal of fine-tuning is to adapt the large model to the specific task and data distribution, thereby improving the model's performance.

[0029] Negative Feedback: In systems theory, negative feedback refers to the situation where the output of a system strengthens the input signal, causing the system to be more inclined to reduce a certain behavior or trend.

[0030] Self-learning: Generally, the self-learning of a large model means that it can autonomously learn and update its own knowledge structure and algorithm model while receiving external input data. The scope of self-learning here is extended to the fact that the data source also partially comes from its own model.

[0031] Self-learning methods in related technologies often face challenges with low data quality and poor model generalization. The training data generated by the model may contain noise or bias. If left uncontrolled, these issues can lead to reduced predictive accuracy during the learning process. Furthermore, the model may overfit to the self-generated data, resulting in poor performance on real-world data.

[0032] To address the above issues, the present application provides a related solution in the embodiments, proposing a self-learning large model fine-tuning scheme combined with negative feedback. First, the large model is trained using a small amount of external annotated data source. After training, the model regenerates the answer on the same dataset and compares the new answer with the original answer. By combining the loss function and the confidence index, the original data ratio is dynamically adjusted to optimize the model's learning process. The following is a detailed description.

[0033] According to an embodiment of the present application, an embodiment of a method for model training is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although a logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that shown here.

[0034] The method embodiments provided in the embodiments of the present application can be executed in a mobile terminal, a computer terminal or a similar computing device. Figure 1 The hardware structure block diagram of a computer terminal (or electronic device) for implementing a model training method or a model question answering method is shown. Figure 1 As shown, the computer terminal 10 (or electronic device) may include one or more (illustrated as 102a, 102b, ..., 102n in the figure) processors 102 (the processor 102 may include but is not limited to a processing device such as a microprocessor MCU or a programmable logic device FPGA), a memory 104 for storing data, and a transmission device 106 for communication functions. In addition, it may also include: a display, an input / output interface (I / O interface), a universal serial bus (USB) port (which may be included as one of the ports of the BUS bus), a network interface, a power supply and / or a camera. It will be understood by those skilled in the art that Figure 1 The structure shown is only for illustration and does not limit the structure of the above electronic device. Figure 1 More or fewer components than shown, or with Figure 1 Different configurations shown.

[0035] It should be noted that the one or more processors 102 and / or other data processing circuits described above may generally be referred to herein as "data processing circuitry". The data processing circuitry may be embodied in whole or in part as software, hardware, firmware, or any other combination thereof. In addition, the data processing circuitry may be a single independent processing module, or may be incorporated in whole or in part into any of the other components of the computer terminal 10 (or electronic device). As described in the embodiments of the present application, the data processing circuitry serves as a processor control (e.g., selection of a variable resistor terminal path connected to an interface).

[0036] The memory 104 can be used to store software programs and modules of application software, such as the program instructions / data storage device corresponding to the model training method or the model question-answering method in the embodiment of the present application. The processor 102 executes various functional applications and data processing by running the software programs and modules stored in the memory 104, that is, realizing the above-mentioned model training method or model question-answering method. The memory 104 may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory 104 may further include a memory remotely arranged relative to the processor 102, and these remote memories may be connected to the computer terminal 10 via a network. Examples of the above-mentioned networks include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and a combination thereof.

[0037] The transmission device 106 is configured to receive or transmit data via a network. A specific example of the aforementioned network may include a wireless network provided by the communications provider of the computer terminal 10. In one embodiment, the transmission device 106 includes a network interface controller (NIC), which can be connected to other network devices via a base station to enable communication with the Internet. In another embodiment, the transmission device 106 may be a radio frequency (RF) module, which is configured to communicate with the Internet wirelessly.

[0038] The display may be, for example, a touch screen liquid crystal display (LCD) that enables a user to interact with a user interface of the computer terminal 10 (or electronic device).

[0039] In the above operating environment, the embodiment of the present application provides a model training method. Figure 2 This is a schematic diagram of a model training method flow according to an embodiment of the present application. Figure 2 As shown, the method includes the following steps:

[0040] Step S202: Analyze the first question-and-answer sample data using the first model to generate second question-and-answer sample data of the same type as the first question-and-answer sample data, thereby obtaining an extended training dataset. The first model is obtained by training a basic model using the original training dataset, wherein the original training dataset contains manually annotated first question-and-answer sample data of multiple different types, each type corresponding to multiple second question-and-answer sample data.

[0041] Step S204: Obtain various types of second question and answer sample data in the extended training data set according to the data matching weights corresponding to each type of data in the extended training data set, and train the basic model based on the obtained second question and answer sample data to obtain a second model, wherein the data matching weights are used to indicate the collection ratio of each type of second question and answer sample data during model training;

[0042] Step S206: Determine the teachability score corresponding to the second model when processing the question text in the original training data set, and adjust the data matching weight based on the teachability score to obtain the target data matching weight. The teachability score is used to represent the learning value of each type of data for the basic model.

[0043] Step S208: training the basic model based on the original training data set, the expanded training data set, and the target data ratio weight to obtain the target model.

[0044] Through the above steps, by using negative feedback indicators to regulate and optimize the learning process of the model on self-generated data, the purpose of improving the performance and generalization ability of the model in real-world tasks is achieved, thereby solving the technical problem of poor application effect of the final trained model caused by the poor quality of training data generated by the self-learning model in related technologies.

[0045] The following further introduces the model training method in steps S202 to S208 of the embodiment of the present application.

[0046] In the embodiment of the present application, a negative feedback mechanism is used to optimize the learning process of the model on self-generated data. In this process, negative feedback enhances the performance of the model by encouraging the model to focus on learning data that can produce high-quality answers. At the same time, by identifying and reducing the impact of data that causes the performance of the model to decline, the model is prevented from overfitting and performance deterioration. By looping this process, the model can continue to learn on self-generated data, constantly adjust the data ratio, and improve its predictive ability in each iteration. Figure 3 This is a schematic diagram of a method flow for fine-tuning a self-learning large model based on data negative feedback provided in accordance with an embodiment of the present application. Figure 3 A detailed introduction to the process is given.

[0047] First, we use a small amount of manually annotated original training dataset D train (Take sft data source as an example), for the basic model M base Perform data fine-tuning training to obtain the first model M sft0 , as shown below:

[0048] M sft0 =Train(M base ,D train , 1)

[0049] Among them, 1 means that during this training, the data ratio weight corresponding to the first question and answer sample data of each type in the original training data set is 1, that is, all the first question and answer sample data of each type in the original training data set are used to fine-tune the basic model.

[0050] In order to improve the computational efficiency of subsequent generation of similar data and other processing steps and reduce resource requirements, in an embodiment of the present application, the data in the original training data set can be clustered and then sampled, and subsequent steps can be performed based on the sampled data, thereby reducing the amount of data required for subsequent processing, as follows.

[0051] In some embodiments of the present application, using the first model to generate second question and answer sample data of the same type as the first question and answer sample data includes the following steps: clustering the first question and answer sample data in the original training data set to obtain multiple cluster clusters, wherein each cluster cluster corresponds to one type; extracting a first preset number of first question and answer sample data from each cluster cluster to form a first training data set, wherein the total number of first question and answer sample data in the first training data set is less than the total number of first question and answer sample data in the original training data set; using the first question and answer sample data of each type in the first training data set as seed data; using the first model to perform diversity generation based on the seed data to obtain second question and answer sample data of the same type as the seed data, and using the second question and answer sample data to form an extended training data set, wherein, in the process of diversity generation, a second preset number of second question and answer sample data are generated based on each type of seed data.

[0052] For example, suppose the original training dataset D train There are 3000 first question and answer sample data in the network, so we can cluster them first. Assuming that after clustering, there are 100 types of data. For each type, we extract one first question and answer sample data. Then the first training data set D obtained after cluster sampling is train_sample There are 100 first question and answer sample data in total.

[0053] After that, the first model M can be used sft0, to generate the first training data set D train_sample (Original training dataset D train ) The same type of data (second question and answer sample data) is used to obtain the expanded training data set D gen0 , as shown in the following formula:

[0054] D gen0 =M sft0 (D train_sample )

[0055] Among them, D train_sample The data is used as seed data (if the cluster sampling step is not performed, the original training data set D can be directly used as the seed data). train As seed data), using the first model M sft0 Generate diversity based on seed data to obtain more data of the same type D gen0 (Each seed data generates x pieces of the same type of question-answer data, x>=3), for example, for D train_sample There are 100 first question and answer sample data of various types, and each type (i.e. each) generates 50 second question and answer sample data of the same type. The new data generated by each seed data is regarded as a small class. There are 100 small classes here, and each small class corresponds to a data matching weight.

[0056] After that, the extended training dataset D generated by the model can be used gen0 , and the data ratio weights corresponding to each data type to fine-tune the basic model. The specific steps are as follows.

[0057] In some embodiments of the present application, the basic model is trained based on the acquired second question and answer sample data to obtain the second model, including: determining the data matching weight corresponding to each type of data in the extended training data set, wherein each type corresponds to a data matching weight; for each type of second question and answer sample data, according to the collection ratio corresponding to the data matching weight of the type, the second question and answer sample data of the collection ratio is randomly obtained from all the second question and answer sample data of the type in the extended training data set to form a second training data set; based on the second training data set, the basic model is trained to obtain the second model.

[0058] Specifically, in this embodiment, the training data set D is expanded gen0 The corresponding initial data ratio weight is θ (w)0 It is represented by, where the data ratio weight corresponding to each type is W c0 ,(0 <W c0 <1, c corresponds to each subcategory), according to the data ratio weight θ (w)0 , get the extended training data set D gen0Various types of second question and answer sample data are obtained, and the basic model is trained based on the obtained second question and answer sample data. base , get the second model M sft1 , as shown below:

[0059] M sft1 =Train(M base ,D gen0 ,θ (w)0 )

[0060] For example, suppose W c0 Take 0.5, D gen0 The CCP uses 100 types of data, and each type corresponds to 50 second question and answer sample data. Then, according to the collection ratio corresponding to the data ratio weight, the second question and answer sample data is obtained from the extended training data set. The second training data set consists of 100*50*0.5=2500 pieces. These 2500 pieces are used to fine-tune the basic model to obtain the second model.

[0061] It should be noted that, as an optional implementation, the initial data allocation weight is θ (w)0 , or you can set W from the unified setting c0 Adjust to The data with high teachability scores are given a large weight for the corresponding sub-category data of the generated data to accelerate convergence.

[0062] Afterwards, the teachability score (STScore) needs to be calculated. This score reflects the learnable value of each type of second question and answer sample data, that is, the potential value that the basic model can learn from this type of data. In other words, for the model, the knowledge conversion ratio of each type of data. The following is an introduction to the calculation process of the teachability score.

[0063] In some embodiments of the present application, determining the teachability score corresponding to the second model when processing the question text in the original training data set includes: using the second model to analyze and process the question text in the first question and answer sample data of each type in the original training data set to obtain the answer text output by the second model; determining a loss value parameter and a confidence parameter based on the answer text, wherein the loss value parameter is used to characterize the degree of error between the answer text output by the second model and the answer text corresponding to the question text in the first question and answer sample data, and the confidence parameter is used to characterize the average confidence of each word in the answer text output by the second model; determining the teachability score corresponding to the type based on the loss value parameter and the confidence parameter, wherein each type of data in the original training data set corresponds to a teachability score.

[0064] Specifically, the calculation formula of the teachability score STScore can be shown as follows:

[0065] STScore=tanh(α*(loss-β*confidence))(-1 <tanh(x)<1)

[0066] Among them, Loss (i.e., loss parameter) is the average token loss value of the model answer (i.e., the answer text output by the second model) and the original answer (i.e., the answer text corresponding to the question text in the first question and answer sample data); Confidence (i.e., confidence parameter) is the average confidence of each token in the second model answer, and the value range is usually between 0 and 1; α is a positive scaling factor used to adjust the sensitivity of the score (because x in tanh(x) is near 0, the change of tanh(x) will be more obvious). β is a positive weight coefficient used to balance the influence of the loss function and the confidence parameter Confidence, and is generally greater than 1; tanh() is a hyperbolic tangent function, and its output range is (-1,1), which can limit the score to the required range.

[0067] In the above calculation formula, when Loss is small (that is, the model fitting effect is good), STScore will be close to 0, because the output of the tanh() function near 0 is also close to 0, which means that the model has fitted the current data very well and does not need to learn too much of this data; when Loss is large and Confidence is also large, Loss-β×ConfidenceLoss will be close to 0 or a negative value, resulting in STScore close to 0 or a negative value, reflecting that the gap between the high-confidence answers generated by the model and the original answers is large, and the model may not be able to learn, or the quality of the original answers is problematic, and the model does not need to learn these data; when Loss is large and Confidence is small, Loss-β×Confidence will be large, resulting in STScore close to 1, indicating that the model does not answer this type of question well and the answer stability is poor, and the model needs to learn these data.

[0068] For example, assume that α = loss, β = 10; in this case, if a D train Data M base The loss value parameter and confidence parameter of inference are: 0.5, 0.5, then STScore = tanh(0.5*(0.5-5)) = tanh(-0.225) (when Loss is small, close to 0); if a D train Data M baseThe loss value parameter and confidence parameter of inference are: 7, 0.8, then STScore = tanh(7*(7-8)) = tanh(-7) (Loss is large and Confidence is also large, close to 0 or negative); if a D train Data M base The loss value parameters and confidence parameters of inference are: 11 and 0.2 respectively, then STScore = tanh(11*10.8), (when Loss is large and Confidence is small, STScore is close to 1).

[0069] In this embodiment, the values ​​of α and β can be adjusted based on actual conditions to fine-tune the model's learning requirements. For example, if you want to more strongly indicate that the model does not need to learn when the confidence level is high, you can increase the value of β. Similarly, α can adjust the sensitivity of the score, making it approach 1 or -1 more quickly.

[0070] In addition, as an optional implementation, after sorting the STScore scores, you can also prioritize screening data with negative scores for manual inspection, because these data have high loss and high confidence in the model. They may be very difficult problems with unclear problem-solving processes, causing learning difficulties, or data quality issues.

[0071] In the embodiment of the present application, the basic model M can be calculated separately base , the first model M sft0 and the second model M sft1 In the original training data set D train The corresponding teachability score STScore score when processing the question text in M ​​is obtained. base _STScore_list、M sft0 _STScore_list and M sft1 _STScore_list. For example, M base _STScore_list=[0.5, 0.4,...], M sft0 _STScore_list=[0.6, 0.2],M sft1 _STScore_list=[0.4, 0.3,...].

[0072] After obtaining the teachability score, you can adjust the data ratio corresponding to each subcategory (i.e., each type) according to the teachability score. Through iterative optimization, you can finally get the target data ratio weight. The specific steps are as follows.

[0073] In some embodiments of the present application, the adjustment of the data matching weight includes multiple training batches; the data matching weight is adjusted according to the teachability score to obtain the target data matching weight, including the following steps: in each training batch, the first teachability score corresponding to the second model of the first training batch and the second teachability score corresponding to the second model of the second training batch are obtained, wherein the second training batch is the training batch immediately before the first training batch; the teachability score of each type between the first teachability score and the second teachability score is compared, and the data matching weight corresponding to each type is adjusted according to the size change; according to the adjusted data matching weight, the second question and answer sample data of various types in the expanded training data set are re-acquired, and the second question and answer sample data of various types in the expanded training data set are re-acquired according to the acquired data matching weight. The second question and answer sample data is used to train the basic model to obtain the second model corresponding to the next batch, and the above steps of adjusting the data matching weight according to the teachability score are iteratively repeated until the preset iteration termination condition is met, wherein the preset iteration termination condition includes at least one of the following: the change amplitude between the teachability scores corresponding to two consecutive training batches is less than the preset amplitude threshold, the number of times the sum of the teachable scores continues to increase exceeds the preset number threshold, and the loss value parameter corresponding to the second model when processing the question text in the original training data set is less than the loss value parameter corresponding to the first model when processing the question text in the original training data set; when the preset iteration termination condition is met, the data matching weight obtained after the last adjustment is determined as the target data matching weight.

[0074] In some embodiments of the present application, adjusting the data allocation weight corresponding to each type based on the size change includes the following steps: when the first teachability score corresponding to the type is greater than the second teachability score, increasing the data allocation weight corresponding to the type by a preset proportional step; when the first teachability score corresponding to the type is less than the second teachability score, reducing the data allocation weight corresponding to the type by a preset proportional step.

[0075] Specifically, the data allocation weight corresponding to each type of data can be adjusted by comparing the changes in the teachability score corresponding to the current training batch with the teachability score corresponding to the previous training batch. Alternatively, the data allocation weight corresponding to each type of data can be adjusted by comparing it with the teachability score corresponding to the basic model. For example, you can combine M base _STScore_list and M sft1 _STScore_list, to M base Study D gen0 Specifically, it is necessary to compare the D train The changes in the STScore scores of the data.

[0076] When the teachability score is in M base Less than M sft1 When , it means that the learning of self-generated data has played a positive role, then the weight of the self-generated data subcategory (C1) corresponding to the data increases, and the increase ratio is γ (that is, the above-mentioned preset proportional step. In this embodiment, γ is taken as 1 / 2 for example). It is shown in the following formula:

[0077]

[0078] When the teachability score is in M base Greater than M sft1 When , it means that the learning of self-generated data has played a negative role, then the weight of the self-generated data subcategory (C2) corresponding to this data decreases by a ratio of γ. As shown in the following formula:

[0079]

[0080] After adjusting the data ratio weight, you can base Study D gen0 The new ratio θ (w)1 , generate a new fine-tuning model M sft2 (i.e., the second model corresponding to the next batch), as shown below:

[0081] M sft2 =Train(M base ,D gen0 ,θ (w)1 )

[0082] By iteratively repeating the above process steps of calculating the teachability score, adjusting the data allocation weight, and training to obtain the second model corresponding to the next batch until the preset iteration termination condition is met, the target data allocation weight can be determined.

[0083] In this embodiment of the present application, the preset iteration termination conditions include but are not limited to:

[0084] 1) When the difference between the sum of the two consecutive teachable scores is not large (indicating stability), that is, the change between the teachable scores corresponding to two consecutive training batches is less than the preset amplitude threshold, satisfying r≈1;

[0085] 2) The sum of the teachability scores increases continuously for a number of times exceeding the preset threshold. For example, the sum of the teachability scores stops when it increases continuously for two times, i.e. sum(M sftx _STScore_list)>sum(M sft(x-1) _STScore_list, and sum(M sft(x-1) _STScore_list)>sum(Msft(x-2) _STScore_list;

[0086] 3) Use the fine-tuned model with its own data to infer D train The total loss under this condition is less than, using D train The fine-tuned model is used in inference D train The total loss under, i.e. sum(loss)_sft0 <sum(loss)_sftx。

[0087] The embodiment of the present application compares the original answer with the answer generated by the model, and dynamically adjusts the ratio of training data in combination with loss value parameters and confidence parameters, thereby improving data quality and promoting model accuracy; through a negative feedback mechanism, it ensures that the model maintains its original generalization ability and sensitivity to new data during the learning process, thereby improving the model's adaptability to unknown data.

[0088] After obtaining the final target data ratio weight, the basic model can be fine-tuned to obtain the target model. The specific steps are as follows.

[0089] In some embodiments of the present application, the basic model is trained based on the original training data set, the extended training data set, and the target data ratio weight, and obtaining the target model includes the following steps: obtaining various types of second question and answer sample data in the extended training data set according to the collection ratio corresponding to the target data ratio weight; combining the obtained second question and answer sample data and the first question and answer sample data in the original training data set to form a target training data set; and training the basic model based on the target training data set to obtain the target model.

[0090] Specifically, the final fine-tuning data is the original SFT data (i.e. the original training dataset D train ) and the expanded training dataset D gen0 , the target data ratio corresponding to each type is θ (w)x , assuming the target data ratio is Then, we can obtain the extended training data set D according to the acquisition ratio corresponding to the target data ratio. gen0 In order to make full use of the original training dataset D train The data in the data, its corresponding data ratio weight is better than 1, for example, it can be set to 4*θ (w)2 .

[0091] This application solution, by introducing the calculation of teachability scores, can intelligently identify which data has higher potential value for the model's learning process, thereby optimizing the data ratio and improving learning efficiency; through an adaptive fine-tuning strategy based on negative feedback of the effect indicator of self-generated data fine-tuning, it can dynamically adjust the ratio of training data according to the real-time performance of the model, so that the model can achieve continuous performance improvement with limited data resources. It improves the accuracy and generalization ability of the model; in addition, the self-learning mechanism can reduce the dependence on external labeled data and reduce the dependence on the cleaning of manually labeled data, so that the model can achieve continuous performance improvement with limited data resources. And when the data quality is average, it can divide the data through the effect of its own learning and accurately locate a small amount of key data that needs to be manually labeled. This application solution is of great value for optimizing large models under resource-constrained conditions, and is particularly suitable for fields such as natural language processing and computer vision, where the quality and diversity of data are crucial to model performance.

[0092] According to an embodiment of the present application, a model question-answering method is also provided. Figure 4 This is a schematic diagram of a model question-answering method flow according to an embodiment of the present application. Figure 4 As shown, the method includes the following steps:

[0093] Step S402: obtaining a question text, wherein the question text is a text containing the question to be solved by the user;

[0094] Step S404: Analyze the question text using the target model to obtain an answer text corresponding to the question text generated by the target model;

[0095] The training step of the target model includes: using a first model to analyze the first question and answer sample data, generating second question and answer sample data of the same type as the first question and answer sample data, and obtaining an extended training data set, wherein the first model is obtained by training a basic model using the original training data set, and the original training data set contains multiple manually annotated first question and answer sample data of different types, and each type corresponds to multiple second question and answer sample data; according to the data ratio weight corresponding to each type of data in the extended training data set, obtaining various types of second question and answer sample data in the extended training data set, and training the basic model based on the obtained second question and answer sample data to obtain a second model, wherein the data ratio weight is used to indicate the collection ratio of each type of second question and answer sample data during model training; determining the teachability score corresponding to the second model when processing the question text in the original training data set, and adjusting the data ratio weight based on the teachability score to obtain a target data ratio weight, wherein the teachability score is used to represent the degree of learning value of each type of data for the basic model; training the basic model based on the original training data set, the extended training data set, and the target data ratio weight to obtain a target model.

[0096] It should be noted that the model question answering method provided in this embodiment is Figure 2 The model training method shown corresponds to a method embodiment at the model application level. Therefore, the relevant explanations and descriptions of the above-mentioned model training method are also applicable to the embodiments of this application and will not be repeated here.

[0097] According to an embodiment of the present application, an embodiment of a model training device is also provided. Figure 5 This is a structural diagram of a model training device provided according to an embodiment of the present application. Figure 5 As shown, the device includes:

[0098] A data generation module 50 is configured to analyze the first question-and-answer sample data using the first model to generate second question-and-answer sample data of the same type as the first question-and-answer sample data, thereby obtaining an extended training data set, wherein the first model is obtained by training a basic model using the original training data set, wherein the original training data set includes manually annotated pieces of first question-and-answer sample data of different types, each type corresponding to a plurality of second question-and-answer sample data;

[0099] A first fine-tuning module 52 is configured to obtain various types of second question-and-answer sample data in the expanded training dataset according to the data allocation weight corresponding to each type of data in the expanded training dataset, and train the basic model based on the obtained second question-and-answer sample data to obtain a second model, wherein the data allocation weight indicates the collection ratio of each type of second question-and-answer sample data during model training;

[0100] A matching adjustment module 54 is configured to determine a teachability score corresponding to the second model when processing the question text in the original training data set, and to adjust the data matching weight based on the teachability score to obtain a target data matching weight, wherein the teachability score is used to represent the degree of learning value of each type of data for the basic model;

[0101] The second fine-tuning module 56 is used to train the basic model based on the original training data set, the expanded training data set, and the target data ratio weight to obtain the target model.

[0102] Optionally, using the first model to generate second question and answer sample data of the same type as the first question and answer sample data includes: clustering the first question and answer sample data in the original training data set to obtain multiple clusters, wherein each cluster corresponds to a type; extracting a first preset number of first question and answer sample data from each cluster to form a first training data set, wherein the total number of first question and answer sample data in the first training data set is less than the total number of first question and answer sample data in the original training data set; using the first question and answer sample data of each type in the first training data set as seed data; using the first model to generate diversity based on the seed data to obtain second question and answer sample data of the same type as the seed data, and using the second question and answer sample data to form an extended training data set, wherein, in the process of diversity generation, a second preset number of second question and answer sample data are generated based on each type of seed data.

[0103] Optionally, the basic model is trained based on the acquired second question and answer sample data to obtain the second model, including: determining the data ratio weight corresponding to each type of data in the extended training data set, wherein each type corresponds to a data ratio weight; for each type of second question and answer sample data, according to the collection ratio corresponding to the data ratio weight of the type, the second question and answer sample data of the collection ratio is randomly obtained from all the second question and answer sample data of the type in the extended training data set to form a second training data set; based on the second training data set, the basic model is trained to obtain the second model.

[0104] Optionally, determining the teachability score corresponding to the question text in the original training data set when the second model processes the question text includes: using the second model to analyze and process the question text in the first question and answer sample data of each type in the original training data set to obtain the answer text output by the second model; determining a loss value parameter and a confidence parameter based on the answer text, wherein the loss value parameter is used to characterize the degree of error between the answer text output by the second model and the answer text corresponding to the question text in the first question and answer sample data, and the confidence parameter is used to characterize the average confidence of each word in the answer text output by the second model; determining the teachability score corresponding to the type based on the loss value parameter and the confidence parameter, wherein each type of data in the original training data set corresponds to a teachability score.

[0105] Optionally, the adjustment of the data matching weight includes multiple training batches; the data matching weight is adjusted according to the teachability score to obtain the target data matching weight, including: in each training batch, obtaining the first teachability score corresponding to the second model of the first training batch, and the second teachability score corresponding to the second model of the second training batch, wherein the second training batch is the training batch immediately before the first training batch; comparing the size changes of each type of teachability score between the first teachability score and the second teachability score, and adjusting the data matching weight corresponding to each type according to the size changes; according to the adjusted data matching weight, re-obtaining the second question and answer sample data of various types in the expanded training data set, and re-obtaining the second question and answer sample data according to the obtained second question and answer sample data. Data, train the basic model to obtain the second model corresponding to the next batch, and iteratively repeat the above steps of adjusting the data matching weight according to the teachability score until the preset iteration termination condition is met, wherein the preset iteration termination condition includes at least one of the following: the change amplitude between the teachability scores corresponding to two consecutive training batches is less than the preset amplitude threshold, the number of times the sum of the teachable scores continues to increase exceeds the preset number threshold, and the loss value parameter corresponding to the second model when processing the question text in the original training data set is less than the loss value parameter corresponding to the first model when processing the question text in the original training data set; when the preset iteration termination condition is met, the data matching weight obtained after the last adjustment is determined as the target data matching weight.

[0106] Optionally, based on the size change, the data allocation weight corresponding to each type is adjusted, including: when the first teachability score corresponding to the type is greater than the second teachability score, the data allocation weight corresponding to the type is increased by a preset proportional step; when the first teachability score corresponding to the type is less than the second teachability score, the data allocation weight corresponding to the type is reduced by a preset proportional step.

[0107] Optionally, the basic model is trained based on the original training data set, the extended training data set, and the target data ratio weight to obtain the target model, including: obtaining various types of second question and answer sample data in the extended training data set according to the collection ratio corresponding to the target data ratio weight; the obtained second question and answer sample data and the first question and answer sample data in the original training data set are combined to form a target training data set; based on the target training data set, the basic model is trained to obtain the target model.

[0108] It should be noted that the various modules in the above-mentioned model training device can be program modules (for example, a set of program instructions that implement a certain specific function) or hardware modules. For the latter, it can be expressed in the following forms, but is not limited to this: the expression form of each of the above-mentioned modules is a processor, or the functions of each of the above-mentioned modules are implemented by a processor.

[0109] It should be noted that the model training device provided in this embodiment can be used to perform Figure 2 The model training method shown, therefore, the relevant explanations of the above model training method are also applicable to the embodiments of this application and will not be repeated here.

[0110] The embodiment of the present application also provides a non-volatile storage medium, the non-volatile storage medium includes a stored computer program, wherein the device where the non-volatile storage medium is located executes the following model training method or model question-answering method by running the computer program: using a first model to analyze the first question-answering sample data, generating second question-answering sample data of the same type as the first question-answering sample data, and obtaining an extended training data set, wherein the first model is obtained by training the basic model using the original training data set, and the original training data set contains multiple manually labeled first question-answering sample data of different types, and each type corresponds to multiple second question-answering sample data; according to the data ratio weight corresponding to each type of data in the extended training data set, obtain Various types of second question and answer sample data in the expanded training data set are used, and the basic model is trained based on the acquired second question and answer sample data to obtain a second model, wherein the data matching weight is used to indicate the collection ratio of each type of second question and answer sample data when training the model; the teachability score corresponding to the second model when processing the question text in the original training data set is determined, and the data matching weight is adjusted based on the teachability score to obtain the target data matching weight, wherein the teachability score is used to characterize the degree of learning value of each type of data for the basic model; the basic model is trained based on the original training data set, the expanded training data set, and the target data matching weight to obtain the target model.

[0111] Alternatively, a question text is obtained, wherein the question text is a text containing a question to be solved by the user; a target model is used to analyze the question text to obtain an answer text corresponding to the question text generated by the target model; wherein the training step of the target model includes: using a first model to analyze the first question and answer sample data, generating second question and answer sample data of the same type as the first question and answer sample data, and obtaining an extended training data set, wherein the first model is obtained by training the basic model using the original training data set, and the original training data set contains a plurality of manually annotated first question and answer sample data of different types, and each type corresponds to a plurality of second question and answer sample data; according to the data ratio weight corresponding to each type of data in the extended training data set, Obtain various types of second question and answer sample data in the extended training data set, and train the basic model based on the obtained second question and answer sample data to obtain a second model, wherein the data matching weight is used to indicate the collection ratio of each type of second question and answer sample data when training the model; determine the teachability score corresponding to the second model when processing the question text in the original training data set, and adjust the data matching weight based on the teachability score to obtain the target data matching weight, wherein the teachability score is used to characterize the degree of learning value of each type of data for the basic model; train the basic model based on the original training data set, the extended training data set, and the target data matching weight to obtain the target model.

[0112] The embodiment of the present application also provides a computer program product, including a computer program, which, when executed by a processor, implements the steps of the model training method or model question-answering method described in each embodiment of the present application: using a first model to analyze the first question-answer sample data, generating second question-answer sample data of the same type as the first question-answer sample data, and obtaining an extended training data set, wherein the first model is obtained by training the basic model using the original training data set, and the original training data set contains multiple manually labeled first question-answer sample data of different types, and each type corresponds to multiple second question-answer sample data; according to the data ratio weight corresponding to each type of data in the extended training data set, the extended training data set is obtained. Various types of second question and answer sample data are gathered, and the basic model is trained based on the acquired second question and answer sample data to obtain a second model, wherein the data matching weight is used to indicate the collection ratio of each type of second question and answer sample data when training the model; the teachability score corresponding to the second model when processing the question text in the original training data set is determined, and the data matching weight is adjusted based on the teachability score to obtain the target data matching weight, wherein the teachability score is used to characterize the degree of learning value of each type of data for the basic model; the basic model is trained based on the original training data set, the expanded training data set, and the target data matching weight to obtain the target model.

[0113] Alternatively, a question text is obtained, wherein the question text is a text containing a question to be solved by the user; a target model is used to analyze the question text to obtain an answer text corresponding to the question text generated by the target model; wherein the training step of the target model includes: using a first model to analyze the first question and answer sample data, generating second question and answer sample data of the same type as the first question and answer sample data, and obtaining an extended training data set, wherein the first model is obtained by training the basic model using the original training data set, and the original training data set contains a plurality of manually annotated first question and answer sample data of different types, and each type corresponds to a plurality of second question and answer sample data; according to the data ratio weight corresponding to each type of data in the extended training data set, Obtain various types of second question and answer sample data in the extended training data set, and train the basic model based on the obtained second question and answer sample data to obtain a second model, wherein the data matching weight is used to indicate the collection ratio of each type of second question and answer sample data when training the model; determine the teachability score corresponding to the second model when processing the question text in the original training data set, and adjust the data matching weight based on the teachability score to obtain the target data matching weight, wherein the teachability score is used to characterize the degree of learning value of each type of data for the basic model; train the basic model based on the original training data set, the extended training data set, and the target data matching weight to obtain the target model.

[0114] The serial numbers of the above embodiments of the present application are for description only and do not represent the advantages or disadvantages of the embodiments.

[0115] In the above embodiments of the present application, the description of each embodiment has its own focus. For parts that are not described in detail in a certain embodiment, please refer to the relevant description of other embodiments.

[0116] In the several embodiments provided in this application, it should be understood that the disclosed technical content can be implemented in other ways. Among them, the device embodiments described above are only exemplary. For example, the division of the units can be a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of units or modules, which can be electrical or other forms.

[0117] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple units. Some or all of the units may be selected according to actual needs to achieve the purpose of the present embodiment.

[0118] In addition, the functional units in the various embodiments of the present application may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.

[0119] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions for enabling a computer device (which can be a personal computer, a server or a network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: various media that can store program codes, such as a USB flash drive, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk or an optical disk.

[0120] The above is only a preferred embodiment of the present application. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present application. These improvements and modifications should also be regarded as the scope of protection of the present application.

Claims

1. A model training method, characterized in that: include: Using a first model, analyzing the first question and answer sample data, generating second question and answer sample data of the same type as the first question and answer sample data, and obtaining an extended training data set, wherein the first model is obtained by training a basic model using an original training data set, and the original training data set contains a plurality of manually annotated first question and answer sample data of different types, and each type corresponds to a plurality of second question and answer sample data; Obtaining the second question-and-answer sample data of various types in the extended training data set according to the data matching weight corresponding to each type of data in the extended training data set, and training the basic model based on the obtained second question-and-answer sample data to obtain a second model, wherein the data matching weight is used to indicate the collection ratio of each type of the second question-and-answer sample data when performing model training; determining a teachability score corresponding to the second model when processing the question text in the original training data set, and adjusting the data matching weight according to the teachability score to obtain a target data matching weight, wherein the teachability score is used to represent the degree of learning value of each type of data for the basic model; The basic model is trained based on the original training data set, the expanded training data set, and the target data ratio weight to obtain a target model.

2. The model training method according to claim 1, characterized in that Using the first model to generate second question and answer sample data of the same type as the first question and answer sample data includes: Clustering the first question-and-answer sample data in the original training data set to obtain a plurality of clusters, wherein each cluster corresponds to one of the types; Extracting a first preset number of the first question and answer sample data from each of the clusters to form a first training data set, wherein the total number of the first question and answer sample data in the first training data set is less than the total number of the first question and answer sample data in the original training data set; Using the first question-and-answer sample data of each type in the first training data set as seed data; The first model is used to generate diversity based on the seed data to obtain the second question and answer sample data of the same type as the seed data, and the second question and answer sample data is used to form the extended training data set, wherein, in the process of diversity generation, a second preset number of second question and answer sample data are generated based on each type of seed data.

3. The model training method according to claim 1, characterized in that The basic model is trained based on the acquired second question and answer sample data to obtain a second model including: Determine the data matching weight corresponding to each type of data in the expanded training data set, wherein each type corresponds to one data matching weight; For each type of the second question-and-answer sample data, randomly obtain the second question-and-answer sample data of the collection ratio corresponding to the data ratio weight of the type from all the second question-and-answer sample data of the type in the expanded training data set to form a second training data set; The basic model is trained based on the second training data set to obtain the second model.

4. The model training method according to claim 1, characterized in that Determining a teachability score corresponding to the second model when processing the question text in the original training dataset includes: Using the second model, analyzing and processing the question text in the first question-and-answer sample data of each type in the original training data set to obtain the answer text output by the second model; Determining a loss value parameter and a confidence parameter based on the answer text, wherein the loss value parameter is used to represent the degree of error between the answer text output by the second model and the answer text corresponding to the question text in the first question and answer sample data, and the confidence parameter is used to represent the average confidence of each word in the answer text output by the second model; The teachability score corresponding to the type is determined based on the loss value parameter and the confidence parameter, wherein each type of data in the original training data set corresponds to one teachability score.

5. The model training method according to claim 4, characterized in that The adjustment of the data allocation weight includes multiple training batches; According to the teachability score, the data matching weight is adjusted to obtain the target data matching weight including: In each training batch, obtaining a first teachability score corresponding to the second model of the first training batch and a second teachability score corresponding to the second model of the second training batch, wherein the second training batch is a training batch immediately preceding the first training batch; comparing changes in the teachability scores of each type between the first teachability score and the second teachability score, and adjusting the data allocation weight corresponding to each type according to the changes in the scores; According to the adjusted data matching weights, the second question and answer sample data of various types in the expanded training data set are reacquired, and the basic model is trained based on the acquired second question and answer sample data to obtain the second model corresponding to the next batch, and the above step of adjusting the data matching weights based on the teachability score is iteratively repeated until a preset iteration termination condition is met, wherein the preset iteration termination condition includes at least one of the following: the amplitude of the change between the teachability scores corresponding to two consecutive training batches is less than a preset amplitude threshold, the number of times the sum of the teachability scores continues to increase exceeds a preset number threshold, and the loss value parameter corresponding to the second model when processing the question text in the original training data set is less than the loss value parameter corresponding to the first model when processing the question text in the original training data set; When the preset iteration termination condition is met, the data allocation weight obtained after the last adjustment is determined as the target data allocation weight.

6. The model training method according to claim 5, characterized in that Adjusting the data allocation weights corresponding to each type according to the size change includes: When the first teachability score corresponding to the type is greater than the second teachability score, increasing the data matching weight corresponding to the type by a preset proportional step; When the first teachability score corresponding to the type is less than the second teachability score, the data allocation weight corresponding to the type is reduced by the preset proportional step.

7. The model training method according to claim 1, characterized in that The basic model is trained based on the original training data set, the expanded training data set, and the target data ratio weight to obtain a target model including: Acquire the second question-and-answer sample data of various types in the expanded training data set according to the collection ratio corresponding to the target data ratio weight; The obtained second question-and-answer sample data and the first question-and-answer sample data in the original training data set are combined to form a target training data set; The basic model is trained according to the target training data set to obtain the target model.

8. A model question answering method, characterized in that: include: Obtaining a question text, wherein the question text is a text containing a question to be solved by the user; Analyzing the question text using a target model to obtain an answer text corresponding to the question text generated by the target model; The training step of the target model includes: using the first model to analyze the first question and answer sample data, generating the second question and answer sample data of the same type as the first question and answer sample data, and obtaining an extended training data set, wherein the first model is obtained by training the basic model using the original training data set, and the original training data set contains a plurality of manually labeled different types of the first question and answer sample data, and each type corresponds to a plurality of second question and answer sample data; according to the data ratio weight corresponding to each type of data in the extended training data set, the second question and answer sample data of various types in the extended training data set are obtained, and based on the obtained number of the second question and answer samples According to, the basic model is trained to obtain a second model, wherein the data matching weight is used to indicate the collection ratio of each type of the second question and answer sample data when performing model training; the teachability score corresponding to the second model when processing the question text in the original training data set is determined, and according to the teachability score, the data matching weight is adjusted to obtain the target data matching weight, wherein the teachability score is used to characterize the degree of learning value of each type of data for the basic model; according to the original training data set, the expanded training data set, and the target data matching weight, the basic model is trained to obtain the target model.

9. A model training device, characterized in that: include: a data generation module, configured to analyze the first question-and-answer sample data using a first model to generate second question-and-answer sample data of the same type as the first question-and-answer sample data, thereby obtaining an extended training data set, wherein the first model is obtained by training a basic model using an original training data set, and the original training data set includes a plurality of manually annotated pieces of the first question-and-answer sample data of different types, with each piece of the second question-and-answer sample data corresponding to a plurality of pieces of the second question-and-answer sample data; a first fine-tuning module, configured to obtain the second question-and-answer sample data of various types in the extended training dataset according to the data allocation weight corresponding to each type of data in the extended training dataset, and train the basic model based on the obtained second question-and-answer sample data to obtain a second model, wherein the data allocation weight is used to indicate the collection ratio of each type of the second question-and-answer sample data during model training; a matching adjustment module, configured to determine a teachability score corresponding to the second model when processing the question text in the original training data set, and adjust the data matching weight according to the teachability score to obtain a target data matching weight, wherein the teachability score is used to represent the degree of learning value of each type of data for the basic model; The second fine-tuning module is used to train the basic model based on the original training data set, the expanded training data set, and the target data ratio weight to obtain a target model.

10. An electronic device, characterized in that: include: A memory and a processor, wherein the processor is used to run a program stored in the memory, wherein when the program is run, the model training method described in any one of claims 1 to 7 or the model question-answering method described in claim 8 is executed.

11. A non-volatile storage medium, characterized in that: The non-volatile storage medium includes a stored computer program, wherein the device where the non-volatile storage medium is located executes the model training method described in any one of claims 1 to 7 or the model question-answering method described in claim 8 by running the computer program.

12. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the model training method described in any one of claims 1 to 7 or the model question-answering method described in claim 8 are implemented.

Citation Information

Patent Citations

  • Sample data acquisition method and device, sample data model acquisition method and device, equipment and storage medium

    CN116401382A

  • Text generation method and device, model training method and device, electronic equipment and medium

    CN117453904A

  • Data processing method for image enhancement model, electronic equipment and storage medium

    CN117746121A

  • Training method of few-sample relation classification model and related equipment

    CN119398195A

  • Dynamic training of Models

    US20240029413A1