Target question and answer model training method and device, equipment, medium and program product
By augmenting and desensitizing bank financial data, a training data set is generated, and a general big model is fine-tuned and trained to obtain a target question-and-answer model, which solves the problems of low security and low efficiency in the fine-tuning of bank financial data, and achieves the effect of improving data security and fine-tuning efficiency.
Patent Information
- Application Number
- CN202510062132.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-15
- Publication Date
- 2025-05-09
Smart Images

Figure CN119961677A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer technology, and in particular to a method, device, equipment, medium and program product for training a target question-answering model. Background Art
[0002] As the digitalization process of the financial industry continues to accelerate, massive financial data is showing explosive growth. Among them, bank financial data involves key business areas such as customer information, transaction records, and risk assessment. Its value mining and analysis are of vital importance to financial institutions' decision-making, risk management, and customer service optimization.
[0003] In recent years, deep learning technology has made significant progress in the field of natural language processing. Large language models (LLMs) have shown certain advantages in general text information extraction tasks due to their powerful language understanding and generation capabilities. However, when general large models are actually applied to the banking and financial fields, due to the wide range of sources, diverse formats, and complex content of banking and financial data, in order to enable general large models to accurately identify key financial information, the practice of fine-tuning general large models using banking and financial data has gradually emerged.
[0004] When using bank financial data to fine-tune the general large model, due to the high sensitivity of financial data and the many security loopholes in each link of the general large model fine-tuning, it is easy to cause data leakage, thereby causing risks such as damage to customer privacy. Summary of the invention
[0005] The present application provides a target question-answering model training method, device, equipment, medium and program product to solve the technical problems of low security and low efficiency in the process of fine-tuning the existing general large model for bank financial data.
[0006] In a first aspect, the present application provides a target question answering model training method, comprising:
[0007] Acquire an initial data set generated based on data in a target domain, perform data enhancement processing on the initial data set, and generate a derived data set;
[0008] Acquire a pre-generated sensitive data set, perform desensitization processing on the sensitive data set, and obtain a fuzzy data set corresponding to the sensitive data set; wherein the sensitive data set is a data set containing preset sensitive information in the target field; the desensitization processing includes reverse expression processing;
[0009] A training data set is obtained according to the initial data set, the derived data set and the fuzzy data set, and at least one round of fine-tuning training is performed on the pre-trained general large model according to the training data set to obtain a target question-answering model; wherein the target question-answering model is used to perform question-answering tasks within the target domain; the input of the target question-answering model is multimodal data, and the multimodal data includes at least one of text, image and video.
[0010] In an optional implementation, the initial data set includes multiple sets of question-answer pairs;
[0011] Performing data enhancement processing on the initial data set to generate a derived data set includes:
[0012] According to the preset enhancement instructions and enhancement model, data enhancement processing is performed on the question data in each question-answer pair to obtain at least one derivative question corresponding to each question data;
[0013] Determine each of the derived questions and its corresponding derived answer to obtain the derived data set.
[0014] In an optional implementation manner, obtaining at least one derived question corresponding to each of the question data includes:
[0015] Inputting the enhancement instruction and each of the question data into the enhancement model to obtain at least one initial derivative question corresponding to each of the question data;
[0016] In a preset optimization model, according to preset optimization instructions, problem evaluation results corresponding to each of the initial derivative problems are determined;
[0017] The initial derived questions whose question evaluation results are less than a preset threshold are optimized to obtain at least one derived question corresponding to each of the question data.
[0018] In an optional implementation, desensitizing the sensitive data set to obtain a fuzzy data set corresponding to the sensitive data set includes:
[0019] Determining that the sensitive data set includes a plurality of sensitive question-answer pairs;
[0020] According to the preset desensitization instruction, the sensitive answers in each of the sensitive question and answer pairs are desensitized to obtain at least one fuzzy answer corresponding to each of the sensitive answers;
[0021] A fuzzy data set is generated according to each of the fuzzy answers and the corresponding sensitive questions.
[0022] In an optional implementation, at least one round of fine-tuning training is performed on the pre-trained general large model according to the training data set to obtain a target question-answering model, including:
[0023] For any round of fine-tuning training, the test problem data in the training data set is input into the general large model after the current round of fine-tuning to obtain the test prediction data corresponding to the test problem data;
[0024] Determine a fine-tuning result corresponding to the current round of fine-tuning according to the data similarity between the test prediction data and the test answer data corresponding to the test question data;
[0025] If the fine-tuning result corresponding to the current round of fine-tuning is better than the fine-tuning result corresponding to the previous round of fine-tuning, continue to perform the next round of fine-tuning on the general large model;
[0026] If the fine-tuning result corresponding to the current round of fine-tuning is worse than the fine-tuning result corresponding to the previous round of fine-tuning, the general large model corresponding to the previous round of fine-tuning is used as the target question-answering model.
[0027] In an optional implementation, determining a fine-tuning result corresponding to the current round of fine-tuning according to the data similarity between the test prediction data and the test answer data corresponding to the test question data includes:
[0028] Obtaining prediction keywords in the test prediction data and answer keywords in the test answer data;
[0029] Determining a first fine-tuning result corresponding to the current round of fine-tuning according to the similarity between the predicted keyword and the answer keyword;
[0030] When the first fine-tuning result satisfies a preset result condition, obtaining a prediction vector corresponding to the test prediction data and an answer vector corresponding to the test answer data;
[0031] Determining a second fine-tuning result corresponding to the current round of fine-tuning according to the similarity between the prediction vector and the answer vector;
[0032] A fine-tuning result corresponding to the current round of fine-tuning is determined according to the first fine-tuning result, or the first fine-tuning result and the second fine-tuning result.
[0033] In a second aspect, the present application provides a target question answering model training device, comprising:
[0034] A derived data set generation module is used to obtain an initial data set generated based on data in the target domain, perform data enhancement processing on the initial data set, and generate a derived data set;
[0035] A fuzzy data set generation module is configured to obtain a pre-generated sensitive data set, perform desensitization processing on the sensitive data set, and obtain a fuzzy data set corresponding to the sensitive data set; wherein the sensitive data set is a data set containing preset sensitive information in the target domain; and the desensitization processing includes reverse expression processing;
[0036] A model fine-tuning module is used to obtain a training data set based on the initial data set, the derived data set and the fuzzy data set, and to perform at least one round of fine-tuning training on the pre-trained general large model based on the training data set to obtain a target question-answering model; wherein the target question-answering model is used to perform question-answering tasks within the target domain; the input of the target question-answering model is multimodal data, and the multimodal data includes at least one of text, image and video.
[0037] In a third aspect, the present application provides an electronic device, comprising: a processor, and a memory communicatively connected to the processor;
[0038] The memory stores computer-executable instructions;
[0039] The processor executes the computer-executable instructions stored in the memory to implement the method according to the first aspect.
[0040] In a fourth aspect, the present application provides a computer-readable storage medium, wherein the computer-readable storage medium stores computer-executable instructions, and when the computer-executable instructions are executed by a processor, they are used to implement the method described in the first aspect.
[0041] In a fifth aspect, the present application provides a computer program product, including a computer program, which implements the method described in the first aspect when executed by a processor.
[0042] The target question-answering model training technology provided in the present application obtains an initial data set by performing a small amount of annotation on the data in the target field, and obtains a derived data set by performing data enhancement on the training data set. In this way, a large amount of training data can be obtained by annotating a small amount of data to reduce the annotation time, thereby achieving the effect of improving the overall efficiency of fine-tuning the general large model; further, a predetermined sensitive data set is obtained, and the sensitive data set is reversely represented, and the fuzzy data set obtained after the reverse representation processing is used together with the above-mentioned initial data set and derivative data set as the final training data, and the general large model is fine-tuned to obtain the target question-answering model. In this way, since the sensitive data in the training data contains two expressions, the general large model considers the sensitive data to be unstable data during the learning process because the sensitive data contains completely different description forms. That is, the sensitive data will not be learned during the learning process. Therefore, the target question-answering model obtained after fine-tuning will not output the sensitive data as an answer when receiving the question input corresponding to the sensitive data. Therefore, data security can be improved. On the other hand, even if the training data is accidentally leaked during the fine-tuning process, due to the multiple expressions of sensitive data, it is difficult for the outside to reversely restore the real sensitive data content, thereby maintaining the privacy security of customers and other relevant subjects to the greatest extent, and ensuring the security of banking and financial data in the training process. BRIEF DESCRIPTION OF THE DRAWINGS
[0043] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present application and, together with the description, serve to explain the principles of the present application.
[0044] Figure 1 An application scenario diagram of the target question-answering model provided in this application;
[0045] Figure 2 A flowchart of a target question-answering model training method provided in an embodiment of the present application;
[0046] Figure 3 A schematic diagram of the structure of a target question-answering model training device provided in an embodiment of the present application;
[0047] Figure 4 It is a block diagram of an electronic device shown in an embodiment of the present application.
[0048] The above drawings have shown clear embodiments of the present application, which will be described in more detail later. These drawings and text descriptions are not intended to limit the scope of the present application in any way, but to illustrate the concept of the present application to those skilled in the art by referring to specific embodiments. DETAILED DESCRIPTION
[0049] Exemplary embodiments will be described in detail herein, examples of which are shown in the accompanying drawings. When the following description refers to the drawings, the same numbers in different drawings represent the same or similar elements unless otherwise indicated. The implementations described in the following exemplary embodiments do not represent all implementations consistent with the present application. Instead, they are merely examples of devices and methods consistent with some aspects of the present application as detailed in the appended claims.
[0050] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with relevant laws, regulations and standards, and provide corresponding operation entrances for users to choose to authorize or refuse.
[0051] The big model, whose full name is Large Language Model (LLM), is based on a deep learning architecture. It ingests massive amounts of text data for training and builds a huge parameter system, which enables it to have super language understanding, text generation, and knowledge reasoning capabilities. It can be widely and deeply applied to natural language processing, intelligent customer service, content creation, assisted translation and many other fields.
[0052] Instructed evolution: In the field of machine learning and natural language processing, instructional evolution of training set problems usually means modifying or extending the original problem to generate new variants. These variants can be used to enhance the training dataset, thereby improving the generalization ability and robustness of the model.
[0053] Large model fine-tuning: By fine-tuning the general large model trained on a large-scale general dataset on a smaller, domain-specific dataset, the performance of the model on a specific task can be optimized.
[0054] Specifically, when using bank financial data to fine-tune the large language model, since financial data itself is highly sensitive, it includes users' privacy data and commercial secrets of financial institutions. Once leaked, the consequences will be disastrous. In addition, there are many potential security loopholes in each link involved in the fine-tuning of general large models, such as data collection, data transmission, storage and even the use stage, which leads to the technical problem of low data security in the fine-tuning training process of general large models.
[0055] Based on this, some existing fine-tuning methods can remove all sensitive information in the training data before fine-tuning the general large model, or add random noise to the sensitive information, so that the sensitive information cannot be distinguished from normal information, and then fine-tune the training based on the processed data. However, the above sensitive information removal method requires finding the data containing sensitive information in a large amount of training data in advance, and then processing the data; since the amount of training data is in the tens of thousands or even hundreds of thousands, the process of data search will lead to the technical problem of reduced overall efficiency in the fine-tuning process of the general large model.
[0056] The target question-answering model training method provided in the present application is intended to solve the above technical problems of the prior art. Specifically, by labeling a small amount of training data set and performing data enhancement on the training data set in the form of instruction evolution to obtain a derivative data set, in this way, a large amount of training data can be obtained by labeling a small amount of data to reduce the labeling time, thereby achieving the effect of improving the overall efficiency of fine-tuning the general large model; further, obtaining the pre-collected sensitive data, and performing reverse expression processing on the sensitive data, and using the reverse expression processed data and the above training data as the final training data to fine-tune the general large model. Since the sensitive data contained in the training data corresponds to at least two expressions, namely affirmative expressions and negative expressions, the general large model will consider the sensitive data to be unstable data when learning the expression of sensitive data during the learning process, that is, the sensitive data will not be learned during the learning process. Therefore, the target question-answering model obtained after fine-tuning will not output the sensitive data as an answer when receiving the question input corresponding to the sensitive data. Therefore, data security can be improved. On the other hand, even if the training data is accidentally leaked during the fine-tuning process, due to the various expressions of sensitive data, it is difficult for the outside to reversely restore the real sensitive data content, thereby maintaining the privacy security of customers and other relevant subjects to the greatest extent, and ensuring the security of data in the banking and financial fields during the training process.
[0057] In summary, the above-mentioned process of processing the training data and fine-tuning the general large model using the processed training data can improve the efficiency and data security of the general large model fine-tuning process, and improve the data security of the fine-tuned target question-answering model during the task execution process.
[0058] Figure 1An application scenario diagram of the target question-answering model provided in this application. The target question-answering model provided in this application can be applied to application scenarios in which target tasks such as information extraction are performed on bank financial data. Since bank financial data involves sensitive information such as customer information and transaction records, in order to further improve data security, the target question-answering model trained by the target question-answering model training method provided in this application can be used to perform target tasks such as data question-answering on bank financial data, thereby providing safe and reliable data services, and providing solid and reliable technical support for the bank's stable operation and the protection of customer rights.
[0059] For ease of understanding, the following Figure 1 The application scenarios to which the embodiments of the present application are applicable are described. Figure 1 The technical solution provided in this application involves two stages, namely, the target question and answer model training stage and the target question and answer model application stage.
[0060] Specifically, in the target question-answering model training stage, a small amount of labeled data in the banking and financial field is obtained, and the labeled data is enhanced through data enhancement strategies and data optimization strategies to obtain a large amount of derivative data. It should be noted that because there are sensitive data containing more and more scattered sensitive information in the labeled data and the derivative data, and the above-mentioned search for the above-mentioned sensitive data takes a long time, in order to avoid the leakage of sensitive information and ensure the fine-tuning efficiency of the general large model, before using the above-mentioned labeled data and derivative data to fine-tune the general large model, the sensitive data containing sensitive information collected in the banking and financial field in advance can be desensitized in advance, that is, reverse expression processing, and the fuzzy data obtained after the desensitization processing as well as the above-mentioned labeled data and derivative data are used as training data, and the pre-trained general large model is fine-tuned to obtain the target question-answering model.
[0061] Specifically, in the target question-answering model application stage, when receiving the target question input by the user, the target question-answering model searches for the answer corresponding to the target question in the target field and outputs it. Since the sensitive data in the training data contains two expressions, the general large model considers the sensitive data to be unstable data during the learning process because the sensitive data contains completely different description forms, that is, it will not learn the sensitive data during the learning process. Therefore, when the target question-answering model obtained after fine-tuning receives the question input corresponding to the sensitive data, it will not output the sensitive data as the answer, so that data security can be improved.
[0062] It should be understood that the target question and answer model provided in this application can be applied to various scenarios such as the medical field, the legal field, and the scientific research field in addition to banking and financial scenarios. The embodiments of this application do not specifically limit the application scenarios of the target question and answer model provided.
[0063] The technical solution of the present application and how the technical solution of the present application solves the above-mentioned technical problems are described in detail below with specific embodiments. The following specific embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments. The embodiments of the present application will be described below in conjunction with the accompanying drawings.
[0064] Figure 2 A flowchart of a method for training a target question answering model provided in an embodiment of the present application. The method can be executed by a target question answering model training device, which can be a server or an electronic device. The following is an example of an electronic device. The method in this embodiment can be implemented by software, hardware, or a combination of software and hardware, such as Figure 2 As shown, the method includes the following steps.
[0065] S201, obtaining an initial data set generated based on data in a target domain, performing data enhancement processing on the initial data set, and generating a derived data set.
[0066] In this application, the target field usually refers to a specific application field or industry background. For example, finance, medical care, and law, etc. Different fields have different definitions and protection requirements for sensitive information. The identification of the target field helps to clarify which information needs special attention and protection. For example, in the field of banking and finance, sensitive information may involve account information, transaction records, customer identity information, etc.
[0067] Specifically, since the general large model obtained after pre-training is a general large model, it lacks a deep understanding of specific fields, which may not be directly applicable to certain field-specific tasks, such as information extraction, risk assessment in the financial field, or diagnostic support in the medical field. Based on this, this solution uses the data of the target field to fine-tune the pre-trained general large model, so that the target question-answering model obtained after fine-tuning can better understand and process the specific terms and contexts in the target field, thereby improving the performance of the target question-answering model when performing question-answering tasks in the target field.
[0068] It should be understood that the initial data set in this application is one of the training data used for fine-tuning the general large model in the target field. For example, taking the target field as the banking and financial field, the initial data set is a data set provided by the FinGLM model containing tens of thousands of annotated data. The above annotated data are all background knowledge in the financial field, and the data structure of the above annotated data is all question-answer pair data.
[0069] Furthermore, the more training is used for fine-tuning the general large model, the better the performance of the target question-answering model obtained by fine-tuning. However, it takes a long time to label the background knowledge in the financial field to obtain training data, which will reduce the overall efficiency of model fine-tuning. In order to balance the relationship between the performance and efficiency of model fine-tuning, in this application, after obtaining a small amount of labeled data in the initial data set, data enhancement processing is performed on the labeled data to obtain a derived data set. It should be understood that the derived data set includes derived questions obtained by data enhancement based on the question data in the above-mentioned labeled data, and the derived answers corresponding to each derived question.
[0070] In this way, by using the derived data set and the initial data set together as training data, it is possible to fine-tune the general large model through a large amount of training data and obtain a target question-answering model with better performance.
[0071] Optionally, there is no specific limitation on the data enhancement method used in the present application. Data enhancement can be performed using an instruction strategy, or it can be performed using a data enhancement model pre-trained based on a deep neural network.
[0072] S202: Obtain a pre-generated sensitive data set, perform desensitization processing on the sensitive data set, and obtain a fuzzy data set corresponding to the sensitive data set.
[0073] In this application, a sensitive data set is a data set that contains preset sensitive information in the target field.
[0074] It should be understood that desensitization can hide or modify the true content of sensitive information without affecting the use of data.
[0075] Specifically, the desensitization process used in this application may include reverse expression processing, which can be specifically understood as converting the original statement into its negative form. Exemplarily, the reverse expression processing method may include but is not limited to a pre-trained reverse expression model, etc., and other language text processing methods may also be used, which are not limited to this.
[0076] Optionally, in some possible cases, the desensitization process may also include synonym replacement, fuzzification or other forms of semantic transformation to cover up or blur sensitive information.
[0077] Specifically, for question-answer data, if a set of question-answer pairs contains sensitive information, the sensitive information will be present in the answer data. Therefore, the sensitive information in the set of question-answer pairs can be desensitized by performing reverse expression processing on the answer data. For example, the sensitive answer data containing sensitive information is: "The GDP of XX unit in the third quarter is XXX million yuan". After desensitization, the fuzzy answer data obtained is: "The GDP of XX unit in the third quarter is not XXX million yuan", or "XXX million yuan is not the GDP of XX unit in the third quarter".
[0078] The desensitized fuzzy answer data and the question data corresponding to the sensitive answer data are combined to obtain a set of fuzzy data, and a fuzzy data set is generated through multiple sets of fuzzy data.
[0079] S203. A training data set is obtained based on the initial data set, the derived data set and the fuzzy data set. The pre-trained general large model is fine-tuned for at least one round based on the training data set to obtain a target question-answering model.
[0080] Specifically, before fine-tuning the general large model, the initial data set, the derived data set and the fuzzy data set obtained in the above implementation are integrated to obtain a training data set. Furthermore, the general large model is fine-tuned using the training data set.
[0081] In this way, not only can a large amount of data be used to fine-tune the training of a general large model to improve the model fine-tuning performance, but the model can also ignore the learning of sensitive data in the derived data set and the initial data set during the fine-tuning process. In this way, the target question-answering model obtained after fine-tuning will not output sensitive data as answers when receiving question input corresponding to sensitive data, thereby improving data security.
[0082] It should be understood that since the training data set is data in the target field, the target question-answering model obtained after fine-tuning has a deeper understanding of the data in the target field, so the target question-answering model can be used to perform question-answering tasks in the target field, so that the task results obtained are better. It should also be understood that the general large model can receive data input in multiple modes, and the training data in this application may also include data in multiple modes, so the input data of the target question-answering model obtained after fine-tuning may also be multimodal data, and the multimodal data includes at least one of text, image and video. In this way, the target question-answering model can flexibly handle different types of question-answering tasks, increasing the flexibility and universality of the model.
[0083] In the above technical scheme, an initial data set is obtained by annotating a small amount of data in the target field, and a derived data set is obtained by data enhancement of the training data set. In this way, a large amount of training data can be obtained by annotating a small amount of data to reduce the annotation time, thereby achieving the effect of improving the overall efficiency of fine-tuning of the general large model; further, a predetermined sensitive data set is obtained, and the sensitive data set is reversely represented, and the fuzzy data set obtained after the reverse representation is used together with the above-mentioned initial data set and derivative data set as the final training data, and the general large model is fine-tuned to obtain the target question-answering model. In this way, since the sensitive data in the training data contains two expressions, the general large model considers the sensitive data to be unstable data during the learning process because the sensitive data contains completely different description forms. That is, sensitive information will not be learned during the learning process. Therefore, the target question-answering model obtained after fine-tuning will not output the sensitive data as an answer when receiving the question input corresponding to the sensitive data. Therefore, data security can be improved. On the other hand, even if the training data is accidentally leaked during the fine-tuning process, due to the multiple expressions of sensitive data, it is difficult for the outside to reversely restore the real sensitive data content, thereby maintaining the privacy security of customers and other relevant subjects to the greatest extent, and ensuring the security of banking and financial data in the training process.
[0084] Next, the process of constructing the training data required for fine-tuning a general large model is described in detail.
[0085] Optionally, since the initial data set obtained by labeling data in the banking and finance field has a small amount of data, in order to increase the data set used for fine-tuning, the present application can perform data enhancement processing on the initial data set to obtain a derived data set.
[0086] In some optional embodiments, the present application performs data enhancement processing on the initial data set to generate a derived data set, including: the initial data set includes multiple sets of question-answer pairs; according to preset enhancement instructions and enhancement models, data enhancement processing is performed on the question data in each question-answer pair to obtain at least one derived question corresponding to each question data; each derived question and its corresponding derived answer are determined to obtain a derived data set.
[0087] In this application, the data structure of the initial data set is question-answer pair data, that is, a set of data includes a question data and an answer data. Based on the instruction strategy, the above question data and answer data are enhanced respectively. And
[0088] During the data enhancement process, in order to ensure the practicality of the data obtained after enhancement, the enhancement instructions generated by the present application may also include at least one enhancement constraint condition.
[0089] Specifically, an enhancement model for data enhancement processing is obtained, the problem data before enhancement and the enhancement instructions are input into the enhancement model, at least one derivative question output by the model is obtained, and the derivative answer corresponding to the derivative question is determined through data in the banking and financial field, and then a set of derivative data is obtained based on the derivative question and the derivative answer. Next, the derivative data set can be obtained by referring to the above implementation method.
[0090] In one possible example, the question data to be enhanced is: "Based on the annual report of XXXX Technology Co., Ltd. in 202X, can you give me a brief introduction to the company's social responsibility work during the reporting period?" The enhancement instruction is: "Based on this question {question}, modify it according to the following requirements to generate a more complex question. Including:
[0091] 1. Add new constraints and requirements, about 5 words;
[0092] 2. If the original problem can be solved with a few logical steps, add more reasoning steps;
[0093] 3. Provide an incorrect example as a reference to increase misleading information;
[0094] 4. Ask questions of higher complexity, but avoid using them too often;
[0095] 5. Do not modify the meaning of the original question;”
[0096] The above content is input into the LLama3.1-8B model (enhanced model), and the derived question output after the model enhancement processing is obtained, namely, "Can you give me a brief introduction to the company's social responsibility work during the reporting period based on the 202X annual report of XXXX Technology Co., Ltd.? Including but not limited to public welfare donations, employee training and welfare, environmental protection measures, etc. Please provide specific data and case support to analyze the impact of these projects on the company and society."
[0097] In some possible cases, the results obtained after data enhancement based on the above implementation may contain low-quality data, including questions that may be too complex to understand, or questions with repeated content and low information gain. For example, "What is the interest rate?", "How to calculate the interest rate swap based on LIBOR?" Another example is repeatedly asking the same interest calculation method. If the above series of data is used as training data to fine-tune the general large model, there may be low fine-tuning efficiency and poor performance of the fine-tuned target question-answering model.
[0098] Based on this, the technical solution of this application takes the enhanced results output by the enhanced model as the initial derivative problem, optimizes it, and uses the optimized data as derivative data, which can improve the efficiency and effect of subsequent fine-tuning of the general large model.
[0099] In some optional embodiments, the present application obtains at least one derivative problem corresponding to each problem data, including: inputting enhancement instructions and each problem data into an enhancement model to obtain at least one initial derivative problem corresponding to each problem data; in a preset optimization model, determining the problem evaluation results corresponding to each initial derivative problem according to preset optimization instructions; optimizing the initial derivative problems whose problem evaluation results are less than a preset parameter threshold to obtain at least one derivative problem corresponding to each problem data.
[0100] In this application, an optimization model and optimization instructions are pre-built to optimize the initial derivative problem and improve the data quality of the training data.
[0101] Specifically, each initial derivative problem and optimization instruction output by the enhanced model are input into the optimization model for processing to obtain at least one optimized derivative problem output by the optimization model.
[0102] In more detail, the optimization process of any initial derivative problem is taken as an example for explanation. The optimization model in the present application includes a problem evaluation module and a problem optimization module. The optimization instruction includes at least one optimization parameter. On this basis, in the problem evaluation module, the initial derivative problem is evaluated by the optimization parameter to obtain a problem evaluation result corresponding to the initial derivative problem. Exemplarily, the problem evaluation result can be a preset score, so that the score is compared with a preset threshold; if the score is less than the preset threshold, it means that the data quality of the initial derivative problem is poor, and the optimization result corresponding to the initial derivative problem can be input into the problem optimization module to optimize the initial derivative problem, that is, the initial derivative problem is not output as a derivative problem.
[0103] Exemplarily, the optimization parameters included in the optimization instructions in this application may be comprehensibility, flow and consistency. Specifically, comprehensibility in the above optimization parameters is used to evaluate the clarity and comprehensibility of the text, which can be reflected in dimensions such as simplicity (lack of unnecessary complexity), accessibility (using language suitable for the audience) and clarity; fluency is used to evaluate the smoothness of the text, focusing on grammar, sentence structure and natural language flow, which can be reflected in dimensions such as grammar (correct use of language rules), sentence structure (diversity and complexity) and naturalness (naturalness of text flow); consistency can be used to evaluate the logical flow of the text and the coherence of ideas, ensuring that the structure of the text is logical and the ideas are connected, which can be reflected in dimensions such as logical flow (clear progression of ideas), transitions (smooth transitions between topics or sentences) and consistency (lack of contradictory or broken ideas).
[0104] It can be explained that the parameter evaluation reference corresponding to the optimization parameter is pre-set, and the parameter evaluation reference is pre-stored in the problem evaluation module to facilitate subsequent evaluation processing. Exemplarily, the parameter evaluation reference may include:
[0105] Comprehensibility (30 points)
[0106] 27-30 points: The text is extremely clear, logically sound, without any ambiguity, and readers can quickly understand its intention.
[0107] 24-26 points: The text is mostly clear, with occasional slight blurriness or ambiguity that does not affect overall comprehension.
[0108] 18-23 points: There are some unclear or ambiguous parts in the text, which requires readers to make some effort to understand.
[0109] 15-17 points: There are many unclear or ambiguous parts in the text, which seriously affect understanding.
[0110] 14 points and below: The text is extremely difficult to understand and conveys little useful information.
[0111] Fluency (30 points)
[0112] 27-30 points: The text is extremely fluent, easy to read, with perfect sentence structure and appropriate wording.
[0113] 24-26 points: The text is mostly fluent, with occasional minor grammatical or spelling errors that do not affect reading.
[0114] 18-23 points: The text contains some grammatical or spelling errors, or the sentence structure is slightly awkward, but it is still readable overall.
[0115] 15-17 points: The text contains multiple grammatical or spelling errors, or the sentence structure is confusing, affecting reading fluency.
[0116] 14 points and below: The text is extremely difficult to read, with many grammatical and spelling errors and poor sentence structure.
[0117] Consistency (40 points)
[0118] 36-40 points: The ideas, information and style in the text are highly consistent, without any logical contradictions or inconsistencies.
[0119] 32-35 points: The text is consistent in most aspects, with occasional minor inconsistencies that do not affect overall understanding.
[0120] 24-31 points: There are some inconsistencies in the text that may be somewhat confusing to the reader, but it is still acceptable overall.
[0121] 16-23 points: There are multiple inconsistencies in the text, which seriously affect the logical coherence and comprehension.
[0122] 15 points and below: The text is extremely inconsistent, the logic is confusing, and it is impossible to form a coherent understanding.
[0123] Through such refinement, the performance of each dimension can be evaluated more accurately and a corresponding score can be given.
[0124] In a possible example, in the problem evaluation module, the parameter scores of the initial derivative problem corresponding to the three optimization parameters are calculated respectively, and the parameter scores are added together to obtain the problem evaluation result corresponding to the derivative problem.
[0125] It should be noted that the above optimization process can exclude data of poor quality and not meeting the requirements, and use the remaining data sets as training data sets. This can encourage the model to learn more general and abstract language patterns during the fine-tuning process, rather than just memorizing the training data, thereby improving the generalization ability of the target question-answering model after fine-tuning.
[0126] When the target domain is data in the banking and financial field, the derived data set and the initial data set obtained contain more sensitive information. In order to protect the data security of the banking and financial field during the fine-tuning training process and the target question-answering model obtained after fine-tuning, this application obtains the pre-determined sensitive data in the banking and financial field, and desensitizes the sensitive data to obtain a fuzzy data set. When the fuzzy data set is used together with the derived data set and the initial data set as the training data set, the sensitive data in the derived data set and the initial data set can be fuzzified, so that the model cannot learn the real sensitive data during the learning process, which can avoid the leakage of sensitive data and improve data security.
[0127] For example, when the model learns sensitive data during the learning process, it will learn different answers to the same question. Therefore, the model cannot obtain a stable answer to the question and will give up answering the question in the subsequent learning process. In this way, when the model receives the input of the question in the subsequent use, it will feedback an uncertain answer, thus avoiding the leakage of sensitive information.
[0128] Moreover, since the sensitive data contained in the derived data set and the initial data set are relatively scattered, it would take a long time to directly search for the sensitive data in the above data set and perform desensitization processing. In this way, the above implementation method of the present application can reduce the time spent in searching for sensitive data in the derived data set and the initial data set, thereby achieving the technical effect of improving the overall efficiency of model fine-tuning.
[0129] Optionally, the present application performs desensitization processing on the sensitive data set to obtain a fuzzy data set corresponding to the sensitive data set, including: determining that the sensitive data set includes multiple sensitive question-answer pairs; desensitizing the sensitive answers in each sensitive question-answer pair according to preset desensitization instructions to obtain at least one fuzzy answer corresponding to each sensitive answer; and generating a fuzzy data set based on each fuzzy answer and its corresponding sensitive question.
[0130] In this application, sensitive data includes sensitive information pre-collected in the target field. Sensitive information can be collected from banking-related applications, business systems, and data reports, etc.
[0131] In order to facilitate the subsequent fine-tuning training of the general large model, this application uses the collected sensitive information as answer data and generates its corresponding question data, so that multiple sensitive question-answer pairs, namely, sensitive data sets, can be obtained.
[0132] On this basis, the preset desensitization instruction is obtained. Exemplarily, if the desensitization process is a reverse expression process, the desensitization instruction can be: a corresponding instruction for performing a reverse expression process on the sentence. On this basis, the desensitization model in this application can also be an LLama3.1-8B model. In this way, the above-mentioned desensitization instruction and the sensitive answer are input into the preset desensitization model to obtain at least one fuzzy answer after the desensitization process.
[0133] Furthermore, a fuzzy data set is generated based on the fuzzy answers and the sensitive questions corresponding to the above sensitive answers, so as to achieve fuzzification of the sensitive data contained in the training data and improve data security.
[0134] It should be noted that although sensitive questions are described as sensitive questions, their problem statements do not contain sensitive information. Therefore, they do not need to be desensitized and can be used directly as training data. This can reduce the amount of desensitized data and improve the overall processing efficiency of model fine-tuning.
[0135] Based on the above implementation methods, the present application merges the obtained fuzzy data set, derived data set and initial data set to obtain a training data set, and fine-tunes the general large model through the training data set to obtain the target question-answering model.
[0136] Optionally, the pre-trained general large model is fine-tuned for at least one round according to the training data set to obtain a target question and answer model, including: for any round of fine-tuning training, the test question data in the training data set is input into the general large model after the current round of fine-tuning to obtain test prediction data corresponding to the test question data; based on the data similarity between the test prediction data and the test answer data corresponding to the test question data, the fine-tuning result corresponding to the current round of fine-tuning is determined; if the fine-tuning result corresponding to the current round of fine-tuning is better than the fine-tuning result corresponding to the previous round of fine-tuning, the general large model is continued to be fine-tuned for the next round; if the fine-tuning result corresponding to the current round of fine-tuning is worse than the fine-tuning result corresponding to the previous round of fine-tuning, the general large model corresponding to the previous round of fine-tuning is used as the target question and answer model.
[0137] In the present application, by adopting at least one round of fine-tuning training and flexibly adjusting whether to continue fine-tuning according to the results of each round of fine-tuning, resources and time can be maximized while improving model performance.
[0138] Specifically, during any round of fine-tuning training, a test data set predetermined in the above training data set is obtained, and the test question data in the test data set are respectively input into the general large model after the current round of fine-tuning to obtain the test prediction data output by the model; and, the test prediction data is compared with the test answer data corresponding to the test question data for similarity to determine the fine-tuning result corresponding to the current round of fine-tuning; optionally, if the fine-tuning result corresponding to the current round of fine-tuning is better than the fine-tuning result corresponding to the previous round of fine-tuning, it means that there is room for further optimization of the general large model and it can continue to be fine-tuned; conversely, if the fine-tuning result corresponding to the current round of fine-tuning is worse than the fine-tuning result corresponding to the previous round of fine-tuning, it means that the model performance obtained by continuing the fine-tuning training will be worse, so the fine-tuning training can be stopped, and in order to obtain a target question and answer model with better performance, the general large model corresponding to the previous round of fine-tuning can be used as the target question and answer model.
[0139] Optionally, in the above-mentioned process of comparing the similarity of the test prediction data with the test answer data corresponding to the test question data, the comparison method may be: obtaining prediction keywords in the test prediction data, and answer keywords in the test answer data; determining a first fine-tuning result corresponding to the current round of fine-tuning based on the similarity between the prediction keywords and the answer keywords; when the first fine-tuning result meets a preset result condition, obtaining a prediction vector corresponding to the test prediction data, and an answer vector corresponding to the test answer data; determining a second fine-tuning result corresponding to the current round of fine-tuning based on the similarity between the prediction vector and the answer vector; determining the fine-tuning result corresponding to the current round of fine-tuning based on the first fine-tuning result, or the first fine-tuning result and the second fine-tuning result.
[0140] The present application can compare similarities in two dimensions, one is the word dimension and the other is the vector dimension, and the final similarity corresponding to the data is obtained by combining the comparison results of the two dimensions.
[0141] Specifically, keyword extraction is performed on the test prediction data and the test answer data respectively to obtain prediction keywords and answer keywords, and then the first fine-tuning result is obtained by comparing the word similarity between the two groups of keywords. Optionally, if the above comparison result does not meet the preset result condition, that is, the similarity between the two groups of keywords is low, the comparison result can be directly used as the similarity comparison result between the test prediction data and the test question data, that is, the fine-tuning result; conversely, if the above comparison result meets the preset condition, a more detailed similarity between the two test data can be further determined to determine whether to continue to fine-tune the general large model.
[0142] In some possible examples, when comparing word similarities, if the number of extracted keywords is 5, and two or more of the two groups of keywords are inconsistent, it means that the word similarity does not meet the preset result conditions.
[0143] Optionally, if the comparison result does not meet the preset result condition, the test prediction data and the test answer data are respectively transformed into vectors to obtain a prediction vector and an answer vector corresponding to the test answer data, and the prediction vector and the answer vector are compared for vector similarity to obtain a comparison result. Furthermore, the keyword comparison result and the vector comparison result are combined as the similarity comparison result between the test prediction data and the test question data, i.e., the fine-tuning result.
[0144] In some possible examples, the vector similarity may be calculated by calculating the cosine similarity, which may achieve accurate and fast results.
[0145] Figure 3 A schematic diagram of the structure of a target question answering model training device provided in an embodiment of the present application. Figure 3 The target question-answering model training device 30 includes: a derivative data set generation module 301, a fuzzy data set generation module 302 and a model fine-tuning module 303; wherein,
[0146] A derivative data set generation module 301 is used to obtain an initial data set generated based on data in a target domain, perform data enhancement processing on the initial data set, and generate a derivative data set;
[0147] The fuzzy data set generation module 302 obtains a pre-generated sensitive data set, performs desensitization processing on the sensitive data set, and obtains a fuzzy data set corresponding to the sensitive data set; wherein the sensitive data set is a data set containing preset sensitive information in the target domain; the desensitization processing includes reverse expression processing;
[0148] The model fine-tuning module 303 is used to obtain a training data set based on the initial data set, the derived data set and the fuzzy data set, and perform at least one round of fine-tuning training on the pre-trained general large model based on the training data set to obtain a target question-answering model; wherein the target question-answering model is used to perform question-answering tasks within the target domain; the input of the target question-answering model is multimodal data, and the multimodal data includes at least one of text, image and video.
[0149] In an optional implementation, the initial data set includes multiple sets of question-answer pairs;
[0150] The derived data set generation module 301 includes:
[0151] A derivative question determination submodule is used to perform data enhancement processing on the question data in each question-answer pair according to a preset enhancement instruction and enhancement model to obtain at least one derivative question corresponding to each question data;
[0152] The derived data set determination submodule is used to determine each derived question and its corresponding derived answer to obtain a derived data set.
[0153] In an optional implementation, the derived problem determination submodule includes:
[0154] An initial derivative question determination unit, used for inputting the enhancement instruction and each question data into the enhancement model to obtain at least one initial derivative question corresponding to each question data;
[0155] A problem evaluation result determination unit is used to determine the problem evaluation results corresponding to each initial derivative problem in a preset optimization model according to a preset optimization instruction;
[0156] The derivative problem determination unit is used to optimize the initial derivative problems whose problem evaluation results are less than a preset threshold value to obtain at least one derivative problem corresponding to each problem data.
[0157] In an optional implementation, the fuzzy data set generating module 302 includes:
[0158] A sensitive question-answer pair determination submodule is used to determine whether a sensitive data set includes multiple sensitive question-answer pairs;
[0159] The fuzzy answer obtaining submodule is used to perform desensitization processing on the sensitive answers in each sensitive question and answer pair according to a preset desensitization instruction, and obtain at least one fuzzy answer corresponding to each sensitive answer; the desensitization processing includes reverse expression processing on the sensitive data set;
[0160] The fuzzy data set generation submodule is used to generate a fuzzy data set according to each fuzzy answer and its corresponding sensitive question.
[0161] In an optional implementation, the model fine-tuning module 303 includes:
[0162] The test prediction data acquisition submodule is used to input the test problem data in the training data set into the general large model after fine-tuning in the current round for any round of fine-tuning training, and obtain the test prediction data corresponding to the test problem data;
[0163] A fine-tuning result obtaining submodule is used to determine the fine-tuning result corresponding to the current round of fine-tuning based on the data similarity between the test prediction data and the test answer data corresponding to the test question data;
[0164] The first fine-tuning submodule is used to continue fine-tuning the general large model for the next round if the fine-tuning result corresponding to the current round of fine-tuning is better than the fine-tuning result corresponding to the previous round of fine-tuning;
[0165] The second fine-tuning submodule is used to use the general large model corresponding to the previous round of fine-tuning as the target question-answering model if the fine-tuning result corresponding to the current round of fine-tuning is worse than the fine-tuning result corresponding to the previous round of fine-tuning.
[0166] In an optional implementation, the fine-tuning result obtaining submodule includes:
[0167] A keyword obtaining unit, used to obtain prediction keywords in the test prediction data and answer keywords in the test answer data;
[0168] A first fine-tuning result obtaining unit, configured to determine a first fine-tuning result corresponding to the current round of fine-tuning according to the similarity between the predicted keyword and the answer keyword;
[0169] A vector obtaining unit, used for obtaining a prediction vector corresponding to the test prediction data and an answer vector corresponding to the test answer data when the first fine-tuning result satisfies a preset result condition;
[0170] A second fine-tuning result obtaining unit, used to determine a second fine-tuning result corresponding to the current round of fine-tuning according to the similarity between the prediction vector and the answer vector;
[0171] The fine-tuning result obtaining unit is used to determine the fine-tuning result corresponding to the current round of fine-tuning according to the first fine-tuning result, or the first fine-tuning result and the second fine-tuning result.
[0172] Figure 4 is a block diagram of an electronic device shown in an embodiment of the present application, and the device may be a computer, a digital broadcast terminal, etc. Figure 4 , the device 800 may include one or more of the following components: a processing component 802 , a memory 804 , a power component 806 , a multimedia component 808 , an audio component 810 , an input / output interface 812 , a sensor component 814 , and a communication component 816 .
[0173] The processing component 802 generally controls the overall operation of the device 800, such as operations associated with display, phone calls, data communications, camera operations, and recording operations. The processing component 802 may include one or more processors 820 to execute instructions to complete all or part of the steps of the above-mentioned method. In addition, the processing component 802 may include one or more modules to facilitate the interaction between the processing component 802 and other components. For example, the processing component 802 may include a multimedia module to facilitate the interaction between the multimedia component 808 and the processing component 802.
[0174] The memory 804 is configured to store various types of data to support operations on the device 800. Examples of such data include instructions for any application or method operating on the device 800, contact data, phone book data, messages, pictures, videos, etc. The memory 804 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk or optical disk.
[0175] The power supply component 806 provides power to the various components of the device 800. The power supply component 806 can include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power for the device 800.
[0176] The multimedia component 808 includes a screen that provides an output interface between the device 800 and the user. In some embodiments, the screen may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen may be implemented as a touch screen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touch, slide, and gestures on the touch panel. The touch sensor may not only sense the boundaries of the touch or slide action, but also detect the duration and pressure associated with the touch or slide operation. In some embodiments, the multimedia component 808 includes a front camera and / or a rear camera. When the device 800 is in an operating mode, such as a shooting mode or a video mode, the front camera and / or the rear camera may receive external multimedia data. Each front camera and rear camera may be a fixed optical lens system or have a focal length and optical zoom capability.
[0177] The audio component 810 is configured to output and / or input audio signals. For example, the audio component 810 includes a microphone (MIC), and when the device 800 is in an operating mode, such as a call mode, a recording mode, and a speech recognition mode, the microphone is configured to receive an external audio signal. The received audio signal can be further stored in the memory 804 or sent via the communication component 816. In some embodiments, the audio component 810 also includes a speaker for outputting audio signals.
[0178] The input / output interface 812 provides an interface between the processing component 802 and the peripheral interface modules, which may be keyboards, click wheels, buttons, etc. These buttons may include but are not limited to: a home button, a volume button, a start button, and a lock button.
[0179] The sensor assembly 814 includes one or more sensors for providing various aspects of status assessment for the device 800. For example, the sensor assembly 814 can detect the open / closed state of the device 800, the relative positioning of components, such as the display and keypad of the device 800, and the sensor assembly 814 can also detect the position change of the device 800 or a component of the device 800, the presence or absence of user contact with the device 800, the orientation or acceleration / deceleration of the device 800, and the temperature change of the device 800. The sensor assembly 814 may include a proximity sensor configured to detect the presence of nearby objects without any physical contact. The sensor assembly 814 may also include an optical sensor, such as a solid image (Complementary Metal Oxide Semiconductor, CMOS) sensor or a semiconductor image (Charge-coupled Device, CCD) sensor, for use in imaging applications. In some embodiments, the sensor assembly 814 may also include an acceleration sensor, a gyroscope sensor, a magnetic sensor, a pressure sensor, or a temperature sensor.
[0180] The communication component 816 is configured to facilitate wired or wireless communication between the device 800 and other devices. The device 800 can access a wireless network based on a communication standard, such as WiFi, 4G or 5G, or a combination thereof. In an exemplary embodiment, the communication component 816 receives a broadcast signal or broadcast-related information from an external broadcast management system via a broadcast channel. In an exemplary embodiment, the communication component 816 also includes a near field communication (NFC) module to facilitate short-range communication. For example, the NFC module can be implemented based on radio frequency identification (RFID) technology, infrared data association (IrDA) technology, ultra-wideband (UWB) technology, Bluetooth (BT) technology and other technologies.
[0181] In an exemplary embodiment, the device 800 may be implemented by one or more application specific integrated circuits (ASIC), digital signal processors (DSP), digital signal processing devices (DSPD), programmable logic devices (PLD), field programmable gate arrays (FPGA), controllers, microcontrollers, microprocessors or other electronic components to perform the above method.
[0182] In an exemplary embodiment, a non-transitory computer-readable storage medium including instructions is also provided, such as a memory 804 including instructions, and the instructions can be executed by a processor 820 of the device 800 to perform the above method. For example, the non-transitory computer-readable storage medium can be a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, an optical data storage device, etc.
[0183] A non-transitory computer-readable storage medium, when the instructions in the storage medium are executed by a processor of a server, enables the server to execute the above-mentioned physical backup method of the database.
[0184] An embodiment of the present application also provides a chip for running instructions, which is used to execute the technical solution of the physical backup method of the database in the above embodiment.
[0185] An embodiment of the present application further provides a computer-readable storage medium, in which computer execution instructions are stored. When the computer execution instructions are executed on a computer, the computer executes the technical solution of the physical backup method of the database in the above embodiment.
[0186] An embodiment of the present application also provides a computer program product, which includes a computer program stored in a computer-readable storage medium. At least one processor can read the computer program from the computer-readable storage medium. When at least one processor executes the computer program, the technical solution of the physical backup method of the database in the above embodiment can be implemented.
[0187] Those skilled in the art will readily appreciate other embodiments of the present application after considering the specification and practicing the invention disclosed herein. The present application is intended to cover any modification, use or adaptation of the present application, which follows the general principles of the present application and includes common knowledge or customary techniques in the art that are not disclosed in the present application. The specification and examples are intended to be exemplary only, and the true scope and spirit of the present application are indicated by the following claims.
[0188] It should be understood that the present application is not limited to the precise structures that have been described above and shown in the drawings, and that various modifications and changes may be made without departing from the scope thereof. The scope of the present application is limited only by the appended claims.
[0189] Those skilled in the art will readily appreciate other embodiments of the present application after considering the specification and practicing the invention disclosed herein. The present application is intended to cover any modification, use or adaptation of the present application, which follows the general principles of the present application and includes common knowledge or customary techniques in the art that are not disclosed in the present application. The specification and examples are intended to be exemplary only, and the true scope and spirit of the present application are indicated by the following claims.
[0190] It should be understood that the present application is not limited to the precise structures that have been described above and shown in the drawings, and that various modifications and changes may be made without departing from the scope thereof. The scope of the present application is limited only by the appended claims.
Claims
1. A target question answering model training method, characterized in that: The method comprises: Acquire an initial data set generated based on data in a target domain, perform data enhancement processing on the initial data set, and generate a derived data set; Acquire a pre-generated sensitive data set, perform desensitization processing on the sensitive data set, and obtain a fuzzy data set corresponding to the sensitive data set; wherein the sensitive data set is a data set containing preset sensitive information in the target field; the desensitization processing includes reverse expression processing; A training data set is obtained according to the initial data set, the derived data set and the fuzzy data set, and at least one round of fine-tuning training is performed on the pre-trained general large model according to the training data set to obtain a target question-answering model; wherein the target question-answering model is used to perform question-answering tasks within the target domain; the input of the target question-answering model is multimodal data, and the multimodal data includes at least one of text, image and video.
2. The method according to claim 1, characterized in that The initial data set includes multiple sets of question-answer pairs; Performing data enhancement processing on the initial data set to generate a derived data set includes: According to the preset enhancement instructions and enhancement model, data enhancement processing is performed on the question data in each question-answer pair to obtain at least one derivative question corresponding to each question data; Determine each of the derived questions and its corresponding derived answer to obtain the derived data set.
3. The method according to claim 2, characterized in that Obtaining at least one derivative question corresponding to each of the problem data, including: Inputting the enhancement instruction and each of the question data into the enhancement model to obtain at least one initial derivative question corresponding to each of the question data; In a preset optimization model, according to preset optimization instructions, problem evaluation results corresponding to each of the initial derivative problems are determined; The initial derived questions whose question evaluation results are less than a preset threshold are optimized to obtain at least one derived question corresponding to each of the question data.
4. The method according to claim 1, characterized in that: Desensitizing the sensitive data set to obtain a fuzzy data set corresponding to the sensitive data set, including: Determining that the sensitive data set includes a plurality of sensitive question-answer pairs; According to the preset desensitization instruction, the sensitive answers in each of the sensitive question and answer pairs are desensitized to obtain at least one fuzzy answer corresponding to each of the sensitive answers; A fuzzy data set is generated according to each of the fuzzy answers and the corresponding sensitive questions.
5. The method according to claim 1, characterized in that Performing at least one round of fine-tuning training on the pre-trained general large model according to the training data set to obtain a target question-answering model, including: For any round of fine-tuning training, the test problem data in the training data set is input into the general large model after the current round of fine-tuning to obtain the test prediction data corresponding to the test problem data; Determine a fine-tuning result corresponding to the current round of fine-tuning according to the data similarity between the test prediction data and the test answer data corresponding to the test question data; If the fine-tuning result corresponding to the current round of fine-tuning is better than the fine-tuning result corresponding to the previous round of fine-tuning, continue to perform the next round of fine-tuning on the general large model; If the fine-tuning result corresponding to the current round of fine-tuning is worse than the fine-tuning result corresponding to the previous round of fine-tuning, the general large model corresponding to the previous round of fine-tuning is used as the target question-answering model.
6. The method according to claim 5, characterized in that Determining a fine-tuning result corresponding to the current round of fine-tuning according to the data similarity between the test prediction data and the test answer data corresponding to the test question data includes: Obtaining prediction keywords in the test prediction data and answer keywords in the test answer data; Determining a first fine-tuning result corresponding to the current round of fine-tuning according to the similarity between the predicted keyword and the answer keyword; When the first fine-tuning result satisfies a preset result condition, obtaining a prediction vector corresponding to the test prediction data and an answer vector corresponding to the test answer data; Determining a second fine-tuning result corresponding to the current round of fine-tuning according to the similarity between the prediction vector and the answer vector; A fine-tuning result corresponding to the current round of fine-tuning is determined according to the first fine-tuning result, or the first fine-tuning result and the second fine-tuning result.
7. A target question answering model training device, characterized in that: The device comprises: A derived data set generation module is used to obtain an initial data set generated based on data in the target domain, perform data enhancement processing on the initial data set, and generate a derived data set; A fuzzy data set generation module is used to obtain a pre-generated sensitive data set, perform desensitization processing on the sensitive data set, and obtain a fuzzy data set corresponding to the sensitive data set; wherein the sensitive data set is a data set containing preset sensitive information in the target field; and the desensitization processing includes reverse expression processing; A model fine-tuning module is used to obtain a training data set based on the initial data set, the derived data set and the fuzzy data set, and to perform at least one round of fine-tuning training on the pre-trained general large model based on the training data set to obtain a target question-answering model; wherein the target question-answering model is used to perform question-answering tasks within the target domain; the input of the target question-answering model is multimodal data, and the multimodal data includes at least one of text, image and video.
8. An electronic device, characterized in that: include: A processor and a memory communicatively connected to the processor; The memory stores computer-executable instructions; When executing the computer-executable instructions, the processor is used to implement the target question-answering model training method as described in any one of claims 1 to 6.
9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer-executable instructions, which, when executed by a processor, are used to implement the target question-answering model training method as described in any one of claims 1 to 6.
10. A computer program product, characterized in that The method comprises a computer program, which implements the method according to any one of claims 1 to 6 when the computer program is executed by a processor.