Model tuning method and device, equipment and storage medium

By retrieving data related to the initial problem from the preset database, determining the target question-and-answer pairs and generating the target data set, the problem of insufficient effective data in vertical field model training is solved, and the tuning effect of the model is improved.

CN120371981AActive Publication Date: 2025-07-25INSPUR SUZHOU INTELLIGENT TECH CO LTD

Patent Information

Application Number
CN202510838221.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-20
Publication Date
2025-07-25
Estimated Expiration
2045-06-20

AI Technical Summary

Technical Problem

In the prior art, the direct use of data sets during vertical field model training results in less effective data and poor tuning effect.

Method used

By retrieving the first data related to the initial question from the preset database, determining the target question-and-answer pairs and generating a target data set where the number of target questions is greater than the initial data set, the target model is tuned using the target data set.

Benefits of technology

The tuning effect of the target model is improved, the proportion of effective data is increased, and the performance of the model in the vertical field is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120371981A_ABST
    Figure CN120371981A_ABST
Patent Text Reader

Abstract

The invention discloses a model tuning method and device, equipment and a storage medium, and relates to the technical field of data processing, and the method comprises the steps: obtaining an initial data set, an initial question and answer pair comprising an initial question and a corresponding initial answer; in order to extract data related to the initial questions, first data corresponding to all the initial questions are retrieved from a preset database, then a target question-answer pair is determined by adopting a target model according to the first data and the initial question-answer pair, and the target question-answer pair is used as a target question-answer pair. The target question is a corresponding initial question which is not answered by the target model. In order to improve the performance of model tuning, the target data set is generated according to the target question and answer pair and the initial data set, then the target data set is used for tuning the target model, and the effective data in the target data set for tuning the target model is improved, so that the tuning effect of the target model is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of data processing technology, and in particular to a model tuning method, device, equipment and storage medium. Background Art

[0002] In today's era of booming artificial intelligence technology, models have shown great application potential in many fields with their powerful language understanding and generation capabilities, especially in vertical fields where application demand is growing. Vertical fields are often highly professional, knowledge-intensive, and complex, which places higher demands on the accuracy, professionalism, and adaptability of models.

[0003] At present, when training models applied to vertical fields, the collected data sets are generally directly input into the model for training, resulting in less effective data and poor tuning effects. Summary of the invention

[0004] The present application provides a model tuning method, apparatus, device and storage medium to at least solve the problem in the related art that the collected data set is directly input into the model for training, resulting in less valid data and poor tuning effect.

[0005] In a first aspect, the present application provides a model tuning method, the method comprising:

[0006] Obtain an initial data set; the initial data set includes multiple initial question-answer pairs; the initial question-answer pairs include initial questions and corresponding initial answers;

[0007] Retrieving first data corresponding to each initial question from a preset database; the first data is data related to the initial question;

[0008] A target question-answer pair is determined using the target model based on the first data and the initial question-answer pair; the target question-answer pair includes a target question and a corresponding initial answer and the corresponding first data; the target question refers to the corresponding initial question that is not answered by the target model;

[0009] Generate a target data set based on the target question-answer pair and the initial data set; the number of target questions included in the target data set is greater than the number of initial questions determined in the initial data set to correspond to the target questions;

[0010] Use the target dataset to tune the target model.

[0011] In a second aspect, the present application also provides a model tuning device, comprising:

[0012] An acquisition module is used to obtain an initial data set;

[0013] A retrieval module, configured to retrieve first data corresponding to each initial question from a preset database; the first data is data related to the initial question;

[0014] A determination module, configured to use a target model and determine a target Q&A pair based on the first data and an initial Q&A pair; the target Q&A pair includes a target question, a corresponding initial answer, and corresponding first data; the target question refers to the corresponding initial question that the target model lacks in answering;

[0015] A generation module, configured to generate a target data set based on the target Q&A pair and an initial data set; the number of target questions included in the target data set is greater than the number of initial questions corresponding to the target questions determined in the initial data set; A tuning module, configured to tune the target model using the target data set.

[0016] In a third aspect, the present application further provides an electronic device, including: a memory, configured to store a computer program; a processor, configured to implement the steps of the model tuning method provided in the first aspect above when executing the computer program.

[0017] In a fourth aspect, the present application further provides a computer-readable storage medium, in which a computer program is stored, and wherein the computer program, when executed by a processor, implements the steps of the model tuning method provided in the first aspect above.

[0018] In a fifth aspect, the present application further provides a computer program product, including a computer program, and the computer program, when executed by a processor, implements the steps of the model tuning method provided in the first aspect above.

[0019] Through the model tuning method, device, equipment, and storage medium provided by the present application, when it is necessary to train a target model, an initial data set is obtained. Among them, the initial data set includes a plurality of initial Q&A pairs, and the initial Q&A pairs include initial questions and corresponding initial answers. In order to identify the questions that the target model lacks in answering, the first data corresponding to each initial question is retrieved from a preset database, where the first data is data related to the initial question. According to each initial Q&A pair and the corresponding first data, and using the target model, a target Q&A pair is determined. The target question in the target Q&A pair is the corresponding initial question that the target model lacks in answering. In order to improve the target model's ability to answer the target question, a target data set is generated according to the target Q&A pair and the initial data set, where the number of target questions included in the target data set is greater than the number of initial questions corresponding to the target questions determined in the initial data set, so as to tune the target model using the target data set, thereby increasing the proportion of valid data in the target data set and improving the tuning effect of the model. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] To more clearly illustrate the embodiments of the present application, the following will briefly introduce the accompanying drawings required in the embodiments. Obviously, the accompanying drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can be obtained based on these drawings.

[0021] Figure 1 It is an application scenario diagram of the model tuning method provided by the embodiments of the present application;

[0022] Figure 2 It is a schematic flowchart of the model tuning method provided by an embodiment of the present application;

[0023] Figure 3 It is a schematic flowchart of the model tuning method provided by another embodiment of the present application;

[0024] Figure 4 It is a schematic structural diagram of the model tuning device provided by an embodiment of the present application;

[0025] Figure 5 It is a schematic structural diagram of the electronic device provided by an embodiment of the present application. Detailed implementation manners

[0026] The following will clearly and completely describe the technical solutions in the embodiments of the present application in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only some embodiments of the present application, rather than all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the protection scope of the present application.

[0027] It should be noted that in the description of the present application, the terms "include", "comprise" or any other variation thereof are intended to cover a non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements but also includes other elements not explicitly listed, or further includes elements inherent to such process, method, article or device. The terms "first", "second", etc. in the present application are used to distinguish similar objects and are not used to describe a specific order or sequence.

[0028] To enable those skilled in the art of the present technology to better understand the solutions of the present application, the following will further elaborate on the present application in conjunction with the accompanying drawings and specific implementation manners.

[0029] In today's era of booming artificial intelligence technology, models have shown great application potential in many fields with their powerful language understanding and generation capabilities, especially in vertical fields where application demand is growing. Vertical fields are often characterized by strong professionalism, knowledge-intensiveness, and complex scenarios, which place higher demands on the accuracy, professionalism, and adaptability of models. In order to improve the performance of large language models in vertical fields, the optimization of training data is crucial. At present, when tuning models applied to vertical fields, the obtained tuning data sets are often directly input into the model for training, without determining the parts that the model actually lacks, and focusing on the parts that the model lacks, resulting in less valid data in the tuning data sets, which in turn leads to poor model tuning effects.

[0030] Therefore, when facing the above technical problems, the data set used for tuning is no longer directly input into the model for tuning, but the data that the model lacks in the tuning data set is determined, so as to increase the proportion of the missing data, focus on the training of the missing part, and improve the performance of the model. Specifically, an initial data set is obtained, and the initial data set includes multiple initial question-answer pairs, and the initial question-answer pairs include initial questions and corresponding initial answers. In order to extract the data related to the initial question, the first data corresponding to each initial question is retrieved from a preset database including the vertical field data corresponding to the target model, and the first data is data related to the initial question. Then, the target model is used and the target question-answer pair is determined based on the first data and the initial question-answer pair, wherein the target question-answer pair includes the target question and the corresponding initial answer and the corresponding first data, and the target question is the corresponding initial question that the target model lacks in answering. In order to improve the performance of model tuning, the missing questions are focused on training, so a target data set is generated according to the target question-answer pair and the initial data set, wherein the number of target questions included in the target data set is greater than the number of initial questions determined as corresponding to the target questions in the initial data set, and then the target data set is used to tune the target model. Therefore, the effective data in the target data set used for tuning the target model is increased, thereby improving the tuning effect of the target model.

[0031] Figure 1 This is an application scenario diagram of the model tuning method provided in the embodiment of the present application. Figure 1As shown in the figure, the application scenario provided in this embodiment includes: a server device 10 and a client device 20. The model tuning method is applied to the server device 10. The server device 10 receives a model tuning request triggered by a user on the client device 20, and after receiving the model tuning request, obtains an initial data set. The initial data set includes a plurality of initial question-and-answer pairs. Each initial question-and-answer pair includes an initial question and a corresponding initial answer. Then, the first data corresponding to each initial question is retrieved from a preset database. The first data is data related to the initial question. The preset database includes data in the vertical domain corresponding to the target model. Further, according to the first data and the initial question-and-answer pairs, the target question-and-answer pairs are determined by using the target model. Among them, each target question-and-answer pair includes a target question, a corresponding initial answer, and corresponding first data. The target question refers to the initial question for which the target model's answer is lacking. Then, according to the target question-and-answer pairs and the initial data set, a target data set is generated, and then the target data set is used to tune the target model. After the tuning of the target model is completed, the tuning result is sent to the client device 20, and the client device 20 displays the tuning result.

[0032] Figure 2 It is a schematic flowchart of the model tuning method provided by an embodiment of the present application, as Figure 2 shown. The model tuning method provided in this embodiment is applied to the server device. The model tuning method provided in this embodiment specifically includes the following steps:

[0033] S201: Obtain an initial data set.

[0034] Among them, the initial data set includes a plurality of initial question-and-answer pairs. Each initial question-and-answer pair includes an initial question and a corresponding initial answer.

[0035] Among them, the initial question refers to the question included in the initial data set, and the initial answer is the answer corresponding to the initial question.

[0036] Among them, the initial answer is the reference answer corresponding to the initial question. If the initial question is a non-open-ended question, the initial answer is the standard answer to the initial question. If the initial question is an open-ended question, the initial answer is the reference answer corresponding to the initial question.

[0037] Specifically, in this embodiment, the server device obtains the initial data set from the preset database.

[0038] S202: Retrieve the first data corresponding to each initial question from the preset database.

[0039] Among them, the first data is data related to the initial question. The preset database includes data in the vertical domain corresponding to the target model.

[0040] Optionally, the vertical domain can be set independently according to requirements, and no limitation is made in this embodiment.

[0041] Optionally, the data in the preset database may include but is not limited to text data, professional knowledge related to the vertical domain, Q&A pairs related to the vertical domain, etc., and no limitation is made in this embodiment.

[0042] Specifically, in this embodiment, the server device adopts a preset technology and retrieves from the preset database according to each initial question, retrieves the data related to the initial question, and uses the data related to the initial question as the first data corresponding to each initial question.

[0043] Optionally, the preset technology can be similarity search (Facebook AI Similarity Search, Faiss), etc., and no limitation is made in this embodiment.

[0044] S203: Use the target model and determine the target Q&A pair based on the first data and the initial Q&A pair.

[0045] Among them, the target Q&A pair includes the target question, the corresponding initial answer, and the corresponding first data. The target question refers to the corresponding initial question that the target model lacks in answering.

[0046] Specifically, in this embodiment, the server device inputs each initial question into the target model for answering, outputs the first initial question answer, then uses the first data corresponding to each initial question as additional context information and adds it to the prompt information of the target model, and then inputs each initial question into the model for answering again, outputs the second initial question answer. Use a preset model to score the first initial question answer and the second initial question answer according to the initial answer. If the score of the second initial question answer is greater than the score of the first initial question answer, then determine the second initial question answer as the target question according to the corresponding initial question. And determine the target question, the corresponding initial answer, and the corresponding first data as the target Q&A pair.

[0047] Optionally, in this embodiment, the server device inputs each initial question into the target model for answering, outputs the first initial question answer, then uses the first data corresponding to each initial question as additional context information and adds it to the prompt information of the target model, and then inputs each initial question into the model for answering again, outputs the second initial question answer. Use a preset model to compare the effects of the first initial question answer and the second initial question answer according to the initial answer, and output the answer with the best effect among each initial question. If the preset model outputs the second initial question answer, then determine the second initial question answer as the target question according to the corresponding initial question. And determine the target question, the corresponding initial answer, and the corresponding first data as the target Q&A pair.

[0048] Among them, the first initial question answer refers to the answer given by the target model when the first data corresponding to each initial question is not added as context information to the target model. The second initial question answer refers to the answer given by the target model when the first data corresponding to each initial question is added as context information to the target model.

[0049] Optionally, the preset model can be a general artificial intelligence model, etc., which is not limited in this embodiment.

[0050] Among them, the general artificial intelligence model refers to an artificial intelligence system with the abilities of cross-domain knowledge understanding, multi-task processing, autonomous learning and reasoning.

[0051] It can be understood that the initial answer is the reference answer corresponding to the initial question. If the initial question is a non-open question, the initial answer is the standard answer to the initial question. If the initial question is an open question, the initial answer is the reference answer corresponding to the initial question. Therefore, when comparing the effects using the preset model, if it is a non-open question, compare which of the first initial question answer and the second initial question answer is closer to the initial answer. If the initial question is an open question, use the preset model to compare which of the first initial question answer and the second initial question answer answers better based on the initial answer.

[0052] Optionally, the target model can be a large language model (LLM), etc., which is not limited in this embodiment.

[0053] S204: Generate a target data set based on the target question-answer pairs and the initial data set. Among them, the number of target questions included in the target data set is greater than the number of initial questions corresponding to the target questions determined in the initial data set. It can be understood that the target questions correspond to the target question-answer pairs, and the number of target question-answer pairs in the target question-answer pairs is greater than the number of initial question-answer pairs corresponding to the initial questions determined as target questions in the initial data set.

[0054] Specifically, in this embodiment, the server device increases the number of target question-answer pairs according to a preset multiple, reads out the initial question-answer pairs of the initial questions that are not determined as target questions from the initial data set, randomly deletes the read-out initial question-answer pairs according to the preset multiple, and determines the target question-answer pairs with the increased number and the initial question-answer pairs with the decreased number as the target data set.

[0055] It can be understood that by increasing the number of target question-answer pairs according to a preset multiple, reading out the initial question-answer pairs of the initial questions that are not determined as target questions from the initial data set, and randomly deleting the read-out initial question-answer pairs according to the preset multiple, the quantity in the initial data set is ensured to remain unchanged.

[0056] Optionally, the preset multiple can be 0.2, 0.3, 0.5, etc., which can be set independently according to requirements and are not limited in this embodiment.

[0057] Optionally, the quantity in the initial dataset can be 100 or other positive integers, which are not limited in this embodiment.

[0058] S205: Optimize the target model using the target dataset.

[0059] Specifically, in this embodiment, the target dataset is divided into a training set and a validation set, input into the target model, and fine-tuned using a preset fine-tuning algorithm. After each training round, the validation set is used to evaluate the target model until the preset evaluation metric meets the preset stopping condition, then the fine-tuning is stopped, indicating that the optimization is successful, and the adjusted target model is saved.

[0060] Optionally, the preset fine-tuning algorithm can be a low-rank adjustment algorithm, etc., which are not limited in this embodiment.

[0061] Among them, the low-rank adjustment algorithm is a parameter-efficient fine-tuning method for pre-trained language models.

[0062] Optionally, the evaluation metric can be accuracy, loss function, etc., which are not limited in this embodiment.

[0063] Optionally, the preset stopping condition can be that the accuracy meets the corresponding preset threshold, the loss function converges, etc., which are not limited in this embodiment.

[0064] Among them, before optimizing the target model, a hyperparameter optimization search algorithm can be used to set the amplification ratio of the target Q&A pairs in the preset interval, and determine the target model with the best training effect during the fine-tuning process.

[0065] Optionally, the preset interval of the amplification ratio of the target model can be [0.1, 0.3], or other range intervals, which are not limited in this embodiment.

[0066] Exemplarily, if the amplification ratio is 0.3, then the quantity of the target Q&A pairs is amplified by 1.3 times.

[0067] Specifically, when vertical domain training of the target model is required, an initial data set is obtained. The initial data set includes a plurality of initial question-and-answer pairs, and each initial question-and-answer pair includes an initial question and a corresponding initial answer. To identify the questions that the target model lacks in answering, first data corresponding to each initial question is retrieved from a preset database including data in the vertical domain corresponding to the target model, where the first data is data related to the initial question. According to each initial question-and-answer pair and the corresponding first data, and using the target model, target question-and-answer pairs are determined. The target questions in the target question-and-answer pairs are the corresponding initial questions that the target model lacks in answering. To improve the target model's ability to answer the target questions, a target data set is generated based on the target question-and-answer pairs and the initial data set. The number of target questions included in the target data set is greater than the number of initial questions corresponding to the target questions determined in the initial data set. Thus, the target model is optimized using the target data set, which increases the proportion of valid data in the target data set and improves the optimization effect of the model.

[0068] As an alternative implementation, based on the above embodiment, retrieving the first data corresponding to each initial question from the preset database includes:

[0069] Obtain the target data in the preset database;

[0070] Divide the target data into first segments of a preset length;

[0071] Perform vector representation on the first segments to obtain vectorized first segments, and determine the vectorized first segments as target segment data;

[0072] Extract keywords from the first segments and construct corresponding indexes to obtain indexed first segments, and determine the indexed first segments as target keyword data;

[0073] Perform vector representation on each initial question to obtain each initial question vector;

[0074] Based on each initial question vector, perform vector retrieval in the target segment data to determine the first preset number of second segments corresponding to each initial question;

[0075] Extract keywords from each initial question to obtain the keywords corresponding to each initial question;

[0076] Based on the keywords corresponding to each initial question, perform keyword retrieval in the target keyword data to determine the second preset number of third segments corresponding to each initial question;

[0077] Determine the second segments and third segments corresponding to each initial question as the first data corresponding to each initial question.

[0078] Among them, the target data is the data included in the preset database, including but not limited to the data in the vertical field corresponding to the target model. The first segment refers to the segment divided from the target data. The target segment data is the data including the vectorized first segment. The target keyword data is the data including the indexed first segment. The second segment is the segment retrieved from the target segment data by using the preset vector retrieval technology. The third segment is the segment retrieved from the target keyword data by using the preset keyword retrieval technology. The initial problem vector is the initial problem represented by a vector.

[0079] Optionally, the first preset quantity can be set independently according to requirements and is not limited in this embodiment.

[0080] Optionally, the second preset quantity can be set independently according to requirements and is not limited in this embodiment.

[0081] Optionally, the preset length can be set independently according to requirements and is not limited in this embodiment.

[0082] Optionally, the preset vector retrieval technology can be a similarity search algorithm, etc., and is not limited in this embodiment.

[0083] Optionally, the preset keyword retrieval technology can be a distributed search engine, etc., and is not limited in this embodiment.

[0084] Among them, the distributed search engine can perform operations such as fast full-text search, structured search, keyword search, and analysis on massive data.

[0085] Specifically, in this embodiment, the server device obtains the target data from the preset database, divides the target data into first segments of a preset length by using a preset text segmentation algorithm, and vectorizes the first segments by using a preset vectorization technology to obtain vectorized first segments, and determines the vectorized first segments as the target segment data. Moreover, a preset keyword extraction algorithm is also used to extract keywords from the first segments and construct corresponding indexes, so as to obtain indexed first segments, and the indexed first segments are determined as the target keyword data. Each initial problem is vectorized by using a preset vectorization technology to obtain each initial problem vector. Further, according to each initial problem vector and by using the preset vector retrieval technology, retrieval is respectively performed in the target segment data, and the first preset quantity of second segments with the highest similarity to each initial problem is retrieved. Moreover, a preset keyword extraction algorithm is used to extract keywords from each initial problem, so as to obtain the keywords corresponding to each initial problem. According to the keywords corresponding to each initial problem and by using the preset keyword retrieval technology, retrieval is respectively performed in the target keyword data, and the second preset quantity of third segments corresponding to each initial problem is determined. The second segments and third segments corresponding to each initial problem are determined as the first data corresponding to each initial problem.

[0086] Among them, the preset text segmentation algorithm refers to an algorithm for dividing text based on the number of characters or sentences, which can be implemented through code tools, etc., and is not limited in this embodiment.

[0087] Optionally, the preset vectorization technology can be word vector conversion (Word to Vector, Word2Vec), document vector conversion (Document to Vector, Doc2Vec), etc., and is not limited in this embodiment.

[0088] Among them, word vector conversion is a technology that can convert words into continuous vector representations. Document vector conversion is a technology for generating vector representations at the sentence, paragraph, or document level.

[0089] It can be understood that the similarity between the initial question and the first fragment can be calculated, and then a preset number of fragments with the highest similarity can be selected as the second fragment.

[0090] Optionally, the preset keyword extraction algorithm can be term frequency-inverse document frequency, etc., and is not limited in this embodiment.

[0091] Optionally, keyword extraction can also be implemented using a search engine, and is not limited in this embodiment.

[0092] Specifically, by obtaining target data from a preset database, the high relevance to the vertical domain is ensured, and the target data is divided into first fragments, which is conducive to sharded retrieval. By performing vector representation and keyword indexing on the fragments, vector retrieval and keyword retrieval of the fragments can be realized, so as to ensure semantic coverage and term accuracy, and further improve the coverage rate and accuracy of the first data.

[0093] As an alternative implementation, based on any of the above embodiments, a target model is adopted and a target question-and-answer pair is determined based on the first data and the initial question-and-answer pair, including:

[0094] Determine the first data corresponding to each initial question as the prompt information for each initial question;

[0095] Input the initial questions including the prompt information and the initial questions not including the prompt information into the target model for questioning respectively to obtain a first answer and a second answer;

[0096] Determine the target question-and-answer pair based on the first answer and the second answer.

[0097] Among them, the first answer is the answer corresponding to the initial question including the prompt information. The second answer is the answer corresponding to the initial question not including the prompt information.

[0098] Specifically, in this embodiment, the server device uses the first data corresponding to each initial question as the hint information for each initial question, and inputs the initial questions including the hint information and the initial questions without the hint information into the target model for questioning respectively, so as to obtain the first answers corresponding to the initial questions including the hint information and the second answers corresponding to the initial questions without the hint information. Then, a preset model is used to compare the effects of the first initial question answers and the second initial question answers based on the initial answers, and outputs the answers with better effects corresponding to each initial question respectively. If the second initial question answer is output by the preset model, the second initial question answer is determined as the target question according to the corresponding initial question. And the target question, the corresponding initial answer and the corresponding first data are determined as the target Q&A pair.

[0099] It can be understood that by adopting the retrieval-augmented generation technology, the first data is obtained from the preset database, and the corresponding first answer is generated by using the first data as the context information, which improves the accuracy and relevance of the model.

[0100] Among them, Retrieval Augmented Generation (RAG) is an artificial intelligence technology framework that combines information retrieval technology with generative models, aiming to improve the quality, accuracy and professionalism of the generated content.

[0101] Specifically, by using the first data as the hint information of the initial question and letting the target model answer the initial questions with and without the hint information respectively, the questions that the target model is not good at can be determined according to the effects of the answers of the target model. If the answer without the hint information is better than the answer with the hint information, it means that the target model has a good grasp of this question. If the answer without the hint information is not as good as the answer with the hint information, it means that the target model lacks a good grasp of this question, so as to achieve the accurate positioning of the target question.

[0102] As an alternative embodiment, based on any of the above embodiments, determining the target Q&A pair based on the first answer and the second answer includes:

[0103] Scoring the first answer and the second answer based on the initial answers corresponding to each initial question to obtain a first score and a second score;

[0104] Subtracting the first score and the second score to obtain a score difference;

[0105] Comparing the score difference with a preset difference threshold;

[0106] If the score difference is greater than the preset difference threshold, the initial question corresponding to the score difference is determined as the target question;

[0107] Determine the target question-and-answer pair by taking the target question, the corresponding initial answer, and the corresponding first data.

[0108] Among them, the first score is the score of the first answer corresponding to the initial question, and the second score is the score of the second answer corresponding to the initial question.

[0109] Specifically, in this embodiment, the server device uses a preset model to score the first initial question answer and the second initial question answer according to the initial answer, and outputs the scores of the first answer and the second answer in each initial question. The server device subtracts the first score from the second score to obtain a score difference, and compares the score difference with a preset difference threshold. If the score difference is greater than the preset difference threshold, it means that the first score is higher than the second score. Determine the initial question corresponding to the score difference as the target question, and determine the target question, the corresponding initial answer, and the corresponding first data as the target question-and-answer pair.

[0110] Optionally, the preset difference threshold can be set independently according to requirements and is not limited in this embodiment.

[0111] Specifically, by scoring the first answer and the second answer corresponding to each initial question according to the initial answer, the effect of the target model's answer is quantified. Then, subtract the first score and the second score corresponding to each initial question, and compare the score difference with the preset difference threshold to determine the corresponding question where the first answer is better than the second answer, thereby determining the target question, and then determining the target question-and-answer pair. Through automated comparison, the judgment efficiency is improved.

[0112] As an alternative implementation, based on any of the above embodiments, generate a target data set based on the target question-and-answer pair and the initial data set, including:

[0113] Obtain a preset multiple;

[0114] Oversample the target question-and-answer pair by a preset multiple to obtain the first question-and-answer pair;

[0115] Undersample the initial question-and-answer pairs in the initial data set that are not determined to be corresponding to the target questions by a preset multiple to obtain the second question-and-answer pair;

[0116] Determine the first question-and-answer pair and the second question-and-answer pair as the target data set.

[0117] Among them, oversampling means increasing the number of target question-and-answer pairs, and undersampling means reducing the number of initial question-and-answer pairs in the initial data set that are not determined to be corresponding to the target questions. The first question-and-answer pair refers to the target question-and-answer pair after oversampling, and the second question-and-answer pair refers to the initial question-and-answer pairs in the initial data set that are not determined to be corresponding to the target questions after undersampling.

[0118] Specifically, in this embodiment, the server device copies the question-and-answer pairs randomly selected from the target question-and-answer pairs according to a preset multiple to obtain the first question-and-answer pairs, and randomly deletes the initial question-and-answer pairs in the initial dataset that are not determined to correspond to the target questions according to the preset multiple to obtain the second question-and-answer pairs. The first question-and-answer pairs and the second question-and-answer pairs are used as the target question-and-answer pairs.

[0119] It can be understood that the target question is the corresponding initial question for which the target model's answer is lacking as determined from the initial dataset.

[0120] Exemplarily, assume that the preset multiple is 0.3, and the initial dataset includes 100 initial question-and-answer pairs, among which 50 are determined as target questions. Then the number of corresponding target question-and-answer pairs is 50; the number of initial question-and-answer pairs that are not determined to correspond to the target questions is also 50. Then the number of target question-and-answer pairs increases to 65, and the number of initial question-and-answer pairs in the initial dataset that are not determined to correspond to the target questions decreases to 35.

[0121] Among them, 65 is 50 plus the product of 50 multiplied by 0.3, and 35 is 50 minus the product of 50 multiplied by 0.3.

[0122] Optionally, the preset multiple can be 0.2, 0.3, 0.5, etc., and can be set independently according to requirements. This embodiment does not make a limitation.

[0123] Specifically, since the target questions in the target question-and-answer pairs are the questions for which the target model's answers are lacking, by oversampling the target question-and-answer pairs according to the preset multiple, and undersampling the initial question-and-answer pairs in the initial dataset that are not determined to correspond to the target questions by the preset multiple, and generating the target dataset, the number of valid data in the target dataset is increased, which can enable the target model to focus on training the lacking questions, thereby improving the optimization effect of the target model and enhancing the performance of the target model in the vertical domain.

[0124] As an alternative implementation, on the basis of any of the above embodiments, the first data includes multiple segments;

[0125] Before determining the target question, the corresponding initial answer, and the corresponding first data as the target question-and-answer pair, it further includes:

[0126] Performing deduplication processing on the second segment and the third segment in the first data corresponding to the target question to obtain the deduplicated first data;

[0127] Constructing a target segment set corresponding to each target question;

[0128] Adding the segments in the deduplicated first data corresponding to each target question to the corresponding target segment set;

[0129] Loop and execute the following operations until all the segments in the deduplicated first data corresponding to the target problem are traversed, and determine the set of target segments corresponding to the target problem as the first data of the target problem;

[0130] The operations include:

[0131] Input the deduplicated first data after removing the currently traversed segment and the corresponding target problem into the target model to obtain a third answer;

[0132] Score the third answer to obtain a third score;

[0133] Compare the third score with the first score;

[0134] If the third score is less than the first score, determine the currently traversed segment as the target segment;

[0135] If the third score is greater than or equal to the first score, delete the currently traversed segment from the set of target segments corresponding to the target problem;

[0136] Continue to traverse the next segment in the set of target segments.

[0137] Among them, the set of target segments refers to the set of the deduplicated first data of the target problem. The third answer refers to the answer obtained after inputting the deduplicated first data after removing the currently traversed segment and the corresponding target problem into the target model. The third score is the score of the third answer corresponding to the initial problem. The target segment refers to the segment whose corresponding third score is less than the first score after removing the traversed segment.

[0138] Specifically, in this embodiment, since there may be an overlap between the first segment and the second segment, the server device performs deduplication processing on the second segment and the third segment in the first data corresponding to each target question, so as to obtain the deduplicated first data corresponding to each target question, construct a target segment set corresponding to each target question, and put the deduplicated first data corresponding to each target question into the target segment set corresponding to each target question. Then, the following loop operations are sequentially performed on each target question until all the segments in the deduplicated first data corresponding to the target question are traversed, and the determined target segment set corresponding to the target question is taken as the first data of the target question. The loop operation includes: deleting the currently traversed segment, and inputting the deduplicated first data with the currently traversed segment deleted and the corresponding target question into the target model to obtain a third answer, and using a preset model to score the third answer according to the initial answer corresponding to the target question, so as to obtain a third score, and comparing the third score with the first score. If the third score is less than the first score, it means that the currently deleted segment is knowledge that the target model does not master, so the currently traversed segment is determined as the target segment. If the third score is greater than or equal to the first score, it means that the currently deleted segment is knowledge that the target model has mastered, and deleting it will not have any impact on the answer. Therefore, the currently traversed segment in the target segment set corresponding to the target question is deleted, and then the next segment in the target segment set is continued to be traversed.

[0139] Specifically, by testing the segments in the deduplicated first data corresponding to each target question, it can be determined which segments have an impact on the quality of the answer generated by the target model and which segments have no impact on the answer generated by the target model. By inputting the deduplicated first data with any segment deleted into the target model for answer output, scoring the answer, and comparing it with the first score, it can be judged which are the key segments corresponding to the target question, so as to directly retain the key segments. Thus, when optimizing the model, computing resources can be saved, but the tuning effect of the target model can also be improved.

[0140] As an optional implementation manner, on the basis of any of the above embodiments, before tuning the target model using the target data set, the method further includes:

[0141] Determine the number of target question-and-answer pairs in the target data set;

[0142] If the number of target question-and-answer pairs is less than the preset number threshold, then re-execute the step of determining the target question-and-answer pairs using the target model and based on the first data and the initial question-and-answer pairs.

[0143] Optionally, the preset number threshold can be set independently according to actual needs and is not limited in this embodiment.

[0144] Specifically, in this embodiment, before the server device tunes the target model using the target model dataset, it reads the number of target question-and-answer pairs included in the target dataset, and compares the number of target question-and-answer pairs with a preset number threshold. If the number of target question-and-answer pairs is less than the preset number threshold, it re-executes the step of using the target model and determining the target question-and-answer pairs based on the first data and the initial question-and-answer pairs, so as to re-obtain the target question-and-answer pairs until the number of target question-and-answer pairs is greater than the preset number threshold, and then tunes the target model using the target dataset.

[0145] Specifically, by determining the number of target question-and-answer pairs included in the target dataset and comparing it with the preset number threshold, it is ensured that when the target model is tuned, the target dataset used contains a sufficient number of target question-and-answer pairs, thereby improving the success rate of model tuning.

[0146] Figure 3 It is a schematic flowchart of a model tuning method provided by another embodiment of the present application, as Figure 3 shown. The model tuning method provided in this embodiment is applied to the server device. The model tuning method provided in this embodiment specifically includes the following steps:

[0147] S301: Obtain the initial dataset.

[0148] S302: Obtain the target data in the preset database, divide the target data into first segments of a preset length, and perform vectorization representation, keyword extraction, and index construction on each first segment respectively, so as to determine the target segment data and the target keyword data.

[0149] S303: Perform vectorization representation on each initial question to obtain each initial question vector.

[0150] S304: Use the preset vector retrieval technology to calculate the similarity between each initial question vector and each segment in the target segment data respectively, sort the similarities corresponding to each initial question from high to low, and determine the segments before the preset ranking as the second segments corresponding to each initial question. Use the preset keyword retrieval technology and retrieve from the target keyword data according to the keywords corresponding to each initial question respectively to determine the third segments of the second preset number corresponding to each initial question. Determine the second segments and the third segments corresponding to each initial question as the first data corresponding to each initial question.

[0151] Among them, the number of segments before the preset ranking can be the first preset number.

[0152] S305: Use the first data corresponding to each initial question as the hint information for each initial question. Input the initial questions including the hint information and the initial questions without the hint information into the target model for questioning respectively, so as to obtain the first answer and the second answer.

[0153] Among them, the first answer is the answer corresponding to the initial question including the hint information; the second answer is the answer corresponding to the initial question without the hint information.

[0154] S306: Input the initial answer, the first answer and the second answer into the preset model. Use the preset model to compare the effects of the first answer and the second answer according to the initial answer, and output the answer with the best effect corresponding to each initial question. If the first answer is output by the preset model, determine the first answer as the target question according to the corresponding initial question. And determine the target question, the corresponding initial answer and the corresponding first data as the target Q&A pair.

[0155] Optionally, the preset model can be an inference model, etc., which is not limited in this embodiment.

[0156] S307: Remove duplicates from the second segment and the third segment in the first data corresponding to the target question, so as to obtain the deduplicated first data.

[0157] S308: Construct the target segment set corresponding to each target question, and add the segments in the deduplicated first data corresponding to each target question to the corresponding target segment set. Execute the preset loop, obtain the target segment set corresponding to each target question after executing the loop, and determine the target segment set corresponding to each target question after executing the loop as the first data of the target question.

[0158] S309: Oversample the target Q&A pair by a preset multiple to obtain the first Q&A pair, and undersample the initial Q&A pair in the initial data set that is not determined to correspond to the target question by a preset multiple to obtain the second Q&A pair. Determine the first Q&A pair and the second Q&A pair as the target data set.

[0159] S310, determine the number of target Q&A pairs in the target data set.

[0160] S311, determine whether the number of target Q&A pairs is less than the preset number threshold.

[0161] S312, if the number of target Q&A pairs is greater than or equal to the preset number threshold, use the target data set to optimize the target model.

[0162] Among them, if the number of target Q&A pairs is less than the preset number threshold, re - execute from the steps corresponding to S306.

[0163] Through the description of the above embodiments, those skilled in the art can clearly understand that the method according to the above embodiments can be implemented by means of software plus a necessary general hardware platform. Of course, it can also be implemented by hardware, but in many cases, the former is a better implementation method.

[0164] Figure 4 This is a schematic structural diagram of a model tuning device provided in an embodiment of the present application. As Figure 4 shown, the execution subject of the above model tuning method is a model tuning device, and this model tuning device can be implemented through a computer program; it can also be implemented through a medium storing relevant computer programs, such as a USB flash drive and / or an optical disc, etc. Or, it can also be implemented through an entity device integrated or installed with relevant computer programs, such as an electronic device. The electronic device can be a computer or a server device, etc. If the model tuning device provided in this embodiment is located in an electronic device, then the model tuning device 40 provided in this embodiment includes: an acquisition module 41, a retrieval module 42, a determination module 43, a generation model 44, and a tuning module 45.

[0165] Specifically, the acquisition module 41 is used to acquire an initial data set; the retrieval module 42 is used to retrieve first data corresponding to each initial question from a preset database; the first data is data related to the initial question; the determination module 43 is used to use a target model and based on the first data and the initial question-and-answer pair to determine a target question-and-answer pair; the target question-and-answer pair includes a target question, a corresponding initial answer, and corresponding first data; the target question refers to the initial question that the target model lacks in answering; the generation module 44 is used to generate a target data set based on the target question-and-answer pair and the initial data set; the number of target questions included in the target data set is greater than the number of initial questions corresponding to the target questions determined in the initial data set; the tuning module 45 is used to tune the target model using the target data set.

[0166] Optionally, when retrieving the first data corresponding to each initial question from the preset database, the retrieval module 42 is specifically configured to: obtain the target data in the preset database; divide the target data into first segments of a preset length; perform vector representation on the first segments to obtain vectorized first segments, and determine the vectorized first segments as target segment data; extract keywords from the first segments and construct corresponding indexes to obtain indexed first segments, and determine the indexed first segments as target keyword data; perform vector representation on each initial question respectively to obtain each initial question vector; perform vector retrieval on each initial question vector in the target segment data respectively to determine the first preset number of second segments corresponding to each initial question; extract keywords from each initial question to obtain the keywords corresponding to each initial question; perform keyword retrieval on each initial question corresponding to the keywords in the target keyword data respectively to determine the second preset number of third segments corresponding to each initial question; and determine the second segments and third segments corresponding to each initial question as the first data corresponding to each initial question.

[0167] Optionally, when determining the target Q&A pairs by using the target model and based on the first data and the initial Q&A pairs, the determination module 43 is configured to determine the first data corresponding to each initial question as the prompt information for each initial question; input the initial questions including the prompt information and the initial questions not including the prompt information into the target model for questioning respectively to obtain a first answer and a second answer; and determine the target Q&A pairs based on the first answer and the second answer.

[0168] Optionally, when determining the target Q&A pairs based on the first answer and the second answer, the determination module 43 is configured to score the first answer and the second answer based on the initial answer corresponding to each initial question to obtain a first score and a second score; subtract the second score from the first score to obtain a score difference; compare the score difference with a preset difference threshold; if the score difference is greater than the preset difference threshold, determine the initial question corresponding to the score difference as the target question; and determine the target question, the corresponding initial answer, and the corresponding first data as the target Q&A pair.

[0169] Optionally, when generating the target data set based on the target Q&A pairs and the initial data set, the generation model 44 is configured to obtain a preset multiple; perform oversampling on the target Q&A pairs by the preset multiple to obtain first Q&A pairs; perform undersampling on the initial Q&A pairs in the initial data set that are not determined to be corresponding to the target questions by the preset multiple to obtain second Q&A pairs; and determine the first Q&A pairs and the second Q&A pairs as the target data set. Optionally, the first data includes multiple segments.

[0170] Wherein, the model tuning device further includes a processing module, a construction model, an adding module, and an execution module.

[0171] Accordingly, before the processing module determines the target question, the corresponding initial answer, and the corresponding first data as the target Q&A pair, it performs deduplication processing on the second segment and the third segment in the first data corresponding to the target question to obtain the deduplicated first data; the construction module is used to construct a target segment set corresponding to each target question; the addition module is used to add the segments in the deduplicated first data corresponding to each target question to the corresponding target segment set; the execution module is used to loop through the following operations until all the segments in the deduplicated first data corresponding to the target question are traversed, and determine the output target segment set corresponding to the target question as the first data of the target question; the operations include: inputting the deduplicated first data with the currently traversed segment removed and the corresponding target question into the target model to obtain a third answer; scoring the third answer to obtain a third score; comparing the third score with the first score; if the third score is less than the first score, determining the currently traversed segment as the target segment; if the third score is greater than or equal to the first score, deleting the currently traversed segment from the target segment set corresponding to the target question; and continuing to traverse the next segment in the target segment set.

[0172] Optionally, the determination module 43 is used to determine the number of target Q&A pairs in the target data set before tuning the target model using the target data set; the execution module is used to, if the number of target Q&A pairs is less than the preset number threshold, re-execute the step of determining the target Q&A pair using the target model and based on the first data and the initial Q&A pair.

[0173] For the description of the features in the corresponding embodiments of the model tuning device, reference can be made to the relevant descriptions in the corresponding embodiments of the model tuning method, which will not be elaborated here one by one.

[0174] Figure 5 This is a schematic structural diagram of an electronic device provided in an embodiment of the present application. As Figure 5 shown, the electronic device 50 provided in the embodiment of the present application includes: a memory 52 and a processor 51.

[0175] The memory 52 stores a computer program, and the processor 51 is configured to run the computer program to execute the steps in any of the above-mentioned embodiments of the model tuning method.

[0176] The specific implementation process of the processor 51 can be referred to in the above method embodiment, and its implementation principle and technical effect are similar, which will not be elaborated here in this embodiment.

[0177] In the above embodiments, it should be understood that the processor may be a central processing unit (CPU), or may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. The steps of the method disclosed in combination with the application can be directly implemented by a hardware processor, or can be implemented by a combination of hardware and software modules in the processor.

[0178] The memory may include a random access memory (RAM), and may also include a non-volatile memory (NVM), such as at least one disk memory.

[0179] The bus may be an industry standard architecture (ISA) bus, a peripheral component interconnect (PCI) bus, an extended industry standard architecture (EISA) bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For the convenience of representation, the bus in the drawings of the present application is not limited to only one bus or one type of bus.

[0180] An embodiment of the present application also provides a computer-readable storage medium, in which a computer program is stored, and the computer program is configured to execute the steps in any of the above model tuning method embodiments when running.

[0181] In an exemplary embodiment, the above computer-readable storage medium may include, but is not limited to: USB flash drives, read-only memories (ROMs), random access memories (RAMs), mobile hard disks, magnetic disks, or optical discs, etc., various media that can store computer programs.

[0182] An embodiment of the present application also provides a computer program product, the above computer program product includes a computer program, and when the computer program is executed by a processor, it implements the steps in any of the above model tuning method embodiments.

[0183] Embodiments of the present application also provide another computer program product, including a non-volatile computer-readable storage medium storing a computer program, which, when executed by a processor, implements the steps in any of the above-described method embodiments for model tuning.

[0184] Those skilled in the art can further realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Skilled professionals can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present application.

[0185] The above has introduced in detail a method for displaying device information provided by the present application. Specific examples are used herein to elaborate on the principle and implementation manner of the present application. The description of the above embodiments is only used to help understand the method and its core idea of the present application. It should be noted that for those of ordinary skill in the art in the technical field, without departing from the principle of the present application, several improvements and modifications can be made to the present application, and these improvements and modifications also fall within the protection scope of the claims of the present application.

Claims

1. A model tuning method, characterized in that, The method includes: Obtaining an initial data set; the initial data set includes a plurality of initial question-and-answer pairs; the initial question-and-answer pairs include an initial question and a corresponding initial answer; Retrieving first data corresponding to each of the initial questions from a preset database; the first data is data related to the initial question; Using a target model and determining target question-and-answer pairs based on the first data and the initial question-and-answer pairs; the target question-and-answer pairs include a target question, the corresponding initial answer, and the corresponding first data; the target question refers to the initial question for which the target model's answer is lacking; Generating a target data set based on the target question-and-answer pairs and the initial data set; the number of target questions included in the target data set is greater than the number of initial questions determined as the target questions in the initial data set; Using the target data set to optimize the target model.

2. The model tuning method according to claim 1, wherein The retrieving first data corresponding to each of the initial questions from a preset database includes: Obtaining target data in the preset database; Dividing the target data into first segments of a preset length; Performing vector representation on the first segments to obtain vectorized first segments, and determining the vectorized first segments as target segment data; Extracting keywords from the first segments and constructing corresponding indexes to obtain indexed first segments, and determining the indexed first segments as target keyword data; Performing vector representation on each of the initial questions to obtain initial question vectors; Performing vector retrieval on each of the initial question vectors in the target segment data to determine a first preset number of second segments corresponding to each of the initial questions; Extracting keywords from each of the initial questions to obtain keywords corresponding to each of the initial questions; Performing keyword retrieval on each of the keywords corresponding to the initial questions in the target keyword data to determine a second preset number of third segments corresponding to each of the initial questions; Determining the second segments and the third segments corresponding to each of the initial questions as the first data corresponding to each of the initial questions.

3. The model tuning method according to claim 2, wherein The using a target model and determining target question-and-answer pairs based on the first data and the initial question-and-answer pairs includes: Determining the first data corresponding to each of the initial questions as prompt information for each of the initial questions; Respectively inputting the initial questions including the prompt information and the initial questions not including the prompt information into the target model for questioning to obtain a first answer and a second answer; Determining the target question-and-answer pairs based on the first answer and the second answer.

4. The model tuning method according to claim 3, wherein The determining the target question-and-answer pairs based on the first answer and the second answer includes: Scoring the first answer and the second answer based on the initial answer corresponding to each of the initial questions to obtain a first score and a second score; Subtracting the second score from the first score to obtain a score difference; Comparing the score difference with a preset difference threshold; If the scoring difference is greater than the preset difference threshold, determine the initial question corresponding to the scoring difference as the target question; Determine the target question, the corresponding initial answer, and the corresponding first data as the target Q&A pair.

5. The model tuning method according to claim 1, wherein Generating the target data set based on the target Q&A pair and the initial data set includes: Obtain a preset multiple; Oversample the target Q&A pair by the preset multiple to obtain the first Q&A pair; Undersample the initial Q&A pairs in the initial data set that are not determined to be corresponding to the target question by the preset multiple to obtain the second Q&A pair; Determine the first Q&A pair and the second Q&A pair as the target data set.

6. The model tuning method according to claim 4, wherein The first data includes multiple segments; Before determining the target question, the corresponding initial answer, and the corresponding first data as the target Q&A pair, the method further includes: Perform deduplication processing on the second segment and the third segment in the first data corresponding to the target question to obtain the deduplicated first data; Construct a target segment set corresponding to each target question; Add the segments in the deduplicated first data corresponding to each target question to the corresponding target segment set; Loop through the following operations until all the segments in the deduplicated first data corresponding to the target question are traversed, and determine the output target segment set corresponding to the target question as the first data of the target question; The operations include: Input the deduplicated first data with the currently traversed segment removed and the corresponding target question into the target model to obtain a third answer; Score the third answer to obtain a third score; Compare the third score with the first score; If the third score is less than the first score, determine the currently traversed segment as the target segment; If the third score is greater than or equal to the first score, delete the currently traversed segment in the target segment set corresponding to the target question; Continue to traverse the next segment in the target segment set.

7. The model tuning method according to claim 1, wherein Before tuning the target model using the target data set, the method further includes: Determine the number of target Q&A pairs in the target data set; If the number of target Q&A pairs is less than the preset number threshold, re-execute the step of using the target model and determining the target Q&A pair based on the first data and the initial Q&A pair.

8. A model tuning device, characterized in that, Includes: An acquisition module for acquiring an initial data set; The initial data set includes multiple initial Q&A pairs; Each initial Q&A pair includes an initial question and a corresponding initial answer; A retrieval module for retrieving the first data corresponding to each initial question from a preset database; the first data is data related to the initial question; A determination module for using a target model and determining a target Q&A pair based on the first data and the initial Q&A pair; Each target Q&A pair includes a target question, the corresponding initial answer, and the corresponding first data; The target problem refers to the corresponding initial problem that the target model lacks in answering; A generation module, configured to generate a target data set based on the target question-answer pair and the initial data set; the number of target questions included in the target data set is greater than the number of initial questions determined as the corresponding initial questions of the target questions in the initial data set; A tuning module, configured to tune the target model by using the target data set.

9. An electronic device, characterized in that, Comprising: A memory, configured to store a computer program; A processor, configured to implement the steps of the model tuning method according to any one of claims 1 to 7 when executing the computer program.

10. A computer-readable storage medium, characterized in that, A computer program is stored in the computer-readable storage medium, wherein the computer program implements the steps of the model tuning method according to any one of claims 1 to 7 when executed by a processor.

Citation Information

Patent Citations

  • Answer generation method based on multi-layer Transformer aggregation encoder

    CN110502627A

  • Question and answer model optimization method and device, computer equipment and storage medium

    CN111078853A

  • Classification model training method and device, electronic equipment and storage medium

    CN112560912A

  • Obtaining method and device of intention recognition model, electronic equipment and storage medium

    CN118467730A

  • Information search method and device, equipment, storage medium and program product

    CN119807368A

Cited By

  • Model illusion relieving method, program product, equipment and medium

    CN121835795A