Acquisition question and answer method and device based on model fine tuning, medium and product
By combining pre-trained models with fine-tuned models, and utilizing structured preprocessing and secondary processing to optimize fine-tuning model parameters, the problems of high cost, long training time, and low accuracy in bank acquiring question-and-answer models were solved, achieving efficient and flexible model training and maintenance.
Patent Information
- Application Number
- CN202511110962.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-08
- Publication Date
- 2025-09-23
AI Technical Summary
After replacing the pre-trained model, the bank's collection question-and-answer model has high training costs, long time, low efficiency, high maintenance difficulty, and low model accuracy.
A model fine-tuning method is adopted. By combining pre-trained models and fine-tuning models, a payment acquisition question-answering model is constructed using general training sample sets and special training sample sets for payment acquisition scenarios. Structured preprocessing and secondary processing are performed, and the parameters of the fine-tuning model are optimized to improve model adaptability and accuracy.
The cost and time of training the acquiring question-and-answer model after replacing the pre-trained model are significantly reduced, the efficiency of model training and the flexibility of maintenance are improved, and the adaptability and accuracy of the model to the acquiring question-and-answer scenario are enhanced.
Smart Images

Figure CN120687578A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the fields of artificial intelligence and financial technology, and in particular to a payment acquisition question-answering method, device, medium, and product based on model fine-tuning. Background Art
[0002] As the bank's acquiring system continues to operate, a large number of problems have accumulated in the system. Due to limited staff, users often cannot get timely answers, resulting in a backlog of questions and affecting the user experience.
[0003] With the continuous development of big model technology, chat question-and-answer models based on big models are gaining popularity among users. The banking industry is also actively leveraging big models, using open source big models to build expert models that meet the needs of acquiring question-and-answer services, thereby improving service efficiency and quality.
[0004] However, large models are updated at an extremely rapid pace, resulting in a significant performance gap between old and new models. In related technologies, after replacing a large model in an expert model with a new pre-trained large model, the expert model often needs to undergo a lengthy retraining process to meet the needs of the acquiring Q&A business. This consumes a significant amount of time and resources, increases the cost and time of training the acquiring Q&A model after replacing the pre-trained model, reduces model training efficiency, and increases the difficulty of model maintenance. Summary of the Invention
[0005] The present invention provides a method, device, medium and product for acquiring question and answer based on model fine-tuning to solve the problems of high cost, long time, low efficiency and high maintenance difficulty in training the acquiring question and answer model after replacing the pre-trained model, as well as the problem of low accuracy of the model acquiring question and answer.
[0006] According to one aspect of an embodiment of the present invention, a method for acquiring Q&A based on model fine-tuning is provided, comprising:
[0007] Obtain the acquiring information to be queried and perform structured preprocessing on the acquiring information to obtain structured acquiring information;
[0008] Input structured acquiring information into the acquiring question-answering model;
[0009] The acquiring question-and-answer model includes a pre-trained model and a fine-tuned model, which are connected end to end. The pre-trained model is trained using a general training sample set, and the fine-tuned model is trained based on the pre-trained model and a special training sample set for the acquiring scenario.
[0010] The structured acquiring information is processed once through the pre-trained model, and the general processing results are provided to the fine-tuning model;
[0011] By fine-tuning the model, the received general data results are processed again to obtain the target answer corresponding to the structured payment information.
[0012] According to another aspect of an embodiment of the present invention, a method for fine-tuning an acquiring question-and-answer model is provided, comprising:
[0013] Select and set a pre-trained model, and build a fine-tuning model based on the pre-trained model and the preset reward function. Then, connect the pre-trained model and the fine-tuning model end to end to obtain the expert model to be trained;
[0014] Generate training samples and test samples based on historical structured acquiring information and corresponding historical answers;
[0015] A cross-entropy loss function is constructed based on the reward function of the fine-tuning model. With the goal of minimizing the cross-entropy loss function, the model parameters of the fine-tuning model are iteratively optimized multiple times based on the training samples to obtain an alternative expert model.
[0016] When it is determined that the prediction accuracy of the alternative expert model for the test sample meets the prediction performance requirements, the alternative expert model is determined as the acquisition question and answer model.
[0017] According to another aspect of an embodiment of the present invention, a payment acquisition question-answering device based on model fine-tuning is provided, comprising:
[0018] The structuring module is used to obtain the acquiring information to be queried and perform structured preprocessing on the acquiring information to obtain structured acquiring information;
[0019] Input module, used to input structured acquiring information into the acquiring question-answering model;
[0020] The acquiring question-and-answer model includes a pre-trained model and a fine-tuned model, which are connected end to end. The pre-trained model is trained using a general training sample set, and the fine-tuned model is trained based on the pre-trained model and a special training sample set for the acquiring scenario.
[0021] The one-time processing module is used to process the structured acquiring information once through the pre-trained model, and obtain the general processing results and provide them to the fine-tuning model;
[0022] The secondary processing module is used to perform secondary processing on the received general data results by fine-tuning the model to obtain the target answer corresponding to the structured payment information.
[0023] According to another aspect of an embodiment of the present invention, a device for fine-tuning a merchant acquisition question-and-answer model is provided, comprising:
[0024] The end-to-end connection module is used to select a pre-trained model, build a fine-tuning model based on the pre-trained model and a preset reward function, and then connect the pre-trained model and the fine-tuning model end-to-end to obtain the expert model to be trained;
[0025] The sample generation module is used to generate training samples and test samples based on historical structured acquiring information and corresponding historical answers;
[0026] The alternative module is used to construct a cross-entropy loss function based on the reward function of the fine-tuning model, and to minimize the cross-entropy loss function. Based on the training samples, the model parameters of the fine-tuning model are optimized multiple times to obtain an alternative expert model;
[0027] The determination module is used to determine the alternative expert model as the acquiring question-and-answer model when it is determined that the prediction accuracy of the alternative expert model for the test sample meets the prediction performance requirements.
[0028] According to another aspect of an embodiment of the present invention, an electronic device is provided, the electronic device comprising:
[0029] At least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor so as to enable the at least one processor to execute the acquisition question-and-answer method based on model fine-tuning or the acquisition question-and-answer model fine-tuning method described in any embodiment of the present invention.
[0030] According to another aspect of an embodiment of the present invention, a computer-readable storage medium is provided, wherein the computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a processor to implement the model fine-tuning-based acquisition question-and-answer method or the acquisition question-and-answer model fine-tuning method described in any embodiment of the present invention when executed.
[0031] According to another aspect of an embodiment of the present invention, a computer program product is provided, including a computer program, wherein when the computer program is executed by a processor, the steps of the method according to any embodiment of the present invention are implemented.
[0032] The technical solution of the embodiments of the present invention obtains the acquiring information to be queried and pre-processes it to obtain structured acquiring information; then inputs the structured acquiring information into an acquiring question-and-answer model; processes the structured acquiring information once using a pre-trained model to obtain a general processing result, which is then provided to a fine-tuned model; and then performs secondary processing on the received general data result using the fine-tuned model to obtain a target answer corresponding to the structured acquiring information. By training the pre-trained model based on a general sample set and the fine-tuned model based on a dedicated acquiring sample set, the adaptability of the acquiring question-and-answer model to acquiring question-and-answer scenarios is improved, significantly reducing the cost and time of training the acquiring question-and-answer model after replacing the pre-trained model, and improving the efficiency of model training and the flexibility of model maintenance. The acquiring information to be queried and pre-processed to obtain structured acquiring information is then input into an acquiring question-and-answer model consisting of a pre-trained model and a fine-tuned model connected end-to-end. After initial processing by the pre-trained model, a general result is obtained, which is then passed to the fine-tuned model for secondary processing, ultimately outputting the target answer, thereby improving the accuracy of acquiring question-and-answer.
[0033] It should be understood that the content described in this section is not intended to identify the key or important features of the embodiments of the present invention, nor is it intended to limit the scope of the present invention. Other features of the present invention will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0034] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0035] Figure 1 This is a flowchart of a Q&A method for acquiring based on model fine-tuning according to the first embodiment of the present invention;
[0036] Figure 2 This is a flowchart of a method for fine-tuning an acquiring question-and-answer model according to the second embodiment of the present invention;
[0037] Figure 3 This is a flowchart of another method for fine-tuning an acquiring question-and-answer model provided in accordance with the third embodiment of the present invention;
[0038] Figure 4 2 is a schematic diagram of the structure of a Q&A device for acquiring based on model fine-tuning according to a fourth embodiment of the present invention;
[0039] Figure 5 2. This is a schematic diagram of the structure of a fine-tuning device for an acquiring question-and-answer model according to a fifth embodiment of the present invention;
[0040] Figure 6 It is a structural diagram of an electronic device for implementing the acquiring question-and-answer method based on model fine-tuning or the acquiring question-and-answer model fine-tuning method according to an embodiment of the present invention. DETAILED DESCRIPTION
[0041] In order to enable those skilled in the art to better understand the solutions of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the embodiments described are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of the present invention.
[0042] It should be noted that the terms "first", "second", etc. in the description and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that the numbers used in this way can be interchanged where appropriate so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0043] Example 1
[0044] Figure 1 This is a flow chart of a method for acquiring question and answer based on model fine-tuning provided in the first embodiment of the present invention. This embodiment is applicable to the case of acquiring question and answer based on the model obtained by model fine-tuning. The method can be executed by an acquiring question and answer device based on model fine-tuning. The acquiring question and answer device based on model fine-tuning can be implemented in the form of hardware and / or software and can generally be configured in an electronic device. Figure 1 As shown, the method includes:
[0045] S110: Acquire the acquiring information to be questioned and answered, perform structured preprocessing on the acquiring information, and obtain structured acquiring information.
[0046] In the embodiments of the present invention, acquiring information can be specifically understood as various data and problem descriptions generated when users interact with a bank's acquiring system. In acquiring services, banks typically provide payment devices, payment gateways, transaction processing, and settlement services to help merchants complete payment transactions with consumers. For example, acquiring information may include information such as user inquiries about reasons for failed bank transactions, details of rate changes, or system operation instructions. While acquiring information typically exists in text format, it may also include images and documents.
[0047] Specifically, the acquiring information to be queried is obtained and structured preprocessed. For example, text information is preprocessed through methods such as word segmentation, removal of meaningless words, stemming, and identification of keywords in the text. Images are preprocessed through methods such as image enhancement, cropping and scaling, and optical character recognition to extract text information from the images and convert it into text format. Documents are preprocessed through methods such as document parsing, page segmentation, table extraction, and optical character recognition to extract document content and convert it into text format. The text format information is then converted into a preset vector form that matches the model input format, thereby obtaining structured acquiring information.
[0048] Optionally, based on the above embodiments, annotations corresponding to the text content may be added to the structured payment information, such as intent annotation and sentiment annotation. Specifically, intent annotation can be understood as: annotating the user's question intention, such as inquiring about the transaction status or applying for a refund. Specifically, sentiment annotation can be understood as: identifying and annotating the emotional tendencies contained in the text based on modal particles, question frequency or keywords in the text, such as complaints or consultations. Through intent annotation, the model can quickly locate problems and generate accurate answers, reducing the communication costs between users and the system, and improving problem-solving efficiency and user experience. Through sentiment annotation, the model can perceive the user's emotional state, thereby adjusting the answer strategy and tone according to the emotion type, and improving user satisfaction with the system.
[0049] S120. Input the structured acquiring information into the acquiring question and answer model.
[0050] Among them, the acquiring question and answer model includes: a pre-trained model and a fine-tuning model connected end to end. The pre-trained model is trained using a general training sample set, and the fine-tuning model is trained based on the pre-trained model and a special training sample set in the acquiring scenario.
[0051] In the embodiment of the present invention, the general training sample set can be specifically understood as: a large-scale data set including different fields and topics. The pre-training model is trained through the general training sample set, thereby obtaining the ability to understand general questions and generate corresponding answers. The acquiring scenario can be specifically understood as: a specific application environment related to the bank's acquiring business with specific business rules, terminology and common problem types, which may include specific scenarios such as merchant transaction processing, fee settlement and customer service. The dedicated training sample set can be specifically understood as: a data set collected and organized specifically for the acquiring scenario, which includes a large number of question and answer pairs related to the acquiring business, which are used to train and fine-tune the model so that the acquiring question and answer model can adapt to the specific needs of the acquiring scenario. The dedicated training sample set can be specifically constructed based on resources such as historical user consultation records, acquiring question and answer business documents, and expert answers to acquiring questions.
[0052] A pre-trained model can be understood as a large model, such as a large language model, pre-trained using a general training sample set. This model possesses the general ability to understand general text and generate general responses. A fine-tuned model can be understood as a model that, based on the general responses generated by the pre-trained model, is further trained on a specialized training sample set and then specifically trained for the Q&A task of acquiring customers. This transforms the general capabilities of the pre-trained model into specialized capabilities for solving problems in the acquiring customer Q&A domain. The form of a fine-tuned model can be selected based on the specific task and application scenario, such as a machine learning model, a neural network model, or a custom, rule-based fine-tuned model. The acquiring customer Q&A model can be understood as an expert model trained based on acquiring customer domain knowledge. Specifically, an expert model can be understood as a model that, based on a pre-trained model, is fine-tuned using acquired customer domain knowledge to achieve better performance in the acquiring customer domain by generating answers that are specific to the task.
[0053] S130: Process the structured acquiring information once through the pre-trained model to obtain a general processing result and provide it to the fine-tuning model.
[0054] Specifically, the pre-trained model performs preliminary processing on the structured acquiring information, and uses the general knowledge and feature extraction capabilities of the pre-trained model to extract general features from the structured acquiring information as general processing results, and provides them to the fine-tuning model.
[0055] S140: Perform secondary processing on the received general data results by fine-tuning the model to obtain a target answer corresponding to the structured payment information.
[0056] Specifically, the fine-tuning model further processes the general processing results output by the pre-trained model. Leveraging the fine-tuning model's specialized knowledge and feature extraction capabilities in the Q&A domain, it extracts and enhances features relevant to the Q&A task. For example, it maps general data results to specific question types and enhances features relevant to those question types. Based on these secondary processed features, it generates a target answer related to the structured Q&A information and presents it to the user.
[0057] The technical solution of the embodiments of the present invention obtains the acquiring information to be queried and pre-processes it to obtain structured acquiring information; then inputs the structured acquiring information into an acquiring question-and-answer model; processes the structured acquiring information once using a pre-trained model to obtain a general processing result, which is then provided to a fine-tuned model; and then performs secondary processing on the received general data result using the fine-tuned model to obtain a target answer corresponding to the structured acquiring information. By training the pre-trained model based on a general sample set and the fine-tuned model based on a dedicated acquiring sample set, the adaptability of the acquiring question-and-answer model to acquiring question-and-answer scenarios is improved, significantly reducing the cost and time of training the acquiring question-and-answer model after replacing the pre-trained model, and improving the efficiency of model training and the flexibility of model maintenance. The acquiring information to be queried and pre-processed to obtain structured acquiring information is then input into an acquiring question-and-answer model consisting of a pre-trained model and a fine-tuned model connected end-to-end. After initial processing by the pre-trained model, a general result is obtained, which is then passed to the fine-tuned model for secondary processing, ultimately outputting the target answer, thereby improving the accuracy of acquiring question-and-answer.
[0058] Example 2
[0059] Figure 2 This is a flow chart of a method for fine-tuning a merchant acquisition question-answering model provided in the second embodiment of the present invention. This embodiment is applicable to the case of training a merchant acquisition question-answering model based on a fine-tuning model. The method can be executed by a fine-tuning device for the merchant acquisition question-answering model. The fine-tuning device for the merchant acquisition question-answering model can be implemented in the form of hardware and / or software and can generally be configured in an electronic device. Figure 2 As shown, the method includes:
[0060] S210 , selecting and setting a pre-trained model, and constructing a fine-tuning model based on the pre-trained model and a preset reward function, and then connecting the pre-trained model and the fine-tuning model end to end to obtain an expert model to be trained.
[0061] In the embodiments of the present invention, a reward function can be specifically understood as a function used to measure performance indicators such as the degree of match, relevance, and accuracy of the answers generated by the model with the true standard answers, and to guide the fine-tuning of the model training process. The reward function can specifically take the form of a reward function based on output probability, a reward function based on output confidence, a reward function based on reinforcement learning, or a custom reward function based on business objectives.
[0062] Specifically, the appropriate pre-trained model can be determined based on the specific requirements of the Q&A task in the field of acquiring transactions, the characteristics of the data, and the performance of pre-trained models that have been trained on large-scale general datasets. For example, the latest version of the open-source pre-trained large language model can be selected as the pre-trained model. A fine-tuning model is constructed based on the pre-trained model and the preset reward function. For example, one or more fully connected layers, classification layers, or other types of neural network layers are added after the output layer of the pre-trained model, and the preset reward function is used to guide the training of the fine-tuning model, thereby expanding on the basis of the pre-trained model so that the generation of answers meets the needs of the Q&A task in acquiring transactions. The model obtained by connecting the pre-trained model and the fine-tuning model end to end is used as the expert model to be trained.
[0063] Optionally, based on the above embodiments, constructing a fine-tuning model according to the pre-trained model and a preset reward function may include:
[0064] Obtaining the model output probability of the pre-trained model, calculating the numerical relationship between the input and output of the fine-tuning model based on the model output eigenvalue, the exponential term of the reward function, and the normalization factor, and constructing the fine-tuning model based on the calculated numerical relationship;
[0065] The reward exponent term is a power operation with a natural constant as the base and the reward function divided by the regularization parameter as the exponent; the normalization factor is the sum of the products of each output probability of the pre-trained model multiplied by the reward exponent term.
[0066] In the embodiment of the present invention, the model output probability of the pre-trained model can be specifically understood as: the probability of the output prediction of the pre-trained model under the condition of a given input, for example, π pt (y|x), that is, the probability of the pre-trained model outputting y given the input x. The reward exponential term can be understood as: taking the natural constant e as the base and the reward function r as the base. θ (x,y) divided by the regularization parameter λ as the exponent of the power operation, that is, exp(r θ (x,y) / λ), which is used to amplify the influence of the reward function so that the output of high reward has a higher probability in the fine-tuning model, that is, to convert the reward function value into a probability value in the probability space. Among them, the reward function r θ (x,y) is a function with θ as a parameter, which is used to measure the quality of the model output when the input is x and the output is y.
[0067] Normalization factor Z θ (x) is the sum of the products of the output probabilities of the pre-trained model and the reward exponential term, that is: Z θ (x)=∑ y π pt(y|x)exp(r θ (x,y) / λ), by calculating the weighted sum of the rewards for all possible outputs y under a given input x, where the weight is the model output probability of the pre-trained model. Adding all weighted exponential terms ensures that the sum of the probability distribution is 1, thereby ensuring that the fine-tuned model is a valid probability distribution.
[0068] It is understandable that the fine-tuning method for the acquiring question-and-answer model proposed in the embodiment of the present invention does not directly optimize the pre-trained model itself during model training, but rather trains the fine-tuned model to minimize the cross-entropy loss between the fine-tuned model and the traditional fine-tuning method. To meet the above requirements, the fine-tuning model used in the embodiment of the present invention is a model that maximizes the explicit reward under the Kulbak-Leibler regularization condition of the pre-trained model. The Kulbak-Leibler regularization condition is a technique that prevents model overfitting by adding a Kulbak-Leibler divergence term to the loss function.
[0069] Accordingly, while keeping the fine-tuning model π θ (y|x) and pre-trained model π pt Under the premise that (y|x) is sufficiently close, maximizing the expected reward can be achieved through the following objective function:
[0070] Among them, π θ (y|x) is the probability of outputting y when the fine-tuned model with θ as the parameter is input x, D KL (π θ (·|x)||π pt (·|x) is the Kulbak-Leibler divergence between the fine-tuned model and the pre-trained model when the input is x, and λ is the regularization parameter that controls the trade-off between the reward term and the Kulbak-Leibler divergence term. The analytical solution of the Kulbak-Leibler divergence formula is: It represents the average reward output by the fine-tuned model under input x, that is, the expected reward.
[0071] Accordingly, the fine-tuning model for reward maximization under Kulbak-Leibler regularization can be viewed as solving the following equation:
[0072] In the above steps, a higher expected reward indicates that the fine-tuned model's output is more consistent with the task requirements. The Kulbak-Leibler divergence term ensures that the difference between the fine-tuned model and the pre-trained model is not too large, thereby maintaining model stability and consistency. By maximizing the expected reward and minimizing the Kulbak-Leibler divergence, the fine-tuned model can maximize the value of the reward function while maintaining close proximity to the pre-trained model.
[0073] Specifically, obtain the model output probability π of the pre-trained model pt (y|x), according to the model output eigenvalue y, the exponential term exp(r θ (x,y) / λ) and the normalization factor Z θ (x), calculate the numerical relationship between the input and output of the fine-tuning model And according to the calculated numerical relationship, a fine-tuning model is constructed. θ (y|x) is the fine-tuned model with θ as the parameter, and the probability of outputting y when the input is x.
[0074] A fine-tuning model is constructed by combining the output probability of the pre-trained model with the reward function and normalization factor. The output probability of the pre-trained model provides general feature extraction and representation capabilities. The reward exponent term and normalization factor guide the fine-tuning model to focus on high-reward outputs based on the specific requirements of the Q&A task. This enables the fine-tuning model to adapt to the task requirements while maintaining the universality of the pre-trained model, thereby improving the accuracy and relevance of the Q&A task.
[0075] S220: Generate training samples and test samples based on the historical structured acquiring information and the corresponding historical answers.
[0076] In the embodiments of the present invention, historical structured acquiring information can be specifically understood as historically collected acquiring data that has been pre-processed and formatted in a unified format, including user questions and background information. Historical answers can be specifically understood as answers that accurately respond to historical questions, corresponding to the historical structured acquiring information.
[0077] Specifically, the historical structured payment information and the corresponding historical answers are combined into samples, and the combined samples can be divided into training samples and test samples according to a preset number or a preset ratio.
[0078] S230. Construct a cross-entropy loss function according to the reward function of the fine-tuning model, and with the goal of minimizing the cross-entropy loss function, perform multiple iterative optimizations on only the model parameters of the fine-tuning model based on the training samples to obtain an alternative expert model.
[0079] In the embodiment of the present invention, the cross entropy loss function can be specifically understood as a function used to measure the difference between the model prediction probability distribution and the true label distribution.
[0080] As you can understand, since the reward function of the fine-tuned model outputs a reward score that measures the quality of the output, this reward score can be converted into a desired probability distribution. For example, by normalizing the reward score, it can be expressed as a probability distribution for outputs belonging to different categories or having different characteristics. The true target probability distribution is determined based on the training samples, and the cross-entropy loss function is constructed based on the reward function of the fine-tuned model. The cross-entropy loss function is used to calculate the difference between the two probability distributions as a metric to measure the gap between the output of the fine-tuned model on the current training sample and the expected output.
[0081] Specifically, a cross-entropy loss function is constructed according to the reward function of the fine-tuning model, and with the goal of minimizing the cross-entropy loss function, only the model parameters of the fine-tuning model are iteratively optimized multiple times based on the training samples. For example, the parameters of the fine-tuning model are adjusted through an optimization algorithm (such as gradient descent) so that the value of the cross-entropy loss function gradually decreases until the difference between the prediction results of the model on the training data and the true label meets the preset difference threshold, thereby obtaining an alternative expert model.
[0082] S240. When it is determined that the prediction accuracy of the candidate expert model for the test sample meets the prediction performance requirements, the candidate expert model is determined as the acquiring question-and-answer model.
[0083] Specifically, the performance of the alternative expert model is evaluated using test samples. By comparing the model's prediction results for the test samples with the actual labels, prediction accuracy evaluation indicators such as accuracy, precision, and recall are calculated. The calculated evaluation indicators are compared with the set performance threshold. If the model's evaluation indicators are less than or greater than the threshold, the model is considered to meet the prediction performance requirements, and the alternative expert model is determined as the final payment collection question and answer model.
[0084] If the evaluation indicators of the model do not reach the set threshold, further optimization can be performed, such as adjusting hyperparameters (modifying learning rate or regularization parameters, etc.), increasing the number of allowed training iterations, and adding training data, etc., to further train and fine-tune the model.
[0085] Furthermore, based on the above embodiments, after the candidate expert model is determined as the acquiring question-and-answer model, the following steps may be further included:
[0086] In response to a pre-trained model update instruction for the acquiring question-and-answer model, extracting a new pre-trained model contained in the pre-trained model update instruction;
[0087] After replacing the original pre-trained model in the acquiring question-answering model with the new pre-trained model, update the fine-tuning model in the acquiring question-answering model based on the new pre-trained model and the preset reward function;
[0088] Return to the process of constructing a cross-entropy loss function based on the reward function of the fine-tuned model until a Q&A model for acquiring customers based on the new pre-trained model is trained that meets the prediction performance requirements.
[0089] Specifically, after receiving a pre-trained model update instruction for the acquiring question and answer model, the system extracts the new pre-trained model contained in the update instruction and replaces the original pre-trained model in the acquiring question and answer model. After replacing the pre-trained model, the fine-tuning model in the acquiring question and answer model is updated based on the new pre-trained model and a pre-set reward function. The system then returns to execute the operation of constructing a cross-entropy loss function based on the reward function of the fine-tuning model. The fine-tuning model is retrained based on the new pre-trained model and a dedicated training sample set for the acquiring scenario until the model meets the predetermined performance standard, thereby obtaining the acquiring question and answer model based on the new pre-trained model.
[0090] By responding to the pre-trained model update instruction of the acquiring question and answer model, the acquiring question and answer model based on the new pre-trained model is trained to ensure that the acquiring question and answer model can utilize the latest pre-trained model with better performance, thereby improving the answer generation efficiency and accuracy of the acquiring question and answer model.
[0091] The technical solution of the embodiment of the present invention is to select and set a pre-trained model, and after constructing a fine-tuning model based on the pre-trained model and a preset reward function, connect the pre-trained model and the fine-tuning model end to end to obtain an expert model to be trained; generate training samples and test samples based on historical structured payment information and corresponding historical answers; construct a cross-entropy loss function based on the reward function of the fine-tuning model, and with the goal of minimizing the cross-entropy loss function, perform multiple iterative optimizations on the model parameters of the fine-tuning model based on the training samples to obtain an alternative expert model; when it is determined that the prediction accuracy of the alternative expert model for the test sample meets the prediction performance requirements, the alternative expert model is determined as the payment acquisition question and answer model. By only retraining the fine-tuning model part after replacing the pre-trained model to obtain a payment acquisition question and answer model that meets the prediction performance requirements, the adaptability of the payment acquisition question and answer model to the payment acquisition question and answer scenario is improved on the basis of the versatility of the pre-trained model, the process of payment acquisition question and answer model training is simplified, the cost and time of model training are significantly reduced, and the efficiency of model training and the flexibility of model maintenance are improved.
[0092] Example 3
[0093] Figure 3This is a flowchart of another method for fine-tuning a payment acquisition question-and-answer model provided in Example 3 of the present invention. This embodiment is a refinement of the "constructing a cross-entropy loss function based on the reward function of the fine-tuning model" in the above embodiment, and can specifically include: calculating the value function of the training sample based on the output value and reward function value of the expert model to be trained for the training sample; converting the value function into a predicted probability distribution through a normalized exponential function; and calculating the cross-entropy loss function based on the value function and the predicted probability distribution.
[0094] Correspondingly, such as Figure 3 As shown, the method includes:
[0095] S310: Select and set a pre-trained model, and after building a fine-tuning model based on the pre-trained model and a preset reward function, connect the pre-trained model and the fine-tuning model end to end to obtain an expert model to be trained.
[0096] S320: Generate training samples and test samples based on the historical structured acquiring information and the corresponding historical answers.
[0097] S330 : Calculate the value function of the training sample based on the output value of the expert model to be trained and the reward function value for the training sample.
[0098] Specifically, based on the output value of the expert model to be trained and the reward function value for the training sample, the value function v of the training sample is calculated. θ =ln softmax(f(x i ,θ pt ))+r θ (x i, f(x i ,θ pt )), where v θ is the value function, f(x i ,θ pt ) is the pre-trained model for the training sample x i The output value, θ pt are the parameters of the pre-trained model, r θ (x i ,f(x i ,θ pt )) is for the training sample x i , the output value f(x i ,θ pt ), and softmax(·) is the normalized exponential function.
[0099] S340. Convert the value function into a predicted probability distribution through a normalized exponential function.
[0100] Specifically, the value function is converted into a predicted probability distribution p through the softmax functionθ , which represents the model's predicted probability for each possible output, that is, p θ =softmax(v θ ).
[0101] S350, calculating the cross entropy loss function based on the value function and the predicted probability distribution, and taking minimizing the cross entropy loss function as the goal, performing multiple iterative optimizations on the model parameters of the fine-tuning model based on the training samples to obtain an alternative expert model.
[0102] Specifically, based on the value function and the predicted probability distribution, the cross entropy loss function T is calculated θ , which is used to measure the difference between the predicted probability distribution and the true label distribution, that is, T θ =CE(p θ ,v θ ), where CE(·) represents the cross entropy loss function.
[0103] With the goal of minimizing the cross entropy loss function, the model parameters of the fine-tuning model are iteratively optimized multiple times based on the training samples to obtain the alternative expert model.
[0104] Optionally, based on the above embodiments, with the goal of minimizing the cross entropy loss function, only the model parameters of the fine-tuning model are iteratively optimized based on the training samples to obtain an alternative expert model, which may include:
[0105] The objective function is constructed based on the cross entropy loss function, and the objective function is transformed based on the normalization factor and the value function to obtain the objective function with the goal of minimizing the cross entropy loss function;
[0106] With the goal of minimizing the objective function, the model parameters of the fine-tuning model are optimized iteratively multiple times based on the training samples to obtain the alternative expert model.
[0107] Specifically, the objective function is constructed based on the cross entropy loss function, and the objective function can be specifically: Among them, S is the training sample set, |S| is the size of the training sample set, that is, the training objective of the fine-tuning model is converted from minimizing the cross entropy loss function to minimizing the sum of the difference between the reward function and the value function, where y* is the correct output value under the input x.
[0108] The value function is defined as the logarithm of the normalization factor multiplied by the regularization parameter, that is, V θ (x)=λlnZ θ (x). Since the normalization factor is the sum of the products of the output probabilities of the pre-trained model and the reward exponent term, it can be expressed in the form of reward expectation. Therefore, the normalized exponential function and the value function are substituted into the objective function, and the objective function is converted to in, Represents the pre-trained model π pt (y|x) is the expected reward exponential term output when the input is x, thus obtaining the objective function with the goal of minimizing the cross-entropy loss function. With the goal of minimizing the objective function, the model parameters of the fine-tuning model are optimized multiple times based on the training samples to obtain the candidate expert model.
[0109] By introducing the value function, the model training process pays more attention to high-value samples, which improves the model's learning effect on key information. The transformation of the objective function based on the normalization factor and the value function makes the model output more in line with the probability distribution requirements, enhances the adaptability and robustness of the model, improves the convergence speed of model parameters, and improves training efficiency.
[0110] Furthermore, based on the above embodiments, after obtaining the objective function with the goal of minimizing the cross entropy loss function, the following steps may be further included:
[0111] The objective function is simplified based on the logarithmic identity and Jensen's inequality, and the simplified function is used as the objective function.
[0112] Specifically, the objective function is simplified based on the logarithmic identity and Jensen's inequality, that is, the expected logarithmic value is greater than or equal to the expectation of the logarithmic value. This simplifies the objective function to
[0113] By simplifying the objective function expression through the logarithmic identity, the computational complexity is reduced and the training efficiency is improved. By determining a reasonable range for the training process through the Jensen inequality, the objective function expression is further simplified, the resource consumption and time cost of model training are reduced, and the model convergence speed is accelerated. At the same time, the numerical stability and interpretability are improved, which facilitates debugging and improves the maintainability of the training model.
[0114] S360. When it is determined that the prediction accuracy of the candidate expert model for the test sample meets the prediction performance requirements, the candidate expert model is determined as the acquiring question-and-answer model.
[0115] The technical solution of the embodiment of the present invention selects and sets a pre-trained model, and after constructing a fine-tuning model based on the pre-trained model and a preset reward function, the pre-trained model and the fine-tuning model are connected end to end to obtain an expert model to be trained; training samples and test samples are generated based on historical structured acquiring information and corresponding historical answers; the value function of the training sample is calculated based on the output value and reward function value of the expert model to be trained for the training sample, and the value function reflects the expected reward under the corresponding input; the value function is converted into a predicted probability distribution using a normalized exponential function to obtain the model's predicted probability for each possible output; the cross-entropy loss function is calculated based on the value function and the predicted probability distribution to measure the difference between the model's predicted probability distribution and the true label distribution; with the goal of minimizing the cross-entropy loss function, the model parameters of the fine-tuning model are iteratively optimized multiple times based on the training samples to obtain an alternative expert model; when it is determined that the prediction accuracy of the alternative expert model for the test sample meets the prediction performance requirements, the alternative expert model is determined as the acquiring question-and-answer model. By minimizing the cross-entropy loss function, the parameters of the fine-tuning model can be optimized so that the model's prediction results on the training sample are close to the true label, thereby improving the model's prediction accuracy. Since this process only optimizes the parameters of the fine-tuning model, it can improve the adaptability of the acquiring question and answer model to acquiring question and answer scenarios while maintaining the universality of the pre-trained model, significantly reduce the computational overhead and time cost during training, and improve the efficiency of model training and the flexibility of model maintenance.
[0116] Example 4
[0117] Figure 4 This is a structural diagram of a Q&A device for acquiring a payment based on model fine-tuning provided by the fourth embodiment of the present invention. Figure 4 As shown, the device includes: a structuring module 410, an input module 420, a primary processing module 430 and a secondary processing module 440, wherein:
[0118] The structuring module 410 is used to obtain the acquiring information to be queried and perform structured pre-processing on the acquiring information to obtain structured acquiring information;
[0119] Input module 420, for inputting structured acquiring information into the acquiring question-answering model;
[0120] The acquiring question-and-answer model includes a pre-trained model and a fine-tuned model, which are connected end to end. The pre-trained model is trained using a general training sample set, and the fine-tuned model is trained based on the pre-trained model and a special training sample set for the acquiring scenario.
[0121] The primary processing module 430 is used to process the structured acquiring information using the pre-trained model, obtain a general processing result and provide it to the fine-tuning model;
[0122] The secondary processing module 440 is used to perform secondary processing on the received general data results by fine-tuning the model to obtain a target answer corresponding to the structured payment information.
[0123] The technical solution of the embodiments of the present invention obtains the acquiring information to be queried and pre-processes it to obtain structured acquiring information; then inputs the structured acquiring information into an acquiring question-and-answer model; processes the structured acquiring information once using a pre-trained model to obtain a general processing result, which is then provided to a fine-tuned model; and then performs secondary processing on the received general data result using the fine-tuned model to obtain a target answer corresponding to the structured acquiring information. By training the pre-trained model based on a general sample set and the fine-tuned model based on a dedicated acquiring sample set, the adaptability of the acquiring question-and-answer model to acquiring question-and-answer scenarios is improved, significantly reducing the cost and time of training the acquiring question-and-answer model after replacing the pre-trained model, and improving the efficiency of model training and the flexibility of model maintenance. The acquiring information to be queried and pre-processed to obtain structured acquiring information is then input into an acquiring question-and-answer model consisting of a pre-trained model and a fine-tuned model connected end-to-end. After initial processing by the pre-trained model, a general result is obtained, which is then passed to the fine-tuned model for secondary processing, ultimately outputting the target answer, thereby improving the accuracy of acquiring question-and-answer.
[0124] The acquiring question-and-answer device based on model fine-tuning provided in an embodiment of the present invention can execute the acquiring question-and-answer method based on model fine-tuning provided in any embodiment of the present invention, and has the corresponding functional modules and beneficial effects of the execution method.
[0125] Example 5
[0126] Figure 5 This is a structural diagram of a fine-tuning device for a question-and-answer model for acquiring provided in the fifth embodiment of the present invention. Figure 5 As shown, the device includes:
[0127] The end-to-end connection module 510 is used to select and set a pre-trained model, and after constructing a fine-tuning model based on the pre-trained model and a preset reward function, connect the pre-trained model and the fine-tuning model end-to-end to obtain an expert model to be trained;
[0128] The sample generation module 520 is used to generate training samples and test samples based on the historical structured acquiring information and the corresponding historical answers;
[0129] An alternative module 530 is configured to construct a cross-entropy loss function according to the reward function of the fine-tuning model, and perform multiple iterative optimizations on the model parameters of the fine-tuning model based on the training samples with the goal of minimizing the cross-entropy loss function to obtain an alternative expert model;
[0130] The determination module 540 is configured to determine the candidate expert model as the acquiring question-and-answer model when it is determined that the prediction accuracy of the candidate expert model for the test sample meets the prediction performance requirement.
[0131] The technical solution of the embodiment of the present invention is to select and set a pre-trained model, and after constructing a fine-tuning model based on the pre-trained model and a preset reward function, connect the pre-trained model and the fine-tuning model end to end to obtain an expert model to be trained; generate training samples and test samples based on historical structured payment information and corresponding historical answers; construct a cross-entropy loss function based on the reward function of the fine-tuning model, and with the goal of minimizing the cross-entropy loss function, perform multiple iterative optimizations on the model parameters of the fine-tuning model based on the training samples to obtain an alternative expert model; when it is determined that the prediction accuracy of the alternative expert model for the test sample meets the prediction performance requirements, the alternative expert model is determined as the payment acquisition question and answer model. By only needing to retrain the fine-tuning model part after replacing the pre-trained model to obtain a payment acquisition question and answer model that meets the prediction performance requirements, the payment acquisition question and answer model training process is simplified, the cost and time of model training are significantly reduced, and the efficiency of model training and the flexibility of model maintenance are improved.
[0132] Based on the above embodiments, the end-to-end connection module 510 is specifically configured to:
[0133] Obtaining the model output probability of the pre-trained model, calculating the numerical relationship between the input and output of the fine-tuning model based on the model output eigenvalue, the exponential term of the reward function, and the normalization factor, and constructing the fine-tuning model based on the calculated numerical relationship;
[0134] The reward exponent term is a power operation with a natural constant as the base and the reward function divided by the regularization parameter as the exponent; the normalization factor is the sum of the products of each output probability of the pre-trained model multiplied by the reward exponent term.
[0135] Based on the above embodiments, the optional module 530 is specifically configured to:
[0136] Calculating the value function of the training sample based on the output value of the to-be-trained expert model and the reward function value for the training sample;
[0137] Convert the value function into a predicted probability distribution through a normalized exponential function;
[0138] Calculate the cross entropy loss function based on the value function and the predicted probability distribution.
[0139] Based on the above embodiments, the optional module 530 is further configured to:
[0140] The objective function is constructed based on the cross entropy loss function, and the objective function is transformed based on the normalization factor and the value function to obtain the objective function with the goal of minimizing the cross entropy loss function;
[0141] With the goal of minimizing the objective function, the model parameters of the fine-tuning model are optimized iteratively multiple times based on the training samples to obtain the alternative expert model.
[0142] Optionally, based on the above embodiments, the alternative module 530 may include a simplification unit, wherein:
[0143] The simplification unit is used to simplify the objective function based on the logarithmic identity and Jensen's inequality after obtaining the objective function with the goal of minimizing the cross entropy loss function, and use the simplified function as the objective function.
[0144] Furthermore, based on the above embodiments, the fine-tuning device for the acquiring question-and-answer model may further include: an extraction module, an update module, and a training module, wherein:
[0145] an extraction module for extracting, after determining the candidate expert model as the acquiring question-and-answer model, a new pre-trained model contained in the pre-trained model update instruction in response to the pre-trained model update instruction for the acquiring question-and-answer model;
[0146] An update module, which is used to replace the original pre-trained model in the acquiring question-answering model with a new pre-trained model and then update the fine-tuning model in the acquiring question-answering model based on the new pre-trained model and a preset reward function;
[0147] The training module is used to return to the operation of constructing the cross-entropy loss function based on the reward function of the fine-tuned model until a Q&A model for acquiring based on the new pre-trained model is trained to meet the prediction performance requirements.
[0148] The fine-tuning device for the acquiring question-and-answer model provided in an embodiment of the present invention can execute the fine-tuning method for the acquiring question-and-answer model provided in any embodiment of the present invention, and has the corresponding functional modules and beneficial effects of the execution method.
[0149] In the technical solutions disclosed herein, the collection, storage, use, processing, transmission, provision and disclosure of user personal information involved comply with the provisions of relevant laws and regulations and do not violate public order and good morals.
[0150] In the technical solution disclosed herein, the information collected is information and data authorized by the user or fully authorized by all parties, and the collection, storage, use, processing, transmission, provision, disclosure and application of the relevant data comply with the relevant laws, regulations and standards of the relevant countries and regions, take necessary confidentiality measures, do not violate public order and good morals, and provide corresponding operation entrances for users to choose to authorize or refuse.
[0151] In the technical solution disclosed herein, if automated decision-making is involved, a corresponding operation entry will be provided for the user to choose to agree or reject the automated decision-making result; if the user chooses to reject, the expert decision-making process will be entered.
[0152] Example 6
[0153] Figure 6 A schematic diagram of the structure of an electronic device 10 that can be used to implement an embodiment of the present invention is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processing, cellular phones, smart phones, wearable devices (such as helmets, glasses, watches, etc.) and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present invention described and / or claimed herein.
[0154] like Figure 6 As shown, the electronic device 10 includes at least one processor 11 and a memory, such as a read-only memory (ROM) 12, a random access memory (RAM) 13, etc., which is communicatively connected to the at least one processor 11. The memory stores a computer program that can be executed by the at least one processor. The processor 11 can perform various appropriate actions and processes according to the computer program stored in the read-only memory (ROM) 12 or the computer program loaded from the storage unit 18 into the random access memory (RAM) 13. Various programs and data required for the operation of the electronic device 10 can also be stored in the RAM 13. The processor 11, ROM 12, and RAM 13 are connected to each other via a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.
[0155] Multiple components in the electronic device 10 are connected to the I / O interface 15, including an input unit 16, such as a keyboard, a mouse, etc.; an output unit 17, such as various types of displays, speakers, etc.; a storage unit 18, such as a magnetic disk, an optical disk, etc.; and a communication unit 19, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 19 allows the electronic device 10 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0156] The processor 11 can be any general-purpose and / or specialized processing component with processing and computing capabilities. Some examples of the processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various specialized artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The processor 11 executes the various methods and processes described above, such as the acquiring question-and-answer method based on model fine-tuning, namely:
[0157] Obtain the acquiring information to be queried and perform structured preprocessing on the acquiring information to obtain structured acquiring information;
[0158] Input structured acquiring information into the acquiring question-answering model;
[0159] The acquiring question-and-answer model includes a pre-trained model and a fine-tuned model, which are connected end to end. The pre-trained model is trained using a general training sample set, and the fine-tuned model is trained based on the pre-trained model and a special training sample set for the acquiring scenario.
[0160] The structured acquiring information is processed once through the pre-trained model, and the general processing results are provided to the fine-tuning model;
[0161] By fine-tuning the model, the received general data results are processed again to obtain the target answer corresponding to the structured payment information.
[0162] Another example is the fine-tuning method of the acquiring question-answering model, which is:
[0163] Select and set a pre-trained model, and build a fine-tuning model based on the pre-trained model and the preset reward function. Then, connect the pre-trained model and the fine-tuning model end to end to obtain the expert model to be trained;
[0164] Generate training samples and test samples based on historical structured acquiring information and corresponding historical answers;
[0165] A cross-entropy loss function is constructed based on the reward function of the fine-tuning model. With the goal of minimizing the cross-entropy loss function, the model parameters of the fine-tuning model are iteratively optimized multiple times based on the training samples to obtain an alternative expert model.
[0166] When it is determined that the prediction accuracy of the alternative expert model for the test sample meets the prediction performance requirements, the alternative expert model is determined as the acquisition question and answer model.
[0167] In some embodiments, the acquirer question-and-answer method based on model fine-tuning or the method for fine-tuning the acquirer question-and-answer model may be implemented as a computer program, which is tangibly contained in a computer-readable storage medium, such as the storage unit 18. In some embodiments, part or all of the computer program may be loaded and / or installed on the electronic device 10 via the ROM 12 and / or the communication unit 19. When the computer program is loaded into the RAM 13 and executed by the processor 11, one or more steps of the acquirer question-and-answer method based on model fine-tuning or the method for fine-tuning the acquirer question-and-answer model described above may be performed. Alternatively, in other embodiments, the processor 11 may be configured to execute the acquirer question-and-answer method based on model fine-tuning or the method for fine-tuning the acquirer question-and-answer model by any other appropriate means (e.g., by means of firmware).
[0168] Various embodiments of the systems and techniques described herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), system-on-chip systems (SOCs), programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include being implemented in one or more computer programs that are executable and / or interpreted on a programmable system comprising at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.
[0169] Computer programs for implementing the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when the computer program is executed by the processor, the functions / operations specified in the flowcharts and / or block diagrams are implemented. The computer program may be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.
[0170] In the context of the present invention, computer-readable storage media can be tangible media that can contain or store a computer program for use with an instruction execution system, device or equipment or used in combination with an instruction execution system, device or equipment. Computer-readable storage media can include but are not limited to electronic, magnetic, optical, electromagnetic, infrared or semiconductor systems, devices or equipment, or any suitable combination of the foregoing. Alternatively, computer-readable storage media can be machine-readable signal media. More specific examples of machine-readable storage media can include electrical connections based on one or more lines, portable computer disks, hard disks, random access memories (RAM), read-only memories (ROM), erasable programmable read-only memories (EPROM or flash memory), optical fibers, portable compact disk read-only memories (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0171] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the electronic device. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).
[0172] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), a blockchain network, and the Internet.
[0173] A computing system may include clients and servers. The clients and servers are typically remote from each other and typically interact via a communication network. This client-server relationship arises through computer programs running on the respective computers, creating a client-server relationship. The server may be a cloud server, also known as a cloud computing server or cloud host. This server is a hosting product within the cloud computing service ecosystem that addresses the management difficulties and limited scalability of traditional physical hosting and VPS services.
[0174] It should be understood that the various forms of the processes shown above can be used to reorder, add, or delete steps. For example, the steps described in the present invention can be performed in parallel, sequentially, or in a different order, as long as the desired results of the technical solution of the present invention can be achieved. This is not limited herein.
[0175] The above specific embodiments do not limit the scope of protection of the present invention. Those skilled in the art will appreciate that various modifications, combinations, sub-combinations, and substitutions may be made based on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention are intended to be included within the scope of protection of the present invention.
Claims
1. A Q&A method for acquiring customers based on model fine-tuning, characterized in that: include: Obtain the acquiring information to be queried and perform structured preprocessing on the acquiring information to obtain structured acquiring information; Input structured acquiring information into the acquiring question-answering model; The acquiring question-and-answer model includes a pre-trained model and a fine-tuned model, which are connected end to end. The pre-trained model is trained using a general training sample set, and the fine-tuned model is trained based on the pre-trained model and a special training sample set for the acquiring scenario. The structured acquiring information is processed once through the pre-trained model, and the general processing results are provided to the fine-tuning model; By fine-tuning the model, the received general data results are processed again to obtain the target answer corresponding to the structured payment information.
2. A fine-tuning method for a payment acquisition question-answering model, characterized in that: include: Select and set a pre-trained model, and build a fine-tuning model based on the pre-trained model and the preset reward function. Then, connect the pre-trained model and the fine-tuning model end to end to obtain the expert model to be trained; Generate training samples and test samples based on historical structured acquiring information and corresponding historical answers; A cross-entropy loss function is constructed based on the reward function of the fine-tuning model. With the goal of minimizing the cross-entropy loss function, the model parameters of the fine-tuning model are iteratively optimized multiple times based on the training samples to obtain an alternative expert model. When it is determined that the prediction accuracy of the alternative expert model for the test sample meets the prediction performance requirements, the alternative expert model is determined as the acquisition question and answer model.
3. The method according to claim 2, characterized in that Build a fine-tuned model based on the pre-trained model and the preset reward function, including: Obtaining the model output probability of the pre-trained model, calculating the numerical relationship between the input and output of the fine-tuning model based on the model output eigenvalue, the exponential term of the reward function, and the normalization factor, and constructing the fine-tuning model based on the calculated numerical relationship; The reward exponent term is a power operation with a natural constant as the base and the reward function divided by the regularization parameter as the exponent; the normalization factor is the sum of the products of each output probability of the pre-trained model and the reward exponent term.
4. The method according to claim 2, characterized in that Construct a cross entropy loss function based on the reward function of the fine-tuned model, including: Calculating the value function of the training sample based on the output value of the to-be-trained expert model and the reward function value for the training sample; Convert the value function into a predicted probability distribution through a normalized exponential function; Calculate the cross entropy loss function based on the value function and the predicted probability distribution.
5. The method according to claim 4, characterized in that With the goal of minimizing the cross entropy loss function, the model parameters of the fine-tuning model are optimized iteratively multiple times based on the training samples to obtain the candidate expert models, including: The objective function is constructed based on the cross entropy loss function, and the objective function is transformed based on the normalization factor and the value function to obtain the objective function with the goal of minimizing the cross entropy loss function; With the goal of minimizing the objective function, the model parameters of the fine-tuning model are optimized iteratively multiple times based on the training samples to obtain the alternative expert model.
6. The method according to claim 5, characterized in that After obtaining the objective function with the goal of minimizing the cross entropy loss function, it also includes: The objective function is simplified based on the logarithmic identity and Jensen's inequality, and the simplified function is used as the objective function.
7. The method according to any one of claims 2 to 6, characterized in that: After the alternative expert model is determined to be the acquisition question-answering model, it also includes: In response to a pre-trained model update instruction for the acquiring question-and-answer model, extracting a new pre-trained model contained in the pre-trained model update instruction; After replacing the original pre-trained model in the acquiring question-answering model with the new pre-trained model, update the fine-tuning model in the acquiring question-answering model based on the new pre-trained model and the preset reward function; Return to the process of constructing a cross-entropy loss function based on the reward function of the fine-tuned model until a Q&A model for acquiring customers based on the new pre-trained model is trained that meets the prediction performance requirements.
8. An electronic device, characterized in that: The electronic device comprises: at least one processor; and a memory communicatively connected to the at least one processor; wherein, The memory stores a computer program that can be executed by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute the acquisition question and answer method based on model fine-tuning according to claim 1 or the acquisition question and answer model fine-tuning method according to any one of claims 2-7.
9. A computer-readable storage medium, characterized in that The computer-readable storage medium stores computer instructions, which are used to enable the processor to implement the model fine-tuning based acquisition question and answer method of claim 1 or the fine-tuning method of the acquisition question and answer model according to any one of claims 2-7 when executed.
10. A computer program product, characterized in that The computer program product includes a computer program, which, when executed by a processor, implements the acquiring question-and-answer method based on model fine-tuning according to claim 1 or the method for fine-tuning the acquiring question-and-answer model according to any one of claims 2-7.