Financial large model training method and business problem processing method
By obtaining financial question-and-answer pairs from different sources and evaluating and screening them based on multiple evaluation dimensions, we solve the data quality and anthropomorphism problems of large financial models in complex scenarios in the financial industry, improve the adaptability and responsiveness of the models, provide personalized professional services, and improve user experience.
Patent Information
- Application Number
- CN202510609891.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-13
- Publication Date
- 2025-09-19
AI Technical Summary
When facing complex scenarios in the financial industry, large financial models face problems such as uneven data quality, uneven distribution of data samples, lack of professional quality control, and insufficient anthropomorphic features. As a result, the model performs poorly when processing low-frequency or edge scenarios, and is unable to effectively understand user needs, affecting the user experience.
By obtaining first and second business question-answer pairs from different sources and evaluating them on multiple evaluation dimensions, high-quality sample business question-answer pairs are screened out for training to ensure the diversity and comprehensiveness of the dataset. Combined with semantic similarity calculation, redundancy and repetition are reduced, thereby improving the generalization ability and adaptability of the model.
It enhances the generalization and adaptability of large financial models, enables more accurate responses to user queries, provides personalized and professional advice and services, improves user experience, avoids the subjectivity and inconsistency of human evaluation, and improves the quality of training data sets and the responsiveness of models.
Smart Images

Figure CN120670545A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of data processing technology, and in particular to a method for training a large financial model and a method for processing business problems. Background Art
[0002] With the rapid development of artificial intelligence technology, financial dialogue systems based on financial big models can significantly improve customer service efficiency, optimize user experience, and support intelligent decision-making. Therefore, they have been widely used in banking, securities, insurance and other fields.
[0003] In financial services, customers expect timely, accurate, and reliable responses. Financial conversational systems must not only efficiently understand user questions but also provide appropriate solutions in highly complex and dynamic scenarios. When interacting with conversational systems, users expect them to respond quickly and maintain a professional service attitude. Therefore, timely and accurate interaction with users is crucial. Summary of the Invention
[0004] This disclosure provides a method for training a large financial model and a method for handling business problems, which at least to some extent solve one of the technical problems in the related art. The technical solution of this disclosure is as follows:
[0005] According to a first aspect of an embodiment of the present disclosure, a method for training a financial big model is provided, including: obtaining a plurality of first business question and answer pairs and a plurality of second business question and answer pairs associated with each financial business; wherein the first business question and answer pairs and the second business question and answer pairs have different sources; obtaining a first evaluation grade obtained by evaluating each first business question and answer pair in a plurality of evaluation dimensions and a second evaluation grade obtained by evaluating each second business question and answer pair in the plurality of evaluation dimensions; screening the plurality of first business question and answer pairs and the plurality of second business question and answer pairs according to the first evaluation grade of each first business question and answer pair in the plurality of evaluation dimensions and the second evaluation grade of each second business question and answer in the plurality of evaluation dimensions to obtain sample business question and answer pairs; and training the financial big model according to the sample business question and answer pairs.
[0006] According to a second aspect of an embodiment of the present disclosure, a method for processing business problems is provided, comprising: obtaining a target business problem associated with a target financial business; inputting the target business problem into a trained financial big model to obtain a target business answer to the target business problem output by the financial big model; wherein the financial big model is trained using the training method for the financial big model described in the embodiment of the first aspect.
[0007] According to a third aspect of an embodiment of the present disclosure, a training device for a financial big model is provided, comprising: a first acquisition module for acquiring a plurality of first business question-and-answer pairs and a plurality of second business question-and-answer pairs associated with each financial business; wherein the first business question-and-answer pairs and the second business question-and-answer pairs have different sources; a second acquisition module for acquiring a first evaluation grade obtained by evaluating each first business question-and-answer pair in a plurality of evaluation dimensions and a second evaluation grade obtained by evaluating each second business question-and-answer pair in the plurality of evaluation dimensions; a screening module for screening a plurality of first business question-and-answer pairs and a plurality of second business question-and-answer pairs according to the first evaluation grade of each first business question-and-answer pair in the plurality of evaluation dimensions and the second evaluation grade of each second business question-and-answer in the plurality of evaluation dimensions, so as to obtain sample business question-and-answer pairs; and a training module for training the financial big model based on the sample business question-and-answer pairs.
[0008] According to a fourth aspect of an embodiment of the present disclosure, a device for processing business problems is provided, comprising: an acquisition module for acquiring a target business problem associated with a target financial business; an output module for inputting the target business problem into a trained financial big model to obtain a target business answer to the target business problem output by the financial big model; wherein the financial big model is trained using the financial big model training device described in the embodiment of the third aspect.
[0009] According to the fifth aspect of the embodiments of the present disclosure, an electronic device is provided, comprising: a processor; a memory for storing instructions executable by the processor; wherein the processor is configured to execute the instructions to implement the training method of the financial big model as described in the embodiment of the first aspect of the present disclosure, or to implement the method for processing business problems as described in the embodiment of the second aspect of the present disclosure.
[0010] According to the sixth aspect of the embodiments of the present disclosure, a computer-readable storage medium is provided. When the instructions in the computer-readable storage medium are executed by the processor of an electronic device, the electronic device is enabled to execute the training method of the financial big model as described in the embodiment of the first aspect of the present disclosure, or to execute the method for processing business problems as described in the embodiment of the second aspect of the present disclosure.
[0011] According to the seventh aspect of the embodiments of the present disclosure, a computer program product is provided, comprising: a computer program, which, when executed by a processor, implements the training method of the financial big model as described in the embodiment of the first aspect of the present disclosure, or implements the method for processing business problems as described in the embodiment of the second aspect of the present disclosure.
[0012] The technical solutions provided by the embodiments of the present disclosure bring at least the following beneficial effects:
[0013] In this technical solution, firstly, by obtaining the first business question and answer pair and the second business question and answer pair from different sources and evaluating them separately in multiple evaluation dimensions, the diversity and comprehensiveness of the data set are ensured, thereby enhancing the generalization ability and adaptability of the model; secondly, by evaluating each question and answer pair in multiple evaluation dimensions and assigning corresponding evaluation grades, it is possible to accurately identify high-quality sample business question and answer pairs, avoid low-quality or misleading information from interfering with the learning process of the model, and ensure the high standard of the training data set; finally, the financial big model is trained based on the sample business question and answer pairs, so that the model can not only understand basic financial knowledge, but also master complex business logic and the latest market dynamics, so that it can respond to user queries more accurately in actual applications and provide personalized and professional advice and services; wherein, when determining the first evaluation grade obtained by evaluating the first business question and answer pair in multiple evaluation dimensions, based on each first business question and answer The first evaluation level of each first business question and answer pair in any evaluation dimension is determined by the semantic similarity between the evaluation standard descriptions of multiple candidate evaluation levels under any evaluation dimension, which realizes the fine quantification of the performance of the question and answer pairs in different dimensions, realizes the automation and standardization of the evaluation, and avoids the subjectivity and inconsistency of human evaluation, thereby improving the reliability of the evaluation results; in addition, when determining the sample business question and answer pairs, the levels of each business question and answer pair in multiple evaluation dimensions are combined for screening, and the semantic similarity between the retained question and answer pairs is calculated to determine the sample business question and answer pairs, ensuring that the quality of the screened question and answer pairs is guaranteed. By calculating the semantic similarity, it is achieved that redundant and repeated question and answer pairs are effectively reduced, and the pertinence and relevance of the sample business question and answer pairs are improved, so that the model training is based on the sample business question and answer pairs, which improves the efficiency of the training model, and at the same time enhances the model's understanding and response capabilities for financial business, and improves the user experience.
[0014] It is to be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] The accompanying drawings herein are incorporated into and constitute a part of the specification, illustrate embodiments consistent with the present disclosure, and together with the description are used to explain the principles of the present disclosure, and do not constitute an improper limitation of the present disclosure.
[0016] Figure 1 This is a flowchart of the training method for a large financial model shown in the first embodiment of the present disclosure;
[0017] Figure 2 1 is a flow chart of a training method for a large financial model according to the second embodiment of the present disclosure;
[0018] Figure 31 is a flowchart of a training method for a large financial model according to the third embodiment of the present disclosure;
[0019] Figure 4 4 is a flowchart of a method for training a large financial model according to the fourth embodiment of the present disclosure;
[0020] Figure 5 is a flowchart of a method for processing a business problem according to the fifth embodiment of the present disclosure;
[0021] Figure 6 1 is a schematic diagram of the structure of a training device for a large financial model according to a sixth embodiment of the present disclosure;
[0022] Figure 7 1 is a schematic diagram of the structure of a device for processing business problems according to the seventh embodiment of the present disclosure;
[0023] Figure 8 It is a schematic structural diagram of an electronic device shown in an exemplary embodiment of the present disclosure. DETAILED DESCRIPTION
[0024] In order to enable ordinary persons in the art to better understand the technical solutions of the present disclosure, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below with reference to the accompanying drawings.
[0025] It should be noted that the terms "first," "second," and the like in the specification and claims of the present disclosure and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or precedence. It should be understood that the numbers used in this manner are interchangeable where appropriate so that the embodiments of the present disclosure described herein can be implemented in an order other than those illustrated or described herein. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present disclosure. Instead, they are merely examples of apparatus and methods consistent with certain aspects of the present disclosure as detailed in the appended claims.
[0026] It should be noted that in the technical solutions disclosed herein, the collection, storage, use, processing, transmission, provision and disclosure of user personal information are all carried out with the user's consent, and are in compliance with relevant laws and regulations and do not violate public order and good morals.
[0027] Currently, most large financial models still rely on open-source models from general domains for fine-tuning. However, these models often face challenges with inconsistent data quality, uneven data sample distribution, and a lack of professional quality control when applied to the complex scenarios of the financial industry. To improve the performance of fine-tuned models, fine-tuning financial data must be meticulously labeled and quality controlled. Fine-tuning can help models better understand financial terminology, user intent, and sentiment, while quality control ensures that the system fully considers user needs and industry standards when generating responses, avoiding incorrect decisions and inappropriate services. Therefore, meticulous labeling, quality control, and data screening of fine-tuning data for large financial models are key to improving their performance.
[0028] However, there are still several technical challenges in improving the quality of conversational data in the financial sector:
[0029] (1) Imbalanced scene data
[0030] Conversational data in financial scenarios varies significantly, especially in terms of diversity and representativeness. Some data may be overly concentrated in a certain type of scenario or conversation. This imbalance in scenario data can result in trained models being unable to process data from low-frequency or marginal scenarios. In practical applications, financial conversational systems may perform poorly in handling uncommon conversational scenarios, making it difficult to adapt to diverse user needs and complex business scenarios. Therefore, balancing the various scenarios in a dataset to ensure that data from all important scenarios is fully represented is a key issue in improving the quality of financial conversational systems.
[0031] (2) Lack of data
[0032] Data scarcity is a common problem in the development and fine-tuning of large financial models. Many financial scenarios, especially complex and low-frequency conversational scenarios, often lack sufficient labeled data. This results in a lack of accuracy and reliability in the systems handling these scenarios, and an inability to effectively understand user needs. Furthermore, the specialized terminology and industry knowledge in the financial field require highly accurate data annotation, but due to data scarcity, many systems struggle to effectively learn and adapt to this knowledge, impacting the overall performance and practical application of the fine-tuned models.
[0033] (3) Insufficient anthropomorphic dialogue
[0034] Many large-scale financial model fine-tuning efforts lack sufficient anthropomorphic features, resulting in mechanical, monotonous conversations that lack emotion and tone, impacting the user experience. This is especially true when dealing with sensitive issues, as users need to feel understood and cared for by the system. A lack of emotional nuance in responses can alienate users and reduce satisfaction. Therefore, financial conversational systems must strike a balance between professionalism and anthropomorphism, providing an interactive experience that is both professional and emotionally caring.
[0035] In response to the above problems, this paper proposes a training method for a large financial model and a method for handling business problems.
[0036] The following describes the training method of the financial big model and the method for handling business problems in the embodiments of the present disclosure with reference to the accompanying drawings.
[0037] Figure 1 This is a flow chart of the method for training a large financial model, as shown in the first embodiment of the present disclosure. It should be noted that the embodiment of the present disclosure can be executed by a device for training a large financial model. This device can be applied to any electronic device with computing capabilities, enabling the electronic device to perform the training function of the large financial model.
[0038] like Figure 1 As shown in Figure 1, the training method of the financial model includes the following steps:
[0039] Step 101: Acquire multiple first business question-answer pairs and multiple second business question-answer pairs associated with each financial business.
[0040] The first business question and answer pair and the second business question and answer pair have different sources.
[0041] In order to ensure the diversity and comprehensiveness of the data, as a possible implementation method, business question and answer pairs related to various financial businesses are obtained from different sources.
[0042] In an embodiment of the present disclosure, multiple first business question-and-answer pairs associated with various financial businesses are obtained from the customer service records of a financial institution, where the financial businesses include but are not limited to: card application, deposits, loans, corporate business, and financial management; the source of the second business question-and-answer pairs is different from that of the first business question-and-answer pairs, and the second business question-and-answer pairs can come from financial web pages, financial service websites, or be generated based on a large model.
[0043] Step 102: Obtain a first evaluation grade obtained by evaluating each first business question-answer pair in multiple evaluation dimensions and a second evaluation grade obtained by evaluating each second business question-answer pair in multiple evaluation dimensions.
[0044] To ensure the quality of the dataset used to train large financial models, in the disclosed embodiment, the obtained first business question-and-answer pairs and second business question-and-answer pairs are evaluated and labeled to obtain a first evaluation grade for each first business question-and-answer pair in multiple evaluation dimensions and a second evaluation grade for each second business question-and-answer pair in multiple evaluation dimensions. It should be noted that the first business question-and-answer pair corresponds to a first evaluation grade under each evaluation dimension, and the second business question-and-answer pair corresponds to a second evaluation grade under each evaluation dimension. The multiple evaluation dimensions include, but are not limited to, professionalism, service attitude, communication skills, and customer sentiment.
[0045] Step 103: Filter the multiple first business question and answer pairs and the multiple second business question and answer pairs according to the first evaluation level of each first business question and answer pair in multiple evaluation dimensions and the second evaluation level of each second business question and answer pair in multiple evaluation dimensions to obtain sample business question and answer pairs.
[0046] In order to effectively obtain high-quality sample business question-answer pairs and improve the quality of the financial big model, as a possible implementation method, sample business question-answer pairs are determined from multiple first business question-answer pairs and multiple second business question-answer pairs.
[0047] In an embodiment of the present disclosure, based on the first evaluation level of each first business question and answer pair in multiple evaluation dimensions and the second evaluation level of each second business question and answer in multiple evaluation dimensions, multiple first business question and answer pairs and multiple second business question and answer pairs are screened to obtain retained first business question and answer pairs and retained second business question and answer pairs, thereby determining sample business question and answer pairs based on the retained first business question and answer pairs and the retained second business question and answer pairs.
[0048] Step 104: Train the financial big model based on the sample business question and answer pairs.
[0049] In an embodiment of the present disclosure, a sample business question in a sample business question-answer pair is input into a financial big model to obtain a predicted business answer output by the financial big model, and a loss function value is generated based on the difference between the predicted business answer and the sample business answer. The financial big model is trained based on the loss function value.
[0050] In summary, by obtaining the first business question-and-answer pairs and the second business question-and-answer pairs from different sources and evaluating them separately on multiple evaluation dimensions, the diversity and comprehensiveness of the data set are ensured, thereby enhancing the generalization ability and adaptability of the model; secondly, by evaluating each question-and-answer pair on multiple evaluation dimensions and assigning corresponding evaluation grades, high-quality sample business question-and-answer pairs are accurately identified, preventing low-quality or misleading information from interfering with the model's learning process, and ensuring the high standards of the training data set; finally, the financial large model is trained based on the screened sample business question-and-answer pairs, so that the model can not only understand basic financial knowledge, but also grasp complex business logic and the latest market trends, so that it can respond to user queries more accurately in actual applications and provide personalized and professional advice and services.
[0051] In order to clearly illustrate how the first evaluation level of each first business question-answer pair evaluated in multiple evaluation dimensions is obtained in the above embodiment, the present disclosure proposes another training method for a large financial model.
[0052] Figure 2 It is a flowchart of the training method of the financial big model shown in the second embodiment of the present disclosure.
[0053] like Figure 2 As shown in Figure 1, the training method of the financial model includes the following steps:
[0054] Step 201: Acquire multiple first business question-answer pairs and multiple second business question-answer pairs associated with each financial business.
[0055] The first business question and answer pair and the second business question and answer pair have different sources.
[0056] Step 202: Obtain descriptions of evaluation criteria for multiple candidate evaluation levels under each evaluation dimension.
[0057] In the disclosed embodiments, each evaluation dimension has multiple candidate rating levels. Each candidate rating level corresponds to an evaluation standard that specifies the specific conditions or characteristics required to achieve the corresponding candidate rating level. It should be noted that the evaluation dimensions include, but are not limited to, professionalism, service attitude, communication skills, and customer sentiment.
[0058] For example, taking the evaluation dimension of "professionalism" as an example, the multiple candidate evaluation dimensions under "professionalism" include: "excellent", "general" and "pass", etc. The evaluation criteria corresponding to "excellent" are described as "accurate use of financial terms and covering necessary professional fields", the evaluation criteria corresponding to "general" are described as "some terms are used inaccurately or some important professional information is omitted", and the evaluation criteria corresponding to "pass" are described as "there are many non-professional terms or inappropriate use of terms in the conversation".
[0059] Step 203 : For any evaluation dimension, determine the semantic matching degree between any first business question-answer pair and the evaluation standard descriptions of multiple candidate evaluation levels under any evaluation dimension.
[0060] In order to accurately determine the evaluation level of each business question and answer pair, in the embodiment of the present disclosure, each business question and answer pair is semantically matched with the descriptions of the evaluation criteria of multiple candidate evaluation levels for each evaluation dimension to determine which evaluation level each business question and answer pair best meets. It should be noted that the semantic matching degree refers to the degree of similarity or fit between a first business question and answer pair and the description of the evaluation criteria of a specific candidate evaluation level under a certain evaluation dimension, that is, a measure of whether the question and answer pair meets the standards of a certain evaluation level.
[0061] As an example, based on the scheduling information assigned to multiple candidate objects, the target object associated with each financial business is determined from the multiple candidate objects; according to the evaluation standard description of any first business question and answer pair and multiple candidate evaluation levels under any evaluation dimension, the target scheduling information is generated; the target scheduling information is scheduled to the target object to evaluate the semantic matching between any first business question and answer pair and the evaluation standard description of multiple candidate evaluation levels under any evaluation dimension through the target object; the evaluation information sent by the target object in response to the target scheduling information is received; wherein, the evaluation information carries the semantic matching between any first business question and answer pair and the evaluation standard description of multiple candidate evaluation levels under any evaluation dimension.
[0062] That is to say, the semantic matching degree between any first business question and answer pair and the evaluation standard descriptions of multiple candidate evaluation levels under any evaluation dimension is manually evaluated; taking the candidate object as an expert in the financial field as an example, according to the scheduling information currently assigned to the financial expert, the target expert associated with each financial business is determined from multiple financial experts, and then, any first business question and answer pair and the evaluation standard descriptions of multiple candidate evaluation levels under any evaluation dimension are packaged into target scheduling information, and the target scheduling information is scheduled to the target expert. After receiving the target scheduling information, the target expert evaluates the semantic matching degree between the question and answer pair and the evaluation standard description of each evaluation level. The target expert will return evaluation information in response to the target scheduling information, which carries the semantic matching degree between the first question and answer pair and the evaluation standard description of each evaluation level.
[0063] As another example, a first prompt template associated with the financial field is queried, wherein the first prompt template is used to indicate the first task information to be performed by the first large language model in the financial field; the first prompt template is updated using the evaluation standard descriptions of multiple candidate evaluation levels under any evaluation dimension and any first business question and answer pair to obtain first prompt information, wherein the first prompt information is used to indicate that the first task information includes a semantic matching task; the first large language model is called to perform semantic matching on the first prompt information to obtain the semantic matching degree between any first business question and answer pair and the evaluation standard descriptions of multiple candidate evaluation levels under any evaluation dimension.
[0064] That is, a large model is used to determine the semantic matching degree between any first business question-answer pair and the evaluation standard descriptions of multiple candidate evaluation levels under any evaluation dimension.
[0065] First, a first prompt template closely related to the financial field is obtained; wherein, the first prompt template is used to indicate the specific task that the first large language model needs to perform in the financial field, that is, the first task information; then, the first prompt template is updated using the evaluation standard descriptions of multiple candidate evaluation levels under any evaluation dimension and any first business question and answer pair to obtain the first prompt information; wherein, the first prompt information is used to indicate that the first large language model needs to perform a semantic matching task, that is, it is necessary to compare the semantic similarity between the first business question and answer pair and the evaluation standard description; finally, the first large language model is called to process the updated first prompt information and perform the semantic matching task, that is, the large language model calculates the semantic matching degree between the first business question and answer pair and the evaluation standard description of each candidate evaluation level.
[0066] Step 204, based on the semantic matching between any first business question and answer pair and the evaluation standard descriptions of multiple candidate evaluation levels under any evaluation dimension, determine the first evaluation level of any first business question and answer pair in any evaluation dimension from the multiple candidate evaluation levels under any evaluation dimension.
[0067] Furthermore, based on the semantic matching degree between any first business question-and-answer pair and the evaluation standard descriptions of multiple candidate evaluation levels under any evaluation dimension, an evaluation level that is adapted to the semantic matching degree is determined from the multiple candidate evaluation levels under any evaluation dimension, that is, the first evaluation level.
[0068] For example, the degree of fitness between the first business question and answer pair and the evaluation standard description of the evaluation level "Excellent" in the evaluation dimension "Professionalism" is 0.8, the degree of fitness between the first business question and answer pair and the evaluation standard description of the evaluation level "General" in the evaluation dimension "Professionalism" is 0.2, the degree of fitness between the first business question and answer pair and the evaluation standard description of the evaluation level "Pass" in the evaluation dimension "Professionalism" is 0.1, and the degree of fitness between the first business question and answer pair and the evaluation standard description of the evaluation level "Poor" in the evaluation dimension "Professionalism" is 0. It can be determined that the evaluation level of the first business question and answer pair in the evaluation dimension "Professionalism" is "Excellent".
[0069] Step 205: Obtain a second evaluation level of each second business question-answer pair evaluated in multiple evaluation dimensions.
[0070] In the embodiment of the present disclosure, the evaluation method of the second evaluation level of each second business question and answer pair in multiple evaluation dimensions can refer to the evaluation method of the second evaluation level of the first business question and answer pair in multiple evaluation dimensions, that is, steps 202 to 204, which will not be repeated in this disclosure.
[0071] Step 206 , based on the first evaluation level of each first business question and answer pair in multiple evaluation dimensions and the second evaluation level of each second business question and answer pair in multiple evaluation dimensions, screen the multiple first business question and answer pairs and the multiple second business question and answer pairs to obtain sample business question and answer pairs.
[0072] Step 207: Train the financial big model based on the sample business question and answer pairs.
[0073] It should be noted that the execution process of step 201 and step 207 can be implemented in any way in the embodiments of the present disclosure, and the embodiments of the present disclosure do not limit this and will not be described in detail.
[0074] In summary, by obtaining the evaluation standard descriptions of multiple candidate evaluation levels under each evaluation dimension; for any evaluation dimension, determining the semantic matching degree between any first business question and answer pair and the evaluation standard descriptions of multiple candidate evaluation levels under any evaluation dimension; based on the semantic matching degree between any first business question and answer pair and the evaluation standard descriptions of multiple candidate evaluation levels under any evaluation dimension, determining the first evaluation level of any first business question and answer pair in any evaluation dimension from the multiple candidate evaluation levels under any evaluation dimension, thereby obtaining the evaluation standard descriptions of multiple candidate evaluation levels under each evaluation dimension, ensuring the clarity and comprehensiveness of the evaluation basis; and then for any evaluation dimension, by calculating the semantic matching degree between any first business question and answer pair and these evaluation standard descriptions, the performance of the question and answer pair in different dimensions can be precisely quantified; finally, based on these semantic matching degrees, determining the first evaluation level of the question and answer pair in each dimension from multiple candidate evaluation levels, not only realizes the automation and standardization of evaluation, but also avoids the subjectivity and inconsistency of human evaluation, thereby improving the reliability of the evaluation results.
[0075] In order to clearly illustrate how in the above embodiment, multiple first business question and answer pairs and multiple second business question and answer pairs are screened according to the first evaluation level of each first business question and answer pair in multiple evaluation dimensions and the second evaluation level of each second business question and answer pair in multiple evaluation dimensions to obtain sample business question and answer pairs, the present disclosure proposes another training method for a large financial model.
[0076] Figure 3 It is a flowchart of the training method of the financial big model shown in the third embodiment of the present disclosure.
[0077] like Figure 3 As shown in Figure 1, the training method of the financial model includes the following steps:
[0078] Step 301: Acquire multiple first business question-answer pairs and multiple second business question-answer pairs associated with each financial business.
[0079] The first business question and answer pair and the second business question and answer pair have different sources.
[0080] Step 302: Obtain a first evaluation grade obtained by evaluating each first business question-answer pair in multiple evaluation dimensions and a second evaluation grade obtained by evaluating each second business question-answer pair in multiple evaluation dimensions.
[0081] Step 303: Filter the multiple first business question and answer pairs and the multiple second business question and answer pairs according to the first evaluation level of each first business question and answer pair in the multiple evaluation dimensions and the second evaluation level of each second business question and answer pair in the multiple evaluation dimensions to obtain the retained first business question and answer pairs and second business question and answer pairs.
[0082] In order to ensure the high quality of the sample business question and answer pairs, as a possible implementation method, the first business question and answer pairs and the second business question and answer pairs are screened to obtain the retained first business question and answer pairs and the retained second business question and answer pairs.
[0083] In order to make the question and answer pairs meet the professionalism and anthropomorphic characteristics at the same time, in the embodiment of the present disclosure, multiple first business question and answer pairs are screened based on the first evaluation level of each first business question and answer pair in multiple evaluation dimensions. For example, if the evaluation levels of a first business question and answer pair in the evaluation dimensions of "professionalism", "service attitude" and "communication skills" are all poor, then the first business question and answer pair will be deleted.
[0084] At the same time, based on the second evaluation level of each second business question and answer in multiple evaluation dimensions, multiple second business question and answer pairs are screened.
[0085] Step 304 : For any financial business, calculate the semantic similarity between the reserved first business question-answer pair and the reserved second business question-answer pair associated with any financial business.
[0086] In order to further improve the quality of the sample business question and answer pairs, as a possible implementation method, based on the semantic similarity between the retained first sample business question and answer pairs and the retained second sample business question and answer pairs, the sample business question and answer pairs are determined from the retained first business question and answer pairs and the retained second business question and answer pairs.
[0087] As an example, for any reserved second business question-answer pair under any financial business, the semantic similarity between the any reserved second business question-answer pair and each reserved first business question-answer pair under the any financial business is determined.
[0088] For example, any retained second business question-answer pair under any financial business is converted into a first vector, and each retained first business question-answer pair is converted into a second vector. Then, the cosine similarity between the first vector and each second vector is calculated, and the cosine similarity between the first vector and each second vector is used as the similarity between any retained second business question-answer pair and each retained first business question-answer pair. The cosine similarity between the first vector and each second vector is expressed as follows:
[0089]
[0090] Cosine_Similarity(A,B) is the cosine similarity between the first vector and each second vector, where A represents the first vector, B represents the second vector, * represents the dot product between the first vector and the second vector, ||A|| represents the modulus of the first vector, and ||B|| represents the modulus of the second vector.
[0091] Step 305 : Determine sample business question-answer pairs from the retained first business question-answer pairs and the retained second business question-answer pairs according to the semantic similarity between the retained first business question-answer pairs and the retained second business question-answer pairs associated with each financial business.
[0092] In order to remove irrelevant or highly similar question-answer pairs, further, based on the semantic similarity between the retained first business question-answer pairs and the retained second business question-answer pairs associated with each financial business, irrelevant or highly similar question-answer pairs are removed from the retained first business question-answer pairs and the retained second business question-answer pairs, and sample business question-answer pairs are determined based on the removed question-answer pairs.
[0093] In the embodiment of the present disclosure, based on the semantic similarity and the first similarity threshold between the retained first business question and answer pairs and the retained second business question and answer pairs associated with each financial business, the retained second business question and answer pairs are first screened to obtain the second business question and answer pairs after the first screening; based on the semantic similarity and the second similarity threshold between the second business question and answer pairs after the first screening and the retained first business question and answer pairs associated with each financial business, the second business question and answer pairs after the first screening are second screened to obtain the second business question and answer pairs after the second screening; based on the retained first business question and answer pairs and the second business question and answer pairs after the second screening, sample business question and answer pairs are generated. It should be noted that the second similarity threshold is greater than the first similarity threshold.
[0094] That is to say, if the semantic similarity between the retained first business question and answer pair and the retained second business question and answer pair associated with each financial business is greater than or equal to the first similarity threshold (e.g., 0.6), it is determined that there is sufficient correlation between the retained first business question and answer pair and the retained second business question and answer pair, and the second business question and answer pair is retained for the next round of screening; otherwise, it is discarded; furthermore, if the semantic similarity between the second business question and answer pair after the first screening of each financial business association and the retained first business question and answer pair is higher than the second similarity threshold (e.g., 0.9), it means that the second business question and answer pair after the first screening and the retained first business question and answer pair are the same or repeated, and the second business question and answer pair after the first screening is discarded, otherwise it is retained; finally, based on the retained first business question and answer pair and the second business question and answer pair after the second screening, a sample business question and answer pair is generated.
[0095] The steps of generating a sample business question-answer pair based on the retained first business question-answer pair and the second business question-answer pair after the second screening are as follows:
[0096] (1) For any financial business, obtain the proportion of first business question-answer pairs associated with the financial business among multiple first business question-answer pairs;
[0097] To ensure that the sample business question-and-answer pairs meet actual user needs and complex business scenarios, in this disclosed embodiment, the proportion of first business question-and-answer pairs associated with any financial business among multiple first business question-and-answer pairs is calculated. For example, financial services include: card application, deposits, loans, corporate services, complaints, financial management, and debt collection. Among them, the proportion of first business question-and-answer pairs associated with card application is 15%, the proportion of first business question-and-answer pairs associated with deposits is 25%, the proportion of first business question-and-answer pairs associated with loans is 22%, the proportion of first business question-and-answer pairs associated with corporate services is 18%, the proportion of first business question-and-answer pairs associated with complaints is 5%, the proportion of first business question-and-answer pairs associated with financial management is 13%, and the proportion of first business question-and-answer pairs associated with debt collection is 2%. It should be noted that the first business question-and-answer pairs associated with each financial business may include business question-and-answer pairs with different evaluation levels based on different evaluation dimensions.
[0098] (2) According to the proportion, determine a sample business question-answer pair associated with any financial business from the retained first business question-answer pairs and the second business question-answer pairs after the second screening.
[0099] Furthermore, for any financial business, the proportion of the sample business Q&A pairs determined from the retained first business Q&A pairs and the second business Q&A pairs after the second screening associated with the financial business is maintained consistent with the proportion of the first business Q&A pairs associated with the financial business. If the proportion of the sample business Q&A pairs associated with the financial business is inconsistent with the proportion of the first business Q&A pairs, the sample business Q&A pairs are adjusted so that the proportion of the adjusted sample business Q&A pairs is consistent with the proportion of the first business Q&A pairs. It should be noted that the proportion of business Q&A pairs with each evaluation level across multiple evaluation dimensions in the sample business Q&A pairs associated with any financial business is also consistent with the proportion of business Q&A pairs with each evaluation level across multiple evaluation dimensions in the first business Q&A pairs of the financial business. For example, the proportion of Q&A pairs with the evaluation dimension "Professionalism" and the evaluation level "Excellent" in the sample business Q&A pairs associated with the financial business "Card Application" is consistent with the proportion of the first business Q&A pairs with the evaluation dimension "Professionalism" and the evaluation level "Excellent" in the first business Q&A pairs associated with the financial business "Card Application."
[0100] Step 306: Train the financial big model based on the sample business question and answer pairs.
[0101] It should be noted that the execution process of step 301 to step 302 and step 306 can be implemented in any way in the embodiments of the present disclosure, and the embodiments of the present disclosure do not limit this and will not be described in detail.
[0102] In summary, multiple first and second business question-answer pairs are screened based on the first evaluation rating of each first business question-answer pair across multiple evaluation dimensions and the second evaluation rating of each second business question-answer across multiple evaluation dimensions to obtain retained first and second business question-answer pairs. For any financial business, the semantic similarity between the retained first and second business question-answer pairs associated with each financial business is calculated. Based on the semantic similarity between the retained first and second business question-answer pairs associated with each financial business, sample business question-answer pairs are determined from the retained first and second business question-answer pairs. Thus, by combining the ratings of each business question-answer pair across multiple evaluation dimensions and calculating the semantic similarity between the retained question-answer pairs to determine the sample business question-answer pairs, the quality of the screened question-answer pairs is guaranteed. By calculating the semantic similarity, redundant and repetitive question-answer pairs are effectively reduced, and the pertinence and relevance of the sample business question-answer pairs are improved. This improves the efficiency of the training model, while enhancing the model's understanding and responsiveness to financial business, and improving the user experience.
[0103] In order to clearly illustrate how to obtain multiple first business question-answer pairs and multiple second business question-answer pairs associated with each financial business in the above embodiment, the present disclosure proposes another training method for a large financial model.
[0104] Figure 4 It is a flowchart of the training method of the financial big model shown in the fourth embodiment of the present disclosure.
[0105] like Figure 4 As shown in Figure 1, the training method of the financial model includes the following steps:
[0106] Step 401: Acquire multiple historical business question-answer pairs associated with various financial businesses in a financial institution.
[0107] In the disclosed embodiments, historical data (e.g., customer service records) is extracted from multiple systems of a financial institution. Question-and-answer pairs directly related to each financial service are screened from the extracted historical data to form multiple historical service question-and-answer pairs associated with each financial service. It should be noted that the multiple historical service question-and-answer pairs associated with each financial service do not include duplicate or invalid question-and-answer pairs.
[0108] Step 402: Use the multiple historical business question-answer pairs as multiple first business question-answer pairs.
[0109] Furthermore, a plurality of historical business question-answer pairs associated with each financial business are used as a plurality of first business question-answer pairs associated with each financial business.
[0110] Step 403: Acquire multiple second business question-answer pairs associated with each financial business from multiple data sources.
[0111] In order to improve the richness of the sample business question-answer pairs, as a possible implementation method, multiple second business question-answer pairs associated with various financial businesses are obtained from multiple different sources.
[0112] In an embodiment of the present disclosure, multiple data sources include: financial web pages and a second largest language model, which capture multiple business question-and-answer pairs associated with each financial business from the financial web pages; query a second prompt template associated with the financial field, wherein the second prompt template is used to indicate the second task information to be performed by the second largest language model in the financial field; use business knowledge and business problems associated with each financial business to update the second prompt template to obtain second prompt information, wherein the second prompt information is used to instruct the second largest language model to perform a question-and-answer pair generation task; call the second largest language model to generate question-and-answer pairs for the first prompt information to obtain multiple output question-and-answer pairs; and use the multiple business question-and-answer pairs and the multiple output question-and-answer pairs as multiple second business question-and-answer pairs associated with each financial business.
[0113] That is to say, multiple collected business question and answer pairs related to financial business are captured from financial web pages. The collected business question and answer pairs directly reflect the questions of customers in actual financial transactions and the answers of financial institutions, providing real and specific materials for the question and answer pair library; then, a second prompt template related to the financial field is queried. The second prompt template is used to indicate the specific tasks that the second language model needs to perform in the financial field, that is, to generate question and answer pairs related to financial business; further, the second prompt template is updated and refined using business knowledge and business problems associated with each financial business, and second prompt information is obtained. The second prompt information describes the content and scope of the question and answer pairs to be generated, and provides detailed guidance for the second language model to perform the question and answer pair generation task; finally, the second language model is called to generate question and answer pairs according to the second prompt information, and multiple output question and answer pairs are obtained. The multiple collected business question and answer pairs and the multiple output question and answer pairs are collectively used as multiple second business question and answer pairs associated with each financial business.
[0114] Step 404 : Obtain a first evaluation grade obtained by evaluating each first business question and answer pair in multiple evaluation dimensions and a second evaluation grade obtained by evaluating each second business question and answer in multiple evaluation dimensions.
[0115] Step 405 , based on the first evaluation level of each first business question and answer pair in multiple evaluation dimensions and the second evaluation level of each second business question and answer pair in multiple evaluation dimensions, screen the multiple first business question and answer pairs and the multiple second business question and answer pairs to obtain sample business question and answer pairs.
[0116] Step 406: Train the financial big model based on the sample business question and answer pairs.
[0117] It should be noted that the execution process of step 404 to step 406 can be implemented in any way in the embodiments of the present disclosure, and the embodiments of the present disclosure do not limit this and will not be described in detail.
[0118] In summary, by obtaining multiple historical business question and answer pairs related to various financial businesses in financial institutions; using the multiple historical business question and answer pairs as multiple first business question and answer pairs; and obtaining multiple second business question and answer pairs related to various financial businesses from multiple data sources, the diversity and comprehensiveness of the first business question and answer pairs and the second business question and answer pairs are enriched, and thus, sample business question and answer pairs are determined from the first business question and answer pairs and the second business question and answer pairs, and the sample business question and answer pairs are used to train the financial big model, thereby improving the accuracy and robustness of the financial big model.
[0119] above Figures 1 to 4 As an embodiment of a training method for a large financial model, the present disclosure also proposes a method for processing business problems.
[0120] Figure 5 It is a flowchart of a method for processing a business problem shown in the fifth embodiment of the present disclosure.
[0121] like Figure 5 As shown, the method for handling this business problem includes the following steps:
[0122] Step 501: Obtain a target business problem associated with a target financial business.
[0123] In the embodiment of the present disclosure, the target financial business may be a designated financial business, for example, the target business is "card application", and the target business question associated with the target financial business may be a question asked by the user in the process of applying for the target financial business, for example, the target business question is "how to apply for a certain card".
[0124] Step 502: Input the target business problem into the trained financial big model to obtain the target business answer to the target business problem output by the financial big model.
[0125] Among them, the financial big model is Figures 1 to 4 The method described in the embodiment is trained.
[0126] In order to improve the accuracy of the target business answer, in the embodiment of the present disclosure, a trained financial big model is used to generate the target business answer to the target business question. For example, the target business question is input into the trained financial big model, so that the target business answer to the target business question output by the financial big model can be obtained. Figures 1 to 4The financial large model obtained by training in the embodiment.
[0127] The method for processing business problems in the embodiment of the present disclosure obtains a target business problem associated with a target financial business; inputs the target business problem into a trained financial big model to obtain a target business answer to the target business problem output by the financial big model, thereby improving the response speed to business needs, while improving the accuracy and professionalism of the target business answer and improving the user experience.
[0128] With the above Figures 1 to 4 Corresponding to the training method of the financial big model provided in the embodiment, the present disclosure also provides a training device for the financial big model. Since the training device for the financial big model provided in the embodiment of the present disclosure corresponds to the training method for the financial big model provided in the above embodiment, the implementation method of the training method for the financial big model is also applicable to the training device for the financial big model provided in the embodiment of the present disclosure, and will not be described in detail in the embodiment of the present disclosure.
[0129] Figure 6 It is a structural diagram of the training device for the financial big model shown in the sixth embodiment of the present disclosure.
[0130] like Figure 6 As shown, the training device 600 for the financial large model includes: a first acquisition module 610, a second acquisition module 620, a screening module 630 and a training module 640.
[0131] Among them, the first acquisition module 610 is used to obtain multiple first business question and answer pairs and multiple second business question and answer pairs associated with each financial business; wherein, the sources of the first business question and answer pairs and the second business question and answer pairs are different; the second acquisition module 620 is used to obtain the first evaluation level obtained by evaluating each first business question and answer pair in multiple evaluation dimensions and the second evaluation level obtained by evaluating each second business question and answer pair in multiple evaluation dimensions; the screening module 630 is used to screen the multiple first business question and answer pairs and the multiple second business question and answer pairs according to the first evaluation level of each first business question and answer pair in multiple evaluation dimensions and the second evaluation level of each second business question and answer in multiple evaluation dimensions to obtain sample business question and answer pairs; the training module 640 is used to train the financial big model based on the sample business question and answer pairs.
[0132] As a possible implementation method of the embodiment of the present disclosure, the second acquisition module 620 is used to obtain the evaluation standard description of multiple candidate evaluation levels under each evaluation dimension; for any evaluation dimension, determine the semantic matching degree between any first business question and answer pair and the evaluation standard description of multiple candidate evaluation levels under any evaluation dimension; based on the semantic matching degree between any first business question and answer pair and the evaluation standard description of multiple candidate evaluation levels under any evaluation dimension, determine the first evaluation level of any first business question and answer pair in any evaluation dimension from the multiple candidate evaluation levels under any evaluation dimension.
[0133] As a possible implementation method of the embodiment of the present disclosure, the second acquisition module 620 is used to determine the target object associated with each financial business from multiple candidate objects based on the scheduling information assigned to multiple candidate objects; generate target scheduling information according to the evaluation standard description of any first business question and answer pair and multiple candidate evaluation levels under any evaluation dimension; schedule the target scheduling information to the target object to evaluate the semantic matching between any first business question and answer pair and the evaluation standard description of multiple candidate evaluation levels under any evaluation dimension through the target object; receive the evaluation information sent by the target object in response to the target scheduling information; wherein the evaluation information carries the semantic matching between any first business question and answer pair and the evaluation standard description of multiple candidate evaluation levels under any evaluation dimension.
[0134] As a possible implementation method of an embodiment of the present disclosure, the second acquisition module 620 is used to query a first prompt template associated with the financial field, wherein the first prompt template is used to indicate the first task information to be performed by the first large language model in the financial field; using the evaluation standard description of multiple candidate evaluation levels under any evaluation dimension and any first business question and answer pair, the first prompt template is updated to obtain first prompt information, wherein the first prompt information is used to indicate that the first task information includes a semantic matching task; calling the first large language model to perform semantic matching on the first prompt information to obtain the semantic matching degree between any first business question and answer pair and the evaluation standard description of multiple candidate evaluation levels under any evaluation dimension.
[0135] As a possible implementation method of the embodiment of the present disclosure, the screening module 630 is used to screen multiple first business question and answer pairs and multiple second business question and answer pairs according to the first evaluation level of each first business question and answer pair in multiple evaluation dimensions and the second evaluation level of each second business question and answer in multiple evaluation dimensions, to obtain retained first business question and answer pairs and second business question and answer pairs; for any financial business, calculate the semantic similarity between the retained first business question and answer pair and the retained second business question and answer pair associated with any financial business; according to the semantic similarity between the retained first business question and answer pair and the retained second business question and answer pair associated with each financial business, determine sample business question and answer pairs from the retained first business question and answer pairs and the retained second business question and answer.
[0136] As a possible implementation method of the embodiment of the present disclosure, the screening module 630 is used to perform a first screening on the retained second business question and answer pairs associated with each financial business based on the semantic similarity and the first similarity threshold between the retained first business question and answer pairs and the retained second business question and answer pairs, so as to obtain the second business question and answer pairs after the first screening; perform a second screening on the second business question and answer pairs after the first screening based on the semantic similarity and the second similarity threshold between the second business question and answer pairs after the first screening and the retained first business question and answer pairs associated with each financial business, so as to obtain the second business question and answer pairs after the second screening; generate sample business question and answer pairs based on the retained first business question and answer pairs and the second business question and answer pairs after the second screening.
[0137] As a possible implementation method of the embodiment of the present disclosure, the screening module 630 is used to obtain, for any financial business, the proportion of first business question and answer pairs associated with any financial business among multiple first business question and answer pairs; based on the proportion, determine the sample business question and answer pairs associated with any financial business from the retained first business question and answer pairs and the second business question and answer pairs after the second screening.
[0138] As a possible implementation method of the embodiment of the present disclosure, the first acquisition module 610 is used to obtain multiple historical business question and answer pairs associated with various financial businesses in a financial institution; use the multiple historical business question and answer pairs as multiple first business question and answer pairs; and obtain multiple second business question and answer pairs associated with each financial business from multiple data sources.
[0139] As a possible implementation method of the embodiment of the present disclosure, multiple data sources include: financial web pages and a second large language model, a first acquisition module 610, which is used to capture multiple collection business question and answer pairs associated with each financial business from the financial web page; query a second prompt template associated with the financial field, wherein the second prompt template is used to indicate the second task information to be performed by the second large language model in the financial field; use the business knowledge and business problems associated with each financial business to update the second prompt template to obtain second prompt information, wherein the second prompt information is used to instruct the second large language model to perform the question and answer pair generation task; call the second large language model to generate question and answer pairs for the first prompt information to obtain multiple output question and answer pairs; use the multiple collection business question and answer pairs and the multiple output question and answer pairs as multiple second business question and answer pairs associated with each financial business.
[0140] The training device for the financial big model of the embodiment of the present disclosure obtains a first business question-and-answer pair and a second business question-and-answer pair from different sources and evaluates them separately on multiple evaluation dimensions, thereby ensuring the diversity and comprehensiveness of the data set, thereby enhancing the generalization ability and adaptability of the model; secondly, by evaluating each question-and-answer pair on multiple evaluation dimensions and assigning corresponding evaluation grades, it is possible to accurately identify high-quality, high-value sample business question-and-answer pairs, avoid low-quality or misleading information from interfering with the learning process of the model, and ensure the high standard of the training data set; finally, the financial big model is trained based on the screened sample business question-and-answer pairs, so that the model can not only understand basic financial knowledge, but also master complex business logic and the latest market trends, so that it can respond to user queries more accurately in actual applications and provide personalized and professional advice and services.
[0141] With the above Figure 5 Corresponding to the method for processing business problems provided in the embodiment, the present disclosure also provides a device for processing business problems. Since the device for processing business problems provided in the embodiment of the present disclosure corresponds to the method for processing business problems provided in the above embodiment, the implementation method of the method for processing business problems is also applicable to the device for processing business problems provided in the embodiment of the present disclosure, and will not be described in detail in the embodiment of the present disclosure.
[0142] Figure 7 It is a structural diagram of a device for processing business problems shown in the seventh embodiment of the present disclosure.
[0143] like Figure 7 As shown, the business problem processing device 700 includes: an acquisition module 710 and an output module 720.
[0144] Among them, the acquisition module 710 is used to obtain the target business problem associated with the target financial business; the output module 720 is used to input the target business problem into the trained financial big model to obtain the target business answer to the target business problem output by the financial big model; wherein, the financial big model is trained using the financial big model training device of the third aspect embodiment.
[0145] The business problem processing device of the embodiment of the present disclosure obtains a target business problem associated with a target financial business; inputs the target business problem into a trained financial big model to obtain a target business answer to the target business problem output by the financial big model, thereby improving the response speed to business needs, while improving the accuracy and professionalism of the target business answer and improving the user experience.
[0146] In an exemplary embodiment, an electronic device is also provided.
[0147] Among them, electronic equipment includes:
[0148] processor;
[0149] a memory for storing processor-executable instructions;
[0150] The processor is configured to execute instructions to implement the training method of the financial big model or the method for processing business problems proposed in any of the aforementioned embodiments.
[0151] As an example, Figure 8 is a structural diagram of an electronic device 800 shown in an exemplary embodiment of the present disclosure, such as Figure 8 As shown, the electronic device 800 may further include:
[0152] The memory 810 and the processor 820, a bus 830 connecting different components (including the memory 810 and the processor 820), the memory 810 stores a computer program, and when the processor 820 executes the program, the training method of the financial big model described in the embodiment of the present disclosure, or the method for processing business problems is implemented.
[0153] Bus 830 represents one or more of several types of bus structures, including a memory bus or memory controller, a peripheral bus, an accelerated graphics port, a processor, or a local bus using any of a variety of bus architectures. Examples of these architectures include, but are not limited to, an Industry Standard Architecture (ISA) bus, a Micro Channel Architecture (MAC) bus, an Enhanced ISA bus, a Video Electronics Standards Association (VESA) local bus, and a Peripheral Component Interconnect (PCI) bus.
[0154] The electronic device 800 typically includes a variety of electronic device-readable media. These media can be any available media that can be accessed by the electronic device 800, including volatile and non-volatile media, removable and non-removable media.
[0155] The memory 810 may also include computer system readable media in the form of volatile memory, such as random access memory (RAM) 840 and / or cache memory 850. The server 800 may further include other removable / non-removable, volatile / non-volatile computer system storage media. By way of example only, the storage system 860 may be used to read and write non-removable, non-volatile magnetic media ( Figure 8 Not shown, often called a "hard drive"). Although Figure 8 Not shown, a disk drive for reading and writing to a removable non-volatile disk (e.g., a "floppy disk"), and an optical drive for reading and writing to a removable non-volatile optical disk (e.g., a CD-ROM, DVD-ROM, or other optical media) may be provided. In these cases, each drive may be connected to bus 830 via one or more data medium interfaces. Memory 810 may include at least one program product having a set (e.g., at least one) of program modules configured to perform the functions of various embodiments of the present disclosure.
[0156] A program / utility 880 having a set (at least one) of program modules 870 may be stored, for example, in memory 810. Such program modules 870 include, but are not limited to, an operating system, one or more application programs, other program modules, and program data, each of which, or some combination thereof, may include an implementation of a network environment. Program modules 870 generally implement the functions and / or methods of the embodiments described herein.
[0157] The electronic device 800 can also communicate with one or more external devices 890 (e.g., a keyboard, a pointing device, a display 791, etc.), one or more devices that enable a user to interact with the electronic device 800, and / or any device that enables the electronic device 800 to communicate with one or more other computing devices (e.g., a network card, a modem, etc.). Such communication can occur via an input / output (I / O) interface 892. Furthermore, the electronic device 800 can communicate with one or more networks (e.g., a local area network (LAN), a wide area network (WAN), and / or a public network such as the Internet) via a network adapter 893. As shown, the network adapter 893 communicates with other modules of the electronic device 800 via a bus 830. It should be understood that, although not shown in the figure, other hardware and / or software modules can be used in conjunction with the electronic device 800, including but not limited to microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems.
[0158] The processor 820 executes various functional applications and data processing by running programs stored in the memory 810 .
[0159] It should be noted that the implementation process and technical principles of the electronic device of this embodiment can be found in the aforementioned explanation of the training method of the financial big model or the method for handling business problems in the embodiment of this disclosure, and will not be repeated here.
[0160] In an exemplary embodiment, a computer-readable storage medium including instructions is also provided, such as a memory including instructions. The instructions can be executed by a processor of an electronic device to implement the method for training a large financial model or the method for processing a business problem proposed in any of the above embodiments. Alternatively, the computer-readable storage medium can be a ROM, random access memory (RAM), CD-ROM, magnetic tape, floppy disk, or optical data storage device.
[0161] In an exemplary embodiment, a computer program product is also provided, including a computer program / instruction, characterized in that when the computer program / instruction is executed by a processor, it implements the training method of the financial big model or the method for processing business problems proposed in any of the above embodiments.
[0162] Other embodiments of the present disclosure will readily occur to those skilled in the art after considering the specification and practicing the invention disclosed herein. This disclosure is intended to cover any variations, uses, or adaptations of the present disclosure that follow the general principles of the present disclosure and include common knowledge or customary techniques in the art not disclosed herein. The description and examples are to be considered as exemplary only, with the true scope and spirit of the present disclosure being indicated by the following claims.
[0163] It should be understood that the present disclosure is not limited to the exact structures that have been described above and shown in the drawings, and that various modifications and changes can be made without departing from the scope thereof. The scope of the present disclosure is limited only by the appended claims.
Claims
1. A training method for a large financial model, characterized in that: include: Acquire multiple first business question-and-answer pairs and multiple second business question-and-answer pairs associated with each financial business; wherein the first business question-and-answer pairs and the second business question-and-answer pairs have different sources; Obtaining a first evaluation grade obtained by evaluating each first business question and answer pair in multiple evaluation dimensions and a second evaluation grade obtained by evaluating each second business question and answer pair in the multiple evaluation dimensions; screening the plurality of first business question-and-answer pairs and the plurality of second business question-and-answer pairs according to the first evaluation level of each first business question-and-answer pair in the plurality of evaluation dimensions and the second evaluation level of each second business question-and-answer pair in the plurality of evaluation dimensions to obtain sample business question-and-answer pairs; The financial big model is trained based on the sample business question and answer pairs.
2. The method according to claim 1, characterized in that The obtaining of a first evaluation grade obtained by evaluating each first business question-answer pair based on multiple evaluation dimensions includes: Obtain descriptions of evaluation criteria for multiple candidate evaluation levels under each evaluation dimension; For any evaluation dimension, determining a semantic matching degree between any first business question-answer pair and the evaluation criteria descriptions of multiple candidate evaluation levels under the any evaluation dimension; Based on the semantic matching degree between the evaluation standard descriptions of any first business question and answer pair and the multiple candidate evaluation levels under any evaluation dimension, the first evaluation level of any first business question and answer pair in any evaluation dimension is determined from the multiple candidate evaluation levels under any evaluation dimension.
3. The method according to claim 2, characterized in that The determining, for any evaluation dimension, the semantic matching degree between any first business question-answer pair and the evaluation standard descriptions of multiple candidate evaluation levels under any evaluation dimension includes: determining, from the plurality of candidate objects, a target object associated with each of the financial services based on the assigned scheduling information of the plurality of candidate objects; Generate target scheduling information according to the evaluation standard description of any first business question-answer pair and multiple candidate evaluation levels under any evaluation dimension; Dispatching the target scheduling information to the target object, so as to evaluate the semantic matching degree between any first business question-answer pair and the evaluation standard descriptions of the multiple candidate evaluation levels under any evaluation dimension through the target object; Receive evaluation information sent by the target object in response to the target scheduling information; wherein, the evaluation information carries the semantic matching degree between any first business question-answer pair and the evaluation standard description of multiple candidate evaluation levels under any evaluation dimension.
4. The method according to claim 2, characterized in that The determining, for any evaluation dimension, the semantic matching degree between any first business question-answer pair and the evaluation standard descriptions of multiple candidate evaluation levels under any evaluation dimension includes: Querying a first prompt template associated with the financial field, wherein the first prompt template is used to indicate first task information to be performed by a first language model in the financial field; Using the evaluation criteria descriptions of the multiple candidate evaluation levels under any evaluation dimension and any first business question-answer pair, the first prompt template is updated to obtain first prompt information, wherein the first prompt information is used to indicate that the first task information includes a semantic matching task; The first large language model is called to perform semantic matching on the first prompt information to obtain a semantic matching degree between any first business question-answer pair and the evaluation standard descriptions of multiple candidate evaluation levels under any evaluation dimension.
5. The method according to claim 1, wherein The screening of the plurality of first business question-answer pairs and the plurality of second business question-answer pairs according to the first evaluation level of each first business question-answer pair in the plurality of evaluation dimensions and the second evaluation level of each second business question-answer in the plurality of evaluation dimensions to obtain sample business question-answer pairs includes: screening the plurality of first business question-and-answer pairs and the plurality of second business question-and-answer pairs according to the first evaluation level of each first business question-and-answer pair in the plurality of evaluation dimensions and the second evaluation level of each second business question-and-answer pair in the plurality of evaluation dimensions to obtain retained first business question-and-answer pairs and second business question-and-answer pairs; For any financial business, calculating the semantic similarity between the reserved first business question-answer pair and the reserved second business question-answer pair associated with the any financial business; According to the semantic similarity between the retained first business question-answer pair and the retained second business question-answer pair associated with each financial business, a sample business question-answer pair is determined from the retained first business question-answer pair and the retained second business question-answer pair.
6. The method according to claim 5, characterized in that The determining, based on the semantic similarity between the retained first business question-answer pair and the retained second business question-answer pair associated with each financial business, a sample business question-answer pair from the retained first business question-answer pair and the retained second business question-answer pair, includes: performing a first screening on the retained second business question-answer pairs according to the semantic similarity between the retained first business question-answer pairs and the retained second business question-answer pairs associated with each financial business and a first similarity threshold, to obtain first-screened second business question-answer pairs; performing a second screening on the first-screened second business question-answer pairs according to the semantic similarity and the second similarity threshold between the first-screened second business question-answer pairs associated with each of the financial businesses and the retained first business question-answer pairs to obtain second-screened second business question-answer pairs; The sample business question-answer pair is generated based on the retained first business question-answer pair and the second business question-answer pair after the second screening.
7. The method according to claim 6, characterized in that Generating the sample business question-answer pair according to the retained first business question-answer pair and the second business question-answer pair after the second screening includes: For any financial business, obtaining a proportion of first business question-answer pairs associated with the any financial business among the multiple first business question-answer pairs; According to the proportion, a sample business question and answer pair associated with any financial business is determined from the retained first business question and answer pairs and the second business question and answer pairs after the second screening.
8. The method according to claim 1, characterized in that The obtaining of a plurality of first business question-answer pairs and a plurality of second business question-answer pairs associated with each financial business includes: Acquire multiple historical business question-answer pairs associated with each of the financial businesses in a financial institution; Using the multiple historical business question-answer pairs as the multiple first business question-answer pairs; A plurality of second business question-answer pairs associated with each of the financial businesses are obtained from a plurality of data sources.
9. The method according to claim 8, characterized in that The multiple data sources include: financial web pages and a second language model, and obtaining multiple second business question-answer pairs associated with each of the financial businesses from the multiple data sources includes: crawling a plurality of collected business question-answer pairs associated with each of the financial businesses from the financial webpage; Querying a second prompt template associated with the financial field, wherein the second prompt template is used to indicate second task information to be performed by the second language model in the financial field; Using business knowledge and business problems associated with each of the financial businesses, the second prompt template is updated to obtain second prompt information, wherein the second prompt information is used to instruct the second language model to perform the question-answer pair generation task; Calling the second largest language model to generate question-answer pairs for the first prompt information to obtain multiple output question-answer pairs; The multiple collected business question-answer pairs and the multiple output question-answer pairs are used as multiple second business question-answer pairs associated with each of the financial businesses.
10. A method for handling business problems, characterized in that: include: Obtaining target business issues associated with target financial business; The target business problem is input into a trained financial big model to obtain a target business answer to the target business problem output by the financial big model; wherein the financial big model is trained using the method described in any one of claims 1-9.
Citation Information
Patent Citations
Financial customer service dialogue generation method and device based on LLM model, equipment and medium
CN117520503A
Machine learning model training method and device and knowledge base generation method and device
CN119782481A
Method for carrying out question and answer pair scoring based on reward mechanism to realize large model fine tuning
CN119940467A