Automatic evaluation method for large model in finance and taxation field

By constructing test sets and evaluation keywords in the financial and taxation field, combining with the , or , non-combination methods and objective multiple-choice conversion, the evaluation effect deviation of the automated evaluation of the financial and taxation model in the professional field is solved, and fast and accurate evaluation and model iteration support is achieved.

CN119938512APending Publication Date: 2025-05-06AISINO CORPORATION
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411765138.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-03
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

The existing automated evaluation methods of fiscal and taxation models have deviations in the field of fiscal and taxation professionals, and are complex in debugging and optimization and high cost.

Method used

By preparing test sets in the financial and taxation field, building evaluation keywords, and using three combinations of versus, or, and non, to evaluate evaluation use cases, and convert subjective questions into objective multiple-choice questions to achieve automated evaluation.

Benefits of technology

It realizes fast and accurate automated evaluation of the fiscal and taxation model, supports rapid iteration and upgrading of the model, simplifies the maintenance and debugging of evaluation use cases, and improves work efficiency and evaluation accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119938512A_ABST
    Figure CN119938512A_ABST
Patent Text Reader

Abstract

An automatic evaluation method for a large model in the finance and taxation field comprises the steps of preparing a test set; aiming at each evaluation case, respectively constructing an evaluation keyword used for evaluating whether the answer result of the to-be-evaluated large model is correct or wrong; for each evaluation case, combining the evaluation keyword corresponding to the test case by using at least one of three combination modes of AND, OR and non, and comparing each combination mode with a corresponding expected result; converting input questions of the evaluation cases into objective selection questions; using the test set to test the to-be-evaluated large model, obtaining an answer result of the to-be-evaluated large model for each evaluation case, and evaluating the answer result in an evaluation keyword combination mode corresponding to the evaluation case to obtain a first evaluation result; and counting the first evaluation result to obtain the evaluation accuracy of the to-be-evaluated large model. According to the method, the large finance and taxation model can be automatically evaluated quickly and accurately, and evaluation result data support can be provided for quick iterative upgrading of the large finance and taxation model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of large model evaluation, and in particular to an automated evaluation method for large models in the finance and taxation field. Background Art

[0002] The finance and taxation big model uses technologies such as continuous pre-training, SFT (SFT, Supervised Fine-Tuning), and retrieval enhancement to provide intelligent finance and taxation consulting, automated finance and taxation business processing, and assisted finance and taxation examination services to corporate and institutional users in the finance and taxation field.

[0003] In terms of automated evaluation of large models, GPT4 is currently generally used as a referee model. The referee model sorts, selects or scores the answers of the large model, and uses the answers with the highest ranking or the highest score as the output of the large model to replace natural persons to complete the evaluation. In basic general fields such as knowledge encyclopedias, text summarization, and code generation, the evaluation effect of GPT4 as a referee model is generally remarkable. However, in the field of finance and taxation, more financial and taxation background knowledge is required, and the evaluation effect of GPT4 as a referee model has a large deviation. At the same time, debugging and optimizing the referee model is complex and costly. Summary of the invention

[0004] The present invention provides an automated evaluation method for large models in the financial and taxation fields, so as to solve the technical problems existing in the above-mentioned prior art.

[0005] To achieve the above object, the present invention provides an automated evaluation method for a large model in the field of finance and taxation, comprising:

[0006] S1: Prepare a test set in the field of finance and taxation. The test set includes multiple evaluation cases. Each evaluation case includes an input question and an expected result.

[0007] S2: for each evaluation case, construct evaluation keywords for evaluating whether the answer result of the large model to be evaluated is correct or incorrect. For the evaluation case whose expected result contains at least one evaluation keyword, execute step S3. For the evaluation case whose expected result does not contain any evaluation keyword, execute step S4.

[0008] S3: For each evaluation case, use at least one of the three combinations of AND, OR, and NOT to combine the evaluation keywords corresponding to the test case and compare each combination with the corresponding expected result, wherein if the expected result is consistent with the combination, the test case is retained and the corresponding evaluation keyword combination is recorded, otherwise the test case is deleted from the test set, wherein the AND in the combination contains one or more evaluation keywords in the expected result corresponding to the test case, the OR in the combination contains at least one evaluation keyword in the expected result corresponding to the test case, and the NOT in the combination does not contain one or more evaluation keywords in the expected result corresponding to the test case, and then execute step S5;

[0009] S4: converting the input question of the evaluation case into an objective multiple-choice question with at least two options, then combining the options in two combinations of AND and NOT and updating the corresponding evaluation case, and then executing step S5;

[0010] S5: Use the test set to test the large model to be evaluated, obtain the answer result of the large model to be evaluated for each evaluation case, and evaluate the answer result in a combination of evaluation keywords corresponding to the evaluation case to obtain a first evaluation result;

[0011] S6: Count the first evaluation results to obtain the evaluation accuracy of the large model to be evaluated.

[0012] In one embodiment of the present invention, step S6 further includes the following steps:

[0013] S7: Use the third-party evaluation big model to test the big model to be evaluated, obtain the answer result of the big model to be evaluated, and call the third-party referee model to evaluate the answer result to obtain a second evaluation result;

[0014] S8: Compare the first evaluation results with the second evaluation results, manually proofread any inconsistencies, and maintain and update the evaluation cases for the next round of evaluation.

[0015] In one embodiment of the present invention, in step S1, the test set involves at least one of the following areas: policies, regulations and basic knowledge, finance and taxation business handling, finance and taxation examinations, and risk control and planning, four major finance and taxation professional capability areas.

[0016] In one embodiment of the present invention, in step S1, the test set is stored in JSON or EXCEL file format.

[0017] In one embodiment of the present invention, in step S2, the evaluation keywords include "Commitment to Issue Invoices with Original Applicable Tax Rates" and red-ink invoices in the field of finance and taxation.

[0018] In one embodiment of the present invention, in step S2, the evaluation keywords further include "can" and "cannot".

[0019] In one embodiment of the present invention, in step S8, when the first evaluation result and the second evaluation result are inconsistent, maintaining and updating the evaluation case is achieved by adjusting the evaluation keyword combination corresponding to the evaluation case.

[0020] The automated evaluation method for the large model in the field of finance and taxation provided by the present invention can not only quickly and accurately perform automated evaluation on the large model of finance and taxation by constructing evaluation keywords for evaluation cases and converting subjective questions into objective questions, but also provide evaluation result data support for the rapid iteration and upgrade of the large model of finance and taxation. At the same time, the present invention is more convenient and intuitive for the maintenance and update of evaluation cases, and can quickly verify whether the evaluation keywords and combination methods are reasonable after fine-tuning, saving debugging time and improving work efficiency and evaluation accuracy. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0022] Figure 1 It is a flow chart of the automated evaluation method of the large model in the finance and taxation field according to the first embodiment of the present invention;

[0023] Figure 2 This is a flow chart of the automated evaluation method for a large model in the finance and taxation field according to the second embodiment of the present invention. DETAILED DESCRIPTION

[0024] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0025] The present invention provides an automated evaluation method for a large model in the field of finance and taxation, comprising:

[0026] S1: Prepare a test set in the field of finance and taxation. The test set includes multiple evaluation cases. Each evaluation case includes an input question and an expected result.

[0027] In one embodiment of the present invention, the test set involves at least one of the following areas: policies, regulations and basic knowledge, financial and tax business handling, financial and tax examinations, and risk control and planning. It should be noted that the test set can be dynamically adjusted and expanded based on business experience, testing experience, industry development and changes, and other factors, and the present invention is not limited to the above-listed areas.

[0028] In one embodiment of the present invention, in step S1, the test set is stored in JSON or EXCEL file format, which is a common file storage format in the testing field. In specific applications, the storage format of the test set can be selected according to actual needs. Generally speaking, it is appropriate to facilitate storage and calling.

[0029] It is well known to technical personnel in this field that test cases include input questions and expected results. When testing a model, the input questions are input into the model. If the output results of the model are consistent with the expected results, the test case passes the test. If the output results of the model are inconsistent with the expected results, the test case fails the test. At this time, it should be considered whether there are problems with the model itself, and attention should be paid to whether the writing of the test case is reasonable.

[0030] S2: for each evaluation case, construct evaluation keywords for evaluating whether the answer result of the large model to be evaluated is correct or incorrect. For the evaluation case whose expected result contains at least one evaluation keyword, execute step S3. For the evaluation case whose expected result does not contain any evaluation keyword, execute step S4.

[0031] In step S2, the evaluation keywords include, for example, "Commitment to Issue Invoices with Original Applicable Tax Rates" and red-ink invoices in the field of finance and taxation. The evaluation keywords may further include "can" and "cannot" and the like.

[0032] For example, the evaluation keywords can be jointly determined by the finance and tax business teachers and testers. The finance and tax business teachers tend to construct evaluation keywords from the business perspective, while the testers tend to construct evaluation keywords from the test perspective. In addition, according to the actual use of the model, the evaluation keywords can be continuously optimized to improve the accuracy of the model.

[0033] S3: For each evaluation case, use at least one of the three combinations of AND, OR, and NOT to combine the evaluation keywords corresponding to the test case and compare each combination with the corresponding expected result, wherein if the expected result is consistent with the combination, the test case is retained and the corresponding evaluation keyword combination is recorded, otherwise the test case is deleted from the test set, wherein the AND in the combination contains one or more evaluation keywords in the expected result corresponding to the test case, the OR in the combination contains at least one evaluation keyword in the expected result corresponding to the test case, and the NOT in the combination does not contain one or more evaluation keywords in the expected result corresponding to the test case, and then execute step S5;

[0034] This step S3 is for the evaluation case that contains at least one evaluation keyword in the expected result. This step S3 can also be determined by the financial and tax business teacher and the tester. Assuming that a test case has 4 evaluation keywords a, b, c, and d, after determination, the evaluation keywords are combined in two combinations of and and not. The final keyword combination is a&b&c and ~d, that is, the first evaluation keyword combination method constructed for the evaluation case is: a, b, c appear at the same time (corresponding to the expected result including a, b, c, a total of 3 evaluation keywords), the second evaluation keyword combination method does not include d (corresponding to the expected result does not include evaluation keyword d), when the expected result meets the above two evaluation keyword combination methods, then retain the test case and record the corresponding evaluation keyword combination method, otherwise delete the test case from the test set.

[0035] After step S3, the test cases and evaluation keywords in the test set

[0036] S4: converting the input question of the evaluation case into an objective multiple-choice question with at least two options, then combining the options in two combinations of AND and NOT and updating the corresponding evaluation case, and then executing step S5;

[0037] This step S4 is for evaluation cases that do not contain any evaluation keywords in the expected results. The input questions of this type of evaluation cases are open-ended and have a large standard range. The range of the results is limited by converting the input questions into objective multiple-choice questions to ensure the accuracy of the evaluation results. Among them, the AND combination method is to combine at least two options "AND" (the corresponding expected results at least meet these two options), and the non-combination method is to combine at least one option "NO" (the corresponding expected results at least do not meet this option).

[0038] This step S converts the input questions of the evaluation case into objective multiple-choice questions with at least two options. First, the large model is called, and the conversion rules are explained in the prompt, such as "the multiple-choice questions need to accurately and comprehensively cover the core points of the original subjective questions while maintaining the clarity and conciseness of the questions", and then manual proofreading is performed to form the final objective multiple-choice questions.

[0039] S5: Use the test set to test the large model to be evaluated, obtain the answer result of the large model to be evaluated for each evaluation case, and evaluate the answer result in a combination of evaluation keywords corresponding to the evaluation case to obtain a first evaluation result;

[0040] S6: Count the first evaluation results to obtain the evaluation accuracy of the large model to be evaluated.

[0041] In one embodiment of the present invention, step S6 further includes the following steps:

[0042] S7: Use the third-party evaluation big model to test the big model to be evaluated, obtain the answer result of the big model to be evaluated, and call the third-party referee model to evaluate the answer result to obtain a second evaluation result;

[0043] S8: Compare the first evaluation results with the second evaluation results, manually proofread any inconsistencies, and maintain and update the evaluation cases for the next round of evaluation.

[0044] The inventor of this case used 100 questions of financial and tax business test data to automatically verify the large model to be evaluated and the large model evaluated by the third party. The accuracy of manual judgment was 61% and 52% respectively, and the accuracy of automated judgment was 62% and 54% respectively. The accuracy of automated judgment was roughly calculated to be 98% and 96%, and the accuracy of automated judgment was accurately calculated to be 94% and 92%. Among them:

[0045] Rough calculation: This calculation method may result in data canceling each other out, that is, correct judgments become wrong, and wrong judgments become correct, and the final accuracy rate is not affected.

[0046] Accurate calculation: This calculation method uses manually proofread data statistics to eliminate situations where mutual cancellation occurs.

[0047] In steps S5 and S7, JAVA or PYTHON engineering can be used to read the test cases in the test set, and the large model to be evaluated can be called according to the model interface parameters.

[0048] Figure 1 , Figure 2 They are respectively flow charts of the automated evaluation method of the large model in the financial and taxation field according to the first embodiment and the second embodiment of the present invention, Figure 1 The method comprises steps S1 to S6, obtaining a first evaluation result, Figure 2The method comprises steps S1 to S4 and steps S7 to S8, obtaining a second evaluation result and updating and maintaining the inconsistencies between the first evaluation result and the second evaluation result.

[0049] The following is a specific example to illustrate the process of evaluating the large model to be evaluated in the present invention:

[0050] Case 1: Combine evaluation keywords using the “and” method for automated evaluation

[0051] Evaluation Questions:

[0052] After the standard for the special additional deduction for supporting the elderly has been raised, my brothers and sisters have not enjoyed this benefit. Can I deduct 3,000 yuan every month?

[0053] Evaluation keyword combination method:

[0054] Evaluation keywords: No, 1500;

[0055] Model to be evaluated answer 1:

[0056] No. After the standard of special additional deduction for supporting the elderly is raised, the only child can deduct 3,000 yuan per month, while the non-only child needs to share the 3,000 yuan deduction with his or her brothers and sisters, and the deduction amount per person per month does not exceed 1,500 yuan.

[0057] Evaluation result 1: Correct

[0058] "Answer 1 of the model to be evaluated" includes both the evaluation keyword "no" and the evaluation keyword "1500", which meets the "and" relationship between the evaluation keywords. Therefore, the evaluation result 1 is "correct".

[0059] Model to be evaluated answer 2:

[0060] According to Article 22 and Article 23 of the Interim Measures for Special Additional Deductions for Individual Income Tax, taxpayers who support one or more dependents can deduct the support expenses at a standard rate of RMB 2,000 per month. If your brothers and sisters do not enjoy this benefit, you cannot deduct RMB 3,000 per month. You can only deduct RMB 2,000 per month.

[0061] Evaluation result 2: Error

[0062] The answer 2 of the model to be evaluated does not include the evaluation keywords "no" and "1500", so the evaluation result 2 is "error".

[0063] Case 2: Use “or” combination to combine evaluation keywords for automated evaluation

[0064] Evaluation question: How to void the invoices issued by a general taxpayer last month in the next month?

[0065] Combination method of evaluation keywords: OR

[0066] Evaluation keywords: Red-ink reversal, Red-letter invoice

[0067] Answer of the model to be evaluated 1:

[0068] For the invoices issued by a general taxpayer last month, if errors are found in the next month, they cannot be directly voided because the tax has been copied and reported. In this case, only the red-letter invoices can be issued for handling.

[0069] To apply for issuing red-letter invoices, the purchaser needs to first issue an information form for red-letter VAT special invoices in the invoicing system and upload it to the IRS server. After the information form number appears, the seller can issue red-letter invoices according to this information form number.

[0070] Evaluation result 1: Correct

[0071] Since the combination method of evaluation keywords is "OR", only one of the evaluation keywords needs to be included in the answer of this model to be judged as correct. The answer of "Answer of the model to be evaluated 1" includes one of the evaluation keywords "Red-letter invoice", so the evaluation result 1 is "Correct".

[0072] Answer of the model to be evaluated 2:

[0073] According to Article 36 of the Measures for the Administration of Tax Certificates, for the invoices issued by a general taxpayer last month, if they need to be voided due to incorrect issuance, the words "void", the reasons for voidance, and the serial number and number of the tax certificate to be reissued shall be indicated on each copy. At the same time, according to Article 37, all kinds of seals on each copy of the paper tax certificate shall be affixed completely, and the seals shall not be overprinted, except as otherwise provided.

[0074] Evaluation result 2: Incorrect

[0075] Neither of the two evaluation keywords "Red-ink reversal" and "Red-letter invoice" is included in the "Answer of the model to be evaluated 2", so the evaluation result is incorrect.

[0076] Case 3: After converting the input question of the evaluation case into an objective multiple-choice question with at least two options, then conduct automated evaluation

[0077] Evaluation question: How should a general taxpayer handle the situation when the income is confirmed first and then the invoice is issued?

[0078] Converted options:

[0079] A. Record the accounts directly according to the invoice amount

[0080] B. First make a red-letter voucher to offset the previously confirmed uninvoiced income, and then make a formal voucher to record the account according to the invoice

[0081] C. In the declaration form, fill in the amount of the supplementary invoice with a positive number.

[0082] D. No processing is required

[0083] Option combination: and: B, C; not: A, D

[0084] Answer 1 of the model to be evaluated: The correct answer is: BC

[0085] Evaluation result 1: Correct

[0086] Answer 2 of the model to be evaluated: The correct answer is: AB

[0087] Evaluation result 2: Error

[0088] From the option combination method, it can be seen that the model answer needs to include options B and C at the same time, and the model answer must not include options A and D, in order to determine that the model answer is correct. The model answer 1 to be evaluated includes options B and C, so it is correct, and other answers are all wrong (such as the model answer 2 to be evaluated).

[0089] In one embodiment of the present invention, in step S8, when the first evaluation result and the second evaluation result are inconsistent, maintaining and updating the evaluation case is achieved by adjusting the evaluation keyword combination corresponding to the evaluation case.

[0090] The automated evaluation method for the large model in the field of finance and taxation provided by the present invention can not only quickly and accurately perform automated evaluation on the large model of finance and taxation by constructing evaluation keywords for evaluation cases and converting subjective questions into objective questions, but also provide evaluation result data support for the rapid iteration and upgrade of the large model of finance and taxation. At the same time, the present invention is more convenient and intuitive for the maintenance and update of evaluation cases, and can quickly verify whether the evaluation keywords and combination methods are reasonable after fine-tuning, saving debugging time and improving work efficiency and evaluation accuracy.

[0091] Those skilled in the art can understand that the accompanying drawings are only schematic diagrams of an embodiment, and the modules or processes in the accompanying drawings are not necessarily required to implement the present invention.

[0092] Those skilled in the art can understand that the modules in the device in the embodiment can be distributed in the device in the embodiment according to the description of the embodiment, or can be changed accordingly and located in one or more devices different from the embodiment. The modules in the above embodiment can be combined into one module, or can be further divided into multiple sub-modules.

[0093] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. An automated evaluation method for a large model in the field of finance and taxation, characterized in that: include: S1: Prepare a test set in the field of finance and taxation. The test set includes multiple evaluation cases. Each evaluation case includes an input question and an expected result. S2: for each evaluation case, construct evaluation keywords for evaluating whether the answer result of the large model to be evaluated is correct or incorrect. For the evaluation case whose expected result contains at least one evaluation keyword, execute step S3. For the evaluation case whose expected result does not contain any evaluation keyword, execute step S4. S3: For each evaluation case, use at least one of the three combinations of AND, OR, and NOT to combine the evaluation keywords corresponding to the test case and compare each combination with the corresponding expected result, wherein if the expected result is consistent with the combination, the test case is retained and the corresponding evaluation keyword combination is recorded, otherwise the test case is deleted from the test set, wherein the AND in the combination contains one or more evaluation keywords in the expected result corresponding to the test case, the OR in the combination contains at least one evaluation keyword in the expected result corresponding to the test case, and the NOT in the combination does not contain one or more evaluation keywords in the expected result corresponding to the test case, and then execute step S5; S4: converting the input question of the evaluation case into an objective multiple-choice question with at least two options, then combining the options in two combinations of AND and NOT and updating the corresponding evaluation case, and then executing step S5; S5: Use the test set to test the large model to be evaluated, obtain the answer result of the large model to be evaluated for each evaluation case, and evaluate the answer result in a combination of evaluation keywords corresponding to the evaluation case to obtain a first evaluation result; S6: Count the first evaluation results to obtain the evaluation accuracy of the large model to be evaluated.

2. The automated evaluation method for a large model in the field of finance and taxation according to claim 1 is characterized in that: Step S6 further includes the following steps: S7: Use the third-party evaluation big model to test the big model to be evaluated, obtain the answer result of the big model to be evaluated, and call the third-party referee model to evaluate the answer result to obtain a second evaluation result; S8: Compare the first evaluation results with the second evaluation results, manually proofread any inconsistencies, and maintain and update the evaluation cases for the next round of evaluation.

3. The automated evaluation method for a large model in the field of finance and taxation according to claim 1 is characterized in that: In step S1, the test set involves at least one of the following areas: policies, regulations and basic knowledge, financial and tax business handling, financial and tax examinations, and risk control and planning, four major financial and tax professional capability areas.

4. The automated evaluation method for a large model in the field of finance and taxation according to claim 1 is characterized in that: In step S1, the test set is stored in JSON or EXCEL file format.

5. The automated evaluation method for a large model in the field of finance and taxation according to claim 1 is characterized in that: In step S2, the evaluation keywords include "Commitment to Issue Invoices with the Original Applicable Tax Rate" and red-ink invoices in the finance and taxation field.

6. The automated evaluation method for a large model in the field of finance and taxation according to claim 5 is characterized in that: In step S2, the evaluation keywords further include "can" and "cannot".

7. The automated evaluation method for a large model in the field of finance and taxation according to any one of claims 1 to 6, characterized in that: In step S8, when the first evaluation result and the second evaluation result are inconsistent, maintaining and updating the evaluation case is achieved by adjusting the evaluation keyword combination corresponding to the evaluation case.