Large language model evaluation method, evaluation device, electronic equipment, computer readable storage medium and program product
By combining structured cross-modal data and multiple types of evaluation tasks, and employing a multi-prompt strategy for evaluating large language models, the problem of unstable evaluation results in the financial field is solved, enabling accurate determination of the overall performance of the model and improving the comprehensiveness and accuracy of the evaluation.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SHANGHAI UNIVERSITY OF FINANCE AND ECONOMICS
- Filing Date
- 2026-03-23
- Publication Date
- 2026-04-17
AI Technical Summary
Existing large language model evaluation methods are difficult to comprehensively, objectively and accurately reflect the overall performance of the model in the financial field. In particular, the evaluation results are unstable in complex financial tasks, and traditional methods rely on a single prompting strategy, which leads to bias in the evaluation results.
By combining structured cross-modal data, multi-type evaluation tasks, and multi-cue strategies, this study acquires structured text and image data from the financial field, designs various pre-set evaluation tasks and cue strategies, and conducts multi-dimensional quantitative analysis to evaluate the comprehensive performance of a large language model.
It enables a comprehensive, objective, and accurate evaluation of large language models in the financial field, improving the accuracy and stability of the evaluation. It can fully reflect the model's capabilities in financial knowledge application, logical analysis, and multimodal processing, and provides clear directions for improvement.
Smart Images

Figure CN121880149A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of large language model technology, and more specifically, to a large language model evaluation method, evaluation device, electronic device, computer-readable storage medium, and program product. Background Technology
[0002] With the rapid development of artificial intelligence technology, Large Language Models (LLMs) have demonstrated broad application potential in various fields such as finance, healthcare, education, and government due to their powerful semantic understanding, logical reasoning, and generative capabilities. Particularly in the financial sector, LLMs can be applied to various business scenarios, including investment research and analysis, risk assessment, intelligent customer service, and compliance review. Therefore, scientifically and systematically evaluating their performance has become a crucial step in ensuring the reliability and practicality of these models. However, given the vast and highly specialized and dynamic nature of financial knowledge, accurately and comprehensively evaluating the overall capabilities of LLMs in the financial field remains a pressing technological challenge.
[0003] Existing large language model evaluation methods mostly construct evaluation sets based on general language tasks or small amounts of single financial text data. These sets typically consist of plain text with low structure, making it difficult to accurately reflect the model's comprehensive processing capabilities in complex financial tasks. Furthermore, in the model testing phase, these techniques often use a single prompt template as input to the model, with output quality evaluated manually. This is not only inefficient but also highly susceptible to the influence of prompt design and subjective judgment, leading to poor stability and repeatability of evaluation results. Especially in financial scenarios, different task types have significantly different reliance on prompts, and fixed prompt strategies cannot cover the model's true performance under multi-dimensional tasks. In other words, existing techniques generally fail to provide a comprehensive, objective, and accurate evaluation system for large language models in the financial field that reflects the model's overall performance. Summary of the Invention
[0004] The summary section of this application introduces a series of simplified concepts, which will be further explained in detail in the detailed description section. The summary section of this application is not intended to limit the key features and essential technical features of the claimed technical solution, nor is it intended to determine the scope of protection of the claimed technical solution.
[0005] The large language model evaluation method and related equipment provided in this application can form a complete and standardized evaluation process for large language models in the financial field by combining structured cross-modal data, multi-type evaluation tasks, multi-prompt strategies and multi-dimensional quantitative analysis. It can comprehensively, objectively and accurately determine the overall performance of the model and effectively improve the accuracy and practical value of large language model evaluation in financial scenarios.
[0006] In a first aspect, this application provides a method for evaluating a large language model, comprising: acquiring benchmark data in the financial field, wherein the benchmark data includes structured text data and structured image data cross-modally associated with the structured text data; generating an evaluation task set for a target large language model based on the benchmark data and preset evaluation tasks, wherein the preset evaluation tasks include at least one of quantitative calculation, logical reasoning, text interpretation, and multimodal comprehensive analysis; inputting the evaluation task set into the target large language model through various preset prompting strategies to obtain the output results of the target large language model; and performing multidimensional analysis on the output results based on the evaluation dimensions and quantitative indicators associated with the evaluation task set to determine the comprehensive performance of the target large language model.
[0007] In some implementations, acquiring assessment benchmark data in the financial field includes: acquiring financial academic knowledge data within a first preset scope, and processing the financial academic knowledge data according to a first preset annotation rule to generate structured multiple-choice question data containing knowledge point affiliation and difficulty level; acquiring financial industry practical data within a second preset scope, and processing the financial industry practical data according to a second preset annotation rule to generate objective short-answer question data containing core entity annotations and subjective open-ended question data containing business constraint annotations; acquiring financial security and compliance data within a third preset scope, and processing the financial security and compliance data according to a third preset annotation rule to generate security test question data containing a risk option and correct solution comparison structure; acquiring financial intelligent agent task data within a fourth preset scope, and processing the financial intelligent agent task data according to a fourth preset annotation rule to generate intelligent agent process test question data containing task process nodes, tool call specifications, and multi-round interaction logic; and integrating the structured multiple-choice question data, the objective short-answer question data, the subjective open-ended question data, the security test question data, and the intelligent agent process test question data to obtain the assessment benchmark data.
[0008] In some implementations, generating an evaluation task set for a target large language model based on the evaluation benchmark data and preset evaluation tasks includes: selecting the evaluation task set from the evaluation benchmark data based on the type of the preset evaluation tasks, wherein the preset prompting strategy includes at least one of a zero-sample prompting strategy, a few-sample prompting strategy, and a thought chain prompting strategy.
[0009] In some implementations, the step of inputting the evaluation task set into a target large language model using multiple preset prompting strategies to obtain the output of the target large language model includes: generating a first prompt template without task examples based on a zero-shot prompting strategy, constructing a first model input sequence from the questions in the evaluation task set based on the first prompt template, and obtaining a first output result generated by the target large language model based on the first model input sequence; generating a second prompt template containing multiple task examples and corresponding answers based on a few-shot prompting strategy, constructing a second model input sequence from the questions in the evaluation task set based on the second prompt template, and obtaining a first output result generated by the target large language model based on the first model input sequence. The second output result generated by the second model input sequence; based on the thought chain prompting strategy, a third prompt template containing step-by-step reasoning requirements is generated, and the questions in the evaluation task set are constructed into a third model input sequence based on the third prompt template, and a third output result containing the reasoning process and final answer generated by the target large language model based on the third model input sequence is obtained; based on the combined prompting strategy of few samples and thought chain, a fourth prompt template containing both task examples and reasoning requirements is generated, and the questions in the evaluation task set are constructed into a fourth model input sequence based on the fourth prompt template, and a fourth output result generated by the target large language model based on the fourth model input sequence is obtained.
[0010] In some implementations, the step of performing multidimensional analysis on the output results based on evaluation dimensions and quantitative indicators associated with the evaluation task set to determine the comprehensive performance of the target large language model includes: for quantitative calculation tasks in the evaluation task set, scoring the output results based on a preset accuracy indicator to obtain a first score; for logical reasoning tasks in the evaluation task set, weighting the output results based on a preset logical coherence dimension and reasoning step completeness dimension to obtain a second score; for text interpretation tasks in the evaluation task set, weighting the output results based on at least several dimensions of semantic relevance, business compliance, and risk warning completeness to obtain a third score; for multimodal comprehensive analysis tasks in the evaluation task set, weighting the output results based on at least one dimension of visual feature recognition accuracy, cross-modal information integration capability, and decision logic rationality to obtain a fourth score; and determining the comprehensive performance score of the target large language model based on the first score, the second score, the third score, and the fourth score through a preset comprehensive scoring strategy to determine the comprehensive performance of the target large language model.
[0011] In some implementations, determining the comprehensive performance score of the target large language model based on the first score, the second score, the third score, and the fourth score using a preset comprehensive scoring strategy includes: weighting the first score, the second score, the third score, and the fourth score based on preset task type weights to obtain an initial comprehensive score, wherein the preset task type weights are adjusted based on the expected application domain of the target large language model; and correcting the initial comprehensive score based on the performance stability of the target large language model under the various preset prompting strategies to obtain the comprehensive performance score.
[0012] Secondly, this application also provides a large language model evaluation device, comprising: a data acquisition unit for acquiring evaluation benchmark data in the financial field, wherein the evaluation benchmark data includes structured text data and structured image data cross-modal associated with the structured text data; a task generation unit for generating an evaluation task set for a target large language model based on the evaluation benchmark data and preset evaluation tasks, wherein the types of the preset evaluation tasks include at least one of quantitative calculation, logical reasoning, text interpretation, and multimodal comprehensive analysis; a model testing unit for inputting the evaluation task set into the target large language model through various preset prompting strategies to obtain the output results of the target large language model; and a performance evaluation unit for performing multidimensional analysis on the output results based on evaluation dimensions and quantitative indicators associated with the evaluation task set to determine the comprehensive performance of the target large language model.
[0013] Thirdly, this application also provides an electronic device, including: a memory and a processor, wherein the processor is configured to implement the steps of the large language model evaluation method described in the first aspect when executing a computer program stored in the memory.
[0014] Fourthly, this application also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps of the large language model evaluation method described in the first aspect.
[0015] Fifthly, this application also provides a computer program product, including a computer program or computer executable instructions, wherein when the computer program or computer executable instructions are executed by a processor, the steps of the large language model evaluation method provided in the embodiments of this application are implemented.
[0016] In summary, this application, by introducing a financial domain evaluation benchmark that includes both structured text and structured image data, enables multimodal comprehensive evaluation of large language models. Compared to traditional methods that rely solely on text data, this application has a broader data source coverage, allowing for simultaneous examination of the model's performance in financial text understanding, chart recognition, and cross-modal reasoning, thereby enhancing the comprehensiveness and accuracy of the evaluation. Furthermore, based on different types of financial tasks, it pre-sets various evaluation tasks, including quantitative calculation, logical reasoning, text interpretation, and multimodal comprehensive analysis, making the evaluation content more targeted and comprehensively reflecting the model's performance in financial knowledge application, logical analysis, semantic understanding, and multimodal capabilities. The performance in key capabilities such as dynamic processing helps to accurately identify the model's strengths and weaknesses in specific capability dimensions. During the task input phase, multiple preset prompting strategies are employed to input the evaluation task set into the target large language model, avoiding bias caused by a single prompting method. The combined application of different prompting strategies allows the model's output performance under various conditions to be detected, thereby improving the accuracy and stability of the evaluation results. In the result analysis phase, based on the evaluation dimensions and quantitative indicators corresponding to the task set, a multi-dimensional analysis of the model output is performed. The evaluation results are not only more detailed and objective but also provide clear directions and basis for model improvement, enhancing the interpretability and operability of the evaluation system. In summary, the large language model evaluation method provided in this application, through the combination of structured cross-modal data, multiple types of evaluation tasks, multiple prompting strategies, and multi-dimensional quantitative analysis, forms a complete and standardized evaluation process for large language models in the financial field. This process can comprehensively, objectively, and accurately determine the overall performance of the model, effectively improving the accuracy and practical value of large language model evaluation in financial scenarios. Attached Figure Description
[0017] Various other advantages and benefits will become apparent to those skilled in the art upon reading the following detailed description of preferred embodiments. The accompanying drawings are for illustrative purposes only and are not intended to limit this specification. Furthermore, the same reference numerals denote the same parts throughout the drawings. In the drawings: Figure 1 A flowchart illustrating a large language model evaluation method provided in this application embodiment; Figure 2 A schematic diagram of the composition structure of a large language model evaluation device provided in this application embodiment; Figure 3 This is a schematic diagram of the composition structure of an electronic device provided in an embodiment of this application. Detailed Implementation
[0018] The terms used in the specification, claims, and drawings of this application, such as "first," "second," "third," "fourth," etc. (if any), are used to distinguish similar objects and not to describe a specific order or sequence. Therefore, it is to be understood that these terms can be used interchangeably where appropriate, allowing the described embodiments to be used in different orders, unless specifically required by the illustrations or description. Furthermore, the terms "is" and "has," and any variations thereof, are intended to cover, non-exclusively, all possible constituent elements. For example, a process, method, system, product, or apparatus comprising several steps or units is not necessarily limited to the steps or units explicitly listed, but may also include other steps or units not explicitly listed, or steps or units inherent to the process, method, product, or apparatus.
[0019] In this application, a "module" or "unit" refers to a computer program or part of a computer program that has a specific function and works in conjunction with other related parts to achieve a predetermined goal. These modules or units can be implemented by software, hardware (e.g., processing circuitry or memory), or a combination of both. One or more processors or memories can implement one or more modules or units. Furthermore, each module or unit can also be part of a larger module or unit.
[0020] The technical solutions of this application will be described in detail below with reference to the accompanying drawings of the embodiments. It should be noted that the described embodiments are only a part of this application, and not all embodiments. In the following description, the "some embodiments" mentioned are only a subset of all possible embodiments, which may be the same or different subsets, and different embodiments can be combined with each other without conflict.
[0021] Figure 1 This is a schematic flowchart illustrating a large language model evaluation method provided in an embodiment of this application. For example, see [link to example]. Figure 1 The large language model evaluation method provided in this application embodiment may include the following steps 101 to 104: Step 101: Obtain assessment benchmark data in the financial field, wherein the assessment benchmark data may include structured text data and structured image data that is cross-modal associated with the structured text data; In some examples, the benchmark data is a standardized dataset that has been systematically collected, processed, and labeled, specifically designed to measure various performance indicators of large language models in the financial field. Its core characteristics are coverage of key dimensions of financial scenarios and evaluability. The benchmark data comprises two core data formats. The first is structured text data, where "structured" means the data is organized according to preset task types, labeling dimensions, and format standards, enabling machines to directly recognize, parse, and verify it. The text data is information in text form, covering questions related to financial academic knowledge, business descriptions in financial industry practice, case analyses in financial security and compliance, and process descriptions in financial intelligent agent tasks. This text data is structured through fine-grained labeling (such as knowledge point attribution, core entities, and business constraints), ensuring that the verification of model outputs during the evaluation process is automated and quantifiable. The second type is cross-model data with structured text data. Structured image data with cross-modal associations refers to the direct correspondence between image data and text data in terms of content, where both serve the same financial assessment scenario. For example, text describing stock price trends is associated with the corresponding stock candlestick chart, and text analyzing a company's financial situation is associated with the corresponding financial statement visualization chart. This association ensures that the model's ability to process multimodal information can be evaluated. The "structured" aspect of structured image data involves labeling the image data with core visual features, metadata, and associated decision-making task information. The image data itself is financially relevant information presented in a visual form, including stock candlestick charts, fund flow heatmaps, and financial statement visualization charts. By structurally labeling the key visual features of these images (such as the opening and closing prices and moving average positions of candlestick charts, and the coordinate axis labels and data point values of financial charts), the model's ability to understand and transform image information can be accurately measured.
[0022] By implementing step 101, a financial evaluation benchmark is obtained, which includes structured text data and structured image data associated with it across modalities. This ensures that the data sources on which the evaluation relies are of high quality and highly relevant. Text data provides the testing basis for the model's language understanding and knowledge application, while image data (such as candlestick charts, financial statements, etc.) supplements the visual information dimension. This enables the evaluation to cover the model's capabilities in multiple aspects, such as text understanding, numerical reasoning, image recognition, and cross-modal information integration. It can effectively overcome the limitations of traditional evaluations that rely solely on single text data and improve the comprehensiveness and authenticity of the model evaluation.
[0023] Step 102: Based on the aforementioned evaluation benchmark data and preset evaluation tasks, generate an evaluation task set for the target large language model. The types of preset evaluation tasks may include at least one of the following: quantitative calculation, logical reasoning, text interpretation, and multimodal comprehensive analysis. In some examples, the pre-defined assessment tasks are pre-planned tasks designed to test specific capabilities of the model, based on actual application needs in the financial field (such as intelligent investment research, risk control, and customer service). These tasks are designed to fully cover the diverse needs of the financial sector for model capabilities, and may include at least one of the following: quantitative calculation, logical reasoning, text interpretation, and multimodal comprehensive analysis. Quantitative calculation tasks focus on numerical computation and mathematical formula reasoning in the financial field. Examples include calculating gross profit margins based on corporate profit and loss statement data, and solving for strike prices using option pricing models (such as the Black-Scholes model). These tasks rely on structured text data containing financial mathematical formulas in LaTeX format (such as financial academic knowledge questions) in the benchmark data. They primarily assess the target model's ability to accurately process financial mathematical calculations, a capability directly related to the accuracy of numerical decisions in financial transactions, thus avoiding investment losses or compliance risks due to calculation errors. Logical reasoning tasks focus on testing a model's ability to make complex logical judgments and perform multi-step deductions in financial scenarios. For example, they may deduce the stock market sector rotation trend based on macroeconomic indicators (such as GDP growth rate and inflation rate), or analyze whether there are vulnerabilities in a payment institution's data transmission process in conjunction with financial security and compliance provisions. These tasks usually link financial security and compliance data in the benchmark data with industry practice data (such as subjective open-ended questions with risk analysis). The evaluation focuses on the coherence of the model's reasoning process, the sufficiency of the arguments, and the rationality of the conclusions. This is a key supporting capability for complex decision-making in the financial field (such as risk assessment of mergers and acquisitions and credit approval decisions). Text interpretation tasks focus on the model's ability to understand professional textual information in the financial field and extract key information. For example, interpreting research reports on listed companies issued by securities firms to extract core investment logic, and analyzing policy announcements issued by financial regulatory agencies (such as the China Securities Regulatory Commission) to clarify business compliance boundaries. These tasks require the use of structured text data (such as financial industry practical questions) containing core entity annotations (such as event type, trigger words, and participating entities) and business constraint annotations (such as customer asset size and service period) in the benchmark data. The main evaluation criteria are the model's accurate grasp of text semantics and its adaptability to financial business scenarios. This capability is the foundation for financial intelligent customer service to respond to customer inquiries and for intelligent investment research systems to generate analysis reports.Multimodal integrated analysis tasks target scenarios in the financial field where multiple information carriers, such as text and images, coexist. These tasks require models to simultaneously process structured text data and associated structured image data. For example, combining daily candlestick charts of stocks (structured image data) with corresponding market interpretation texts can help determine short-term stock trends and provide recommendations for portfolio adjustments. Alternatively, based on visualized bar charts of a company's annual financial statements (structured image data) and accompanying text descriptions, these tasks can help analyze the core reasons for fluctuations in a company's net profit. These tasks rely on cross-modal text and image data in the benchmark data to evaluate the model's ability to integrate information from different modalities, the accuracy of visual feature recognition (such as opening and closing prices in candlestick charts and data point values in financial charts), and its comprehensive decision-making ability based on multi-source information. This capability is a core requirement for complex scenarios such as financial technical analysis and the generation of comprehensive investment research reports.
[0024] By implementing step 102, multiple types of assessment task sets are generated based on the assessment benchmark data and preset tasks, which can specifically examine the core capabilities of the large language model in financial scenarios. Quantitative calculation tasks are used to test the model's numerical reasoning and computational accuracy, logical reasoning tasks are used to verify its analytical and judgment capabilities, text interpretation tasks assess the model's semantic understanding level, and multimodal comprehensive analysis tasks test the model's ability to integrate textual and graphical information. The design of multiple task types makes the assessment content more targeted and hierarchical, and can comprehensively reflect the model's performance in different dimensions.
[0025] Step 103: Using various preset prompting strategies, the evaluation task set is input into the target large language model to obtain the output results of the target large language model; In some examples, the preset prompting strategy is a framework of instructions designed in advance to guide the model to understand the task objectives, standardize the output format, or demonstrate the reasoning process, based on the diverse usage needs of large language models in the financial field (such as instant consultation, case adaptation, complex decision-making, etc.). Its core value lies in simulating different ways in which the model receives information in actual business, avoiding misjudgment of model capabilities due to a single input mode. For example, there may be no reference materials in the financial customer service scenario, there may be a few case references in the investment research scenario, and detailed reasoning is required in the risk control scenario. Different strategies correspond to these real scenarios.
[0026] Specifically, various preset prompting strategies can cover core types such as zero-shot prompting strategies, few-shot prompting strategies, and chain-of-thought prompting (CoT) strategies, and can be combined according to task complexity. Zero-shot prompting strategies refer to not providing any examples or reference information related to the evaluation task to the target large language model, but only directly inputting question instructions from the evaluation task set (such as "Please explain the composition of core tier 1 capital in the 'Commercial Bank Capital Management Measures'"). This guides the model to autonomously generate output based entirely on the financial knowledge learned during the pre-training phase. This strategy mainly simulates the need for "instant response without prior reference materials" in financial scenarios, such as customers suddenly inquiring about compliance issues or business personnel quickly querying basic financial knowledge. Its output can reflect the depth of the model's autonomous mastery of basic financial knowledge. The few-sample hint strategy refers to providing the model with a small number (usually 3-5) of "question-answer" examples consistent with the current task type before inputting the question instruction for the evaluation task (e.g., when evaluating the "corporate accounts receivable risk analysis" task, provide 3 risk analysis cases and conclusions of similar companies first). This guides the model to generate the output of the current question by learning the problem-solving logic, output structure, or financial business rules of the examples. This strategy corresponds to the need in financial scenarios to "quickly adapt to new business areas by relying on a small number of cases", such as when the model first handles green finance project evaluation, cross-border payment compliance review, etc. Its output can measure the model's scenario learning efficiency and adaptability. The thought chain prompt strategy simultaneously adds instructions to the model to "step by step break down the reasoning process" when inputting the evaluation task (such as "when analyzing the reasons for the decline in the stock price of a listed company, please explain the impact of macroeconomic factors, industry policies, company operating data, etc."). It requires the model not only to output the final conclusion, but also to show the key logical nodes and evidence in the derivation process. This strategy mainly corresponds to the need for "transparent reasoning in complex decisions" in the financial field, such as risk assessment of mergers and acquisitions, multi-factor portfolio allocation, and tracing the source of financial security vulnerabilities. The reasoning process in its output can reflect the rigor of the model's decision-making logic and its fit with the financial business logic.
[0027] In actual implementation, different combinations of preset prompting strategies can be selected for different types of tasks in the assessment task set (such as quantitative calculation and multimodal comprehensive analysis). For example, for the multimodal comprehensive analysis task of "judging stock trends based on candlestick charts," a combination of small sample and thought chain strategies can be used. Two "candlestick chart-trend analysis" examples are provided first, and then the model is asked to explain step by step the visual feature recognition (such as the position of MA5 and MA10 moving averages), the basis for trend judgment, and operational suggestions. Through the application of various preset prompting strategies, the assessment task input to the target large language model is no longer an isolated problem, but an information interaction form that fits the actual financial business scenario. The model output results also cover various forms such as "direct answer," "reasoning steps," and "formatted output." These outputs not only contain information on the correctness of the final conclusion, but also contain the model's learning ability, reasoning logic, scenario adaptability, and other deep capabilities, providing rich and accurate raw data for the comprehensive analysis based on multidimensional assessment dimensions and quantitative indicators in subsequent steps.
[0028] By implementing step 103, multiple preset prompting strategies are used to input the evaluation task into the target model. The generalization and adaptability of the model can be tested under different prompting conditions. Different prompting methods can guide the model to complete the task through different reasoning paths, thereby revealing the differences in the model's performance under no examples, limited examples, or complex thinking guidance. This can effectively avoid the result bias caused by a single prompt and improve the objectivity of the evaluation process and the stability of the results.
[0029] Step 104: Based on the evaluation dimensions and quantitative indicators associated with the evaluation task set, perform multidimensional analysis on the output results to determine the comprehensive performance of the target large language model. In some examples, the evaluation dimensions associated with the evaluation task set are specific assessment directions set according to the core capability requirements of different task types within the evaluation task set (such as quantitative calculation, logical reasoning, text interpretation, and multimodal comprehensive analysis). These dimensions directly correspond to the actual capability requirements of models in the financial field. For example, the evaluation dimension for quantitative calculation tasks focuses on "calculation accuracy" because the accuracy of numerical calculations in financial scenarios (such as calculating net profit in financial statements and option pricing results) is directly related to the correctness of investment decisions and risk assessment. The evaluation dimension for logical reasoning tasks emphasizes "completeness of reasoning steps" and "logical coherence" because complex financial decisions (such as M&A risk analysis and credit approval) require a rigorous derivation process, and relying solely on conclusions is insufficient. The evaluation dimensions for text interpretation tasks include "semantic relevance," "business compliance," and "completeness of risk warnings." The former ensures that the model understands the core meaning of financial texts (such as research reports and regulatory announcements), while the latter two align with the principle of "compliance first" in the financial field, preventing the model from outputting non-compliant suggestions or omitting key risk points. The evaluation dimensions for multimodal comprehensive analysis tasks include "accuracy of visual feature recognition" and "cross-modal information integration capability." The former examines the model's accuracy in extracting key information from financial images (such as candlestick charts and financial statement charts), while the latter verifies whether the model can effectively combine image information with text information to form a unified decision-making logic (such as generating investment suggestions based on candlestick chart features and market text). Quantitative indicators transform the aforementioned evaluation dimensions into calculable and comparable numerical standards. Specific calculation rules eliminate subjective judgment errors, ensuring the objectivity of the evaluation results. For example, for the "calculation accuracy" dimension of quantitative calculation tasks, the quantitative indicator can be set as "accuracy rate," which is the percentage of correct calculation results output by the model out of the total number of tasks in this category. This indicator directly reflects the reliability of the model's financial numerical calculations. For the "completeness of reasoning steps" and "logical coherence" dimensions of logical reasoning tasks, the quantitative indicator can be designed as a "weighted score," that is, assigning a 40% weight to "completeness of steps" (such as whether all key reasoning nodes are covered) and a 60% weight to "logical coherence" (such as whether there are causal contradictions in the reasoning process). The overall score considers both the completeness of the process and the rigor of the logic. For the "business compliance" dimension of text interpretation tasks, the quantitative indicator can be "compliance compliance score," which sets compliance points based on financial regulatory rules (such as the Securities Law and the Data Security Law). The score is the proportion of compliance points covered by the model output to the total number of points, ensuring that the assessment meets financial compliance requirements. For the "cross-modal information integration capability" dimension of multimodal comprehensive analysis tasks, the quantitative indicator can be set as "information relevance score," which judges the degree of matching between image features (such as the position of moving averages in candlestick charts) and text conclusions (such as trend judgments) in the model output. The higher the matching degree, the higher the score, avoiding the problem of "disconnect between image and text conclusions" in the model.
[0030] In the specific multidimensional analysis process, the first step is to calculate the score for each task category based on its output results within the evaluation task set, comparing them with the corresponding evaluation dimensions and quantitative indicators. For example, in the financial security and compliance task (belonging to the logical reasoning category), if 8 out of 10 questions output by a certain model have complete and logically coherent reasoning steps, the score for this task category is calculated to be 80 points according to the weighted scoring rules. Subsequently, the scores for each type of task are integrated using a "preset task weight" based on the expected application areas of the target large language model. The preset task weight is determined according to the actual needs of financial institutions (e.g., when evaluating models used for risk control scenarios, the weight for financial security and compliance tasks is set at 40%; when evaluating models used for investment...). When developing models for specific scenarios, the weight of practical tasks in the financial industry is set at 35%. The proportion of each task in the comprehensive score is set to ensure that the comprehensive performance evaluation focuses on the core application scenarios of the model. Finally, the comprehensive performance score of the model is obtained through weighted calculation. For example, the weighted total score of a certain investment research model in the tasks of quantitative calculation (weight 25%, score 90), logical reasoning (weight 30%, score 85), text interpretation (weight 35%, score 88), and multimodal comprehensive analysis (weight 10%, score 82) is 90×25%+85×30%+88×35%+82×10%=86.8 points. This score is the quantitative manifestation of the model's comprehensive performance.
[0031] By implementing step 104, multidimensional analysis of the output results is performed based on the evaluation dimensions and quantitative indicators corresponding to the task set, enabling a systematic and refined evaluation of model performance. The combined effect of indicators from different dimensions can quantify the model's performance in various financial tasks, ensuring that the evaluation results are more objective and reliable, and providing a clear direction for model optimization, thereby achieving an accurate determination of the comprehensive performance of the target large language model.
[0032] In summary, this application's embodiments, by introducing a financial domain evaluation benchmark that includes structured text data and structured image data, enable multimodal comprehensive evaluation of large language models. Compared to traditional methods that rely solely on text data, this application has a broader data source coverage, allowing for simultaneous examination of the model's performance in financial text understanding, chart recognition, and cross-modal reasoning, thereby improving the comprehensiveness and accuracy of the evaluation. Furthermore, based on different types of financial tasks, it pre-sets various evaluation tasks, including quantitative calculation, logical reasoning, text interpretation, and multimodal comprehensive analysis, making the evaluation content more targeted and comprehensively reflecting the model's performance in financial knowledge application, logical analysis, semantic understanding, and... Performance in key capabilities such as multimodal processing helps to accurately identify the model's strengths and weaknesses in specific capability dimensions. During the task input phase, multiple preset prompting strategies are employed to input the evaluation task set into the target large language model, avoiding bias caused by a single prompting method. The combined application of different prompting strategies allows the model's output performance under various conditions to be detected, thereby improving the accuracy and stability of the evaluation results. In the result analysis phase, based on the evaluation dimensions and quantitative indicators corresponding to the task set, a multidimensional analysis of the model output is performed. The evaluation results are not only more detailed and objective but also provide clear directions and basis for model improvement, enhancing the interpretability and operability of the evaluation system. In summary, the large language model evaluation method provided in this application, through the combination of structured cross-modal data, multiple types of evaluation tasks, multiple prompting strategies, and multidimensional quantitative analysis, forms a complete and standardized evaluation process for large language models in the financial field. This process can comprehensively, objectively, and accurately determine the overall performance of the model, effectively improving the accuracy and practical value of large language model evaluation in financial scenarios.
[0033] In some embodiments, step 101 may include: acquiring financial academic knowledge data within a first preset scope, and processing the financial academic knowledge data according to a first preset annotation rule to generate structured multiple-choice question data containing knowledge point attribution and difficulty level; acquiring financial industry practical data within a second preset scope, and processing the financial industry practical data according to a second preset annotation rule to generate objective short-answer question data containing core entity annotations and subjective open-ended question data containing business constraint annotations; acquiring financial security compliance data within a third preset scope, and processing the financial security compliance data according to a third preset annotation rule to generate security test question data containing a risk option and correct solution comparison structure; acquiring financial intelligent agent task data within a fourth preset scope, and processing the financial intelligent agent task data according to a fourth preset annotation rule to generate intelligent agent process test question data containing task process nodes, tool call specifications, and multi-round interaction logic; and integrating the structured multiple-choice question data, objective short-answer question data, subjective open-ended question data, security test question data, and intelligent agent process test question data to obtain evaluation benchmark data.
[0034] In some examples, the first step is to acquire and structure financial academic knowledge data within a predefined scope. This predefined scope specifically refers to 34 related disciplines covering the core theoretical system of the financial field, including fundamental disciplines such as finance, accounting, and economics, as well as financial industry qualification certification subjects such as certified public accountant and securities practitioner qualifications. The data sources are mainly public examination questions and exercises from university finance textbooks, ensuring the authority and comprehensiveness of the financial academic knowledge data. The predefined annotation rules are exclusive annotation standards designed for the theoretical attributes of the financial academic knowledge data. In practice, each collected question needs to be labeled with "knowledge point affiliation" and "difficulty level," and the knowledge point affiliation must clearly indicate the corresponding question. The specific subfields under the discipline, such as labeling "calculating the impact of the reserve requirement ratio on the money supply" as "monetary economics - monetary policy tools", ensure that subsequent assessments can accurately locate the model's mastery of a certain theoretical module; the difficulty level is divided into 1-5 levels, with level 1 corresponding to basic concept questions (such as "explaining the definition of price-earnings ratio") and level 5 corresponding to complex formula reasoning questions (such as "calculating option prices based on the Black-Scholes model"). At the same time, all questions are required to be multiple choice and include financial mathematical formulas in LaTeX format to ensure that the data is suitable for quantitative calculation assessment tasks. Finally, this processing generates structured multiple choice data containing knowledge point affiliation and difficulty level, which is used to assess the model's financial theoretical foundation and mathematical formula reasoning ability.
[0035] Secondly, it involves acquiring and processing practical financial industry data within the second preset scope. This second preset scope focuses on actual business scenarios of financial institutions, with data sources including publicly released investment research reports from securities firms, internal business manuals from banks, securities and fund companies, and policy announcements issued by financial regulatory agencies (such as the China Securities Regulatory Commission). Practical financial industry data directly reflects the practical processes and compliance requirements of financial businesses. The second preset annotation rules categorize industry data into two types based on their business attributes: For objective short-answer question data (such as financial text classification and event extraction tasks), "core entities" need to be annotated. Core entities specifically include fine-grained information such as event type, trigger words, and participating entities. For example, when processing "a listed company announces the acquisition of a peer company,"... When extracting questions, label them with "Event Type: Company Merger and Acquisition; Triggering Word: Announcement of Acquisition; Participating Entities: Company A (Acquirer), Company B (Acquired Party)" to ensure the accuracy of the model output can be verified using automated tools (such as regular expressions). For subjective open-ended questions (such as customer profile building and investment advice generation tasks), label them with "Business Constraints." Business constraints specifically refer to key limitations in actual business operations, such as "Customer asset size of 500,000 yuan, low to medium risk tolerance, and investment period of 3 years" to ensure that the model output aligns with actual business needs. Finally, this process generates objective short-answer questions with core entity labels and subjective open-ended questions with business constraint labels to evaluate the model's practical financial business processing capabilities.
[0036] Furthermore, the process involves acquiring and processing financial security compliance data within a third pre-defined scope. This third pre-defined scope revolves around the core security and compliance needs of the financial sector. Data sources include financial security vulnerability databases (such as SecEval data adapted to financial scenarios), typical cases of bank user privacy protection, and financial system encryption protocol documents (such as the application specifications of SSL / TLS in payment systems). All data must be screened and rewritten by experts with more than 5 years of experience in financial security to ensure the professionalism and scenario authenticity of the financial security compliance data. The core of the third pre-defined labeling rules is to construct a comparison structure of "risk options - correct solutions." For example, for the question of "data storage of sensitive user information," the design would be "risk option: unencrypted storage of user bank card information; correct solution: storage of user bank card information using AES-256 encryption algorithm." Each question must also be labeled with a "security risk level" (high / medium / low). High risk levels correspond to scenarios that may lead to major compliance incidents (such as the leakage of a large amount of user privacy data), while low risk levels correspond to minor operational violations (such as an outdated encryption protocol version). This processing generates security test question data containing a comparison structure of risk options and correct solutions, filling the gap in the existing assessment of financial security compliance capabilities.
[0037] Finally, the fourth preset scope of financial agent task data is acquired and processed. This fourth preset scope targets the complex task processing capabilities of the financial big language model. Data sources include mainstream financial application programming interface (API) documents (such as the yfinance API for stock data acquisition and the Wind API for financial data terminals), actual financial business processes (such as portfolio construction and the organization of financial industry forums), and records of multiple rounds of financial consultation dialogues (such as continuous dialogues between clients and financial advisors regarding asset allocation). Financial agent task data is used to describe complex task scenarios in the financial field that require the big language model to complete through tool calls, multi-step process execution, or multi-round dialogue interactions. The fourth preset annotation rules must cover all elements of the agent task, specifically including annotating "task process nodes" (such as annotating the node order of "data acquisition - risk assessment - asset allocation - report generation" for "portfolio management tasks") and "tool call specifications" (such as calling yfinance). When using the API, the parameters must be formatted as "ticker='AAPL', period='1y'" and the "multi-turn interaction logic" must be specified (e.g., the dialogue record must be marked with "the customer previously consulted about bond risks - subsequent answers must be related to this risk point"). All reference answers to the questions must be jointly written by financial engineers and industry consultants, and after the initial draft is generated by a large language model (such as GPT-4o), it must be optimized through multiple rounds of expert review to ensure the practicality and accuracy of the answers. Finally, through this process, intelligent agent process test question data containing task flow nodes, tool call specifications and multi-turn interaction logic is generated to evaluate the model's tool use and complex task planning capabilities.
[0038] After completing the collection and structuring of the above four types of data, the generated structured multiple-choice question data, objective short-answer question data, subjective open-ended question data, security test question data, and intelligent agent process test question data need to be integrated. During the integration process, it is necessary to ensure that the different types of data have the same format (such as using standardized JSON format) and complementary dimensions (covering the entire fields of academia, industry, security, and intelligent agents). The final result is evaluation benchmark data with full-dimensional coverage, high degree of structure, and strong business relevance, which lays the data foundation for the subsequent comprehensive and objective evaluation of the target large language model.
[0039] Through the implementation of the above embodiments, financial assessment benchmark data is subdivided into four dimensions: financial academic knowledge, industry practice, security compliance, and intelligent agent tasks. Corresponding annotation rules are formulated for different data types, which can achieve full-scenario coverage from theoretical knowledge and business operations to security standards and intelligent interaction. This enables assessment tasks to not only have high professionalism and real-world relevance, but also to simulate the multi-task performance of models in real financial business environments. The assessment benchmark data constructed in this way can effectively improve the relevance and diversity of model testing content, thereby improving the representativeness and authenticity of assessment results.
[0040] In some embodiments, step 102 may include: selecting an evaluation task set from the evaluation benchmark data based on the type of the preset evaluation task, wherein the preset prompting strategy may include at least one of the zero-sample prompting strategy, the few-sample prompting strategy, and the mind chain prompting strategy.
[0041] In some examples, firstly, the operation of selecting the assessment task set from the assessment benchmark data based on the preset assessment task type requires establishing a corresponding mapping relationship between the preset assessment task type and each sub-data of the assessment benchmark data. Quantitative calculation tasks need to be selected from the structured multiple-choice question data of the assessment benchmark data. This type of data contains LaTeX format financial mathematical formulas (such as option pricing formulas and variance estimation formulas) and has clear difficulty levels, which can accurately match the assessment objective of "numerical calculation accuracy". When selecting, the difficulty level range can be set according to the assessment needs. For example, when assessing the basic calculation ability of the model, focus on level 1-3 questions, and when assessing the complex reasoning calculation ability, focus on level 4-5 questions. Logical reasoning tasks need to be selected from security test question data and agent process test question data. The "risk option-correct solution" comparison structure of security test questions and the "task process node" annotation of agent process test questions can effectively examine the model's logical performance. For tasks involving analysis and step decomposition, priority should be given to high-risk security test questions or agent tasks involving multiple tool calls during the screening process to highlight the assessment of complex reasoning. For text interpretation tasks, selection should be made from objective short-answer question data and subjective open-ended question data. Core entity annotations (such as event trigger words and participating entities) for objective short-answer questions can be used to test the accuracy of information extraction, while business constraint annotations (such as budget and customer scale) for subjective open-ended questions can be used to test the compliance and practicality of text generation. During the screening process, different text types such as financial research reports and regulatory announcements should be considered to ensure coverage of text interpretation needs in multiple scenarios. For multimodal comprehensive analysis tasks, it is necessary to screen and evaluate structured text data and structured image data with cross-modal correlation in the benchmark data, such as combined data such as "stock candlestick chart + market analysis text" and "financial statement chart + performance interpretation text". During the screening process, the closeness of the correlation between images and text should be verified to avoid evaluation errors due to low correlation.
[0042] During the screening process, screening weights can be introduced. These weights are determined by the proportion of different task types within the task set, based on the expected application domain of the target large language model. For example, when evaluating models for intelligent risk control scenarios, the screening weight for financial security and compliance tasks (belonging to the logical reasoning category) can be set to 40%, ensuring that the task set focuses on the model's core application capabilities. When evaluating models for comprehensive investment research scenarios, the combined screening weight for text interpretation and multimodal comprehensive analysis tasks can be set to 60%, aligning with the core needs of text analysis and chart interpretation in investment research. Through this screening method that combines type matching and weight adjustment, the generated evaluation task set can cover all capability dimensions in the financial field while highlighting key evaluation points and avoiding redundant irrelevant tasks.
[0043] Meanwhile, the preset prompting strategies can include at least one of zero-shot prompting strategies, few-shot prompting strategies, and thought chain prompting strategies. This is to establish the "task type-prompt strategy" adaptation relationship during the task set generation stage, preparing for the model input in the subsequent step 103. Specifically, the zero-shot prompting strategy adapts to task types that do not require prior example guidance, such as basic financial academic knowledge Q&A (belonging to the quantitative calculation category). For this type of task, the model only needs to answer directly based on its own knowledge reserves. The few-shot prompting strategy adapts to task types that require a small number of case references, such as text interpretation in new business scenarios in the financial industry (belonging to the text interpretation category). By providing 3-5 similar business case examples, it guides the model to quickly adapt to the task requirements. The thought chain prompting strategy adapts to task types that require the demonstration of the reasoning process, such as complex financial security vulnerability analysis (belonging to the logical reasoning category) or multimodal investment trend judgment (belonging to the multimodal comprehensive analysis category). By requiring the model to output the reasoning logic step by step, it ensures that the evaluation can deeply examine the model's decision-making basis. This design, which associates the prompting strategy during the task set generation stage, avoids the problem of mismatch between the strategy and the task in the subsequent input stage. It ensures that each task can stimulate the model performance through the most suitable guidance method, laying the foundation for obtaining comprehensive and realistic output results in the future.
[0044] Through the implementation of the above embodiments, in the process of generating the evaluation task set, intelligent filtering is performed from the evaluation benchmark data according to the preset task type. Combined with various prompting strategies such as zero-sample, small-sample, and thought chain, the optimal test mode can be automatically matched for different task characteristics. This not only improves the scientificity and adaptability of the evaluation tasks, but also tests the generalization ability and robustness of the model under different prompting conditions. Through flexible task generation and prompting strategy matching, the comprehensiveness and accuracy of the evaluation of large language models in the financial field can be improved.
[0045] In some embodiments, step 103 may include: generating a first prompt template without task examples based on a zero-shot prompt strategy, constructing a first model input sequence based on the first prompt template for questions in the evaluation task set, and obtaining a first output result generated by the target large language model based on the first model input sequence; generating a second prompt template containing multiple task examples and corresponding answers based on a few-shot prompt strategy, constructing a second model input sequence based on the second prompt template for questions in the evaluation task set, and obtaining a second output result generated by the target large language model based on the second model input sequence; generating a third prompt template containing step-by-step reasoning requirements based on a thought chain prompt strategy, constructing a third model input sequence based on the third prompt template for questions in the evaluation task set, and obtaining a third output result generated by the target large language model based on the third model input sequence, containing the reasoning process and the final answer; and generating a fourth prompt template containing both task examples and reasoning requirements based on a combined few-shot and thought chain prompt strategy, constructing a fourth model input sequence based on the fourth prompt template for questions in the evaluation task set, and obtaining a fourth output result generated by the target large language model based on the fourth model input sequence.
[0046] In some examples, the zero-sample prompt strategy is primarily designed to simulate the need for immediate response in financial scenarios without prior reference materials (such as a customer suddenly inquiring about financial compliance issues or business personnel quickly searching for basic financial knowledge). Therefore, a first prompt template without any task examples needs to be generated. The structure of the first prompt template should be concise and clear, containing only the task instruction and the question to be evaluated. For example, a template for quantitative calculation tasks could be designed as "Please calculate the result of the following financial question and directly output the final numerical answer: [Evaluate the specific question in the task set]", while a template for text interpretation tasks could be designed as "Please interpret the core meaning of the following financial text and extract the key information: [Evaluate the specific text in the task set]". This example-free template design avoids interference from example information on the model's autonomous judgment. When constructing the first model input sequence, the questions corresponding to zero-sample scenarios in the evaluation task set (such as basic financial academic knowledge questions and simple compliance Q&A) need to be substituted one by one into the first prompt template. At the same time, it is necessary to ensure that the input sequence meets the token length limit of the target large language model (such as reasonably segmenting long text questions when the limit is exceeded). Then, the first model input sequence is passed to the target large language model through the model's inference interface (such as API interface or locally deployed inference function) to obtain the first output result generated directly by the model based on its own pre-trained knowledge. This result can be used to evaluate the model's depth of autonomous mastery of basic financial knowledge and the accuracy of its immediate response.
[0047] Secondly, for the small sample prompting strategy, which corresponds to the need in financial scenarios to quickly adapt to new businesses based on a small number of cases (such as the model's first processing of green finance project assessment and cross-border payment compliance review), it is necessary to generate a second prompt template containing multiple task examples and corresponding answers. The second prompt template should include example prompts, multiple sets of examples, and questions to be evaluated. The number of examples is usually set to 3-5 sets (this number provides sufficient learning reference for the model while avoiding excessively long input sequences due to too many examples). Each set of examples must strictly follow the format "Question: [Example Question] Answer: [Example Answer]", and the examples must belong to the same task type as the questions to be evaluated (e.g., when evaluating investment advice generation tasks, the examples must all be investment advice questions and standard answers). For example, the template design for subjective questions in the financial industry is as follows: "The following are three examples of financial investment advice tasks. Please refer to the answer logic of the examples to answer the subsequent questions: Example 1 - Question: A client has a low to medium risk tolerance and assets of 500,000 yuan. Recommend a suitable investment portfolio; Answer: Recommend '60% bond fund + 30% money market fund + 10% stable equity fund', with a hint of 'market volatility risk'... Example 3 - [Omitted] Please answer: A client has a high risk tolerance and assets of 2 million yuan. Recommend a suitable investment portfolio." When constructing the second model input sequence, it is necessary to combine the questions in the evaluation task set that require small sample guidance (such as industry practice new business questions, complex text classification questions) with the corresponding second prompt template to ensure the logical coherence between the examples and the questions to be evaluated. Then, the input sequence is passed into the target large language model to obtain the second output result generated by the model after learning the examples. This result can be used to measure the model's learning efficiency and adaptability to the scenario of new financial business.
[0048] Furthermore, regarding the thought chain prompting strategy, it focuses on the need for transparent reasoning in complex financial decisions (such as M&A risk analysis, multi-factor portfolio allocation, and financial security vulnerability tracing). For these tasks, the rationality of the decision cannot be verified by the final answer alone. Therefore, a third prompt template containing step-by-step reasoning requirements needs to be generated. The core of the third prompt template is to explicitly add "reasoning process guidance" to the task instructions. The structure is designed as follows: "Please analyze the following financial issues step by step, first explaining the reasoning basis for each step (in conjunction with financial knowledge or business rules), and then giving the final answer: [Evaluate the specific issues in the task set]". For example, the template design for financial security compliance reasoning questions is: "Please analyze whether a bank's behavior of 'not encrypting and storing user ID information' is compliant, explain the reasoning process step by step (referencing relevant regulatory rules), and then give the conclusion: [Specific scenario description]". When constructing the input sequence for the third model, complex reasoning problems (such as security vulnerability analysis questions and multimodal trend judgment questions) from the evaluation task set should be selected and substituted into the third prompt template to ensure that the reasoning requirements are clear and conform to the financial business logic. After being input into the model, the third output result containing the reasoning steps and the final answer should be obtained. The reasoning steps should be able to reflect the model's application of financial rules (such as citing the provisions of the Data Security Law) and the decomposition of complex factors (such as decomposing policy, market, and financial factors when analyzing M&A risks). This result can be used to evaluate the rigor of the model's decision-making logic and its fit with the financial business logic, avoiding misjudgments where the answer is correct but the reasoning is wrong.
[0049] Finally, regarding the combined prompting strategy of small samples and thought chains, which corresponds to complex tasks in financial scenarios and requires reference to high-difficulty cases (such as analyzing the compliance risks of new M&A cases based on historical M&A cases, judging stock trends and deriving logic by combining multiple sets of K-line chart cases), these tasks require examples to guide the model to understand the task boundaries and the reasoning process to verify the depth of decision-making. Therefore, a fourth prompting template that includes both task examples and reasoning requirements needs to be generated. The structure of the fourth prompting template needs to integrate the core elements of small samples and thought chains, designed as follows: "The following are examples of 2-3 financial tasks. Each example includes a question, reasoning process, and answer. Please refer to the reasoning logic and answer structure of the example, analyze the subsequent questions step by step, and give the answer: Example 1 - Question: [Example Question 1] Reasoning process: [Step-by-step reasoning of Example 1, such as 'First step to judge the impact of industry policies: ... Second step to analyze company financial data: ...'] Answer: [Answer to Example 1] ... Example 3 - [Omitted] Please answer: [Evaluate the specific questions in the task set], requiring step-by-step reasoning first, and then giving the answer." The reasoning process in the examples needs to be detailed and consistent with the financial business logic, providing the model with a clear reasoning paradigm. When constructing the fourth model input sequence, it is necessary to select high-complexity comprehensive problems (such as agent tool call process planning problems and cross-modal comprehensive analysis problems) from the evaluation task set and combine them with the fourth prompt template. After inputting them into the model, the fourth output result is obtained. This result can be used to evaluate the comprehensive ability of the model in scenarios with case references and requiring deep reasoning. In particular, it can reflect the model's learning and transfer ability to complex financial tasks and is a key basis for measuring whether the model can adapt to high-difficulty financial business scenarios.
[0050] Through the implementation of the above embodiments, zero-shot, few-shot, thought chain, and combined prompt templates are constructed respectively, and corresponding input sequences and output results are generated. The performance of the model can be comprehensively tested under different inference conditions and task examples. Zero-shot prompts can measure the model's knowledge generalization ability, few-shot prompts are used to verify its example learning ability, thought chain prompts can test its inference depth and logical consistency, and combined strategies can evaluate the model's comprehensive processing ability under complex tasks, which can improve the robustness and discriminativeness of the evaluation results, so that the model's ability can be more objectively reflected in multi-dimensional scenarios.
[0051] In some embodiments, step 104 may include: for quantitative calculation tasks in the evaluation task set, scoring the output results based on a preset accuracy index to obtain a first score; for logical reasoning tasks in the evaluation task set, weighting the output results based on a preset logical coherence dimension and reasoning step completeness dimension to obtain a second score; for text interpretation tasks in the evaluation task set, weighting the output results based on at least several dimensions of semantic relevance, business compliance, and risk warning completeness to obtain a third score; for multimodal comprehensive analysis tasks in the evaluation task set, weighting the output results based on at least one dimension of visual feature recognition accuracy, cross-modal information integration capability, and decision logic rationality to obtain a fourth score; and determining the comprehensive performance score of the target large language model based on the first score, second score, third score, and fourth score through a preset comprehensive scoring strategy to determine the comprehensive performance of the target large language model.
[0052] In some examples, the evaluation first targets quantitative calculation tasks within the task set. These tasks focus on numerical computation and formula reasoning in the financial field (such as calculating corporate gross profit margins and applying the Black-Scholes option pricing model). The core assessment is the accuracy of the model's calculation results, so a pre-defined accuracy index is used for scoring. The pre-defined accuracy index is a pre-set quantitative standard used to measure the correctness of the calculation results. Specifically, it is defined as the percentage of correct calculation results output by the model out of the total number of tasks in this category. When judging the correctness of the results, it is necessary to consider the numerical precision requirements in the financial field (such as retaining two decimal places for monetary calculations and one decimal place for percentage calculations). For example, if a quantitative calculation task contains 10 questions, and the model outputs results for 8 questions that are completely consistent with the standard answers (verified by financial experts) (numerical precision meets the requirements), then the pre-defined accuracy index calculation result is 80%. This result is the first score for quantitative calculation tasks, which directly reflects the reliability of the model's financial numerical computation.
[0053] Secondly, for the logical reasoning tasks in the assessment task set, these tasks cover complex decision-making scenarios such as financial security vulnerability analysis and M&A risk derivation. It is necessary to examine both the completeness and rigor of the reasoning process. Therefore, a weighted score is applied based on the pre-set logical coherence dimension and the reasoning step completeness dimension. The pre-set logical coherence dimension assesses whether the causal relationships between each step in the model's reasoning process are reasonable and consistent. For example, when analyzing the "impact of interest rate increases on banking performance," if the model first points out that "interest rate increases improve the deposit-loan interest rate spread," and then subsequently concludes that "banking performance declines" without supplementing other interfering factors (such as an increase in the bad debt ratio), then the logical coherence is deemed insufficient. The pre-set reasoning step completeness dimension assesses whether the model's reasoning process covers the key analytical nodes of the task. For example, in the task of identifying financial system vulnerabilities, if the model omits the crucial step of "assessing the scope of vulnerability impact," then the step completeness is deemed insufficient. When weighting the scores, the weights for "process completeness" and "logical rigor" should be set according to the financial scenario's requirements. Typically, the completeness of the reasoning steps accounts for 40% of the weight, and logical coherence accounts for 60%. For example, if a logic reasoning question has a full score of 100 points, the model's reasoning steps covering all key nodes will get 40 points (40% of the full score), and the reasoning process without causal contradictions and in line with financial business logic will get 54 points (90 points under 60% weight). After weighting, the question will get 94 points. The average score of all logic reasoning tasks is the second score, which accurately reflects the model's reasoning ability for complex decisions.
[0054] Furthermore, for text interpretation tasks within the assessment task set, these tasks involve scenarios such as financial research report analysis, regulatory policy interpretation, and customer consultation response. It is necessary to simultaneously examine the model's accuracy in understanding the text, its compliance awareness, and its risk warning awareness. Therefore, a weighted score is applied based on at least several dimensions, including semantic relevance, business compliance, and the completeness of risk warnings. The semantic relevance dimension assesses the degree of matching between the model's output and the core meaning of the text to be interpreted. For example, when interpreting the text "A listed company's annual report shows a 10% increase in net profit," if the model's output focuses on "the reasons for the net profit growth and future profit forecasts," then the semantic relevance is high. The business compliance dimension assesses whether the model's output complies with financial regulatory rules (such as the "Guidelines for Compliance Risk Management of Commercial Banks"). For example, investment advice outputs should not contain illegal statements such as "guaranteeing principal and returns." The completeness of risk warnings dimension assesses whether the model covers the key risk points related to the text to be interpreted. For example, when interpreting "the investment value of high-yield bonds," it is necessary to warn of core risks such as "credit default risk" and "interest rate fluctuation risk." When weighting the scores, the weights of each dimension can be assigned according to the needs of the financial scenario. For example, semantic relevance accounts for 33%, business compliance accounts for 34%, and risk warning completeness accounts for 33%. If the model scores 30 points (out of 33) for semantic relevance, 34 points (out of 34) for business compliance, and 28 points (out of 33) for risk warning completeness in a certain text interpretation task, then the weighted score is (30+34+28)=92 points. The average score of all text interpretation tasks is the third score, which reflects the model's ability to understand financial texts and its awareness of compliance output.
[0055] Then, for the multimodal comprehensive analysis tasks in the evaluation task set, these tasks involve the collaborative analysis of financial images and text (such as trend judgment combining candlestick charts and market data, and risk analysis combining financial statement charts and textual explanations). The evaluation needs to assess the model's accuracy in recognizing visual information, its ability to integrate cross-modal information, and the rationality of its decisions. Therefore, a weighted score is applied based on at least one of the following dimensions: accuracy in visual feature recognition, ability to integrate cross-modal information, and rationality of decision logic. The accuracy in visual feature recognition assesses the model's precision in extracting key visual information from financial images, such as accurately capturing core features like opening price, closing price, and the position of the MA5 moving average when recognizing candlestick charts. The ability to integrate cross-modal information assesses whether the model forms a unified conclusion with related textual information; for example, if the candlestick chart shows a "bullish trend" and the textual interpretation supports this judgment, there is no disconnect between the image and textual conclusions. The rationality of decision logic assesses whether the decision recommendations generated by the model based on multimodal information conform to financial business logic; for example, combining "bullish candlestick trend + company net profit growth text" to give a "short-term hold" recommendation, rather than an unfounded "sell" recommendation. When weighting the scores, visual feature recognition accuracy typically accounts for 40% of the weight, cross-modal information integration capability accounts for 30%, and decision logic rationality accounts for 30%. If a model scores 36 points (out of 40) for visual feature recognition, 27 points (out of 30) for cross-modal integration, and 27 points (out of 30) for decision logic in a multimodal task, the weighted score is 90 points. The average score of all multimodal comprehensive analysis tasks is the fourth score, which measures the model's comprehensive ability to process financial multimodal information.
[0056] Finally, based on the first, second, third, and fourth scores, the overall performance score of the target large language model is determined through a pre-set comprehensive scoring strategy. This strategy pre-determines the weighting of various tasks according to the expected application areas of the target large language model (such as intelligent risk control, intelligent investment research, and financial customer service), ensuring the comprehensive score focuses on the model's core application capabilities. For example, when evaluating a model for intelligent risk control scenarios, the weighting of logical reasoning tasks (including security and compliance reasoning) can be set at 40%, quantitative calculation at 20%, text interpretation at 30%, and multimodal comprehensive analysis at 10%. When evaluating a model for intelligent investment research scenarios, the weighting of text interpretation (research report analysis) is set at 35%, multimodal comprehensive analysis (chart and graph analysis) at 25%, logical reasoning (investment logic derivation) at 30%, and quantitative calculation at 10%. Taking investment research scenarios as an example, if the first score is 80 (weight 10%), the second score is 94 (weight 30%), the third score is 88 (weight 35%), and the fourth score is 90 (weight 25%), then the comprehensive performance score is calculated as 80×10%+94×30%+88×35%+90×25%=8+28.2+30.8+22.5=89.5 points. This comprehensive performance score directly quantifies the model's overall capabilities in investment research scenarios. Combining the score, we can further analyze the model's strengths (such as strong logical reasoning and text interpretation capabilities) and weaknesses (such as the need to optimize quantitative calculations), and finally determine whether the comprehensive performance of the target large language model meets the needs of actual financial applications.
[0057] By implementing the above embodiments, a correspondence is established between different types of tasks and corresponding evaluation dimensions and quantitative indicators, and scores for quantitative calculation, logical reasoning, text interpretation, and multimodal analysis are calculated respectively, which can achieve a highly targeted multidimensional quantitative evaluation. The scoring dimensions of each task category are designed to match the actual business needs of finance. For example, calculation tasks focus on accuracy, reasoning tasks focus on logical integrity, text tasks emphasize semantics and compliance, and multimodal tasks examine visual recognition and cross-modal fusion capabilities, which can make the evaluation results more detailed and objective, and can comprehensively reflect the model's actual business capabilities and overall performance.
[0058] In some embodiments, the aforementioned determination of the comprehensive performance score of the target large language model based on the first score, second score, third score, and fourth score through a preset comprehensive scoring strategy may include: weighting the first score, second score, third score, and fourth score based on preset task type weights to obtain an initial comprehensive score, wherein the preset task type weights are adjusted based on the expected application domain of the target large language model; and correcting the initial comprehensive score based on the performance stability of the target large language model under various preset prompting strategies to obtain the comprehensive performance score.
[0059] In some examples, an initial comprehensive score can first be calculated based on preset task type weights. These preset task type weights are determined according to the expected application areas of the target large language model (such as intelligent risk control systems, intelligent investment research platforms, financial customer service robots, etc.), and are set as the proportion of quantitative calculation, logical reasoning, text interpretation, and multimodal comprehensive analysis tasks in the comprehensive score. The core function of these weights is to ensure that the evaluation focuses on the core application capabilities of the model and avoids the scoring from being out of touch with actual needs due to a one-size-fits-all average weight. The determination of the weights for preset task types requires research by domain experts and analysis of business needs. For example, if the target large language model is expected to be applied to intelligent risk control scenarios, which have the highest requirements for the model's ability to identify risks and deduce compliance vulnerabilities, then the weight of logical reasoning tasks (including financial security compliance reasoning) can be set at 40%, the weight of text interpretation tasks (such as interpreting risk control requirements in regulatory policies) can be set at 30%, and the weights of quantitative calculation tasks (such as numerical calculation of risk exposure) and multimodal comprehensive analysis tasks (such as judging risk levels by combining risk charts and text) can be set at 20% and 10%, respectively. If the model is expected to be applied to intelligent investment research scenarios, which emphasize the model's ability to analyze research reports and financial charts, then the weights of text interpretation tasks can be set at 35%, multimodal comprehensive analysis tasks (such as combining candlestick charts and market data analysis) at 25%, logical reasoning tasks (such as investment logic deduction) at 30%, and quantitative calculation tasks (such as valuation model calculation) at 10%. In the specific weighted calculation, the first score (quantitative calculation), the second score (logical reasoning), the third score (text interpretation), and the fourth score (multimodal comprehensive analysis) need to be multiplied by the corresponding preset task type weights and then summed. For example, in the intelligent investment research scenario, if the first score is 85, the second score is 90, the third score is 88, and the fourth score is 92, the initial comprehensive score calculated according to the above weights is 85×10%+90×30%+88×35%+92×25%=8.5+27+30.8+23=89.3 points. This initial comprehensive score reflects the average capability level of the model in the target application scenario.
[0060] Secondly, the initial comprehensive score can be corrected based on the performance stability of the target large language model under various preset prompting strategies. Performance stability refers to the consistency and volatility of the model's output results for the same type of task under zero-shot prompting strategy, few-shot prompting strategy, thought chain prompting strategy, and combined prompting strategy. If the model's scores for the same type of task are small under different prompting strategies (e.g., scores of 88, 89, 87, and 88 for quantitative calculation tasks under the four strategies), then the performance stability is high. If the score differences are large (e.g., scores of 70, 90, 85, and 65 respectively), then the performance stability is low. The financial field has extremely high requirements for model stability (e.g., robo-advisors need to stably output reasonable suggestions under different customer consultation scenarios to avoid contradictory conclusions due to different prompting methods). Therefore, stability correction is needed to reflect the reliability of the model. The stability of performance can be quantified by calculating the coefficient of variation (COP) of scores for the same type of task under different prompting strategies. The COP is the ratio of the standard deviation of the score to the mean. The smaller the ratio, the higher the stability. For example, if the average score of a certain type of task under four strategies is 85 and the standard deviation is 2, then the COP is 2 / 85≈0.023, indicating high stability. If the mean is 80 and the standard deviation is 10, then the COP is 10 / 80=0.125, indicating low stability. During the calibration process, corresponding stability calibration coefficients need to be set for different coefficient of variation ranges. For example, when the coefficient of variation is ≤0.05, the calibration coefficient is 1.0 (no adjustment), when 0.05 < coefficient of variation ≤0.1, the calibration coefficient is 0.95, and when the coefficient of variation >0.1, the calibration coefficient is 0.9. The calibration coefficient is multiplied by the initial comprehensive score to obtain the calibrated comprehensive performance score. For example, the initial comprehensive score of the above intelligent investment research scenario is 89.3 points. If the coefficient of variation of the model's stability is 0.08 (corresponding to the calibration coefficient of 0.95), then the comprehensive performance score is 89.3 × 0.95 ≈ 84.8 points. This score reflects both the model's average ability in the target scenario and the reliability of its output through stability adjustment, avoiding the misjudgment of a model with a high initial score but low stability as high performance.
[0061] By implementing the above embodiments, the scores of different task categories are weighted and fused, and the model's performance stability under multi-prompt strategies is corrected to generate a comprehensive performance score that better meets the needs of actual applications. The weight allocation can be adjusted according to the model's expected financial application scenarios to ensure that the evaluation results match business needs. The stability correction can reduce the deviation caused by differences in prompting methods, thereby improving the fairness and reliability of the scoring system. It can more accurately quantify the overall performance of the financial big language model and provide a scientific basis for model selection, deployment and continuous improvement.
[0062] Furthermore, as an implementation of the foregoing method embodiments, this application also provides a large language model evaluation apparatus for implementing the foregoing method embodiments. This apparatus embodiment corresponds to the foregoing method embodiments. For ease of reading, this large language model evaluation apparatus embodiment will not repeat the details of the foregoing method embodiments one by one, but it should be clear that the apparatus in this application embodiment can correspondingly implement all the contents of the foregoing method embodiments. For example... Figure 2 As shown, the large language model evaluation device 20 includes: a data acquisition unit 201, a task generation unit 202, a model testing unit 203, and a performance evaluation unit 204. The data acquisition unit 201 acquires benchmark data in the financial field, which may include structured text data and structured image data cross-modally associated with the structured text data. The task generation unit 202 generates an evaluation task set for the target large language model based on the benchmark data and preset evaluation tasks. The preset evaluation tasks may include at least one of quantitative calculation, logical reasoning, text interpretation, and multimodal comprehensive analysis. The model testing unit 203 inputs the evaluation task set into the target large language model using various preset prompting strategies to obtain the output results of the target large language model. The performance evaluation unit 204 performs multidimensional analysis on the output results based on the evaluation dimensions and quantitative indicators associated with the evaluation task set to determine the comprehensive performance of the target large language model.
[0063] This application also provides a computer-readable storage medium storing computer-executable instructions or a computer program, which, when executed by a processor, will cause the processor to perform any step of the large language model evaluation method provided in this application.
[0064] In some embodiments, the computer-readable storage medium may be a random access memory (RAM), a read-only memory (ROM), flash memory, a magnetic surface memory, an optical disc, or a compact disc read-only memory (CD-ROM); or it may be a variety of devices that include one or any combination of the above-mentioned memories.
[0065] In some embodiments, computer-executable instructions may take the form of programs, software, software modules, scripts, or code, written in any form of programming language (including compiled or interpreted languages, or declarative or procedural languages), and may be deployed in any form, including as stand-alone programs or as modules, components, subroutines, or other units suitable for use in a computing environment.
[0066] In some embodiments, computer-executable instructions may, but do not necessarily, correspond to files in a file system, and may be stored as part of a file that holds other programs or data, for example, in one or more scripts in a HyperText Markup Language (HTML) document, in a single file dedicated to the program in question, or in multiple co-located files (e.g., files that store one or more modules, subroutines, or code sections).
[0067] In some embodiments, computer-executable instructions may be deployed to execute on an electronic device, or on multiple electronic devices located at one location, or on multiple electronic devices distributed across multiple locations and interconnected via a communication network.
[0068] like Figure 3 As shown, this application also provides an electronic device 30, including a memory 310, a processor 320, and a computer program 311 stored in the memory 310 and executable on the processor. When the processor 320 executes the computer program 311, it implements any step of the above-described large language model evaluation method.
[0069] This application also provides a computer program product comprising a computer program or computer-executable instructions stored in a computer-readable storage medium. A processor of an electronic device reads the computer program or computer-executable instructions from the computer-readable storage medium and executes the computer program or computer-executable instructions, causing the electronic device to perform any step of the large language model evaluation method described above.
[0070] The above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application.
Claims
1. A method for evaluating large language models, characterized in that, include: Acquire benchmark data in the financial field, wherein the benchmark data includes structured text data and structured image data that is cross-modal associated with the structured text data; Based on the evaluation benchmark data and the preset evaluation tasks, an evaluation task set for the target large language model is generated, wherein the types of the preset evaluation tasks include at least one of quantitative calculation, logical reasoning, text interpretation and multimodal comprehensive analysis. The evaluation task set is input into the target large language model through a variety of preset prompting strategies to obtain the output results of the target large language model; Based on the evaluation dimensions and quantitative indicators associated with the evaluation task set, a multidimensional analysis is performed on the output results to determine the comprehensive performance of the target large language model.
2. The large language model evaluation method according to claim 1, characterized in that, The acquisition of benchmark data in the financial field includes: Acquire financial academic knowledge data within a first preset range, and process the financial academic knowledge data according to a first preset annotation rule to generate structured multiple-choice question data containing knowledge point affiliation and difficulty level; Acquire practical data of the financial industry within a second preset range, and process the practical data of the financial industry according to the second preset annotation rules to generate objective short-answer question data containing core entity annotations and subjective open-ended question data containing business constraint annotations; Acquire financial security compliance data within a third preset range, and process the financial security compliance data according to the third preset annotation rules to generate security test question data containing a comparison structure of risk options and correct solutions; Acquire financial agent task data within a fourth preset range, and process the financial agent task data according to the fourth preset annotation rules to generate agent process test question data containing task process nodes, tool call specifications and multi-round interaction logic; The structured multiple-choice question data, the objective short-answer question data, the subjective open-ended question data, the security test question data, and the intelligent agent process test question data are integrated to obtain the evaluation benchmark data.
3. The large language model evaluation method according to claim 1, characterized in that, The step of generating an evaluation task set for the target large language model based on the evaluation benchmark data and preset evaluation tasks includes: Based on the type of the preset evaluation task, the evaluation task set is obtained by filtering from the evaluation benchmark data, wherein the preset prompting strategy includes at least one of the zero-sample prompting strategy, the small-sample prompting strategy, and the thinking chain prompting strategy.
4. The large language model evaluation method according to claim 3, characterized in that, The process of inputting the evaluation task set into the target large language model through multiple preset prompting strategies to obtain the output results of the target large language model includes: Based on the zero-shot prompting strategy, a first prompt template without task examples is generated. The questions in the evaluation task set are constructed as a first model input sequence based on the first prompt template, and the first output result generated by the target large language model based on the first model input sequence is obtained. Based on the few-sample prompting strategy, a second prompting template containing multiple task examples and corresponding answers is generated. The questions in the evaluation task set are constructed as a second model input sequence based on the second prompting template, and the second output result generated by the target large language model based on the second model input sequence is obtained. Based on the thought chain prompting strategy, a third prompt template containing step-by-step reasoning requirements is generated. The questions in the evaluation task set are constructed as a third model input sequence based on the third prompt template. The third output result, which contains the reasoning process and the final answer, is generated by the target large language model based on the third model input sequence. Based on a combination of small sample size and thought chain prompting strategy, a fourth prompting template is generated that simultaneously contains task examples and reasoning requirements. The questions in the evaluation task set are constructed as a fourth model input sequence based on the fourth prompting template, and the fourth output result generated by the target large language model based on the fourth model input sequence is obtained.
5. The large language model evaluation method according to claim 1, characterized in that, The multidimensional analysis of the output results, based on the evaluation dimensions and quantitative indicators associated with the evaluation task set, to determine the comprehensive performance of the target large language model includes: For quantitative calculation tasks in the evaluation task set, the output results are scored based on a preset accuracy index to obtain a first score; For logical reasoning tasks in the evaluation task set, the output results are weighted and scored based on a preset logical coherence dimension and reasoning step completeness dimension to obtain a second score. For text interpretation tasks in the assessment task set, the output results are weighted and scored based on at least several dimensions, including semantic relevance, business compliance, and risk warning completeness, to obtain a third score. For multimodal comprehensive analysis tasks in the evaluation task set, the output results are weighted and scored based on at least one dimension of visual feature recognition accuracy, cross-modal information integration capability, and decision logic rationality to obtain a fourth score; Based on the first score, the second score, the third score, and the fourth score, the comprehensive performance score of the target large language model is determined through a preset comprehensive scoring strategy, thereby determining the comprehensive performance of the target large language model.
6. The large language model evaluation method according to claim 5, characterized in that, The determination of the comprehensive performance score of the target large language model based on the first score, the second score, the third score, and the fourth score using a preset comprehensive scoring strategy includes: Based on preset task type weights, the first score, the second score, the third score, and the fourth score are weighted and calculated to obtain an initial comprehensive score. The preset task type weights are adjusted based on the expected application domain of the target large language model. Based on the performance stability of the target large language model under the various preset prompting strategies, the initial comprehensive score is corrected to obtain the comprehensive performance score.
7. A large language model evaluation device, characterized in that, include: A data acquisition unit is used to acquire assessment benchmark data in the financial field, wherein the assessment benchmark data includes structured text data and structured image data that is cross-modal associated with the structured text data; The task generation unit is used to generate an evaluation task set for the target large language model based on the evaluation benchmark data and the preset evaluation tasks. The preset evaluation tasks include at least one of the following types: quantitative calculation, logical reasoning, text interpretation, and multimodal comprehensive analysis. The model testing unit is used to input the evaluation task set into the target large language model through various preset prompting strategies in order to obtain the output results of the target large language model. The performance evaluation unit is used to perform multidimensional analysis on the output results based on the evaluation dimensions and quantitative indicators associated with the evaluation task set, and to determine the comprehensive performance of the target large language model.
8. An electronic device, comprising: A memory and a processor, characterized in that the processor, when executing a computer program stored in the memory, implements the steps of the large language model evaluation method as described in any one of claims 1 to 6.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the large language model evaluation method as described in any one of claims 1 to 6.
10. A computer program product comprising a computer program or computer-executable instructions, characterized in that, When the computer program or computer-executable instructions are executed by a processor, they implement the steps of the large language model evaluation method as described in any one of claims 1 to 6.
Citation Information
Patent Citations
AI model dynamic evaluation method and system based on multi-agent collaboration
CN121658336A