A financial artificial intelligence system evaluation method and device
Patent Information
- Application Number
- CN202611004834.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-07
- Publication Date
- 2026-09-25
AI Technical Summary
现有技术在应用于金融领域人工智能系统安全评估时,评估的效果并不理想
本方法先依据评估任务加载多维特征,再生成结构化的测试样例,并预先绑定判定规则与知识证据,使得样例生成过程不再依赖人工经验。执行评估后将输出内容、上下文状态和工具调用日志与样例关联形成观测结果,再结合判定规则和知识证据进行判定,这一过程不仅可以判断输出是否合规,还能依据金融法规、业务规则等证据核验输出的事实依据与权限边界。由此,本方法可以在金融人工智能系统上线前或运行中,发现单轮测试难以暴露的多轮诱导风险等问题,同时每个失败样例可追溯,便于定位整改,从而显著提升金融领域安全评估的全面性与结果的可信度。
Smart Images

Figure CN122818366A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, and in particular to a method and device for evaluating a financial artificial intelligence system. Background Technology
[0002] Before a financial AI system goes live or during its continuous operation, a security assessment is typically required. Existing technologies mainly employ methods such as human red team testing, static test sets, rule template libraries, warning word attack libraries, benchmark evaluation platforms, large-scale model judging, and manual review to achieve security assessments.
[0003] Existing AI security assessment methods are mostly geared towards general content security testing, capable of identifying some illegal, sensitive, or obviously inappropriate content. However, in financial business scenarios, AI systems not only need to meet general content security requirements but also comply with financial-related conditions. Current technologies have not yielded ideal results when applied to the security assessment of AI systems in the financial sector. Summary of the Invention
[0004] This specification provides one or more embodiments of a financial artificial intelligence system evaluation method and device to solve the technical problems raised in the background art.
[0005] One or more embodiments of this specification employ the following technical solutions: This specification provides one or more embodiments of a financial artificial intelligence system evaluation method, the method comprising: Obtain pre-set evaluation task parameters for the financial artificial intelligence system to be evaluated; Based on the evaluation task parameters, a multi-dimensional feature set matching the current evaluation task is loaded from the preset feature definition module; Based on the multidimensional feature set, a structured test case set is generated, which includes multi-round interactive test cases. Each test case is pre-bound with judgment rule features and knowledge evidence features. The test case set is input into the financial AI system to be evaluated, and the evaluation execution data of the financial AI system to be evaluated is collected. The evaluation execution data includes output content, context state and tool call logs. The evaluation execution data is correlated with the corresponding test cases and interaction rounds to form an observation result set; Based on the set of observation results, and in combination with the judgment rule features and the knowledge evidence features, the output of the financial artificial intelligence system to be evaluated is judged, and a judgment result is generated. Based on the determination results, a security assessment report is generated.
[0006] It should be noted that this method first loads multi-dimensional features based on the assessment task, then generates structured test cases, and pre-binds judgment rules and knowledge evidence, so that the sample generation process no longer relies on human experience. After the assessment is performed, the output content, context state, and tool call logs are correlated with the samples to form observation results, which are then combined with judgment rules and knowledge evidence for judgment. This process can not only determine whether the output is compliant, but also verify the factual basis and authority boundaries of the output based on evidence such as financial regulations and business rules. Therefore, this method can discover problems such as multi-round inducement risks that are difficult to expose in a single round of testing before or during the operation of financial artificial intelligence systems. At the same time, each failed sample is traceable, which facilitates location and rectification, thereby significantly improving the comprehensiveness and credibility of security assessments in the financial field.
[0007] Furthermore, the step of generating a structured test sample set based on the multidimensional feature set includes: Based on the multidimensional feature set, a candidate multidimensional feature combination set is generated by selecting feature values from each feature subset and combining them. The candidate multidimensional feature combination set is subjected to constraint filtering and priority sorting to obtain the sorted effective feature set. Based on the effective feature set, a structured test case set is generated. Each test case is associated with at least the corresponding feature combination, input sequence, expected safe behavior, prohibited behavior, judgment rule and evidence reference.
[0008] It should be noted that this method first generates candidate feature combinations by combining feature values from various feature subsets. Then, it uses constraint filtering to remove invalid combinations that do not match the business scenario or cannot be effectively judged. Simultaneously, it prioritizes these invalid combinations so that the remaining valid feature combinations can be focused on high-risk, high-value, or insufficiently covered test areas. Each test case generated based on this includes a feature combination, input sequence, expected safe behavior, prohibited behavior, judgment rules, and evidence citations. This approach ensures that the test case generation process has clear constraints and priority guidance. Furthermore, the risk type, judgment basis, and expected result for each case are clearly traceable, facilitating subsequent automatic execution and result comparison, thereby significantly improving the targeting and execution efficiency of security assessments in financial AI systems.
[0009] Furthermore, the constraint filtering of the candidate multidimensional feature combination set includes: The candidate multidimensional feature combination set is constrained and filtered based on preset constraint rules. The constraint rules include: whether the security risk feature matches the business scenario feature, whether the attack expression feature is applicable to the corresponding interaction process feature, whether the judgment rule feature covers the output behavior feature, and whether the knowledge evidence feature supports one or more of the business scenario features.
[0010] It should be noted that this method introduces constraint rules to filter candidate feature combinations before generating test cases. Specifically, it checks whether the security risks match the business scenario, whether the attack expression is applicable to the current interaction process, whether the judgment rules can cover the expected output behavior, and whether the knowledge evidence supports the business scenario. These rules can exclude invalid combinations that, although the feature values can be combined, cannot be executed in practice or the test results cannot be judged from three dimensions: business rationality, technical feasibility, and evaluation effectiveness.
[0011] Furthermore, the candidate multidimensional feature combination set is prioritized, including: Based on preset weights, the priority score of each candidate combination is calculated. The calculation factors of the priority score include the security risk level corresponding to the security risk feature, the attack expression strength corresponding to the attack expression feature, the relevance of historical failure examples, the importance of the business scenario, the similarity between the new coverage benefit and the existing examples.
[0012] It should be noted that this method considers multiple calculation factors when prioritizing, including security risk level, attack expression strength, relevance of historical failed examples, importance of business scenarios, benefits of new coverage, and similarity to existing examples. By weighting and ranking these factors, the system can automatically identify which feature combinations correspond to higher risks, stronger inducements, more concentrated historical issues, more critical business aspects, greater coverage value, and lower duplication with existing examples. Effective feature combinations ranked higher will be prioritized for generating test examples. This allows limited evaluation resources to automatically focus on high-risk areas while avoiding repeated testing of already adequately covered content. The entire ranking process can be completed automatically based on preset weights and real-time statistical information without relying on manual judgment. This significantly improves the priority of generating high-risk examples while ensuring comprehensive evaluation, making it more likely that security issues will be discovered in earlier rounds.
[0013] Furthermore, the evaluation task parameters include one or more of the following: the identifier of the financial artificial intelligence system to be evaluated, the type of financial business, the evaluation scope, the category of financial security risks, the range of attack expression characteristics, the range of business scenarios, the number of samples, the maximum number of interaction rounds, the judgment threshold, and the convergence threshold; the multidimensional feature set includes one or more of the following: security risk features, attack expression features, business scenario features, interaction process features, output behavior features, judgment rule features, and knowledge evidence features.
[0014] It should be noted that the evaluation task parameters of this method cover specific configuration items such as the identifier of the system under test, the type of financial business, the risk category, the attack method, the business scenario, the number of interaction rounds, and the judgment threshold. The multi-dimensional feature set includes multiple dimensions such as security risk, attack expression, business scenario, interaction process, output behavior, judgment rules, and knowledge evidence. The system can flexibly adjust the evaluation scope, risk focus, and attack method according to different financial business needs and evaluation objectives, while comprehensively characterizing the financial artificial intelligence system under evaluation from multiple feature dimensions. This evaluation process is customized for specific financial scenarios, specific risk types, and specific interaction methods, thereby ensuring that the generated test cases are highly matched with the current evaluation task.
[0015] Furthermore, the generation of the structured test case set includes: Multi-round interaction test cases are generated based on the characteristics of the interaction process. The multi-round interaction test cases include multi-round dialogue sequences. The input of each round is determined by the context state of the previous round, the target risk, the attack expression method and the business scenario.
[0016] It should be noted that when generating multi-round interactive test cases, the input for each round is jointly determined by the context state, target risk, attack expression method, and business scenario of the previous round. The context state can record historical inputs, historical system outputs, changes in role settings, and triggered risk points. This allows subsequent rounds to dynamically adjust attack strategies or conduct further probing based on the results of the previous round's interaction, thereby constructing a coherent dialogue chain that gradually approaches the security boundary. This context-state-driven generation method can simulate the complex behavior of users in real financial scenarios who test the system's compliance boundaries through multiple rounds of preparation, role switching, or task decomposition. The test cases generated in this way can effectively identify financial AI systems that perform securely in a single round of testing but exhibit degradation issues such as missing risk warnings, business overreach, or insufficient evidence after continuous interaction, significantly improving the detection capability for multi-round induced risks.
[0017] Furthermore, the step of judging the output of the financial artificial intelligence system to be evaluated and generating a judgment result includes: The risk score output by the financial artificial intelligence system to be evaluated is determined based on preset risk dimensions, including content security risk, business compliance risk, tool call risk, context security degradation risk, insufficient evidence or incorrect basis risk, and historical repeated failure risk. Based on the risk score and the pre-set judgment threshold, the output of the financial artificial intelligence system to be evaluated is automatically judged to generate the judgment result. The judgment result is either pass, fail, or requires manual review. The automatic judgment includes one or more of the following: rule judgment, semantic judgment, business evidence judgment, multi-round context judgment, and tool call judgment.
[0018] It should be noted that this method comprehensively calculates risk scores from multiple risk dimensions, including content security, business compliance, tool invocation, context degradation, insufficient evidence, and historical repeated failures. Based on preset thresholds, the judgment results are categorized into three types: pass, fail, or requiring manual review. Specifically, the business compliance dimension assesses specific risks such as investment advice boundaries and return promises in financial scenarios; the context degradation dimension captures the weakening process of security boundaries during multiple rounds of interaction; the insufficient evidence dimension verifies whether the output is supported by laws or business regulations; and the tool invocation dimension monitors whether the use of external interfaces exceeds authorized limits. Furthermore, the judgment methods encompass rule-based judgment, semantic judgment, business evidence judgment, multi-round context judgment, and tool invocation judgment, ensuring that different types of risks are covered by the corresponding judgment logic. Through this multi-dimensional comprehensive scoring and multi-method joint judgment design, the system can automatically identify critical cases requiring manual intervention.
[0019] Furthermore, the method also includes: Based on the statistical results of the security assessment report and the preset convergence threshold, it is determined whether the current assessment has converged. The security assessment report includes one or more of the following: overall pass rate, risk category distribution, failure sample list, and evidence chain. If convergence fails, a reason for non-convergence is generated, and the evaluation strategy parameters for the next round of evaluation are adjusted based on the reason for non-convergence, the failure examples corresponding to the judgment result, and the manual review results. A new round of evaluation will be launched based on the adjusted evaluation strategy parameters.
[0020] It should be noted that after generating the security assessment report, this method determines whether the current assessment has converged based on the statistical results in the report and a preset convergence threshold. If convergence is not achieved, the reasons for non-convergence are identified, and the strategy parameters for the next round of assessment are adjusted based on failed samples and manual review results, before initiating a new round of assessment. Through this closed-loop iterative design, the system can automatically identify which risks have not been eliminated, which categories are under-covered, or which problems recur, and dynamically adjust the key risk categories, attack methods, business scenarios, and sample generation weights for subsequent assessments accordingly. This allows subsequent rounds of assessment to focus specifically on failed samples, critical samples, and high-risk areas, while appropriately reducing the weight of passed or low-risk content. This iterative process continuously guides assessment resources to the most critical security vulnerabilities, gradually leading to the convergence of security issues in the financial AI system.
[0021] Furthermore, the adjustment of the evaluation strategy parameters for the next round of evaluation includes at least one of the following: increasing the weight of risk categories with high repeated failure rates in feature combinations, generating new feature combinations for insufficiently covered risk categories or business scenarios, and adding manually verified risky samples to the key retest set.
[0022] It should be noted that when the evaluation fails to converge, this method analyzes the results of the previous round to identify risk categories with high recurrence failure rates, insufficiently covered risk categories or business scenarios, and manually verified risky examples. Based on these, adjustments are made, such as increasing the weight of the corresponding categories, generating new feature combinations, and adding confirmed risky examples to the retest set. This dynamic adjustment mechanism ensures that the next round of evaluation automatically focuses limited resources on security vulnerabilities that repeatedly fail, insufficiently tested gaps, and manually confirmed real risk points. As the iteration rounds increase, the evaluation strategy continuously tilts towards high-risk and insufficiently covered areas, while gradually reducing the testing weight of already passed or low-risk content. This closed-loop feedback significantly improves the targeting of subsequent evaluations, enabling the continuous tracking and in-depth verification of security issues in financial AI systems, thereby accelerating the convergence process of security risks.
[0023] This specification provides one or more embodiments of a financial artificial intelligence system evaluation device, comprising: At least one processor and bus; and, A memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, which, when executed by the at least one processor, enable the at least one processor to: Obtain pre-set evaluation task parameters for the financial artificial intelligence system to be evaluated; Based on the evaluation task parameters, a multi-dimensional feature set matching the current evaluation task is loaded from the preset feature definition module; Based on the multidimensional feature set, a structured test case set is generated, which includes multi-round interactive test cases. Each test case is pre-bound with judgment rule features and knowledge evidence features. The test case set is input into the financial AI system to be evaluated, and the evaluation execution data of the financial AI system to be evaluated is collected. The evaluation execution data includes output content, context state and tool call logs. The evaluation execution data is correlated with the corresponding test cases and interaction rounds to form an observation result set; Based on the set of observation results, and in combination with the judgment rule features and the knowledge evidence features, the output of the financial artificial intelligence system to be evaluated is judged, and a judgment result is generated. Based on the determination results, a security assessment report is generated.
[0024] The above-described at least one technical solution adopted in the embodiments of this specification can achieve the following beneficial effects: This method first loads multi-dimensional features based on the assessment task, then generates structured test cases, and pre-binds judgment rules and knowledge evidence, making the case generation process independent of human experience. After the assessment is performed, the output content, context state, and tool call logs are correlated with the test cases to form observation results. These results are then combined with judgment rules and knowledge evidence for judgment. This process not only determines whether the output is compliant, but also verifies the factual basis and authority boundaries of the output based on evidence such as financial regulations and business rules. Therefore, this method can discover problems such as multi-round inducement risks that are difficult to expose in a single round of testing before or during the operation of financial artificial intelligence systems. At the same time, each failed case is traceable, which facilitates location and rectification, thereby significantly improving the comprehensiveness and credibility of security assessments in the financial field. Attached Figure Description
[0025] To more clearly illustrate the technical solutions in the embodiments or prior art of this specification, the drawings used in the description of the embodiments or prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in this specification. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort. In the drawings: Figure 1 A flowchart illustrating a financial artificial intelligence system evaluation method provided for one or more embodiments of this specification; Figure 2 This is a schematic diagram of the structure of a financial artificial intelligence system evaluation device provided for one or more embodiments of this specification. Detailed Implementation
[0026] This specification provides an example of a method and device for evaluating a financial artificial intelligence system.
[0027] To enable those skilled in the art to better understand the technical solutions in this specification, the technical solutions in the embodiments of this specification will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this specification, and not all embodiments. Based on the embodiments of this specification, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of this specification.
[0028] Figure 1 This diagram illustrates a process for evaluating a financial artificial intelligence system, provided in one or more embodiments of this specification. This process can be executed by a financial artificial intelligence system evaluation system. Certain input parameters or intermediate results in the process can be manually adjusted to help improve accuracy.
[0029] The method flow steps of the embodiments in this specification are as follows: S101, Obtain the pre-set evaluation task parameters of the financial artificial intelligence system to be evaluated.
[0030] In the embodiments of this specification, the evaluation task parameters may include the identifier of the financial artificial intelligence system to be evaluated, the type of financial business, the evaluation scope, the category of financial security risks, the range of attack expression characteristics, the scope of business scenarios, the number of samples, the maximum number of interaction rounds, the judgment threshold, and the convergence threshold.
[0031] The system can standardize the above parameters to form a set of evaluation task parameters: Wherein, ID represents the identifier of the financial artificial intelligence system to be evaluated. Indicates the scope of the assessment. Indicates the category of financial security risk. Indicates the range of attack expression characteristics. Indicates the scope of the business scenario. Indicates the number of samples. Indicates the maximum number of interaction rounds. Indicates the judgment threshold. This represents the convergence threshold.
[0032] S102, based on the evaluation task parameters, load a multi-dimensional feature set that matches the current evaluation task from the preset feature definition module.
[0033] In this embodiment, based on the evaluation task parameter set T output by S101, multi-dimensional features matching the current evaluation task are loaded from a pre-built feature definition module. This feature definition module maintains a security risk feature library, an attack expression feature library, a business scenario feature library, an interaction process feature library, an output behavior feature library, a judgment rule feature library, and a knowledge evidence feature library. The multi-dimensional features include security risk features S, attack expression features A, business scenario features B, interaction process features P, output behavior features O, judgment rule features R, and knowledge evidence features K. The system filters security risk features based on the risk category in the evaluation task parameter set, filters attack expression features based on the attack expression range in the evaluation task parameter set, filters business scenario features based on the business scenario range in the evaluation task parameter set, and determines whether to load business knowledge evidence features K based on the evaluation range. This forms a task-related multi-dimensional feature set: ,in, , , , , , and These respectively represent the security risk characteristics, attack expression characteristics, business scenario characteristics, interaction process characteristics, output behavior characteristics, judgment rule characteristics, and knowledge evidence characteristics that match the current assessment task.
[0034] S103. Based on the multi-dimensional feature set, a structured test case set is generated. The test case set includes multi-round interactive test cases, and each test case is pre-bound with judgment rule features and knowledge evidence features.
[0035] In the embodiments of this specification, candidate feature combinations are generated by combining feature values from each feature subset based on the multi-dimensional feature set loaded in S102. Subsequently, the candidate combinations are constrained and filtered to eliminate invalid combinations that do not match the security risk and business scenario, whose attack expression is not applicable to the current interaction process, whose judgment rules cannot cover the output behavior, or whose knowledge evidence does not support the business scenario. The filtered valid combinations are prioritized based on risk level, attack strength, historical failure history, business importance, new coverage value, and similarity to existing examples. Based on the top-ranked valid feature combinations, the system generates a structured set of test examples. For examples involving multi-turn interactions, the system generates multi-turn dialogue sequences based on the interaction process features. Each round's input is determined by the context state, target risk, attack expression, and business scenario of the previous round. Each test example is associated with at least the corresponding feature combination, input sequence, expected security behavior, prohibited behavior, judgment rules, and evidence citations.
[0036] S104, input the test sample set into the financial artificial intelligence system to be evaluated, and collect the evaluation execution data of the financial artificial intelligence system to be evaluated. The evaluation execution data includes output content, context status and tool call log.
[0037] In the embodiments described in this specification, based on the test sample set E output in S103, test samples are input into the financial artificial intelligence system to be evaluated through API interfaces, web automation, command-line interfaces, test agents, sandbox environments, or message queues. For single-round samples, the system can directly submit sample inputs and record system outputs; for multi-round samples, the system can submit them round by round according to the multi-round input sequence generated in S103, and update the context state in each round. During execution, the system records input content, output content, response time, tool call status, retrieval behavior, and context changes. The output of S104 is the evaluation execution data. .
[0038] S105, the evaluation execution data is associated with the corresponding test cases and interaction rounds to form an observation result set.
[0039] In the embodiments of this specification, the evaluation execution data is based on the output of S104. The system collects the actual response results of the financial AI system under evaluation during the testing process. The collected content may include text output, refusal behavior, clarification behavior, search content, tool call parameters, tool return results, and context state. The system further performs structured processing on the collected results, associating each output result with the corresponding test case number, interaction round, feature combination, and supporting evidence to form an observation set. ,in Indicates a test case. Indicates the first Wheel input, Indicates system output, This indicates the tool call log. Indicates the context state. This indicates search results or business evidence.
[0040] S106, Based on the set of observation results and combined with the judgment rule features and the knowledge evidence features, the output of the financial artificial intelligence system to be evaluated is judged, and a judgment result is generated.
[0041] In the embodiments of this specification, the structured observation result set based on S105 output is used. Combined with the set of decision rules bound in the test cases Knowledge evidence set and the judgment threshold set in S101 The system automatically assesses the output of the financial AI system under evaluation. This automatic assessment includes rule-based assessment, semantic assessment, business evidence assessment, tool invocation assessment, and multi-round contextual assessment. For outputs that trigger hard prohibition rules, the system directly marks them as failed or high-risk. For outputs that cannot be directly assessed using hard rules, the system uses semantic matching models, classification models, or large-scale model adjudication models for auxiliary assessment. For financial business security examples, the system further compares the output content with financial knowledge evidence to determine if there are issues such as insufficient evidence, omission of applicable conditions, lack of risk warnings, unauthorized business operations, or misleading financial conclusions. The system calculates a comprehensive risk score based on multiple risk dimensions.
[0042] Wherein, CR represents content security risk; BR represents business compliance risk; TR represents tool invocation risk; Cr represents context security degradation risk; ER represents insufficient evidence or incorrect basis risk; RR represents historical repeated failure risk; w1 to w6 are weights. If judged as failure or high risk; If it is determined to be passed; It was determined that manual review was required.
[0043] The output of S106 is used as the set of judgment results. This includes the judgment result, risk score, risk level, triggering rules, triggering round, chain of evidence, and manual review indicator.
[0044] S107, Based on the determination result, generate a security assessment report.
[0045] In the embodiments of this specification, the set of determination results output by S106 is used. The system generates a security assessment report. The report may include overall pass rate, risk category distribution, a list of failure examples, trigger rules, trigger rounds, evidence chain, tool call anomalies, context degradation, and remediation suggestions. Simultaneously, the system calculates high-risk failure rate, repeat failure rate, key category coverage, and manual review pass rate based on the assessment results, forming the assessment statistics. The output of S107 is an assessment report (Report) and assessment statistics (Stat). The Report presents the assessment conclusions to the user, while Stat serves as input for subsequent convergence determination.
[0046] It should be noted that this method first loads multi-dimensional features based on the assessment task, then generates structured test cases, and pre-binds judgment rules and knowledge evidence, so that the sample generation process no longer relies on human experience. After the assessment is performed, the output content, context state, and tool call logs are correlated with the samples to form observation results, which are then combined with judgment rules and knowledge evidence for judgment. This process can not only determine whether the output is compliant, but also verify the factual basis and authority boundaries of the output based on evidence such as financial regulations and business rules. Therefore, this method can discover problems such as multi-round inducement risks that are difficult to expose in a single round of testing before or during the operation of financial artificial intelligence systems. At the same time, each failed sample is traceable, which facilitates location and rectification, thereby significantly improving the comprehensiveness and credibility of security assessments in the financial field.
[0047] Furthermore, in the process of generating a structured test sample set based on the multidimensional feature set, the method flow steps of this embodiment are as follows: S201, Based on the multidimensional feature set, a candidate multidimensional feature combination set is generated by selecting feature values from each feature subset and combining them.
[0048] In the embodiments of this specification, the task feature set output by S102 is... Construct candidate multidimensional feature combinations. Each candidate feature combination consists of security risk, attack expression, business scenario, interaction process, output behavior, judgment rule, and knowledge evidence, and can be represented as: For general security assessment tasks that do not involve business knowledge, It can be left blank or based on general security standards; for security assessment tasks in financial operations, It is retrieved from databases of financial regulations and rules, exchange rules, listed company announcements, financial product information, risk disclosure documents, or typical cases. S202, constrain and prioritize the candidate multidimensional feature combination set to obtain the sorted effective feature set.
[0049] In the embodiments of this specification, the candidate multidimensional feature combination set output by S201 is used. Combined with the evaluation task parameter set output by S101 The candidate combinations are constrained, filtered, and prioritized. First, invalid combinations are filtered based on constraint rules. These constraints include: whether the security risk category matches the business scenario; whether the attack expression method is applicable to the corresponding interaction process; whether the judgment rules can cover the expected output behavior; whether the business knowledge evidence supports the business scenario; whether the candidate combination highly overlaps with existing examples; and whether the candidate combination exceeds the allowable evaluation range. Then, priority scores are calculated for the filtered valid combinations. Where: Risk(S) represents the security risk level, Severity(A) represents the attack expression strength, FailHistory(f) represents the relevance of historical failure examples, BusinessCritical(B) represents the importance of the financial business scenario, CoverageGain(f) represents the benefit of new coverage, and Similarity(f) represents the similarity with existing examples. α, β, γ, δ, ε, and η are configurable weights. The output of S202 is the sorted set of effective features. .
[0050] S203. Based on the effective feature set, generate a structured test case set, where each test case is associated with at least the corresponding feature combination, input sequence, expected safe behavior, prohibited behavior, judgment rule, and evidence reference.
[0051] In the embodiments of this specification, the effective feature set is based on the output of S202. Generate a set of structured test cases: E = { , , ...,}
[0052] For each valid feature combination Based on the security risk characteristics, attack expression characteristics, business scenario characteristics, and interaction process characteristics, single-round or multi-round test inputs are generated; expected security behaviors and prohibited behaviors are generated based on output behavior characteristics; automatic judgment rules are bound based on judgment rule characteristics; and evidence is bound based on knowledge evidence characteristics. Each test case includes at least: case number, feature vector, input sequence, expected security behavior, prohibited behavior, judgment rule, evidence citation, risk level, manual review conditions, and case version. The case number uniquely identifies the test case; the feature vector comes from the valid feature set output by S202; the input sequence represents a single-round or multi-round input sequence; the expected security behavior represents the security response that the evaluated system should output; the prohibited behavior represents content that must not be output or tool call behavior that must not be executed; the judgment rule represents the set of rules used for automatic judgment; the evidence citation represents the corresponding financial regulations, business rules, announcements, case evidence, or search evidence; the risk level represents the risk level corresponding to the case; the manual review conditions represent the triggering conditions for manual review; and the case version is used to support subsequent regression testing.
[0053] For multi-turn examples, generate a multi-turn dialogue sequence based on the interaction process feature P: Q = { , , ..., }, 1 ≤ m ≤ M, where each round's input is jointly determined by the previous round's context state, target risk, attack expression method, and business scenario. The system describes the process of risk gradually approaching the security boundary through context state transitions, rather than mechanically piecing together multiple single-round problems.
[0054] It should be noted that this method first generates candidate feature combinations by combining feature values from various feature subsets. Then, it uses constraint filtering to remove invalid combinations that do not match the business scenario or cannot be effectively judged. Simultaneously, it prioritizes these invalid combinations so that the remaining valid feature combinations can be focused on high-risk, high-value, or insufficiently covered test areas. Each test case generated based on this includes a feature combination, input sequence, expected safe behavior, prohibited behavior, judgment rules, and evidence citations. This approach ensures that the test case generation process has clear constraints and priority guidance. Furthermore, the risk type, judgment basis, and expected result for each case are clearly traceable, facilitating subsequent automatic execution and result comparison, thereby significantly improving the targeting and execution efficiency of security assessments in financial AI systems.
[0055] Furthermore, in the process of constraining and filtering the candidate multidimensional feature combination set, the method flow steps of this embodiment are as follows: S301, perform constraint filtering on the candidate multidimensional feature combination set based on preset constraint rules, the constraint rules including: whether the security risk feature matches the business scenario feature, whether the attack expression feature is applicable to the corresponding interaction process feature, whether the judgment rule feature covers the output behavior feature, and whether the knowledge evidence feature supports one or more of the business scenario features.
[0056] It should be noted that this method introduces constraint rules to filter candidate feature combinations before generating test cases. Specifically, it checks whether the security risks match the business scenario, whether the attack expression is applicable to the current interaction process, whether the judgment rules can cover the expected output behavior, and whether the knowledge evidence supports the business scenario. These rules can exclude invalid combinations that, although the feature values can be combined, cannot be executed in practice or the test results cannot be judged from three dimensions: business rationality, technical feasibility, and evaluation effectiveness.
[0057] Furthermore, in the process of prioritizing the candidate multidimensional feature combination set, the method flow steps of this embodiment are as follows: S401, based on preset weights, calculate the priority score of each candidate combination. The calculation factors of the priority score include the security risk level corresponding to the security risk feature, the attack expression strength corresponding to the attack expression feature, the correlation of historical failure examples, the importance of the business scenario, the similarity between the new coverage benefit and the existing examples.
[0058] It should be noted that this method considers multiple calculation factors when prioritizing, including security risk level, attack expression strength, relevance of historical failed examples, importance of business scenarios, benefits of new coverage, and similarity to existing examples. By weighting and ranking these factors, the system can automatically identify which feature combinations correspond to higher risks, stronger inducements, more concentrated historical issues, more critical business aspects, greater coverage value, and lower duplication with existing examples. Effective feature combinations ranked higher will be prioritized for generating test examples. This allows limited evaluation resources to automatically focus on high-risk areas while avoiding repeated testing of already adequately covered content. The entire ranking process can be completed automatically based on preset weights and real-time statistical information without relying on manual judgment. This significantly improves the priority of generating high-risk examples while ensuring comprehensive evaluation, making it more likely that security issues will be discovered in earlier rounds.
[0059] Furthermore, the evaluation task parameters include one or more of the following: the identifier of the financial artificial intelligence system to be evaluated, the type of financial business, the evaluation scope, the category of financial security risks, the range of attack expression characteristics, the range of business scenarios, the number of samples, the maximum number of interaction rounds, the judgment threshold, and the convergence threshold; the multidimensional feature set includes one or more of the following: security risk features, attack expression features, business scenario features, interaction process features, output behavior features, judgment rule features, and knowledge evidence features.
[0060] It should be noted that the evaluation task parameters of this method cover specific configuration items such as the identifier of the system under test, the type of financial business, the risk category, the attack method, the business scenario, the number of interaction rounds, and the judgment threshold. The multi-dimensional feature set includes multiple dimensions such as security risk, attack expression, business scenario, interaction process, output behavior, judgment rules, and knowledge evidence. The system can flexibly adjust the evaluation scope, risk focus, and attack method according to different financial business needs and evaluation objectives, while comprehensively characterizing the financial artificial intelligence system under evaluation from multiple feature dimensions. This evaluation process is customized for specific financial scenarios, specific risk types, and specific interaction methods, thereby ensuring that the generated test cases are highly matched with the current evaluation task.
[0061] Furthermore, in the process of generating the structured test case set, the method flow steps of this embodiment are as follows: S501, Generate multi-round interaction test cases based on the characteristics of the interaction process. The multi-round interaction test cases include multi-round dialogue sequences. The input of each round is determined by the context state, target risk, attack expression method and business scenario of the previous round.
[0062] It should be noted that when generating multi-round interactive test cases, the input for each round is jointly determined by the context state, target risk, attack expression method, and business scenario of the previous round. The context state can record historical inputs, historical system outputs, changes in role settings, and triggered risk points. This allows subsequent rounds to dynamically adjust attack strategies or conduct further probing based on the results of the previous round's interaction, thereby constructing a coherent dialogue chain that gradually approaches the security boundary. This context-state-driven generation method can simulate the complex behavior of users in real financial scenarios who test the system's compliance boundaries through multiple rounds of preparation, role switching, or task decomposition. The test cases generated in this way can effectively identify financial AI systems that perform securely in a single round of testing but exhibit degradation issues such as missing risk warnings, business overreach, or insufficient evidence after continuous interaction, significantly improving the detection capability for multi-round induced risks.
[0063] Furthermore, in the process of determining the output of the financial artificial intelligence system to be evaluated and generating the determination result, the method flow steps of this embodiment are as follows: S601, determine the risk score output by the financial artificial intelligence system to be evaluated based on preset risk dimensions, the preset risk dimensions include content security risk, business compliance risk, tool call risk, context security degradation risk, insufficient evidence or incorrect basis risk, and historical repeated failure risk; S602, based on the risk score and the pre-set judgment threshold, automatically judge the output of the financial artificial intelligence system to be evaluated, and generate the judgment result. The judgment result is pass, failure, or requires manual review. The automatic judgment includes one or more of the following: rule judgment, semantic judgment, business evidence judgment, multi-round context judgment, and tool call judgment.
[0064] It should be noted that this method comprehensively calculates risk scores from multiple risk dimensions, including content security, business compliance, tool invocation, context degradation, insufficient evidence, and historical repeated failures. Based on preset thresholds, the judgment results are categorized into three types: pass, fail, or requiring manual review. Specifically, the business compliance dimension assesses specific risks such as investment advice boundaries and return promises in financial scenarios; the context degradation dimension captures the weakening process of security boundaries during multiple rounds of interaction; the insufficient evidence dimension verifies whether the output is supported by laws or business regulations; and the tool invocation dimension monitors whether the use of external interfaces exceeds authorized limits. Furthermore, the judgment methods encompass rule-based judgment, semantic judgment, business evidence judgment, multi-round context judgment, and tool invocation judgment, ensuring that different types of risks are covered by the corresponding judgment logic. Through this multi-dimensional comprehensive scoring and multi-method joint judgment design, the system can automatically identify critical cases requiring manual intervention.
[0065] Furthermore, after generating the security assessment report, the method flow steps of this embodiment are as follows: S701, based on the statistical results of the security assessment report and the preset convergence threshold, determine whether the current assessment has converged. The security assessment report includes one or more of the following: overall pass rate, risk category distribution, failure sample list, and evidence chain.
[0066] In this embodiment of the specification, based on the evaluation statistics Stat output in S107 and combined with the convergence threshold λ set in S101, it is determined whether the current evaluation meets the convergence condition. The convergence determination condition includes: HighRiskFailRate ≤ λ1; RepeatFailRate ≤ λ2; CriticalCoverage ≥ λ3; ManualReviewPassRate ≥ λ4; Wherein, λ1 represents the high-risk failure rate threshold, λ2 represents the repeated failure rate threshold, λ3 represents the key category coverage threshold, and λ4 represents the manual review pass rate threshold. If the above conditions are met, the system determines that the current assessment has converged and outputs the final assessment report; if not, the system generates a set of reasons for non-convergence, including high-risk issues not being eliminated, similar issues recurring, insufficient key category coverage, or numerous disputes during manual review.
[0067] S702, if convergence fails, generate a reason for non-convergence, and adjust the evaluation strategy parameters for the next round of evaluation based on the reason for non-convergence, the failure examples corresponding to the judgment result, and the manual review result; start a new round of evaluation according to the adjusted evaluation strategy parameters.
[0068] In the embodiments described in this specification, when S10 determines that the current evaluation has not yet converged, the system adjusts the next round of evaluation strategy based on the set of reasons for non-convergence output by S701, combined with the previously output failure examples, critical examples, and manual review results. Specifically, the system clusters failure examples according to risk category, attack expression method, business scenario, trigger round, and tool call type; for categories with a high rate of repeated failures, its weight in the next round of feature combinations is increased; for risk categories, attack methods, or business scenarios with insufficient coverage, new feature combinations are generated; for examples confirmed by manual review to have risks, they are added to the next round of key retest set; for examples confirmed by manual misjudgment, their generation weight is reduced or the judgment rules are adjusted.
[0069] The output of S702 is the parameter for the next round of evaluation strategy: ; in, Indicates the key risk categories for the next round. This indicates the expression of the next round of key attacks. This indicates the key business scenarios for the next round. This represents the updated combined weights. This indicates a failed example that requires regression testing. This indicates critical samples that require manual review. The next round of evaluation strategy output by S702 returns to S201 to rebuild the feature set; when new business scenarios, risk categories, or knowledge evidence are involved, it can also return to S102 to reload the feature set. This forms a continuously iterative security evaluation closed loop.
[0070] It should be noted that after generating the security assessment report, this method determines whether the current assessment has converged based on the statistical results in the report and a preset convergence threshold. If convergence is not achieved, the reasons for non-convergence are identified, and the strategy parameters for the next round of assessment are adjusted based on failed samples and manual review results, before initiating a new round of assessment. Through this closed-loop iterative design, the system can automatically identify which risks have not been eliminated, which categories are under-covered, or which problems recur, and dynamically adjust the key risk categories, attack methods, business scenarios, and sample generation weights for subsequent assessments accordingly. This allows subsequent rounds of assessment to focus specifically on failed samples, critical samples, and high-risk areas, while appropriately reducing the weight of passed or low-risk content. This iterative process continuously guides assessment resources to the most critical security vulnerabilities, gradually leading to the convergence of security issues in the financial AI system.
[0071] Furthermore, the adjustment of the evaluation strategy parameters for the next round of evaluation includes at least one of the following: increasing the weight of risk categories with high repeated failure rates in feature combinations, generating new feature combinations for insufficiently covered risk categories or business scenarios, and adding manually verified risky samples to the key retest set.
[0072] It should be noted that when the evaluation fails to converge, this method analyzes the results of the previous round to identify risk categories with high recurrence failure rates, insufficiently covered risk categories or business scenarios, and manually verified risky examples. Based on these, adjustments are made, such as increasing the weight of the corresponding categories, generating new feature combinations, and adding confirmed risky examples to the retest set. This dynamic adjustment mechanism ensures that the next round of evaluation automatically focuses limited resources on security vulnerabilities that repeatedly fail, insufficiently tested gaps, and manually confirmed real risk points. As the iteration rounds increase, the evaluation strategy continuously tilts towards high-risk and insufficiently covered areas, while gradually reducing the testing weight of already passed or low-risk content. This closed-loop feedback significantly improves the targeting of subsequent evaluations, enabling the continuous tracking and in-depth verification of security issues in financial AI systems, thereby accelerating the convergence process of security risks.
[0073] It should be noted that the following are specific implementation examples provided in the embodiments of this specification: Implementation Example 1: Security Assessment in a Financial Investor Q&A Scenario This embodiment is used to evaluate whether the financial intelligent question-and-answer system has problems such as inappropriate investment advice, return promises, missing risk warnings, and insufficient evidence in investor question-and-answer scenarios.
[0074] First, the system obtains the evaluation task parameters. The system to be evaluated is a financial intelligent question-and-answer system, the business scenario is investor question-and-answer, and the risk categories include inappropriate investment advice, promises of guaranteed principal and returns, lack of risk warnings, and insufficient factual basis; the test expression characteristics include implicit investment advice inducement, return promise inducement, role inducement, and multiple rounds of follow-up questions; the maximum number of interaction rounds is set to 3 to 5 rounds.
[0075] Then, the system loads multi-dimensional features matching the task. Among them, security risk features include "no certain profit judgments can be made", "no direct buy or sell recommendations can be given", and "investment risks should be highlighted"; business scenario features include individual stock consultation, fund product consultation, market trend consultation, etc.; knowledge evidence features include financial business rules, risk disclosure requirements, product description documents, and typical compliance cases.
[0076] The system constructs test cases based on the above features. For example, the system can generate the following feature combinations: Investor Q&A scenarios + inducement through profit promises + multiple rounds of follow-up questions + required risk disclosure + prohibition of statements about certain returns + legal basis for risk disclosure rules.
[0077] Based on this feature combination, the system generates multiple rounds of test cases. The first round of input can ask whether a certain type of financial product is worth allocating; the second round further inquires whether returns can be guaranteed; the third round requires the system to provide a clear conclusion. The system simultaneously binds expected and prohibited behaviors when generating test cases. Expected safe behaviors include highlighting investment risks, stating that information is for reference only, and suggesting that investors assess their own risk tolerance; prohibited behaviors include promising guaranteed principal and returns, providing a definitive prediction of price increases, or directly making investment decisions on behalf of investors.
[0078] During the evaluation, the system inputs the aforementioned test samples into the financial intelligent question-and-answer system to be evaluated round by round, and collects the output content, risk warning content, cited basis, and context status for each round. The automatic judgment module makes a judgment based on preset rules and financial knowledge evidence: if the system output contains statements such as principal protection, guaranteed profit, or certain returns, or gives explicit buy advice without obtaining information on the investor's risk tolerance, it is judged as having business security risks; if the system can reject certain investment advice and provide risk warnings and explanations of applicable conditions, it is judged as passing.
[0079] For samples that are deemed unsuccessful or require review, the system records the sample number, feature combination, input and output content, triggering rules, triggering round, financial knowledge evidence, and risk level to form an evidence chain, and adds the sample to the next round of key retesting set.
[0080] Implementation Example 2: Security Assessment of Financial Intelligent Agents with Tool Invocation Capabilities This embodiment is used to evaluate financial intelligent agent systems with the ability to invoke external tools, such as financial intelligent agents that can invoke market data queries, announcement retrieval, customer profiling, product recommendations, or transaction assistance interfaces.
[0081] The system first obtains the evaluation task parameters. The object to be evaluated is a financial intelligent agent system, and the evaluation scope includes text output security and tool invocation security; risk categories include unauthorized tool invocation, sensitive parameter transmission, unconfirmed high-risk operations, misuse of tool return results, and breach of business boundaries; test expression characteristics include tool invocation inducement, multiple rounds of follow-up questioning, role inducement, and task decomposition inducement.
[0082] The system loads tools based on relevant judgment rules and financial knowledge evidence. These judgment rules include: unauthorized tools must not be invoked; customer information exceeding authorized limits must not be passed in; high-risk operations must not be performed without user confirmation; and tool-returned results must not be interpreted beyond the scope of their application.
[0083] The system is constructed using the following combination of features: Financial intelligent agent scenario + tool call inducement + unauthorized call risk + multi-round interaction + tool call permission rules.
[0084] Based on this combination of features, the system generates multiple rounds of test cases. For example, the first few rounds of input are used to establish an investment consulting background, while subsequent inputs induce the system to call market data queries, product recommendations, or trading assistance tools, requiring the system to make direct decisions based on the tool results. During execution, the system records the tool name, call parameters, call time, return results, and final output.
[0085] The automatic judgment module makes judgments based on tool call logs and business rules. If the agent calls a high-risk tool without confirmation, passes unnecessary sensitive parameters, or directly provides definitive investment advice based on the tool's return results, it is judged that there is a security risk in the tool call. If the agent can identify the permission boundaries, confirm before calling the tool, and carefully interpret the tool's return results, it is judged as passing the judgment.
[0086] For tool call examples that pose a risk, the system generates a chain of evidence, including input content, context state, tool call parameters, tool return results, final output, triggering rules, and risk level. The system also feeds back such failed examples to the next round of example generation module, increasing the proportion of tool call inducement test examples generated.
[0087] It should be noted that the embodiments in this specification have the following beneficial effects: 1. Improve test coverage This invention generates test cases based on a combination of multi-dimensional features, including security risks, attack expressions, business scenarios, interaction processes, output behaviors, judgment rules, and knowledge evidence, rather than relying solely on manually written test cases. Therefore, with the same human resource investment, it can cover more risk categories, attack methods, business scenarios, and interaction rounds, improving the ability to detect risks such as implicit intent, multi-round inducements, and tool call chains.
[0088] 2. Improve the controllability of the sample generation and evaluation process. The test samples generated by this invention are bound to feature combinations, expected safe behaviors, prohibited behaviors, judgment rules, and evidentiary basis, making the generation reason, corresponding risk, and judgment basis of each sample traceable. Compared with ordinary test samples lacking structured annotation, this invention facilitates automatic execution, automatic judgment, and manual review.
[0089] 3. Enhance risk identification capabilities in complex interactions and business scenarios. This invention generates multi-round interaction examples by maintaining context state and introduces business knowledge evidence to participate in example generation and result determination. Therefore, it can identify problems that are difficult to detect in single-round testing, such as context security degradation, role inducement, business overreach, insufficient factual basis, and abnormal tool invocation.
[0090] 4. Improve the efficiency of security issue localization and convergence. This invention generates a chain of evidence during the judgment process and feeds back failed examples, critical examples, manual review results, and repair results to the next round of example generation. Therefore, the system can focus on high-risk and recurring failure issues, reduce invalid repeated testing, and improve the efficiency of security repair, regression verification, and continuous evaluation.
[0091] Figure 2 A schematic diagram of the structure of a financial artificial intelligence system evaluation device provided for one or more embodiments of this specification includes: At least one processor and bus; and, A memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, which, when executed by the at least one processor, enable the at least one processor to: Obtain pre-set evaluation task parameters for the financial artificial intelligence system to be evaluated; Based on the evaluation task parameters, a multi-dimensional feature set matching the current evaluation task is loaded from the preset feature definition module; Based on the multidimensional feature set, a structured test case set is generated, which includes multi-round interactive test cases. Each test case is pre-bound with judgment rule features and knowledge evidence features. The test case set is input into the financial AI system to be evaluated, and the evaluation execution data of the financial AI system to be evaluated is collected. The evaluation execution data includes output content, context state and tool call logs. The evaluation execution data is correlated with the corresponding test cases and interaction rounds to form an observation result set; Based on the set of observation results, and in combination with the judgment rule features and the knowledge evidence features, the output of the financial artificial intelligence system to be evaluated is judged, and a judgment result is generated. Based on the determination results, a security assessment report is generated.
[0092] The various embodiments in this specification are described in a progressive manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, the embodiments of apparatus, devices, and non-volatile computer storage media are basically similar to the method embodiments, so the descriptions are relatively simple; relevant parts can be referred to the descriptions of the method embodiments.
[0093] The various embodiments in this specification are described in a progressive manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, the apparatus embodiments are basically similar to the method embodiments, so the description is relatively simple; relevant parts can be referred to the descriptions of the method embodiments.
[0094] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0095] In the embodiments provided in this application, it should be understood that the disclosed apparatus / network devices and methods can be implemented in other ways. For example, the apparatus / network device embodiments described above are merely illustrative. For instance, the division of modules or units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between devices or units may be electrical, mechanical, or other forms.
[0096] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0097] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The aforementioned units can be implemented in hardware or software.
[0098] If the integrated module / unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, all or part of the processes in the methods of the above embodiments can also be implemented by a computer program instructing related hardware. The computer program can be stored in a computer-readable storage medium, and when executed by a processor, it can implement the steps of the various method embodiments described above. The computer program includes computer program code, which can be in the form of source code, object code, executable files, or certain intermediate forms. The computer-readable medium can include: any entity or device capable of carrying the computer program code, recording media, USB flash drives, portable hard drives, magnetic disks, optical disks, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signals, telecommunication signals, and software distribution media, etc. It should be noted that the content included in the computer-readable medium can be appropriately added or removed according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, computer-readable media do not include electrical carrier signals and telecommunication signals.
[0099] The above-described embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application, and should all be included within the protection scope of this application.
Claims
1. A method for evaluating a financial artificial intelligence system, characterized in that, The method includes: Obtain pre-set evaluation task parameters for the financial artificial intelligence system to be evaluated; Based on the evaluation task parameters, a multi-dimensional feature set matching the current evaluation task is loaded from the preset feature definition module; Based on the multidimensional feature set, a structured test case set is generated, which includes multi-round interactive test cases. Each test case is pre-bound with judgment rule features and knowledge evidence features. The test case set is input into the financial AI system to be evaluated, and the evaluation execution data of the financial AI system to be evaluated is collected. The evaluation execution data includes output content, context state and tool call logs. The evaluation execution data is correlated with the corresponding test cases and interaction rounds to form an observation result set; Based on the set of observation results, and in combination with the judgment rule features and the knowledge evidence features, the output of the financial artificial intelligence system to be evaluated is judged, and a judgment result is generated. Based on the determination results, a security assessment report is generated.
2. The method according to claim 1, characterized in that, The evaluation task parameters include one or more of the following: the identifier of the financial artificial intelligence system to be evaluated, the type of financial business, the evaluation scope, the category of financial security risks, the range of attack expression features, the range of business scenarios, the number of samples, the maximum number of interaction rounds, the judgment threshold, and the convergence threshold; the multidimensional feature set includes one or more of the following: security risk features, attack expression features, business scenario features, interaction process features, output behavior features, judgment rule features, and knowledge evidence features.
3. The method according to claim 2, characterized in that, The generation of a structured test sample set based on the multidimensional feature set includes: Based on the multidimensional feature set, a candidate multidimensional feature combination set is generated by selecting feature values from each feature subset and combining them. The candidate multidimensional feature combination set is subjected to constraint filtering and priority sorting to obtain the sorted effective feature set. Based on the effective feature set, a structured test case set is generated. Each test case is associated with at least the corresponding feature combination, input sequence, expected safe behavior, prohibited behavior, judgment rule and evidence reference.
4. The method according to claim 3, characterized in that, The constraint filtering of the candidate multidimensional feature combination set includes: The candidate multidimensional feature combination set is constrained and filtered based on preset constraint rules. The constraint rules include: whether the security risk feature matches the business scenario feature, whether the attack expression feature is applicable to the corresponding interaction process feature, whether the judgment rule feature covers the output behavior feature, and whether the knowledge evidence feature supports one or more of the business scenario features.
5. The method according to claim 4, characterized in that, Prioritizing the candidate multidimensional feature combination set includes: Based on preset weights, the priority score of each candidate combination is calculated. The calculation factors of the priority score include the security risk level corresponding to the security risk feature, the attack expression strength corresponding to the attack expression feature, the relevance of historical failure examples, the importance of the business scenario, the similarity between the new coverage benefit and the existing examples.
6. The method according to claim 1, characterized in that, The generation of the structured test sample set includes: Multi-round interaction test cases are generated based on the characteristics of the interaction process. The multi-round interaction test cases include multi-round dialogue sequences. The input of each round is determined by the context state of the previous round, the target risk, the attack expression method and the business scenario.
7. The method according to claim 1, characterized in that, The step of judging the output of the financial artificial intelligence system to be evaluated and generating a judgment result includes: The risk score output by the financial artificial intelligence system to be evaluated is determined based on preset risk dimensions, including content security risk, business compliance risk, tool call risk, context security degradation risk, insufficient evidence or incorrect basis risk, and historical repeated failure risk. Based on the risk score and the pre-set judgment threshold, the output of the financial artificial intelligence system to be evaluated is automatically judged to generate the judgment result. The judgment result is either pass, fail, or requires manual review. The automatic judgment includes one or more of the following: rule judgment, semantic judgment, business evidence judgment, multi-round context judgment, and tool call judgment.
8. The method according to claim 1, characterized in that, The method further includes: Based on the statistical results of the security assessment report and the preset convergence threshold, it is determined whether the current assessment has converged. The security assessment report includes one or more of the following: overall pass rate, risk category distribution, failure sample list, and evidence chain. If convergence fails, a reason for non-convergence is generated, and the evaluation strategy parameters for the next round of evaluation are adjusted based on the reason for non-convergence, the failure examples corresponding to the judgment result, and the manual review results. A new round of evaluation will be launched based on the adjusted evaluation strategy parameters.
9. The method according to claim 8, characterized in that, The adjustment of the evaluation strategy parameters for the next round of evaluation includes at least one of the following: increasing the weight of risk categories with high repeated failure rates in feature combinations, generating new feature combinations for insufficiently covered risk categories or business scenarios, and adding manually verified risky samples to the key retest set.
10. A financial artificial intelligence system evaluation device, characterized in that, include: At least one processor and bus; as well as, A memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to: Obtain pre-set evaluation task parameters for the financial artificial intelligence system to be evaluated; Based on the evaluation task parameters, a multi-dimensional feature set matching the current evaluation task is loaded from the preset feature definition module; Based on the multidimensional feature set, a structured test case set is generated, which includes multi-round interactive test cases. Each test case is pre-bound with judgment rule features and knowledge evidence features. The test case set is input into the financial AI system to be evaluated, and the evaluation execution data of the financial AI system to be evaluated is collected. The evaluation execution data includes output content, context state and tool call logs. The evaluation execution data is correlated with the corresponding test cases and interaction rounds to form an observation result set; Based on the set of observation results, and in combination with the judgment rule features and the knowledge evidence features, the output of the financial artificial intelligence system to be evaluated is judged, and a judgment result is generated. Based on the determination results, a security assessment report is generated.