Double-track evaluation driving optimization processing method, device, equipment and medium
By constructing a dual-track evaluation system and a multi-role intelligent agent simulation environment, we analyze training data and real business feedback indicators, generate training adjustment information, and drive the retraining of the target model. This solves the problem of the disconnect between evaluation and model performance improvement, and enhances the compliance and reliability of the model in actual business.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- PING AN TECH (SHENZHEN) CO LTD
- Filing Date
- 2026-02-06
- Publication Date
- 2026-05-12
AI Technical Summary
Existing evaluation methods cannot be continuously and in a closed loop used to optimize training data, resulting in a disconnect between evaluation and model performance improvement. Furthermore, the evaluation results are difficult to reflect the model's real business capabilities and cannot support the continuous optimization needs for real business applications.
A dual-track evaluation system is constructed, which includes generating an evaluation set and a business scenario evaluation set. Optimized training data is generated by analyzing training data, and the target model is analyzed using a multi-role intelligent agent simulation environment. Base capability indicators and interactive task execution indicators are generated, real business feedback indicators are obtained, and the attribution link between the training evaluation indicator set and the real business feedback indicators is determined. Training adjustment information is generated based on the attribution link to drive the retraining of the target model.
It enables the linkage between evaluation results and business feedback, which can identify and correct gaps in model capabilities, and improve the compliance, reliability and application effectiveness of the model in actual business.
Smart Images

Figure CN122020091A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of model building technology, and in particular to a dual-track evaluation-driven optimization processing method, apparatus, equipment, and medium. Background Technology
[0002] As large language models are gradually introduced into financial business scenarios such as banking, securities, insurance, and investment consulting, the industry has begun to explore their application in tasks such as customer service Q&A, business guidance, marketing outreach, risk analysis, and compliance assistance. To ensure the reliable operation of these models in environments with high regulatory requirements, the industry generally adopts some form of model evaluation to verify knowledge coverage and task responsiveness. However, most of these evaluations follow the testing methods of general models and do not fully align with the system requirements of financial scenarios.
[0003] Existing evaluation methods generally suffer from a disconnect between training and evaluation. Evaluation often occurs after model training, serving only as a standard for verifying effectiveness, and is difficult to use to help optimize training data or training strategies. At the same time, the existing evaluation dimensions in the industry are relatively narrow, focusing mostly on the accuracy of knowledge-based question answering or single business indicators, lacking a systematic evaluation of models from multiple capability dimensions such as risk control, compliance constraints, business reasoning, and multi-turn interactions. As a result, the evaluation results fail to reflect the model's true business capabilities.
[0004] Furthermore, existing evaluation methods typically employ fixed question sets and static scoring mechanisms, which still differ significantly from real-world business interactions. For example, when faced with real user input, complex decision-making processes, and regulatory constraints, model performance often deviates from evaluation scores. Simultaneously, there is a lack of clear correlation between evaluation results and operational metrics, making it difficult to explain the relationship between model capabilities and conversion rates, human intervention rates, or compliance compliance performance, and failing to provide effective evidence for training adjustments. These issues collectively result in evaluations failing to cover the entire model lifecycle and struggling to support continuous optimization needs for real-world business applications. Summary of the Invention
[0005] The main objective of this invention is to provide a dual-track evaluation-driven optimization processing method, apparatus, device, and storage medium, aiming to solve the technical problem that the existing technology cannot continuously and in a closed loop use the evaluation results in reverse for training data optimization, model capability diagnosis, and retraining adjustment, resulting in a disconnect between evaluation and model performance improvement.
[0006] To achieve the above objectives, the present invention provides a dual-track evaluation-driven optimization processing method, comprising: Construct a dual-track evaluation system that includes generated evaluation sets and business scenario evaluation sets; The training data is analyzed using the dual-track evaluation system to form a data health analysis result. Based on the data health analysis result, a data optimization strategy is executed to generate optimized training data. The target model is then established using the optimized training data. The target model is analyzed using the dual-track evaluation system and multi-role intelligent agent simulation environment to generate base capability indicators and interactive task execution indicators. Analyze the balance between the compliance and security performance and business availability performance of the target model during the reinforcement learning phase, and generate alignment analysis results; Obtain real business feedback metrics from the return flow, and determine the attribution link between the training evaluation metric set and the real business feedback metrics, wherein the training evaluation metric set consists of the base capability metrics, the interaction task execution metrics, and the alignment analysis results; Based on the attribution link, the capability gap of the target model is determined, and the capability gap is used to generate training adjustment information including parameter configuration, data ratio and reward weight. The training adjustment information and the optimized training data are used to drive the target model to be retrained to obtain the optimized target model. The optimized target model is used to process business input data and generate business processing results.
[0007] Furthermore, to achieve the above objectives, the present invention provides a dual-track evaluation-driven optimization processing device, comprising: The dual-track evaluation construction module is used to build a dual-track evaluation system that includes a generated evaluation set and a business scenario evaluation set; The training data optimization module is used to analyze the training data using the dual-track evaluation system to form a data health analysis result, execute a data optimization strategy based on the data health analysis result to generate optimized training data, and use the optimized training data to build a target model. The model capability evaluation module is used to analyze the target model using the dual-track evaluation system and the multi-role intelligent agent simulation environment, and to generate base capability indicators and interactive task execution indicators. The alignment analysis module is used to analyze the balance between the compliance and security performance and business availability performance of the target model during the reinforcement learning phase, and generate alignment analysis results. The business feedback attribution module is used to obtain the real business feedback indicators of the return flow and determine the attribution link between the training evaluation indicator set and the real business feedback indicators, wherein the training evaluation indicator set consists of the base capability indicators, the interaction task execution indicators and the alignment analysis results. The model retraining module is used to determine the capability gap of the target model based on the attribution link, and use the capability gap to generate training adjustment information including parameter configuration, data ratio and reward weight. The training adjustment information and the optimized training data are used to drive the retraining of the target model to obtain the optimized target model. The business processing execution module is used to process business input data using the optimized target model and generate business processing results.
[0008] Furthermore, to achieve the above objectives, the present invention also provides a computer device, the computer device including a memory, a processor, and a dual-track evaluation-driven optimization processing program stored in the memory and executable on the processor, wherein when the dual-track evaluation-driven optimization processing program is executed by the processor, it implements the steps of the dual-track evaluation-driven optimization processing method as described above.
[0009] Furthermore, to achieve the above objectives, the present invention also provides a computer-readable storage medium storing a dual-track evaluation driver optimization processing program, wherein when the dual-track evaluation driver optimization processing program is executed by a processor, it implements the steps of the dual-track evaluation driver optimization processing method as described above.
[0010] Beneficial Effects: This invention relates to the field of model building technology and can be applied to business scenarios such as fintech. It discloses a dual-track evaluation-driven optimization processing method, apparatus, equipment, and medium, comprising: constructing a dual-track evaluation system and analyzing training data to generate optimized training data, and establishing a target model; using the dual-track evaluation system and multi-role intelligent agent simulation to generate capability indicators, and combining reinforcement learning performance to generate alignment analysis results; obtaining real business feedback indicators to form a training evaluation indicator set and determining attribution links; generating training adjustment information based on the attribution links and driving the target model to retrain to obtain an optimized target model, used to process business input data and output business processing results. This invention achieves a training closed loop by linking evaluation results with business feedback, enabling model capability gaps to be identified and corrected, and improving the compliance, reliability, and application effectiveness of the model in actual business. Attached Figure Description
[0011] The present invention will be further described below with reference to the accompanying drawings and embodiments. In the accompanying drawings: Figure 1 This is a schematic diagram of an application environment for the dual-track evaluation-driven optimization processing method in one embodiment of the present invention; Figure 2 This is a flowchart illustrating an embodiment of the dual-track evaluation-driven optimization processing method of the present invention; Figure 3 This is a schematic diagram of the functional modules of a preferred embodiment of the dual-track evaluation-driven optimization processing device of the present invention; Figure 4 This is a schematic diagram of the structure of a computer device according to an embodiment of the present invention; Figure 5 This is another structural schematic diagram of a computer device according to one embodiment of the present invention. Detailed Implementation
[0012] It should be understood that the specific embodiments described herein are for illustrative purposes only and are not intended to limit the scope of the invention.
[0013] The dual-track evaluation-driven optimization processing method provided in this embodiment of the invention can be applied to, for example... Figure 1 In this application environment, the client communicates with the server via a network. The server can construct a dual-track evaluation system through the client and analyze training data to generate optimized training data and establish a target model; it uses the dual-track evaluation system and multi-role intelligent agent simulation to generate capability indicators, and combines reinforcement learning performance to generate alignment analysis results; it obtains real business feedback indicators to form a training evaluation indicator set and determines the attribution link; based on the attribution link, it generates training adjustment information and drives the target model to retrain, obtaining an optimized target model for processing business input data and outputting business processing results. This invention achieves a training closed loop by linking evaluation results with business feedback, enabling model capability gaps to be identified and corrected, and improving the model's compliance, reliability, and application effectiveness in actual business. The client can be, but is not limited to, various personal computers, laptops, smartphones, tablets, and portable wearable devices. The server can be implemented using a standalone server or a server cluster consisting of multiple servers. The invention will be described in detail below through specific embodiments.
[0014] Please see Figure 2 , Figure 2 This is a flowchart illustrating an embodiment of the dual-track evaluation-driven optimization processing method provided by the present invention. It should be noted that although the logical order is shown in the flowchart, in some cases, the steps shown or described may be performed in a different order than that shown here.
[0015] like Figure 2 As shown, the dual-track evaluation-driven optimization processing method proposed in this invention includes the following steps: S10, construct a dual-track evaluation system that includes generated evaluation sets and business scenario evaluation sets; In this embodiment, the dual-track evaluation system relies on two data sources: a generated evaluation set and a business scenario evaluation set. The system revolves around knowledge extraction, data generation, data filtering, and data structure transformation. The knowledge graph is parsed to extract entity names, conditional logic, and constraint entries from structured nodes and semantic relationships, enabling business rules to form mechanically computable test-driven elements. Regulatory text libraries and product documentation provide legal obligations, business rules, and exception clauses. Through text parsing, restrictive conditions, triggering conditions, and business process information are extracted and combined with the rule entries generated from the knowledge graph to form a common knowledge domain set. Based on this set, structured evaluation data is generated. Through template filling, logical combination, and rule verification, entries covering definition judgment, rule reasoning, and condition matching are compiled into a generated evaluation set, ensuring that the model capability verification has complete rule coverage.
[0016] In contrast, the business scenario evaluation set originates from real business interaction data. Privacy processing is used to eliminate fields containing personal characteristics while maintaining semantic structural continuity, ensuring secure data use. The filtering process focuses on identifying interaction records containing high-risk or complex logical tags. These tags consist of common compliance errors, business misjudgments, or multi-step logical derivations found in model interactions. Data reconstruction standardizes these records through sentence formatting, condition completion, and step streamlining, transforming the original dialogues into evaluation tasks that can be directly input. Finally, the generated evaluation set and the business scenario evaluation set are merged through format standardization, field alignment, and classification labeling, forming a parallel data system that creates complementary verification dimensions between rule testing and real-world scenario simulation.
[0017] This embodiment uses a dual-track system to simultaneously absorb rule knowledge and real interaction content, enabling model verification to cover both explicit rule requirements and reflect real business language and logical characteristics. It establishes a data correspondence between artificial knowledge structures and natural business scenarios, thereby improving the breadth of evaluation coverage and the effectiveness of verification.
[0018] S20, the training data is analyzed using the dual-track evaluation system to form a data health analysis result, a data optimization strategy is executed based on the data health analysis result to generate optimized training data, and a target model is established using the optimized training data; In this embodiment, a dual-track evaluation system is used to analyze training data, comprising two parallel analysis processes: a knowledge structure perspective and a business corpus perspective. After entering the processing chain, the data first undergoes entity recognition, semantic segmentation, and tag extraction, breaking down text, fields, or dialogue sequences into labelable elements, enabling the content to establish a mapping relationship with knowledge items or scene features. An evaluation set is generated to verify the coverage between the training data and business rules. By matching entities, conditional relationships, and logical inference fragments, it identifies missing topics, weakly related domains, or knowledge gaps, forming structural coverage difference information. The business scenario evaluation set is responsible for characterizing the fit between the data and real semantic scenarios. By analyzing contextual coherence, logical completeness, and risk performance, it identifies semantically ambiguous fragments, reasoning breakpoints, or misleading expressions, generating expression quality difference information.
[0019] The data health analysis results consist of a combination of coverage difference information and expression quality difference information. A set of structural indicators is formed through counting, weight integration, or structural proportion calculation to reflect the distribution of training data across three dimensions: knowledge coverage, semantic performance, and quality credibility. The data optimization strategy relies on this indicator set. It involves supplementing data to fill topic gaps, removing low-trust corpora through noise elimination, and extending the inference chain or supplementing context through semantic enhancement to ensure the input content has a more robust expression structure. The optimized training data then enters the training system, undergoing vector encoding, network parameter initialization, and weight updates to execute the learning process, outputting a model with the target capabilities.
[0020] This embodiment utilizes a dual-track perspective to conduct training data quality analysis and automatically generate optimized data, enabling the learning content to converge between knowledge coverage and realistic expression. At the same time, it reduces the impact of noise, improves the reliability of training sources, and provides more stable capability support for the trained model.
[0021] S30, using the dual-track evaluation system and multi-role intelligent agent simulation environment to analyze the target model and generate base capability indicators and interactive task execution indicators; In this embodiment, a dual-track evaluation system and a multi-role intelligent agent simulation environment are used to analyze the target model, which includes two complementary measurement paths. The dual-track evaluation system provides static input materials, inputting text, rule fragments, or inference variables extracted from the generated evaluation set and the business scenario evaluation set into the model. It generates basic capability representation data by parsing the output judgments, explanations, inference chains, and answer consistency. These results reflect the model's performance in concept understanding, factual association, logical inference, and knowledge coherence. The multi-role intelligent agent simulation environment constructs an interaction loop, simulating users, adversaries, rule checkers, and outcome adjudicators that may appear in real business processes. After entering the dialogue loop, the model receives continuous inquiries, challenges, or interaction instructions from simulated roles. The simulation environment records the behavioral trajectory, including the coherence of the model's answers, the degree of logical compliance, and response deviations.
[0022] The analysis process extracts static representation data and dynamic trajectory data, structuring the output into numerical values, labels, or state sequences. Then, it generates two types of metrics based on task type and interaction goals. The foundational capability metrics, derived from the static input output, reflect the model's ability to correctly understand input semantics, maintain an uninterrupted logical chain, and complete knowledge matching. The interaction task execution metrics, generated from dynamic trajectory analysis, measure the task completion rate, response consistency, behavioral strategy robustness, and risk exposure probability in continuous dialogue. These two types of metrics jointly characterize the model's actual performance in reasoning-driven capabilities and business interaction capabilities.
[0023] This embodiment combines static evaluation and interactive simulation to output the model's performance in two dimensions: understanding and task execution. This results in a more business-explanatory capability profile and provides quantifiable evidence for subsequent capability improvements.
[0024] S40, Analyze the balance between the compliance and security performance and business availability performance of the target model in the reinforcement learning phase, and generate alignment analysis results; In this embodiment, analyzing the balance between the compliance and security performance and business availability performance of the target model during the reinforcement learning phase relies on decomposing and evaluating the model output generated during training. In each iteration, the training engine collects the model's responses to input instructions, including structured elements such as completion level, rejection cases, content bias, violation statements, and inference consistency, and records them as a time-series dataset. The system extracts response features related to security requirements from this dataset, such as whether sensitive guidelines are violated, whether regulatory boundaries are ignored, and whether outputs lacking factual support are generated; simultaneously, it extracts behavioral features related to actual business effectiveness, such as task completion probability, information sufficiency, output coverage of needs, and assistance efficiency.
[0025] Compliance and security performance and business availability performance are quantified using a separate indicator system. The former can be derived from the number of violation triggers, the approval rate of reviews, or the success rate of risk screening, while the latter can be derived from the completion rate, response success rate, or the progress rate of tasks. To measure the trade-off between these two types of performance, the analysis engine constructs a computational structure that includes positive reward and negative reduction factors. This structure is used to compare how much conservative behavior the model introduces to improve security, or how much self-restraint it sacrifices to improve business effectiveness. By analyzing the trend of the indicators during training rounds, it can be inferred whether the model is gradually biased towards refusing to answer, whether it is reducing the reasonable output space, and whether convergent behavior collapse or strategy degradation has occurred. Finally, an alignment analysis result reflecting the dynamic balance between the two is generated and used as a basis for subsequent adjustments.
[0026] This embodiment simultaneously quantifies security constraint capabilities and business output capabilities and reveals the relationship between the two, enabling timely detection of capability shifts during the training phase. This provides a clear fulcrum for subsequent improvements and prevents the model from becoming overly contracted or exhibiting behavioral imbalances.
[0027] S50, obtain the real business feedback indicators of the return flow, and determine the attribution link between the training evaluation indicator set and the real business feedback indicators, wherein the training evaluation indicator set consists of the base capability indicators, the interactive task execution indicators and the alignment analysis results. In this embodiment, obtaining real business feedback metrics establishes an external data source for the training closed loop. Performance metrics after model execution are obtained through data channels generated by the online-deployed model in the actual business environment. These feedback metrics originate from statistical records of user interaction platforms, tool systems, or process systems, including user satisfaction, interaction success rate, processing progress, and human intervention triggers. Each data point is stored in a measurable format. To ensure consistency with the training and evaluation system, the system performs time alignment, format conversion, noise filtering, and anomaly removal on the feedback metrics after collection to avoid offsets caused by sampling errors.
[0028] The training evaluation metric set consists of foundational capability metrics, interactive task execution metrics, and alignment analysis results. All three types of metrics are calculated internally during the training phase. Foundational capability metrics reflect basic comprehension, reasoning ability, and knowledge consistency; interactive task execution metrics reflect multi-turn interaction performance, execution completeness, and the likelihood of achieving interaction goals; alignment analysis results present the trade-offs between security and business considerations, describing whether the model exhibits behavioral deviations. To establish the correlation between the two types of information, it is necessary to extract time overlap windows from the two sets of data and segment them according to business type, request source, and usage path, ensuring that the compared objects fall within the same semantic space.
[0029] The attribution chain consists of the mapping relationship between the training evaluation metric set and the real business feedback metrics. Potential causal chains are determined by identifying the degree of mutual influence between the metrics. The system uses statistical modeling, causal comparison, or weighted sensitivity calculation to analyze the contribution of different training evaluation metrics to changes in business performance, and maps the contribution to an interpretable sequence, forming a sequence chain representing how training metrics affect real business output. This chain reflects the path of action of internal metrics formed during training in real business performance, providing a foundation for subsequent identification of the sources of weaknesses.
[0030] This embodiment introduces real business feedback metrics and establishes a mapping relationship with training and evaluation metrics, which can reveal the source structure of the model's online performance, making the metrics generated during the training phase interpretable and providing basic support for identifying behavioral deviations and locating capability deficiencies.
[0031] S60, determine the capability gap of the target model based on the attribution link, and use the capability gap to generate training adjustment information including parameter configuration, data ratio and reward weight, and use the training adjustment information and the optimized training data to drive the target model to retrain, so as to obtain the optimized target model; In this embodiment, the attribution link provides a mapping between training metrics and business performance. By analyzing the link content, the trajectory of training dimension changes affecting model behavior is extracted, and the magnitude of change is calculated to form a capability gap. This gap represents the range between the current model capability and the ideal performance, defined as a capability gap. Capability gaps can manifest as weak knowledge understanding, insufficient interaction advancement, or strategy selection deviations. Different capability gaps are represented by quantitative identifiers or difference structures.
[0032] The parameter configuration, derived from the capability gap, includes control variables for the learning process, determining the parameter update rate, optimization rules, and local adjustments to the network structure. If the gap reflects fundamental comprehension issues, the parameter configuration focuses on underlying updates; if the gap involves behavioral biases, the parameter configuration focuses on policy parameters. Data allocation represents the sampling ratio of different training sample types, guiding the model to accumulate experience in the target capability dimension by adjusting the proportion of sampling sources. Reward weights are used to adjust behavioral preferences, redirecting output tendencies by increasing rewards or penalties, thereby enhancing the representation of the target capability during reinforcement training.
[0033] The training adjustment information consists of parameter configuration, data allocation, and reward weights, which are encapsulated into a unified execution instruction. This information is input into the training execution module, driving the training component to read the optimized training data and select a sample set that meets the data allocation requirements. The training execution process iteratively updates under the guidance of capability gaps, adjusting parameters through gradient flow to gradually move the model towards the target range. After the final update is completed, the model weights reach a new convergence point, and the optimized target model is output from the training execution module.
[0034] This embodiment drives training execution by addressing capability gaps, enabling parameter configuration, data allocation, and reward weights to focus on strengthening the weak capabilities of the model. This allows the model's capabilities to be adjusted to meet business needs and shortens the performance improvement cycle.
[0035] S70, the optimized target model is used to process the business input data and generate the business processing result.
[0036] In this embodiment, business input data is received from an external business system, parsed through an interface, and transformed from text, structured fields, or sequence information into an expressive format that the model can process. The parsing process identifies the request content and potential intent, such as consultation, judgment, or information completion, and extracts key fields to prevent irrelevant content from affecting inference execution.
[0037] The optimized target model receives the processed input and performs inference based on existing parameters and knowledge, generating inference results through weight propagation and feature association. Inference may include natural language generation, information extraction, recommendation decisions, or classification judgments, all completed by the model based on the input.
[0038] To avoid generating distorted or misleading information, the output is checked by a content validation module, including consistency rules, sensitive word scanning, or logical checks, to identify response content that does not meet constraints. Output that passes validation is organized according to the format requirements of the calling system, encapsulating text, tags, or structure fields into standardized business processing results before returning them to the external system.
[0039] Business input data can originate from customer service dialogue systems, internal work order systems, or business process nodes. Content parsing can employ keyword extraction or classifier-based intent recognition. Model inference can use a single-round inference approach or a multi-round inference approach to generate more complete results. The verification module can use static rule matching or model-assisted checking to filter abnormal or risky content. Result encapsulation can use a text response format or a structured markup format.
[0040] This implementation utilizes an optimized target model to process actual business inputs, enabling the model's capabilities to directly impact real-world scenarios, providing structured or textual outputs, and reducing the risk of error propagation through parsing, reasoning, verification, and encapsulation processes, thereby improving the reliability of business outputs.
[0041] In one embodiment, step S10 above includes: S101: Analyze the knowledge graph, regulatory text library and product description documents to extract core entities and constraint logic to form a basic knowledge domain set; S102, Generate evaluation data for stability testing based on the set of basic knowledge domains, and compile the evaluation data to form a generated evaluation set; S103, acquire historical business interaction data, use the privacy processing module to remove sensitive information from the historical business interaction data and retain the business context logic, and obtain desensitized business interaction data; S104, using preset business scenario screening conditions, feature extraction is performed on the desensitized business interaction data to identify records containing high-risk tags and complex logic tags, and the records containing high-risk tags and complex logic tags are reconstructed to form the business scenario evaluation set. S105, the generated evaluation set is integrated with the business scenario evaluation set to construct a dual-track evaluation system.
[0042] In this embodiment, the dual-track evaluation system is built upon two sources and two types of datasets: one for knowledge and rule coverage, and the other for real-world business scenario reconstruction. When parsing the knowledge graph, regulatory text library, and product description documents, the system first reads the node and edge structure from the knowledge graph, extracting nodes representing financial products, customer attributes, business actions, and risk events as candidate entity sets. Then, based on edge types, it extracts constraints such as limit limits, suitability conditions, prohibited situations, and liability division. The regulatory text library stores clause-type texts. The system segments these texts into sentences and clauses, identifying segments containing numerical ranges, obligations, responsibilities, and prohibited behaviors. It further uses template matching or sequence labeling models to identify the regulatory agency name, clause number, triggering conditions, and consequences. Product description documents typically consist of explanatory sections, risk disclosures, and fee structures. In these documents, the system identifies key fields such as product name, underwriting conditions, fee structure, and cancellation rules, aligning them with similar concepts in the knowledge graph and regulatory clauses. Through the above multi-source parsing process, entity names in various documents are standardized, redundant aliases are removed, and they are unified into standardized terms. At the same time, the restrictive descriptions, triggering conditions, and combination relationships surrounding these terms are extracted into logical expressions, forming a set of basic knowledge domains that can cover product rules, compliance constraints, and risk boundaries in financial business.
[0043] After the basic knowledge domain set is constructed, the system generates a dedicated dataset for model stability testing. Utilizing each entity and its constraint logic within this set, question-answer pairs, reasoning chains, or judgment tasks are generated through template combination or structured sampling. For example, definition verification questions are generated based on a product definition, boundary value test questions are generated based on quota constraints, and multi-premise reasoning questions are generated based on multi-condition triggering clauses. Parameter control can be introduced during the generation process, such as controlling the length of single-round questions, the level of logical nesting, and the number of entities involved, to examine the model's stability and consistency at different difficulty levels. The generated entries undergo deduplication and coverage checks, compiling data entries covering different product lines, different regulatory clauses, and different risk categories to form a generated evaluation set focused on basic capabilities and rule understanding.
[0044] The other dataset comes from historical business interaction data. The system batch-reads raw interaction content from customer service records, online consultation dialogues, business processing dialogues, and internal review communication records, and organizes the customer input, system responses, and supplementary explanations from human customer service into a unified format. Before entering the evaluation and construction chain, the privacy processing module performs de-identification processing on these records, identifying sensitive information fields such as names, ID numbers, contact numbers, bank card numbers, and specific addresses, and deleting or replacing them through masking, generalization, or deletion, while maintaining the chronological order, the order of business actions, and the correlation between important business fields, so that the processed interaction still retains the complete business context logic. The privacy-processed interaction can map the semantic chain of a consultation or transaction from initiation, clarification, decision-making to conclusion, providing raw material for subsequent scenario screening.
[0045] Based on anonymized business interaction data, the system introduces pre-defined business scenario filtering conditions to extract features from the data. These filtering conditions can be pre-configured by domain experts or generated by the strategy system based on historical risk event statistics. The conditions include risk level identification, business type classification, whether cross-selling is involved, whether complex clause interpretations are involved, and whether regulatory keywords are triggered. The system calculates a feature vector for each interaction based on these conditions and identifies interaction records containing high-risk and complex logic tags through classification or retrieval. High-risk tags can represent situations such as policy cancellation disputes, misleading sales, and inappropriate risk matching; complex logic tags can represent situations such as cross-product portfolios, cross-institutional information citations, and multi-round negation reasoning. The selected records are not used directly but reconstructed according to evaluation needs. Key rounds, key questions, and key responses are extracted, and redundant small talk or irrelevant content is trimmed when necessary to generate business scenario entries that centrally reflect risk exposure points and complex reasoning chains, while maintaining the original semantic order and business causal relationships. Through this reconstruction process, a business scenario evaluation set is formed specifically for testing the model's performance in complex real-world business scenarios.
[0046] After both the generated evaluation set and the business scenario evaluation set are constructed, the system integrates the two datasets to form a dual-track evaluation system. The integration process first establishes unified identifiers and metadata descriptions for both types of evaluation items, including source type, business domain involved, type of regulatory clauses involved, risk level, and logical complexity level, to facilitate selective invocation for different training stages or model versions. Then, a unified input / output format is designed for both types of data, allowing rule understanding items in the generated evaluation set and dialogue or processing items in the business scenario evaluation set to be submitted to the model for testing using the same API. The integrated structure supports the construction of combined evaluation tasks using track type, risk level, and complexity level as filtering conditions. This aligns the evaluation results of basic rule understanding capabilities with the interaction performance in complex scenarios within a unified coordinate system, providing a stable and reusable evaluation foundation for subsequent data health analysis, capability diagnosis, and training optimization.
[0047] This embodiment analyzes knowledge graphs, regulatory text libraries, and product documentation to form a set of foundational knowledge domains covering rules and constraints. This transforms scattered compliance clauses and product logic in financial transactions into structured testing resources. Stability evaluation data is generated from this set, forming a generative evaluation set focused on fundamental capabilities. This allows for the quantitative measurement of the model's rule understanding ability under controlled conditions. By performing privacy processing, scenario filtering, and record reconstruction on historical business interaction data, high-risk and complex logical scenarios are extracted, forming a business scenario evaluation set closely resembling real-world business. Finally, these two types of evaluation data are integrated into a dual-track evaluation system, linking rule understanding capability assessment and real-world scenario performance assessment within a unified framework. This provides a more comprehensive and traceable evaluation foundation for subsequent training data analysis and capability optimization, thereby enhancing the explanatory power and guiding value of the evaluation results for real-world business performance.
[0048] In one embodiment, step S20 above includes: S201, Obtain training data, and use the knowledge graph and compliance detection model in the dual-track evaluation system to perform entity linking and rule scanning on the training data to determine the knowledge coverage and compliance sample ratio; S202, perform fact consistency verification on the training data to determine the proportion of noisy samples, and form a data health analysis result based on the knowledge coverage, the proportion of compliant samples and the proportion of noisy samples; S203, compare the data health analysis results with preset thresholds, and determine a data optimization strategy that includes data supplementation, noise removal and compliance enhancement based on the comparison results; S204, The data optimization strategy is executed to clean and enhance the training data to generate optimized training data; S205, the optimized training data is input into the basic network architecture to be trained for parameter initialization and pre-training to establish the target model.
[0049] In this embodiment, the training data is obtained from multi-source corpora accumulated in financial business, typically including structured or semi-structured content such as customer service dialogue records, business processing records, risk control review texts, investment research report summaries, and historical Q&A pairs. To facilitate unified processing later, this data is standardized before entering the analysis process, aligning log fields exported from different systems to a unified field set. For example, fields such as request text, reply text, business tags, timestamps, and channel identifiers are used uniformly. Basic filtering is performed on abnormal encoding, garbled content, and empty fields to obtain a training data set with a consistent structure.
[0050] In the dual-track evaluation system, the knowledge graph serves as a semantic reference, covering various entities such as financial products, customer attributes, business actions, risk categories, and regulatory clauses, as well as the adaptation, constraint, and deductive relationships between entities. Entity linking operations, through word segmentation, named entity recognition, and phrase matching, align natural language fragments in the training data with standard entity entries in the knowledge graph. For example, "whole life insurance" and "participating insurance" are mapped to a unified product node, and "high-risk customer" is mapped to a customer attribute node. This mapping allows for the attachment of a set of graph entity labels to each training sample, facilitating the statistical analysis of the coverage of different entities in the training data.
[0051] The compliance detection model undertakes the task of rule scanning, with rules derived from regulatory provisions, internal compliance policies, and patterns extracted from historical violation cases. After entity linking is completed, the training data enters the rule scanning process. During scanning, fragments potentially involving non-compliant expressions are identified based on entity labels, syntactic structure, and keyword patterns, with a focus on checking statements such as profit promises, product suitability suggestions, and risk disclosures. Each sample is categorized as compliant or non-compliant based on whether it conforms to the rules. By accumulating statistics across the entire dataset, two metrics are obtained: knowledge coverage and the proportion of compliant samples. Knowledge coverage reflects the number of graph entities and relationships contained in the training data, while the proportion of compliant samples reflects the percentage of samples consistent with regulatory and internal rules. Both metrics jointly characterize the basic quality of the training data in terms of knowledge completeness and compliance.
[0052] Fact consistency verification targets training data containing references to external facts, such as interest rate figures, term descriptions, fee ratios, and key points of product terms. During verification, knowledge graphs, authoritative data source snapshots, and product description libraries are used as benchmarks. The numerical values and conditional expressions appearing in the training data are compared with these benchmarks, and entries with obvious mismatches or logical contradictions are marked as noise samples. The noise sample ratio is calculated as the ratio of the number of noise samples to the total amount of training data, used to quantify the proportion of erroneous information, outdated information, or content that seriously deviates from actual business rules in the dataset. The knowledge coverage rate, compliance sample ratio, and noise sample ratio are combined to constitute the data health analysis result, providing a quantitative description of the training data in terms of coverage, compliance, and credibility.
[0053] When comparing data health analysis results with preset thresholds, different dimensions correspond to different parameter ranges. These preset thresholds can be predetermined based on historical project experience and business needs, such as requiring knowledge coverage to be higher than a certain percentage, the proportion of compliant samples to be higher than a certain percentage, and the proportion of noisy samples to be lower than a certain percentage. During the comparison, if a dimension falls below the lower limit or rises above the upper limit, the direction and intensity of adjustment required for that dimension are recorded. This multi-dimensional comparison generates data optimization strategies, including whether to introduce new domain data, whether to remove data fragments from certain sources, and whether to rewrite or downsample samples containing sensitive expressions. For ease of execution, the strategies are organized into combinations of three types of operations: data augmentation, noise removal, and compliance enhancement. Each type of operation further specifies its scope and priority, such as specifying which business tags to perform data augmentation on and which source logs to perform noise removal on.
[0054] The data optimization strategy is translated into a series of actual data processing pipelines during the execution phase. Data augmentation involves retrieving additional samples from product lines, scenario types, or customer groups with insufficient knowledge coverage by searching external data sources or internal cold data warehouses. These samples are then cleaned and labeled according to the same rules as the existing training data before being incorporated into the training set. Noise removal uses filtering rules or scoring models to batch screen the training data, deleting or downweighting samples deemed high-noise to reduce the model's exposure to erroneous information during training. Compliance enhancement increases the proportion of compliant samples in the training data through reweighting, resampling, or template rewriting. For example, samples containing compliant and correct expressions are given a higher sampling probability, or non-compliant expressions are rewritten to comply with regulatory requirements while maintaining their original meaning. After cleaning and enhancement, the optimized training data is formed. This dataset is structurally compatible with the original training data but significantly improved in terms of coverage, compliance, and noise levels.
[0055] The optimized training data is input into the underlying network architecture to be trained, used for parameter initialization and pre-training. The underlying network architecture can be a large language model backbone based on a transformer structure, or a multi-layered structure with domain adaptation layers or instruction tuning heads. During parameter initialization, the optimized training data is iterated through several times to adjust word embeddings, attention weights, and intermediate layer parameters to a state that can basically characterize the distribution and knowledge structure of financial business language. In the pre-training phase, self-supervised learning or autoregressive prediction training continues on the same data, allowing the network to gradually learn structured language patterns such as product descriptions, risk disclosure texts, and business dialogue patterns, until the loss on the training set decreases to a stable range. After pre-training, a stable mapping relationship is established between the underlying network architecture and the optimized training data, forming a target model that can support subsequent reinforcement learning and alignment adjustments, providing a starting model for the subsequent evaluation-driven optimization process.
[0056] This embodiment introduces knowledge graphs and compliance detection models into the training data analysis process to achieve entity linking and rule scanning. This allows for a quantitative characterization of the training data across two dimensions: economic and business knowledge and compliance constraints. This enables the direct identification of areas with insufficient knowledge coverage and concentrated compliance risks. By introducing comparisons with authoritative external data through fact consistency verification, the proportion of noisy samples is explicitly measured, ensuring that the data health analysis results simultaneously reflect coverage, compliance, and credibility. By comparing these analysis results with preset thresholds, a data optimization strategy is generated, incorporating data supplementation, noise removal, and compliance enhancement operations. After cleaning and enhancement, optimized training data is formed. This dataset is then used for parameter initialization and pre-training of the basic network architecture. This ensures that the target model is built on a foundation of controllable data quality from the initial modeling stage, reducing the burden on subsequent alignment and reinforcement learning stages and improving the model's knowledge completeness, compliance robustness, and language expression reliability in financial business scenarios.
[0057] In one embodiment, step S30 above includes: S301, extract definitional test data, comparative test data, and counterfactual test data from the generated test set of the dual-track evaluation system; S302, input the definitional test data, the comparative test data, and the counterfactual test data into the target model; S303, Analyze the output response of the target model to the definitional test data, the comparative test data and the counterfactual test data, determine the accuracy of concept understanding and the consistency of logical reasoning, and obtain the foundation capability index; S304, in a multi-role intelligent agent simulation environment, a simulation evaluation group is configured that includes simulated user intelligent agents, adversarial attack intelligent agents, compliance review intelligent agents and referee intelligent agents. S305, Establish a multi-round dialogue channel between the simulation evaluation group and the target model, and record the interaction trajectory data; S306, using the compliance review agent and the referee agent to analyze the interaction trajectory data, determine the task execution success rate, multi-round dialogue consistency, logical completeness and risk exposure degree, and obtain the interaction task execution indicators.
[0058] In this embodiment, the generated evaluation set of the dual-track evaluation system serves as the input source for the static capability verification of the target model. It needs to be labeled and stratified for different testing purposes during the construction phase to facilitate accurate extraction later. The generated evaluation set pre-distinguishes three types of data: definitional test data, comparative test data, and counterfactual test data. Definitional test data is typically constructed around fundamental knowledge such as financial concepts, product terms, and regulatory terminology. Each data point contains a clear question statement and a single correct answer, used to test the target model's understanding of concept definitions, attribute ranges, and basic relationships. Comparative test data exists in pairs or groups, designed with semantically similar questions but differing key conditions, such as changes in limit amounts, risk levels, or term conditions. This examines whether the target model can provide sufficiently discriminative and consistent answers when faced with subtle changes in conditions. Counterfactual test data, on the other hand, deliberately modifies the preconditions or hypothetical scenarios, such as replacing real interest rates with fictitious interest rates or rewriting regulatory permitting conditions with prohibiting conditions, to construct inputs that are inconsistent with the real world. This is used to test whether the target model can identify anomalies and give a rejection or corrective response when faced with unrealistic preconditions.
[0059] Definitional, comparative, and counterfactual test data are extracted, packaged separately, and input into the target model for inference calculations. To ensure the comparability of evaluation results, the input process maintains a consistent format and decoding parameters, such as a uniform temperature coefficient and upper limit for decoding length, to avoid introducing unexpected randomness. The target model's output response to the three types of test data is fully recorded, including the generated text, confidence distribution, and candidate answer order. Based on this, the system calculates the accuracy of concept understanding and the consistency of logical reasoning by aligning and comparing with the standard answer, reference explanation, and predefined logical constraints. The accuracy of concept understanding can be achieved by statistically analyzing the proportion of correct answers from the target model on the definitional test data, combined with the completeness of explanations of key concepts, such as whether the necessary elements in the risk definition are covered and whether product features are accurately listed, thus forming a comprehensive indicator that includes both selection accuracy and explanation quality. Consistency in logical reasoning can be achieved by analyzing the output behavior of the target model on comparative and counterfactual test data. Specifically, this includes whether the output conclusion changes reasonably with slight parameter variations, whether the model can identify and adjust its response direction when premises are reversed or fabricated, and whether multiple questions and answers within the same logical chain are consistent. By quantifying and summarizing these performance metrics, foundational capability indicators are formed to characterize the target model's basic capabilities in both conceptual understanding and logical reasoning.
[0060] A multi-role intelligent agent simulation environment is used to simulate interaction processes in real business scenarios. This requires configuring simulated user agents, adversarial attack agents, compliance review agents, and adjudicator agents within the environment, and combining these agents into a simulation evaluation group. The simulated user agent constructs natural language interactions according to predetermined business processes and customer behavior patterns. For example, it poses a series of questions, counter-questions, or supplementary explanations related to tasks such as account opening, claims, and asset allocation to examine the target model's coherence and problem-solving capabilities in ordinary business interactions. The adversarial attack agent constructs input through methods such as injecting hints, bypassing keywords, and disguising normal inquiries, attempting to guide the target model to output illegal suggestions, over-promises, or inappropriate information to assess the target model's robustness in an adversarial environment. The compliance review agent performs rule matching and semantic review on each round of responses during the interaction process, marking content involving compliance boundaries, missing risk warnings, or ambiguous expressions. The adjudicator agent, considering the task setting, context, and compliance review results, assigns labels and scores to each round of responses based on whether they advance task completion, meet customer needs, and identify potential risks.
[0061] The simulation evaluation team establishes a multi-turn dialogue channel with the target model through a unified interface. This channel maintains information such as conversation context, turn number, and role labels to support complex dialogue processes. During the dialogue, each request and response is recorded as interaction trajectory data, which includes time sequence, input and output text, role identity, labels, and environmental state information. After the interaction, the compliance review agent and the adjudicator agent perform offline analysis of the interaction trajectory data. Task execution success rate is calculated by statistically analyzing whether the target task is completed within the limited number of turns and constraints, such as whether the risk assessment is successfully completed or whether a compliant product recommendation is provided. Multi-turn dialogue consistency is quantified by analyzing whether the target model's stance is stable across different turns, whether its explanations are self-consistent, and whether its responses are mutually supportive. Logical completeness is measured by evaluating whether the model provides sufficient reasoning chains, explains key premises, and proactively provides supplementary information when necessary. Risk exposure is comprehensively calculated based on dimensions such as the frequency of violations, ambiguous expressions, and high-risk expressions marked by compliance review, as well as the error rate under adversarial attack inputs. The above indicators are combined to form interactive task execution indicators, which characterize the target model's task completion ability, dialogue quality, and risk control level in dynamic interactive scenarios.
[0062] This embodiment extracts definitional test data, comparative test data, and counterfactual test data from the generated evaluation set and inputs them uniformly into the target model. Then, by combining standard answers and logical constraints to calculate the accuracy of conceptual understanding and the consistency of logical reasoning, it can finely characterize the target model's basic level of financial knowledge and reasoning ability in a static offline environment. Simultaneously, by configuring simulated user agents, adversarial attack agents, compliance review agents, and referee agents in a multi-role intelligent agent simulation environment, a simulation evaluation group is constructed, and multi-round dialogue channels are established. Interaction trajectory data is recorded and analyzed, forming interactive task execution indicators from multiple dimensions such as task execution success rate, multi-round dialogue consistency, logical completeness, and risk exposure. This allows static foundational capability indicators to complement dynamic interactive performance, quantifying both the target model's knowledge and reasoning ability, and reflecting task performance and compliance risks in complex financial business dialogues. This enables a quantifiable evaluation of the comprehensive capabilities of large-scale financial models, providing a structured basis for subsequent evaluation-driven adjustments.
[0063] In one embodiment, step S40 above includes: S401, Monitor the real-time response data of the target model during the reinforcement learning phase, and statistically analyze the compliance and security hit rate, business completion rate, refusal rate, and hallucination occurrence rate; S402, use the compliance and security hit rate to quantify compliance and security performance, and use the business completion rate, the rejection rate and the hallucination occurrence rate to quantify business availability performance; S403, Establish a balance analysis relationship including positive gain factors and negative penalty factors to define the balance between the compliance and security performance and the business availability performance; S404, using the balance analysis relationship to analyze the compliance and security hit rate, the business completion rate, the refusal rate and the hallucination occurrence rate to analyze the balance state, and track the changing trend of the balance state; S405, based on the changing trend, determine whether the current training strategy has the risk of over-defense or reward collapse, and generate alignment analysis results.
[0064] In this embodiment, during the reinforcement learning phase, the output behavior of the target model needs to be continuously monitored to quantify its performance in both compliance and business dimensions. During the monitoring phase, the training control component collects real-time response data of the model to training samples, evaluation samples, and online replay samples before and after each human feedback reinforcement, reward model scoring, or policy update. This real-time response data includes at least the fields of input content, generated text, rejection flags, internal scores, and reward signals. Based on these response records, compliance hit rates, task completion rates, rejection behaviors, and the generation of false content can be statistically analyzed in batches or by time windows, transforming the raw logs into statistical indicators that can be used for quantitative analysis. The compliance and security hit rate can be obtained by statistically analyzing the proportion of model outputs that meet policy rules, regulatory constraints, and internal risk control rules in samples marked with compliance check tags; the business completion rate can be obtained by statistically analyzing the proportion of model outputs that meet task success conditions in samples marked with task objectives, such as whether risk warnings were completed or whether a complete business path was provided; the refusal rate can be obtained by identifying outputs containing signs such as refusal to answer or prompts for manual processing, and statistically analyzing the proportion relative to the total number of requests; the hallucination occurrence rate can be obtained by statistically analyzing the number of times false facts, fabricated data, and fictitious clauses are marked by the fact verification module and knowledge base comparison module, and statistically analyzing the proportion in the relevant samples.
[0065] After obtaining the compliance and security hit rate, business completion rate, refusal rate, and hallucination rate, these statistical indicators need to be mapped to two aggregated dimensions, forming compliance and security performance and business availability performance, respectively. Compliance and security performance can be defined as the overall level of compliance hit rate over a period of time. This can be achieved by directly using the compliance and security hit rate, or by combining a weighted approach based on the severity of violations, assigning different weights to serious violations, minor inappropriate expressions, and missing potential risk warnings, and then adjusting them to a unified scoring range. Business availability performance needs to consider both task completion and user access to services smoothly. Therefore, a comprehensive metric can be constructed based on the business completion rate, refusal rate, and hallucination rate. For example, while maintaining task completion capabilities, unnecessary refusal rates can be reduced, and the proportion of answers containing false information can be limited. Through this mapping process, the originally scattered statistical indicators are organized into two relatively independent yet coupled performance metrics, providing input for subsequent balanced analysis.
[0066] To characterize the relationship between compliance and security performance and business availability performance, a balance analysis structure needs to be introduced. This balance analysis relationship can be understood as a functional description of the mutual influence between the two types of performance and their constituent indicators. Positive gain factors reflect the benefits of moving towards the target direction, while negative penalty factors reflect the losses incurred when deviating from constraints or introducing risks. Positive gain factors can be linked to improvements in compliance and security hit rates, business completion rates, etc., for example, granting a gain when business completion rate increases and compliance hit rates remain above a threshold. Negative penalty factors can be linked to situations such as increased illusion occurrence rates, abnormally high refusal rates, and decreased compliance hit rates, applying penalty weights to these deviations. By configuring parameters for gain and penalty factors, the interrelationship between compliance enhancement, risk reduction, and business availability can be clearly expressed, forming numerically comparable and ranking equilibrium indicators.
[0067] After establishing the balance analysis relationship, it is necessary to use this relationship to calculate the compliance and security hit rate, business completion rate, rejection rate, and hallucination occurrence rate to obtain the balance state under the current training cycle or the current time window. Specifically, at the end of each training cycle or at each preset time interval, the proportions of each item statistically obtained over a recent period can be input into the balance analysis relationship to generate a balance value representing the current state. This balance value, along with historical records, is then stored in a time-series storage area. By tracking the balance values over multiple consecutive cycles, trends can be observed. For example, one could observe a situation where compliance and security performance continues to rise but business availability declines significantly, or where business availability rapidly improves while compliance and security performance fluctuates greatly. These trends can be obtained through simple differencing, moving averages, or fitted curves to identify structural shifts during the training process.
[0068] After understanding the changing trends of the equilibrium state, pre-defined judgment rules can be used to identify whether the training strategy has entered an unreasonable zone. Over-defense can be defined as a range where compliance and security performance has reached a predetermined threshold, while the rejection rate continues to rise and the business completion rate continues to decline. This manifests as the model classifying a large number of requests as high-risk and refusing to respond, thus sacrificing service availability. Reward collapse risk can be defined as a range where, during reinforcement learning, reward signals concentrate on limited behavioral patterns, leading to a seemingly high business completion rate in the short term, but the illusion rate begins to rise or compliance hit rate declines. This manifests as the model exchanging excessively risky or provocative answers for higher short-term rewards. In actual implementation, alarm ranges and boundary conditions can be set for the equilibrium state and changing trends. When a combination of indicators is detected entering these ranges, alignment analysis results are generated, structurally outputting the current equilibrium state, change trajectory, and possible risk types for subsequent strategy adjustment decisions.
[0069] This embodiment continuously collects real-time response data during reinforcement learning and statistically analyzes compliance and security hit rate, business completion rate, refusal rate, and hallucination occurrence rate. It summarizes multidimensional behavioral data into two directions: compliance and security performance and business availability performance. Then, it combines the balance analysis relationship including positive gain factors and negative penalty factors to quantify the state and track the trend of change. It can identify the imbalance between compliance and availability in real time during the training phase. It expands the traditional evaluation that only focuses on a single indicator to the alignment analysis results that can distinguish between excessive defense and reward collapse risk. This allows subsequent training strategy adjustments to be made in a targeted manner based on clear balance signals, reducing the situation where financial large models have high scores but low utilization or hidden accumulation of compliance risks before and after going live.
[0070] In one embodiment, step S50 above includes: S501, collects the satisfaction index, conversion rate index and manual takeover rate index of the target model in real time through the business monitoring interface after the target model is launched, and forms real business feedback indicators. S502, extract the base capability indicators, interactive task execution indicators and alignment analysis results corresponding to the target model, and perform time-series alignment processing on the extracted data to construct a training evaluation indicator set; S503, use causal correlation analysis to analyze the correlation mapping relationship between the training evaluation index set and the real business feedback index, and identify the key evaluation index that has a dominant influence on the real business feedback index. S504. Based on the correlation mapping relationship between the key evaluation indicators and the real business feedback indicators, locate the specific training capability dimensions that cause fluctuations in business performance and determine the attribution path.
[0071] In this embodiment, during the model's online operation phase, the business system continuously generates a large amount of operational data related to customer interactions, business processes, and review procedures. To extract quantifiable feedback signals from this operational data, a dedicated business monitoring interface needs to be configured. This interface can be implemented as a data collection component embedded in the online business system. It receives information such as customer requests, model responses, human agent operations, and transaction results through methods like event tracking, log subscriptions, and message queue subscriptions, and archives this information according to session identifiers, user identifiers, business order numbers, and timestamps. Based on the archived data, various business-side evaluation indicators can be constructed. The satisfaction indicator can be derived from customer ratings, questionnaire feedback, or complaint tags submitted after a session ends. The conversion rate indicator can be obtained by statistically analyzing the percentage of sessions with business result identifiers such as "business acceptance successful," "order successful," and "contract completed." The human intervention rate indicator can be obtained by identifying the percentage of times a session is transferred to a human agent or undergoes manual review. Through this collection and calculation process, the originally scattered operational logs are aggregated into real business feedback indicators reflecting actual business operations, used to characterize the model's true performance after its online launch.
[0072] To correlate online performance with the evaluation results obtained during the training phase, it is necessary to extract the underlying capability indicators, interactive task execution indicators, and alignment analysis results corresponding to the target model from the evaluation and analysis system. Underlying capability indicators can include quantitative values calculated from the dual-track evaluation system, such as accuracy of concept understanding, consistency of logical reasoning, and coverage of financial knowledge, reflecting the model's basic capabilities in the knowledge and reasoning dimensions. Interactive task execution indicators can include task execution success rate, consistency of multi-turn dialogue, logical completeness, and risk exposure, obtained from a multi-role intelligent agent simulation environment, reflecting the model's task completion capabilities in simulated business interaction scenarios. Alignment analysis results are derived from the analysis output of the balance between compliance and security performance and business usability performance during the reinforcement learning phase, typically including compliance and security hit rates, distribution of refusal behaviors, distribution of hallucination behaviors, and balance state trend information. Since business feedback data and evaluation data are generated on different timelines and may have different statistical periods, time-series alignment processing is required to ensure the effectiveness of subsequent analysis. Temporal alignment processing can aggregate base capability metrics, interactive task execution metrics, alignment analysis results, and real business feedback metrics within each time period into the same time window by setting a unified time granularity (e.g., day, week, or training round), thereby constructing a one-to-one correspondence of records. After temporal alignment, multi-source and multi-granularity metrics are uniformly integrated into a training evaluation metric set, providing structured input for subsequent connection establishment.
[0073] Complex multivariate relationships exist between the training evaluation metric set and real business feedback metrics. To identify the explanatory components of these relationships, causal association analysis can be introduced. Causal association analysis can employ interpretive statistical analysis techniques such as structural equation modeling, partial least squares regression, and causal graph reasoning to model the dependency structure between each metric in the training evaluation metric set and the real business feedback metrics. In practice, the training evaluation metric set can first be standardized and pre-screened for relevance, eliminating highly collinear or irrelevant variables. Then, a joint variable set including knowledge capability, interaction capability, and alignment balance dimensions can be constructed. Satisfaction, conversion rate, and manual takeover rate metrics are used as target variables. Causal association analysis is then used to estimate the direction and strength of the influence of each training metric on the target variables. During the analysis, a set of key evaluation metrics can be obtained. These metrics have a statistically significant impact on the real business feedback metrics. For example, a positive influence can be found between logical reasoning consistency and satisfaction, risk exposure and manual takeover rate, and a non-linear correlation between alignment balance and conversion rate.
[0074] After identifying key evaluation metrics, it's necessary to further map these metrics back to training capability dimensions to form an interpretable attribution structure. Training capability dimensions can be categorized into areas such as knowledge understanding, rule adherence, risk identification, dialogue management, and task planning. Each area can find one or more corresponding metrics in the training evaluation metric set. By establishing a mapping relationship between key evaluation metrics and training capability dimensions, it's possible to identify which capability dimensions' changes drive fluctuations in real-world business feedback metrics. For example, when data analysis shows a strong and consistent correlation between risk exposure and human intervention rates, business performance fluctuations can be attributed to the risk control capability dimension; when analysis shows a strong positive correlation between multi-turn dialogue consistency and satisfaction, business performance fluctuations can be partially attributed to the dialogue continuity capability dimension. By combining these mapping relationships and changes over time, an attribution chain can be formed, starting from real-world business feedback metrics, passing through key evaluation metrics and training capability dimensions, and ultimately tracing back to the model's internal capability structure. The attribution chain records "which type of business indicator has changed", "how the corresponding evaluation indicators have changed", and "which capability dimensions have shifted as represented by these changes" in the form of directed relationships, providing a traceable basis for subsequent capability gap identification and training adjustments.
[0075] This embodiment extracts satisfaction, conversion rate, and manual takeover rate metrics through a business monitoring interface, aggregating discrete business events during online operation into real business feedback metrics. Combined with a pre-constructed training evaluation metric set, a causal correlation analysis method is introduced to identify key evaluation metrics that have a dominant impact on business performance. These key evaluation metrics are then mapped to the training capability dimension, constructing an attribution link from the training evaluation metric set to the real business feedback metrics. This establishes a stable and interpretable connection between the evaluation layer and real business, ensuring that the underlying capability metrics, interactive task execution metrics, and alignment analysis results generated during the training phase no longer remain at the offline scoring level but directly contribute to explaining the fluctuations in satisfaction, conversion rate, and manual takeover rate. This provides a clear basis for subsequent targeted optimization based on the capability dimension, reducing the disconnect between evaluation results and actual business performance in financial scenarios.
[0076] In one embodiment, step S60 above includes: S601, Analyze the attribution chain to locate the training capability dimension that causes the deviation of business indicators, and quantify the deficiencies of the target model in the training capability dimension to determine the capability gap. S602, Based on the capability gap, determine the parameter configuration for adjusting the model hyperparameters, the data ratio for sample sampling, and the reward weight for adjusting the alignment strength; S603, integrate the parameter configuration, the data ratio, and the reward weight to generate executable training adjustment information; S604, the training adjustment information is loaded into the training controller, and retraining samples are selected from the optimized training data according to the data ratio; S605, using the selected retraining samples to drive the target model to perform parameter update and retraining operations, thereby obtaining the optimized target model.
[0077] In this embodiment, the attribution chain expresses the causal relationship between business metrics and evaluation results during the training phase. It is used to identify which type of training capability changes will cause abnormal fluctuations in business metrics such as satisfaction, conversion rate, and manual takeover rate. In actual implementation, the attribution chain can be represented as a set of weighted directed association records. Each record contains a business metric identifier, an associated evaluation metric identifier, a training capability dimension identifier, and an influence strength coefficient. The system first retrieves records with significant deviations in recent business metrics from the attribution chain, and filters out training capability dimensions whose influence strength exceeds a threshold, such as risk identification capability, compliance understanding capability, and multi-turn dialogue management capability. For each training capability dimension, it is necessary to extract the capability score time series from historical evaluation results and online performance, construct a target capability range or target capability level, and then calculate the difference between the current capability level and the target range to obtain the capability gap in that training capability dimension. The size and direction of the gaps in multiple dimensions can be represented in vector form.
[0078] After identifying the capability gap, this abstract representation needs to be transformed into a concrete, executable training configuration for implementation in the training system. Parameter configuration mainly targets adjustable parameters in the model structure and optimization process, such as learning rate, weight decay coefficient, gradient pruning threshold, layer freeze range, incremental training layers, and the rank of low-rank adaptation modules. The system can pre-maintain a set of parameter adjustment rules for each training capability dimension, mapping the capability gap value to the magnitude and direction of parameter adjustment. For example, when the capability gap related to logical reasoning is large, parameter updates can be enhanced by increasing the learning rate of the corresponding layer and expanding the range of trainable layers; when the capability gap related to compliance is large, risk output can be limited by lowering the learning rate of freely generated modules and increasing the training intensity of compliance modules. Data allocation is used to control the proportion of different types of training samples in the retraining batch. The weights of various samples need to be adjusted based on the training capability dimension to which the capability gap belongs. For example, if the capability gap is concentrated in compliance understanding, the sampling ratio of samples such as compliance Q&A, interpretation of regulatory provisions, and boundary scenario inquiries can be increased, while the proportion of casual chat samples unrelated to the current capability gap can be reduced. Reward weights are used in reinforcement learning or preference-based training scenarios. By adjusting the weight of reward signals on different behaviors, the system guides the model to shift towards directions that better align with business objectives during generation. For example, it increases the reward weights for compliant responses, robust explanations, and conservative decision-making, while reducing the reward weights for hastily drawing conclusions in uncertain scenarios. The system unifies parameter configuration, data allocation, and reward weights into a structured configuration object, adding meta-information such as version number, applicable model identifier, and applicable training capability dimension identifier to form training adjustment information, facilitating direct loading and rollback within the training cluster.
[0079] The training controller is responsible for implementing training adjustments into the specific training process. Upon receiving new training adjustments, the controller first takes a snapshot of the current training state for later restoration if necessary. Then, it updates the relevant parameters of the optimizer and model according to the parameter configuration, such as resetting the learning rate scheduling strategy, adjusting the mask matrix of trainable parameters, and reinitializing or loading the incremental training module. Next, the controller selects retraining samples from the optimized training data based on the data allocation. This involves labeling the sample metadata with scene tags, risk tags, and logical complexity levels, and then performing stratified random sampling according to the proportions set in the data allocation to ensure sufficient coverage of the target capability-related dimensions while keeping the total sample size within an acceptable range. For training scenarios that include reward weights, the controller loads the reward configuration when constructing training batches, injecting the reward weights into the loss function calculation or comparison preference scoring process, ensuring that the model is constrained by capability gaps with each gradient update. After the above configuration takes effect, the target model completes parameter updates through multiple iterations. At the end of training, the evaluation metrics can be recalculated on the internal validation set to obtain a new set of evaluation results. The updated model weights are then solidified as the optimized target model and archived along with the corresponding training adjustment information and capability gap records for subsequent analysis of the differences in effects brought about by each adjustment.
[0080] This embodiment traces back the business indicator deviations recorded in the attribution chain to the training capability dimension, formally quantifies the capability gap using difference calculation, and then generates three types of training adjustment information based on the capability gap: parameter configuration, data allocation, and reward weight. A training controller is introduced to synchronously implement these configurations at three levels: parameter update, sample selection, and reward injection, driving the target model to complete a round of controlled retraining. This tightly links offline evaluation results, online business performance, and the training process, achieving a closed-loop adjustment from business deviation to capability gap, and then to training configuration and parameter update. This makes model optimization behavior no longer dependent on manual experience-based parameter tuning, but directly driven by quantified attribution results, improving the targeting and interpretability of model optimization in financial scenarios, and reducing the training costs and business risks caused by multiple trial and error.
[0081] In one embodiment, a dual-track evaluation-driven optimization processing device is provided, which corresponds one-to-one with the dual-track evaluation-driven optimization processing method described in the above embodiments. (Refer to...) Figure 3 , Figure 3 This is a schematic diagram of the functional modules of a preferred embodiment of the dual-track evaluation-driven optimization processing device of the present invention. The modules include a dual-track evaluation construction module 10, a training data optimization module 20, a model capability evaluation module 30, an alignment relationship analysis module 40, a business feedback attribution module 50, a model retraining module 60, and a business processing execution module 70. Detailed descriptions of each functional module are as follows: The dual-track evaluation construction module 10 is used to build a dual-track evaluation system that includes a generated evaluation set and a business scenario evaluation set; The training data optimization module 20 is used to analyze the training data using the dual-track evaluation system to form a data health analysis result, execute a data optimization strategy based on the data health analysis result to generate optimized training data, and use the optimized training data to build a target model. The model capability evaluation module 30 is used to analyze the target model using the dual-track evaluation system and the multi-role intelligent agent simulation environment, and generate base capability indicators and interactive task execution indicators. Alignment relationship analysis module 40 is used to analyze the balance between the compliance and security performance and business availability performance of the target model in the reinforcement learning stage, and generate alignment analysis results; The business feedback attribution module 50 is used to obtain the real business feedback indicators of the return flow and determine the attribution link between the training evaluation indicator set and the real business feedback indicators, wherein the training evaluation indicator set consists of the base capability indicators, the interaction task execution indicators and the alignment analysis results. The model retraining module 60 is used to determine the capability gap of the target model based on the attribution link, and use the capability gap to generate training adjustment information including parameter configuration, data ratio and reward weight. The training adjustment information and the optimized training data are used to drive the retraining of the target model to obtain the optimized target model. The business processing execution module 70 is used to process business input data using the optimized target model and generate business processing results.
[0082] Specific limitations regarding the dual-track evaluation-driven optimization processing device can be found in the aforementioned limitations of the dual-track evaluation-driven optimization processing method, and will not be repeated here. Each module in the aforementioned dual-track evaluation-driven optimization processing device can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device in hardware form, or stored in the memory of a computer device in software form, so that the processor can call and execute the operations corresponding to each module.
[0083] In one embodiment, a computer device is provided, which may be a server, and its internal structure diagram may be as follows: Figure 4As shown, the computer device includes a processor, memory, network interface, and database connected via a system bus. The processor provides determination and control capabilities. The memory includes non-volatile and / or volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and database. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage media. The network interface is used to communicate with external clients via a network connection. When executed by the processor, the computer program implements the functions or steps of a dual-track evaluation-driven optimization processing method on the server side.
[0084] In one embodiment, a computer device is provided, which may be a client, and its internal structure diagram may be as follows: Figure 5 As shown, the computer device includes a processor, memory, network interface, display screen, and input devices connected via a system bus. The processor provides determination and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage media. The network interface is used to communicate with an external server via a network connection. When executed by the processor, the computer program implements the client-side functions or steps of a dual-track evaluation-driven optimization processing method.
[0085] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to perform the following steps: Construct a dual-track evaluation system that includes generated evaluation sets and business scenario evaluation sets; The training data is analyzed using the dual-track evaluation system to form a data health analysis result. Based on the data health analysis result, a data optimization strategy is executed to generate optimized training data. The target model is then established using the optimized training data. The target model is analyzed using the dual-track evaluation system and multi-role intelligent agent simulation environment to generate base capability indicators and interactive task execution indicators. Analyze the balance between the compliance and security performance and business availability performance of the target model during the reinforcement learning phase, and generate alignment analysis results; Obtain real business feedback metrics from the return flow, and determine the attribution link between the training evaluation metric set and the real business feedback metrics, wherein the training evaluation metric set consists of the base capability metrics, the interaction task execution metrics, and the alignment analysis results; Based on the attribution link, the capability gap of the target model is determined, and the capability gap is used to generate training adjustment information including parameter configuration, data ratio and reward weight. The training adjustment information and the optimized training data are used to drive the target model to be retrained to obtain the optimized target model. The optimized target model is used to process business input data and generate business processing results.
[0086] In one embodiment, a computer-readable storage medium is provided, which may be non-volatile or volatile, and a computer program is stored thereon, which, when executed by a processor, performs the following steps: Construct a dual-track evaluation system that includes generated evaluation sets and business scenario evaluation sets; The training data is analyzed using the dual-track evaluation system to form a data health analysis result. Based on the data health analysis result, a data optimization strategy is executed to generate optimized training data. The target model is then established using the optimized training data. The target model is analyzed using the dual-track evaluation system and multi-role intelligent agent simulation environment to generate base capability indicators and interactive task execution indicators. Analyze the balance between the compliance and security performance and business availability performance of the target model during the reinforcement learning phase, and generate alignment analysis results; Obtain real business feedback metrics from the return flow, and determine the attribution link between the training evaluation metric set and the real business feedback metrics, wherein the training evaluation metric set consists of the base capability metrics, the interaction task execution metrics, and the alignment analysis results; Based on the attribution link, the capability gap of the target model is determined, and the capability gap is used to generate training adjustment information including parameter configuration, data ratio and reward weight. The training adjustment information and the optimized training data are used to drive the target model to be retrained to obtain the optimized target model. The optimized target model is used to process business input data and generate business processing results.
[0087] It should be noted that the functions or steps that can be implemented by the computer-readable storage medium or computer device described above can be referred to the relevant descriptions on the server side and client side in the foregoing method embodiments. To avoid repetition, they will not be described one by one here.
[0088] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.
[0089] It should be noted that if any AI models, software tools, or components not belonging to this company appear in the embodiments of this application, they are merely illustrative examples and do not represent actual use. The above-described embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included within the protection scope of the present invention.
[0090] The user personal information involved in this application embodiment is all authorized (knowing and consenting) by the relevant parties or fully authorized by all parties, and the executing entity can obtain it through various open, legal and compliant means. The collection, storage, use, processing, transmission, provision and disclosure of the information, data and signals involved all comply with the relevant laws and regulations of the relevant countries and regions, and do not violate public order and good morals.
Claims
1. A dual-track evaluation-driven optimization processing method, characterized in that, Includes the following steps: Construct a dual-track evaluation system that includes generated evaluation sets and business scenario evaluation sets; The training data is analyzed using the dual-track evaluation system to form a data health analysis result. Based on the data health analysis result, a data optimization strategy is executed to generate optimized training data. The target model is then established using the optimized training data. The target model is analyzed using the dual-track evaluation system and multi-role intelligent agent simulation environment to generate base capability indicators and interactive task execution indicators. Analyze the balance between the compliance and security performance and business availability performance of the target model during the reinforcement learning phase, and generate alignment analysis results; Obtain real business feedback metrics from the return flow, and determine the attribution link between the training evaluation metric set and the real business feedback metrics, wherein the training evaluation metric set consists of the base capability metrics, the interaction task execution metrics, and the alignment analysis results; Based on the attribution link, the capability gap of the target model is determined, and the capability gap is used to generate training adjustment information including parameter configuration, data ratio and reward weight. The training adjustment information and the optimized training data are used to drive the target model to be retrained to obtain the optimized target model. The optimized target model is used to process business input data and generate business processing results.
2. The dual-track evaluation-driven optimization processing method as described in claim 1, characterized in that, Construct a dual-track evaluation system that includes generated evaluation sets and business scenario evaluation sets, including: Analyze knowledge graphs, regulatory text libraries, and product documentation to extract core entities and constraint logic to form a set of basic knowledge domains; Evaluation data for stability testing is generated based on the aforementioned set of basic knowledge domains, and the evaluation data is compiled to form a generated evaluation set. Acquire historical business interaction data, use a privacy processing module to remove sensitive information from the historical business interaction data while retaining the business context logic, and obtain desensitized business interaction data; The desensitized business interaction data is subjected to feature extraction using preset business scenario screening conditions. Records containing high-risk tags and complex logic tags are identified and reconstructed to form the business scenario evaluation set. The generated evaluation set is integrated with the business scenario evaluation set to construct a dual-track evaluation system.
3. The dual-track evaluation-driven optimization processing method as described in claim 1, characterized in that, The training data is analyzed using the aforementioned dual-track evaluation system to generate data health analysis results. Based on these results, a data optimization strategy is executed to generate optimized training data. Finally, a target model is established using the optimized training data, including: Acquire training data, and use the knowledge graph and compliance detection model in the dual-track evaluation system to perform entity linking and rule scanning on the training data to determine the knowledge coverage and compliance sample ratio. The training data is subjected to fact consistency verification to determine the proportion of noisy samples, and a data health analysis result is formed based on the knowledge coverage, the proportion of compliant samples, and the proportion of noisy samples. The data health analysis results are compared with preset thresholds, and a data optimization strategy including data supplementation, noise removal and compliance enhancement is determined based on the comparison results. The data optimization strategy is executed to clean and enhance the training data, generating optimized training data. The optimized training data is input into the basic network architecture to be trained for parameter initialization and pre-training to establish the target model.
4. The dual-track evaluation-driven optimization processing method as described in claim 1, characterized in that, The target model is analyzed using the aforementioned dual-track evaluation system and multi-role intelligent agent simulation environment to generate base capability indicators and interactive task execution indicators, including: Definition-based test data, comparative test data, and counterfactual test data are extracted from the generated test set of the dual-track evaluation system. Input the definitional test data, the comparative test data, and the counterfactual test data into the target model; By analyzing the output response of the target model to the definitional test data, the comparative test data, and the counterfactual test data, the accuracy of concept understanding and the consistency of logical reasoning are determined, and the foundation capability index is obtained. In a multi-role intelligent agent simulation environment, a simulation evaluation group is configured, which includes simulated user intelligent agents, adversarial attack intelligent agents, compliance review intelligent agents, and referee intelligent agents. Establish a multi-turn dialogue channel between the simulation evaluation group and the target model, and record the interaction trajectory data; By analyzing the interaction trajectory data using the compliance review agent and the adjudication agent, the success rate of task execution, consistency of multi-round dialogue, logical completeness, and degree of risk exposure are determined, thereby obtaining the interaction task execution indicators.
5. The dual-track evaluation-driven optimization processing method as described in claim 1, characterized in that, Analyze the balance between compliance and security performance and business availability performance of the target model during the reinforcement learning phase, and generate alignment analysis results, including: Monitor the real-time response data of the target model during the reinforcement learning phase, and statistically analyze the compliance and security hit rate, business completion rate, refusal rate, and hallucination occurrence rate. The compliance and security hit rate is used to quantify compliance and security performance, and the business completion rate, the refusal rate, and the hallucination occurrence rate are used to quantify business availability performance. Establish a balance analysis relationship that includes positive gain factors and negative penalty factors to define the balance between the compliance and security performance and the business availability performance; The balance analysis relationship is used to analyze the compliance and security hit rate, the business completion rate, the refusal rate, and the hallucination occurrence rate to analyze the balance state and track the changing trend of the balance state; Based on the changing trend, determine whether the current training strategy has the risk of over-defense or reward collapse, and generate alignment analysis results.
6. The dual-track evaluation-driven optimization processing method as described in claim 1, characterized in that, Obtain real business feedback metrics from the backflow, determine the attribution link between the training evaluation metric set and the real business feedback metrics, wherein the training evaluation metric set consists of the base capability metrics, the interaction task execution metrics, and the alignment analysis results, including: The business monitoring interface collects the satisfaction index, conversion rate index and manual takeover rate index of the target model in real time after it goes online, forming real business feedback indicators. Extract the base capability indicators, interactive task execution indicators and alignment analysis results corresponding to the target model, and perform time-series alignment processing on the extracted data to construct a training evaluation indicator set; The correlation mapping relationship between the training evaluation index set and the real business feedback index is analyzed using the causal association analysis method to identify the key evaluation index that has a dominant influence on the real business feedback index. Based on the correlation and mapping relationship between the key evaluation indicators and the real business feedback indicators, the specific training capability dimensions that cause fluctuations in business performance are identified, and the attribution path is determined.
7. The dual-track evaluation-driven optimization processing method as described in claim 1, characterized in that, Based on the attribution path, the capability gap of the target model is determined, and the capability gap is used to generate training adjustment information including parameter configuration, data allocation, and reward weights. The training adjustment information and the optimized training data are used to drive the retraining of the target model to obtain an optimized target model, including: The attribution chain is analyzed to locate the training capability dimension that causes the deviation of business indicators, and the deficiencies of the target model in the training capability dimension are quantified to determine the capability gap. Based on the aforementioned capability gap, determine the parameter configuration for adjusting model hyperparameters, the data ratio for sample sampling, and the reward weight for adjusting alignment strength; By integrating the parameter configuration, the data allocation, and the reward weights, executable training adjustment information is generated; The training adjustment information is loaded into the training controller, and retraining samples are selected from the optimized training data according to the data ratio; The selected retraining samples are used to drive the target model to perform parameter updates and retraining operations, resulting in an optimized target model.
8. A dual-track evaluation-driven optimization processing device, characterized in that, The dual-track evaluation-driven optimization processing device includes: The dual-track evaluation construction module is used to build a dual-track evaluation system that includes a generated evaluation set and a business scenario evaluation set; The training data optimization module is used to analyze the training data using the dual-track evaluation system to form a data health analysis result, execute a data optimization strategy based on the data health analysis result to generate optimized training data, and use the optimized training data to build a target model. The model capability evaluation module is used to analyze the target model using the dual-track evaluation system and the multi-role intelligent agent simulation environment, and to generate base capability indicators and interactive task execution indicators. The alignment analysis module is used to analyze the balance between the compliance and security performance and business availability performance of the target model during the reinforcement learning phase, and generate alignment analysis results. The business feedback attribution module is used to obtain the real business feedback indicators of the return flow and determine the attribution link between the training evaluation indicator set and the real business feedback indicators, wherein the training evaluation indicator set consists of the base capability indicators, the interaction task execution indicators and the alignment analysis results. The model retraining module is used to determine the capability gap of the target model based on the attribution link, and use the capability gap to generate training adjustment information including parameter configuration, data ratio and reward weight. The training adjustment information and the optimized training data are used to drive the retraining of the target model to obtain the optimized target model. The business processing execution module is used to process business input data using the optimized target model and generate business processing results.
9. A computer device, characterized in that, The computer device includes a memory, a processor, and a dual-track evaluation-driven optimization processing program stored in the memory and executable on the processor. When the dual-track evaluation-driven optimization processing program is executed by the processor, it implements the steps of the dual-track evaluation-driven optimization processing method as described in any one of claims 1-7.
10. A computer-readable storage medium, characterized in that, The storage medium stores a dual-track evaluation driver optimization processing program, which, when executed by the processor, implements the steps of the dual-track evaluation driver optimization processing method as described in any one of claims 1-7.