Model multi-level evaluation and online decision-making method and device, equipment and medium
By employing a multi-level evaluation and deployment decision-making method, the objectivity and consistency issues of model evaluation in financial telemarketing scenarios were resolved, achieving automation and quantification of model evaluation and ensuring the security and compliance of the model in financial telemarketing scenarios.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- PING AN TECH (SHENZHEN) CO LTD
- Filing Date
- 2026-02-06
- Publication Date
- 2026-05-15
AI Technical Summary
Existing technologies lack objective and consistent model evaluation methods in financial telemarketing scenarios, failing to accurately reflect the model's performance in actual task execution. Furthermore, they lack rigorous performance comparison mechanisms and risk and compliance verification, leading to reliance on experience-based judgments for large-scale model deployments and a lack of quantitative basis.
This paper presents a multi-level evaluation and deployment decision-making method for models. By acquiring the large-scale task-oriented dialogue model to be evaluated, the benchmark task-oriented dialogue model, and the supporting test dataset, the method generates multi-dimensional evaluation results using language quality analysis, task execution capability evaluation, security risk scanning, and simulated adversarial dialogue. Based on preset decision conditions, the method generates deployment decision instructions.
It automates and quantifies model evaluation, improves the accuracy and consistency of deployment decisions, and ensures the security and compliance of the model in financial telemarketing scenarios.
Smart Images

Figure CN122045012A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of functional testing technology, and in particular to a method, apparatus, equipment, and medium for multi-level evaluation and online decision-making of models. Background Technology
[0002] In telemarketing scenarios within the financial industry, institutions typically fine-tune general-purpose language models to support automated script generation, customer intent recognition, and interactive engagement. However, pre-deployment capability assessments still largely rely on manual methods, such as sampling calls or subjective scoring by quality control personnel. These evaluation methods lack objectivity and consistency, and struggle to cover diverse business scenarios, resulting in an inaccurate reflection of the model's performance in actual tasks.
[0003] Some existing solutions attempt to introduce automated evaluation metrics, such as perplexity, keyword matching, or average number of rounds. However, most of these metrics only focus on the language itself or the outward appearance of the call, failing to characterize the most critical capabilities of a financial telemarketing model, including guiding users through the business process, handling rejection and doubts, and maintaining conversational progress. Furthermore, these methods lack rigorous performance comparison mechanisms, making it impossible to determine, under identical conditions, whether model fine-tuning truly leads to improved performance.
[0004] Financial telemarketing is subject to strict regulation, but existing evaluation systems generally lack systematic verification mechanisms for risk and compliance. Current technologies often fail to identify whether models contain exaggerated promises, misleading statements, or deviations from reality, and lack the ability to output clear deployment conclusions based on unified indicators. Therefore, in practical applications, whether a large-scale model can replace the existing model often still relies on experience-based judgment, lacking quantitative evidence and reliable deployment decision-making basis. Summary of the Invention
[0005] The main objective of this invention is to provide a method, apparatus, device, and storage medium for multi-level evaluation and online decision-making of models, aiming to solve the technical problem that existing technologies cannot quantitatively evaluate the multi-dimensional performance of task-oriented dialogue models and provide clear online decision instructions in a unified process.
[0006] To achieve the above objectives, this invention provides a multi-level model evaluation and deployment decision-making method, comprising: Obtain the large-scale dialogue model for the task to be evaluated, the large-scale dialogue model for the benchmark task, and the corresponding test dataset; Using the test dataset, perform basic language quality analysis on the text generated by the task-oriented dialogue model to be evaluated, and generate basic language quality evaluation results. The task execution capability evaluation results are generated by using a simulated user intelligent agent to perform task flow interaction tests on the task-oriented dialogue model to be evaluated, and by performing a performance comparison analysis of the task-oriented dialogue model to be evaluated relative to the benchmark task-oriented dialogue model. Perform a security risk and compliance scan on the text generated by the task-oriented dialogue model to be evaluated, and generate security risk evaluation results; The simulated user agent is used to control the task-type dialogue model under evaluation to conduct offline simulated adversarial dialogue with the benchmark task-type dialogue model and compare the task execution effects to generate simulation task effect evaluation results. Based on preset decision conditions, the basic language quality evaluation results, the task execution capability evaluation results, the security risk evaluation results, and the simulation task effect evaluation results are comprehensively verified to generate a decision to go online for the large-scale dialogue model to be evaluated.
[0007] Furthermore, to achieve the above objectives, the present invention provides a multi-level model evaluation and deployment decision-making device, comprising: The model and data loading module is used to acquire the large-scale dialogue model for the task to be evaluated, the large-scale dialogue model for the benchmark task, and the corresponding test dataset. The language quality analysis module is used to perform basic language quality analysis on the text generated by the task-oriented dialogue model to be evaluated using the test dataset, and generate basic language quality evaluation results. The task capability evaluation module is used to perform task flow interaction tests on the task-type dialogue model to be evaluated using a simulated user intelligent agent, and to perform a performance comparison analysis of the task-type dialogue model to be evaluated relative to the benchmark task-type dialogue model, and generate task execution capability evaluation results. The security detection module is used to perform security risk and compliance scanning on the text generated by the task-oriented dialogue model to be evaluated, and generate security risk evaluation results. The simulation dialogue evaluation module is used to control the task-type dialogue model under evaluation and the benchmark task-type dialogue model to conduct offline simulated adversarial dialogue and compare the task execution effect to generate simulation task effect evaluation results. The online decision module is used to comprehensively verify the basic language quality evaluation results, the task execution capability evaluation results, the security risk evaluation results, and the simulation task effect evaluation results based on preset decision conditions, and generate an online decision instruction for the large-scale dialogue model of the task to be evaluated.
[0008] Furthermore, to achieve the above objectives, the present invention also provides a computer device, the computer device including a memory, a processor, and a model multilevel evaluation and deployment decision program stored in the memory and executable on the processor, wherein when the model multilevel evaluation and deployment decision program is executed by the processor, it implements the steps of the model multilevel evaluation and deployment decision method as described above.
[0009] Furthermore, to achieve the above objectives, the present invention also provides a computer-readable storage medium storing a multi-level model evaluation and deployment decision program, wherein the multi-level model evaluation and deployment decision program, when executed by a processor, implements the steps of the multi-level model evaluation and deployment decision method as described above.
[0010] Beneficial Effects: This invention relates to the field of functional testing technology and can be applied to business scenarios such as fintech. It discloses a multi-level model evaluation and deployment decision-making method, apparatus, equipment, and medium, including: acquiring a large-scale task-oriented dialogue model to be evaluated, a benchmark task-oriented dialogue model, and a test dataset; using the test dataset to perform basic language quality analysis on the text generated by the large-scale task-oriented dialogue model to be evaluated to obtain basic language quality evaluation results; using a simulated user agent to perform task flow interaction testing and comparing it with the benchmark task-oriented dialogue model to obtain task execution capability evaluation results; performing security risk and compliance scanning on the generated text to obtain security risk evaluation results; in an offline environment, having a simulated user agent drive the large-scale task-oriented dialogue model to be evaluated and the benchmark task-oriented dialogue model to perform simulated adversarial dialogue to obtain simulated task effect evaluation results; and performing comprehensive verification on each evaluation result according to preset decision conditions to generate deployment decision instructions. This invention, through joint language quality evaluation, task execution capability evaluation, security risk evaluation, and simulated task effect evaluation, and performing centralized verification based on preset conditions, forms results that can be directly used for deployment decisions, achieving automated and quantifiable model evaluation, and improving the accuracy and consistency of deployment judgments. Attached Figure Description
[0011] The present invention will be further described below with reference to the accompanying drawings and embodiments. In the accompanying drawings: Figure 1 This is a schematic diagram of an application environment for a multi-level evaluation and deployment decision method for models according to an embodiment of the present invention; Figure 2 This is a flowchart illustrating an embodiment of the multi-level evaluation and deployment decision-making method for the model of the present invention; Figure 3 This is a schematic diagram of the functional modules of a preferred embodiment of the multi-level evaluation and online decision-making device for the model of the present invention; Figure 4 This is a schematic diagram of the structure of a computer device according to an embodiment of the present invention; Figure 5 This is another structural schematic diagram of a computer device according to one embodiment of the present invention. Detailed Implementation
[0012] It should be understood that the specific embodiments described herein are for illustrative purposes only and are not intended to limit the scope of the invention.
[0013] The multi-level evaluation and deployment decision-making method for models provided in this invention can be applied to, for example... Figure 1 In this application environment, the client communicates with the server via a network. The server can obtain the large-scale task-oriented dialogue model to be evaluated, the benchmark task-oriented dialogue model, and the test dataset from the client. Using the test dataset, it performs basic language quality analysis on the text generated by the large-scale task-oriented dialogue model to obtain basic language quality evaluation results; it uses a simulated user agent to perform task flow interaction tests and compares them with the benchmark task-oriented dialogue model to obtain task execution capability evaluation results; it performs security risk and compliance scanning on the generated text to obtain security risk evaluation results; in an offline environment, a simulated user agent drives the large-scale task-oriented dialogue model to be evaluated and the benchmark task-oriented dialogue model to conduct simulated adversarial dialogue to obtain simulated task effect evaluation results; and it performs comprehensive verification of each evaluation result based on preset decision conditions and generates a deployment decision instruction. This invention, through joint language quality evaluation, task execution capability evaluation, security risk evaluation, and simulated task effect evaluation, and performs centralized verification based on preset conditions, forms results that can be directly used for deployment decisions, achieving automated and quantifiable model evaluation, and improving the accuracy and consistency of deployment judgments. The client can be, but is not limited to, various personal computers, laptops, smartphones, tablets, and portable wearable devices. The server can be implemented using a standalone server or a server cluster consisting of multiple servers. The invention will be described in detail below through specific embodiments.
[0014] Please see Figure 2 , Figure 2 This is a flowchart illustrating an embodiment of the multi-level evaluation and deployment decision-making method for models provided by the present invention. It should be noted that although the logical order is shown in the flowchart, in some cases, the steps shown or described may be performed in a different order than that shown here.
[0015] like Figure 2 As shown, the multi-level evaluation and deployment decision-making method for models proposed in this invention includes the following steps: S10: Obtain the large-scale dialogue model for the task to be evaluated, the benchmark large-scale dialogue model for the task, and the corresponding test dataset. In this embodiment, when acquiring the large-scale task-oriented dialogue model to be evaluated, the weight file of the target version needs to be selected from the model version management repository and loaded in the inference environment. The model version management repository can store weight data generated through multiple iterations. Each version is indexed by a unique identifier. The weight file contains information such as network structure parameters, vocabulary configuration, and inference-related scaling factors. After receiving the weight file, the inference environment sequentially completes file decompression, parameter mapping, computation graph construction, and GPU or memory allocation, so that the large-scale task-oriented dialogue model to be evaluated exists as a callable instance, capable of accepting input and outputting text. The acquisition process of the benchmark task-oriented dialogue model is similar to the former, but the weight version selected is used as a comparison benchmark. It generally corresponds to a historical stable version or a live version. It is instantiated through the same loading process to ensure that the running environment and resource configuration of the two are consistent, so that subsequent differences in results are only caused by the model itself. The supporting test dataset comes from the historical dialogue data storage system. The data content includes multi-turn dialogue text, role tags, timestamp information, and tags related to dialogue quality. During the acquisition process, the raw data needs to be anonymized. Sensitive fields such as names, contact information, and account numbers are replaced with substitution symbols or pseudo-random markers. Records with broken formats or severely missing fields are removed, and the encoding format, time format, and role field are standardized to transform dialogue records from different sources into a uniformly structured sample set. To enable subsequent use of the data across different evaluation dimensions, the samples can be divided into subsets based on scenario type, dialogue turn intervals, or tag distribution according to preset screening criteria at this stage. These subsets are then combined using a unified index structure to form a matching test dataset, ensuring stable data sources and clear boundaries in the same batch of model evaluations.
[0016] This embodiment obtains the large-scale dialogue model for the task to be evaluated, the benchmark large-scale dialogue model for the task, and the corresponding test dataset through a unified process. This ensures consistency and traceability of model weight versions, operating environments, and input data, thereby reducing the interference of environmental differences on the evaluation results. By performing desensitization, cleaning, and structural unification processing on historical dialogue data, the corresponding test dataset meets data security requirements and has reusability and scalability, providing a stable and reliable input foundation for subsequent multi-round automatic evaluation.
[0017] S20, using the test dataset, perform basic language quality analysis on the text generated by the task-oriented dialogue model to be evaluated, and generate basic language quality evaluation results; In this embodiment, when performing basic language quality analysis on the text generated by the large-scale dialogue model to be evaluated using the test dataset, it is necessary to first clarify the data structure and calling method used in this step within the test dataset. The test dataset serves as a unified input set, with each sample containing at least dialogue context, role labels, and necessary scene description information. The system inputs the dialogue context into the large-scale dialogue model to be evaluated according to a predetermined batch strategy, and the model outputs one or more text responses. The generated text can be a single-turn response or a multi-turn continuous output. The system associates these outputs with the original context through conversation identifiers and temporal order, and stores them in the intermediate result storage area using a unified encoding format. Basic language quality analysis extracts quantifiable indicators from multiple dimensions of the generated text, including fluency, semantic coherence, structural integrity, and naturalness of expression. For example, to reflect fluency, the proportion of repeated segments, the number of abnormal interruptions, and the appropriateness of punctuation usage can be statistically analyzed based on character sequences or word segmentation sequences. To reflect semantic coherence, the vector similarity between the generated text and its corresponding context can be calculated using a semantic vector model, and obviously contradictory entity combinations can be detected. To reflect the naturalness of expression, overly stiff, overly negative, or unnatural sentence structures can be identified based on preset colloquial and emotional feature templates. The basic language quality assessment results can be structurally designed as a composite record containing multi-dimensional scores, such as fluency scores, coherence scores, natural expression scores, and several auxiliary markers. Each score is derived from the aforementioned indicators through weighting, interval mapping, or rule combination methods, and is archived along with the model version identifier and data source batch for easy subsequent comparison and tracking.
[0018] This embodiment performs basic language quality analysis on the text generated by the large dialogue model for the evaluation task based on the test dataset, and generates structured basic language quality evaluation results. Without relying on subjective human scoring, it can establish a unified quantitative characterization of the model's language output from multiple dimensions such as fluency, coherence, and natural expression. This allows for a clear comparison of the language performance differences between different versions, and at the same time, it can filter out model configurations with obviously substandard language quality before the subsequent evaluation process begins, thereby improving the reliability and efficiency of subsequent evaluation and decision-making.
[0019] S30, using a simulated user agent to perform task flow interaction tests on the task-oriented dialogue model to be evaluated, and performing a performance comparison analysis of the task-oriented dialogue model to be evaluated relative to the benchmark task-oriented dialogue model, generating task execution capability evaluation results. In this embodiment, when using a simulated user agent to perform task flow interaction testing, it is first necessary to establish a task-completion-oriented dialogue flow structure for the task-based dialogue model. The task, from initiation, information gathering, solution provision, to result confirmation, is abstracted into identifiable flow nodes, and the progress of the dialogue between these nodes is tracked during the interaction. The simulated user agent, as the behavior-driven entity on the other end of the dialogue, internally maintains parameters such as user goals, preferences, patience thresholds, and rejection tendencies. Based on the model's responses, it dynamically selects the next round of questions, challenges, additional conditions, or termination of the dialogue, thus forming multi-round interaction records covering the entire task flow. The goal of task flow interaction testing is to identify from these records whether the task is completed, abandoned midway, or experiences flow stagnation, characterizing the performance of the large model in driving the task towards the target state from a behavioral sequence perspective. In the performance comparison and analysis phase, under the same task set and the same simulated user agent configuration, the task-oriented dialogue model under evaluation and the benchmark task-oriented dialogue model were driven to complete the same batch of tasks. By comparing the statistical differences between the two in terms of task completion, dialogue rounds, and process interruptions, the results were summarized into a structured task execution capability evaluation result, which is used to reflect the improvement or degradation of the new model's task execution capability compared to the benchmark model.
[0020] In one implementation, the evaluation system pre-defines processes for multiple typical tasks, describing the task's path from initiation to completion through node identifiers and transition conditions. The process definitions are then associated with the dialogue context and task objectives provided in the test dataset. During runtime, the system selects a task instance from the test dataset, generates a simulated user agent object, loads user profile parameters and a behavior policy table onto it, enabling it to respond with different behaviors such as follow-up questions, additional requests, challenges, or termination under varying feedback conditions. The system inputs the initial context into the large-scale dialogue model of the task to be evaluated. After receiving the model's response, the process control module parses the response content, determines the current task node and whether the process can proceed to the next node, and records node changes, number of rounds, and reasons for interruption. Subsequently, the simulated user agent updates its internal state based on the response content and generates the next round of input, continuously interacting with the large-scale model until the task's final state is reached or the termination condition is triggered. After completing a round of the task, the system extracts indicators from the interaction log, such as whether the task is completed, the number of rounds required for completion, and process rollback or stall events. Next, with identical task instances and simulated user agent configurations, the system switches the dialogue interface to the baseline task-based dialogue model, repeats the above interaction process, and generates another log. The evaluation module aligns the two logs, compares the numerical values of indicators such as task completion, number of rounds, and process interruptions, and writes the comparison results for each task instance and the aggregated statistics across multiple tasks into the task execution capability evaluation result data structure.
[0021] This embodiment utilizes simulated user agents to perform task flow interaction tests and conducts performance comparison analysis between the large-scale task-oriented dialogue model under evaluation and the benchmark task-oriented dialogue model under a unified configuration. This transforms the behavioral characteristics of the model during task completion into quantifiable statistical results, enabling the evaluation to move beyond the quality of single-round responses and objectively measure key dimensions such as whether the task is completed smoothly, whether the process proceeds smoothly, and whether the interaction is frequently interrupted. This provides a more direct and traceable basis for judging whether the new model is superior to the benchmark model in terms of task execution capabilities.
[0022] S40, Perform a security risk and compliance scan on the text generated by the task-oriented dialogue model to be evaluated, and generate a security risk evaluation result; In this embodiment, a security risk and compliance scan is performed on the text generated by the large-scale dialogue model to be evaluated. After obtaining the model-generated text, a structured examination is conducted on potential violations, factual biases, and expressions within the text content. Risk manifestations are mapped to quantitative indicators and summarized into security risk assessment results. The security risk scan is based on three types of risk sources. First, the text contains high-risk triggering content. It is necessary to identify words and phrases in the language that may lead to violations, including sensitive words, prohibited expressions, and words that induce violations. During processing, the word segmentation sequence or sentence fragments of the generated text are matched with a sensitive word lexicon, and the number of text rounds in which the triggering behavior occurs and the frequency of recurrence are counted as risk signals. Second, the text may contain factual description biases. Due to the knowledge generalization characteristics of model-generated text, it may deviate from objective standards when answering questions involving specific values, rule requirements, or conditional constraints. To identify such biases, entity names, values, or logical descriptions in the generated text are extracted and compared with an external set of factual information for consistency judgment, outputting the number of rounds that deviate from dialogue requirements or factual constraints. Third, the text may contain over-promising language. Some expressions, while not triggering sensitive words or involving objective facts, can be identified as potentially risky behaviors through language pattern recognition, such as semantic structures like explicit promises, without exception, and 100%. These can be statistically analyzed from the generated text. The above content is considered the basic carrier of high-risk language performance. By analyzing the generated text round by round, the number of sensitive word matching triggers, the number of factual deviations, and the number of inappropriate promises are accumulated and mapped to evaluation data representing the proportion of violation rounds, the proportion of factual deviations, and the frequency of promise behaviors. This constitutes the basic indicator set for the security risk assessment results.
[0023] In one implementation, the evaluation system pre-creates a sensitive word lexicon, collecting expressions that violate communication norms through manual compilation or corpus mining. The lexicon is organized into a structure containing direct violations, high-risk warnings, and context-limited triggers. After acquiring multi-turn texts generated by the large-scale dialogue model for the task to be evaluated, the system performs word segmentation or fragment analysis on each round's response, matching each segmented result or phrase with entries in the lexicon. If a match is successful, the round number is recorded, and the percentage of violation rounds is updated. During factual consistency detection, the system extracts entity information from the generated text, using an entity recognition model to extract key content such as amounts, rule descriptions, time conditions, and subject behavior requirements. Then, it calls the fact verification module, inputting the extracted entities into an external factual information set. Through string matching, pattern recognition, or template parsing, it determines whether the expression belongs to the valid fact set, and records each deviation to update the fact deviation ratio. For commitment-related expression detection, the system constructs a set of commitment language pattern expressions, identifying absolute terms in the text through syntactic structure matching or high-confidence phrase recognition, accumulating their frequency, and mapping them to a commitment behavior frequency index. After completing the above analysis, the system converts the number of violation rounds, factual deviation rounds, and commitment behavior into percentage values based on the number of text rounds, and combines them to form the security risk assessment results.
[0024] This embodiment, by performing sensitive word detection, factual consistency verification, and commitment behavior identification on the generated text, can systematically decompose potential risk expression behaviors into quantifiable indicators, so that security risks no longer rely on manual reading and judgment, thereby constructing stable assessment results on multi-dimensional risk dimensions and improving the breadth of scanning coverage and the accuracy of risk identification.
[0025] S50, using the simulated user agent to control the task-type dialogue model to be evaluated and the benchmark task-type dialogue model to conduct offline simulated adversarial dialogue and compare the task execution effect to generate simulation task effect evaluation results. In this embodiment, a simulated user agent is used to control an offline simulated adversarial dialogue between a large-scale task-oriented dialogue model under evaluation and a benchmark large-scale task-oriented dialogue model. The performance of the tasks is compared to generate simulation task performance evaluation results. The goal is to examine the performance differences between the two models in completing the same task in a controlled environment without consuming real business traffic. The simulated user agent can be understood as a "virtual user" running in the program. By configuring user profile parameters, behavioral preference parameters, and feedback strategy parameters, it simulates the behaviors of real users in task dialogues, such as asking questions, hesitating, refusing, and probing. The source of such agents can be statistical analysis of user behavior in historical dialogue data or typical user types abstracted by experts. The large-scale task-oriented dialogue model under evaluation is a model instance that needs to be verified after being updated or fine-tuned. The benchmark large-scale task-oriented dialogue model is a model instance that has been put into use or is regarded as a reference standard. Both receive completely identical task starting conditions and user behavior inputs in the same simulation environment, which facilitates horizontal comparison. Offline simulated adversarial dialogue refers to pre-generating a complete task dialogue script on computing resources outside of an online production environment. This simulates the entire process from task initiation, requirement clarification, mid-process resistance, result confirmation, to the end of the dialogue. The adversarial aspect involves simulating the user's intelligent experience dynamically adjusting questions and attitudes based on the model's responses, such as using stronger negative tones, introducing additional restrictions, or directly interrupting the communication, to examine the model's ability to maintain task progress in complex scenarios. Task performance can be broken down into multiple dimensions, including whether the task was completed, whether obvious logical errors occurred during task completion, whether the dialogue rounds were too long or too short, and the proportion of users dropping out midway. Each dimension is recorded as a quantitative indicator or labeled result. The simulation task performance evaluation results are a comprehensive summary of these dimensions, summarizing the records from multiple rounds of offline simulated adversarial dialogue into a structured output. This output reflects the difference in task performance between the evaluated task-oriented dialogue model and the benchmark task-oriented dialogue model. The results can be presented as scores or as a set of multi-dimensional indicators for easy retrieval by subsequent decision-making modules.
[0026] This embodiment utilizes simulated user agents to drive a high-intensity simulated adversarial dialogue between the large-scale task-oriented dialogue model under evaluation and a benchmark task-oriented dialogue model in an offline environment. The task execution effect is quantified and organized into simulated task effect evaluation results. This can expose problems such as difficulties in task progress, logical errors, and dialogue interruptions in advance without occupying real business traffic. At the same time, it forms comparable results on the performance differences of the two models in the same task scenario, which helps to reduce the cost of trial and error in online deployment and improve the reliability of decision-making for updating task-oriented dialogue models.
[0027] S60, based on preset decision conditions, comprehensively verify the basic language quality evaluation results, the task execution capability evaluation results, the security risk evaluation results, and the simulation task effect evaluation results, and generate an online decision instruction for the task-type dialogue model to be evaluated.
[0028] In this embodiment, the basic language quality evaluation results, task execution capability evaluation results, security risk evaluation results, and simulation task effect evaluation results are comprehensively verified based on preset decision conditions. This generates a decision to deploy the large-scale dialogue model for the task to be evaluated. The process revolves around converging evaluation results from different sources into a unified judgment process, filtering them through a rule set to determine whether the model is ready for deployment in real-world tasks. The basic language quality evaluation results reflect the model's ability to complete comprehensible dialogues, including abstract indicators such as fluency, semantic consistency, and naturalness of expression. The task execution capability evaluation results reflect the model's performance in goal-oriented tasks, such as stage recognition accuracy, proactive guidance capability, and adversarial test success rate. The security risk evaluation results focus on violation triggers, factual bias, and absolute commitment markers to identify potential risky behaviors. The simulation task effect evaluation results, derived from offline simulation execution, reflect the model's transformation performance under near-real-world task conditions, including indicators such as task completion rate, dialogue interruption rate, and user feedback tendency. These evaluation results can all be represented as a set of structured fields, with each set containing a score, tags, and necessary additional markers. The pre-defined decision conditions can be organized into a hierarchical threshold set, including blocking thresholds, basic compliance thresholds, and performance improvement thresholds. Blocking thresholds are used to identify red lines that disallow deployment, such as a risk indicator not being zero. Basic compliance thresholds confirm basic usability, such as language indicators exceeding the minimum threshold. Performance improvement thresholds compare whether the model has replacement value, such as simulation performance meeting requirements. The comprehensive verification logic can be sequentially chained, checking risk, basic language performance, and finally task performance and simulation performance. A rejection branch can be triggered if any level fails to meet the requirements. When all dimensions meet the minimum requirements, the decision module can determine that the model meets the deployment conditions, thereby generating a deployment instruction. The deployment decision instruction can be a simple instruction marker or include a summary of decision results and reference information.
[0029] This implementation unifies the multi-dimensional evaluation results into the threshold judgment process and generates clear online instructions, which can realize the connection from the model detection process to the model use decision, greatly reduce the bias of human judgment, make the model online decision-making process structured, avoid erroneous decisions caused by single index driving, and improve the efficiency and security of model iteration.
[0030] In one embodiment, step S10 above includes: S101, retrieve the weight file of the model to be evaluated, which has been trained and fine-tuned with supervision of data in a specific domain, from the model version control repository, and load the weight file of the model to be evaluated in the inference environment to instantiate the large dialogue model of the task to be evaluated. S102, retrieve the baseline model weight file from the model version control repository, and load the baseline model weight file in the inference environment to instantiate the baseline task-oriented dialogue large model; S103, access the business history dialogue database and extract historical business interaction logs containing multi-round dialogue text, user intent annotation tags and quality inspection results; S104, perform privacy information desensitization and format cleaning on the historical business interaction logs to obtain desensitized and cleaned logs; S105, based on preset classification and filtering conditions, divide the desensitized and cleaned logs into a general corpus subset for language quality analysis, a task flow interaction test subset for task flow interaction testing, and an adversarial sample subset for security risk and compliance scanning. S106, combine the general corpus subset, the task flow interaction test subset, and the adversarial sample subset into a matching test dataset.
[0031] In this embodiment, the large-scale task-oriented dialogue model to be evaluated can be understood as a parameter snapshot formed after supervised fine-tuning using task-related corpora. This snapshot is indexed and managed in a model version control repository using version number, timestamp, and tag information. The version control repository can be a centralized storage service. Each version's weight file is equipped with an integrity checksum and metadata description, used to mark the training data range, training epochs, and training configuration used. During actual retrieval, the weight file is pulled from the repository by searching for the target version identifier. After integrity verification, the weights are loaded into the inference environment. The inference environment can be a runtime environment deployed on a physical server or in a virtual container, containing computing resources, runtime library versions, and computational frameworks compatible with the training phase. By loading the weight file and configuring inference parameters, a model instance capable of performing dialogue inference is instantiated, forming the large-scale task-oriented dialogue model entity to be evaluated.
[0032] The baseline task-oriented dialogue model corresponds to a reference model used for comparison. This can be a general model version before fine-tuning or a historically released version. Weight files and metadata are also maintained in the model version control repository. The system retrieves the baseline model's weight file using different version identifiers and loads it into the inference environment following the same process as the previous model, ensuring that both models run under consistent hardware and software conditions. This approach avoids interference from environmental differences during the evaluation phase, allowing the differences to primarily originate from the parameters themselves.
[0033] On the data side, it is necessary to extract real interaction records from the business history dialogue database. This database stores user-server interaction information accumulated in the runtime environment. Each record may include fields such as session identifier, initiation time, participating roles, text content, intent labeling tags, and quality inspection results. Multi-round dialogue texts are linked together using session identifiers to form a complete interaction chain. User intent labels, generated by a manual labeling team or upstream systems, describe the intent category of each user's statement, such as inquiry, confirmation, rejection, or supplementary information. Quality inspection results, derived from a manual quality inspection system or semi-automatic quality inspection process, provide evaluation conclusions and labels for historical dialogues in terms of compliance, wording standardization, and service quality. The system extracts historical business interaction logs from the historical dialogue database in batches, meeting time intervals, business types, and quality requirements, forming the raw data set.
[0034] To use historical records without exposing sensitive information, the extracted historical business interaction logs need to undergo privacy desensitization and format cleaning. Privacy desensitization may involve processing sensitive fields such as names, phone numbers, accounts, and addresses, using methods such as replacing them with placeholders, hash mapping, or field deletion, while ensuring that the same entity within the same session is consistently replaced to avoid disrupting the dialogue logic. Format cleaning includes standardizing text encoding, removing records with missing key fields, correcting abnormal timestamps, standardizing role labels, and normalizing abnormal symbols. The dialogue is also broken down into a data structure based on rounds, providing standardized input for subsequent classification and sampling. The log set obtained after desensitization and cleaning maintains the original dialogue relationships in semantic structure and meets processing requirements in terms of data security and structural standardization.
[0035] Based on this, the system categorizes the anonymized and cleaned logs according to preset classification and filtering conditions. These conditions can be defined as a set of field-based filtering expressions, selecting corresponding records for different usage scenarios. The general corpus subset for language quality analysis can be biased towards samples covering rich linguistic phenomena, such as dialogues with various question-and-answer formats, expression styles, and lengths. Records with extremely singular intent or obviously abnormal quality are excluded by filtering intent tags, quality control tags, and topic tags. The task flow interaction testing subset can be biased towards conversations with complete task chains, requiring the dialogue to show the complete stages from requirement presentation to goal achievement or failure, and historical quality control records to include task result-related markers, facilitating subsequent evaluation of task completion capabilities. The adversarial sample subset for security risk and compliance scanning can construct a higher-risk-density dataset by selecting dialogues containing sensitive topics, boundary expressions, controversial expressions, or those marked as having risky tendencies in historical quality control, to examine the behavior of generated content near risk boundaries.
[0036] The three subsets are structurally consistent, with each sample record labeled with its subset affix. The system combines the general corpus subset, the task flow interaction test subset, and the adversarial example subset to form a matching test dataset. This combination process is achieved through a unified index structure, storing all samples in the same data store, while recording the subset type, original session identifier, and sampling source in the metadata. In subsequent calls, the evaluation process can select the corresponding sample set based on the subset type field, enabling flexible switching between language quality assessment, task execution assessment, and risk scanning. This organizational approach allows different evaluation dimensions to share the same historical data foundation while also enabling the extraction of subsets with different characteristics as needed, ensuring that the test data is both unified and hierarchical.
[0037] This embodiment loads the model to be evaluated and the benchmark model into a unified version repository, and performs anonymization, cleaning and classification from historical dialogue data. It combines language quality assessment data, task interaction assessment data and risk assessment data into a structured test dataset, which can compare different models in the same environment. At the same time, it ensures that the input data covers a variety of dialogue characteristics and risk scenarios, making the subsequent evaluation results more credible and stable, and reducing subjective bias caused by human intervention.
[0038] In one embodiment, step S20 above includes: S201, Extract dialogue context data from the test dataset and input it into the task-oriented dialogue model to be evaluated, and obtain the generated text output by the task-oriented dialogue model to be evaluated based on the dialogue context data; S202, Analyze the perplexity index of the generated text to obtain the word repetition frequency and abnormal sentence break frequency of the generated text, and obtain the fluency index based on the perplexity index, the word repetition frequency and the abnormal sentence break frequency. S203, extract the semantic vectors of the dialogue context data and the generated text, analyze the similarity relationship between the semantic vectors, and detect whether the key entity information contained in the generated text has a logical conflict with the dialogue context data. Based on the similarity relationship and the logical conflict detection results, obtain the semantic consistency index. S204, Based on preset colloquial and emotional features, the emotional appropriateness and natural expression of the generated text are analyzed to obtain a personification score. S205, Perform a comprehensive analysis of the fluency index, the semantic consistency index, and the anthropomorphism score to generate basic language quality evaluation results.
[0039] In this embodiment, the dialogue context data in the test dataset can be understood as a combination of multi-turn dialogue history and current-turn input, including user's historical statements, system's historical responses, and the current-turn user input text, organized into a structured sequence through session identifiers and time order. The system extracts dialogue context data from the test dataset in batches and sends it to the large-scale dialogue model to be evaluated. Under a unified inference configuration, an automatic response process is triggered to obtain a generated text corresponding to each context, which is used for subsequent analysis of various language quality indicators.
[0040] To address the generated text, a perplexity metric is first introduced. This metric is derived from the word probability distribution output by the model during generation. For each word position, the model internally provides a corresponding probability. The system normalizes the logarithmic probability sequence of the entire text, mapping the result to a single numerical value to reflect the uncertainty of predictions during generation. An excessively high perplexity value often indicates a lack of stability in word selection. To further characterize local language phenomena, the system statistically analyzes the repetition frequency of phrases in the generated text. By segmenting the text into phrases or consecutive phrases of several lengths, the system counts the frequency of different phrases to obtain the repetition percentage, which is used to identify mechanical repetition and excessive occurrence of the same phrase. Simultaneously, abnormal sentence segmentation events are detected using information such as punctuation distribution, sentence length thresholds, and intra-sentence dependencies. Phenomena such as excessively short segments, high-frequency inappropriate punctuation segmentation, and premature termination before semantic completion are recorded as abnormal sentence segmentation. The frequency of abnormal sentence segmentation is obtained based on the percentage of abnormal events. After normalization, the perplexity index, phrase repetition frequency, and abnormal sentence break frequency can be combined according to preset weights to form a fluency index, which reflects the overall performance of the generated text in terms of coherence, readability, and rhythm.
[0041] Semantic aspects are characterized jointly through semantic vectors and logical consistency. On both the dialogue context and the generated text sides, the system invokes a semantic encoding network to map the entire text or segmented text into high-dimensional semantic vectors. The semantic vector components can originate from the encoding layer output of a pre-trained language model or a specially configured text encoder. For a single sample, a set of contextual semantic vectors and generated text semantic vectors can be constructed. The closeness between the two is analyzed using cosine similarity or other vector similarity metrics to obtain a semantic closeness measure. Simultaneously, key entity information is extracted at the entity level. Amount, time, quantity, location, entity name, and business object are extracted from the context and generated text sides respectively using named entity recognition models and template matching. Corresponding entities are compared to identify logical conflicts such as inconsistent values, opposite expression polarities, and reversed time sequences, and the types and quantities of conflicts are statistically analyzed. The semantic closeness and logical conflict statistics are jointly mapped to a semantic consistency index. When the semantic closeness remains within a set range and logical conflict events are scarce, the semantic consistency index receives a high evaluation; otherwise, it is marked as deviating from the context.
[0042] At the level of expression style and emotion, anthropomorphic scores are introduced for quantification. The system first defines a set of colloquial features, including sentence length distribution, pronoun usage ratio, frequency of interjections, sentence structure diversity, and conjunction usage. The generated text is analyzed through a linguistic feature extraction module to obtain quantified values for these features. The emotion feature part can include sentiment polarity, sentiment intensity, politeness level, and appeasement tendency, which are output simultaneously by the sentiment analysis model and the politeness recognition model, giving each generated text a corresponding emotion label and score. Emotional appropriateness is evaluated by matching these emotion scores with the emotional state of the corresponding dialogue context and the current task type. For example, if the user shows obvious negative emotions in the context, the generated text that is too cold or completely ignores emotion will be marked as inappropriate. The naturalness of expression is evaluated by analyzing the collocation relationship between interjections and main content, the richness of modifier structures, and the connection between adjacent sentences. Texts with stiff structure, too many filler phrases, or serious logical jumps are classified as having poor naturalness. The results of the colloquial features, along with the emotional appropriateness and the degree of natural expression, are combined into a humanization score after being converted using a unified scale. This score is used to represent the performance of the generated text in terms of its resemblance to a natural conversational style.
[0043] After obtaining the fluency index, semantic consistency index, and anthropomorphism score, the system performs a comprehensive analysis of the three to form a basic language quality assessment result. The synthesis process can construct a multi-dimensional vector, using the three indicators as coordinate dimensions and mapping each generated text to a point. Different regions are then divided into several quality levels using threshold intervals and segmented scoring tables. Alternatively, the three indicators can be combined according to task weights to form a single language quality score for sorting and filtering. By associating with sample identifiers in the test dataset, the basic language quality assessment result can be output as a list of single-sample scores or as a statistical summary, such as the average language quality level under different dialogue types, providing directly usable quantitative input for subsequent modules.
[0044] This embodiment constructs three analysis paths around the generated text: perplexity, repetition and punctuation features, semantic vectors and logical conflicts, and colloquialism and emotional expression. It integrates fluency indicators, semantic consistency indicators, and anthropomorphism scores in a unified structure, which can quantitatively evaluate the output quality at multiple levels of syntax, semantics, and expression style. This allows the basic language quality evaluation results to accurately reflect the differences in model stability, context fit, and the naturalness of human-computer interaction, thus providing a more reliable and comparable language quality metric for subsequent evaluation stages.
[0045] In one embodiment, step S30 above includes: S301, Obtain a preset set of standard task flow definitions, perform semantic matching between the response text output by the task-type dialogue model to be evaluated during the interaction process and the set of standard task flow definitions, and determine the stage recognition accuracy of the task-type dialogue model to be evaluated in recognizing and advancing the task stage. S302, control the simulated user agent to send a resistance test statement containing the intention to refuse or evade to the task-type dialogue model to be evaluated, and use the referee agent to analyze whether the response content generated by the task-type dialogue model to be evaluated in response to the resistance test statement effectively resolves the resistance, and count the proportion of times the resistance is effectively resolved to obtain the resistance response success rate. S303, track the dialogue state, obtain the first guidance tendency value of the task-oriented dialogue model to be evaluated on the guidance behavior of the preset key task goal, and obtain the second guidance tendency value of the benchmark task-oriented dialogue model on the guidance behavior of the same key task goal. S304, Based on the first guidance tendency value and the second guidance tendency value, obtain guidance tendency difference data; S305, Based on the phase identification accuracy, the resistance response success rate, and the guidance tendency difference data, a comprehensive analysis is performed to generate a task execution capability evaluation result.
[0046] In this embodiment, a controlled dialogue is conducted between a simulated user agent and two large task-oriented dialogue models to quantitatively evaluate the execution capability of the model in a complete task flow. To this end, a standard task flow definition set needs to be pre-constructed. Typical task processes in the target business are abstracted into structured flows consisting of multiple task stages. Each task stage includes fields such as stage name, stage objective, key semantic points, required information slots, and acceptable dialogue transition conditions, and is stored in a flow definition library in the form of a state machine, flowchart, or table structure. During online evaluation, a target flow is selected from the standard task flow definition set according to task type, and a unique identifier is assigned to each flow for association with subsequent dialogue records and evaluation results.
[0047] In the task flow interaction test, the task-oriented dialogue model under evaluation receives dialogue content sent by simulated user agents through a unified dialogue interface and returns response text. After each round of interaction, the system semantically matches the response text of the current round with the stage descriptions in the standard task flow definition set. The matching process can be completed collaboratively based on an intent classification model, a slot filling model, and a text similarity model. The intent classification model outputs which stage intent the current response is closer to, the slot filling model determines whether the key information required by the current stage has been given in the response, and the text similarity model is used to help determine the degree of consistency between the response and the stage template expression. Combining the three outputs, each round of response is assigned a task stage label that best matches the current stage, and it is determined whether the stage proceeds in the order defined in the flow, and whether there are any jumps, regressions, or stagnations. The accuracy of stage identification is obtained by statistically analyzing the correctness of stage identification and whether the stage progress conforms to the flow path for all sample dialogues. This accuracy is used to measure the task-oriented dialogue model under evaluation's understanding of the task structure and its ability to advance.
[0048] To examine the model's ability to handle user resistance, a set of resistance test statement templates is maintained within the simulated user agent. These templates cover different intent types such as rejection, hesitation, questioning, and delay, and can be dynamically populated with details based on the current task stage and historical interaction content to generate specific resistance test statements. When the set trigger conditions are met, the simulated user agent inserts these resistance test statements into the dialogue rounds, guiding the task-oriented dialogue model under evaluation to provide a response. The referee agent, acting as an independent evaluation component, receives the resistance test statements and corresponding responses as input. Using both as input, it determines whether the response effectively resolves the resistance, such as whether it directly addresses the question, provides a reasonable explanation, or guides the user back to the task objective rather than simply ending the dialogue. The evaluation results are output in binary or multi-level labeled form and recorded in the evaluation record of the current dialogue round. The system calculates the ratio of rounds deemed to have effectively resolved resistance to the total number of rounds that triggered resistance, obtaining the resistance handling success rate, which measures the model's ability to handle objections and maintain task progress.
[0049] In the task-goal advancement behavior evaluation section, a comparison needs to be made between the task-oriented dialogue model under evaluation and the benchmark task-oriented dialogue model. The system continuously tracks the dialogue status, maintaining state variables such as the current task stage, completed sub-goals, and user response type for each dialogue. For predefined key task goals, such as guiding users to provide necessary information, confirming agreement terms, or completing a key action, the system checks for relevant guidance content in each round of responses. Guiding behavior detection can be performed by combining keyword matching, semantic matching, and dialogue behavior classification models. Each detected guiding behavior is assigned an intensity score, considering factors such as clarity of expression, timing appropriateness, and relevance to the current stage. Throughout the dialogue, these guiding behavior scores are accumulated or statistically analyzed over time, and normalized to obtain a first guiding tendency value, which represents the overall guiding initiative and effectiveness of the task-oriented dialogue model under evaluation in achieving key task goals.
[0050] For the baseline task-oriented dialogue model, the same dialogue context and simulated user agent configuration are used, the same interaction process is repeated, and the corresponding guidance behavior scores are recorded. These scores are then subjected to the same statistical and normalization processing to obtain a second guidance tendency value. The difference between the first and second guidance tendency values is transformed into guidance tendency difference data through numerical calculation or segmented comparison. This difference data can be represented by difference, ratio, or hierarchical labels to directly reflect whether the task-oriented dialogue model under evaluation is enhanced, on par with, or declining in key objective guidance compared to the baseline model. Finally, the stage identification accuracy, resistance handling success rate, and guidance tendency difference data are integrated into a unified evaluation logic. Through preset weights, threshold ranges, or multi-dimensional scoring rules, task execution capability evaluation results are generated, unifying task structure mastery, resistance handling capability, and objective guidance capability onto a set of quantitative indicators that can be used for decision-making.
[0051] This embodiment introduces a standard task flow definition set to perform semantic matching on the response text and statistically analyze the stage identification accuracy. This can reflect the reliability of task-oriented dialogue in stage identification and progress path from a structural perspective. By injecting resistance test statements containing rejection or deflection intentions using a simulated user agent and having the response content evaluated by a referee agent, the resistance handling success rate can quantify the model's performance in dealing with objections and resuming task progress. By obtaining the first guidance tendency value and the second guidance tendency value in a unified dialogue environment and forming guidance tendency difference data, it can directly reflect the gain of the task-oriented dialogue model under evaluation compared with the benchmark task-oriented dialogue model in guiding key task objectives. The task execution capability evaluation results formed by the comprehensive analysis of the three types of quantitative results not only cover multiple dimensions such as process compliance, resistance handling, and objective progress, but also provide a detailed and comparable basis for whether to adopt a new task-oriented dialogue model in the future.
[0052] In one embodiment, step S40 above includes: S401, Obtain the generated text of the task-oriented dialogue model to be evaluated; S402, construct a sensitive word feature library containing illegal inducement words and sensitive topic words, use the sensitive word feature library to match and scan the generated text, and count the number of dialogue rounds that trigger illegal words to determine the proportion of illegal rounds; S403, access an external factual knowledge base containing standard business parameters and legal and regulatory provisions, verify the consistency between the entity information contained in the generated text and the external factual knowledge base, and count the number of dialogue rounds with factual deviations to determine the proportion of factual illusions. S404, Semantic analysis is used to identify statements in the generated text that contain absolute promises or unconditional guarantees, and the frequency of occurrence of the statements is counted to determine the frequency of over-guarantee behavior; S405, compare the percentage of violation rounds, the percentage of factual illusions, and the frequency of over-guarantee behavior with a preset risk blocking threshold to generate a security risk assessment result.
[0053] In this embodiment, during the security risk and compliance scanning phase, the system first continuously acquires generated text from the inference interface of the large-scale dialogue model to be evaluated. The generated text corresponds to the output of one or more dialogue rounds, and can be either a single-round response or a complete dialogue fragment. During the generation phase, the system binds a session identifier, round number, and time information to each generated text, facilitating subsequent statistical analysis and tracing by session dimension. Simultaneously, the text content is uniformly encoded into a standard character set format, handling line breaks, emoticons, and control characters, allowing the subsequent scanning process to proceed on a unified text representation.
[0054] The sensitive word feature library contains risk-related terms and phrases. Content sources can include lists of risk terms compiled by the internal compliance department, examples of negative statements issued by regulatory agencies, and phrases marked as violations during historical quality inspections. To improve coverage, the sensitive word feature library can group different expressions of the same risk concept into the same category. For example, phrases like "guaranteed pass" and "100% success" can be grouped into an absolute commitment category, while phrases like "ignore fees" and "no need to consider costs" can be grouped into a risk masking category. Each term can be accompanied by a category label, severity level, and applicable scenario. The feature library is stored internally using a prefix tree structure, a list of key phrases, or a vectorized representation. During scanning, the system generates text and sequentially inputs it into the sensitive word matching engine. It identifies whether the text contains terms from the sensitive word feature library through string matching, subsequence matching, or vector similarity matching, and records the session identifier and round number for each match. For the same conversation, we can count whether there is at least one hit by each round, and then count the ratio between the number of hit rounds and the total number of rounds for all conversations. This ratio is defined as the percentage of violating rounds, which is used to reflect the degree of exposure of risky language in the conversation process.
[0055] An external fact knowledge base supports fact consistency verification. Its content includes standard business parameters and normative texts relevant to the business. Data in the knowledge base can come from structured configuration tables, product specification document parsing results, and policy text extraction results. For example, fields such as interest rates, rates, time intervals, and object names are uniformly structured and stored, with a unique identifier generated for each record. When generated text enters the fact verification process, entity information is first extracted from the text using an entity recognition component, including numerical entities, time entities, organizational entities, and product names. Then, normalization is used to map different representations to a unified format; for example, "January next year" and "January 2026" are mapped to the same time standard, and "three percent" and "3%" are mapped to a unified percentage representation. Subsequently, based on entity type and semantic context, entity information is mapped to corresponding records in the external fact knowledge base, comparing the numerical values, times, or objects mentioned in the text with the standard values recorded in the knowledge base. For each dialogue round, if there is at least one entity information that is significantly inconsistent with the knowledge base record, the round is marked as a factual deviation round. The ratio of the number of rounds marked as factual deviation to the total number of rounds is calculated to obtain the factual illusion ratio, which is used to measure the degree of deviation of the generated text from the factual level.
[0056] High-risk statements related to absolute promises or unconditional guarantees are identified using a semantic analysis component. This involves first constructing a set of absolute expression patterns, organizing commonly used absolute terms, unconditional guarantee terms, and zero-risk implications into a pattern set, and then expanding the set using semantic similarity to obtain similar sentence structures. When processing the generated text, the text is segmented into sentence-level or phrase-level units, each of which is semantically encoded and then matched against the set of absolute expression patterns to identify statements that semantically belong to absolute promises or unconditional guarantees. The identification results are recorded along with session identifiers and round numbers, and the number of times such statements are included across all rounds is counted, expressed as the frequency of over-guarantee behavior, thus reflecting the degree of exposure of exaggerated promises in the generated text.
[0057] Risk blocking thresholds are used to associate the aforementioned multiple risk indicators with pre-defined security boundaries. Risk blocking thresholds can be a single threshold or a set of thresholds containing multiple dimensions, such as setting maximum allowable ranges for the percentage of violation rounds, the proportion of factual illusions, and the frequency of over-guarantee behavior. After obtaining the percentage of violation rounds, the proportion of factual illusions, and the frequency of over-guarantee behavior, the system compares each indicator with its corresponding risk blocking threshold, generating a judgment result on whether each indicator has exceeded the boundary, and then determines the overall risk level based on combination logic. The risk level can be output through level labels, score ranges, or by associating with state values defined by upper-level control logic, ultimately recorded in the form of security risk assessment results. These results are then associated with the corresponding version of the task-oriented dialogue model to be evaluated and the test dataset, providing a quantitative basis from a risk perspective for subsequent adoption of the model.
[0058] This embodiment matches the generated text from the large-scale dialogue model of the task to be evaluated against a sensitive word feature library to form the proportion of violation rounds, and the frequency of risky terms can be reflected in a quantifiable form. Combined with an external factual knowledge base, consistency verification of entity information is carried out to obtain the proportion of factual illusions, which can identify deviations from key facts in the generated text and provide a basis for controlling the accuracy of content. By identifying absolute promises or unconditional guarantees through semantic analysis and statistically analyzing the frequency of over-guarantee behavior, potential high-risk expressions can be captured. The three proportions are compared with the risk blocking threshold and summarized into a security risk assessment result, so that language risk, factual deviation and promise risk can be summarized and quantified within a unified framework, thereby providing a clear and comprehensive quantitative basis for security risk for subsequent assessment and decision-making.
[0059] In one embodiment, step S50 above includes: S501 is configured with multiple simulated user agents with different personality traits, decision-making logic, and task acceptance thresholds, and initializes a unified dialogue start scenario. S502, respectively control the large-scale dialogue model of the task to be evaluated and the large-scale dialogue model of the benchmark task to conduct multiple rounds of offline simulated adversarial dialogue with the simulated user agent in the dialogue start scenario, and record the dialogue interaction log; S503, calculate the proportion of conversations that reach the preset task final state in the dialogue interaction log to obtain the task conversion rate, and obtain the trust score based on the emotional feedback data in the dialogue interaction log. S504. Based on the performance of the task-oriented dialogue model to be evaluated and the benchmark task-oriented dialogue model on the task conversion rate, obtain task conversion rate difference data, and based on the performance of the task-oriented dialogue model to be evaluated and the benchmark task-oriented dialogue model on the trust score, obtain trust score difference data. S505, Based on the task conversion rate difference data and the trust score difference data, generate a simulation task effect evaluation result that includes effect prediction data.
[0060] In this embodiment, during the offline simulated adversarial dialogue phase, a set of simulated user agents needs to be configured in the simulation environment. These simulated user agents can be understood as personalized dialogue opponents implemented in the program. Each agent instance carries a set of personality trait parameters, a set of decision logic parameters, and a task acceptance willingness threshold. The personality trait parameters control attributes such as language style, tolerance for waiting time, and sensitivity to information redundancy. The decision logic parameters describe the transition conditions for whether to continue the dialogue, enter a hesitant state, or enter a rejection state under different rounds, emotions, and information sufficiency levels. The task acceptance willingness threshold limits the minimum trust level or information sufficiency required to elevate from the current level of interest to accepting the target task. Simultaneously, a unified dialogue starting scenario is established in the simulation environment. This scenario includes initial system prompts, opening greeting text, a task objective description, and an initial context state, ensuring that the large-scale task-oriented dialogue model under evaluation and the benchmark large-scale task-oriented dialogue model conduct dialogue under completely identical initial conditions.
[0061] During the simulation execution phase, the control logic triggers multiple rounds of offline simulated adversarial dialogue between the large-scale dialogue model to be evaluated and the baseline large-scale dialogue model, both against a unified dialogue initiation scenario, and the same group of simulated user agents. In each round of interaction, the current model first generates a response text based on the output of the simulated user agent in the previous round and the dialogue history, and then inputs the response text into the corresponding simulated user agent instance. The simulated user agent updates its internal state and generates the next round of user-side response based on its internal decision-making logic, current emotional state, task acceptance threshold, and the content of the received response text. It can also maintain a task progress variable internally to mark whether the target task has been completed, whether it has entered a strong rejection state, or whether the conversation has been terminated prematurely. Throughout the dialogue process, the simulation engine generates dialogue interaction log entries for each request and response. The dialogue interaction log records at least the model identifier, simulated user agent identifier, round number, input text, output text, internal state label, and time stamp for subsequent offline statistical analysis of task completion and trust changes.
[0062] After the simulation ends, the dialogue interaction logs can be organized by session dimension, checking whether each session has reached a pre-agreed task final state. Task final states can be marked with internal state labels, such as task successful completion, user explicit rejection, and dialogue timeout termination. Final states related to achieving the task objective are considered successful final states. By counting the number of sessions that reached successful final states and comparing it to the total number of sessions, the task conversion rate can be obtained, representing the model's ability to guide simulated users to complete the target task under uniform starting conditions. The dialogue interaction logs can also store sentiment feedback data, such as a sequence of emotional polarities generated by a sentiment analysis model scoring each round of simulated user output, or a sequence of satisfaction labels generated internally by the simulated user agent. Based on this type of sentiment feedback data, sentiment changes can be aggregated by session or round to form a trust score, reflecting the accumulation or decay trend of user-side trust during the dialogue.
[0063] After obtaining the task conversion rate and trust score for each of the two sets of models, the results need to be compared. For the task conversion rate, the difference or ratio between the task conversion rate of the model under evaluation and the benchmark model can be compared to obtain task conversion rate difference data, which represents the improvement or decrease in task completion effectiveness in offline simulation scenarios. For the trust score, a similar approach can be used to compare the trust distributions of the two sets, such as comparing the average trust score, the final round trust score, or the median trust score within a unified scoring range, to obtain trust score difference data, which describes changes in the user's sense of trust. The task conversion rate difference data and trust score difference data can be encapsulated as part of the effect prediction data. The effect prediction data can include numerical fields, trend fields, and discrete label fields, such as "significantly improved," "basically the same," and "significantly decreased," for intuitive presentation of results in subsequent display and analysis. Finally, the effect prediction data is associated with the version identifier and simulation scenario identifier of the model under evaluation to form the simulation task effect evaluation results, providing quantitative basis at the offline simulation level for subsequent decision-making processes.
[0064] This embodiment configures multiple simulated user agents with different personality traits, decision-making logic, and task acceptance thresholds in a unified dialogue starting scenario. It then triggers multiple rounds of offline simulated adversarial dialogue between the large-scale task-oriented dialogue model under evaluation and the benchmark model. This allows for the reproduction of diverse user behaviors without consuming real data. By combining task final state identification and sentiment feedback analysis of dialogue interaction logs, task conversion rate and trust score are obtained, clearly quantifying the model's performance in task completion capability and user trust dimensions. Furthermore, by comparing and generating task conversion rate and trust score difference data and encapsulating them into effect prediction data, the embodiment can simultaneously provide the changing trends of efficiency and experience indicators within a single simulation process. This provides a reproducible and quantifiable basis for determining whether the large-scale task-oriented dialogue model under evaluation has the potential to replace the benchmark model in real-world deployment.
[0065] In one embodiment, step S60 above includes: S601, Load preset decision conditions, which include risk blocking threshold, basic access threshold and effect improvement benchmark; S602, determine whether the security risk assessment result meets the risk blocking threshold; S603, if the security risk assessment result does not meet the risk blocking threshold, then generate a denial-on-line instruction for the task-type dialogue model to be assessed. S604, if the security risk assessment result meets the risk blocking threshold, then determine whether the basic language quality assessment result meets the basic access threshold; S605, if the basic language quality assessment result does not meet the basic admission threshold, then the rejection instruction is generated; S606, if the basic language quality evaluation result meets the basic admission threshold, then determine whether the task execution capability evaluation result and the simulation task effect evaluation result simultaneously meet the effect improvement benchmark; S607, if either the task execution capability evaluation result or the simulation task effect evaluation result does not meet the effect improvement benchmark, then the rejection instruction is generated. S608, if the basic language quality evaluation result, the task execution capability evaluation result, the security risk evaluation result, and the simulation task effect evaluation result all meet the preset decision conditions, then an online permission instruction is generated for the large-scale dialogue model of the task to be evaluated.
[0066] In this embodiment, during the comprehensive verification phase, a pre-defined set of decision conditions needs to be loaded into the configuration storage. This set of decision conditions can be maintained through configuration files, parameter management services, or database tables, and includes three interrelated constraint entries: a risk blocking threshold, a basic admission threshold, and an effectiveness improvement benchmark. The risk blocking threshold limits the upper bound of risks allowed in the security risk assessment results. For example, it can exist as a Boolean field representing a security risk classification label, a risk score upper limit, or a risk indicator hit rate. Once the security risk assessment result exceeds the threshold constraint, it is no longer allowed to proceed to subsequent judgments. The basic admission threshold limits the minimum acceptable range for basic language quality assessment results. It can be formed by combining lower limits for fluency indicators, semantic consistency indicators, and anthropomorphism scores to ensure that the large model entering the subsequent business evaluation phase has reached a usable level in terms of language expression and semantic coherence. The effectiveness improvement benchmark constrains the improvement or relative performance between the task execution capability assessment results and the simulation task effectiveness assessment results. It can be defined as the lower limit of the difference data relative to the benchmark model, the range of the task completion rate improvement ratio, or the trust change range, used to determine whether there is sufficient performance gain to support the model replacement decision.
[0067] After loading the decision condition set, the decision engine sequentially reads the security risk assessment results, basic language quality assessment results, task execution capability assessment results, and simulation task effect assessment results in a fixed order. The processing of security risk assessment results employs a priority judgment logic, matching security-related fields with risk blocking thresholds. When any security indicator triggers a risk blocking threshold—for example, a high-risk label is marked as true, the risk level reaches the unacceptable level, or the violation hit rate exceeds the allowable range—the decision engine immediately terminates the current round of comprehensive verification and generates a rejection instruction for the current task-oriented dialogue model to be evaluated. The rejection instruction can carry the model version identifier, the triggered risk item identifier, and a timestamp, and is sent to the model management platform or deployment control module via an interface to prevent that version from entering the deployment process.
[0068] Provided the security risk assessment results meet the risk blocking threshold, the decision engine continues to process the basic language quality assessment results, comparing language quality-related indicators with the basic admission threshold. This comparison process can include a combined judgment of three dimensions: fluency, semantic consistency, and anthropomorphism, or it can use a weighted scoring method, as long as it ultimately provides a Boolean conclusion on whether the basic admission level has been reached. If any dimension of the basic language quality assessment results falls below the threshold, it indicates a significant deficiency in the dialogue output's fluency or semantic coherence. Even if subsequent business indicators perform well, it is not suitable for deployment. In this case, the decision engine also generates a rejection instruction, and can mark the reason for the language quality failure in the instruction's additional fields, supporting subsequent targeted optimization.
[0069] Once both the security risk assessment and the basic language quality assessment results meet their corresponding thresholds, the decision-making process enters the business effectiveness judgment stage. At this point, the task execution capability assessment results and the simulation task effectiveness assessment results are used as a set of joint inputs and compared simultaneously with the effectiveness improvement benchmark. The task execution capability assessment results may include aggregated indicators such as stage identification accuracy, obstacle response success rate, and guidance tendency difference data. The simulation task effectiveness assessment results may include effect prediction fields corresponding to task conversion rate difference data and trust score difference data. To avoid triggering deployment when only a slight improvement occurs in a single dimension, the decision engine requires both assessment results to reach the effectiveness improvement benchmark. At the configuration level, the effectiveness improvement benchmark can be defined as a set of lower limits, such as a stage identification accuracy improvement of no less than a certain percentage, a task conversion rate difference of no less than zero, and a trust score change of no less than zero. During the actual comparison, a Boolean flag indicating whether the improvement is met is first generated for the task execution capability evaluation result, and another Boolean flag is generated for the simulation task effect evaluation result. Only when both flags are true is the effect improvement benchmark considered to be met. Once either result fails to meet the effect improvement benchmark, the decision engine will generate a rejection instruction and can record the specific dimensions that are not met, supporting subsequent iterations of the task process or simulation strategy.
[0070] The decision engine will only output a deployment license instruction when all three conditions are met: the security risk assessment result meets the risk blocking threshold, the basic language quality assessment result meets the basic access threshold, and the task execution capability assessment result and the simulation task effect assessment result together meet the effect improvement benchmark. The deployment license instruction can include the target model identifier, the allowed deployment environment level, the validity period marker, and the identifiers of the dependent assessment batches, which drive the model release system to perform the actual deployment action. Simultaneously, to maintain traceability, the various assessment results and corresponding thresholds used in the comprehensive verification process can also be stored in the audit record storage area along with the deployment license instruction, allowing for retrospective review of the assessment basis and decision-making path should problems arise during model operation.
[0071] Example Description: In an offline evaluation task, the evaluation platform retrieves the weight file of the model to be evaluated, which has been fine-tuned and trained with supervised training on domain-specific data, from the model version control repository. It then loads this weight file into the inference environment to instantiate the large-scale dialogue model for the evaluation task. Simultaneously, it retrieves the weight file of the benchmark model from the model version control repository and loads it into the inference environment to instantiate the benchmark large-scale dialogue model for the benchmark task. The evaluation platform accesses a business history dialogue database, extracts historical business interaction logs containing multi-turn dialogue text, user intent annotations, and quality inspection results. It then performs privacy desensitization and format cleaning on these historical business interaction logs to obtain desensitized and cleaned logs. For example, it masks and replaces contact information, address fragments, and account fragments in the logs, normalizes abnormal encodings, garbled symbols, and repeated whitespace, and reconstructs the boundaries of multi-turn conversations. The evaluation platform, based on preset classification and filtering criteria, divided the anonymized and cleaned logs into three subsets: a general corpus for language quality analysis, a task flow interaction test subset for task flow interaction testing, and an adversarial sample subset for security risk and compliance scanning. The general corpus subset focuses on covering diverse expression styles and contextual coherence, with dialogue turns spanning both short and long conversations. The task flow interaction test subset retains intent annotation tags and key task node markings. The adversarial sample subset contains a higher proportion of leading phrases, sensitive topic triggers, and extreme follow-up questions. The evaluation platform then combined these subsets into a matching test dataset, recording the test dataset version number and sampling conditions to ensure the reproducibility of the same evaluation at different times.
[0072] In the language quality analysis phase, the evaluation platform extracts dialogue context data from the test dataset and inputs it into the large-scale dialogue model of the task to be evaluated. This yields the generated text output by the large-scale dialogue model based on the dialogue context data. For example, the dialogue context data includes multiple turns of content such as users making requests, asking questions, and expressing hesitation. The generated text includes response segments such as replies, explanations, follow-up questions, and guidance. The evaluation platform analyzes the perplexity index of the generated text and simultaneously obtains the word repetition frequency and abnormal sentence break frequency. The perplexity index can be obtained from the language model scoring output, the word repetition frequency can be obtained by counting the repetition of consecutive word segments, and the abnormal sentence break frequency can be obtained by judging sentence length distribution, punctuation distribution, and dependency structure integrity. Subsequently, based on the perplexity index, word repetition frequency, and abnormal sentence break frequency, a fluency index is obtained, enabling the fluency index to simultaneously reflect the naturalness of the sentences, the degree of redundancy, and abnormal sentence breakage. The evaluation platform extracts semantic vectors from the dialogue context data and the generated text. These semantic vectors are obtained by a vectorized encoder that encodes the multi-turn context and generated text separately. The similarity relationship between the semantic vectors is then analyzed to characterize continuity consistency. Simultaneously, it detects whether key entity information in the generated text logically conflicts with the dialogue context data. For example, it checks whether product names, time ranges, and cost ranges confirmed in the dialogue context are inconsistently expressed in the generated text. A semantic consistency index is obtained based on the similarity relationship and logical conflict detection results. The evaluation platform analyzes the emotional appropriateness and naturalness of the generated text based on preset colloquial and emotional features. Colloquial features include the occurrence of pause words, transition words, polite language, and natural transition sentences. Emotional features include the distribution of reassuring, empathetic, and oppressive expressions. A personification score is obtained through these features. The evaluation platform comprehensively analyzes the fluency index, semantic consistency index, and anthropomorphism score to generate basic language quality evaluation results. For example, the basic language quality evaluation results record that the fluency index is passed, the semantic consistency index is passed, and the anthropomorphism score reaches the preset range, thus providing traceable language quality conclusions for subsequent evaluations.
[0073] During the task flow interaction testing phase, the evaluation platform acquires a pre-defined set of standard task flow definitions. These definitions semantically describe task stages and inter-stage progression conditions, such as typical sentence intentions and key semantic slots for stages like opening confirmation, requirement mining, information explanation, objection handling, progress confirmation, and conclusion. The evaluation platform semantically matches the response text output by the task-oriented dialogue model under evaluation during the interaction with the set of standard task flow definitions. Semantic matching is achieved through sentence vector similarity, consistency of intent classification results, and slot filling consistency, thereby determining the stage recognition accuracy of the task-oriented dialogue model under evaluation in identifying and advancing task stages. The evaluation platform controls a simulated user agent to send resistance test statements containing rejection or deflection intentions to the large-scale task-oriented dialogue model under evaluation. For example, the simulated user agent outputs resistance test statements such as "It's not convenient now," "I don't want it for now," "I'll think about it," and "It's too expensive." The judge agent analyzes whether the response generated by the large-scale task-oriented dialogue model under evaluation in response to the resistance test statements effectively resolves the resistance. The judge agent can give a judgment result on the effectiveness of the resolution from the perspectives of whether the response content completes the emotional easing, whether it provides alternative options, whether it returns to the task progress path, and whether it avoids triggering security risks. Based on this, the evaluation platform counts the proportion of times the resistance is effectively resolved to obtain the resistance response success rate. The evaluation platform tracks the dialogue state, which can be composed of stage tags, user intent tags, and key slot filling status. Based on this, the platform obtains the first guidance tendency value of the task-oriented dialogue model under evaluation on the preset key task objective guidance behavior. For example, if the key task objective guidance behavior is set to guide the user to complete a confirmation action or guide the user to the next stage, the first guidance tendency value can be obtained by comprehensively considering the frequency of guidance statements, the intensity distribution of guidance statements, and the positional distribution of guidance statements. Simultaneously, the platform obtains the second guidance tendency value of the benchmark task-oriented dialogue model on the same key task objective guidance behavior, and obtains guidance tendency difference data based on the first and second guidance tendency values. This guidance tendency difference data can intuitively reflect the change in the progress tendency of the task-oriented dialogue model under evaluation relative to the benchmark task-oriented dialogue model. The evaluation platform performs comprehensive analysis based on stage recognition accuracy, resistance response success rate, and guidance tendency difference data to generate task execution capability evaluation results. For example, the task execution capability evaluation results record that the stage recognition accuracy reaches the preset range, the resistance response success rate is higher than the benchmark level, and the guidance tendency difference data shows a positive improvement, thus forming a quantifiable conclusion on task execution capability.
[0074] During the security risk and compliance scanning phase, the evaluation platform acquires the generated text of the task-oriented dialogue model to be evaluated and constructs a sensitive word feature library containing prohibited inducement words and sensitive topic words. Prohibited inducement words can cover expressions such as exaggerated claims, inducement promises, and avoidance prompts, while sensitive topic words can cover high-risk subject words and combined trigger words. The evaluation platform uses the sensitive word feature library to perform matching scans on the generated text. The matching scan supports synonym rewriting, variant spelling, word segmentation and combination, and cross-sentence triggering modes. The evaluation platform counts the number of dialogue rounds that trigger prohibited words to determine the proportion of prohibited rounds, so that the proportion of prohibited rounds can reflect the degree of risk exposure of the generated text in multiple rounds of dialogue. The evaluation platform accesses an external factual knowledge base containing standard business parameters and legal and regulatory clauses. This external factual knowledge base stores verifiable entity attributes and key clause points in structured entries. The evaluation platform performs consistency checks between the entity information in the generated text and the external factual knowledge base. Consistency checks can include entity attribute comparison, numerical range comparison, condition restriction comparison, and scope of application comparison. The evaluation platform counts the number of dialogue turns with factual discrepancies to determine the proportion of factual illusions, thus reflecting the degree of deviation in the factual statements of the generated text. The evaluation platform uses semantic analysis to identify statements containing absolute promises or unconditional guarantees in the generated text. Semantic analysis can identify statements around absolute semantic patterns such as "inevitable," "certain," "risk-free," and "100%," as well as strong promise sentences that omit conditions. The evaluation platform counts the frequency of these statements to determine the frequency of over-guarantee behavior. The evaluation platform compares the percentage of violation rounds, the proportion of factual illusions, and the frequency of over-guarantee behavior with preset risk blocking thresholds to generate security risk evaluation results. For example, the risk blocking thresholds stipulate that the percentage of violation rounds must be lower than a certain upper limit, the proportion of factual illusions must be lower than a certain upper limit, and the frequency of over-guarantee behavior must be close to zero. If any indicator triggers the threshold, it will be marked as failing in the security risk evaluation results and the trigger item will be recorded.
[0075] In the offline simulated adversarial dialogue phase, the evaluation platform configures multiple simulated user agents with different personality traits, decision-making logic, and task acceptance thresholds, and initializes a unified dialogue starting scenario. Personality traits can be reflected in patience, skepticism, and emotional fluctuation, while decision-making logic can be reflected in different preferences for information sufficiency, risk sensitivity, and cost sensitivity. The task acceptance threshold describes the minimum set of conditions that the simulated user agent needs to meet to transition from a rejection state to an acceptance state. The dialogue starting scenario can consist of a unified opening, a unified user initial intent, and unified constraints, ensuring that the large-scale task-oriented dialogue model under evaluation and the benchmark large-scale task-oriented dialogue model face the same type of starting input. The evaluation platform controls the large-scale task-oriented dialogue model under evaluation and the benchmark large-scale task-oriented dialogue model to conduct multiple rounds of offline simulated adversarial dialogue with the simulated user agents in the dialogue starting scenario, and records the dialogue interaction logs. The dialogue interaction logs include at least the round number, text from both parties, stage status, emotional feedback tags, and a final state marker. The final state marker is used to indicate whether the preset task final state has been reached. The evaluation platform calculates the task conversion rate by statistically analyzing the proportion of conversations reaching a preset task final state in the dialogue interaction log. The preset task final state can be defined as a definable termination state such as confirmation, appointment completion, or submission of intent. The platform also obtains a trust score based on emotional feedback data in the dialogue interaction log. This emotional feedback data can be provided by the simulated user agent after each round of interaction, indicating emotional tendencies and satisfaction levels, allowing the trust score to reflect the cumulative changes in the interactive experience. The platform obtains task conversion rate difference data based on the performance of the task-oriented dialogue model under evaluation and the benchmark task-oriented dialogue model, and trust score difference data based on their performance in trust scores. Based on the task conversion rate difference data and trust score difference data, the platform generates simulated task performance evaluation results that include effect prediction data. For example, the effect prediction data can simultaneously include positive task conversion rate difference data and non-negative trust score difference data, thus characterizing the overall trend of the task-oriented dialogue model under evaluation in offline simulated adversarial dialogue.
[0076] In the comprehensive verification and instruction generation phase, the evaluation platform loads preset decision conditions, including risk blocking thresholds, basic access thresholds, and performance improvement benchmarks. The platform first determines whether the security risk assessment result meets the risk blocking threshold. If the result does not, it directly generates a rejection instruction for the large-scale dialogue model to be evaluated. This rejection instruction can carry a trigger identifier to pinpoint the source of the risk. If the security risk assessment result meets the risk blocking threshold, the platform continues to determine whether the basic language quality assessment result meets the basic access threshold. If the result does not, it generates a rejection instruction and retains the information regarding the language quality failure. When the basic language quality assessment results meet the basic admission threshold, the assessment platform determines whether the task execution capability assessment results and the simulation task effect assessment results simultaneously meet the effect improvement benchmarks. The effect improvement benchmarks can be set as a combination of conditions, such as the stage recognition accuracy and resistance response success rate reaching the lower limit, the guidance tendency difference data being positive, the task conversion rate difference data being positive or not lower than zero, and the trust score difference data not being negative. If either the task execution capability assessment results or the simulation task effect assessment results do not meet the effect improvement benchmarks, a rejection order is generated. When the basic language quality assessment results, task execution capability assessment results, security risk assessment results, and simulation task effect assessment results all meet the preset decision conditions, the assessment platform generates an online permission order for the large-scale dialogue model to be assessed. The online permission order can include a model version identifier and an assessment batch identifier, making the online permission order and the corresponding assessment criteria traceable.
[0077] This embodiment introduces a set of decision conditions consisting of risk blocking thresholds, basic access thresholds, and performance improvement benchmarks. Within a unified decision engine, security risk assessment results, basic language quality assessment results, task execution capability assessment results, and simulation task performance assessment results are sequentially screened in the order of security priority, language quality second, and business performance last. This allows for security compliance interception, language quality control, and business performance gain confirmation within the same decision-making chain, thereby generating rejection or approval instructions. This transforms deployment decisions from relying on experience to automated decision-making based on multi-source assessment results and threshold configurations. This reduces the probability of high-risk models entering the runtime environment and prevents models with insufficient language quality or diminishing business performance from being incorrectly deployed.
[0078] In one embodiment, a multi-level model evaluation and deployment decision-making device is provided, which corresponds one-to-one with the multi-level model evaluation and deployment decision-making method described in the above embodiments. (Refer to...) Figure 3 , Figure 3This is a schematic diagram of the functional modules of a preferred embodiment of the multi-level evaluation and deployment decision-making device for the model of the present invention. The modules include: model and data loading module 10, language quality analysis module 20, task capability evaluation module 30, security detection module 40, simulation dialogue evaluation module 50, and deployment decision-making module 60. Detailed descriptions of each functional module are as follows: The model and data loading module 10 is used to acquire the large-scale dialogue model for the task to be evaluated, the large-scale dialogue model for the benchmark task, and the accompanying test dataset. Language quality analysis module 20 is used to perform basic language quality analysis on the text generated by the task-oriented dialogue model to be evaluated using the test dataset, and generate basic language quality evaluation results. The task capability evaluation module 30 is used to perform task flow interaction tests on the task-type dialogue model to be evaluated using a simulated user intelligent agent, and to perform a performance comparison analysis of the task-type dialogue model to be evaluated relative to the benchmark task-type dialogue model, and generate task execution capability evaluation results. Security detection module 40 is used to perform security risk and compliance scanning on the text generated by the task-oriented dialogue model to be evaluated, and generate security risk evaluation results; The simulation dialogue evaluation module 50 is used to control the task-type dialogue model to be evaluated and the benchmark task-type dialogue model to conduct offline simulated adversarial dialogue and compare the task execution effect to generate simulation task effect evaluation results. The online decision module 60 is used to comprehensively verify the basic language quality evaluation results, the task execution capability evaluation results, the security risk evaluation results, and the simulation task effect evaluation results based on preset decision conditions, and generate an online decision instruction for the large-scale dialogue model of the task to be evaluated.
[0079] Specific limitations regarding the multi-level model evaluation and deployment decision-making device can be found in the aforementioned limitations on the multi-level model evaluation and deployment decision-making method, and will not be repeated here. Each module in the aforementioned multi-level model evaluation and deployment decision-making device can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device in hardware form, or stored in the memory of a computer device in software form, so that the processor can call and execute the corresponding operations of each module.
[0080] In one embodiment, a computer device is provided, which may be a server, and its internal structure diagram may be as follows: Figure 4As shown, the computer device includes a processor, memory, network interface, and database connected via a system bus. The processor provides determination and control capabilities. The memory includes non-volatile and / or volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and database. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage media. The network interface is used to communicate with external clients via a network connection. When executed by the processor, the computer program implements the functions or steps of a multi-level evaluation and deployment decision-making method for a model on the server side.
[0081] In one embodiment, a computer device is provided, which may be a client, and its internal structure diagram may be as follows: Figure 5 As shown, the computer device includes a processor, memory, network interface, display screen, and input devices connected via a system bus. The processor provides determination and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage media. The network interface is used to communicate with an external server via a network connection. When executed by the processor, the computer program implements client-side functions or steps of a multi-level evaluation and deployment decision-making method for a model.
[0082] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to perform the following steps: Obtain the large-scale dialogue model for the task to be evaluated, the large-scale dialogue model for the benchmark task, and the corresponding test dataset; Using the test dataset, perform basic language quality analysis on the text generated by the task-oriented dialogue model to be evaluated, and generate basic language quality evaluation results. The task execution capability evaluation results are generated by using a simulated user intelligent agent to perform task flow interaction tests on the task-oriented dialogue model to be evaluated, and by performing a performance comparison analysis of the task-oriented dialogue model to be evaluated relative to the benchmark task-oriented dialogue model. Perform a security risk and compliance scan on the text generated by the task-oriented dialogue model to be evaluated, and generate security risk evaluation results; The simulated user agent is used to control the task-type dialogue model under evaluation to conduct offline simulated adversarial dialogue with the benchmark task-type dialogue model and compare the task execution effects to generate simulation task effect evaluation results. Based on preset decision conditions, the basic language quality evaluation results, the task execution capability evaluation results, the security risk evaluation results, and the simulation task effect evaluation results are comprehensively verified to generate a decision to go online for the large-scale dialogue model to be evaluated.
[0083] In one embodiment, a computer-readable storage medium is provided, which may be non-volatile or volatile, and a computer program is stored thereon, which, when executed by a processor, performs the following steps: Obtain the large-scale dialogue model for the task to be evaluated, the large-scale dialogue model for the benchmark task, and the corresponding test dataset; Using the test dataset, perform basic language quality analysis on the text generated by the task-oriented dialogue model to be evaluated, and generate basic language quality evaluation results. The task execution capability evaluation results are generated by using a simulated user intelligent agent to perform task flow interaction tests on the task-oriented dialogue model to be evaluated, and by performing a performance comparison analysis of the task-oriented dialogue model to be evaluated relative to the benchmark task-oriented dialogue model. Perform a security risk and compliance scan on the text generated by the task-oriented dialogue model to be evaluated, and generate security risk evaluation results; The simulated user agent is used to control the task-type dialogue model under evaluation to conduct offline simulated adversarial dialogue with the benchmark task-type dialogue model and compare the task execution effects to generate simulation task effect evaluation results. Based on preset decision conditions, the basic language quality evaluation results, the task execution capability evaluation results, the security risk evaluation results, and the simulation task effect evaluation results are comprehensively verified to generate a decision to go online for the large-scale dialogue model to be evaluated.
[0084] It should be noted that the functions or steps that can be implemented by the computer-readable storage medium or computer device described above can be referred to the relevant descriptions on the server side and client side in the foregoing method embodiments. To avoid repetition, they will not be described one by one here.
[0085] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.
[0086] It should be noted that if any AI models, software tools, or components not belonging to this company appear in the embodiments of this application, they are merely illustrative examples and do not represent actual use. The above-described embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included within the protection scope of the present invention.
[0087] The user personal information involved in this application embodiment is all authorized (knowing and consenting) by the relevant parties or fully authorized by all parties, and the executing entity can obtain it through various open, legal and compliant means. The collection, storage, use, processing, transmission, provision and disclosure of the information, data and signals involved all comply with the relevant laws and regulations of the relevant countries and regions, and do not violate public order and good morals.
Claims
1. A multi-level evaluation and deployment decision-making method for a model, characterized in that, Includes the following steps: Obtain the large-scale dialogue model for the task to be evaluated, the large-scale dialogue model for the benchmark task, and the corresponding test dataset; Using the test dataset, perform basic language quality analysis on the text generated by the task-oriented dialogue model to be evaluated, and generate basic language quality evaluation results. The task execution capability evaluation results are generated by using a simulated user intelligent agent to perform task flow interaction tests on the task-oriented dialogue model to be evaluated, and by performing a performance comparison analysis of the task-oriented dialogue model to be evaluated relative to the benchmark task-oriented dialogue model. Perform a security risk and compliance scan on the text generated by the task-oriented dialogue model to be evaluated, and generate security risk evaluation results; The simulated user agent is used to control the task-type dialogue model under evaluation to conduct offline simulated adversarial dialogue with the benchmark task-type dialogue model and compare the task execution effects to generate simulation task effect evaluation results. Based on preset decision conditions, the basic language quality evaluation results, the task execution capability evaluation results, the security risk evaluation results, and the simulation task effect evaluation results are comprehensively verified to generate a decision to go online for the large-scale dialogue model to be evaluated.
2. The multi-level evaluation and deployment decision-making method for models as described in claim 1, characterized in that, Obtain the large-scale dialogue model for the task to be evaluated, the benchmark large-scale dialogue model for the task, and the corresponding test dataset, including: Retrieve the weight file of the model to be evaluated, which has been fine-tuned and trained with supervision of domain-specific data, from the model version control repository, and load the weight file of the model to be evaluated in the inference environment to instantiate the large-scale dialogue model of the task to be evaluated. Retrieve the baseline model weight file from the model version control repository, and load the baseline model weight file into the inference environment to instantiate the baseline task-oriented dialogue large model; Access the business history dialogue database and extract historical business interaction logs containing multi-round dialogue text, user intent annotation tags, and quality inspection results; The historical business interaction logs are subjected to privacy information desensitization and format cleaning processes to obtain desensitized and cleaned logs; Based on preset classification and filtering conditions, the de-identified and cleaned logs are divided into a general corpus subset for language quality analysis, a task flow interaction test subset for task flow interaction testing, and an adversarial sample subset for security risk and compliance scanning. The general corpus subset, the task flow interaction test subset, and the adversarial sample subset are combined into a matching test dataset.
3. The multi-level evaluation and deployment decision-making method for models as described in claim 1, characterized in that, Using the test dataset, perform basic language quality analysis on the text generated by the task-oriented dialogue model to be evaluated, and generate basic language quality evaluation results, including: Extract dialogue context data from the test dataset and input it into the large-scale dialogue model to be evaluated, then obtain the generated text output by the large-scale dialogue model to be evaluated based on the dialogue context data; The perplexity index of the generated text is analyzed to obtain the word repetition frequency and abnormal sentence break frequency of the generated text. The fluency index is obtained based on the perplexity index, the word repetition frequency and the abnormal sentence break frequency. Extract the semantic vectors of the dialogue context data and the generated text, analyze the similarity relationship between the semantic vectors, and detect whether the key entity information contained in the generated text has a logical conflict with the dialogue context data. Based on the similarity relationship and the logical conflict detection results, obtain the semantic consistency index. The emotional appropriateness and naturalness of the generated text are analyzed based on preset colloquial and emotional features to obtain a personification score. The fluency index, the semantic consistency index, and the anthropomorphism score are comprehensively analyzed to generate basic language quality assessment results.
4. The multi-level evaluation and deployment decision-making method for models as described in claim 1, characterized in that, The task-oriented dialogue model under evaluation is tested by simulating a user agent to perform task flow interaction tests. A performance comparison analysis is then performed between the model and a benchmark task-oriented dialogue model to generate task execution capability evaluation results, including: Obtain a preset set of standard task flow definitions, and perform semantic matching between the response text output by the task-type dialogue model to be evaluated during the interaction process and the set of standard task flow definitions to determine the stage recognition accuracy of the task-type dialogue model to be evaluated in terms of task stage recognition and progress. The simulated user agent is controlled to send resistance test statements containing rejection or deflection intentions to the task-oriented dialogue model to be evaluated. The referee agent analyzes whether the response content generated by the task-oriented dialogue model to the resistance test statements effectively resolves the resistance, and the proportion of times the resistance is effectively resolved is statistically analyzed to obtain the resistance response success rate. Track the dialogue state to obtain the first guidance tendency value of the task-oriented dialogue model to be evaluated on the guidance behavior of the preset key task goal, and obtain the second guidance tendency value of the benchmark task-oriented dialogue model on the guidance behavior of the same key task goal. Based on the first guidance tendency value and the second guidance tendency value, guidance tendency difference data is obtained; Based on a comprehensive analysis of the stage identification accuracy, the resistance response success rate, and the guidance tendency difference data, a task execution capability evaluation result is generated.
5. The multi-level evaluation and deployment decision-making method for models as described in claim 1, characterized in that, The text generated by the task-oriented dialogue model to be evaluated is subjected to a security risk and compliance scan, and security risk evaluation results are generated, including: Obtain the generated text of the large-scale task-oriented dialogue model to be evaluated; Construct a sensitive word feature library containing illegal inducement words and sensitive topic words, use the sensitive word feature library to match and scan the generated text, and count the number of dialogue rounds that trigger illegal words to determine the proportion of illegal rounds; Access an external factual knowledge base containing standard business parameters and legal and regulatory provisions, verify the consistency between the entity information contained in the generated text and the external factual knowledge base, and count the number of dialogue rounds with factual deviations to determine the proportion of factual illusions; Semantic analysis is used to identify statements in the generated text that contain absolute promises or unconditional guarantees, and the frequency of these statements is counted to determine the frequency of over-guarantee behavior. The percentage of violation rounds, the percentage of factual illusions, and the frequency of over-guarantee behavior are compared with preset risk blocking thresholds to generate security risk assessment results.
6. The multi-level evaluation and deployment decision-making method for models as described in claim 1, characterized in that, The simulated user agent controls the large-scale dialogue model of the task under evaluation to conduct offline simulated adversarial dialogue with the benchmark large-scale dialogue model of the task, and compares the task execution effects to generate simulation task effect evaluation results, including: Configure multiple simulated user agents with different personality traits, decision-making logic, and task acceptance thresholds, and initialize a unified dialogue start scenario; The task-oriented dialogue model to be evaluated and the benchmark task-oriented dialogue model are respectively controlled to conduct multiple rounds of offline simulated adversarial dialogue with the simulated user agent in the dialogue initiation scenario, and the dialogue interaction logs are recorded. The percentage of conversations that reach the preset task final state in the dialogue interaction log is statistically analyzed to obtain the task conversion rate, and the trust score is obtained based on the emotional feedback data in the dialogue interaction log. Based on the performance of the task-oriented dialogue model to be evaluated and the benchmark task-oriented dialogue model on the task conversion rate, the task conversion rate difference data is obtained, and based on the performance of the task-oriented dialogue model to be evaluated and the benchmark task-oriented dialogue model on the trust score, the trust score difference data is obtained. Based on the task conversion rate difference data and the trust score difference data, a simulation task performance evaluation result containing performance prediction data is generated.
7. The multi-level evaluation and deployment decision-making method for models as described in claim 1, characterized in that, Based on preset decision conditions, the basic language quality evaluation results, the task execution capability evaluation results, the security risk evaluation results, and the simulation task effect evaluation results are comprehensively verified to generate a deployment decision instruction for the large-scale dialogue model to be evaluated, including: Load preset decision conditions, which include risk blocking threshold, basic access threshold and effect improvement benchmark; Determine whether the security risk assessment result meets the risk blocking threshold; If the security risk assessment result does not meet the risk blocking threshold, a denial-of-access instruction is generated for the task-based dialogue model to be assessed. If the security risk assessment result meets the risk blocking threshold, then determine whether the basic language quality assessment result meets the basic access threshold; If the basic language quality assessment result does not meet the basic admission threshold, then the rejection instruction is generated. If the basic language quality evaluation result meets the basic admission threshold, then determine whether the task execution capability evaluation result and the simulation task effect evaluation result simultaneously meet the effect improvement benchmark. If either the task execution capability evaluation result or the simulation task effect evaluation result fails to meet the effect improvement benchmark, then the rejection order is generated. If the basic language quality evaluation results, the task execution capability evaluation results, the security risk evaluation results, and the simulation task effect evaluation results all meet the preset decision conditions, then an online permission instruction is generated for the large-scale dialogue model to be evaluated.
8. A multi-level evaluation and deployment decision-making device for a model, characterized in that, The multi-level evaluation and deployment decision-making device for the model includes: The model and data loading module is used to acquire the large-scale dialogue model for the task to be evaluated, the large-scale dialogue model for the benchmark task, and the corresponding test dataset. The language quality analysis module is used to perform basic language quality analysis on the text generated by the task-oriented dialogue model to be evaluated using the test dataset, and generate basic language quality evaluation results. The task capability evaluation module is used to perform task flow interaction tests on the task-type dialogue model to be evaluated using a simulated user intelligent agent, and to perform a performance comparison analysis of the task-type dialogue model to be evaluated relative to the benchmark task-type dialogue model, and generate task execution capability evaluation results. The security detection module is used to perform security risk and compliance scanning on the text generated by the task-oriented dialogue model to be evaluated, and generate security risk evaluation results. The simulation dialogue evaluation module is used to control the task-type dialogue model under evaluation and the benchmark task-type dialogue model to conduct offline simulated adversarial dialogue and compare the task execution effect to generate simulation task effect evaluation results. The online decision module is used to comprehensively verify the basic language quality evaluation results, the task execution capability evaluation results, the security risk evaluation results, and the simulation task effect evaluation results based on preset decision conditions, and generate an online decision instruction for the large-scale dialogue model of the task to be evaluated.
9. A computer device, characterized in that, The computer device includes a memory, a processor, and a multi-level model evaluation and deployment decision program stored in the memory and executable on the processor. When the multi-level model evaluation and deployment decision program is executed by the processor, it implements the steps of the multi-level model evaluation and deployment decision method as described in any one of claims 1-7.
10. A computer-readable storage medium, characterized in that, The storage medium stores a multi-level model evaluation and deployment decision program, which, when executed by a processor, implements the steps of the multi-level model evaluation and deployment decision method as described in any one of claims 1-7.