Multi-agent based agent evaluation dataset generation method, device and medium
Patent Information
- Application Number
- CN202611257300.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-08-19
- Publication Date
- 2026-09-18
AI Technical Summary
但是,直接Prompt大模型主要追求文本生成流畅性,缺少对具体业务领域的本体语义约束和业务逻辑规范约束,生成的测评数据集易存在业务语义偏离、逻辑规则冲突、领域特征缺失、推理场景单一等问题
[0016] The technical solution of this application generates candidate evaluation samples for multiple agents through full-link constraints using domain business ontology and a small number of real test data. Combined with the determination of preset difference values corresponding to probability distribution difference values, it realizes closed-loop iterative optimization of candidate evaluation sample distribution. For candidate evaluation samples with unqualified distribution deviation, iterative updates are performed or the domain business ontology adaptively evolves and regenerates candidate evaluation samples. Qualified candidate evaluation samples are processed hierarchically to obtain multi-level agent evaluation datasets. This realizes the automated generation, quality verification and distribution alignment optimization of agent evaluation datasets, significantly improving the authenticity, rationality and domain adaptability of agent evaluation datasets.
Smart Images

Figure CN122777409A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence, and in particular to a method, apparatus and medium for generating intelligent agent evaluation datasets based on multi-agent systems. Background Technology
[0002] With the large-scale deployment of business intelligent agents (IT operations, customer service, data analysis, and enterprise digital employees) driven by large language models, the industry urgently needs standardized, domain-adaptable benchmark datasets that can cover all business capabilities to identify shortcomings in intelligent agent reasoning, tool invocation, exception handling, multi-round planning, and other capabilities.
[0003] Currently, the main approach to generating intelligent agent evaluation datasets in related technologies is to use large-scale direct prompt models to generate datasets in batches. However, these large-scale direct prompt models primarily prioritize the fluency of text generation, lacking constraints on ontological semantics and business logic specifications specific to the business domain. This often results in evaluation datasets with issues such as deviations from business semantics, conflicts in logical rules, missing domain features, and limited reasoning scenarios. Furthermore, the generation of evaluation datasets needs to simultaneously meet multiple quality requirements, including domain realism, consistency in semantic distribution, compliance with business logic, richness of scenario layers, completeness of anomaly cases, and gradient of reasoning difficulty. Large-scale direct prompt models typically struggle to uniformly model and precisely constrain these dimensions.
[0004] Therefore, existing assessment dataset generation schemes suffer from poor domain adaptability in generating assessment datasets. Summary of the Invention
[0005] The main objective of this application is to propose a method, apparatus, and medium for generating intelligent agent evaluation datasets based on multi-agent systems, aiming to improve the domain adaptability of the generated evaluation datasets.
[0006] To achieve the above objectives, this application proposes a method for generating an agent evaluation dataset based on multi-agent systems, comprising: Obtain the domain business ontology of the target business scenario and a small number of real test data, and extract the task semantic skeleton based on the domain business ontology from the small number of real test data. Multiple generative agents are invoked to generate multiple sub-datasets based on the task semantic skeleton, and the multiple sub-datasets are integrated to obtain candidate evaluation samples; Calculate the probability distribution difference between the candidate evaluation samples and the few sample real test data, and determine whether the probability distribution difference is greater than a preset difference value; If the probability distribution difference value is greater than the preset difference value, then the candidate evaluation samples are iteratively updated according to the domain business ontology or the ontology introspective agent is called to update the domain business ontology and regenerate the candidate evaluation samples, and the probability distribution difference value between the candidate evaluation samples and the few sample real test data is recalculated. If the probability distribution difference value is less than or equal to the preset difference value, the candidate evaluation samples are processed in a hierarchical manner according to the domain business ontology to obtain a multi-level intelligent agent evaluation dataset.
[0007] In some embodiments, the step of extracting the task semantic skeleton from the few-sample real test data based on the domain business ontology includes: Based on the class hierarchy and attribute definition of the domain business ontology, semantic parsing is performed on each piece of real test data in the few sample real test data, and the business domain, user intent, involved entities and task types are identified to obtain the identification results of each piece of real test data. Each of the identification results is instantiated into an entity in the domain business ontology, and the abstract concept class to which each entity belongs is inferred to obtain an ontology semantic profile; Identify the task pattern class to which the ontology semantic profile belongs in the conceptual hierarchy of the domain business ontology, extract the process steps, participating roles, required resource types and preconditions associated with the task pattern class, and convert them into a formal skeleton representation to obtain the task semantic skeleton.
[0008] In some embodiments, the step of invoking multiple generative agents to generate multiple sub-datasets based on the task semantic skeleton, and integrating the multiple sub-datasets to obtain candidate evaluation samples includes: Multiple generated intelligent agents are invoked to perform at least one of the following expansion operations—scenario expansion, expression transformation, anomaly construction, and reasoning enhancement—within the semantic constraint space of the domain business ontology, respectively, to obtain multiple subset datasets. The multiple subset datasets are deduplicated, format-aligned, and merged to obtain initial candidate evaluation samples; The verification agent is invoked to perform semantic consistency parsing on the initial candidate evaluation samples based on the conceptual axioms and attribute constraints of the domain business ontology, removing invalid sample entries that contain semantic illusions or conflicting business logic, thereby obtaining the candidate evaluation samples.
[0009] In some embodiments, the plurality of generated agents include a scene extension agent, an expression variant agent, an anomaly construction agent, and a reasoning enhancement agent; the invocation of the plurality of generated agents, within the semantic constraint space of the domain business ontology, performs at least one of the following expansion operations: scene extension, expression transformation, anomaly construction, and reasoning enhancement, respectively, based on the task semantic skeleton, to obtain the plurality of the following subsets: The scenario extension agent is invoked to expand the business domain in the task semantic skeleton based on the concept classification tree and dependency relationship of the domain business ontology, and to expand the related services with mutual dependencies to obtain the scenario extension subset. The expression variant agent is invoked to instantiate the ontology terms in the task semantic skeleton into various natural language expressions according to the language layer mapping rules of the domain business ontology, thereby obtaining a subset of expression variants. The exception construction agent is invoked to construct business boundary exception samples based on the legitimate exception sub-states of the domain business ontology and the task semantic skeleton, thereby obtaining a subset of exception construction data. The reasoning-enhanced agent is invoked to automatically combine and nest the multi-dimensional dependency conditions of the task semantic skeleton based on the cross-knowledge source dependency relationships of the domain business ontology, thereby generating complex task samples with high reasoning depth and obtaining a reasoning-enhanced subset.
[0010] In some embodiments, calculating the probability distribution difference between the candidate evaluation samples and the few-sample real test data includes: Obtain multi-dimensional semantic features of the domain business ontology, wherein the multi-dimensional semantic features include at least the subject domain class hierarchy, intent class hierarchy, complexity attribute, and user role class; Construct a multi-dimensional semantic feature space based on the aforementioned multi-dimensional semantic features; The candidate evaluation samples and the few-sample real test data are respectively mapped to the multi-dimensional semantic feature space to obtain the first probability distribution of the candidate evaluation samples and the second probability distribution of the few-sample real test data. The probability distribution difference value is obtained by calling the Jensen-Shannon divergence algorithm or the Wasserstein distance algorithm to calculate the probability distribution difference between the first probability distribution and the second probability distribution in the multi-dimensional semantic feature space.
[0011] In some embodiments, iteratively updating candidate evaluation samples based on the domain business ontology or calling the ontology introspective agent to update the domain business ontology and regenerate candidate evaluation samples includes: Check whether the number of iterations for updating candidate evaluation samples exceeds the preset number; If the number of iterations is less than or equal to the preset number, then the candidate evaluation samples are iteratively updated according to the domain business ontology. If the number of iterations is greater than the preset number, then the ontology introspection agent is invoked to update the domain business ontology and regenerate candidate evaluation samples.
[0012] In some embodiments, iteratively updating the candidate evaluation samples based on the domain business ontology includes: In the instance space of the domain business ontology, data resampling, concept correction, or targeted regeneration operations are performed on the candidate evaluation samples to iteratively update the candidate evaluation samples.
[0013] In some embodiments, invoking the ontology introspection agent to update the domain business ontology and regenerate candidate evaluation samples includes: Based on the probability distribution difference value, a dimension-by-dimensional difference attribution analysis is performed on the multi-dimensional semantic feature space to determine the bottleneck dimension with the highest difference contribution. The bottleneck dimension is then mapped back to the corresponding concept class in the domain business ontology, and a failure diagnosis report is generated. The ontology introspective agent is invoked to generate a constrained extended domain business ontology based on the failure diagnosis report and the OWL description logic specification. The extended domain business ontology is checked for logical consistency, maintainability of existing instance classification, and monotonically decreasing distribution differences. If any verification fails, the extended domain business ontology is discarded, while the domain business ontology remains unchanged. If all verifications are successful, the ontology changes that differ between the extended domain business ontology and the domain business ontology are cascaded and propagated to the ontology semantic profile re-instantiation, task semantic skeleton re-extraction, and dataset difficulty stratification rule update to obtain a new task semantic skeleton, and candidate evaluation samples are regenerated based on the new task semantic skeleton.
[0014] This application further proposes a device for generating an intelligent agent evaluation dataset based on multi-agent systems, including: At least one processor; and, A memory communicatively connected to the at least one processor; wherein, The memory stores instructions that are executed by the at least one processor, which enable the at least one processor to perform the multi-agent-based agent evaluation dataset generation method described above.
[0015] This application further proposes a storage medium storing a computer program, the computer program including program instructions, which, when executed by a processor, enable the processor to execute the above-described method for generating an intelligent agent evaluation dataset based on multi-agent systems.
[0016] The technical solution of this application generates candidate evaluation samples for multiple agents through full-link constraints using domain business ontology and a small number of real test data. Combined with the determination of preset difference values corresponding to probability distribution difference values, it realizes closed-loop iterative optimization of candidate evaluation sample distribution. For candidate evaluation samples with unqualified distribution deviation, iterative updates are performed or the domain business ontology adaptively evolves and regenerates candidate evaluation samples. Qualified candidate evaluation samples are processed hierarchically to obtain multi-level agent evaluation datasets. This realizes the automated generation, quality verification and distribution alignment optimization of agent evaluation datasets, significantly improving the authenticity, rationality and domain adaptability of agent evaluation datasets. Attached Figure Description
[0017] Figure 1 This is a flowchart illustrating an embodiment of the multi-agent intelligent agent evaluation dataset generation method of this application; Figure 2 This is a flowchart illustrating another embodiment of the multi-agent intelligent agent evaluation dataset generation method of this application; Figure 3 This is a flowchart illustrating another embodiment of the multi-agent intelligent agent evaluation dataset generation method of this application; Figure 4 This is a flowchart illustrating another embodiment of the multi-agent intelligent agent evaluation dataset generation method of this application; Figure 5 This is a flowchart illustrating another embodiment of the multi-agent intelligent agent evaluation dataset generation method of this application; Figure 6 This is a flowchart illustrating another embodiment of the multi-agent intelligent agent evaluation dataset generation method of this application; Figure 7 This is a flowchart illustrating another embodiment of the multi-agent intelligent agent evaluation dataset generation method of this application; Figure 8 This is a schematic diagram of an embodiment of the intelligent agent evaluation dataset generation device based on multi-agent systems according to this application. Detailed Implementation
[0018] The solutions in the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments in this application, and not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of this application.
[0019] It should be noted that all directional indicators (such as up, down, left, right, front, back, etc.) in the embodiments of this application are only used to explain the relative positional relationship and movement of each component in a certain specific posture (as shown in the figure). If the specific posture changes, the directional indicator will also change accordingly.
[0020] It should also be noted that when a component is referred to as "fixed to" or "set on" another component, it can be directly on the other component or may have an intervening component present. When a component is referred to as "connected to" another component, it can be directly connected to the other component or may have an intervening component present.
[0021] Furthermore, the use of terms such as "first" and "second" in this application is for descriptive purposes only and should not be construed as indicating or implying their relative importance or implicitly specifying the number of technical features indicated. Therefore, a feature defined as "first" or "second" may explicitly or implicitly include at least one of that feature. Additionally, the technical solutions of the various embodiments can be combined with each other, but only on the basis of being achievable by those skilled in the art. When the combination of technical solutions is contradictory or impossible to implement, such a combination of technical solutions should be considered non-existent and not within the scope of protection claimed in this application.
[0022] This application proposes a method for generating agent evaluation datasets based on multi-agent systems, referring to... Figure 1 , Figure 1 This is a flowchart illustrating an embodiment of the multi-agent-based agent evaluation dataset generation method of this application. In some embodiments, the multi-agent-based agent evaluation dataset generation method includes: Step S110: Obtain the domain business ontology and a small number of real test data for the target business scenario, and extract the task semantic skeleton from the small number of real test data based on the domain business ontology. Step S120: Invoke multiple generative agents to generate multiple subset datasets based on the task semantic skeleton, and integrate the multiple subset datasets to obtain candidate evaluation samples. Step S130: Calculate the probability distribution difference between the candidate evaluation sample and the small sample real test data, and determine whether the probability distribution difference is greater than the preset difference value. Step S140: If the probability distribution difference value is greater than the preset difference value, the candidate evaluation samples are iteratively updated according to the domain business ontology or the ontology introspective agent is called to update the domain business ontology and regenerate the candidate evaluation samples, and the probability distribution difference value between the candidate evaluation samples and the few sample real test data is recalculated. Step S150: If the probability distribution difference value is less than or equal to the preset difference value, the candidate evaluation samples are processed in layers according to the domain business ontology to obtain a multi-level intelligent agent evaluation dataset.
[0023] In this embodiment, as Figure 1 As shown, the multi-agent-based intelligent agent evaluation dataset generation method proposed in this application can be configured as a running program, functional software, or packaged as firmware, control driver, or independent functional module; then it can be deployed to a multi-agent-based intelligent agent evaluation dataset generation device, so that the multi-agent-based intelligent agent evaluation dataset generation device can run the multi-agent-based intelligent agent evaluation dataset generation method. Alternatively, it can be configured into other computer devices to enable the computer devices to run the multi-agent-based intelligent agent evaluation dataset generation method. The multi-agent-based intelligent agent evaluation dataset generation device can be simply referred to as the generation device.
[0024] When running a multi-agent-based intelligent agent evaluation dataset generation method, the generation device can acquire the domain business ontology of the target business scenario and a small sample of real test data. Based on the domain business ontology, it extracts the task semantic skeleton from the small sample of real test data. For example, when a user needs to generate an intelligent agent evaluation dataset, the user can first construct the domain business ontology based on the target business scenario and acquire a small sample of real test data. Then, the user inputs the domain business ontology and the small sample of real test data into the generation device. At this point, the generation device can acquire the domain business ontology and the small sample of real test data for the target business scenario. The target business scenario can include scenarios such as IT operations and maintenance, customer service, data analysis, and enterprise digital employees. The domain business ontology includes at least classes, attributes, constraints, and axioms. The small sample of real test data refers to real test data with a small number of samples, originating from real business interactions, and possessing real business semantics and task characteristics.
[0025] After acquiring the domain business ontology and a small number of real test data, the generation device can perform refined semantic analysis on each piece of real test data in the small number of real test data according to the class hierarchy and attribute definitions of the domain business ontology. This accurately identifies the business domain, user intent, involved entities, and task type corresponding to each piece of real test data, obtaining the recognition result for each piece of real test data. Each recognition result is instantiated into an individual instance within the domain business ontology, and the upper-level abstract concept class corresponding to each entity is inferred by combining ontology reasoning rules. Based on this, an ontology semantic profile that can accurately represent the real business characteristics is constructed. On this basis, the task pattern class corresponding to the ontology semantic profile in the domain business ontology concept hierarchy is matched, and the process steps, participating roles, required resource types, and pre-constraints bound to the task pattern class are extracted. The colloquial expressions, personalized redundant information, and concrete random entities in the original samples are stripped away and abstracted into standardized, formalized structured expressions, ultimately obtaining a task semantic skeleton with general extensibility that can be reused by multiple intelligent agents.
[0026] After obtaining the task semantic skeleton, multiple generative agents can be invoked to generate multiple subsets based on the task semantic skeleton, and these subsets are then integrated to obtain candidate evaluation samples. For example, the generation device invokes multiple generative agents to perform differentiated data augmentation operations based on a unified task semantic skeleton within the semantic constraint space defined by the domain business ontology. These data augmentation operations can include four types: scenario expansion, expression transformation, anomaly construction, and reasoning enhancement, generating multiple subsets respectively. The augmentation process of each generative agent strictly adheres to the concept classification, dependency relationships, language mapping rules, and anomaly definition specifications of the domain business ontology, ensuring that the subsets do not deviate from the actual business logic and do not exceed semantic boundaries. The generation device performs batch deduplication, unified format alignment, and data merging on the multiple subsets to obtain initial candidate evaluation samples. Subsequently, a verification agent is invoked to perform global semantic consistency verification on the initial candidate evaluation samples based on the concept axioms and attribute constraints of the domain business ontology, screening and eliminating invalid sample entries with semantic illusions, business logic conflicts, or violations of ontology rules, ultimately obtaining compliant and reliable candidate evaluation samples.
[0027] After obtaining candidate evaluation samples, the probability distribution difference between the candidate evaluation samples and a small number of real test data can be calculated, and it can be determined whether the probability distribution difference is greater than a preset difference value. For example, the preset difference value can be customized by the user according to the actual situation. The generation device extracts multi-dimensional semantic features defined by the domain business ontology. Multi-dimensional semantic features can include subject domain class hierarchy, intent class hierarchy, complexity attribute, and user role class. A unified multi-dimensional semantic feature space is built based on the multi-dimensional semantic features. The candidate evaluation samples and the small number of real test data are mapped to the multi-dimensional semantic feature space respectively, and the probability distributions corresponding to the two sets of data are constructed respectively. Then, the degree of difference between the two sets of probability distributions is quantified by the Jensen-Shannon (JS) divergence or Wasserstein distance algorithm to obtain the probability distribution difference value, and the probability distribution difference value is compared with the preset difference value.
[0028] If the probability distribution difference value is greater than the preset difference value, the candidate evaluation samples are updated iteratively according to the domain business ontology or the ontology introspective agent is called to update the domain business ontology and regenerate the candidate evaluation samples. The probability distribution difference value between the candidate evaluation samples and the few samples of real test data is recalculated. For example, if the probability distribution difference value is greater than a preset difference value, it indicates that the current candidate evaluation samples have a distribution drift compared to the few samples of real test data, failing to meet the data distribution consistency requirements of the agent evaluation. In this case, the generation device can distinguish the two-layer optimization path by counting the number of consecutive iterations of the candidate evaluation samples: if the number of iterations is within the preset range, it prioritizes performing inner-loop optimization operations such as data resampling, concept correction, and targeted regeneration on the candidate evaluation samples within the existing instance space of the domain business ontology, iteratively correcting the candidate evaluation samples; if the number of iterations exceeds the preset number, it is determined that the existing concept system of the domain business ontology has insufficient coverage, triggering the ontology introspection mechanism, calling the ontology introspection agent to complete the ontology's restricted expansion and multi-dimensional consistency verification based on the probability distribution difference value, and completing the update of the domain business ontology after the verification is passed, and re-executing the agent expansion process based on the updated domain business ontology to generate new candidate evaluation samples. After each round of updates, it returns to step S130 to recalculate the probability distribution difference value, achieving closed-loop iterative convergence.
[0029] If the probability distribution difference is less than or equal to a preset difference value, the candidate evaluation samples are stratified according to the domain business ontology to obtain a multi-level intelligent agent evaluation dataset. For example, when the probability distribution difference is less than or equal to the preset difference value, it indicates that the semantic distribution of the current candidate evaluation samples is highly consistent with a small number of real test data, and the data quality meets the standards. At this time, the generation device can perform multi-dimensional difficulty stratification and classification of the compliant and converged candidate evaluation samples according to the stratification rules built into the domain business ontology, such as task complexity attributes, inference level, and resource dependency difficulty, and finally obtain a multi-level intelligent agent evaluation dataset with clear hierarchy that adapts to the evaluation needs of different intelligent agent capabilities.
[0030] The technical solution of this application generates candidate evaluation samples for multiple agents through full-link constraints using domain business ontology and a small number of real test data. Combined with the determination of preset difference values corresponding to probability distribution difference values, it realizes closed-loop iterative optimization of candidate evaluation sample distribution. For candidate evaluation samples with unqualified distribution deviation, iterative updates are performed or the domain business ontology adaptively evolves and regenerates candidate evaluation samples. Qualified candidate evaluation samples are processed hierarchically to obtain multi-level agent evaluation datasets. This realizes the automated generation, quality verification and distribution alignment optimization of agent evaluation datasets, significantly improving the authenticity, rationality and domain adaptability of agent evaluation datasets.
[0031] Reference Figure 2 , Figure 2 This is a flowchart illustrating another embodiment of the method for generating intelligent agent evaluation datasets based on multi-agent systems according to this application. In some embodiments, the aforementioned extraction of the task semantic skeleton from a small number of real test data based on the domain business ontology includes: Step S160: Based on the class hierarchy and attribute definition of the domain business ontology, perform semantic parsing on each piece of real test data in the few sample real test data, and identify the business domain, user intent, involved entities and task types to obtain the identification results of each piece of real test data. Step S161: Instantiate each recognition result into an entity in the domain business ontology, and infer the abstract concept class to which each entity belongs to obtain the ontology semantic profile. Step S162: Identify the task pattern class to which the ontology semantic profile belongs in the conceptual hierarchy of the domain business ontology, extract the process steps, participating roles, required resource types and preconditions associated with the task pattern class, and convert them into a formal skeleton representation to obtain the task semantic skeleton.
[0032] In this embodiment, as Figure 2As shown, during step S110, semantic parsing can be performed on each piece of real test data in the few-sample real test data based on the class hierarchy and attribute definitions of the domain business ontology. Semantic parsing is performed on each piece of real test data in the few-sample real test data based on the class hierarchy and attribute definitions of the domain business ontology, identifying the business domain, user intent, involved entities, and task type, thus obtaining the recognition results for each piece of real test data. For example, the generation device uses the class hierarchy and attribute definitions of the domain business ontology as the semantic constraint benchmark to perform refined semantic parsing on each piece of real test data in the few samples. For each piece of real test data, the generation device discards the colloquial sentences, personalized expressions, and random concrete content of the original real test data, accurately deconstructs and identifies the business domain, user intent, involved entities, and task type corresponding to the real test data from the business dimension, achieves multi-dimensional structured information extraction for a single piece of real test data, and finally outputs the recognition results corresponding to each piece of real test data.
[0033] After obtaining the recognition results of each real test data, each recognition result is instantiated into an entity within the domain business ontology, and the abstract concept class to which each entity belongs is inferred, resulting in an ontology semantic profile. For example, the generation device maps and instantiates each recognition result one by one into a corresponding individual within the domain business ontology, ensuring that each piece of real test data corresponds to a unique individual within the domain business ontology. Simultaneously, relying on the inference rules and class hierarchy inheritance relationships built into the domain business ontology, semantic reasoning is performed on each instantiated individual, automatically deriving and determining the upper-level abstract concept class to which each entity belongs, clarifying the semantic attribution, business level, and concept association relationship of the real test data, and obtaining an ontology semantic profile.
[0034] After obtaining the ontology semantic profile, the task pattern class to which the ontology semantic profile belongs in the conceptual hierarchy of the domain business ontology can be identified. The associated process steps, participating roles, required resource types, and preconditions of the task pattern class are extracted and converted into a formal skeleton representation, resulting in the task semantic skeleton. For example, the generation device matches the ontology semantic profile to the global conceptual hierarchy of the domain business ontology, accurately locating the standardized task pattern class to which the ontology semantic profile belongs. Then, it extracts the core business elements pre-bound to the task pattern class. These business elements may include standardized business process steps, task participating role types, resource types required for task execution, and preconditions for task initiation. The generation device then standardizes, removes redundancy, and abstracts the extracted business elements, stripping away personalized random variable information and converting fragmented business elements into a unified, standardized, and extensible formal skeleton expression structure, ultimately generating a task semantic skeleton that combines business realism with general extensibility.
[0035] Reference Figure 3 , Figure 3This is a flowchart illustrating another embodiment of the multi-agent intelligent agent evaluation dataset generation method of this application. In some embodiments, the aforementioned invocation of multiple generating agents to generate multiple sub-datasets based on the task semantic skeleton, and the integration of multiple sub-datasets to obtain candidate evaluation samples includes: Step S170: Invoke multiple generating agents to perform at least one of the following expansion operations in the semantic constraint space of the domain business ontology: scenario expansion, expression transformation, anomaly construction, and reasoning enhancement, respectively, to obtain multiple subset datasets. Step S171: Perform deduplication, format alignment and merging on multiple subset datasets to obtain initial candidate evaluation samples; Step S172: Invoke the verification agent to perform semantic consistency parsing on the initial candidate evaluation samples based on the conceptual axioms and attribute constraints of the domain business ontology, remove invalid sample entries with semantic illusions or business logic conflicts, and obtain candidate evaluation samples.
[0036] In this embodiment, as Figure 3 As shown, during step S120, multiple generative agents can be invoked to generate multiple subset datasets. These agents, within the semantic constraint space of the domain business ontology, perform at least one expansion operation—scenario expansion, expression transformation, anomaly construction, and reasoning enhancement—based on the task semantic skeleton, resulting in multiple subset datasets. For example, the generation device invokes multiple generative agents within the semantic constraint space defined by the domain business ontology, using a unified task semantic skeleton as the generation benchmark, and combining the concept hierarchy, attribute rules, and business constraints of the domain business ontology to perform differentiated sample expansion operations. These expansion operations can include various types such as scenario expansion, expression transformation, anomaly construction, and reasoning enhancement. Each generative agent completes the sample expansion operation based on the semantic constraints of the domain business ontology, enriching the sample content from different dimensions such as business scenario coverage, text expression forms, boundary anomaly cases, and task reasoning depth. This effectively overcomes the limitations of limited real test data quantity, single scenario, and insufficient type coverage, ultimately generating multiple multi-dimensional and differentiated subset datasets.
[0037] After obtaining multiple subsets of data, deduplication, format alignment, and merging can be performed on these subsets to obtain initial candidate evaluation samples. For example, the generation device performs global deduplication on all subsets, removing redundant samples with completely identical or highly similar content; then, it performs unified format alignment on subsets generated by different generative agents with different formats, standardizing text structure, task fields, entity representations, and output paradigms; finally, it summarizes and merges all processed subsets to obtain initial candidate evaluation samples with sufficient sample quantity, rich scene types, and diverse expression forms.
[0038] After obtaining the initial candidate evaluation samples, the verification agent can be invoked to perform semantic consistency analysis on the initial candidate evaluation samples based on the conceptual axioms and attribute constraints of the domain business ontology. Invalid sample entries with semantic illusions or business logic conflicts are removed, resulting in the final candidate evaluation samples. For example, the generation device invokes the verification agent, using the conceptual axioms, hierarchical constraints, and attribute association rules of the domain business ontology as the sole verification standard, to perform semantic consistency analysis and business logic verification on the initial candidate evaluation samples. This accurately identifies invalid sample entries that deviate from ontology semantic constraints, have conflicting business logic, or exhibit semantic illusions. By eliminating all conflicting, distorted, and non-compliant invalid samples, only valid samples that conform to ontology specifications, have rigorous business logic, and accurate semantic expression are retained, ultimately yielding high-quality, highly compliant candidate evaluation samples.
[0039] Reference Figure 4 , Figure 4 This is a flowchart illustrating another embodiment of the multi-agent agent evaluation dataset generation method of this application. In some embodiments, the multiple generating agents include a scene extension agent, an expression variant agent, an anomaly construction agent, and a reasoning enhancement agent. The aforementioned invocation of multiple generating agents, within the semantic constraint space of the domain business ontology, performs at least one of the following expansion operations: scene extension, expression transformation, anomaly construction, and reasoning enhancement, based on the task semantic skeleton, to obtain multiple subset datasets, including: Step S180: Invoke the scene extension agent to extend the business domain in the task semantic skeleton according to the concept classification tree and dependency relationship of the domain business ontology, and expand the related services with mutual dependencies to obtain the scene extension subset. Step S181: Invoke the expression variant agent to instantiate the ontology terms in the task semantic skeleton into various natural language expressions according to the language layer mapping rules of the domain business ontology, and obtain a subset of expression variants. Step S182: Invoke the exception construction agent to construct business boundary exception samples based on the legal exception sub-states of the domain business ontology and the task semantic skeleton, and obtain the exception construction subset. Step S183: Invoke the reasoning enhancement agent to automatically combine and nest the multi-dimensional dependency conditions of the task semantic skeleton based on the cross-knowledge source dependency relationship of the domain business ontology to generate complex task samples with high reasoning depth, and obtain the reasoning enhancement subset.
[0040] In this embodiment, as Figure 4As shown, during step S170, the scenario extension agent, expression variant agent, anomaly construction agent, and reasoning enhancement agent can be invoked to generate subset datasets respectively. The scenario extension agent is invoked to expand the business domain in the task semantic skeleton based on the concept classification tree and dependencies of the domain business ontology, and to expand related services with interdependent relationships, resulting in a scenario-extended subset dataset. For example, the generation device invokes the scenario extension agent, using the concept classification tree and dependencies of the domain business ontology as constraints, and based on the task semantic skeleton, to perform horizontal scenario extension and business branch expansion of the target business domain corresponding to the task semantic skeleton without deviating from the original core business semantics. This matches and expands auxiliary business services and related task scenarios that have dependencies on the core task, enriching the scenario coverage of the original task skeleton and forming a scenario-extended subset dataset with more comprehensive scenario coverage.
[0041] The expression variant agent is invoked to instantiate ontology terms in the task semantic skeleton into various natural language expressions based on the language layer mapping rules of the domain business ontology, resulting in a subset of expression variants. For example, the generation device invokes the expression variant agent, following the language layer mapping rules of the domain business ontology, keeping the business structure, user intent, task logic, and core entities of the task semantic skeleton unchanged, and only performs diversified natural language instantiation on the standardized ontology terms and normative sentence structures in the task semantic skeleton, generating multiple versions of task expressions that conform to real user expression habits, have different sentence styles, and are semantically equivalent, ultimately resulting in a subset of expression variants with rich expression forms and consistent semantics.
[0042] The anomaly construction agent is invoked to construct business boundary anomaly samples based on the legitimate anomaly sub-states of the domain business ontology and the task semantic skeleton, resulting in a subset of anomaly constructions. For example, the generation device invokes the anomaly construction agent to construct compliant boundary anomaly task samples within the anomaly range allowed by the semantic rules of the domain business ontology, based on the legitimate anomaly sub-states defined in the domain business ontology (e.g., approval timeout, permission denial), business boundary constraints, and anomaly triggering conditions, combined with the standard business process of the task semantic skeleton. This compensates for the deficiencies of the limited number and incomplete coverage of anomaly scenarios in the few samples of real test data, resulting in a subset of anomaly constructions covering various legitimate business anomaly scenarios.
[0043] The inference enhancement agent is invoked to automatically combine and nest multi-dimensional dependency conditions of the task semantic skeleton based on the cross-knowledge source dependency relationships of the domain business ontology, generating complex task samples with high inference depth, thus obtaining an inference-enhanced subset. For example, the generation device invokes the inference enhancement agent to reasonably combine and nest multi-dimensional dependency elements such as process conditions, resource constraints, role permissions, and pre-rules contained in the task semantic skeleton according to the cross-knowledge source dependency relationships of the domain business ontology, constructing complex business task samples with multi-layered logical relationships and high inference difficulty, effectively improving the inference complexity coverage of the final evaluation dataset, and obtaining an inference-enhanced subset with deeper inference levels and more complex logical structures.
[0044] Reference Figure 5 , Figure 5 This is a flowchart illustrating another embodiment of the multi-agent intelligent agent evaluation dataset generation method of this application. In some embodiments, the aforementioned calculation of the probability distribution difference between candidate evaluation samples and a small number of real test data includes: Step S190: Obtain multi-dimensional semantic features of the domain business ontology. The multi-dimensional semantic features include at least the subject domain class hierarchy, intent class hierarchy, complexity attribute, and user role class. Step S191: Construct a multi-dimensional semantic feature space based on multi-dimensional semantic features; Step S192: Map the candidate evaluation samples and the few sample real test data to the multi-dimensional semantic feature space respectively to obtain the first probability distribution of the candidate evaluation samples and the second probability distribution of the few sample real test data. Step S193: Call the Jensen-Shannon divergence algorithm or the Wasserstein distance algorithm to calculate the probability distribution difference between the first probability distribution and the second probability distribution in the multi-dimensional semantic feature space to obtain the probability distribution difference value.
[0045] In this embodiment, as Figure 5 As shown, when calculating the probability distribution difference between candidate evaluation samples and a small number of real test data in step S130, the multi-dimensional semantic features of the domain business ontology can be obtained first. The multi-dimensional semantic features of the domain business ontology include at least the subject domain class hierarchy, intent class hierarchy, complexity attribute, and user role class. For example, the generation device extracts and obtains the multi-dimensional semantic features inherent to the domain business ontology. These four types of semantic features are standardized semantic dimensions predefined by the domain business ontology and strongly bound to the target business scenario, comprehensively representing the scenario affiliation, user demands, difficulty level, and execution subject attributes of the business task.
[0046] A multi-dimensional semantic feature space is constructed based on multi-dimensional semantic features. For example, the generation device uses subject domain class hierarchy, intent class hierarchy, complexity attribute and user role class as basic dimensions to construct a unified multi-dimensional semantic feature space adapted to the target business scenario.
[0047] Candidate evaluation samples and a small number of real test data are mapped to a multi-dimensional semantic feature space, respectively, to obtain the first probability distribution of the candidate evaluation samples and the second probability distribution of the small number of real test data. For example, the generation device parses and extracts the multi-dimensional semantic features corresponding to each sample from the candidate evaluation samples and the small number of real test data, and maps them uniformly to the constructed multi-dimensional semantic feature space; then, by performing global feature statistics on the two types of sample sets, the first probability distribution corresponding to the candidate evaluation samples and the second probability distribution corresponding to the small number of real test data are fitted respectively.
[0048] The Jensen-Shannon divergence algorithm or the Wasserstein distance algorithm is used to calculate the probability distribution difference between the first and second probability distributions in the multi-dimensional semantic feature space, obtaining the probability distribution difference value. For example, the generation device uses the Jensen-Shannon divergence algorithm or the Wasserstein distance algorithm as a distribution difference measurement tool to quantitatively calculate the overall deviation between the first and second probability distributions in the multi-dimensional semantic feature space, accurately obtaining the probability distribution difference value between the two sets of data. The probability distribution difference value can objectively reflect the degree of fit between the candidate evaluation samples and a small number of real test data in core dimensions such as business semantics, scenario type, task difficulty, and role adaptation.
[0049] Reference Figure 6 , Figure 6 This is a flowchart illustrating another embodiment of the multi-agent agent evaluation dataset generation method of this application. In some embodiments, the aforementioned step of iteratively updating candidate evaluation samples based on the domain business ontology or calling the ontology introspective agent to update the domain business ontology and regenerating candidate evaluation samples includes: Step S200: Check whether the number of iterations for updating the candidate evaluation samples is greater than the preset number; Step S201: If the number of iterations is less than or equal to the preset number, then perform iterative updates of candidate evaluation samples based on the domain business ontology. Step S202: If the number of iterations is greater than the preset number, then execute the call to the ontology self-introspection agent to update the domain business ontology and regenerate candidate evaluation samples.
[0050] In this embodiment, as Figure 6As shown, when performing step S140, which involves iteratively updating candidate evaluation samples based on the domain business ontology or calling the ontology introspective agent to update the domain business ontology and regenerate candidate evaluation samples, it is possible to first check whether the number of iterations for updating candidate evaluation samples is greater than a preset number. For example, the preset number can be customized by the user according to actual conditions. The generation device first counts the number of consecutive iterations for the current candidate evaluation samples and compares the number of iterations with the preset number to check whether the number of iterations is greater than the preset number.
[0051] If the number of iterations is less than or equal to a preset number, then the candidate evaluation samples are iteratively updated based on the domain business ontology. For example, when the number of iterations is less than or equal to a preset number, the generation device can enter the inner iterative optimization branch and perform targeted iterative updates on the candidate evaluation samples based on the existing stable domain business ontology.
[0052] If the number of iterations exceeds a preset number, the ontology introspection agent is invoked to update the domain business ontology and regenerate candidate evaluation samples. For example, when the number of iterations exceeds a preset number, the generation device can switch to the outer ontology evolution branch, initiate the ontology introspection optimization mechanism, invoke the ontology introspection agent to adaptively update and expand the existing domain business ontology, and based on the updated and improved domain business ontology, re-execute the multi-agent dataset generation process to generate new candidate evaluation samples that adapt to the characteristics of real data distribution.
[0053] In some embodiments, the aforementioned iterative update of candidate evaluation samples based on domain business ontology includes: Within the instance space of the domain business ontology, data resampling, concept correction, or targeted regeneration operations are performed on candidate evaluation samples to iteratively update the candidate evaluation samples.
[0054] In this embodiment, during step S201, data resampling, concept correction, or targeted regeneration operations can be performed on the candidate evaluation samples. Within the instance space of the domain business ontology, these operations are performed to iteratively update the candidate evaluation samples. For example, the generation device selectively performs data resampling, concept correction, or targeted regeneration operations based on the distribution differences in the multi-dimensional semantic feature space. Data resampling primarily addresses the issue of uneven sample distribution across semantic dimensions in the candidate evaluation samples, supplementing scarce sample dimensions and downsampling redundant enriched samples to optimize the overall sample distribution density. Concept correction corrects issues such as concept shifts, inaccurate semantic matching, and incorrect ontology instance correspondences in the candidate evaluation samples, ensuring that the business concepts, entity affiliations, and intent classifications of each sample strictly conform to the domain business ontology definition. Targeted regeneration targets the imbalanced semantic dimensions with the highest contribution to the difference, supplementing them with high-quality, highly matched compliant samples based on the task semantic skeleton and ontology semantic constraints.
[0055] Reference Figure 7 , Figure 7 This is a flowchart illustrating another embodiment of the multi-agent agent evaluation dataset generation method of this application. In some embodiments, the aforementioned invocation of the ontology introspective agent to update the domain business ontology and regenerate candidate evaluation samples includes: Step S210: Based on the probability distribution difference value, perform a dimension-by-dimensional difference attribution analysis on the multi-dimensional semantic feature space to determine the bottleneck dimension with the highest difference contribution, map the bottleneck dimension back to the corresponding concept class in the domain business ontology, and generate a failure diagnosis report. Step S211: Invoke the ontology introspective agent to generate a constrained extended domain business ontology based on the failure diagnosis report and the OWL description logic specification. Step S212: Perform logical consistency, maintainability of existing instance classification, and monotonically decreasing distribution difference checks on the extended domain business ontology. Step S213: If any verification fails, discard the extended domain business ontology and keep the domain business ontology unchanged. Step S214: If all verifications are successful, the ontology changes that differ between the extended domain business ontology and the domain business ontology will be cascaded and propagated to the ontology semantic profile re-instantiation, task semantic skeleton re-extraction, and dataset difficulty stratification rule update to obtain a new task semantic skeleton. Candidate evaluation samples will then be regenerated based on the new task semantic skeleton.
[0056] In this embodiment, as Figure 7As shown, when executing step S202, a dimension-by-dimensional difference attribution analysis can be performed on the multi-dimensional semantic feature space based on the probability distribution difference value to determine the bottleneck dimension with the highest difference contribution. The bottleneck dimension is then mapped back to the corresponding concept class in the domain business ontology, and a failure diagnosis report is generated. For example, the generation device performs a dimension-by-dimensional attribution analysis on the distribution deviation of each semantic dimension in the multi-dimensional semantic feature space based on the probability distribution difference value, quantifies the contribution of each dimension to the overall distribution drift, and accurately locates the bottleneck dimension with the highest difference contribution that causes the sample distribution to fail to converge. Then, the bottleneck dimension is mapped back to the concept hierarchy structure of the domain business ontology to locate the corresponding ontology concept class, attribute constraint, or relation rule. Simultaneously, combined with the sample characteristics of long-standing sample adaptation anomalies and semantic matching failures, a standardized alignment failure diagnosis report is generated.
[0057] The ontology introspection agent is invoked to generate a constrained extended domain business ontology based on the failure diagnosis report and the OWL (Web Ontology Language) description logic specification. For example, the generation device invokes the ontology introspection agent, using the failure diagnosis report as the optimization basis, and strictly follows the OWL description logic specification. Without disrupting the original domain business ontology infrastructure and legitimate business constraints, it generates a restricted and controllable ontology extension optimization scheme, resulting in the extended domain business ontology. The ontology introspection agent precisely supplements and corrects only the weak conceptual systems corresponding to bottleneck dimensions, including adding subclasses, supplementing attribute relationships, and improving constraint axioms, ensuring that the ontology extension is targeted and business-reasonable, and avoiding arbitrary modifications to the ontology without constraints or basis.
[0058] The extended domain business ontology undergoes checks for logical consistency, the maintainability of existing instance classification, and the monotonically decreasing nature of distribution differences. For example, the generation device performs triple compliance checks on the extended domain business ontology: logical consistency check, maintainability of existing instance classification check, and monotonically decreasing nature of distribution differences check. Logical consistency check ensures that the extended domain business ontology does not contain conceptual conflicts, axiomatic contradictions, or semantic paradoxes. Maintainability of existing instance classification check ensures that the classification and semantic relationships of historically valid samples and original business instances are not disruptively altered. Monotonically decreasing nature of distribution differences check ensures that the current upgrade of the extended domain business ontology effectively improves sample distribution bias and achieves a positive iterative effect.
[0059] If any verification fails, the extended domain business ontology is discarded, while the original domain business ontology remains unchanged. For example, if any one of the three verifications fails, the extended domain business ontology is determined to be defective and unsuitable for deployment. The generation device then discards the extended domain business ontology, keeping the original domain business ontology unchanged.
[0060] If all verifications are successful, the changes to the extended domain business ontology that differ from the original domain business ontology are cascaded and propagated to the re-instantiation of the ontology semantic profile, the re-extraction of the task semantic skeleton, and the update of the dataset difficulty stratification rules to obtain a new task semantic skeleton. Candidate evaluation samples are then regenerated based on this new task semantic skeleton. For example, if all verification items are successful, the extension of the extended domain business ontology is confirmed to be compliant, effective, and possesses optimization gains. The generation device performs end-to-end cascading propagation of all changes between the old and new ontology, simultaneously completing the re-instantiation of the ontology semantic profile, the re-extraction of the task semantic skeleton, and the synchronous update of the dataset difficulty stratification rules, forming a new task semantic skeleton adapted to the new ontology system. Finally, based on the updated domain business ontology and the new task semantic skeleton, the multi-agent collaborative generation process is restarted to complete the generation of a new round of candidate evaluation samples.
[0061] In addition, the generation device can also call the description logic inference engine to perform consistency checks on the extended domain business ontology, checking that all existing semantic profile instances before the extension can still obtain valid concept classifications in the extended ontology; the distribution difference after the extension is strictly smaller than the distribution difference before the extension, ensuring that the ontology evolution sequence converges monotonically.
[0062] The generation device can also generate a new ontology version snapshot for each extended domain business ontology that passes verification, recording the differences in class sets, attribute sets, and corresponding distribution differences before and after the extension; the ontology version snapshot is used to support backtracking, so that when subsequent ontology extensions cause the distribution differences to increase, it can revert to the most recently verified version.
[0063] The technical solution of this application generates candidate evaluation samples for multiple agents through full-link constraints using domain business ontology and a small number of real test data. Combined with the determination of preset difference values corresponding to probability distribution difference values, it realizes closed-loop iterative optimization of candidate evaluation sample distribution. For candidate evaluation samples with unqualified distribution deviation, iterative updates are performed or the domain business ontology adaptively evolves and regenerates candidate evaluation samples. Qualified candidate evaluation samples are processed hierarchically to obtain multi-level agent evaluation datasets. This realizes the automated generation, quality verification and distribution alignment optimization of agent evaluation datasets, significantly improving the authenticity, rationality and domain adaptability of agent evaluation datasets.
[0064] This application further proposes a device for generating intelligent agent evaluation datasets based on multi-agent systems, referring to... Figure 8 , Figure 8 This is a schematic diagram of an embodiment of the multi-agent-based intelligent agent evaluation dataset generation device of this application. In some embodiments, the multi-agent-based intelligent agent evaluation dataset generation device includes: At least one processor; and, A memory that is communicatively connected to at least one processor; wherein, The memory stores instructions that are executed by the at least one processor, which enable the at least one processor to perform the multi-agent-based agent evaluation dataset generation method described above.
[0065] In this embodiment, as Figure 8 As shown, the multi-agent-based intelligent agent evaluation dataset generation device in this application embodiment can be a processor capable of running a multi-agent-based intelligent agent evaluation dataset generation method; there is at least one processor. Figure 8 As shown, the multi-agent-based intelligent agent evaluation dataset generation device may include: a processor 1001 (e.g., CPU), a network interface 1004, a user interface 1003, a memory 1005, and a communication bus 1002. The communication bus 1002 is used to establish communication between these components. The user interface 1003 may include a display screen and an input unit, such as a keyboard; optionally, the user interface 1003 may also include a standard wired interface or a wireless interface. The network interface 1004 may optionally include a standard wired interface or a wireless interface (such as a Wi-Fi interface). The memory 1005 may be a high-speed RAM or a stable, non-volatile memory, such as a disk drive. Optionally, the memory 1005 may also be a storage device independent of the aforementioned processor 1001.
[0066] Those skilled in the art will understand that Figure 8 The structure of the multi-agent-based intelligent agent evaluation dataset generation device shown does not constitute a limitation on the multi-agent-based intelligent agent evaluation dataset generation device. It may include more or fewer components than shown, or combine certain components, or have different component arrangements.
[0067] like Figure 8 As shown, the memory 1005, which serves as a computer storage medium, may include an operating system, a network communication module, a user interface module, and computer-executable instructions.
[0068] exist Figure 8In the multi-agent-based intelligent agent evaluation dataset generation device shown, the network interface 1004 is mainly used to connect to the backend server and communicate with the backend server; the user interface 1003 is mainly used to connect to the client (user end) and communicate with the client; and the processor 1001 can be used to call the computer-executable instructions stored in the memory 1005. When the instructions are called and executed by the processor 1001, they implement the steps of the multi-agent-based intelligent agent evaluation dataset generation method described above.
[0069] This application further proposes a storage medium storing a computer program, the computer program including program instructions, which, when executed by a processor, enable the processor to execute the above-described method for generating an intelligent agent evaluation dataset based on multi-agent systems.
[0070] The above description is only a part or preferred embodiment of this application. Neither the text nor the drawings should limit the scope of protection of this application. All equivalent structural transformations made using the content of this application's specification and drawings under the overall concept of this application, or direct / indirect applications in other related technical fields, are included within the scope of protection of this application.
Claims
1. A method for generating an intelligent agent evaluation dataset based on multi-agent systems, characterized in that, include: Obtain the domain business ontology of the target business scenario and a small number of real test data, and extract the task semantic skeleton based on the domain business ontology from the small number of real test data. Multiple generative agents are invoked to generate multiple sub-datasets based on the task semantic skeleton, and the multiple sub-datasets are integrated to obtain candidate evaluation samples; Calculate the probability distribution difference between the candidate evaluation samples and the few sample real test data, and determine whether the probability distribution difference is greater than a preset difference value; If the probability distribution difference value is greater than the preset difference value, then the candidate evaluation samples are iteratively updated according to the domain business ontology or the ontology introspective agent is called to update the domain business ontology and regenerate the candidate evaluation samples, and the probability distribution difference value between the candidate evaluation samples and the few sample real test data is recalculated. If the probability distribution difference value is less than or equal to the preset difference value, the candidate evaluation samples are processed in layers according to the domain business ontology to obtain a multi-level intelligent agent evaluation dataset. The step of iteratively updating candidate evaluation samples based on the domain business ontology or calling the ontology introspective agent to update the domain business ontology and regenerate candidate evaluation samples includes: Check whether the number of iterations for updating candidate evaluation samples exceeds the preset number; If the number of iterations is less than or equal to the preset number, then the candidate evaluation samples are iteratively updated according to the domain business ontology. If the number of iterations is greater than the preset number, then the ontology self-introspection agent is invoked to update the domain business ontology and regenerate candidate evaluation samples. The iterative update of candidate evaluation samples based on the domain business ontology includes: In the instance space of the domain business ontology, data resampling, concept correction, or targeted regeneration operations are performed on the candidate evaluation samples to iteratively update the candidate evaluation samples; The step of invoking the ontology introspection agent to update the domain business ontology and regenerate candidate evaluation samples includes: Based on the probability distribution difference value, a dimension-by-dimensional difference attribution analysis is performed on the multi-dimensional semantic feature space to determine the bottleneck dimension with the highest difference contribution. The bottleneck dimension is then mapped back to the corresponding concept class in the domain business ontology, and a failure diagnosis report is generated. The ontology introspective agent is invoked to generate a constrained extended domain business ontology based on the failure diagnosis report and the OWL description logic specification. The extended domain business ontology is checked for logical consistency, maintainability of existing instance classification, and monotonically decreasing distribution differences. If any verification fails, the extended domain business ontology is discarded, while the domain business ontology remains unchanged. If all verifications are successful, the ontology changes that differ between the extended domain business ontology and the domain business ontology are cascaded and propagated to the ontology semantic profile re-instantiation, task semantic skeleton re-extraction, and dataset difficulty stratification rule update to obtain a new task semantic skeleton, and candidate evaluation samples are regenerated based on the new task semantic skeleton.
2. The method for generating an agent evaluation dataset based on multi-agent systems according to claim 1, characterized in that, The step of extracting the task semantic skeleton from the few-sample real test data based on the domain business ontology includes: Based on the class hierarchy and attribute definition of the domain business ontology, semantic parsing is performed on each piece of real test data in the few sample real test data, and the business domain, user intent, involved entities and task types are identified to obtain the identification results of each piece of real test data. Each of the identification results is instantiated into an entity in the domain business ontology, and the abstract concept class to which each entity belongs is inferred to obtain an ontology semantic profile; Identify the task pattern class to which the ontology semantic profile belongs in the conceptual hierarchy of the domain business ontology, extract the process steps, participating roles, required resource types and preconditions associated with the task pattern class, and convert them into a formal skeleton representation to obtain the task semantic skeleton.
3. The method for generating an agent evaluation dataset based on multi-agent systems according to claim 2, characterized in that, The process of calling multiple generative agents to generate multiple sub-datasets based on the task semantic skeleton, and integrating the multiple sub-datasets to obtain candidate evaluation samples includes: Multiple generated intelligent agents are invoked to perform at least one of the following expansion operations—scenario expansion, expression transformation, anomaly construction, and reasoning enhancement—within the semantic constraint space of the domain business ontology, respectively, to obtain multiple subset datasets. The multiple subset datasets are deduplicated, format-aligned, and merged to obtain initial candidate evaluation samples; The verification agent is invoked to perform semantic consistency parsing on the initial candidate evaluation samples based on the conceptual axioms and attribute constraints of the domain business ontology, removing invalid sample entries that contain semantic illusions or conflicting business logic, thereby obtaining the candidate evaluation samples.
4. The method for generating an agent evaluation dataset based on multi-agent systems according to claim 3, characterized in that, The multiple generative agents include a scene extension agent, an expression variant agent, an anomaly construction agent, and a reasoning enhancement agent; the invocation of the multiple generative agents, within the semantic constraint space of the domain business ontology, performs at least one of the following expansion operations: scene extension, expression transformation, anomaly construction, and reasoning enhancement, respectively, based on the task semantic skeleton, to obtain multiple subset datasets, including: The scenario extension agent is invoked to expand the business domain in the task semantic skeleton based on the concept classification tree and dependency relationship of the domain business ontology, and to expand the related services with mutual dependencies to obtain the scenario extension subset. The expression variant agent is invoked to instantiate the ontology terms in the task semantic skeleton into various natural language expressions according to the language layer mapping rules of the domain business ontology, thereby obtaining a subset of expression variants. The exception construction agent is invoked to construct business boundary exception samples based on the legitimate exception sub-states of the domain business ontology and the task semantic skeleton, thereby obtaining a subset of exception construction data. The reasoning-enhanced agent is invoked to automatically combine and nest the multi-dimensional dependency conditions of the task semantic skeleton based on the cross-knowledge source dependency relationships of the domain business ontology, thereby generating complex task samples with high reasoning depth and obtaining a reasoning-enhanced subset.
5. The method for generating an agent evaluation dataset based on multi-agent systems according to claim 4, characterized in that, The calculation of the probability distribution difference between the candidate evaluation samples and the few-sample real test data includes: Obtain multi-dimensional semantic features of the domain business ontology, wherein the multi-dimensional semantic features include at least the subject domain class hierarchy, intent class hierarchy, complexity attribute, and user role class; Construct a multi-dimensional semantic feature space based on the aforementioned multi-dimensional semantic features; The candidate evaluation samples and the few-sample real test data are respectively mapped to the multi-dimensional semantic feature space to obtain the first probability distribution of the candidate evaluation samples and the second probability distribution of the few-sample real test data. The probability distribution difference value is obtained by calling the Jensen-Shannon divergence algorithm or the Wasserstein distance algorithm to calculate the probability distribution difference between the first probability distribution and the second probability distribution in the multi-dimensional semantic feature space.
6. A device for generating an intelligent agent evaluation dataset based on multi-agent systems, characterized in that, include: At least one processor; as well as, A memory communicatively connected to the at least one processor; wherein, The memory stores instructions that are executed by the at least one processor to enable the at least one processor to perform the multi-agent-based agent evaluation dataset generation method according to any one of claims 1 to 5.
7. A storage medium, characterized in that, The storage medium stores a computer program, which includes program instructions that, when executed by a processor, enable the processor to perform the multi-agent-based agent evaluation dataset generation method according to any one of claims 1 to 5.