A dataset-oriented multi-source parsing operator evaluation method

By adopting a multi-source parsing operator-based evaluation method, the problem of difficulty in horizontal comparison and vertical tracking of dataset evaluation results is solved, the auditability and comparability of evaluation results are achieved, the evaluation cost is reduced, the iteration efficiency is improved, and the compliance requirements of sensitive data are met.

CN122309492APending Publication Date: 2026-06-30BEIJING ZHONGLIANG INTELLIGENT NUMBER TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-03-18
Publication Date
2026-06-30

AI Technical Summary

Technical Problem

Existing dataset evaluation techniques make it difficult to compare and track evaluation results across standards, versions, and deployment environments, lacking verifiability and failing to achieve auditability and sustainable quality governance under sensitive data governance.

Method used

The multi-source parsing operator-based evaluation method is adopted, which generates an auditable evaluation report through access registration, evaluation protocol instance determination, capability component screening and consistency verification, calibrated signature generation, execution plan binding, result consistency verification and comparability calibration.

Benefits of technology

It achieves the solidification and verifiability of evaluation criteria, reduces evaluation costs, improves iteration efficiency, ensures the comparability and interpretability of evaluation results across different environments, and meets the compliance requirements for sensitive data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122309492A_ABST
    Figure CN122309492A_ABST
Patent Text Reader

Abstract

This invention discloses a multi-source parsing operator-based evaluation method for datasets, belonging to the field of dataset evaluation technology. The method includes: based on the acquired target dataset and authorization information, forming auditable access records; constructing a resource list, a basic feature summary, and an evaluation boundary description; determining evaluation protocol instances; determining a capability component list; generating a caliber signature; generating an execution plan; generating auditable execution records and a capability component output set; performing result aggregation, comparability calibration, and issue ledger generation to obtain evaluation results, an issue ledger, and result verification records; and generating an evaluation report based on the evaluation results, issue ledger, execution plan, caliber signature, capability component list, evaluation boundary description, and result verification records. This invention enables the evaluation caliber to be fixed and verifiable, and the evaluation results to be reproducible and auditable. It can reduce the evaluation cost during updates and iterations, improve iteration efficiency, and provide comparability across cross-environment scenarios.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to a multi-source parsing operator-based evaluation method for datasets, belonging to the field of dataset evaluation technology. Background Technology

[0002] Existing dataset quality assessment systems are typically built around data access, rule or model detection, result aggregation, and report output, enabling a degree of automation in assessment. However, existing technologies still generally suffer from problems in large-scale assessment scenarios that cross standards, versions, and deployment environments.

[0003] The existing technologies have the following characteristics when evaluating datasets: (1) The evaluation criteria mainly rely on manual configuration, project document agreements or solidified individual experience. The meaning of indicators, thresholds, and labels are prone to drift between different projects or different stages, making it difficult to compare evaluation results horizontally and track them vertically. (2) The evaluation process is organized around task scheduling or process orchestration, but it lacks a formal definition of "evaluation identity" and a verifiable description of reproducible boundaries, making it difficult to provide verifiable evidence in audit, review, or dispute resolution scenarios. (3) The system often tightly couples parsing, reasoning, scoring, and suggestion generation capabilities with the platform process. When capabilities are upgraded, replaced, or parameters are changed, it is difficult to attribute and explain the differences between historical and new results, affecting the continuity and comparability of results. (4) Under the increasingly stringent conditions of sensitive data governance and compliance requirements, constraints such as data not leaving the domain, minimum visibility, desensitization and trimming, and on-site execution are difficult to incorporate into the formal rule system of the evaluation process, making it difficult to audit and prove compliance boundaries. (5) The evaluation output is mostly presented in the form of reports or statistical summaries, lacking a structured problem list, regression verification conditions and status flow mechanism for rectification closure, making it difficult to form a sustainable quality governance process.

[0004] Based on the shortcomings of the existing technologies, there is an urgent need in the field of dataset evaluation technology for a dataset evaluation method that can solidify and verify evaluation criteria, reproduce and audit evaluation results, reduce evaluation costs during updates and iterations, improve iteration efficiency, and be comparable across different environments. Summary of the Invention

[0005] The technical problem to be solved by the present invention is to provide a multi-source parsing operator-based evaluation method and system for datasets, which solves the problems of existing technologies in evaluating datasets, such as difficulty in horizontal comparison and vertical tracking of evaluation results, difficulty in verification and interpretation, and difficulty in forming a sustainable quality governance process.

[0006] To achieve the above objectives, the present invention employs the following technical solution: This invention discloses a multi-source analytical operator-based evaluation method for datasets, comprising: Based on the acquired target dataset and authorization information, the target dataset is registered for access, and the data source, access method, permission constraints and version boundaries are recorded to form an auditable access record. Based on the auditable access record, a resource list, basic feature summary and evaluation boundary description are constructed. The evaluation protocol instance is determined based on the obtained evaluation objectives and applicable standards; Candidate capability components are selected based on the evaluation protocol instance, resource list, and basic feature summary, and the capability component list is determined after consistency verification of the interchange interface. A caliber signature is generated based on the evaluation protocol instance, the list of capability components, and the execution constraints and consistency verification constraints corresponding to the evaluation protocol instance; the caliber signature is used for execution plan binding, result consistency verification, and change impact analysis; An execution plan is generated based on the evaluation protocol instance, caliber signature, evaluation boundary description, and capability component interface information determined according to the capability component list. The evaluation protocol instance is executed according to the execution plan, the list of capability components, and the compliance constraints determined by the evaluation protocol instance, generating auditable execution records and a set of capability component outputs; Based on the output set of the capability components, the caliber signature, the execution record, and the comparability calibration requirements determined by the evaluation protocol instance, the results are aggregated, the comparability calibration is performed, and the problem ledger is generated to obtain the evaluation results, the problem ledger, and the result verification record. An evaluation report is generated based on the evaluation results, issue ledger, execution plan, signature, capability component list, evaluation boundary description, and result verification record.

[0007] Furthermore, the process of registering access to the target dataset, recording the data source, access method, permission constraints, and version boundaries to form an auditable access record, and constructing a resource list, basic feature summary, and evaluation boundary description based on the auditable access record, includes: Record the data source, access method, permission constraints, and version boundaries of the target dataset to form an auditable access record; the data source includes files, directories, object storage, databases, or dataset version references; A resource list is constructed based on the resource objects identified in the auditable access records. The resource list is a set of resource objects in the target dataset that can be evaluated and processed. Based on the format, scale, modality, and partition information of the resource objects corresponding to the resource list, extract basic feature summaries; The evaluation boundary is determined based on the resource list, permission constraints, and version boundaries.

[0008] Furthermore, the step of determining the evaluation protocol instance based on the acquired evaluation objectives and applicable standards includes: The evaluation objectives and applicable standards include: indicator definitions, threshold definitions, label definitions, aggregation rules, compliance constraints, budget constraints, coverage objectives, confidence objectives, and comparability calibration requirements; wherein, the comparability calibration requirements are calibration rule requirements pre-configured in the evaluation protocol instance according to the evaluation objectives and applicable standards. Based on the indicator definition, threshold definition, label definition, aggregation rules, compliance constraints, budget constraints, coverage targets, confidence targets and comparability calibration requirements, select or generate evaluation protocol instances from the protocol library; In scenarios where multiple standards coexist, mapping and conflict resolution are performed between different standards, and the mapping basis and difference descriptions are recorded to obtain a protocol difference description; the multiple standards include industry specifications, enterprise specifications, and project agreements.

[0009] Further, the step of filtering candidate capability components based on the evaluation protocol instance, resource list, and basic feature summary, and determining the capability component list after consistency verification of the interchange interface, includes: The input requirements of the evaluation protocol are determined based on the evaluation protocol instance, and candidate external capability components that can process the corresponding resource objects are screened by combining the resource object types represented by the resource list and the format, scale, modality and partition information represented by the basic feature summary. Perform interface consistency checks on external capability components, including: input view constraints, output structure constraints, error caliber constraints, version declarations and compatibility rules; when the checks are not met, trigger replacement, degradation or isolation strategies. Based on the candidate external capability components that have passed the consistency verification of the interchange interface, determine the list of capability components and record the interface verification results and the triggering of replacement / degradation / isolation strategies.

[0010] Further, the step of generating a signature based on the evaluation protocol instance, the capability component list, and the execution constraints and consistency verification constraints corresponding to the evaluation protocol instance includes: Generate caliber signatures for key caliber elements of the evaluation protocol instance. These caliber signatures are used for execution plan binding, result consistency verification, and change impact analysis. The execution constraints corresponding to the evaluation protocol instance include resource constraints, timeout constraints, and budget constraints; the consistency verification constraints include input view constraints, output structure constraints, error caliber constraints, and version declaration and compatibility rules. Based on the execution constraints, input view constraints, output structure constraints, error caliber constraints, and version declarations and compatibility rules, determine whether the results are comparable and whether they need to be re-evaluated or marked as incomparable when the data version, protocol version, capability version, or execution constraints change, and obtain a list of consistency verification points.

[0011] Further, the step of generating an execution plan based on the evaluation protocol instance, caliber signature, evaluation boundary description, and capability component interface information determined according to the capability component list includes: Based on the requirements of the evaluation protocol instance, and combined with the evaluation scope and version boundaries defined by the evaluation boundary description, as well as the calling methods, input-output mappings and compatibility relationships determined by the capability component interface information, an auditable execution plan is generated. The execution plan includes: determining the execution path, parallel granularity, triggering conditions, budget constraints, sampling strategies, compliance trimming rules, failure handling and degradation strategies, and incremental re-evaluation boundaries. Bind auditable execution plans to caliber signatures.

[0012] Further, the step of executing the evaluation protocol instance according to the execution plan, the list of capability components, and the compliance constraints determined by the evaluation protocol instance, and generating auditable execution records and a set of capability component outputs, includes: The evaluation of the evaluation protocol instance is carried out according to the execution plan, including: determining the set of capability component outputs for execution based on the capability component list and the execution plan; dynamically governing budget constraints, coverage targets and confidence targets in combination with the set of capability component outputs; and implementing compliance tailoring and on-site execution strategies based on the compliance constraints. Record the strategy triggers, degradation paths, error reasons, retry information and execution trajectory during the execution process to form an auditable execution log.

[0013] Furthermore, the process of aggregating results, performing comparability calibration, and generating an issue ledger based on the output set of the capability components, the caliber signature, the execution record, and the comparability calibration requirements determined by the evaluation protocol instance, to obtain evaluation results, an issue ledger, and result verification records, includes: Based on the signature, the structural and caliber consistency of the output set of the capability components is verified to obtain the result verification record. Based on the results of structural consistency and caliber consistency verification, a comparable score and a comparable conclusion are generated according to the calibration rules corresponding to the scale anchoring and the aforementioned comparability calibration requirements. Based on comparable scores and comparable conclusions, hierarchical evaluation results are generated by aggregating them according to the evaluation protocol. Issues identified during structural consistency verification, calibrator consistency verification, comparability calibration, and aggregation according to the evaluation protocol are solidified into an issue ledger in the form of structured entries. The issue entries include label calibrator, scope of impact, location clues, rectification suggestions, and regression verification conditions, and are associated with the evaluation snapshot. The evaluation snapshot includes dataset version boundary, evaluation protocol version, calibrator signature, execution plan, capability component version set, execution constraints, and execution records.

[0014] Furthermore, the generation of the evaluation report based on the evaluation results, issue ledger, execution plan, caliber signature, capability component list, evaluation boundary description, and result verification record includes: Solidification evaluation snapshot; When a change in dataset version, evaluation protocol version, or capability component version is detected, change impact analysis is performed based on consistency checkpoints and an impact set is generated. Incremental review is triggered on the impact set, consistency verification is performed on the reuse boundary, and difference attribution information is output. The reuse boundary is the resource range, result range, and corresponding constraint range that can be directly reused after the change impact analysis. The difference attribution information is structured information that characterizes the source and impact range of the difference between the review result and the historical result, and is used to distinguish whether the difference comes from data change, caliber change, capability version change, or compliance trimming change. An evaluation report is generated based on the evaluation snapshot and differential attribution information.

[0015] Furthermore, it also includes: A small-scale trial run of the execution plan was conducted to estimate anomaly density, coverage achievement, and resource consumption. Under the premise of meeting the budget constraints, coverage targets, confidence targets and compliance constraints preset in the evaluation protocol instance, the sampling ratio, budget upper limit and trigger threshold parameters are calibrated, and the basis for cost and confidence interpretation is formed; if adjustments are made, the reasons for the adjustments and the differences before and after the adjustments are recorded.

[0016] The beneficial effects achieved by this invention are as follows: The method provided by this invention, through evaluation protocols and caliber signatures, enables the evaluation caliber to be fixed and verifiable, reducing the risk of incomparability and disputes caused by caliber drift; through evaluation snapshots and execution records, it supports consistency verification of re-evaluation and audit arbitration, clarifies the sources of differences, and reduces evaluation costs and improves iteration efficiency during updates and iterations through change impact analysis and incremental evaluation; through scale anchoring and comparability calibration, it ensures that scores and conclusions are comparable and interpretable across projects, time periods, and environmental scenarios; through compliance execution boundary control, it meets the requirements of data not leaving the domain and minimum visibility in sensitive data scenarios, and the process and output are auditable; through issue ledgers and regression verification mechanisms, it ensures that the evaluation results are executable, traceable, and verifiable. Attached Figure Description

[0017] Figure 1 This is a flowchart illustrating a multi-source analytical operator-based evaluation method for datasets provided by the present invention.

[0018] Figure 2 This invention provides an architecture diagram of a multi-source parsing operator-based evaluation system for datasets. Detailed Implementation

[0019] The present invention will be further described below with reference to the accompanying drawings. The following embodiments are only used to more clearly illustrate the technical solution of the present invention, and should not be used to limit the scope of protection of the present invention.

[0020] Example 1, as Figure 1 As shown, this invention discloses a multi-source analytical operator-based evaluation method for datasets, comprising: Based on the acquired target dataset and authorization information, the target dataset is registered for access, and the data source, access method, permission constraints and version boundaries are recorded to form an auditable access record. Based on the auditable access record, a resource list, basic feature summary and evaluation boundary description are constructed. The evaluation protocol instance is determined based on the obtained evaluation objectives and applicable standards; Candidate capability components are selected based on the evaluation protocol instance, resource list, and basic feature summary, and the capability component list is determined after consistency verification of the interchange interface. A caliber signature is generated based on the evaluation protocol instance, the list of capability components, and the execution constraints and consistency verification constraints corresponding to the evaluation protocol instance; the caliber signature is used for execution plan binding, result consistency verification, and change impact analysis; An execution plan is generated based on the evaluation protocol instance, caliber signature, evaluation boundary description, and capability component interface information determined according to the capability component list. The evaluation protocol instance is executed according to the execution plan, the list of capability components, and the compliance constraints determined by the evaluation protocol instance, generating auditable execution records and a set of capability component outputs; Based on the output set of the capability components, the caliber signature, the execution record, and the comparability calibration requirements determined by the evaluation protocol instance, the results are aggregated, the comparability calibration is performed, and the problem ledger is generated to obtain the evaluation results, the problem ledger, and the result verification record. An evaluation report is generated based on the evaluation results, issue ledger, execution plan, signature, capability component list, evaluation boundary description, and result verification record.

[0021] The setup process for the evaluation protocol layer includes: Based on the evaluation objectives and applicable standards, select or generate evaluation protocol instances from the protocol library and configure evaluation elements; the evaluation elements include evaluation scope, indicator caliber, label caliber, aggregation rules, compliance constraints, budget constraints, coverage objectives, confidence objectives, and comparability calibration requirements; the comparability calibration requirements are calibration rule requirements pre-configured according to the evaluation objectives and applicable standards; In scenarios where multiple standards coexist, such as industry norms, corporate norms, and project agreements, the mapping and conflict resolution between different standards should be completed, and the basis for the mapping and explanations of the differences should be recorded.

[0022] The process of driving the evaluation run with an evaluation protocol instance to obtain the corresponding execution plan includes: Based on the input requirements of the evaluation protocol, and combined with the resource object types represented by the resource list and the format, scale, modality, and partition information represented by the basic feature summary, candidate external capability components capable of processing the corresponding resource objects are screened. The external capability components are then subjected to interface consistency verification to obtain a capability component list and interface verification results. The verification includes input view constraints, output structure constraints, error caliber constraints, and version declarations and compatibility rules. When the verification is not satisfied, replacement, degradation, or isolation strategies are triggered and a record is generated. Generate caliber signatures for key caliber elements of the evaluation protocol. These caliber signatures are used for execution plan binding, result consistency verification, and change impact analysis, and consistency checkpoints are set. An execution plan is generated, and the execution plan is tested and calibrated to obtain a calibrated execution plan.

[0023] The generated execution plan includes: Clearly define the execution path, parallel granularity, triggering conditions, budget constraints, sampling strategies, compliance trimming rules, failure handling and degradation strategies, and incremental review boundaries; The execution plan is linked to the signature of the relevant authority, serving as the basis for subsequent execution and auditing.

[0024] The trial run and calibration of the execution plan includes: Conduct small-scale trial runs to estimate abnormal density, coverage achievement, and resource consumption; Under the premise of meeting the preset budget constraints, coverage targets, confidence targets, and compliance constraints in the evaluation protocol instance, the evaluation parameters are calibrated to form the basis for cost and confidence interpretation; the evaluation parameters include sampling ratio, budget upper limit, and trigger threshold; For parameters that have been calibrated, record the reasons for the adjustment and the differences before and after the adjustment; Output calibration records and post-calibration execution plans.

[0025] The evaluation performed according to the execution plan, and the resulting execution record, includes: The evaluation is carried out according to the execution plan, and the budget constraints, coverage targets and confidence targets are dynamically managed during runtime. Implement compliance tailoring and on-site execution strategies based on the aforementioned compliance constraints; Record policy triggers, degradation paths, error reasons, retry information, and execution traces to obtain an auditable set of execution records and capability component outputs.

[0026] The output results of the convergence evaluation are verified and analyzed to obtain a difference attribution report, including: The evaluation outputs are aggregated, and the outputs are calibrated for comparability to generate an issue ledger. The evaluation outputs include a set of capability component outputs, caliber signatures, execution records, and comparability calibration requirements determined by the evaluation protocol instance. Solidify the evaluation snapshot, and conduct change impact analysis and incremental review.

[0027] The process of performing comparability calibration on the output results and generating a problem ledger includes, The structural consistency and caliber consistency of the evaluation output results are verified. Based on the calibration rules corresponding to the scale anchoring and comparability calibration requirements, generate comparable scores and comparable conclusions, and retain the calibration basis; The hierarchical evaluation results are generated by aggregating according to the evaluation protocol. The issues identified during the process of structural consistency verification, calibrator consistency verification, comparability calibration, and aggregation according to the evaluation protocol are solidified into an issue ledger in the form of structured entries. The issue entries include label calibrators, scope of impact, location clues, rectification suggestions and regression verification conditions, and are associated with the evaluation snapshot. Output stratified evaluation results, problem log, and result verification records, including comparability explanations and calibration basis.

[0028] The solidified evaluation snapshot is used for change impact analysis and incremental review, including: The evaluation snapshot is fixed, and the evaluation snapshot includes the dataset version boundary, evaluation protocol version, caliber signature, execution plan, capability component version set, execution constraints, and execution records; When a change in data version, evaluation protocol version, or capability component version is detected, a change impact analysis is performed based on the consistency checkpoint, and an impact set is generated. Incremental review is triggered for the aforementioned impact set, and consistency verification is performed on the reuse boundary; wherein, the reuse boundary is the resource range, result range, and corresponding constraint range that are determined to be directly reusable from historical evaluation results after the change impact analysis; Output evaluation snapshots, reproduction verification instructions, incremental re-evaluation results, and difference attribution reports.

[0029] Example 2, as Figure 2 This embodiment introduces a multi-source analytical operator-based evaluation method for datasets provided by the present invention, including: S101. Dataset access and evaluation boundary determination, including: Input: Data source and authorization information, including files, directories, object storage, databases, or system dataset version references.

[0030] Processing: Register data sources, access methods, permission constraints, and version boundaries to form auditable access records; construct a resource list based on the resource objects identified in the auditable access records, where the resource list is a set of resource objects in the target dataset that can be evaluated and processed; extract basic feature summaries based on the format, scale, modality, and partition information of the resource objects corresponding to the resource list; and combine the resource list, permission constraints, and version boundaries to form an evaluation boundary description.

[0031] Output: Access records, resource list, basic feature summary, and evaluation boundary description.

[0032] S102, Evaluation Protocol Loading and Standard Mapping, including: Input: Evaluation objectives and applicable standards, including industry standards, corporate standards, or project agreements.

[0033] Processing: Select or generate evaluation protocol instances from the protocol library; configure indicator caliber, threshold caliber, label caliber, aggregation rules, compliance constraints, budget constraints, coverage targets, confidence targets, and comparability calibration requirements. Among them, the comparability calibration requirements are calibration rule requirements pre-configured in the evaluation protocol instance according to the evaluation objectives and applicable standards; in scenarios where multiple standards such as industry specifications, enterprise specifications, and project agreements coexist, complete the mapping and conflict handling between different standards, and record the mapping basis and difference descriptions.

[0034] Output: Evaluation protocol example, protocol difference description, including migration or update.

[0035] S103, Verification of Capability Component Selection and Interchange Interface, including: Input: Evaluation protocol instance, resource list, and basic feature summary.

[0036] Processing: Based on the input requirements of the evaluation protocol, and combined with the resource object types represented by the resource list and the format, scale, modality, and partition information represented by the basic feature summary, candidate external capability components that can process the corresponding resource objects are selected; the external capability components are subjected to interface consistency verification, including input view constraints, output structure constraints, error caliber constraints, version declarations, and compatibility rules; when the verification is not satisfied, replacement, degradation, or isolation strategies are triggered and a record is generated.

[0037] Output: a list of capability components, interface verification results, and records of replacement / downgrade / isolation strategies. The list of capability components includes version and compatibility information.

[0038] S104. Criterion signature generation and consistency verification point solidification, including: Input: Evaluation protocol instance, capability component list, execution constraints and consistency verification constraints. The execution constraints include resources, timeouts, budgets, etc., and the consistency verification constraints include input view constraints, output structure constraints, error caliber constraints, and version declarations and compatibility rules.

[0039] Processing: Generate caliber signatures for key caliber elements of the evaluation protocol. These caliber signatures are used for execution plan binding, result consistency verification, and change impact analysis. Consistency verification points are solidified based on execution constraints and consistency verification constraints. These points are used to determine whether the results are comparable and whether re-evaluation or marking as incomparable is required when data version, protocol version, capability version, or execution constraints change.

[0040] Output: caliber signature, list of consistency checkpoints.

[0041] S105, Execution plan generation, including: Input: Evaluation protocol instance, caliber signature, evaluation boundary description, capability component interface information.

[0042] Processing: Based on the requirements of the evaluation protocol instance, and combined with the evaluation scope and version boundaries defined in the evaluation boundary description, as well as the calling methods, input / output mappings, and compatibility relationships determined by the capability component interface information, an auditable execution plan is generated, including: determining the execution path, parallel granularity, triggering conditions, budget constraints, sampling strategies, compliance trimming rules, failure handling and degradation strategies, and incremental review boundaries; the execution plan is bound to the caliber signature, serving as the basis for subsequent execution, result consistency verification, and change impact analysis.

[0043] Output: Execution plan, plan description, which includes key constraints and degradation paths.

[0044] S106. Trial run calibration and cost explainability; Input: Execution plan, representative subset of data, or sampling rules.

[0045] Handling: Conduct small-scale trial runs to estimate anomaly density, coverage achievement, and resource consumption; under the premise of meeting the budget constraints, coverage targets, confidence targets, and compliance constraints preset in the evaluation protocol instance, calibrate parameters such as sampling ratio, budget upper limit, and trigger threshold, and form the basis for cost and confidence interpretation; if adjustments are made, record the reasons for the adjustments and the differences before and after the adjustments.

[0046] Output: Calibration record, calibrated execution plan (if adjusted).

[0047] S107. Controlled execution and audit records, including: Inputs: Execution plan (including calibrated execution plan), list of capability components, and compliance constraints determined by the evaluation protocol instance.

[0048] Processing: Conduct evaluation according to the execution plan; dynamically manage budget constraints, coverage targets, and confidence targets during runtime; implement compliance tailoring and on-site execution strategies based on the aforementioned compliance constraints; record strategy triggers, degradation paths, error reasons, retry information, and execution trajectories to form auditable execution records.

[0049] Output: Execution records and a collection of capability components that conform to the evaluation protocol constraints.

[0050] S108, Results Aggregation, Comparability Calibration, and Problem Ledger Generation, including: Inputs: Capability component output set, caliber signature, execution record, and comparability calibration requirements determined by the evaluation protocol instance.

[0051] Processing: Verify the structural consistency and caliber consistency of the output; generate comparable scores and comparable conclusions according to the calibration rules corresponding to the scale anchoring and comparability calibration requirements, and retain the calibration basis; aggregate and generate hierarchical evaluation results according to the evaluation protocol; solidify the problems identified in the process of structural consistency verification, caliber consistency verification, comparability calibration, and aggregation according to the evaluation protocol into a problem ledger in the form of structured entries. The problem entries should at least include the label caliber, scope of impact, location clues, rectification suggestions, and regression verification conditions, and be associated with the evaluation snapshot.

[0052] Outputs: stratified evaluation results (including comparability explanations and calibration basis), problem ledger, and result verification records.

[0053] S109, Evaluation snapshot fixation, change impact analysis and incremental review, including: Inputs: tiered evaluation results, issue ledger, execution plan, caliber signature, list of capability components and their version information, and evaluation boundary description.

[0054] Processing: A snapshot of the evaluation is fixed, containing at least the dataset version boundary, evaluation protocol version, caliber signature, execution plan, capability component version set, execution constraints, and execution records. When a change in the data version, evaluation protocol version, or capability component version is detected, a change impact analysis is performed based on consistency checkpoints, generating an impact set. An incremental re-evaluation is triggered on the impact set, and consistency checks are performed on the reuse boundary. The reuse boundary is defined after the change impact analysis as the scope of resources, results, and corresponding constraints from which historical evaluation results can be directly reused. Difference attribution information is output; this information is structured information characterizing the source and scope of difference between the re-evaluation results and historical results, distinguishing whether the difference originates from data changes, caliber changes, capability version changes, or compliance tailoring changes.

[0055] Output: Evaluation snapshot, reproduction verification instructions, incremental re-evaluation results, and difference attribution report.

[0056] Example 3, taking protocol-based evaluation for multi-source heterogeneous datasets as an example, shows a dataset containing structured tables, unstructured text, image samples, and corresponding annotation files, along with documentation and a field dictionary. The platform implementation process is as follows: (1) Access the data source and record the access method, permission constraints and version boundaries to form an auditable access record; construct a resource list based on the resource objects identified in the auditable access record; and extract the format, scale, modality and partition information of the resource objects to form a basic feature summary; (2) Select an applicable evaluation protocol instance, configure the indicator scope, threshold scope, label scope and aggregation rules, and write compliance constraints and budget-coverage-confidence targets; (3) Select external capability components and complete the consistency verification of interchangeable interfaces to form a list of capability components and a compatibility record; (4) Generate caliber signatures and solidify consistency verification points; (5) Generate an execution plan, and optionally implement a trial run calibration to determine the sampling ratio and budget threshold; (6) Execute the plan in a controlled manner and generate execution records; (7) Aggregate the output and perform comparability calibration according to the scale to generate stratified evaluation results; (8) Generate problem ledger entries and provide regression verification conditions; (9) Solidify evaluation snapshots to support re-evaluation, auditing and comparison.

[0057] Example 4, taking the impact analysis and incremental review under version change conditions as an example, when a new version of the dataset is released or the evaluation protocol is upgraded, the platform implements the following: (1) Identify the type of change (data version change, protocol definition change, capability version change, etc.); (2) Perform change impact analysis based on consistency checkpoints and generate an impact set; (3) Trigger incremental re-evaluation for the affected set, reuse historical results for non-affected areas and perform consistency verification on the reuse boundary; the reuse boundary is the resource range, result range and corresponding constraint range that can be directly reused after the change impact analysis. (4) Output a variance attribution report to distinguish the sources of variance and support audit and decision-making.

[0058] In summary, the multi-source parsing operator-based evaluation method and system for datasets provided by this invention achieves solidified and verifiable evaluation criteria through evaluation protocols and caliber signatures, reducing the risk of incomparability and disputes caused by caliber drift; it supports consistency verification of re-evaluation and audit arbitration through evaluation snapshots and execution records, clarifies the sources of differences, and reduces evaluation costs and improves iteration efficiency during updates and iterations through change impact analysis and incremental evaluation; it ensures comparability and interpretability of scores and conclusions across projects, time periods, and environments through scale anchoring and comparability calibration; it meets the requirements of data not leaving the domain and minimum visibility in sensitive data scenarios through compliance execution boundary control, and the process and output are auditable through a problem ledger and regression verification mechanism; and it makes the evaluation results executable, traceable, and verifiable through a problem ledger and regression verification mechanism.

[0059] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0060] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart... Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0061] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0062] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0063] The embodiments of the present invention have been described above with reference to the accompanying drawings. However, the present invention is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of the present invention without departing from the spirit and scope of the claims. All of these forms are within the protection scope of the present invention.

Claims

1. A data set oriented multi-source resolving operatorized evaluation method, characterized in that, include: Based on the acquired target dataset and authorization information, the target dataset is registered for access, and the data source, access method, permission constraints and version boundaries are recorded to form an auditable access record. Based on the auditable access record, a resource list, basic feature summary and evaluation boundary description are constructed. The evaluation protocol instance is determined based on the obtained evaluation objectives and applicable standards; Candidate capability components are selected based on the evaluation protocol instance, resource list, and basic feature summary, and the capability component list is determined after consistency verification of the interchange interface. A caliber signature is generated based on the evaluation protocol instance, the list of capability components, and the execution constraints and consistency verification constraints corresponding to the evaluation protocol instance; the caliber signature is used for execution plan binding, result consistency verification, and change impact analysis; An execution plan is generated based on the evaluation protocol instance, caliber signature, evaluation boundary description, and capability component interface information determined according to the capability component list. The evaluation protocol instance is executed according to the execution plan, the list of capability components, and the compliance constraints determined by the evaluation protocol instance, generating auditable execution records and a set of capability component outputs; Based on the output set of the capability components, the caliber signature, the execution record, and the comparability calibration requirements determined by the evaluation protocol instance, the results are aggregated, the comparability calibration is performed, and the problem ledger is generated to obtain the evaluation results, the problem ledger, and the result verification record. An evaluation report is generated based on the evaluation results, issue ledger, execution plan, signature, capability component list, evaluation boundary description, and result verification record.

2. The multi-source analytical operator-based evaluation method for datasets according to claim 1, characterized in that, The process of registering access to the target dataset, recording the data source, access method, permission constraints, and version boundaries to form an auditable access record, and constructing a resource list, basic feature summary, and evaluation boundary description based on the auditable access record, includes: Record the data source, access method, permission constraints, and version boundaries of the target dataset to form an auditable access record; the data source includes files, directories, object storage, databases, or dataset version references; A resource list is constructed based on the resource objects identified in the auditable access records. The resource list is a set of resource objects in the target dataset that can be evaluated and processed. Based on the format, scale, modality, and partition information of the resource objects corresponding to the resource list, extract basic feature summaries; The evaluation boundary is determined based on the resource list, permission constraints, and version boundaries.

3. The multi-source analytical operator-based evaluation method for datasets according to claim 1, characterized in that, The method for determining an evaluation protocol instance based on the acquired evaluation objectives and applicable standards includes: The evaluation objectives and applicable standards include: indicator definitions, threshold definitions, label definitions, aggregation rules, compliance constraints, budget constraints, coverage objectives, confidence objectives, and comparability calibration requirements; wherein, the comparability calibration requirements are calibration rule requirements pre-configured in the evaluation protocol instance according to the evaluation objectives and applicable standards. Based on the indicator definition, threshold definition, label definition, aggregation rules, compliance constraints, budget constraints, coverage targets, confidence targets and comparability calibration requirements, select or generate evaluation protocol instances from the protocol library; In scenarios where multiple standards coexist, mapping and conflict resolution are performed between different standards, and the mapping basis and difference descriptions are recorded to obtain a protocol difference description; the multiple standards include industry specifications, enterprise specifications, and project agreements.

4. The multi-source analytical operator-based evaluation method for datasets according to claim 1, characterized in that, The process of filtering candidate capability components based on the evaluation protocol instance, resource list, and basic feature summary, and determining the capability component list after consistency verification of the interchange interface, includes: The input requirements of the evaluation protocol are determined based on the evaluation protocol instance, and candidate external capability components that can process the corresponding resource objects are screened by combining the resource object types represented by the resource list and the format, scale, modality and partition information represented by the basic feature summary. Perform interface consistency checks on external capability components, including: input view constraints, output structure constraints, error caliber constraints, version declarations and compatibility rules; when the checks are not met, trigger replacement, degradation or isolation strategies. Based on the candidate external capability components that have passed the consistency verification of the interchange interface, determine the list of capability components and record the interface verification results and the triggering of replacement / degradation / isolation strategies.

5. The multi-source analytical operator-based evaluation method for datasets according to claim 1, characterized in that, The step of generating a signature based on the evaluation protocol instance, the capability component list, and the execution constraints and consistency verification constraints corresponding to the evaluation protocol instance includes: Generate caliber signatures for key caliber elements of the evaluation protocol instance. These caliber signatures are used for execution plan binding, result consistency verification, and change impact analysis. The execution constraints corresponding to the evaluation protocol instance include resource constraints, timeout constraints, and budget constraints; the consistency verification constraints include input view constraints, output structure constraints, error caliber constraints, and version declaration and compatibility rules. Based on the execution constraints, input view constraints, output structure constraints, error caliber constraints, and version declarations and compatibility rules, determine whether the results are comparable and whether they need to be re-evaluated or marked as incomparable when the data version, protocol version, capability version, or execution constraints change, and obtain a list of consistency verification points.

6. The multi-source analytical operator-based evaluation method for datasets according to claim 1, characterized in that, The step of generating an execution plan based on the evaluation protocol instance, caliber signature, evaluation boundary description, and capability component interface information determined according to the capability component list includes: Based on the requirements of the evaluation protocol instance, and combined with the evaluation scope and version boundaries defined by the evaluation boundary description, as well as the calling methods, input-output mappings and compatibility relationships determined by the capability component interface information, an auditable execution plan is generated. The execution plan includes: determining the execution path, parallel granularity, triggering conditions, budget constraints, sampling strategies, compliance trimming rules, failure handling and degradation strategies, and incremental re-evaluation boundaries. Bind auditable execution plans to caliber signatures.

7. The multi-source analytical operator-based evaluation method for datasets according to claim 1, characterized in that, The step of executing the evaluation protocol instance according to the execution plan, the list of capability components, and the compliance constraints determined by the evaluation protocol instance, and generating auditable execution records and a set of capability component outputs, includes: The evaluation of the evaluation protocol instance is carried out according to the execution plan, including: determining the set of capability component outputs for execution based on the capability component list and the execution plan; dynamically governing budget constraints, coverage targets and confidence targets in combination with the set of capability component outputs; and implementing compliance tailoring and on-site execution strategies based on the compliance constraints. Record the strategy triggers, degradation paths, error reasons, retry information and execution trajectory during the execution process to form an auditable execution log.

8. The multi-source analytical operator-based evaluation method for datasets according to claim 1, characterized in that, The process of aggregating results, performing comparability calibration, and generating an issue ledger based on the output set of the capability components, caliber signatures, execution records, and comparability calibration requirements determined by the evaluation protocol instance, yields evaluation results, an issue ledger, and result verification records, including: Based on the signature, the structural and caliber consistency of the output set of the capability components is verified to obtain the result verification record. Based on the results of structural consistency and caliber consistency verification, a comparable score and a comparable conclusion are generated according to the calibration rules corresponding to the scale anchoring and the aforementioned comparability calibration requirements. Based on comparable scores and comparable conclusions, hierarchical evaluation results are generated by aggregating them according to the evaluation protocol. Issues identified during structural consistency verification, calibrator consistency verification, comparability calibration, and aggregation according to the evaluation protocol are solidified into an issue ledger in the form of structured entries. The issue entries include label calibrator, scope of impact, location clues, rectification suggestions, and regression verification conditions, and are associated with the evaluation snapshot. The evaluation snapshot includes dataset version boundary, evaluation protocol version, calibrator signature, execution plan, capability component version set, execution constraints, and execution records.

9. The multi-source analytical operator-based evaluation method for datasets according to claim 8, characterized in that, The evaluation report, generated based on evaluation results, issue ledger, execution plan, signature criteria, capability component list, evaluation boundary description, and result verification records, includes: Solidification evaluation snapshot; When a change in dataset version, evaluation protocol version, or capability component version is detected, change impact analysis is performed based on consistency checkpoints and an impact set is generated. Incremental review is triggered on the impact set, consistency verification is performed on the reuse boundary, and difference attribution information is output. The reuse boundary is the resource range, result range, and corresponding constraint range that can be directly reused after the change impact analysis. The difference attribution information is structured information that characterizes the source and impact range of the difference between the review result and the historical result, and is used to distinguish whether the difference comes from data change, caliber change, capability version change, or compliance trimming change. An evaluation report is generated based on the evaluation snapshot and differential attribution information.

10. The multi-source analytical operator-based evaluation method for datasets according to claim 1, characterized in that, Also includes: A small-scale trial run of the execution plan was conducted to estimate anomaly density, coverage achievement, and resource consumption. Under the premise of meeting the budget constraints, coverage targets, confidence targets and compliance constraints preset in the evaluation protocol instance, the sampling ratio, budget upper limit and trigger threshold parameters are calibrated, and the basis for cost and confidence interpretation is formed; if adjustments are made, the reasons for the adjustments and the differences before and after the adjustments are recorded.