Model evaluation method, platform, equipment, storage medium and computer program product

By customizing task configuration and metadata processing mechanisms, the limitations of existing model evaluation schemes are addressed, providing a full-process, automated model evaluation method and platform that enables comprehensive evaluation of models and is suitable for diverse service needs and complex scenarios.

CN121636987APending Publication Date: 2026-03-10ALIBABA CLOUD COMPUTING CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-09-09
Publication Date
2026-03-10

AI Technical Summary

Technical Problem

Existing model evaluation schemes can only evaluate fixed and limited performance indicators, making it difficult to adapt to the ever-changing service needs. They also have low automation and are not comprehensive enough.

Method used

This paper provides a model evaluation method and platform that allows users to customize task configuration information. Through the metadata management module, operator management module, and task management module, it automatically selects suitable models, datasets, and operators, covering model inference, single evaluation, and comprehensive evaluation. It supports full-process evaluation and achieves a high degree of automation and flexibility.

Benefits of technology

It provides comprehensive and scalable model evaluation services, which can flexibly respond to diverse service needs, improve evaluation efficiency and accuracy, and are suitable for various complex application scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121636987A_ABST
    Figure CN121636987A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a model evaluation method, platform and equipment, a storage medium and a computer program product, in the embodiment of the invention, a user is allowed to customize task configuration information, then the user is supported to select different metadata, data sets and operators in a model evaluation process, continuously changing service requirements are met, and the user experience is improved. And the flexibility of model evaluation is improved. The adaptive data processing mechanism based on the metadata can automatically match appropriate models, data sets and operators based on the metadata, manual intervention is reduced, and evaluation efficiency is improved. In the model evaluation process, three main links of model reasoning, single evaluation and comprehensive evaluation can be covered, limitation to a single performance index is avoided, various performance indexes of the model are comprehensively investigated, a more comprehensive evaluation result is provided, and the overall performance of the model can be better reflected. Therefore, a comprehensive, extensible and highly automatic model evaluation service is provided, and diversified service requirements can be flexibly met.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of model evaluation technology, and in particular to a model evaluation method, platform, device, storage medium and computer program product. Background Technology

[0002] A converged media platform is a comprehensive content publishing and interaction platform that integrates multiple media formats (such as text, images, audio, and video). By applying the "large language model," converged media platforms can not only improve content quality and production efficiency, but also enhance user experience and interactivity, thereby improving the overall competitiveness of the converged media platform.

[0003] To improve the effectiveness of large language models in deployment, evaluating them is a crucial step. Existing model evaluation schemes often only assess a fixed number of performance metrics (such as accuracy and recall), which is not comprehensive enough, makes it difficult to flexibly respond to constantly changing service needs, and has a low degree of automation. Summary of the Invention

[0004] This application provides a model evaluation method, platform, device, storage medium, and computer program product for providing comprehensive, scalable, and highly automated model evaluation services.

[0005] This application provides a model evaluation method, comprising: receiving a full-process evaluation request from a front end, the full-process evaluation request being used to instruct the execution of model inference, single-item evaluation, and comprehensive evaluation in the model evaluation process, the full-process evaluation request including first task configuration information; determining the metadata of a first model, the metadata of a first domain dataset, and the metadata of multiple first performance indicators based on multiple existing metadata and the first task configuration information; obtaining the first domain dataset and its corresponding first dataset processing operator from multiple existing datasets and multiple existing operators based on the metadata of the first domain dataset, and processing the first domain dataset using the first dataset processing operator; triggering the first model to perform inference based on the processed first domain dataset based on the metadata of the first model, obtaining the first model inference result; obtaining the first single-item evaluation operator corresponding to the first performance indicator from multiple existing operators based on the metadata of the first performance indicator, and performing a single-item evaluation on the first model inference result using the first single-item evaluation operator, obtaining a single-item evaluation result of the first performance indicator; obtaining a comprehensive evaluation operator from multiple existing operators, and performing a comprehensive evaluation on the single-item evaluation results of multiple first performance indicators using the comprehensive evaluation operator, obtaining a first comprehensive evaluation result and returning it to the front end.

[0006] This application embodiment also provides a model evaluation platform, including: a front-end and a back-end server, wherein the back-end server includes a metadata management module, an operator management module and a task management module;

[0007] The metadata management module is used to provide multiple existing metadata and multiple datasets;

[0008] The operator management module provides multiple existing operators; the task management module receives a full-process evaluation request triggered by the user on the front end. This full-process evaluation request instructs the execution of model inference, individual evaluation, and comprehensive evaluation within the model evaluation process. The full-process evaluation request includes first task configuration information. Based on multiple existing metadata and the first task configuration information, the module determines the metadata of the first model, the metadata of the first domain dataset, and the metadata of multiple first performance metrics. Based on the metadata of the first domain dataset, the module retrieves the first domain dataset and its corresponding first dataset processing operator from multiple existing datasets and multiple existing operators, and utilizes... The first dataset processing operator processes the first domain dataset; based on the metadata of the first model, the first model is triggered to perform inference based on the processed first domain dataset to obtain the inference result of the first model; based on the metadata of the first performance index, the first single evaluation operator corresponding to the first performance index is obtained from multiple existing operators, and the first single evaluation operator is used to perform single evaluation on the inference result of the first model to obtain the single evaluation result of the first performance index; the comprehensive evaluation operator is obtained from multiple existing operators, and the comprehensive evaluation operator is used to perform comprehensive evaluation on the single evaluation results of multiple first performance indices to obtain the first comprehensive evaluation result and return it to the front end.

[0009] This application also provides an electronic device, including: a memory and a processor; the memory for storing a computer program; and the processor coupled to the memory for executing the computer program to perform steps in a model evaluation method.

[0010] This application also provides a computer-readable storage medium storing a computer program, which, when executed by a processor, enables the processor to implement the steps in the model evaluation method.

[0011] This application also provides a computer program product, including a computer program / instructions, which, when executed by a processor, enable the processor to implement the steps in the model evaluation method.

[0012] The technical solution provided in this application allows users to customize task configuration information, thereby supporting the selection of different metadata, datasets, and operators in the model evaluation process. This meets constantly changing service needs and improves the flexibility of model evaluation. The metadata-based adaptive data processing mechanism can automatically match suitable models, datasets, and operators based on metadata, reducing manual intervention and improving evaluation efficiency. The model evaluation process covers three main stages: model inference, individual evaluation, and comprehensive evaluation. It is no longer limited to a single performance indicator but comprehensively examines various performance indicators of the model, providing more comprehensive evaluation results that better reflect the overall performance of the model. Therefore, it provides a comprehensive, scalable, and highly automated model evaluation service that can flexibly respond to diverse service needs and is suitable for various complex application scenarios. Attached Figure Description

[0013] The accompanying drawings, which are included to provide a further understanding of this application and form part of this application, illustrate exemplary embodiments and are used to explain this application, but do not constitute an undue limitation of this application. In the drawings:

[0014] Figure 1 A schematic diagram illustrating an application scenario provided in an embodiment of this application;

[0015] Figure 2 A flowchart of a model evaluation method provided in an embodiment of this application;

[0016] Figure 3 A diagram illustrating the generation process of an exemplary robust dataset;

[0017] Figure 4 A schematic diagram illustrating another application scenario provided by an embodiment of this application;

[0018] Figure 5 A system architecture diagram of a model evaluation platform provided in this application embodiment;

[0019] Figure 6 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation

[0020] To make the objectives, technical solutions, and advantages of this application clearer, the technical solutions of this application will be clearly and completely described below in conjunction with specific embodiments and corresponding drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of them. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0021] In the embodiments of this application, "at least one" refers to one or more, and "more than one" refers to two or more. "And / or" describes the access relationship between associated objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, and B existing alone, where A and B can be singular or plural. In the textual description of this application, the character " / " generally indicates that the preceding and following associated objects have an "or" relationship. Furthermore, in the embodiments of this application, "first," "second," "third," etc., are only used to distinguish the content of different objects and have no other special meaning.

[0022] The following is a description of the terms used in the embodiments of this application:

[0023] Large Language Models (LLMs), also known as Large Language Models, refer to a class of natural language processing models with an extremely large number of parameters. LLMs are typically based on deep learning architectures, especially the Transformer architecture, and learn the complex structure and rich contextual information of language through pre-training on massive amounts of text data.

[0024] Metadata refers to data about data; that is, information used to describe, interpret, locate, or help manage information. The meaning and purpose of metadata may vary in different fields and application scenarios.

[0025] An operator is a functional unit or function that performs a specific operation.

[0026] Inference is the process of using a trained model to process new data and produce output.

[0027] A converged media platform is a comprehensive content publishing and interaction platform that integrates multiple media formats (such as text, images, audio, and video). By applying the "large language model," converged media platforms can not only improve content quality and production efficiency, but also enhance user experience and interactivity, thereby improving the overall competitiveness of the converged media platform.

[0028] To improve the effectiveness of large language models in deployment, evaluating them is a crucial step. Existing model evaluation schemes often only assess a fixed number of performance metrics (such as accuracy and recall), which is not comprehensive enough, makes it difficult to flexibly respond to constantly changing service needs, and has a low degree of automation.

[0029] To address this, embodiments of this application provide a model evaluation method, platform, device, storage medium, and computer program product. These embodiments allow users to customize task configuration information, enabling them to select different metadata, datasets, and operators during the model evaluation process. This meets evolving service needs and improves the flexibility of model evaluation. The metadata-based adaptive data processing mechanism automatically matches suitable models, datasets, and operators based on metadata, reducing manual intervention and improving evaluation efficiency. The model evaluation process covers three main stages: model inference, individual evaluation, and comprehensive evaluation. It is no longer limited to a single performance indicator but comprehensively examines various performance indicators of the model, providing more comprehensive evaluation results that better reflect the overall performance of the model. Therefore, this provides a comprehensive, scalable, and highly automated model evaluation service that can flexibly respond to diverse service needs and is suitable for various complex application scenarios.

[0030] The technical solutions of this application and how they solve the aforementioned technical problems are described in detail below with specific embodiments. These specific embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments. The technical solutions provided by each embodiment of this application are described in detail below with reference to the accompanying drawings.

[0031] Figure 1 This diagram illustrates an application scenario provided by an embodiment of this application. In this scenario, the model evaluation platform can provide a comprehensive, scalable, and highly automated model evaluation service. The model evaluation platform includes a front-end and a back-end server, and users interact with the back-end server through the front-end.

[0032] In practical applications, model evaluation platforms can evaluate various models, including but not limited to: large language models, sequence labeling models, text classification models, and machine translation models.

[0033] In practical applications, relevant personnel can perform metadata management operations on the front-end interface, while the back-end server performs metadata management operations such as adding, deleting, or updating. For example, users can register various metadata on the back-end server, and can also delete or update existing metadata. Multiple existing metadata provided by the back-end server include, but are not limited to: multiple model metadata, multiple dataset metadata, multiple performance metric metadata, and multiple service metadata.

[0034] For example, model metadata describes the model and includes, but is not limited to: model name, model version, model API endpoint, and default configuration. The model name is a unique identifier used to distinguish different models. The model API endpoint is the model's network location used to send requests to the model. The model version is the specific version number of the model, used to track and manage differences between versions. The default configuration may include default configuration options for the model at runtime.

[0035] For example, dataset metadata describes the metadata of a dataset. Dataset metadata includes, but is not limited to: dataset name, dataset version, dataset local path, the mapping between datasets and operators, and the mapping between datasets and performance metrics. Here, the dataset name is a unique identifier used to distinguish different datasets. The dataset version refers to the specific version number of the dataset, used to track and manage differences between different versions. The dataset local path refers to the storage location of the dataset, used to retrieve the dataset.

[0036] For example, performance metric metadata refers to the metadata describing a performance metric. Performance metric metadata includes, but is not limited to: metric name, metric version, and the mapping relationship between performance metrics and operators. The metric name is a unique identifier for the performance metric, used to distinguish different performance metrics. The metric version refers to the specific version number of the performance metric, used to track and manage the differences between different versions.

[0037] For example, service metadata refers to metadata that describes a service. Service metadata includes, but is not limited to, service names, mappings between services and performance metrics, and mappings between services and datasets. Taking the evaluation of a large language model as an example, services include, but are not limited to, dialogue services, text title generation services, or text summarization services, etc.

[0038] In practical applications, relevant personnel can perform operator management operations on the front-end interface, while the back-end server performs operations such as adding, deleting, or updating operators. For example, users can register various operators on the back-end server, and can also delete or update existing operators. The back-end server provides multiple existing operators, including but not limited to: multiple dataset processing operators, multiple single-item evaluation operators, multiple comprehensive evaluation operators, and multiple robust operators.

[0039] Among them, dataset processing operators are operators that process datasets. Dataset processing operators include, but are not limited to, service processing operators and data file processing operators.

[0040] Service processing operators are primarily used to perform specific service logic processing on data. Common service processing operators include, but are not limited to: title generation operators, summary generation operators, and intelligent question-answering operators. Title generation operators automatically generate an appropriate title from given text content. Summary generation operators generate a concise and representative summary from longer text content. Intelligent question-answering operators search for or generate answers from an existing knowledge base based on user-submitted questions.

[0041] Data file processing operators are primarily used to process data files of various formats. Common data file processing operators include, but are not limited to, single-file processing operators and recursive file processing operators. Single-file processing operators are used to read and parse files of a single format based on a file path, such as JSONL or CSV (Comma-Separated Values) files. JSONL is a file format where each line contains a single JSON (JavaScript Object Notation) object; CSV is a simple file format for storing tabular data, typically used for data exchange and backup. Recursive file processing operators are used to process folders or compressed files containing multiple files, recursively reading all related files. Generally, a recursive file processing operator can output the contents of all files based on the input folder or compressed file path.

[0042] Among them, a single evaluation operator is an operator that evaluates the data to be evaluated for a single performance indicator. The single evaluation operators corresponding to different performance indicators can be the same or different. Single evaluation operators include, but are not limited to: Bleu (Bilingual Evaluation Understudy) operator, single-word Rouge (Recall-Oriented Understudy for Giving Evaluation) operator, jieba segmentation Rouge operator, precision operator, and keyword coverage operator.

[0043] The BLEU operator is an operator used to evaluate the quality of machine translation. It calculates the similarity between the reference translation and the machine translation result based on n-gram overlap. An n-gram, often referred to as an "n-tuple," is a sequence of n consecutive words or characters extracted from text or speech, where n is a positive integer.

[0044] The Rouge operator is a commonly used operator for evaluating the quality of automatic text summarization, primarily used to measure the similarity between automatically generated summaries and human-written reference summaries. The Rouge operator includes different variants, such as ROUGE-N and ROUGE-L, to measure overlap and similarity in different aspects. ROUGE-N measures the degree of overlap among n consecutive words; ROUGE-L uses the Longest Common Subsequence (LCS) to calculate similarity.

[0045] The single-word Rouge operator, also known as the ROUGE-1 operator, primarily measures the overlap of individual words or characters by comparing words appearing in the automatically generated summary with those appearing in the reference summary to calculate similarity.

[0046] The Jieba segmentation ROUGE operator combines Jieba segmentation with the ROUGE operator to evaluate the quality of Chinese text summarization. Specifically, it first uses Jieba to segment the Chinese text, and then uses the ROUGE operator to evaluate the quality of the Chinese text summary based on the segmentation results. Jieba is a commonly used Chinese word segmentation tool that can divide a continuous sequence of Chinese characters into individual words.

[0047] The accuracy operator is used to measure the degree of match between a model's predictions and the actual results. In natural language processing and other machine learning tasks, the accuracy operator is commonly used to evaluate the performance of a model on tasks such as classification and label prediction.

[0048] The keyword coverage operator is used to evaluate how many keywords are contained in the text generated by a model. In natural language processing tasks, especially for scenarios such as summarization, dialogue systems, or text generation, keyword coverage is an important performance metric that helps assess whether the text generated by the model covers the important information in the original text.

[0049] Among them, the comprehensive evaluation operator refers to the operator used in the comprehensive evaluation, which is mainly used to perform weighted summation on multiple data to obtain the comprehensive evaluation result.

[0050] Robust operators are used to generate robust datasets based on domain datasets. These robust datasets simulate diverse input data that may occur in the real world to test the model's performance under different conditions, thereby verifying the model's robustness. Robust operators include, but are not limited to, basic text manipulation operators, object manipulation operators, and combinatorial operators.

[0051] The basic text manipulation operators are used to perform specific types of modifications on the input text to test the robustness of the model. These basic text manipulation operators include, for example, one or more of the following:

[0052] 1. Random deletion operator: Used to randomly delete certain words, phrases, or sentences from a text.

[0053] 2. Random insertion operator: Used to insert new characters, words, or sentences at random positions in the text.

[0054] 3. Random Replacement Operator: Used to randomly replace certain words, phrases, or sentences in a text.

[0055] 4. Minor word order shuffling operator: Used to slightly change the order of words or phrases in a text.

[0056] 5. Synonym Replacement Operator: Used to replace certain words in the original text with synonyms.

[0057] 6. Repeat Segment Multiple Operator: Used to repeat a paragraph or sentence in the text.

[0058] 7. Input truncation operator: Used to truncate a portion of the text, keeping the rest.

[0059] 8. Construct meaningless text operator: used to generate irrelevant text to test the model's resistance to irrelevant information.

[0060] The operation object operator is used to extract operation objects from the text content. Operation objects include, but are not limited to, characters, words, sentences, and paragraphs.

[0061] Combinatorial operators are used to process operands using basic operators. In implementation, they can be obtained by multi-dimensional combination of basic operators and operands. Optionally, a rule cube mechanism can be used to obtain combinatorial operators by multi-dimensional combination of basic operators and operands. The rule cube mechanism can provide a configurable set of rules, allowing users to define various abnormal input conditions, which can improve test coverage when applied to dataset generation.

[0062] In practical applications, relevant personnel can perform dataset management operations on the front end, while the backend server performs operations such as adding, deleting, or updating datasets. For example, users can register various datasets on the backend server (i.e., add various datasets on the backend server), and can also delete or update existing datasets. The backend server provides several existing datasets, including but not limited to: general language understanding datasets, dialogue and interaction datasets, medical domain datasets, sentiment analysis datasets, machine translation datasets, single-sentence summarization datasets for news articles, story creation datasets, etc. The datasets mentioned above can be understood as domain-specific datasets. By using robust operators to process these domain-specific datasets, corresponding robust datasets can be obtained and stored on the backend server for later use.

[0063] In practical applications, model evaluation platforms can provide various services, including but not limited to: full-process evaluation services, inference services, single-item evaluation services, and comprehensive evaluation services. Full-process evaluation services support the sequential execution of model inference, single-item evaluation, and comprehensive evaluation within the model evaluation process. Inference services support model inference, and the results can be saved for subsequent single-item evaluations. Single-item evaluation services support using existing model inference results for single-item evaluations, and these results can be saved for subsequent comprehensive evaluations. Comprehensive evaluation services support using existing single-item evaluation results for comprehensive evaluations.

[0064] See Figure 1 As shown in ① and ②, in practical applications, users first submit task requests on the front-end interface, and then the back-end server executes the task and returns the task execution result. Task requests include, but are not limited to: full-process evaluation requests, reasoning requests, single-item evaluation requests, or comprehensive evaluation requests.

[0065] Specifically, users configure task configuration information on the front-end interface. This configuration information varies depending on the task and includes, but is not limited to: task name, basic model attributes (e.g., model name, model version, model API address), basic dataset attributes (e.g., dataset name, dataset version), basic performance metric attributes, number of evaluations, and task mode. After configuring the task configuration information, the user submits a task request through the front-end interface. Upon receiving the task request, the back-end server begins executing the task. Based on the task configuration information, multiple existing metadata sets, multiple existing datasets, and multiple existing operators, the back-end server executes the task and returns the execution result to the front-end. This satisfies diverse user needs, including end-to-end evaluation, inference, single-item evaluation, and comprehensive evaluation, achieving a more comprehensive and highly automated model evaluation.

[0066] It should be noted that, Figure 1 The application scenario shown is merely an exemplary one, and the embodiments of this application do not limit the application scenario.

[0067] Figure 2 This is a flowchart illustrating a model evaluation method provided in an embodiment of this application. The method can be executed by a model evaluation platform, which may consist of software and / or hardware, and is generally configured in an electronic device. See also... Figure 2 The model evaluation method may include the following steps:

[0068] 201. Receive a full-process evaluation request from the front end. The full-process evaluation request is used to instruct the execution of model inference, individual evaluation and comprehensive evaluation in the model evaluation process. The full-process evaluation request includes the first task configuration information.

[0069] In practical applications, various metadata, datasets, and operators can be registered in the model evaluation platform. This allows the platform to provide multiple existing metadata sets, datasets, and operators, enabling users to evaluate various models across different datasets using various performance metrics. This makes model evaluation more comprehensive, flexibly adaptable to evolving service needs, and highly scalable and automated.

[0070] Optionally, the model evaluation platform can receive management operation requests triggered by users on the front-end interface (i.e., the interface displayed on the front end), and perform corresponding management operations on multiple existing operators, datasets, or metadata. These management operations include any of the following: add, delete, and update operations. This allows for the flexible expansion of various operators, datasets, or metadata within the model evaluation platform, achieving automated management of operators, datasets, or metadata. It can flexibly respond to diverse service needs and effectively improve the scalability of the model evaluation service.

[0071] Further optionally, the management operation request includes a robust operator addition request, which includes: basic text operation operators, operation object operators, and composite operators. Correspondingly, the implementation of performing corresponding management operations on multiple existing operators is as follows: adding basic text operation operators, operation object operators, and composite operators to multiple existing operators. Among them, basic text operation operators are used to modify text content, and include at least one of the following: random deletion operator, random insertion operator, random replacement operator, slightly shuffling word order operator, synonym replacement operator, repeated segment multiple times operator, input truncation operator, and meaningless text construction operator; operation object operators are used to extract operation objects from text content, where the operation objects are characters, words, sentences, or paragraphs; composite operators are used to process the operation objects using basic operation operators.

[0072] It is worth noting that the model evaluation platform provides robust operators, which can be used to process domain datasets to generate robust datasets. This not only greatly improves the coverage of automated tests, but also effectively enhances the robustness of the model in the face of adversarial attacks by simulating real-world data noise or anomalous inputs, ensuring the accuracy and practicality of robustness evaluation.

[0073] In this embodiment, the model evaluation platform can meet the requirements of the entire evaluation process, that is, it sequentially executes model inference, individual evaluation, and comprehensive evaluation in the model evaluation process. Specifically, the user inputs the task configuration information required for the entire evaluation process on the front-end interface and triggers the entire evaluation request through the front-end interface. For ease of understanding and distinction, the task configuration information carried in the entire evaluation request is referred to as the first task configuration information.

[0074] In practical applications, the first task configuration information can be configured as needed. The first task configuration information includes, but is not limited to: task name, basic attribute information of the first model, basic attribute information of the first service, basic attribute information of multiple first performance indicators, basic attribute information of the first domain dataset, number of target evaluations, and task mode.

[0075] The first model is the model that requires full-process evaluation, and the first service is any service, such as a dialogue service, a text title generation service, or a text summary generation service. There is no limit to the number of first models; there can be one or more.

[0076] The first performance metric can be any performance metric, such as keyword coverage, accuracy, text summarization quality, machine translation quality, etc.

[0077] The first domain dataset can be any dataset, such as a general language understanding dataset, a dialogue and interaction dataset, a medical dataset, a sentiment analysis dataset, a machine translation dataset, a single-sentence summary dataset for news articles, a story creation dataset, etc. There is no limit to the number of first domain datasets; there can be one or more.

[0078] In practical applications, users can configure the target number of evaluations as needed. The target number of evaluations refers to the number of times a single evaluation will be performed. Optionally, if the user configures the target number of evaluations in the configuration information of the first task, the default target number of evaluations can be set to 1.

[0079] In practical applications, selectable task modes include, but are not limited to: full-process evaluation mode, inference mode, single-item evaluation model, and comprehensive evaluation mode. If the user's requirement is full-process evaluation, the user will configure the task mode in the first task configuration information as full-process evaluation mode. Optionally, if the user does not configure a task mode in the first task configuration information, the full-process evaluation mode can be used as the default task mode.

[0080] 202. Based on multiple existing metadata and the first task configuration information, determine the metadata of the first model, the metadata of the first domain dataset, and the metadata of multiple first performance metrics.

[0081] In practical applications, multiple existing metadata sets include, but are not limited to, multiple model metadata sets, multiple dataset metadata sets, multiple performance metric metadata sets, and multiple service metadata sets. By searching the first task configuration information among these existing metadata sets, the metadata sets of the first model, the first domain dataset, and multiple first performance metrics can be found.

[0082] Optionally, to support users in flexibly configuring task configuration information and meeting flexible and ever-changing service needs, step 202 is implemented as follows: based on the basic attribute information of the first model and the basic attribute information of the first service in the first task configuration information, the metadata of the first model and the metadata of the first service are searched in multiple existing metadata; based on the existence of the metadata of the first service and the basic attribute information of the first domain dataset in the first task configuration information, the metadata of the first domain dataset is searched in multiple existing metadata; based on the existence of the metadata of the first service and the basic attribute information of multiple first performance indicators in the first task configuration information, the metadata of multiple first performance indicators is searched in multiple existing metadata.

[0083] Specifically, when configuring the first task configuration information for the full-cycle evaluation task, users can configure the basic attribute information of the first model, the basic attribute information of the first service, the basic attribute information of multiple first performance indicators, and the basic attribute information of the first domain dataset in the first task configuration information, so as to search for the metadata of the first model, the metadata of the first domain dataset, and the metadata of multiple first performance indicators from multiple existing metadata based on this information.

[0084] In practical applications, based on the basic attribute information of the first model and the basic attribute information of the first service in the configuration information of the first task, and by utilizing the mapping relationship between the service and the dataset and the mapping relationship between the service and the performance indicators in the metadata of the first service, it is also possible to find the metadata of the first model, the metadata of the first domain dataset, and the metadata of multiple first performance indicators.

[0085] For example, the implementation of searching for the metadata of the first domain dataset among multiple existing metadata based on the existence of the metadata of the first service and the basic attribute information of the first domain dataset in the first task configuration information is as follows: if the first task configuration information does not include the basic attribute information of the first domain dataset, then the metadata of the first domain dataset is searched among multiple existing metadata based on the mapping relationship between the service and the dataset in the metadata of the first service; if the first task configuration information includes the basic attribute information of the first domain dataset, then the metadata of the first domain dataset is searched among multiple existing metadata based on the basic attribute information of the first domain dataset.

[0086] Understandably, when searching for the metadata of the first domain dataset among multiple existing metadata sources, based on the mapping relationship between services and datasets in the metadata of the first service, the basic attribute information of the first domain dataset can be determined according to the mapping relationship between services and datasets in the metadata of the first service. The metadata of the first domain dataset can then be searched among multiple existing metadata sources based on this basic attribute information. Specifically, the mapping relationship between services and datasets can be a mapping relationship between service names and dataset names. The name of the first domain dataset is determined based on the mapping relationship between services and datasets in the metadata of the first service, and this name is used as the basic attribute information of the first domain dataset.

[0087] For example, the implementation of searching for the metadata of multiple first performance indicators in multiple existing metadata based on the existence of the basic attribute information of multiple first performance indicators in the metadata of the first service and the configuration information of the first task is as follows: if the configuration information of the first task does not include the basic attribute information of multiple first performance indicators, then the metadata of multiple first performance indicators is searched in multiple existing metadata based on the mapping relationship between services and performance indicators in the metadata of the first service; if the configuration information of the first task includes the basic attribute information of multiple first performance indicators, then the metadata of multiple first performance indicators is searched in multiple existing metadata based on the basic attribute information of multiple first performance indicators.

[0088] Understandably, when searching for the metadata of multiple first performance metrics among multiple existing metadata sources, based on the mapping relationship between services and performance metrics in the metadata of the first service, the basic attribute information of the first performance metric can be determined according to the mapping relationship between services and performance metrics in the metadata of the first service. The metadata of the first performance metric can then be searched among multiple existing metadata sources based on this basic attribute information. Specifically, the mapping relationship between services and performance metrics can be a mapping relationship between the service name and the metric name of the performance metric. The metric name of the first performance metric is determined based on the mapping relationship between services and performance metrics in the metadata of the first service, and this metric name serves as the basic attribute information of the first performance metric.

[0089] 203. Based on the metadata of the first domain dataset, obtain the first domain dataset and its corresponding first dataset processing operator from multiple existing datasets and multiple existing operators, and process the first domain dataset using the first dataset processing operator; trigger the first model to perform inference based on the processed first domain dataset according to the metadata of the first model, and obtain the inference result of the first model.

[0090] Specifically, during the inference phase of the model evaluation process, the first-domain dataset is obtained from multiple existing datasets based on its metadata. The basic attribute information (e.g., operator name) of the first-domain dataset's processing operator is determined based on the mapping relationship between datasets and operators in the first-domain dataset's metadata. The first-domain dataset processing operator is then obtained from multiple existing operators based on this basic attribute information, and the first-domain dataset is processed using this operator. The first-domain dataset processing operator can be, for example, a single-file processing operator or a recursive file processing operator, but is not limited to these.

[0091] In practical applications, the model evaluation platform can access the first model based on its API address in the model's metadata to request the first model to perform inference based on the processed first-domain dataset and obtain the inference result. During the inference process, the first model may use various first-domain dataset processing operators, such as title generation operators, summary generation operators, and intelligent question-answering operators, to process the first-domain dataset. Thus, the inference result of the first model may include, but is not limited to, text titles, text summaries, or dialogue interaction results.

[0092] In some optional embodiments, the first task configuration information includes: the first robust dataset label corresponding to the first domain dataset; the model evaluation platform can also obtain the first robust dataset corresponding to the first robust dataset label; accordingly, the implementation of triggering the first model to perform inference based on the processed first domain dataset according to the metadata of the first model and obtaining the inference result of the first model is as follows: triggering the first model to perform inference based on the first robust dataset and the processed first domain dataset according to the metadata of the first model and obtaining the inference result of the first model.

[0093] Specifically, the first robust dataset is the robust dataset corresponding to the first domain dataset. During the model inference phase, controlling the first model to use the first robust dataset and the processed first domain dataset for inference can greatly improve test coverage and model robustness.

[0094] In practical applications, robust datasets can be pre-generated and stored for later reuse. If robust datasets are not pre-stored, they can be generated in real-time to ensure the reliability of robustness testing. Based on this, the implementation of obtaining the first robust dataset corresponding to the first robust dataset label includes: searching multiple existing datasets for the existence of the first robust dataset based on its label; if the first robust dataset does not exist, obtaining a robust operator from multiple existing operators; and processing the first domain dataset using the robust operator to obtain the first robust dataset. Optionally, the first robust dataset and its corresponding first robust dataset label can be stored together for later reuse.

[0095] The following is combined with Figure 3 The process of generating a robust dataset is described. When a first robust dataset is not available in multiple existing datasets, several texts can be randomly selected from the multiple texts included in the first domain dataset; for example, 20% of the texts can be selected. For each of the selected texts, the text can be segmented into multiple sentences, and operations such as word insertion, synonym replacement, or word deletion can be performed on each sentence to obtain new texts. Alternatively, for each text, repeated sentences can be added, or the last sentence can be truncated to obtain new texts. These new texts are then used as the first robust dataset, and the first robust dataset is registered in existing datasets for subsequent reuse.

[0096] 204. Based on the metadata of the first performance index, obtain the first single evaluation operator corresponding to the first performance index from multiple existing operators, and use the first single evaluation operator to perform a single evaluation on the reasoning result of the first model to obtain the single evaluation result of the first performance index.

[0097] Specifically, in the single-item evaluation stage of the model evaluation process, the first model inference result obtained in the inference stage is obtained; for any first performance index, the content to be evaluated of the first performance index is determined according to the first model inference result, and the content to be evaluated of the first performance index is processed using the first single-item evaluation operator to obtain the single-item evaluation result of the first performance index.

[0098] Optionally, the model evaluation platform can also determine the target number of evaluations based on the first task configuration information. Correspondingly, the method for evaluating the first model inference result using the first single-item evaluation operator to obtain the single-item evaluation result of the first performance index is as follows: The first single-item evaluation operator is used to evaluate the first model inference result for the target number of evaluations, obtaining sub-single-item evaluation results for each single-item evaluation; the sub-single-item evaluation results of each single-item evaluation are then weighted to obtain the single-item evaluation result of the first performance index. Therefore, by performing multiple single-item evaluations and fusing the results, the reliability of the single-item evaluation results can be effectively improved.

[0099] 205. Obtain a comprehensive evaluation operator from multiple existing operators, and use the comprehensive evaluation operator to comprehensively evaluate the individual evaluation results of multiple first performance indicators, obtain the first comprehensive evaluation result, and return it to the front end.

[0100] Specifically, in the comprehensive evaluation phase of the model evaluation process, the individual evaluation results of multiple first performance indicators obtained in the single evaluation phase are acquired. A comprehensive evaluation operator is then used to comprehensively evaluate these individual evaluation results, yielding a first comprehensive evaluation result which is returned to the front end. This process further integrates and comprehensively evaluates the results of multiple single evaluations, moving beyond a single performance indicator to fully examine various performance indicators of the model and obtain a comprehensive and accurate evaluation result.

[0101] Optionally, the comprehensive evaluation operator can be used to comprehensively evaluate the individual evaluation results of multiple first performance indicators to obtain the comprehensive evaluation result. This can be achieved by: obtaining the manual evaluation results of multiple first performance indicators; and then using the comprehensive evaluation operator to comprehensively evaluate both the individual evaluation results and the manual evaluation results of the multiple first performance indicators to obtain the comprehensive evaluation result. Therefore, comprehensively evaluating both the individual evaluation results of multiple first performance indicators obtained through automated evaluation using the model evaluation platform and the manual evaluation results of multiple first performance indicators ensures a high degree of scientific rigor and reliability in the evaluation process.

[0102] The technical solution provided in this application allows users to customize task configuration information, thereby supporting the selection of different metadata, datasets, and operators in the model evaluation process. This meets constantly changing service needs and improves the flexibility of model evaluation. The metadata-based adaptive data processing mechanism can automatically match suitable models, datasets, and operators based on metadata, reducing manual intervention and improving evaluation efficiency. The model evaluation process covers three main stages: model inference, individual evaluation, and comprehensive evaluation. It is no longer limited to a single performance indicator but comprehensively examines various performance indicators of the model, providing more comprehensive evaluation results that better reflect the overall performance of the model. Therefore, it provides a comprehensive, scalable, and highly automated model evaluation service that can flexibly respond to diverse service needs and is suitable for various complex application scenarios.

[0103] In some optional embodiments, the model evaluation platform can meet model inference requirements, and users can trigger model inference requests on the front-end interface as needed. Based on this, the model evaluation platform can also receive model inference requests from the front-end, which include second task configuration information; determine the metadata of the second model and the metadata of the second domain dataset based on multiple existing metadata and the second task configuration information; obtain the second domain dataset and its corresponding second dataset processing operator from multiple existing datasets and multiple existing operators based on the metadata of the second domain dataset, and process the second domain dataset using the second dataset processing operator; trigger the second model to perform inference based on the processed second domain dataset according to the metadata of the second model, and obtain the second model inference result; and associate and store the second model inference result and its corresponding identifier.

[0104] For details on how the model evaluation platform can independently perform model inference tasks, please refer to the relevant description of the model inference stage in the full-process evaluation process in the aforementioned embodiments, which will not be repeated here.

[0105] For model evaluation platforms that execute model inference tasks separately, the inference results of the second model and its corresponding identifiers can be stored together for subsequent reuse, thereby accelerating the evaluation efficiency of single-item evaluation.

[0106] In some optional embodiments, the model evaluation platform can meet individual evaluation needs, allowing users to trigger individual evaluation requests on the front-end interface as needed. Based on this, the model evaluation platform can also receive individual evaluation requests from the front-end, which include third task configuration information; obtain the third model inference result from multiple existing model inference results based on the identifier of the third model inference result in the third task configuration information; determine the metadata of the second performance indicator based on multiple existing metadata and the third task configuration information; obtain the second individual evaluation operator corresponding to the second performance indicator from multiple existing operators based on the metadata of the second performance indicator, and use the second individual evaluation operator to perform an individual evaluation on the third model inference result to obtain the individual evaluation result of the second performance indicator and return it to the front-end; and associate and store the individual evaluation result of the second performance indicator and its corresponding identifier.

[0107] For details on how the model evaluation platform can perform individual evaluation tasks, please refer to the relevant description of the individual evaluation stage in the full-process evaluation process in the aforementioned embodiments, which will not be repeated here.

[0108] For individual evaluation tasks performed on the model evaluation platform, if reusable model inference results exist, the individual evaluation is performed based on these results, and the individual evaluation results and their corresponding identifiers are associated and stored for subsequent reuse, thus accelerating the efficiency of comprehensive evaluation. If no reusable model inference results are available, a prompt message can be output to indicate that the individual evaluation task cannot be executed.

[0109] In some optional embodiments, the model evaluation platform can meet comprehensive evaluation requirements, allowing users to trigger comprehensive evaluation requests on the front-end interface as needed. Based on this, the model evaluation platform can also receive comprehensive evaluation requests from the front-end, which include identifiers of multiple target individual evaluation results; retrieve multiple target individual evaluation results from multiple existing individual evaluation results based on these identifiers; obtain a comprehensive evaluation operator from multiple existing operators; and use the comprehensive evaluation operator to perform a comprehensive evaluation of the multiple target individual evaluation results to obtain a second comprehensive evaluation result, which is then returned to the front-end.

[0110] For details on how the model evaluation platform can independently perform the comprehensive evaluation task, please refer to the relevant introduction in the comprehensive evaluation stage of the full-process evaluation process in the aforementioned embodiments, which will not be repeated here.

[0111] For the model evaluation platform to perform a comprehensive evaluation task independently, if reusable individual evaluation results exist, the comprehensive evaluation is performed based on these individual results to achieve efficient comprehensive evaluation. If no reusable individual evaluation results exist, a prompt message can be output to indicate that the comprehensive evaluation task cannot be executed.

[0112] Figure 4 This is a schematic diagram illustrating another application scenario provided by an embodiment of this application.

[0113] Existing model evaluation solutions often focus on single-dimensional metric testing, lacking in-depth examination of the model's overall performance and robustness. They typically do not support fully automated dataset management and automatic generation of robust datasets, making it difficult to flexibly respond to constantly changing service requirements and customize new evaluation metrics. Therefore, this paper proposes a new model evaluation platform that provides a comprehensive, scalable, and highly automated model evaluation service to adapt to diverse service needs and datasets. The core of the model evaluation platform mainly includes the following aspects:

[0114] 1. Adaptive data processing mechanism: Utilizing various metadata and operators, it can intelligently identify and process datasets from multiple sources and at multiple levels, flexibly responding to different types and structures of data, and achieving efficient data processing and optimized configuration.

[0115] 2. Dynamic Robust Dataset Generation Mechanism: This mechanism can simulate real-world data noise and anomalies using various text perturbation strategies, thereby generating robust datasets. It allows users to configure different anomalous input conditions as needed, improving the model's test coverage and robustness under adversarial attacks.

[0116] 3. Dynamic registration and expansion mechanism for performance metrics: It supports the instant addition of new performance metrics and can automatically integrate them into the model evaluation platform, which enables the model evaluation platform to better adapt to ever-changing service needs.

[0117] 4. Comprehensive Evaluation Mechanism: This mechanism enables in-depth integration and comprehensive evaluation of multiple individual evaluation results, moving beyond single indicators to comprehensively examine model performance. Furthermore, it can incorporate expert evaluation results to enhance the reliability of the assessment. Thus, this multi-source comprehensive evaluation mechanism strengthens the scientific rigor and reliability of the evaluation.

[0118] 5. Dynamic operator registration and extension mechanism: Allows users to extend various operators as needed (such as dataset operators, performance metric operators, etc.) to quickly respond to changes in service requirements.

[0119] 6. Fully automated dataset management: Dynamically register metadata for various datasets to achieve automated dataset management.

[0120] 7. Simplified configuration and process management: The configuration work required when adding new datasets or models is simplified, making the entire model evaluation process more efficient.

[0121] See Figure 4 The model evaluation platform comprises a front-end and a back-end server. The back-end server includes a model evaluation service API, an operator management module, a metadata management module, and a task management module. The platform relies on various model APIs, a dataset SDK (Software Development Kit), and a human evaluation platform. The complete model evaluation process of the platform includes inference, individual evaluation, and comprehensive evaluation phases. The platform supports full-cycle evaluation, inference, individual evaluation, and comprehensive evaluation modes. Inference results generated during the recommendation phase are stored for use in subsequent individual evaluation phases; individual evaluation results from the individual evaluation phases are stored for use in subsequent comprehensive evaluation phases.

[0122] For the metadata management module: Users can trigger metadata management requests on the front end, sending these requests to the task management module via the model evaluation service API. The task management module then requests the metadata management module to perform metadata management operations. For example, registering new metadata in the metadata management module, or deleting or modifying existing metadata. The metadata management module provides several existing metadata types, including but not limited to: model metadata, performance metric metadata, service metadata, and dataset metadata. Service metadata may include mappings between services and performance metrics or between services and datasets.

[0123] Alternatively, the task management module can also register multiple datasets in the metadata management module.

[0124] For the operator management module: Users can trigger operator management requests on the front end, sending these requests to the task management module via the model evaluation service API. The task management module then requests the operator management module to perform operator management operations. For example, registering new operators in the operator management module, or deleting or modifying existing operators, etc. The metadata management module provides several existing operators, including but not limited to: dataset processing operators, robust operators, single-item evaluation operators, comprehensive evaluation operators, etc. Robust operators can be combined in multiple dimensions.

[0125] For the task management module: it can execute automatic evaluation tasks, robust dataset generation tasks, or request manual evaluation from a human evaluation platform. Robust dataset generation tasks can be executed independently or during the execution of automatic evaluation tasks. When executing tasks, the task management module can perform necessary preprocessing operations. These preprocessing operations include: generating task status files and task locks to prevent the frontend from frequently initiating the same task in a short period (e.g., preventing repeated execution of a task while it is in progress); registering model APIs or datasets in the metadata module, etc.

[0126] When executing automated evaluation tasks, the task management module can first configure an execution script template to generate an execution script, and then execute the script to perform the automated evaluation task. Automated evaluation tasks can be full-cycle evaluation tasks, inference tasks, single-item evaluation tasks, or comprehensive evaluation tasks. The task management module can also perform status monitoring, such as monitoring whether the inference phase, single-item evaluation phase, or comprehensive evaluation phase has been completed.

[0127] The following section introduces several specific tasks performed by the model evaluation platform:

[0128] Full-cycle evaluation task: Users configure the full-cycle evaluation task on the front-end interface, setting the task configuration information and triggering the full-cycle evaluation request. The front-end sends the full-cycle evaluation request to the task management module via the model evaluation service API. The task management module interacts with the metadata management module, operator management module, or manual evaluation platform to perform inference, individual evaluation, and comprehensive evaluation sequentially according to the model evaluation process. In the comprehensive evaluation stage, the results of multiple evaluations can be summarized, and the final comprehensive evaluation result is returned to the front-end for user viewing.

[0129] Inference Task: Users perform task configuration operations on the front-end interface, configure the task configuration information of the inference task, and trigger the inference request. The front-end sends the inference request to the task management module through the model evaluation service API. The task management module performs inference through interaction with the metadata management module and the operator management module, and stores the model inference results for subsequent individual evaluation.

[0130] Individual Evaluation Task: Users configure individual evaluation tasks on the front-end interface, setting the task configuration information and triggering an evaluation request. The front-end sends the evaluation request to the task management module via the model evaluation service API. The task management module performs the individual evaluation through interaction with the metadata management module, operator management module, or manual evaluation platform, and stores the evaluation results for subsequent comprehensive evaluation.

[0131] Comprehensive evaluation task: Users perform task configuration operations on the front-end interface to configure the task configuration information and trigger a comprehensive evaluation request. The front-end sends the comprehensive evaluation request to the task management module through the model evaluation service API. The task management module performs a comprehensive evaluation by interacting with the metadata management module, operator management module or manual evaluation platform, and returns the comprehensive evaluation result to the front-end.

[0132] Robust Dataset Generation Task: Users configure the robust dataset generation task on the front-end interface, setting the task configuration information and triggering a robust dataset generation request. The front-end sends the robust dataset generation request to the task management module via the model evaluation service API. The task management module interacts with the metadata management module and operator management module to generate the robust dataset and register it in the metadata management module. Subsequent testing of model robustness can extract the corresponding robust dataset from the metadata management module for testing.

[0133] Figure 5 A system architecture diagram of a model evaluation platform provided in an embodiment of this application. See also... Figure 5 The model evaluation platform may include a front-end 10 and a back-end server 20. The back-end server 20 may include a metadata management module 21, an operator management module 22, and a task management module 23.

[0134] The metadata management module is used to provide multiple existing metadata and multiple datasets;

[0135] The operator management module is used to provide multiple existing operators;

[0136] The task management module receives end-to-end evaluation requests from the front end. These requests instruct the execution of model inference, individual evaluations, and comprehensive evaluations within the model evaluation process. The end-to-end evaluation request includes first task configuration information. Based on multiple existing metadata sources provided by the metadata management module and the first task configuration information, the module determines the metadata of the first model, the metadata of the first domain dataset, and the metadata of multiple first performance metrics. Based on the metadata of the first domain dataset, the module retrieves the first domain dataset and its corresponding first dataset processing operators from multiple existing datasets provided by the metadata management module and multiple existing operators provided by the operator management module. The system processes the first domain dataset using the first dataset processing operator; it triggers the first model to perform inference based on the processed first domain dataset according to the metadata of the first model, and obtains the inference result of the first model; based on the metadata of the first performance index, it obtains the first single evaluation operator corresponding to the first performance index from multiple existing operators, and uses the first single evaluation operator to perform single evaluation on the inference result of the first model, and obtains the single evaluation result of the first performance index; it obtains the comprehensive evaluation operator from multiple existing operators, and uses the comprehensive evaluation operator to perform comprehensive evaluation on the single evaluation results of multiple first performance indices, and obtains the first comprehensive evaluation result and returns it to the front end.

[0137] Optionally, the task management module is also used to determine the target number of evaluations based on the first task configuration information; correspondingly, when the task management module evaluates the reasoning result of the first model using the first single-item evaluation operator to obtain the single-item evaluation result of the first performance index, it is specifically used to: evaluate the target number of evaluations of the reasoning result of the first model using the first single-item evaluation operator to obtain the sub-single-item evaluation result of each single-item evaluation; and perform weighted processing on the sub-single-item evaluation results of each single-item evaluation to obtain the single-item evaluation result of the first performance index.

[0138] Optionally, the first task configuration information includes: the first robust dataset label corresponding to the first domain dataset.

[0139] The task management module is also used to obtain the first robust dataset corresponding to the first robust dataset label.

[0140] Accordingly, when the task management module triggers the first model to perform inference based on the processed first domain dataset according to the metadata of the first model, and obtains the inference result of the first model, it is specifically used to: trigger the first model to perform inference based on the first robust dataset and the processed first domain dataset according to the metadata of the first model, and obtain the inference result of the first model.

[0141] Optionally, when the task management module obtains the first robust dataset corresponding to the first robust dataset label, it specifically performs the following: based on the first robust dataset label, it searches multiple existing datasets to see if the first robust dataset exists; if the first robust dataset does not exist, it obtains a robust operator from multiple existing operators; and it processes the first domain dataset using the robust operator to obtain the first robust dataset.

[0142] Optionally, when the task management module uses the comprehensive evaluation operator to comprehensively evaluate the individual evaluation results of multiple first performance indicators and obtain the first comprehensive evaluation result, it is specifically used to: obtain the manual evaluation results of multiple first performance indicators; and use the comprehensive evaluation operator to comprehensively evaluate the individual evaluation results and manual evaluation results of multiple first performance indicators to obtain the first comprehensive evaluation result.

[0143] Optionally, when the task management module determines the metadata of the first model, the metadata of the first domain dataset, and the metadata of multiple first performance metrics based on multiple existing metadata and the first task configuration information, it specifically performs the following: searching for the metadata of the first model and the metadata of the first service in multiple existing metadata based on the basic attribute information of the first model and the basic attribute information of the first service in the first task configuration information; searching for the metadata of the first domain dataset in multiple existing metadata based on the existence of the metadata of the first service and the basic attribute information of the first domain dataset in the first task configuration information; and searching for the metadata of multiple first performance metrics in multiple existing metadata based on the existence of the metadata of the first service and the basic attribute information of the multiple first performance metrics in the first task configuration information.

[0144] Optionally, when the task management module searches for the metadata of the first domain dataset among multiple existing metadata based on the existence of the metadata of the first service and the basic attribute information of the first domain dataset in the first task configuration information, it specifically performs the following: if the first task configuration information does not include the basic attribute information of the first domain dataset, then it searches for the metadata of the first domain dataset among multiple existing metadata based on the mapping relationship between the service and the dataset in the metadata of the first service; if the first task configuration information includes the basic attribute information of the first domain dataset, then it searches for the metadata of the first domain dataset among multiple existing metadata based on the basic attribute information of the first domain dataset.

[0145] Optionally, when the task management module searches for the metadata of multiple first performance indicators in multiple existing metadata based on the existence of the basic attribute information of multiple first performance indicators in the metadata of the first service and the configuration information of the first task, it specifically performs the following: if the configuration information of the first task does not include the basic attribute information of multiple first performance indicators, then it searches for the metadata of multiple first performance indicators in multiple existing metadata based on the mapping relationship between services and performance indicators in the metadata of the first service; if the configuration information of the first task includes the basic attribute information of multiple first performance indicators, then it searches for the metadata of multiple first performance indicators in multiple existing metadata based on the basic attribute information of multiple first performance indicators.

[0146] Optionally, the task management module is also used to: receive a model inference request from the front end, the model inference request including second task configuration information; determine the metadata of the second model and the metadata of the second domain dataset based on multiple existing metadata and the second task configuration information; obtain the second domain dataset and its corresponding second dataset processing operator from multiple existing datasets and multiple existing operators based on the metadata of the second domain dataset, and process the second domain dataset using the second dataset processing operator; trigger the second model to perform inference based on the processed second domain dataset based on the metadata of the second model, and obtain the second model inference result; and associate and store the second model inference result and its corresponding identifier.

[0147] Optionally, the task management module is also used to: receive a single evaluation request from the front end, the single evaluation request including third task configuration information; obtain the third model inference result from multiple existing model inference results according to the identifier of the third model inference result in the third task configuration information; determine the metadata of the second performance indicator according to multiple existing metadata and the third task configuration information; obtain the second single evaluation operator corresponding to the second performance indicator from multiple existing operators according to the metadata of the second performance indicator, and use the second single evaluation operator to perform a single evaluation on the third model inference result to obtain the single evaluation result of the second performance indicator and return it to the front end; and associate and store the single evaluation result of the second performance indicator and its corresponding identifier.

[0148] Optionally, the task management module is also used to: receive a comprehensive evaluation request from the front end, the comprehensive evaluation request including the identifiers of multiple target individual evaluation results; obtain multiple target individual evaluation results from multiple existing individual evaluation results based on the identifiers of multiple target individual evaluation results; obtain a comprehensive evaluation operator from multiple existing operators, and use the comprehensive evaluation operator to perform a comprehensive evaluation on the multiple target individual evaluation results to obtain a second comprehensive evaluation result and return it to the front end.

[0149] Optionally, the task management module is also used to: receive management operation requests from the front end, and perform corresponding management operations on multiple existing operators, datasets or metadata. The management operations include any of the following: add operation, delete operation and update operation.

[0150] Optionally, the management operation request includes a robust operator addition request, which includes: basic text operation operators, operation object operators, and composite operators;

[0151] Accordingly, when the task management module performs corresponding management operations on multiple existing operators, it is specifically used to: add basic text operation operators, operation object operators, and combination operators to multiple existing operators;

[0152] Among them, basic text operation operators are used to modify text content. Basic text operation operators include at least one of the following: random deletion operator, random insertion operator, random replacement operator, slightly shuffling word order operator, synonym replacement operator, repeated segment multiple times operator, input truncation operator, and construction of meaningless text operator; operation object operators are used to extract operation objects from text content, and the operation objects are characters, words, sentences, and paragraphs; combination operators are used to process the operation objects using basic operation operators.

[0153] Figure 5 The model evaluation platform shown can execute the model evaluation method in the foregoing embodiments, and its implementation principle and technical effects will not be elaborated here.

[0154] It should be noted that the execution subject of each step of the method provided in the above embodiments can be the same device, or the method can be executed by different devices. For example, the execution subject of steps 201 to 203 can be device A; or the execution subject of steps 201 and 202 can be device A, and the execution subject of step 203 can be device B, etc.

[0155] Furthermore, in some of the processes described in the above embodiments and accompanying drawings, multiple operations appear in a specific order. However, it should be clearly understood that these operations may not be executed in the order they appear herein, or they may be executed in parallel. The operation numbers, such as 201, 202, etc., are merely used to distinguish different operations and do not represent any execution order. Additionally, these processes may include more or fewer operations, and these operations may be executed sequentially or in parallel. It should be noted that the descriptions such as "first" and "second" in this document are used to distinguish different messages, devices, modules, etc., and do not represent a sequential order, nor do they limit "first" and "second" to different types.

[0156] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties. Furthermore, the collection, use and processing of the relevant data must comply with the relevant laws, regulations and standards of the relevant countries and regions, and corresponding operation entry points are provided for users to choose to authorize or refuse.

[0157] Figure 6 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Figure 6 As shown, the electronic device includes a memory 61 and a processor 62.

[0158] Memory 61 is used to store computer programs and can be configured to store various other data to support operation on the computing platform. Examples of this data include instructions for any application or method operating on the computing platform, contact data, phone book data, messages, pictures, videos, etc.

[0159] The memory 61 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random-access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk.

[0160] Processor 62, coupled to memory 61, is used to execute computer programs in memory 61 for: performing steps in the model evaluation method.

[0161] Optional, such as Figure 6 As shown, the electronic device also includes other components such as a communication component 63, a display 64, a power supply component 65, and an audio component 66. Figure 6 The diagram only shows some components and does not mean that the electronic device includes only these components. Figure 6 The components shown. Additionally... Figure 6The components within the dashed box are optional, not mandatory, and their specific requirements depend on the product form of the electronic device. The electronic device in this embodiment can be a desktop computer, laptop computer, smartphone, or IoT (Internet of Things) device, or a server-side device such as a conventional server, cloud server, or server array. If the electronic device in this embodiment is a desktop computer, laptop computer, smartphone, or other terminal device, it may include... Figure 6 The components within the dashed box; if the electronic device in this embodiment is implemented as a conventional server, cloud server, or server array, etc., it may be omitted. Figure 6 The component within the dashed box.

[0162] For a detailed description of the implementation process of each action by the processor, please refer to the relevant descriptions in the foregoing method embodiments or device embodiments, which will not be repeated here.

[0163] Accordingly, embodiments of this application also provide a computer-readable storage medium storing a computer program, which, when executed, can implement the steps that can be performed by an electronic device in the above method embodiments.

[0164] Accordingly, this application also provides a computer program product, including a computer program / instructions, which, when executed by a processor, enables the processor to perform the steps that can be executed by an electronic device in the above method embodiments.

[0165] The aforementioned communication component is configured to facilitate wired or wireless communication between the device containing the communication component and other devices. The device containing the communication component can access wireless networks based on communication standards, such as WiFi (Wireless Fidelity), 2G (2nd Generation), 3G (3rd Generation), 4G (4th Generation) / LTE (Long Term Evolution), 5G (5th Generation), or combinations thereof. In one exemplary embodiment, the communication component receives broadcast signals or broadcast-related information from an external broadcast management system via a broadcast channel. In one exemplary embodiment, the communication component also includes a Near Field Communication (NFC) module to facilitate short-range communication. For example, NFC modules can be implemented based on Radio Frequency Identification (RFID) technology, Infrared Data Association (IRDA) technology, Ultra Wide Band (UWB) technology, Bluetooth (BT) technology, and other technologies.

[0166] The aforementioned display includes a screen, which may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen can be implemented as a touchscreen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touches, swipes, and gestures on the touch panel. The touch sensors can sense not only the boundaries of touch or swipe actions but also the duration and pressure associated with the touch or swipe operation.

[0167] The aforementioned power supply components provide power to various components within the device in which they reside. These power supply components may include a power management system, one or more power sources, and other components associated with generating, managing, and distributing power to the device in which they reside.

[0168] The aforementioned audio component can be configured to output and / or input audio signals. For example, the audio component includes a microphone (MIC) configured to receive external audio signals when the device containing the audio component is in an operating mode, such as call mode, recording mode, and voice recognition mode. The received audio signals can be further stored in memory or transmitted via a communication component. In some embodiments, the audio component also includes a speaker for outputting audio signals.

[0169] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-readable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0170] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart... Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0171] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0172] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0173] In a typical configuration, an electronic device includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.

[0174] Memory may include non-persistent storage in computer-readable media, such as random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash memory. Memory is an example of computer-readable media.

[0175] Computer-readable media, including both permanent and non-permanent, removable and non-removable media, can store information using any method or technology. Information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase-change RAM (PRAM), static random-access memory (SRAM), dynamic random-access memory (DRAM), other types of random-access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other non-transfer medium that can be used to store information accessible by electronic devices. As defined in this article, computer-readable media do not include transient computer-readable media, such as modulated data signals and carrier waves.

[0176] It should also be noted that the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.

[0177] The above are merely embodiments of this application and are not intended to limit the scope of this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of the claims of this application.

Claims

1. A model evaluation method, characterized by, The method comprises the following steps: receiving a full-process evaluation request from a front end, the full-process evaluation request being used to indicate that model inference, single-item evaluation and comprehensive evaluation are performed in a model evaluation process, and the full-process evaluation request comprising first task configuration information; determining metadata of a first model, metadata of a first domain data set and metadata of a plurality of first performance indexes according to a plurality of existing metadata and the first task configuration information; acquiring the first domain data set and a corresponding first data set processing operator from a plurality of existing data sets and a plurality of existing operators respectively according to the metadata of the first domain data set, and processing the first domain data set by using the first data set processing operator; triggering the first model to perform inference based on the processed first domain data set according to the metadata of the first model, and obtaining a first model inference result; acquiring a first single-item evaluation operator corresponding to the first performance index from a plurality of existing operators according to the metadata of the first performance index, and performing single-item evaluation on the first model inference result by using the first single-item evaluation operator, and obtaining a single-item evaluation result of the first performance index; acquiring a comprehensive evaluation operator from a plurality of existing operators, and performing comprehensive evaluation on the single-item evaluation results of the plurality of first performance indexes by using the comprehensive evaluation operator, and obtaining a first comprehensive evaluation result and returning the first comprehensive evaluation result to the front end.

2. The method of claim 1, wherein, The method further comprises the following steps: determining a target evaluation times according to the first task configuration information; performing evaluation on the first model inference result by using the first single-item evaluation operator to obtain the single-item evaluation result of the first performance index, comprising the following steps: performing evaluation on the first model inference result by using the first single-item evaluation operator for the target evaluation times to obtain sub-single-item evaluation results of each single-item evaluation; performing weighted processing on the sub-single-item evaluation results of each single-item evaluation to obtain the single-item evaluation result of the first performance index.

3. The method of claim 1, wherein, The first task configuration information comprises a first robust data set label corresponding to the first domain data set; acquiring a first robust data set corresponding to the first robust data set label; triggering the first model to perform inference based on the processed first domain data set according to the metadata of the first model to obtain a first model inference result, comprising the following steps: triggering the first model to perform inference based on the first robust data set and the processed first domain data set according to the metadata of the first model to obtain a first model inference result.

4. The method of claim 3, wherein, acquiring a first robust data set corresponding to the first robust data set label, comprising the following steps: searching a plurality of existing data sets according to the first robust data set label to determine whether the first robust data set exists in the plurality of existing data sets; if the first robust data set does not exist, acquiring a robust operator from a plurality of existing operators; processing the first domain data set by using the robust operator to obtain the first robust data set.

5. The method of claim 3, wherein, performing comprehensive evaluation on the single-item evaluation results of the plurality of first performance indexes by using the comprehensive evaluation operator to obtain a first comprehensive evaluation result, comprising the following steps: acquiring artificial evaluation results of the plurality of first performance indexes; The single evaluation results of the plurality of first performance indicators and the artificial evaluation results are comprehensively evaluated by using the comprehensive evaluation operator to obtain a first comprehensive evaluation result.

6. The method according to any one of claims 1 to 5, characterized in that, According to the plurality of existing metadata and the first task configuration information, metadata of the first model, metadata of the first domain data set, and metadata of the plurality of first performance indicators are determined, including: According to the basic attribute information of the first model and the basic attribute information of the first service in the first task configuration information, the metadata of the first model and the metadata of the first service are searched in the plurality of existing metadata respectively; According to the metadata of the first service and the existence of the basic attribute information of the first domain data set in the first task configuration information, the metadata of the first domain data set is searched in the plurality of existing metadata; According to the metadata of the first service and the existence of the basic attribute information of the plurality of first performance indicators in the first task configuration information, the metadata of the plurality of first performance indicators is searched in the plurality of existing metadata.

7. The method of claim 6, wherein, According to the metadata of the first service and the existence of the basic attribute information of the first domain data set in the first task configuration information, the metadata of the first domain data set is searched in the plurality of existing metadata, including: If the basic attribute information of the first domain data set is not included in the first task configuration information, the metadata of the first domain data set is searched in the plurality of existing metadata according to the mapping relationship between the service and the data set in the metadata of the first service; If the basic attribute information of the first domain data set is included in the first task configuration information, the metadata of the first domain data set is searched in the plurality of existing metadata according to the basic attribute information of the first domain data set.

8. The method of claim 6, wherein, According to the metadata of the first service and the existence of the basic attribute information of the plurality of first performance indicators in the first task configuration information, the metadata of the plurality of first performance indicators is searched in the plurality of existing metadata, including: If the basic attribute information of the plurality of first performance indicators is not included in the first task configuration information, the metadata of the plurality of first performance indicators is searched in the plurality of existing metadata according to the mapping relationship between the service and the performance indicator in the metadata of the first service; If the basic attribute information of the plurality of first performance indicators is included in the first task configuration information, the metadata of the plurality of first performance indicators is searched in the plurality of existing metadata according to the basic attribute information of the plurality of first performance indicators.

9. The method of claim 1, wherein, Further comprising: receiving a model inference request from the front end, the model inference request including second task configuration information; According to the plurality of existing metadata and the second task configuration information, metadata of the second model and metadata of the second domain data set are determined; According to the metadata of the second domain data set, the second domain data set and its corresponding second data set processing operator are obtained from the plurality of existing data sets and the plurality of existing operators respectively, and the second domain data set is processed by using the second data set processing operator; Trigger the second model to perform inference based on the processed second domain data set according to the metadata of the second model, to obtain a second model inference result; Store the second model inference result and its corresponding identifier in association.

10. The method of claim 9, wherein, Further comprising: Receiving a single-item evaluation request from the front end, the single-item evaluation request including third task configuration information; According to the identifier of the third model inference result in the third task configuration information, obtaining the third model inference result from a plurality of existing model inference results; According to a plurality of existing metadata and the third task configuration information, determining metadata of a second performance indicator; According to the metadata of the second performance indicator, obtaining a second single-item evaluation operator corresponding to the second performance indicator from a plurality of existing operators, and using the second single-item evaluation operator to perform single-item evaluation on the third model inference result, to obtain a single-item evaluation result of the second performance indicator and return it to the front end; Store the single-item evaluation result of the second performance indicator and its corresponding identifier in association.

11. The method of claim 10, wherein, Further comprising: Receiving a comprehensive evaluation request from the front end, the comprehensive evaluation request including identifiers of a plurality of target single-item evaluation results; According to the identifiers of the plurality of target single-item evaluation results, obtaining the plurality of target single-item evaluation results from a plurality of existing single-item evaluation results; Obtaining a comprehensive evaluation operator from a plurality of existing operators, and using the comprehensive evaluation operator to perform comprehensive evaluation on the plurality of target single-item evaluation results, to obtain a second comprehensive evaluation result and return it to the front end.

12. The method according to any one of claims 1 to 5, characterized in that, Further comprising: Receiving a management operation request from the front end, and performing corresponding management operations on a plurality of existing operators, data sets or metadata, the management operations including any one of the following: addition operation, deletion operation and update operation.

13. The method of claim 12, wherein, The management operation request includes a robust operator addition request, and the robust operator addition request includes a basic text operation operator, an operation object operator and a combination type operator; Performing corresponding management operations on a plurality of existing operators, including: Adding the basic text operation operator, the operation object operator and the combination type operator to the plurality of existing operators; The basic text operation operator is used to modify text content, and the basic text operation operator includes at least one of the following: random deletion operator, random insertion operator, random replacement operator, light word order disordering operator, synonym replacement operator, repeated fragment multiple times operator, truncated input operator and meaningless text construction operator; The operation object operator is used to extract operation objects from the text content, and the operation objects are words, phrases, sentences and paragraphs; The combination type operator is used to process the operation objects using the basic operation operator.

14. A model evaluation platform, characterized in that, Comprising: A front end and a back-end server, the back-end server including a metadata management module, an operator management module and a task management module; The metadata management module is used to provide a plurality of existing metadata and a plurality of data sets; The operator management module is used to provide a plurality of existing operators; The task management module is configured to receive a full-process evaluation request from a front end, the full-process evaluation request being used to indicate that model inference, single-item evaluation, and comprehensive evaluation are performed in a model evaluation process, the full-process evaluation request including first task configuration information; determine metadata of a first model, metadata of a first domain data set, and metadata of a plurality of first performance indicators according to a plurality of existing metadata and the first task configuration information; According to the metadata of the first domain data set, a first domain data set and a corresponding first data set processing operator are respectively acquired from a plurality of existing data sets and a plurality of existing operators, and the first domain data set is processed by using the first data set processing operator; According to the metadata of the first model, the first model is triggered to perform inference based on the processed first domain data set, and a first model inference result is obtained; According to the metadata of the first performance indicators, a first single-item evaluation operator corresponding to the first performance indicators is acquired from a plurality of existing operators, and the first single-item evaluation operator is used to perform single-item evaluation on the first model inference result, and a single-item evaluation result of the first performance indicators is obtained; a comprehensive evaluation operator is acquired from a plurality of existing operators, and the comprehensive evaluation operator is used to perform comprehensive evaluation on the single-item evaluation results of the plurality of first performance indicators, and a first comprehensive evaluation result is obtained and returned to the front end.

15. An electronic device, comprising: Comprising: a memory and a processor; the memory is configured to store a computer program; the processor is coupled to the memory and is configured to execute the computer program to perform the steps in the method of any one of claims 1-13.

16. A computer readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, the processor is enabled to implement the steps in the method of any one of claims 1-13.

17. A computer program product, characterised in that, comprising computer programs / instructions, when the computer programs / instructions are executed by the processor, the processor is enabled to implement the steps in the method of any one of claims 1-13.