Model evaluation method and device, electronic equipment and storage medium
By using model evaluation methods to migrate the target model from the graphics card environment to the domestic chip environment, and using correctness and performance evaluation tools for automated evaluation, the problem of lack of systematic tools in the migration process of domestic chips is solved, and standardized evaluation and data accumulation are achieved, thereby improving the adaptation efficiency and reliability.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- BEIJING ACAD OF ARTIFICIAL INTELLLIGENCE
- Filing Date
- 2026-01-09
- Publication Date
- 2026-06-05
AI Technical Summary
Due to significant differences between domestic AI chips and NVIDIA chips in hardware architecture, instruction sets, runtime, and operator implementation, there is a lack of systematic tools for correctness verification when migrating from an NVIDIA environment to domestic chips. Performance evaluation is inconsistent, data is difficult to accumulate, and there is a lack of end-to-end automated toolchains, resulting in low efficiency of chip ecosystem adaptation.
This paper provides a model evaluation method that migrates the target model from the graphics card environment to the target chip environment through model management and adaptation tools, and performs automated evaluation using correctness evaluation and performance evaluation tools, including correctness evaluation and performance evaluation, to generate standardized performance data and form a long-term queryable chip performance database.
It has achieved standardization and automation of correctness evaluation, reduced human error and repetitive work, generated reproducible and quantifiable correctness judgments, improved the efficiency of domestic chip ecosystem adaptation, reduced labor costs and shortened the delivery cycle.
Smart Images

Figure CN122152649A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of testing technology, and in particular to a model evaluation method, apparatus, electronic device, and storage medium. Background Technology
[0002] With the development of large models, open-source model-as-a-service sharing platforms such as ModelScope and model developer social platforms such as Hugging Face have become the mainstream channels for model distribution and sharing.
[0003] The vast majority of models are trained and validated in an NVIDIA graphics processing unit (GPU) environment, that is, in an environment based on the Compute Unified Device Architecture (CUDA).
[0004] Because domestically produced artificial intelligence (AI) chips differ significantly from NVIDIA chips in hardware architecture, instruction sets, runtime, and operator implementation, there is a lack of systematic tools for correctness verification when migrating from an NVIDIA environment to domestically produced chips. Furthermore, inconsistent performance evaluation methods lead to difficulties in data accumulation. Therefore, an effective solution is urgently needed to address these issues. Summary of the Invention
[0005] To address the above problems, the present invention provides a model evaluation method, apparatus, electronic device, and storage medium.
[0006] This invention provides a model evaluation method, comprising: migrating a target model from a graphics card environment to a target chip environment based on a model management and adaptation tool; Based on the inference services of the target model in the graphics card environment and the target chip environment respectively, the correctness of the target model is evaluated, and the correctness evaluation results are obtained. If the correctness evaluation result is passed, the target model is evaluated based on the inference performance evaluation tool to obtain the performance evaluation result of the target model.
[0007] According to a model evaluation method provided by the present invention, the correctness evaluation of the target model based on the inference services of the target model in the graphics card environment and the target chip environment respectively includes: With the same configuration, a unified dataset is input into the target model in the graphics card environment and the target chip environment respectively for inference service, to obtain the first correctness index of the target model in the graphics card environment and the second correctness index of the target model in the target chip environment; Based on the first correctness indicator and the second correctness indicator, calculate the difference in correctness indicators; If the difference in the correctness index is less than or equal to the difference threshold, the correctness evaluation result is determined to be passed; If the difference in the correctness index is greater than the difference threshold, the correctness evaluation result is determined to be unsuccessful.
[0008] According to a model evaluation method provided by the present invention, the step of calculating the difference in correctness indicators based on the first correctness indicator and the second correctness indicator includes: Calculate the absolute value of the difference between the first correctness index and the second correctness index; The ratio of the absolute value to the first correctness index is determined as the difference in the correctness index.
[0009] According to a model evaluation method provided by the present invention, the step of evaluating the performance of the target model based on an inference performance evaluation tool to obtain the performance evaluation result of the target model includes: Based on the aforementioned inference performance evaluation tool, a request load is generated; Based on the requested load, inference load is initiated on the target model in the graphics card environment and the target chip environment respectively, and performance indicators are collected to obtain the performance evaluation results.
[0010] According to a model evaluation method provided by the present invention, after evaluating the performance of the target model based on an inference performance evaluation tool and obtaining the performance evaluation result of the target model, the method further includes: The correctness evaluation results and / or the performance evaluation results are associated with the target migration task to form a traceable and complete migration link record, wherein the target migration task represents the migration of the target model from the graphics card environment to the target chip environment.
[0011] According to a model evaluation method provided by the present invention, after evaluating the performance of the target model based on an inference performance evaluation tool and obtaining the performance evaluation result of the target model, the method further includes: Generate a model weight package that adapts the target model to the target chip environment; The target chip environment can be constructed to directly deploy the runtime image of the target model; Generate documentation to adapt the target chip environment to the target chip environment; The publishing tool is invoked to synchronously upload the model weight package, the running image, and the documentation to the model publishing platform in order to publish the target model adapted to the target chip environment.
[0012] According to a model evaluation method provided by the present invention, after evaluating the correctness of the target model based on the inference services of the target model in the graphics card environment and the target chip environment respectively, and obtaining the correctness evaluation result, the method further includes: If the correctness evaluation result is unsuccessful, the reason for the failure is recorded, and the reason for failure is used to repair the target chip.
[0013] The present invention also provides a model evaluation device, comprising: The migration module is configured to migrate the target model from the graphics card environment to the target chip environment based on model management and adaptation tools. The correctness evaluation module is configured to evaluate the correctness of the target model based on the inference services of the target model in the graphics card environment and the target chip environment, respectively, and obtain the correctness evaluation result. The performance evaluation module is configured to perform performance evaluation on the target model based on an inference performance evaluation tool, provided that the correctness evaluation result is passed, and obtain the performance evaluation result of the target model.
[0014] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the model evaluation method as described above.
[0015] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the model evaluation method as described above.
[0016] The present invention also provides a computer program product, including a computer program that, when executed by a processor, implements the model evaluation method as described above.
[0017] The model evaluation method, apparatus, electronic device, and storage medium provided by this invention migrate the target model from the graphics card environment to the target chip environment based on model management and adaptation tools. Based on the inference services of the target model in both the graphics card environment and the target chip environment, a correctness evaluation is performed on the target model to obtain a correctness evaluation result. If the correctness evaluation result is satisfactory, a performance evaluation is performed on the target model based on an inference performance evaluation tool to obtain a performance evaluation result. This invention achieves standardization and automation of correctness evaluation through a correctness evaluation mechanism based on the graphics card environment and the target chip environment, obtaining reproducible and quantifiable correctness judgments, reducing human error and repetitive work. The automated evaluation mechanism of the inference performance evaluation tool can generate standardized performance data, thus forming a long-term queryable chip performance database, achieving consistency and storability of performance evaluation. The systematic correctness evaluation mechanism and automated evaluation mechanism can support expansion and customization for different chip manufacturers, different models, and different testing strategies, improving practicality and applicability. Attached Figure Description
[0018] To more clearly illustrate the technical solutions in this invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.
[0019] Figure 1 This is one of the flowcharts of the model evaluation method provided by the present invention.
[0020] Figure 2 This is the second flowchart of the model evaluation method provided by the present invention.
[0021] Figure 3 This is a schematic diagram of the model evaluation device provided by the present invention.
[0022] Figure 4 This is a schematic diagram of the structure of the electronic device provided by the present invention. Detailed Implementation
[0023] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.
[0024] First, a brief description of the relevant content involved in this invention will be given.
[0025] Due to significant differences between domestically produced AI chips and NVIDIA chips, common problems encountered when migrating from an NVIDIA environment to a domestically produced target chip environment include: 1) Lack of systematic tools for correctness verification. The inference results of the model on the chip need to be compared with the NVIDIA baseline to determine whether they are correct. However, currently, comparisons are mostly done manually or with ad-hoc scripts, lacking unified input, unified comparison rules, and automatic judgment standards. This results in evaluations that are not reproducible, time-consuming, and prone to errors. 2) Inconsistent performance evaluation and difficulty in data accumulation. Scattered scripts are often used to test throughput (Tokens Per Second, TPS), latency, etc. The results are scattered and inconsistent in format, making it impossible to form a queryable chip performance database, which is difficult to use for long-term comparison and regression analysis. 3) Lack of end-to-end automated toolchain. The lack of an integrated platform from model acquisition, automatic adaptation to target chips, automated testing, result storage, and automatic release results in low chip ecosystem adaptation efficiency, high manual intervention, and slow expansion.
[0026] To address at least one of the aforementioned problems, this invention provides a model evaluation method, apparatus, electronic device, and storage medium to achieve the goals of evaluation standardization, data structuring, process automation, and result traceability.
[0027] Next, a brief explanation of the relevant terminology involved in this invention will be given.
[0028] FlagOS is an open-source system software stack for various AI chips, designed to solve technical challenges such as heterogeneous computing and high-speed interconnection in large model training and inference, focusing on solving problems such as the fragmentation of the AI computing power ecosystem and the difficulty of migrating large models; it is a framework for model management and adaptation, referred to as "FlagOS" in this article.
[0029] Domestic AI chips: refers to artificial intelligence accelerators provided by domestic manufacturers, including but not limited to chips such as Kunlunxin, Ascend, Iluvatar, MThreads, Hygon, and Enflame.
[0030] Vectorized Large Language Model Inference (vLLM) is a lightweight inference and load generation framework used to simulate inference requests and collect performance data. It can also serve as a vLLM benchmark suite.
[0031] Correctness Test: The inference output of the transferred model on the target chip is compared with the NVIDIA baseline output to determine functional / numerical consistency.
[0032] Performance Testing: Collect metrics such as throughput and latency under a uniform workload to measure inference performance.
[0033] TPS stands for requests per second.
[0034] A token is the basic unit into which language is broken down when AI processes language; Tokens / s is the number of tokens generated per second.
[0035] The README document serves as the entry point and business card for a project, showcasing its goals, features, and usage methods to visitors. A comprehensive and clear README helps developers quickly understand the project's value and purpose, thereby deciding whether to use or participate in it.
[0036] README document, model weight package, and image: Metadata and downloadable distribution files prepared for model release.
[0037] The following is combined with Figures 1 to 4 The present invention describes the model evaluation method, apparatus, electronic device, and storage medium.
[0038] Figure 1 This is one of the flowcharts illustrating the model evaluation method provided by this invention, such as... Figure 1 As shown, the method includes the following: Step 101: Using model management and adaptation tools, migrate the target model from the graphics card environment to the target chip environment; Step 102: Based on the inference services of the target model in the graphics card environment and the target chip environment respectively, perform a correctness evaluation on the target model to obtain the correctness evaluation results; Step 103: If the correctness evaluation result is passed, the target model is evaluated based on the inference performance evaluation tool to obtain the performance evaluation result of the target model.
[0039] First, the entity executing the model evaluation method can be a model evaluation system, hereinafter referred to as the system. This system can be installed on a server equipped with a graphics card and a target chip, or it can operate independently of the server.
[0040] Specifically, the model management and adaptation tool can be FlagOS, or other tools that enable model migration. The graphics card environment refers to the environment provided by NVIDIA, such as an NVIDIA GPU environment or an NVIDIA environment. The target chip environment refers to the environment provided by the target chip, which is the chip that the target model needs to be adapted to. The inference performance evaluation tool can be vLLM.
[0041] In practical applications, you can select the target model from a model publishing platform, such as ModelScope or Hugging Face, and the system will automatically download the necessary resources such as the model weight package, configuration file, and word segmenter.
[0042] After downloading, based on FlagOS's runtime loading mechanism, the corresponding chip runtime (mechanism), kernel / operator library, and adaptation layer components are dynamically loaded according to the target chip type (such as Kunlunxin, Ascend, Hygon, Enflame, etc.), and the inference service is built. Once the inference service starts successfully, the system generates a unique Uniform Resource Locator (URL) for subsequent correctness and performance evaluation.
[0043] Inference service construction refers to downloading the target model to the server. The runtime mechanism is a crucial intermediate layer connecting the front-end model compiler and the underlying hardware execution unit, responsible for core functions such as model scheduling, task switching, and resource management.
[0044] Specifically, after downloading the target model, an initial migration task needs to be performed. The system will create independent migration task instances for the NVIDIA environment and the target chip environment respectively, to record the loading, running, and evaluation information of the two sets of inference services. Each migration task includes key fields such as deployment server information, inference service port, running parameters, and runtime version, and is persistently stored to form the basic data for the model migration chain.
[0045] For example, a MigrationTask_NVIDIA instance is created for an NVIDIA environment; and a MigrationTask_Chip instance is created for a target chip environment.
[0046] After the migration task is initialized, a correctness evaluation is performed. The goal of the correctness evaluation is to verify whether the output of the target model on the target chip environment is consistent with or nearly consistent with the output on the NVIDIA environment, ensuring that operator replacement, frame replacement, etc., do not compromise the correctness results.
[0047] Specifically, the correctness evaluation process can be as follows: based on the correctness evaluation mechanism of unified input, dual-end operation and automatic difference comparison, under the unified dataset and configuration, inference is run in the NVIDIA environment and the target chip environment respectively, and the output is automatically compared; the output after the transfer is judged by the predefined difference metric and threshold to determine whether it is acceptable, that is, the correctness evaluation result is obtained.
[0048] After the correctness evaluation is completed, if the correctness evaluation result is passed, the automatic performance evaluation system based on vLLM, that is, based on the inference performance evaluation tool, generates inference load and schedules requests, and the automatic performance indicators are based on the performance evaluation results.
[0049] The model evaluation method provided by this invention achieves standardization and automation of correctness evaluation through a correctness evaluation mechanism based on the graphics card environment and the target chip environment, obtaining reproducible and quantifiable correctness judgments, reducing human error and repetitive work; through the automated evaluation mechanism of the inference performance evaluation tool, standardized performance data can be generated to form a long-term queryable chip performance database, achieving consistency and storability of performance evaluation; the systematic correctness evaluation mechanism and automated evaluation mechanism can support the expansion and customization of different chip manufacturers, different models and different testing strategies, improving practicality and applicability.
[0050] In one or more optional embodiments of the present invention, the correctness evaluation of the target model based on the inference services of the target model in the graphics card environment and the target chip environment respectively includes: With the same configuration, a unified dataset is input into the target model in the graphics card environment and the target chip environment respectively for inference service, to obtain the first correctness index of the target model in the graphics card environment and the second correctness index of the target model in the target chip environment; Based on the first correctness indicator and the second correctness indicator, calculate the difference in correctness indicators; If the difference in the correctness index is less than or equal to the difference threshold, the correctness evaluation result is determined to be passed; If the difference in the correctness index is greater than the difference threshold, the correctness evaluation result is determined to be unsuccessful.
[0051] In practical applications, the system automatically selects a unified dataset or a specified validation set and initiates inference requests to two sets of inference services—NVIDIA service (the inference service provided by the target model in the NVIDIA environment) and target chip service (the inference service provided by the target model in the target chip environment). The system supports different types of evaluation methods, such as language models and multimodal models (i.e., the dataset or specified validation set contains data with at least one modality of information, such as text, images, or speech). For each dataset, the system calculates a final correctness metric, namely the first correctness metric and the second correctness metric.
[0052] Both the first and second correctness metrics can be correctness scores. The dataset or specified validation set can carry a standard answer. Based on the standard answer and the service output results in various environments, the first and second correctness metrics can be obtained.
[0053] After the evaluation is completed, the system compares the first and second correctness metrics at the dataset level to obtain the difference in correctness metrics. The system presets a pass / fail threshold, i.e., a difference threshold. If the difference in correctness metrics is less than or equal to the difference threshold, it means that the output of the transferred target model is acceptable, i.e., the correctness evaluation result is determined to be passed; if the difference in correctness metrics is greater than the difference threshold, it means that the output of the transferred target model is unacceptable, and the correctness evaluation result is determined to be failed.
[0054] In this embodiment of the invention, by standardizing and automating the correctness evaluation, and by using a unified input and automatic alignment mechanism and comparing it with NVIDIA output, a reproducible and quantifiable correctness judgment is obtained, reducing human error and repetitive work.
[0055] In one or more optional embodiments of the present invention, calculating the difference in correctness indicators based on the first correctness indicator and the second correctness indicator includes: Calculate the absolute value of the difference between the first correctness index and the second correctness index; The ratio of the absolute value to the first correctness index is determined as the difference in the correctness index.
[0056] In practical applications, the evaluation results corresponding to NVIDIA (the first correctness indicator) are recorded as nvidia, and the evaluation results corresponding to the target chip (the second correctness indicator) are recorded as chip.
[0057] Then, the system calculates the absolute difference between the first correctness indicator (nvidia) and the second correctness indicator (chip), i.e., the absolute value of the difference A = |nvidia - chip|. Next, it calculates the percentage difference, i.e., the ratio of the absolute value A to the first correctness indicator (nvidia). A%=A / nvidia. when If A% ≤ the difference threshold (e.g., 5%), the dataset is considered to have passed the "correctness evaluation". If A% > the difference threshold, the subsequent performance evaluation process and automatic release process will be terminated.
[0058] In this embodiment of the invention, by calculating the differences in correctness indicators in the standardized correctness evaluation, a reproducible and quantifiable correctness determination is achieved, reducing human error and repetitive work.
[0059] In one or more optional embodiments of the present invention, after performing a correctness evaluation on the target model based on the inference services of the target model in the graphics card environment and the target chip environment respectively, and obtaining the correctness evaluation result, the method further includes: If the correctness evaluation result is unsuccessful, the reason for the failure is recorded, and the reason for failure is used to repair the target chip.
[0060] In practical applications, if the correctness evaluation result is "failed", such as If A% > the difference threshold, the reason for failure (reason for not passing) can be recorded and the subsequent performance evaluation process and automatic release process can be stopped, waiting for the manufacturer to fix the chip problem.
[0061] Therefore, if the correctness evaluation fails, it indicates that there is a problem with the target chip. Recording the reason for the failure helps the manufacturer to locate the fault and improve the repair efficiency. In addition, stopping the subsequent process when the correctness evaluation fails avoids invalid evaluation and release, which not only reduces the data processing volume of the system, but also ensures the reliability of the released model.
[0062] In one or more optional embodiments of the present invention, the step of evaluating the performance of the target model based on the inference performance evaluation tool to obtain the performance evaluation result of the target model includes: Based on the aforementioned inference performance evaluation tool, a request load is generated; Based on the requested load, inference load is initiated on the target model in the graphics card environment and the target chip environment respectively, and performance indicators are collected to obtain the performance evaluation results.
[0063] Specifically, the request load includes at least one of input length, output length, and concurrency. The performance metrics include at least one of throughput, latency, token generation rate, and resource usage data.
[0064] In practical applications, vLLM is used as a unified inference performance evaluation tool. By simulating real large-scale model inference load, it collects metrics such as throughput, latency, token generation rate, and resource consumption data to measure the service performance after adaptation to the target chip.
[0065] The system applies high load to the inference service by setting multiple test parameters, including input (Token) length, output (Token) length, and number of concurrent requests (concurrency).
[0066] Specifically, performance metrics include at least one of the following: (Token) throughput, Time To First Token (TTFT), End-to-End Latency, Token generation rate, and resource usage data.
[0067] Among them, (Token) throughput represents the number of output tokens generated per unit time; first token latency represents the average time taken from sending each request to outputting the first token; overall request latency represents the time required to complete the entire request.
[0068] The process of determining the first token delay can be as follows: for each sending request, determine the time taken from the sending request to the output of the first token, and then take the average of the overall statistics to obtain the first token delay.
[0069] In this embodiment of the invention, standardized performance data can be generated and stored in a database through automated evaluation based on vLLM, forming a long-term queryable chip performance database, which facilitates performance comparison and regression analysis by manufacturers, and achieves consistency and retention of performance evaluation data.
[0070] In one or more optional embodiments of the present invention, after evaluating the performance of the target model based on the inference performance evaluation tool and obtaining the performance evaluation result of the target model, the method further includes: The correctness evaluation results and / or the performance evaluation results are associated with the target migration task to form a traceable and complete migration link record, wherein the target migration task represents the migration of the target model from the graphics card environment to the target chip environment.
[0071] In practical applications, after the correctness evaluation is completed, the system can upload and store the correctness evaluation results (such as a CorrectnessReport) and associate them with the target migration task, MigrationTask. Similarly, after the performance evaluation is completed, the system can upload and store the performance evaluation results (such as a PerformanceReport) and associate them with the target migration task, MigrationTask. Alternatively, the system can associate the CorrectnessReport and / or PerformanceReport with the MigrationTask after the performance evaluation is completed. This data storage approach creates a complete and traceable migration chain record.
[0072] In one or more optional embodiments of the present invention, after evaluating the performance of the target model based on the inference performance evaluation tool and obtaining the performance evaluation result of the target model, the method further includes: Generate a model weight package that adapts the target model to the target chip environment; The target chip environment can be constructed to directly deploy the runtime image of the target model; Generate documentation to adapt the target chip environment to the target chip environment; The publishing tool is invoked to synchronously upload the model weight package, the running image, and the documentation to the model publishing platform in order to publish the target model adapted to the target chip environment.
[0073] Specifically, the runtime image includes at least one of a runtime environment, operator library, dependent environment, and startup script. The documentation can be a README document. The documentation includes at least one of the following: chip support information, correctness test results, image download instructions, and service deployment instructions.
[0074] In practical applications, once the correctness assessment is passed and the performance test is completed, the system enters the automatic deployment phase.
[0075] Generate model weight packages adapted to the target model, such as weight.bin, config.json, and tokenizer.model. The format of the model weight packages must be supported by the target chip.
[0076] Build a runtime image that can be deployed directly, including runtime, operator library, dependency environment, startup script, etc.
[0077] Automatically generate a README document, which includes: chip support information, correctness test results, image download, and service deployment instructions.
[0078] After generating the model weight package, runtime image, and README document, the system calls the publishing tool to synchronously upload the model weight package, runtime image, and README document to the model publishing platform, achieving one-click publishing.
[0079] Meanwhile, the system records the migration process, correctness evaluation process, performance evaluation process, and release process in the background to ensure that the migration path and adaptation results can be audited and reproduced.
[0080] This invention, through an end-to-end automated process, covers correctness verification, performance evaluation, result storage, and automatic publishing after model migration and deployment, significantly improving the efficiency of domestic chip ecosystem adaptation, reducing labor costs, and shortening the delivery cycle. Furthermore, all test configurations, script versions, hardware environments, and original logs are structured and stored, meeting acceptance, auditing, and backtracking requirements, achieving the goals of traceability and reproducibility.
[0081] The following is combined with Figure 2 The model evaluation method provided by this invention will be further explained.
[0082] First, the model evaluation system consists of the following core units: The correctness evaluation unit is used to unify the input calls to the inference interface between NVIDIA and the target chip and collect the output, and perform automatic alignment and difference calculation; The performance evaluation unit is used to generate request load and collect performance metrics based on the vLLM benchmark suite. The data storage unit is used to store structured evaluation records and raw logs based on a relational database (such as MySQL) and object storage (for model weight packages and logs).
[0083] The release unit is used to execute the model packaging, image generation, and upload process (Moda / Hugging Face upload adapter) when the correctness is passed. Visualization units are used to display performance comparison curves, historical records, etc.
[0084] See Figure 2 , Figure 2 This is the second flowchart of the model evaluation method provided by the present invention: The system first completes the model migration based on FlagOS; then, it performs correctness evaluation in the NVIDIA environment and in the target chip environment; the correctness evaluation results are compared to determine whether the difference is within the threshold; if not, the manufacturer modifies the problem and then performs correctness evaluation again; if so, performance evaluation, data archiving and model release are performed.
[0085] The model evaluation method provided in this invention obtains reproducible and quantifiable correctness judgments by using a unified input and automatic alignment mechanism and comparing with NVIDIA output, reducing human error and repetitive work. The vLLM-based automated evaluation can generate standardized performance data and store it in a database, forming a long-term queryable chip performance database, facilitating performance comparison and regression analysis by manufacturers. It covers correctness verification, performance evaluation, result storage, and automatic release after model migration and deployment, significantly improving the efficiency of domestic chip ecosystem adaptation, reducing labor costs, and shortening the delivery cycle. All test configurations, script versions, hardware environments, and original logs are structured and stored to meet acceptance, auditing, and backtracking requirements. The systematic data structure and modular design support expansion and customization for different chip manufacturers, different models, and different testing strategies.
[0086] The model evaluation device provided by the present invention is described below. The model evaluation device described below can be referred to in correspondence with the model evaluation method described above.
[0087] Figure 3 This is a schematic diagram of the model evaluation device provided by the present invention, as shown below. Figure 3 As shown, the model evaluation device includes: Migration module 301 is configured to migrate the target model from the graphics card environment to the target chip environment based on model management and adaptation tools. The correctness evaluation module 302 is configured to evaluate the correctness of the target model based on the inference services of the target model in the graphics card environment and the target chip environment, respectively, and obtain the correctness evaluation result. The performance evaluation module 303 is configured to perform a performance evaluation on the target model based on an inference performance evaluation tool, provided that the correctness evaluation result is passed, and to obtain the performance evaluation result of the target model.
[0088] The model evaluation device provided by this invention, based on the correctness evaluation mechanism of the graphics card environment and the target chip environment, realizes the standardization and automation of correctness evaluation, obtains reproducible and quantifiable correctness judgments, and reduces human error and repetitive work. Through the automated evaluation mechanism of the inference performance evaluation tool, standardized performance data can be generated to form a long-term queryable chip performance database, realizing the consistency and storability of performance evaluation. The systematic correctness evaluation mechanism and automated evaluation mechanism can support the expansion and customization of different chip manufacturers, different models and different testing strategies, improving practicality and applicability.
[0089] In one or more optional embodiments of the present invention, the correctness evaluation module 302 is specifically configured as follows: With the same configuration, a unified dataset is input into the target model in the graphics card environment and the target chip environment respectively for inference service, to obtain the first correctness index of the target model in the graphics card environment and the second correctness index of the target model in the target chip environment; Based on the first correctness indicator and the second correctness indicator, calculate the difference in correctness indicators; If the difference in the correctness index is less than or equal to the difference threshold, the correctness evaluation result is determined to be passed; If the difference in the correctness index is greater than the difference threshold, the correctness evaluation result is determined to be unsuccessful.
[0090] In one or more optional embodiments of the present invention, the correctness evaluation module 302 is specifically configured as follows: Calculate the absolute value of the difference between the first correctness index and the second correctness index; The ratio of the absolute value to the first correctness index is determined as the difference in the correctness index.
[0091] In one or more optional embodiments of the present invention, the performance evaluation module 303 is specifically configured as follows: Based on the aforementioned inference performance evaluation tool, a request load is generated; Based on the requested load, inference load is initiated on the target model in the graphics card environment and the target chip environment respectively, and performance indicators are collected to obtain the performance evaluation results.
[0092] In one or more optional embodiments of the present invention, the device further includes an association module configured to: The correctness evaluation results and / or the performance evaluation results are associated with the target migration task to form a traceable and complete migration link record, wherein the target migration task represents the migration of the target model from the graphics card environment to the target chip environment.
[0093] In one or more optional embodiments of the invention, the apparatus further includes a publishing module configured to: Generate a model weight package that adapts the target model to the target chip environment; The target chip environment can be constructed to directly deploy the runtime image of the target model; Generate documentation to adapt the target chip environment to the target chip environment; The publishing tool is invoked to synchronously upload the model weight package, the running image, and the documentation to the model publishing platform in order to publish the target model adapted to the target chip environment.
[0094] In one or more optional embodiments of the invention, the apparatus further includes a recording module configured to: If the correctness evaluation result is unsuccessful, the reason for the failure is recorded, and the reason for failure is used to repair the target chip.
[0095] Figure 4 This is a schematic diagram of the structure of the electronic device provided by the present invention, such as... Figure 4 As shown, the electronic device may include a processor 410, a communications interface 420, a memory 430, and a communication bus 440, wherein the processor 410, communications interface 420, and memory 430 communicate with each other via the communication bus 440. The processor 410 can call logical instructions in the memory 430 to execute a model evaluation method. This method includes: migrating a target model from a graphics card environment to a target chip environment based on a model management and adaptation tool; performing a correctness evaluation on the target model based on the inference services of the target model in both the graphics card environment and the target chip environment, obtaining a correctness evaluation result; and, if the correctness evaluation result is passed, performing a performance evaluation on the target model based on an inference performance evaluation tool, obtaining a performance evaluation result for the target model.
[0096] Furthermore, the logical instructions in the aforementioned memory 430 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0097] On the other hand, the present invention also provides a computer program product, which includes a computer program that can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the model evaluation method provided by the above methods. The method includes: migrating a target model from a graphics card environment to a target chip environment based on a model management and adaptation tool; performing a correctness evaluation on the target model based on the inference services of the target model in the graphics card environment and the target chip environment respectively, and obtaining a correctness evaluation result; and, if the correctness evaluation result is passed, performing a performance evaluation on the target model based on an inference performance evaluation tool, and obtaining a performance evaluation result of the target model.
[0098] In another aspect, the present invention also provides a non-transitory computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the model evaluation method provided by the above methods. This method includes: migrating a target model from a graphics card environment to a target chip environment based on a model management and adaptation tool; performing a correctness evaluation on the target model based on inference services of the target model in both the graphics card environment and the target chip environment, to obtain a correctness evaluation result; and, if the correctness evaluation result is passed, performing a performance evaluation on the target model based on an inference performance evaluation tool, to obtain a performance evaluation result for the target model.
[0099] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.
[0100] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.
[0101] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A model evaluation method, characterized in that, include: Based on model management and adaptation tools, the target model is migrated from the graphics card environment to the target chip environment; Based on the inference services of the target model in the graphics card environment and the target chip environment respectively, the correctness of the target model is evaluated, and the correctness evaluation results are obtained. If the correctness evaluation result is passed, the target model is evaluated based on the inference performance evaluation tool to obtain the performance evaluation result of the target model.
2. The model evaluation method according to claim 1, characterized in that, The correctness evaluation of the target model based on the inference services of the target model in the graphics card environment and the target chip environment respectively includes: With the same configuration, a unified dataset is input into the target model in the graphics card environment and the target chip environment respectively for inference service, to obtain the first correctness index of the target model in the graphics card environment and the second correctness index of the target model in the target chip environment; Based on the first correctness indicator and the second correctness indicator, calculate the difference in correctness indicators; If the difference in the correctness index is less than or equal to the difference threshold, the correctness evaluation result is determined to be passed; If the difference in the correctness index is greater than the difference threshold, the correctness evaluation result is determined to be unsuccessful.
3. The model evaluation method according to claim 2, characterized in that, The step of calculating the difference in correctness indicators based on the first correctness indicator and the second correctness indicator includes: Calculate the absolute value of the difference between the first correctness index and the second correctness index; The ratio of the absolute value to the first correctness index is determined as the difference in the correctness index.
4. The model evaluation method according to claim 1, characterized in that, The inference performance evaluation tool is used to evaluate the performance of the target model, and the performance evaluation results of the target model are obtained, including: Based on the aforementioned inference performance evaluation tool, a request load is generated; Based on the requested load, inference load is initiated on the target model in the graphics card environment and the target chip environment respectively, and performance indicators are collected to obtain the performance evaluation results.
5. The model evaluation method according to any one of claims 1-4, characterized in that, After evaluating the performance of the target model using the inference performance evaluation tool and obtaining the performance evaluation result of the target model, the method further includes: The correctness evaluation results and / or the performance evaluation results are associated with the target migration task to form a traceable and complete migration link record, wherein the target migration task represents the migration of the target model from the graphics card environment to the target chip environment.
6. The model evaluation method according to any one of claims 1-4, characterized in that, After evaluating the performance of the target model using the inference performance evaluation tool and obtaining the performance evaluation result of the target model, the method further includes: Generate a model weight package that adapts the target model to the target chip environment; The target chip environment can be constructed to directly deploy the runtime image of the target model; Generate documentation to adapt the target chip environment to the target chip environment; The publishing tool is invoked to synchronously upload the model weight package, the running image, and the documentation to the model publishing platform in order to publish the target model adapted to the target chip environment.
7. The model evaluation method according to any one of claims 1-4, characterized in that, After evaluating the correctness of the target model based on the inference services of the target model in the graphics card environment and the target chip environment respectively, and obtaining the correctness evaluation results, the method further includes: If the correctness evaluation result is unsuccessful, the reason for the failure is recorded, and the reason for failure is used to repair the target chip.
8. A model evaluation device, characterized in that, include: The migration module is configured to migrate the target model from the graphics card environment to the target chip environment based on model management and adaptation tools. The correctness evaluation module is configured to evaluate the correctness of the target model based on the inference services of the target model in the graphics card environment and the target chip environment, respectively, and obtain the correctness evaluation result. The performance evaluation module is configured to perform performance evaluation on the target model based on an inference performance evaluation tool, provided that the correctness evaluation result is passed, and obtain the performance evaluation result of the target model.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that, When the processor executes the computer program, it implements the model evaluation method as described in any one of claims 1 to 7.
10. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the model evaluation method as described in any one of claims 1 to 7.