A model training benchmark evaluation method, device, program product and medium
By configuring benchmark evaluation options and error verification instructions in the training startup script of the model fine-tuning tool, using loop statements to start multi-scene evaluation in one click, and automatically comparing errors, the problems of cumbersome evaluation steps and unintuitive results in the existing technology are solved, and efficient and intuitive model training benchmark evaluation is achieved.
Patent Information
- Application Number
- CN202411931726.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-26
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2044-12-26
AI Technical Summary
The prior art has cumbersome steps in model training benchmark evaluation, low testing efficiency, and unintuitive results, resulting in difficulty in judgment.
By modifying the training startup script in the model fine-tuning tool, configure benchmark evaluation options for implementing different evaluation scenarios, and add error verification instructions to the script, use loop statements to start hundreds of evaluation scenarios in one click, and automatically compare training errors with preset benchmark errors.
It effectively improves the benchmark evaluation efficiency of the artificial intelligence computing power execution framework, realizes intuitive presentation of evaluation results, simplifies the evaluation process, and improves the accuracy and efficiency of evaluation.
Smart Images

Figure CN119357025B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer technology, and in particular to a model training benchmark evaluation method, device, program product and medium. Background Art
[0002] With the rapid development of artificial intelligence models, the number of model parameters of tens of billions, hundreds of billions, or even trillions, coupled with the high cost and shortage of high-performance chip resources, has led to an explosive growth in training costs. In order to solve the problem of scarce model training resources and high costs, some of the currently used artificial intelligence acceleration chips have some pain points, such as incomplete software architecture, many performance bottlenecks, and inconsistent development progress of different manufacturers. The use of an artificial intelligence computing power execution framework decouples the upper-level software required for model training from the underlying chip interface, thereby reducing the difficulty of relevant chip manufacturers in adapting model training.
[0003] The framework is designed to connect to a variety of low-cost chips for users to choose from. After each chip is connected to the AI computing execution framework, it is necessary to eliminate errors that may be introduced due to the immaturity of the chip software stack or the incomplete connection with the AI computing execution framework, so as to ensure the correctness of the training function and verify the efficiency of the chip training. Therefore, each time a new accelerator card is connected to the AI computing execution framework, extensive model training benchmark evaluation is required. Summary of the invention
[0004] In view of this, the purpose of the present invention is to provide a model training benchmark evaluation method, device, program product and medium, which can improve the benchmark evaluation efficiency of the artificial intelligence computing power execution framework. The specific scheme is as follows:
[0005] In a first aspect, the present application discloses a model training benchmark evaluation method, comprising:
[0006] Determine the model fine-tuning tool and obtain the training startup script preset in the model fine-tuning tool;
[0007] Configure the training startup script with benchmark options for different evaluation scenarios, and add error verification instructions in the training startup script to automatically process the training error of the model after training in each evaluation scenario;
[0008] Based on the benchmark evaluation options in the training startup script, a loop statement is used to start the model training task on the target hardware device. The error verification instruction added in the training startup script is used to compare the training error of the model training task with the preset benchmark error to generate the benchmark evaluation task results for each evaluation scenario.
[0009] Optionally, configure the training startup script with benchmark options to implement different benchmark scenarios, including:
[0010] Configure the benchmarking options of any one or a combination of fine-tuning strategy options, fine-tuning type options, model type options, data type options, and optimization strategy options for the training startup script.
[0011] Optionally, based on the benchmark option in the training startup script, after starting the model training task on the target hardware device using a loop statement, it also includes:
[0012] Record the target training error obtained in each step of the model training task and the number of samples used in the model training task per unit time in text form to generate a first record log;
[0013] Accordingly, the error verification instructions added in the training startup script are used to compare the training error of the model training task with the preset benchmark error, including:
[0014] Trigger the error verification instruction added in the training startup script and obtain the second record log generated when the model training task is trained on the benchmark hardware device;
[0015] The target training error in the first record log is compared with the preset benchmark error in the second record log.
[0016] Optionally, the model training benchmark evaluation method of the present application further includes:
[0017] Determine a first performance corresponding to the model on the target hardware device according to the number of samples used for training per unit time in the model training task in the first record log;
[0018] The first performance is compared with the corresponding second performance when the model is trained on other supported hardware devices, and the comparison results are displayed in a visual interface.
[0019] Optionally, the training error of the model training task is compared with the preset benchmark error to generate benchmark evaluation task results for each evaluation scenario, including:
[0020] determining a target programming language for writing error verification instructions;
[0021] Based on the text log processing function of the target programming language, respectively obtain a first record log generated when the model training task is trained on the target hardware device, and a second record log generated when the model training task is trained on the reference hardware device;
[0022] The target training error in the first record log is compared with the preset benchmark error in the second record log, and the comparison and judgment results are displayed in the form of images on a visual interface using a drawing tool implemented based on the target programming language to generate benchmark evaluation task results for each evaluation scenario.
[0023] Optionally, the training error of the model training task is compared with the preset benchmark error to generate benchmark evaluation task results for each evaluation scenario, including:
[0024] Determine the absolute value of the difference between the training error of the model training task and the preset benchmark error to obtain a reference value, and determine whether the reference value is within the preset error range;
[0025] When the reference value is within the preset error range, the first benchmark evaluation task results for demonstrating the correctness of the training in each evaluation scenario are generated on the visual interface;
[0026] When the reference value is not within the preset error range, a second benchmark evaluation task result is generated in the visualization interface for displaying a training error comparison chart in each evaluation scenario.
[0027] Optionally, the model training benchmark evaluation method of the present application further includes:
[0028] A visual interface is built through a web interface creation tool, and the benchmark evaluation options are displayed in the form of a text box on the visual interface through a text box component in the web interface creation tool.
[0029] Optionally, the model training benchmark evaluation method of the present application further includes:
[0030] Through the button component in the web interface creation tool, a training start button is created in the visual interface, and a trigger event is assigned to the training start button to start the model training task based on the trigger event.
[0031] Optionally, assign a trigger event to the training start button to start the model training task based on the trigger event, including:
[0032] Use the listening method for processing button click events in the web interface creation tool to bind the training start button to the training run function;
[0033] When the training start button is clicked, the training run function is triggered to start the model training task.
[0034] Optionally, trigger the training run function to start the model training task, including:
[0035] In the training run function, the subprocess startup function is used to create interactive instructions that are executed through the Shell script, and the interactive instructions are used to trigger the steps of starting the model training task on the target hardware device using a loop statement based on the benchmark evaluation options in the training startup script.
[0036] Optionally, based on the benchmark option in the training startup script, use a loop statement to start the model training task on the target hardware device, including:
[0037] Enter the parameters to be evaluated in the text box of the benchmark evaluation option through the visual interface;
[0038] When the training start button is clicked, the model training task is started on the target hardware device using a loop statement based on the parameters to be evaluated.
[0039] Optionally, the model training benchmark evaluation method of the present application further includes:
[0040] Provide the target user with the network address for model training benchmark evaluation so that the target user can enter the network address through the browser to start the model training task.
[0041] In a second aspect, the present application discloses an electronic device, comprising:
[0042] Memory for storing computer programs;
[0043] A processor is used to load and execute a computer program to implement the aforementioned model training benchmark evaluation method.
[0044] In a third aspect, the present application discloses a computer program product, including a computer program / instructions, which implement the steps of the aforementioned model training benchmark evaluation method when executed by a processor.
[0045] In a fourth aspect, the present application discloses a computer-readable storage medium for storing a computer program; wherein the computer program implements the aforementioned model training benchmark evaluation method when executed by a processor.
[0046] The present application provides a model training benchmark evaluation method, including: determining a model fine-tuning tool and obtaining a training startup script preset in the model fine-tuning tool; configuring the training startup script with benchmark evaluation options for implementing different evaluation scenarios, and adding error verification instructions in the training startup script for automatically processing the training error of the model after training in each evaluation scenario; based on the benchmark evaluation options in the training startup script, using a loop statement to start the model training task on the target hardware device, and using the error verification instructions added in the training startup script to compare and judge the training error of the model training task with the preset benchmark error, so as to generate the benchmark evaluation task results in each evaluation scenario.
[0047] The beneficial technical effects of this application are: by modifying the training startup script in the model fine-tuning tool, configuring the benchmark evaluation options for implementing different evaluation scenarios, and using loop statements to achieve the effect of starting hundreds of evaluation scenarios with one click; combined with the error verification instructions added in the training startup script, the training error of the model training task is automatically compared and judged with the preset benchmark error, and the benchmark evaluation task results under each evaluation scenario are generated, so that the evaluation results can be presented intuitively. In this way, the benchmark evaluation efficiency of the artificial intelligence computing power execution framework can be effectively improved.
[0048] In addition, the present application provides a model training benchmark evaluation device, equipment and storage medium, which correspond to the above-mentioned model training benchmark evaluation method and have the same effect as above. BRIEF DESCRIPTION OF THE DRAWINGS
[0049] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without paying creative work.
[0050] Figure 1 A flow chart of a model training benchmark evaluation method disclosed in this application;
[0051] Figure 2 This is a benchmark evaluation training error comparison effect diagram disclosed in this application;
[0052] Figure 3 A visualization diagram of a model training benchmark evaluation disclosed in this application;
[0053] Figure 4 A schematic diagram of the structure of a model training benchmark evaluation device disclosed in this application;
[0054] Figure 5 This is a structural diagram of an electronic device disclosed in this application. DETAILED DESCRIPTION
[0055] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0056] Pre-trained models refer to models that contain tens of billions or even more parameters. These models are trained on a large amount of text data to obtain general models. General models are powerful, but may not perform well in specific fields. In order to make the model better adapt to and complete tasks in specific fields, it is usually possible to use data sets in specific fields to fine-tune the trained general models to give the general models more customized functions.
[0057] At present, in order to solve the problem of resource shortage and high cost of chip adaptation model during model training, attention is focused on some low-cost artificial intelligence acceleration chips. Compared with mature but expensive artificial intelligence acceleration cards, these chips have many advantages such as ample quantity, no restrictions and low technical barriers. However, due to the relatively short rise of these chips, their shortcomings are also obvious: the software architecture is not perfect, there are many performance bottlenecks, and the development progress of different manufacturers is inconsistent, resulting in the adaptation of model training on these chips is not very smooth. In order to solve some pain points of such chip adaptation models, an artificial intelligence computing power execution framework is proposed. By proposing a universal chip access standard and platform, the framework decouples the upper-level software required for model training, such as PyTorch and DeepSpeed, from the underlying chip interface, which can greatly reduce the difficulty of chip manufacturers in adapting model training.
[0058] As more and more chips are connected to the AI computing power execution framework, it is essential to evaluate the correctness and efficiency of model training for these chips in a large number of different scenarios. Currently, training benchmarks are usually performed based on the open source model fine-tuning tool platform (Llama-Factory). Llama-Factory calls the interfaces of software such as PyTorch and DeepSpeed provided by the AI computing power execution framework to complete the deployment of model training tasks on the chip. There are usually three steps:
[0059] (1) Start training on the chip that needs to be benchmarked. The script is as follows:
[0060] deepspeed, / LLaMA-Factory / train.py \
[0061] --deepspeed, / LLaMA-Factory / examples / deepspeed / ds_z2_config.json\
[0062] The optimization strategy is z2
[0063] --stage sft \ Fine-tune the strategy to sft
[0064] --do_train \
[0065] --finetuning_type lora \ Fine tuning type is lora
[0066] --template llama2 \ Model type is llama2
[0067] --fp16 Data type is half precision.
[0068] The above script uses the DeepSpeed ZeRO-2 optimization strategy to perform supervised fine-tuning (SFT) on the Llama2 model. After the training is completed, a log is obtained, including the error (loss) of each step of training and the number of samples trained per second (train_samples_per_second).
[0069] (2) Verify whether the training error in step (1) is calculated correctly: Take the model training error on an AI accelerator card GPU (Graphics Processing Unit) that has been widely used in the model training field and has guaranteed training correctness in the same training scenario as a benchmark to determine whether the training error obtained by the chip connected to the AI computing power execution framework in (1) is accurate. If the training errors of the two models are consistent, the training is considered to be correct.
[0070] (3) Verify whether the training performance in step (1) is superior: The number of samples trained per second can be used as one of the indicators to measure the training performance. After the training is completed, the LlamaFactory log tool will record this indicator.
[0071] The above technology has the following defects:
[0072] (1) The steps are complicated and the test efficiency is low: When verifying whether a chip model is trained correctly, it is often necessary to conduct comprehensive verification of multiple scenarios, such as selecting eight to ten model types, two to three fine-tuning strategies, two to three optimization strategies, two data types, and two fine-tuning types. This will result in the need to manually modify and start the training script hundreds of times, which is cumbersome and prone to omissions and errors. If there are multiple chips to be verified, the workload will be even greater.
[0073] (2) Results are not intuitive: As described in (1), the benchmark evaluation scenarios may reach hundreds of times, and the final training results are presented in hundreds of reports. Each report contains dozens or hundreds of steps of training errors, which need to be manually compared with the training errors on the benchmark chip GPU, which is very unintuitive and easy to lead to misjudgment.
[0074] To this end, this application provides a model training benchmark evaluation solution that can solve the pain points of the current model training benchmark, such as cumbersome steps and non-intuitive results, and improve the benchmark evaluation efficiency of the artificial intelligence computing power execution framework.
[0075] The embodiment of the present invention discloses a model training benchmark evaluation method, which performs benchmark evaluation on models trained on large-scale data sets. These models can be applied in natural language processing (such as realizing automatic translation between multiple languages and improving the efficiency of cross-language communication), computer vision (such as used for face recognition, license plate recognition, medical image analysis, etc., to improve safety and diagnostic accuracy), and other scenarios. For details, see Figure 1 As shown, the method includes:
[0076] Step S11: determine the model fine-tuning tool, and obtain the training startup script preset in the model fine-tuning tool.
[0077] In the embodiment of the present application, a model fine-tuning tool is used to perform a benchmark test of model training. Among them, the open source tool Llama-Factory is used as an example for explanation. The bottom layer of Llama-Factory completes the deployment of model training tasks on the chip by calling the software interface provided by the artificial intelligence computing power execution framework. After determining the model fine-tuning tool used, obtain the training startup script preset in the tool. Model training can be started through the training startup script.
[0078] Step S12: configuring the training startup script with benchmark evaluation options for implementing different evaluation scenarios, and adding error verification instructions in the training startup script for automatically processing the training error of the model after training in each evaluation scenario.
[0079] In an embodiment of the present application, the training startup script is optimized, and the one-click start of the benchmark evaluation task of model training in multiple scenarios is realized by looping the execution of the training command. Specifically, before the training starts, several benchmark evaluation options for realizing different evaluation scenarios are added to the training startup script. In a specific implementation, the training startup script is configured with any one or several combinations of benchmark evaluation options among fine-tuning strategy options, fine-tuning type options, model type options, data type options, and optimization strategy options. In this way, different benchmark evaluation options are combined and cross-executed with each other, which can generate hundreds of scenarios required for benchmark evaluation.
[0080] After the current benchmark evaluation, it is often necessary to manually compare the training error of the model on the benchmark hardware device with each training result report, which is very unintuitive. Therefore, in the embodiment of the present application, an error verification instruction is added to the training startup script. When the instruction is executed, the training error of the model after training in each evaluation scenario can be automatically processed.
[0081] Step S13: Based on the benchmark evaluation option in the training startup script, use a loop statement to start the model training task on the target hardware device, and use the error verification instruction added in the training startup script to compare and judge the training error of the model training task with the preset benchmark error to generate the benchmark evaluation task results under each evaluation scenario.
[0082] In the embodiment of the present application, for the configured benchmark evaluation options, a loop statement is used to start the model training task on the target hardware device to achieve cross-generation of hundreds of scenarios required for the benchmark evaluation. Among them, the target hardware device is used to process artificial intelligence computing tasks, usually a low-cost and abundant artificial intelligence acceleration chip.
[0083] For example, the open source tool Llama-Factory is used as an example to explain that after optimizing and modifying the training startup script of Llama-Factory, it is modified to the following form:
[0084] Fine-tune strategy option = 'sft pt'
[0085] Fine-tune type option = 'lora full'
[0086] Model type option = 'llama2 llama3 glm4 chatglm3 yuan baichuan qwen'
[0087] Data type option = 'fp16 bf16'
[0088] Optimization strategy option = 'ds_z2 ds_z3'
[0089] Five-layer loop:
[0090] for fine-tuning strategy in $fine-tuning strategy options; do
[0091] for fine-tuning type in $fine-tuning type option; do
[0092] for model type in $model type option; do
[0093] for data type in $data type option; do
[0094] for optimization strategy in $optimization strategy option; do
[0095] Execute the training start command:
[0096] deepspeed . / LLaMA-Factory / train.py \
[0097] --deepspeed . / LLaMA-Factory / examples / deepspeed / $optimization_strategy_config.json\
[0098] --stage $fine-tune-strategy\
[0099] --do_train \
[0100] --finetuning_type $finetuning type\
[0101] --template $model type\
[0102] --$Data type
[0103] Execute the command to process the training log:
[0104] python loss_plot.py
[0105] The loop ends.
[0106] As shown in the above script, configure the fine-tuning strategy options, fine-tuning type options, model type options, data type options, and optimization strategy options in the script before training begins; use the loop statements in the script to cross-generate hundreds of scenarios required for benchmark evaluation. By starting the above script once, you can complete model training for hundreds of different scenarios, solving the problem of cumbersome steps caused by repeated modification of the training startup script. In addition, after the command statement for deepspeed to start training, the added python command is the error verification command used to process the training error after each training.
[0107] In an embodiment of the present application, by executing an error verification instruction, the training error of the model training task is compared and judged with the preset benchmark error, and the benchmark evaluation task results under each evaluation scenario are generated. Among them, the preset benchmark error is the error generated when the model is trained on a benchmark hardware device. The benchmark hardware device is a hardware device that has been widely used in the field of model training and whose training correctness is guaranteed. The model training error on the benchmark hardware device GPU in the same training scenario is used as the preset benchmark error to determine whether the training error obtained on the target hardware device connected to the artificial intelligence computing power execution framework is accurate. If it is consistent with the preset benchmark error, it is considered that the training is correct.
[0108] The present application provides a model training benchmark evaluation method, including: determining a model fine-tuning tool and obtaining a training startup script preset in the model fine-tuning tool; configuring the training startup script with benchmark evaluation options for implementing different evaluation scenarios, and adding error verification instructions in the training startup script for automatically processing the training error of the model after training in each evaluation scenario; based on the benchmark evaluation options in the training startup script, using a loop statement to start the model training task on the target hardware device, and using the error verification instructions added in the training startup script to compare and judge the training error of the model training task with the preset benchmark error, so as to generate the benchmark evaluation task results in each evaluation scenario.
[0109] The beneficial technical effects of this application are: by modifying the training startup script in the model fine-tuning tool, configuring the benchmark evaluation options for implementing different evaluation scenarios, and using loop statements to achieve the effect of starting hundreds of evaluation scenarios with one click; combined with the error verification instructions added in the training startup script, the training error of the model training task is automatically compared and judged with the preset benchmark error, and the benchmark evaluation task results under each evaluation scenario are generated, so that the evaluation results can be presented intuitively. In this way, the benchmark evaluation efficiency of the artificial intelligence computing power execution framework can be effectively improved.
[0110] In a feasible implementation, the model fine-tuning tool will record the training error in text form after the model training is completed to obtain a log, including the error (loss) of each step of training and the number of samples trained per unit time, such as the number of samples trained per second train_samples_per_second. Specifically, based on the benchmark evaluation option in the training startup script, after using the loop statement to start the model training task on the target hardware device, the following steps are also included: the target training error obtained from each step of the model training task and the number of samples used for training per unit time of the model training task are recorded in text form to generate a first record log.
[0111] In the embodiment of the present application, since the model fine-tuning tool records the training error generated when the model training task is trained on the target hardware device in the first record log in text form, when the training error of the model training task is compared with the preset benchmark error through the error verification instruction added in the training startup script, the following steps are specifically included:
[0112] Trigger the error verification instruction added in the training startup script and obtain the second record log generated when the model training task is trained on the benchmark hardware device;
[0113] The target training error in the first record log is compared with the preset benchmark error in the second record log.
[0114] Among them, the benchmark hardware device is a hardware device that has been widely used in the field of model training and whose training correctness is guaranteed. It can be seen that after the error verification instruction is triggered, the training errors recorded in the two logs will be compared to determine whether the training error obtained when the artificial intelligence computing power execution framework is connected to the target hardware device is accurate. If it is consistent with the preset benchmark error, it is considered that the training is correct.
[0115] In a specific implementation, the error verification instruction added for automatically processing the training error of the model after training in each evaluation scenario is explained by taking a python instruction as an example. Specifically, the training error of the model training task is compared with the preset benchmark error to generate the benchmark evaluation task results in each evaluation scenario, including the following steps:
[0116] determining a target programming language for writing error verification instructions;
[0117] Based on the text log processing function of the target programming language, respectively obtain a first record log generated when the model training task is trained on the target hardware device, and a second record log generated when the model training task is trained on the reference hardware device;
[0118] The target training error in the first record log is compared with the preset benchmark error in the second record log, and the comparison and judgment results are displayed in the form of images on a visual interface using a drawing tool implemented based on the target programming language to generate benchmark evaluation task results for each evaluation scenario.
[0119] In the embodiment of the present application, the target programming language for writing error verification instructions is Python, and the error verification instructions are Python instructions. The function of processing text logs in Python is used to read the record logs of the model fine-tuning tool after training in each benchmark evaluation scenario on the target hardware device and the benchmark hardware device, and the error data and performance indicators therein are captured, and each set of error data is compared and judged; further, the drawing tool matplotlib.pyplot in Python is used to display the comparison and judgment results in the form of images on a visual interface, and the benchmark evaluation task results in each evaluation scenario are generated.
[0120] Specifically, when comparing and judging each set of errors, determine the absolute value of the difference between the training error of the model training task and the preset benchmark error to obtain a reference value, and determine whether the reference value is within the preset error range; if the reference value exceeds the acceptable error range, an image is formed to intuitively show that the training result of the chip in this evaluation scenario is incorrect, and a training error comparison chart will be displayed. If the reference value is within the acceptable error range, the evaluation result of the training correctness will be displayed, and the words "passed the evaluation" will be displayed. The relevant processing code is as follows:
[0121] import matplotlib.pyplot as plt
[0122] metrics = []
[0123] with open('. / MLU_loss.log', 'r') as f:
[0124] lines = f.readlines()
[0125] for line in lines:
[0126] if line.startswith('{\'loss\':'):
[0127] metrics.append(float(line.split(",")[0].split(" ")[-1]))
[0128] plt.figure()
[0129] plt.plot(steps, metrics).
[0130] In addition, to facilitate intuitive comparison for users, after training is completed, the training performance of the current target hardware device will be displayed in the visualization interface, specifically the number of samples used for training per unit time. That is, the first performance of the model on the target hardware device. The first performance is compared with the second performance of the model when it is trained on other supported hardware devices, and the comparison results are displayed in the visualization interface. Figure 2 The figure shows the effect of the implementation, which shows the results of the benchmark evaluation of chip 1 and chip 2 on a certain model. Among them, the benchmark error (BASELINEloss) is the training result of the benchmark hardware device GPU that has been widely used in the field of model training and has guaranteed training correctness. Figure 2 It can be seen intuitively that the training result of chip 2 is correct, while the training result of chip 1 is wrong.
[0131] Based on the foregoing embodiments, the present invention improves the training startup script and utilizes loop statements to achieve the effect of starting hundreds of training scenarios with one click. It then combines the function of processing text logs in Python and the drawing tool matplotlib.pyplot to automatically read the text-based logs for training result determination. Finally, it is presented in the form of an image, which solves the pain points of cumbersome steps and non-intuitive results caused by repeated modification of the training startup script. On this basis, in order to achieve intuitive display of the visual interface, a visual interface is built through a web interface creation tool in this embodiment, wherein the web interface creation tool can be a Gradio component. Gradio is an interactive Web application Python library for quickly building and deploying machine learning models. A visual interface can be built with the help of the component layout provided by Gradio. That is, a Web application device is built based on the Gradio component to provide a visual model training benchmark evaluation service.
[0132] According to the content disclosed in the above embodiments, users need to modify various benchmark evaluation options in the training startup script, such as fine-tuning strategies, model types, etc. Then, after visualization, these options become input boxes on the service interface, which are implemented by Gradio's text box component (Textbox). Specifically, the benchmark evaluation options are displayed in the form of text boxes on the visualization interface through the text box component in the web interface creation tool.
[0133] In addition, the start of the evaluation task is based on the Gradio button component (Button). Specifically, through the button component in the web interface creation tool, a training start button is created in the visual interface, and a trigger event is assigned to the training start button so that the model training task can be started based on the trigger event.
[0134] Furthermore, the parameters to be evaluated are entered in the text box of the benchmark evaluation option through the visual interface; when the training start button is clicked, the model training task is started on the target hardware device using a loop statement according to the parameters to be evaluated. In this way, users can start benchmark evaluation tasks for hundreds of scenarios with one click, and present the evaluation results in a visual form. This can effectively improve the benchmark evaluation efficiency of the AI computing power execution framework.
[0135] In the embodiment of the present application, a training start button for starting training is created in an interactive interface through a button component. A trigger event, namely a training run function (run_train), is assigned to it through the listener click method, and the training start button is bound to the training run function. When the user clicks the start training button, the function is triggered to run.
[0136] In the embodiment of the present application, in the training run function, Python's subprocess.Popen function is used to create an instruction executed by the shell, which is the instruction to start the optimized training startup script. Specifically, in the training run function, an interactive instruction executed by the Shell script is created through the subprocess startup function, and the interactive instruction is triggered based on the benchmark evaluation option in the training startup script, and the step of starting the model training task on the target hardware device using a loop statement.
[0137] In a specific implementation, the visualization of the model training benchmark evaluation can be started through a python script, and the implementation method is as follows:
[0138] with gr.Blocks() as demo:
[0139] gr.HTML(" <h1> <center> Model training benchmark evaluation device< / center> < / h1> ")
[0140] with gr.Row():
[0141] stage = gr.Textbox(label="Fine-tuning strategy", value="sft", interactive=True)
[0142] start_btn = gr.Button(value="Training evaluation starts!")
[0143] start_btn.click(fn=run_train, inputs=args, outputs=train_samples_per_sec).
[0144] Run the above python script to start the model training benchmark evaluation. In a feasible implementation, the above implementation method is provided to the user in the form of a network address, and the user opens the device through a browser. Specifically, the network address for performing the model training benchmark evaluation is provided to the target user, so that the target user can enter the network address through the browser to start the model training task.
[0145] like Figure 3The final visualization provided in this embodiment is shown. The user enters the parameters to be tested in the text box of the benchmark evaluation options such as chip name, fine-tuning strategy, fine-tuning type, model type, data type, optimization strategy, etc. in the upper left corner, and then clicks the start button to start the benchmark evaluation task for the chip to be tested. After the training is completed, the training performance of the chip is displayed by the number of training samples per second in the lower left corner, and the performance data of the artificial intelligence chips supported by the current artificial intelligence computing power execution framework will be displayed above it, which is convenient for users to compare intuitively. Figure 3 On the right side of the page, the evaluation results of the training correctness are displayed. If the training results are incorrect, the training error comparison chart will be displayed here. If the training results are correct, the words "Evaluation Passed" will be displayed here.
[0146] In a feasible implementation, model training can be model training for a large language model (Large Language Model, LLM). LLM, as an artificial intelligence model designed to understand and generate human language, is also commonly referred to as a large model. A service call interface can be added to verify whether the large model is adapted on the target hardware device, and the interface is called to determine whether the large model is adapted to the target hardware device. In addition, after generating the benchmark evaluation task results in each evaluation scenario, each evaluation scenario can be scored to obtain a score result of the degree of adaptation to the model in each evaluation scenario. The higher the score, the better the training performance of the model on the corresponding chip. In this way, the results are displayed more intuitively. At the same time, the present invention can also be used in server performance evaluation scenarios, optimal parameter search scenarios, etc.
[0147] Correspondingly, the present application embodiment also discloses a model training benchmark evaluation device, see Figure 4 As shown, the device comprises:
[0148] The script acquisition module 11 is used to determine the model fine-tuning tool and obtain the training startup script preset in the model fine-tuning tool;
[0149] The script optimization module 12 is used to configure the benchmark evaluation options for implementing different evaluation scenarios for the training startup script, and to add error verification instructions for automatically processing the training error of the model after training in each evaluation scenario in the training startup script;
[0150] The benchmark evaluation module 13 is used to start the model training task on the target hardware device based on the benchmark evaluation option in the training startup script using a loop statement, and compare the training error of the model training task with the preset benchmark error through the error verification instruction added in the training startup script to generate the benchmark evaluation task results under each evaluation scenario.
[0151] Among them, for more specific working processes of the above-mentioned modules, please refer to the corresponding contents disclosed in the aforementioned embodiments, which will not be repeated here.
[0152] It can be seen that the above scheme of this embodiment includes: determining a model fine-tuning tool and obtaining a training startup script preset in the model fine-tuning tool; configuring a benchmark evaluation option for implementing different evaluation scenarios for the training startup script, and adding an error verification instruction in the training startup script for automatically processing the training error of the model after training in each evaluation scenario; based on the benchmark evaluation option in the training startup script, using a loop statement to start the model training task on the target hardware device, and through the error verification instruction added in the training startup script, comparing and judging the training error of the model training task with the preset benchmark error, so as to generate the benchmark evaluation task results in each evaluation scenario.
[0153] The beneficial technical effects of this application are: by modifying the training startup script in the model fine-tuning tool, configuring the benchmark evaluation options for implementing different evaluation scenarios, and using loop statements to achieve the effect of starting hundreds of evaluation scenarios with one click; combined with the error verification instructions added in the training startup script, the training error of the model training task is automatically compared and judged with the preset benchmark error, and the benchmark evaluation task results under each evaluation scenario are generated, so that the evaluation results can be presented intuitively. In this way, the benchmark evaluation efficiency of the artificial intelligence computing power execution framework can be effectively improved.
[0154] Furthermore, the present application also discloses an electronic device. Figure 5 This is a structural diagram of an electronic device 20 according to an exemplary embodiment, and the content in the diagram cannot be considered as any limitation on the scope of use of the present application.
[0155] Figure 5 A schematic diagram of the structure of an electronic device 20 provided in an embodiment of the present application. The electronic device 20 may specifically include: at least one processor 21, at least one memory 22, a power supply 23, a communication interface 24, an input / output interface 25, and a communication bus 26. The memory 22 is used to store a computer program, which is loaded and executed by the processor 21 to implement the relevant steps in the model training benchmark evaluation method disclosed in any of the aforementioned embodiments. In addition, the electronic device 20 in this embodiment may specifically be a computer.
[0156] In this embodiment, the power supply 23 is used to provide working voltage for each hardware device on the electronic device 20; the communication interface 24 can create a data transmission channel between the electronic device 20 and the external device, and the communication protocol it follows is any communication protocol that can be applied to the technical solution of the present application, and is not specifically limited here; the input and output interface 25 is used to obtain external input data or output data to the outside world, and its specific interface type can be selected according to specific application needs and is not specifically limited here.
[0157] In addition, the memory 22 as a carrier for resource storage may be a read-only memory, a random access memory, a disk or an optical disk, etc. The resources stored thereon may include an operating system 221, a computer program 222 and data 223, etc. The data 223 may include various data. The storage method may be temporary storage or permanent storage.
[0158] The operating system 221 is used to manage and control the hardware devices and computer program 222 on the electronic device 20, which can be Windows Server, Netware, Unix, Linux, etc. In addition to including a computer program that can be used to complete the model training benchmark evaluation method performed by the electronic device 20 disclosed in any of the aforementioned embodiments, the computer program 222 can further include a computer program that can be used to complete other specific tasks.
[0159] Furthermore, the embodiment of the present application also discloses a computer-readable storage medium, where the computer-readable storage medium includes a random access memory (RAM), a memory, a read-only memory (ROM), an electrically programmable ROM, an electrically erasable programmable ROM, a register, a hard disk, a magnetic disk or an optical disk, or any other form of storage medium known in the technical field. Among them, when the computer program is executed by the processor, the aforementioned model training benchmark evaluation method is implemented. For the specific steps of the method, reference can be made to the corresponding contents disclosed in the aforementioned embodiments, which will not be repeated here.
[0160] Furthermore, an embodiment of the present application also provides a computer program product, including a computer program / instruction, which, when executed by a processor, implements any one of the above-mentioned model training benchmark evaluation methods.
[0161] In this specification, each embodiment is described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the embodiments can be referred to each other. For the device disclosed in the embodiment, since it corresponds to the method disclosed in the embodiment, the description is relatively simple, and the relevant parts can be referred to the method part.
[0162] The steps of the model training benchmark evaluation method or algorithm described in conjunction with the embodiments disclosed herein can be implemented directly using hardware, a software module executed by a processor, or a combination of the two. The software module can be placed in a random access memory (RAM), a memory, a read-only memory (ROM), an electrically programmable ROM, an electrically erasable programmable ROM, a register, a hard disk, a removable disk, a CD-ROM, or any other form of storage medium known in the art.
[0163] Finally, it should be noted that, in this article, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, the elements defined by the sentence "comprise a ..." do not exclude the presence of other identical elements in the process, method, article or device including the elements.
[0164] The above is a detailed introduction to a model training benchmark evaluation method, device, program product and medium provided by the present invention. Specific examples are used in this article to illustrate the principles and implementation methods of the present invention. The description of the above embodiments is only used to help understand the method of the present invention and its core idea; at the same time, for those skilled in the art, according to the idea of the present invention, there will be changes in the specific implementation method and application scope. In summary, the content of this specification should not be understood as a limitation on the present invention.
Claims
1. A model training benchmark evaluation method, characterized in that: include: Determine a model fine-tuning tool, and obtain a training startup script preset in the model fine-tuning tool; Configuring the training startup script with benchmark evaluation options for implementing different evaluation scenarios, and adding an error verification instruction in the training startup script for automatically processing the training error of the model after training in each of the evaluation scenarios; Based on the benchmark evaluation option in the training startup script, a model training task is started on the target hardware device using a loop statement, and the training error of the model training task is compared and judged with a preset benchmark error through the error verification instruction added in the training startup script to generate the benchmark evaluation task results under each evaluation scenario; The target hardware device is a device used to process artificial intelligence computing tasks; the preset benchmark error is a parameter value used to compare with the training error to determine whether the training corresponding to the model training task is correct.
2. The model training benchmark evaluation method according to claim 1, characterized in that: The configuration of the training startup script to implement benchmark evaluation options for different evaluation scenarios includes: The training startup script is configured with any one or a combination of benchmark evaluation options of fine-tuning strategy options, fine-tuning type options, model type options, data type options, and optimization strategy options.
3. The model training benchmark evaluation method according to claim 1, characterized in that: After starting the model training task on the target hardware device using a loop statement based on the benchmark evaluation option in the training startup script, the method further includes: Recording the target training error obtained in each step of the model training task and the number of samples used in the model training task per unit time in text form to generate a first record log; Correspondingly, the error verification instruction added in the training startup script compares and judges the training error of the model training task with the preset benchmark error, including: Trigger the error verification instruction added in the training startup script, and obtain a second record log generated when the model training task is trained on the benchmark hardware device; The target training error in the first record log is compared with the preset benchmark error in the second record log.
4. The model training benchmark evaluation method according to claim 3, characterized in that: Also includes: Determine a first performance corresponding to the model on the target hardware device according to the number of samples used for training per unit time in the model training task in the first record log; The first performance is compared with the second performance corresponding to the model when it is trained on other supported hardware devices, and the comparison result is displayed in a visual interface.
5. The model training benchmark evaluation method according to claim 3, characterized in that: The comparing and judging the training error of the model training task with the preset benchmark error to generate the benchmark evaluation task results under each evaluation scenario includes: determining a target programming language for writing the error verification instructions; Based on the text log processing function of the target programming language, respectively obtain the first record log generated when the model training task is trained on the target hardware device, and the second record log generated when the model training task is trained on the reference hardware device; The target training error in the first record log is compared and judged with the preset benchmark error in the second record log, and the comparison and judgment results are displayed in the form of images on a visual interface using a drawing tool implemented based on the target programming language to generate benchmark evaluation task results under each of the evaluation scenarios.
6. The model training benchmark evaluation method according to claim 1, characterized in that: The comparing and judging the training error of the model training task with the preset benchmark error to generate the benchmark evaluation task results under each evaluation scenario includes: Determine the absolute value of the difference between the training error of the model training task and a preset reference error to obtain a reference value, and determine whether the reference value is within a preset error range; When the reference value is within the preset error range, generating a first benchmark evaluation task result for demonstrating the correctness of the training in each evaluation scenario on a visual interface; When the reference value is not within the preset error range, a second benchmark evaluation task result for displaying a training error comparison chart under each of the evaluation scenarios is generated on the visualization interface.
7. The model training benchmark evaluation method according to claim 1, characterized in that: Also includes: A visual interface is constructed by using a web page interface creation tool, and the benchmark evaluation options are displayed in the form of a text box on the visual interface by using a text box component in the web page interface creation tool.
8. The model training benchmark evaluation method according to claim 7, characterized in that: Also includes: A training start button is created in the visual interface through a button component in the web page interface creation tool, and a trigger event is assigned to the training start button so as to start the model training task based on the trigger event.
9. The model training benchmark evaluation method according to claim 8, characterized in that: The assigning a trigger event to the training start button so as to start the model training task based on the trigger event includes: Using the monitoring method for processing button click events in the web page interface creation tool, the training start button is bound to the training run function; When the training start button is clicked, the training run function is triggered to start the model training task.
10. The model training benchmark evaluation method according to claim 9, characterized in that: The triggering of the training run function to start the model training task includes: In the training run function, an interactive instruction executed by a Shell script is created through a subprocess startup function, and the interactive instruction is triggered to start the model training task on the target hardware device based on the benchmark evaluation option in the training startup script using a loop statement.
11. The model training benchmark evaluation method according to claim 8, characterized in that: Based on the benchmark evaluation option in the training startup script, a model training task is started on the target hardware device using a loop statement, including: Inputting the parameter to be evaluated in the text box of the benchmark evaluation option through the visual interface; When the training start button is clicked, a model training task is started on the target hardware device using a loop statement according to the parameters to be evaluated.
12. The model training benchmark evaluation method according to any one of claims 1 to 11, characterized in that: Also includes: Provide the target user with a network address for performing a model training benchmark evaluation so that the target user can enter the network address through a browser to start the model training task.
13. An electronic device, characterized in that: include: Memory for storing computer programs; A processor, configured to load and execute the computer program to implement the model training benchmark evaluation method according to any one of claims 1 to 12.
14. A computer program product comprising a computer program / instructions, characterized in that When the computer program / instructions are executed by a processor, the steps of the model training benchmark evaluation method described in any one of claims 1 to 12 are implemented.
15. A computer-readable storage medium, characterized in that: Used to store computer programs; wherein the computer program, when executed by a processor, implements the model training benchmark evaluation method described in any one of claims 1 to 12.
Citation Information
Patent Citations
Low-resource computing power inference method of embedded device deployment model
CN118171741A
Methods and apparatus for detecting a side channel attack using hardware performance counters
US20190130101A1