Method and device for performing asynchronous parallel fine adjustment on large model based on computing power of intelligent computing center

Through the powerful computing resources and asynchronous parallel processing method of the Intelligent Computing Center, the problem of inefficient fine-tuning of large models is solved, rapid iteration and efficient development are achieved, and model performance is significantly improved.

CN120104277APending Publication Date: 2025-06-06DATACANVAS LTD

Patent Information

Application Number
CN202510161944.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-13
Publication Date
2025-06-06

AI Technical Summary

Technical Problem

The existing large model fine-tuning methods are inefficient and difficult to meet the needs of rapid iteration and efficient development of models.

Method used

Using the asynchronous parallel fine-tuning method of computing power based on the intelligent computing center, we automatically search for the best hyperparameter combination by determining the hyperparameter space, parallel fine-tuning and evaluating the big model, and maximize the utilization of computing resources.

Benefits of technology

It significantly shortens the model tuning time, improves the efficiency of model tuning, and ensures the performance of large models in vertical fields.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120104277A_ABST
    Figure CN120104277A_ABST
Patent Text Reader

Abstract

The invention provides a computing power asynchronous parallel fine-tuning large model method and device based on an intelligent computing center, and the method comprises the steps: obtaining a to-be-processed task, and determining a training set, a test set, a base large model and an evaluation index; determining a hyper-parameter space; based on the hyper-parameter space, the training set and the computing power resource of the intelligent computing center, performing parallel fine tuning on the base large model; based on the test set, the evaluation index and the computing power resource of the intelligent computing center, performing parallel evaluation on each fine-tuned large model to obtain an evaluation result corresponding to each fine-tuned large model; and determining an optimal evaluation result, determining the fine-tuned large model corresponding to the optimal evaluation result as a target large model, and recording a hyper-parameter group corresponding to the target large model. Through powerful computing power resources of the intelligent computing center and in combination with an asynchronous parallel processing mode, large model fine tuning and evaluation tasks based on multiple hyper-parameter combinations can be processed at the same time, the model tuning time is remarkably shortened, and the model tuning efficiency is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present application relate to the technical field of computing power infrastructure, and in particular, to a method and device for asynchronous parallel fine-tuning of a large model based on the computing power of an intelligent computing center. Background Art

[0002] With the development of artificial intelligence technology and computing power technology, the concept of intelligent computing center has emerged. "Intelligent computing center" refers to the use of large-scale heterogeneous computing power resources, including general computing power and intelligent computing power, mainly for artificial intelligence applications (such as artificial intelligence deep learning model development, model training and model reasoning and other scenarios) to provide the required computing power, data and algorithms. Intelligent computing center covers facilities, hardware, software, and can provide full-stack capabilities from bottom-level computing power to top-level application enablement.

[0003] With the rapid development of artificial intelligence technology, especially in the fields of natural language processing and computer vision, the capabilities of large models, namely large language models (LLM) and multimodal large models (Multimodal LargeModels), are increasing. However, these large models still have significant room for improvement when facing special tasks in vertical fields. In order to enhance the performance of large models in specific fields, researchers usually prepare relevant text training sets or image-text pair training sets, and optimize the model through supervised fine-tuning (SFT) and asynchronous parallel fine-tuning.

[0004] In the traditional supervised fine-tuning (SFT) and asynchronous parallel fine-tuning process, researchers need to repeatedly experiment with multiple hyperparameters [such as learning rate, rank of LoRA (Low-Rank Adaptation), scaling factor of LoRA, weight decay, batch size and rounds, etc.] to determine the best hyperparameter combination. This process is not only time-consuming, but also due to the complexity of the hyperparameter combination, it may lead to waste of resources and unpredictability of model performance. In addition, traditional fine-tuning methods often use limited computing resources, resulting in inefficient model training and difficulty in meeting the needs of rapid model iteration and efficient development.

[0005] In summary, the existing large model fine-tuning methods are inefficient and cannot meet the needs of rapid model iteration and efficient development. Summary of the invention

[0006] The embodiments of the present application provide a method and device for asynchronously and parallelly fine-tuning a large model based on the computing power of an intelligent computing center, so as to solve the technical problems that the existing large model fine-tuning method is inefficient and difficult to meet the needs of rapid model iteration and efficient development.

[0007] In order to solve the above technical problems, this application is implemented as follows:

[0008] In a first aspect, an embodiment of the present application provides a method for asynchronously and parallelly fine-tuning a large model based on the computing power of an intelligent computing center, the method comprising:

[0009] Step S1: Obtain the task to be processed, and determine the corresponding training set, test set, base large model and evaluation index based on the task to be processed;

[0010] Step S2: determining a hyperparameter space, wherein the hyperparameter space includes a plurality of hyperparameter groups, each hyperparameter group includes at least one type of hyperparameter, and each type of hyperparameter has only one value in each hyperparameter group, and in different hyperparameter groups, the combination of values ​​of each type of hyperparameter is different;

[0011] Step S3: executing a model fine-tuning task, the model fine-tuning task comprising: fine-tuning the base large model in parallel based on the hyperparameter space, the training set and the computing power resources of the intelligent computing center to obtain a plurality of fine-tuned large models, wherein each fine-tuned large model corresponds to a hyperparameter group in the hyperparameter space;

[0012] Step S4: executing a model evaluation task, wherein the model evaluation task includes: based on the test set, the evaluation index, and the computing power resources of the intelligent computing center, performing parallel evaluation on each fine-tuned large model to obtain an evaluation result corresponding to each fine-tuned large model, wherein the evaluation result is the value of the evaluation index;

[0013] Step S5: Execute a model determination task, which includes: determining the best evaluation result in the evaluation results, determining the fine-tuned large model corresponding to the best evaluation result as the target large model, and recording the hyperparameter group corresponding to the target large model; wherein the best evaluation result is the evaluation result corresponding to the best value of the evaluation indicator.

[0014] Optionally, the hyperparameters include at least one of the following: learning rate, rank and round of LoRA.

[0015] Optionally, before step S3, the method further includes:

[0016] Step S6: construct a directed acyclic graph based on the Airflow tool, and the directed acyclic graph is used to define the execution order of tasks. The execution order is: first execute the model fine-tuning task, then execute the model evaluation task, and finally execute the model determination task.

[0017] Optionally, the computing resources of the intelligent computing center include at least one graphics processing unit (GPU), and only one model fine-tuning task or one model evaluation task is executed on one GPU at a time.

[0018] In a second aspect, an embodiment of the present application provides a device for asynchronously and parallelly fine-tuning a large model based on the computing power of an intelligent computing center, the device comprising:

[0019] An acquisition module is used to acquire tasks to be processed, and determine corresponding training sets, test sets, base large models, and evaluation indicators based on the tasks to be processed;

[0020] An execution module is used to determine a hyperparameter space, wherein the hyperparameter space includes multiple hyperparameter groups, each hyperparameter group includes at least one type of hyperparameter, and each type of hyperparameter has only one value in each hyperparameter group, and in different hyperparameter groups, the combination of values ​​of each type of hyperparameter is different;

[0021] Executing a model fine-tuning task, the model fine-tuning task comprising: fine-tuning the base large model in parallel based on the hyperparameter space, the training set, and the computing power resources of the intelligent computing center to obtain a plurality of fine-tuned large models, wherein each fine-tuned large model corresponds to a hyperparameter group in the hyperparameter space;

[0022] Execute a model evaluation task, the model evaluation task comprising: based on the test set, the evaluation index and the computing power resources of the intelligent computing center, perform parallel evaluation on each fine-tuned large model to obtain an evaluation result corresponding to each fine-tuned large model, wherein the evaluation result is the value of the evaluation index;

[0023] Execute a model determination task, the model determination task comprising: determining the best evaluation result in the evaluation results, determining the fine-tuned large model corresponding to the best evaluation result as the target large model, and recording the hyperparameter group corresponding to the target large model; wherein the best evaluation result is the evaluation result corresponding to the best value of the evaluation indicator.

[0024] Optionally, the hyperparameters include at least one of the following: learning rate, LoRA rank and round. Optionally, the execution module is also used to build a directed acyclic graph based on the Airflow tool before executing the model fine-tuning task, and the directed acyclic graph is used to define the order of execution of tasks, and the order of execution is: first execute the model fine-tuning task, then execute the model evaluation task, and finally execute the model determination task.

[0025] In a third aspect, an embodiment of the present application provides an electronic device, comprising: a processor, a memory, and a program stored in the memory and executable on the processor, wherein the program, when executed by the processor, implements the steps of a method for asynchronously and parallelly fine-tuning a large model based on the computing power of an intelligent computing center as described in the first aspect.

[0026] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium having a computer program stored thereon. When the computer program is executed by the processor, the steps of a method for asynchronously and parallelly fine-tuning a large model based on the computing power of an intelligent computing center as described in the first aspect are implemented.

[0027] In a fifth aspect, an embodiment of the present application provides a computer program product, comprising computer instructions, which, when executed by the processor, implement the steps of a method for asynchronously and parallelly fine-tuning a large model based on the computing power of an intelligent computing center as described in the first aspect.

[0028] The method provided in the embodiment of the present application is equipped with the powerful computing power of the intelligent computing center, which can maximize the use of computing resources and improve the resource utilization of large model training and evaluation; the intelligent computing center can realize automated hyperparameter search and evaluation, reduce the complexity of manual debugging, and ensure that the best hyperparameter combination can be found, thereby improving the performance of large models in vertical fields; by allocating fine-tuning tasks of different hyperparameter combinations to the heterogeneous computing nodes of the intelligent computing center for asynchronous and parallel execution, the fine-tuning and evaluation of all combinations can be completed at one time, which can avoid serial resource waste and shorten the tuning time.

[0029] In summary, through the powerful computing resources of the intelligent computing center, combined with asynchronous parallel processing, it is possible to simultaneously handle large model fine-tuning and evaluation tasks based on multiple hyperparameter combinations, significantly shortening the model tuning time and improving the efficiency of model tuning. BRIEF DESCRIPTION OF THE DRAWINGS

[0030] Various other advantages and benefits will become apparent to those of ordinary skill in the art by reading the detailed description of the preferred embodiments below. The accompanying drawings are only for the purpose of illustrating the preferred embodiments and are not to be considered as limiting the present application. Also, the same reference symbols are used throughout the accompanying drawings to represent the same components. In the accompanying drawings:

[0031] Figure 1 A flowchart of a method for asynchronously and parallelly fine-tuning a large model based on the computing power of an intelligent computing center provided in an embodiment of the present application;

[0032] Figure 2 A flowchart of a method for asynchronously and parallelly fine-tuning a large model based on the computing power of an intelligent computing center provided in an embodiment of the present application;

[0033] Figure 3 A flowchart of a method for asynchronously and parallelly fine-tuning a large model based on the computing power of an intelligent computing center provided in an embodiment of the present application;

[0034] Figure 4 A structural block diagram of a device for asynchronously and parallelly fine-tuning a large model based on the computing power of an intelligent computing center provided in an embodiment of the present application;

[0035] Figure 5 A schematic diagram of the structure of an electronic device according to an embodiment of the present application. DETAILED DESCRIPTION

[0036] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.

[0037] The following is a brief description of the technical terms involved in this application.

[0038] The "computing power" mentioned in the present invention is the ability of computer equipment or computing / data center to process information. It is the ability of computer hardware and software to work together to execute certain computing requirements. It is the computing power to achieve target result output by processing information data. It is a new type of productivity that integrates information computing power, network carrying capacity, and data storage capacity. It mainly provides services to the society through computing power infrastructure.

[0039] The "computing power" (Computational Power, CP) described in the present invention is the ability of a data center server to process data and output results. It is a comprehensive indicator to measure the computing power of a data center, including general computing power, super computing power and intelligent computing power. The commonly used unit of measurement is the number of floating point operations performed per second (FLOPS: Floating Point Operations Per Second, 1EFLOPS = 10^18FLOPS). The larger the value, the stronger the comprehensive computing power. According to calculations, 1EFLOPS is approximately the computing power output of 5 Tianhe-2A or 500,000 mainstream server CPUs or 2 million mainstream notebooks. The calculation formula is: CP = CP 通用 +CP 智能 +CP 超级 .

[0040] The "Network Power" (NP) described in the present invention is a manifestation of the data transmission capability of computing facilities, including comprehensive capabilities such as network architecture, network bandwidth, transmission latency, intelligent management and scheduling, etc. The carrying capacity involves network transmission within and between data centers, and is a comprehensive indicator for measuring network transmission scheduling capabilities. In the embodiment of the present invention, the carrying capacity uses the memory bandwidth.

[0041] The "Storage Power" (SP) described in the present invention is the comprehensive ability of a data center in terms of data storage capacity, performance, safety and reliability, and green and low-carbon. It is a comprehensive indicator for measuring the data storage capacity of a data center, including external storage devices such as storage arrays and server built-in storage devices. The commonly used unit of measurement for storage capacity is exabyte (EB, 1EB = 2^60bytes), and the commonly used unit of measurement for performance is the number of reads and writes per second per unit capacity (IOPS / TB, Input / Output Operations Per Second / TB). The disaster recovery ratio is an important manifestation of safety and reliability.

[0042] The "computing power infrastructure" described in the present invention is a new type of information infrastructure that integrates information computing power, network carrying capacity, and data storage capacity. It can realize centralized computing, storage, transmission and application of information, and presents characteristics such as diversity and ubiquity, intelligence and agility, security and reliability, and green and low-carbon.

[0043] The "new information infrastructure" mentioned in the present invention refers to: mainly including network infrastructure such as 5G network, fiber broadband network, backbone network, international communication network, satellite Internet, computing power infrastructure such as data center, general computing power center, intelligent computing center, supercomputing center, and new technology facilities such as artificial intelligence, blockchain, and quantum computing. With the emergence and promotion and application of new general technologies, the new information infrastructure will become more diverse.

[0044] The “computing power” mentioned in the present invention includes general computing power, intelligent computing power and super computing power.

[0045] The “general computing power” mentioned in the present invention refers to the computing power provided by servers based on CPU (Central Processing Unit) chips, which is used to support basic general computing such as cloud computing and edge computing.

[0046] The "intelligent computing power" described in the present invention is a computing platform for large-scale deployment of special chips such as GPU (Graphics Processing Unit), FPGA (Field Programmable Gate Array), ASIC (Application Specific Integrated Circuit) for various innovative applications of artificial intelligence, such as natural language processing and machine vision.

[0047] The "super computing power" mentioned in the present invention is mainly the computing power provided by high-performance computing clusters such as supercomputers. It utilizes the centralized computing resources of multiple computer systems working in parallel and uses a dedicated operating system to handle extremely complex or data-intensive problems. It is mainly used for calculations in cutting-edge scientific fields, such as planetary simulation, drug molecule design, and gene analysis.

[0048] The "intelligent computing center" described in the present invention refers to a facility that provides the required computing power, data and algorithms for artificial intelligence applications (such as artificial intelligence deep learning model development, model training and model reasoning scenarios) by using large-scale heterogeneous computing power resources, including general computing power (CPU: Central Processing Unit) and intelligent computing power (GPU: Graphics Processing Unit, FPGA: Field Programmable Gate Array, ASIC: Application Specific Integrated Circuit, etc.). The intelligent computing center covers facilities, hardware, and software, and can provide full-stack capabilities from bottom-level computing power to top-level application enablement.

[0049] The "intelligent computing center" mentioned in the present invention includes but is not limited to an intelligent computing center.

[0050] The "intelligent computing center" mentioned in the present invention is an artificial intelligence computing center, which is a type of computing power infrastructure that is based on artificial intelligence theory, adopts artificial intelligence computing architecture, and provides computing power services, data services, and algorithm services required for artificial intelligence applications.

[0051] The "computing power center" mentioned in the present invention refers to: a facility that is mainly composed of infrastructure such as wind, fire, water, electricity, and IT hardware and software equipment, and has computing power, transportation capacity, and storage capacity, including general data centers, intelligent computing centers, supercomputing centers, etc.

[0052] The "supercomputing center" mentioned in the present invention refers to: a supercomputing data center, which is a data center based on a supercomputer or a large-scale computing cluster, which can provide functions such as large-scale computing, storage and network services, and is widely used in application scenarios such as aerospace, national defense, oil exploration, climate modeling and genome sequencing.

[0053] The "computing resources" mentioned in the present invention refer to the technologies and facilities with information calculation, transmission, storage and application capabilities required for the development of the digital society, including but not limited to computing resources such as CPU (Central Processing Unit), GPU (Graphics Processing Unit), network resources such as switches and routers, storage resources such as storage arrays and distributed storage, security resources such as firewalls and intrusion detection systems, as well as supporting and guarantee resources such as wind, fire, water and electricity.

[0054] The "big model" described in the present invention includes: a big language model and a multimodal big model.

[0055] The “large language model” mentioned in the present invention refers to a large language model (LLM), which is a language model with a large parameter scale. It is designed to understand and generate human language. It is trained with a large amount of text data and can perform a wide range of tasks including text summarization, translation, sentiment analysis, etc.

[0056] The "Multimodal Large Models" mentioned in the present invention refer to models that combine multimodal information such as text, images, videos, audio, etc. for training, including but not limited to multimodal large language models.

[0057] Figure 1 A method for asynchronously and parallelly fine-tuning a large model based on the computing power of an intelligent computing center according to an embodiment of the present application is shown, and the method includes:

[0058] Step S1: Obtain the task to be processed, and determine the corresponding training set, test set, base large model and evaluation index based on the task to be processed;

[0059] Step S2: determine the hyperparameter space;

[0060] The hyperparameter space includes multiple hyperparameter groups. Each hyperparameter group includes at least one type of hyperparameter. Each type of hyperparameter has only one value in each hyperparameter group. In different hyperparameter groups, the combination of values ​​of each type of hyperparameter is different.

[0061] Step S3: Execute the model fine-tuning task, which includes: fine-tuning the base large model in parallel based on the hyperparameter space, the training set, and the computing power resources of the intelligent computing center to obtain multiple fine-tuned large models;

[0062] Among them, each fine-tuned large model corresponds to a hyperparameter group in the hyperparameter space;

[0063] Step S4: Execute the model evaluation task, which includes: based on the test set, evaluation indicators and computing resources of the intelligent computing center, perform parallel evaluation on each fine-tuned large model to obtain the evaluation results corresponding to each fine-tuned large model, and the evaluation results are the values ​​of the evaluation indicators;

[0064] Step S5: executing a model determination task, the model determination task including: determining the best evaluation result in the evaluation results, determining the fine-tuned large model corresponding to the best evaluation result as the target large model, and recording the hyperparameter group corresponding to the target large model;

[0065] The best evaluation result is the evaluation result corresponding to the best value of the evaluation indicator.

[0066] In step S1, the base large model refers to the large model that needs to be fine-tuned. In order to distinguish the large model after fine-tuning, the large model before fine-tuning is generally called the base large model. The training set refers to the training data used to fine-tune the large model, such as plain text and / or text-image pairs. The test set refers to the data used to evaluate the indicators of the fine-tuned large model, and its data type is consistent with the data type of the training set. The evaluation index refers to the indicator value finally obtained after using the test set, through the fine-tuned large model weight file, and inferring each sample. For example, accuracy.

[0067] In step S1, we first need to identify the task to be processed, such as text classification, image recognition, etc. Based on this task, we can determine the required training set and test set, select a suitable base model (such as GPT, etc.), and define evaluation indicators (such as accuracy, etc.) for subsequent model evaluation. This step ensures the pertinence and effectiveness of subsequent work. By clarifying the task and selecting the appropriate data and model, the efficiency and effectiveness of model fine-tuning can be improved.

[0068] In step S2, a hyperparameter space needs to be defined, which contains multiple hyperparameter groups. Each hyperparameter group contains different types of hyperparameters (such as learning rate, batch size, regularization coefficient, etc.), and each type of hyperparameter has only one value in each group. Different hyperparameter groups have different hyperparameter combinations, so that the performance of the model under different hyperparameter settings can be explored. By systematically defining the hyperparameter space, the impact of different hyperparameter combinations on model performance can be comprehensively evaluated, providing diversity and flexibility for subsequent model fine-tuning.

[0069] In one possible implementation, the hyperparameters include at least one of the following: learning rate, rank of LoRA, and rounds. In an exemplary scenario, if there are three types of hyperparameters, learning rate, rank of LoRA, and rounds, and the learning rate can choose 3 different values: 1e-5, 5e-5, 1e-4; the rank of LoRA can choose 3 different values: 8, 16, 32, and the round can choose 3 different values: 1, 3, 6. Then there can be 3^3=27 hyperparameter groups. And if a parameter is not determined in the range, the default value of the current parameter will be used.

[0070] In step S3, the defined hyperparameter space and training set can be used in combination with the computing power resources of the intelligent computing center to perform parallel fine-tuning on the base large model. Each fine-tuned model corresponds to a hyperparameter group, and multiple fine-tuned large models are finally obtained. Among them, the computing power resources of the intelligent computing center include at least one GPU, and only one model fine-tuning task is executed on a GPU at a time. It should be noted that the use of the computing power resources of the intelligent computing center for parallel fine-tuning can significantly shorten the model training time. At the same time, through the combination of different hyperparameter groups, a variety of fine-tuned models can be generated, which is convenient for subsequent evaluation and selection of the best model.

[0071] In a specific application scenario, there are 27 hyperparameter groups, corresponding to 27 model fine-tuning tasks, each task requires 1 GPU, and the computing power resources of the intelligent computing center are: currently there are 10 idle schedulable GPUs, which can execute 10 tasks in parallel, and use K8S to start 10 Pods at the same time, each Pod is bound to 1 GPU, and 10 tasks are run at the same time. After each task is completed, 1 GPU can be released for the next task. The total time consumption is about 3 rounds (10, 10, 7), which is 3 times that of a single task. And if the hardware resources are sufficient, the overall time consumption is theoretically equivalent to a fine-tuning task that takes the longest time, which reduces the overall fine-tuning time and improves the efficiency of model tuning.

[0072] In step S4, the test set and evaluation indicators are used to evaluate each fine-tuned large model, and the computing resources of the intelligent computing center can be used to process the evaluation of multiple models in parallel, and finally obtain the evaluation results of each model. Similarly, the computing resources of the intelligent computing center include at least one GPU, and only one model evaluation task is executed on a GPU at a time. By evaluating multiple fine-tuned models, the performance indicators of each model can be quickly obtained, providing a basis for subsequent model selection. Parallel evaluation improves the efficiency of model evaluation, reduces the overall evaluation time, and improves the efficiency of model tuning.

[0073] In step S5, the evaluation results need to be analyzed to determine the best evaluation result, and the corresponding fine-tuned large model is determined as the target large model. At the same time, the hyperparameter group information corresponding to the model is recorded. Thus, the fine-tuning evaluation task of the model is completed, and the large model with the best performance (target large model) in this fine-tuning evaluation can be obtained.

[0074] It should be noted that in the embodiment of the present application, the fine-tuning assessment task is performed asynchronously, that is, when the user's needs are sent to the system, there is no need to maintain a connection with the system. The task runs in the background and the user is notified after completion. Specifically, the user can submit a task request, the system receives the request and starts to execute the task, and the user does not need to wait and can disconnect or perform other operations. After the task is completed, the system informs the user of the result through a notification (such as an email or message). As a result, the user experience can be improved, and the user end does not need to maintain a connection, further reducing resource usage.

[0075] The method provided in the embodiment of the present application is equipped with the powerful computing power of the intelligent computing center, which can maximize the use of computing resources and improve the resource utilization of large model training and evaluation; the intelligent computing center can realize automated hyperparameter search and evaluation, reduce the complexity of manual debugging, and ensure that the best hyperparameter combination can be found, thereby improving the performance of large models in vertical fields; by allocating fine-tuning tasks of different hyperparameter combinations to the heterogeneous computing nodes of the intelligent computing center for asynchronous and parallel execution, the fine-tuning and evaluation of all combinations can be completed at one time, which can avoid serial resource waste and shorten the tuning time.

[0076] In summary, through the powerful computing resources of the intelligent computing center, combined with asynchronous parallel processing, it is possible to simultaneously handle large model fine-tuning and evaluation tasks based on multiple hyperparameter combinations, significantly shortening the model tuning time and improving the efficiency of model tuning.

[0077] In one possible implementation, Figure 2 As shown, a method for asynchronous parallel fine-tuning of a large model based on the computing power of an intelligent computing center is provided, and the method includes:

[0078] Step S6: Build a directed acyclic graph based on the Airflow tool;

[0079] The directed acyclic graph is used to define the execution order of tasks, and the execution order is: first execute the model fine-tuning task, then execute the model evaluation task, and finally execute the model determination task;

[0080] Step S3: executing a model fine-tuning task, the model fine-tuning task comprising: fine-tuning the base large model in parallel based on the hyperparameter space, the training set and the computing power resources of the intelligent computing center to obtain a plurality of fine-tuned large models, wherein each fine-tuned large model corresponds to a hyperparameter group in the hyperparameter space;

[0081] Step S4: executing a model evaluation task, wherein the model evaluation task includes: based on the test set, the evaluation index, and the computing power resources of the intelligent computing center, performing parallel evaluation on each fine-tuned large model to obtain an evaluation result corresponding to each fine-tuned large model, wherein the evaluation result is the value of the evaluation index;

[0082] Step S5: Execute a model determination task, which includes: determining the best evaluation result in the evaluation results, determining the fine-tuned large model corresponding to the best evaluation result as the target large model, and recording the hyperparameter group corresponding to the target large model; wherein the best evaluation result is the evaluation result corresponding to the best value of the evaluation indicator.

[0083] It should be noted that Airflow is an open source workflow scheduling tool, which is mainly used to orchestrate and manage complex data workflows. It is designed to help users define workflows in the form of code and provides rich functions to schedule, monitor and manage these workflows. A directed acyclic graph (DAG) is a graph structure consisting of a set of vertices and a set of directed edges, which meets the following two conditions: Each edge in the graph has a direction, that is, it points from one vertex to another. This means that the edges are ordered, indicating a certain dependency or order. There is no path in the graph that starts from a vertex, passes through several edges, and then returns to the vertex, that is, there is no cycle in the directed acyclic graph, which ensures that the execution order of tasks or events is linear.

[0084] In step S6, first, a directed acyclic graph needs to be constructed so that the corresponding evaluation task can be automatically executed after the batch fine-tuning task, and after the evaluation task is completed, the best fine-tuning large model can be automatically triggered. Airflow can be used to construct and schedule the execution of a directed acyclic graph to automatically execute a batch fine-tuning evaluation task.

[0085] And steps S3 to S5 are performed in Figure 1 The embodiments shown have been described in detail and will not be described in detail here.

[0086] The method provided in the embodiment of the present application is equipped with the powerful computing power of the intelligent computing center, which can maximize the use of computing resources and improve the resource utilization of large model training and evaluation; the intelligent computing center can realize automated hyperparameter search and evaluation, reduce the complexity of manual debugging, and ensure that the best hyperparameter combination can be found, thereby improving the performance of large models in vertical fields; by allocating fine-tuning tasks of different hyperparameter combinations to the heterogeneous computing nodes of the intelligent computing center for asynchronous and parallel execution, the fine-tuning and evaluation of all combinations can be completed at one time, which can avoid serial resource waste and shorten the tuning time.

[0087] In summary, through the powerful computing resources of the intelligent computing center, combined with asynchronous parallel processing, it is possible to simultaneously handle large model fine-tuning and evaluation tasks based on multiple hyperparameter combinations, significantly shortening the model tuning time and improving the efficiency of model tuning.

[0088] From the perspective of specific application scenarios, a method for asynchronous parallel fine-tuning of a large model based on the computing power of an intelligent computing center as shown in an embodiment of the present application is now introduced (see Figure 3 ).

[0089] The specific steps are as follows:

[0090] 1. Prepare the corresponding training set and test set according to the task situation in the vertical field, and determine the base model and evaluation indicators suitable for the task.

[0091] 2. Select a set of hyperparameter spaces, such as 3 different values ​​for learning rate: 1e-5, 5e-5, 1e-4, 3 different values ​​for LoRA rank: 8, 16, 32, and 3 different values ​​for rounds: 1, 3, 6. If a parameter is not in a certain range, the default value of the current parameter will be used.

[0092] 3. Use Airflow to build DAG.

[0093] (1) Based on the hyperparameter space, Airflow constructs this batch fine-tuning evaluation task. For each set of hyperparameters, Airflow constructs a corresponding fine-tuning evaluation task (equivalent to Figure 1 , Figure 2 The fine-tuning tasks and evaluation tasks in the embodiment shown, and the execution order is determined by using DAG), each fine-tuning evaluation task here will use the computing resources of the K8S scheduling intelligent computing center, such as GPU resources, to run the fine-tuning evaluation task, for example: 3^3=27 fine-tuning evaluation tasks can be started. And this step can be implemented based on the container operator in Airflow. Through the container operator, Airflow can trigger Kubernetes to start a Pod to run the fine-tuning evaluation task script.

[0094] (2) Create a task to select the best hyperparameters and fine-tune the large model. This step can be implemented based on Python operators. Through Python operators, an operation that will start a python script is created in Airflow. Based on the evaluation indicators, the best hyperparameters are selected and the large model is fine-tuned.

[0095] 4. Execute the DAG.

[0096] (1) All fine-tuning evaluation tasks are executed in parallel. After the fine-tuning task is completed, the evaluation task will be performed to generate the numerical value corresponding to the selected evaluation indicator.

[0097] (2) After the batch fine-tuning evaluation task is completed, the fine-tuned large model and its corresponding hyperparameter group corresponding to the best evaluation indicator can be automatically selected based on all evaluation indicators of the current batch fine-tuning evaluation task.

[0098] Therefore, relying on the computing power resources of the intelligent computing center, when fine-tuning a large model, it is only necessary to start a batch fine-tuning task once, without the need to manually start the fine-tuning task individually each time. The operation is simple, less time-consuming and has high operating efficiency. Multiple fine-tuning tasks can be executed in parallel. In theory, the total time is equivalent to the longest fine-tuning task, which significantly reduces the overall fine-tuning time and improves operating efficiency.

[0099] Figure 4 A device for asynchronously and parallelly fine-tuning a large model based on the computing power of an intelligent computing center according to an embodiment of the present application is shown. Figure 4 As shown, the device 40 includes:

[0100] An acquisition module 401 is used to acquire tasks to be processed, and determine corresponding training sets, test sets, base large models, and evaluation indicators based on the tasks to be processed;

[0101] An execution module 402 is used to determine a hyperparameter space, where the hyperparameter space includes multiple hyperparameter groups, each hyperparameter group includes at least one type of hyperparameter, and each type of hyperparameter has only one value in each hyperparameter group. In different hyperparameter groups, the combination of values ​​of each type of hyperparameter is different.

[0102] Execute the model fine-tuning task, which includes: based on the hyperparameter space, training set and computing resources of the intelligent computing center, parallel fine-tuning the base large model to obtain multiple fine-tuned large models, where each fine-tuned large model corresponds to a hyperparameter group in the hyperparameter space;

[0103] Execute the model evaluation task, which includes: based on the test set, evaluation indicators and computing resources of the intelligent computing center, perform parallel evaluation on each fine-tuned large model to obtain the evaluation results corresponding to each fine-tuned large model, and the evaluation results are the values ​​of the evaluation indicators;

[0104] Execute the model determination task, which includes: determining the best evaluation result in the evaluation results, determining the fine-tuned large model corresponding to the best evaluation result as the target large model, and recording the hyperparameter group corresponding to the target large model; wherein the best evaluation result is the evaluation result corresponding to the best value of the evaluation indicator.

[0105] In one possible implementation, the hyperparameters include at least one of the following: learning rate, rank and round of LoRA.

[0106] In a possible implementation, the execution module 402 is also used to build a directed acyclic graph based on the Airflow tool before executing the model fine-tuning task. The directed acyclic graph is used to define the execution order of the tasks. The execution order is: first execute the model fine-tuning task, then execute the model evaluation task, and finally execute the model confirmation task.

[0107] In one possible implementation, the computing resources of the intelligent computing center include at least one GPU, and only one model fine-tuning task or model evaluation task is executed on one GPU at a time.

[0108] The device provided in the embodiment of the present application is equipped with the powerful computing power of the intelligent computing center, which can maximize the use of computing resources and improve the resource utilization of large model training and evaluation; the intelligent computing center can realize automated hyperparameter search and evaluation, reduce the complexity of manual debugging, and ensure that the best hyperparameter combination can be found, thereby improving the performance of large models in vertical fields; by allocating fine-tuning tasks of different hyperparameter combinations to the heterogeneous computing nodes of the intelligent computing center for asynchronous and parallel execution, the fine-tuning and evaluation of all combinations can be completed at one time, which can avoid serial resource waste and shorten the tuning time.

[0109] In summary, through the powerful computing resources of the intelligent computing center, combined with asynchronous parallel processing, it is possible to simultaneously handle large model fine-tuning and evaluation tasks based on multiple hyperparameter combinations, significantly shortening the model tuning time and improving the efficiency of model tuning.

[0110] Please refer to Figure 5The embodiment of the present application also provides an electronic device 50, including a processor 501, a memory 502, and a computer program stored in the memory 502 and executable on the processor 501. When the computer program is executed by the processor 501, each process of the above-mentioned method embodiment of asynchronous parallel fine-tuning of a large model based on the computing power of an intelligent computing center is implemented, and the same technical effect can be achieved. To avoid repetition, it will not be described here.

[0111] The embodiment of the present application also provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, each process of the above-mentioned method embodiment for asynchronously and parallelly fine-tuning a large model based on the computing power of an intelligent computing center is implemented, and the same technical effect can be achieved. To avoid repetition, it is not repeated here. Among them, the computer-readable storage medium is, for example, a read-only memory (ROM), a random access memory (RAM), a disk or an optical disk, etc.

[0112] An embodiment of the present application also provides a computer program product, including computer instructions, which, when executed by a processor, implement the various processes of the above-mentioned method embodiment of asynchronous parallel fine-tuning of a large model based on the computing power of an intelligent computing center, and can achieve the same technical effect. To avoid repetition, it will not be repeated here.

[0113] It should be noted that, in this article, the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, an element defined by the sentence "comprises a ..." does not exclude the existence of other identical elements in the process, method, article or device including the element.

[0114] Through the description of the above implementation methods, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be implemented by means of software plus a necessary general hardware platform, and of course by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of the present application, or the part that contributes to the prior art, can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, a magnetic disk, or an optical disk), and includes a number of instructions for a terminal (which can be a mobile phone, a computer, a server, an air conditioner, or a network device, etc.) to execute the methods described in each embodiment of the present application.

[0115] The embodiments of the present application are described above in conjunction with the accompanying drawings, but the present application is not limited to the above-mentioned specific implementation methods. The above-mentioned specific implementation methods are merely illustrative and not restrictive. Under the guidance of the present application, ordinary technicians in this field can also make many forms without departing from the purpose of the present application and the scope of protection of the claims, all of which are within the protection of the present application.

Claims

1. A method for fine-tuning a large model asynchronously and in parallel based on the computing power of an intelligent computing center, characterized in that: The method comprises: Step S1: Obtain the task to be processed, and determine the corresponding training set, test set, base large model and evaluation index based on the task to be processed; Step S2: determining a hyperparameter space, wherein the hyperparameter space includes a plurality of hyperparameter groups, each hyperparameter group includes at least one type of hyperparameter, and each type of hyperparameter has only one value in each hyperparameter group, and in different hyperparameter groups, the combination of values ​​of each type of hyperparameter is different; Step S3: executing a model fine-tuning task, the model fine-tuning task comprising: fine-tuning the base large model in parallel based on the hyperparameter space, the training set and the computing power resources of the intelligent computing center to obtain a plurality of fine-tuned large models, wherein each fine-tuned large model corresponds to a hyperparameter group in the hyperparameter space; Step S4: executing a model evaluation task, wherein the model evaluation task includes: based on the test set, the evaluation index, and the computing power resources of the intelligent computing center, performing parallel evaluation on each fine-tuned large model to obtain an evaluation result corresponding to each fine-tuned large model, wherein the evaluation result is the value of the evaluation index; Step S5: Execute a model determination task, which includes: determining the best evaluation result in the evaluation results, determining the fine-tuned large model corresponding to the best evaluation result as the target large model, and recording the hyperparameter group corresponding to the target large model; wherein the best evaluation result is the evaluation result corresponding to the best value of the evaluation indicator.

2. The method according to claim 1, characterized in that The hyperparameters include at least one of the following: learning rate, rank and round of LoRA.

3. The method according to claim 1, characterized in that Before step S3, the method further includes: Step S6: construct a directed acyclic graph based on the Airflow tool, and the directed acyclic graph is used to define the execution order of tasks. The execution order is: first execute the model fine-tuning task, then execute the model evaluation task, and finally execute the model determination task.

4. The method according to any one of claims 1 to 3, characterized in that The computing resources of the intelligent computing center include at least one graphics processing unit (GPU), and only one model fine-tuning task or one model evaluation task is executed on one GPU at a time.

5. A device for asynchronously and parallelly fine-tuning a large model based on the computing power of an intelligent computing center, characterized in that: The device comprises: An acquisition module is used to acquire tasks to be processed, and determine corresponding training sets, test sets, base large models, and evaluation indicators based on the tasks to be processed; An execution module is used to determine a hyperparameter space, wherein the hyperparameter space includes multiple hyperparameter groups, each hyperparameter group includes at least one type of hyperparameter, and each type of hyperparameter has only one value in each hyperparameter group, and in different hyperparameter groups, the combination of values ​​of each type of hyperparameter is different; Executing a model fine-tuning task, the model fine-tuning task comprising: fine-tuning the base large model in parallel based on the hyperparameter space, the training set, and the computing power resources of the intelligent computing center to obtain a plurality of fine-tuned large models, wherein each fine-tuned large model corresponds to a hyperparameter group in the hyperparameter space; Execute a model evaluation task, the model evaluation task comprising: based on the test set, the evaluation index and the computing power resources of the intelligent computing center, perform parallel evaluation on each fine-tuned large model to obtain an evaluation result corresponding to each fine-tuned large model, wherein the evaluation result is the value of the evaluation index; Execute a model determination task, the model determination task comprising: determining the best evaluation result in the evaluation results, determining the fine-tuned large model corresponding to the best evaluation result as the target large model, and recording the hyperparameter group corresponding to the target large model; wherein the best evaluation result is the evaluation result corresponding to the best value of the evaluation indicator.

6. The device according to claim 5, characterized in that The hyperparameters include at least one of the following: learning rate, rank and round of LoRA.

7. The device according to claim 5, characterized in that The execution module is also used to build a directed acyclic graph based on the Airflow tool before executing the model fine-tuning task. The directed acyclic graph is used to define the order of execution of tasks. The order of execution is: first execute the model fine-tuning task, then execute the model evaluation task, and finally execute the model determination task.

8. An electronic device, characterized in that: include: A processor, a memory, and a program stored in the memory and executable on the processor, wherein when the program is executed by the processor, the method for asynchronously and parallelly fine-tuning a large model based on the computing power of an intelligent computing center as described in any one of claims 1 to 4 is implemented.

9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by the processor, the method for asynchronously and parallelly fine-tuning a large model based on the computing power of an intelligent computing center as described in any one of claims 1 to 4 is implemented.

10. A computer program product, characterized in that It includes computer instructions, which, when executed by the processor, implement the method for asynchronous parallel fine-tuning of a large model based on the computing power of an intelligent computing center as described in any one of claims 1 to 4.

Citation Information

Patent Citations

  • Hyper-parameter determination method and device, computer equipment and storage medium

    CN112529211A

  • Airflow-based model training scheduling method and device

    CN113837697A

  • Model hyper-parameter valuing method and device, processing core, equipment, chip and medium

    CN116542286A

  • Large model training deployment method and system combining fine tuning technology and distributed scheduling

    CN117632381A

  • Question and answer model-based reply information searching method and device, question and answer model training method and device, electronic equipment and storage medium

    CN118332084A

Cited By

  • Fine adjustment method, device, equipment and medium

    CN120851086A