Fine tuning method and device for hybrid expert model, equipment, medium and program product
By dividing the data loading process into multiple stages, combining expert popularity and hardware device affinity, the problems of high memory demand and low efficiency in fine-tuning of hybrid expert models are solved, and a more efficient fine-tuning process is achieved.
Patent Information
- Application Number
- CN202510291051.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-12
- Publication Date
- 2025-07-18
AI Technical Summary
The prior art has high memory requirements and low efficiency when fine-tuning hybrid expert models, mainly due to the large communication overhead caused by unloading all model parameters, output activation and then loading on demand.
The data loading process is divided into interstage loading, model interlayer loading and inter-expert loading, and combined with expert popularity and hardware device affinity, loading decisions are determined and loading processes are scheduled to reduce redundant unloading and loading operations.
It reduces communication overhead, improves fine-tuning efficiency, reduces video memory requirements, optimizes hardware resource utilization, and improves fine-tuning efficiency of hybrid expert models.
Smart Images

Figure CN120337983A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of model fine-tuning, and in particular to a fine-tuning method for a mixture of experts model, a fine-tuning device for a mixture of experts model, a fine-tuning device for a mixture of experts model, a storage medium, and a computer program product. Background Art
[0002] In recent years, the mixture of experts model has received extensive attention. It can increase the number of model parameters while maintaining a comparable amount of computation, thereby improving the model's performance. The mixture of experts model uses multiple sub-models, called "experts", and each sub-model specializes in different sub-tasks. In order to efficiently utilize the capabilities of the mixture of experts model, it is necessary to fine-tune the mixture of experts model and update the model weights. When using the existing offloading strategy to fine-tune the mixture of experts model, it is necessary to offload all model parameters and output activations and then load them as needed, which will significantly increase the communication overhead of the mixture of experts model and result in low fine-tuning efficiency. Summary of the Invention
[0003] The main purpose of this application is to provide a fine-tuning method for a mixture of experts model, a fine-tuning device for a mixture of experts model, a fine-tuning device for a mixture of experts model, a storage medium, and a computer program product, aiming to solve the technical problems of high video memory requirements for fine-tuning the mixture of experts model and low fine-tuning efficiency of the mixture of experts model under the offloading strategy.
[0004] To achieve the above object, this application proposes a fine-tuning method for a mixture of experts model. The mixture of experts model is composed of at least one layer of mixture of experts layer. The mixture of experts layer is composed of an attention model block and a mixture of experts model block. The mixture of experts model block is composed of a gating operation and at least one expert. The fine-tuning method for the mixture of experts model includes:
[0005] Under the determined mapping relationship, the data loading process of loading the data related to model fine-tuning of the mixture of experts layer is divided into inter-stage loading, inter-model-layer loading, and inter-expert loading according to the overlapping relationship with the next calculation process.
[0006] Combining the expert popularity and affinity, determine the loading decisions for inter-stage loading, inter-model-layer loading, and inter-expert loading, where the expert popularity is the frequency of different experts being activated in the mixture of experts model, and the affinity is the affinity between the calculated load and the hardware device determined in the offline stage.
[0007] Schedule the loading decisions, and load the data related to model fine-tuning of the mixture of experts layer during inter-stage loading, inter-model-layer loading, and inter-expert loading respectively.
[0008] In one embodiment, the steps of determining the loading decisions for inter-stage loading, inter-model-layer loading, and inter-expert loading by combining expert popularity and affinity include:
[0009] During inter-stage loading, according to the computational intensity of each model block in the mixture-of-experts layer, determine the loading priority of each model block in the mixture-of-experts layer. The loading decision for inter-stage loading is to determine the model blocks to be loaded in the next pipeline stage based on the loading priority and the historical expert popularity of the experts.
[0010] During inter-model-layer loading, predict the expert popularity of the experts in the mixture-of-experts model blocks to be loaded according to the expert popularity predictor. The loading decision for inter-model-layer loading is to select the popular experts and delete the unpopular experts from the mixture-of-experts model blocks to be loaded according to the expert popularity and the affinity.
[0011] During inter-expert loading, the loading decision for inter-expert loading is to determine whether to load the popular experts according to the real-time expert popularity generated by the gating operation.
[0012] In one embodiment, before the step of predicting the expert popularity of the experts in the mixture-of-experts model blocks to be loaded according to the expert popularity predictor during inter-model-layer loading, it includes:
[0013] Initialize the to-be-trained expert popularity predictor through the weights of the gating operation of the mixture-of-experts layer where the expert popularity predictor is located.
[0014] Use the target fine-tuning dataset as the training dataset to train the initialized to-be-trained expert popularity predictor to obtain the expert popularity predictor.
[0015] In one embodiment, the steps of scheduling the loading decisions include:
[0016] Determine the loading order of the mixture-of-experts layer according to the data loading process status of the hardware device running the mixture-of-experts layer.
[0017] According to the loading order, add or delete the loading decisions of the mixture-of-experts layer in the data loading queues for inter-stage loading, inter-model-layer loading, and inter-expert loading.
[0018] Allocate priorities to the data loading queues according to the urgency of the task computing requirements for inter-stage loading, inter-model-layer loading, and inter-expert loading.
[0019] Schedule the loading decisions in the data loading queue with the highest priority to fine-tune the mixture-of-experts model.
[0020] In one embodiment, before the step of determining the loading order of the mixture-of-experts layer according to the data loading process status of the hardware device running the mixture-of-experts layer, the following steps are included:
[0021] In the data loading queues of inter-phase loading, inter-model-layer loading, and inter-expert loading, determine the number of movements of data related to model fine-tuning in the loading decision;
[0022] Transmit the number of movements to the communication stream in the GPU;
[0023] Obtain the data loading process status from the communication stream through CUDA events.
[0024] In one embodiment, before the step of dividing the data loading process of loading the data related to model fine-tuning of the mixture-of-experts layer into inter-phase loading, inter-model-layer loading, and inter-expert loading according to the overlapping relationship with the next calculation process under the determined mapping relationship, the following steps are included:
[0025] Under the preset conditions in the offline phase, fine-tune the mixture-of-experts layer to obtain the memory usage and execution time of the mixture-of-experts layer during fine-tuning;
[0026] According to the memory usage and the execution time, respectively obtain the mapping relationship between the mixture-of-experts layer and the hardware device running the mixture-of-experts layer and the affinity of the hardware device.
[0027] In addition, to achieve the above object, the present application also proposes a fine-tuning device for a mixture-of-experts model. The fine-tuning device for the mixture-of-experts model includes: a dividing module, configured to divide the data loading process of loading the data related to model fine-tuning of the mixture-of-experts layer into inter-phase loading, inter-model-layer loading, and inter-expert loading according to the overlapping relationship with the next calculation process under the determined mapping relationship;
[0028] A determining module, configured to determine the loading decisions of inter-phase loading, inter-model-layer loading, and inter-expert loading in combination with expert popularity and affinity, where the mapping relationship and the affinity are the mapping relationship between the mixture-of-experts layer and the hardware device running the mixture-of-experts layer and the affinity of the hardware device determined in the offline phase;
[0029] A scheduling module, configured to schedule the loading decisions and load the data related to model fine-tuning of the mixture-of-experts layer in inter-phase loading, inter-model-layer loading, and inter-expert loading respectively.
[0030] In addition, to achieve the above object, the present application further provides a fine-tuning device for a mixture-of-experts model, the device including: a memory, a processor, and a computer program stored on the memory and executable on the processor, the computer program being configured to implement the steps of the fine-tuning method for the mixture-of-experts model as described above.
[0031] In addition, to achieve the above object, the present application further provides a storage medium, the storage medium being a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it implements the steps of the fine-tuning method for the mixture-of-experts model as described above.
[0032] In addition, to achieve the above object, the present application further provides a computer program product, the computer program product including a computer program, and when the computer program is executed by a processor, it implements the steps of the fine-tuning method for the mixture-of-experts model as described above.
[0033] One or more technical solutions proposed by the present application have at least the following technical effects:
[0034] Under the existing offloading strategy, when fine-tuning the mixture-of-experts model, all model parameters and output activations need to be offloaded and then loaded on demand. During this process, a large number of data offloading and reloading operations will cause a large communication overhead, consume a lot of time and resources, and lead to low fine-tuning efficiency.
[0035] However, the present application first divides the existing data loading process into inter-stage loading, inter-model-layer loading, and inter-expert loading according to the overlapping relationship with the next calculation process under the determined mapping relationship, and combines the expert popularity and affinity to determine the loading decisions for each loading process. Finally, the loading decisions are scheduled to load the data related to the mixture-of-experts layer and model fine-tuning at different stages.
[0036] By dividing the data loading process into multiple stages and performing data loading in a targeted manner in stages, it is not necessary to offload all model parameters and output activations and then load on demand as in the existing offloading strategy, reducing frequent data transmission, thereby reducing the communication overhead; by combining the expert popularity and affinity to determine the loading decisions, the characteristics of the hardware device can be better utilized, and reasonable loading decisions can ensure that the data loaded in each stage is the most relevant data to the current fine-tuning task, avoiding a large number of redundant offloading and loading operations, making the data processing in the fine-tuning process more efficient, and thus improving the fine-tuning efficiency of the mixture-of-experts model; since only a small number of experts are highly activated when fine-tuning the mixture-of-experts model, storing most of the low-frequency activated experts in the CPU can reduce the video memory requirement for fine-tuning the mixture-of-experts model. BRIEF DESCRIPTION OF THE DRAWINGS
[0037] The accompanying drawings here are incorporated into the description and form a part of this description, showing embodiments consistent with this application, and are used together with the description to explain the principles of this application.
[0038] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the following will briefly introduce the accompanying drawings required for use in the description of the embodiments or the prior art. Obviously, for those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0039] Figure 1 It is a schematic structural diagram of the mixture-of-experts model of this application;
[0040] Figure 2 It is a schematic flowchart provided by the first embodiment of the fine-tuning method of the mixture-of-experts model of this application;
[0041] Figure 3 It is a schematic diagram of each stage of the fine-tuning method of the mixture-of-experts model of this application;
[0042] Figure 4 It is a schematic diagram of the expert popularity predictor in the fine-tuning method of the mixture-of-experts model of this application;
[0043] Figure 5 It is a schematic diagram of the overall fine-tuning method of the mixture-of-experts model of this application;
[0044] Figure 6 It is a schematic diagram of the module structure of the fine-tuning device of the mixture-of-experts model of this application;
[0045] Figure 7 It is a schematic diagram of the device structure of the hardware operating environment involved in the fine-tuning method of the mixture-of-experts model of this application.
[0046] The implementation, functional features, and advantages of this application will be further described with reference to the embodiments and the accompanying drawings. Detailed implementation manners
[0047] It should be understood that the specific embodiments described herein are only used to explain the technical solutions of this application and are not used to limit this application.
[0048] To better understand the technical solutions of this application, the following will be described in detail in combination with the accompanying drawings of the description and the specific implementation manners.
[0049] In recent years, the mixture-of-experts model has received extensive attention. It can increase the number of model parameters while maintaining a comparable amount of computation, thereby improving the model's performance. The mixture-of-experts model uses multiple sub-models, called "experts", and each sub-model specializes in different sub-tasks.
[0050] Figure 1 shows the mixture-of-experts layer of the mixture-of-experts model, which is composed of multiple mixture-of-experts layers.
[0051] The mixture-of-experts model block contains two main components: a gating operation and multiple sub-models, i.e., experts. The purpose of the gating operation component is to route each input token in the input data to each expert. Specifically, it includes: the gating operation assigns all the given tokens in the input data to different experts; and then the experts perform calculations on the tokens. Each expert has the same structure as the multi-layer perceptron model block in the Transformer model.
[0052] To make the mixture-of-experts model have better performance, it is necessary to fine-tune the mixture-of-experts model. However, due to the sparse activation characteristics of the mixture-of-experts model, all the expert weight data of one layer need to be moved at a time, but only a few experts are calculated. The existing fine-tuning methods based on traditional model offloading strategies are not applicable to the mixture-of-experts model. This method needs to offload all model parameters and their output activations and then load them on demand. Due to the characteristic of the increasing ratio of data transmission to calculation in the mixture-of-experts model, the data loading process will block the calculation of experts.
[0053] Based on this, the present application proposes a new offloading strategy to fine-tune the mixture-of-experts model, specifically including: under the determined mapping relationship, dividing the existing data loading process into inter-stage loading, inter-model-layer loading, and inter-expert loading according to the overlapping relationship with the next calculation process, and combining expert popularity and affinity to determine the loading decisions of each loading process. Finally, scheduling the loading decisions to load the data related to the fine-tuning of the mixture-of-experts layer at different stages to complete the fine-tuning of the mixture-of-experts model.
[0054] It should be noted that the execution subject of this embodiment can be the fine-tuning device of the mixture-of-experts model, or a computing service device with data processing, network communication, and program running functions, such as a tablet computer, a personal computer, a mobile phone, etc., or an electronic device, a processor, etc. that can implement the above functions. Hereinafter, the fine-tuning device of the mixture-of-experts model is taken as an example to illustrate this embodiment and the following embodiments.
[0055] Based on this, the embodiment of the present application provides a method for fine-tuning a mixture-of-experts model, referring to Figure 2 , Figure 2 which is the schematic flowchart of the first embodiment of the method for fine-tuning the mixture-of-experts model of the present application.
[0056] The mixture-of-experts model is composed of mixture-of-experts layers, and the mixture-of-experts layers are composed of attention model blocks and mixture-of-experts model blocks;
[0057] In this embodiment, the fine-tuning method of the mixture-of-experts model includes steps S10 to S30:
[0058] Step S10, under the determined mapping relationship, the data loading process of loading the data related to the mixture-of-experts layer and model fine-tuning is divided into inter-stage loading, inter-model-layer loading, and inter-expert loading according to the overlapping relationship with the next calculation process.
[0059] It should be noted that the determined mapping relationship refers to which GPU device each layer of the mixture-of-experts layer will run on. When loading the data related to model fine-tuning, the data related to model fine-tuning includes model parameters and intermediate activations. The pipeline parallelism method with lower requirements for inter-device interconnection is adopted. In the pipeline parallelism method, several layers of the mixture-of-experts model are divided into several pipeline stages, and the input data enters different stages in sequence, and different stages of calculations are performed by GPU devices. In this application, a new data loading strategy, i.e., hierarchical loading, is used. All the data related to model fine-tuning is not loaded in the pipeline stage. The data loading process is divided into three stages, namely inter-stage loading, inter-model-layer loading, and inter-expert loading. The schematic diagrams of the three stages are as Figure 3 shown. Inter-stage loading means loading the data of the next pipeline stage when calculating the current pipeline stage on the same device. Inter-model-layer loading means loading the data of the next layer when calculating the current mixture-of-experts layer in the same pipeline stage. Inter-expert data loading means loading the data of other experts when calculating the current expert in the same mixture-of-experts layer.
[0060] Step S20, combining the expert popularity and affinity, determine the loading decisions for inter-stage loading, inter-model-layer loading, and inter-expert loading, where the expert popularity is the frequency of different experts being activated in the mixture-of-experts model, and the affinity is the affinity between the calculated load and the hardware device determined in the offline stage;
[0061] It should be noted that the expert popularity is the frequency of different experts being activated in the mixture-of-experts model; in the offline stage (i.e., the preparation stage before the actual operation of the model), the affinity of each model block of the mixture-of-experts layer to the hardware device is determined, which is a kind of adaptability between the calculated load and the hardware device, specifically referring to whether different model blocks (attention model blocks, gating operations, and experts) are more suitable for calculation on the CPU or more suitable for calculation on the GPU under a specific workload.
[0062] Based on the above expert popularity and affinity, it is necessary to determine the loading decisions in the processes of inter-stage loading, inter-model-layer loading, and inter-expert loading, that is, in different loading processes, determine the loading decisions for each stage. Exemplarily, the loading decision can be to load the attention model block and gating operation first and then load the experts, or to decide which experts are loaded first according to the expert popularity.
[0063] Step S30, scheduling loading decisions, for inter-stage loading, inter-model-layer loading, and inter-expert loading, load data related to the mixture-of-experts layer and model fine-tuning respectively.
[0064] It should be noted that there are loading decisions for each loading stage. Therefore, it is necessary to manage these loading decisions. Exemplarily, a queue can be used to manage the loading decisions of different stages. By adding or deleting the names of model blocks (attention model blocks, gating operations, and experts) in the corresponding queue, the loading decisions can be hierarchically managed.
[0065] In this embodiment, under the determined mapping relationship, the existing data loading process is divided into inter-stage loading, inter-model-layer loading, and inter-expert loading according to the overlapping relationship with the next calculation process. Combining the popularity and affinity of experts, the loading decisions for each loading process are determined. Finally, the loading decisions are scheduled to load data related to the mixture-of-experts layer and model fine-tuning at different stages.
[0066] By dividing the data loading process into multiple stages and performing data loading in a targeted manner in stages, it is not necessary to unload all model parameters and output activations and then load them on demand as in the existing unloading strategy, reducing frequent data transmission, thereby reducing communication overhead; by combining the popularity and affinity of experts to determine the loading decisions, the characteristics of the hardware device can be better utilized. A reasonable loading decision can ensure that the data loaded at each stage is the most relevant data for the current fine-tuning task, avoiding a large number of redundant unloading and loading operations, making the data processing during fine-tuning more efficient, thereby improving the fine-tuning efficiency of the mixture-of-experts model; since only a small number of experts are highly activated when fine-tuning the mixture-of-experts model, storing most of the low-frequency activated experts in the CPU can reduce the video memory requirements for fine-tuning the mixture-of-experts model.
[0067] In one implementation, before fine-tuning the mixture-of-experts model, analyze the memory situation and execution time of the mixture-of-experts model fine-tuning through a static analyzer under given conditions and use them for fine-tuning the mixture-of-experts model; specifically, it includes steps A10 to A20:
[0068] Step A10, under the preset conditions in the offline stage, fine-tune the mixture-of-experts layer to obtain the memory usage and execution time of the mixture-of-experts layer during fine-tuning;
[0069] It should be noted that the preset conditions include the hardware conditions for running the mixture-of-experts model (GPU video memory, GPU computing power, CPU main memory, CPU computing power, memory bandwidth), the model configuration of the mixture-of-experts model (number of layers, hidden layer dimension, model structure), the input data configuration of the mixture-of-experts model (batch size, sentence length size), and fine-tuning the mixture-of-experts layer under the preset conditions; analyzing the memory usage and execution time during fine-tuning through an analyzer. Exemplarily, a single mixture-of-experts layer can be fine-tuned to obtain the fine-tuning data of a single mixture-of-experts layer.
[0070] Step A20, obtaining the mapping relationship from the mixture-of-experts layer to the hardware device running the mixture-of-experts layer and the affinity of the hardware device according to the memory usage and the execution time respectively.
[0071] It should be noted that the analyzer generates the mapping relationship from the mixture-of-experts layer to the hardware device running the mixture-of-experts layer according to the memory occupation of a single mixture-of-experts layer. Since there are many mixture-of-experts layers and many hardware devices, it is possible to obtain on which hardware device each mixture-of-experts layer runs according to the mapping relationship. In addition, the analyzer records the execution time of a single expert (a single mixture-of-experts layer includes at least one expert) on the CPU and GPU devices respectively under the same number of input tokens, that is, representing the affinity of the hardware device by comparing the execution time under the same workload. The affinity of the hardware device can obtain whether a certain expert is more suitable for computing on the CPU or more suitable for computing on the GPU under a specific workload.
[0072] In this embodiment, when fine-tuning the mixture-of-experts layer in the offline stage, different mixture-of-experts layers will occupy different amounts of memory when processing data. According to the mapping relationship of the memory usage, it is possible to avoid the performance degradation or data processing failure caused by insufficient memory of the hardware device, improve the accuracy and stability of the model, and also avoid resource waste and improve the utilization rate of hardware resources; according to the affinity, it can be determined whether an expert is more suitable for computing on the CPU or more suitable for computing on the GPU under a specific workload. When the workload of the expert matches the characteristics of the computing device, the computing efficiency will be significantly improved, and resource waste can also be avoided.
[0073] Based on the first embodiment of the present application, in the second embodiment of the present application, the same or similar content as the above first embodiment can be referred to the above introduction and will not be repeated hereinafter. Step S20, the fine-tuning method of the mixture-of-experts model further includes steps D10 to D30:
[0074] Step D10: During inter-stage loading, based on the computing intensity of each model block in the Mixture of Experts (MoE) layer under the mapping relationship, determine the loading priority of each model block in the MoE layer. The loading decision for inter-stage loading is to determine the model block to be loaded in the next pipeline stage based on the loading priority and the historical popularity of the experts.
[0075] It should be noted that during inter-stage loading, the gating operation in the MoE layer serves to allocate input data to different experts. The computing intensity of an expert reflects the amount of computation required to process the data allocated to it after passing through the gating operation. Based on this, the loading priority of different experts in the MoE layer is determined. If the data has a high computing intensity, the relevant expert is preferentially loaded to ensure that the data can be processed in a timely manner. The loading decision for inter-stage loading is to comprehensively consider the loading priority and historical popularity of the experts to determine the target expert to be loaded in the MoE model. This decision-making method helps to more reasonably select the experts to be loaded during inter-stage loading, improving the processing efficiency and accuracy of the MoE model. For the attention model blocks, gating operations, and experts that are not loaded at this stage, their loading decisions are postponed to the subsequent inter-model-layer loading and inter-expert loading to utilize more accurate expert popularity.
[0076] Step D20: During inter-model-layer loading, predict the popularity of the experts in the MoE model blocks to be loaded using an expert popularity predictor. The loading decision for inter-model-layer loading is to select the popular experts and delete the unpopular experts from the MoE model blocks to be loaded based on the expert popularity and the affinity.
[0077] It should be noted that the expert popularity predictor is used to predict the popularity of the target experts. If an expert has a high popularity, it means that it has a higher computing intensity and is more suitable to be loaded when processing the current task. Affinity represents the relationship between an expert and the hardware device under a specific workload. During inter-model-layer loading, experts will be dynamically added to or deleted from the queue based on the predicted expert popularity and computing affinity. Among them, popular experts refer to those with an expert popularity greater than or equal to a preset threshold. For example, the preset threshold can be set to 80%. Unpopular experts refer to those with an expert popularity less than the preset threshold.
[0078] Under the decision of inter-model-layer loading, in order to save resources (such as computing resources, memory resources, etc.) and improve the overall model operation efficiency, these unpopular experts will be deleted and not participate in the subsequent data loading process. This can ensure that the model can focus on the more urgently needed experts during inter-model-layer loading, optimizing the overall computing performance and resource utilization efficiency.
[0079] It can be understood that, due to the dynamic nature of the expert activation in the mixture-of-experts model, the true expert popularity remains unknown before the gating operation is completed, making the data prefetching process in the offloading strategy difficult. Therefore, this application predicts the expert popularity of experts in advance through an expert popularity predictor and pre-determines the allocation of experts between the GPU and the CPU.
[0080] In a feasible implementation, before step D20, it includes steps E10 to E20:
[0081] Step E10, initialize the to-be-trained expert popularity predictor through the weights of the gating operation in the mixture-of-experts layer where the expert popularity predictor is located;
[0082] It should be noted that the expert popularity predictor is an independently operating structure additionally introduced in the mixture-of-experts model and does not change the original mixture-of-experts model. As Figure 4 shown, before the gating operation, there is an expert popularity predictor in each mixture-of-experts layer, and the expert popularity predictor adopts the same structure as the gating operation.
[0083] In the mixture-of-experts layer, the gating operation plays a role in allocating input data to different experts; the gating operation can determine which expert each input data should be allocated to for calculation according to the characteristics of the input data; the weights of the gating operation reflect the importance of different experts in processing data.
[0084] Since the weights of the gating operation already reflect the relative importance of different experts in processing data, using these weights as the initial values of the expert popularity predictor can provide a relatively reasonable starting point for the predictor, which enables the expert popularity predictor to have a prior estimate based on the existing information within the model at the beginning of training, helping to improve the training efficiency and accuracy. During the subsequent training process, the expert popularity predictor can further adjust these initial values according to more data and feedback, so as to more accurately predict the expert popularity.
[0085] Step E20, use the target fine-tuning dataset as the training dataset to train the initialized to-be-trained expert popularity predictor to obtain the expert popularity predictor.
[0086] It should be noted that the expert popularity predictor is used to predict the expert popularity of each expert (i.e., sub-modules with different specializations) in the mixture-of-experts model, where the expert popularity represents the frequency of a certain expert being used when processing a specific task or data.
[0087] Before training, the expert popularity predictor needs to be initialized by assigning initial parameter values to it. When using the target fine-tuning dataset to train the initialized expert popularity predictor, the expert popularity predictor adjusts its parameters according to the actual performance of each expert in different situations in the dataset to more accurately predict the popularity of experts.
[0088] In this embodiment, the expert popularity predictor to be trained is initialized through the weights of the gating operation in the mixture-of-experts layer where the expert popularity predictor is located, and then trained using the target fine-tuning dataset, enabling the expert popularity predictor to learn the activation status of each expert under different tasks or data types, so as to more accurately predict which experts are more popular; by optimizing the selection of experts, the mixture-of-experts model can reduce unnecessary computational overhead while maintaining or even improving performance, thereby enhancing the fine-tuning performance of the mixture-of-experts model.
[0089] Step D30, during expert inter-loading, the loading decision for expert inter-loading is to determine whether to load popular experts based on the real-time expert popularity generated by the gating operation.
[0090] It should be noted that after the input data enters the current mixture-of-experts layer of the mixture-of-experts model, the corresponding gating operation generates real-time expert popularity, and based on the real-time expert popularity, it is determined whether to load popular experts.
[0091] In this embodiment, different levels of expert popularity are used to plan data flow and computing placement, reducing data flow overhead, increasing throughput, and fine-tuning the mixture-of-experts model under limited resources. Inter-stage loading determines the loading priorities of the attention model block and the gating operation in the mixture-of-experts layer according to the computational intensity of different model blocks in the mixture-of-experts layer, which can ensure that computing resources (such as the computing capabilities of CPUs and GPUs) are preferentially allocated to the attention model block and the gating operation with stronger data processing requirements to improve the overall data processing efficiency; in addition, determining the model blocks of the next pipeline stage based on the loading priority and the historical expert popularity of experts can select experts with high historical expert popularity during inter-stage loading, which helps to avoid loading unnecessary experts, reduce memory occupancy and waste of communication resources, and reduce the video memory requirements for fine-tuning the mixture-of-experts model.
[0092] Model inter-layer loading selects popular experts and deletes unpopular experts according to the expert popularity predicted by the expert popularity predictor and affinity, which can achieve more precise resource allocation, reduce unnecessary expert loading, and enable more efficient computing resources to be concentrated on experts with higher computational intensity, thereby improving the efficiency of fine-tuning the mixture-of-experts model.
[0093] The expert - level loading determines whether to load popular experts based on the real - time expert popularity generated by the gating operation. It can instantaneously adjust the experts that need to be loaded onto the GPU for calculation, avoiding the limitations of fixedly loading all experts, improving the flexibility and adaptability of the loading process of the mixture - of - experts model, being able to maintain good performance in different data scenarios, and reducing the waste of communication resources and performance degradation caused by loading unnecessary experts.
[0094] Based on the first embodiment of the present application, in the third embodiment of the present application, the same or similar content as the above - mentioned first embodiment can be referred to the above introduction and will not be repeated hereinafter. Step S30, the fine - tuning method of the mixture - of - experts model further includes steps T10 - T40:
[0095] Three loading phases generate different loading decisions. These loading decisions are made at different times. The inter - stage loading makes loading decisions on the data of the model blocks in the next pipeline stage. The inter - model - layer loading makes loading decisions on the attention model blocks and mixture - of - experts model blocks of the mixture - of - experts layer. And the inter - expert loading makes loading decisions for fewer experts in the current mixture - of - experts layer. The decision - making time and the execution time of data loading are inconsistent. However, since these loading phases depend on the same PCIe (Peripheral Component Interconnect Express, high - speed serial computer expansion bus standard) link in the same direction and cannot be executed simultaneously, they may interfere with each other. Therefore, it is necessary to schedule the loading decisions of each stage.
[0096] Step T10, determine the loading order of the mixture - of - experts layer according to the data loading process status of the hardware device running the mixture - of - experts layer;
[0097] It should be noted that the data loading process has different states, such as being in the initial loading state, partially loaded, etc. Exemplarily, the loading process status of the GPU can be queried regularly or irregularly by the main program in the CPU. Different mixture - of - experts layers have different requirements and dependencies on hardware resources. According to the data loading process status of the hardware device, it can be determined which mixture - of - experts layers to load first. Dynamically determine the loading order of the mixture - of - experts layers during operation according to the data loading process status.
[0098] In a feasible implementation manner, before step T10, it includes steps L10 - L30:
[0099] Step L10, in the data loading queues of the inter - stage loading, inter - model - layer loading, and inter - expert loading, determine the moving quantity of the data related to model fine - tuning in the loading decision;
[0100] It should be noted that, in order to coordinate the loading decisions at different stages, the number of data movement operations pre-determined from the queue in the loading stage, that is, according to experience, the number of data related to model fine-tuning that is pre-determined to be moved to the GPU from the loading decisions at each stage.
[0101] Step L20, emit the number of movements into the communication stream in the GPU;
[0102] Step L30, obtain the data loading process status from the communication stream through a CUDA event.
[0103] It should be noted that the number of movements is emitted into the communication stream in the GPU, and then the CUDA event (CUDAEvent, a mechanism for recording timestamps at specific moments during GPU execution) is used to query the status of GPU loading, and the CPU-GPU synchronization mechanism is used to query whether the CUDA event is triggered. Specifically: insert the CUDA event before the last movement operation.
[0104] In this embodiment, the number of movements is emitted into the communication stream in the GPU, and the data loading process status is obtained from the communication stream through a CUDA event. By inserting the CUDA event into the communication stream, the key time points and status information of the data loading process can be recorded, ensuring that when the CUDA event is triggered, the data loading kernel is still executing, effectively hiding the kernel launch overhead, and can also achieve fine control and optimization of data loading during model fine-tuning, thereby improving the performance and fine-tuning efficiency of the model.
[0105] Step T20, according to the loading order, add or delete the loading decisions for the mixture-of-experts layer in the data loading queues for inter-stage loading, inter-model-layer loading, and inter-expert loading;
[0106] It should be noted that according to the previously determined loading order of the mixture-of-experts layer, operate on the loading decisions in the data loading queue. Specifically: add or delete the name of the model block to the queue. Add it to the queue if it needs to be loaded, and delete it from the queue if it doesn't need to be loaded.
[0107] Step T30, allocate priorities to the data loading queues according to the urgency of the task computing requirements for inter-stage loading, inter-model-layer loading, and inter-expert loading;
[0108] It should be noted that based on the urgency of the task computing requirements at different loading stages, different priorities are allocated to the queues where each stage is located. Specifically, the highest priority is allocated to inter-expert loading, the medium priority is allocated to inter-model-layer loading, and the low priority is allocated to inter-stage loading.
[0109] Step T40, schedule the loading decisions in the data loading queue with the highest priority to fine-tune the mixture-of-experts model.
[0110] It should be noted that during the data loading process, there are multiple data loading queues. Scheduling the loading decisions in the data loading queue with the highest priority to preferentially load the data in this stage, and then fine-tuning the mixture-of-experts model to improve its performance on specific tasks. Exemplarily, the highest priority is generally the loading decision in the inter-expert loading stage.
[0111] In this embodiment, determine the loading order of the mixture-of-experts layer according to the data loading process status of the hardware device (such as GPU) running the mixture-of-experts layer. Then, according to the determined loading order, add or delete the loading decisions in the data loading queues for inter-stage loading, inter-model-layer loading, and inter-expert loading. After that, allocate priorities to the data loading queues according to the computing requirements of the mixture-of-experts layer, load the mixture-of-experts layer with high computing requirements, and finally schedule the loading decisions in the data loading queue with the highest priority to fine-tune the mixture-of-experts model. Based on the hardware device status, loading order adjustment, priority allocation, and fine-tuning, by preferentially loading key data, experts with higher computational intensity can be processed faster, which can improve the running efficiency of the mixture-of-experts model.
[0112] Exemplarily, to help understand the implementation process of the fine-tuning method of the mixture-of-experts model obtained by combining this embodiment with the above first embodiment, please refer to Figure 5 , Figure 5 which provides a brief schematic diagram of the fine-tuning method of the mixture-of-experts model. Specifically:
[0113] First, the fine-tuning method of the mixture-of-experts model in this application divides the whole process into two parts. One part is the static part, namely the offline stage, and the other part is the runtime part, namely the fine-tuning stage;
[0114] In the offline stage, analyze the mixture-of-experts model, input data, fine-tuning parameters for the mixture-of-experts model, and information about the system (such as a computer) for fine-tuning the mixture-of-experts model through an analyzer device to obtain the memory analysis results, i.e., the memory usage, and the time analysis results, i.e., the execution time, of the mixture-of-experts model during simulated fine-tuning. Then, obtain the mapping relationship between the mixture-of-experts layer and the hardware device running the mixture-of-experts layer according to the memory analysis results, calculate the affinity of the hardware device according to the time analysis results, i.e., compare the execution times on the CPU and GPU under the same load, and store the mapping relationship and affinity in the form of a lookup table or other forms.
[0115] In the fine-tuning stage, data loading is carried out in stages. When loading between stages, the model blocks of the next pipeline stage to be loaded are determined according to historical popularity; when loading at the model layer, the predicted expert popularity is obtained from an additional expert popularity predictor, and popular experts are selected from the mixed expert model blocks to be loaded and unpopular experts are deleted according to the predicted expert popularity and affinity; when loading between experts, it is determined whether to load popular experts according to the real-time expert popularity generated by the gating operation. Additionally, each of these three stages corresponds to a loading queue, and the loading decisions of these three loading stages are coordinated according to CUDA events to schedule the loading execution and alleviate the situation of mutual interference.
[0116] It should be noted that the above examples are only for understanding this application and do not constitute a limitation on the fine-tuning method of the mixed expert model of this application. Based on this technical concept, more forms of simple transformations, such as the interaction and combination of various embodiments, are within the protection scope of this application.
[0117] This application also provides a fine-tuning device for a mixed expert model. Please refer to Figure 6 , the fine-tuning device for the mixed expert model includes:
[0118] A partitioning module 10, configured to partition the data loading process of loading the data related to the mixed expert layer and model fine-tuning into inter-stage loading, inter-model-layer loading, and inter-expert loading according to the overlapping relationship with the next calculation process under the determined mapping relationship;
[0119] A determination module 20, configured to determine the loading decisions for inter-stage loading, inter-model-layer loading, and inter-expert loading in combination with expert popularity and affinity, where the expert popularity is the frequency of different experts being activated in the mixed expert model, and the affinity is the affinity between the calculated load and the hardware device determined in the offline stage;
[0120] A scheduling module 30, configured to schedule the loading decisions to load the data related to the mixed expert layer and model fine-tuning during inter-stage loading, inter-model-layer loading, and inter-expert loading respectively.
[0121] Optionally, the determination module 20 is further configured to, during inter-stage loading, determine the loading priorities of the model blocks in the mixed expert layer according to the calculation intensities of the model blocks in the mixed expert layer, and the loading decision for inter-stage loading is to determine the model blocks to be loaded in the next pipeline stage based on the loading priorities and the historical expert popularity of the experts;
[0122] During inter-model-layer loading, predict the expert popularity of the experts in the mixed expert model blocks to be loaded by an expert popularity predictor, and the loading decision for inter-model-layer loading is to select popular experts from the mixed expert model blocks to be loaded and delete unpopular experts according to the expert popularity and the affinity;
[0123] When loading among experts, the loading decision for loading among experts is to determine whether to load popular experts according to the real-time expert popularity generated by the gating operation.
[0124] Optionally, the determination module 20 is further configured to initialize the to-be-trained expert popularity predictor by the weights of the gating operations of the mixture-of-experts layer where the expert popularity predictor is located;
[0125] Using the target fine-tuning dataset as the training dataset, training the initialized to-be-trained expert popularity predictor to obtain the expert popularity predictor.
[0126] Optionally, the scheduling module 30 is further configured to determine the loading order of the mixture-of-experts layer according to the data loading process status of the hardware device running the mixture-of-experts layer;
[0127] According to the loading order, add or delete the loading decision for the mixture-of-experts layer in the data loading queues of inter-stage loading, inter-model-layer loading, and loading among experts;
[0128] According to the urgency of the task calculation requirements for inter-stage loading, inter-model-layer loading, and loading among experts, assign priorities to the data loading queues;
[0129] Schedule the loading decision in the data loading queue with the highest priority to fine-tune the mixture-of-experts model.
[0130] Optionally, the scheduling module 30 is further configured to determine the number of data movements related to model fine-tuning in the loading decision in the data loading queues of inter-stage loading, inter-model-layer loading, and loading among experts;
[0131] Transmit the number of movements to the communication stream in the GPU;
[0132] Obtain the data loading process status from the communication stream through CUDA events.
[0133] Optionally, the partitioning module 10 is further configured to fine-tune the mixture-of-experts layer under preset conditions in the offline stage to obtain the memory usage and execution time of the mixture-of-experts layer during fine-tuning;
[0134] According to the memory usage and the execution time, respectively obtain the mapping relationship from the mixture-of-experts layer to the hardware device running the mixture-of-experts layer and the affinity of the hardware device.
[0135] The fine-tuning device of the mixture-of-experts model provided by the present application adopts the fine-tuning method of the mixture-of-experts model in the above embodiment, and can solve the technical problems of high video memory requirements for fine-tuning the mixture-of-experts model and low fine-tuning efficiency of the mixture-of-experts model under the offloading strategy. Compared with the prior art, the beneficial effects of the fine-tuning device of the mixture-of-experts model provided by the present application are the same as those of the fine-tuning method of the mixture-of-experts model provided by the above embodiment, and other technical features in the fine-tuning device of the mixture-of-experts model are the same as the features disclosed in the method of the above embodiment, and will not be elaborated herein.
[0136] The present application provides a fine-tuning device for a mixture-of-experts model. The fine-tuning device for a mixture-of-experts model includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the fine-tuning method of the mixture-of-experts model in the above first embodiment.
[0137] Refer to the following Figure 7 , which shows a schematic structural diagram of a fine-tuning device for a mixture-of-experts model suitable for implementing the embodiments of the present application. The fine-tuning device for a mixture-of-experts model in the embodiments of the present application may include, but is not limited to, mobile terminals such as mobile phones, laptop computers, PDAs (Personal Digital Assistants), PADs (Portable Application Descriptions: tablet computers), etc., and fixed terminals such as digital TVs, desktop computers, etc. Figure 7 The fine-tuning device for a mixture-of-experts model shown is merely an example and should not impose any limitation on the functions and usage scope of the embodiments of the present application.
[0138] As Figure 7As shown, the fine-tuning device of the mixture-of-experts model may include a processing device 1001 (such as a central processing unit, a graphics processing unit, etc.), which may perform various appropriate actions and processes according to the program stored in the read-only memory (ROM: Read Only Memory) 1002 or the program loaded from the storage device 1003 into the random access memory (RAM: Random Access Memory) 1004. In the RAM 1004, various programs and data required for the operation of the fine-tuning device of the mixture-of-experts model are also stored. The processing device 1001, the ROM 1002, and the RAM 1004 are connected to each other through a bus 1005. An input / output (I / O) interface 1006 is also connected to the bus. Generally, the following systems may be connected to the I / O interface 1006: an input device 1007 including, for example, a touch screen, a touchpad, a keyboard, a mouse, an image sensor, a microphone, an accelerometer, a gyroscope, etc.; an output device 1008 including, for example, a liquid crystal display (LCD: Liquid Crystal Display), a speaker, a vibrator, etc.; a storage device 1003 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 1009. The communication device 1009 may allow the fine-tuning device of the mixture-of-experts model to communicate with other devices wirelessly or wiredly to exchange data. Although the figure shows a fine-tuning device of the mixture-of-experts model having various systems, it should be understood that it is not required to implement or have all the shown systems. More or fewer systems may be implemented or had alternatively.
[0139] In particular, according to the embodiments disclosed in the present application, the processes described above with reference to the flowcharts may be implemented as computer software programs. For example, the embodiments disclosed in the present application include a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program contains program codes for executing the methods shown in the flowcharts. In such an embodiment, the computer program may be downloaded and installed from the network through the communication device, or installed from the storage device 1003, or installed from the ROM 1002. When the computer program is executed by the processing device 1001, the above-mentioned functions defined in the methods of the embodiments disclosed in the present application are executed.
[0140] The fine-tuning device of the mixture-of-experts model provided by the present application adopts the fine-tuning method of the mixture-of-experts model in the above embodiments, and can solve the technical problems of high video memory requirements for fine-tuning the mixture-of-experts model and low fine-tuning efficiency of the mixture-of-experts model under the offloading strategy. Compared with the prior art, the beneficial effects of the fine-tuning device of the mixture-of-experts model provided by the present application are the same as those of the fine-tuning method of the mixture-of-experts model provided by the above embodiments, and other technical features in the fine-tuning device of the mixture-of-experts model are the same as the features disclosed in the method of the previous embodiment, and will not be elaborated here.
[0141] It should be understood that each part disclosed in this application can be implemented by hardware, software, firmware, or a combination thereof. In the description of the above embodiments, specific features, structures, materials, or characteristics can be combined in a suitable manner in any one or more embodiments or examples.
[0142] As described above, only the specific embodiments of this application are provided, but the protection scope of this application is not limited thereto. Any person skilled in the art can easily think of changes or substitutions within the technical scope disclosed in this application, and all should be covered within the protection scope of this application. Therefore, the protection scope of this application should be subject to the protection scope of the claims.
[0143] This application provides a computer-readable storage medium with computer-readable program instructions (i.e., computer programs) stored thereon, and the computer-readable program instructions are used to execute the fine-tuning method of the mixture-of-experts model in the above embodiments.
[0144] The computer-readable storage medium provided by this application can be, for example, a USB flash drive, but is not limited to electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems or devices, or any combination of the above. More specific examples of the computer-readable storage medium may include, but are not limited to: electrical connections with one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM) or flash memory, optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the above. In this embodiment, the computer-readable storage medium can be any tangible medium that contains or stores a program, and this program can be used by or combined with an instruction execution system or device. The program code contained on the computer-readable storage medium can be transmitted by any appropriate medium, including but not limited to: wires, optical cables, RF (Radio Frequency), etc., or any suitable combination of the above.
[0145] The above computer-readable storage medium can be included in the fine-tuning device of the mixture-of-experts model; it can also exist separately without being assembled into the fine-tuning device of the mixture-of-experts model.
[0146] The above computer-readable storage medium carries one or more programs, which, when executed by a fine-tuning device of a mixture-of-experts model, cause the fine-tuning device of the mixture-of-experts model to: under a determined mapping relationship, divide the data loading process for loading data related to model fine-tuning of the mixture-of-experts layer into inter-stage loading, inter-model-layer loading, and inter-expert loading according to the overlapping relationship with the next computing process;
[0147] Combine the expert popularity and affinity to determine the loading decisions for inter-stage loading, inter-model-layer loading, and inter-expert loading, where the expert popularity is the frequency of different experts being activated in the mixture-of-experts model, and the affinity is the affinity between the computing load determined in the offline stage and the hardware device;
[0148] Schedule the loading decisions to load the data related to model fine-tuning of the mixture-of-experts layer in inter-stage loading, inter-model-layer loading, and inter-expert loading respectively.
[0149] Computer program code for performing the operations of this application can be written in one or more programming languages or combinations thereof. The above programming languages include object-oriented programming languages - such as Java, Smalltalk, C++, and also include conventional procedural programming languages - such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, executed as an independent software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer can be connected to the user's computer through any type of network - including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computer (for example, by using an Internet service provider to connect through the Internet).
[0150] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present application. In this regard, each block in the flowchart or block diagram may represent a module, a segment of a program, or a part of code that contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than that marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, as well as combinations of blocks in the block diagram and / or flowchart, may be implemented by a dedicated hardware-based system that performs the specified functions or operations, or may be implemented by a combination of dedicated hardware and computer instructions.
[0151] The modules described in the embodiments of the present application can be implemented in software or in hardware. Among them, the name of the module does not constitute a limitation on the unit itself in some cases.
[0152] The readable storage medium provided by the present application is a computer-readable storage medium that stores computer-readable program instructions (i.e., computer programs) for executing the above-mentioned fine-tuning method of the mixture of experts model, and can solve the technical problems of high video memory requirements for fine-tuning the mixture of experts model and low fine-tuning efficiency of the mixture of experts model under the offloading strategy. Compared with the prior art, the beneficial effects of the computer-readable storage medium provided by the present application are the same as those of the fine-tuning method of the mixture of experts model provided in the above embodiments, and will not be elaborated here.
[0153] The present application also provides a computer program product, including a computer program, and when the computer program is executed by a processor, it implements the steps of the above-mentioned fine-tuning method of the mixture of experts model.
[0154] The computer program product provided by the present application can solve the technical problems of high video memory requirements for fine-tuning the mixture of experts model and low fine-tuning efficiency of the mixture of experts model under the offloading strategy. Compared with the prior art, the beneficial effects of the computer program product provided by the present application are the same as those of the fine-tuning method of the mixture of experts model provided in the above embodiments, and will not be elaborated here.
[0155] The above are only some embodiments of the present application, and do not limit the patent scope of the present application. Any equivalent structural transformation made under the technical concept of the present application by using the content of the specification and drawings of the present application, or direct / indirect application in other related technical fields, is included in the patent protection scope of the present application.
Claims
1. A fine-tuning method for a mixture of experts model, characterized in that, The mixture-of-experts model consists of at least one layer of mixture-of-experts layers, and each mixture-of-experts layer consists of an attention model block and a mixture-of-experts model block. The mixture-of-experts model block consists of a gating operation and at least one expert. The fine-tuning method of the mixture-of-experts model includes: Under the determined mapping relationship, the data loading process for loading data related to the mixture-of-experts layer and model fine-tuning is divided into inter-stage loading, inter-model-layer loading, and inter-expert loading according to the overlapping relationship with the next calculation process. Combining expert popularity and affinity to determine the loading decisions for inter-stage loading, inter-model-layer loading, and inter-expert loading, where the expert popularity is the frequency of different experts being activated in the mixture-of-experts model, and the affinity is the affinity between the calculated load determined in the offline stage and the hardware device. Schedule the loading decisions to load the data related to the mixture-of-experts layer and model fine-tuning during inter-stage loading, inter-model-layer loading, and inter-expert loading respectively.
2. The fine-tuning method of the mixture of experts model according to claim 1, wherein The step of combining expert popularity and affinity to determine the loading decisions for inter-stage loading, inter-model-layer loading, and inter-expert loading includes: During inter-stage loading, according to the calculation intensity of each model block in the mixture-of-experts layer, determine the loading priority of each model block in the mixture-of-experts layer. The loading decision for inter-stage loading is to determine the model blocks to be loaded in the next pipeline stage based on the loading priority and the historical expert popularity of the experts. During inter-model-layer loading, predict the expert popularity of the experts in the mixture-of-experts model blocks to be loaded according to the expert popularity predictor. The loading decision for inter-model-layer loading is to select the popular experts and delete the unpopular experts from the mixture-of-experts model blocks to be loaded according to the expert popularity and the affinity. During inter-expert loading, the loading decision for inter-expert loading is to determine whether to load the popular experts according to the real-time expert popularity generated by the gating operation.
3. The fine-tuning method of the mixture of experts model according to claim 2, wherein Before the step of predicting the expert popularity of the experts in the mixture-of-experts model blocks to be loaded according to the expert popularity predictor during inter-model-layer loading, it includes: Initialize the to-be-trained expert popularity predictor through the weights of the gating operation of the mixture-of-experts layer where the expert popularity predictor is located. Use the target fine-tuning dataset as the training dataset to train the initialized to-be-trained expert popularity predictor to obtain the expert popularity predictor.
4. The fine-tuning method of the mixture of experts model according to claim 1, characterized in that, The step of scheduling the loading decisions includes: Determine the loading order of the mixture-of-experts layer according to the data loading process status of the hardware device running the mixture-of-experts layer. According to the loading order, add or delete the loading decisions of the mixture-of-experts layer in the data loading queues for inter-stage loading, inter-model-layer loading, and inter-expert loading. Allocate priorities to the data loading queues according to the urgency of the task calculation requirements for inter-stage loading, inter-model-layer loading, and inter-expert loading. Schedule the loading decisions in the data loading queue with the highest priority to fine-tune the mixture-of-experts model.
5. The fine-tuning method of the mixture of experts model according to claim 4, characterized in that Before the step of determining the loading order of the mixture-of-experts layer according to the data loading process status of the hardware device running the mixture-of-experts layer, it includes: In the data loading queues for inter-stage loading, inter-model-layer loading, and inter-expert loading, determine the number of data movements related to model fine-tuning in the loading decision. Transmit the number of movements to the communication stream in the GPU. Obtain the data loading process status from the communication stream through CUDA events.
6. The fine-tuning method of the mixture of experts model according to claim 1, characterized in that Before the step of dividing the data loading process for loading the data related to the mixture-of-experts layer and model fine-tuning into inter-stage loading, inter-model-layer loading, and inter-expert loading according to the overlapping relationship with the next calculation process under the determined mapping relationship, it includes: Under the preset conditions in the offline stage, fine-tune the mixture-of-experts layer to obtain the memory usage and execution time of the mixture-of-experts layer during fine-tuning. According to the memory usage and the execution time, respectively obtain the mapping relationship between the mixture-of-experts layer and the hardware device running the mixture-of-experts layer and the affinity between the computational load and the hardware device.
7. A fine-tuning device for a mixture of experts model, characterized in that, The fine-tuning device for the mixture-of-experts model includes: A division module, configured to divide the data loading process for loading the data related to the mixture-of-experts layer and model fine-tuning into inter-stage loading, inter-model-layer loading, and inter-expert loading according to the overlapping relationship with the next calculation process under the determined mapping relationship. A determination module, configured to determine the loading decisions for inter-stage loading, inter-model-layer loading, and inter-expert loading by combining expert popularity and affinity, where the mapping relationship and the affinity are the mapping relationship between the mixture-of-experts layer and the hardware device running the mixture-of-experts layer and the affinity between the computational load and the hardware device determined in the offline stage. A scheduling module, configured to schedule the loading decisions to respectively load the data related to the mixture-of-experts layer and model fine-tuning in inter-stage loading, inter-model-layer loading, and inter-expert loading.
8. A fine-tuning device for a mixture of experts model, characterized in that, The device includes: a memory, a processor, and a computer program stored on the memory and executable on the processor, and the computer program is configured to implement the steps of the fine-tuning method for the mixture-of-experts model according to any one of claims 1 to 6.
9. A storage medium, characterized in that, The storage medium is a computer-readable storage medium, and a computer program is stored on the storage medium, and when the computer program is executed by a processor, it implements the steps of the fine-tuning method for the mixture-of-experts model according to any one of claims 1 to 6.
10. A computer program product, characterized in that, The computer program product includes a computer program, and when the computer program is executed by a processor, it implements the steps of the fine-tuning method for the mixture-of-experts model according to any one of claims 1 to 6.
Citation Information
Cited By
Model processing method, electronic equipment, storage medium and program product
CN122285303A
Model processing method, electronic device, storage medium, and program product
CN122285303B