Load balancing method and device for hybrid expert model, and electronic equipment

By determining the load to be discarded based on the actual total load and the preset total load of each sub-model in the hybrid expert model, the model performance degradation caused by load imbalance is solved, and load balancing and computing power optimization are achieved.

CN120029788AActive Publication Date: 2025-05-23SHANGHAI XIYU TECHNOLOGY CO LTD

Patent Information

Application Number
CN202510512128.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-23
Publication Date
2025-05-23
Estimated Expiration
2045-04-23

AI Technical Summary

Technical Problem

In the hybrid expert model architecture, there are inherent threshold constraints on the parallel processing capabilities of each expert submodel, which leads to insufficient update of submodel parameters when the local load distribution is unbalanced, which reduces model representation capabilities, and causes computing power loss.

Method used

By obtaining the actual total load and the preset total load for each submodel, the load to be discarded is determined from the pending load and discarded to achieve load balancing. This method requires only one global communication synchronization, significantly reducing the number of load drops.

Benefits of technology

It effectively reduces the number of load discards, avoids the impact on the model characterization ability, and prevents waste of computing power, ensuring the stable performance of the hybrid expert model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120029788A_ABST
    Figure CN120029788A_ABST
Patent Text Reader

Abstract

The invention provides a load balancing method and device for a hybrid expert model and electronic equipment, and the load balancing method is applied to a training stage of the hybrid expert model, and comprises the steps: obtaining the number of to-be-processed loads of each sub-model under each calculation node; for each sub-model, summing the number of the to-be-processed loads of the sub-model under each computing node to obtain the total number of actual loads received by the sub-model, and determining the total number of preset loads of the sub-model under a plurality of computing nodes; and determining a load to be discarded from the loads to be processed based on the total number of actual loads and the total number of preset loads of each sub-model, and discarding the load to be discarded, so that each sub-model in the hybrid expert model achieves load balance. According to the method and the device, the number of discarded loads can be remarkably reduced only through one-time global communication synchronization, the characterization capability of the hybrid expert model is not affected, and meanwhile, the computing power waste of the hybrid expert model is also prevented.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of hybrid expert models, and in particular to a load balancing method, device and electronic device of a hybrid expert model. Background Art

[0002] In recent years, deep neural networks with Transformer architecture as the core have achieved breakthrough progress. In this context, distributed training technology (such as Mixture-of-Experts) has become a key technical path to achieve efficient training of ultra-large-scale models.

[0003] In the hybrid expert model architecture, the parallel processing capabilities of each expert sub-model have inherent threshold constraints. When the local load distribution is uneven, the conventional local discarding method will directly discard the load that exceeds the sub-model processing upper limit. If too much load is discarded on the training side, it will lead to insufficient update of sub-model parameters, reduce the model representation ability, and also cause more serious computing power loss. If the load distribution is more uneven, the number of discards will be greater, which will eventually lead to the performance degradation of the overall model.

[0004] Therefore, during the training phase, how to build a dynamic and adaptive expert load balancing mechanism to ensure that each sub-model obtains equivalent optimization intensity in distributed training has become a technical issue that cannot be underestimated. Summary of the invention

[0005] In view of this, the purpose of the present application is to provide a load balancing method, device and electronic device for a hybrid expert model, which determines the load to be discarded from the load to be processed according to the total actual load of each sub-model and the total preset load, so that each sub-model in the hybrid expert model can achieve load balancing. Only one global communication synchronization is required to significantly reduce the number of load discards, and the number of discarded loads is low, which will not affect the characterization ability of the hybrid expert model, and also prevents the waste of computing power of the hybrid expert model.

[0006] In a first aspect, an embodiment of the present application provides a load balancing method of a hybrid expert model, which is applied in a hybrid expert model training phase. The load balancing method includes: Get the number of pending loads for each sub-model under each computing node; For each sub-model, the number of loads to be processed by the sub-model at each computing node is summed to obtain the total number of actual loads received by the sub-model, and the total number of preset loads of the sub-model at multiple computing nodes is determined; Based on the total actual load of each sub-model and the total preset load, the load to be discarded is determined from the load to be processed, and the load to be discarded is discarded, so that each sub-model in the hybrid expert model achieves load balance.

[0007] Further, the determining the load to be discarded from the load to be processed based on the total actual load of each sub-model and the total preset load, includes: For each sub-model, determine whether the total actual load of the sub-model is greater than the total preset load of the sub-model; If yes, the sub-model is determined as the target sub-model, and the difference between the actual total load of the target sub-model and the preset total load is calculated; Based on the difference value corresponding to the target sub-model, determining the load to be discarded from the load to be processed under each computing node of the target sub-model; or, For each sub-model, determine whether the total actual load of the sub-model is greater than the total preset load of the sub-model; If yes, the sub-model is determined as the target sub-model, and a load difference ratio is calculated based on the difference between the actual total load of the target sub-model and the preset total load; wherein the load difference ratio is the ratio between the difference and the preset total load; Based on the load difference ratio corresponding to the target sub-model, the load to be discarded is determined from the load to be processed under each computing node of the target sub-model.

[0008] Further, the determining the load to be discarded from the load to be processed under each computing node of the target sub-model based on the difference corresponding to the target sub-model includes: Determining a first discard quantity based on a ratio of the difference to the number of computing nodes; For each computing node, based on the first discard quantity, determine the to-be-discarded load corresponding to the first discard quantity from the to-be-processed load of the target sub-model under the computing node.

[0009] Further, the determining the load to be discarded from the load to be processed under each computing node of the target sub-model based on the difference corresponding to the target sub-model includes: Based on the difference, determine the second discard quantity corresponding to each computing node; wherein the sum of the second discard quantity of each computing node is equal to the difference; For each computing node, based on the second discard quantity corresponding to the computing node, the load to be discarded corresponding to the second discard quantity is determined from the load to be processed under the computing node of the target sub-model.

[0010] Furthermore, the load to be discarded determined from the load to be processed under each computing node of the target sub-model based on the load difference ratio corresponding to the target sub-model includes: For each computing node, based on the product of the number of to-be-processed loads of the target sub-model under the computing node and the load difference ratio, determine a third discard quantity corresponding to the computing node; Based on the third discard quantity corresponding to the computing node, the to-be-discarded load corresponding to the third discard quantity is determined from the to-be-processed load of the target submodel under the computing node.

[0011] Furthermore, the load balancing method further includes: When it is detected that there is a sub-model in which the total number of actual loads is greater than the total number of preset loads, a score corresponding to each load to be processed is obtained; wherein the score represents the degree of matching between each load to be processed and the sub-model corresponding to the load to be processed, and a higher score indicates a higher degree of matching; The plurality of to-be-processed loads are sorted in order from low to high based on the scores, and a first preset number of to-be-processed loads are determined from the sorting results as to-be-discarded loads.

[0012] Furthermore, the load balancing method further includes: When it is detected that there is a sub-model whose total actual load is greater than the total preset load, a second preset number of loads to be processed is determined from the loads to be processed of multiple sub-models under multiple computing nodes as loads to be discarded.

[0013] Furthermore, obtaining the number of loads to be processed of each sub-model under each computing node includes: The hybrid expert model includes a communication layer, which sends an acquisition request to each computing node, and each computing node returns meta information including the amount of load to be processed to the communication layer.

[0014] In a second aspect, an embodiment of the present application further provides a load balancing device of a hybrid expert model, which is applied in a hybrid expert model training phase, and the load balancing device includes: A load quantity acquisition module is used to obtain the quantity of load to be processed of each sub-model under each computing node; A total load calculation module is used to sum the number of loads to be processed of each sub-model under each computing node, obtain the total actual load received by the sub-model, and determine the total preset load of the sub-model under multiple computing nodes; The first to-be-discarded load determination module is used to determine the to-be-discarded load from the to-be-processed load based on the total actual load of each sub-model and the total preset load, and discard the to-be-discarded load to achieve load balance for each sub-model in the hybrid expert model.

[0015] In a third aspect, an embodiment of the present application further provides an electronic device, comprising: a processor, a memory and a bus, wherein the memory stores machine-readable instructions executable by the processor, and when the electronic device is running, the processor and the memory communicate through the bus, and when the machine-readable instructions are executed by the processor, the steps of the load balancing method of the hybrid expert model as described above are performed.

[0016] The embodiments of the present application provide a load balancing method, device and electronic device for a hybrid expert model. The load balancing method is applied to the training stage of the hybrid expert model. First, the number of loads to be processed of each sub-model under each computing node is obtained; then, for each sub-model, the number of loads to be processed of the sub-model under each computing node is summed to obtain the total number of actual loads received by the sub-model, and the total number of preset loads of the sub-model under multiple computing nodes is determined; finally, based on the total number of actual loads and the total number of preset loads of each sub-model, the load to be discarded is determined from the loads to be processed, and the load to be discarded is discarded, so that each sub-model in the hybrid expert model achieves load balancing.

[0017] In the process of training the hybrid expert model, the present application counts the number of loads received by each sub-model, and determines the load to be discarded from the load to be processed according to the total actual load of each sub-model and the total preset load, so that each sub-model in the hybrid expert model can achieve load balancing. According to the load balancing method provided by the present application, only one global communication synchronization is required to significantly reduce the number of load discards, and the number of discarded loads is low, which will not affect the characterization ability of the hybrid expert model, and also prevents the waste of computing power of the hybrid expert model.

[0018] In order to make the above-mentioned objects, features and advantages of the present application more obvious and easy to understand, preferred embodiments are specifically cited below and described in detail with reference to the attached drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings required for use in the embodiments will be briefly introduced below. It should be understood that the following drawings only show certain embodiments of the present application and therefore should not be regarded as limiting the scope. For ordinary technicians in this field, other related drawings can be obtained based on these drawings without paying creative work.

[0020] Figure 1A flow chart of a load balancing method of a hybrid expert model provided in an embodiment of the present application; Figure 2 One of the structural schematic diagrams of a load balancing device of a hybrid expert model provided in an embodiment of the present application; Figure 3 A second structural diagram of a load balancing device of a hybrid expert model provided in an embodiment of the present application; Figure 4 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0021] To make the purpose, technical scheme and advantages of the embodiments of the present application clearer, the technical scheme in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all of the embodiments. The components of the embodiments of the present application usually described and shown in the drawings here can be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present application provided in the drawings is not intended to limit the scope of the application claimed for protection, but merely represents the selected embodiments of the present application. Based on the embodiments of the present application, each other embodiment obtained by those skilled in the art without making creative work belongs to the scope of protection of the present application.

[0022] First, the application scenarios to which the present application is applicable are introduced. The present application can be applied to the technical field of hybrid expert models.

[0023] In recent years, deep neural networks with Transformer architecture as the core have achieved breakthrough progress. In this context, distributed training technology (such as Mixture-of-Experts) has become a key technical path to achieve efficient training of ultra-large-scale models.

[0024] Research has found that in the hybrid expert model architecture, the parallel processing capabilities of each expert sub-model have inherent threshold constraints. When the local load distribution is uneven, the conventional local discarding method will directly discard the load that exceeds the sub-model processing upper limit. If too much load is discarded on the training side, it will lead to insufficient update of sub-model parameters, reduce the model representation ability, and also cause more serious computing power loss. If the load distribution is more uneven, the number of discards will be greater, which will eventually lead to the performance degradation of the overall model.

[0025] Therefore, during the training phase, how to build a dynamic and adaptive expert load balancing mechanism to ensure that each sub-model obtains equivalent optimization intensity in distributed training has become a technical issue that cannot be underestimated.

[0026] Based on this, an embodiment of the present application provides a load balancing method for a hybrid expert model, which can significantly reduce the number of load discards by only requiring one global communication synchronization. The number of discarded loads is low and will not affect the characterization capability of the hybrid expert model. At the same time, it also prevents the waste of computing power of the hybrid expert model.

[0027] See also Figure 1 , Figure 1 The load balancing method provided by the embodiment of the present application is applied in the training stage of the hybrid expert model. Figure 1 As shown in , the load balancing method includes: S101, obtaining the number of loads to be processed of each sub-model under each computing node.

[0028] It should be noted that in distributed computing, a computing node usually represents an independent process or computing unit, which is responsible for executing a part of the computing tasks of the hybrid expert model. Each computing node is composed of multiple sub-models to perform corresponding computing tasks. The sub-model is the expert model in the hybrid expert model. Each expert model is an independent neural network (such as a fully connected layer, Transformer module, etc.), which is responsible for processing inputs of specific patterns or features.

[0029] For the above step S101, in the specific implementation, the number of to-be-processed loads of each sub-model in the hybrid expert model under each computing node is obtained. Here, the hybrid expert model includes a gating module and softmax, which cooperate with each other to determine which sub-model each acquired to-be-processed load is allocated to; wherein the gating module is used to generate the allocation probability of each to-be-processed load for each sub-model, and softmax normalizes the above allocation probabilities, selects the model with the highest allocation probability to allocate the current to-be-processed load, and obtains the allocation results of all to-be-processed loads after all to-be-processed loads are allocated, forming the number of to-be-processed loads of each sub-model under each computing node in this application.

[0030] As an optional embodiment, with respect to the above step S101, obtaining the number of to-be-processed loads of each sub-model under each computing node includes: The hybrid expert model includes a communication layer, which sends an acquisition request to each computing node, and each computing node returns meta information including the amount of load to be processed to the communication layer.

[0031] Here, in the distributed implementation of the hybrid expert model, the interaction between the communication layer, computing nodes, and load meta-information is the key to achieving efficient task scheduling and load balancing. The communication layer coordinates the data transmission between different computing nodes in the hybrid expert model and obtains the number of loads to be processed by each computing node in real time. The meta-information includes the number of loads currently queued for processing in each computing node.

[0032] For the above steps, in the specific implementation, the communication layer in the hybrid expert model is responsible for obtaining the number of loads to be processed, and the communication layer sends an acquisition request to all computing nodes at a fixed time interval. After receiving the acquisition request, the computing node returns the meta information containing the number of loads to be processed to the communication layer based on the acquisition request. Here, the present application takes the hybrid expert model containing two computing nodes and five sub-models as an example for explanation. For example, each computing node returns the meta information containing the number of loads to be processed to the communication layer as [6, 6, 9, 15, 14], [13, 13, 9, 8, 7], and the number of loads to be processed of each sub-model under each computing node can be determined according to the above meta information. Specifically, the number of loads to be processed of the first sub-model under the first computing node is 6, and the number of loads to be processed of the first sub-model under the second computing node is 13; the number of loads to be processed of the second sub-model under the first computing node is 6, and the number of loads to be processed of the second sub-model under the second computing node is 13, and so on. Here, it should be noted that the above load distribution example is only for ease of understanding. The actual load is at least millions or even hundreds of millions.

[0033] S102, for each sub-model, summing up the number of loads to be processed of the sub-model under each computing node to obtain the total number of actual loads received by the sub-model, and determining the total number of preset loads of the sub-model under multiple computing nodes.

[0034] Here, the preset total load refers to the upper limit of the load that the sub-model can handle under multiple computing nodes, and the actual total load refers to the total number of loads actually received by the sub-model under multiple computing nodes.

[0035] With respect to the above step S102, in the specific implementation, for each sub-model, the number of pending loads of the sub-model under each computing node is summed to obtain the total number of actual loads received by the sub-model. Specifically, continuing the example in the above step S101, when the number of pending loads of the first sub-model under the first computing node is 6, and the number of pending loads of the first sub-model under the second computing node is 13, the two numbers are added together to determine that the total number of actual loads received by the sub-model is 19. Then, the number of preset loads of the sub-model under each computing node is determined, and the number of preset loads under each computing node is added together to obtain the total number of preset loads of the sub-model under multiple computing nodes. Here, as an example, when the number of preset loads of the sub-model under each computing node is 10, and there are two computing nodes, it can be determined that the total number of preset loads of the sub-model under multiple computing nodes is 20.

[0036] S103, based on the total actual load of each sub-model and the total preset load, determine the load to be discarded from the load to be processed, and discard the load to be discarded, so that each sub-model in the hybrid expert model achieves load balance.

[0037] Regarding the above step S103, in a specific implementation, after the total actual load and the total preset load of each sub-model are determined in the above step S102, based on the total actual load and the total preset load of each sub-model, the load to be discarded that needs to be discarded is determined from the load to be processed received by the sub-model. By discarding the determined load to be discarded, each sub-model in the hybrid expert model can achieve load balancing.

[0038] Here, in the load balancing method provided in the embodiment of the present application, two methods are provided to determine the load to be discarded from the load to be processed. The first method is to determine the load to be discarded based on the difference between the total number of actual loads and the total number of preset loads, and the second method is to determine the load to be discarded based on the load difference ratio between the total number of actual loads and the total number of preset loads.

[0039] First, the first method of determining the load to be discarded according to the difference between the total number of actual loads and the total number of preset loads is explained. Specifically, for the above step S103, the load to be discarded is determined from the load to be processed based on the total number of actual loads and the total number of preset loads of each sub-model, including: I: For each sub-model, determine whether the actual total load of the sub-model is greater than the preset total load of the sub-model.

[0040] II: If yes, the sub-model is determined as the target sub-model, and the difference between the actual total load of the target sub-model and the preset total load is calculated.

[0041] For the above steps I-II, in the specific implementation, for each sub-model, first determine whether the total actual load of the sub-model is greater than the total preset load of the sub-model. If the total actual load of the sub-model is less than or equal to the total preset load of the sub-model, it can be considered that the sub-model can handle the load it receives. If the total actual load of the sub-model is greater than the total preset load of the sub-model, it is considered that the load received by the sub-model exceeds its processing upper limit, and the load needs to be discarded to ensure the load balancing of the sub-model. At this time, the sub-model is used as the target sub-model, and the difference between the total actual load of the target sub-model and the total preset load is calculated. Continuing the above example, according to the number of loads to be processed of each sub-model under each computing node obtained, it is determined that the total actual load of the fourth sub-model is 23, and the total actual load of the fifth sub-model is 21, both of which exceed the total preset load of 20. Then the fourth sub-model is determined as the first target sub-model, and the fifth sub-model is determined as the second target sub-model. The difference corresponding to the first target sub-model is 3, and the difference corresponding to the second target sub-model is 1.

[0042] III: Based on the difference value corresponding to the target sub-model, determine the load to be discarded from the load to be processed under each computing node of the target sub-model.

[0043] For the above step III, in the specific implementation, after determining the difference between the actual total load of the target sub-model and the preset total load, based on the difference, determine the load to be discarded from the load to be processed under each computing node of the target sub-model.

[0044] According to the load balancing method provided by the present application, the load to be processed can be discarded on average according to the calculated difference. As an optional embodiment, for the above step III, in the specific implementation, the load to be discarded is determined from the load to be processed under each computing node of the target submodel based on the difference corresponding to the target submodel, including: i: Determine a first discard quantity based on the ratio of the difference to the quantity of computing nodes.

[0045] For the above step i, in the specific implementation, the first discarded number is determined based on the ratio between the difference value corresponding to the target sub-model and the number of multiple computing nodes. Here, continuing the example of the above steps, for the first target sub-model, the difference value is 3, and the ratio of the difference value 3 to the number of computing nodes 2 is determined to be 1.5. When it cannot be divided evenly, it is rounded up, that is, the first discarded number corresponding to the first target sub-model is determined to be 2. For the second target sub-model, the difference value is 1, and the ratio of the difference to the number of computing nodes is 0.5. When it cannot be divided evenly, it is rounded up, that is, the first discarded number corresponding to the second target sub-model is determined to be 1.

[0046] ii: For each computing node, based on the first discard quantity, determine the to-be-discarded load corresponding to the first discard quantity from the to-be-processed load of the target sub-model under the computing node.

[0047] For the above step ii, in the specific implementation, for each computing node, based on the first discard quantity determined in the above step i, the load to be discarded corresponding to the first discard quantity is determined from the load to be processed under the computing node of the target submodel. Here, continuing the above example, for the first target submodel, the first discard quantity is 2, so two loads to be discarded are determined under each computing node, that is, two loads to be processed are discarded from the load to be processed under each computing node of the first target submodel. For the second target submodel, the first discard quantity is 1, so one load to be discarded is determined under each computing node, that is, one load to be processed is discarded from the load to be processed under each computing node of the second target submodel. After discarding, the meta information of the number of loads to be processed of each computing node is [6, 6, 9, 13, 13], [13, 13, 9, 6, 6]. A total of 6 loads are discarded by this discarding method, which is acceptable in the training phase, and this method does not require additional communication.

[0048] According to the load balancing method provided by the present application, the load to be processed can also be selectively discarded or randomly discarded according to the calculated difference. As an optional embodiment, for the above step III, in the specific implementation, the load to be discarded is determined from the load to be processed under each computing node of the target submodel based on the difference corresponding to the target submodel, including: (1) Based on the difference, determine a second discard quantity corresponding to each computing node.

[0049] For the above step (1), in the specific implementation, based on the difference corresponding to the target sub-model, the second discard number corresponding to each computing node is determined, wherein the sum of the second discard numbers of each computing node is equal to the difference. Here, continuing the example of the above steps, for the first target sub-model, the difference is 3, then you can randomly choose to discard 3 loads in one computing node, or randomly select one computing node to discard 1 load, another computing node to discard two loads, or select a computing node with a larger number of loads to discard 3 loads, as long as the total number of discarded loads is equal to the difference 3. For the second target sub-model, the difference is 1, you can randomly choose to discard 1 load in one computing node, or select a computing node with a larger number of loads to discard 1 load, as long as the total number of discarded loads is equal to the difference 1.

[0050] (2) For each computing node, based on the second discard quantity corresponding to the computing node, determine the to-be-discarded load corresponding to the second discard quantity from the to-be-processed load of the target submodel under the computing node.

[0051] For the above step (2), in the specific implementation, for each computing node, based on the second discard quantity corresponding to the computing node, the load to be discarded corresponding to the second discard quantity is determined from the load to be processed under the computing node of the target sub-model. Here, continuing the example in the above steps, for the first target sub-model, 2 loads are discarded from the load to be processed under the first computing node, and 1 load is discarded from the load to be processed under the second computing node. For the second target sub-model, 1 load is discarded from the load to be processed under the first computing node. After discarding, the metadata of the number of loads to be processed of each computing node is [6, 6, 9, 13, 13], [13, 13, 9, 7, 7]. A total of 4 loads are discarded through this discarding method, and the number of discarded loads is the lowest.

[0052] Here, the second method of determining the load to be discarded according to the load difference ratio between the total number of actual loads and the total number of preset loads is explained. Specifically, for the above step S103, the load to be discarded is determined from the load to be processed based on the total number of actual loads and the total number of preset loads of each sub-model, including: A: For each sub-model, determine whether the actual total load of the sub-model is greater than the preset total load of the sub-model.

[0053] B: If yes, the sub-model is determined as the target sub-model, and the load difference ratio is calculated based on the difference between the actual total load of the target sub-model and the preset total load.

[0054] For the above steps AB, during the specific implementation, for each sub-model, determine whether the actual total load of the sub-model is greater than the preset total load of the sub-model. If it is greater, it is considered that the load received by the sub-model exceeds its processing upper limit. At this time, the sub-model is determined as the target sub-model, and the load difference ratio is calculated based on the difference between the actual total load of the target sub-model and the preset total load. Specifically, the load difference ratio is the ratio between the difference and the preset total load. Here, continuing the example in the above step II, the difference corresponding to the first target sub-model is 3, and the load difference ratio is 0.15, and the difference corresponding to the second target sub-model is 1, and the load difference ratio is 0.05.

[0055] C: based on the load difference ratio corresponding to the target sub-model, the load to be discarded is determined from the load to be processed under each computing node of the target sub-model.

[0056] Regarding the above step C, in the specific implementation, after determining the load difference ratio corresponding to the target sub-model, based on the load difference ratio, the load to be discarded is determined from the load to be processed under each computing node of the target sub-model.

[0057] According to the load balancing method provided by the present application, proportional discarding can be performed according to the calculated load difference ratio. As an optional embodiment, for the above step C, in the specific implementation, the load to be discarded determined from the load to be processed under each computing node of the target submodel based on the load difference ratio corresponding to the target submodel includes: a: For each computing node, based on the product of the number of to-be-processed loads of the target sub-model under the computing node and the load difference ratio, determine the third discard quantity corresponding to the computing node.

[0058] For the above step a, in the specific implementation, for each computing node, based on the product between the number of to-be-processed loads of the target sub-model under the computing node and the load difference ratio, determine the third discard quantity corresponding to the computing node. Here, continuing the example in the above steps, the load difference ratio of the first target sub-model is 0.15, so under the first computing node, 15*0.15=2.25, the third discard quantity rounded up is 3, and under the second computing node, 8*0.15=1.2, the third discard quantity rounded up is 2. For the second target sub-model, the third discard quantity is calculated similarly to be 1 under the first computing node and the second computing node.

[0059] b: Based on the third discard quantity corresponding to the computing node, determine the to-be-discarded load corresponding to the third discard quantity from the to-be-processed load of the target submodel under the computing node.

[0060] For the above step b, in the specific implementation, based on the third discard quantity corresponding to the computing node, the load to be discarded corresponding to the third discard quantity is determined from the load to be processed under the computing node of the target sub-model. Here, continuing the example in the above steps, for the first target sub-model, 3 loads are discarded under the first computing node, 2 loads are discarded under the second computing node, and for the second target sub-model, 1 load is discarded under each computing node. After discarding, the meta-information of the number of loads to be processed of each computing node is [6, 6, 9, 12, 13], [13, 13, 9, 6, 6]. A total of 7 loads are discarded through this discarding method, which is acceptable in the training phase, and this method does not require additional communication.

[0061] According to the load balancing method provided by the present application, when the number of loads received by the sub-model is too large, the loads with lower scores can be discarded. As an optional embodiment, the load balancing method provided by the embodiment of the present application also includes: When it is detected that there is a sub-model in which the total number of actual loads is greater than the total number of preset loads, the score corresponding to each load to be processed is obtained; based on the scores, multiple loads to be processed are sorted in order from low to high, and a first preset number of loads to be processed are determined from the sorting results as loads to be discarded.

[0062] The score represents the degree of matching between each load to be processed and the sub-model corresponding to the load to be processed, and a higher score indicates a higher degree of matching.

[0063] For the above two steps, in the specific implementation, when it is detected that there is a sub-model whose total actual load is greater than the total preset load, the score corresponding to each load to be processed is obtained. Here, the score corresponding to the load to be processed can be the score obtained by the expert scoring each load to be processed that needs to be processed, and this application does not make specific restrictions on this. Then, based on the obtained scores, the multiple loads to be processed are sorted from low to high, and the first preset number of loads to be processed are determined from the sorting results as the loads to be discarded.

[0064] According to the load balancing method provided by the present application, when the number of loads received by the sub-model is too large, the loads can be randomly discarded. As an optional embodiment, the load balancing method further includes: When it is detected that there is a sub-model whose total actual load is greater than the total preset load, a second preset number of loads to be processed is determined from the loads to be processed of multiple sub-models under multiple computing nodes as loads to be discarded.

[0065] Regarding the above steps, in the specific implementation, when it is detected that there is a sub-model whose actual total load is greater than the preset total load, a second preset number of loads to be processed is determined from the loads to be processed of multiple sub-models under multiple computing nodes as loads to be discarded.

[0066] The load balancing method of the hybrid expert model provided in the embodiment of the present application is applied to the training stage of the hybrid expert model. First, the number of loads to be processed of each sub-model under each computing node is obtained; then, for each sub-model, the number of loads to be processed of the sub-model under each computing node is summed to obtain the total actual load received by the sub-model, and the total preset load of the sub-model under multiple computing nodes is determined; finally, based on the total actual load and the total preset load of each sub-model, the load to be discarded is determined from the load to be processed, and the load to be discarded is discarded, so that each sub-model in the hybrid expert model achieves load balancing.

[0067] In the process of training the hybrid expert model, the present application counts the number of loads received by each sub-model, and determines the load to be discarded from the load to be processed according to the total actual load of each sub-model and the total preset load, so that each sub-model in the hybrid expert model can achieve load balancing. According to the load balancing method provided by the present application, only one global communication synchronization is required to significantly reduce the number of load discards, and the number of discarded loads is low, which will not affect the characterization ability of the hybrid expert model, and also prevents the waste of computing power of the hybrid expert model.

[0068] See also Figure 2 , Figure 3 , Figure 2 This is one of the structural diagrams of a load balancing device of a hybrid expert model provided in an embodiment of the present application. Figure 3 The second structural diagram of a load balancing device of a hybrid expert model provided in an embodiment of the present application. Figure 3 As shown in , the load balancing device 200 includes: The load quantity acquisition module 201 is used to acquire the quantity of the load to be processed of each sub-model under each computing node; The total load calculation module 202 is used to sum the number of loads to be processed of each sub-model at each computing node to obtain the total actual load received by the sub-model, and determine the total preset load of the sub-model at multiple computing nodes; The first to-be-discarded load determination module 203 is used to determine the to-be-discarded load from the to-be-processed load based on the total actual load of each sub-model and the total preset load, and to discard the to-be-discarded load so that each sub-model in the hybrid expert model achieves load balance.

[0069] Further, when the first to-be-discarded load determination module 203 is used to determine the to-be-discarded load from the to-be-processed load based on the total actual load of each sub-model and the total preset load, the first to-be-discarded load determination module 203 is further used to: For each sub-model, determine whether the total actual load of the sub-model is greater than the total preset load of the sub-model; If yes, the sub-model is determined as the target sub-model, and the difference between the actual total load of the target sub-model and the preset total load is calculated; Based on the difference value corresponding to the target sub-model, determining the load to be discarded from the load to be processed under each computing node of the target sub-model; or, For each sub-model, determine whether the total actual load of the sub-model is greater than the total preset load of the sub-model; If yes, the sub-model is determined as the target sub-model, and a load difference ratio is calculated based on the difference between the actual total load of the target sub-model and the preset total load; wherein the load difference ratio is the ratio between the difference and the preset total load; Based on the load difference ratio corresponding to the target sub-model, the load to be discarded is determined from the load to be processed under each computing node of the target sub-model.

[0070] Further, when the first to-be-discarded load determination module 203 is used to determine the to-be-discarded load from the to-be-processed load under each computing node of the target submodel based on the difference corresponding to the target submodel, the first to-be-discarded load determination module 203 is also used to: Determining a first discard quantity based on a ratio of the difference to the number of computing nodes; For each computing node, based on the first discard quantity, determine the to-be-discarded load corresponding to the first discard quantity from the to-be-processed load of the target sub-model under the computing node.

[0071] Further, when the first to-be-discarded load determination module 203 is used to determine the to-be-discarded load from the to-be-processed load under each computing node of the target submodel based on the difference corresponding to the target submodel, the first to-be-discarded load determination module 203 is also used to: Based on the difference, determine the second discard quantity corresponding to each computing node; wherein the sum of the second discard quantity of each computing node is equal to the difference; For each computing node, based on the second discard quantity corresponding to the computing node, the load to be discarded corresponding to the second discard quantity is determined from the load to be processed under the computing node of the target sub-model.

[0072] Further, when the first to-be-discarded load determination module 203 is used to determine the to-be-discarded load from the to-be-processed load under each computing node of the target submodel based on the load difference ratio corresponding to the target submodel, the first to-be-discarded load determination module 203 is further used to: For each computing node, based on the product of the number of to-be-processed loads of the target sub-model under the computing node and the load difference ratio, determine a third discard quantity corresponding to the computing node; Based on the third discard quantity corresponding to the computing node, the to-be-discarded load corresponding to the third discard quantity is determined from the to-be-processed load of the target submodel under the computing node.

[0073] Further, such as Figure 3 As shown, the load balancing device 200 further includes a second to-be-discarded load determination module 204, and the second to-be-discarded load determination module 204 is used to: When it is detected that there is a sub-model in which the total number of actual loads is greater than the total number of preset loads, a score corresponding to each load to be processed is obtained; wherein the score represents the degree of matching between each load to be processed and the sub-model corresponding to the load to be processed, and a higher score indicates a higher degree of matching; The plurality of to-be-processed loads are sorted in order from low to high based on the scores, and a first preset number of to-be-processed loads are determined from the sorting results as to-be-discarded loads.

[0074] Further, such as Figure 3 As shown, the load balancing device 200 further includes a third to-be-discarded load determination module 205, and the second to-be-discarded load determination module 205 is used to: When it is detected that there is a sub-model whose total actual load is greater than the total preset load, a second preset number of loads to be processed is determined from the loads to be processed of multiple sub-models under multiple computing nodes as loads to be discarded.

[0075] Furthermore, when the load quantity acquisition module 201 is used to acquire the quantity of loads to be processed of each sub-model under each computing node, the load quantity acquisition module 201 is also used to: The hybrid expert model includes a communication layer, which sends an acquisition request to each computing node, and each computing node returns meta information including the amount of load to be processed to the communication layer. See also Figure 4 , Figure 4 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present application. Figure 4 As shown in , the electronic device 400 includes a processor 410 , a memory 420 and a bus 430 .

[0076] The memory 420 stores machine-readable instructions executable by the processor 410. When the electronic device 400 is running, the processor 410 communicates with the memory 420 via the bus 430. When the machine-readable instructions are executed by the processor 410, the above-mentioned Figure 1 The steps of the load balancing method of the hybrid expert model in the method embodiment shown, the specific implementation method can be found in the method embodiment, and will not be repeated here.

[0077] The present application also provides a computer-readable storage medium on which a computer program is stored. When the computer program is executed by a processor, the computer program can execute the above-mentioned Figure 1The steps of the load balancing method of the hybrid expert model in the method embodiment shown, the specific implementation method can be found in the method embodiment, and will not be repeated here.

[0078] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.

[0079] In the several embodiments provided in the present application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. The device embodiments described above are merely schematic. For example, the division of the units is only a logical function division. There may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some communication interfaces, and the indirect coupling or communication connection of devices or units can be electrical, mechanical or other forms.

[0080] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0081] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.

[0082] If the functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a non-volatile computer-readable storage medium that is executable by a processor. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium, including several instructions for a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM), random access memory (RAM), disk or optical disk, and other media that can store program codes.

[0083] Finally, it should be noted that the above-described embodiments are only specific implementation methods of the present application, which are used to illustrate the technical solutions of the present application, rather than to limit them. The protection scope of the present application is not limited thereto. Although the present application is described in detail with reference to the above-mentioned embodiments, ordinary technicians in the field should understand that any technician familiar with the technical field can still modify the technical solutions recorded in the above-mentioned embodiments within the technical scope disclosed in the present application, or can easily think of changes, or make equivalent replacements for some of the technical features therein; and these modifications, changes or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application, and should be included in the protection scope of the present application. Therefore, the protection scope of the present application shall be based on the protection scope of the claims.

Claims

1. A load balancing method for a hybrid expert model, applied in the hybrid expert model training phase, characterized in that: The load balancing method comprises: Get the number of pending loads for each sub-model under each computing node; For each sub-model, the number of loads to be processed by the sub-model at each computing node is summed to obtain the total number of actual loads received by the sub-model, and the total number of preset loads of the sub-model at multiple computing nodes is determined; Based on the total actual load of each sub-model and the total preset load, the load to be discarded is determined from the load to be processed, and the load to be discarded is discarded, so that each sub-model in the hybrid expert model achieves load balance.

2. The load balancing method according to claim 1, characterized in that: The step of determining the load to be discarded from the load to be processed based on the total actual load of each sub-model and the total preset load, comprises: For each sub-model, determine whether the total actual load of the sub-model is greater than the total preset load of the sub-model; If yes, the sub-model is determined as the target sub-model, and the difference between the actual total load of the target sub-model and the preset total load is calculated; Based on the difference value corresponding to the target sub-model, determining the load to be discarded from the load to be processed under each computing node of the target sub-model; or, For each sub-model, determine whether the total actual load of the sub-model is greater than the total preset load of the sub-model; If yes, the sub-model is determined as the target sub-model, and a load difference ratio is calculated based on the difference between the actual total load of the target sub-model and the preset total load; wherein the load difference ratio is the ratio between the difference and the preset total load; Based on the load difference ratio corresponding to the target sub-model, the load to be discarded is determined from the load to be processed under each computing node of the target sub-model.

3. The load balancing method according to claim 2, characterized in that: The step of determining the load to be discarded from the load to be processed under each computing node of the target sub-model based on the difference corresponding to the target sub-model includes: Determining a first discard quantity based on a ratio of the difference to the number of computing nodes; For each computing node, based on the first discard quantity, determine the to-be-discarded load corresponding to the first discard quantity from the to-be-processed load of the target sub-model under the computing node.

4. The load balancing method according to claim 2, characterized in that: The step of determining the load to be discarded from the load to be processed under each computing node of the target sub-model based on the difference corresponding to the target sub-model includes: Based on the difference, determine the second discard quantity corresponding to each computing node; wherein the sum of the second discard quantity of each computing node is equal to the difference; For each computing node, based on the second discard quantity corresponding to the computing node, the load to be discarded corresponding to the second discard quantity is determined from the load to be processed under the computing node of the target sub-model.

5. The load balancing method according to claim 2, characterized in that: The load to be discarded determined from the load to be processed under each computing node of the target sub-model based on the load difference ratio corresponding to the target sub-model includes: For each computing node, based on the product of the number of to-be-processed loads of the target sub-model under the computing node and the load difference ratio, determine a third discard quantity corresponding to the computing node; Based on the third discard quantity corresponding to the computing node, the to-be-discarded load corresponding to the third discard quantity is determined from the to-be-processed load of the target submodel under the computing node.

6. The load balancing method according to claim 1, characterized in that: The load balancing method further includes: When it is detected that there is a sub-model in which the total number of actual loads is greater than the total number of preset loads, a score corresponding to each load to be processed is obtained; wherein the score represents the degree of matching between each load to be processed and the sub-model corresponding to the load to be processed, and a higher score indicates a higher degree of matching; The plurality of to-be-processed loads are sorted in order from low to high based on the scores, and a first preset number of to-be-processed loads are determined from the sorting results as to-be-discarded loads.

7. The load balancing method according to claim 1, characterized in that: The load balancing method further includes: When it is detected that there is a sub-model whose total actual load is greater than the total preset load, a second preset number of loads to be processed is determined from the loads to be processed of multiple sub-models under multiple computing nodes as loads to be discarded.

8. The load balancing method according to claim 1, characterized in that: The obtaining of the number of loads to be processed of each sub-model under each computing node includes: The hybrid expert model includes a communication layer, which sends an acquisition request to each computing node, and each computing node returns meta information including the amount of load to be processed to the communication layer.

9. A load balancing device for a hybrid expert model, applied in the hybrid expert model training phase, characterized in that: The load balancing device comprises: A load quantity acquisition module is used to obtain the quantity of load to be processed of each sub-model under each computing node; A total load calculation module is used to sum the number of loads to be processed of each sub-model under each computing node, obtain the total actual load received by the sub-model, and determine the total preset load of the sub-model under multiple computing nodes; The first to-be-discarded load determination module is used to determine the to-be-discarded load from the to-be-processed load based on the total actual load of each sub-model and the total preset load, and discard the to-be-discarded load to achieve load balance for each sub-model in the hybrid expert model.

10. An electronic device, characterized in that: include: A processor, a memory and a bus, wherein the memory stores machine-readable instructions executable by the processor, and when the electronic device is running, the processor and the memory communicate through the bus, and the machine-readable instructions are executed by the processor to execute the steps of the load balancing method of the hybrid expert model as described in any one of claims 1 to 8.

Citation Information

Patent Citations

  • Data processing method and device

    CN110209492A

  • Load balancing method and device for distributed model training

    CN115629879A

  • Data processing method and device, computer equipment, storage medium and program product

    CN117764116A

  • Hybrid expert model distributed training method based on dynamic load balancing

    CN118838711A

  • Data balanced distribution method based on MOE scene, electronic equipment and storage medium

    CN118966275A

Cited By

  • Load adjustment method and device for container instance and readable storage medium

    CN120560785A