Phased Hybrid Parallel Inference Method and System for MoE Sparse Large Model
By adopting different parallel strategies for the pre-filling and decoding stages in the MoE sparse big model, the problem of poor adaptability in the prior art is solved, and the inference efficiency and communication performance are improved.
Patent Information
- Application Number
- CN202510542935.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-28
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2045-04-28
AI Technical Summary
The existing MoE sparse big model inference technology has poor adaptability in parallel strategies in the pre-filling stage and the decoding stage, resulting in large communication overhead and affecting the inference efficiency.
The staged mixed parallel inference method is adopted, and different parallel strategies are adopted for the pre-filling stage and the decoding stage respectively. The pre-filling stage uses expert parallel strategies, and the decoding stage uses tensor parallel strategies, and hides communication overhead through layer-by-layer conversion to avoid the increase in device resource occupancy.
It improves the adaptability and efficiency of the MoE sparse big model inference process, reduces the communication overhead between devices, and optimizes the performance of the pre-filling and decoding stages.
Smart Images

Figure CN120069097B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of computing systems based on specific computing models, and particularly to a phased hybrid parallel inference method and system for MoE sparse large models. Background Art
[0002] The MoE sparse large model (which can also be called: MoE hybrid expert sparse large model) refers to introducing a sparse activation mechanism on the basis of the mixture-of-expert model MoE (Mixture-of-Expert). Compared with the mixture-of-expert model MoE whose computational complexity is close to that of a dense model, the MoE sparse large model replaces the multilayer perceptron MLP (MultilayerPerceptron) layer in the mixture-of-expert model MoE with a gating function and multiple experts, while the multi-head attention layer is the same as that of the mixture-of-expert model MoE. In the MoE sparse large model, each character only activates some experts for calculation, so as to achieve a sublinear growth of the computational cost with the model capacity, and achieve the effect of rapid expansion of model parameters. Compared with traditional dense large models, under the same number of parameters, the MoE sparse large model can significantly reduce the computational cost. The computational process of large model inference can be divided into two processes, namely the prefill stage (Prefill) and the decoding stage (Decoding).
[0003] Currently, the parallel strategy adopted by the existing distributed inference method for MoE sparse large model inference has poor adaptability to the prefill stage and the decoding stage. For example, DeepSpeed-MoE proposes a hybrid parallel strategy of "data parallelism + tensor parallelism + expert parallelism", and restricts expert parallelism within a node, so as to improve the communication efficiency of all-to-all communication by using the high bandwidth within the node. Since DeepSpeed-MoE does not distinguish the parallel strategies for the prefill stage and the decoding stage, and the prefill stage processes a large amount of data, while the decoding stage has a low computational density and processes a small amount of data, it will cause a large number of all-gather and reduce-scatter communication operations on data when using the tensor parallel strategy in the prefill stage, thereby increasing the communication overhead, which is more serious when dealing with a large batch of data; at the same time, it will also make the all-to-all communication overhead of the decoding stage more significant when using the expert parallel strategy.
[0004] Therefore, there is a need to design a MoE sparse large model inference method that can solve the problems of poor adaptability of the parallel strategy adopted by the existing MoE sparse large model inference technology to the prefill stage and the decoding stage and large communication overhead. Summary of the Invention
[0005] In view of this, the embodiments of the present application provide a phased hybrid parallel inference method and system for MoE sparse large models to eliminate or improve one or more defects existing in the prior art.
[0006] One aspect of the present application provides a phased hybrid parallel inference method for MoE sparse large models, including:
[0007] In the pre-fill stage of data inference based on the MoE sparse large model, controlling each architecture layer of the MoE sparse large model to execute the first step layer by layer; wherein, the first step includes: based on the multi-head attention layer model parameters and gating functions in each device, obtaining the respective expert numbers of each character corresponding to the prompt data sequence currently located at its initial position in its respective device, and adding the first mixture-of-experts layer model parameters operating based on the tensor parallel strategy to each device; then performing expert parallel calculation on each character according to the respective expert numbers of each character and the respective second mixture-of-experts layer model parameters operating based on the expert parallel strategy in each device; and then restoring each character after expert parallel calculation to its initial position in its respective device and releasing the second mixture-of-experts layer model parameters in each device;
[0008] Sending the predicted characters output by the last architecture layer of the MoE sparse large model to the multi-head attention layer in the first architecture layer of the MoE sparse large model for performing the decoding stage of data inference of the MoE sparse large model according to the predicted characters, the multi-head attention layer model parameters in each device, the gating function, and the first mixture-of-experts layer model parameters operating based on the tensor parallel strategy in each device.
[0009] In some embodiments of the present application, the phased hybrid parallel inference method for the MoE sparse large model further includes:
[0010] In each iteration round of the decoding stage of data inference based on the MoE sparse large model, controlling each architecture layer of the MoE sparse large model to execute the second step layer by layer, and sending the predicted characters output by the last architecture layer of the MoE sparse large model in each iteration round to the multi-head attention layer in the first architecture layer of the MoE sparse large model in a full-swap communication manner;
[0011] Wherein, the second step includes:
[0012] In the current iteration round, based on the multi-head attention layer model parameters and gating functions in each of the devices, obtain the expert numbers of each character corresponding to the latest data sequence, where each character is currently located at its respective initial position in its own device; wherein, if the current iteration round is the first round, the latest data sequence includes the prompt data sequence cached in the pre-fill stage and the predicted character output by the last architecture layer of the MoE sparse large model in the pre-fill stage; if the current iteration round is not the first round, the latest data sequence includes the prompt data sequence, the predicted character output in the pre-fill stage, and the predicted characters output in the previous historical iteration rounds before the current iteration round;
[0013] Collect each of the characters to each device in an all-gather communication manner, and according to the expert number of each character, enable each device to perform tensor parallel computing on the collected characters according to the first mixture-of-experts layer model parameters corresponding to each of them and running based on the tensor parallel strategy;
[0014] Restore each of the characters after tensor parallel computing to their respective initial positions in the devices in a reduce-scatter communication manner.
[0015] In some embodiments of the present application, the obtaining of the expert numbers of each character corresponding to the prompt data sequence, where each character is currently located at its respective initial position in its own device, based on the multi-head attention layer model parameters and gating functions in each device, includes:
[0016] Divide the prompt data sequence into multiple first data subsequences evenly according to the total number of each of the devices;
[0017] Input each of the first data subsequences into each device one-to-one, so that each device performs multi-head attention computing on the first data subsequence received by it respectively based on the multi-head attention layer model parameters stored in it, to obtain each character corresponding to the first data subsequence received by it respectively, and determine the expert number of each character on each device respectively based on the gating function;
[0018] Correspondingly, the obtaining of the expert numbers of each character corresponding to the latest data sequence, where each character is currently located at its respective initial position in its own device, based on the multi-head attention layer model parameters and gating functions in each device, includes:
[0019] Divide the latest data sequence into multiple second data subsequences evenly according to the total number of each of the devices;
[0020] Input each of the second data subsequences one-to-one into each device, so that each device performs multi-head attention calculation on the received second data subsequence based on the multi-head attention layer model parameters stored in each device respectively, to obtain each character corresponding to the received second data subsequence respectively, and collect all the characters to each device in a all-gather communication manner, so that each device determines the expert number of each character on each device based on the gating function respectively.
[0021] In some embodiments of the present application, the adding the first mixture-of-experts layer model parameters operating based on the tensor parallel strategy to each device includes:
[0022] If the time required to add the first mixture-of-experts layer model parameters operating based on the tensor parallel strategy to each device is shorter than or equal to the time required to obtain the expert number of each character currently located at the initial position of its respective device corresponding to the prompt data sequence, then while obtaining the expert number of each character currently located at the initial position of its respective device corresponding to the prompt data sequence, add the first mixture-of-experts layer model parameters operating based on the tensor parallel strategy to each device.
[0023] In some embodiments of the present application, before restoring each of the characters after the expert parallel calculation to their respective initial positions of the devices in the first step, it further includes:
[0024] If the time required to add the first mixture-of-experts layer model parameters operating based on the tensor parallel strategy to each device is longer than the time required to obtain the expert number of each character currently located at the initial position of its respective device corresponding to the prompt data sequence, then divide the first mixture-of-experts layer model parameters corresponding to each device operating based on the tensor parallel strategy into a first part of parameters and a second part of parameters;
[0025] Based on the multi-head attention layer model parameters and the gating function in each device, obtain the expert number of each character currently located at the initial position of its respective device corresponding to the prompt data sequence, and add the first part of parameters to each device at the same time;
[0026] Then perform expert parallel calculation on each character according to the expert number of each character and the second mixture-of-experts layer model parameters corresponding to each device operating based on the expert parallel strategy, and add the second part of parameters to each device at the same time.
[0027] In some embodiments of the present application, the expert parallel computing of each of the characters according to the respective expert numbers of each character and the model parameters of the second hybrid expert layer running based on the expert parallel strategy in each of the devices includes:
[0028] According to the respective expert numbers of each character, aiming at equalizing the number of characters processed by each of the devices during the expert parallel computing process and preferentially arranging the characters with the smallest expert numbers within the global scope locally in each of the devices, perform load balancing processing on each of the characters in each of the devices based on the greedy algorithm;
[0029] In a full exchange communication manner, route each of the characters after the load balancing processing to each of the devices where their respective expert numbers are currently located, so that each of the devices performs expert parallel computing on the characters received by each of them according to the model parameters of the second hybrid expert layer running based on the expert parallel strategy corresponding to each of them.
[0030] In some embodiments of the present application, the load balancing processing on each of the characters in each of the devices based on the greedy algorithm according to the respective expert numbers of each character, aiming at equalizing the number of characters processed by each of the devices and preferentially arranging the characters with the smallest expert numbers within the global scope locally in each of the devices, includes:
[0031] Control each of the devices to sort the local characters in ascending order of the expert numbers respectively to determine the distribution information of the local characters in each of the devices in each expert, and obtain the distribution information of the local characters in each of the other devices in each expert in a all-gather communication manner to determine the global distribution information of each character in each expert;
[0032] Control each of the devices to determine the character distribution information after each of the devices routes the characters to the respective devices where their expert numbers are currently located respectively according to the global distribution information of each character in each expert;
[0033] Control each of the devices to determine the start address and offset of the characters sent by each of the devices and the start address and offset of the characters received by each of the devices according to the distribution information of the local characters in each of the devices in each expert and the character distribution information after each of the devices routes the characters to the respective devices where their expert numbers are currently located;
[0034] Take the start address and offset of the characters sent by each of the devices, the start address and offset of the characters received by each of the devices, the send data buffer, and the receive data buffer as input parameters, and input them into the full exchange communication function corresponding to the full exchange communication manner.
[0035] In some embodiments of the present application, after determining the start addresses and offsets of characters sent by each of the devices and the start addresses and offsets of characters received by each of the devices, the following steps are further included:
[0036] If it is determined that there is a device that misses the expert after full exchange communication based on the offsets of characters sent by each of the devices and the offsets of characters received by each of the devices, before the device performs expert parallel computing on each of the characters based on the second hybrid expert layer model parameters corresponding to the missed expert, load the second hybrid expert layer model parameters corresponding to the missed expert from the processor memory to the device.
[0037] Another aspect of the present application provides a phased hybrid parallel inference device for a MoE sparse large model, including:
[0038] A pre-fill stage execution module, configured to control each architecture layer of the MoE sparse large model to sequentially execute the first step in the pre-fill stage of data inference based on the MoE sparse large model; wherein, the first step includes: based on the multi-head attention layer model parameters and gating functions in each device, obtain the expert numbers of each character corresponding to the prompt data sequence that are currently at their respective initial positions in their respective devices, and simultaneously add the first hybrid expert layer model parameters operating based on the tensor parallel strategy to each of the devices; then perform expert parallel computing on each of the characters according to the expert numbers of each character and the second hybrid expert layer model parameters corresponding to each of the devices operating based on the expert parallel strategy; then restore each of the characters after expert parallel computing to their respective device initial positions and release the second hybrid expert layer model parameters in each of the devices;
[0039] A stage transition execution module, configured to send the predicted characters output by the last architecture layer of the MoE sparse large model to the multi-head attention layer in the first architecture layer of the MoE sparse large model, for performing the decoding stage of data inference on the MoE sparse large model according to the predicted characters, the multi-head attention layer model parameters in each device, the gating function, and the first hybrid expert layer model parameters operating based on the tensor parallel strategy in each of the devices.
[0040] In some embodiments of the present application, the phased hybrid parallel inference device for the MoE sparse large model further includes:
[0041] The decoding stage execution module is used to control each architecture layer of the MoE sparse large model to execute the second step layer by layer in each iteration round during the decoding stage of data inference based on the MoE sparse large model, and send the predicted characters output by the last architecture layer of the MoE sparse large model in each iteration round to the multi-head attention layer in the first architecture layer of the MoE sparse large model in a full exchange communication manner;
[0042] Wherein, the second step includes:
[0043] In the current iteration round, based on the multi-head attention layer model parameters and gating functions in each device, obtain the expert numbers of each character corresponding to the latest data sequence currently located at their respective initial positions in their respective devices; wherein, if the current iteration round is the first round, the latest data sequence includes the prompt data sequence cached in the pre-filling stage and the predicted characters output by the last architecture layer of the MoE sparse large model in the pre-filling stage; if the current iteration round is not the first round, the latest data sequence includes the prompt data sequence, the predicted characters output in the pre-filling stage, and the predicted characters output in the previous historical iteration rounds before the current iteration round;
[0044] According to the expert numbers of each character, route each character to the respective devices where their expert numbers are currently located in a full exchange communication manner, so that each device respectively performs tensor parallel computing on the characters received by it according to the first mixture-of-experts layer model parameters operating based on the tensor parallel strategy;
[0045] Restore each character after tensor parallel computing to its respective initial position in the device in a reduction and distribution communication manner.
[0046] The third aspect of the present application provides a phased hybrid parallel inference system for a MoE sparse large model, including: a scheduling device and each device respectively communicatively connected to the scheduling device; the scheduling device includes a memory, a processor, and a computer program stored on the memory and executable on the processor, and when the processor executes the computer program, it implements the phased hybrid parallel inference method of the MoE sparse large model described in the foregoing first aspect.
[0047] The fourth aspect of the present application provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, and when the processor executes the computer program, it implements the phased hybrid parallel inference method of the MoE sparse large model.
[0048] The fifth aspect of the present application provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it implements the phased hybrid parallel inference method of the MoE sparse large model described above.
[0049] The sixth aspect of the present application provides a computer program product, including a computer program, and when the computer program is executed by a processor, it implements the phased hybrid parallel inference method of the MoE sparse large model described above.
[0050] For the pre-filling stage of data inference based on the MoE sparse large model in the phased hybrid parallel inference method of the MoE sparse large model provided by the present application, control each architecture layer of the MoE sparse large model to execute the first step layer by layer; wherein, the first step includes: based on the multi-head attention layer model parameters and gating function in each device, obtain the respective expert numbers of each character corresponding to the prompt data sequence currently located at the initial position of its own device, and at the same time add the first hybrid expert layer model parameters operating based on the tensor parallel strategy to each device; then perform expert parallel calculation on each character according to the respective expert numbers of each character and the second hybrid expert layer model parameters operating based on the expert parallel strategy in each device; then restore each character after expert parallel calculation to its initial position in the respective device and release the second hybrid expert layer model parameters in each device; send the predicted character output by the last architecture layer of the MoE sparse large model to the multi-head attention layer in the first architecture layer of the MoE sparse large model for executing the decoding stage of the MoE sparse large model data inference according to the predicted character, the multi-head attention layer model parameters in each device, the gating function, and the first hybrid expert layer model parameters operating based on the tensor parallel strategy in each device; that is to say, for the pre-filling stage with a high data volume, the hybrid expert layer adopts the expert parallel strategy with relatively lower communication volume to effectively reduce the communication overhead between devices in the pre-filling stage; and for the decoding stage with a low data volume, the hybrid expert layer adopts a more efficient tensor parallel strategy to effectively reduce the communication overhead between devices in the decoding stage, which can effectively improve the adaptability to the pre-filling stage and the decoding stage and improve the efficiency of the MoE sparse large model inference process; and, by adding the first hybrid expert layer model parameters operating based on the tensor parallel strategy to each device while performing multi-head attention layer calculation in the pre-filling stage, the communication overhead of parallel strategy conversion can be hidden in a layer-by-layer conversion manner without significantly increasing the device resource occupancy rate; and after the calculation of the hybrid expert layer of each structural layer in the pre-filling stage is completed, the second hybrid expert layer model parameters operating based on the expert parallel strategy can be released layer by layer, which can further avoid the increase of device resource occupancy rate.
[0051] Additional advantages, objects, and features of the present application will be partly set forth in the description which follows, and will partly become apparent to those of ordinary skill in the art upon examination of the following, or may be learned by practice of the present application. The objects and other advantages of the present application may be realized and attained by the structure particularly pointed out in the specification and the drawings.
[0052] Those skilled in the art will understand that the objects and advantages that can be achieved by the present application are not limited to those specifically described above, and the above and other objects that the present application can achieve will be more clearly understood according to the following detailed description. BRIEF DESCRIPTION OF THE DRAWINGS
[0053] The drawings described herein are used to provide a further understanding of the present application, and constitute a part of the present application, but do not limit the present application. The components in the drawings are not drawn to scale, but are only for showing the principles of the present application. For the convenience of showing and describing some parts of the present application, the corresponding parts in the drawings may be enlarged, that is, may become larger relative to other components in the exemplary device actually manufactured according to the present application. In the drawings:
[0054] Figure 1 It is a schematic diagram of the architecture layer of the MoE sparse large model.
[0055] Figure 2 It is a schematic diagram of the character routing (1) from data parallelism of the multi-head attention layer to expert parallelism of the mixture-of-experts layer with 4 devices and 8 experts as an example in the prior art.
[0056] Figure 3 It is a schematic diagram of the character routing (2) from expert parallelism of the mixture-of-experts layer to data parallelism of the multi-head attention layer with 4 devices and 8 experts as an example in the prior art.
[0057] Figure 4 It is the first flow schematic diagram of the staged hybrid parallel inference method for the MoE sparse large model in an embodiment of the present application.
[0058] Figure 5 It is an example diagram of the MoE sparse large model with two basic architecture layers in series and 8 experts in each basic architecture layer.
[0059] Figure 6 It is the first flow schematic diagram of the first step in the staged hybrid parallel inference method for the MoE sparse large model in an embodiment of the present application.
[0060] Figure 7 It is the second flow schematic diagram of the staged hybrid parallel inference method for the MoE sparse large model in an embodiment of the present application.
[0061] Figure 8 Schematic diagram of the first process of the second step in the phased hybrid parallel inference method of the MoE sparse large model in an embodiment of the present application.
[0062] Figure 9 Schematic diagram of the second process of the first step in the phased hybrid parallel inference method of the MoE sparse large model in an embodiment of the present application.
[0063] Figure 10 Schematic diagram of the second process of the second step in the phased hybrid parallel inference method of the MoE sparse large model in an embodiment of the present application.
[0064] Figure 11 Schematic diagram of the third process of the first step in the phased hybrid parallel inference method of the MoE sparse large model in an embodiment of the present application.
[0065] Figure 12 Schematic diagram of the fourth process of the first step in the phased hybrid parallel inference method of the MoE sparse large model in an embodiment of the present application.
[0066] Figure 13 Schematic diagram of character routing (1) from data parallelism of the multi-head attention layer to expert parallelism of the mixture-of-experts layer with 4 devices and 8 experts as an example in an application example of the present application.
[0067] Figure 14 Schematic diagram of the phased hybrid parallel strategy of the MoE mixture-of-experts layer provided by the application example of the present application.
[0068] Figure 15 Schematic diagram of the parallel strategy in the prefill stage with 4 devices as an example provided by the application example of the present application.
[0069] Figure 16 Schematic diagram of the parallel strategy in the decoding stage with 4 devices as an example provided by the application example of the present application.
[0070] Figure 17 Schematic diagram of character routing (2) from expert parallelism of the mixture-of-experts layer to data parallelism of the multi-head attention layer with 4 devices and 8 experts as an example provided by the application example of the present application.
[0071] Figure 18 Schematic diagram of the data transfer overhead hiding technology provided by the application example of the present application.
[0072] Figure 19(a) is a schematic diagram comparing the performance test results of the Prefill stage between this application and DeepSeek-MoE when batchsize = 128 and sequence_length = 2048 provided by the application example of this application.
[0073] Figure 19(b) is a schematic diagram comparing the performance test results of the Prefill stage between this application and DeepSeek-MoE when batchsize = 256 and sequence_length = 2048 provided by the application example of this application.
[0074] Figure 20(a) is a schematic diagram comparing the performance test results of the Decoding stage between this application and DeepSeek-MoE when batchsize = 128 and 100 tokens are output during inference provided by the application example of this application.
[0075] Figure 20(b) is a schematic diagram comparing the performance test results of the Decoding stage between this application and DeepSeek-MoE when batchsize = 256 and 100 tokens are output during inference provided by the application example of this application.
[0076] Figure 21(a) is a schematic diagram comparing the performance test results of the end-to-end inference (Prefill + Decoding) between this application and DeepSeek-MoE when batchsize = 128, sequence_length = 2048, and 100 tokens are output during inference provided by the application example of this application.
[0077] Figure 21(b) is a schematic diagram comparing the performance test results of the end-to-end inference (Prefill + Decoding) between this application and DeepSeek-MoE when batchsize = 256, sequence_length = 2048, and 100 tokens are output during inference provided by the application example of this application. Detailed implementation manners
[0078] To make the objectives, technical solutions, and advantages of this application clearer and more understandable, the following further elaborates on this application in combination with the implementation manners and the accompanying drawings. Herein, the illustrative implementation manners of this application and their descriptions are used to explain this application, but do not limit this application.
[0079] Here, it should also be noted that, in order to avoid obscuring the present application with unnecessary details, only the structures and / or processing steps closely related to the solution according to the present application are shown in the drawings, while other details less relevant to the present application are omitted.
[0080] It should be emphasized that the term "comprising / including" as used herein refers to the presence of features, elements, steps or components, but does not exclude the presence or addition of one or more other features, elements, steps or components.
[0081] Here, it should also be noted that, unless otherwise specified, the term "connection" in this document can not only refer to direct connection, but also represent indirect connection with an intermediate.
[0082] In the following, embodiments of the present application will be described with reference to the drawings. In the drawings, the same reference numerals represent the same or similar components, or the same or similar steps.
[0083] In recent years, the parameter scale of large models has grown rapidly. Training large models requires a large amount of computing resources. For example, taking GPT-4 with 1.8 trillion parameters as an example, the computing power required to train GPT-4 is equivalent to running for 90 to 100 days on 25,000 NVIDIA A100 GPUs. The MoE sparse large model can effectively reduce the training cost of large models and achieve sub-linear growth of the computing cost with respect to the model capacity. The MoE sparse large model consists of multiple sequentially connected architecture layers (which can also be referred to as: the basic architecture layer of the MoE large model). Among them, as Figure 1 shown, an architecture layer is composed of an adjacent multi-head attention layer and a mixture-of-experts layer (which can also be referred to as: the MoE mixture-of-experts layer). This architecture layer can be repeated multiple times and connected in series to form a multi-layer deep neural network. FFN0, FFN1 and FFN E-1They respectively represent the model parameters of the mixture-of-experts layer corresponding to different experts. Among them, Query represents the "question" or "target" feature to be attended to, which can be abbreviated as Q; Key represents the reference feature providing "position" or "matching basis", which can be abbreviated as K; Value represents carrying actual information for weighted aggregation, which can be abbreviated as V; QKV LinearProjection (QKV linear projection) is the core operation in the attention mechanism to generate Query (Q), Key (K), and Value (V). It maps the input data to different semantic spaces through linear transformation. In the MoE sparse large model, after each character (token) passes through the gating function, the affinity score of this character for each expert will be calculated, and then the top K experts with the highest affinity scores will be selected for calculation, where K ≥ 1. Each character does not need to be calculated by all experts, thus significantly reducing the computational cost. Current mainstream large models such as DeepSeek and Mixtral adopt the MoE mixture-of-experts architecture to reduce costs.
[0084] The MoE sparse large model introduces a kind of expert parallel model. For the mixture-of-experts layer in the MoE sparse large model, expert parallelism realizes the segmentation of expert parameters by evenly dividing all experts onto different devices; for the multi-head attention layer, data parallelism can be adopted, that is, the complete model parameters of the multi-head attention layer are saved on each device. The input characters are divided onto each device for multi-head attention layer calculation, and then according to the calculation result of the affinity score of the gating function, through an all-to-all communication operation, the characters are routed to the devices where the corresponding experts are located for MoE mixture-of-experts layer calculation, as Figure 2 shown. In Figure 2 , 8 experts are evenly divided onto 4 devices. Figure 2 In it, each square represents a character (token), and the number in the square represents the expert number to which it is routed (which can also be called: the destination expert number or expert number, etc.), and each of the said expert numbers corresponds one-to-one with each of the said experts. After the calculation of the mixture-of-experts layer is completed, through an all-to-all communication operation, the characters are routed back to the original devices, as Figure 3 shown, that is, each character after expert parallel calculation is respectively restored to its initial position on the device in the multi-head attention layer calculation, so as to perform data parallel calculation of the multi-head attention layer of the next layer.
[0085] The Mixture of Experts (MoE) model can significantly reduce the computational cost of large models. However, sparse large MoE models still face challenges in inference performance. The computational process of large model inference can be divided into two processes, namely the prefill stage and the decoding stage. The prefill stage refers to the computational process of generating the first character by inputting prompt words (such as user questions) to the model. The decoding stage refers to the computational process of iteratively generating the next character based on the current input sequence and the newly generated character after generating the first character. Decoding is an iterative computational process where a new character is generated in each iteration until the maximum generation length is reached or a termination character is encountered. In the prefill-decoding integrated (such as DeepSpeed-MoE) inference scheme, the prefill stage and the decoding stage share the same set of computing devices. However, the prefill stage is a compute-intensive load, while the decoding stage is a memory-intensive load, and the two stages interfere with each other in the prefill-decoding integrated inference scheme. Therefore, some research has proposed the prefill-decoding separation technique, where different computing devices are used for the prefill stage and the decoding stage respectively, achieving physical resource separation, eliminating the interference between the two stages, and enabling each stage to focus on its respective optimization goals, that is, the prefill stage focuses on optimizing the first character latency metric, while the decoding stage focuses on optimizing the output character latency metric. However, in the prefill-decoding separation technique, different computing devices are used for the prefill stage and the decoding stage respectively, so the prefill-decoding separation technique requires more computing resources. In addition, in the prefill-decoding separation technique, after the prefill stage calculation is completed, the KV-cache calculated in the prefill stage needs to be transmitted to the decoding device before the decoding device can perform subsequent calculations, thus incurring additional data transmission overhead. Therefore, both the prefill-decoding separation and prefill-decoding integrated schemes have their advantages and disadvantages.
[0086] In the expert parallelism of MoE sparse large model distributed inference, the load imbalance problem caused by the dynamic routing of characters is one of the key factors affecting inference performance. As Figure 2 shown, in the character routing from data parallelism in the multi-head attention layer to expert parallelism in the MoE layer, the dynamic routing of characters will result in different numbers of characters processed on each device, and at the same time, the all-to-all communication load is also unbalanced, that is, the number of characters received on each device is different. As Figure 3As shown, in the character routing from expert parallelism in the MoE layer to data parallelism in the multi-head attention layer, the dynamic routing of characters leads to an unbalanced communication load in the AlltoAll operation, that is, the number of characters sent on each device is different. Therefore, in the inference of MoE sparse large models, the dynamic routing of characters results in unbalanced computational and communication loads. To alleviate the load imbalance problem, Gshard (an efficient data processing technology designed based on the concept of distributed computing) proposed a design for the upper limit of expert capacity, that is, characters exceeding the expert capacity will directly skip the mixture-of-experts layer using a residual network, while experts that do not reach the capacity limit will be padded. Although this strategy can balance the computational and communication loads, its disadvantage is that it will generate useless computations, and at the same time, there may be a loss of accuracy due to the discarding of characters during inference, which is unacceptable during inference. DeepSeekV3 adopts a load balancing strategy without auxiliary loss. This strategy introduces an adjustable bias value for each expert, and this bias value is added to the expert affinity score calculated by the gating function to adjust the load on each expert. The bias value will be decreased for experts with high load and increased for experts with low load, so as to make the load on each expert as balanced as possible. However, this strategy cannot completely solve the load imbalance problem.
[0087] That is to say, for the distributed inference of MoE sparse large models, the parallel strategies adopted by existing technologies have poor adaptability to the pre-filling stage and the decoding stage, especially the PD integrated inference scheme. DeepSpeed-MoE is an optimization technology for the mixture-of-experts (MoE) network, aiming to improve the training and inference efficiency of large language models (LLMs). DeepSpeed-MoE proposed a hybrid parallel strategy of "data parallelism + tensor parallelism + expert parallelism" and restricted expert parallelism within a node, so as to utilize the high bandwidth within the node to improve the communication efficiency of the AlltoAll operation. However, the parallel strategy adopted by DeepSpeed-MoE does not distinguish between the pre-filling stage and the decoding stage, and the parallel strategy remains unchanged throughout the inference process, which is not compatible with the computational load characteristics of the two stages.
[0088] Based on this, in order to solve the problems that the parallel strategies adopted by existing MoE sparse large model inference technologies have poor adaptability to the pre-filling stage and the decoding stage and large communication overheads, etc., the embodiments of the present application respectively provide a phased hybrid parallel inference method for a MoE sparse large model, a phased hybrid parallel inference device for a MoE sparse large model for executing the phased hybrid parallel inference method of the MoE sparse large model, a phased hybrid parallel inference system for a MoE sparse large model, an electronic device, a computer-readable storage medium, and a computer program product. By adopting a phased hybrid parallel inference method, different parallel strategies adapted to different stages are used, and efficient connection is carried out through a layer-by-layer parallel strategy conversion method, which can achieve better performance compared with existing MoE sparse large model inference technologies.
[0089] Specifically, it will be described in detail through the following embodiments.
[0090] Based on this, the embodiments of the present application provide a phased hybrid parallel inference method for a MoE sparse large model that can be implemented by a phased hybrid parallel inference device for a MoE sparse large model. Refer to Figure 4 , the phased hybrid parallel inference method for the MoE sparse large model specifically includes the following contents:
[0091] Step 100: In the pre-filling stage of data inference based on the MoE sparse large model, control each architecture layer of the MoE sparse large model to sequentially execute the first step layer by layer; wherein, the first step includes: based on the multi-head attention layer model parameters and gating functions in each device, obtain the respective expert numbers of each character corresponding to the prompt data sequence currently located at the initial position of its own device, and at the same time add the first hybrid expert layer model parameters operating based on the tensor parallel strategy to each device; then perform expert parallel calculation on each character according to the respective expert numbers of each character and the respective second hybrid expert layer model parameters operating based on the expert parallel strategy in each device; and then restore each character after expert parallel calculation to its own device initial position and release the second hybrid expert layer model parameters in each device.
[0092] In one or more embodiments of the present application, refer to Figure 5 , the architecture layer refers to the basic architecture layer of the MoE sparse large model, and can also be called the basic architecture layer of the MoE large model. The MoE sparse large model is composed of multiple architecture layers connected in series in sequence. The first step is sequentially executed for each architecture layer layer by layer. Here, the layer-by-layer is based on the architecture layer as a unit. Each architecture layer contains a multi-head attention layer and a hybrid expert layer, and the output end of the multi-head attention layer is connected to the input end of the hybrid expert layer, and the output end of the hybrid expert layer is connected to the input end of the multi-head attention layer in the next architecture layer.
[0093] It should be noted that, for the pre-filling stage with a high volume of processed data, the embodiments of the present application adopt an expert parallel strategy with relatively lower communication volume for the mixture-of-experts layer to effectively reduce the communication overhead between devices during the pre-filling stage; while for the decoding stage with a low volume of processed data, the embodiments of the present application adopt a more efficient tensor parallel strategy for the mixture-of-experts layer to effectively reduce the communication overhead between devices during the decoding stage, which can effectively improve the adaptability to the pre-filling stage and the decoding stage and improve the efficiency of the inference process of the MoE sparse large model. However, since the mixture-of-experts layer adopts different parallel strategies in the pre-filling stage and the decoding stage respectively, therefore, it is necessary to change the parallel strategy of the mixture-of-experts layer from the expert parallel strategy to the tensor parallel strategy before executing the decoding stage. And changing the expert parallel strategy to the tensor parallel strategy will result in additional communication overhead between devices; if this communication process is to be avoided, it is necessary to store both the first mixture-of-experts layer model parameters running based on the tensor parallel strategy and the second mixture-of-experts layer model parameters running based on the expert parallel strategy corresponding to each mixture-of-experts layer in each architecture layer in each device simultaneously, which will in turn cause a significant increase in the data occupancy rate of each device.
[0094] Based on this, in order to avoid the additional communication overhead between devices caused by changing the expert parallel strategy to the tensor parallel strategy and avoid a significant increase in the data occupancy rate of each device, in step 100 of the phased hybrid parallel inference method of the MoE sparse large model provided by the embodiments of the present application, a first step is designed layer by layer for each architecture layer during the pre-filling stage, see Figure 6 , and the first step specifically includes the following contents:
[0095] Step 110: Based on the multi-head attention layer model parameters and the gating function in each device, obtain the expert numbers of each character corresponding to the prompt data sequence that are currently at their respective initial positions in their respective devices, and add the first mixture-of-experts layer model parameters running based on the tensor parallel strategy to each of the devices.
[0096] It can be understood that the prompt data sequence refers to a text sequence used as prompt information for predicting characters, and its data type can also be set according to actual application needs. In one example, the prompt data sequence can be user question text data. The device initial position refers to the position of each character in the device at this time, and for the convenience of subsequent description, it is referred to as the device initial position.
[0097] Among them, the multi-head attention layer model parameters refer to the model parameters of the multi-head attention layer, and each device stores the model parameters of each multi-head attention layer in each architecture layer.
[0098] In one or more embodiments of the present application, the tensor parallel strategy refers to: evenly dividing the high feature dimension H of all E experts onto P devices, where each device has [E, H / P] experts, that is, each device has H / P part of all E experts, and each device needs to process all characters. The model parameters of the first mixture-of-experts layer running based on the tensor parallel strategy refer to the model parameters of the mixture-of-experts layer corresponding to [E, H / P] owned by each device; and the model parameters of the first mixture-of-experts layer running based on the tensor parallel strategy are only used in the decoding phase, and the current prefill phase is only used to hide the communication overhead of setting the model parameters of the first mixture-of-experts layer running based on the tensor parallel strategy in the mixture-of-experts layer.
[0099] Step 120: Perform expert parallel computing on each of the characters according to the expert numbers of the respective characters and the model parameters of the second mixture-of-experts layer running based on the expert parallel strategy in the respective devices.
[0100] It can be understood that when step 110 is executed, the model parameters of the second mixture-of-experts layer running based on the expert parallel strategy corresponding to at least one expert (or expert number) already exist in the respective devices, and thus the model parameters of the second mixture-of-experts layer can be directly called to perform mixture-of-experts layer computing on each of the received characters when step 120 is executed. Among them, the expert parallel strategy refers to: evenly dividing all E experts onto P devices. Suppose the high feature dimension of each expert is H, and each device has [E / P, H] experts, where E / P is the number of experts on each device. And the model parameters of the second mixture-of-experts layer running based on the expert parallel strategy refer to: the model parameters of the mixture-of-experts layer corresponding to [E / P, H] owned by each device.
[0101] Step 130: Restore each of the characters after expert parallel computing to their respective initial positions in the devices and release the model parameters of the second mixture-of-experts layer in the respective devices.
[0102] In step 130, restoring each of the characters after expert parallel computing to their respective initial positions in the devices means that after the expert parallel computing in the mixture-of-experts layer is completed, the characters are routed back to the original devices through a full exchange communication operation, so as to perform data parallel computing on the multi-head attention layer of the next layer.
[0103] That is to say, in the above first step, during the process of performing multi-head attention calculation by the multi-head attention layer in the current architecture layer, the model parameters of the first mixture-of-experts layer running based on the tensor parallelism strategy corresponding to the multi-head attention layer in this architecture layer are added to each of the devices at the same time, so that the model parameters of each first mixture-of-experts layer added in each of the devices form a mixture-of-experts layer running based on the tensor parallelism strategy that is uniquely corresponding to the multi-head attention layer in this architecture layer. It should be noted that in the current state, each of the devices stores at the same time the model parameters of the multi-head attention layer corresponding to each architecture layer respectively, the model parameters of the second mixture-of-experts layer running based on the expert parallelism strategy corresponding to the current architecture layer and subsequent architecture layers respectively, and the model parameters of the first mixture-of-experts layer running based on the tensor parallelism strategy corresponding to the current architecture layer and previous architecture layers respectively. The subsequent architecture layers refer to all architecture layers located after the current architecture layer. If the current architecture layer is the last layer of the MoE sparse large model, its subsequent architecture layers are considered non-existent. The previous architecture layers refer to all architecture layers located before the current architecture layer. If the current architecture layer is the first layer of the MoE sparse large model, its previous architecture layers are considered non-existent.
[0104] In one example, if the MoE sparse large model includes 60 sequentially connected architecture layers, and the 30th architecture layer is currently executing the first step, then in step 110, while controlling each device to respectively obtain the expert numbers of each character corresponding to the prompt data sequence at their respective initial positions in the device based on the multi-head attention layer model parameters and gating functions of the 30th layer locally, the first hybrid expert layer model parameters corresponding to the 30th layer's hybrid expert layer operating based on the tensor parallel strategy are added to each device. Then, in step 120, each device is controlled to perform expert parallel computing on each character based on the second hybrid expert layer model parameters of the 30th layer operating based on the expert parallel strategy, and in step 130 after step 120, the second hybrid expert layer model parameters of the 30th layer operating based on the expert parallel strategy in each device are released. At this time, each device stores the multi-head attention layer model parameters of the 60-layer multi-head attention layer, the second hybrid expert layer model parameters corresponding to the hybrid expert layers of the 31st to 60th layers operating based on the expert parallel strategy, and the first hybrid expert layer model parameters corresponding to the hybrid expert layers of the 1st to 30th layers operating based on the tensor parallel strategy. That is to say, for the 60 hybrid expert layers in the 60 sequentially connected architecture layers included in the MoE sparse large model at this time, each device only needs to store the first hybrid expert layer model parameters operating based on the tensor parallel strategy or the second hybrid expert layer model parameters operating based on the expert parallel strategy. Each device may only store the first hybrid expert layer model parameters and the second hybrid expert layer model parameters corresponding to the same hybrid expert layer operating based on the tensor parallel strategy and the expert parallel strategy respectively during the execution of step 120. This can not only hide the communication overhead between devices caused by adding the first hybrid expert layer model parameters operating based on the tensor parallel strategy through the multi-head attention calculation process in the current architecture layer, but also effectively avoid a significant increase in the data occupancy rate of each device.
[0105] Step 200: Send the predicted characters output by the last architecture layer of the MoE sparse large model to the multi-head attention layer in the first architecture layer of the MoE sparse large model for performing the decoding stage of data inference of the MoE sparse large model based on the predicted characters, the multi-head attention layer model parameters in each device, the gating function, and the first hybrid expert layer model parameters of the tensor parallel strategy in each device.
[0106] In step 200, the predicted character output by the last architecture layer of the MoE sparse large model refers to: the output sequence corresponding to the prompt data sequence finally output by the mixture-of-experts layer in the last architecture layer of the MoE sparse large model based on the respective different second mixture-of-experts layer model parameters in each device, and then the last character is extracted from this output sequence as the predicted character obtained in the prefill stage. This predicted character will jointly form a predicted string with the predicted characters output in each iteration of the decoding stage, so as to serve as the response text data corresponding to the prompt data sequence such as the user question text data.
[0107] As can be seen from the above description, for the prefill stage with a high data volume processed by the phased hybrid parallel inference method of the MoE sparse large model provided in the embodiments of the present application, the mixture-of-experts layer adopts an expert parallel strategy with relatively lower communication volume to effectively reduce the communication overhead between devices in the prefill stage; while for the decoding stage with a low data volume, the mixture-of-experts layer adopts a more efficient tensor parallel strategy to effectively reduce the communication overhead between devices in the decoding stage, which can effectively improve the adaptability to the prefill stage and the decoding stage and improve the efficiency of the inference process of the MoE sparse large model; and, by adding the first mixture-of-experts layer model parameters operating based on the tensor parallel strategy to each device while performing the multi-head attention layer calculation in the prefill stage, the communication overhead of the parallel strategy conversion can be hidden in a layer-by-layer conversion manner without significantly increasing the device resource occupancy rate; and after the calculation of the mixture-of-experts layer of each structural layer in the prefill stage is completed, the second mixture-of-experts layer model parameters operating based on the expert parallel strategy can be released layer by layer, which can further avoid the increase in device resource occupancy rate.
[0108] In order to further improve the adaptability of the decoding stage in the phased hybrid parallel inference process of the MoE sparse large model to further reduce the communication overhead between devices, in a phased hybrid parallel inference method of the MoE sparse large model provided in the embodiments of the present application, see Figure 7 , after step 200 in the phased hybrid parallel inference method of the MoE sparse large model, the following specific content is further included:
[0109] Step 300: In each iteration round in the decoding stage of data inference based on the MoE sparse large model, control each architecture layer of the MoE sparse large model to execute the second step layer by layer, and send the predicted character output by the last architecture layer of the MoE sparse large model in each iteration round to the multi-head attention layer in the first architecture layer of the MoE sparse large model in a full-swap communication manner.
[0110] After the iteration ends, the predicted characters are output to jointly form a predicted string.
[0111] Among them, referring to Figure 8 , the specific content of the second step includes the following:
[0112] Step 310: In the current iteration round, based on the multi-head attention layer model parameters and gating functions in each of the devices, obtain the expert numbers of each character corresponding to the latest data sequence currently located at its respective initial position in its own device; wherein, if the current iteration round is the first round, the latest data sequence includes the prompt data sequence cached in the pre-fill stage and the predicted character output by the last architecture layer of the MoE sparse large model in the pre-fill stage; if the current iteration round is not the first round, the latest data sequence includes the prompt data sequence, the predicted character output in the pre-fill stage, and the predicted characters output in the previous historical iteration rounds before the current iteration round.
[0113] Step 320: Collect each of the characters to each device in an all-gather communication manner, and according to the expert number of each character, enable each device to perform tensor parallel computing on the collected characters according to the first mixture-of-experts layer model parameters running based on the tensor parallel strategy corresponding to each of them.
[0114] Step 330: Restore each of the characters after tensor parallel computing to their respective initial positions in their devices in a reduce-scatter communication manner.
[0115] To further improve the reliability of multi-head attention calculation in the staged hybrid parallel inference process of the MoE sparse large model, in a staged hybrid parallel inference method of a MoE sparse large model provided in an embodiment of the present application, referring to Figure 9 , the specific content of step 110 in the staged hybrid parallel inference method of the MoE sparse large model includes the following:
[0116] Step 111: Divide the prompt data sequence into multiple first data subsequences evenly according to the total number of each of the devices.
[0117] Step 112: Input each of the first data subsequences into each of the devices one-to-one, so that each device performs multi-head attention calculation on the first data subsequence received by it based on the multi-head attention layer model parameters stored in it respectively, to obtain each character corresponding to the first data subsequence received by it respectively, and determine the expert number of each character on each device based on the gating function respectively;
[0118] While performing step 112, perform step 113: Add the first mixture-of-experts layer model parameters running based on the tensor parallel strategy to each of the devices simultaneously.
[0119] Specifically, in the pre-filling stage, the multi-head attention layer adopts a data parallelism strategy, that is, all input characters are divided among P devices for parallel processing, and the model parameters of the multi-head attention layer are replicated on each device. After the data parallel computing of the multi-head attention layer is completed, the calculation result of the affinity score of the gating function is calculated.
[0120] To further improve the reliability of the multi-head attention calculation in the staged hybrid parallel inference process of the MoE sparse large model, in a staged hybrid parallel inference method of the MoE sparse large model provided in the embodiments of the present application, refer to Figure 10 Step 310 in the staged hybrid parallel inference method of the MoE sparse large model specifically includes the following contents:
[0121] Step 311: In the current iteration round, according to the total number of each device, the latest data sequence is evenly divided into multiple second data subsequences.
[0122] Step 312: Input each of the second data subsequences one-to-one into each device, so that each device performs multi-head attention calculation on the received second data subsequence based on the multi-head attention layer model parameters stored in it, respectively obtaining each character corresponding to the received second data subsequence, and collecting all the characters to each device in an all-gather communication manner, so that each device determines the expert number of each character on each device based on the gating function.
[0123] Specifically, in the decoding stage, the multi-head attention layer of the model adopts data parallelism, that is, all input characters are divided among each device for parallel processing, and the model parameters of the multi-head attention layer are replicated on each device. After the data parallel computing of the multi-head attention layer is completed, all the characters are collected to each device through an all-gather communication operation, and the activated experts are selected according to the calculation result of the affinity score of the gating function.
[0124] To further improve the effectiveness and reliability of the communication overhead of the hidden parallelism strategy conversion, in a staged hybrid parallel inference method of the MoE sparse large model provided in the embodiments of the present application, refer to Figure 11 Step 113 in the staged hybrid parallel inference method of the MoE sparse large model specifically includes the following contents:
[0125] Step 1131: If the time required to add the first mixture-of-experts layer model parameters running based on the tensor parallelism strategy to each of the devices is less than or equal to the time required to obtain the expert numbers of the respective characters currently located at their initial positions in their respective devices for the prompt data sequence, then while obtaining the expert numbers of the respective characters currently located at their initial positions in their respective devices for the prompt data sequence, add the first mixture-of-experts layer model parameters running based on the tensor parallelism strategy to each of the devices.
[0126] To further improve the applicability and reliability of the communication overhead of the hidden parallelism strategy conversion, in a phased hybrid parallel inference method for a MoE sparse large model provided in an embodiment of the present application, refer to Figure 12 , in the phased hybrid parallel inference method for the MoE sparse large model, steps 113 and 120 in the first step can also be replaced with steps 140 to 160, and specifically include the following contents:
[0127] Step 140: If the time required to add the first mixture-of-experts layer model parameters running based on the tensor parallelism strategy to each of the devices is longer than the time required to obtain the expert numbers of the respective characters currently located at their initial positions in their respective devices for the prompt data sequence, then divide the first mixture-of-experts layer model parameters corresponding to each of the devices into a first part of parameters and a second part of parameters.
[0128] Step 150: Based on the multi-head attention layer model parameters and gating functions in each device, obtain the expert numbers of the respective characters currently located at their initial positions in their respective devices for the prompt data sequence, and at the same time add the first part of parameters to each of the devices.
[0129] Step 160: Perform expert parallel computing on each of the characters according to the expert numbers of the respective characters and the second mixture-of-experts layer model parameters corresponding to each of the devices running based on the expert parallelism strategy, and at the same time add the second part of parameters to each of the devices.
[0130] That is to say, in the pre-fill stage, in addition to the addition process of the first mixture-of-experts layer model parameters running based on the tensor parallelism strategy that can be hidden during the multi-head attention layer calculation process, the addition process of the first mixture-of-experts layer model parameters running based on the tensor parallelism strategy can also be hidden during the mixture-of-experts layer calculation process running based on the expert parallelism strategy, and specifically can be determined according to the estimated communication duration required for the addition process of the first mixture-of-experts layer model parameters running based on the tensor parallelism strategy.
[0131] It is understandable that the design of the expert capacity upper limit means that characters exceeding the expert capacity will directly skip the MoE mixture-of-experts layer using a residual network, while experts that do not reach the capacity upper limit will be padded to complete. Although this strategy can balance the computational and communication loads, its drawback is that it will generate useless computations, and at the same time, due to the discarding of characters during inference, there may be a loss of accuracy, which is difficult to accept during inference. In DeepSeek V3, a load balancing strategy without auxiliary loss is adopted. This strategy introduces an adjustable bias value for each expert, and this bias value is added to the expert affinity score calculated by the gating function, thereby adjusting the load on each expert. For experts with high load, the bias value will be reduced, and for experts with low load, the bias value will be increased, but it can only make the load on each expert as balanced as possible. The prior art mainly optimizes load balancing at the algorithm level and cannot completely solve the problem of load imbalance. That is to say, for the distributed inference of MoE sparse large models, the prior art cannot completely solve the problem of unbalanced expert parallel load.
[0132] Based on this, in order to further solve the problem of unbalanced expert parallel load in the pre-padding stage, this application proposes a load balancing strategy based on the greedy algorithm. Specifically, in a phased hybrid parallel inference method for a MoE sparse large model provided in an embodiment of this application, refer to Figure 9 In step 120 of the phased hybrid parallel inference method for the MoE sparse large model, it specifically includes the following content:
[0133] Step 121: According to the respective expert numbers of each character, with the goal of equalizing the number of characters processed by each device during expert parallel computing and preferentially arranging the characters with the smallest expert numbers within the global scope locally in each device, perform load balancing processing on each character in each device based on the greedy algorithm.
[0134] Step 122: In a full-exchange communication manner, route each character after load balancing processing to the respective devices where their current expert numbers are located, so that each device respectively performs expert parallel computing on the characters it receives according to the second mixture-of-experts layer model parameters corresponding to its respective expert parallel strategy.
[0135] That is to say, in the phased hybrid parallel inference method proposed in this application, the mixture-of-experts layer in the decoding stage adopts tensor parallelism, which inherently has the characteristic of load balancing. However, the expert parallelism in the prefill stage still faces the problem of load imbalance. Therefore, for the expert parallelism in the prefill stage, this application proposes a load balancing strategy based on the greedy algorithm. While achieving load balancing for computing and communication, it preferentially arranges the characters with the smallest destination expert numbers within the global scope locally on the devices, so that the experts required for the characters processed on each device are as local as possible (i.e., hitting experts).
[0136] To further improve the application effectiveness and reliability of the load balancing strategy based on the greedy algorithm, in a phased hybrid parallel inference method of a MoE sparse large model provided in an embodiment of this application, refer to Figure 11 In step 121 of the phased hybrid parallel inference method of the MoE sparse large model, it specifically includes the following content:
[0137] Step 1211: Control each of the devices to sort the local characters in ascending order of expert numbers respectively to determine the distribution information of each local character on each expert for each device, and obtain the distribution information of the local characters of other devices on each expert in a all-gather communication manner to determine the global distribution information of each character on each expert.
[0138] Step 1212: Control each of the devices to determine the character distribution information after routing the characters to the devices where their respective expert numbers are currently located respectively according to the global distribution information of each character on each expert.
[0139] Step 1213: Control each of the devices to determine the start address and offset of sending characters and the start address and offset of receiving characters for each device according to the distribution information of each local character on each expert of each device and the character distribution information after routing the characters to the devices where their respective expert numbers are currently located.
[0140] Step 1214: Input the start address and offset of sending characters, the start address and offset of receiving characters, the send data buffer, and the receive data buffer of each device as input parameters into the all-to-all communication function corresponding to the all-to-all communication method.
[0141] That is to say, while achieving computing and communication load balancing, the present application preferentially arranges characters with the smallest destination expert numbers in the global scope locally on the device, so that the experts required for the characters processed on each device are as local as possible (i.e., expert hits). If the required experts are not local to the device (i.e., expert misses), the missing experts are loaded from the memory of a processor such as a CPU or GPU to the device video memory, and the loading overhead of the missing experts is hidden by the computing on the hit experts. The above strategy solves the load imbalance problem at the system level and is a load balancing optimization method without loss of accuracy. In addition, in the staged hybrid parallel inference method proposed in the present application, the hybrid expert layer in the decoding stage uses tensor parallelism and inherently has load balancing characteristics. Therefore, the present application can achieve full-stage lossless load balancing in the prefill stage and the decoding stage.
[0142] However, there is a phenomenon of "miss" of local experts for characters after the above load balancing character routing. For example Figure 13 In the right box (i.e., the grid diagram on the right side of the arrow where "Character Routing (1)" is located), after character routing, the characters processed by "Device 1" require "Expert 1" and "Expert 2", but the local experts of "Device 1" are "Expert 2" and "Expert 3", and "Expert 1" misses. Therefore, to solve the problem of expert misses, the present application proposes the following technical solution: all experts are stored in the CPU memory in advance. When an expert miss occurs on a certain device, the missing expert is loaded from the CPU memory to the device video memory for subsequent calculation, and the expert loading overhead can be hidden by the computing and communication overhead on the hit experts.
[0143] Specifically, in a staged hybrid parallel inference method for a MoE sparse large model provided in an embodiment of the present application, refer to Figure 11 , after step 1213 in the staged hybrid parallel inference method for the MoE sparse large model, the following specific content is further included:
[0144] Step 1215: If it is determined that there is a device with a missing expert after full exchange communication according to the offsets of the characters sent by each device and the offsets of the received characters, before performing expert parallel calculation on each character based on the model parameters of the second hybrid expert layer corresponding to the missing expert on this device, load the model parameters of the second hybrid expert layer corresponding to the missing expert from the processor memory to this device.
[0145] That is to say, the embodiment of the present application proposes a phased hybrid parallel inference method for MoE sparse large model inference. The core idea is that the expert parallelism is adopted in the mixture-of-experts layer in the prefill stage, the tensor parallelism is adopted in the mixture-of-experts layer in the decoding stage, and the data parallelism is adopted for the multi-head attention layers in both stages. In the scenario where the computing devices are shared between the prefill stage and the decoding stage, since the model parameters used by each device are different under expert parallelism and tensor parallelism, the present application proposes an efficient layer-by-layer parallel strategy conversion method to hide the communication overhead of parallel strategy conversion through layer-by-layer conversion, so as to complete the efficient conversion from expert parallelism in the prefill stage to tensor parallelism in the decoding stage. The principle that the phased hybrid parallel inference method proposed in the present application is more advantageous than the prior art is as follows: The amount of data processed in the prefill stage is large. Using expert parallelism in the mixture-of-experts layer only brings an all-to-all communication operation, and the communication volume of all-to-all is only the size of the input data, and the communication overhead is lower than the all-gather and reduce-scatter communication operations adopted by tensor parallelism. Therefore, expert parallelism is adopted in the mixture-of-experts layer in the prefill stage of the present application; the amount of data processed in the decoding stage is small and the computational density is low. If expert parallelism is adopted in the mixture-of-experts layer, it will bring serious load imbalance problems. Therefore, tensor parallelism is adopted in the mixture-of-experts layer in the decoding stage of the present application, which can achieve perfect load balance. At the same time, the input data volume in the decoding stage is low, and the communication is latency-limited. Using tensor parallelism will also result in a lower communication latency. The above phased hybrid parallel strategy can combine the performance advantages of adopting expert parallelism in the mixture-of-experts layer in the prefill stage and adopting tensor parallelism in the mixture-of-experts layer in the decoding stage. In the scenario where the computing devices are shared between the prefill stage and the decoding stage, the present application proposes an efficient layer-by-layer parallel strategy conversion method to achieve efficient conversion of the parallel strategy between the two stages.
[0146] To solve the load imbalance problem of expert parallelism in the prefill stage, the embodiment of the present application also proposes a load balancing strategy based on the greedy algorithm to solve the load imbalance problem from the system level. The principle of this load balancing strategy is as Figure 2As shown, in the expert parallelism during the pre-filling stage, the dynamic routing of characters can cause load imbalance problems. To achieve load balance, if some characters on high-load devices are directly distributed to low-load devices for processing, it is inevitable that the corresponding experts also need to be migrated to low-load devices, resulting in expert migration overhead. This application proposes a load balancing strategy based on the greedy algorithm. While achieving load balance in computing and communication, it preferentially arranges characters with the smallest destination expert number within the global scope locally on the device, so that the experts required for the characters processed on each device are as local as possible (i.e., hit experts). If the required experts are not local to the device (i.e., missed experts), the missed experts are loaded from the CPU memory to the device video memory, and the loading overhead of the missed experts is hidden through the computing on the hit experts. The above strategy solves the load imbalance problem at the system level and is a load balancing optimization method without loss of accuracy.
[0147] To further illustrate the staged hybrid parallel inference method of the MoE sparse large model provided in the above embodiments, this application also provides a specific application example of the staged hybrid parallel inference method of the MoE sparse large model, which is used to solve the following technical problems existing in the prior art:
[0148] (1) For the distributed inference of the MoE sparse large model, the parallel strategies adopted by the prior art have poor adaptability to the pre-filling stage and the decoding stage, especially the PD integrated inference scheme. A typical example is the hybrid parallel strategy of "data parallelism + tensor parallelism + expert parallelism" adopted by DeepSpeed-MoE. This parallel strategy does not distinguish between the pre-filling stage and the decoding stage, and the parallel strategy remains unchanged throughout the inference process. However, during the pre-filling stage, a large amount of data is processed, and using tensor parallelism will bring communication operations of all-gather and reduce-scatter in data, with high communication overhead, which is more serious when dealing with a large batch of data. In the decoding stage, the amount of data processed is relatively small. Using expert parallelism will cause load imbalance problems, and the decoding stage has a low computational density and a small amount of data processed, so the all-to-all communication overhead brought by expert parallelism is also more significant.
[0149] To this end, this application proposes a phased hybrid parallel inference method for MoE sparse large models. For the prefill stage with a high volume of processed data, the mixture-of-experts layer adopts expert parallelism with relatively lower communication volume. For the decoding stage with a low volume of processed data, the mixture-of-experts layer adopts more efficient tensor parallelism. For the multi-head attention layer, data parallelism is adopted in both stages. Since the parallel strategies adopted by the mixture-of-experts layer in the prefill stage and the decoding stage are different, and the model parameters used by each device in the two stages are also different, it is necessary to perform a parallel strategy conversion between the two stages, that is, the model parameters required for the decoding stage need to be loaded after the prefill execution ends. This application proposes an efficient way to perform parallel strategy conversion layer by layer, hiding the communication overhead of parallel strategy conversion through layer-by-layer conversion.
[0150] (2) For the distributed inference of MoE sparse large models, the prior art cannot completely solve the problem of uneven load in expert parallelism. Gshard proposed a design for the expert capacity limit, that is, characters exceeding the expert capacity will directly skip the MoE mixture-of-experts layer using a residual network, while experts that do not reach the capacity limit will be padded. Although this strategy can balance the computational and communication loads, its disadvantage is that it will generate useless computations, and at the same time, there may be a loss of accuracy due to the discarding of characters during inference, which is difficult to accept during inference. Deepseek V3 adopted a load balancing strategy without auxiliary loss. This strategy introduces an adjustable bias value for each expert, and this bias value will be added to the expert affinity score calculated by the gating function to adjust the load on each expert. The bias value will be reduced for experts with high load and increased for experts with low load, but it can only make the load on each expert as balanced as possible. The prior art mainly optimizes load balancing at the algorithm level and cannot completely solve the problem of uneven load.
[0151] To solve the problem of uneven load in expert parallelism during the prefill stage, this application proposes a load balancing strategy based on the greedy algorithm. While achieving balanced computational and communication loads, it preferentially arranges characters with the smallest destination expert number within the global scope locally on the device, so that the experts required for the characters processed on each device are as likely as possible to be local (i.e., hit experts). If the required expert is not local to the device (i.e., a missed expert), the missed expert will be loaded from the CPU memory to the device video memory, and the loading overhead of the missed expert is hidden by the computations on the hit experts. The above strategy solves the problem of uneven load at the system level and is a load balancing optimization method without loss of accuracy. In addition, in the phased hybrid parallel inference method proposed in this application, the mixture-of-experts layer in the decoding stage adopts tensor parallelism, which inherently has the characteristics of load balancing. Therefore, this application can achieve lossless load balancing in the entire prefill stage and decoding stage.
[0152] Based on this, seeFigure 14 The core MoE hybrid expert layer phased hybrid parallel strategy involved in the phased hybrid parallel inference method of the MoE sparse large model provided by the application example of this application is as follows:
[0153] 1. Phased hybrid parallel strategy for MoE sparse large model inference
[0154] The application example of this application proposes a phased hybrid parallel strategy for MoE sparse large model inference. The overall scheme is as Figure 14 shown. Among them, the hybrid expert layer in the prefill stage adopts expert parallelism, the hybrid expert layer in the decoding stage adopts tensor parallelism, and the multi-head attention layer in both stages adopts data parallelism. Between the two stages, an efficient layer-by-layer parallel strategy conversion method is adopted to achieve efficient conversion of the parallel strategy between the two stages. The specific technical solution of the phased hybrid parallel strategy is as follows:
[0155] 1-1. Parallel strategy in the prefill stage: Refer to Figure 15 , for the prefill stage, the hybrid expert layer of the model adopts expert parallelism, that is, all E experts are evenly divided into P devices. Let the high feature dimension of each expert be H, and each device has [E / P, H] experts, where E / P is the number of experts on each device, and each device processes the characters after routing; the multi-head attention layer of the model adopts data parallelism, that is, all input characters are divided into P devices for parallel processing, and the model parameters of the multi-head attention layer are replicated on each device. After the data parallel calculation of the multi-head attention layer is completed, according to the calculation result of the affinity score of the gating function, the characters are routed to the devices where the corresponding experts are located through an AlltoAll all-to-all communication operation for expert parallel calculation of the hybrid expert layer. After the expert parallel calculation of the hybrid expert layer is completed, the characters are routed back to the original devices through an AlltoAll all-to-all communication operation, so as to connect to the data parallel calculation of the multi-head attention layer of the next layer.
[0156] 1-2. Parallel strategy in the decoding stage: Refer to Figure 16, for the decoding stage, the mixture-of-experts layer of the model uses tensor parallelism, that is, the high feature dimension H of all E experts is evenly divided among P devices. Each device has [E, H / P] experts, that is, each device has H / P parts of all E experts, and each device needs to process all characters; the multi-head attention layer of the model uses data parallelism, that is, all input characters are divided among devices for parallel processing, and the model parameters of the multi-head attention layer are replicated on each device. After the data parallel computing of the multi-head attention layer is completed, all characters are gathered on each device through an AllGather communication operation, and the activated experts are selected according to the calculation results of the affinity scores of the gating function. Then, all characters and the activated experts are used as inputs for the local computing of the tensor parallelism of the mixture-of-experts layer. After the local computing of the tensor parallelism of the mixture-of-experts layer is completed, the local computing results are reduced and distributed through a ReduceScatter communication operation, so as to connect to the data parallel computing of the multi-head attention layer of the next layer.
[0157] 1-3. Method for converting layer-by-layer parallel strategies between the prefill and decoding stages: The mixture-of-experts layer in the prefill stage uses expert parallelism, and the mixture-of-experts layer is divided according to the number of experts dimension; the mixture-of-experts layer in the decoding stage uses tensor parallelism, and the mixture-of-experts layer is divided according to the high feature dimension of the experts. The model parameters used by the mixture-of-experts layer in the prefill and decoding stages are different, and preparing two copies of the model parameters in the video memory at the same time will significantly increase the video memory overhead. Therefore, the application example of this application proposes an efficient way to convert layer-by-layer parallel strategies, and hides the communication overhead of the parallel strategy conversion through the way of layer-by-layer conversion. Suppose the model has a total of L layers of the basic architecture of the MoE mixture-of-experts model (as Figure 1 shown), when performing the calculation of each layer of the MoE basic architecture, the distribution of the model parameters of the expert parallelism of the mixture-of-experts layer of the current basic architecture is converted into the distribution of the model parameters of the tensor parallelism of the mixture-of-experts layer of the current basic architecture through the AlltoAll all-to-all communication operation. After the AlltoAll all-to-all communication operation and the calculation of the current mixture-of-experts layer are completed, the model parameters required for the expert parallelism of the mixture-of-experts layer in the video memory are released, so as to complete the parallel strategy conversion of the current layer. By sequentially completing the parallel strategy conversion of L layers according to the above method, the parallel strategy conversion of the entire model from the prefill stage to the decoding stage can be completed.
[0158] 2. Load balancing strategy based on greedy algorithm
[0159] In the phased hybrid parallel inference method proposed in the application example of this application, the hybrid expert layer in the decoding stage uses tensor parallelism, which inherently has the characteristic of load balancing. However, the expert parallelism in the prefill stage still faces the problem of load imbalance. Therefore, the application example of this application proposes a load balancing strategy based on the greedy algorithm for the expert parallelism in the prefill stage. While achieving load balancing of computing and communication, it preferentially arranges the characters with the smallest destination expert numbers within the global scope locally on the device, so that the experts required for the characters processed on each device are as local as possible (i.e., hitting the experts). Specifically: First step, calculate the destination expert numbers of each input character according to the gating function, sort the input characters locally on each device, and arrange the characters with smaller destination expert numbers in the front, as shown in Figure 13 the left box in Figure 13 (i.e., the grid diagram to the left of the arrow where "Character Routing (1)" is located), so as to determine the distribution information of the local input characters among each expert before the AlltoAll character routing; Second step, determine the character distribution targets after the AlltoAll character routing for each device. Here, two goals need to be achieved. One is to ensure that the number of characters processed by each device is equal to achieve load balancing, and at the same time, preferentially arrange the characters with the smallest destination expert numbers within the global scope locally on the device, so that the characters processed on each device can hit the local experts as much as possible; Third step, in order to achieve the above goals, each device counts the distribution information of the local input characters among each expert, and each device uses the AllGather communication operation to collect the distribution information of all input characters among each expert, so as to calculate the total number of characters processed by each expert within the global scope; Fourth step, according to the global distribution information of the characters among each expert, each device determines the character distribution information after the AlltoAll character routing. Let the total number of input characters be N and the number of devices be P. Starting from the device with the smallest number, preferentially arrange the characters with the smallest destination expert numbers within the global scope locally on the device until N / P characters are filled, and then arrange the next device until P devices are arranged. The character distribution effect after routing is as shown in the right box in Figure 13 (i.e., the grid diagram to the right of the arrow where "Character Routing (1)" is located); Fifth step, according to the distribution information of the local input characters among each expert before the AlltoAll character routing and the character distribution information after the AlltoAll character routing obtained above, calculate the start address and offset of the characters sent by each device and the start address and offset of the characters received in the AlltoAll communication operation; Sixth step, use the start address and offset of the sent and received data calculated in the fifth step, as well as the send data buffer and receive data buffer as input parameters, input them into the AlltoAll communication function, and then perform the AlltoAll communication operation to complete the load-balanced character routing.
[0160] An example of load - balanced character routing implemented according to the above technical solution is as follows Figure 13 as shown. After routing, the number of characters processed by each device is the same, achieving computational load balance. At the same time, for the AlltoAll communication operation, the total amount of data sent by each device is the same, and the total amount of data received by each device is the same, achieving communication load balance at the endpoints. However, after the above load - balanced character routing, there is a phenomenon that characters "miss" local experts. For example Figure 13 in the right - hand box of Figure 13 (i.e., "the grid diagram on the right side of the arrow where 'Character Routing (1)' is located"), after character routing, the characters processed by "Device 1" require "Expert 1" and "Expert 2", but the local experts of "Device 1" are "Expert 2" and "Expert 3", and "Expert 1" misses. To solve the problem of expert miss, the application example of this application proposes the following technical solution: Store all experts in the CPU memory in advance. When an expert miss occurs in a certain device, the missing expert is loaded from the CPU memory to the device video memory for subsequent calculations, and the expert loading overhead can be hidden by the computational and communication overhead on the hit experts.
[0161] After the expert parallel calculation in the hybrid expert layer is completed, the characters are routed back to the original devices through an inverse AlltoAll full - exchange communication operation, so as to connect to the data parallel calculation of the multi - head attention layer in the next layer, as Figure 17 shown. For this AlltoAll communication operation, the total amount of data sent by each device is the same, and the total amount of data received by each device is the same, also achieving communication load balance at the endpoints.
[0162] The above load - balanced technical solution proposed by the application example of this application achieves computational and communication load balance while ensuring that the model accuracy is not damaged, and completely solves the load - balance problem at the system level. At the same time, in the staged hybrid parallel inference method proposed by the application example of this application, the hybrid expert layer in the decoding stage uses tensor parallelism, which inherently has load - balancing characteristics. Therefore, the application example of this application can achieve full - stage lossless load balance in the pre - fill stage and the decoding stage.
[0163] 3. Data transmission overhead hiding technology
[0164] In the above phased hybrid parallel technical solution for the inference of MoE sparse large models, it is necessary to complete the layer-by-layer parallel strategy conversion of the mixture-of-experts layer from the prefill stage to the decoding stage through an all-to-all communication operation. In the above technical solution of the load balancing strategy based on the greedy algorithm, it is necessary to load the missed experts from the CPU memory to the device video memory. The layer-by-layer parallel strategy conversion and the loading of missed experts will bring certain data transmission overheads. The application example of this application proposes a data transmission overhead hiding technology, which hides the data transmission overheads brought by the layer-by-layer parallel strategy conversion and the loading of missed experts through irrelevant computations and communication loads. The technical solution is as Figure 18 shown. For the all-to-all communication overhead brought by the parallel strategy conversion (from expert parallelism to tensor parallelism) of the MoE mixture-of-experts layer, it is masked by computational loads such as multi-head attention layer computation, gate function computation, and MoE mixture-of-experts layer computation. For the data transmission overhead brought by the loading of missed experts, it is masked by loads such as the all-to-all communication of character routing from multi-head attention layer data parallelism to MoE layer expert parallelism, the mixture-of-experts layer computation on the hit experts, and offset computation.
[0165] In addition, the application example of this application also conducted corresponding experimental tests, and the specific content is as follows:
[0166] The experimental tests were carried out on a 2-machine 16-card Ascend 910B distributed machine, where the single-card video memory capacity of Ascend 910B is 64GB. The test object is the DeepSeek V2 mixture-of-experts MoE sparse large model (with 236 billion model parameters), and the comparison baseline is the DeepSpeed-MoE parallel framework.
[0167] Figures 19(a) and 19(b) respectively show the performance results of the DeepSeek-V2 prefill stage under the conditions of batchsize (batch size) = 128 and batchsize = 256, where the abscissa is the degree of load imbalance (represented by variance). The test results show that the load balancing strategy based on the greedy algorithm proposed in the application example of this application can effectively solve the load imbalance problem of the mixture-of-experts sparse large model, and the execution time of the prefill stage is hardly affected by the degree of load imbalance. Compared with the baseline DeepSpeed-MoE framework, as the degree of load imbalance increases, the performance advantage of the application example of this application becomes more obvious, and the highest performance speedup ratio of 2.7 times is obtained. The above performance test results verify the effectiveness of the load balancing strategy based on the greedy algorithm proposed in this application.
[0168] Figures 20(a) and 20(b) respectively show the performance results of the DeepSeek-V2 decoding stage under the conditions of batchsize = 128 and batchsize = 256, where the abscissa is the degree of load imbalance. The test results show that in the phased hybrid parallel strategy proposed in the application example of this application, tensor parallelism is adopted in the decoding stage, and a performance speedup ratio of 2.3 times is obtained compared with the expert parallel strategy adopted by the baseline DeepSpeed-MoE, thus proving the rationality and effectiveness of adopting tensor parallelism in the decoding stage of the phased hybrid parallel strategy proposed in the application example of this application.
[0169] Figures 21(a) and 21(b) respectively show the performance results of the end-to-end inference (including two stages of pre-filling and decoding) of DeepSeek-V2 under the conditions of batchsize = 128 and batchsize = 256, where the abscissa is the degree of load imbalance. The test results show that the phased hybrid parallel inference method and system of the MoE sparse large model proposed in the application example of this application can obtain a maximum performance speedup ratio of 2.4 times compared with the parallel strategy with the unchanged entire inference process adopted by the baseline DeepSpeed-MoE. In the end-to-end inference, the system implementation of the application example of this application integrates all the proposed technical solutions, including the phased hybrid parallel strategy for MoE sparse large model inference, the load balancing strategy based on the greedy algorithm, and the data transmission overhead hiding technology. The above end-to-end inference performance results can verify the effectiveness of the technical solutions proposed in the application example of this application.
[0170] In summary, compared with the prior art, the phased hybrid parallel inference method of the MoE sparse large model provided by the application example of this application has the following beneficial effects:
[0171] (1) The parallel strategy adopted by the prior art does not fully consider the computational load characteristics of the prefill and decoding stages. The proposed phased hybrid parallel strategy in this application better adapts to the computational load characteristics of the prefill and decoding stages during the inference process of the MoE sparse large model. Compared with the prior art, it has the following advantages: In the phased hybrid parallel strategy proposed in this application, the hybrid expert layer in the prefill stage adopts expert parallelism. The communication volume of the Alltoall operation brought by expert parallelism is only the size of the input data, and the communication overhead is lower than that of the AllGather and ReduceScatter communication operations used in tensor parallelism. In this application, the hybrid expert layer in the decoding stage adopts tensor parallelism, which can achieve full load balancing in terms of computing, memory access, and communication. At the same time, the input data volume in the decoding stage is low, and the communication is latency-limited. Using tensor parallelism will also result in lower communication latency. An efficient layer-by-layer parallel strategy conversion method is adopted to hide the communication overhead of parallel strategy conversion through layer-by-layer conversion, thereby completing the efficient conversion from expert parallelism in the prefill stage to tensor parallelism in the decoding stage, and combining the performance advantages of expert parallelism in the hybrid expert layer in the prefill stage and tensor parallelism in the hybrid expert layer in the decoding stage. In summary, the proposed phased hybrid parallel inference method in this application adopts different parallel strategies adapted to different stages and is efficiently connected through the layer-by-layer parallel strategy conversion method, achieving better performance compared with the prior art.
[0172] (2) Regarding the load balancing problem of the MoE sparse large model, the prior art mainly optimizes from the model algorithm level, such as the expert capacity upper limit strategy of Gshard and the load balancing strategy without auxiliary loss of Deepseek V3. However, the prior art cannot fundamentally solve the load balancing problem, either causing model accuracy loss or still having a certain degree of load imbalance. For the load imbalance problem of expert parallelism in the prefill stage, this application proposes a load balancing strategy based on the greedy algorithm, which ensures model accuracy lossless while achieving computational and communication load balancing, and completely solves the load balancing problem from the system level. At the same time, in the phased hybrid parallel inference method proposed in this application, the hybrid expert layer in the decoding stage adopts tensor parallelism, which inherently has the load balancing characteristic. Therefore, this application can achieve full-stage lossless load balancing in the prefill stage and the decoding stage.
[0173] From a software perspective, this application also provides a phased hybrid parallel inference device for the MoE sparse large model to execute all or part of the content in the phased hybrid parallel inference method of the MoE sparse large model. The phased hybrid parallel inference device for the MoE sparse large model specifically includes the following content:
[0174] A pre - filling stage execution module, which is used to control each architecture layer of the MoE sparse large model to execute the first step layer by layer during the pre - filling stage of data inference based on the MoE sparse large model; wherein, the first step includes: based on the multi - head attention layer model parameters and gating functions in each device, obtaining the respective expert numbers of each character corresponding to the prompt data sequence that are initially located at their respective device positions, and adding the first mixture - of - experts layer model parameters operating based on the tensor parallel strategy to each device; then performing expert parallel computing on each character according to the respective expert numbers of each character and the second mixture - of - experts layer model parameters operating based on the expert parallel strategy in each device; and then restoring each character after expert parallel computing to its respective device initial position and releasing the second mixture - of - experts layer model parameters in each device.
[0175] A stage transition execution module, which is used to send the predicted character output by the last architecture layer of the MoE sparse large model to the multi - head attention layer in the first architecture layer of the MoE sparse large model, so as to execute the decoding stage of data inference of the MoE sparse large model according to the predicted character, the multi - head attention layer model parameters in each device, the gating function, and the first mixture - of - experts layer model parameters operating based on the tensor parallel strategy in each device.
[0176] In some embodiments of the present application, the stage - based hybrid parallel inference device of the MoE sparse large model further includes the following modules executed after the stage transition execution module:
[0177] A decoding stage execution module, which is used to control each architecture layer of the MoE sparse large model to execute the second step layer by layer in each iteration round during the decoding stage of data inference based on the MoE sparse large model, and send the predicted character output by the last architecture layer of the MoE sparse large model in each iteration round to the multi - head attention layer in the first architecture layer of the MoE sparse large model in a full - exchange communication manner.
[0178] The embodiments of the stage - based hybrid parallel inference device of the MoE sparse large model provided in the present application can be specifically used to execute the processing flow of the embodiments of the stage - based hybrid parallel inference method of the MoE sparse large model in the above embodiments. Its functions will not be elaborated here, and reference can be made to the detailed description of the embodiments of the stage - based hybrid parallel inference method of the MoE sparse large model above.
[0179] The part of the phased hybrid parallel inference device for the MoE sparse large model to perform phased hybrid parallel inference of the MoE sparse large model can be completed in a server or a client device. Specifically, it can be selected according to the processing capacity of the client device and the limitations of the user usage scenario, etc. This application does not make any limitations in this regard. If all operations are completed in the client device, the client device may further include a processor for specific processing of the phased hybrid parallel inference of the MoE sparse large model.
[0180] The above-mentioned client device may have a communication module (i.e., a communication unit) and can communicate with a remote server to achieve data transmission with the server. The server may include a server on the task scheduling center side, and in other implementation scenarios, it may also include a server of an intermediate platform, such as a server of a third-party server platform having a communication link with the task scheduling center server. The server may include a single computer device, or a server cluster composed of multiple servers, or a server structure of a distributed device.
[0181] Any suitable network protocol can be used for communication between the above-mentioned server and the client device side, including network protocols not yet developed on the filing date of this application. The network protocol may, for example, include TCP / IP protocol, UDP / IP protocol, HTTP protocol, HTTPS protocol, etc. Of course, the network protocol may, for example, also include the RPC protocol (Remote Procedure Call Protocol) and the REST protocol (Representational State Transfer) used on top of the above protocols, etc.
[0182] As can be seen from the above description, for the pre-fill stage that processes a large amount of data, the mixture-of-experts layer of the phased hybrid parallel inference device for the MoE sparse large model provided by the embodiments of the present application adopts an expert parallel strategy with relatively lower communication volume to effectively reduce the communication overhead between devices in the pre-fill stage; for the decoding stage that processes a small amount of data, the mixture-of-experts layer adopts a more efficient tensor parallel strategy to effectively reduce the communication overhead between devices in the decoding stage, which can effectively improve the adaptability to the pre-fill stage and the decoding stage and improve the efficiency of the inference process of the MoE sparse large model; and, by adding the first mixture-of-experts layer model parameters operating based on the tensor parallel strategy to each device while performing the multi-head attention layer calculation in the pre-fill stage, the communication overhead of the parallel strategy conversion can be hidden in a layer-by-layer conversion manner without significantly increasing the device resource occupancy rate; and after the calculation of the mixture-of-experts layer of each structural layer in the pre-fill stage is completed, the second mixture-of-experts layer model parameters operating based on the expert parallel strategy can be released layer by layer, which can further avoid the increase in device resource occupancy rate.
[0183] Embodiments of the present application also provide an electronic device, which may include a processor, a memory, a receiver, and a transmitter. The processor is configured to execute the phased hybrid parallel inference method for the MoE sparse large model mentioned in the above embodiments. The processor and the memory may be connected through a bus or other means. Taking the bus connection as an example. The receiver can be connected to the processor and the memory in a wired or wireless manner.
[0184] The processor may be a central processing unit (CPU). The processor may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. chips, or combinations of the above types of chips.
[0185] As a non-transitory computer-readable storage medium, the memory can be used to store non-transitory software programs, non-transitory computer-executable programs, and modules, such as program instructions / modules corresponding to the phased hybrid parallel inference method for the MoE sparse large model in the embodiments of the present application. The processor executes various functional applications and data processing of the processor by running the non-transitory software programs, instructions, and modules stored in the memory, that is, implements the phased hybrid parallel inference method for the MoE sparse large model in the above method embodiments.
[0186] The memory may include a program storage area and a data storage area. Among them, the program storage area can store the operating system and application programs required for at least one function; the data storage area can store data created by the processor and the like. In addition, the memory may include high-speed random access memory, and may also include non-transitory memory, such as at least one disk storage device, a flash memory device, or other non-transitory solid-state storage devices. In some embodiments, the memory may optionally include a memory remotely set relative to the processor, and these remote memories can be connected to the processor through a network. Examples of the above-mentioned network include but are not limited to the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0187] The one or more modules are stored in the memory and, when executed by the processor, implement the phased hybrid parallel inference method of the MoE sparse large model in the embodiments.
[0188] In some embodiments of the present application, the user equipment may include a processor, a memory, and a transceiver unit. The transceiver unit may include a receiver and a transmitter. The processor, the memory, the receiver, and the transmitter may be connected through a bus system. The memory is used to store computer instructions, and the processor is used to execute the computer instructions stored in the memory to control the transceiver unit to transmit and receive signals.
[0189] As an implementation manner, the functions of the receiver and the transmitter in the present application may be considered to be implemented through a transceiver circuit or a dedicated chip for transceiver, and the processor may be considered to be implemented through a dedicated processing chip, a processing circuit, or a general-purpose chip.
[0190] As another implementation manner, it may be considered to use a general-purpose computer to implement the server provided in the embodiments of the present application. That is, the program codes for implementing the functions of the processor, the receiver, and the transmitter are stored in the memory, and the general-purpose processor implements the functions of the processor, the receiver, and the transmitter by executing the codes in the memory.
[0191] The embodiments of the present application also provide a phased hybrid parallel inference system for the MoE sparse large model. The phased hybrid parallel inference system for the MoE sparse large model at least includes a scheduling device and various devices respectively communicatively connected to the scheduling device; the scheduling device includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, it implements the phased hybrid parallel inference method of the MoE sparse large model provided in the foregoing embodiments.
[0192] An embodiment of the present application further provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the foregoing phased hybrid parallel inference method for the MoE sparse large model are implemented. The computer-readable storage medium may be a tangible storage medium, such as a random access memory (RAM), internal memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, register, floppy disk, hard disk, removable storage disk, CD-ROM, or any other form of storage medium well-known in the technical field.
[0193] An embodiment of the present application further provides a computer program product, including a computer program. When the computer program is executed by a processor, the steps of the foregoing phased hybrid parallel inference method for the MoE sparse large model are implemented.
[0194] Those of ordinary skill in the art should understand that the various exemplary components, systems, and methods described in conjunction with the embodiments disclosed herein can be implemented in hardware, software, or a combination of both. Specifically, whether to implement in hardware or software depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present application. When implemented in hardware, it can be, for example, an electronic circuit, an application-specific integrated circuit (ASIC), appropriate firmware, a plug-in, a functional card, and so on. When implemented in software, the elements of the present application are programs or code segments used to perform the required tasks. The program or code segment can be stored in a machine-readable medium or transmitted via a data signal carried in a carrier wave on a transmission medium or a communication link.
[0195] It should be clear that the present application is not limited to the specific configurations and processes described above and illustrated in the figures. For the sake of brevity, detailed descriptions of known methods are omitted here. In the above embodiments, several specific steps are described and illustrated as examples. However, the method process of the present application is not limited to the specific steps described and illustrated. Those skilled in the art can make various changes, modifications, and additions, or change the order between steps after understanding the spirit of the present application.
[0196] In the present application, the features described and / or illustrated for one embodiment can be used in the same way or in a similar way in one or more other embodiments, and / or combined with the features of other embodiments or replace the features of other embodiments.
[0197] The above are only the preferred embodiments of the present application and are not intended to limit the present application. For those skilled in the art, various changes and modifications can be made to the embodiments of the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included within the protection scope of the present application.
Claims
1. A phased hybrid parallel inference method for MoE sparse large models, characterized in that, Including: In the pre-filling stage of data inference based on the MoE sparse large model, control each architecture layer of the MoE sparse large model to sequentially execute the first step layer by layer; wherein, the first step includes: based on the multi-head attention layer model parameters and gating functions in each device, obtain the expert numbers of each character corresponding to the prompt data sequence currently located at their respective initial positions in their respective devices, and at the same time add the first mixture-of-experts layer model parameters operating based on the tensor parallel strategy to each of the devices; then perform expert parallel computing on each of the characters according to the expert numbers of each character and the second mixture-of-experts layer model parameters operating based on the expert parallel strategy in each of the devices; then restore each of the characters after expert parallel computing to their respective device initial positions and release the second mixture-of-experts layer model parameters in each of the devices. Send the predicted characters output by the last architecture layer of the MoE sparse large model to the multi-head attention layer in the first architecture layer of the MoE sparse large model for performing the decoding stage of data inference of the MoE sparse large model according to the predicted characters, the multi-head attention layer model parameters in each device, the gating function, and the first mixture-of-experts layer model parameters operating based on the tensor parallel strategy in each of the devices.
2. The phased hybrid parallel inference method of the MoE sparse large model according to claim 1, wherein Also including: In each iteration round of the decoding stage of data inference based on the MoE sparse large model, control each architecture layer of the MoE sparse large model to sequentially execute the second step layer by layer, and send the predicted characters output by the last architecture layer of the MoE sparse large model in each iteration round to the multi-head attention layer in the first architecture layer of the MoE sparse large model in a full-swap communication manner. Wherein, the second step includes: In the current iteration round, based on the multi-head attention layer model parameters and gating functions in each of the devices, obtain the expert numbers of each character corresponding to the latest data sequence currently located at their respective initial positions in their respective devices; wherein, if the current iteration round is the first round, the latest data sequence includes the prompt data sequence cached in the pre-filling stage and the predicted characters output by the last architecture layer of the MoE sparse large model in the pre-filling stage; if the current iteration round is not the first round, the latest data sequence includes the prompt data sequence, the predicted characters output in the pre-filling stage, and the predicted characters output in the historical iteration rounds before the current iteration round. Collect each of the characters to each device in a all-gather communication manner, and according to the expert numbers of each character, enable each device to perform tensor parallel computing on the collected characters according to the first mixture-of-experts layer model parameters operating based on the tensor parallel strategy corresponding to each of them. Restore each of the characters after tensor parallel computing to their respective device initial positions in a reduce-scatter communication manner.
3. The phased hybrid parallel inference method for the MoE sparse large model according to claim 2, wherein Obtaining the respective expert numbers of each character currently located at the initial position of its respective device corresponding to the prompt data sequence based on the multi-head attention layer model parameters and gating functions in each device, includes: Dividing the prompt data sequence into multiple first data subsequences evenly according to the total number of each device; Inputting each of the first data subsequences one-to-one into each device respectively, so that each device performs multi-head attention calculation on the first data subsequence it receives based on the multi-head attention layer model parameters stored in itself, to obtain the respective characters corresponding to the first data subsequence it receives respectively, and determining the respective expert numbers of each character on each device based on the gating function respectively; Correspondingly, obtaining the respective expert numbers of each character currently located at the initial position of its respective device corresponding to the latest data sequence based on the multi-head attention layer model parameters and gating functions in each device, includes: Dividing the latest data sequence into multiple second data subsequences evenly according to the total number of each device; Inputting each of the second data subsequences one-to-one into each device respectively, so that each device performs multi-head attention calculation on the second data subsequence it receives based on the multi-head attention layer model parameters stored in itself, to obtain the respective characters corresponding to the second data subsequence it receives respectively, and collecting all the characters to each device in an all-gather communication manner, so that each device determines the respective expert numbers of each character on each device based on the gating function respectively.
4. The phased hybrid parallel inference method for the MoE sparse large model according to claim 1, wherein Adding the first mixture-of-experts layer model parameters operating based on the tensor parallel strategy to each device respectively, includes: If the time required to add the first mixture-of-experts layer model parameters operating based on the tensor parallel strategy to each device is shorter than or equal to the time required to obtain the respective expert numbers of each character currently located at the initial position of its respective device corresponding to the prompt data sequence, then while obtaining the respective expert numbers of each character currently located at the initial position of its respective device corresponding to the prompt data sequence, adding the first mixture-of-experts layer model parameters operating based on the tensor parallel strategy to each device.
5. The phased hybrid parallel inference method for the MoE sparse large model according to claim 1, wherein Before restoring each of the characters after expert parallel calculation to its respective device initial position in the first step, further includes: If the time required to add the first mixture-of-experts layer model parameters operating based on the tensor parallel strategy to each device is longer than the time required to obtain the respective expert numbers of each character currently located at the initial position of its respective device corresponding to the prompt data sequence, then dividing the first mixture-of-experts layer model parameters corresponding to each device operating based on the tensor parallel strategy into a first part of parameters and a second part of parameters; Obtaining the respective expert numbers of each character currently located at the initial position of its respective device corresponding to the prompt data sequence based on the multi-head attention layer model parameters and gating functions in each device, and adding the first part of parameters to each device simultaneously; Then, perform expert parallel computing on each of the characters according to the respective expert numbers of each character and the second hybrid expert layer model parameters corresponding to each of the devices and running based on the expert parallel strategy, and add the second part of the parameters to each of the devices at the same time.
6. The phased hybrid parallel inference method of the MoE sparse large model according to claim 1, characterized in that The performing expert parallel computing on each of the characters according to the respective expert numbers of each character and the second hybrid expert layer model parameters corresponding to each of the devices and running based on the expert parallel strategy includes: According to the respective expert numbers of each character, with the goal that the number of characters processed by each of the devices during the expert parallel computing process is equal and the character with the smallest expert number within the global scope is preferentially arranged locally in each of the devices, perform load balancing processing on each of the characters in each of the devices based on the greedy algorithm; In a full exchange communication manner, route each of the characters after the load balancing processing to each of the devices where their respective expert numbers are currently located, so that each of the devices respectively performs expert parallel computing on the characters received by each of them according to the second hybrid expert layer model parameters corresponding to each of them and running based on the expert parallel strategy.
7. The phased hybrid parallel inference method for the MoE sparse large model according to claim 6, characterized in that The performing load balancing processing on each of the characters in each of the devices according to the respective expert numbers of each character, with the goal that the number of characters processed by each of the devices during the expert parallel computing process is equal and the character with the smallest expert number within the global scope is preferentially arranged locally in each of the devices, based on the greedy algorithm includes: Control each of the devices to sort the respective characters locally in ascending order of the expert numbers to respectively determine the distribution information of each of the characters in each of the devices locally in each expert, and obtain the distribution information of each of the characters in each of the other devices locally in each expert in a all-gather communication manner to determine the global distribution information of each of the characters in each expert; Control each of the devices to respectively determine the character distribution information of each of the devices after routing the characters to each of the devices where their respective expert numbers are currently located according to the global distribution information of each of the characters in each expert; Control each of the devices to respectively determine the start address and offset of the characters sent by each of the devices and the start address and offset of the characters received by each of the devices according to the distribution information of each of the characters in each of the devices locally and the character distribution information of each of the devices after routing the characters to each of the devices where their respective expert numbers are currently located; Take the start address and offset of the characters sent by each of the devices, the start address and offset of the characters received by each of the devices, the send data buffer, and the receive data buffer as input parameters and input them into the full exchange communication function corresponding to the full exchange communication manner.
8. The phased hybrid parallel inference method for the MoE sparse large model according to claim 7, wherein, After determining the start address and offset of the characters sent by each of the devices and the start address and offset of the characters received by each of the devices, it further includes: If it is determined that there is a device with a missed expert after full exchange communication according to the offsets of the characters sent by each of the devices and the offsets of the received characters, before the device performs expert parallel computing on each of the characters based on the second hybrid expert layer model parameters corresponding to the missed expert, the second hybrid expert layer model parameters corresponding to the missed expert are loaded from the processor memory to the device.
9. A phased hybrid parallel inference system for MoE sparse large models, characterized in that, It includes: a scheduling device and each device that is communicatively connected to the scheduling device respectively; The scheduling device includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, it implements the phased hybrid parallel inference method of the MoE sparse large model according to any one of claims 1 to 8.
10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the phased hybrid parallel inference method of the MoE sparse large model according to any one of claims 1 to 8.
Citation Information
Patent Citations
Data processing method, end-side device, storage medium, chip system and computer program product
CN119831056A
Data processing method, apparatus and system, and medium and program product
WO2024066791A1