Staged hybrid parallel reasoning method and system of MoE sparse large model

In the inference process of MoE sparse big model, the pre-filling stage adopts expert parallelism, the decoding stage adopts tensor parallelism, and through layer-by-layer parallelism strategy transformation, the problem of insufficient adaptability of parallel strategies to the pre-filling and decoding stages in the existing technology is solved, and more efficient communication and computing performance is achieved.

CN120069097AActive Publication Date: 2025-05-30BEIJING UNIV OF POSTS & TELECOMM

Patent Information

Application Number
CN202510542935.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-28
Publication Date
2025-05-30
Estimated Expiration
2045-04-28

AI Technical Summary

Technical Problem

In the existing MoE sparse big model inference technology, the parallel strategy adopted by the distributed inference method has poor adaptability to the pre-filling stage and the decoding stage, resulting in a large communication overhead.

Method used

A phased mixed parallel inference method for MoE sparse big model is proposed. By adopting expert parallel strategy in the pre-filling stage, tensor parallel strategy is adopted in the decoding stage, and the communication overhead of parallel strategy transformation is hidden through layer-by-layer parallel strategy transformation.

Benefits of technology

It effectively reduces the communication overhead between devices in the pre-filling stage and the decoding stage, improves the adaptability to the two stages, and improves the efficiency of the MoE sparse big model inference process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120069097A_ABST
    Figure CN120069097A_ABST
Patent Text Reader

Abstract

The invention provides a staged hybrid parallel reasoning method and system for a MoE sparse large model, and relates to the technical field of specific calculation model systems.The method comprises the steps that the MoE sparse large model is controlled to execute layer by layer in the pre-filling stage; adding a first hybrid expert layer model parameter running based on a tensor parallel strategy to each device; performing expert parallel calculation based on second hybrid expert layer model parameters running based on an expert parallel strategy in each device; the characters are recovered to the initial position of the device, and the second hybrid expert layer model parameters are released; and sending a prediction character output by the last layer of the model to the first layer so as to execute reasoning of a decoding stage according to the prediction character and the first hybrid expert layer model parameter in each device. The problems that a parallel strategy adopted by an existing MoE sparse large model reasoning technology is poor in adaptability to a pre-filling stage and a decoding stage and large in communication overhead can be solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of computing systems based on specific computing models, and particularly to a phased hybrid parallel inference method and system for MoE sparse large models. Background Art

[0002] The MoE sparse large model (which can also be called: MoE hybrid expert sparse large model) refers to introducing a sparse activation mechanism on the basis of the mixture-of-expert model MoE (Mixture-of-Expert). Compared with the MoE mixture-of-expert model with a computational complexity close to that of a dense model, the MoE sparse large model replaces the multilayer perceptron MLP (MultilayerPerceptron) layer in the MoE mixture-of-expert model with a gating function and multiple experts, while the multi-head attention layer is the same as that of the MoE mixture-of-expert model. In the MoE sparse large model, each character only activates some experts for calculation, so as to achieve a sublinear growth of the computational overhead with the model capacity and achieve the effect of rapid expansion of model parameters. Compared with traditional dense large models, the MoE sparse large model can significantly reduce the computational cost under the same number of parameters. The computational process of large model inference can be divided into two processes, namely the prefill stage (Prefill) and the decoding stage (Decoding).

[0003] Currently, the parallel strategies adopted by the existing distributed inference methods for MoE sparse large model inference have poor adaptability to the prefill stage and the decoding stage. For example, DeepSpeed-MoE proposes a hybrid parallel strategy of "data parallelism + tensor parallelism + expert parallelism" and restricts expert parallelism within a node, so as to improve the communication efficiency of all-to-all communication by using the high bandwidth within the node. Since DeepSpeed-MoE does not distinguish the parallel strategies for the prefill stage and the decoding stage, and the prefill stage processes a large amount of data, while the decoding stage has a low computational density and processes a small amount of data, it will cause a large number of all-gather and reduce-scatter communication operations on data when using the tensor parallel strategy in the prefill stage, thus increasing the communication overhead, which is more serious when dealing with a large batch of data; at the same time, it will also make the all-to-all communication overhead of the decoding stage more significant when using the expert parallel strategy.

[0004] Therefore, there is a need to design a MoE sparse large model inference method that can solve the problems of poor adaptability of the parallel strategy adopted by the existing MoE sparse large model inference technology to the prefill stage and the decoding stage and large communication overhead. Summary of the Invention

[0005] In view of this, the embodiments of the present application provide a phased hybrid parallel inference method and system for MoE sparse large models to eliminate or improve one or more defects existing in the prior art.

[0006] One aspect of the present application provides a phased hybrid parallel inference method for MoE sparse large models, including: In the pre-filling stage of data inference based on the MoE sparse large model, controlling each architecture layer of the MoE sparse large model to execute the first step layer by layer; wherein, the first step includes: based on the multi-head attention layer model parameters and gating functions in each device, obtaining the respective expert numbers of each character corresponding to the prompt data sequence currently located at its respective initial position in the device, and adding the first mixture-of-experts layer model parameters operating based on the tensor parallel strategy to each device; then performing expert parallel computing on each character according to the respective expert numbers of each character and the second mixture-of-experts layer model parameters operating based on the expert parallel strategy in each device; and then restoring each character after expert parallel computing to its respective initial position in the device and releasing the second mixture-of-experts layer model parameters in each device; Sending the predicted characters output by the last architecture layer of the MoE sparse large model to the multi-head attention layer in the first architecture layer of the MoE sparse large model for executing the decoding stage of data inference of the MoE sparse large model according to the predicted characters, the multi-head attention layer model parameters in each device, the gating function, and the first mixture-of-experts layer model parameters operating based on the tensor parallel strategy in each device.

[0007] In some embodiments of the present application, the phased hybrid parallel inference method for the MoE sparse large model further includes: In each iteration round of the decoding stage of data inference based on the MoE sparse large model, controlling each architecture layer of the MoE sparse large model to execute the second step layer by layer, and sending the predicted characters output by the last architecture layer of the MoE sparse large model in each iteration round to the multi-head attention layer in the first architecture layer of the MoE sparse large model in a full-swap communication manner; Wherein, the second step includes: In the current iteration round, based on the multi-head attention layer model parameters and gating functions in each of the devices, obtain the expert numbers of each character corresponding to the latest data sequence, where each character is currently located at its respective initial position in its corresponding device; wherein, if the current iteration round is the first round, the latest data sequence includes the prompt data sequence cached in the pre-fill stage and the predicted character output by the last architecture layer of the MoE sparse large model in the pre-fill stage; if the current iteration round is not the first round, the latest data sequence includes the prompt data sequence, the predicted character output in the pre-fill stage, and the predicted characters output in the previous historical iteration rounds before the current iteration round; Collect each of the characters to each device in an all-gather communication manner, and according to the expert number of each character, enable each device to perform tensor parallel computing on the collected characters respectively according to the first mixture-of-experts layer model parameters running based on the tensor parallel strategy corresponding to each device; Restore each of the characters after tensor parallel computing to their respective initial positions in their corresponding devices in a reduce-scatter communication manner.

[0008] In some embodiments of the present application, the obtaining of the expert numbers of each character corresponding to the prompt data sequence, where each character is currently located at its respective initial position in its corresponding device, based on the multi-head attention layer model parameters and gating functions in each of the devices, includes: Divide the prompt data sequence into multiple first data subsequences evenly according to the total number of each of the devices; Input each of the first data subsequences into each device one-to-one, so that each device performs multi-head attention calculation on the first data subsequence received by it respectively based on the multi-head attention layer model parameters stored in it, to obtain each character corresponding to the first data subsequence received by it respectively, and determine the expert number of each character on each device respectively based on the gating function; Correspondingly, the obtaining of the expert numbers of each character corresponding to the latest data sequence, where each character is currently located at its respective initial position in its corresponding device, based on the multi-head attention layer model parameters and gating functions in each of the devices, includes: Divide the latest data sequence into multiple second data subsequences evenly according to the total number of each of the devices; Input each of the second data subsequences one-to-one into each device, so that each device performs multi-head attention calculation on the received second data subsequence based on the multi-head attention layer model parameters stored in each device respectively, to obtain each character corresponding to the received second data subsequence respectively, and collect all the characters to each device in an all-gather communication manner, so that each device determines the expert number of each of the characters on each device based on the gating function respectively.

[0009] In some embodiments of the present application, the adding the first mixture-of-experts layer model parameters operating based on the tensor parallel strategy to each device includes: If the time required to add the first mixture-of-experts layer model parameters operating based on the tensor parallel strategy to each device is shorter than or equal to the time required to obtain the expert number of each character currently located at the initial position of its respective device corresponding to the prompt data sequence, then while obtaining the expert number of each character currently located at the initial position of its respective device corresponding to the prompt data sequence, add the first mixture-of-experts layer model parameters operating based on the tensor parallel strategy to each device.

[0010] In some embodiments of the present application, before restoring each of the characters after the expert parallel calculation to its respective device initial position in the first step, it further includes: If the time required to add the first mixture-of-experts layer model parameters operating based on the tensor parallel strategy to each device is longer than the time required to obtain the expert number of each character currently located at the initial position of its respective device corresponding to the prompt data sequence, then divide the first mixture-of-experts layer model parameters corresponding to each device operating based on the tensor parallel strategy into a first part of parameters and a second part of parameters; Based on the multi-head attention layer model parameters and the gating function in each device, obtain the expert number of each character currently located at the initial position of its respective device corresponding to the prompt data sequence, and add the first part of parameters to each device at the same time; Then perform expert parallel calculation on each character according to the expert number of each character and the second mixture-of-experts layer model parameters corresponding to each device operating based on the expert parallel strategy, and add the second part of parameters to each device at the same time.

[0011] In some embodiments of the present application, the performing expert parallel calculation on each character according to the expert number of each character and the second mixture-of-experts layer model parameters corresponding to each device operating based on the expert parallel strategy includes: Based on the respective expert numbers of each character, with the goal of equalizing the number of characters processed by each device during expert parallel computing and preferentially arranging the character with the smallest expert number within the global scope locally in each device, perform load balancing processing on each character in each device based on the greedy algorithm; In a full exchange communication manner, route each character after load balancing processing to the respective device where its expert number is currently located, so that each device performs expert parallel computing on the characters received by it according to the second hybrid expert layer model parameters corresponding to its respective expert parallel strategy.

[0012] In some embodiments of the present application, the load balancing processing of each character in each device based on the greedy algorithm, with the goal of equalizing the number of characters processed by each device and preferentially arranging the character with the smallest expert number within the global scope locally in each device, based on the respective expert numbers of each character, includes: Control each device to sort the respective characters locally in ascending order of expert number to respectively determine the distribution information of each character locally in each device among each expert, and obtain the distribution information of each character locally in other devices among each expert in a collective gather communication manner to determine the global distribution information of each character among each expert; Control each device to respectively determine the character distribution information of each device after routing the characters to the respective devices where their expert numbers are currently located according to the global distribution information of each character among each expert; Control each device to determine the starting address and offset of the characters sent by each device and the starting address and offset of the characters received by each device according to the distribution information of each character locally in each device and the character distribution information of each device after routing the characters to the respective devices where their expert numbers are currently located; Take the starting address and offset of the characters sent by each device, the starting address and offset of the characters received by each device, the send data buffer, and the receive data buffer as input parameters and input them into the full exchange communication function corresponding to the full exchange communication manner.

[0013] In some embodiments of the present application, after determining the starting address and offset of the characters sent by each device and the starting address and offset of the characters received by each device, it further includes: If it is determined that there is a device that misses the expert after full exchange communication based on the offset of the characters sent by each device and the offset of the received characters, before the device performs expert parallel calculation on each character based on the second mixture-of-experts layer model parameters corresponding to the missed expert, the second mixture-of-experts layer model parameters corresponding to the missed expert are loaded from the processor memory to the device.

[0014] Another aspect of the present application provides a phased hybrid parallel inference device for a MoE sparse large model, including: A prefill stage execution module, configured to control each architecture layer of the MoE sparse large model to sequentially execute a first step in the prefill stage of data inference based on the MoE sparse large model; wherein, the first step includes: based on the multi-head attention layer model parameters and gating functions in each device, obtaining the respective expert numbers of each character corresponding to the prompt data sequence currently located at their respective initial positions in their respective devices, and adding the first mixture-of-experts layer model parameters operating based on the tensor parallel strategy to each device; then performing expert parallel calculation on each character according to the respective expert numbers of each character and the second mixture-of-experts layer model parameters operating based on the expert parallel strategy in each device; and then restoring each character after expert parallel calculation to its respective device initial position and releasing the second mixture-of-experts layer model parameters in each device. A stage transition execution module, configured to send the predicted characters output by the last architecture layer of the MoE sparse large model to the multi-head attention layer in the first architecture layer of the MoE sparse large model for performing the decoding stage of data inference of the MoE sparse large model according to the predicted characters, the multi-head attention layer model parameters in each device, the gating function, and the first mixture-of-experts layer model parameters operating based on the tensor parallel strategy in each device.

[0015] In some embodiments of the present application, the phased hybrid parallel inference device for the MoE sparse large model further includes: A decoding stage execution module, configured to control each architecture layer of the MoE sparse large model to sequentially execute a second step in each iteration round of the decoding stage of data inference based on the MoE sparse large model, and send the predicted characters output by the last architecture layer of the MoE sparse large model in each iteration round to the multi-head attention layer in the first architecture layer of the MoE sparse large model in a full exchange communication manner. Wherein, the second step includes: In the current iteration round, based on the multi-head attention layer model parameters and the gating function in each of the devices, obtain the respective expert numbers of each character corresponding to the latest data sequence and initially located at its respective device initial position; wherein, if the current iteration round is the first round, the latest data sequence includes the prompt data sequence cached in the pre-fill stage and the predicted character output by the last architecture layer of the MoE sparse large model in the pre-fill stage; if the current iteration round is not the first round, the latest data sequence includes the prompt data sequence, the predicted character output in the pre-fill stage, and the predicted characters output in the historical iteration rounds before the current iteration round. According to the respective expert numbers of each character, route each of the characters to the respective devices where their expert numbers are currently located in a full exchange communication manner, so that each of the devices respectively performs tensor parallel computing on the characters received by them according to the first mixture-of-experts layer model parameters operating based on the tensor parallel strategy. Restore each of the characters after tensor parallel computing to their respective device initial positions in a reduction distribution communication manner.

[0016] The third aspect of the present application provides a phased hybrid parallel inference system for a MoE sparse large model, including: a scheduling device and each device respectively communicatively connected to the scheduling device; the scheduling device includes a memory, a processor, and a computer program stored on the memory and executable on the processor, and when the processor executes the computer program, it implements the phased hybrid parallel inference method of the MoE sparse large model described in the foregoing first aspect.

[0017] The fourth aspect of the present application provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, and when the processor executes the computer program, it implements the phased hybrid parallel inference method of the MoE sparse large model.

[0018] The fifth aspect of the present application provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it implements the phased hybrid parallel inference method of the MoE sparse large model.

[0019] The sixth aspect of the present application provides a computer program product, including a computer program, and when the computer program is executed by a processor, it implements the phased hybrid parallel inference method of the MoE sparse large model.

[0020] The phased hybrid parallel inference method for the MoE sparse large model provided by this application controls each architecture layer of the MoE sparse large model to execute the first step layer by layer during the pre-filling stage of data inference based on the MoE sparse large model; wherein, the first step includes: based on the multi-head attention layer model parameters and gating functions in each device, obtaining the respective expert numbers of each character corresponding to the prompt data sequence currently located at their respective initial positions in their respective devices, and adding the first hybrid expert layer model parameters operating based on the tensor parallel strategy to each device; then performing expert parallel calculation on each character according to the respective expert numbers of each character and the second hybrid expert layer model parameters operating based on the expert parallel strategy in each device; then restoring each character after expert parallel calculation to its respective initial position in the device and releasing the second hybrid expert layer model parameters in each device; sending the predicted character output by the last architecture layer of the MoE sparse large model to the multi-head attention layer in the first architecture layer of the MoE sparse large model for executing the decoding stage of data inference of the MoE sparse large model according to this predicted character, the multi-head attention layer model parameters in each device, the gating function, and the first hybrid expert layer model parameters operating based on the tensor parallel strategy in each device; that is to say, for the pre-filling stage with a high data volume, the hybrid expert layer adopts the expert parallel strategy with relatively lower communication volume to effectively reduce the communication overhead between devices during the pre-filling stage; and for the decoding stage with a low data volume, the hybrid expert layer adopts the more efficient tensor parallel strategy to effectively reduce the communication overhead between devices during the decoding stage, which can effectively improve the adaptability to the pre-filling stage and the decoding stage and improve the efficiency of the MoE sparse large model inference process; and, by adding the first hybrid expert layer model parameters operating based on the tensor parallel strategy to each device while performing multi-head attention layer calculation during the pre-filling stage, the communication overhead of parallel strategy conversion can be hidden in a layer-by-layer conversion manner without significantly increasing the device resource occupancy rate; and the second hybrid expert layer model parameters operating based on the expert parallel strategy can be released layer by layer after the calculation of the hybrid expert layer in each structural layer during the pre-filling stage, which can further avoid the increase in device resource occupancy rate.

[0021] Additional advantages, objects, and features of this application will be partially described below and will become partially apparent to those of ordinary skill in the art after studying the following text, or may be learned from the practice of this application. The objects and other advantages of this application can be achieved and obtained by the structures specifically pointed out in the specification and the drawings.

[0022] Those skilled in the art will understand that the objectives and advantages achievable with this application are not limited to those specifically described above, and the above and other objectives achievable with this application will be more clearly understood based on the following detailed description. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] The drawings described herein are used to provide a further understanding of this application, form a part of this application, and do not limit this application. The components in the drawings are not drawn to scale, but only to illustrate the principles of this application. To facilitate the illustration and description of some parts of this application, the corresponding parts in the drawings may be enlarged, that is, may become larger relative to other components in the exemplary device actually manufactured according to this application. In the drawings: Figure 1 It is a schematic diagram of the architecture layer of the MoE sparse large model.

[0024] Figure 2 It is a schematic diagram of the character routing (1) from data parallelism of the multi-head attention layer to expert parallelism of the mixture-of-experts layer with 4 devices and 8 experts as an example in the prior art.

[0025] Figure 3 It is a schematic diagram of the character routing (2) from expert parallelism of the mixture-of-experts layer to data parallelism of the multi-head attention layer with 4 devices and 8 experts as an example in the prior art.

[0026] Figure 4 It is the first flow schematic diagram of the phased hybrid parallel inference method of the MoE sparse large model in an embodiment of this application.

[0027] Figure 5 It is an example diagram of the MoE sparse large model with two basic architecture layers in series and 8 experts in each basic architecture layer.

[0028] Figure 6 It is the first flow schematic diagram of the first step in the phased hybrid parallel inference method of the MoE sparse large model in an embodiment of this application.

[0029] Figure 7 It is the second flow schematic diagram of the phased hybrid parallel inference method of the MoE sparse large model in an embodiment of this application.

[0030] Figure 8 It is the first flow schematic diagram of the second step in the phased hybrid parallel inference method of the MoE sparse large model in an embodiment of this application.

[0031] Figure 9 It is the second flow schematic diagram of the first step in the phased hybrid parallel inference method of the MoE sparse large model in an embodiment of this application.

[0032] Figure 10 It is a schematic diagram of the second process of the second step in the phased hybrid parallel inference method of the MoE sparse large model in an embodiment of the present application.

[0033] Figure 11 It is a schematic diagram of the third process of the first step in the phased hybrid parallel inference method of the MoE sparse large model in an embodiment of the present application.

[0034] Figure 12 It is a schematic diagram of the fourth process of the first step in the phased hybrid parallel inference method of the MoE sparse large model in an embodiment of the present application.

[0035] Figure 13 It is a schematic diagram of character routing (1) from data parallelism of the multi-head attention layer to expert parallelism of the mixture of experts layer with 4 devices and 8 experts as an example in an embodiment of the present application.

[0036] Figure 14 It is a schematic diagram of the phased hybrid parallel strategy of the MoE mixture of experts layer provided by the application example of the present application.

[0037] Figure 15 It is a schematic diagram of the parallel strategy in the prefill stage with 4 devices as an example provided by the application example of the present application.

[0038] Figure 16 It is a schematic diagram of the parallel strategy in the decoding stage with 4 devices as an example provided by the application example of the present application.

[0039] Figure 17 It is a schematic diagram of character routing (2) from expert parallelism of the mixture of experts layer to data parallelism of the multi-head attention layer with 4 devices and 8 experts as an example provided by the application example of the present application.

[0040] Figure 18 It is a schematic diagram of the data transmission overhead hiding technology provided by the application example of the present application.

[0041] Figure 19 (a) is a schematic diagram comparing the performance test results of the prefill stage between the present application and DeepSeek-MoE when batchsize (batch size) = 128 and sequence_length (sequence length) = 2048 provided by the application example of the present application.

[0042] Figure 19 (b) is a schematic diagram comparing the performance test results of the prefill stage between the present application and DeepSeek-MoE when batchsize (batch size) = 256 and sequence_length (sequence length) = 2048 provided by the application example of the present application.

[0043] Figure 20(a) is a schematic diagram comparing the performance test results of the decoding stage between this application and DeepSeek-MoE when batchsize (batch size) = 128 and 100 tokens are output during inference in the application example of this application.

[0044] Figure 20(b) is a schematic diagram comparing the performance test results of the decoding stage between this application and DeepSeek-MoE when batchsize (batch size) = 256 and 100 tokens are output during inference in the application example of this application.

[0045] Figure 21(a) is a schematic diagram comparing the performance test results of the end-to-end inference (prefill + decoding) between this application and DeepSeek-MoE when batchsize (batch size) = 128, sequence_length (sequence length) = 2048, and 100 tokens are output during inference in the application example of this application.

[0046] Figure 21(b) is a schematic diagram comparing the performance test results of the end-to-end inference (prefill + decoding) between this application and DeepSeek-MoE when batchsize (batch size) = 256, sequence_length (sequence length) = 2048, and 100 tokens are output during inference in the application example of this application. Detailed implementation manners

[0047] To make the objectives, technical solutions, and advantages of this application clearer and more understandable, the following further elaborates on this application in combination with the implementation manners and the accompanying drawings. Herein, the illustrative implementation manners and descriptions of this application are used to explain this application, but do not limit this application.

[0048] Herein, it should also be noted that to avoid obscuring this application with unnecessary details, only the structures and / or processing steps closely related to the solution of this application are shown in the drawings, while other details less relevant to this application are omitted.

[0049] It should be emphasized that the term "including / containing" when used herein refers to the existence of features, elements, steps, or components, but does not exclude the existence or addition of one or more other features, elements, steps, or components.

[0050] Herein, it should also be noted that if not otherwise specified, the term "connection" in this text can not only refer to a direct connection but also represent an indirect connection with an intermediate.

[0051] In the following, embodiments of the present application will be described with reference to the accompanying drawings. In the drawings, the same reference numerals represent the same or similar components, or the same or similar steps.

[0052] In recent years, the parameter scale of large models has grown rapidly. Training large models requires a large amount of computing resources. For example, taking GPT-4 with 1.8 trillion parameters as an example, the computing power required to train GPT-4 is equivalent to running on 25,000 NVIDIA A100 GPUs for 90 to 100 days. The MoE sparse large model can effectively reduce the training overhead of large models and achieve sublinear growth of the computing overhead with the model capacity. The MoE sparse large model consists of multiple sequentially connected architecture layers (which can also be called: the basic architecture layer of the MoE large model). Among them, as Figure 1 shown, an architecture layer is composed of a connected multi-head attention layer and a mixture-of-experts layer (which can also be called: the MoE mixture-of-experts layer). This architecture layer can be repeated multiple times and connected in series to form a multi-layer deep neural network. FFN 0 、FFN 1 and FFN E-1 respectively represent the mixture-of-experts layer model parameters corresponding to different experts. Among them, Query represents the "question" or "target" feature to be focused on, which can be abbreviated as Q; Key represents the reference feature providing the "position" or "matching basis", which can be abbreviated as K; Value represents carrying the actual information for weighted aggregation, which can be abbreviated as V; QKV LinearProjection (QKV linear projection) is the core operation in the attention mechanism to generate Query (Q), Key (K), and Value (V). It maps the input data to different semantic spaces through linear transformation. In the MoE sparse large model, after each character (token) passes through the gating function, the affinity score of the character for each expert will be calculated, and then the top K experts with the highest affinity scores will be selected for calculation, K≥1. Each character does not need to be calculated by all experts, thus significantly reducing the computing cost. Current mainstream large models such as DeepSeek and Mixtral all adopt the MoE mixture-of-experts architecture to reduce costs.

[0053] The MoE sparse large model introduces an expert parallel model. For the mixture-of-experts layer in the MoE sparse large model, expert parallelism is achieved by evenly dividing all experts onto different devices to split the expert parameters; while for the multi-head attention layer, it can be in the form of data parallelism, that is, the complete multi-head attention layer model parameters are stored on each device. The input characters are divided onto each device for multi-head attention layer calculation, and then according to the calculation result of the affinity score of the gating function, through an all-to-all communication operation, the characters are routed to the devices where the corresponding experts are located for MoE mixture-of-experts layer calculation, asFigure 2 As shown in Figure 2 , 8 experts are evenly distributed to 4 devices. Figure 2 In , each grid represents a character (token), and the number in the grid represents the expert number to be routed (also known as: destination expert number or expert ID, etc.), and each of the said expert numbers corresponds one-to-one with each of the said experts. After the calculation of the mixture-of-experts layer is completed, the characters are routed back to the original devices through an all-to-all communication operation, as Figure 3 shown, that is, each character after the expert parallel calculation is restored to its initial device position in the multi-head attention layer calculation, so as to perform data parallel calculation in the next multi-head attention layer.

[0054] The MoE mixture-of-experts model can significantly reduce the calculation cost of large models. However, the MoE sparse large model still faces challenges in inference performance. The calculation process of large model inference can be divided into two processes, namely the prefill stage and the decoding stage. The prefill stage refers to the calculation process of generating the first character by inputting prompt words (such as user questions) to the model. The decoding stage refers to the calculation process of iteratively generating the next character based on the current input sequence and the latest generated character after the first character is generated. Decoding is an iterative calculation process, generating a new character in each iteration until the maximum generation length is reached or a termination character is encountered. In the PD-integrated (such as DeepSpeed-MoE) inference scheme, the prefill stage and the decoding stage share the same set of calculation devices. However, the prefill stage is a compute-intensive load, while the decoding stage is a memory-intensive load, and the two stages will interfere with each other in the PD-integrated inference scheme. Therefore, some research has proposed the PD-separation technology, where the prefill stage and the decoding stage use different calculation devices respectively, separating them physically, eliminating the interference between the two stages, enabling each stage to focus on its respective optimization goals, that is, the prefill stage focuses on optimizing the first-character latency metric, while the decoding stage focuses on optimizing the output-character latency metric. However, in the PD-separation technology, the prefill stage and the decoding stage use different calculation devices respectively, so the PD-separation technology requires more computing resources. In addition, in the PD-separation technology, after the prefill stage calculation is completed, the KV-cache calculated in the prefill stage needs to be transmitted to the decoding device before the decoding device can perform subsequent calculations, so there will be additional data transmission overhead. Therefore, both the PD-separation and PD-integrated schemes have their own advantages and disadvantages.

[0055] In the expert parallelism of MoE sparse large model distributed inference, the load imbalance problem caused by the dynamic routing of characters is one of the key factors affecting inference performance. As Figure 2As shown, in the character routing from data parallelism in the multi-head attention layer to expert parallelism in the MoE layer, the dynamic routing of characters causes the number of characters processed on each device to be different, and at the same time, the all-to-all communication load is also unbalanced, that is, the number of characters received on each device is different. As Figure 3 As shown, in the character routing from expert parallelism in the MoE layer to data parallelism in the multi-head attention layer, the dynamic routing of characters causes the all-to-all communication load to be unbalanced, that is, the number of characters sent on each device is different. Therefore, in the inference of the MoE sparse large model, the dynamic routing of characters causes the computational load and the communication load to be unbalanced. To alleviate the load imbalance problem, Gshard (an efficient data processing technology designed based on the concept of distributed computing) proposed the design of the expert capacity upper limit, that is, characters exceeding the expert capacity will directly skip the mixture-of-experts layer using the residual network, while experts that do not reach the capacity upper limit will be padded to make up the difference. Although this strategy can balance the computational and communication loads, the disadvantage is that it will generate useless computations, and at the same time, due to the discarding of characters during inference, there may be a loss of accuracy, which is difficult to accept during inference. DeepSeekV3 adopts a load balancing strategy without auxiliary loss. This strategy introduces an adjustable bias value for each expert, and this bias value is added to the expert affinity score calculated by the gating function, so as to adjust the load on each expert. For experts with high load, the bias value will be reduced, and for experts with low load, the bias value will be increased, so as to make the load on each expert as balanced as possible. However, this strategy cannot completely solve the load imbalance problem.

[0056] That is to say, for the distributed inference of the MoE sparse large model, the parallel strategies adopted by the existing technologies have poor adaptability to the pre-padding stage and the decoding stage, especially the PD integrated inference scheme. DeepSpeed-MoE is an optimization technology for the mixture-of-experts (MoE) network, aiming to improve the training and inference efficiency of large language models (LLMs). DeepSpeed-MoE proposes a hybrid parallel strategy of "data parallelism + tensor parallelism + expert parallelism", and limits the expert parallelism within the node, so as to improve the all-to-all communication efficiency by using the high bandwidth within the node. However, the parallel strategy adopted by DeepSpeed-MoE does not distinguish between the pre-padding stage and the decoding stage, and the parallel strategy remains unchanged throughout the inference process, which is not suitable for the computational load characteristics of the two stages.

[0057] Based on this, in order to solve the problems that the parallel strategies adopted by existing MoE sparse large model inference technologies have poor adaptability to the pre-filling stage and the decoding stage and large communication overhead, etc., the embodiments of the present application respectively provide a phased hybrid parallel inference method for MoE sparse large models, a phased hybrid parallel inference device for MoE sparse large models for executing the phased hybrid parallel inference method of MoE sparse large models, a phased hybrid parallel inference system for MoE sparse large models, an electronic device, a computer-readable storage medium, and a computer program product. By adopting a phased hybrid parallel inference method, different parallel strategies adapted to different stages are used, and efficient connection is carried out through a layer-by-layer parallel strategy conversion method, which can achieve better performance compared with existing MoE sparse large model inference technologies.

[0058] Specific details are described in detail through the following embodiments.

[0059] Based on this, the embodiments of the present application provide a phased hybrid parallel inference method for MoE sparse large models that can be implemented by a phased hybrid parallel inference device for MoE sparse large models. Refer to Figure 4 The phased hybrid parallel inference method for MoE sparse large models specifically includes the following contents: Step 100: In the pre-filling stage of data inference based on the MoE sparse large model, control each architecture layer of the MoE sparse large model to execute the first step layer by layer; wherein, the first step includes: based on the multi-head attention layer model parameters and gating functions in each device, obtain the expert numbers of each character corresponding to the prompt data sequence currently located at their respective initial positions in their respective devices, and at the same time add the first hybrid expert layer model parameters operating based on the tensor parallel strategy to each device; then perform expert parallel calculation on each character according to the expert numbers of each character and the second hybrid expert layer model parameters operating based on the expert parallel strategy in each device; then restore each character after expert parallel calculation to its respective initial position in the device and release the second hybrid expert layer model parameters in each device.

[0060] In one or more embodiments of the present application, refer to Figure 5 An architecture layer refers to the basic architecture layer of the MoE sparse large model, and can also be called the basic architecture layer of the MoE large model. The MoE sparse large model is composed of multiple architecture layers connected in series in sequence. The first step is executed layer by layer for each architecture layer. Here, the layer-by-layer is based on the architecture layer as a unit. Each architecture layer contains a multi-head attention layer and a hybrid expert layer, and the output end of the multi-head attention layer is connected to the input end of the hybrid expert layer, and the output end of the hybrid expert layer is connected to the input end of the multi-head attention layer in the next architecture layer.

[0061] It should be noted that, for the pre-filling stage with a high data volume to be processed, the embodiments of the present application adopt an expert parallel strategy with relatively lower communication volume for the mixture-of-experts layer, so as to effectively reduce the communication overhead between devices in the pre-filling stage; while for the decoding stage with a low data volume to be processed, the embodiments of the present application adopt a more efficient tensor parallel strategy for the mixture-of-experts layer, so as to effectively reduce the communication overhead between devices in the decoding stage, which can effectively improve the adaptability to the pre-filling stage and the decoding stage and improve the efficiency of the inference process of the MoE sparse large model. However, since the mixture-of-experts layer adopts different parallel strategies in the pre-filling stage and the decoding stage respectively, it is necessary to change the parallel strategy of the mixture-of-experts layer from the expert parallel strategy to the tensor parallel strategy before executing the decoding stage. And changing the expert parallel strategy to the tensor parallel strategy will result in additional communication overhead between devices; if this communication process is to be avoided, it is necessary to store the first mixture-of-experts layer model parameters based on the tensor parallel strategy and the second mixture-of-experts layer model parameters based on the expert parallel strategy corresponding to each mixture-of-experts layer in each architecture layer in each device at the same time, which will in turn cause a significant increase in the data occupancy rate of each device.

[0062] Based on this, in order to avoid the additional communication overhead between devices caused by changing the expert parallel strategy to the tensor parallel strategy and avoid a significant increase in the data occupancy rate of each device, in step 100 of the phased hybrid parallel inference method of the MoE sparse large model provided by the embodiments of the present application, a first step is designed layer by layer for each architecture layer in the pre-filling stage, see Figure 6 The first step specifically includes the following contents: Step 110: Based on the multi-head attention layer model parameters and gating functions in each device, obtain the expert numbers of each character corresponding to the prompt data sequence currently located at its respective initial position in its own device, and add the first mixture-of-experts layer model parameters based on the tensor parallel strategy to each device.

[0063] It can be understood that the prompt data sequence refers to a text sequence used as prompt information for predicting characters, and its data type can also be set according to actual application needs. In one example, the prompt data sequence can be user question text data. The device initial position refers to the position of each character in the device at this time, and for the convenience of subsequent description, it is referred to as the device initial position.

[0064] Among them, the multi-head attention layer model parameters refer to the model parameters of the multi-head attention layer, and the model parameters of each multi-head attention layer in each architecture layer are stored separately in each device.

[0065] In one or more embodiments of the present application, the tensor parallel strategy refers to: evenly dividing the high feature dimension H of all E experts onto P devices, where each device has [E, H / P] experts, that is, each device has H / P parts of all E experts, and each device needs to process all characters. The first mixture-of-experts layer model parameters operating based on the tensor parallel strategy refer to the mixture-of-experts layer model parameters corresponding to [E, H / P] owned by each device; and the first mixture-of-experts layer model parameters operating based on the tensor parallel strategy are only used in the decoding phase, and the current prefill phase is only used to hide the communication overhead of setting the first mixture-of-experts layer model parameters operating based on the tensor parallel strategy in the mixture-of-experts layer.

[0066] Step 120: Perform expert parallel computing on each of the characters according to the expert numbers of each character and the second mixture-of-experts layer model parameters operating based on the expert parallel strategy in each of the devices corresponding thereto.

[0067] It can be understood that when step 110 is executed, the second mixture-of-experts layer model parameters operating based on the expert parallel strategy corresponding to at least one expert (or expert number) already exist in each of the devices, and thus the second mixture-of-experts layer model parameters can be directly called to perform mixture-of-experts layer computing on each of the received characters when step 120 is executed. Among them, the expert parallel strategy refers to: evenly dividing all E experts onto P devices. Assuming the high feature dimension of each expert is H, each device has [E / P, H] experts, where E / P is the number of experts on each device. And the second mixture-of-experts layer model parameters operating based on the expert parallel strategy refer to: the mixture-of-experts layer model parameters corresponding to [E / P, H] owned by each device.

[0068] Step 130: Restore each of the characters after expert parallel computing to their respective initial positions in the devices and release the second mixture-of-experts layer model parameters in each of the devices.

[0069] In step 130, restoring each of the characters after expert parallel computing to their respective initial positions in the devices means that after the expert parallel computing in the mixture-of-experts layer is completed, the characters are routed back to the original devices through a full exchange communication operation, so as to perform data parallel computing on the multi-head attention layer in the next layer.

[0070] That is to say, in the above first step, during the process of performing multi-head attention calculation by the multi-head attention layer in the current architecture layer, the model parameters of the first mixture-of-experts layer operating based on the tensor parallel strategy corresponding to the multi-head attention layer in this architecture layer are added to each of the devices simultaneously, so that the model parameters of the first mixture-of-experts layer added to each of the devices form the mixture-of-experts layer operating based on the tensor parallel strategy that is uniquely corresponding to the multi-head attention layer in this architecture layer. It should be noted that in the current state, each of the devices stores simultaneously the model parameters of the multi-head attention layer corresponding to each architecture layer respectively, the model parameters of the second mixture-of-experts layer operating based on the expert parallel strategy corresponding to the current architecture layer and the subsequent architecture layers respectively, and the model parameters of the first mixture-of-experts layer operating based on the tensor parallel strategy corresponding to the current architecture layer and the previous architecture layers respectively. The subsequent architecture layers refer to all the architecture layers after the current architecture layer. If the current architecture layer is the last layer of the MoE sparse large model, then its subsequent architecture layers are considered non-existent. The previous architecture layers refer to all the architecture layers before the current architecture layer. If the current architecture layer is the first layer of the MoE sparse large model, then its previous architecture layers are considered non-existent.

[0071] In an example, if the MoE sparse large model includes 60 sequentially connected architecture layers, and the 30th architecture layer is currently executing the first step, then in step 110, while controlling each device to respectively obtain the expert numbers of each character corresponding to the prompt data sequence at the initial positions of their respective devices based on the multi-head attention layer model parameters and gating functions of the 30th layer locally, the first hybrid expert layer model parameters corresponding to the 30th layer's hybrid expert layer running based on the tensor parallel strategy are added to each device. Then, in step 120, each device is controlled to perform expert parallel computing on each character based on the second hybrid expert layer model parameters of the 30th layer running based on the expert parallel strategy, and in step 130 after step 120, the second hybrid expert layer model parameters of the 30th layer running based on the expert parallel strategy in each device are released. At this time, each device stores the multi-head attention layer model parameters of the 60-layer multi-head attention layer, the second hybrid expert layer model parameters corresponding to the hybrid expert layers of the 31st to 60th layers running based on the expert parallel strategy, and the first hybrid expert layer model parameters corresponding to the hybrid expert layers of the 1st to 30th layers running based on the tensor parallel strategy. That is to say, for the 60-layer hybrid expert layers in the 60 sequentially connected architecture layers included in the MoE sparse large model at this time, each device only needs to store the first hybrid expert layer model parameters running based on the tensor parallel strategy or the second hybrid expert layer model parameters running based on the expert parallel strategy. Each device may only store the first hybrid expert layer model parameters running based on the tensor parallel strategy and the second hybrid expert layer model parameters running based on the expert parallel strategy corresponding to the same hybrid expert layer simultaneously during the execution of step 120. This can not only hide the communication overhead between devices caused by adding the first hybrid expert layer model parameters running based on the tensor parallel strategy through the multi-head attention calculation process in the current architecture layer, but also effectively avoid a significant increase in the data occupancy rate of each device.

[0072] Step 200: Send the predicted characters output by the last architecture layer of the MoE sparse large model to the multi-head attention layer in the first architecture layer of the MoE sparse large model for performing the decoding stage of data inference of the MoE sparse large model according to the predicted characters, the multi-head attention layer model parameters in each device, the gating function, and the first hybrid expert layer model parameters of each device running based on the tensor parallel strategy.

[0073] In step 200, the predicted character output by the last architecture layer of the MoE sparse large model refers to: the output sequence corresponding to the prompt data sequence finally output by the mixture-of-experts layer in the last architecture layer of the MoE sparse large model based on the respective different second mixture-of-experts layer model parameters in each device, and then the last character is extracted from this output sequence as the predicted character obtained in the prefill stage. This predicted character will jointly form a predicted string with the predicted characters output in each iteration of the decoding stage to serve as the response text data corresponding to the prompt data sequence such as the user question text data.

[0074] As can be seen from the above description, in the staged hybrid parallel inference method of the MoE sparse large model provided by the embodiments of the present application, for the prefill stage with a high data volume, the mixture-of-experts layer adopts an expert parallel strategy with relatively lower communication volume to effectively reduce the communication overhead between devices in the prefill stage; while for the decoding stage with a low data volume, the mixture-of-experts layer adopts a more efficient tensor parallel strategy to effectively reduce the communication overhead between devices in the decoding stage, which can effectively improve the adaptability to the prefill stage and the decoding stage and improve the efficiency of the inference process of the MoE sparse large model; and, by adding the first mixture-of-experts layer model parameters operating based on the tensor parallel strategy to each device while performing the multi-head attention layer calculation in the prefill stage, the communication overhead of the parallel strategy conversion can be hidden in a layer-by-layer conversion manner without significantly increasing the device resource occupancy rate; and the second mixture-of-experts layer model parameters operating based on the expert parallel strategy can be released layer by layer after the calculation of the mixture-of-experts layer in each structural layer of the prefill stage, which can further avoid the increase of the device resource occupancy rate.

[0075] To further improve the adaptability of the decoding stage in the staged hybrid parallel inference process of the MoE sparse large model to further reduce the communication overhead between devices, in a staged hybrid parallel inference method of the MoE sparse large model provided by the embodiments of the present application, see Figure 7 , after step 200 in the staged hybrid parallel inference method of the MoE sparse large model, the following specific content is further included: Step 300: In each iteration round of the decoding stage of data inference based on the MoE sparse large model, control each architecture layer of the MoE sparse large model to execute the second step layer by layer, and send the predicted character output by the last architecture layer of the MoE sparse large model in each iteration round to the multi-head attention layer in the first architecture layer of the MoE sparse large model in a full exchange communication manner.

[0076] After the iteration ends, the predicted characters are output together to form a predicted string.

[0077] Among them, referring to Figure 8 , the second step specifically includes the following content: Step 310: In the current iteration round, based on the multi-head attention layer model parameters and gating functions in each of the devices, obtain the expert numbers of each character corresponding to the latest data sequence, which are initially located at their respective device positions; wherein, if the current iteration round is the first round, the latest data sequence includes the prompt data sequence cached in the pre-fill stage and the predicted character output by the last architecture layer of the MoE sparse large model in the pre-fill stage; if the current iteration round is not the first round, the latest data sequence includes the prompt data sequence, the predicted character output in the pre-fill stage, and the predicted characters output in the previous historical iteration rounds before the current iteration round.

[0078] Step 320: Collect each of the characters to each device in an all-gather communication manner, and according to the expert number of each character, enable each device to perform tensor parallel computing on the collected characters according to the first mixture-of-experts layer model parameters running based on the tensor parallel strategy corresponding to each device.

[0079] Step 330: Restore each of the characters after tensor parallel computing to their respective device initial positions in a reduce-scatter communication manner.

[0080] In order to further improve the reliability of multi-head attention calculation in the staged hybrid parallel inference process of the MoE sparse large model, in a method for staged hybrid parallel inference of a MoE sparse large model provided in an embodiment of the present application, referring to Figure 9 , step 110 in the method for staged hybrid parallel inference of the MoE sparse large model specifically includes the following content: Step 111: Divide the prompt data sequence into multiple first data subsequences evenly according to the total number of each of the devices.

[0081] Step 112: Input each of the first data subsequences into each device one-to-one, so that each device performs multi-head attention calculation on the first data subsequence received by itself based on the multi-head attention layer model parameters stored in itself, to respectively obtain each character corresponding to the first data subsequence received by itself, and respectively determine the expert number of each character on each device based on the gating function; While executing step 112, execute step 113: Add the first mixture-of-experts layer model parameters running based on the tensor parallel strategy to each of the devices simultaneously.

[0082] Specifically, in the pre-filling stage, the multi-head attention layer adopts a data parallelism strategy, that is, all input characters are divided among P devices for parallel processing, and the model parameters of the multi-head attention layer are replicated on each device. After the data parallel computing of the multi-head attention layer is completed, the calculation result of the affinity score of the gating function is calculated.

[0083] To further improve the reliability of multi-head attention calculation in the phased hybrid parallel inference process of the MoE sparse large model, in a phased hybrid parallel inference method of the MoE sparse large model provided in an embodiment of the present application, refer to Figure 10 In step 310 of the phased hybrid parallel inference method of the MoE sparse large model specifically includes the following contents: Step 311: In the current iteration round, according to the total number of each device, the latest data sequence is evenly divided into multiple second data subsequences.

[0084] Step 312: Each of the second data subsequences is input into each device one by one, so that each device performs multi-head attention calculation on the received second data subsequence based on the model parameters of the multi-head attention layer stored in each device, respectively obtaining each character corresponding to the received second data subsequence, and collecting all the characters to each device in an all-gather communication manner, so that each device determines the expert number of each character on each device based on the gating function.

[0085] Specifically, in the decoding stage, the multi-head attention layer of the model adopts data parallelism, that is, all input characters are divided among each device for parallel processing, and the model parameters of the multi-head attention layer are replicated on each device. After the data parallel computing of the multi-head attention layer is completed, all the characters are collected to each device through an all-gather communication operation, and the activated experts are selected according to the calculation result of the affinity score of the gating function.

[0086] To further improve the effectiveness and reliability of the communication overhead of the hidden parallelism strategy conversion, in a phased hybrid parallel inference method of the MoE sparse large model provided in an embodiment of the present application, refer to Figure 11 In step 113 of the phased hybrid parallel inference method of the MoE sparse large model specifically includes the following contents: Step 1131: If the time required to add the first mixture-of-experts layer model parameters running based on the tensor parallel strategy to each of the devices is shorter than or equal to the time required to obtain the expert numbers of the respective characters currently located at the initial positions of their respective devices for the prompt data sequence, then while obtaining the expert numbers of the respective characters currently located at the initial positions of their respective devices for the prompt data sequence, add the first mixture-of-experts layer model parameters running based on the tensor parallel strategy to each of the devices.

[0087] To further improve the applicability and reliability of the communication overhead of the hidden parallel strategy conversion, in a phased hybrid parallel inference method for a MoE sparse large model provided in an embodiment of the present application, refer to Figure 12 , in the phased hybrid parallel inference method for the MoE sparse large model, steps 113 and 120 in the first step can also be replaced with steps 140 to 160, which specifically include the following content: Step 140: If the time required to add the first mixture-of-experts layer model parameters running based on the tensor parallel strategy to each of the devices is longer than the time required to obtain the expert numbers of the respective characters currently located at the initial positions of their respective devices for the prompt data sequence, then divide the first mixture-of-experts layer model parameters corresponding to each of the devices into a first part of parameters and a second part of parameters.

[0088] Step 150: Based on the multi-head attention layer model parameters and gating functions in each device, obtain the expert numbers of the respective characters currently located at the initial positions of their respective devices for the prompt data sequence, and at the same time add the first part of parameters to each of the devices.

[0089] Step 160: Perform expert parallel calculation on each of the characters according to the expert numbers of the respective characters and the second mixture-of-experts layer model parameters corresponding to each of the devices running based on the expert parallel strategy, and at the same time add the second part of parameters to each of the devices.

[0090] That is to say, in the pre-fill stage, in addition to the process of adding the first mixture-of-experts layer model parameters running based on the tensor parallel strategy that can be hidden during the multi-head attention layer calculation process, the process of adding the first mixture-of-experts layer model parameters running based on the tensor parallel strategy can also be hidden during the calculation process of the mixture-of-experts layer running based on the expert parallel strategy, and it can be specifically determined according to the estimated communication duration required for the process of adding the first mixture-of-experts layer model parameters running based on the tensor parallel strategy.

[0091] It is understandable that the design of the expert capacity upper limit means that characters exceeding the expert capacity will directly skip the MoE mixture-of-experts layer using a residual network, while experts that do not reach the capacity upper limit will be padded. Although this strategy can balance the computational and communication loads, its disadvantage is that it will generate useless computations, and at the same time, due to the discarding of characters during inference, there may be a loss of accuracy, which is difficult to accept during inference. In DeepSeek V3, a load balancing strategy without auxiliary losses is adopted. This strategy introduces an adjustable bias value for each expert, and this bias value is added to the expert affinity score calculated by the gating function, thereby adjusting the load on each expert. For experts with high load, the bias value will be reduced, and for experts with low load, the bias value will be increased, but it can only make the loads on each expert as balanced as possible. Existing technologies mainly optimize load balancing at the algorithm level and cannot completely solve the problem of load imbalance. That is to say, for the distributed inference of MoE sparse large models, existing technologies cannot completely solve the problem of unbalanced expert parallel loads.

[0092] Based on this, in order to further solve the problem of unbalanced expert parallel loads in the pre-padding stage, the present application proposes a load balancing strategy based on the greedy algorithm. Specifically, in a phased hybrid parallel inference method for a MoE sparse large model provided in an embodiment of the present application, refer to Figure 9 In step 120 of the phased hybrid parallel inference method for the MoE sparse large model, it specifically includes the following content: Step 121: According to the expert numbers of each character, with the goal of equalizing the number of characters processed by each device during expert parallel computing and preferentially arranging the character with the smallest expert number within the global scope locally on each device, perform load balancing processing on each character in each device based on the greedy algorithm.

[0093] Step 122: Route each character after load balancing processing to the respective devices where their expert numbers are currently located in a full-exchange communication manner, so that each device respectively performs expert parallel computing on the received characters according to the second mixture-of-experts layer model parameters corresponding to their respective expert parallel strategies.

[0094] That is to say, in the phased hybrid parallel inference method proposed in the present application, the mixture-of-experts layer in the decoding stage uses tensor parallelism and inherently has load balancing characteristics. However, the expert parallelism in the pre-padding stage still faces the problem of load imbalance. Therefore, the present application proposes a load balancing strategy based on the greedy algorithm for the expert parallelism in the pre-padding stage. While achieving load balancing of computation and communication, it preferentially arranges the character with the smallest destination expert number within the global scope locally on the device, so that as many experts required for the characters processed on each device as possible are local (i.e., hit experts).

[0095] In order to further improve the application effectiveness and reliability of the load balancing strategy based on the greedy algorithm, in a phased hybrid parallel inference method for a MoE sparse large model provided in an embodiment of the present application, refer to Figure 11 , step 121 in the phased hybrid parallel inference method for the MoE sparse large model specifically includes the following contents: Step 1211: Control each of the devices to sort the respective local characters in ascending order of expert numbers to respectively determine the distribution information of each local character of each device in each expert, and obtain the distribution information of each local character of other devices in each expert in a full collection communication manner to determine the global distribution information of each character in each expert.

[0096] Step 1212: Control each of the devices to respectively determine the character distribution information of each device after routing the characters to the devices where the current expert numbers of the characters are located according to the global distribution information of each character in each expert.

[0097] Step 1213: Control each of the devices to determine the start address and offset of sending characters and the start address and offset of receiving characters of each device according to the distribution information of each local character of each device in each expert and the character distribution information of each device after routing the characters to the devices where the current expert numbers of the characters are located.

[0098] Step 1214: Input the start address and offset of sending characters of each device, the start address and offset of receiving characters of each device, the sending data buffer, and the receiving data buffer as input parameters into the full exchange communication function corresponding to the full exchange communication method.

[0099] That is to say, while achieving computational and communication load balancing, the present application preferentially arranges the characters with the smallest destination expert numbers in the global scope locally on the device, so that the experts required for the characters processed on each device are as much as possible local (i.e., hitting the experts). If the required expert is not local to the device (i.e., not hitting the expert), the non-hitting expert is loaded from the memory of a processor such as a CPU or GPU to the device video memory, and the loading overhead of the non-hitting expert is hidden by the calculation on the hitting expert. The above strategy solves the load imbalance problem at the system level and is a load balancing optimization method without loss of accuracy. In addition, in the phased hybrid parallel inference method proposed in the present application, the hybrid expert layer in the decoding stage adopts tensor parallelism and inherently has load balancing characteristics. Therefore, the present application can achieve full-stage lossless load balancing in the pre-fill stage and the decoding stage.

[0100] However, after the above load balancing character routing, there is a phenomenon that characters "miss" local experts. For example Figure 13 In the right box (i.e., the grid diagram on the right side of the arrow where "Character Routing (1)" is located), the characters processed by "Device 1" after character routing require "Expert 1" and "Expert 2". However, the local experts of "Device 1" are "Expert 2" and "Expert 3", and "Expert 1" misses. Therefore, to solve the problem of expert miss, the present application proposes the following technical solution: All experts are stored in the CPU memory in advance. When an expert miss occurs in a certain device, the missing expert is loaded from the CPU memory to the device video memory, so as to perform subsequent calculations, and the expert loading overhead can be hidden by the calculation and communication overhead on the hit experts.

[0101] Specifically, in a phased hybrid parallel inference method of a MoE sparse large model provided in an embodiment of the present application, refer to Figure 11 After step 1213 in the phased hybrid parallel inference method of the MoE sparse large model, the following specific content is further included: Step 1215: If it is determined according to the offsets of the characters sent by each device and the offsets of the received characters that there is a device with a missing expert after all-to-all communication, then before performing expert parallel calculation on each character based on the second hybrid expert layer model parameters corresponding to the missing expert in this device, load the second hybrid expert layer model parameters corresponding to the missing expert from the processor memory to this device.

[0102] That is to say, for the inference of MoE sparse large models, the embodiments of this application propose a phased hybrid parallel inference method. The core idea is that the mixture-of-experts layer in the prefill stage adopts expert parallelism, the mixture-of-experts layer in the decoding stage adopts tensor parallelism, and the multi-head attention layers in both stages adopt data parallelism. In the scenario where the computing devices are shared between the prefill stage and the decoding stage, since the model parameters used by each device under expert parallelism and tensor parallelism are different, this application proposes an efficient layer-by-layer parallel strategy conversion method to hide the communication overhead of parallel strategy conversion through layer-by-layer conversion, thereby completing the efficient conversion from expert parallelism in the prefill stage to tensor parallelism in the decoding stage. The principle that the phased hybrid parallel inference method proposed in this application is more advantageous than the prior art is as follows: In the prefill stage, the amount of data processed is large. Using expert parallelism in the mixture-of-experts layer only brings an all-to-all communication operation, and the communication volume of all-to-all is only the size of the input data. The communication overhead is lower than the all-gather and reduce-scatter communication operations used by tensor parallelism. Therefore, expert parallelism is adopted in the mixture-of-experts layer in the prefill stage of this application; in the decoding stage, the amount of data processed is small and the computational density is low. If expert parallelism is used in the mixture-of-experts layer, it will bring serious load imbalance problems. Therefore, in this application, tensor parallelism is adopted in the mixture-of-experts layer in the decoding stage, which can achieve perfect load balance. At the same time, the input data volume in the decoding stage is low, and the communication is latency-limited. Using tensor parallelism will also result in a lower communication latency. The above phased hybrid parallel strategy can combine the performance advantages of using expert parallelism in the mixture-of-experts layer in the prefill stage and using tensor parallelism in the mixture-of-experts layer in the decoding stage. In the scenario where the computing devices are shared between the prefill stage and the decoding stage, this application proposes an efficient layer-by-layer parallel strategy conversion method to achieve the efficient conversion of parallel strategies between the two stages.

[0103] To solve the load imbalance problem of expert parallelism in the prefill stage, the embodiments of this application also propose a load balancing strategy based on the greedy algorithm to solve the load imbalance problem at the system level. The principle of this load balancing strategy is as follows: As Figure 2As shown, in the expert parallelism during the pre-filling stage, the dynamic routing of characters can lead to load imbalance problems. To achieve load balancing, if some characters on a high-load device are directly distributed to a low-load device for processing, it is inevitable to migrate the corresponding experts to the low-load device, resulting in expert migration overhead. This application proposes a load balancing strategy based on the greedy algorithm. While achieving computational and communication load balancing, it preferentially arranges the characters with the smallest destination expert numbers globally on the local device, so that the experts required for the characters processed on each device are as local as possible (i.e., hitting the experts). If the required experts are not local to the device (i.e., missing the experts), the missing experts are loaded from the CPU memory to the device video memory, and the loading overhead of the missing experts is hidden by the calculations on the hitting experts. The above strategy solves the load imbalance problem at the system level and is a load balancing optimization method without loss of accuracy.

[0104] To further illustrate the phased hybrid parallel inference method of the MoE sparse large model provided in the above embodiments, this application also provides a specific application example of the phased hybrid parallel inference method of the MoE sparse large model, which is used to solve the following technical problems existing in the prior art: (1) For the distributed inference of the MoE sparse large model, the parallel strategies adopted in the prior art have poor adaptability to the pre-filling stage and the decoding stage, especially the PD integrated inference scheme. A typical example is the "data parallelism + tensor parallelism + expert parallelism" hybrid parallel strategy adopted by DeepSpeed-MoE. This parallel strategy does not distinguish between the pre-filling stage and the decoding stage, and the parallel strategy remains unchanged throughout the inference process. However, during the pre-filling stage, a large amount of data is processed, and using tensor parallelism will bring communication operations of all-gather and reduce-scatter on data, with high communication overhead, which is more serious when dealing with a large batch of data. In the decoding stage, the amount of data processed is relatively small. Using expert parallelism will lead to load imbalance problems, and the decoding stage has a low computational density and processes a small amount of data, so the all-to-all communication overhead brought by expert parallelism is also more significant.

[0105] To this end, the present application proposes a phased hybrid parallel inference method for MoE sparse large models. For the prefill stage with a high data volume to be processed, the mixture-of-experts layer adopts expert parallelism with relatively lower communication volume. For the decoding stage with a low data volume to be processed, the mixture-of-experts layer adopts more efficient tensor parallelism. For the multi-head attention layer, data parallelism is adopted in both stages. Since the parallel strategies adopted by the mixture-of-experts layer in the prefill stage and the decoding stage are different, and the model parameters used by each device in the two stages are also different, it is necessary to perform a parallel strategy conversion between the two stages, that is, the model parameters required for the decoding stage need to be loaded after the prefill execution ends. The present application proposes an efficient method for layer-by-layer parallel strategy conversion to hide the communication overhead of the parallel strategy conversion through layer-by-layer conversion.

[0106] (2) For the distributed inference of MoE sparse large models, the prior art cannot completely solve the problem of unbalanced expert parallel load. Gshard proposed a design of expert capacity limit, that is, characters exceeding the expert capacity will directly skip the MoE mixture-of-experts layer using a residual network, while experts that do not reach the capacity limit will be padded to make up the difference. Although this strategy can balance the computing and communication loads, its disadvantage is that it will generate useless computations, and at the same time, due to the discarding of characters during inference, there may be a loss of accuracy, which is difficult to accept during inference. Deepseek V3 adopted a load balancing strategy without auxiliary loss. This strategy introduces an adjustable bias value for each expert, and this bias value will be added to the expert affinity score calculated by the gating function to adjust the load on each expert. The bias value will be reduced for experts with high load and increased for experts with low load, but it can only make the load on each expert as balanced as possible. The prior art mainly optimizes load balancing from the algorithm level and cannot completely solve the problem of unbalanced load.

[0107] To solve the problem of unbalanced expert parallel load in the prefill stage, the present application proposes a load balancing strategy based on the greedy algorithm. While achieving balanced computing and communication loads, it preferentially arranges characters with the smallest destination expert number within the global scope locally on the device, so that the experts required for the characters processed on each device are as likely as possible to be local (i.e., hit experts). If the required expert is not local to the device (i.e., a missed expert), the missed expert will be loaded from the CPU memory to the device video memory, and the loading overhead of the missed expert is hidden by the computing on the hit experts. The above strategy solves the problem of unbalanced load from the system level and is a load balancing optimization method without loss of accuracy. In addition, in the phased hybrid parallel inference method proposed in the present application, the mixture-of-experts layer in the decoding stage adopts tensor parallelism, which inherently has the characteristics of load balancing. Therefore, the present application can achieve lossless load balancing in the entire prefill stage and decoding stage.

[0108] Based on this, seeFigure 14 , the core MoE hybrid expert layer phased hybrid parallel strategy involved in the phased hybrid parallel inference method of the MoE sparse large model provided by the application example of this application is as follows: 1. Phased hybrid parallel strategy for MoE sparse large model inference The application example of this application proposes a phased hybrid parallel strategy for MoE sparse large model inference. The overall scheme is as Figure 14 shown. In the prefill stage, the hybrid expert layer adopts expert parallelism. In the decoding stage, the hybrid expert layer adopts tensor parallelism. For the multi-head attention layers in both stages, data parallelism is adopted. Between the two stages, an efficient layer-by-layer parallel strategy conversion method is used to achieve efficient conversion of the parallel strategies between the two stages. The specific technical solutions of the phased hybrid parallel strategy are as follows: 1-1. Parallel strategy in the prefill stage: Refer to Figure 15 , for the prefill stage, the hybrid expert layer of the model adopts expert parallelism, that is, all E experts are evenly divided into P devices. Let the high feature dimension of each expert be H, and each device has [E / P, H] experts, where E / P is the number of experts on each device, and each device processes the characters after routing; the multi-head attention layer of the model adopts data parallelism, that is, all input characters are divided into P devices for parallel processing, and the model parameters of the multi-head attention layer are replicated on each device. After the data parallel calculation of the multi-head attention layer is completed, according to the calculation result of the affinity score of the gating function, the characters are routed to the devices where the corresponding experts are located through an AlltoAll all-to-all communication operation for expert parallel calculation of the hybrid expert layer. After the expert parallel calculation of the hybrid expert layer is completed, the characters are routed back to the original devices through an AlltoAll all-to-all communication operation, so as to connect to the data parallel calculation of the multi-head attention layer of the next layer.

[0109] 1-2. Parallel strategy in the decoding stage: Refer to Figure 16, for the decoding stage, the mixture-of-experts layer of the model adopts tensor parallelism, that is, the high feature dimension H of all E experts is evenly divided among P devices. Each device has [E, H / P] experts, that is, each device has H / P parts of all E experts, and each device needs to process all characters; the multi-head attention layer of the model adopts data parallelism, that is, all input characters are divided among devices for parallel processing, and the model parameters of the multi-head attention layer are replicated on each device. After the data parallel calculation of the multi-head attention layer is completed, all characters are collected on each device through an AllGather communication operation, and the activated experts are selected according to the calculation results of the affinity scores of the gating function. Then, local calculations of tensor parallelism of the mixture-of-experts layer are performed with all characters and the activated experts as inputs. After the local calculations of tensor parallelism of the mixture-of-experts layer are completed, the local calculation results are reduced and distributed through a ReduceScatter communication operation, so as to connect to the data parallel calculation of the multi-head attention layer of the next layer.

[0110] 1-3. Method for converting layer-by-layer parallel strategies between the prefill and decoding stages: The mixture-of-experts layer in the prefill stage adopts expert parallelism, and the mixture-of-experts layer is divided according to the dimension of the number of experts; the mixture-of-experts layer in the decoding stage adopts tensor parallelism, and the mixture-of-experts layer is divided according to the high feature dimension of the experts. The model parameters used by the mixture-of-experts layer in the prefill and decoding stages are different, and preparing two copies of the model parameters in the video memory at the same time will significantly increase the video memory overhead. Therefore, the application example of this application proposes an efficient method for converting layer-by-layer parallel strategies to hide the communication overhead of parallel strategy conversion. Suppose the model has a total of L layers of the basic MoE mixture-of-experts model architecture (as Figure 1 shown). When performing the calculation of each layer of the MoE basic architecture, the distribution of the model parameters of the expert parallelism of the mixture-of-experts layer of the current basic architecture is converted into the distribution of the model parameters of the tensor parallelism of the mixture-of-experts layer of the current basic architecture through an AlltoAll all-to-all communication operation. After the AlltoAll all-to-all communication operation and the calculation of the current mixture-of-experts layer are completed, the model parameters required for the expert parallelism of the mixture-of-experts layer in the video memory are released, so as to complete the parallel strategy conversion of the current layer. By sequentially completing the parallel strategy conversion of L layers according to the above method, the parallel strategy conversion of the entire model from the prefill stage to the decoding stage can be completed.

[0111] 2. Load balancing strategy based on the greedy algorithm In the staged hybrid parallel inference method proposed in the application example of this application, the hybrid expert layer in the decoding stage uses tensor parallelism, which inherently has the characteristics of load balancing. However, the expert parallelism in the prefill stage still faces the problem of load imbalance. Therefore, the application example of this application proposes a load balancing strategy based on the greedy algorithm for the expert parallelism in the prefill stage. While achieving load balancing for computing and communication, it preferentially arranges the characters with the smallest destination expert number within the global scope locally on the device, so that the experts required for the characters processed on each device are as local as possible (i.e., hitting the experts). Specifically: First step, calculate the destination expert numbers of each input character according to the gating function, sort the input characters locally on each device, and arrange the characters with smaller destination expert numbers in the front, as shown in Figure 13 the left box in Figure 13 (i.e., the grid graph to the left of the arrow where "Character Routing (1)" is located), so as to determine the distribution information of the local input characters among each expert before the AlltoAll character routing; Second step, determine the character distribution targets after the AlltoAll character routing for each device. Here, two goals need to be achieved. One is to ensure that the number of characters processed by each device is equal to achieve load balancing, and at the same time, preferentially arrange the characters with the smallest destination expert number within the global scope locally on the device, so that the characters processed on each device can hit the local experts as much as possible; Third step, in order to achieve the above goals, each device counts the distribution information of the local input characters among each expert, and each device uses the AllGather communication operation to collect the distribution information of all input characters among each expert, so as to calculate the total number of characters processed by each expert within the global scope; Fourth step, according to the global distribution information of the characters among each expert, each device determines the character distribution information after the AlltoAll character routing. Let the total number of input characters be N and the number of devices be P. Starting from the device with the smallest number, preferentially arrange the characters with the smallest destination expert number within the global scope locally on the device until N / P characters are filled, and then arrange the next device until P devices are arranged. The character distribution effect after routing is as shown in the right box in Figure 13 (i.e., the grid graph to the right of the arrow where "Character Routing (1)" is located); Fifth step, according to the distribution information of the local input characters among each expert before the AlltoAll character routing obtained above and the character distribution information after the AlltoAll character routing, calculate the start address and offset of the characters sent by each device and the start address and offset of the characters received in the AlltoAll communication operation; Sixth step, use the start address and offset of the sent and received data calculated in the fifth step, as well as the send data buffer and receive data buffer as input parameters, input them into the AlltoAll communication function, and then perform the AlltoAll communication operation to complete the load-balanced character routing.

[0112] An example of load-balanced character routing implemented according to the above technical solution is as follows Figure 13 As shown. After routing, the number of characters processed by each device is the same, achieving computational load balance. At the same time, for the AlltoAll communication operation, the total amount of data sent by each device is the same, and the total amount of data received by each device is the same, achieving communication load balance at the endpoints. However, after the above load-balanced character routing, there is a phenomenon that characters "miss" local experts. For example Figure 13 In the right box (i.e., "the grid diagram on the right side of the arrow where 'Character Routing (1)' is located"), after character routing, the characters processed by "Device 1" require "Expert 1" and "Expert 2", but the local experts of "Device 1" are "Expert 2" and "Expert 3", and "Expert 1" misses. To solve the problem of expert miss, the application example of this application proposes the following technical solution: Store all experts in the CPU memory in advance. When an expert miss occurs in a certain device, the missing expert is loaded from the CPU memory to the device video memory for subsequent calculations, and the expert loading overhead can be hidden by the computational and communication overhead on the hit experts.

[0113] After the expert parallel calculation in the hybrid expert layer is completed, the characters are routed back to the original devices through an inverse AlltoAll full exchange communication operation, so as to connect to the data parallel calculation of the multi-head attention layer in the next layer, as Figure 17 shown. For this AlltoAll communication operation, the total amount of data sent by each device is the same, and the total amount of data received by each device is the same, also achieving communication load balance at the endpoints.

[0114] The above load-balanced technical solution proposed by the application example of this application achieves computational and communication load balance while ensuring no loss of model accuracy, and completely solves the load-balanced problem at the system level. At the same time, in the phased hybrid parallel inference method proposed by the application example of this application, the hybrid expert layer in the decoding stage adopts tensor parallelism, which inherently has load-balanced characteristics. Therefore, the application example of this application can achieve full-stage lossless load balance in the pre-fill stage and the decoding stage.

[0115] 3. Data transmission overhead hiding technology In the above phased hybrid parallel technical solution for MoE sparse large model inference, it is necessary to complete the layer-by-layer parallel strategy conversion of the mixture-of-experts layer from the prefill stage to the decoding stage through an all-to-all communication operation. In the above technical solution of the load balancing strategy based on the greedy algorithm, it is necessary to load the missed experts from the CPU memory to the device video memory. The layer-by-layer parallel strategy conversion and the loading of missed experts will bring certain data transmission overheads. The application example of this application proposes a data transmission overhead hiding technology, which hides the data transmission overheads brought by the layer-by-layer parallel strategy conversion and the loading of missed experts through irrelevant computations and communication loads. The technical solution is as Figure 18 shown. For the all-to-all communication overhead brought by the parallel strategy conversion (from expert parallelism to tensor parallelism) of the MoE mixture-of-experts layer, it is masked by computation loads such as multi-head attention layer computation, gate function computation, and MoE mixture-of-experts layer computation. For the data transmission overhead brought by the loading of missed experts, it is masked by loads such as the all-to-all communication of character routing from multi-head attention layer data parallelism to MoE layer expert parallelism, the mixture-of-experts layer computation on the hit experts, and offset computation.

[0116] In addition, the application example of this application also conducted corresponding experimental tests, and the specific content is as follows: The experimental tests were carried out on a 2-machine 16-card Ascend 910B distributed machine, where the single-card video memory capacity of Ascend 910B is 64GB. The test object is the DeepSeek V2 mixture-of-experts MoE sparse large model (with 236 billion model parameters), and the comparison baseline is the DeepSpeed-MoE parallel framework.

[0117] Figures 19(a) and 19(b) respectively show the performance results of the DeepSeek-V2 prefill stage in the cases of batchsize (batch size) = 128 and batchsize = 256, where the abscissa is the degree of load imbalance (represented by variance). The test results show that the load balancing strategy based on the greedy algorithm proposed in the application example of this application can effectively solve the load imbalance problem of the mixture-of-experts sparse large model, and the execution time of the prefill stage is hardly affected by the degree of load imbalance. Compared with the baseline DeepSpeed-MoE framework, as the degree of load imbalance increases, the performance advantage of the application example of this application becomes more obvious, and the highest performance speedup ratio of 2.7 times is obtained. The above performance test results verify the effectiveness of the load balancing strategy based on the greedy algorithm proposed in this application.

[0118] Figures 20(a) and 20(b) respectively show the performance results of the DeepSeek-V2 decoding stage under the conditions of batchsize = 128 and batchsize = 256, where the abscissa is the degree of load imbalance. The test results show that in the phased hybrid parallel strategy proposed by the application example of this application, tensor parallelism is adopted in the decoding stage, and a 2.3-fold performance speedup ratio is obtained compared with the expert parallel strategy adopted by the baseline DeepSpeed-MoE, thus proving the rationality and effectiveness of adopting tensor parallelism in the decoding stage of the phased hybrid parallel strategy proposed by the application example of this application.

[0119] Figures 21(a) and 21(b) respectively show the performance results of the DeepSeek-V2 end-to-end inference (including two stages of pre-filling and decoding) under the conditions of batchsize = 128 and batchsize = 256, where the abscissa is the degree of load imbalance. The test results show that the phased hybrid parallel inference method and system of the MoE sparse large model proposed by the application example of this application can obtain a maximum performance speedup ratio of 2.4 times compared with the parallel strategy with the entire inference process unchanged adopted by the baseline DeepSpeed-MoE. In the end-to-end inference, the system of the application example of this application integrates all the proposed technical solutions, including the phased hybrid parallel strategy for MoE sparse large model inference, the load balancing strategy based on the greedy algorithm, and the data transmission overhead hiding technology. The above end-to-end inference performance results can verify the effectiveness of the technical solutions proposed by the application example of this application.

[0120] In summary, compared with the prior art, the phased hybrid parallel inference method of the MoE sparse large model provided by the application example of this application has the following beneficial effects: (1) The parallel strategies adopted in the prior art do not fully consider the computational load characteristics of the prefill and decoding stages. The proposed phased hybrid parallel strategy in this application better adapts to the computational load characteristics of the prefill and decoding stages during the inference process of the MoE sparse large model. Compared with the prior art, it has the following advantages: In the phased hybrid parallel strategy proposed in this application, the hybrid expert layer in the prefill stage adopts expert parallelism. The communication volume of the Alltoall operation brought by expert parallelism is only the size of the input data, and the communication overhead is lower than that of the AllGather and ReduceScatter communication operations used in tensor parallelism. In this application, the hybrid expert layer in the decoding stage adopts tensor parallelism, which can achieve full load balancing in terms of computing, memory access, and communication. At the same time, the input data volume in the decoding stage is low, and the communication is latency-limited. Using tensor parallelism will also result in lower communication latency. An efficient layer-by-layer parallel strategy conversion method is adopted to hide the communication overhead of parallel strategy conversion through layer-by-layer conversion, thereby completing the efficient conversion from expert parallelism in the prefill stage to tensor parallelism in the decoding stage, and combining the performance advantages of using expert parallelism in the hybrid expert layer in the prefill stage and tensor parallelism in the hybrid expert layer in the decoding stage. In summary, the proposed phased hybrid parallel inference method in this application adopts different parallel strategies adapted to different stages and is efficiently connected through the layer-by-layer parallel strategy conversion method, achieving better performance compared with the prior art.

[0121] (2) Regarding the load balancing problem of the MoE sparse large model, the prior art mainly optimizes from the model algorithm level, such as the expert capacity upper limit strategy of Gshard and the load balancing strategy without auxiliary loss of Deepseek V3. However, the prior art cannot fundamentally solve the load balancing problem, either causing model accuracy loss or still having a certain degree of load imbalance. For the load imbalance problem of expert parallelism in the prefill stage, this application proposes a load balancing strategy based on the greedy algorithm, which ensures model accuracy lossless while achieving computational and communication load balancing, and completely solves the load balancing problem from the system level. At the same time, in the phased hybrid parallel inference method proposed in this application, the hybrid expert layer in the decoding stage adopts tensor parallelism, which inherently has the load balancing characteristic. Therefore, this application can achieve lossless load balancing in both the prefill stage and the decoding stage.

[0122] From a software perspective, this application also provides a phased hybrid parallel inference device for the MoE sparse large model that executes all or part of the content in the phased hybrid parallel inference method of the MoE sparse large model. The phased hybrid parallel inference device for the MoE sparse large model specifically includes the following: The pre - filling stage execution module is used to control each architecture layer of the MoE sparse large model to execute the first step layer by layer during the pre - filling stage of data inference based on the MoE sparse large model. Wherein, the first step includes: based on the multi - head attention layer model parameters and gating functions in each device, obtaining the expert numbers of each character corresponding to the prompt data sequence, which are initially located at their respective positions in their respective devices, and adding the first mixture - of - experts layer model parameters operating based on the tensor parallel strategy to each device; then performing expert parallel computing on each character according to the expert numbers of each character and the second mixture - of - experts layer model parameters operating based on the expert parallel strategy in each device; and then restoring each character after expert parallel computing to its initial position in its respective device and releasing the second mixture - of - experts layer model parameters in each device. The stage transition execution module is used to send the predicted character output by the last architecture layer of the MoE sparse large model to the multi - head attention layer in the first architecture layer of the MoE sparse large model, for executing the decoding stage of data inference of the MoE sparse large model according to the predicted character, the multi - head attention layer model parameters in each device, the gating function, and the first mixture - of - experts layer model parameters operating based on the tensor parallel strategy in each device.

[0123] In some embodiments of the present application, the staged hybrid parallel inference device of the MoE sparse large model further includes the following modules executed after the stage transition execution module: The decoding stage execution module is used to control each architecture layer of the MoE sparse large model to execute the second step layer by layer in each iteration round during the decoding stage of data inference based on the MoE sparse large model, and send the predicted character output by the last architecture layer of the MoE sparse large model in each iteration round to the multi - head attention layer in the first architecture layer of the MoE sparse large model in a full - exchange communication manner.

[0124] The embodiments of the staged hybrid parallel inference device of the MoE sparse large model provided in the present application can specifically be used to execute the processing flow of the embodiments of the staged hybrid parallel inference method of the MoE sparse large model in the above embodiments. Its functions will not be elaborated here, and reference can be made to the detailed description of the embodiments of the staged hybrid parallel inference method of the MoE sparse large model above.

[0125] The part of the phased hybrid parallel inference device for the MoE sparse large model to perform phased hybrid parallel inference of the MoE sparse large model can be completed in a server or a client device. Specifically, it can be selected according to the processing capacity of the client device and the limitations of the user usage scenario, etc. This application does not make a limitation in this regard. If all operations are completed in the client device, the client device may further include a processor for specific processing of the phased hybrid parallel inference of the MoE sparse large model.

[0126] The above-mentioned client device may have a communication module (i.e., communication unit), which can communicate with a remote server to achieve data transmission with the server. The server may include a server on the task scheduling center side, and in other implementation scenarios, it may also include a server of an intermediate platform, such as a server of a third-party server platform having a communication link with the task scheduling center server. The server may include a single computer device, or may include a server cluster composed of multiple servers, or a server structure of a distributed device.

[0127] Any suitable network protocol can be used for communication between the above-mentioned server and the client device side, including network protocols not yet developed on the filing date of this application. The network protocol may, for example, include TCP / IP protocol, UDP / IP protocol, HTTP protocol, HTTPS protocol, etc. Of course, the network protocol may, for example, also include the RPC protocol (Remote Procedure Call Protocol) and the REST protocol (Representational State Transfer) used on top of the above protocols, etc.

[0128] As can be seen from the above description, for the pre-fill stage with a high volume of data to be processed, the stage-based hybrid parallel inference device of the MoE sparse large model provided by the embodiments of the present application adopts an expert parallel strategy with relatively lower communication volume in the mixture-of-experts layer to effectively reduce the communication overhead between devices in the pre-fill stage; while for the decoding stage with a low volume of data to be processed, the mixture-of-experts layer adopts a more efficient tensor parallel strategy to effectively reduce the communication overhead between devices in the decoding stage, which can effectively improve the adaptability to the pre-fill stage and the decoding stage and improve the efficiency of the inference process of the MoE sparse large model; and, by adding the model parameters of the first mixture-of-experts layer operating based on the tensor parallel strategy to each of the devices while performing the multi-head attention layer calculation in the pre-fill stage, the communication overhead of the parallel strategy conversion can be hidden in a layer-by-layer conversion manner without significantly increasing the device resource occupancy rate; and after the calculation of the mixture-of-experts layer of each structural layer in the pre-fill stage is completed, the model parameters of the second mixture-of-experts layer operating based on the expert parallel strategy can be released layer by layer, which can further avoid the increase in device resource occupancy rate.

[0129] The embodiments of the present application further provide an electronic device, which may include a processor, a memory, a receiver, and a transmitter. The processor is configured to execute the stage-based hybrid parallel inference method of the MoE sparse large model mentioned in the above embodiments. The processor and the memory may be connected through a bus or other means. Taking the connection through the bus as an example, the receiver may be connected to the processor and the memory in a wired or wireless manner.

[0130] The processor may be a central processing unit (CPU). The processor may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc., or a combination of the above types of chips.

[0131] As a non-transitory computer-readable storage medium, the memory can be used to store non-transitory software programs, non-transitory computer-executable programs, and modules, such as the program instructions / modules corresponding to the stage-based hybrid parallel inference method of the MoE sparse large model in the embodiments of the present application. By running the non-transitory software programs, instructions, and modules stored in the memory, the processor can perform various functional applications and data processing of the processor, that is, implement the stage-based hybrid parallel inference method of the MoE sparse large model in the above method embodiments.

[0132] The memory may include a program storage area and a data storage area. Among them, the program storage area can store the operating system and application programs required for at least one function; the data storage area can store data created by the processor and the like. In addition, the memory may include high-speed random access memory, and may also include non-transitory memory, such as at least one disk storage device, a flash memory device, or other non-transitory solid-state storage devices. In some embodiments, the memory optionally includes a memory remotely set relative to the processor, and these remote memories can be connected to the processor through a network. Examples of the above-mentioned network include but are not limited to the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0133] The one or more modules are stored in the memory and, when executed by the processor, implement the phased hybrid parallel inference method of the MoE sparse large model in the embodiments.

[0134] In some embodiments of the present application, the user equipment may include a processor, a memory, and a transceiver unit. The transceiver unit may include a receiver and a transmitter. The processor, the memory, the receiver, and the transmitter may be connected through a bus system. The memory is used to store computer instructions, and the processor is used to execute the computer instructions stored in the memory to control the transceiver unit to transmit and receive signals.

[0135] As an implementation manner, the functions of the receiver and the transmitter in the present application may be considered to be implemented through a transceiver circuit or a dedicated chip for transceiver, and the processor may be considered to be implemented through a dedicated processing chip, a processing circuit, or a general-purpose chip.

[0136] As another implementation manner, it may be considered to use a general-purpose computer to implement the server provided in the embodiments of the present application. That is, the program codes for implementing the functions of the processor, the receiver, and the transmitter are stored in the memory, and the general-purpose processor implements the functions of the processor, the receiver, and the transmitter by executing the codes in the memory.

[0137] The embodiments of the present application also provide a phased hybrid parallel inference system for the MoE sparse large model. The phased hybrid parallel inference system for the MoE sparse large model at least includes a scheduling device and each device communicatively connected to the scheduling device; the scheduling device includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, it implements the phased hybrid parallel inference method of the MoE sparse large model provided in the foregoing embodiments.

[0138] The embodiments of the present application also provide a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the foregoing phased hybrid parallel inference method for the MoE sparse large model are implemented. The computer-readable storage medium may be a tangible storage medium, such as a random access memory (RAM), internal memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, floppy disks, hard disks, removable storage disks, CD-ROMs, or any other form of storage medium known in the art.

[0139] The embodiments of the present application also provide a computer program product, including a computer program. When the computer program is executed by a processor, the steps of the foregoing phased hybrid parallel inference method for the MoE sparse large model are implemented.

[0140] Those of ordinary skill in the art should understand that the various exemplary components, systems, and methods described in connection with the embodiments disclosed herein can be implemented in hardware, software, or a combination of both. Specifically, whether to implement in hardware or software depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present application. When implemented in hardware, it can be, for example, an electronic circuit, an application-specific integrated circuit (ASIC), appropriate firmware, a plug-in, a functional card, etc. When implemented in software, the elements of the present application are programs or code segments used to execute the required tasks. The program or code segment can be stored in a machine-readable medium or transmitted via a data signal carried in a carrier wave on a transmission medium or a communication link.

[0141] It should be clear that the present application is not limited to the specific configurations and processes described above and shown in the figures. For the sake of brevity, the detailed description of known methods is omitted here. In the above embodiments, several specific steps are described and shown as examples. However, the method process of the present application is not limited to the specific steps described and shown. Those skilled in the art can make various changes, modifications, and additions, or change the order between steps after understanding the spirit of the present application.

[0142] In the present application, the features described and / or illustrated for one embodiment can be used in the same or similar manner in one or more other embodiments, and / or combined with the features of other embodiments or replace the features of other embodiments.

[0143] The above are only the preferred embodiments of the present application and are not intended to limit the present application. For those skilled in the art, various changes and modifications can be made to the embodiments of the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included within the protection scope of the present application.

Claims

1. A phased hybrid parallel inference method for MoE sparse large models, characterized by: include: In the pre-filling stage of data reasoning based on the MoE sparse large model, each architecture layer of the MoE sparse large model is controlled to execute the first step layer by layer; wherein the first step includes: based on the multi-head attention layer model parameters and the gating function in each device, obtaining the expert number of each character corresponding to the prompt data sequence and currently located at the initial position of each device, and adding the first hybrid expert layer model parameters based on the tensor parallel strategy to each device; then, according to the expert number of each character and the corresponding second hybrid expert layer model parameters based on the expert parallel strategy in each device, expert parallel calculation is performed on each character; then, each character after expert parallel calculation is restored to the initial position of each device and the second hybrid expert layer model parameters in each device are released; The predicted characters output by the last architectural layer of the MoE sparse large model are sent to the multi-head attention layer in the first architectural layer of the MoE sparse large model, so as to execute the decoding stage of data inference of the MoE sparse large model according to the predicted characters, the multi-head attention layer model parameters in each device, the gating function and the first hybrid expert layer model parameters running based on the tensor parallel strategy in each of the devices.

2. The phased hybrid parallel reasoning method for the MoE sparse large model according to claim 1, characterized in that: Also includes: In each iteration round of the decoding phase of data reasoning based on the MoE sparse large model, each architecture layer of the MoE sparse large model is controlled to perform the second step layer by layer, and the predicted characters output by the last architecture layer of the MoE sparse large model in each iteration round are sent to the multi-head attention layer in the first architecture layer of the MoE sparse large model in a full exchange communication mode; Wherein, the second step comprises: In the current iteration round, based on the multi-head attention layer model parameters and gating functions in each of the devices, obtain the expert numbers of the respective characters currently located at the initial positions of the respective devices corresponding to the latest data sequence; wherein, if the current iteration round is the first round, the latest data sequence includes the prompt data sequence cached in the pre-filling stage and the predicted characters output by the last architecture layer of the MoE sparse large model in the pre-filling stage; if the current iteration round is not the first round, the latest data sequence includes the prompt data sequence, the predicted characters output in the pre-filling stage, and the predicted characters output in the historical iteration rounds before the current iteration round; Collecting each of the characters into each device in a full collection communication mode, and according to the respective expert numbers of the characters, enabling each of the devices to perform tensor parallel calculations on the collected characters according to the respective corresponding first hybrid expert layer model parameters running based on the tensor parallel strategy; The characters after tensor parallel calculation are restored to their respective initial positions of the devices in a reduction-distribution communication manner.

3. The phased hybrid parallel reasoning method for the MoE sparse large model according to claim 2, characterized in that: The method of obtaining the expert number of each character currently located at the initial position of each device corresponding to the prompt data sequence based on the multi-head attention layer model parameters and the gating function in each device includes: According to the total number of each of the devices, the prompt data sequence is evenly divided into a plurality of first data subsequences; Input each of the first data subsequences into each device one by one, so that each of the devices performs multi-head attention calculation on the first data subsequence received by each device based on the multi-head attention layer model parameters stored by each device, so as to obtain each character corresponding to the first data subsequence received by each device, and determine the expert number of each character on each device based on the gating function; Correspondingly, based on the multi-head attention layer model parameters and gating functions in each of the devices, obtaining the expert numbers of the respective characters currently located at the initial positions of the respective devices corresponding to the latest data sequence includes: According to the total number of each of the devices, the latest data sequence is evenly divided into a plurality of second data subsequences; Input each of the second data subsequences into each device one by one, so that each of the devices performs multi-head attention calculation on the second data subsequence received based on the multi-head attention layer model parameters stored respectively, so as to obtain each character corresponding to the second data subsequence received respectively, and collect all the characters on each device in a full-collection communication manner, so that each of the devices determines the expert number of each of the characters on each device based on the gating function.

4. The phased hybrid parallel reasoning method for the MoE sparse large model according to claim 1, characterized in that: The simultaneously adding the first hybrid expert layer model parameters based on the tensor parallel strategy to each of the devices includes: If the time required to add the first hybrid expert layer model parameters based on the tensor parallel strategy to each of the devices is shorter than or equal to the time required to obtain the expert numbers of the respective characters corresponding to the prompt data sequence and currently located at the initial positions of the respective devices, then the first hybrid expert layer model parameters based on the tensor parallel strategy are added to each of the devices while obtaining the expert numbers of the respective characters corresponding to the prompt data sequence and currently located at the initial positions of the respective devices.

5. The staged hybrid parallel reasoning method for the MoE sparse large model according to claim 1, characterized in that: Before the characters calculated in parallel by the expert are restored to their respective initial positions of the device in the first step, the method further includes: If the time required to add the first hybrid expert layer model parameters based on the tensor parallel strategy to each of the devices is longer than the time required to obtain the expert numbers of the respective characters corresponding to the prompt data sequence and currently located at the initial positions of the respective devices, then the first hybrid expert layer model parameters based on the tensor parallel strategy corresponding to each of the devices are divided into a first part of parameters and a second part of parameters respectively; Based on the multi-head attention layer model parameters and the gating function in each device, obtain the expert number of each character currently located at the initial position of each device corresponding to the prompt data sequence, and add the first part of parameters to each device; Then, expert parallel calculation is performed on each of the characters according to the expert number of each character and the corresponding second hybrid expert layer model parameters based on the expert parallel strategy in each of the devices, and the second part of parameters is added to each of the devices.

6. The phased hybrid parallel reasoning method for the MoE sparse large model according to claim 1, characterized in that: The expert parallel calculation of each character according to the expert number of each character and the corresponding second hybrid expert layer model parameters based on the expert parallel strategy in each device includes: According to the expert number of each character, with the goal of equalizing the number of characters processed by each device in the expert parallel computing process and preferentially arranging the character with the smallest expert number in the global scope locally on each device, load balancing processing based on the greedy algorithm is performed on each character in each device; In a full exchange communication mode, each character after load balancing processing is routed to each device where the respective expert number is currently located, so that each of the devices performs expert parallel calculation on the characters received by each device according to its corresponding second hybrid expert layer model parameters running based on the expert parallel strategy.

7. The phased hybrid parallel reasoning method for the MoE sparse large model according to claim 6, characterized in that: The process of performing load balancing processing based on a greedy algorithm on each character in each device according to the respective expert number of each character, with the goal of equalizing the number of characters processed by each device and preferentially arranging the character with the smallest expert number in the global scope in each local area, includes: Control each of the devices to sort the local characters in the order of expert number from small to large to respectively determine the distribution information of each local character of each device in each expert, and obtain the distribution information of each local character of other devices in each expert in a full collection communication mode to determine the global distribution information of each character in each expert; Controlling each of the devices to respectively determine, according to the global distribution information of each character in each expert, the character distribution information after each of the devices routes the character to each device where the respective expert number is currently located; Controlling each of the devices to determine the starting address and offset of each of the devices sending characters and the starting address and offset of each of the devices receiving characters according to the distribution information of each of the local characters in each of the experts and the character distribution information after each of the devices routes the characters to each of the devices where the respective expert numbers are currently located; The starting address and offset of each device sending characters, the starting address and offset of each device receiving characters, the sending data buffer and the receiving data buffer are input as input parameters into the full switching communication function corresponding to the full switching communication mode.

8. The staged hybrid parallel reasoning method for the MoE sparse large model according to claim 7, characterized in that: After determining the starting address and offset of each of the devices for sending characters and the starting address and offset of each of the devices for receiving characters, the method further includes: If it is determined that there is a device with missed experts after full exchange communication based on the offset of the characters sent and the offset of the characters received by each of the devices, then before the device performs expert parallel calculation on each of the characters based on the second mixed expert layer model parameters corresponding to the missed experts, the second mixed expert layer model parameters corresponding to the missed experts are loaded from the processor memory to the device.

9. A staged hybrid parallel inference system for MoE sparse large models, characterized by: It includes: a scheduling device and various devices respectively connected to the scheduling device for communication; The scheduling device includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the phased hybrid parallel reasoning method for the MoE sparse large model as described in any one of claims 1 to 8 is implemented.

10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the method for phased hybrid parallel reasoning of the MoE sparse large model as claimed in any one of claims 1 to 8 is implemented.

Citation Information

Patent Citations

  • Data processing method, end-side device, storage medium, chip system and computer program product

    CN119831056A

  • Data processing method, apparatus and system, and medium and program product

    WO2024066791A1

Cited By

  • Expert parallel computing method for hybrid expert model and computer program product

    CN120493998A

  • Model loading and unloading method and electronic equipment

    CN120909807A