Optimization method and device of hybrid expert operator, computer equipment, storage medium and program product

CN122840140APending Publication Date: 2026-09-29SHANGHAI BIREN TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610991828.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-03
Publication Date
2026-09-29

AI Technical Summary

Technical Problem

[0004]但在小Shape场景下(如16 Token),模型计算量大幅缩减,上述数据重排操作产生的固定调度开销占比会显著抬升,导致无法有效降低整体推理延迟

Benefits of technology

[0041]上述混合专家算子的优化方法、装置、计算机设备、存储介质和程序产品,采用路由算子对输入张量进行路由计算,得到路由权重掩码,该路由权重掩码表示每个路由专家网络对应输入张量中每一词元的权重值,在路由专家网络未被词元激活时,路由权重掩码中路由专家网络对应该词元的权重值为0,其中输入张量的张量尺寸小于或者等于预置张量尺寸。各路由专家网络分别对输入张量进行运算处理,得到各路由专家网络的第一输出张量,并对各路由专家网络的第一输出张量与路由权重掩码进行乘积累加处理,得到混合专家算子的最终运算结果。采用本申请实施例提供的混合专家算子的优化方法、装置、计算机设备、存储介质和程序产品,在小张量尺寸的场景下,不再执行拆分分发词元的操作,所有路由专家网络均读取完整输入张量进行运算,并通过路由权重掩码实现稀疏加权融合,彻底消除了Permutation/Unpermutation的数据重排以及All-To-All的多对多通信,从而有效降低了小张量尺寸场景下混合专家算子的整体推理延迟,也即提高了混合专家算子的推理效率。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122840140A_ABST
    Figure CN122840140A_ABST
Patent Text Reader

Abstract

The application relates to an optimization method and device of a hybrid expert operator, a computer device, a storage medium and a program product. The method comprises the following steps: performing routing calculation on an input tensor by using a routing operator to obtain a routing weight mask, the routing weight mask representing a weight value of each routing expert network corresponding to each word in the input tensor, and the weight value of the routing expert network corresponding to the word in the routing weight mask being 0 when the routing expert network is not activated by the word, wherein the tensor size of the input tensor is less than or equal to a preset tensor size; each routing expert network performs operation processing on the input tensor to obtain a first output tensor of each routing expert network; and the first output tensor of each routing expert network and the routing weight mask are subjected to product accumulation processing to obtain a final operation result of the hybrid expert operator. The method can improve the performance of the hybrid expert operator.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence technology, and in particular to an optimization method, apparatus, computer device, storage medium, and program product using hybrid expert operators. Background Technology

[0002] Hybrid expert architecture has become the mainstream structure for large models. The MOE (Mixture of Experts) operator within the model contains a large number of routing expert modules, resulting in high overall computational overhead. Therefore, during the inference phase, the MOE operator is the core bottleneck restricting model latency. However, inference scenarios with short inputs and a small number of tokens (or shapes) are common requirements in online question answering, single-turn interactions, and other businesses. These scenarios are more sensitive to inference latency and urgently require targeted optimization of the MOE operator's execution flow.

[0003] Among related technologies, the parallel optimization method for MOE operators is mainly EP (Experts Parallel). This scheme distributes all routing experts across multiple GPUs, with each GPU performing expert FFN (Feed Forward Network) calculations in parallel, which can shorten the expert computation time to some extent. However, EP parallelism relies on data rearrangement operations such as Permutation and Unpermutation, combined with All-To-All (many-to-many full interconnection) communication between GPUs to form a complete Dispatch / Combine link.

[0004] However, in small shape scenarios (such as 16 tokens), the computational load of the model is greatly reduced, and the proportion of fixed scheduling overhead generated by the above data rearrangement operation will increase significantly, making it impossible to effectively reduce the overall inference latency. Summary of the Invention

[0005] Therefore, it is necessary to provide an optimization method, apparatus, computer equipment, storage medium, and program product for hybrid expert operators that can effectively reduce the overall inference latency of hybrid expert operators, addressing the aforementioned technical problems.

[0006] In a first aspect, this application provides an optimization method using a hybrid expert operator, wherein the hybrid expert operator includes a routing operator and multiple routing expert networks, and the method includes:

[0007] The routing operator is used to perform routing calculation on the input tensor to obtain a routing weight mask. The routing weight mask represents the weight value of each term in the input tensor corresponding to each routing expert network. When the routing expert network is not activated by the term, the weight value of the term corresponding to the routing expert network in the routing weight mask is 0. The tensor size of the input tensor is less than or equal to a preset tensor size.

[0008] Each of the routing expert networks performs operations on the input tensor to obtain the first output tensor of each of the routing expert networks;

[0009] The first output tensor of each routing expert network is multiplied and summed with the routing weight mask to obtain the final operation result of the hybrid expert operator.

[0010] In one embodiment, each of the routing expert networks is deployed in multiple artificial intelligence chips using an expert parallel strategy. The step of multiplying and summing the first output tensor of each of the routing expert networks with the routing weight mask to obtain the final computation result of the hybrid expert operator includes:

[0011] Within any of the AI ​​chips, a multiplication-accumulation process is performed on the first output tensor of each of the local routing expert networks of the AI ​​chip and the routing weight mask to obtain the second output tensor of the AI ​​chip.

[0012] The second output tensor of each of the aforementioned artificial intelligence chips is subjected to full reduction communication processing to obtain the final computation result of the hybrid expert operator.

[0013] In one embodiment, each of the routing expert networks is deployed in the same artificial intelligence chip. The step of multiplying and summing the first output tensor of each of the routing expert networks with the routing weight mask to obtain the final computation result of the hybrid expert operator includes:

[0014] Within the artificial intelligence chip, multiplication and accumulation processing is performed based on the first output tensor of each routing expert network and the routing weight mask to obtain the final computation result of the hybrid expert operator.

[0015] In one embodiment, the hybrid expert operator further includes a shared expert network, which is deployed in each of the AI ​​chips using a tensor parallel strategy. The second output tensor is obtained by performing multiplication and accumulation processing on the first output tensor of each of the routing expert networks local to the AI ​​chip and the routing weight mask, including:

[0016] Within any of the aforementioned AI chips, the shared expert network performs arithmetic operations on the input tensor to obtain a third output tensor. Based on the first output tensor of each of the routing expert networks local to the AI ​​chip and the routing weight mask, a multiplication and accumulation process is performed to obtain a fourth output tensor. The third output tensor and the fourth output tensor are then accumulated to obtain a second output tensor.

[0017] In one embodiment, the hybrid expert operator further includes a shared expert network, which is fully loaded into each of the artificial intelligence chips. The step of performing full reduction communication processing on the second output tensors of each of the artificial intelligence chips to obtain the final computation result of the hybrid expert operator includes:

[0018] The second output tensor of each of the aforementioned artificial intelligence chips is subjected to full reduction communication processing to obtain the full reduction result;

[0019] Within any of the aforementioned AI chips, the shared expert network performs computational processing on the input tensor to obtain a third output tensor. Then, the full reduction result is summed with the third output tensor to obtain the final computational result of the hybrid expert operator.

[0020] In one embodiment, the hybrid expert operator further includes a shared expert network, which is fully loaded into each of the artificial intelligence chips. Within each artificial intelligence chip, based on the first output tensor of each routing expert network and the routing weight mask, multiplication and accumulation processing is performed to obtain the final computation result of the hybrid expert operator, including:

[0021] Within the AI ​​chip, the shared expert network performs computational processing on the input tensor to obtain a third output tensor. Based on the first output tensor of each routing expert network and the routing weight mask, multiplication and accumulation processing is performed to obtain a fifth output tensor. An accumulation operation is performed on the third output tensor and the fifth output tensor to obtain the final computation result of the hybrid expert operator.

[0022] Secondly, this application also provides an optimization apparatus for a hybrid expert operator, wherein the hybrid expert operator includes a routing operator and multiple routing expert networks, and the apparatus includes:

[0023] The first operation module is used to perform routing calculations on the input tensor using the routing operator to obtain a routing weight mask. The routing weight mask represents the weight value of each term in the input tensor corresponding to each routing expert network. When the routing expert network is not activated by the term, the weight value of the term corresponding to the routing expert network in the routing weight mask is 0. The tensor size of the input tensor is less than or equal to a preset tensor size.

[0024] The second computation module is used for each of the routing expert networks to perform computational processing on the input tensor to obtain the first output tensor of each of the routing expert networks.

[0025] The third computation module is used to perform multiplication and summation on the first output tensor of each of the routing expert networks and the routing weight mask to obtain the final computation result of the hybrid expert operator.

[0026] In one embodiment, each of the routing expert networks is deployed in multiple artificial intelligence chips using an expert parallel strategy. The step of multiplying and summing the first output tensor of each of the routing expert networks with the routing weight mask to obtain the final computation result of the hybrid expert operator includes:

[0027] Within any of the AI ​​chips, a multiplication-accumulation process is performed on the first output tensor of each of the local routing expert networks of the AI ​​chip and the routing weight mask to obtain the second output tensor of the AI ​​chip.

[0028] The second output tensor of each of the aforementioned artificial intelligence chips is subjected to full reduction communication processing to obtain the final computation result of the hybrid expert operator.

[0029] In one embodiment, each of the routing expert networks is deployed in the same artificial intelligence chip. The step of multiplying and summing the first output tensor of each of the routing expert networks with the routing weight mask to obtain the final computation result of the hybrid expert operator includes:

[0030] Within the artificial intelligence chip, multiplication and accumulation processing is performed based on the first output tensor of each routing expert network and the routing weight mask to obtain the final computation result of the hybrid expert operator.

[0031] In one embodiment, the hybrid expert operator further includes a shared expert network, which is deployed in each of the AI ​​chips using a tensor parallel strategy. The second output tensor is obtained by performing multiplication and accumulation processing on the first output tensor of each of the routing expert networks local to the AI ​​chip and the routing weight mask, including:

[0032] Within any of the aforementioned AI chips, the shared expert network performs arithmetic operations on the input tensor to obtain a third output tensor. Based on the first output tensor of each of the routing expert networks local to the AI ​​chip and the routing weight mask, a multiplication and accumulation process is performed to obtain a fourth output tensor. The third output tensor and the fourth output tensor are then accumulated to obtain a second output tensor.

[0033] In one embodiment, the hybrid expert operator further includes a shared expert network, which is fully loaded into each of the artificial intelligence chips. The step of performing full reduction communication processing on the second output tensors of each of the artificial intelligence chips to obtain the final computation result of the hybrid expert operator includes:

[0034] The second output tensor of each of the aforementioned artificial intelligence chips is subjected to full reduction communication processing to obtain the full reduction result;

[0035] Within any of the aforementioned AI chips, the shared expert network performs computational processing on the input tensor to obtain a third output tensor. Then, the full reduction result is summed with the third output tensor to obtain the final computational result of the hybrid expert operator.

[0036] In one embodiment, the hybrid expert operator further includes a shared expert network, which is fully loaded into each of the artificial intelligence chips. Within each artificial intelligence chip, based on the first output tensor of each routing expert network and the routing weight mask, multiplication and accumulation processing is performed to obtain the final computation result of the hybrid expert operator, including:

[0037] Within the AI ​​chip, the shared expert network performs computational processing on the input tensor to obtain a third output tensor. Based on the first output tensor of each routing expert network and the routing weight mask, multiplication and accumulation processing is performed to obtain a fifth output tensor. An accumulation operation is performed on the third output tensor and the fifth output tensor to obtain the final computation result of the hybrid expert operator.

[0038] Thirdly, this application also provides a computer device, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the optimization method of the hybrid expert operator of any of the above.

[0039] Fourthly, this application also provides a computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the optimization method of the hybrid expert operator of any of the above.

[0040] Fifthly, this application also provides a computer program product, including a computer program that, when executed by a processor, implements an optimization method for a hybrid expert operator as described above.

[0041] The aforementioned optimization method, apparatus, computer equipment, storage medium, and program product of the hybrid expert operator employs a routing operator to perform routing calculations on the input tensor, obtaining a routing weight mask. This routing weight mask represents the weight value of each term in the input tensor corresponding to each routing expert network. When a routing expert network is not activated by a term, the weight value of the routing expert network corresponding to that term in the routing weight mask is 0. The tensor size of the input tensor is less than or equal to a preset tensor size. Each routing expert network performs calculations on the input tensor to obtain the first output tensor of each routing expert network. The first output tensor of each routing expert network is then multiplied and accumulated with the routing weight mask to obtain the final calculation result of the hybrid expert operator. The optimization method, apparatus, computer equipment, storage medium, and program product of the hybrid expert operator provided in the embodiments of this application eliminate the operation of splitting and distributing tokens in scenarios with small tensor size. All routing expert networks read the complete input tensor for computation and achieve sparse weighted fusion through routing weight mask. This completely eliminates data rearrangement of permutation / unpermutation and all-to-all many-to-many communication, thereby effectively reducing the overall inference latency of the hybrid expert operator in scenarios with small tensor size, which also improves the inference efficiency of the hybrid expert operator. Attached Figure Description

[0042] To more clearly illustrate the technical solutions in the embodiments of this application or related technologies, the drawings used in the description of the embodiments of this application or related technologies will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.

[0043] Figure 1 This is a schematic diagram of the logic of a hybrid expert operator in one embodiment that does not employ the EP parallel strategy;

[0044] Figure 2 This is a schematic diagram of a hybrid expert operator employing the EP parallel strategy in one embodiment;

[0045] Figure 3 This is a flowchart illustrating an optimization method using hybrid expert operators in one embodiment;

[0046] Figure 4 This is a schematic diagram of the structure of an artificial intelligence chip in one embodiment;

[0047] Figure 5 This is a flowchart illustrating the full reduction operation steps in an EP parallel strategy implementation example.

[0048] Figure 6 This is a flowchart illustrating the full reduction operation steps in one embodiment when the TP strategy is not used;

[0049] Figure 7 This is a schematic diagram of the optimization logic of a hybrid expert algorithm in one embodiment;

[0050] Figure 8 This is a schematic diagram of a performance test in one embodiment;

[0051] Figure 9 This is a structural block diagram of an optimization device that uses hybrid expert operators in one embodiment;

[0052] Figure 10 This is an internal structural diagram of a computer device in one embodiment. Detailed Implementation

[0053] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.

[0054] It should be noted that the terms "first," "second," etc., used in this application can be used to describe various elements, but these elements are not limited by these terms. These terms are only used to distinguish the first element from the second element. The terms "comprising" and "having," and any variations thereof, used in this application, are intended to cover non-exclusive inclusion. The term "multiple" used in this application refers to two or more. The term "and / or" used in this application refers to one of the embodiments, or any combination of multiple embodiments.

[0055] Reference Figure 1 As shown, a typical logical diagram of a hybrid expert operator without employing an EP parallel strategy is illustrated. Figure 1 As shown: After the input tensor enters the hybrid expert operator, routing calculation is performed first, selecting the routing expert network activated by each term and the weights of the routing expert networks participating in the output. Then, a permutation operation is performed, filtering out the terms that the current routing expert network within the current AI chip needs to calculate, and inputting them into the selected routing expert network for calculation. After the routing expert network has completed its calculation, an unpermutation operation is performed to restore the calculation results of all different routing expert networks activated by all terms to the same tensor size as the input. Finally, an accumulation (or Add) operator is used to sum all routing expert results for all terms, as well as the shared expert results, to obtain the final output.

[0056] Reference Figure 2 As shown, a typical logical diagram of a hybrid expert operator employing an EP parallel strategy is illustrated, such as... Figure 2 As shown, the EP parallel strategy distributes each routing expert network across different AI chips, prepares corresponding inputs for each network, and then performs routing expert computations in parallel across the AI ​​chips. This can, to some extent, mask the time-consuming computation of routing expert networks in hybrid expert operators, thereby improving the inference performance of the hybrid expert operators. In the EP parallel strategy scenario, permutation / unpermutation operations still exist, combining with the all-to-all operations between operators to form chip-to-chip dispatch / combination operations.

[0057] The Dispatch / Permutation and Combine / Unpermutation operations are the performance bottlenecks of the hybrid expert operator. These operations are time-consuming, especially in computational scenarios with small tensor sizes (such as 16 tokens). The fixed scheduling overhead generated by these operations will increase significantly, making it impossible to effectively reduce the overall inference latency.

[0058] To address the aforementioned technical issues, this application provides an optimization method for hybrid expert operators. In scenarios with small tensor sizes, the splitting and distribution of tokens is no longer performed. All routing expert networks read the complete input tensor for computation and achieve sparse weighted fusion through the routing weight mask calculated by the routing operator. This completely eliminates data rearrangement under permutation / unpermutation and all-to-all many-to-many communication, thereby effectively reducing the overall inference latency of hybrid expert operators in scenarios with small tensor sizes, which in turn improves the inference efficiency of hybrid expert operators.

[0059] like Figure 3As shown, a hybrid expert operator optimization method is provided and applied to artificial intelligence chips. In this embodiment, the artificial intelligence chip is any one of GPU (Graphics Processing Unit), TPU (Tensor Processing Unit), NPU (Neural Network Processing Unit), DPU (Deep Learning Processing Unit), APU (Accelerated Processing Unit), and GPGPU (General-Purpose Graphics Processing Unit). This embodiment does not specifically limit the specific type of chip, and the following description uses GPGPU as an example.

[0060] Reference Figure 4 The diagram shows a schematic of a GPGPU. A GPGPU is actually an array of Streaming Processor Clusters (SPCs), including, for example,... Figure 4 The diagram shows streaming processor clusters 1, ..., M, where M is a positive integer greater than 1. In a graphics processing unit (GPU), one streaming processor cluster processes one computational task, or multiple streaming processor clusters process one computational task. Multiple streaming processor clusters share data through a global cache or global memory.

[0061] like Figure 4 As shown, taking streaming processor cluster 1 as an example, one streaming processor cluster includes multiple computing units, such as... Figure 4 The system is structured as Computation Unit 1, Computation Unit 2, ..., Computation Unit N, where N is a positive integer. Each Computation Unit (CU) performs arithmetic and logical operations other than matrix calculations such as matrix multiplication and convolution, including operations like accumulation, reduction, and standard addition, subtraction, multiplication, and division. A Computation Unit contains multiple cores (also called computational kernels), each including an Arithmetic Logic Unit (ALU), a floating-point unit, etc., which are used to execute specific computational tasks. Furthermore, the Computation Unit also includes registers (e.g., ...). Figure 4 The register file and shared cache in a computing unit are used to store source and destination data related to computing tasks in a hierarchical manner. The shared cache in a computing unit is used to share data between the cores of that computing unit.

[0062] In parallel computing, computational tasks are typically executed by multiple threads. These threads are divided into multiple thread blocks before execution in a general-purpose graphics processor (or parallel computing processor), and then dispatched via a thread block distribution module. Figure 4 (Not shown in the image) Multiple thread blocks are distributed to various computation units. All threads in a thread block must be assigned to the same computation unit for execution. Simultaneously, thread blocks are broken down into minimum execution thread bundles (or simply warps), each containing a fixed number (or less than this fixed number) of threads, for example, 32 threads. Multiple thread blocks can execute in the same computation unit or in different computation units.

[0063] In each computing unit, the thread beam scheduling / distribution module ( Figure 4 (Not shown in the diagram) Thread bundles are scheduled and allocated so that multiple computing cores within the computing unit can run thread bundles. Depending on the number of computing cores in the computing unit, multiple thread bundles within a thread block can be executed concurrently or in a time-sharing manner. Multiple threads within each thread bundle execute the same instructions. Memory execution instructions are issued to the shared cache within the computing unit or further issued to intermediate-level caches, global caches, or global memory for read and write operations, etc.

[0064] like Figure 4 As shown, the streaming processor cluster 1 also includes a tensor operation unit, which is used to perform tensor calculations, such as matrix multiplication, convolution operations, etc.

[0065] In this embodiment of the application, the hybrid expert operator includes a routing operator and multiple routing expert networks, as referred to... Figure 3 As shown in the embodiment of this application, an optimization method using a hybrid expert operator includes the following steps: 302, 304, and 306. Wherein:

[0066] Step 302: Use a routing operator to perform routing calculations on the input tensor to obtain a routing weight mask. The routing weight mask represents the weight value of each term in the input tensor corresponding to each routing expert network. When the routing expert network is not activated by a term, the weight value of the term corresponding to the routing expert network in the routing weight mask is 0. The tensor size of the input tensor is less than or equal to the preset tensor size.

[0067] The hybrid expert operator optimization method provided in this application can be applied to artificial intelligence fields such as natural language processing and image recognition, adapting to online inference interaction scenarios and reducing inference latency under small tensor size inputs. The following embodiments will use text interaction inference in natural language processing tasks as an example for illustration. Assume that the input tensor is the hidden state tensor after model encoding. This hidden state tensor is composed of multiple tokens (which can also be represented as tokens) obtained from text segmentation, forming the token sequence [x1, x2, ..., xS] corresponding to the input tensor.

[0068] In this embodiment, an input tensor with a size less than or equal to a preset tensor threshold can be input into the routing operator (the preset tensor threshold is a value determined by those skilled in the art based on experience or experimentation; an input tensor with a size less than or equal to the preset tensor threshold can be considered a small-shape input tensor). The routing operator can calculate the matching degree for each term in the input tensor based on pre-trained routing parameters, determining which routing expert network among multiple routing expert networks is suitable for processing that term. For example, the routing operator filters out the active routing expert corresponding to a single term by calculating the matching degree between the term features and the features of each routing expert network, and assigns a corresponding weight. Routing experts that are not suitable for processing that term are determined to be inactive. The routing operator can ultimately output a routing weight mask, which carries the weight information of each term corresponding to each routing expert network.

[0069] In one exemplary embodiment, the data shape of the route weight mask is [E, N, S], where E represents the total number of routing expert networks, N represents the batch concurrency, and S represents the length of a single batch sequence. That is, each element M in the route weight mask M... ij This represents the activation state of the i-th routing expert network for the j-th word. When it is in an activated state for the j-th word, this M... ij For the corresponding weighted weight value, if it corresponds to the j-th word element being in an inactive state, then M... ij If the value is assigned to 0, the routing operator in this embodiment does not need to output auxiliary tensors such as expert index and activation number.

[0070] Step 304: Each routing expert network performs calculations on the input tensor to obtain the first output tensor of each routing expert network.

[0071] In this embodiment, all routing expert networks read the complete input tensor to perform feedforward operations. That is, each routing expert network can independently expand and operate on the entire input tensor word by word. After each routing expert network completes its operation, it outputs a first output tensor with dimensions and shape completely identical to the input tensor. Even if some routing expert networks are not activated by the current word, they still complete the feedforward operation on the entire input tensor.

[0072] Step 306: Perform multiplication and summation on the first output tensor of each routing expert network and the routing weight mask to obtain the final operation result of the hybrid expert operator.

[0073] In this embodiment, the first output tensor corresponding to each routing expert network can be obtained, and multiplication and summation processing can be performed in combination with the routing weight mask. This includes: multiplying the first output tensor corresponding to a single routing expert network element by element with the corresponding weight in the routing weight mask. The weight corresponding to an inactive routing expert network is 0, and the multiplication can directly cancel the operation output of that expert. After completing the multiplication of all routing expert networks, the element-wise summation of all product tensors is performed to finally obtain the final operation result of the hybrid expert operator.

[0074] The optimization method of the hybrid expert operator provided in this application embodiment uses a routing operator to perform routing calculations on the input tensor to obtain a routing weight mask. This routing weight mask represents the weight value of each term in the input tensor corresponding to each routing expert network. When a routing expert is not activated, the weight value corresponding to the routing expert network in the routing weight mask is 0. The tensor size of the input tensor is less than or equal to a preset tensor size. Each routing expert network performs calculations on the input tensor to obtain the first output tensor of each routing expert network. The first output tensor of each routing expert network is then multiplied and accumulated with the routing weight mask to obtain the final calculation result of the hybrid expert operator. The optimization method of hybrid expert operators provided in this application eliminates the splitting and distribution of tokens in scenarios with small tensor sizes. All routing expert networks read the complete input tensor for computation and achieve sparse weighted fusion through routing weight masks. This completely eliminates data rearrangement of permutation / unpermutation and all-to-all many-to-many communication, thereby effectively reducing the overall inference latency of hybrid expert operators in scenarios with small tensor sizes, which also improves the inference efficiency of hybrid expert operators.

[0075] In one exemplary embodiment, each routing expert network is deployed in multiple artificial intelligence chips using an expert parallel strategy, referring to... Figure 5As shown, in step 306, the first output tensor of each routing expert network is multiplied and accumulated with the routing weight mask to obtain the final operation result of the hybrid expert operator. This may include the following steps 502 and 504, wherein:

[0076] Step 502: Inside any AI chip, based on the first output tensor of each routing expert network in the local AI chip and the routing weight mask, perform multiplication and accumulation processing to obtain the second output tensor of the AI ​​chip.

[0077] Step 504: Perform full reduction communication processing on the second output tensor of each artificial intelligence chip to obtain the final computation result of the hybrid expert operator.

[0078] In this embodiment, an EP (Electronic Programming) expert parallel deployment strategy is adopted, distributing all routing expert networks evenly across multiple AI chips, with each AI chip only supporting a portion of the routing expert networks. Each AI chip reads the complete input tensor, and the locally deployed routing expert networks independently perform feedforward operations to generate the corresponding first output tensor. Each AI chip can retrieve a globally unified routing weight mask, matching only the weights corresponding to its local assigned routing expert network. The first output tensor of each local routing expert network is then multiplied element-wise with the matched weights to obtain the corresponding product tensor. The weights of inactive routing expert networks are 0, and these invalid outputs are automatically discarded after multiplication. The product tensors of all local routing expert networks are further summed element-wise within the chip to obtain the second output tensor corresponding to that AI chip. The second output tensor generated by each AI chip represents only the local feature tensor after weighted fusion of the local routing expert networks.

[0079] After all AI chips complete local multiplication and accumulation, generating their respective second output tensors, AllReduce communication is performed on the second output tensors of all AI chips. AllReduce communication enables global accumulation of tensors across chips, integrating the second output tensors of the routing expert networks within all AI chips. After accumulation, each AI chip obtains a complete tensor integrating the outputs of all routing expert networks, which is the final result of the hybrid expert operator.

[0080] The optimization method for hybrid expert operators provided in this application can be adapted to multi-chip distributed inference scenarios. After completing the expert weighted accumulation operation locally on the chip, the global feature aggregation is completed through a single AllReduce communication process. That is, this application only retains a single lightweight reduction communication and discards All-To-All communication and the corresponding data rearrangement operator, which can reduce the communication scheduling time in multi-chip small-size inference scenarios and improve the inference efficiency of hybrid expert operators.

[0081] In one exemplary embodiment, each routing expert network is deployed in the same artificial intelligence chip. The first output tensor of each routing expert network is multiplied and accumulated with the routing weight mask to obtain the final computation result of the hybrid expert operator, which may include:

[0082] Within the AI ​​chip, based on the first output tensor of each routing expert network and the routing weight mask, multiplication and accumulation processing is performed to obtain the final computation result of the hybrid expert operator.

[0083] In this embodiment, a single-chip centralized deployment strategy is adopted, with all routing expert networks deployed within the same AI chip, eliminating the need to split expert computing power or perform cross-chip data interaction. This AI chip can read the complete input tensor, and all routing expert networks within the chip synchronously perform feedforward operations on the input tensor, generating their respective first output tensors. After retrieving the routing weight mask, the chip multiplies the first output tensor of each routing expert network element-wise with the weights at matching positions in the routing weight mask, obtaining the product tensor of each routing expert network. The weights corresponding to inactive routing expert networks are 0, and multiplication directly eliminates the invalid computation output of that expert. The chip can directly perform element-wise summation on all the product tensors of the expert networks; the summed tensor is the final computation result of the hybrid expert operator.

[0084] The optimization method for hybrid expert operators provided in this application can be adapted to single-chip localized inference scenarios. All routing expert networks perform operations in parallel and output filtering is completed based on routing weight masks. This eliminates operations such as word sorting and rearrangement scheduling within the chip, which can minimize scheduling overhead, simplify the execution process of hybrid expert operators, reduce the inference time of single-chip hybrid expert operators in small tensor input scenarios, and improve the inference efficiency of hybrid expert operators.

[0085] In one exemplary embodiment, the hybrid expert operator further includes a shared expert network, which is deployed in each AI chip using a tensor parallel strategy. Based on the first output tensor of each routing expert network on the AI ​​chip and the routing weight mask, multiplication and accumulation processing is performed to obtain a second output tensor, including:

[0086] Within any AI chip, the shared expert network performs operations on the input tensor to obtain a third output tensor. Based on the first output tensor of each routing expert network in the AI ​​chip and the routing weight mask, multiplication and accumulation processing is performed to obtain a fourth output tensor. The third output tensor and the fourth output tensor are then accumulated to obtain a second output tensor.

[0087] This application embodiment adopts a hybrid parallel deployment architecture of parallel routing expert network (EP) and parallel shared expert network (TP). Each artificial intelligence chip simultaneously carries a portion of the routing expert network and a segmented shared expert network (the shared expert network is divided and deployed to each artificial intelligence chip, and the segmented shared expert network is the part deployed in the artificial intelligence chip). The routing expert network belonging to the chip reads the complete input tensor to complete the feedforward operation, and after obtaining the first output tensor of the routing expert network, it combines the routing weight mask to complete the local multiplication and accumulation processing (the specific process can be referred to the relevant description in the previous embodiment, and will not be repeated here in this application embodiment), and then integrates to obtain the fourth output tensor of the artificial intelligence chip.

[0088] Simultaneously, the AI ​​chip can load the slice weights of shared experts (i.e., a slice shared expert network). This slice shared expert network can independently perform feedforward operations on the complete input tensor, generating the corresponding slice operation results, i.e., the third output tensor. Finally, the AI ​​chip locally adds the third output tensor and the fourth output tensor element by element to obtain the final second output tensor of the AI ​​chip.

[0089] In traditional hybrid parallel architectures, routing expert aggregation and shared expert aggregation each require a separate AllReduce full reduction communication, incurring overhead from two communication initialization and data transmission / reception operations. In small-shape scenarios, this communication overhead further amplifies inference latency. The optimized hybrid expert operator method provided in this application, by eliminating operator overhead such as data rearrangement, pre-merges the computation results of the routing expert network and the shared expert network locally on the chip. Only one global AllReduce full reduction is needed to simultaneously complete the global aggregation of the routing expert network and the fragmented aggregation of the shared expert network. Compared to traditional solutions, this reduces the number of cross-chip communications by half, further compressing the time consumption of distributed inference communication and improving the inference speed of the hybrid expert operator with small-size input tensors.

[0090] In one exemplary embodiment, the hybrid expert operator further includes a shared expert network, which is fully loaded within each AI chip, as described above. Figure 6 As shown, the second output tensor of each artificial intelligence chip undergoes full reduction communication processing to obtain the final computation result of the hybrid expert operator, which may include steps 602 and 604, wherein:

[0091] Step 602: Perform full reduction communication processing on the second output tensor of each artificial intelligence chip to obtain the full reduction result;

[0092] Step 604: Within any AI chip, execute the shared expert network to process the input tensor to obtain the third output tensor, and then perform the full reduction result and the accumulation process of the third output tensor to obtain the final calculation result of the hybrid expert operator.

[0093] This application embodiment adopts a deployment architecture of parallel routing expert networks (EPs) and parallel shared expert networks (TPs) without TPs. The routing expert networks are evenly deployed across chips, with each AI chip carrying only a portion of the routing expert network, and all chips fully loading the weights of the complete shared expert network. Each AI chip first performs local multiplication and accumulation based on its local routing expert network and routing weight mask (the specific process can be referred to the relevant description in the previous embodiment, and will not be repeated here), generating the second output tensor of each AI chip. Subsequently, AllReduce full reduction communication is performed on the second output tensors of all AI chips to complete the accumulation and fusion of the output tensors of the global routing expert network. At this point, each AI chip can obtain the full reduction result integrating the outputs of all routing expert networks.

[0094] Simultaneously, each AI chip can independently access its complete local shared expert network, read the original complete input tensor to perform feedforward operations, and directly generate a complete shared feature tensor, i.e., the third output tensor. Each AI chip can then element-wise accumulate the full reduction result with the locally generated third output tensor, fusing routing expert features and shared expert features to ultimately obtain the final computation result of the hybrid expert operator.

[0095] The optimization method for hybrid expert operators provided in this application still abandons permutation, unpermutation and other rearrangement operators as well as all-to-all distribution and merging communication. It only performs a single AllReduce full reduction for local routing features. Shared experts are calculated independently on each chip without participating in cross-chip reduction, which can reduce inference latency in small shape scenarios and improve the performance of hybrid expert operators.

[0096] In an exemplary embodiment, the hybrid expert operator further includes a shared expert network. Each AI chip fully loads the shared expert network. Within the AI ​​chip, based on the first output tensor of each routing expert network and the routing weight mask, multiplication and accumulation processing is performed to obtain the final computation result of the hybrid expert operator, including:

[0097] Within the AI ​​chip, the shared expert network performs operations on the input tensor to obtain the third output tensor. Based on the first output tensor of each routing expert network and the routing weight mask, multiplication and accumulation processing is performed to obtain the fifth output tensor. The third and fifth output tensors are then accumulated to obtain the final result of the hybrid expert operator.

[0098] This application embodiment adopts a single-chip deployment architecture integrating routing expert networks and shared expert networks. All routing expert networks and the complete shared expert network are deployed within the same artificial intelligence chip, without expert computing power splitting or cross-chip data interaction and communication operations. The artificial intelligence chip performs two types of operations simultaneously and in parallel: On one hand, all routing expert networks within the chip read the complete input tensor and complete feedforward operations to obtain their respective first output tensors. These are then combined with the global routing weight mask through element-wise multiplication and accumulation (the specific process is described in the aforementioned embodiments and will not be repeated here), resulting in a fifth output tensor. On the other hand, the chip fully loads the weights of the shared expert network and independently performs feedforward operations on the original input tensor, directly generating a third output tensor with complete shared features. The chip locally accumulates the fifth and third output tensors element-wise to fuse the two types of expert features; the resulting tensor is the final result of the hybrid expert operator.

[0099] The embodiments of this application can be adapted to local inference scenarios of extremely simple chips. It can eliminate the complete set of scheduling operators for word distribution, sorting and merging within the chip. That is, there is no data rearrangement or cross-chip communication throughout the process. Feature fusion is completed solely by tensor multiplication and addition and tensor accumulation within the chip. This can reduce the memory scheduling time of MoE operators under small input size and reduce edge inference latency.

[0100] To enable those skilled in the art to better understand the embodiments of this application, the embodiments of this application are described below through specific examples.

[0101] Reference Figure 7 The diagram illustrates the optimization logic of the hybrid expert algorithm in this embodiment, including: after the input features enter the hybrid expert operator, routing calculation is performed first, selecting the routing expert network (Index) activated by each term and the weights (Scale) of the routing expert networks participating in the output. In this embodiment, the Index and Scale are merged into a routing weight mask. For terms not selected by the current expert, the Mask is set to 0. For each routing expert network, the hidden state of all terms is calculated, and the calculated routing expert result is multiplied by the Mask to obtain the final result of each routing expert network.

[0102] For parallel EP scenarios, the accumulation between routing expert networks can be combined with AllReduce to accumulate the results of selected routing experts. This actually accumulates the results of all experts, but the results of unselected experts are 0, so the accumulation has no impact. For scenarios without parallel EP, the results of the routing expert networks can be directly accumulated.

[0103] For EP parallel scenarios, the shared expert network can be parallelized using TP. For example, assuming there are 256 routing expert networks, EP parallelization distributes the 256 routing expert networks to 8 GPUs, and then TP parallelization distributes one shared expert network to 8 GPUs. Each GPU is responsible for the computation of 32 routing expert networks and 1 / 8 of the shared expert network. All computation results are shared by AllReduce to accumulate the results and obtain the final result.

[0104] The optimization method for hybrid expert operators provided in this application combines the scaling (weights of the output) of the routing expert network with the mask operation to eliminate permutation / unpermutation operations. The mask is used to control the computation result of the routing expert network for each token. Here, the routing expert network does not distinguish which tokens are selected during computation; all computations are performed through masking to achieve permutation filtering. Because there is no token reorganization, there is no need to implement unpermutation. Eliminating permutation / unpermutation reduces computational load and achieves latency optimization.

[0105] In the small-shape scenario, this application's embodiments eliminate the performance overhead of Permutation / Unpermutation or Dispatch / Combine, significantly improving performance. When using EP parallelism in the small-shape scenario, the Add operator is merged with AllReduce to save computation and reduce latency. In the scenario where the shared expert network uses the TP parallel scheme, the shared expert network and the routing expert network share a single AllRedue communication operation to obtain the final hybrid expert operator's calculation result. TP parallelism for shared experts can significantly improve the performance of shared experts in the MoE module. Figure 8 The inference performance of the embodiments of this application implemented on a test chip is shown. Compared with the solution that does not eliminate Dispatch / Combine, the technical solution provided by the embodiments of this application has a significant performance improvement under a small shape input with 192 concurrent connections.

[0106] It should be understood that although the steps in the flowcharts of the embodiments described above are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the embodiments described above may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages in other steps. It is understood that the steps in different embodiments can be freely combined as needed, and all non-contradictory solutions formed by such combinations are within the scope of protection of this application.

[0107] Based on the same inventive concept, this application also provides an optimization apparatus for implementing the above-described optimization method of hybrid expert operators. The solution provided by this apparatus is similar to the implementation described in the above method. Therefore, the specific limitations in one or more embodiments of the optimization apparatus for hybrid expert operators provided below can be found in the limitations of the optimization method for hybrid expert operators described above, and will not be repeated here.

[0108] In one exemplary embodiment, such as Figure 9 As shown, an optimization device 900 with a hybrid expert operator is provided. The hybrid expert operator includes a routing operator and multiple routing expert networks, comprising: a first operation module 902, a second operation module 904, and a third operation module 906, wherein:

[0109] The first operation module 902 is used to perform routing calculations on the input tensor using the routing operator to obtain a routing weight mask. The routing weight mask represents the weight value of each word in the input tensor corresponding to each routing expert network. When the routing expert network is not activated by a word, the weight value of the word corresponding to the routing expert network in the routing weight mask is 0. The tensor size of the input tensor is less than or equal to a preset tensor size.

[0110] The second operation module 904 is used for each of the routing expert networks to perform operations on the input tensor to obtain the first output tensor of each of the routing expert networks.

[0111] The third operation module 906 is used to perform multiplication and summation on the first output tensor of each of the routing expert networks and the routing weight mask to obtain the final operation result of the hybrid expert operator.

[0112] The aforementioned optimization device for the hybrid expert operator uses a routing operator to perform routing calculations on the input tensor to obtain a routing weight mask. This routing weight mask represents the weight value of each term in the input tensor corresponding to each routing expert network. When a routing expert is not activated, the weight value corresponding to the routing expert network in the routing weight mask is 0. The tensor size of the input tensor is less than or equal to a preset tensor size. Each routing expert network performs calculations on the input tensor to obtain the first output tensor of each routing expert network. The first output tensor of each routing expert network is then multiplied and accumulated with the routing weight mask to obtain the final calculation result of the hybrid expert operator. The optimization device for hybrid expert operators provided in this application eliminates the splitting and distribution of tokens in scenarios with small tensor sizes. All routing expert networks read the complete input tensor for computation and achieve sparse weighted fusion through routing weight masks. This completely eliminates data rearrangement for permutation / unpermutation and many-to-many communication for all-to-all, thereby effectively reducing the overall inference latency of hybrid expert operators in scenarios with small tensor sizes, which also improves the inference efficiency of hybrid expert operators.

[0113] In one embodiment, each of the routing expert networks is deployed in multiple artificial intelligence chips using an expert parallel strategy. The step of multiplying and summing the first output tensor of each of the routing expert networks with the routing weight mask to obtain the final computation result of the hybrid expert operator includes:

[0114] Within any of the AI ​​chips, a multiplication-accumulation process is performed on the first output tensor of each of the local routing expert networks of the AI ​​chip and the routing weight mask to obtain the second output tensor of the AI ​​chip.

[0115] The second output tensor of each of the aforementioned artificial intelligence chips is subjected to full reduction communication processing to obtain the final computation result of the hybrid expert operator.

[0116] In one embodiment, each of the routing expert networks is deployed in the same artificial intelligence chip. The step of multiplying and summing the first output tensor of each of the routing expert networks with the routing weight mask to obtain the final computation result of the hybrid expert operator includes:

[0117] Within the artificial intelligence chip, multiplication and accumulation processing is performed based on the first output tensor of each routing expert network and the routing weight mask to obtain the final computation result of the hybrid expert operator.

[0118] In one embodiment, the hybrid expert operator further includes a shared expert network, which is deployed in each of the AI ​​chips using a tensor parallel strategy. The second output tensor is obtained by performing multiplication and accumulation processing on the first output tensor of each of the routing expert networks local to the AI ​​chip and the routing weight mask, including:

[0119] Within any of the aforementioned AI chips, the shared expert network performs arithmetic operations on the input tensor to obtain a third output tensor. Based on the first output tensor of each of the routing expert networks local to the AI ​​chip and the routing weight mask, a multiplication and accumulation process is performed to obtain a fourth output tensor. The third output tensor and the fourth output tensor are then accumulated to obtain a second output tensor.

[0120] In one embodiment, the hybrid expert operator further includes a shared expert network, which is fully loaded into each of the artificial intelligence chips. The step of performing full reduction communication processing on the second output tensors of each of the artificial intelligence chips to obtain the final computation result of the hybrid expert operator includes:

[0121] The second output tensor of each of the aforementioned artificial intelligence chips is subjected to full reduction communication processing to obtain the full reduction result;

[0122] Within any of the aforementioned AI chips, the shared expert network performs computational processing on the input tensor to obtain a third output tensor. Then, the full reduction result is summed with the third output tensor to obtain the final computational result of the hybrid expert operator.

[0123] In one embodiment, the hybrid expert operator further includes a shared expert network, which is fully loaded into each of the artificial intelligence chips. Within each artificial intelligence chip, based on the first output tensor of each routing expert network and the routing weight mask, multiplication and accumulation processing is performed to obtain the final computation result of the hybrid expert operator, including:

[0124] Within the AI ​​chip, the shared expert network performs computational processing on the input tensor to obtain a third output tensor. Based on the first output tensor of each routing expert network and the routing weight mask, multiplication and accumulation processing is performed to obtain a fifth output tensor. An accumulation operation is performed on the third output tensor and the fifth output tensor to obtain the final computation result of the hybrid expert operator.

[0125] The modules in the aforementioned hybrid expert operator optimization device can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device, or stored in the computer device's memory as software, so that the processor can call and execute the operations corresponding to each module.

[0126] In one exemplary embodiment, a computer device is provided, which may be a terminal, and its internal structure diagram may be as follows: Figure 10 As shown, the computer device includes a processor, memory, input / output interfaces, a communication interface, a display unit, and an input device. The processor, memory, and input / output interfaces are connected via a system bus, and the communication interface, display unit, and input device are also connected to the system bus via the input / output interfaces. The processor provides computational and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The input / output interfaces are used for exchanging information between the processor and external devices. The communication interface is used for wired or wireless communication with external terminals; wireless communication can be achieved through Wi-Fi, mobile cellular networks, Near Field Communication (NFC), or other technologies. When the computer program is executed by the processor, it implements a hybrid expert operator optimization method.

[0127] Those skilled in the art will understand that Figure 10 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.

[0128] In one embodiment, a computer device is also provided, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps in the above method embodiments.

[0129] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon that, when executed by a processor, implements the steps in the above method embodiments.

[0130] In one embodiment, a computer program product is provided, including a computer program that, when executed by a processor, implements the steps in the above method embodiments.

[0131] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of the relevant data must comply with relevant regulations.

[0132] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile memory and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM). The databases involved in the embodiments provided in this application may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided in this application may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, quantum computing-based data processing logic devices, artificial intelligence (AI) processors, etc., and are not limited to these.

[0133] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this application.

[0134] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of this patent application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this application should be determined by the appended claims.

Claims

1. An optimization method using hybrid expert operators, characterized in that, The hybrid expert operator includes a routing operator and multiple routing expert networks, and the method includes: The routing operator is used to perform routing calculation on the input tensor to obtain a routing weight mask. The routing weight mask represents the weight value of each term in the input tensor corresponding to each routing expert network. When the routing expert network is not activated by the term, the weight value of the term corresponding to the routing expert network in the routing weight mask is 0. The tensor size of the input tensor is less than or equal to a preset tensor size. Each of the routing expert networks performs operations on the input tensor to obtain the first output tensor of each of the routing expert networks; The first output tensor of each routing expert network is multiplied and summed with the routing weight mask to obtain the final operation result of the hybrid expert operator.

2. The method according to claim 1, characterized in that, Each of the routing expert networks is deployed in multiple artificial intelligence chips using an expert parallel strategy. The step of multiplying and summing the first output tensor of each of the routing expert networks with the routing weight mask to obtain the final computation result of the hybrid expert operator includes: Within any of the AI ​​chips, a multiplication-accumulation process is performed on the first output tensor of each of the local routing expert networks of the AI ​​chip and the routing weight mask to obtain the second output tensor of the AI ​​chip. The second output tensor of each of the aforementioned artificial intelligence chips is subjected to full reduction communication processing to obtain the final computation result of the hybrid expert operator.

3. The method according to claim 1, characterized in that, Each of the routing expert networks is deployed in the same artificial intelligence chip. The step of multiplying and summing the first output tensor of each of the routing expert networks with the routing weight mask to obtain the final computation result of the hybrid expert operator includes: Within the artificial intelligence chip, multiplication and accumulation processing is performed based on the first output tensor of each routing expert network and the routing weight mask to obtain the final computation result of the hybrid expert operator.

4. The method according to claim 2, characterized in that, The hybrid expert operator further includes a shared expert network, which is deployed in each of the AI ​​chips using a tensor parallel strategy. The second output tensor is obtained by performing multiplication and accumulation processing on the first output tensor of each of the routing expert networks local to the AI ​​chip and the routing weight mask, including: Within any of the aforementioned AI chips, the shared expert network performs arithmetic operations on the input tensor to obtain a third output tensor. Based on the first output tensor of each of the routing expert networks local to the AI ​​chip and the routing weight mask, a multiplication and accumulation process is performed to obtain a fourth output tensor. The third output tensor and the fourth output tensor are then accumulated to obtain a second output tensor.

5. The method according to claim 2, characterized in that, The hybrid expert operator further includes a shared expert network, which is fully loaded into each of the artificial intelligence chips. The process of performing full reduction communication processing on the second output tensor of each of the artificial intelligence chips to obtain the final computation result of the hybrid expert operator includes: The second output tensor of each of the aforementioned artificial intelligence chips is subjected to full reduction communication processing to obtain the full reduction result; Within any of the aforementioned AI chips, the shared expert network performs computational processing on the input tensor to obtain a third output tensor. Then, the full reduction result is summed with the third output tensor to obtain the final computational result of the hybrid expert operator.

6. The method according to claim 3, characterized in that, The hybrid expert operator further includes a shared expert network, which is fully loaded into each of the artificial intelligence chips. Within each artificial intelligence chip, based on the first output tensor of each routing expert network and the routing weight mask, multiplication and accumulation processing is performed to obtain the final computation result of the hybrid expert operator, including: Within the AI ​​chip, the shared expert network performs computational processing on the input tensor to obtain a third output tensor. Based on the first output tensor of each routing expert network and the routing weight mask, multiplication and accumulation processing is performed to obtain a fifth output tensor. An accumulation operation is performed on the third output tensor and the fifth output tensor to obtain the final computation result of the hybrid expert operator.

7. An optimization device using hybrid expert operators, characterized in that, The hybrid expert operator includes a routing operator and multiple routing expert networks, and the apparatus includes: The first operation module is used to perform routing calculations on the input tensor using the routing operator to obtain a routing weight mask. The routing weight mask represents the weight value of each term in the input tensor corresponding to each routing expert network. When the routing expert network is not activated by the term, the weight value of the term corresponding to the routing expert network in the routing weight mask is 0. The tensor size of the input tensor is less than or equal to a preset tensor size. The second computation module is used for each of the routing expert networks to perform computational processing on the input tensor to obtain the first output tensor of each of the routing expert networks. The third computation module is used to perform multiplication and summation on the first output tensor of each of the routing expert networks and the routing weight mask to obtain the final computation result of the hybrid expert operator.

8. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 6.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 6.

10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 6.