Operator running method, electronic device, storage medium, and program product
Patent Information
- Application Number
- CN202611031242.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-10
- Publication Date
- 2026-09-25
AI Technical Summary
[0004]本发明提供一种算子运行方法、电子设备、存储介质和程序产品,用以解决相关技术中基于多裸片芯片的MoE架构下的算子运算性能较低的缺陷
[0015]本发明提供的算子运行方法、电子设备、存储介质和程序产品,结合算子所处混合专家层中的各专家的待处理数据的数据量、隐层维度和输出维度,评估各专家在多裸片上的各种排布策略的负载均衡项,并基于负载均衡项从各排布策略中选择目标排布策略,由此基于目标排布策略运行算子,实现了算子在多裸片架构下计算性能的全局最大化。
Smart Images

Figure CN122816892A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of high-performance computing technology, and in particular to an operator operation method, electronic device, storage medium, and program product. Background Technology
[0002] In the MoE (Mixture of Experts) architecture, input data is routed to different expert networks for processing. Because the data distribution is dynamic and uneven, the amount of data allocated to each expert varies. When each expert performs its own GEMM (General Matrix Multiplication) operation, the shape of their respective matrices differs, thus forming multiple GroupGEMMs (Group General Matrix Multiplications) with different shapes.
[0003] When performing GroupGEMM operations on a multi-die chip, the GEMM for each expert can be executed sequentially on each die in parallel, based on a pre-defined execution order for each expert. How to further optimize the computational performance of multi-die chips remains a problem that urgently needs to be solved by those skilled in the art. Summary of the Invention
[0004] This invention provides an operator operation method, electronic device, storage medium, and program product to address the shortcomings of low operator operation performance in related technologies based on the MoE architecture with multiple bare chips.
[0005] This invention provides an operator operation method, comprising: Obtain the data to be processed of each expert in the hybrid expert layer where the operator is located, the hidden layer dimension and output dimension of the data to be processed, and the various arrangement strategies of each expert on multiple bare dies; Based on the data volume, hidden layer dimension, and output dimension of the data to be processed by each expert, the load balancing items for various arrangement strategies are determined. Based on the load balancing items of the various arrangement strategies, a target arrangement strategy is determined from the various arrangement strategies. The operator is run based on the target layout strategy and the data to be processed by each expert.
[0006] According to an operator operation method provided by the present invention, the step of determining the load balancing item for various arrangement strategies based on the data volume, hidden layer dimension, and output dimension of the data to be processed by each expert includes: Based on the amount of data to be processed, the hidden layer dimension, and the output dimension of the data to be processed by each expert, the computational load of each expert when running the operator is determined. Based on the computational load of each expert when running the operator, the total computational load of each die under each of the various arrangement strategies is determined. Based on the difference between the maximum and minimum total computational cost of each die under the various arrangement strategies, the load balancing term for each arrangement strategy is determined.
[0007] According to an operator operation method provided by the present invention, the step of determining a target arrangement strategy from the multiple arrangement strategies based on the load balancing item of the multiple arrangement strategies includes: Based on the data volume, hidden layer dimension, and output dimension of the data to be processed by each expert, the peak memory usage of each arrangement strategy is determined. Based on the alignment granularity corresponding to the multiple bare dies and the amount of data to be processed by each expert, the hardware alignment penalty term for each arrangement strategy is determined. Based on the degree of disruption of each arrangement strategy compared to the original arrangement strategy, the scheduling overhead of each arrangement strategy is determined. Based on at least one of the memory peak usage, hardware alignment penalty, and scheduling overhead of the various arrangement strategies, and the load balancing item of the various arrangement strategies, a target arrangement strategy is determined from the various arrangement strategies.
[0008] According to an operator execution method provided by the present invention, determining the peak memory usage of various arrangement strategies based on the data volume, hidden layer dimension, and output dimension of the data to be processed by each expert includes: Based on the amount of data to be processed, the hidden layer dimension, and the output dimension of each expert, the memory usage of each expert when running the operator is determined. Based on the memory usage of each expert when running the operator, the maximum memory usage of each die under each of the various layout strategies is determined. Based on the maximum memory usage of each die under the various layout strategies, and the total memory usage of all experts when running the operators, the peak memory usage of each layout strategy is determined.
[0009] According to an operator operation method provided by the present invention, the step of determining the hardware alignment penalty term for various layout strategies based on the alignment granularity corresponding to the multiple bare dies and the data volume of the data to be processed by each expert includes: Based on the alignment granularity corresponding to the multiple bare wafers and the amount of data to be processed by each expert, the alignment penalty value of each expert is determined; Based on the alignment penalty values of each expert and the computational cost of each expert when running the operator, the hardware alignment penalty terms for the various layout strategies are determined.
[0010] According to an operator execution method provided by the present invention, determining the scheduling overhead of each arrangement strategy based on the degree of disruption of each arrangement strategy compared to the original arrangement strategy includes: Based on the degree of disruption of each arrangement strategy compared to the original arrangement strategy, and the number of experts included in the hybrid expert layer, the scheduling overhead of each arrangement strategy is determined.
[0011] An operator operation method provided by the present invention further includes: The experts are arranged in descending order of the amount of data to be processed to obtain the initial sequence; Based on the current load of each die, the experts in the initial sequence are sequentially assigned to each die to obtain the various arrangement strategies.
[0012] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the operator running method as described above.
[0013] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the operator running method as described above.
[0014] The present invention also provides a computer program product, including a computer program that, when executed by a processor, implements the operator execution method as described above.
[0015] The operator running method, electronic device, storage medium, and program product provided by this invention combine the data volume, hidden layer dimension, and output dimension of the data to be processed by each expert in the hybrid expert layer where the operator is located, evaluate the load balancing term of various arrangement strategies of each expert on multiple dies, and select the target arrangement strategy from each arrangement strategy based on the load balancing term. Thus, the operator is run based on the target arrangement strategy, thereby achieving global maximization of the computational performance of the operator in a multi-die architecture.
[0016] Furthermore, the above method can be directly adapted to the current multi-die architecture's split storage and execution mode without modifying the hardware, reconstructing the underlying operators, or altering the gating network and expert network structure of the hybrid expert layer. It only requires adding a step of selecting the target arrangement strategy after the data to be processed is allocated and before the operator runs, which can effectively improve the utilization of computing resources for the operator running under the multi-die architecture, thereby effectively improving the operator's computing performance. Attached Figure Description
[0017] To more clearly illustrate the technical solutions in this invention or related technologies, the accompanying drawings used in the description of the embodiments or related technologies will be briefly introduced below. Obviously, the accompanying drawings described below are some embodiments of this invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0018] Figure 1 This is a schematic diagram of tensor distribution in the dual-die amortized computation mode of related technologies.
[0019] Figure 2 This is a schematic diagram of tensor distribution in the dual-die tensor storage mode of related technologies.
[0020] Figure 3 This is a flowchart illustrating the operator execution method provided by the present invention.
[0021] Figure 4 This is a schematic diagram of the target arrangement strategy provided by the present invention.
[0022] Figure 5 This is a schematic diagram of the operator running device provided by the present invention.
[0023] Figure 6 This is a schematic diagram of the structure of the electronic device provided by the present invention. Detailed Implementation
[0024] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.
[0025] All actions involving the acquisition of signal information or data in this invention are carried out in compliance with the relevant data protection laws and policies of the country where the device is located, and with the authorization granted by the owner of the device.
[0026] In the context of artificial intelligence computing and high-performance computing technologies, operators such as GroupGEMM are widely used in deep learning model training, inference, and scientific computing. Here, GroupGEMM is a computational pattern that packages multiple independent GEMM subtasks for parallel execution, and it is an important execution operator in the MoE model. As a core architecture for improving model capacity and computational efficiency in large language models, the MoE model can decompose the traditional FFN (Feedforward Neural Network) into multiple parallel expert networks, where each expert can execute its own GEMM operations, thus realizing the execution of GroupGEMM operations.
[0027] To meet the ever-increasing demand for computing power, multi-die chips, with their advanced chiplet packaging technology offering enhanced computing power, have gradually become one of the mainstream hardware architectures supporting GroupGEMM-type operators. Here, a multi-die chip can be understood as integrating multiple independent dies within a single AI chip, with each die possessing its own dedicated computing unit, memory controller, and on-chip cache. A typical multi-die chip can be a dual-die chip.
[0028] It's important to note that when performing GroupGEMM operations in a multi-die chip architecture, there's a crucial pre-processing step: zero-padding. Taking a dual-die chip architecture as an example, since the computational units on both dies need to maintain consistent computational behavior and adapt to the same computational interface, zero-padding must first be applied to the grouping matrices of different shapes, filling them to the same length. Only after this can the computational tasks be distributed across the two dies for parallel computation.
[0029] While the zero-padding process described above can satisfy the unified computation behavior of multi-die architecture, it will increase the memory usage space. In addition, the zero-padding and subsequent de-padding operations will also increase a lot of computational redundancy.
[0030] In related technologies, there exists an operator operation mode that features amortized computation across two dies and unified memory management. Specifically, in a dual-die architecture, read access operations on tensors and memory must ensure that the size read from both dies remains consistent. This mode employs a shared memory design, distributing computational tasks of different shapes evenly across the two dies in a fixed proportion, and aligning the tensor shapes on both dies by padding with zeros. For example, Figure 1 This is a schematic diagram of tensor distribution in the dual-die amortized computation mode of related technologies, such as... Figure 1As shown, tensor A and tensor C are assigned to bare disk 0, and tensor B and tensor D are assigned to bare disk 1. To align tensor A and tensor B, zero-padding is applied to tensor B, thus aligning tensor B and padding0 as a whole with tensor A; similarly, to align tensor C and tensor D, zero-padding is applied to tensor C, thus aligning tensor C and padding1 as a whole with tensor D.
[0031] While the computational power allocation logic is clear in this mode, in practical applications, the memory requirements for tensor operations of different shapes vary greatly. If very small tensors are grouped with very large tensors, zero-padding will introduce a large amount of invalid data. Furthermore, since unified memory management cannot dynamically allocate memory resources based on tensor shape, this leads to severe memory redundancy and is prone to memory overflow or low memory utilization. In addition, in the case of dual-die parallel processing, if there is a very small tensor and a very large tensor in the same group, subsequent calculations must wait until the calculation of the large tensor is completed. This waiting mechanism causes significant computational redundancy, increases computational latency, and leads to a significant decrease in overall computational efficiency.
[0032] Furthermore, some related technologies employ operator operation modes that directly store the same tensor on two separate dies. In this mode, computation can be performed using only the tensor on one die, while tensors stored on the other die occupy video memory and do not participate in the computation. For example, Figure 2 This is a schematic diagram of tensor distribution in the dual-die tensor storage mode of related technologies, such as... Figure 2 As shown, each tensor is stored on bare die 0 and bare die 1.
[0033] Although this mode does not require zero-padding, one of the dies not only does not participate in the calculation, but also continuously occupies video memory resources, completely deviating from the original intention of dual-die parallel computing and further exacerbating the waste of computing resources.
[0034] Furthermore, when performing GroupGEMM operations on multiple bare chips, the execution order of each expert on the bare chip is predetermined, and the GEMM of each expert is executed sequentially on each bare chip in parallel. Neither of these two operator operation modes is adapted to the amount of data that different experts need to process. Therefore, they cannot adaptively adjust the execution order of experts on the bare chip according to the dynamic changes in the amount of data that different experts need to process, further leading to the low computational efficiency of variable-shape operators like GroupGEMM in MoE scenarios.
[0035] To address the above problems, this invention provides an operator execution method. Figure 3 This is a flowchart illustrating the operator execution method provided by the present invention, as follows: Figure 3 As shown, the method includes: Step 310: Obtain the data to be processed of each expert in the hybrid expert layer where the operator is located, the hidden layer dimension and output dimension of the data to be processed, and the various arrangement strategies of each expert on multiple bare dies.
[0036] The operator here refers to the operator to be executed, specifically the operator running under the MoE architecture based on multiple bare chips. The operator can be an operator used to process multiple dynamically shaped matrices, or it can be understood as a group of operators that process multiple dynamically shaped matrices separately, such as GroupGEMM, or other operators that need to be routed to various experts in the MoE architecture for separate computation. This embodiment of the invention does not specifically limit this.
[0037] The hybrid expert layer where the operator resides refers to the MoE architecture in the model where the operator is located. The hybrid expert layer contains multiple parallel expert networks, referred to as multiple experts in this embodiment. Specifically, each expert in the hybrid expert layer can be one of the first preset number of activated experts with the highest weight. For each expert in the hybrid expert layer, there is corresponding data to be processed when running that operator. This can also be understood as the data routed from the operator's input data to each expert, or the input data corresponding to each expert when running the operator separately. Furthermore, the amount of data to be processed for each expert can be different. In some embodiments, the number of tokens can be used as the unit of measurement for data volume. For example, if there are four experts, namely Expert1, Expert2, Expert3, and Expert4, the amount of data to be processed for each expert can be expressed as the number of tokens received by each expert. For example, Expert1 receives 100 tokens, Expert2 receives 50 tokens, Expert3 receives 200 tokens, and Expert4 receives 45 tokens.
[0038] Hidden dimensions of the data to be processed This can also be understood as the data dimension of the data to be processed when it is input into the hybrid expert layer, or the input dimension of the data to be processed. The output dimension of the data to be processed. This refers to the dimension of the processed data output by the expert after running the operator to process the data to be processed. In some embodiments, the shape of the GEMM operator for each expert can be represented as follows: .
[0039] Furthermore, to optimize the performance of operators in a multi-die architecture, various arrangement strategies are pre-defined in this embodiment of the invention. Here, the arrangement strategy refers to the allocation method and order of experts in the hybrid expert layer across multiple dies. Specifically, it can include which experts are run on each die and the order in which the experts are run on each die.
[0040] For example, considering a scenario with four experts (Expert1, Expert2, Expert3, and Expert4) and two bare dies (Die0 and Die1), one deployment strategy could be to run Expert1 and Expert2 sequentially on Die0, and Expert3 and Expert4 sequentially on Die1. Another strategy could be to run Expert3 on Die0, and Expert1, Expert2, and Expert4 sequentially on Die1. This document does not exhaustively list all possible deployment strategies for the above scenario, nor does it limit the possible deployment strategies.
[0041] Understandably, the performance of operators may vary depending on the arrangement strategy used when running operators in a multi-die architecture.
[0042] Step 320: Based on the data volume, hidden layer dimension, and output dimension of the data to be processed by each expert, determine the load balancing item for various arrangement strategies.
[0043] Specifically, for each layout strategy, the load of each die executing the operator can be calculated based on the amount of data to be processed, the hidden layer dimension, and the output dimension of each expert's data. This allows for the measurement of the load balance of each die under the layout strategy, thus obtaining the load balancing term of the layout strategy.
[0044] Here, for any arrangement strategy, the load balancing term reflects the load balancing of each die under that arrangement strategy. It can be understood that the smaller the difference in load between different dies, and the more balanced the load among different dies, the better the utilization of computing resources in the multi-die architecture; conversely, the greater the difference in load between different dies, and the more unbalanced the load among different dies, the worse the utilization of computing resources in the multi-die architecture.
[0045] Step 330: Based on the load balancing items of the multiple arrangement strategies, determine the target arrangement strategy from the multiple arrangement strategies.
[0046] Specifically, after obtaining the load balancing item for each deployment strategy, a deployment strategy can be selected as the target deployment strategy based on the load balancing item of each deployment strategy. It can be understood that the target deployment strategy is the one among the various deployment strategies mentioned above that has the best load balancing among the individual dies; it can also be understood as the deployment strategy among the various deployment strategies mentioned above that has the best utilization of computing resources when running on a multi-die architecture.
[0047] Step 340: Based on the target layout strategy and the data to be processed by each expert, run the operator.
[0048] Specifically, after determining the target arrangement strategy, the experts in the hybrid expert layer can be arranged on multiple bare dies based on the target arrangement strategy. Thus, the experts arranged by each bare die are executed in parallel on each of the multiple bare dies. The data to be processed by each expert is processed through the operation of each expert, thereby realizing the overall operation of the operator.
[0049] In the method provided in this embodiment of the invention, the data volume, hidden layer dimension and output dimension of the data to be processed of each expert in the hybrid expert layer where the operator is located are combined to evaluate the load balancing term of various arrangement strategies of each expert on multiple dies, and select the target arrangement strategy from each arrangement strategy based on the load balancing term. Thus, the operator is run based on the target arrangement strategy, thereby realizing the global maximization of the computational performance of the operator in the multi-die architecture.
[0050] Furthermore, the above method can be directly adapted to the current multi-die architecture's split storage and execution mode without modifying the hardware, reconstructing the underlying operators, or altering the gating network and expert network structure of the hybrid expert layer. It only requires adding a step of selecting the target arrangement strategy after the data to be processed is allocated and before the operator runs, which can effectively improve the utilization of computing resources for the operator running under the multi-die architecture, thereby effectively improving the operator's computing performance.
[0051] Based on the above embodiments, step 320, which involves determining the load balancing items for various arrangement strategies based on the data volume, hidden layer dimension, and output dimension of the data to be processed by each expert, includes: Based on the amount of data to be processed, the hidden layer dimension, and the output dimension of the data to be processed by each expert, the computational load of each expert when running the operator is determined. Based on the computational load of each expert when running the operator, the total computational load of each die under each of the various arrangement strategies is determined. Based on the difference between the maximum and minimum total computational cost of each die under the various arrangement strategies, the load balancing term for each arrangement strategy is determined.
[0052] Specifically, after obtaining the data volume, hidden layer dimension, and output dimension of the data to be processed for each expert, the computational cost for each expert when running the operator can be calculated. Here, for each expert, the computational cost when running the operator can be expressed by the following formula: In the formula, For the first The computational cost when an expert runs the operator; For the first The amount of data to be processed by each expert, for example, could be the first... The number of tokens received by each expert; and These are the hidden dimension and the output dimension of the data to be processed, respectively.
[0053] Based on this, for various layout strategies, the total computational cost of each die under that strategy can be determined based on the computational cost of each expert running the operator. It can be understood that for any layout strategy, the strategy assigns a specific expert to each die. For each die, the computational cost of its own experts running the operator can be summed to obtain the die's total computational cost.
[0054] For example, in a dual-die architecture, Among the experts to An expert was assigned to the bare die Die0, which will to One expert is assigned to Die1. Therefore, the total computational cost of Die0 and Die1 can be expressed by the following formulas: In the formula, and These represent the total computational cost of the bare die Die0 and Die1, respectively.
[0055] Subsequently, the load balancing of the arrangement strategy can be measured by the difference between the maximum and minimum total computational loads of each die under that strategy. Under a given arrangement strategy, the total computational load of different dies reflects the load on each die when running the operator. The maximum total computational load represents the load on the die with the highest load, and the minimum total computational load represents the load on the die with the lowest load. The difference between the maximum and minimum total computational loads reflects the maximum load difference between dies under the arrangement strategy and can be used as a factor to measure whether the load is balanced among the dies. For example, a difference of zero indicates that the load on all dies is consistent, and the load is perfectly balanced among the dies under this arrangement strategy; the larger the difference between the maximum and minimum, the more unbalanced the load is among the dies under this arrangement strategy.
[0056] In some embodiments, the load balancing term can be calculated based on the following formula: In the formula, This refers to the load balancing factor, or the degree of load imbalance. For the number of nude films, express The total computational cost for each individual numeric disc.
[0057] In this embodiment of the invention, the load balancing term under various arrangement strategies is measured based on the total computational cost of each die under various arrangement strategies, thereby providing a basis for selecting a target arrangement strategy with load balancing and high computing resource utilization.
[0058] Based on any of the above embodiments, in step 330, determining the target arrangement strategy from the multiple arrangement strategies based on the load balancing item of the multiple arrangement strategies includes: Based on the data volume, hidden layer dimension, and output dimension of the data to be processed by each expert, the peak memory usage of each arrangement strategy is determined. Based on the alignment granularity corresponding to the multiple bare dies and the amount of data to be processed by each expert, the hardware alignment penalty term for each arrangement strategy is determined. Based on the degree of disruption of each arrangement strategy compared to the original arrangement strategy, the scheduling overhead of each arrangement strategy is determined. Based on at least one of the memory peak usage, hardware alignment penalty, and scheduling overhead of the various arrangement strategies, and the load balancing item of the various arrangement strategies, a target arrangement strategy is determined from the various arrangement strategies.
[0059] Specifically, in addition to referring to the load balancing items of each arrangement strategy to select the target arrangement strategy, at least one of the memory peak usage items, hardware alignment penalty items, and scheduling overhead items of each arrangement strategy can also be referred to.
[0060] For each layout strategy, the peak memory usage term reflects the memory usage of each die when the operator is run based on the layout strategy, especially the peak memory usage of each die within the total memory usage of all dies. Based on the data volume, hidden layer dimension, and output dimension of each expert's data to be processed, the memory usage of each expert during operator execution can be calculated. Then, under the given layout strategy, the maximum value of the allocated memory usage for each expert is calculated for each die, yielding the peak memory usage of each die. Combining this with the peak memory usage of all dies, the peak memory usage term of the layout strategy is measured. It is understood that a higher peak memory usage term indicates a higher proportion of peak memory usage within the total memory usage of all dies, resulting in more uneven memory usage throughout the operator's operation; conversely, a lower peak memory usage term indicates a lower proportion of peak memory usage within the total memory usage of all dies, more even memory usage throughout the operator's operation, and a lower likelihood of the operator exceeding its memory limit.
[0061] The hardware alignment penalty term reflects the cost incurred by each expert assigned to a die in zero-padding to align the data to be processed with the hardware-required alignment granularity when running an operator based on a layout strategy. The data to be processed for each expert is compared with the alignment granularity to obtain the amount of zero-padding required for each expert. The total amount of zero-padding required is then calculated to measure the hardware alignment penalty term. It is understandable that the higher the total amount of data requiring padding, the higher the hardware alignment penalty term, the more irrelevant data is introduced during operator execution, and the more easily operator performance is affected.
[0062] The scheduling overhead item reflects the degree to which the layout strategy is shuffled compared to the original layout strategy. The original layout strategy here refers to the layout strategy that can run directly without expert scheduling; it can be understood as the initial layout strategy, and is also one of several layout strategies. It is understandable that the higher the degree of shuffling of the layout strategy compared to the original layout strategy, the higher the scheduling overhead introduced when running operators based on that layout strategy; conversely, the lower the degree of shuffling and the closer the layout strategy is to the original layout strategy, the lower the scheduling overhead introduced when running operators based on that layout strategy. When selecting a target layout strategy from various layout strategies, for each layout strategy, at least one of the peak memory usage, hardware alignment penalty, and scheduling overhead items of the layout strategy, as well as the load balancing item of the layout strategy, can be weighted and summed. The result of the weighted sum is used as the total cost of the layout strategy. Thus, the layout strategy with the lowest total cost is selected from all layout strategies as the target layout strategy.
[0063] For example, the total cost of the layout strategy can be determined based on the following formula: In the formula, Indicates the arrangement strategy The total cost, , , , These are load balancing, peak memory usage, hardware alignment penalty, and scheduling overhead. , , , All are weights. .
[0064] In some embodiments, The value can be between 0.6 and 0.7. The value can be between 0.2 and 0.3. The value can be between 0.05 and 0.1. The value can be between 0.05 and 0.1.
[0065] In the method provided in this embodiment of the invention, four evaluation indicators—load balancing, peak memory usage, hardware alignment penalty, and scheduling overhead—are incorporated into the evaluation system for the deployment strategy. Furthermore, the weights can be dynamically adjusted based on hardware bottlenecks, model characteristics, and business scenarios reflected in the load and memory. Thus, without exceeding the memory limit or introducing excessive scheduling overhead, the global performance of the operator is maximized, taking into account the comprehensiveness, flexibility, and engineering practicality of the operator optimization effect.
[0066] Based on any of the above embodiments, in step 330, determining the peak memory usage of each arrangement strategy based on the data volume, hidden layer dimension, and output dimension of the data to be processed by each expert includes: Based on the amount of data to be processed, the hidden layer dimension, and the output dimension of each expert, the memory usage of each expert when running the operator is determined. Based on the memory usage of each expert when running the operator, the maximum memory usage of each die under each of the various layout strategies is determined. Based on the maximum memory usage of each die under the various layout strategies, and the total memory usage of all experts when running the operators, the peak memory usage of each layout strategy is determined.
[0067] Specifically, after obtaining the data volume, hidden layer dimension, and output dimension of each expert's data to be processed, the memory usage of each expert when running the operator can be calculated. Here, for each expert, the memory usage when running the operator can be expressed by the following formula: In the formula, For the first Memory usage when an expert runs an operator.
[0068] Based on this, for various layout strategies, the maximum memory usage of each die under that strategy can be determined based on the memory usage of each expert when running the operator. It can be understood that for any layout strategy, the strategy assigns a specific expert to each die, and for each die, the maximum memory usage can be selected from the memory usage of its own experts when running the operator.
[0069] Subsequently, the peak memory usage of this layout strategy can be measured based on the maximum memory usage of each die under this strategy and the sum of the memory usage of all experts when running the operator. For example, the peak memory usage can be obtained by dividing the sum of the maximum memory usage of all dies by the sum of the memory usage of all experts when running the operator. Therefore, the peak memory usage can be calculated using the following formula: In the formula, This refers to peak memory usage. For the number of nude films, for The first nude film Maximum memory usage per bare die; For the number of experts, for The first of the experts The amount of memory used by an expert when running an operator.
[0070] In this embodiment of the invention, the peak memory usage under the arrangement strategy is measured based on the maximum memory usage of each die and the total memory usage of all experts when running the operator, thereby providing a basis for selecting a target arrangement strategy with reasonable memory usage.
[0071] Based on any of the above embodiments, in step 330, determining the hardware alignment penalty term for the various layout strategies based on the alignment granularity corresponding to the multiple bare dies and the amount of data to be processed by each expert includes: Based on the alignment granularity corresponding to the multiple bare wafers and the amount of data to be processed by each expert, the alignment penalty value of each expert is determined; Based on the alignment penalty values of each expert and the computational cost of each expert when running the operator, the hardware alignment penalty terms for the various layout strategies are determined.
[0072] Specifically, the alignment granularity corresponding to multiple bare dies is the pre-defined tensor core alignment granularity under the multi-battery hardware architecture. This can be the smallest alignment unit required for the shape and size of the input matrix, such as 128 or 256. Understandably, if the dimensions of the input matrix are not aligned to this alignment granularity, the hardware will automatically pad the matrix with zeros.
[0073] For each expert, the alignment penalty value can be calculated based on the alignment granularity and the amount of data to be processed. The alignment penalty value here is the product of the amount of zero-padding data required to achieve alignment and the hidden layer dimension.
[0074] For example, suppose the hidden layer dimension With an alignment granularity of 128, and the data volumes to be processed for the four experts (Expert1, Expert2, Expert3, and Expert4) being 100, 50, 200, and 45 respectively, then the alignment penalty value for Expert1 is: The alignment penalty value for Expert2 is: The alignment penalty value for Expert3 is: The alignment penalty value for Expert4 is: .
[0075] Based on this, the alignment penalty values of all experts can be summed, and the computational cost of each expert running the operator can be summed. The total alignment penalty value can then be divided by the total computational cost to obtain the hardware alignment penalty term. Therefore, the hardware alignment penalty term can be calculated based on the following formula: In the formula, Hardware alignment penalty; For the number of experts, for The first of the experts Alignment penalty value for each expert, for The first of the experts The computational load for each expert when running the operator.
[0076] In this embodiment of the invention, the hardware alignment penalty term of the layout strategy is measured based on the alignment granularity and the amount of data to be processed by each expert, thereby providing a basis for selecting a target layout strategy that introduces less redundant and invalid data to achieve hardware alignment.
[0077] Based on any of the above embodiments, in step 330, determining the scheduling overhead of each arrangement strategy based on the degree of disruption of each arrangement strategy compared to the original arrangement strategy includes: Based on the degree of disruption of each arrangement strategy compared to the original arrangement strategy, and the number of experts included in the hybrid expert layer, the scheduling overhead of each arrangement strategy is determined.
[0078] Specifically, the degree of shuffling in the arrangement strategy compared to the original arrangement strategy can be reflected in the number of changes in the expert positions in the arrangement strategy compared to the original arrangement strategy; that is, the degree of shuffling can be represented by the number of inversions. Furthermore, the number of all possible expert rankings can be estimated based on the number of experts included in the hybrid expert layer.
[0079] Based on this, the degree of shuffling can be divided by the number of all possible expert rankings to obtain the scheduling cost. Therefore, the scheduling cost can be calculated using the following formula: In the formula, This is a scheduling overhead item; The arrangement of experts in the layout strategy. For arrangement Compared to the degree of disruption in the arrangement of experts in the original layout strategy; For the number of experts, It represents the maximum possible number of inversions.
[0080] In this embodiment of the invention, the scheduling overhead of the arrangement strategy is measured based on the degree of disruption of the arrangement strategy compared to the original arrangement strategy, thereby providing a basis for selecting the target arrangement strategy and avoiding the introduction of excessive scheduling overhead.
[0081] Based on any of the above embodiments, the operator operation method further includes: The experts are arranged in descending order of the amount of data to be processed to obtain the initial sequence; Based on the current load of each die, the experts in the initial sequence are sequentially assigned to each die to obtain the various arrangement strategies.
[0082] Specifically, before step 310 is executed, multiple arrangement strategies can be generated in advance. Here, all possible arrangements of experts can be enumerated to form all possible expert sequences, and all split points are traversed on each expert sequence. By splitting the expert sequence, experts are assigned to each die, thereby obtaining various arrangement strategies.
[0083] Furthermore, experts can be sorted from largest to smallest based on the amount of data they have to process, thus obtaining an initial sequence. It can be understood that the initial sequence is a sequence that sorts the experts, with the amount of data to be processed decreasing sequentially for each expert.
[0084] After obtaining the initial sequence, the current load of each die can be updated in real time. Based on the current load of each die, each expert in the initial sequence is sequentially assigned to a die with a lower current load, thereby achieving a rearrangement of the initial sequence. It can be understood that each time an expert in the initial sequence is assigned to a die, the current load of that die increases accordingly. Afterward, the current load of that die is updated, and then the next expert in the initial sequence is assigned to a die with a lower current load, and so on, until all experts in the initial sequence are assigned, thus obtaining the placement strategy.
[0085] In some embodiments, the die with the lowest current load may be the die with the lowest current load among all dies, or it may be one of the top dies sorted from lowest to highest current load among all dies. This embodiment of the invention does not specifically limit this.
[0086] In other embodiments, when the number of experts is less than or equal to a preset threshold, all possible expert sequences can be enumerated to form an arrangement strategy; when the number of experts is greater than the preset threshold, the experts can be arranged in descending order of the amount of data to be processed based on the idea of a greedy algorithm, and the experts with the largest amount of unallocated data can be assigned to the bare dies with the lowest current load, thereby obtaining an arrangement strategy.
[0087] In the method provided in this embodiment of the invention, experts are arranged in descending order of the amount of data to be processed, and the expert with the largest amount of unallocated data is assigned to the die with the lowest current load. This forms a better arrangement strategy for subsequent selection, effectively controlling the total number of arrangement strategies while ensuring the rationality of the arrangement strategy.
[0088] Based on any of the above embodiments, the operator execution method can specifically be a method for executing the GroupGEMM operator. This method includes the following steps: First, obtain the data to be processed, the hidden dimension and output dimension of the data to be processed for each expert in the hybrid expert layer where the operator is located, as well as the various arrangement strategies of each expert on multiple bare wafers.
[0089] For example, the hybrid expert layer in which the GroupGEMM operator is located includes four experts: Expert1, Expert2, Expert3, and Expert4. The amount of data to be processed by each expert can be represented by the number of tokens received by each expert. For example, Expert1 receives 100 tokens, Expert2 receives 50 tokens, Expert3 receives 200 tokens, and Expert4 receives 45 tokens.
[0090] In addition, the hidden dimension of the data to be processed Output dimension .
[0091] Furthermore, the GroupGEMM operator operates in a dual-die architecture, with the two dies being Die0 and Die1.
[0092] The original arrangement strategy of GroupGEMM operators in a dual-wafer architecture includes expert sequences. and the dividing point This means that Die0 executes Expert1 and Expert2 sequentially, and Die1 executes Expert3 and Expert4 sequentially.
[0093] Furthermore, all expert sequences can be enumerated and the split points traversed to obtain other arrangement strategies. For example, there is also an arrangement strategy that includes expert sequences. and the dividing point This means that Die0 executes Expert3, and Die1 executes Expert1, Expert2, and Expert4 in sequence.
[0094] Secondly, the parameters of each expert can be preprocessed separately.
[0095] Specifically, the computational cost, memory usage, and alignment penalty value for each expert can be calculated. Assuming an alignment granularity of 128, the calculation process can be represented in the following table: Next, for each arrangement strategy, the load balancing item, peak memory usage item, hardware alignment penalty item, and scheduling overhead item can be calculated separately, and then the total cost can be obtained by weighting them.
[0096] For example, the calculation process for the load balancing term is as follows for the original layout strategy: =5.87×10 9 +2.93×10 9 =8.8×10 9 , =1.17×10 10 +2.64×10 9 =1.43×10 10 =(1.43×10 10 -8.8×10 9 ) / ( 5.87×10 9 +2.93×10 9 +1.17×10 10 +2.64×10 9 )=0.238 The calculation process for peak memory usage is as follows: =1.84×10 6 , =3.69×10 6 =(1.84×10 6 +3.69×10 6 ) / (1.84×10 6 +9.22×10 5 +3.69×10 6 +8.30×10 5 )=0.760 The calculation process for the hardware alignment penalty is as follows: =(1.15×10 5 +3.19×10 5 +2.29×10 5 +3.40×10 5 ) / ( 5.87×10 9 +2.93×10 9 +1.17×10 10 +2.64×10 9 = 4.33 × 10 -5 The calculation process for scheduling overhead items is as follows: Inversion number = 0, =0 Therefore, the total cost of the original layout strategy is calculated as follows: =0.65×0.238+0.25×0.760+0.05×4.33×10 -5+0.05×0=0.345.
[0097] For the load balancing strategy where Die0 executes Expert3, and Die1 sequentially executes Expert1, Expert2, and Expert4, the calculation process for the load balancing items is as follows: =1.17×10 10 , =5.87×10 9 +2.93×10 9 +2.64×10 9 =1.14×10 10 =(1.17×10 10 -1.14×10 10 ) / (5.87×10 9 +2.93×10 9 +1.17×10 10 +2.64×10 9 )=0.013 The calculation process for peak memory usage is as follows: =3.69×10 6 , =1.84×10 6 =(3.69×10 6 +1.84×10 6 ) / (1.84×10 6 +9.22×10 5 +3.69×10 6 +8.30×10 5 )=0.760 The calculation process for the hardware alignment penalty is as follows: =(1.15×10 5 +3.19×10 5 +2.29×10 5 +3.40×10 5 ) / (5.87×10 9 +2.93×10 9 +1.17×10 10 +2.64×10 9 = 4.33 × 10 -5 The calculation process for scheduling overhead items is as follows: Inversion number = 3, =3 / (4×3 / 2)=0.5 Therefore, the total cost of this arrangement strategy is calculated to be: =0.65×0.013+0.25×0.760+0.05×4.33×10 -5 +0.05×0.5=0.223.
[0098] Finally, based on the total cost of each layout strategy, a target layout strategy is determined from the layout strategies, and operators are run on multiple bare dies based on the target layout strategy.
[0099] For example, the layout strategy with the minimum total cost is one that means Die0 executes Expert3 and Die1 executes Expert1, Expert2 and Expert4 in sequence. This layout strategy can be used as the target layout strategy.
[0100] Understandably, compared to the original layout strategy, the total cost of the target layout strategy is reduced by (0.345-0.223) / 0.345=35.4%, the computational load imbalance is reduced by (0.238-0.013) / 0.238=94.5%, and the dual-die resource utilization is improved by about 25%~30%.
[0101] Figure 4 This is a schematic diagram of the target arrangement strategy provided by the present invention, such as... Figure 4 As shown, to ensure that the computing units of the two bare dies maintain consistent computing behavior, Die0 and Die1 can sequentially read data of the size to be processed from Expert1 for GEMM calculation, read data of the size to be processed from Expert1 for GEMM calculation, and read all remaining data to be processed for GEMM calculation. Die1 only needs to pad with a small number of zeros to achieve alignment with Die0.
[0102] In this embodiment of the invention, the total cost is obtained by weighting the load balancing item, memory peak usage item, hardware alignment penalty item and scheduling overhead item of various arrangement strategies, and the target arrangement strategy is selected with the goal of minimizing the total cost, thereby achieving the global optimal utilization of multi-die resources.
[0103] The operator running apparatus provided by the present invention is described below. The operator running apparatus described below can be referred to in correspondence with the operator running method described above.
[0104] Figure 5 This is a schematic diagram of the operator operation device provided by the present invention, as shown below. Figure 5 As shown, the device includes: The input unit 510 is used to obtain the data to be processed of each expert in the hybrid expert layer where the operator is located, the hidden layer dimension and output dimension of the data to be processed, and the various arrangement strategies of each expert on multiple bare dies. The load balancing unit 520 is used to determine the load balancing items for various arrangement strategies based on the data volume, hidden layer dimension, and output dimension of the data to be processed by each expert. The strategy selection unit 530 is used to determine a target arrangement strategy from the multiple arrangement strategies based on the load balancing items of the multiple arrangement strategies. The operator execution unit 540 is used to run the operator based on the target layout strategy and the data to be processed by each expert.
[0105] In the apparatus provided in this embodiment of the invention, the load balancing term of each expert's data to be processed, hidden layer dimension, and output dimension in the hybrid expert layer where the operator is located are combined to evaluate the load balancing term of each expert's various arrangement strategies on multiple dies. Based on the load balancing term, a target arrangement strategy is selected from each arrangement strategy. Thus, the operator is run based on the target arrangement strategy, thereby achieving global maximization of the operator's computational performance in a multi-die architecture.
[0106] Furthermore, the aforementioned device can be directly adapted to the current multi-die architecture's Split storage and execution mode without modifying the hardware, reconstructing the underlying operators, or altering the gating network and expert network structure of the hybrid expert layer. It only requires adding a step of selecting the target arrangement strategy after the data to be processed is allocated and before the operator runs, which can effectively improve the utilization of computing resources for the operator running under the multi-die architecture, thereby effectively improving the operator's computing performance.
[0107] Based on any of the above embodiments, the load balancing unit is specifically used for: Based on the amount of data to be processed, the hidden layer dimension, and the output dimension of the data to be processed by each expert, the computational load of each expert when running the operator is determined. Based on the computational load of each expert when running the operator, the total computational load of each die under each of the various arrangement strategies is determined. Based on the difference between the maximum and minimum total computational cost of each die under the various arrangement strategies, the load balancing term for each arrangement strategy is determined.
[0108] Based on any of the above embodiments, the strategy selection unit is specifically used for: Based on the data volume, hidden layer dimension, and output dimension of the data to be processed by each expert, the peak memory usage of each arrangement strategy is determined. Based on the alignment granularity corresponding to the multiple bare dies and the amount of data to be processed by each expert, the hardware alignment penalty term for each arrangement strategy is determined. Based on the degree of disruption of each arrangement strategy compared to the original arrangement strategy, the scheduling overhead of each arrangement strategy is determined. Based on at least one of the memory peak usage, hardware alignment penalty, and scheduling overhead of the various arrangement strategies, and the load balancing item of the various arrangement strategies, a target arrangement strategy is determined from the various arrangement strategies.
[0109] Based on any of the above embodiments, the strategy selection unit is specifically used for: Based on the amount of data to be processed, the hidden layer dimension, and the output dimension of each expert, the memory usage of each expert when running the operator is determined. Based on the memory usage of each expert when running the operator, the maximum memory usage of each die under each of the various layout strategies is determined. Based on the maximum memory usage of each die under the various layout strategies, and the total memory usage of all experts when running the operators, the peak memory usage of each layout strategy is determined.
[0110] Based on any of the above embodiments, the strategy selection unit is specifically used for: Based on the alignment granularity corresponding to the multiple bare wafers and the amount of data to be processed by each expert, the alignment penalty value of each expert is determined; Based on the alignment penalty values of each expert and the computational cost of each expert when running the operator, the hardware alignment penalty terms for the various layout strategies are determined.
[0111] Based on any of the above embodiments, the strategy selection unit is specifically used for: Based on the degree of disruption of each arrangement strategy compared to the original arrangement strategy, and the number of experts included in the hybrid expert layer, the scheduling overhead of each arrangement strategy is determined.
[0112] Based on any of the above embodiments, the device further includes a strategy generation unit, used for: The experts are arranged in descending order of the amount of data to be processed to obtain the initial sequence; Based on the current load of each die, the experts in the initial sequence are sequentially assigned to each die to obtain the various arrangement strategies.
[0113] Figure 6 An example is a schematic diagram of the physical structure of an electronic device, such as... Figure 6As shown, the electronic device may include: a processor 610, a communications interface 620, a memory 630, and a communication bus 640, wherein the processor 610, the communications interface 620, and the memory 630 communicate with each other via the communication bus 640. The processor 610 can call logical instructions in the memory 630 to execute an operator execution method, which includes: Obtain the data to be processed of each expert in the hybrid expert layer where the operator is located, the hidden layer dimension and output dimension of the data to be processed, and the various arrangement strategies of each expert on multiple bare dies; Based on the data volume, hidden layer dimension, and output dimension of the data to be processed by each expert, the load balancing items for various arrangement strategies are determined. Based on the load balancing items of the various arrangement strategies, a target arrangement strategy is determined from the various arrangement strategies. The operator is run based on the target layout strategy and the data to be processed by each expert.
[0114] Furthermore, the logical instructions in the aforementioned memory 630 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to related technologies, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0115] On the other hand, the present invention also provides a computer program product, the computer program product comprising a computer program, the computer program being able to be stored on a non-transitory computer-readable storage medium, and when the computer program is executed by a processor, the computer being able to execute the operator execution method provided by the above methods, the method comprising: Obtain the data to be processed of each expert in the hybrid expert layer where the operator is located, the hidden layer dimension and output dimension of the data to be processed, and the various arrangement strategies of each expert on multiple bare dies; Based on the data volume, hidden layer dimension, and output dimension of the data to be processed by each expert, the load balancing items for various arrangement strategies are determined. Based on the load balancing items of the various arrangement strategies, a target arrangement strategy is determined from the various arrangement strategies. The operator is run based on the target layout strategy and the data to be processed by each expert.
[0116] In another aspect, the present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, is implemented to perform the operator execution methods provided by the above methods, the method comprising: Obtain the data to be processed of each expert in the hybrid expert layer where the operator is located, the hidden layer dimension and output dimension of the data to be processed, and the various arrangement strategies of each expert on multiple bare dies; Based on the data volume, hidden layer dimension, and output dimension of the data to be processed by each expert, the load balancing items for various arrangement strategies are determined. Based on the load balancing items of the various arrangement strategies, a target arrangement strategy is determined from the various arrangement strategies. The operator is run based on the target layout strategy and the data to be processed by each expert.
[0117] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.
[0118] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the parts that contribute to the related technology, can be embodied in the form of software products. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.
[0119] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for operating an operator, characterized in that, include: Obtain the data to be processed of each expert in the hybrid expert layer where the operator is located, the hidden layer dimension and output dimension of the data to be processed, and the various arrangement strategies of each expert on multiple bare dies; Based on the data volume, hidden layer dimension, and output dimension of the data to be processed by each expert, the load balancing items for various arrangement strategies are determined. Based on the load balancing items of the various arrangement strategies, a target arrangement strategy is determined from the various arrangement strategies. The operator is run based on the target layout strategy and the data to be processed by each expert.
2. The operator operation method according to claim 1, characterized in that, The process of determining load balancing items for various deployment strategies based on the data volume, hidden layer dimension, and output dimension of the data to be processed by each expert includes: Based on the amount of data to be processed, the hidden layer dimension, and the output dimension of the data to be processed by each expert, the computational load of each expert when running the operator is determined. Based on the computational load of each expert when running the operator, the total computational load of each die under each of the various arrangement strategies is determined. Based on the difference between the maximum and minimum total computational cost of each die under the various arrangement strategies, the load balancing term for each arrangement strategy is determined.
3. The operator operation method according to claim 1, characterized in that, The load balancing item based on the multiple arrangement strategies determines the target arrangement strategy from the multiple arrangement strategies, including: Based on the data volume, hidden layer dimension, and output dimension of the data to be processed by each expert, the peak memory usage of each arrangement strategy is determined. Based on the alignment granularity corresponding to the multiple bare dies and the amount of data to be processed by each expert, the hardware alignment penalty term for each arrangement strategy is determined. Based on the degree of disruption of each arrangement strategy compared to the original arrangement strategy, the scheduling overhead of each arrangement strategy is determined. Based on at least one of the memory peak usage, hardware alignment penalty, and scheduling overhead of the various arrangement strategies, and the load balancing item of the various arrangement strategies, a target arrangement strategy is determined from the various arrangement strategies.
4. The operator operation method according to claim 3, characterized in that, The peak memory usage of each arrangement strategy is determined based on the data volume, hidden layer dimension, and output dimension of the data to be processed by each expert, including: Based on the amount of data to be processed, the hidden layer dimension, and the output dimension of each expert, the memory usage of each expert when running the operator is determined. Based on the memory usage of each expert when running the operator, the maximum memory usage of each die under each of the various layout strategies is determined. Based on the maximum memory usage of each die under the various layout strategies, and the total memory usage of all experts when running the operators, the peak memory usage of each layout strategy is determined.
5. The operator operation method according to claim 3, characterized in that, The hardware alignment penalty term for each arrangement strategy is determined based on the alignment granularity corresponding to the multiple bare dies and the amount of data to be processed by each expert, including: Based on the alignment granularity corresponding to the multiple bare wafers and the amount of data to be processed by each expert, the alignment penalty value of each expert is determined; Based on the alignment penalty values of each expert and the computational cost of each expert when running the operator, the hardware alignment penalty terms for the various layout strategies are determined.
6. The operator operation method according to claim 3, characterized in that, The determination of scheduling overhead items for each arrangement strategy based on the degree of disruption compared to the original arrangement strategy includes: Based on the degree of disruption of each arrangement strategy compared to the original arrangement strategy, and the number of experts included in the hybrid expert layer, the scheduling overhead of each arrangement strategy is determined.
7. The operator operation method according to any one of claims 1 to 6, characterized in that, Also includes: The experts are arranged in descending order of the amount of data to be processed to obtain the initial sequence; Based on the current load of each die, the experts in the initial sequence are sequentially assigned to each die to obtain the various arrangement strategies.
8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that, When the processor executes the computer program, it implements the operator running method as described in any one of claims 1 to 7.
9. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the operator running method as described in any one of claims 1 to 7.
10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the operator running method as described in any one of claims 1 to 7.