Method, computing device, medium and program product for performing computations on neural network model structures

By adjusting the calculation task sequence and the allocation of kernel functions of the neural network model structure, the full utilization of the calculation core of the artificial intelligence chip is achieved, the problem of resource waste in traditional solutions is solved, and the computing efficiency and resource utilization are improved.

CN120373382AActive Publication Date: 2025-07-25SHANGHAI BIREN TECH CO LTD

Patent Information

Application Number
CN202510873024.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-26
Publication Date
2025-07-25
Estimated Expiration
2045-06-26

AI Technical Summary

Technical Problem

Traditional neural network model structure computing solutions are difficult to make full use of the computing cores of artificial intelligence chips, especially when executing Router TopK and Shared Expert modules, resulting in waste of computing resources.

Method used

By adjusting the execution order of the calculation tasks, the fusion kernel functions can perform the calculation tasks of the first and second network modules without data dependence between each other in parallel, and select part of the calculation kernel to perform these tasks separately according to the calculation kernel identification, so as to realize partition allocation of the calculation kernel.

Benefits of technology

Make full use of the computing cores of artificial intelligence chips to improve computing efficiency and resource utilization, and ensure the integrity of computing logic and accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120373382A_ABST
    Figure CN120373382A_ABST
Patent Text Reader

Abstract

The invention relates to a method, a computing device, a medium and a program product for performing computations on a neural network model structure. The method comprises the following steps: adjusting an execution sequence of calculation tasks related to a neural network model structure so as to enable a constructed fusion kernel function to be used for executing calculation tasks related to a first network module and calculation tasks related to a second network module in parallel, the first network module and the second network module are included in the neural network model structure and do not have a data dependency relationship; and according to the calculation core identifier, selecting a part of calculation cores in the corresponding artificial intelligence chip so as to respectively execute a calculation task about the first network module and a calculation task about the second network module. According to the invention, the calculation core of the artificial intelligence chip can be fully utilized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present invention generally relate to the field of artificial intelligence, and more particularly to a method, a computing device, a computer-readable storage medium, and a computer program product for performing computations on a neural network model structure. Background Art

[0002] Traditional solutions for performing computations on a neural network model structure typically require sequential processing of multiple network modules included in the neural network model. The neural network model is, for example, of a Mixture of Experts (MoE) structure for a Large Language Model (LLM). It should be understood that the Feed-Forward Network (FFN) in the MoE network structure is one of the core computational modules of the model. The FFN module has a special interconnection relationship, and the computational time it requires also occupies a very large proportion of the overall computational time of the MoE network structure. Specifically, when performing computations, it is necessary to first compute the Router TopK module, then compute the Activated Routed Expert module, and additionally, it is also necessary to compute the Shared Expert module, resulting in a relatively long overall computational time.

[0003] In traditional solutions for performing computations on a neural network model structure, different kernel functions are typically used to compute the Router TopK, Activated Routed Expert, and Shared Expert modules respectively. However, when computing the Router TopK module, since the computational task volume of the Router TopK module is small, not all computing cores of the artificial intelligence chip can be fully utilized. Additionally, when computing the Shared Expert module, because the number of Shared Experts is small, the computing cores of the artificial intelligence chip cannot be fully utilized either, thus resulting in unfriendliness in task partitioning and execution on the artificial intelligence chip.

[0004] In summary, the deficiencies of traditional solutions for performing computations on a neural network model structure are as follows: it is difficult to fully utilize the computing cores of the artificial intelligence chip when performing computational tasks on the neural network model. Summary of the Invention

[0005] The present invention provides a method, a computing device, a computer-readable storage medium, and a computer program product for performing computations on a neural network model structure, which can fully utilize the computing cores of the artificial intelligence chip.

[0006] According to a first aspect of the present invention, there is provided a method for performing calculations on a neural network model structure. The method includes: adjusting an execution order of calculation tasks on the neural network model structure so that a constructed fusion kernel function is used to perform calculation tasks on a first network module and calculation tasks on a second network module in parallel, where the first network module and the second network module are included in the neural network model structure and there is no data dependency between them; and selecting a part of calculation cores in a corresponding artificial intelligence chip according to a calculation core identifier to respectively perform the calculation tasks on the first network module and the calculation tasks on the second network module.

[0007] In some embodiments, the method for performing calculations on a neural network model structure further includes: calculating a ratio between the amount of calculation tasks on the first network module and the amount of calculation tasks on the second network module so as to determine the number of calculation cores required for the calculation tasks on the first network module based on the ratio; and constructing a fusion kernel function, where the fusion kernel function is configured to: at least based on a current calculation core identifier, select between the calculation tasks on the first network module and the calculation tasks on the second network module.

[0008] In some embodiments, the fusion kernel function being configured to: at least based on a current calculation core identifier, select between the calculation tasks on the first network module and the calculation tasks on the second network module includes: in response to determining that the current calculation core identifier is less than a task splitting identifier, selecting the calculation tasks on the first network module; and in response to determining that the current calculation core identifier is greater than or equal to the task splitting identifier, selecting the calculation tasks on the second network module.

[0009] In some embodiments, determining the number of calculation cores required for the calculation tasks on the first network module based on the ratio includes: based on the ratio, determining the number of calculation cores required for the calculation tasks on the first network module, thereby determining a task splitting identifier, where the task splitting identifier indicates a position division status when the calculation tasks are executed on the calculation cores of the corresponding artificial intelligence chip.

[0010] In some embodiments, constructing a fused kernel function includes: modifying a first original kernel function for calculating a first network module into a first device function, and making the number of calculation cores of the first device function equal to the split task identifier, where the first original kernel function is a global function; and modifying a second original kernel function for calculating a second network module into a second device function, and making the number of calculation cores of the second device function equal to the difference between the artificial intelligence chip calculation core identifier and the split task identifier, and making the task identifier of the second device function equal to the difference between the artificial intelligence chip calculation core identifier and the split task identifier, where the second original kernel function is a global function.

[0011] In some embodiments, the neural network model structure is a mixture-of-experts structure of a large language model, the first network module is a pre-routing pre-positioned expert module, and the second network module is a shared expert module.

[0012] In some embodiments, constructing a fused kernel function includes: configuring the following parameters for the fused kernel function: output hidden data, activation expert list, input hidden data, and split task identifier.

[0013] In some embodiments, selecting a calculation task regarding the first network module includes: selecting to call the first device function; and selecting a calculation task regarding the second network module includes: selecting to call the second device function.

[0014] According to a second aspect of the present invention, there is also provided a computing device. The computing device includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the computing device can execute the method of the first aspect of the present invention.

[0015] According to a third aspect of the present invention, there is also provided a computer-readable storage medium. A computer program is stored on the computer-readable storage medium, and when the computer program is executed by a machine, it executes the method of the first aspect of the present invention.

[0016] According to a fourth aspect of the present invention, there is also provided a computer program product, including a computer program, and when the computer program is executed by a machine, it executes the method of the first aspect of the present invention.

[0017] The present invention enables the fused kernel function to allocate the calculation cores of the artificial intelligence chip to the calculation tasks of the first network module and the second network module, so that the calculation cores are partitioned for different calculation tasks regarding different network models, thereby enabling the calculation cores of the artificial intelligence chip to be fully utilized. Therefore, the present invention can fully utilize the calculation cores of the artificial intelligence chip.

[0018] It should be understood that the content described in this section is not intended to identify the key or important features of the embodiments of the present invention, nor is it used to limit the scope of the present invention. Other features of the present invention will become readily understood from the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] In conjunction with the accompanying drawings and with reference to the following detailed description, the above and other features, advantages, and aspects of the embodiments of the present invention will become more apparent. In the drawings, the same or similar reference numerals denote the same or similar elements.

[0020] Figure 1 A schematic diagram schematically showing a typical neural network model structure is shown.

[0021] Figure 2 A schematic diagram schematically showing a conventional method for performing calculations on a neural network model structure is shown.

[0022] Figure 3 A schematic diagram schematically showing a computing device for implementing a method for performing calculations on a neural network model structure according to an embodiment of the present invention is shown.

[0023] Figure 4 A flowchart showing a method for performing calculations on a neural network model structure according to an embodiment of the present invention is shown.

[0024] Figure 5 A schematic diagram showing a fusion kernel function and a graphics processing unit according to an embodiment of the present invention is shown.

[0025] Figure 6 A flowchart showing a method for constructing a fusion kernel function according to an embodiment of the present invention is shown.

[0026] Figure 7 A schematic diagram schematically showing code for constructing a fusion kernel function according to an embodiment of the present invention is shown.

[0027] In the respective drawings, the same or corresponding reference numerals denote the same or corresponding parts. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0028] Preferred embodiments of the present invention will be described in more detail below with reference to the accompanying drawings. Although the preferred embodiments of the present invention are shown in the drawings, it should be understood that the present invention can be implemented in various forms and should not be limited by the embodiments set forth herein. On the contrary, these embodiments are provided so that the present invention will be more thorough and complete, and will fully convey the scope of the present invention to those skilled in the art.

[0029] As used herein, the term "comprising" and its variations denote open inclusion, i.e., "including but not limited to". The term "based on" means "at least partially based on". The terms "an example embodiment" and "an embodiment" mean "at least one example embodiment". The term "another embodiment" means "at least one additional embodiment". The terms "first", "second", etc. may refer to different or the same objects.

[0030] As described above, Figure 1 A schematic diagram schematically shows a typical neural network model structure 100. This neural network model structure 100 is, for example, a feed-forward network (FFN) model structure in a mix of experts (MoE) model structure in a large language model (LLM). Specifically, when performing a computational task regarding the FFN model structure in MoE, multiple network modules need to be processed sequentially. For example, it is necessary to first calculate the pre-routing positioning (i.e., Router TopK) module 120, which selects 8 experts from 256 experts according to the input 110. Then, calculations are performed on the selected 8 activated experts via the Activated Routed Expert module. Additionally, it is also necessary to calculate the Shared Expert module. The calculation results of the shared expert network module and the calculation results of the activated routed expert module are superimposed to generate the output 130.

[0031] Figure 2 A schematic diagram schematically shows a conventional method 200 for performing calculations regarding a neural network model structure. Research shows that in a conventional method for performing calculations regarding a neural network model structure, different kernel functions are generally used to sequentially calculate the Router TopK, Activated Routed Expert, and Shared Expert modules. For example, as Figure 2 shown, first, calculations regarding the pre-routing positioning expert module are performed using kernel function 1; then, calculations regarding the routed activation expert module are performed using kernel function 2; and then, calculations regarding the shared expert module are performed using kernel function 3. Table 1 below schematically shows the number of shared experts, the number of activated experts, and the total number of routed experts for several typical neural network model structures (e.g., DeepSeek V3 / R1, QWen3, and Llama4).

[0032] Table 1

[0033] When performing calculations on the Router TopK module, the Router TopK module is responsible for selecting the top K experts with the highest weights from all experts. The computational task of this step is relatively simple, and its computational characteristic is memory-bound. Therefore, the computing cores of the GPGPU are not fully utilized, resulting in a relatively low utilization rate of the GPGPU. When calculating the Activated Routed Expert module, multiple Routed Experts are activated simultaneously. Therefore, all the computing cores of the GPGPU can be fully utilized, resulting in a relatively high utilization rate of the GPGPU. When calculating the Shared Expert module, usually only a small number of Shared Experts are activated, and the computing cores of the GPGPU are not fully utilized. Therefore, the overall utilization rate of the GPGPU is relatively low.

[0034] In summary, the deficiencies of the traditional solution for performing calculations on the neural network model structure are as follows: it is difficult to fully utilize the computing cores of the artificial intelligence chip when performing computational tasks on the neural network model.

[0035] To at least partially address the above and one or more other potential problems, exemplary embodiments of the present invention propose a solution for performing calculations on the neural network model structure. In this solution, by adjusting the execution order of the computational tasks of the neural network model structure, so that the fusion kernel function parallelly executes the computational tasks of the first network module and the second network module that have no data dependencies between each other; and according to the computing core identifier, selecting some computing cores to separately execute the computational tasks of the first network module and the second network module, the present invention enables the fusion kernel function to allocate the computing cores of the artificial intelligence chip to the computational tasks of the first network module and the second network module, so that the computing cores are partitioned for different computational tasks of different network models, and further enables the computing cores of the artificial intelligence chip to be fully utilized. Therefore, the present invention can fully utilize the computing cores of the artificial intelligence chip. It should be understood that the artificial intelligence chip is, for example but not limited to: Graphics Processing Unit (GPU), General purpose computing on graphics processing units (GPGPU), or Tensor Processing Unit (TPU), etc.

[0036] Figure 3 Schematically shows a schematic diagram of a computing device 300 for implementing a method for performing calculations on a neural network model structure according to an embodiment of the present invention. As Figure 3As shown, the computing device 300 may have one or more processing units, including dedicated processing units such as a Graphics Processing Unit (GPU), a Field Programmable Gate Array (FPGA), an Application Specific Integrated Circuit (ASIC), or a General-purpose computing on graphics processing units (GPGPU), as well as a general-purpose processing unit such as a CPU. The computing device 300 further includes at least: a neural network model structure computing task execution order adjustment unit 302, and a computing core selection unit 304. It should be understood that the neural network model structure computing task execution order adjustment unit 302 and the computing core selection unit 304 may be software modules that run, for example, on one or more processing units configured in the computing device 300.

[0037] Regarding the neural network model structure computing task execution order adjustment unit 302, it is used to adjust the execution order of the computing tasks regarding the neural network model structure, so that the constructed fusion kernel function is used to execute the computing tasks regarding the first network module and the computing tasks regarding the second network module in parallel. The first network module and the second network module are included in the neural network model structure and there is no data dependency between them.

[0038] Regarding the computing core selection unit 304, it is used to select some computing cores in the corresponding artificial intelligence chip according to the computing core identifier to respectively execute the computing tasks regarding the first network module and the computing tasks regarding the second network module.

[0039] The following will be combined with Figure 4 and Figure 5 to describe the method 400 for performing the computation regarding the neural network model structure in the embodiments of the present invention. Figure 4 FIG. shows a flowchart of the method 400 for performing the computation regarding the neural network model structure according to an embodiment of the present invention. Figure 5 FIG. shows a schematic diagram of a fusion kernel function and a graphics processor according to an embodiment of the present invention. It should be understood that the method 400 may be executed, for example, at Figure 3 the computing device 300 described above. The method 400 may further include additional actions not shown and / or may omit the actions shown, and the scope of the present invention is not limited in this regard.

[0040] At step 402, the computing device 300 adjusts the execution order of the computing tasks regarding the neural network model structure so that the constructed fusion kernel function is used to execute the computing tasks regarding the first network module and the computing tasks regarding the second network module in parallel. The first network module and the second network module are included in the neural network model structure and there is no data dependency between them.

[0041] Regarding the neural network model structure, it includes, for example, at least: a first network module and a second network module. There is no data dependency between the first network module and the second network module. It should be understood that the neural network model structure may also include: other network modules in addition to the first network module and the second network module, for example, a third network module. In some embodiments, the neural network model structure is a mixture-of-experts structure of a large language model, the first network module is a pre-routing pre-positioning expert module, and the second network module is a shared expert module.

[0042] For example, on the computation graph, there is no sequential data dependency between the computing tasks of the pre-routing pre-positioning expert module and the computing tasks of the shared expert module. Therefore, the computing device 300 can adjust the execution order of the computing tasks (for example, by adjusting the computation graph) so that the constructed fusion kernel function executes the computing tasks regarding the pre-routing pre-positioning expert module and the computing tasks regarding the shared expert module in parallel. Thus, the present invention can improve the utilization rate of the artificial intelligence chip while ensuring the original computing logic and accuracy.

[0043] Such as Figure 5As shown, the fused kernel function 510 is formed by fusing two kernel functions for performing the computing tasks of the first network module and the second network module. Specifically, the fused kernel function 510 is configured to include: a first subtask execution unit 514 and a second subtask execution unit 516. Among them, the first subtask execution unit is used to perform the computing tasks regarding the first network module. The second subtask execution unit 516 is used to perform the computing tasks regarding the second network module. In some embodiments, the computing device 300 adjusts the execution order of the computing tasks by adjusting the computing graph, so that the first original kernel function for computing the Router TopK module and the third original kernel function for computing the Shared Expert module are fused into the fused kernel function 510, such that the first subtask execution unit 514 of the fused kernel function 510 is used to perform the computing tasks regarding the Router TopK module; and the second subtask execution unit 516 is used to perform the computing tasks regarding the Shared Expert module. After the fused kernel function 510, the second kernel function 520 performs the computing tasks regarding the Routed Expert module.

[0044] In some embodiments, the fused kernel function is constructed by the following method: modifying the first original kernel function for computing the first network module into a first device function, and making the number of computing cores of the first device function equal to the split task identifier, the first original kernel function being a global function; and modifying the second original kernel function for computing the second network module into a second device function, and making the number of computing cores of the second device function the difference between the computing core identifier of the artificial intelligence chip and the split task identifier, and making the task identifier of the second device function equal to the difference between the computing core identifier of the artificial intelligence chip and the split task identifier, the second original kernel function being a global function.

[0045] At step 404, the computing device 300 selects a part of the computing cores in the corresponding artificial intelligence chip according to the computing core identifier for respectively performing the computing tasks regarding the first network module and the computing tasks regarding the second network module.

[0046] For example, taking an artificial intelligence chip as a graphics processing unit, through a subtask selection unit 512 configured at the entrance of the fusion kernel function 510, at least based on the computing core identifier, a first part of computing cores 532 is selected in the graphics processing unit 530 corresponding to the fusion kernel function 510 to execute the computing task for the first network module (for example, the computing task for the Router TopK module executed by the first subtask execution unit 514); and a second part of computing cores 534 is selected in the corresponding graphics processing unit 530 to execute the computing task for the second network module (that is, the computing task for the Shared Expert module executed by the second subtask execution unit 516).

[0047] Table 2 below exemplarily shows the code for the computing core allocation of the traditional computing task for the first network module (for example, Router TopK). The first original kernel function for executing the computing for the first network module (for example, Router TopK) is a global function, for example, indicated by "_global__ void router(ActivatedExpertList ae, InputHiden ih)" in Table 2. In addition, through the code "computeCoreNum = gpu_compute_core_num()", the number of computing cores is made equal to the computing core identifier of the artificial intelligence chip (for example, the graphics processing unit). Taking the graphics processing unit as an example. For example, the number of computing cores is made equal to 16. Furthermore, through the code "taskID = gpu_compute_core_id()", the task identifier is made equal to the computing core identifier of the image processor. After that, the current computing core task is obtained, and on 16 computing cores, the computing task for the first network module (for example, Router TopK) is executed starting from the corresponding task identifier. Thus, it can be seen that although the computing task volume of the first network module (for example, Router TopK) is small, the computing task of the first network module (for example, Router TopK) is assigned to 16 computing cores of the image processor, so that the computing task volume assigned to each computing core is even smaller, and further the computing power of the image processor is not fully utilized, resulting in a low utilization rate and computing efficiency of the image processor.

[0048] Table 2

[0049] Table 3 below schematically shows the code for the calculation core allocation of the computational tasks of the traditional second network module (e.g., Shared Expert). The second original core function for performing the computational tasks of the second network module (e.g., Shared Expert) is a global function, as indicated by, for example, "_global__ void sharedExpert(OutputHidden oh, InputHiden ih)" in Table 2. In addition, through the code "computCoreNum = gpu_compute_core_num()", the number of calculation cores is made equal to the calculation core identifier of the image processor. For example, the number of calculation cores is made equal to 16. Through the code "taskID = gpu_compute_core_id()", the task identifier is made equal to the calculation core identifier of the image processor. Then, the current calculation core task is obtained, and on 16 calculation cores, the computational tasks regarding the second network module (e.g., Shared Expert) are executed starting from the corresponding task identifier.

[0050] Table 3

[0051] As can be seen from the above code, in the traditional method for performing the calculations regarding the neural network model structure, if the computational tasks regarding the first network module or the computational tasks regarding the second network module are run separately on the corresponding artificial intelligence chip (e.g., graphics processor), the amount of computational tasks executed on each of the 16 calculation cores of the artificial intelligence chip is small, making it difficult to fully utilize the 16 calculation cores of the artificial intelligence chip.

[0052] Table 4 below exemplarily shows the code for the calculation core allocation of the computational tasks of the first network module (e.g., RouterTopK) according to an embodiment of the present invention. As shown in Table 4, the fused core function for calculating the computational tasks of the first network module (e.g., Router TopK) is configured as a first device function, as indicated by, for example, "__device__void router(ActivatedExpertList ae, InputHiden ih, int taskSplitID)" in Table 4. Regarding the method of configuring it as a first device function, it includes, for example: modifying the first original core function for calculating the first network module (as shown in Table 2) from a global function to a first device function. It should be understood that a global function cannot be called by other core functions. Modifying the first original core function for calculating the first network module from a global function to a first device function with the attribute of a device function can enable the call of the first device function.

[0053] As shown in Table 4, through the code "computeCoreNum = taskSplitID", the number of computing cores is made equal to the corresponding task split identifier. Here, taskSplitID represents the corresponding task split identifier, which is the input parameter of the first device function. For example, when taskSplitID = 8, the number of computing cores is equal to 8. In addition, through the code "taskID = gpu_compute_core_id()" in Table 4, the task identifier is made equal to the computing core identifier of the graphics processor. Then, the current computing core task is obtained, and starting from the corresponding task identifier of 8 computing cores, the computing task of the pre-routing pre-positioning expert module is executed. As Figure 5 shown, for computing core 5, gpu_compute_core_id() = 5, and its taskID = 5, so computing core 5 is used to execute the computing task with taskID = 5 for the first network module.

[0054] Table 4

[0055] Table 5 below schematically shows the code for computing core allocation regarding the second network module (e.g., SharedExpert) according to an embodiment of the present invention. As shown in Table 5, the fusion kernel function used for computing the second network module (e.g., SharedExpert) is configured as a second device function, for example, as indicated by "__device__ void router(ActivatedExpertList ae, InputHiden ih, int taskSplitID)" in Table 5. Regarding the method of configuring it as a second device function, for example, it includes modifying the second original kernel function (as shown in Table 3) used for computing the second network module (e.g., Shared Expert) from a global function to a second device function.

[0056] In addition, as shown in Table 5, the relationships among the number of compute cores (computeCoreNum) for configuring the second device function, the compute core identifier of the corresponding graphics processing unit (gpu_compute_core_num()), the task split identifier (taskSplitID), and the task identifier (taskID) are configured. Specifically, for example, through the code "computeCoreNum = gpu_compute_core_num() – taskSplitID", the number of compute cores is made equal to the number of compute cores of the graphics processing unit minus the corresponding task split identifier; and through the code "taskID = gpu_compute_core_id() - taskSplitID", the task identifier is made equal to the compute core identifier of the corresponding graphics processing unit minus the corresponding task split identifier. For example, when taskSplitID = 8, as Figure 5 shown, for compute core 12, gpu_compute_core_id() = 12, and its taskID = 12 – 8 = 4, so compute core 12 is used to execute the computational task with taskID 4 for the second network module.

[0057] With the above means of the present invention, two different computational tasks for different network modules are assigned to different partitions of the artificial intelligence chip (for example, graphics processing unit) (for example, compute cores 0 to 8 are used to execute the computational tasks for the first network module, and compute cores 9 to 15 are used to execute the computational tasks for the second network module), so that the computations for the first network module and the second network module can make full use of the compute cores of the processor, and the spatial partitioning of different compute cores can execute different computational tasks in parallel in time, thus significantly improving the computational efficiency.

[0058] Table 5

[0059] In the above solution, by adjusting the execution order of the computational tasks regarding the neural network model structure, so that the fusion kernel function parallelly executes the computational tasks of the first network module and the second network module that have no data dependency relationships with each other; and according to the compute core identifier, selecting some compute cores to respectively execute the computational tasks of the first network module and the second network module, the present invention enables the fusion kernel function to allocate the compute cores of the artificial intelligence chip to the computational tasks of the first network module and the second network module, making the compute cores partitioned for different computational tasks of different network models, and further enabling the compute cores of the artificial intelligence chip to be fully utilized. Therefore, the present invention can make full use of the compute cores of the artificial intelligence chip.

[0060] In some embodiments, method 400 further includes method 600 for constructing a fused kernel function.

[0061] The following will be combined with Figure 6 and Figure 7 to describe method 600 for constructing a fused kernel function according to an embodiment of the present invention. Figure 6 FIG. shows a flowchart of method 600 for constructing a fused kernel function according to an embodiment of the present invention. Figure 7 FIG. schematically shows a schematic diagram of code for constructing a fused kernel function according to an embodiment of the present invention. It should be understood that method 600 can be executed, for example, at Figure 3 the computing device 300 described. Method 600 may further include additional actions not shown and / or may omit the actions shown, and the scope of the present invention is not limited in this regard.

[0062] At step 602, the computing device 300 calculates the ratio between the amount of computing tasks of the first network module and the amount of computing tasks of the second network module, so as to determine the number of computing cores required for the computing tasks of the first network module based on the ratio.

[0063] For example, the computing device 300 calculates the number of computing cores required for the computing tasks of the network structure of the first network module. For example, if the amount of computing tasks of the first network module is different, the number of computing cores required may also be different. For example, the number of computing cores of a graphics calculator is 16. If the ratio between the amount of computing tasks of the Router TopK module and the Shared Expert module for the DeepSeek model is 1:1, then the number of computing cores required by the computing device 300 for the computing tasks of the first network module is 8. For example, the computing tasks of the Router TopK module for QWen3 are less, and the ratio between the amount of computing tasks of its Router TopK module and the Shared Expert module is 1:4. Then the number of computing cores required for the Router TopK module of QWen3 is 4.

[0064] In some embodiments, determining the number of computing cores required for the computing tasks of the first network module based on the ratio includes: determining the number of computing cores required for the computing tasks of the first network module based on the ratio, and determining a task splitting identifier, where the task splitting identifier indicates the location division status of task execution. For example, it indicates how many computing cores are used to execute the computing tasks of the first network module and how many computing cores are used to execute the computing tasks of the second network module. The task splitting identifier is represented as "taskSplitID" for example.

[0065] At step 604, the computing device 300 constructs a fused kernel function.

[0066] Table 6 below exemplarily shows the code for constructing the fused kernel function according to an embodiment of the present invention. As shown in Table 6, the fused kernel function is defined, for example, as indicated by "__global__ void fuseWork(OutputHidden oh,ActivatedExpertList ae,InputHiden ih,int taskSplitID)". This fused kernel function is configured with parameters: output hidden data "OutputHidden oh", activated expert list "ActivatedExpertList ae", input hidden data "InputHiden ih", and task split identifier "int taskSplitID". For example, through the code "coreID = gpu_compute_core_id()", the current compute core identifier is made equal to the graphics processor compute core identifier.

[0067] It should be understood that the fused kernel function is configured to select between a computing task for a first network module and a computing task for a second network module based at least on the current compute core identifier. Specifically, the method of selecting between a computing task for a first network module and a computing task for a second network module, for example, includes steps 606 to 610.

[0068] At step 606, the computing device 300 determines whether the current compute core identifier is less than the task split identifier.

[0069] At step 608, if the computing device 300 determines that the current compute core identifier is less than the task split identifier, it selects the computing task for the first network module.

[0070] As Figure 7 shown, if the computing device 300 determines that the current compute core identifier is less than the task split identifier, as indicated by arrow 710, it selects to call the first device function 712.

[0071] For example, through the code "if(coreID < taskSplitID) { router(ae, ih,taskSplitID)" in Table 6, when the current computing core identifier (coreID) is less than the task splitting identifier (taskSplitID), the calculation of the first device function of the computing task for the routing pre-positioning expert (Router TopK) module is performed, as indicated by, for example, "router(ae, ih, taskSplitID)" in Table 6. Another example is that when the previous computing core identifier (coreID) is "0", which is less than the task splitting identifier (taskSplitID, for example, 8), the first device function for the computing task of the Router TopK module is selected.

[0072] At step 610, if the computing device 300 determines that the current computing core identifier is greater than or equal to the task splitting identifier, the computing task of the second network module is selected.

[0073] As Figure 7 shown, if the computing device 300 determines that the current computing core identifier is greater than or equal to the task splitting identifier, as indicated by arrow 720, the second device function 722 is selected for invocation.

[0074] For example, through the code "} else sharedExpert(oh, ih, taskSplitID);}" in Table 6, when the current computing core identifier (coreID) is greater than or equal to the task splitting identifier (taskSplitID), the calculation of the second device function of the computing task for the shared expert (Shared Expert) module is performed, as indicated by, for example, "sharedExpert(oh, ih, taskSplitID)" in Table 6. For example, when the current computing core identifier (coreID) is "9", which is greater than the task splitting identifier (taskSplitID, for example, 8), the calculation of the second device function for the computing task of the Shared Expert module is selected.

[0075] Table 6

[0076] In the above solution, the present invention can reuse the codes of the existing first and second original kernel functions without affecting any network structure and completely maintaining the accuracy of the network.

[0077] The various processes and treatments described above, such as methods 400 and 600, can be executed at a computing device. The computing device includes, for example: at least one processor (at least one graphics processor and at least one central processor); and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor. In some embodiments, methods 400 and 600 can be implemented as a computer software program or program product, which is tangibly contained in a machine-readable medium. In some embodiments, part or all of the computer program can be loaded and / or installed onto the computing device via a read-only memory (ROM) and / or a communication unit. When the computer program is loaded into a random-access memory (RAM) and executed by a GPU and a CPU, one or more actions of methods 400 and 600 described above can be executed.

[0078] The present invention can be a method, an apparatus, a system, and / or a computer program product. The computer program product can include a computer-readable storage medium having thereon computer-readable program instructions for performing various aspects of the present invention. The computer-readable storage medium can be a tangible device that can hold and store instructions used by an instruction execution device. The computer-readable storage medium can be, for example, but not limited to, an electrical storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing.

[0079] The computer-readable program instructions described herein can be downloaded from the computer-readable storage medium to various computing / processing devices, or downloaded to an external computer or external storage device via a network, such as the Internet, a local area network, a wide area network, and / or a wireless network. Aspects of the present invention are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the present invention. It should be understood that each block of the flowchart illustrations and / or block diagrams, and the combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.

[0080] These computer-readable program instructions can be provided to a central processing unit of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine such that the instructions, when executed by the central processing unit of the computer or other programmable data processing apparatus, create means for implementing the functions / acts specified in one or more boxes of the flowchart and / or block diagram. These computer-readable program instructions may also be stored in a computer-readable storage medium that causes a computer, a programmable data processing apparatus, and / or other devices to function in a particular manner, such that the computer-readable medium storing the instructions comprises a manufacture including instructions for implementing various aspects of the functions / acts specified in one or more boxes of the flowchart and / or block diagram.

[0081] The flowcharts and block diagrams in the figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each box in the flowchart or block diagram may represent a module, a segment of a program, or a part of an instruction, which contains one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions noted in the boxes may occur in a different order than noted in the figures. For example, two consecutive boxes may actually be executed substantially in parallel, or they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each box in the block diagrams and / or flowcharts, and combinations of boxes in the block diagrams and / or flowcharts, can be implemented by a dedicated hardware-based system that performs the specified functions or acts, or by a combination of dedicated hardware and computer instructions.

[0082] It should be understood that various forms of the flows shown above can be used, steps reordered, added, or deleted. For example, the steps recited in this application can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this application can be achieved, and no limitations are imposed herein.

[0083] The above specific embodiments do not constitute a limitation on the scope of protection of this application. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors.

Claims

1. A method for performing calculations regarding a neural network model structure, characterized in that, including: Adjust the execution order of the computational tasks regarding the neural network model structure so that the constructed fusion kernel function is used to execute the computational tasks regarding the first network module and the computational tasks regarding the second network module in parallel, where the first network module and the second network module are included in the neural network model structure and there is no data dependency between them; and According to the computational core identifier, select some computational cores in the corresponding artificial intelligence chip to be used to execute the computational tasks regarding the first network module and the computational tasks regarding the second network module respectively.

2. The method according to claim 1, wherein It also includes: Calculate the ratio between the amount of computational tasks of the first network module and the amount of computational tasks of the second network module so as to determine the number of computational cores required for the computational tasks regarding the first network module based on the ratio; and Construct a fusion kernel function, where the fusion kernel function is configured to: select at least based on the current computational core identifier between the computational tasks regarding the first network module and the computational tasks regarding the second network module.

3. The method according to claim 2, wherein The fusion kernel function being configured to: select at least based on the current computational core identifier between the computational tasks regarding the first network module and the computational tasks regarding the second network module includes: In response to determining that the current computational core identifier is less than the task splitting identifier, select the computational tasks regarding the first network module; and In response to determining that the current computational core identifier is greater than or equal to the task splitting identifier, select the computational tasks regarding the second network module.

4. The method according to claim 3, wherein Determining the number of computational cores required for the computational tasks regarding the first network module based on the ratio includes: Based on the ratio, determine the number of computational cores required for the computational tasks regarding the first network module, thereby determining the task splitting identifier, where the task splitting identifier indicates the position division status when the computational tasks are executed on the computational cores of the corresponding artificial intelligence chip.

5. The method according to claim 2, wherein Constructing the fusion kernel function includes: Modify the first original kernel function for calculating the first network module into a first device function, and make the number of computational cores of the first device function equal to the task splitting identifier, where the first original kernel function is a global function; and Modify the second original kernel function for calculating the second network module into a second device function, and make the number of computational cores of the second device function equal to the difference between the computational core identifier of the artificial intelligence chip and the task splitting identifier, and make the task identifier of the second device function equal to the difference between the computational core identifier of the artificial intelligence chip and the task splitting identifier, where the second original kernel function is a global function.

6. The method according to claim 1, wherein The neural network model structure is a mixture-of-experts structure of a large language model, the first network module is a pre-routing pre-positioning expert module, and the second network module is a shared expert module.

7. The method according to claim 3, characterized in that, Constructing the fusion kernel function includes: Configure the following parameters for the fusion kernel function: output hidden data, activation expert list, input hidden data, and task splitting identifier.

8. The method according to claim 5, characterized in that Selecting the computational tasks regarding the first network module includes: selecting to call the first device function; and Selecting the computational tasks regarding the second network module includes: selecting to call the second device function.

9. A computing device, characterized in that, including: at least one processor; and a memory communicatively connected to the at least one processor; where The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method according to any one of claims 1-8.

10. A computer-readable storage medium, characterized in that, A computer program is stored on the computer-readable storage medium, and when the computer program is executed by a machine, it executes the method according to any one of claims 1-8.

11. A computer program product, characterized in that, It includes a computer program, and when the computer program is executed by a machine, it executes the method according to any one of claims 1-8.

Citation Information

Patent Citations

  • Hierarchical parallelism in a network of distributed neural network cores

    CN112384935A

  • Text data reasoning method and device based on hybrid expert model

    CN119443279A

Cited By

  • Method for issuing tasks about kernel functions, artificial intelligence chip, computing device, medium and program product

    CN121210146A