Acceleration method for executing operation task by expert hybrid model and related equipment

By storing expert network parameters in the accelerator chip and optimizing resource utilization, the problem of excessive memory usage of expert hybrid models is solved, and efficient computing task acceleration and optimized hardware resources are achieved.

CN120297430AActive Publication Date: 2025-07-11北京汤谷软件技术有限公司

Patent Information

Application Number
CN202510779045.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-12
Publication Date
2025-07-11
Estimated Expiration
2045-06-12

AI Technical Summary

Technical Problem

As the number of expert networks in the expert hybrid model increases, the accelerator chip memory space is used up too much, affecting operational efficiency, and becoming a bottleneck restricting the acceleration of model deployment and computing tasks.

Method used

Deploy expert hybrid model in the accelerator chip, divide expert network parameters into on-chip memory and off-chip memory storage, select target expert networks through parallel calculations, and update the weight matrix in real time using distributed lookup tables and boundary scanning test interfaces, and optimize resource utilization in combination with time division multiplexing algorithms.

Benefits of technology

It improves the efficiency of computing tasks acceleration of expert hybrid models, reduces data access latency, improves hardware resource utilization and system performance, and supports efficient inference and training of large-scale complex models.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120297430A_ABST
    Figure CN120297430A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of artificial intelligence chip acceleration, and provides an acceleration method for an expert hybrid model to execute an operation task and related equipment. The method comprises the steps that a to-be-processed operation task is acquired, a target expert network used for executing the to-be-processed operation task in an expert hybrid model is determined, and network parameters of all first expert networks in the expert hybrid model are stored in an on-chip memory of an accelerator chip, network parameters of each second expert network in the expert hybrid model are stored in an off-chip memory of the accelerator chip; calling a first network parameter of a first expert network in the target expert network from the on-chip memory, and calling a second network parameter of a second expert network in the target expert network from the off-chip memory; and executing the to-be-processed operation task through a target expert network in the expert hybrid model based on the first network parameter and the second network parameter. Through the technical scheme provided by the invention, the acceleration efficiency of executing the operation task by the expert hybrid model can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application belongs to the field of artificial intelligence chip acceleration, and in particular, relates to an acceleration method and related equipment for executing computing tasks using an expert hybrid model. Background Art

[0002] With the rapid development of technologies such as artificial intelligence and deep learning, the scale and complexity of neural network models continue to increase, and the computing resource consumption in the model reasoning and training process is also increasing. In order to meet the needs of high-performance computing, various types of dedicated accelerator chips are widely used in the reasoning and training tasks of deep neural network models. At the same time, the Mixture of Experts (MoE) model, as an innovative model structure that can improve model capacity and expression capabilities and reduce computing resource consumption, has gradually become the focus of attention in the industry and academia. The Mixture of Experts model introduces multiple sub-networks (i.e., expert networks) and dynamically selects some expert networks to participate in reasoning according to the characteristics of the input task, thereby realizing on-demand allocation and efficient utilization of computing resources.

[0003] However, as the number of expert networks increases, the total number of network parameters in the expert mixture model also expands dramatically, resulting in an increasing amount of memory space occupied in the accelerator chip, which restricts the operating efficiency of the accelerator chip and becomes a key bottleneck restricting the actual deployment of the expert mixture model and accelerating the expert mixture model to perform computing tasks. Based on this, how to improve the acceleration efficiency of the expert mixture model to perform computing tasks is a technical problem that needs to be solved urgently. Summary of the invention

[0004] The embodiments of the present application provide a method, device, computer program product, computer-readable storage medium and electronic device for accelerating an expert mixture model to perform computing tasks, thereby improving the acceleration efficiency of the expert mixture model to perform computing tasks to a certain extent.

[0005] Other features and advantages of the present application will become apparent from the following detailed description, or may be learned in part by the practice of the present application.

[0006] According to the first aspect of the embodiments of the present application, an acceleration method for an expert mixture model to execute an arithmetic task is provided. The expert mixture model is deployed in an accelerator chip, and the method includes: obtaining an arithmetic task to be processed, and determining a target expert network in the expert mixture model for executing the arithmetic task to be processed. The network parameters of each first expert network in the expert mixture model are stored in the on-chip memory of the accelerator chip, and the network parameters of each second expert network in the expert mixture model are stored in the off-chip memory of the accelerator chip; calling the first network parameters of the first expert network in the target expert network from the on-chip memory, and calling the second network parameters of the second expert network in the target expert network from the off-chip memory; based on the first network parameters and the second network parameters, executing the arithmetic task to be processed through the target expert network in the expert mixture model.

[0007] In some embodiments of the present application, based on the foregoing solution, the determining a target expert network in the expert mixture model for executing the arithmetic task to be processed includes: calculating the similarity between the arithmetic task to be processed and each expert network in the expert mixture model in parallel in the accelerator chip; selecting the set number of expert networks with the highest similarity as the target expert network for executing the arithmetic task to be processed.

[0008] In some embodiments of the present application, based on the foregoing solution, the calculating the similarity between the arithmetic task to be processed and each expert network in the expert mixture model in parallel in the accelerator chip includes: obtaining the feature matrix of the arithmetic task to be processed and the expert weight matrix of each expert network in the expert mixture model; based on the feature matrix and the expert weight matrix, calculating the similarity between the arithmetic task to be processed and each expert network in the expert mixture model in parallel in the accelerator chip.

[0009] In some embodiments of the present application, based on the foregoing solution, the calculating the similarity between the arithmetic task to be processed and each expert network in the expert mixture model in parallel in the accelerator chip includes: calculating the similarity between the arithmetic task to be processed and each expert network in the expert mixture model in parallel in the accelerator chip according to the Hamming distance algorithm or the cosine similarity algorithm or the gating network.

[0010] In some embodiments of the present application, based on the foregoing solution, the expert weight matrix of each expert network in the expert mixture model is pre-stored in the distributed lookup table of the accelerator chip.

[0011] In some embodiments of the present application, based on the foregoing solution, the method further includes: updating the expert weight matrix stored in the distributed lookup table in real time based on the boundary scan test interface.

[0012] In some embodiments of the present application, based on the foregoing solution, the method further includes: in response to an update instruction for network parameters in the on-chip memory, counting the frequencies at which each expert network in the expert mixture model has performed arithmetic tasks historically; defining a set number of expert networks with the highest frequencies as the first expert networks, and defining the expert networks in the expert mixture model other than the first expert networks as the second expert networks; saving the network parameters of the first expert networks in the on-chip memory of the accelerator chip, and saving the network parameters of the second expert networks in the off-chip memory of the accelerator chip.

[0013] In some embodiments of the present application, based on the foregoing solution, the method further includes: before calling the first network parameters of the first expert networks in the target expert network from the on-chip memory and calling the second network parameters of the second expert networks in the target expert network from the off-chip memory, triggering an update instruction for network parameters in the on-chip memory.

[0014] In some embodiments of the present application, based on the foregoing solution, the method further includes: abstracting common operators in the accelerator chip into configurable modules; and running each target expert network through the configurable modules based on a time-division multiplexing algorithm.

[0015] In some embodiments of the present application, based on the foregoing solution, the accelerator chip includes any one of a field-programmable gate array chip, a graphics processing chip, a Google Tensor Processing Unit chip, and an application-specific integrated circuit.

[0016] According to a second aspect of the embodiments of the present application, there is provided an acceleration device for an expert mixture model to perform arithmetic tasks. The expert mixture model is deployed in an accelerator chip. The device includes: an acquisition unit, configured to acquire an arithmetic task to be processed and determine a target expert network in the expert mixture model for performing the arithmetic task to be processed. The network parameters of each first expert network in the expert mixture model are saved in the on-chip memory of the accelerator chip, and the network parameters of each second expert network in the expert mixture model are saved in the off-chip memory of the accelerator chip; a call unit, configured to call the first network parameters of the first expert networks in the target expert network from the on-chip memory and call the second network parameters of the second expert networks in the target expert network from the off-chip memory; and an execution unit, configured to perform the arithmetic task to be processed through the target expert network in the expert mixture model based on the first network parameters and the second network parameters.

[0017] According to a third aspect of the embodiments of the present application, there is provided a computer program product, which includes computer instructions stored in a computer-readable storage medium and adapted to be read and executed by a processor, so that a computer device having the processor executes to implement the operations performed by the method described in the first aspect above.

[0018] According to a fourth aspect of the embodiments of the present application, there is provided a computer-readable storage medium, in which at least one computer program instruction is stored, and the at least one computer program instruction is loaded and executed by a processor to implement the operations performed by the method described in the first aspect above.

[0019] According to a fifth aspect of the embodiments of the present application, there is provided an electronic device, which includes one or more processors and one or more memories, and at least one computer program instruction is stored in the one or more memories, and the at least one computer program instruction is loaded and executed by the one or more processors to implement the operations performed by the method described in the first aspect above.

[0020] Based on the technical solution proposed in the present application, by deploying a mixture-of-experts model in an accelerator chip and storing the parameters of different expert networks in on-chip memory and off-chip memory respectively, the acceleration efficiency of the model in performing computing tasks can be effectively improved. Specifically, the parameters of the first expert network that are commonly used or frequently calculated can be stored in on-chip memory, and the characteristics of high-speed and low-latency access of on-chip memory can be fully utilized to achieve fast invocation of key network parameters, significantly reducing data access latency and improving the overall inference and training speed. And storing the parameters of the second expert network in off-chip memory can break through the capacity limit of on-chip memory, support the flexible deployment of more and larger-scale expert networks, and provide guarantee for the scalability and diversity of the model. During the execution of specific computing tasks, the network parameters in on-chip memory and off-chip memory can be flexibly invoked according to actual needs to give full play to the advantages of the two-level storage structure. On the one hand, the high bandwidth and low power consumption characteristics of on-chip memory ensure the efficient processing of high-frequency access to network parameters, reducing data transfer and access bottlenecks; on the other hand, the capacity advantage of off-chip memory supports the complex structure and large-scale parameter storage requirements of the mixture-of-experts model. Through this hierarchical storage and efficient scheduling mechanism, the efficient acceleration of the computing tasks of the mixture-of-experts model can be finally achieved, improving the utilization rate of hardware resources and the overall performance of the system, and meeting the efficient inference and training requirements of large-scale and complex mixture-of-experts models in practical applications.

[0021] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present application. Description of the Drawings

[0022] The accompanying drawings herein are incorporated into and constitute a part of this specification, showing embodiments consistent with the present application, and are used together with the specification to explain the principles of the present application. Obviously, the accompanying drawings in the following description are only some embodiments of the present application, and those of ordinary skill in the art can obtain other drawings based on these drawings without creative efforts. In the drawings: Figure 1 shows a flowchart of an acceleration method for an expert mixture model to execute an operation task in an embodiment of the present application; Figure 2 shows a hardware architecture diagram for executing the acceleration method in an embodiment of the present application; Figure 3 shows a block diagram of an acceleration device for an expert mixture model to execute an operation task in an embodiment of the present application; Figure 4 shows a schematic structural diagram of an electronic device in an embodiment of the present application. Detailed implementation manners

[0023] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part rather than all of the embodiments of the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts fall within the scope of protection of the present application.

[0024] In addition, the described features, structures or characteristics can be combined in any suitable manner in one or more embodiments. In the following description, many specific details are provided to give a full understanding of the embodiments of the present application. However, those skilled in the art will realize that the technical solutions of the present application can be practiced without one or more of the specific details, or other methods, components, devices, steps, etc. can be adopted. In other cases, well-known methods, devices, implementations or operations are not shown or described in detail to avoid obscuring aspects of the present application.

[0025] The block diagrams shown in the drawings are only functional entities and do not necessarily correspond to physically independent entities. That is, these functional entities can be implemented in software form, or in one or more hardware modules or integrated circuits, or in different networks and / or processor devices and / or microcontroller devices. It should also be noted that in the drawings, for the sake of simplicity of the drawings, some components that do not affect the explanation of the technical solutions of the present application are adaptively omitted.

[0026] The flowcharts shown in the accompanying drawings are only illustrative and not necessarily include all contents and operations / steps, nor are they necessarily executed in the order described. For example, some operations / steps can be decomposed, while some operations / steps can be combined or partially combined. Therefore, the actual execution order may change according to the actual situation.

[0027] In the description of this application, it should be understood that the terms "first" and "second" are only used for descriptive purposes and cannot be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include one or more of such features. In the description of this application, unless otherwise specified, the meaning of "a plurality" is two or more.

[0028] To enable those skilled in the art to better understand this application, the technical concepts and application backgrounds involved in this application will be briefly described first.

[0029] Mixture of Experts (MoE): The mixture of experts model is an ensemble learning method. Its core idea is to combine multiple "expert models" together and dynamically select or weight the outputs of different expert models according to the characteristics of the input data, thereby improving the expressive power and generalization performance of the overall model. That is, the mixture of experts model introduces multiple sub-networks (i.e., expert networks) and dynamically selects some expert networks to participate in the inference according to the characteristics of the input task, so as to achieve the on-demand allocation and efficient utilization of computing resources. This dynamic selection mechanism enables the model to flexibly adjust the allocation of computing resources when processing different types of data, thereby improving the overall performance. The mixture of experts model is widely used in fields such as natural language processing and computer vision, especially in large-scale models (such as GPT-4, SwitchTransformer, etc.) to improve parameter utilization and inference efficiency.

[0030] Accelerator Chip: An accelerator chip refers to an integrated circuit (chip) designed specifically for specific computing tasks or application scenarios, aiming to be superior to general-purpose processors (such as CPUs) in terms of energy consumption, speed, or efficiency. Accelerator chips are usually optimized at the hardware level for a certain type of computing to achieve higher parallelism, lower latency, or better energy efficiency ratio.

[0031] In this application, the accelerator chip may include any one of a Field Programmable Gate Array (FPGA), a Graphics Processing Unit (GPU), a Google Tensor Processing Unit (TPU), and an Application-Specific Integrated Circuit (ASIC).

[0032] With the rapid development of technologies such as artificial intelligence and deep learning, the scale and complexity of neural network models have been continuously increasing, and the computational resource consumption during model inference and training has also been growing day by day. To meet the requirements of high-performance computing, various types of dedicated accelerator chips have been widely used in the inference and training tasks of deep neural network models. These accelerator chips can significantly improve the execution efficiency and response speed of the models through parallel computing and optimized hardware architectures.

[0033] At the same time, the mixture-of-experts model, as an innovative model structure that can enhance model capacity and expressive power while reducing computational resource consumption, has gradually become the focus of attention in the industry and academia.

[0034] However, with the increase in the number of expert networks, the total number of parameters in the mixture-of-experts model has also increased sharply, which leads to an increasing occupation of the memory space in the accelerator chip. This increase in memory occupation not only affects the operating efficiency of the accelerator chip but also becomes a key bottleneck restricting the actual deployment of the mixture-of-experts model and accelerating the execution of arithmetic tasks of the mixture-of-experts model. Especially in an environment with limited resources, how to effectively manage and optimize these parameters has become an important challenge.

[0035] Currently, researchers are exploring various technical means to address this challenge, including but not limited to model compression, parameter sharing, and introducing more efficient storage management strategies, etc. The purpose of these methods is to improve the feasibility and efficiency of the mixture-of-experts model in practical applications. In this context, this application also proposes an acceleration scheme for the execution of arithmetic tasks of the mixture-of-experts model to improve the acceleration efficiency of the execution of arithmetic tasks of the mixture-of-experts model.

[0036] The implementation details of the technical solution of the embodiments of this application are elaborated below: Referring to Figure 1 , a flowchart of the acceleration method for the execution of arithmetic tasks of the mixture-of-experts model in the embodiments of this application is shown. The acceleration method for the execution of arithmetic tasks of the mixture-of-experts model can be executed by a device with computing and processing capabilities, where the mixture-of-experts model can be deployed in an accelerator chip. Referring to Figure 1As shown, the method for accelerating the operation task execution of the expert hybrid model at least includes steps 110 to 130, which are introduced in detail as follows: Referring to Figure 1 , in step 110, a to-be-processed operation task is obtained, and a target expert network for executing the to-be-processed operation task in the expert hybrid model is determined. The network parameters of each first expert network in the expert hybrid model are stored in the on-chip memory of the accelerator chip, and the network parameters of each second expert network in the expert hybrid model are stored in the off-chip memory of the accelerator chip.

[0037] In this application, obtaining the to-be-processed operation task may be receiving specific task data that needs to be inferred or trained from an external input or an upper-layer application. For example, in a natural language processing scenario, the task may be a piece of text to be analyzed. In an image recognition scenario, it may be a picture to be classified. By performing feature extraction on the to-be-processed operation task, a feature vector or a feature matrix that can be used for subsequent model processing is formed.

[0038] Specifically, in one embodiment, please refer to Figure 2 , which shows the hardware architecture diagram for executing the acceleration method in the embodiment of this application.

[0039] As Figure 2 shown, the feature vector or the feature matrix is input into the attention network model loaded on the accelerator chip in the form of an operation task. The attention network model performs feature extraction on the feature vector or the feature matrix to obtain a new feature vector or a new feature matrix. The new feature vector or the new feature matrix is input into the expert hybrid model loaded on the accelerator chip in the form of an operation task, and the operation result is obtained after being processed by the expert hybrid model.

[0040] In this application, the expert hybrid model contains a large number of expert networks, and each expert network has an independent set of network parameters, resulting in a large overall parameter scale of the model. For different input operation tasks, it is necessary to dynamically select the most suitable expert network and schedule the corresponding network parameters to participate in the operation, that is, dynamically select several expert networks that are most suitable for the current operation task as the "target expert network".

[0041] In this application, in order to optimize the memory usage of the accelerator chip and improve the inference efficiency, the expert networks in the expert mixture model are divided into two categories: the first expert network and the second expert network. Among them, the network parameters of the first expert network can be stored in the on-chip memory of the accelerator chip. For example, the Static Random Access Memory (SRAM) in the accelerator chip, and also the Block RAM (BRAM). The network parameters of the second expert network are stored in the off-chip memory outside the accelerator chip. For example, the Double Data Rate (DDR) synchronous dynamic random access memory, and also the High Bandwidth Memory (HBM), and also the Dynamic Random Access Memory (DRAM), and also the Static Random Access Memory (SRAM), and also the Flash memory.

[0042] In this application, using HBM to replace DDR storage can achieve a 5-fold increase in bandwidth, but the cost increases by 30%. To balance cost and performance, a hybrid cache hierarchy can be adopted, combining BRAM (storing high-frequency experts) with HBM (storing low-frequency experts).

[0043] In this application, to determine the target expert network in the expert mixture model for performing the to-be-processed operation task, the following steps 111 to 112 can be executed: Step 111, parallelly calculate the similarity between the to-be-processed operation task and each expert network in the expert mixture model in the accelerator chip.

[0044] Step 112, select the set number of expert networks with the highest similarity as the target expert network for performing the to-be-processed operation task.

[0045] In this application, the similarity between the to-be-processed operation task and each expert network in the expert mixture model is an indicator to measure the "fitting degree" between the current to-be-processed operation task and each expert network, and can be calculated based on the characteristics of the input operation task and the capabilities or historical performances of the expert networks.

[0046] Specifically, in this application, to parallelly calculate the similarity between the to-be-processed operation task and each expert network in the expert mixture model in the accelerator chip, the following steps 1111 to 1112 can be calculated: Step 1111: Obtain the feature matrix of the operation task to be processed and the expert weight matrices of each expert network in the expert mixture model.

[0047] Step 1112: Based on the feature matrix and the expert weight matrices, parallelly calculate the similarities between the operation task to be processed and each expert network in the expert mixture model in the accelerator chip.

[0048] In this application, the feature matrix is a vector or matrix representation formed after feature extraction of the current operation task to be processed (such as input samples, requests, data blocks, etc.). It can be a real-valued vector obtained by embedding, encoding, etc. of the original data of the operation task to be processed, or an intermediate representation obtained through a preprocessing network (such as convolution, transformer, MLP, etc.).

[0049] In this application, the expert weight matrix is a set of parameters used by each expert network to characterize its "expertise direction" or "capability space", usually a feature vector or parameter vector of each expert. It can be part of the weights of the expert network (such as the weight vector of the last layer), or an "expert description vector" designed specifically for similarity calculation.

[0050] In this application, the expert weight matrices of each expert network in the expert mixture model can be pre-stored in the distributed look-up table (LUT) of the accelerator chip. The accelerator chip (such as FPGA) contains a large number of programmable LUT resources, which can not only implement logic functions but also serve as efficient on-chip storage units. Through a distributed method, each LUT can store the expert weight matrix of an expert network separately. For example, if there are 128 expert networks, the 128 expert weight matrices can be stored in 128 independent LUTs respectively to achieve physical isolation and parallel access. When calculating the similarities between the operation task to be processed and each expert network in the expert mixture model, the weight matrices of the 128 expert networks can be read from their respective LUTs simultaneously and independently without serial waiting, greatly improving the calculation efficiency. It can be seen that using the distributed look-up table of the accelerator chip to save the expert weight matrices can greatly reduce the memory access latency, support high-frequency and low-latency expert network selection requests, realize expert network selection decisions at the microsecond (μs) level, and is suitable for large-scale inference or online real-time decision-making scenarios.

[0051] In this application, the accelerator chip has a large number of parallel processing units, and the feature matrix of the operation task to be processed can be simultaneously calculated for similarity with the expert weight matrices of each expert network in the expert hybrid model in the accelerator chip, so that the similarities between the operation task to be processed and all expert networks can be output at one time, improving the calculation speed. In the traditional solution, the selection of the expert network generally relies on the CPU, but this has the disadvantages of fragmented hardware resources and the separation of the calculation unit and the selection logic, resulting in a decision-making delay exceeding 50 μs, far from meeting the calculation requirements of real-time scenarios of edge devices (such as business requirements that require ensuring calculation real-time, such as autonomous driving).

[0052] In this application, the expert weight matrix stored in the distributed look-up table can be updated in real time based on the boundary scan test interface.

[0053] In this application, the expert weight matrices of each expert network in the expert hybrid model can be flexibly, online, seamlessly upgraded and maintained, and dynamically adjusted according to requirements such as model optimization, online learning, and model upgrade. If the weight matrix is hard-coded inside the accelerator chip, it will not be able to flexibly adapt to model changes. Therefore, in this application, the expert weight matrix stored in the distributed look-up table can be updated in real time based on the boundary scan test interface.

[0054] Specifically, the boundary scan test interface (Boundary Scan, JTAG) is a chip-level debugging and data access interface used to connect the hardware devices between the accelerator chip and the CPU. Through the boundary scan test interface, the expert weight matrix stored in the distributed look-up table (LUT) can be accessed and modified in real time during the operation of the accelerator chip.

[0055] In this application, the operation process of updating the expert weight matrix in the distributed look-up table can be as follows: First, the external master (such as the CPU) establishes communication with the accelerator chip through the boundary scan test interface, and then writes the new expert weight matrix item by item or in batches into the specified LUT of the accelerator chip through JTAG instructions. The LUT content is refreshed in real time, and the expert weight matrix takes effect immediately without restarting or reconfiguring the chip, thus providing a solid technical guarantee for the efficient deployment, continuous optimization, and intelligent evolution of the expert hybrid model. In addition, the boundary scan interface also supports regional, batch, and atomic writing to ensure data consistency and stability during the update process of the expert weight matrix.

[0056] In this application, further, the parallel calculation of the similarity between the operation task to be processed and each expert network in the expert hybrid model in the accelerator chip can also be calculated according to the following step 1113: Step 1113: According to the Hamming distance algorithm, or cosine similarity algorithm, or gating network, calculate the similarity between the operation task to be processed and each expert network in the expert mixture model in parallel in the accelerator chip.

[0057] In this application, the Hamming distance algorithm is used in the scenario where the feature matrix of the operation task to be processed and the expert weight matrix are binary (0 / 1) or fixed-length discrete encoded. Its calculation method is to count the number of differences in the corresponding bits between the feature matrix of the operation task to be processed and the expert weight matrix, that is, the Hamming distance. In terms of hardware implementation, the accelerator chip can quickly calculate the Hamming distance between the operation task to be processed and the expert network through a parallel XOR gate and adder array, thus greatly improving the processing speed.

[0058] In this application, the cosine similarity algorithm can be used in the scenario where the feature matrix of the operation task to be processed and the expert weight matrix are real numbers or high-dimensional floating-point vectors. Its calculation method is to calculate the cosine value of the included angle between the vector of the feature matrix of the operation task to be processed and the expert weight matrix through dot product and norm normalization. The larger the value, the higher the similarity. In terms of hardware implementation, multiple parallel multiply-accumulate units can be configured inside the accelerator chip to achieve the dot product and modulus length calculation of the input vector and multiple expert weight matrices, support pipeline parallel processing, and significantly shorten the calculation delay.

[0059] In this application, the gating network is suitable for scenarios that require a more complex expert selection mechanism. Its calculation method is to use the gating network as a small neural network, with the input being the feature matrix of the operation task to be processed and the output being the selection probability or weight of each expert network. In terms of hardware implementation, a lightweight gating network inference unit is integrated in the accelerator chip, and the parallel computing resources are used to score multiple expert networks simultaneously, thus realizing efficient expert network allocation.

[0060] In this application, by implementing the calculation methods of Hamming distance, or cosine similarity, or gating network similarity in parallel in the accelerator chip, high-speed similarity evaluation between the operation task to be processed and the expert network can be achieved, giving full play to the advantages of hardware parallel computing, greatly improving the efficiency and flexibility of expert selection, and providing support for the high-performance inference of the expert mixture model.

[0061] In the above step 112, the set number of expert networks with the highest similarity is selected as the target expert network for performing the operation task to be processed. Specifically, first, the set number N of the target expert network can be determined according to the actual model design, hardware resources, and task requirements. For example, 2, 4, or 8 target expert networks with the highest similarity can be selected, and this parameter can be dynamically adjusted to adapt to different task complexities or hardware load states. Then, after obtaining the similarity scores of all expert networks, it is necessary to efficiently screen out the N expert networks with the highest scores. For this, it can be achieved by adopting efficient algorithms such as parallel sorting, priority queue, and local maximum reduction on the accelerator chip to further shorten the selection time. For example, the FPGA chip can implement a parallel comparison tree, the GPU chip can use a block reduction algorithm, and the ASIC chip can customize a dedicated selection circuit. Finally, the selected N expert networks are the target expert networks for performing the operation task to be processed, and thus the corresponding network parameters can be scheduled to the computing unit according to the storage locations of the network parameters of these target expert networks (such as on-chip memory, off-chip memory) for subsequent execution of the operation task to be processed.

[0062] Continue to refer to Figure 1 , in step 120, the first network parameters of the first expert network in the target expert network are called from the on-chip memory, and the second network parameters of the second expert network in the target expert network are called from the off-chip memory.

[0063] Please continue to refer to Figure 2 , in this application, after determining the target expert network for performing the operation task to be processed, when the target expert network includes the first expert network, the first network parameters can be directly read at high speed from the on-chip memory. When the target expert network includes the second expert network, the required second network parameters can be called from the off-chip memory through the on-chip controller (such as DMA, AXI bus).

[0064] In this application, the network parameters of some expert networks are stored in the on-chip memory, and the network parameters of some other expert networks are stored in the off-chip memory. The advantage is that it can reduce the occupancy of the memory space in the accelerator chip, and thus effectively relieve the on-chip memory pressure caused by the increase in the number of expert networks, realizing the efficient utilization of memory resources. Further, by releasing the memory resources of the accelerator chip, more memory resources can be allocated to the computing units, thereby enhancing the overall parallel computing ability and task processing throughput of the accelerator chip. In addition, the reduction of the on-chip memory pressure also helps to reduce memory access conflicts and bandwidth bottlenecks, optimize the data flow path, and improve the data processing efficiency. At the same time, it also provides greater flexibility and scalability for subsequent model upgrades, expert network expansions, or multi-model collaborative deployments, helping to meet diverse application requirements.

[0065] Further, in the present application, the first expert network may be an expert network with a higher usage frequency, and the second expert network may be an expert network with a lower usage frequency. That is, preferentially storing the parameters of the expert network with a higher invocation frequency in the on-chip memory can significantly improve the access speed of these expert networks during inference or training, reduce the parameter loading latency, and ensure the real-time performance and efficiency of high-frequency tasks. While storing the parameters of the expert network with a lower invocation frequency in the off-chip memory and loading them on demand when needed can avoid wasting on-chip memory resources. Through the hierarchical storage mechanism, not only can the dependence of the accelerator chip on a large-capacity on-chip memory be reduced, the hardware cost be lowered, but also the efficient deployment and expansion of the expert mixture model can be achieved, and the acceleration efficiency of the expert mixture model for performing computing tasks can be improved.

[0066] In the present application, the following steps 113 to 115 may also be executed: Step 113, in response to an update instruction for the network parameters in the on-chip memory, statistically count the frequencies of each expert network in the expert mixture model for performing computing tasks in history.

[0067] Step 114, define the set number of expert networks with the highest frequencies as the first expert network, and define the expert networks in the expert mixture model other than the first expert network as the second expert network.

[0068] Step 115, save the network parameters of the first expert network in the on-chip memory of the accelerator chip, and save the network parameters of the second expert network in the off-chip memory of the accelerator chip.

[0069] In the present application, after receiving an update instruction for the network parameters in the on-chip memory, the historical computing task execution frequencies of each expert network in the expert mixture model can be statistically counted. Specifically, the number of times each expert network is scheduled, invoked, or participates in an inference / training task can be recorded by a counter to form a statistical table of the access frequencies of the expert networks. This statistical process can be dynamically updated based on a fixed time window (such as the most recent N tasks) or a sliding window to reflect the activity levels of each expert network under the current actual workload. In this way, a data basis is provided for the subsequent hierarchical storage and resource allocation of the expert networks.

[0070] After completing the frequency statistics, according to the statistical results, a set number (e.g., the first K) of expert networks with the highest historical operation frequencies can be defined as the "first expert network", and the remaining expert networks as the "second expert network". This partitioning method can ensure that the expert networks that are frequently called and have a greater impact on the overall model performance are given priority attention. The set number can be flexibly adjusted according to the actual on-chip memory capacity, the scale of the expert network, and application requirements, so as to achieve a balance between performance and resource utilization.

[0071] According to the above partitioning results, the network parameters of the first expert network can be preferentially stored in the on-chip memory of the accelerator chip, so as to enable fast access in subsequent operation tasks, achieve microsecond-level response, and at the same time reduce the number of accesses to off-chip memory, reduce the network parameter loading delay, and improve the response speed and execution efficiency of high-frequency tasks. While the network parameters of the second expert network are stored in off-chip memory and are only loaded into on-chip memory for calculation when needed, which can significantly reduce the pressure on on-chip memory, release more memory resources for other key modules or calculation tasks, and thus improve the utilization rate of on-chip memory resources, achieving the effect of trading time for space. For example, in a large-scale language model inference server, some expert networks are frequently selected for processing mainstream semantic tasks, and their parameters are resident in on-chip memory to achieve microsecond-level response; while the expert networks for processing special domain corpora are stored in off-chip memory and are only dynamically loaded when specific inputs appear. Another example is in edge AI devices, where the on-chip memory is more limited. Implementing hierarchical storage management for expert network parameters not only ensures the fast response of high-frequency expert networks in inference tasks but also takes into account the balance between model capacity and hardware resources, effectively balancing the model capacity and real-time requirements.

[0072] Based on the technical solutions of the above steps 113 to 115, the expert mixture model can dynamically adjust the parameter storage strategy according to the actual running situation, realizing hierarchical management of "hot experts resident and cold experts loaded on demand". This mechanism can not only improve the utilization efficiency of on-chip memory, avoid low-frequency expert networks occupying precious on-chip resources, but also ensure the efficient access of high-frequency expert networks, optimizing the overall system performance. In addition, this solution has good adaptability and scalability, can flexibly handle changes in the scale of expert networks or fluctuations in workload, and meet the dual requirements of high performance and high resource utilization in practical applications.

[0073] In this application, the following step 1131 can also be executed: Step 1131, before calling the first network parameters of the first expert network in the target expert network from the on-chip memory and calling the second network parameters of the second expert network in the target expert network from the off-chip memory, trigger the update instruction of the network parameters in the on-chip memory.

[0074] In this application, before the target expert network parameters are about to be called (i.e., the first network parameters of the first expert network are called from on-chip memory, and the second network parameters of the second expert network are called from off-chip memory), by actively triggering the update instruction of the network parameters in the on-chip memory, it can be ensured that before each actual operation task arrives, based on the latest expert network usage situation, the storage and scheduling strategies of the parameters can be dynamically optimized.

[0075] Specifically, triggering the update instruction can re-stat the call frequencies of each expert network during the recent execution of operation tasks, and based on this, determine which expert networks are frequently used (the first expert network) and which are infrequently used (the second expert network). According to the statistical results, the storage locations of the expert network parameters can be dynamically adjusted. For example, some frequently used expert network parameters can be migrated to on-chip memory, and the infrequently used expert network parameters can be transferred to off-chip memory. As the types of operation tasks, input data, or business loads change, the activities of different expert networks may change. Through the update instruction before each parameter call, the distribution of the expert network parameters can be adjusted in a timely manner, avoiding waste of on-chip memory resources or the occurrence of bottlenecks, thereby achieving dynamic optimization of resource utilization and system performance. For example, in scenarios such as multi-task inference and online model learning, the activities of different expert networks may change with the input data or task types. By triggering the update instruction before each parameter call, it can be ensured that the system always responds to the current task with the optimal layout of expert network parameters, improving the system response ability and throughput.

[0076] Continue to refer to Figure 1 In step 130, based on the first network parameter and the second network parameter, the to-be-processed operation task is executed by the target expert network in the expert mixture model.

[0077] In this application, first, the first network parameter and the second network parameter can be loaded into the computing unit. Then, through the routing mechanism of the mixture-of-experts model, the input data corresponding to the operation task to be processed is assigned to the corresponding expert subnetworks in the target expert network. Each expert subnetwork independently or collaboratively completes the calculation of subtasks based on its own network parameters, and finally outputs the calculation results of each expert network, and performs fusion or selection according to the model design to obtain the final operation result output. In this application, the high-frequency expert network parameters are resident in the on-chip memory, while the low-frequency expert network parameters are loaded from the off-chip memory as needed, which can ensure efficient parameter access and utilization of computing resources, thereby ensuring that the system can efficiently respond to diverse task requirements even under limited hardware resources. In scenarios such as large-scale multi-task inference and personalized model services, different tasks will activate different expert networks, and the system can flexibly and efficiently schedule the required expert networks to complete the processing of complex tasks. Therefore, by combining the first network parameter and the second network parameter to drive the target expert network in the mixture-of-experts model to work collaboratively, the efficient execution of the operation task to be processed is achieved, ensuring the flexibility, scalability of the model and the optimal utilization of system resources.

[0078] In this application, the following steps 131 to 132 can also be performed: Step 131, abstract the common operators in the accelerator chip into configurable modules.

[0079] Step 132, based on the time-division multiplexing algorithm, run each target expert network through the configurable module.

[0080] In this application, the common operators in the deep learning model (such as GELU activation function, Layer-Norm normalization layer, convolutional layer, fully connected layer, etc.) can be abstracted and designed as configurable modules. These modules have flexible function and parameter interfaces, and can be configured and reused according to the actual needs of different expert networks, thereby greatly reducing the redundancy of hardware resources, improving the chip utilization rate, and facilitating the adaptation and expansion of subsequent diverse models. After verification, by serving 32 expert networks through time-division multiplexing (TDM), the resource utilization rate of a single GELU activation function can be increased by 4 times.

[0081] Furthermore, the configurable module can be dynamically scheduled by adopting the time-division multiplexing algorithm. Specifically, the right to use the configurable module can be allocated in turn according to time slices according to the target expert network that needs to be run currently. In each time slot, the module loads the parameters and configurations of the corresponding expert network to complete the specified calculation task, and then switches to the next expert network after the time slot ends. In this way, the shared operator can achieve the efficient reuse and dynamic allocation of limited hardware resources among multiple expert networks, and can also greatly reduce the occupancy rate of memory resources.

[0082] Based on the technical solutions of the above steps 131 to 132, by abstracting common operators into configurable modules and combining time-division multiplexing algorithms for resource scheduling, the utilization rate of hardware resources and the flexibility of the system can be significantly improved. In this way, not only can it support the high-concurrency and high-throughput operation of multiple expert networks, but also it provides an efficient and scalable hardware foundation for the large-scale inference and training of expert hybrid models. Taking the 128-expert model of DeepSeek R1 as an example, after the input text vector is matched by the LUT selection engine, the top 8 experts are selected. The popularity statistics module marks these 8 experts as high-frequency experts and preloads their weights from DDR to BRAM. At the same time, the shared GELU unit serves these two experts through TDM time-sharing, and the final output is passed to the next layer after weighted aggregation. After verification, the latency of the entire process can be controlled within 8 μs.

[0083] Based on the technical solution proposed in this application, by deploying an expert hybrid model in an accelerator chip and storing the parameters of different expert networks in on-chip memory and off-chip memory respectively, the acceleration efficiency of the model in performing arithmetic tasks can be effectively improved. Specifically, the parameters of the first expert network that are commonly used or frequently calculated can be stored in on-chip memory, which can make full use of the characteristics of high-speed and low-latency access of on-chip memory to achieve fast invocation of key network parameters, significantly reduce data access latency, and improve the overall inference and training speed. And storing the parameters of the second expert network in off-chip memory can break through the capacity limit of on-chip memory, support the flexible deployment of more and larger-scale expert networks, and provide guarantees for model scalability and diversity. During the execution of specific arithmetic tasks, the network parameters of on-chip memory and off-chip memory can be flexibly invoked according to actual needs to give full play to the advantages of the two-level storage structure. On the one hand, the high bandwidth and low power consumption characteristics of on-chip memory ensure the efficient processing of high-frequency access to network parameters and reduce data transfer and access bottlenecks; on the other hand, the capacity advantage of off-chip memory supports the complex structure and large-scale parameter storage requirements of the expert hybrid model. Through this hierarchical storage and efficient scheduling mechanism, the efficient acceleration of the arithmetic tasks of the expert hybrid model can finally be achieved, improving the utilization rate of hardware resources and the overall performance of the system, and meeting the high-efficiency inference and training needs of large-scale and complex expert hybrid models in practical applications.

[0084] The accelerator chip arithmetic task acceleration method based on the expert hybrid model proposed in this application can be widely applied to the following typical application scenarios: 1. Large-scale intelligent inference platforms. In scenarios such as cloud AI inference servers and data centers, which need to efficiently process a large number of concurrent inference requests, there are high requirements for system throughput, latency, and energy efficiency.

[0085] 2. Intelligent edge devices. Embedded devices such as intelligent cameras, drones, autonomous vehicles, and smart homes are limited by hardware resources and power consumption constraints and need to achieve high-performance model inference with limited chip resources.

[0086] 3. High-speed image / video processing. In scenarios such as real-time video analysis, image recognition, and object detection, it is necessary to dynamically select the optimal expert network for different input features to improve the accuracy and processing speed of model inference.

[0087] 4. Natural language processing and speech recognition. In fields such as machine translation, speech recognition, and dialogue systems, the expert mixture model can dynamically allocate computing resources according to different text / speech features to improve the system response speed and intelligence level.

[0088] 5. Intelligent manufacturing and industrial automation. In fields such as industrial robots and intelligent detection equipment, it is necessary to efficiently process complex and changing input tasks. The expert mixture model can dynamically adapt according to task characteristics to improve production efficiency and intelligence level.

[0089] 6. Financial risk control and intelligent recommendation. In scenarios such as financial anti-fraud and intelligent recommendation systems, the expert mixture model can achieve efficient processing of diverse input features and improve the model's risk recognition and personalized recommendation capabilities.

[0090] The following describes the device embodiments of the present application, which can be used to execute the acceleration method for the expert mixture model to perform arithmetic tasks in the above embodiments of the present application. The expert mixture model is deployed in an accelerator chip. For details not disclosed in the device embodiments of the present application, please refer to the embodiments of the acceleration method for the expert mixture model to perform arithmetic tasks in the above of the present application.

[0091] See Figure 3 , which shows a block diagram of an acceleration device for an expert mixture model to perform arithmetic tasks in an embodiment of the present application.

[0092] As Figure 3 shown, the acceleration device 300 for an expert mixture model to perform arithmetic tasks according to an embodiment of the present application includes: an acquisition unit 301, a call unit 302, and an execution unit 303.

[0093] Among them, an obtaining unit 301 is configured to obtain an operation task to be processed and determine a target expert network in the expert mixture model for executing the operation task to be processed. Network parameters of each first expert network in the expert mixture model are stored in on-chip memory of the accelerator chip, and network parameters of each second expert network in the expert mixture model are stored in off-chip memory of the accelerator chip. A calling unit 302 is configured to call first network parameters of the first expert network in the target expert network from the on-chip memory and call second network parameters of the second expert network in the target expert network from the off-chip memory. An execution unit 303 is configured to execute the operation task to be processed through the target expert network in the expert mixture model based on the first network parameters and the second network parameters.

[0094] In some embodiments of the present application, based on the foregoing solution, the obtaining unit 301 is configured to: calculate similarities between the operation task to be processed and each expert network in the expert mixture model in parallel in the accelerator chip; and select a set number of expert networks with the highest similarities as the target expert network for executing the operation task to be processed.

[0095] In some embodiments of the present application, based on the foregoing solution, the obtaining unit 301 is configured to: obtain a feature matrix of the operation task to be processed and an expert weight matrix of each expert network in the expert mixture model; and calculate similarities between the operation task to be processed and each expert network in the expert mixture model in parallel in the accelerator chip based on the feature matrix and the expert weight matrix.

[0096] In some embodiments of the present application, based on the foregoing solution, the obtaining unit 301 is configured to: calculate similarities between the operation task to be processed and each expert network in the expert mixture model in parallel in the accelerator chip according to a Hamming distance algorithm, a cosine similarity algorithm, or a gating network.

[0097] In some embodiments of the present application, based on the foregoing solution, the expert weight matrix of each expert network in the expert mixture model is pre-stored in a distributed lookup table of the accelerator chip.

[0098] In some embodiments of the present application, based on the foregoing solution, the apparatus further includes an updating unit configured to update the expert weight matrix stored in the distributed lookup table in real time based on a boundary scan test interface.

[0099] In some embodiments of the present application, based on the foregoing solution, the device further includes: a saving unit, configured to, in response to an update instruction for network parameters in the on-chip memory, count the frequencies at which each expert network in the expert mixture model has performed arithmetic tasks historically; define the set number of expert networks with the highest frequencies as the first expert networks, and define the expert networks in the expert mixture model other than the first expert networks as the second expert networks. Save the network parameters of the first expert networks in the on-chip memory of the accelerator chip, and save the network parameters of the second expert networks in the off-chip memory of the accelerator chip.

[0100] In some embodiments of the present application, based on the foregoing solution, the device further includes: a triggering unit, configured to trigger an update instruction for network parameters in the on-chip memory before calling the first network parameters of the first expert network in the target expert network from the on-chip memory and calling the second network parameters of the second expert network in the target expert network from the off-chip memory.

[0101] In some embodiments of the present application, based on the foregoing solution, the device further includes: an abstracting unit, configured to abstract common operators in the accelerator chip into configurable modules; and run each target expert network through the configurable modules based on a time-division multiplexing algorithm.

[0102] In some embodiments of the present application, based on the foregoing solution, the accelerator chip includes any one of a field-programmable gate array chip, a graphics processing chip, a Google Tensor Processing Unit chip, and an application-specific integrated circuit.

[0103] Based on the same inventive concept, an embodiment of the present application provides a computer program product, which includes computer instructions stored in a computer-readable storage medium and adapted to be read and executed by a processor, so that a computer device having the processor performs operations executed by the acceleration method for an expert mixture model to perform arithmetic tasks as described above.

[0104] Based on the same inventive concept, an embodiment of the present application provides a computer-readable storage medium, in which at least one computer program instruction is stored, and the at least one computer program instruction is loaded and executed by a processor to implement the operations executed by the acceleration method for an expert mixture model to perform arithmetic tasks as described above.

[0105] Based on the same inventive concept, an embodiment of the present application further provides an electronic device, refer to Figure 4, which shows a schematic structural diagram of an electronic device in an embodiment of the present application. The electronic device includes one or more memories 404, one or more processors 402, and at least one computer program (computer program instructions) stored on the memory 404 and executable on the processor 402. When the processor 402 executes the computer program, it implements the acceleration method for the expert mixture model to perform arithmetic tasks as described above.

[0106] Among them, in Figure 4 , the bus architecture (represented by bus 400), bus 400 may include any number of interconnected buses and bridges. Bus 400 links together various circuits including one or more processors represented by processor 402 and memories represented by memory 404. Bus 400 may also link together various other circuits such as peripheral devices, voltage regulators, and power management circuits, etc., which are well known in the art. Therefore, they will not be further described herein. Bus interface 405 provides an interface between bus 400 and receiver 401 and transmitter 403. Receiver 401 and transmitter 403 may be the same element, i.e., a transceiver, which provides a unit for communicating with various other devices on the transmission medium. Processor 402 is responsible for managing bus 400 and general processing, while memory 404 may be used to store data used by processor 402 when performing operations.

[0107] The functions described herein may be implemented in hardware, software executed by a processor, firmware, or any combination thereof. If implemented in software executed by a processor, the functions may be stored on a computer-readable medium or transmitted via a computer-readable medium as one or more instructions or codes. Other examples and implementations are within the scope and spirit of the present application and the appended claims. For example, due to the nature of software, the functions described above may be implemented using software executed by a processor, hardware, firmware, hardwiring, or any combination of these. In addition, each functional unit may be integrated in one processing unit, or each unit may exist physically alone, or two or more units may be integrated in one unit.

[0108] In several embodiments provided by the present application, it should be understood that the disclosed technical content can be implemented in other ways. Among them, the device embodiments described above are merely illustrative. For example, the division of the units may be a logical function division, and there may be other division methods in actual implementation. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection to each other may be through some interfaces, and the indirect coupling or communication connection of units or modules may be in an electrical or other form.

[0109] The unit described as a separating component may or may not be physically separated. The component as a control device may or may not be a physical unit, that is, it may be located in one place or may be distributed among multiple units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0110] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of this application. The foregoing storage medium includes: various media such as USB flash drives, read-only memories (ROMs), random access memories (RAMs), mobile hard disks, magnetic disks or optical discs that can store computer program instructions.

[0111] The above are only the embodiments of this application and are not used to limit this application. For those skilled in the art, this application can have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of this application shall be included within the scope of the claims of this application.

Claims

1. An acceleration method for an expert mixture model to perform arithmetic tasks, characterized in that, The expert mixture model is deployed in an accelerator chip, and the method includes: Obtain a computing task to be processed, and determine a target expert network in the expert mixture model for executing the computing task to be processed. The network parameters of each first expert network in the expert mixture model are stored in the on-chip memory of the accelerator chip, and the network parameters of each second expert network in the expert mixture model are stored in the off-chip memory of the accelerator chip; Call the first network parameters of the first expert network in the target expert network from the on-chip memory, and call the second network parameters of the second expert network in the target expert network from the off-chip memory; Based on the first network parameters and the second network parameters, execute the computing task to be processed through the target expert network in the expert mixture model.

2. The method according to claim 1, characterized in that, The determining the target expert network in the expert mixture model for executing the computing task to be processed includes: Parallelly calculate the similarity between the computing task to be processed and each expert network in the expert mixture model in the accelerator chip; Select the set number of expert networks with the highest similarity as the target expert network for executing the computing task to be processed.

3. The method according to claim 2, characterized in that, The parallelly calculating the similarity between the computing task to be processed and each expert network in the expert mixture model in the accelerator chip includes: Obtain the feature matrix of the computing task to be processed and the expert weight matrix of each expert network in the expert mixture model; Based on the feature matrix and the expert weight matrix, parallelly calculate the similarity between the computing task to be processed and each expert network in the expert mixture model in the accelerator chip.

4. The method according to claim 3, wherein The parallelly calculating the similarity between the computing task to be processed and each expert network in the expert mixture model in the accelerator chip includes: Parallelly calculate the similarity between the computing task to be processed and each expert network in the expert mixture model in the accelerator chip according to the Hamming distance algorithm or the cosine similarity algorithm or the gating network.

5. The method according to claim 3, characterized in that The expert weight matrix of each expert network in the expert mixture model is pre-stored in the distributed lookup table of the accelerator chip.

6. The method according to claim 5, wherein The method further includes: Based on the boundary scan test interface, update the expert weight matrix stored in the distributed lookup table in real time.

7. The method according to claim 1, characterized in that The method further includes: In response to the update instruction of the network parameters in the on-chip memory, count the frequencies of each expert network in the expert mixture model performing computing tasks historically; Define the set number of expert networks with the highest frequency as the first expert network, and define the expert networks in the expert mixture model other than the first expert network as the second expert network; Store the network parameters of the first expert network in the on-chip memory of the accelerator chip, and store the network parameters of the second expert network in the off-chip memory of the accelerator chip.

8. The method according to claim 7, wherein The method further includes: Before calling the first network parameters of the first expert network in the target expert network from the on-chip memory and calling the second network parameters of the second expert network in the target expert network from the off-chip memory, trigger an update instruction for the network parameters in the on-chip memory.

9. The method according to claim 1, wherein The method further includes: Abstracting common operators in the accelerator chip into configurable modules; Based on the time-division multiplexing algorithm, running each target expert network through the configurable module.

10. The method according to claim 1, characterized in that, The accelerator chip includes any one of a field programmable gate array chip, a graphics processing chip, a Google tensor processing chip, and an application specific integrated circuit.

11. An acceleration device for an expert mixture model to perform arithmetic tasks, characterized in that, The expert mixture model is deployed in the accelerator chip, and the device includes: An acquisition unit, configured to acquire a to-be-processed operation task and determine a target expert network in the expert mixture model for executing the to-be-processed operation task, wherein network parameters of each first expert network in the expert mixture model are stored in the on-chip memory of the accelerator chip, and network parameters of each second expert network in the expert mixture model are stored in the off-chip memory of the accelerator chip; A calling unit, configured to call the first network parameters of the first expert network in the target expert network from the on-chip memory and call the second network parameters of the second expert network in the target expert network from the off-chip memory; An execution unit, configured to execute the to-be-processed operation task through the target expert network in the expert mixture model based on the first network parameters and the second network parameters.

12. A computer program product, characterized in that, The computer program product includes computer instructions, which are stored in a computer-readable storage medium and are adapted to be read and executed by a processor, so that a computer device having the processor executes the method according to any one of claims 1 to 10.

13. A computer-readable storage medium, characterized in that, At least one program code is stored in the computer-readable storage medium, and the at least one program code is loaded and executed by a processor to implement the operations performed by the method according to any one of claims 1 to 10.

14. An electronic device, characterized in that, The electronic device includes one or more processors and one or more memories, and at least one program code is stored in the one or more memories, and the at least one program code is loaded and executed by the one or more processors to implement the method according to any one of claims 1 to 10.

Citation Information

Patent Citations

  • Deep learning-oriented FPGA universal heterogeneous acceleration method and equipment

    CN118014022A

  • Hybrid expert model reasoning acceleration method, device, equipment, medium and program

    CN118761472A

  • Server board card, performance adjusting method of server board card and chip

    CN119127632A

  • Method and device for executing operation task on multi-core system and related product

    CN119806807A

  • Deployment of resources in mixture of experts processing

    US20250086424A1

Cited By

  • Model loading and unloading method and electronic equipment

    CN120909807A

  • Tensor processing method and device based on hybrid expert network, and storage medium

    CN121072622A

  • Expert model preloading method and device, chip, electronic equipment, storage medium and computer program product

    CN121116656A

  • Data processing method, terminal equipment and storage medium

    CN121255115A

  • Data processing method, terminal device, and storage medium

    CN121255115B