Expert network processing method and device based on hybrid expert model, and electronic equipment

By selecting a target expert network in a hybrid expert model and combining it with target CPU and GPU to process tokens, the problem of unbalanced GPU load is solved, achieving load balancing and improved processing speed, and avoiding inference pauses caused by copying model parameters.

CN121008905APending Publication Date: 2025-11-25BEIJING BAIDU NETCOM SCI & TECH CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510990435.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-17
Publication Date
2025-11-25

AI Technical Summary

Technical Problem

In hybrid expert models, the difference in the number of sequence units processed by different expert networks leads to uneven GPU load, resulting in excessive GPU computing power being occupied by popular expert networks, which affects token processing speed.

Method used

By selecting a target expert network and combining it with the target CPU and GPU to process tokens, the model parameters of all expert networks are pre-stored on the candidate CPU, avoiding model parameter copying, thus achieving load balancing and improving processing speed.

Benefits of technology

It reduces GPU load, improves token processing speed, avoids slowdowns caused by excessive GPU computing power, and reduces inference pauses for candidate CPUs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121008905A_ABST
    Figure CN121008905A_ABST
Patent Text Reader

Abstract

The invention provides an expert network processing method based on a hybrid expert model, and relates to the technical field of artificial intelligence such as large models, deep learning, natural language processing and computer vision. The expert network processing method based on the hybrid expert model comprises the following steps: selecting a target expert network from a plurality of expert networks according to a first processing number of target sequence unit tokens respectively processed by the plurality of expert networks; obtaining a second processing number corresponding to the target expert network according to a subtraction result between the first processing number and the target processing number; selecting a target CPU corresponding to the target expert network from the plurality of candidate CPUs according to the second processing quantity; using the target GPU to process the first target token corresponding to the target processing quantity, and using the target CPU to process the second target token corresponding to the second processing quantity; and obtaining a target processing result according to the processing result output by the target GPU and the processing result output by the target CPU. The processing speed of the target token can be improved, and the processing time delay of the target token can be reduced.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present disclosure relates to the technical field of computer, and particularly relates to the technical field of artificial intelligence such as large model, deep learning, natural language processing and computer vision. A method, device and electronic equipment for processing expert network based on mixed expert model and a readable storage medium are provided. BACKGROUND

[0002] In the mixed expert model (Moe) architecture, the number of sequence units (tokens) to be processed by different expert networks is obviously different, thereby directly leading to the unbalanced load of GPUs (graphics processing units) of different expert networks. For example, when the number of tokens to be processed by a certain expert network is too large, the computing power of the GPU where the expert network is located will be excessively occupied, thereby affecting the processing speed of the tokens. Therefore, how to reduce the load of the GPU where the popular expert is located to improve the processing speed of the tokens is a technical problem to be solved at present. SUMMARY

[0003] According to a first aspect of the present disclosure, a method for processing expert network based on mixed expert model is provided, comprising: selecting a target expert network from a plurality of expert networks in a mixed expert model according to a first processing number of target sequence units (tokens) processed by the plurality of expert networks respectively; obtaining a second processing number corresponding to the target expert network according to a subtraction result between the first processing number and a target processing number; selecting a target CPU corresponding to the target expert network from a plurality of candidate CPUs according to the second processing number, wherein the candidate CPUs include model parameters of all expert networks; processing first target tokens corresponding to the target processing number using a target GPU corresponding to the target expert network, and processing second target tokens corresponding to the second processing number using the target CPU; and obtaining a target processing result output by the target expert network when processing the target tokens according to a processing result output by the target GPU and a processing result output by the target CPU.

[0004] According to a second aspect of the present disclosure, an expert network processing apparatus based on a hybrid expert model is provided, comprising: a first selection unit configured to select a target expert network from a plurality of expert networks in the hybrid expert model according to a first processing quantity of target sequence units tokens processed by the plurality of expert networks respectively; a processing unit configured to obtain a second processing quantity corresponding to the target expert network according to a subtraction result between the first processing quantity and a target processing quantity; a second selection unit configured to select a target CPU corresponding to the target expert network from a plurality of candidate CPUs according to the second processing quantity, wherein the candidate CPUs comprise model parameters of all expert networks; a running unit configured to process first target tokens corresponding to the target processing quantity using a target GPU corresponding to the target expert network, and process second target tokens corresponding to the second processing quantity using the target CPU; and an output unit configured to obtain a target processing result output by the target expert network when processing the target tokens according to a processing result output by the target GPU and a processing result output by the target CPU.

[0005] According to a third aspect of the present disclosure, an electronic device is provided, comprising: at least one processor; and a memory connected with the at least one processor in communication; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method as described above.

[0006] According to a fourth aspect of the present disclosure, a non-transitory computer readable storage medium storing computer instructions is provided, wherein the computer instructions are used to enable the computer to perform the method as described above.

[0007] According to a fifth aspect of the present disclosure, a computer program product is provided, comprising a computer program which, when executed by a processor, implements the method as described above.

[0008] It should be understood that the content described in this part is not intended to identify key or important features of the embodiments of the present disclosure, nor to limit the scope of the present disclosure. Other features of the present disclosure will become apparent from the following description. BRIEF DESCRIPTION OF DRAWINGS

[0009] The accompanying drawings are used to better understand the present scheme, and do not limit the present disclosure. Among them:

[0010] Figure 1 is a schematic diagram according to the first embodiment of the present disclosure;

[0011] Figure 2 is a schematic diagram according to the second embodiment of the present disclosure;

[0012] Figure 3 This is a schematic diagram according to the third embodiment of the present disclosure;

[0013] Figure 4 This is a schematic diagram according to the fourth embodiment of the present disclosure;

[0014] Figure 5 This is a schematic diagram according to the fifth embodiment of the present disclosure;

[0015] Figure 6 This is a block diagram of an electronic device used to implement the expert network processing method based on a hybrid expert model according to the embodiments of this disclosure. Detailed Implementation

[0016] The exemplary embodiments of this disclosure are described below with reference to the accompanying drawings, including various details of the embodiments to aid understanding, and should be considered merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of this disclosure. Similarly, for clarity and brevity, descriptions of well-known functions and mechanisms are omitted in the following description.

[0017] Figure 1 This is a schematic diagram based on the first embodiment of this disclosure. (See diagram below.) Figure 1 As shown, the expert network processing method based on a hybrid expert model in this embodiment specifically includes the following steps:

[0018] S101. Select a target expert network from the multiple expert networks based on the first processing quantity of the target sequence unit token processed by the multiple expert networks in the hybrid expert model.

[0019] S102. Based on the subtraction between the first processing quantity and the target processing quantity, obtain the second processing quantity corresponding to the target expert network;

[0020] S103. Based on the second processing quantity, select the target CPU corresponding to the target expert network from multiple candidate CPUs, wherein the candidate CPUs include the model parameters of all expert networks.

[0021] S104. Use the target GPU corresponding to the target expert network to process the first target token corresponding to the target processing quantity, and use the target CPU to process the second target token corresponding to the second processing quantity;

[0022] S105. Based on the processing results output by the target GPU and the processing results output by the target CPU, obtain the target processing result output by the target expert network when processing the target token.

[0023] The expert network processing method based on a hybrid expert model in this embodiment addresses two issues. First, after selecting a target expert network from multiple expert networks based on the first number of target tokens to be processed by the expert network, the target GPU and target CPU corresponding to that expert network are used to jointly process the target tokens. This achieves full utilization of the computing power provided by the CPU and avoids the problem of excessive GPU computing power being over-occupied when using the GPU alone when the number of target tokens to be processed by the expert network is large (excessive GPU computing power will affect the processing speed of the target tokens). This reduces the load on the GPU where the target expert network is located, thereby improving the processing speed of the target tokens. Second, the model parameters of all expert networks are pre-stored in each candidate CPU, so that each candidate CPU can play the role of any expert network without copying the model parameters when processing the target tokens. This avoids the inference pause problem caused by the candidate CPU copying the model parameters and reduces the processing latency of the target tokens.

[0024] In this embodiment, the Mixture of Experts (MoE) model is a machine learning architecture that can be used for reasoning in multimodal large models or large language models. Its core idea is to decompose complex tasks into multiple sub-tasks, which are processed by multiple expert networks respectively. Then, the processing results output by multiple expert networks are combined through a gating network to obtain the final processing result.

[0025] In this embodiment, the hybrid expert model includes multiple expert networks, which are multiple independent sub-models (usually small neural networks). Each expert network is used to process a specific subset of the input data or a specific task.

[0026] In this embodiment, a sequence unit (i.e., a token) is essentially the smallest unit of the input sequence. Different expert networks are used to process one or more tokens. In natural language processing, a token usually refers to a character or word after the input text has been segmented. In computer vision, a token usually refers to an image patch after the input image has been segmented. In speech processing, a token usually refers to an audio frame after the input audio has been segmented. The target token to be processed by the expert network in this embodiment is specifically the feature vector of the corresponding character, word, image patch, or audio frame.

[0027] The hybrid expert model in this embodiment is deployed in a distributed cluster, which includes multiple computing nodes. Each computing node consists of a CPU (Central Processing Unit) and multiple GPUs (Graphics Processing Units), for example, each computing node includes four GPUs. In this embodiment, one GPU is used to handle the computation of multiple expert networks, for example, two expert networks.

[0028] In this embodiment, when executing S101, the first number of target tokens processed by each expert network in the hybrid expert model is first obtained, and then at least one target expert network is selected from the multiple expert networks based on the first number of target tokens processed by the multiple expert networks.

[0029] In this embodiment, when executing S101, the first processing number of target tokens processed by each expert network can be obtained based on the output of the gating network in the hybrid expert model. The first processing number is the total number of tokens to be processed by each expert network. The gating network is used to determine the expert network that processes each sequence unit based on the feature vector of each sequence unit in the input sequence.

[0030] Specifically, in this embodiment, when executing S101 to select a target expert network from multiple expert networks based on the first processing quantity of the target token processed by multiple expert networks in the hybrid expert model, the following implementation method can be adopted: sort the multiple expert networks in descending order of the first processing quantity; select the top N expert networks as the target expert network based on the sorting result, where N is a positive integer greater than or equal to 1.

[0031] In other words, this embodiment uses a sorting method based on a first processing quantity to select the target expert network, ensuring that at least one selected target expert network needs to process a large number of tokens, thereby improving the accuracy of target expert network selection.

[0032] In order to maximize the utilization of the computing power provided by the CPU and to balance the GPU load as much as possible, this embodiment may also include the following when executing S101: obtaining the number of candidate CPUs, which is the total number of CPUs included in the distributed cluster; taking the obtained number of CPUs as the value of N; for example, if the number of CPUs is 32, this embodiment will select the top 32 expert networks as the target expert network.

[0033] In addition, when performing S101 to select a target expert network from multiple expert networks based on the first processing quantity of the target tokens processed by multiple expert networks in the hybrid expert model, this embodiment can also adopt the following method: based on the first processing quantity of the target tokens processed by multiple expert networks and the target GPUs corresponding to the multiple expert networks, the load quantity of each target GPU is obtained. In this embodiment, the load quantity is the total number of tokens to be processed by the target GPU; the target GPUs whose load quantity meets the preset requirements are selected as the GPUs to be processed. For example, the target GPUs with the top N load quantities are selected as the GPUs to be processed, or the target GPUs whose load quantity exceeds the preset threshold are selected as the GPUs to be processed; the target expert network is selected from the expert networks corresponding to the GPUs to be processed.

[0034] In other words, this embodiment first selects a GPU to be processed from multiple target GPUs based on the first processing quantity corresponding to the expert network and the target GPU, and then selects a target expert network from the expert networks corresponding to the GPU to be processed, so that the selected target expert network corresponds to the GPU with a high load, thereby effectively reducing the load of the GPU with a high load and avoiding the problem of slow processing speed caused by excessive occupation of the computing power of the GPU with a high load.

[0035] For example, if the first processing quantity of expert network 0 is 10, the first processing quantity of expert network 1 is 20, the first processing quantity of expert network 2 is 8, and the first processing quantity of expert network 3 is 10, if the target GPU corresponding to expert network 0 and expert network 1 is GPU0, and the target GPU corresponding to expert network 2 and expert network 3 is GPU1, then the load quantity of GPU0 obtained by executing S101 in this embodiment is 30, and the load quantity of GPU1 is 18. If the load quantity of GPU0 meets the preset requirements, then GPU0 is selected as the GPU to be processed, and then expert network 1 in GPU0 (i.e., the expert network with the larger first processing quantity in the GPU to be processed) can be used as the target expert network, or expert network 0 and expert network 1 (i.e., all expert networks in the GPU to be processed) can be used as target expert networks respectively.

[0036] It is understood that if N in this embodiment is the number of candidate CPUs, then when executing S101, this embodiment can select the expert network with a higher processing volume among the GPUs to be processed as the target expert network to ensure that there are enough CPUs to process the target token corresponding to the selected target expert network.

[0037] In this embodiment, after selecting the target expert network from multiple expert networks in S101, S102 is executed to obtain the second processing number of the corresponding target expert network based on the subtraction result between the first processing number and the target processing number.

[0038] In this embodiment, the target processing number can be preset. The preset target processing number is the number of tokens processed when the expert network has good performance, as obtained through actual testing of the GPU.

[0039] In this embodiment, when executing S102 to obtain the target processing quantity, the following method can also be adopted: obtain the first processing quantity of the target token to be processed by the expert network ranked N+1, and use it as the target processing quantity; for example, when N is 32, this embodiment obtains the first processing quantity of the expert network ranked 33 as the target processing quantity.

[0040] In other words, after sorting the expert networks according to the first processing quantity, this embodiment can also obtain the target processing quantity in real time based on the sorting result, thereby simplifying the steps for obtaining the target processing quantity and improving the speed of obtaining the target processing quantity. Furthermore, by using the first processing quantity of the expert network ranked N+1 as the target processing quantity, the problem of not being able to obtain the second processing quantity based on the first processing quantity of the target expert network and the target processing quantity is effectively avoided, thus improving the accuracy of obtaining the target processing quantity.

[0041] In addition, when executing S102, this embodiment can also obtain the average value among the first processing quantities of the expert networks after the Nth position as the target processing quantity.

[0042] If multiple target expert networks are selected in S101 of this embodiment, then in S102, for each target expert network, the second processing number corresponding to that target expert network can be obtained by subtracting the first processing number of the target expert network from the obtained target processing number.

[0043] In this embodiment, after executing S102 to obtain the second processing quantity of the corresponding target expert network, S103 is executed to select the target CPU of the corresponding target expert network from multiple candidate CPUs according to the second processing quantity; wherein, each candidate CPU in this embodiment includes the model parameters of all expert networks.

[0044] In this embodiment, when executing S103, multiple target expert networks can first be sorted in descending order of the second processing quantity. Then, according to the sorting result, a candidate CPU is assigned to each target expert network in turn, so that the assigned candidate CPU is used as the target CPU for each target expert network.

[0045] In other words, this embodiment prioritizes allocating candidate CPUs to target expert networks with a higher second processing quantity to ensure that candidate CPUs can be used to process a portion of the tokens corresponding to these target expert networks, thereby reducing the load on the GPUs where these target expert networks are located and achieving the purpose of load balancing between different GPUs.

[0046] In this embodiment, a candidate CPU is only used to process a portion of the target tokens corresponding to a target expert network, that is, a candidate CPU can only play the role of a target expert network.

[0047] In this embodiment, after executing S103 to select the target CPU corresponding to the target expert network from multiple candidate CPUs, S104 is executed to use the target GPU of the corresponding target expert network to process the first target token corresponding to the target processing quantity, and to use the target CPU of the corresponding target expert network to process the second target token corresponding to the second processing quantity.

[0048] In other words, after determining the target CPU corresponding to the target expert network, this embodiment can divide the target token to be processed by the target expert network into a first target token corresponding to the target processing quantity and a second target token corresponding to the second processing quantity, so that the target GPU processes the first target token and the target CPU processes the second target token.

[0049] In this embodiment, when executing S104, a token corresponding to the target processing quantity can be randomly selected from the target tokens to be processed by the target expert network as the first target token, and then the remaining tokens in the target tokens can be used as the second target token corresponding to the second processing quantity.

[0050] In this embodiment, when executing S104, the target GPU of the corresponding target expert network processes the first target token corresponding to the target processing quantity, the target GPU corresponding to the target expert network can be determined first, for example, by determining it according to the correspondence table between expert networks and GPUs. Then, the identification information of the target expert network and the first target token are sent to the target GPU so that the target GPU can run the model parameters of the corresponding target expert network to process the first target token according to the identification information.

[0051] In this embodiment, when the target CPU of the corresponding target expert network processes the second target token corresponding to the second processing quantity in S104, the identification information of the target expert network and the second target token can be sent to the target CPU so that the target CPU can run the model parameters of the corresponding target expert network to process the second target token according to the identification information.

[0052] Since the candidate CPUs pre-store the model parameters of all expert networks, in this embodiment, when executing S104, the target CPU corresponding to the target expert network does not need to copy the model parameters, thereby avoiding inference pauses caused by the copying process.

[0053] It is understood that if multiple target CPUs corresponding to the target expert network are determined when S103 is executed in this embodiment, the number of second target tokens that each target CPU can process can also be determined when S104 is executed (for example, determined according to the third processing capacity of the target CPU), and then the identification information of the target expert network and the corresponding number of second target tokens are sent to different target CPUs so that the second target tokens can be processed by different target CPUs.

[0054] For example, if the target expert network is expert network 1, if the number of target tokens to be processed by expert network 1 is 20, and if the target processing quantity is 15, then the second processing quantity is 5; if the target GPU corresponding to expert network 1 is GPU0, and the target CPU corresponding to expert network 1 is CPU1, then in this embodiment, when executing S104, 15 tokens selected from the target tokens will be used as the first target tokens and processed by GPU0, and the remaining 5 tokens will be used as the second target tokens and processed by CPU1.

[0055] For another example, if the target expert network is expert network 1, and the number of target tokens to be processed by expert network 1 is 20, and the target processing quantity is 15, then the second processing quantity is 5. If the target GPU corresponding to expert network 1 is GPU0, and the target CPUs corresponding to expert network 1 are CPU1 and CPU2, and the third processing quantity for each CPU is 3, then in this embodiment, when executing S104, 15 tokens selected from the target tokens will be used as the first target tokens and processed by GPU0. The remaining 5 tokens will be further divided into 3 second target tokens and 2 second target tokens, so that CPU1 can process 3 second target tokens and CPU2 can process 2 second target tokens.

[0056] It is understandable that for non-target expert networks in the expert network, or target expert networks that cannot determine the target CPU, this embodiment can directly send the target token to be processed and the identification information to the corresponding target GPU when executing S104, so that the target GPU can process the received target token.

[0057] In this embodiment, after executing S104, which uses the target GPU to process the first target token and the target CPU to process the second target token, S105 is executed to obtain the target processing result output by the target expert network when processing the target token based on the processing result output by the target GPU and the processing result output by the target CPU.

[0058] In this embodiment, when executing S105, the processing results output by the target GPU and the processing results output by the target CPU can be concatenated according to the order of the target token in the input sequence, so as to obtain the target processing result output by the target expert network when processing the target token based on the concatenation result.

[0059] Figure 2 This is a schematic diagram according to the second embodiment of this disclosure. (See diagram below.) Figure 2 As shown in the figure, when executing S103 "selecting the target CPU corresponding to the target expert network from multiple candidate CPUs according to the second processing quantity", this embodiment can be implemented in the following way:

[0060] S201. Obtain the third processing quantity corresponding to the candidate CPU;

[0061] S202. Determine the target number of CPUs corresponding to the target expert network based on the second processing quantity and the third processing quantity;

[0062] S203. Select candidate CPUs corresponding to the target number of CPUs as the target CPUs for the target expert network.

[0063] In other words, this embodiment fully considers the token processing capabilities of candidate CPUs. The target CPU is selected based on the second processing quantity of the corresponding target expert network and the third processing quantity of the corresponding candidate CPU. This ensures that at least one selected target CPU can meet the token processing requirements while avoiding the problem that the second processing quantity exceeds the processing capability of a single target CPU, thus improving the accuracy of target CPU selection.

[0064] In this embodiment, when executing S201, the CPU computing power of the candidate CPU can be obtained first, and then the processing quantity corresponding to the obtained CPU computing power can be used as the third processing quantity according to the preset correspondence table between CPU computing power and processing quantity.

[0065] In this embodiment, the third processing quantity is the maximum number of tokens that the candidate CPU can process.

[0066] In this embodiment, when executing S202 to determine the target CPU number of the corresponding target expert network based on the second processing quantity and the third processing quantity, the multiple target expert networks can first be sorted in descending order of the second processing quantity. Then, based on the sorting result, the target CPU number of each target expert network can be determined sequentially based on the second processing quantity and the third processing quantity of each target expert network.

[0067] In other words, when there are multiple target expert networks, this embodiment prioritizes allocating candidate CPUs to target expert networks with a higher second processing quantity to ensure that candidate CPUs can be used to process some tokens corresponding to these target expert networks, thereby reducing the load on the GPUs where these target expert networks are located and achieving the purpose of load balancing between different GPUs.

[0068] In this embodiment, when S203 is executed, each candidate CPU can only be selected once, that is, the selected candidate CPU cannot be used as the target CPU of other target expert networks.

[0069] Figure 3 This is a schematic diagram according to the third embodiment of this disclosure. (See diagram below.) Figure 3 As shown in the diagram, this embodiment illustrates a structural diagram of an expert network processing method based on a hybrid expert model. In this embodiment, a computing node includes one CPU and four GPUs. GPU0 is responsible for the computation of expert network 0 and expert network 1, GPU1 is responsible for the computation of expert network 2 and expert network 3, GPU2 is responsible for the computation of expert network 4 and expert network 5, and GPU3 is responsible for the computation of expert network 6 and expert network 7. CPU1 contains the model parameters of all expert networks in the corresponding distributed cluster.

[0070] If this embodiment determines that expert network 1 is the target expert network, and if the target CPU corresponding to the target expert network is determined to be CPU1, then this embodiment can split the target tokens to be processed by expert network 1. The first target token corresponding to the target processing quantity is processed by GPU0, while the second target token corresponding to the second processing quantity is processed by CPU1.

[0071] In other words, this embodiment uses the CPU in the computing node to be responsible for the computation of an expert network, which can reduce the number of computing nodes in the distributed cluster.

[0072] For example, if there are currently 256 expert networks and 32 redundant expert networks, the distributed cluster needs to handle the computation of 288 expert networks. If the CPU in the distributed cluster does not participate in expert network computation, and each computing node is used for the computation of 8 expert networks (the computing node includes 4 GPUs responsible for the computation of 8 expert networks), then the distributed cluster needs 36 computing nodes. In this embodiment, by introducing a CPU to participate in the computation of one expert network, while the distributed cluster still needs to handle the computation of 288 expert networks, each computing node can be used to handle the computation of 9 expert networks (the computing node includes 4 GPUs responsible for the computation of 8 expert networks, and 1 CPU responsible for the computation of 1 expert network), then the distributed cluster only needs 32 computing nodes.

[0073] Therefore, when the number of expert networks that the distributed cluster needs to be responsible for is the same, this embodiment saves 4 computing nodes (i.e., from the original 36 computing nodes to 32 computing nodes) by introducing CPUs to participate in expert network calculations, thereby effectively reducing the demand for expensive GPU resources.

[0074] Figure 4 This is a schematic diagram according to the fourth embodiment of this disclosure. (See diagram below.) Figure 4As shown in the figure, this embodiment illustrates a flowchart of an expert network processing method based on a hybrid expert model: S401, obtaining the first processing quantity of the target tokens processed by multiple expert networks in the hybrid expert model; S402, sorting the multiple expert networks in descending order of the first processing quantity, and selecting the top 32 expert networks as multiple target expert networks, where 32 is the number of candidate CPUs included in the distributed cluster, which includes 32 computing nodes; S403, obtaining the first processing quantity of the target tokens to be processed by the expert network ranked 33rd, as the target processing quantity; S404, for each target expert network, according to the first processing quantity and the target processing quantity... S405. Subtract the processing quantities from each target expert network to obtain the second processing quantity corresponding to the target expert network; S406. For each target expert network, select the target CPU corresponding to the target expert network from multiple candidate CPUs based on the second processing quantity and the third processing quantity of the corresponding candidate CPU; S407. For each target expert network, process the first target token using the target GPU corresponding to the target expert network and process the second target token using the target CPU corresponding to the target expert network; S408. For each target expert network, obtain the target processing result output by the target expert network when processing the target token based on the processing result output by the target GPU and the processing result output by the target CPU.

[0075] In other words, this embodiment selects the target expert network and determines the target processing quantity based on the number of candidate CPUs included in the distributed cluster. This makes the selection of the target expert network and the determination of the target processing quantity more compatible with the distributed cluster in which the hybrid expert model is located, thereby improving the accuracy of the selected target expert network and the accuracy of the determined target processing quantity. Furthermore, by combining the third processing quantity of the corresponding candidate CPUs to complete the selection of the target CPU, the embodiment can make full use of the computing power provided by the candidate CPUs and avoid the problem of excessive tokens affecting the running efficiency of the candidate CPUs, effectively reducing the latency of the expert network when processing tokens.

[0076] Figure 5 This is a schematic diagram according to the fifth embodiment of this disclosure. (See diagram below.) Figure 5 As shown, the expert network processing device 500 based on a hybrid expert model in this embodiment includes:

[0077] The first selection unit 501 is used to select a target expert network from the multiple expert networks based on the first processing quantity of the target sequence unit token processed by the multiple expert networks in the hybrid expert model.

[0078] Processing unit 502 is used to obtain a second processing quantity corresponding to the target expert network based on the subtraction result between the first processing quantity and the target processing quantity;

[0079] The second selection unit 503 is used to select a target CPU corresponding to the target expert network from a plurality of candidate CPUs according to the second processing quantity, wherein the candidate CPUs include the model parameters of all expert networks.

[0080] The execution unit 504 is configured to use the target GPU corresponding to the target expert network to process a first target token corresponding to the target processing quantity, and use the target CPU to process a second target token corresponding to the second processing quantity;

[0081] Output unit 505 is used to obtain the target processing result output by the target expert network when processing the target token based on the processing result output by the target GPU and the processing result output by the target CPU.

[0082] The hybrid expert model in this embodiment is deployed in a distributed cluster, which includes multiple computing nodes. Each computing node consists of a CPU (Central Processing Unit) and multiple GPUs (Graphics Processing Units). In this embodiment, one GPU is used to handle the computation of multiple expert networks.

[0083] The expert network processing device 500 based on the hybrid expert model in this embodiment is located in a distributed cluster and is used to process the expert network included in the distributed cluster.

[0084] The first selection unit 501 can first obtain the first number of target tokens processed by each expert network in the hybrid expert model, and then select at least one target expert network from the multiple expert networks based on the first number of target tokens processed by the multiple expert networks.

[0085] The first selection unit 501 can obtain the first processing number of target tokens processed by each expert network according to the output of the gating network in the hybrid expert model. The first processing number is the total number of tokens to be processed by each expert network. The gating network is used to determine the expert network to process each sequence unit according to the feature vector of each sequence unit in the input sequence.

[0086] Specifically, when the first selection unit 501 selects a target expert network from multiple expert networks based on the first processing quantity of the target token processed by the multiple expert networks in the hybrid expert model, the following implementation method can be adopted: sort the multiple expert networks in descending order of the first processing quantity; select the top N expert networks as the target expert network based on the sorting result, where N is a positive integer greater than or equal to 1.

[0087] In other words, the first selection unit 501 selects the target expert network by sorting multiple expert networks according to the first processing quantity, ensuring that at least one selected target expert network needs to process a large number of tokens, thereby improving the selection accuracy of the target expert network.

[0088] In order to maximize the utilization of the computing power provided by the CPU and to balance the GPU load as much as possible, the first selection unit 501 may also perform the following: obtain the number of candidate CPUs, which is the total number of CPUs included in the distributed cluster; use the obtained number of CPUs as the value of N; for example, if the number of CPUs is 32, the first selection unit 501 will select the top 32 expert networks as the target expert networks.

[0089] In addition, when the first selection unit 501 selects a target expert network from multiple expert networks based on the first processing quantity of the target tokens processed by the multiple expert networks in the hybrid expert model, it can also adopt the following method: based on the first processing quantity of the target tokens processed by the multiple expert networks and the target GPUs corresponding to the multiple expert networks, the load quantity of each target GPU is obtained. In this embodiment, the load quantity is the total number of tokens to be processed by the target GPU; the target GPUs whose load quantity meets the preset requirements are selected as the GPUs to be processed. For example, the target GPUs with the top N load quantities are selected as the GPUs to be processed, or the target GPUs whose load quantity exceeds the preset threshold are selected as the GPUs to be processed; the target expert network is selected from the expert networks corresponding to the GPUs to be processed.

[0090] In other words, the first selection unit 501 first selects a GPU to be processed from multiple target GPUs based on the first processing quantity corresponding to the expert network and the target GPU, and then selects a target expert network from the expert networks corresponding to the GPU to be processed, so that the selected target expert network corresponds to the GPU with a high load, thereby effectively reducing the load of the GPU with a high load and avoiding the problem of slow processing speed caused by excessive occupation of the computing power of the GPU with a high load.

[0091] It is understood that if N in this embodiment is the number of candidate CPUs, then the first selection unit 501 can select the expert network with a higher processing volume among the GPUs to be processed as the target expert network, so as to ensure that there are enough CPUs to process the target token corresponding to the selected target expert network.

[0092] In this embodiment, after the first selection unit 501 selects the target expert network from multiple expert networks, the processing unit 502 obtains the second processing quantity of the corresponding target expert network based on the subtraction result between the first processing quantity and the target processing quantity.

[0093] In this embodiment, the target processing number can be preset. The preset target processing number is the number of tokens processed when the expert network has good performance, as obtained through actual testing of the GPU.

[0094] When obtaining the target processing quantity, the processing unit 502 may also use the following method: obtain the first processing quantity of the target token to be processed by the expert network ranked N+1, and use it as the target processing quantity; for example, when N is 32, this embodiment obtains the first processing quantity of the expert network ranked 33 as the target processing quantity.

[0095] In other words, after sorting the expert networks according to the first processing quantity, the processing unit 502 can also obtain the target processing quantity in real time according to the sorting result, thereby simplifying the steps to obtain the target processing quantity, improving the speed of obtaining the target processing quantity, and using the first processing quantity of the expert network ranked N+1 as the target processing quantity, it also effectively avoids the problem of not being able to obtain the second processing quantity based on the first processing quantity of the target expert network and the target processing quantity, thus improving the accuracy of obtaining the target processing quantity.

[0096] In addition, the processing unit 502 can also obtain the average value among the first processing quantities of the expert networks after the Nth position as the target processing quantity.

[0097] If the first selection unit 501 selects multiple target expert networks, the processing unit 502 can obtain the second processing number for each target expert network based on the difference between the first processing number of the target expert network and the obtained target processing number.

[0098] In this embodiment, after the processing unit 502 obtains the second processing quantity of the corresponding target expert network, the second selection unit 503 selects the target CPU of the corresponding target expert network from multiple candidate CPUs according to the second processing quantity; wherein, each candidate CPU in this embodiment includes the model parameters of all expert networks.

[0099] The second selection unit 503 can first sort the multiple target expert networks in descending order of the second processing quantity, and then, according to the sorting result, assign a candidate CPU to each target expert network in turn, so that the assigned candidate CPU is used as the target CPU for each target expert network.

[0100] In other words, the second selection unit 503 prioritizes allocating candidate CPUs to target expert networks with a higher second processing quantity, so as to ensure that the candidate CPUs can be used to process some tokens corresponding to these target expert networks, thereby reducing the load on the GPUs where these target expert networks are located and achieving the purpose of load balancing between different GPUs.

[0101] In this embodiment, a candidate CPU is only used to process a portion of the target tokens corresponding to a target expert network, that is, a candidate CPU can only play the role of a target expert network.

[0102] When the second selection unit 503 selects the target CPU of the corresponding target expert network from multiple candidate CPUs according to the second processing quantity, it may also adopt the following implementation method: obtain the third processing quantity of the corresponding candidate CPU; determine the target CPU quantity of the corresponding target expert network according to the second processing quantity and the third processing quantity; select the candidate CPUs corresponding to the target CPU quantity as the target CPU of the corresponding target expert network.

[0103] In other words, the second selection unit 503 fully considers the token processing capabilities of the candidate CPUs. It selects the target CPU based on the second processing quantity of the corresponding target expert network and the third processing quantity of the corresponding candidate CPU. This ensures that at least one selected target CPU can meet the token processing requirements while avoiding the problem that the second processing quantity exceeds the processing capability of a single target CPU, thus improving the accuracy of target CPU selection.

[0104] In this embodiment, after the second selection unit 503 selects the target CPU corresponding to the target expert network from multiple candidate CPUs, the running unit 504 uses the target GPU of the corresponding target expert network to process the first target token corresponding to the target processing quantity, and uses the target CPU of the corresponding target expert network to process the second target token corresponding to the second processing quantity.

[0105] In other words, after determining the target CPU corresponding to the target expert network, this embodiment can divide the target token to be processed by the target expert network into a first target token corresponding to the target processing quantity and a second target token corresponding to the second processing quantity, so that the target GPU processes the first target token and the target CPU processes the second target token.

[0106] The running unit 504 can first randomly select a token corresponding to the target processing quantity from the target tokens to be processed by the target expert network as the first target token, and then use the remaining tokens in the target tokens as the second target token corresponding to the second processing quantity.

[0107] When the execution unit 504 processes the first target token corresponding to the target processing quantity using the target GPU of the corresponding target expert network, it can first determine the target GPU corresponding to the target expert network, for example, by determining it according to the correspondence table between expert networks and GPUs. Then, it sends the identification information of the target expert network and the first target token to the target GPU, so that the target GPU can run the model parameters of the corresponding target expert network to process the first target token according to the identification information.

[0108] When the execution unit 504 executes S104 and uses the target CPU of the corresponding target expert network to process the second target token corresponding to the second processing quantity, it can send the identification information of the target expert network and the second target token to the target CPU, so that the target CPU can run the model parameters of the corresponding target expert network to process the second target token according to the identification information.

[0109] Since the candidate CPUs pre-store the model parameters of all expert networks, the target CPU corresponding to the target expert network does not need to copy the model parameters, thus avoiding inference pauses caused by the copying process.

[0110] Understandably, if multiple target CPUs corresponding to the target expert network are determined, the running unit 504 can also determine the number of second target tokens that each target CPU can process (for example, based on the third processing capacity of the target CPU), and then send the identification information of the target expert network and the corresponding number of second target tokens to different target CPUs so that the different target CPUs can process the second target tokens.

[0111] It is understandable that for non-target expert networks in expert networks, or target expert networks that cannot determine the target CPU, the running unit 504 can directly send the target token to be processed and the identification information to the corresponding target GPU, so that the target GPU can process the received target token.

[0112] In this embodiment, after the running unit 504 processes the first target token using the target GPU and the second target token using the target CPU, the output unit 505 obtains the target processing result output by the target expert network when processing the target token based on the processing result output by the target GPU and the processing result output by the target CPU.

[0113] The output unit 505 can concatenate the processing results output by the target GPU and the processing results output by the target CPU according to the order of the target token in the input sequence, so as to obtain the target processing result output by the target expert network when processing the target token based on the concatenation result.

[0114] The acquisition, storage, and application of user personal information involved in the technical solution disclosed herein comply with the provisions of relevant laws and regulations and do not violate public order and good morals.

[0115] According to embodiments of this disclosure, this disclosure also provides an electronic device, a readable storage medium, and a computer program product.

[0116] like Figure 6 The diagram shown is a block diagram of an electronic device according to an embodiment of the expert network processing method based on a hybrid expert model, as described in this disclosure. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0117] like Figure 6As shown, device 600 includes a computing unit 601, which can perform various appropriate actions and processes based on a computer program stored in read-only memory (ROM) 602 or a computer program loaded from storage unit 608 into random access memory (RAM) 603. RAM 603 may also store various programs and data required for the operation of device 600. The computing unit 601, ROM 602, and RAM 603 are interconnected via bus 604. Input / output (I / O) interface 605 is also connected to bus 604.

[0118] Multiple components in device 600 are connected to I / O interface 605, including: input unit 606, such as keyboard, mouse, etc.; output unit 607, such as various types of displays, speakers, etc.; storage unit 608, such as disk, optical disk, etc.; and communication unit 609, such as network card, modem, wireless transceiver, etc. Communication unit 609 allows device 600 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.

[0119] The computing unit 601 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 601 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 601 performs the various methods and processes described above, such as the expert network processing method based on a hybrid expert model. For example, in some embodiments, the expert network processing method based on a hybrid expert model can be implemented as a computer software program that is tangibly contained in a machine-readable medium, such as storage unit 608.

[0120] In some embodiments, part or all of the computer program may be loaded and / or installed on the device 600 via ROM 602 and / or communication unit 609. When the computer program is loaded into RAM 603 and executed by computing unit 601, one or more steps of the expert network processing method based on a hybrid expert model described above may be performed. Alternatively, in other embodiments, computing unit 601 may be configured to perform the expert network processing method based on a hybrid expert model by any other suitable means (e.g., by means of firmware).

[0121] Various implementations of the systems and techniques described herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various implementations may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transferring data and instructions to the storage system, the at least one input device, and the at least one output device.

[0122] The program code used to implement the methods of this disclosure may be written in any combination of one or more programming languages. This program code may be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable expert network processing device based on a hybrid expert model, such that when executed by the processor or controller, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.

[0123] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0124] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for showing information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).

[0125] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as a data server), or computing systems that include middleware components (e.g., an application server), or computing systems that include frontend components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with embodiments of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., a communication network). Examples of communication networks include local area networks (LANs), wide area networks (WANs), and the Internet.

[0126] Computer systems can include clients and servers. Clients and servers are generally geographically separated and typically interact via communication networks. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. A server can be a cloud server, also known as a cloud computing server or cloud host, a hosting product within the cloud computing service system that addresses the management difficulties and weak business scalability inherent in traditional physical hosts and VPS (Virtual Private Server) services. Servers can also be servers for distributed systems or servers incorporating blockchain technology.

[0127] It should be understood that the various forms of processes shown above can be used to rearrange, add, or delete steps. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution disclosed in this disclosure can be achieved, and this is not limited herein.

[0128] The specific embodiments described above do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure should be included within the scope of protection of this disclosure.

Claims

1. An expert network processing method based on a hybrid expert model, comprising: Based on the first number of target sequence unit tokens processed by multiple expert networks in the hybrid expert model, a target expert network is selected from the multiple expert networks; The second processing number corresponding to the target expert network is obtained by subtracting the first processing number from the target processing number. Based on the second processing quantity, a target CPU corresponding to the target expert network is selected from multiple candidate CPUs, wherein the candidate CPUs include the model parameters of all expert networks; The first target token corresponding to the target processing quantity is processed using the target GPU corresponding to the target expert network, and the second target token corresponding to the second processing quantity is processed using the target CPU. Based on the processing results output by the target GPU and the processing results output by the target CPU, the target processing result output by the target expert network when processing the target token is obtained.

2. The method according to claim 1, wherein, The step of selecting a target expert network from the multiple expert networks based on the first number of target sequence unit tokens processed by the multiple expert networks in the hybrid expert model includes: The multiple expert networks are sorted in descending order of the number of processes processed. Based on the ranking results, the top N expert networks are selected as the target expert network, where N is a positive integer greater than or equal to 1.

3. The method according to claim 2, further comprising: Get the number of candidate CPUs; The number of CPUs is used as the value of N.

4. The method according to claim 1, wherein, The step of selecting a target expert network from the multiple expert networks based on the first number of target sequence unit tokens processed by the multiple expert networks in the hybrid expert model includes: The load on each target GPU is obtained based on the first number of target tokens processed by the multiple expert networks and the target GPUs corresponding to the multiple expert networks. Select the target GPUs whose workload meets the preset requirements as the GPUs to be processed; The target expert network is selected from the expert networks corresponding to the GPU to be processed.

5. The method according to claim 2, wherein, Obtaining the target processing quantity includes: Obtain the first number of target tokens to be processed by the expert network ranked N+1, and use it as the target processing number.

6. The method according to claim 1, wherein, The step of selecting the target CPU corresponding to the target expert network from multiple candidate CPUs based on the second processing quantity includes: The multiple target expert networks are sorted in descending order of the number of second processes; Based on the ranking results, a candidate CPU is assigned to each target expert network in turn, and the assigned candidate CPU is used as the target CPU for each target expert network.

7. The method according to claim 1, wherein, The step of selecting the target CPU corresponding to the target expert network from multiple candidate CPUs based on the second processing quantity includes: Obtain the third processing quantity corresponding to the candidate CPU; The target number of CPUs corresponding to the target expert network is determined based on the second processing quantity and the third processing quantity. Select candidate CPUs corresponding to the target number of CPUs as the target CPUs for the target expert network.

8. The method according to claim 7, wherein, Determining the target number of CPUs corresponding to the target expert network based on the second processing quantity and the third processing quantity includes: The multiple target expert networks are sorted in descending order of the number of second processes; Based on the sorting results, the target CPU count for each target expert network is determined sequentially according to the second processing count and the third processing count for each target expert network.

9. The method according to claim 1, wherein, The first target token, which uses the target GPU corresponding to the target expert network to process the target number of processes, includes: Determine the target GPU corresponding to the target expert network; The identification information of the target expert network and the first target token are sent to the target GPU, so that the target GPU processes the first target token by running the model parameters corresponding to the target expert network according to the identification information.

10. The method according to claim 1, wherein, The step of using the target CPU to process the second target token corresponding to the second processing quantity includes: The identification information of the target expert network and the second target token are sent to the target CPU, so that the target CPU processes the second target token by running the model parameters corresponding to the target expert network according to the identification information.

11. The method according to claim 1, wherein, Determining the first target token and the second target token includes: From the target tokens to be processed by the target expert network, a token corresponding to the target processing quantity is randomly selected as the first target token; The remaining tokens in the target token are used as the second target token corresponding to the second processing quantity.

12. An expert network processing device based on a hybrid expert model, comprising: The first selection unit is used to select a target expert network from the multiple expert networks based on the first processing quantity of the target sequence unit token processed by the multiple expert networks in the hybrid expert model. The processing unit is configured to obtain a second processing quantity corresponding to the target expert network based on the subtraction result between the first processing quantity and the target processing quantity; The second selection unit is used to select a target CPU corresponding to the target expert network from a plurality of candidate CPUs according to the second processing quantity, wherein the candidate CPUs include the model parameters of all expert networks. The execution unit is configured to process a first target token corresponding to the target processing quantity using a target GPU corresponding to the target expert network, and to process a second target token corresponding to the second processing quantity using the target CPU; The output unit is used to obtain the target processing result output by the target expert network when processing the target token, based on the processing result output by the target GPU and the processing result output by the target CPU.

13. An electronic device, comprising: At least one processor; as well as A memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the method of any one of claims 1-11.

14. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to perform the method according to any one of claims 1-11.

15. A computer program product comprising a computer program that, when executed by a processor, implements the method according to any one of claims 1-11.

Citation Information

Cited By

  • Data processing method, data processing device, electronic equipment and storage medium

    CN121210152A

  • Model scheduling method and device, computer equipment and storage medium

    CN121523863A