Expert parallel computing method and computer program product for hybrid expert model

By dynamically allocating video memory in the hybrid expert model and controlling the parallel processing of expert sub-networks on the computing card, the problem of excessive video memory usage is solved, and the model inference efficiency and maximum input number are improved.

CN120493998BActive Publication Date: 2025-09-12INSPUR SUZHOU INTELLIGENT TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510965269.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-14
Publication Date
2025-09-12
Estimated Expiration
2045-07-14

AI Technical Summary

Technical Problem

In a multi-machine, multi-GPU scenario, the hybrid expert model occupies too much video memory, resulting in a reduction in the maximum number of inputs supported by inference, which in turn reduces model inference efficiency.

Method used

By obtaining the input data blocks processed on each computing card, dynamically allocating the minimum target memory space, and controlling multiple expert sub-networks on the computing card to process the input data in parallel, including allocating memory space to the computing card based on the input fragments, and ensuring the correctness of data routing through expert identification sequence and storage address mapping.

Benefits of technology

The maximum number of inputs supported during model inference has been increased, which improves the model's inference efficiency, reduces graphics memory usage, and improves the parallel computing efficiency of computing cards.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120493998B_ABST
    Figure CN120493998B_ABST
Patent Text Reader

Abstract

The present application discloses an expert parallel computing method and computer program product for a hybrid expert model, belonging to the field of artificial intelligence technology. The method comprises: obtaining the storage address of each input data block in a global sequence based on the input data blocks processed by each expert subnetwork and the expert identification sequences corresponding to the multiple expert subnetworks; obtaining the input fragments processed by all expert subnetworks deployed on each computing card based on the number of occurrences of each expert subnetwork in the expert identification sequence and the storage address corresponding to each input data block; allocating target video memory space to each computing card based on the input fragments corresponding to each computing card; controlling each expert subnetwork deployed on each computing card to read the input data block to be processed from the target video memory space corresponding to each computing card, and calculating the input data block to be processed to obtain sub-result data output by each computing card; and fusing at least one sub-result data to obtain target result data output by the hybrid expert model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application belongs to the field of artificial intelligence technology, and in particular relates to an expert parallel computing method and computer program product of a hybrid expert model. Background Art

[0002] The Mixture of Experts (MoE) model is a special neural network architecture designed to improve model flexibility and efficiency by combining multiple "expert" sub-networks. Each "expert" is typically a small, independent neural network or sub-module that specializes in processing different parts or patterns of the input data. In related technologies, such as in multi-machine and multi-GPU scenarios, the storage space allocated to each card for the input data blocks (tokens) to be processed exceeds the space occupied by the tokens actually processed by the experts on the current card. This results in excessive video memory capacity being occupied in the expert parallel mode, which reduces the maximum number of inputs (batch) supported for inference and, in turn, reduces model inference efficiency. Summary of the Invention

[0003] This application aims to address at least one of the technical problems existing in the related art. To this end, this application proposes an expert parallel computing method and computer program product for a hybrid expert model, which solves the problem of excessive video memory usage in the expert parallel mode and increases the maximum number of inputs supported during inference, thereby improving the inference efficiency of the model.

[0004] In a first aspect, the present application provides an expert parallel computing method for a hybrid expert model, wherein the hybrid expert model includes multiple expert sub-networks, and the multiple expert sub-networks are deployed on at least one computing card; the method includes:

[0005] Based on the input data blocks processed by each of the expert sub-networks and the expert identification sequences corresponding to the multiple expert sub-networks, obtaining the storage address of each of the input data blocks in the global sequence; the multiple input data blocks correspond one-to-one to the multiple expert sub-networks;

[0006] Based on the number of occurrences of each expert sub-network in the expert identification sequence and the storage address corresponding to each input data block, obtaining input segments processed by all expert sub-networks deployed on each computing card; the input segments include multiple input data blocks;

[0007] Allocating target video memory space to each computing card based on the input fragments corresponding to each computing card;

[0008] Controlling each expert sub-network deployed on each computing card, reading an input data block to be processed from a target memory space corresponding to each computing card, and performing calculations on the input data block to be processed to obtain sub-result data output by each computing card;

[0009] At least one of the sub-result data is fused to obtain target result data output by the hybrid expert model.

[0010] According to the expert parallel computing method of the hybrid expert model provided in the embodiment of the present application, by obtaining all input data blocks (input fragments) processed on each computing card, the minimum target video memory space is dynamically allocated to it according to the input fragment corresponding to the computing card, and multiple expert sub-networks on the computing card are controlled to process the input data in parallel, thereby solving the problem of excessive video memory usage in the expert parallel mode, increasing the maximum number of inputs supported during inference, and thus improving the inference efficiency of the model.

[0011] In an embodiment of the present application, an expert parallel computing method for a hybrid expert model allocates target video memory space to each computing card based on the input fragments corresponding to each computing card, including:

[0012] Based on the input fragments corresponding to the computing cards, a target number of input data blocks to be processed corresponding to the computing cards is obtained;

[0013] Based on the target number corresponding to each computing card, the index information corresponding to each input data block in the global space is mapped to the local space corresponding to each computing card;

[0014] Based on the index information of the input data block in the local space corresponding to each computing card, a target video memory space is allocated to each computing card.

[0015] An expert parallel computing method for a hybrid expert model according to an embodiment of the present application, mapping the index information corresponding to each input data block in the global space to the local space corresponding to each computing card based on the target number corresponding to each computing card, includes:

[0016] Based on the target number corresponding to the computing card, the index information of each input data block in the global sequence is mapped to the local space corresponding to the computing card; and the index information of each input data block in the input segment corresponding to the computing card is mapped to the local space corresponding to the computing card.

[0017] An expert parallel computing method for a hybrid expert model according to an embodiment of the present application, mapping the index information corresponding to each input data block in the global space to the local space corresponding to each computing card based on the target number corresponding to each computing card, includes:

[0018] Based on the target number corresponding to each computing card, obtaining the intra-card starting offset corresponding to each computing card;

[0019] The index information corresponding to the input data block corresponding to each computing card in the global space is subtracted from the starting offset within the card to obtain the index information corresponding to the input data block corresponding to each computing card in the local space.

[0020] In an embodiment of the present application, an expert parallel computing method for a hybrid expert model, obtaining input segments processed by all expert subnetworks deployed on each computing card based on the number of occurrences of each expert subnetwork in the expert identification sequence and the storage address corresponding to each input data block, includes:

[0021] Based on the number of occurrences of each of the expert sub-networks in the expert identification sequence, an interval tag array is generated; the interval tag array is used to represent the storage addresses of all input data blocks processed by each of the expert sub-networks in the global sequence;

[0022] Based on identification information corresponding to each expert sub-network deployed on the computing card, an input segment corresponding to the computing card is obtained from the interval tag array.

[0023] In an expert parallel computing method for a hybrid expert model according to an embodiment of the present application, obtaining a storage address of each input data block in a global sequence based on the input data blocks processed by each expert sub-network and expert identification sequences corresponding to the multiple expert sub-networks includes:

[0024] Based on the expert sub-networks corresponding to the original input data blocks, obtaining the copy input data blocks corresponding to the expert sub-networks;

[0025] Arranging the copy input data blocks corresponding to each of the expert sub-networks based on the identification information corresponding to each of the expert sub-networks to obtain the global sequence;

[0026] Based on the arrangement position information of each expert sub-network corresponding to the original input data block in the expert identification sequence, the storage address of each copy input data block in the global sequence is obtained.

[0027] An expert parallel computing method for a hybrid expert model according to an embodiment of the present application, for obtaining expert identification sequences corresponding to the multiple expert sub-networks, includes:

[0028] Obtaining expert identification matrices corresponding to all expert sub-networks for processing the plurality of input data blocks;

[0029] Based on the identification information corresponding to each expert sub-network in the expert identification matrix, the identification information corresponding to the multiple expert sub-networks is sorted in ascending order of the identification information to obtain the expert identification sequence.

[0030] In an expert parallel computing method for a hybrid expert model according to an embodiment of the present application, obtaining expert identification matrices corresponding to a plurality of expert sub-networks for processing a plurality of input data blocks includes:

[0031] The expert identification matrix is ​​constructed based on the number of the multiple input data blocks and the number of expert sub-networks required for processing each of the input data blocks; the first dimension of the expert identification matrix corresponds to the number of the multiple input data blocks, and the second dimension of the expert identification matrix corresponds to the number of expert sub-networks required for each of the input data blocks.

[0032] In an embodiment of the present application, the expert parallel computing method of the hybrid expert model, wherein the computing of the input data block to be processed to obtain sub-result data output by each computing card, includes:

[0033] Performing a gating operation, an upper projection operation, an activation function operation, and a lower projection operation on the input data block to be processed to obtain the sub-result data output by each of the computing cards.

[0034] In an embodiment of the present application, an expert parallel computing method for a hybrid expert model controls each expert sub-network deployed on each computing card, reads an input data block to be processed from a target memory space corresponding to each computing card, and computes the input data block to be processed to obtain sub-result data output by each computing card, including:

[0035] Controlling each expert sub-network deployed on each computing card, reading an input data block to be processed from the target video memory space, and performing the gating operation and the up-projection operation on the input data block to be processed to obtain first-stage result data;

[0036] Performing the activation function operation on the first-stage result data to obtain the second-stage result data;

[0037] The lower projection operation is performed on the second-stage result data to obtain the sub-result data output by each of the computing cards.

[0038] In an expert parallel computing method of a hybrid expert model according to an embodiment of the present application, performing the down-projection operation on the second-stage result data to obtain the sub-result data output by each computing card includes:

[0039] Controlling each expert sub-network deployed on each of the computing cards, reading the input data block to be processed from the second-stage result data, and performing the down-projection operation on the second-stage result data corresponding to the input data block to be processed to obtain the sub-result data output by each of the computing cards.

[0040] In an embodiment of the present application, the expert parallel computing method of the hybrid expert model, wherein the fusion processing of at least one sub-result data to obtain the target result data output by the hybrid expert model, includes:

[0041] Performing splicing processing on at least one of the sub-result data to obtain first result data; the first result data includes result data output by each expert sub-network corresponding to each input data block;

[0042] Based on the weight information corresponding to each of the expert sub-networks, weighted fusion processing is performed on the result data output by each of the expert sub-networks in the first result data to obtain the target result data.

[0043] In a second aspect, the present application provides a computer program product, wherein the hybrid expert model includes multiple expert sub-networks, and the multiple expert sub-networks are deployed on at least one computing card; including:

[0044] a first processing module, configured to obtain a storage address of each input data block in a global sequence based on the input data blocks processed by each of the expert sub-networks and the expert identification sequences corresponding to the multiple expert sub-networks; wherein the multiple input data blocks correspond one-to-one to the multiple expert sub-networks;

[0045] a second processing module, configured to obtain input segments processed by all expert subnetworks deployed on each of the computing cards based on the number of occurrences of each of the expert subnetworks in the expert identification sequence and the storage address corresponding to each of the input data blocks; the input segments comprising a plurality of input data blocks;

[0046] A third processing module is configured to allocate target video memory space to each computing card based on the input fragments corresponding to each computing card;

[0047] a fourth processing module, configured to control each expert sub-network deployed on each computing card, read an input data block to be processed from a target video memory space corresponding to each computing card, and perform calculations on the input data block to be processed to obtain sub-result data output by each computing card;

[0048] The fifth processing module is used to perform fusion processing on at least one of the sub-result data to obtain the target result data output by the hybrid expert model.

[0049] According to the computer program product provided in the embodiment of the present application, by obtaining all input data blocks (input fragments) processed on each computing card, the minimum target video memory space is dynamically allocated to it according to the input fragments corresponding to the computing card, and multiple expert sub-networks on the computing card are controlled to process the input data in parallel, thereby solving the problem of excessive video memory usage in the expert parallel mode, increasing the maximum number of inputs supported during inference, and thus improving the inference efficiency of the model.

[0050] In a third aspect, the present application provides an electronic device comprising a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the expert parallel computing method of the hybrid expert model as described in the first aspect above is implemented.

[0051] In a fourth aspect, the present application provides a non-volatile computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the expert parallel computing method of the hybrid expert model as described in the first aspect above.

[0052] The above one or more technical solutions in the embodiments of the present application have at least one of the following technical effects:

[0053] By obtaining all input data blocks (input fragments) processed on each computing card, the minimum target video memory space is dynamically allocated to it according to the corresponding input fragment of the computing card, and multiple expert sub-networks on the computing card are controlled to process the input data in parallel. This solves the problem of excessive video memory usage in expert parallel mode, increases the maximum number of inputs supported during inference, and thus improves the inference efficiency of the model.

[0054] Furthermore, based on the number of all tokens processed by each computing card, the starting offset within the card corresponding to each computing card is obtained. Then, based on the starting offset within the card, the index information of the token corresponding to each computing card in the global space is mapped to the local space. This allows the minimum video memory space to be dynamically allocated according to the number of tokens processed by the computing card, which can solve the problem of excessive video memory usage in expert parallel mode, thereby increasing the maximum number of inputs supported during model inference.

[0055] Furthermore, by using the arrangement position information of each expert sub-network corresponding to the original input data block in the expert identification sequence, the storage address of each copy input data block in the global sequence is obtained, and the expert sub-network and the token it processes can be matched, which makes it easier for each expert sub-network to obtain the token it processes, ensuring the correctness of data routing, enabling each expert sub-network to calculate in parallel, and improving the efficiency of model reasoning.

[0056] Additional aspects and advantages of the present application will be given in part in the description below, and in part will become obvious from the description below, or will be learned through practice of the present application. BRIEF DESCRIPTION OF THE DRAWINGS

[0057] The above and / or additional aspects and advantages of the present application will become apparent and easily understood from the description of the embodiments in conjunction with the following drawings, in which:

[0058] Figure 1 This is one of the flow charts of the expert parallel computing method of the hybrid expert model provided in the embodiment of the present application;

[0059] Figure 2 This is the second flow chart of the expert parallel computing method of the hybrid expert model provided in the embodiment of the present application;

[0060] Figure 3 This is one of the principle diagrams of the expert parallel computing method of the hybrid expert model provided in the embodiment of the present application;

[0061] Figure 4 This is the second schematic diagram of the principle of the expert parallel computing method of the hybrid expert model provided in the embodiment of the present application;

[0062] Figure 5 is a schematic diagram of the structure of a computer program product provided in an embodiment of the present application;

[0063] Figure 6 It is a structural diagram of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0064] The following will be combined with the accompanying drawings in the embodiments of the present application to clearly describe the technical solutions in the embodiments of the present application. Obviously, the embodiments described are part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field are within the scope of protection of this application.

[0065] The terms "first," "second," and the like in the specification and claims of this application are used to distinguish similar objects, and are not used to describe a specific order or precedence. It should be understood that the terms used in this manner are interchangeable where appropriate, so that the embodiments of this application can be implemented in an order other than that illustrated or described herein, and that the objects distinguished by "first," "second," and the like are generally of the same type, and do not limit the number of objects; for example, the first object can be one or more. In addition, the term "and / or" in the specification and claims refers to at least one of the connected objects, and the character " / " generally indicates that the objects connected are in an "or" relationship.

[0066] The expert parallel computing method, computer program product, electronic device and readable storage medium of the hybrid expert model provided by the embodiments of the present application are described in detail below with reference to specific embodiments and their application scenarios in conjunction with the accompanying drawings.

[0067] The expert parallel computing method of the hybrid expert model can be applied to a terminal, and can be specifically executed by hardware or software in the terminal.

[0068] The terminal includes, but is not limited to, a portable communication device such as a mobile phone or tablet computer having a touch-sensitive surface (e.g., a touch screen display and / or a touch pad). It should also be understood that in some embodiments, the terminal may not be a portable communication device, but rather a desktop computer having a touch-sensitive surface (e.g., a touch screen display and / or a touch pad).

[0069] In the following embodiments, a terminal including a display and a touch-sensitive surface is described. However, it should be understood that the terminal may include one or more other physical user interface devices such as a physical keyboard, a mouse, and a joystick.

[0070] The embodiment of the present application provides an expert parallel computing method for a hybrid expert model. The execution subject of the expert parallel computing method for the hybrid expert model can be an electronic device or a functional module or functional entity in the electronic device that can implement the expert parallel computing method for the hybrid expert model. The electronic devices mentioned in the embodiment of the present application include but are not limited to mobile phones, tablet computers, computers, cameras, and wearable devices. The expert parallel computing method for the hybrid expert model provided in the embodiment of the present application is described below using an electronic device as an example of the execution subject.

[0071] like Figure 1 As shown, the expert parallel computing method of the hybrid expert model includes: step 110, step 120, step 130, step 140 and step 150.

[0072] It should be noted that the Mixture of Experts (MoE) model can improve the flexibility and efficiency of the model by combining multiple "expert" sub-networks.

[0073] The mixture of experts model can dynamically select which experts should respond to a given input through a gating network, thereby achieving efficient resource utilization and performance improvement.

[0074] Each "expert" is usually an independent small neural network or sub-module. Each expert sub-network can be used to specifically process different parts or patterns of the input data. The expert sub-network can be a fully connected layer, convolutional layer or other types of neural network structures.

[0075] The hybrid expert model can activate only some experts through conditional calculation to reduce resource consumption; different experts can focus on different types of tasks or data distributions, so that the entire model can better adapt to complex and changing data environments; as new tasks or new fields emerge, the model's capabilities can be expanded by simply adding new experts without retraining the entire system.

[0076] In a multi-machine, multi-card scenario, multiple servers and multiple computing cards on each server can work together to complete model training or inference tasks.

[0077] Multiple expert sub-networks can be deployed on at least one computing card.

[0078] Step 110: Based on the input data blocks processed by each expert sub-network and the expert identification sequences corresponding to the multiple expert sub-networks, obtain the storage address of each input data block in the global sequence;

[0079] In this step, multiple expert sub-networks can be assigned to process each original input data block (token), and for each original input data block, a copy of the input data block can be made for each expert sub-network participating in the calculation.

[0080] Multiple (copy) input data blocks correspond one-to-one to multiple expert sub-networks.

[0081] Each expert sub-network has corresponding identification information. The identification information corresponding to multiple expert sub-networks can be sorted, and the sorting results can be saved as an expert identification sequence.

[0082] The input data blocks processed by each expert can be arranged in sequence according to the expert identification sequence to obtain a global sequence, thereby obtaining the number (storage address) of the copy of each original input data block in the global sequence.

[0083] Step 120: Based on the number of occurrences of each expert sub-network in the expert identification sequence and the storage address corresponding to each input data block, obtain the input segments processed by all expert sub-networks deployed on each computing card;

[0084] In this step, different expert sub-networks correspond to different identification information. Based on the identification information corresponding to the i-th expert sub-network, the number of occurrences of the i-th expert sub-network in the expert identification sequence can be determined.

[0085] Each expert sub-network can process a portion of the tokens in the global sequence, and the tokens and the number of tokens processed by each expert sub-network are also different.

[0086] Multiple expert sub-networks can be deployed on each computing card. Based on the number of appearances of each expert sub-network in the expert identification sequence and the storage address corresponding to each input data block, the storage address of the first token processed by each expert sub-network in the global sequence and the number of tokens processed by the expert sub-network can be determined. In this way, all tokens processed by all expert sub-networks deployed on each computing card, that is, the input fragments, can be obtained.

[0087] A sequence seg_indptr can be designed. There are N+1 elements in seg_indptr (dimension is [1, N+1]). The left-closed and right-open interval [seg(i), seg(i+1)) composed of the i-th element and the i+1-th element is the number (storage address) of all tokens processed by expert sub-network i in the complete token sequence (global sequence). If seg(i)=seg(i+1), expert i does not participate in the calculation.

[0088] By slicing seg_indptr, we can get all the tokens processed by the expert sub-network on the current computing card, that is, the input fragment, which includes multiple input data blocks.

[0089] Step 130: Allocate target video memory space to each computing card based on the input fragments corresponding to each computing card;

[0090] In this step, the local space of the computing card that each expert sub-network can access is not the complete seq_len*top_k tokens (where seq_len is the number of tokens to be processed and top_k is the number of expert sub-networks corresponding to each token), but cur_num_token tokens (where cur_num_token is the total number of tokens on the current computing card). The global sequence and seg_indptr are designed based on the accessible storage space of seq_len*top_k, and need to be adjusted. For example, the global sequence and seg_indptr can be mapped to the local space of the computing card to obtain the target video memory space corresponding to each computing card.

[0091] Based on the number of tokens currently processed on the card, the minimum space required for input to the local space of the computing card, that is, the target video memory space, can be allocated to save video memory.

[0092] After allocating target video memory space to each computing card, each computing card can only process the input data blocks processed by the experts on the current computing card.

[0093] Step 140: Control each expert sub-network deployed on each computing card to read the input data block to be processed from the target memory space corresponding to each computing card, and perform calculations on the input data block to be processed to obtain sub-result data output by each computing card;

[0094] In this step, the input data block to be processed can be read from the target video memory space corresponding to the computing card according to the index information of the input fragment corresponding to the computing card mapped to the local space.

[0095] Then control the parallel computing of each expert sub-network deployed on the computing card.

[0096] Each expert sub-network calculates the input data block that needs to be processed and obtains the sub-result data output by the calculation card.

[0097] Step 150: Perform fusion processing on at least one sub-result data to obtain target result data output by the hybrid expert model.

[0098] In this step, the sub-result data on each card is part of the complete output result. The sub-result data output by each calculation card can be fused to obtain the final target result data output by the hybrid expert model.

[0099] According to the expert parallel computing method of the hybrid expert model provided in the embodiment of the present application, by obtaining all input data blocks (input fragments) processed on each computing card, the minimum target video memory space is dynamically allocated to it according to the input fragment corresponding to the computing card, and multiple expert sub-networks on the computing card are controlled to process the input data in parallel, thereby solving the problem of excessive video memory usage in the expert parallel mode, increasing the maximum number of inputs supported during inference, and thus improving the inference efficiency of the model.

[0100] In some embodiments, step 130 may include:

[0101] Based on the input fragments corresponding to each computing card, obtain the target number of input data blocks to be processed by each computing card;

[0102] Based on the target number corresponding to each computing card, the index information corresponding to each input data block in the global space is mapped to the local space corresponding to each computing card;

[0103] Based on the index information of the input data block in the local space corresponding to each computing card, the target video memory space is allocated to each computing card.

[0104] In this embodiment, the input fragment includes the storage address (i.e., number) of the first input data block processed by the computing card in the global sequence, and the storage address of the last input data block processed by the computing card in the global sequence. Based on the number range of the input data blocks processed by the current computing card, the total number (target number) of input data blocks processed on the current computing card can be obtained.

[0105] Then, according to the target number corresponding to each computing card, the index information corresponding to each input data block in the global space is mapped to the local space corresponding to each computing card, so as to adjust the index information corresponding to each input data block to the address range accessible to the current computing card.

[0106] Based on the index information of the input data block in the local space corresponding to each computing card, the target video memory space can be allocated to each computing card.

[0107] In some embodiments, based on the target number corresponding to each computing card, mapping the index information corresponding to each input data block in the global space to the local space corresponding to each computing card may include:

[0108] Based on the target number corresponding to the computing card, the index information of each input data block in the global sequence is mapped to the local space corresponding to the computing card; and the index information of each input data block in the input fragment corresponding to the computing card is mapped to the local space corresponding to the computing card.

[0109] In this embodiment, the index information corresponding to the input data block in the global space may include: index information of each input data block in the global sequence, and index information of each input data block in the input segment.

[0110] The sequence length of the global space is different from the sequence length that can be processed by the local space, and the markers of the global interval can be mapped to the starting position of the local space in the card.

[0111] In some embodiments, based on the target number corresponding to each computing card, mapping the index information corresponding to each input data block in the global space to the local space corresponding to each computing card may include:

[0112] Based on the target number corresponding to each computing card, obtain the card starting offset corresponding to each computing card;

[0113] The index information corresponding to the input data block corresponding to each computing card in the global space is subtracted from the starting offset within the card to obtain the index information corresponding to the input data block corresponding to each computing card in the local space.

[0114] In this embodiment, the processing range of each computing card in the global sequence can be determined according to the target number corresponding to each computing card, and the range of identification information of the expert subnetwork deployed on each computing card can be determined.

[0115] For example, the global index information of the first input data block processed by the current computing card can be obtained and determined as the starting offset within the card.

[0116] Then, the index information corresponding to the input data block in the global space can be subtracted from the starting offset within the card to obtain the index information corresponding to the input data block in the local space corresponding to each computing card, so that the index information corresponding to each input data block in the global space can be mapped to the local space corresponding to each computing card.

[0117] During the actual execution process, each expert sub-network processes only a part of gateup_input_all (the set of all tokens processed by all expert sub-networks). The minimum space required for its input gateup_input (local space) can be allocated according to the number of tokens processed on the current computing card (the target number). The storage space required for each subsequent step can be determined by gateup_input to save video memory.

[0118] Assuming that the total number of tokens processed on the current computing card is cur_num_token, the dimension of gateup_input is [cur_num_token, hidden_size], and cur_seg_indptr represents the number range (storage address) of the tokens processed by the current computing card. From this, cur_num_token can be obtained:

[0119] cur_num_token=cur_seg_indptr[-1]-cur_seg_indptr[0].

[0120] Since the gateup_input space accessible to each expert is not the complete seq_len*top_k tokens, but cur_num_token tokens, and the global sequence src2dst and the interval mark array seg_indptr are both designed based on the accessible storage space of seq_len*top_k, they need to be adjusted.

[0121] The number of tokens processed on each computing card is different, and the number of tokens in gateup_input is also different. Each card needs to generate the address number in src2dst within the gateup_input space of the current computing card:

[0122] The number of the first token processed on each card is token_id_start = seg_indptr[start_expert_id], and the number of the last token is token_id_end = seg_indptr[end_expert_id+1]-1. When the element token_id_start in src2dst <= src2dst[i, j] <= token_id_end, the adjusted src2dst_new[i, j] = src2dst[i, j]-token_id_start.

[0123] Similarly, the elements in cur_seg_indptr also need to be within the gateup_input address range accessible to the current card. After adjustment, new_cur_seg_indptr[i]=cur_seg_indptr[i]-cur_seg_indptr[0].

[0124] When constructing the input data gateup_input for the current card, the original input token needs to be copied and saved according to the corresponding expert id in topk_ids and the corresponding number in src2dst_new. Assume that the jth expert id in topk_ids corresponding to the i-th token is expert_id[i, j], and the corresponding number in src2dst_new is addr[i, j]. First, determine whether expert_id[i, j] is on the current card. If so, copy the i-th token and save it to the addr[i, j] row of gateup_input. Otherwise, do nothing. That is, each card only processes tokens processed by the experts on the current card.

[0125] like Figure 4 As shown, the input data src2dst_new and new_cur_seg_indptr obtained in this embodiment are described by taking seq_len=7, top_k=6, M=2 (the number of at least one computing card), and N=64 (the number of multiple expert sub-networks) as an example:

[0126] The expert on card 0 processes tokens 0 through 24, a total of 25 tokens. The expert on card 1 processes tokens 25 through 41, a total of 17 tokens. Elements with values ​​0 through 24 in src2dst remain unchanged (i.e., the starting offset within the card corresponding to card 0 is 0). Elements with values ​​25 through 41 are converted as follows: src2dst_new = src2dst – 25 (i.e., the starting offset within the card corresponding to card 1 is 25). The new_cur_seg_indptr is processed similarly.

[0127] According to the expert parallel computing method of the hybrid expert model provided in the embodiment of the present application, based on the number of all tokens processed by each computing card, the card-internal starting offset corresponding to each computing card is obtained, and then based on the card-internal starting offset, the index information of the token corresponding to each computing card in the global space is mapped to the local space, so that the minimum video memory space can be dynamically allocated according to the number of tokens processed by the computing card, which can solve the problem of excessive video memory usage in the expert parallel mode, thereby increasing the maximum number of inputs supported during model inference.

[0128] In some embodiments, step 120 may include:

[0129] Generate an interval label array based on the number of times each expert sub-network appears in the expert identification sequence;

[0130] Based on the identification information corresponding to each expert sub-network deployed on the computing card, the input fragment corresponding to the computing card is obtained from the interval label array.

[0131] In this embodiment, the interval tag array is used to represent the storage addresses of all input data blocks processed by each expert sub-network in the global sequence.

[0132] By counting the number of times each expert sub-network appears in the expert identification sequence, we can generate the interval tag array seg_indptr. There are N+1 elements in seg_indptr (the dimension is [1, N+1]). The left-closed and right-open interval [seg(i), seg(i+1)) composed of the i-th element and the i+1-th element is the number of all tokens processed by the expert sub-network i in the complete token sequence.

[0133] By slicing the interval marker array, we can get the input fragments corresponding to each calculation card.

[0134] In some embodiments, based on identification information corresponding to each expert sub-network deployed on the computing card, obtaining an input segment corresponding to the computing card from the interval tag array may include:

[0135] Based on the number of at least one computing card and the number of the plurality of expert sub-networks, obtain identification information of a first expert sub-network and identification information of a last expert sub-network deployed on each computing card by calculation;

[0136] Based on the identification information of the first expert sub-network deployed on each computing card and the identification information of the last expert sub-network, the input fragment corresponding to each computing card is obtained from the interval tag array.

[0137] In this embodiment, a corresponding number of expert sub-networks may be allocated to each computing card for computing according to the number of at least one computing card and the number of the plurality of expert sub-networks.

[0138] Thus, the identification information of the first expert sub-network deployed on each computing card and the identification information of the last expert sub-network can be obtained.

[0139] Once the identification information range of the expert sub-network deployed on the computing card is determined, the interval tag array can be sliced ​​to obtain all tokens processed by the expert sub-network on the current card.

[0140] According to the expert parallel computing method of the hybrid expert model provided in the embodiment of the present application, the expert range deployed on each card is dynamically calculated based on the number of computing cards and the total number of experts, ensuring that expert resources are evenly distributed among multiple cards and avoiding overload of a single card. Each card independently processes the assigned expert subnetwork and its corresponding token, thereby improving the overall computing throughput. Each card only loads the parameters of its assigned expert subnetwork, significantly reducing the memory usage of a single card.

[0141] In actual implementation, taking multiple machines and multiple graphics cards as an example, assume there are M computing cards in total. There are N experts in the hybrid expert model, and the number of experts on each card is N / M. Experts 0 to N / M-1 are located on card 0, experts N / M to 2*N / M-1 are located on card 1, and so on.

[0142] Each expert sub-network processes a part of the global sequence, and the tokens and the number of tokens processed by each expert sub-network are different. It is necessary to specify the number of the first token processed by each expert sub-network in the token sequence (global sequence) and the number of tokens processed by the expert.

[0143] When seg(i) = seg(i+1), expert i does not participate in the calculation. The zeroth element of the interval marker array seg_indptr is always 0, that is, seg_indptr[0] = 0. If 0 appears p times in the expert identifier sequence reorder_topk_ids, then seg_indptr[1] = p; if seg_indptr[i] = p, and expert i appears q times in reorder_topk_ids, then seg_indptr[i+1] = p+q... and so on.

[0144] By slicing seg_indptr, we can get all the tokens processed by the experts on the current card. For the i-th card, the identification information of the first expert on the card is: start_expert_id=i*N / M, and the identification information of the last expert is: end_expert_id=(i+1)*N / M-1. All the tokens processed by the experts on the current card are the fragments corresponding to cur_seg_indptr=seg_indptr[start_expert_id, end_expert_id+2].

[0145] According to the expert parallel computing method of the hybrid expert model provided in the embodiment of the present application, the range of the input data block processed by each expert sub-network is recorded by an interval marking array, which can divide the global task into local tasks, so that each computing card only stores and processes the data of the expert sub-network it is responsible for, thereby improving the efficiency of parallel computing.

[0146] In some embodiments, step 110 may include:

[0147] Based on the expert sub-networks corresponding to the original input data blocks, obtain the copy input data blocks corresponding to the expert sub-networks;

[0148] Based on the identification information corresponding to each expert sub-network, the copy input data blocks corresponding to each expert sub-network are arranged to obtain a global sequence;

[0149] Based on the arrangement position information of each expert sub-network corresponding to the original input data block in the expert identification sequence, the storage address of each copy input data block in the global sequence is obtained.

[0150] In this embodiment, for each original input data block, a copy of the token, ie, a copy input data block, may be copied for each expert sub-network participating in the calculation.

[0151] The copy input data blocks of each expert sub-network can be arranged in sequence from small to large according to the identification information of the expert sub-network to obtain a copy sequence of tokens. This sequence is the set of all tokens processed by all expert sub-networks. The copies of each original input data block obtained can be numbered in the copy sequence of the above tokens, and the results can be saved in the global sequence.

[0152] According to the arrangement position of each expert sub-network corresponding to the original data block in the expert identification sequence, the storage address of each copy input data block in the global sequence can be obtained.

[0153] In the actual execution process, the token (copy) sequence can be recorded as gateup_input_all, with a dimension of [seq_len*top_k, hidden_size] (hidden_size is the hidden layer size of the model).

[0154] The number of the copies of each original input token in the above token (copy) sequence is saved in the sequence src2dst (dimension is [1, seq_len*top_k]). Assume that the identification information of the j-th expert sub-network in the top_k expert sub-networks selected for the i-th token is expert_id[i, j], and the position number of expert_id[i, j] in reorder_topk_ids is addr[i, j]. Then the element at the corresponding position in src2dst is src2dst[i*top_k+j]=addr[i, j].

[0155] like Figure 3 As shown, taking seq_len=7, top_k=6, M=2, N=64 as an example, steps 110 and 120 are explained:

[0156] The 0th token is processed by expert 8, so expert_id[0,0]=8, and the position number in reorder_topk_ids is 7, that is, addr[0,0]=7. Then a copy of token0 needs to be stored at the 7th token in the complete token sequence, that is, src2dst[0]=7.

[0157] Expert id (identification information) 0 appears once in reorder_topk_ids, so seg_indptr[1]=1; expert ids 1 and 2 do not appear, so seg_indptr[2]=1, seg_indptr[3]=1; expert id 3 appears 3 times, so seg_indptr[4]=4.

[0158] According to the expert parallel computing method of the hybrid expert model provided in the embodiment of the present application, the storage address of each copy input data block in the global sequence is obtained through the arrangement position information of each expert sub-network corresponding to the original input data block in the expert identification sequence. The expert sub-network can be matched with the token it processes, which makes it easier for each expert sub-network to obtain the token it processes, ensures the correctness of data routing, enables each expert sub-network to calculate in parallel, and improves the efficiency of model reasoning.

[0159] In some embodiments, obtaining expert identification sequences corresponding to multiple expert sub-networks may include:

[0160] Obtaining expert identification matrices corresponding to all expert sub-networks for processing multiple input data blocks;

[0161] Based on the identification information corresponding to each expert sub-network in the expert identification matrix, the identification information corresponding to multiple expert sub-networks is sorted in ascending order to obtain an expert identification sequence.

[0162] In this embodiment, multiple expert sub-networks may be selected for each token for calculation.

[0163] For multiple input data blocks, expert identification matrices corresponding to all expert sub-networks for processing the multiple input data blocks can be obtained.

[0164] According to the identification information of each expert sub-network in the expert identification matrix, the identification information corresponding to the multiple expert sub-networks can be sorted in ascending order of the identification information to obtain an expert identification sequence.

[0165] In some embodiments, obtaining expert identification matrices corresponding to multiple expert sub-networks for processing multiple input data blocks may include:

[0166] An expert identification matrix is ​​constructed based on the number of the plurality of input data blocks and the number of expert sub-networks required for processing each input data block.

[0167] In this embodiment, the first dimension of the expert identification matrix corresponds to the number of the plurality of input data blocks, and the second dimension of the expert identification matrix corresponds to the number of expert sub-networks required for each input data block.

[0168] In the actual execution process, it is assumed that the input length to be processed by the hybrid expert model is seq_len (that is, there are seq_len tokens to be processed). After the expert selection is completed, the top_k expert sub-networks are selected for each token, and the expert identification matrix of the selected expert sub-network is topk_ids (dimension is [seq_len, top_k]).

[0169] like Figure 3 As shown, when seq_len=7 and top_k=6, the dimension of topk_ids is [7, 6].

[0170] Sort the expert identification information in topk_ids from small to large, and save the result as the expert identification sequence reorder_topk_ids (dimension is [1, seq_len*top_k]), such as Figure 3 As shown, the dimension of reorder_topk_ids is [1, 42].

[0171] According to the expert parallel computing method of the hybrid expert model provided in the embodiment of the present application, by sorting the identification information of multiple expert sub-networks selected by each token, an expert identification sequence is generated, which can make the expert processing order continuous and facilitate the subsequent fragmented storage of data according to the expert identification information.

[0172] like Figure 2 As shown, in some embodiments, step 140 may include:

[0173] Perform gate operation, up-projection operation, activation function operation and down-projection operation on the input data block to be processed to obtain the sub-result data output by each computing card.

[0174] In this embodiment, a gating operation may be used to determine which expert sub-networks should process an input data block.

[0175] The up-projection operation can map the input features to a high-dimensional space, which can improve the feature expression capability.

[0176] The activation function operation can introduce nonlinearity while retaining some linear features, which can alleviate the gradient disappearance problem.

[0177] The down-projection operation can map high-dimensional features back to the original dimension, compress the feature dimension, and match the output requirements of the model.

[0178] In some embodiments, step 140 may include:

[0179] Control each expert sub-network deployed on each computing card, read the input data block to be processed from the target video memory space, and perform gating and up-projection operations on the input data block to obtain the first-stage result data;

[0180] Perform activation function operation on the result data of the first stage to obtain the result data of the second stage;

[0181] Perform down-projection operation on the second stage result data to obtain the sub-result data output by each calculation card.

[0182] In this embodiment, the expert subnetwork on each computing card can independently read the token segment it is responsible for from the target video memory space according to new_cur_seg_indptr. For each token, the expert subnetwork can perform gating operations (gate operations) and up-projection operations (up operations). The gating operations can calculate the gate value, and the up-projection operations can perform dimensionality increase transformation.

[0183] The result data of the first stage is the concatenation of the result of the gating operation (gate value) and the result of the up-projection operation (up value).

[0184] A nonlinear activation function can be used to perform activation function operation (silu operation) on the first stage result data. The nonlinear activation function can be: silu(x)=x*σ(x) (σ is Sigmoid).

[0185] An activation function can be used to perform activation function operations on the results of the gating operations in the first stage result data. For example, for each token, the gate value, the gate value after silu activation, and the up value can be multiplied element by element to obtain the second stage result data.

[0186] The expert sub-network on each card can independently read the token fragments it is responsible for and perform dimensionality reduction transformation on each token. The final feature representation of each token only contains the results processed by the expert sub-network on the current computing card, thereby obtaining the sub-result data output by each computing card.

[0187] In some embodiments, performing a down-projection operation on the second-stage result data to obtain sub-result data output by each computing card may include:

[0188] Control each expert sub-network deployed on each computing card, read the input data block to be processed from the second-stage result data, and perform a down-projection operation on the second-stage result data corresponding to the input data block to be processed to obtain the sub-result data output by each computing card.

[0189] In this embodiment, before the operation starts, each expert sub-network needs to read the input data block to be processed from the second-stage result data according to new_cur_seg_indptr, and then perform a down projection operation (down operation) on the second result data corresponding to the input data block to be processed to output sub-result data.

[0190] In the actual execution process, first, the gate operation (gating operation) and up operation (upper projection operation) of the MoE layer can be performed on gateup_input, which is realized by matrix multiplication. The output is gateup_output (first stage result data). The dimension of gateup_output on each card is [cur_num_token, gateup_output.shape[1]], where cur_num_token is the total number of input data blocks processed by the computing card.

[0191] Before the operation starts, each expert sub-network can read the token it needs to process from gateup_input according to new_cur_seg_indptr, and then each expert sub-network calculates in parallel.

[0192] gateup_output is composed of the gate output gate_output and the up output up_output.

[0193] By performing the activation function operation on the result data of the first stage, the result data of the second stage can be obtained: silu_output=gete_output*silu(gate_output)*up_output. The dimension of silu_output on each computing card is [cur_num_token, silu_outpout.shape[1]], where cur_num_token is the total number of input data blocks processed by the computing card, gete_output is the result of the gate operation output, and up_output is the result of the up-projection operation output.

[0194] The down-projection operation (down operation) of the MoE layer is performed on the second-stage result data silu_output, which is realized by matrix multiplication. The output is the sub-result data down_output. The dimension of down_output on each computing card is [cur_num_token, down_output.shape[1]], where cur_num_token is the total number of input data blocks processed by the computing card.

[0195] Before the operation begins, each expert sub-network also needs to read the token it needs to process from silu_output according to new_cur_seg_indptr.

[0196] According to the expert parallel computing method of the hybrid expert model provided in the embodiment of the present application, by performing gating operations, upper projection operations, activation function operations and lower projection operations on the input data blocks to be processed, the sub-result data output by each computing card is obtained, thereby realizing the complete process of dynamic routing, feature transformation and nonlinear activation, and improving the reasoning efficiency in a multi-card environment.

[0197] In some embodiments, step 150 may include:

[0198] Performing splicing processing on at least one sub-result data to obtain first result data; the first result data includes result data output by each expert sub-network corresponding to each input data block;

[0199] Based on the weight information corresponding to each expert sub-network, the result data output by each expert sub-network in the first result data is weighted and fused to obtain the target result data.

[0200] In this embodiment, the gating network can be used to determine how to distribute data to different expert sub-networks. The gating network can generate weights based on the input features, that is, the weight information corresponding to each expert sub-network. The weights are used to determine the contribution of each expert sub-network to the final output.

[0201] The sub-result data of each computing card can be spliced ​​into a globally complete intermediate result, and the results corresponding to multiple expert sub-networks of each token can be arranged in a fixed order.

[0202] The order of arrangement of the results corresponding to the expert sub-network may be consistent with the order of the expert identifications in the expert identification matrix.

[0203] According to the weight information, the result data output by each expert sub-network in the first result data can be weighted and fused to obtain the target result data.

[0204] In actual execution, the sub-result data down_output on each calculation card is part of the complete output result. The final output can be obtained in two steps:

[0205] First, the valid outputs on each computing card can be spliced ​​into a complete comb_output (the dimension is [seq_len*top_k, hidden_size], where seq_len is the number of tokens to be processed and top_k is the number of expert sub-networks corresponding to each token);

[0206] Secondly, the results of each token processed by the top_k expert sub-networks can be weighted and fused according to the corresponding weight information topk_weights (the dimension is [seq_len, top_k], where seq_len is the number of tokens to be processed and top_k is the number of expert sub-networks corresponding to each token). The output is the final result (target result data) output by the hybrid expert model.

[0207] Among them, when performing weighted fusion, the arrangement order of the top_k results of each expert in comb_output is required to be consistent with the weight order in topk_weights (that is, the expert ID of each row in cpmb_output is arranged according to the expert identification matrix topk_ids).

[0208] The output of the i-th token after being processed by the j-th expert among the selected top_k experts is comb_output[i, j], and the corresponding number in src2dst_new is src2dst_new[i, j]. Then the card where the expert is located needs to read the src2dst_new[i, j]-th token from down_output and save it to the i*top_k+j-th token in comb_output.

[0209] According to the expert parallel computing method of the hybrid expert model provided in the embodiment of the present application, each computing card obtains sub-result data by independently processing part of the token. By splicing at least one sub-result data, the first result data can be obtained, avoiding the synchronization of the full data and reducing the communication overhead. The result data output by each expert sub-network in the first result data is weightedly fused through the weight information corresponding to each expert sub-network to obtain the final target result data, thereby improving the feature richness and the precision and accuracy of the model output results.

[0210] The computer program product provided by the present application is described below. The computer program product described below and the expert parallel computing method of the hybrid expert model described above can be referenced to each other.

[0211] The embodiment of the present application provides an expert parallel computing method for a hybrid expert model, and the execution subject may be a computer program product. In the embodiment of the present application, the computer program product provided by the embodiment of the present application is described by taking the computer program product executing the expert parallel computing method for a hybrid expert model as an example.

[0212] An embodiment of the present application also provides a computer program product.

[0213] like Figure 5 As shown, the computer program product, the hybrid expert model includes multiple expert sub-networks, and the multiple expert sub-networks are deployed on at least one computing card; including: a first processing module 510, a second processing module 520, a third processing module 530, a fourth processing module 540 and a fifth processing module 550.

[0214] A first processing module 510 is configured to obtain a storage address of each input data block in the global sequence based on the input data blocks processed by each expert sub-network and the expert identification sequences corresponding to the multiple expert sub-networks; the multiple input data blocks correspond one-to-one to the multiple expert sub-networks;

[0215] The second processing module 520 is configured to obtain input segments processed by all expert subnetworks deployed on each computing card based on the number of occurrences of each expert subnetwork in the expert identification sequence and the storage address corresponding to each input data block; the input segments include multiple input data blocks;

[0216] A third processing module 530 is configured to allocate target video memory space to each computing card based on the input fragments corresponding to each computing card;

[0217] The fourth processing module 540 is used to control each expert sub-network deployed on each computing card, read the input data block to be processed from the target memory space corresponding to each computing card, and calculate the input data block to be processed to obtain the sub-result data output by each computing card;

[0218] The fifth processing module 550 is configured to perform fusion processing on at least one sub-result data to obtain target result data output by the hybrid expert model.

[0219] According to the computer program product provided in the embodiment of the present application, by obtaining all input data blocks (input fragments) processed on each computing card, the minimum target video memory space is dynamically allocated to it according to the input fragments corresponding to the computing card, and multiple expert sub-networks on the computing card are controlled to process the input data in parallel, thereby solving the problem of excessive video memory usage in the expert parallel mode, increasing the maximum number of inputs supported during inference, and thus improving the inference efficiency of the model.

[0220] In some embodiments, the third processing module 530 may also be configured to:

[0221] Based on the input fragments corresponding to each computing card, obtain the target number of input data blocks to be processed by each computing card;

[0222] Based on the target number corresponding to each computing card, the index information corresponding to each input data block in the global space is mapped to the local space corresponding to each computing card;

[0223] Based on the index information of the input data block in the local space corresponding to each computing card, the target video memory space is allocated to each computing card.

[0224] In some embodiments, the third processing module 530 may also be configured to:

[0225] Based on the target number corresponding to the computing card, the index information of each input data block in the global sequence is mapped to the local space corresponding to the computing card; and the index information of each input data block in the input fragment corresponding to the computing card is mapped to the local space corresponding to the computing card.

[0226] In some embodiments, the third processing module 530 may also be configured to:

[0227] Based on the target number corresponding to each computing card, obtain the card starting offset corresponding to each computing card;

[0228] The index information corresponding to the input data block corresponding to each computing card in the global space is subtracted from the starting offset within the card to obtain the index information corresponding to the input data block corresponding to each computing card in the local space.

[0229] In some embodiments, the second processing module 520 may also be configured to:

[0230] Based on the number of times each expert sub-network appears in the expert identification sequence, an interval tag array is generated; the interval tag array is used to represent the storage address of all input data blocks processed by each expert sub-network in the global sequence;

[0231] Based on the identification information corresponding to each expert sub-network deployed on the computing card, the input fragment corresponding to the computing card is obtained from the interval label array.

[0232] In some embodiments, the first processing module 510 may also be configured to:

[0233] Based on the expert sub-networks corresponding to the original input data blocks, obtain the copy input data blocks corresponding to the expert sub-networks;

[0234] Based on the identification information corresponding to each expert sub-network, the copy input data blocks corresponding to each expert sub-network are arranged to obtain a global sequence;

[0235] Based on the arrangement position information of each expert sub-network corresponding to the original input data block in the expert identification sequence, the storage address of each copy input data block in the global sequence is obtained.

[0236] In some embodiments, the computer program product may further include a sixth processing module configured to:

[0237] Obtaining expert identification matrices corresponding to all expert sub-networks for processing multiple input data blocks;

[0238] Based on the identification information corresponding to each expert sub-network in the expert identification matrix, the identification information corresponding to multiple expert sub-networks is sorted in ascending order to obtain an expert identification sequence.

[0239] In some embodiments, the sixth processing module may further be configured to:

[0240] An expert identification matrix is ​​constructed based on the number of multiple input data blocks and the number of expert sub-networks required to process each input data block; the first dimension of the expert identification matrix corresponds to the number of multiple input data blocks, and the second dimension of the expert identification matrix corresponds to the number of expert sub-networks required for each input data block.

[0241] In some embodiments, the fourth processing module 540 may also be configured to:

[0242] Perform gate operation, up-projection operation, activation function operation and down-projection operation on the input data block to be processed to obtain the sub-result data output by each computing card.

[0243] In some embodiments, the fourth processing module 540 may also be configured to:

[0244] Control each expert sub-network deployed on each computing card, read the input data block to be processed from the target video memory space, and perform gating and up-projection operations on the input data block to obtain the first-stage result data;

[0245] Perform activation function operation on the result data of the first stage to obtain the result data of the second stage;

[0246] Perform down-projection operation on the second stage result data to obtain the sub-result data output by each calculation card.

[0247] In some embodiments, the fourth processing module 540 may also be configured to:

[0248] Control each expert sub-network deployed on each computing card, read the input data block to be processed from the second-stage result data, and perform a down-projection operation on the second-stage result data corresponding to the input data block to be processed to obtain the sub-result data output by each computing card.

[0249] In some embodiments, the fifth processing module 550 may also be configured to:

[0250] Performing splicing processing on at least one sub-result data to obtain first result data; the first result data includes result data output by each expert sub-network corresponding to each input data block;

[0251] Based on the weight information corresponding to each expert sub-network, the result data output by each expert sub-network in the first result data is weighted and fused to obtain the target result data.

[0252] The computer program product in the embodiments of the present application may be an electronic device or a component of an electronic device, such as an integrated circuit or chip. The electronic device may be a terminal or other device other than a terminal. For example, the electronic device may be a mobile phone, a tablet computer, a laptop computer, a PDA, an in-vehicle electronic device, a mobile internet device (MID), an augmented reality (AR) / virtual reality (VR) device, a robot, a wearable device, an ultra-mobile personal computer (UMPC), a netbook, or a personal digital assistant (PDA). It may also be a server, a network attached storage (NAS), a personal computer (PC), a television (TV), a teller machine (ATM), or a self-service machine, etc., and the embodiments of the present application do not specifically limit this.

[0253] The computer program product in the embodiments of the present application may be a device having an operating system. The operating system may be an Android operating system, an iOS operating system, or other possible operating systems, which are not specifically limited in the embodiments of the present application.

[0254] The computer program product provided in the embodiments of the present application can achieve Figures 1 to 4 To avoid repetition, the various processes implemented in the method embodiment are not described here.

[0255] In some embodiments, as Figure 6 As shown, an embodiment of the present application also provides an electronic device 600, including a processor 601, a memory 602, and a computer program stored in the memory 602 and executable on the processor 601. When the program is executed by the processor 601, each process of the embodiment of the expert parallel computing method of the above-mentioned hybrid expert model is implemented, and the same technical effect can be achieved. To avoid repetition, it will not be described here.

[0256] It should be noted that the electronic devices in the embodiments of the present application include the mobile electronic devices and non-mobile electronic devices mentioned above.

[0257] On the other hand, the present application also provides a non-transitory computer-readable storage medium on which a computer program is stored. When the computer program is executed by a processor, it is implemented to execute the various processes of the expert parallel computing method embodiment of the above-mentioned hybrid expert model and can achieve the same technical effect. To avoid repetition, it will not be repeated here.

[0258] On the other hand, an embodiment of the present application further provides a chip, which includes a processor and a communication interface, the communication interface is coupled to the processor, and the processor is used to run programs or instructions to implement the various processes of the expert parallel computing method embodiment of the above-mentioned hybrid expert model, and can achieve the same technical effect. To avoid repetition, it will not be repeated here.

[0259] It should be understood that the chip mentioned in the embodiments of the present application can also be called a system-level chip, a system chip, a chip system or a system-on-chip chip, etc.

[0260] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, i.e., they may be located in one location or distributed across multiple network units. Some or all of the modules may be selected based on actual needs to achieve the objectives of the present embodiment. Persons of ordinary skill in the art will be able to understand and implement the present invention without inventive effort.

[0261] Through the description of the above embodiments, those skilled in the art will clearly understand that each embodiment can be implemented using software plus a necessary general-purpose hardware platform, or of course, hardware. Based on this understanding, the essence of the above technical solution, or the portion that contributes to the relevant technology, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a magnetic disk, or an optical disk, and includes a number of instructions for causing a computer device (such as a personal computer, server, or network device) to execute the methods described in each embodiment or certain portions of the embodiments.

[0262] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present application.

Claims

1. An expert parallel computing method for a hybrid expert model, characterized in that: The hybrid expert model includes multiple expert sub-networks, and the multiple expert sub-networks are deployed on at least one computing card; the method includes: Based on the input data blocks processed by each of the expert sub-networks and the expert identification sequences corresponding to the multiple expert sub-networks, obtaining the storage address of each of the input data blocks in the global sequence; the multiple input data blocks correspond one-to-one to the multiple expert sub-networks; Based on the number of occurrences of each expert sub-network in the expert identification sequence and the storage address corresponding to each input data block, obtaining input segments processed by all expert sub-networks deployed on each computing card; the input segments include multiple input data blocks; Allocating target video memory space to each computing card based on the input fragments corresponding to each computing card; Controlling each expert sub-network deployed on each computing card, reading an input data block to be processed from a target memory space corresponding to each computing card, and performing calculations on the input data block to be processed to obtain sub-result data output by each computing card; At least one of the sub-result data is fused to obtain target result data output by the hybrid expert model.

2. The expert parallel computing method of the hybrid expert model according to claim 1, characterized in that: Allocating target video memory space to each computing card based on the input fragments corresponding to each computing card includes: Based on the input fragments corresponding to the computing cards, a target number of input data blocks to be processed corresponding to the computing cards is obtained; Based on the target number corresponding to each computing card, the index information corresponding to each input data block in the global space is mapped to the local space corresponding to each computing card; Based on the index information of the input data block in the local space corresponding to each computing card, a target video memory space is allocated to each computing card.

3. The expert parallel computing method of the hybrid expert model according to claim 2, characterized in that: The mapping of the index information corresponding to each input data block in the global space to the local space corresponding to each computing card based on the target number corresponding to each computing card includes: Based on the target number corresponding to the computing card, the index information of each input data block in the global sequence is mapped to the local space corresponding to the computing card; and the index information of each input data block in the input segment corresponding to the computing card is mapped to the local space corresponding to the computing card.

4. The expert parallel computing method of the hybrid expert model according to claim 2, characterized in that: The mapping of the index information corresponding to each input data block in the global space to the local space corresponding to each computing card based on the target number corresponding to each computing card includes: Based on the target number corresponding to each computing card, obtaining the intra-card starting offset corresponding to each computing card; The index information corresponding to the input data block corresponding to each computing card in the global space is subtracted from the starting offset within the card to obtain the index information corresponding to the input data block corresponding to each computing card in the local space.

5. The expert parallel computing method of the hybrid expert model according to any one of claims 1 to 4, characterized in that: The step of obtaining input segments processed by all expert sub-networks deployed on each computing card based on the number of occurrences of each expert sub-network in the expert identification sequence and the storage address corresponding to each input data block includes: Based on the number of occurrences of each of the expert sub-networks in the expert identification sequence, an interval tag array is generated; the interval tag array is used to represent the storage addresses of all input data blocks processed by each of the expert sub-networks in the global sequence; Based on identification information corresponding to each expert sub-network deployed on the computing card, an input segment corresponding to the computing card is obtained from the interval tag array.

6. The expert parallel computing method of the hybrid expert model according to any one of claims 1 to 4, characterized in that: The acquiring, based on the input data blocks processed by each of the expert sub-networks and the expert identification sequences corresponding to the multiple expert sub-networks, the storage address of each of the input data blocks in the global sequence comprises: Based on the expert sub-networks corresponding to the original input data blocks, obtaining the copy input data blocks corresponding to the expert sub-networks; Arranging the copy input data blocks corresponding to each of the expert sub-networks based on the identification information corresponding to each of the expert sub-networks to obtain the global sequence; Based on the arrangement position information of each expert sub-network corresponding to the original input data block in the expert identification sequence, the storage address of each copy input data block in the global sequence is obtained.

7. The expert parallel computing method of the hybrid expert model according to any one of claims 1 to 4, characterized in that: Obtaining expert identification sequences corresponding to the multiple expert sub-networks includes: Obtaining expert identification matrices corresponding to all expert sub-networks for processing the plurality of input data blocks; Based on the identification information corresponding to each expert sub-network in the expert identification matrix, the identification information corresponding to the multiple expert sub-networks is sorted in ascending order of the identification information to obtain the expert identification sequence.

8. The expert parallel computing method of the hybrid expert model according to claim 7, characterized in that: The obtaining of expert identification matrices corresponding to a plurality of expert sub-networks for processing a plurality of the input data blocks includes: The expert identification matrix is ​​constructed based on the number of the multiple input data blocks and the number of expert sub-networks required for processing each of the input data blocks; the first dimension of the expert identification matrix corresponds to the number of the multiple input data blocks, and the second dimension of the expert identification matrix corresponds to the number of expert sub-networks required for each of the input data blocks.

9. The expert parallel computing method of the hybrid expert model according to any one of claims 1 to 4, characterized in that: The step of calculating the input data block to be processed to obtain sub-result data output by each calculation card includes: Performing a gating operation, an upper projection operation, an activation function operation, and a lower projection operation on the input data block to be processed to obtain the sub-result data output by each of the computing cards.

10. The expert parallel computing method of the hybrid expert model according to claim 9, characterized in that: The controlling of each expert sub-network deployed on each computing card, reading an input data block to be processed from a target memory space corresponding to each computing card, and performing calculations on the input data block to be processed to obtain sub-result data output by each computing card, includes: Controlling each expert sub-network deployed on each computing card, reading an input data block to be processed from the target video memory space, and performing the gating operation and the up-projection operation on the input data block to be processed to obtain first-stage result data; Performing the activation function operation on the first-stage result data to obtain the second-stage result data; The lower projection operation is performed on the second-stage result data to obtain the sub-result data output by each of the computing cards.

11. The expert parallel computing method of the hybrid expert model according to claim 10, characterized in that: The performing the down-projection operation on the second-stage result data to obtain the sub-result data output by each computing card includes: Controlling each expert sub-network deployed on each of the computing cards, reading the input data block to be processed from the second-stage result data, and performing the down-projection operation on the second-stage result data corresponding to the input data block to be processed to obtain the sub-result data output by each of the computing cards.

12. The expert parallel computing method of the hybrid expert model according to any one of claims 1 to 4, characterized in that: The fusing of at least one of the sub-result data to obtain the target result data output by the hybrid expert model includes: Performing splicing processing on at least one of the sub-result data to obtain first result data; the first result data includes result data output by each expert sub-network corresponding to each input data block; Based on the weight information corresponding to each of the expert sub-networks, weighted fusion processing is performed on the result data output by each of the expert sub-networks in the first result data to obtain the target result data.

13. A computer program product, characterized in that The hybrid expert model includes multiple expert sub-networks, which are deployed on at least one computing card; including: a first processing module, configured to obtain a storage address of each input data block in a global sequence based on the input data blocks processed by each of the expert sub-networks and the expert identification sequences corresponding to the multiple expert sub-networks; wherein the multiple input data blocks correspond one-to-one to the multiple expert sub-networks; a second processing module, configured to obtain input segments processed by all expert subnetworks deployed on each of the computing cards based on the number of occurrences of each of the expert subnetworks in the expert identification sequence and the storage address corresponding to each of the input data blocks; the input segments comprising a plurality of input data blocks; A third processing module is configured to allocate target video memory space to each computing card based on the input fragments corresponding to each computing card; a fourth processing module, configured to control each expert sub-network deployed on each computing card, read an input data block to be processed from a target video memory space corresponding to each computing card, and perform calculations on the input data block to be processed to obtain sub-result data output by each computing card; The fifth processing module is used to perform fusion processing on at least one of the sub-result data to obtain the target result data output by the hybrid expert model.

14. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the expert parallel computing method of the hybrid expert model according to any one of claims 1 to 12 is implemented.

15. A non-volatile computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the expert parallel computing method of the hybrid expert model according to any one of claims 1 to 12 is implemented.

Citation Information

Patent Citations

  • Unified anomaly detection model and detection method based on hybrid expert system

    CN117911328A

  • Video memory optimization online method and system based on FastMoE model

    CN118505491A