Data processing method, processor, electronic equipment and storage medium

By reading the input tensor into the cache in the MOE structure and aggregating the eigenvectors using the converted routing information, the problem of low aggregation efficiency in the prior art is solved, and efficient memory utilization and bandwidth optimization are achieved.

CN120181131AActive Publication Date: 2025-06-20SHANGHAI BIREN TECH CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510562020.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-30
Publication Date
2025-06-20
Estimated Expiration
2045-04-30

AI Technical Summary

Technical Problem

When processing data in the MOE structure, it is difficult to efficiently aggregate the feature vectors in the input tensor according to the network model, resulting in an increase in the number of memory accesses, low bandwidth utilization, and the on-chip cache resources cannot be fully utilized.

Method used

By reading the input tensor from memory to the cache, the routing information is obtained and converted to aggregate the word eigenvectors according to the network model. The specific steps include reading the first tensor to the cache, obtaining routing information, and converting routing information is used to indicate the linear index position of the word element feature vector of each network model, and performing a first arrangement operation to aggregate the feature vectors.

Benefits of technology

It realizes that when storing into memory, the eigenvectors have been arranged, which reduces the number of memory access times, improves bandwidth utilization, and makes full use of on-chip cache resources, reducing resource waste.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120181131A_ABST
    Figure CN120181131A_ABST
Patent Text Reader

Abstract

The invention discloses a data processing method, a processor, electronic equipment and a storage medium. The data processing method comprises the following steps: reading a first tensor from a memory to a cache; obtaining routing information, the routing information comprising n pieces of routing sub-information in one-to-one correspondence with the n lexical element feature vectors, and each piece of routing sub-information being used for indicating m network models input by the corresponding lexical element feature vector; the routing information is converted to obtain converted routing information, the converted routing information is used for indicating a linear index position corresponding to a lexical element feature vector input into each network model with the network model as a unit, and the linear index position is used for indicating the position of the corresponding lexical element feature vector in the first tensor; on the basis of the converted routing information, lexical element feature vectors input into each network model are extracted from the cache, and the extracted lexical element feature vectors input into the same network model form a tensor whole to be stored in a memory.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present disclosure relate to a data processing method, a processor, an electronic device, and a non-transitory computer-readable storage medium. Background Art

[0002] The MOE structure, short for Mixture of Experts, is an advanced neural network architecture that improves the overall model performance by integrating the predictions of multiple network models (such as "expert networks"). Its core idea is to allocate input data to different expert networks for processing, and then a gating network is used to weight and combine the outputs of the expert networks to generate the final result. The MOE network is widely used in fields such as natural language processing (NLP), computer vision (CV), and recommendation systems. For example, in large language models (LLMs), the MOE structure can replace the feed-forward network (FFN) layer in the traditional Transformer architecture to improve the efficiency and performance of the model. Summary of the Invention

[0003] At least one embodiment of the present disclosure provides a data processing method for sorting an input tensor to aggregate multiple feature vectors included in the input tensor according to a network model, and storing the sorted tensor obtained after sorting in a memory. The data processing method includes: reading a first tensor from the memory to a cache, where the first tensor includes n token feature vectors, each token feature vector is input to m different network models for subsequent processing, the first tensor is continuously stored in the memory, the first tensor includes at least part of the data in the input tensor, each token feature vector is at least part of the content in a corresponding one of the feature vectors in the input tensor, and m and n are positive integers; obtaining routing information, where the routing information includes n routing sub-information corresponding one-to-one to the n token feature vectors, and each routing sub-information is used to indicate the m network models input by the corresponding token feature vector; converting the routing information to obtain converted routing information, where the converted routing information is used to indicate, in units of network models, the linear index positions corresponding to the token feature vectors input to each network model, and the linear index positions are used to indicate the positions of the corresponding token feature vectors in the first tensor; and based on the converted routing information, performing a first permutation operation to aggregate the n token feature vectors included in the first tensor according to the network model, where the token feature vectors input to the same network model are grouped into a tensor as a whole and stored in the memory.

[0004] For example, in a data processing method provided by at least one embodiment of the present disclosure, converting the routing information to obtain converted routing information includes: storing the linear index positions corresponding to the n token feature vectors into an array corresponding to the assigned network model based on each routing sub-information, so as to obtain the converted routing information.

[0005] For example, in a data processing method provided by at least one embodiment of the present disclosure, each routing sub-information includes m model identification information for indicating the identification numbers of m network models into which the corresponding token feature vector is input. Storing the linear index positions corresponding to the n token feature vectors into an array corresponding to the assigned network model based on each routing sub-information includes: traversing the n routing sub-informations, and for any one of the routing sub-informations, determining the linear index position corresponding to the token feature vector based on the index position of the token feature vector corresponding to the any one of the routing sub-informations in the first tensor, wherein the linear index position corresponding to the token feature vector includes m linear index positions corresponding one-to-one to the m model identification information; for any one of the m linear index positions, writing the any one of the linear index positions into the array corresponding to the network model indicated by the corresponding model identification information.

[0006] For example, in a data processing method provided by at least one embodiment of the present disclosure, performing a first permutation operation based on the converted routing information includes: based on the converted routing information, for each network model, extracting the token feature vector input to the network model from the cache, and storing the extracted token feature vectors input to the network model as a tensor as a whole in the memory.

[0007] For example, in a data processing method provided by at least one embodiment of the present disclosure, the converted routing information includes v arrays corresponding one-to-one to v network models, and each array includes a plurality of valid elements for indicating the linear index positions corresponding to the token feature vectors input to the corresponding network model, where v is a positive integer. Performing, based on the converted routing information, for each network model, extracting the token feature vector input to the network model from the cache, and storing the extracted token feature vectors input to the network model as a tensor as a whole in the memory includes: for any one of the v arrays: traversing each valid element in the any one of the arrays, extracting the token feature vector indicated by each valid element from the cache and caching it in the buffer component; in response to traversing all the valid elements in the any one of the arrays, storing the token feature vectors extracted in the buffer component as a tensor as a whole in the memory.

[0008] For example, in a data processing method provided by at least one embodiment of the present disclosure, extracting the token feature vectors indicated by each valid element from the cache and caching them in a buffer component includes: for each valid element, determining the index position of the token feature vector corresponding to the linear index position in the first tensor based on the linear index position indicated by the valid element; extracting the corresponding token feature vector from the cache based on the index position and caching it in the buffer component.

[0009] For example, a data processing method provided by at least one embodiment of the present disclosure further includes: determining a block size based on the capacity of the cache; dividing the input tensor into multiple blocks in units of the block size, where the number of token feature vectors included in each block is n, and any one token feature vector in each block is at least part of the corresponding feature vector in the input tensor; where reading the first tensor from the memory to the cache includes: reading one block from the memory as the first tensor and storing it in the cache; in response to the current first tensor in the cache being processed, reading the next block into the cache until all the multiple blocks are processed.

[0010] For example, in a data processing method provided by at least one embodiment of the present disclosure, based on the block size, the multiple routing sub-information corresponding to the multiple feature vectors is divided into multiple groups, the number of routing sub-information in each group is n, and for any one network model, the token feature vectors input to the any one network model extracted based on different groups of routing information are continuously stored in the memory.

[0011] For example, a data processing method provided by at least one embodiment of the present disclosure further includes: performing a second permutation operation on the result tensors respectively output by each network model based on the transformed routing information to merge and sort the result tensors respectively output by each network model and store them in the memory, where the result tensors respectively output by each network model are obtained by processing the feature vectors input to each network model; where in the second permutation operation, the result feature vectors written into the cache based on the same transformed routing information form a tensor and are stored in the memory as a whole, and the tensor is continuously stored in the memory.

[0012] For example, in a data processing method provided by at least one embodiment of the present disclosure, a second permutation operation is performed on the result tensors respectively output by each network model based on the converted routing information, including: based on the converted routing information, reading a second tensor from the memory to a buffer component, where the second tensor is at least part of the result tensor output by a network model, and the second tensor is continuously stored in the memory; based on the converted routing information, performing a merge sorting process on the second tensor, where in the merge sorting process, based on the converted routing information, writing each token result vector included in the second tensor to the cache, and the relative writing position relationship between the token result vectors is the same as the relative position relationship of the token feature vectors corresponding to the token result vectors in the input tensor; in response to all the second tensors obtained based on the converted routing information having completed the merge sorting process, storing the tensor formed by the token result vectors obtained based on the merge sorting process currently stored in the cache to the memory as a whole.

[0013] For example, in a data processing method provided by at least one embodiment of the present disclosure, the converted routing information includes v arrays corresponding one-to-one to v network models. The valid elements in each array indicate the linear index positions corresponding to the token feature vectors input to the corresponding network model. The number of valid elements in each array is the number of token feature vectors in the network model corresponding to the input. Based on the converted routing information, reading a second tensor from the memory to a buffer component includes: sequentially traversing the v arrays. For the j-th array, according to the number x of valid elements in the j-th array, reading at least part of x result feature vectors from the memory as the second tensor, and caching the second tensor to the buffer component, where the x result feature vectors come from the result tensor output by the network model corresponding to the j-th array, the second tensor includes x token result vectors, and each token result vector includes at least part of the corresponding result feature vector content, j is a positive integer and less than or equal to v, and x is an integer; in response to the current second tensor in the buffer component being processed, reading a new second tensor from the memory to the buffer component based on the (j + 1)-th array, where the new second tensor is at least part of the result tensor output by the network model corresponding to the (j + 1)-th array.

[0014] For example, in a data processing method provided by at least one embodiment of the present disclosure, performing a merge sorting process on the second tensor based on the converted routing information includes: determining the writing positions of each token result vector in the second tensor in the cache based on the valid elements in the array corresponding to the second tensor; writing each token result vector in the second tensor to the corresponding writing position in the cache according to the writing positions.

[0015] For example, in a data processing method provided by at least one embodiment of the present disclosure, according to the writing position, writing each token result vector in the second tensor to the corresponding writing position in the cache includes: for any one of the token result vectors included in the second tensor, in response to the fact that the corresponding writing position of the any one token result vector in the cache already stores data, performing an arithmetic operation on the any one token result vector and the data, and writing the operation result to the corresponding writing position; in response to the fact that the corresponding writing position of the any one token result vector in the cache does not store data, writing the any one token result vector to the corresponding writing position.

[0016] For example, in a data processing method provided by at least one embodiment of the present disclosure, in response to the input tensor being divided into multiple blocks based on a block size, the routing information is divided into multiple groups of routing information based on the block size, the multiple groups of routing information include a first group of routing information and a second group of routing information, the first group of routing information corresponds to a first block, the second group of routing information corresponds to a second block, performing the second permutation operation on the result tensors respectively output by each network model based on the transformed routing information obtained by transforming the first group of routing information, so as to write a first intermediate tensor into the memory, performing the second permutation operation on the result tensors respectively output by each network model based on the transformed routing information obtained by transforming the second group of routing information, so as to write a second intermediate tensor into the memory, when writing the first intermediate tensor and the second intermediate tensor into the memory, the storage position relationship between the first intermediate tensor and the second intermediate tensor in the memory is the same as the relative position relationship between the first block and the second block.

[0017] For example, in a data processing method provided by at least one embodiment of the present disclosure, the data processing method is used for data processing of a mixture of experts model structure, the mixture of experts model structure includes a plurality of expert networks as the multiple network models and a routing module, the routing information is determined by the routing module according to the input tensor input to the mixture of experts model structure, and each expert network is trained to process specific tasks and data features.

[0018] At least one embodiment of the present disclosure provides a data processing device for sorting an input tensor to aggregate a plurality of feature vectors included in the input tensor according to a network model, and storing the sorted tensor obtained after sorting into a memory. The data processing device includes: a reading module configured to read a first tensor from the memory into a cache, where the first tensor includes n token feature vectors, each token feature vector is input into m different network models for subsequent processing, the first tensor is continuously stored in the memory, the first tensor includes at least part of the data in the input tensor, each token feature vector is at least part of the content of a feature vector in the input tensor, and m and n are positive integers; an obtaining module configured to obtain routing information, where the routing information includes n routing sub-information corresponding one-to-one to the n feature vectors, and each routing sub-information is used to indicate the m network models into which the corresponding token feature vector is input; a conversion module configured to convert the routing information to obtain converted routing information, where the converted routing information is used to indicate, in units of network models, the linear index positions corresponding to the token feature vectors input into each network model, and the linear index positions are used to indicate the positions of the corresponding token feature vectors in the first tensor; a first operation module configured to perform a first permutation operation based on the converted routing information to aggregate the n token feature vectors included in the first tensor according to the network model, where the token feature vectors input into the same network model form a tensor as a whole and are stored in the memory.

[0019] At least one embodiment of the present disclosure provides a processor, including an instruction parsing unit, an execution unit, a memory, and a cache. Wherein, the instruction parsing unit is configured to receive and parse data processing instructions, and the data processing instructions include an input tensor and routing information as input parameters. The data processing instructions are used to sort the input tensor based on the routing information to aggregate multiple feature vectors included in the input tensor according to a network model, and store the sorted tensor obtained after sorting into the memory. The routing information includes n routing sub-information corresponding one-to-one to the n feature vectors, and each routing sub-information is used to indicate m network models for input of the corresponding token feature vector; after the instruction parsing unit parses the data processing instructions, the execution unit executes the data processing instructions. When the execution unit executes the data processing instructions, the following operations are included: reading a first tensor from the memory to the cache, where the first tensor includes n token feature vectors, each token feature vector is input to m different network models for subsequent processing, the first tensor is continuously stored in the memory, the first tensor includes at least part of the data in the input tensor, and each token feature vector is at least part of the content of a feature vector in the input tensor, and m and n are positive integers; converting the routing information to obtain converted routing information, where the converted routing information is used to indicate, in units of network models, the linear index positions corresponding to the token feature vectors input to each network model, and the linear index positions are used to indicate the positions of the corresponding token feature vectors in the first tensor; based on the converted routing information, performing a first permutation operation to aggregate the n token feature vectors included in the first tensor according to the network model, where the token feature vectors input to the same network model form a tensor as a whole and are stored in the memory.

[0020] At least one embodiment of the present disclosure provides an electronic device, including: a memory that stores computer-executable instructions non-transiently; a processor configured to run the computer-executable instructions, where the computer-executable instructions, when run by the processor, implement the data processing method according to any one of the embodiments of the present disclosure.

[0021] At least one embodiment of the present disclosure provides a non-transient computer-readable storage medium, where the non-transient computer-readable storage medium stores computer-executable instructions, and the computer-executable instructions, when executed by a processor, implement the data processing method according to any one of the embodiments of the present disclosure.

[0022] In the data processing method provided by at least one embodiment of the present disclosure, the first tensor is first read into the cache, and intermediate operations are all performed in the cache. Finally, the token feature vectors input to the same network model form a tensor and are stored in the memory as a whole. That is, when stored in the memory, the token feature vectors have been arranged, that is, the token feature vectors are aggregated according to the network model. For the first tensor, only one data loading from the memory is required, without repeated loading. Compared with the current implementation method of the arrangement module, the number of memory accesses can be significantly reduced, the access overhead can be reduced, and the on-chip cache resources can be fully utilized, reducing resource waste.

[0023] Moreover, in the present disclosure, the first tensor is continuously stored in the memory. When reading the first tensor from the memory into the cache, consecutive memory blocks are accessed. For example, consecutive read and write operations can be performed in units of storage cells, which can significantly improve the bandwidth utilization rate. Further, the present disclosure converts the routing information and performs a sorting operation according to the converted routing information. The converted routing information is in units of network models and indicates the linear index positions corresponding to the token feature vectors input to each network model. The converted routing information can intuitively provide a reference for aggregating the token feature vectors according to the network model. Based on this reference, the token feature vectors input to the same network model can form a tensor and be stored in the memory as a whole, so that consecutive memory blocks are also accessed when writing to the memory. For example, consecutive read and write operations can be performed in units of storage cells, significantly improving the bandwidth utilization rate. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] In order to more clearly illustrate the technical solutions of the embodiments of the present disclosure, the drawings of the embodiments will be briefly introduced below. Obviously, the drawings in the following description only relate to some embodiments of the present disclosure and do not limit the present disclosure.

[0025] Figure 1 It is a schematic structural diagram of a general-purpose graphics processing unit (GPGPU); Figure 2 It is a schematic structural diagram of a MOE layer; Figure 3 It is a schematic diagram of the routing information output by the routing module; Figure 4A It is a schematic diagram of the process of the arrangement module; Figure 4B It is a schematic diagram of the process of the inverse arrangement module; Figure 5 It is a schematic flowchart of the data processing method provided by at least one embodiment of the present disclosure; Figure 6 It is a schematic diagram of the relationship between the first tensor and the input tensor provided by at least one embodiment of the present disclosure; Figure 7A It is a schematic diagram of the converted routing information provided by an embodiment of the present disclosure; Figure 7B Schematic diagram of the converted routing information provided by another embodiment of the present disclosure; Figure 8 Schematic diagram of the first arrangement operation provided by at least one embodiment of the present disclosure; Figure 9 Schematic process diagram of the second arrangement operation provided by at least one embodiment of the present disclosure; Figure 10 Schematic block diagram of a data processing device provided by at least one embodiment of the present disclosure; Figure 11 Schematic structural diagram of a processor provided by at least one embodiment of the present disclosure; Figure 12 Schematic block diagram of an electronic device provided by an embodiment of the present disclosure; Figure 13 Schematic diagram of a non-transitory computer-readable storage medium provided by at least one embodiment of the present disclosure. Detailed implementation manners

[0026] In order to make the objectives, technical solutions, and advantages of the embodiments of the present disclosure clearer, the technical solutions of the embodiments of the present disclosure will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present disclosure. Obviously, the described embodiments are some, but not all, of the embodiments of the present disclosure. All other embodiments obtained by those of ordinary skill in the art based on the described embodiments of the present disclosure without creative efforts shall fall within the protection scope of the present disclosure.

[0027] Unless otherwise defined, the technical terms or scientific terms used in the present disclosure shall have the ordinary meanings understood by those of ordinary skill in the art to which the present disclosure pertains. The "first", "second", and similar terms used in the present disclosure do not denote any order, quantity, or importance, but are only used to distinguish different components. Words such as "including" or "comprising" mean that the elements or items appearing before the word cover the elements or items listed after the word and their equivalents, without excluding other elements or items. Words such as "connected" or "coupled" are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect. "Up", "down", "left", "right", etc. are only used to represent relative positional relationships, and when the absolute position of the object being described changes, the relative positional relationship may also change accordingly. To keep the following description of the embodiments of the present disclosure clear and concise, some details of known functions and known components are omitted in the present disclosure.

[0028] Figure 1 Schematic structural diagram of a general-purpose graphics processing unit (GPGPU).

[0029] As Figure 1 shown, a general-purpose graphics processing unit is actually an array of programmable multi-processors. For example, the programmable multi-processor can be a Streaming Processor Cluster (SPC), such as including Figure 1 the streaming processor clusters 1 shown, ..., streaming processor clusters M, where M is a positive integer greater than 1. In the general-purpose graphics processing unit, one streaming processor cluster processes one computing task, or multiple streaming processor clusters process one computing task. Data sharing occurs between multiple streaming processor clusters through global caches or global memory.

[0030] As Figure 1 shown, taking the streaming processor cluster 1 as an example, one streaming processor cluster includes multiple computing units, such as Figure 1 the computing units 1, computing unit 2, ..., computing unit N in, where N is a positive integer. Each computing unit (Compute Unit, CU for short) is used to perform arithmetic and logical operations, such as accumulation, reduction, conventional addition, subtraction, multiplication, division, etc. A computing unit includes multiple cores (also called computing cores or calculation cores), and each core includes an arithmetic logic unit (ALU), a floating-point computing unit, etc. The core is used to perform specific computing tasks. In addition, the computing unit also includes registers (such as Figure 1 the register file in) and shared memory, which are used to hierarchically store source data and destination data related to the computing task. The shared memory in a computing unit is used to share data between the cores of the computing unit.

[0031] In parallel computing, computing tasks are generally executed by multiple threads. These threads are divided into multiple thread blocks before execution in the general-purpose graphics processing unit (or called a parallel computing processor), and then multiple thread blocks are distributed to each computing unit via a thread block distribution module ( Figure 1 not shown in). All threads in a thread block must be assigned to the same computing unit for execution. At the same time, the thread block is split into the smallest execution thread bundles (or simply called thread bundles, warps), and each thread bundle contains a fixed number (or less than this fixed number) of threads. For example, 32 threads. Multiple thread blocks can be executed in the same computing unit or in different computing units.

[0032] In each computing unit, the thread bundle scheduling / distribution module ( Figure 1The warp is scheduled and allocated (not shown in the figure) so that multiple computing cores in the computing unit can run the warp. According to the number of computing cores in the computing unit, multiple warps in a thread block can be executed simultaneously or time-shared. Multiple threads in each warp will execute the same instructions. Memory execution instructions will be issued to the shared memory in the computing unit or further issued to the middle-level cache, global cache, or global memory for read and write operations, etc.

[0033] The Transformer model is a classic neural network-based text processing model that has spawned important models such as BERT (Bidirectional Encoder Representation from Transformers), GPT (Generative Pre-trained Transformer), and LLM (Large Language Model), which have significantly promoted the development of the NLP (Natural Language Processing) field.

[0034] The MOE structure is a very important structure in the Transformer model. It makes the model perform better by allocating different tokens (also known as token feature vectors) to appropriate experts.

[0035] Figure 2 It is a schematic diagram of the structure of an MOE layer.

[0036] As Figure 2 shown, the MOE layer includes multiple expert networks and a routing module.

[0037] The expert network is the core computing unit in the MOE layer. Each expert network is an independent neural network, usually in the form of a feed-forward network (FFN). Each expert network is designed to focus on processing specific tasks or data features. For example, in natural language processing, some expert networks may focus on syntactic analysis, while others focus on semantic understanding.

[0038] The role of the routing module is to determine which expert networks the input each token should be sent to for processing. For example, the number of expert networks each token is sent to can be preset, such as 2, 4, 6, 8, etc.

[0039] Figure 3 It is a schematic diagram of the routing information output by the routing module. As Figure 3As shown, the routing information output by the routing module is represented as [total_token_num, m], where total_token_num is the total number of all tokens in an input, and m represents to which expert networks each token is assigned.

[0040] As Figure 3 shown, 0 and 2 in the first row indicate that the first token is assigned to expert network 0 and expert network 2, 2 and 1 in the second row indicate that the second token is assigned to expert network 2 and expert network 1, 3 and 1 in the third row indicate that the third token is assigned to expert network 3 and expert network 1, and so on.

[0041] In addition, as Figure 2 shown, the MOE layer also includes a permutation module and an inverse permutation module.

[0042] The permutation module aggregates the feature vectors of the tokens according to the routing information by expert network. Specifically, the permutation module traverses the m expert networks to which each token will be input according to the indexes in the routing information, aggregates the feature vectors of each token by expert network, and stores them in memory (such as HBM, high bandwidth memory). Aggregating by expert network, for example, aggregates the feature vectors of expert network 0 together, the feature vectors of expert network 1 together, and so on. The feature vectors of n expert networks are arranged in a tensor by expert network and stored in memory. After that, the feature vectors assigned to each expert network can be distributed to different expert networks for processing.

[0043] Since the feature vectors of each token need to be input to 2 expert networks, this process will also generate copy operations. For example, the feature vectors of the first token are copied twice, one for input to expert network 0 and one for input to expert network 2.

[0044] Figure 4A is a schematic diagram of the process of the permutation module.

[0045] The feature vectors of all tokens of a certain input are stored in order in memory, for example, called the input tensor. For example, the first row of the input tensor is the feature vector of the first token, the second row is the feature vector of the second token, and so on. Through the permutation module, the feature vectors of the tokens at the corresponding positions in the input tensor are read from memory to the cache, and then written to the corresponding positions in the permuted tensor according to the routing information.

[0046] Take Figure 3Taking the routing information shown as an example, the feature vectors of each token are assigned to two expert networks for processing. The feature vector of the first token is assigned to expert network 0 and expert network 2. Specifically, when implemented, the feature vector of the first token is extracted from the memory and stored in the cache, and then the feature vector of the first token is written to the corresponding position in the permuted tensor in the memory. Refer to Figure 4A , the feature vector of the first token is written to the position of the first black-on-white box in the memory representing expert network 0, and at the same time, the feature vector of the first token is also written to the position of the first diagonal shaded box in the memory representing expert network 2.

[0047] Refer to Figure 4A , the feature vector of the third token is assigned to expert network 3 and expert network 1. Specifically, when implemented, the feature vector of the third token is extracted from the memory and stored in the buffer area, and then the feature vector of the third token is written to the corresponding position in the permuted tensor in the memory. Refer to Figure 4A , the feature vector of the third token is written to the position of the first white-on-black box in expert network 1 in the memory, and at the same time, the feature vector of the third token is also written to the position of the first grid shaded box in expert network 3 in the memory.

[0048] Taking the high-bandwidth memory of the graphics processor as an example, in the high-bandwidth memory, each data read and write is processed at the granularity of a storage unit (BLOCK). The storage capacity of the storage unit is 512B. Even if the actual amount of data to be read / written is less than the size of a storage unit, the entire storage unit will be read / written together.

[0049] Refer to Figure 4A As shown in the implementation method of the permutation module, it sequentially assigns individual tokens to different expert networks according to the indexes of the expert networks indicated in the routing information. Therefore, in each write process, due to the uncertain write position, and the physical positions of the two memories where the feature vector of a certain token is written each time are much larger than the granularity of the storage unit. For example, the feature vector of the first token is written to the first position of expert network 0 and the first position of expert network 2, and the span of the two positions is very large. Therefore, the current implementation method of the permutation module will cause other data adjacent to the required data that is currently not required to be written to the memory to be carried each time when writing to the memory. For example, when writing to the first position of expert network 0, other data that is not required at this time in the storage unit containing the feature vector of the first token is also written at the same time, which results in low bandwidth utilization and ineffective use of the data bandwidth.

[0050] Moreover, since the intervals between the two physical locations written are large, other data in the read cache except for the required data cannot be used currently. By the time these data can be used, they have been flushed out of the buffer, resulting in low cache utilization, wasting resources. Also, each time a feature vector of a certain token is written to memory, a memory read / write operation is required, leading to high read / write overhead. The permutation module belongs to the memory bound operator and will have poor performance if the bandwidth and on-chip cache cannot be effectively utilized.

[0051] Figure 4B It is a schematic diagram of the process of the inverse permutation module.

[0052] As Figure 4B shown, in the result tensor, the processing results of expert network 0 are aggregated and permuted, the processing results of expert network 1 are aggregated and permuted, and so on. When implementing the inverse permutation module, it traverses the token direction of the routing information according to the routing information, extracts the corresponding results from the expert networks indicated in the routing information of each token, and writes them to the corresponding positions in the output tensor after calculations such as accumulation.

[0053] For example, as Figure 4B shown, the first result of expert network 0 and the first result of expert network 2 in the result tensor are read into the cache, and after calculations such as accumulation, they are written to the first position of the output tensor as the processing result of the first token. The first result of expert network 1 and the first result of expert network 3 in the result tensor are read into the cache, and after calculations such as accumulation, they are written to the third position of the output tensor as the processing result of the third token. And so on.

[0054] Similar to the implementation process of the permutation module, the data reading of the inverse permutation module has a large span. For example, it is necessary to read the first result of expert network 0 and the first result of expert network 2 for operations such as accumulation. Similar to the permutation module, a large span during reading will lead to low bandwidth utilization, inability to effectively utilize the data bandwidth, low cache utilization, and waste of resources. Since the inverse permutation module also belongs to the memory bound operator, it will have poor performance if the bandwidth and on-chip cache cannot be effectively utilized.

[0055] At least one embodiment of the present disclosure provides a data processing method, apparatus, electronic device, and non-transitory storage medium. The data processing method is used to sort an input tensor to aggregate multiple feature vectors included in the input tensor according to a network model, and store the sorted tensor obtained after sorting in a memory. The data processing method includes: reading a first tensor from the memory to a cache, where the first tensor includes n token feature vectors, each token feature vector is input to m different network models for subsequent processing, the first tensor is continuously stored in the memory, the first tensor includes at least part of the data in the input tensor, and m and n are positive integers; obtaining routing information, where the routing information includes n routing sub-information corresponding one-to-one to the n token feature vectors, and each routing sub-information is used to indicate the m network models input by the corresponding token feature vector; converting the routing information to obtain converted routing information, where the converted routing information is used to indicate, in units of network models, the linear index positions corresponding to the token feature vectors input to each network model, and the linear index position is used to indicate the position of the corresponding token feature vector in the first tensor; based on the converted routing information, extracting the token feature vectors input to each network model from the cache, and the token feature vectors input to the same network model after extraction are combined into a tensor as a whole and stored in the memory.

[0056] In the data processing method provided by at least one embodiment of the present disclosure, the first tensor is first read into the cache, and intermediate operations are all performed in the cache. Finally, the token feature vectors input to the same network model are combined into a tensor and stored in the memory as a whole. That is, when storing in the memory, the token feature vectors have completed the arrangement, that is, the token feature vectors are aggregated according to the network model. For the first tensor, only one data loading from the memory is required, without repeated loading. Compared with the current implementation method of the arrangement module, the number of memory accesses can be greatly reduced, the access overhead can be reduced, and the on-chip cache resources can be fully utilized, reducing resource waste.

[0057] Moreover, in the present disclosure, the first tensor is continuously stored in the memory. When reading the first tensor from the memory to the cache, consecutive memory blocks are accessed. For example, consecutive read and write operations can be performed in units of storage cells, which can greatly improve the bandwidth utilization rate. Further, the present disclosure converts the routing information and performs a sorting operation according to the converted routing information. The converted routing information indicates, in units of network models, the linear index positions corresponding to the token feature vectors input to each network model. The converted routing information can intuitively provide a reference for aggregating the token feature vectors according to the network model. Based on this reference, the token feature vectors input to the same network model can be combined into a tensor as a whole and stored in the memory, so that consecutive memory blocks are also accessed when writing to the memory. For example, consecutive read and write operations can be performed in units of storage cells, greatly improving the bandwidth utilization rate.

[0058] The data processing method provided by the embodiments of the present disclosure can be applied to the data processing device provided by the embodiments of the present disclosure, and the data processing device can be configured on an electronic device. The electronic device can be a personal computer, a mobile terminal, etc., and the mobile terminal can be a hardware device such as a mobile phone, a tablet computer, a laptop computer, etc.

[0059] The embodiments of the present disclosure will be described in detail below with reference to the accompanying drawings, but the present disclosure is not limited to these specific embodiments.

[0060] For example, the data processing method provided by at least one embodiment of the present disclosure is used to sort an input tensor to aggregate multiple feature vectors included in the input tensor according to a network model, and store the sorted tensor obtained after sorting in a memory.

[0061] For example, in one embodiment, the data processing method is used for data processing of a mixture-of-experts model structure, and the mixture-of-experts model structure can refer to Figure 2 . As Figure 2 shown, the mixture-of-experts model structure includes multiple expert networks as multiple network models and a routing module, and the routing module determines routing information according to the input tensor input to the mixture-of-experts model structure. Regarding the routing module and routing information, reference can be made to the foregoing content and will not be elaborated here.

[0062] Each expert network is trained to process specific tasks and data features. For example, in the field of natural language processing, some expert networks may focus on syntactic analysis, while other expert networks focus on semantic understanding.

[0063] For example, the routing module determines the activation probability of each expert network according to the input tensor. According to these probabilities, the routing module will select a few expert networks to process the input tensor, and perform weighted merging on the processing results based on the activation probabilities.

[0064] For example, the input parameter of the MOE structure, that is, the input tensor, which can include multiple feature vectors, is an intermediate result obtained by processing by the network structure before the MOE structure.

[0065] The sorted tensor obtained after sorting can be aggregated in units of network models. For example, the feature vectors input to the same network model are put together and stored in the memory. Here, reference can be made to the foregoing Figure 4A related description and will not be elaborated here.

[0066] Of course, the present disclosure is not limited to this, and network models of other network structures with similar functions and structures can also use the data processing method provided by the present disclosure.

[0067] The data processing method provided by at least one embodiment of the present disclosure can be applied to application fields involving neural networks with a similar MOE structure, such as natural language processing (NLP), computer vision (CV), recommendation systems, speech processing, reinforcement learning, edge computing and resource optimization, medical diagnosis, etc.

[0068] Figure 5 It is a schematic flowchart of the data processing method provided by at least one embodiment of the present disclosure.

[0069] As Figure 5 shown, the data processing method includes steps S10 - S40.

[0070] As Figure 5 shown, in step S10, the first tensor is read from the memory to the cache.

[0071] For example, the first tensor includes n token feature vectors, each token feature vector is input into m different network models for subsequent processing, the first tensor is continuously stored in the memory, the first tensor includes at least part of the data in the input tensor, and m and n are positive integers.

[0072] For example, since the first tensor needs to be read to the cache first, and when the cache capacity permits, the first tensor can be the input tensor itself.

[0073] For example, assume that the cache capacity is small and the entire input tensor cannot be stored in the cache. Then the input tensor can be divided into blocks, and each time a block is loaded into the cache as the first tensor for processing. Therefore, the present disclosure is beneficial to the processing of long text scenarios and has a wider application range.

[0074] For example, the data processing method provided by at least one embodiment of the present disclosure further includes: determining the block size based on the cache capacity; dividing the input tensor into multiple blocks (Tiles) with the block size as the unit, where each block includes n token feature vectors, and any one token feature vector in each block is at least part of the feature vectors in the input tensor.

[0075] For example, the block size can be determined on the premise of making the best use of the cache capacity as much as possible, and a block needs to be able to be completely stored in the cache. The block size can be determined according to the cache capacity, and the present disclosure does not make specific limitations on this.

[0076] For example, in some embodiments, each hardware execution unit processes the relevant processing of a block, so that multiple hardware execution units can perform the relevant processing of multiple first tensors, such as the data processing method provided by at least one embodiment of the present disclosure, in parallel.

[0077] For example, in one embodiment, refer to Figure 2For the general graphics processing unit shown, the hardware execution unit can be a programmable multi-processor, such as a stream processor cluster. The hardware execution unit includes multiple computing units, each computing unit includes multiple execution units, and each execution unit includes multiple thread hardwares, and one thread runs in one thread hardware. Each hardware execution unit includes a vector operation unit buffer for caching data of the thread hardwares in the hardware execution unit. For example, N computing units belonging to one hardware execution unit can share the vector operation unit buffer. The vector operation unit buffer has a caching mechanism different from that of a conventional L1 cache or L2 cache. The vector operation unit buffer is closer to the computing core than the memory and can provide a data storage function. For example, each hardware execution unit can read the first block from the memory and cache it in its own vector operation unit buffer. The block size can be determined according to the capacity of the vector operation unit buffer.

[0078] Figure 6 Schematic diagram of the relationship between the first tensor and the input tensor provided by at least one embodiment of the present disclosure.

[0079] As Figure 6 shown, the input tensor includes multiple feature vectors. The multiple white boxes in the input tensor represent one feature vector, and the bit width of each feature vector is L'. The first tensor includes part of the content of n feature vectors in the input tensor. As Figure 6 shown, each token feature vector in the first tensor includes L bits, where L is less than L'. For example, Figure 6 the shaded part in

[0080] is a token feature vector, which is part of the corresponding feature vector in the input tensor. Both L and L' are positive integers. The block size L and n can be determined according to the cache capacity.

[0081] For example, the processing completion here can mean that all the data in the currently cached first tensor has been stored in the memory. The specific process can be referred to the following description.

[0082] In at least one embodiment of the present disclosure, setting an appropriate block size according to the size of the on-chip cache and dividing the input tensor into multiple first tensors for sequential processing is more beneficial for processing long text scenarios. And since the input tensor has been divided into multiple blocks for batch processing, in a distributed scenario, regardless of the slicing strategy in each dimension, it can be uniformly processed, having the advantage of generalization adaptation in a distributed scenario.

[0083] For example, in step S20, routing information is obtained.

[0084] For example, the routing information includes n routing sub-information corresponding one-to-one to n token feature vectors, and each routing sub-information is used to indicate m network models into which the corresponding token feature vector is input.

[0085] For example, referring to Figure 3 the shown routing information, one line represents one routing sub-information, each line corresponds to one feature vector, and describes which network models the corresponding feature vector can be input into. Similarly, one line corresponds to one token feature vector, and the token feature vector can include at least some bit positions in the feature vector. For example, m can take values such as 2, 4, 6, 8, etc., and the present disclosure does not make specific limitations on this. More content about the routing information can be referred to the foregoing description and will not be elaborated here.

[0086] For example, in step S30, the routing information is converted to obtain the converted routing information.

[0087] For example, the converted routing information is used to indicate, in units of network models, the linear index positions corresponding to the token feature vectors input into each network model. For example, the linear index position is used to indicate the position of the corresponding token feature vector in the first tensor.

[0088] For example, in some embodiments, step S30 may include: based on each routing sub-information, storing the linear index positions corresponding to the n token feature vectors into the array corresponding to the allocated network model respectively to obtain the converted routing information.

[0089] For example, each routing sub-information includes m model identification information, which is used to indicate the identification numbers of the m network models into which the corresponding token feature vector is input. Taking Figure 3 as an example, each routing sub-information includes 2 model identification information. For example, taking Figure 3 the first routing sub-information in the first row in as an example, the two model identification information are 0 and 2 respectively, indicating that the first token feature vector is input into expert network 0 and expert network 2; the two model identification information in the second routing sub-information in the second row are 2 and 1 respectively, indicating that the second token feature vector is input into expert network 2 and expert network 1.

[0090] For example, in some embodiments, based on each routing sub-information, storing the linear index positions respectively corresponding to n token feature vectors into an array corresponding to the allocated network model to obtain the transformed routing information may include: traversing the n routing sub-informations, for any one of the routing sub-informations, determining the linear index position corresponding to the token feature vector based on the index position of the token feature vector corresponding to any one of the routing sub-informations in the first tensor, wherein the linear index position corresponding to the token feature vector includes m linear index positions corresponding one-to-one to m model identification information; for any one of the m linear index positions, writing any one of the linear index positions into the array corresponding to the network model indicated by the corresponding model identification information.

[0091] For example, each network model may correspond to an array, and the array includes a plurality of elements arranged in sequence, which can be used to store / sequentially store the linear index positions corresponding to the token feature vectors to be input into the network model.

[0092] For example, the linear index position corresponding to the token feature vector includes m linear index positions corresponding one-to-one to m model identification information. Since the token feature vector may need to be copied and sent to different network models, the linear index position may include m linear index positions corresponding one-to-one to m model identification information.

[0093] For example, in one embodiment, the m linear index positions corresponding to the same token feature vector may be different so as to be able to distinguish different model identification information during subsequent weighting. Assume that the index position of the r-th token feature vector in the first tensor in the first tensor is r - 1, and the m linear index positions corresponding to the r-th token feature vector may be (r - 1)×m, (r - 1)×m + 1,..., r×m + 1, so as to be able to indicate the absolute index positions of the token feature vectors input into different network models, where r is a positive integer.

[0094] For example, in this embodiment, the m linear index positions corresponding to the r-th token feature vector are (r - 1)×m, (r - 1)×m + 1,..., r×m + 1. If the r-th token feature vector needs to be written into network models 1, 2,.., m, then (r - 1)×m can be written into the array corresponding to network model 1, (r - 1)×m + 1 can be written into the array corresponding to network model 2,..., and r×m + 1 can be written into the array corresponding to network model m.

[0095] For example, in another embodiment, the m linear index positions corresponding to the same token feature vector may be the same. Suppose the index position of the r-th token feature vector in the first tensor in the first tensor is r - 1, and the m linear index positions corresponding to the r-th token feature vector may all be r - 1, that is, all equal to the index position of the r-th token feature vector in the first tensor, which is r - 1.

[0096] Figure 7A Schematic diagram of the converted routing information provided by an embodiment of the present disclosure.

[0097] For example, referring to Figure 3 the routing information shown, the routing information includes multiple routing sub-information. Suppose each routing sub-information includes 2 model identification information, 0 represents network model 0, 1 represents network model 1, 2 represents network model 2, and 3 represents network model 3.

[0098] In Figure 7A , array 0 corresponds to network model 0, array 1 corresponds to network model 1, array 2 corresponds to network model 2, and array 3 corresponds to network model 3. Each array includes multiple elements for sequentially storing linear index positions.

[0099] Taking Figure 3 the routing information shown as an example for conversion, as Figure 3 shown, suppose the routing sub-information (0, 2) in the first row corresponds to the first token feature vector in the first tensor. The linear index positions corresponding to the first token feature vector are 0 and 1. The linear index position corresponding to the model identification information with a value of 0 is 0, and the linear index position corresponding to the model identification information with a value of 2 is 1. Store the linear index position corresponding to the model identification information with a value of 0, that is, 0, into array 0 corresponding to network model 0, and store the linear index position corresponding to the model identification information with a value of 1, that is, 1, into array 2 corresponding to network model 2.

[0100] For the second routing sub-information (2, 1) corresponding to the second token feature vector, the linear index positions corresponding to the second token feature vector are 2 and 3. The linear index position corresponding to the model identification information with a value of 2 is 2, and the linear index position corresponding to the model identification information with a value of 1 is 3. Store the linear index position corresponding to the model identification information with a value of 2, that is, 2, into array 2 corresponding to network model 2, and store the linear index position corresponding to the model identification information with a value of 1, that is, 3, into array 1 corresponding to network model 1.

[0101] For the third routing sub - information (3,1) corresponding to the third token feature vector, the linear index positions corresponding to the third token feature vector are 4 and 5, the linear index position corresponding to the model identification information with a value of 3 is 4, and the linear index position corresponding to the model identification information with a value of 1 is 5. The linear index position corresponding to the model identification information with a value of 3, that is, 4, is stored in array 3 corresponding to network model 3, and the linear index position corresponding to the model identification information with a value of 1, that is, 5, is stored in array 1 corresponding to network model 1.

[0102] The same applies to other routing sub - information and will not be elaborated here.

[0103] Thus, the routing information is converted into converted routing information in units of network models, which indicates the linear index positions corresponding to the token feature vectors input to each network model. The converted routing information intuitively provides a reference for aggregating token feature vectors according to network models. Based on this reference, the token feature vectors input to the same network model can be stored as a tensor whole in memory, so that when writing to memory, continuous memory blocks are also accessed. For example, continuous read - write can be performed in units of storage cells, greatly improving bandwidth utilization.

[0104] For example, in some embodiments, before storing the linear index position corresponding to each token feature vector into the array corresponding to the allocated network model based on each routing sub - information, the data - processing method provided by at least one embodiment of the present disclosure further includes: filling multiple elements included in all arrays in the converted routing information with a preset value, where the preset value is used to indicate that the element is invalid.

[0105] For example, first initialize all arrays, that is, write the preset value to all elements in the arrays. For example, the preset value can be INF, indicating that the element is invalid. Then, based on each routing sub - information, store the linear index position corresponding to each token feature vector into the array corresponding to the allocated network model. Since the arrangement of the model identification information in the first tensor is not determined in advance, the lengths of the arrays are also not determined. Therefore, the lengths of the arrays can be set longer and the arrays can be initialized first to make the redundant elements invalid, so as to avoid incorrect extraction when using the converted routing information later.

[0106] For example, assume that the first tensor is part of the input tensor, for example, a block of the input tensor. Then, every n routing sub - information is converted into a converted routing information.

[0107] For example, the routing information includes a plurality of routing sub-information corresponding one-to-one to a plurality of feature vectors, for example, including total_token_num routing sub-information. For example, in some embodiments, based on the chunk size, the total_token_num routing sub-information is divided into multiple groups, where each group includes n routing sub-information; a conversion operation is performed on each group of routing information to obtain the converted routing information corresponding to each group.

[0108] Figure 7B FIG. is a schematic diagram of the converted routing information provided by another embodiment of the present disclosure. As Figure 7B shown, it is assumed that the n routing sub-information covered by the upper brace is the routing sub-information corresponding to the first to nth feature vectors in the input tensor. After processing it with reference to the above step S30, the converted routing information 1 is obtained. The specific process is as Figure 7A shown, and this converted routing information 1 is used for the first permutation operation and the second permutation operation related to the first to nth feature vectors.

[0109] It is assumed that the n routing sub-information covered by the lower brace is the routing sub-information corresponding to the (n + 1)th to 2×nth feature vectors in the input tensor. After processing it, the converted routing information 2 is obtained. The specific process is similar to the Figure 7A process described above and will not be elaborated here. This converted routing information 2 is used for the first permutation operation and the second permutation operation related to the (n + 1)th to 2×nth feature vectors.

[0110] The subsequent process is similar to the above and will not be elaborated here.

[0111] For example, each first tensor has its own corresponding converted routing information. It is assumed that the first to nth feature vectors in the input tensor are chunked into L1 first tensors, where L1 = L’ / L, and these L1 first tensors all correspond to the converted routing information 1.

[0112] For example, in step S40, based on the converted routing information, a first permutation operation is performed to aggregate the n token feature vectors included in the first tensor according to the network model.

[0113] For example, the token feature vectors input to the same network model are combined into a tensor as a whole and stored in the memory.

[0114] For example, in some embodiments, step S40 may include: based on the converted routing information, for each network model, extracting the token feature vectors input to the network model from the cache, and combining the extracted token feature vectors input to the network model into a tensor as a whole and storing it in the memory.

[0115] For example, the converted routing information includes v arrays corresponding one-to-one to v network models, and each array includes multiple valid elements for indicating the linear index positions corresponding to the token feature vectors of the input corresponding network models. For example, a valid element is an element whose value is not a preset value. For example, v is the total number of network models in the network, and v is a positive integer.

[0116] For example, in some embodiments, based on the converted routing information, the token feature vectors input to each network model are extracted from the cache, and the token feature vectors input to the same network model that are extracted form a tensor and are stored in the memory as a whole. This may include: for any one of the v arrays, traversing each valid element in the array, extracting the token feature vector indicated by each valid element from the cache and caching it in a buffer component, and in response to traversing all the valid elements in any one of the arrays, storing the token feature vectors that have been extracted in the buffer component in the memory as a tensor as a whole.

[0117] For example, traverse the v arrays, read each array in sequence, and for each array, read the valid elements in the array in sequence, that is, the elements that are not preset values. For each valid element read, extract the corresponding data from the first tensor in the cache according to the valid element and temporarily store it in a buffer component (such as a buffer Buffer or an extended thread-local register, Thread-Local-Register, abbreviated as TLR). After reading all the valid elements in an array, store the data that has been extracted in the buffer component in the memory as a tensor as a whole.

[0118] For example, the corresponding data extracted here can be a complete feature vector in the input tensor or partial bits in the feature vector according to different block sizes.

[0119] For example, in some embodiments, extracting the token feature vectors indicated by each valid element from the cache and caching them in a buffer component may include: for each valid element, based on the linear index position indicated by the valid element, determining the index position of the token feature vector corresponding to the linear index position in the first tensor; and extracting the corresponding token feature vector from the cache based on the index position and caching it in the buffer component.

[0120] For example, the result of dividing the linear index position by m and rounding down is used as the index position of the token feature vector corresponding to the linear index position in the first tensor.

[0121] For example, Figure 7A taking [as an example], for array 0, first read the valid element 0 from it. According to the valid element 0, determine that the index position of the token feature vector corresponding to the linear index position 0 in the first tensor is = 0, then extract the data at index position 0 in the first tensor (e.g., the first token feature vector in the first tensor) from the cache and store it in the buffer component.

[0122] Then, read the next valid element in array 0, which is 7. According to the valid element 7, determine that the index position of the token feature vector corresponding to the linear index position 7 in the first tensor is = 3, then extract the data at index position 3 in the first tensor (e.g., the fourth token feature vector in the first tensor) from the cache and store it in the buffer component.

[0123] Then, read the next valid element in array 0, which is 14. According to the valid element 14, determine that the index position of the token feature vector corresponding to the linear index position 14 in the first tensor is = 7, then extract the data at index position 7 in the first tensor (e.g., the eighth token feature vector in the first tensor) from the cache and store it in the buffer component.

[0124] Then, read the next valid element in array 0 and perform the above operations until the first array is traversed. After that, store the data stored in the buffer component as a whole tensor in the memory. At this time, the aggregation of the data of the first tensor input to the network model 0 is completed, and the tensor is continuously stored in the memory. Therefore, continuous reading and writing with the storage unit as the granularity can greatly improve the bandwidth utilization rate.

[0125] After that, for array 1, first read the valid element 3 from it. According to the valid element 3, determine that the index position of the token feature vector corresponding to the linear index position 3 in the first tensor is = 1, then extract the data at index position 1 in the first tensor (e.g., the second token feature vector in the first tensor) from the cache and store it in the buffer component.

[0126] Then, read the next valid element in array 0, which is 5. According to the valid element 5, determine that the index position of the token feature vector corresponding to the linear index position 5 in the first tensor is = 2, then extract the data at index position 2 in the first tensor (e.g., the third token feature vector in the first tensor) from the cache and store it in the buffer component.

[0127] The subsequent process is carried out in the same way and will not be described repeatedly.

[0128] As described above, when the first tensor is a block of the input tensor, multiple routing sub-information corresponding to multiple eigenvectors is divided into multiple groups, and each group includes n routing sub-information. In memory, for any network model, the token feature vectors input to the any network model extracted based on different groups of routing information are continuously stored in memory. That is to say, in the process of storing tensors extracted based on different blocks into memory, they are stored in memory according to the positional relationship between the blocks, and the token feature vectors used to input the same network model are continuously stored in memory.

[0129] Figure 8 Schematic diagram of the first permutation operation provided by at least one embodiment of the present disclosure.

[0130] As Figure 8 shown, the routing information is divided into multiple groups according to n in the block size, each group includes n routing sub-information, and every n routing sub-information is converted according to the process of step S30 above to obtain the converted routing information. For example, the first group of routing information is converted to obtain the converted routing information 1, and the second group of routing information is converted to obtain the converted routing information 2. The same applies to other groups of routing information, which will not be elaborated here one by one.

[0131] After that, referring to the process described above, as Figure 8 shown, for the token feature vectors input to network model 0, tensor 0_0 input to network model 0 is obtained through array 0 in the converted routing information 1, and tensor 1_0 input to network model 0 is obtained through array 0 in the converted routing information 2. Although tensors 0_0 and 1_0 are stored in memory in batches, tensors 0_0 and 1_0 are continuously stored in memory.

[0132] Similarly, as Figure 8 shown, for the token feature vectors input to network model 1, tensor 0_1 input to network model 0 is obtained through array 1 in the converted routing information 1, and tensor 1_1 input to network model 0 is obtained through array 1 in the converted routing information 2. Although tensors 0_1 and 1_1 are stored in memory in batches, tensors 0_1 and 1_1 are continuously stored in memory.

[0133] As Figure 8 shown, for the token feature vectors input to network model 2, tensor 0_2 input to network model 0 is obtained through array 2 in the converted routing information 1, and tensor 1_2 input to network model 0 is obtained through array 2 in the converted routing information 2. Although tensors 0_2 and 1_2 are stored in memory in batches, tensors 0_2 and 1_2 are continuously stored in memory.

[0134] As Figure 8As shown, for the token feature vectors of the input network model 3, the tensor 0_3 input to the network model 0 is obtained through the array 3 in the transformed routing information 1, and the tensor 1_3 input to the network model 0 is obtained through the array 3 in the transformed routing information 2. Although the tensor 0_3 and the tensor 1_3 are stored in batches in the memory, they are continuously stored in the memory.

[0135] In addition, when partitioning the input tensor, assuming that each feature vector is divided into L1 parts, that is, n feature vectors are divided into L1 blocks. When these L1 blocks are stored in the memory, they are concatenated according to the positional relationship during partitioning to form a feature vector with a length of L' bits.

[0136] Thus, in the present disclosure, the first tensor is continuously stored in the memory. When reading the first tensor from the memory to the cache, consecutive memory blocks are accessed. For example, consecutive read and write operations can be performed at the granularity of storage units, which can greatly improve the bandwidth utilization rate. Moreover, in the first permutation operation, based on the intuitive reference for aggregating token feature vectors according to the network model provided by the transformed routing information, the token feature vectors input to each network model are first extracted and cached in the buffer component, and then the token feature vectors input to the same network model in the first tensor are stored in the memory as a whole. Therefore, consecutive memory blocks are also accessed when writing to the memory. For example, consecutive read and write operations can be performed at the granularity of storage units, greatly improving the bandwidth utilization rate.

[0137] Furthermore, the first tensor is first read into the cache, and intermediate operations are all performed in the cache. Finally, the token feature vectors input to the same network model form a tensor and are stored in the memory as a whole. That is, the token feature vectors have been permuted when stored in the memory, that is, the token feature vectors are aggregated according to the network model. For the first tensor, only one data loading from the memory is required, without repeated loading. Compared with the current implementation method of the permutation module, the number of memory accesses can be greatly reduced, the memory access overhead can be reduced, and the on-chip cache resources can be fully utilized, reducing resource waste.

[0138] In addition, combined with the hardware characteristics, the bandwidth utilization rate of the on-chip high-speed storage can be fully exerted, improving the overall performance of the operator.

[0139] Moreover, in the present disclosure, the output of the first permutation operation is directly v independent tensors, which can facilitate subsequent direct input to each network model or other operations without further splitting, reducing programming complexity.

[0140] After that, the sorted tensors obtained through the above process are input to the corresponding network models for processing. For example, they are input to the corresponding expert networks for processing. Each expert network outputs a result tensor, and the result tensor includes multiple result feature vectors. The number of result feature vectors is the same as the number of feature vectors input to the expert network.

[0141] Corresponding to the above first permutation operation, the data processing method provided by at least one embodiment of the present disclosure further includes: performing a second permutation operation on the result tensors respectively output by each network model to merge and sort the result tensors respectively output by each network model and store them in memory, where the result tensors are obtained by processing the feature vectors input to each network model.

[0142] The second permutation operation is the inverse operation of the first permutation operation, that is, merging and reordering the result tensors output by each network model to obtain an output tensor corresponding to the input tensor as the final processing result, and the result feature vectors in the output tensor are arranged in the order of the feature vectors in the input tensor.

[0143] For example, the data processing method provided by at least one embodiment of the present disclosure further includes: based on the converted routing information, performing a second permutation operation on the result tensors respectively output by each network model to merge and sort the result tensors respectively output by each network model and store them in memory, where the result tensors are obtained by processing the feature vectors input to each network model.

[0144] For example, in the second permutation operation, the result feature vectors written into the cache based on the same converted routing information form a tensor, which is stored in memory as a whole, and the tensor is stored continuously in memory.

[0145] The above tensor is obtained based on the same converted routing information, that is, it is the processing result of the same first tensor, and the relative writing position relationship between the token result vectors in the above tensor is the same as the relative position relationship of the token feature vectors corresponding to the token result vectors in the input tensor. Therefore, the tensor can be directly stored in memory as the final result, and when writing to memory, it also accesses continuous memory blocks. For example, it can be continuous read and write with the storage unit as the granularity, which greatly improves the bandwidth utilization rate.

[0146] Similar to the foregoing first permutation operation, in the second permutation operation, the token result vectors have been permuted when stored in memory, that is, the token feature vectors are aggregated according to the network model. For the second tensor, only a limited number of data need to be loaded from memory, without repeated loading. Compared with the current implementation method of the inverse permutation module, it can greatly reduce the number of memory accesses, reduce the memory access overhead, and can make full use of the on-chip cache resources to reduce resource waste.

[0147] For example, in some embodiments, performing a second permutation operation on the result tensors respectively output by each network model based on the transformed routing information may include: reading a second tensor from a memory to a buffer component based on the transformed routing information, where the second tensor is at least part of the result tensor output by one network model; performing a merge sorting process on the second tensor based on the transformed routing information, where in the merge sorting process, each token result vector included in the second tensor is written into a cache based on the transformed routing information, and the relative writing position relationship between each token result vector is the same as the relative position relationship of the token feature vectors corresponding to each token result vector in the input tensor; in response to completing the merge sorting process for all second tensors obtained based on the transformed routing information, storing the tensor formed by the token result vectors obtained based on the merge sorting process currently stored in the cache into the memory as a whole.

[0148] For example, the number of token result vectors included in the second tensor is determined based on the number of token feature vectors input to one network model, and the number of token feature vectors input to one network model is indicated by the transformed routing information.

[0149] For example, for one piece of transformed routing information, the number of valid elements in each array indicates the number of feature vectors or token feature vectors written into the network model. For example, referring to the foregoing content, assume that there are x valid elements in array 0 of the transformed routing information, and at least part of x result feature vectors are read from the result tensor output by network model 0 to the buffer component as the second tensor. For example, the second tensor includes x token feature vectors, each token feature vector includes at least part of the corresponding result feature vector. For example, the width of the result feature vector is L', and the width of the token result vector is L. The token result vector can be part of the bits in the result feature vector. The specific relationship can be referred to the foregoing Figure 6 related content.

[0150] For example, the converted routing information includes v arrays corresponding one-to-one to v network models. The valid elements in each array indicate the linear index positions corresponding to the token feature vectors of the corresponding network model of the input. The number of valid elements in each array is the number of token feature vectors in the network model corresponding to the input. Based on the converted routing information, reading the second tensor from the memory to the buffer component may include: sequentially traversing the v arrays. For the j-th array, according to the number x of valid elements in the j-th array, reading at least part of the x result feature vectors from the memory as the second tensor, and caching the second tensor to the buffer component, where the x result feature vectors are from the result tensor output by the network model corresponding to the j-th array. The second tensor includes x token result vectors, and each token result vector includes at least part of the content of the corresponding result feature vector. j is a positive integer and less than or equal to v, and x is an integer; in response to the current second tensor in the buffer component being processed, reading at least part of the content of the result tensor output by the network model corresponding to the (j + 1)-th array to the buffer component based on the (j + 1)-th array.

[0151] For example, based on the converted routing information, performing a merge sort process on the second tensor may include: determining the write positions in the cache of each token result vector in the second tensor based on the valid elements in the array corresponding to the second tensor; writing each token result vector in the second tensor to the corresponding write position in the cache according to the write positions.

[0152] As described above, each valid element is the linear index position of the corresponding token feature vector in the first tensor. According to different ways of obtaining the linear index position, the way of calculating the write position is also different.

[0153] For example, as described above, assume that the m linear index positions corresponding to the same token feature vector can be the same and all equal to the index position of the token feature vector in the first tensor. In this case, the write position can be equal to the linear index position.

[0154] For example, as described above, assume that the m linear index positions corresponding to the same token feature vector are different. The write position is equal to the floor result of the quotient of the linear index position and m.

[0155] For example, writing each token result vector in the second tensor to the corresponding write position in the cache according to the write positions may include: for any one of the token result vectors included in the second tensor, in response to the write position corresponding to any one of the token result vectors in the cache already having data, performing an arithmetic operation on the any one of the token result vectors and the data, and writing the operation result to the corresponding write position; in response to the write position corresponding to any one of the token result vectors in the cache not having data, writing the any one of the token result vectors to the corresponding write position.

[0156] For example, the arithmetic operation can be set as needed. For example, a weighted summation operation can be performed.

[0157] As described above, by grouping the routing information, multiple transformed routing information can be obtained. When the results of the batch processing of the multiple transformed routing information are stored in memory, the relative positional relationship is the same as the positional relationship of the corresponding blocks of the transformed routing information.

[0158] In response to the input tensor being divided into multiple blocks based on the block size, the routing information is divided into multiple groups of routing information based on the block size. The multiple groups of routing information include a first group of routing information and a second group of routing information. The first group of routing information corresponds to the first block, and the second group of routing information corresponds to the second block. A second permutation operation is performed on the result tensors respectively output by each network model based on the transformed routing information obtained from the first group of routing information to write the first intermediate tensor into memory. A second permutation operation is performed on the result tensors respectively output by each network model based on the transformed routing information obtained from the second group of routing information to write the second intermediate tensor into memory. When writing the first intermediate tensor and the second intermediate tensor into memory, the storage positional relationship of the first intermediate tensor and the second intermediate tensor in memory is the same as the relative positional relationship between the first block and the second block.

[0159] Figure 9 It is a schematic process diagram of the second permutation operation provided by at least one embodiment of the present disclosure.

[0160] The following combines Figure 9 to specifically illustrate the specific process of the second permutation operation in the data processing method provided by at least one embodiment of the present disclosure.

[0161] As Figure 9 shown, the result tensor 0 is the processing result output by the network model 0, which includes sub-tensors 0_0, sub-tensor 1_0, etc. Among them, the sub-tensor 0_0 is the processing result based on the transformed routing information 1, and the sub-tensor 1_0 is the processing result based on the transformed routing information 2. The same applies to other sub-tensors. For more detailed content about the transformed routing information 1, the transformed routing information 2, etc., reference can be made to the relevant descriptions in the foregoing Figure 8 , which will not be elaborated here.

[0162] The result tensor 1 is the processing result output by the network model 1, which includes sub-tensors 0_1, sub-tensor 1_1, etc. Among them, the sub-tensor 0_1 is the processing result based on the transformed routing information 1, and the sub-tensor 1_1 is the processing result based on the transformed routing information 2. The same applies to other tensors.

[0163] The result tensor 2 is the processing result output by the network model 2, and the result tensor 3 is the processing result output by the network model 3. Its structure is similar to that of the result tensor 0, which will not be elaborated here.

[0164] First, based on the converted routing information 1, the second tensor 0 is read from the memory to the buffer component. For example, as mentioned before, the buffer component can be a buffer Buffer or an extended thread local register.

[0165] For example, based on the array 0 in the converted routing information 1, the second tensor 0 is read from the sub-tensor 0_0. The second tensor 0 includes x token result vectors, where x is the number of valid elements in the array 0 in the converted routing information 1. x can be pre-recorded or obtained by traversing the array 0. The bit width of each token result vector can be smaller than the bit width of each result feature vector in the sub-tensor 0_0, and the specific size can be determined according to the block size.

[0166] Then, based on the converted routing information 1, the second tensor 0 is subjected to a merge sorting process.

[0167] Specifically, based on the valid elements in the converted routing information 1, the write positions of each token result vector in the second tensor 0 in the cache are determined.

[0168] For example, referring to Figure 8 , the first valid element in the array 0 in the converted routing information 1 is 0, and it is determined that the write position of the first token result vector in the second tensor 0 in the cache is =0, and the first token result vector in the second tensor 0 is written to this write position.

[0169] For example, the second valid element in the array 0 in the converted routing information 1 is 7, and it is determined that the write position of the second token result vector in the second tensor 0 in the cache is =3, and the second token result vector in the second tensor 0 is written to this write position.

[0170] The same applies to other valid elements, which will not be elaborated here.

[0171] After traversing all the valid elements in the array 0 in the converted routing information 1, all the token result vectors in the second tensor 0 have been written to the cache. It is determined that the current second tensor 0 in the buffer component has been processed. Based on the array 1 in the converted routing information 1, the second tensor 1 is read from the sub-tensor 0_1 to the buffer component.

[0172] Similar to the second tensor 0, the number of token result tensors in the second tensor 1 is the same as the number of valid elements in the array 1 in the converted routing information 1. The specific description will not be elaborated here.

[0173] Then, based on the transformed routing information 1, perform a merge sorting process on the second tensor 1.

[0174] For example, as Figure 8 shown, the first valid element in array 1 of the transformed routing information 1 is 3, determining that the write position of the first token result vector in the second tensor 1 in the cache is = 1, and write the first token result vector in the second tensor 1 to this write position.

[0175] For example, the second valid element in array 1 of the transformed routing information 1 is 5, determining that the write position of the second token result vector in the second tensor 1 in the cache is = 2, and write the second token result vector in the second tensor 1 to this write position.

[0176] And so on for other valid elements, which will not be elaborated here.

[0177] After traversing all the valid elements in array 1 of the transformed routing information 1, all the token result vectors in the second tensor 1 have been written to the cache. It is determined that the current second tensor 1 in the buffer component has been processed. Then, based on array 2 in the transformed routing information 1, read the second tensor 2 from the sub-tensor 0_2 to the buffer component.

[0178] Similar to the second tensor 0 and the second tensor 1, the number of token result tensors in the second tensor 2 is the same as the number of valid elements in array 2 of the transformed routing information 1. The specific description will not be elaborated here.

[0179] Then, based on the transformed routing information 1, perform a merge sorting process on the second tensor 2.

[0180] For example, the first valid element in array 2 of the transformed routing information 1 is 1, indicating that the write position of the first token result vector in the second tensor 2 in the cache is = 0. Write the first token result vector in the second tensor 2 to this write position. It should be noted that the first token result vector in the second tensor 0 has been written to this write position before. At this time, an arithmetic operation, such as weighted summation, needs to be performed on the first token result vector in the second tensor 2 and the first token result vector in the second tensor 0, and write the operation result to this write position.

[0181] For example, the second valid element in array 2 of the transformed routing information 1 is 2, indicating that the write position of the second token result vector in the second tensor 2 in the cache is = 1, write the second token result vector in the second tensor 2 to this write position. It should be noted that the first token result vector in the second tensor 1 has been written to this write position before. At this time, an operation, such as weighted summation, needs to be performed on the second token result vector in the second tensor 2 and the first token result vector in the second tensor 1, and the operation result is written to this write position.

[0182] And so on for other valid elements, which will not be elaborated here.

[0183] After traversing all the valid elements in the array 2 of the converted routing information 1, all the token result vectors in the second tensor 2 have been written to the cache. It is determined that the current second tensor 2 in the buffer component has been processed. Based on the array 3 in the converted routing information 1, the second tensor 3 is read from the sub-tensor 0_3 to the buffer component.

[0184] Similar to the second tensor 0 and the second tensor 1, the number of token result tensors in the second tensor 3 is the same as the number of valid elements in the array 3 of the converted routing information 1. The specific description will not be elaborated here.

[0185] Then, based on the converted routing information 1, the second tensor 3 is subjected to merge sorting processing.

[0186] For example, the first valid element in the array 3 of the converted routing information 1 is 4, and it is determined that the write position of the first token result vector in the second tensor 3 in the cache is = 2, write the first token result vector in the second tensor 3 to this write position. It should be noted that the second token result vector in the second tensor 1 has been written to this write position before. At this time, an operation, such as weighted summation, needs to be performed on the first token result vector in the second tensor 3 and the second token result vector in the second tensor 1, and the operation result is written to this write position.

[0187] For example, the second valid element in the array 3 of the converted routing information 1 is 8, and it is determined that the write position of the second token result vector in the second tensor 3 in the cache is = 4, write the second token result vector in the second tensor 3 to this write position. It should be noted that the second token result vector in the second tensor 0 has been written to this write position before. At this time, an operation, such as weighted summation, needs to be performed on the second token result vector in the second tensor 3 and the second token result vector in the second tensor 0, and the operation result is written to this write position.

[0188] And so on for other valid elements, which will not be elaborated here.

[0189] After traversing all valid elements in array 3 of the transformed routing information 1, all token result vectors in the second tensor 3 have been written to the cache. It is determined that the current second tensor 3 in the buffer component has been processed, and the tensor 0 composed of the token result vectors in the cache at this time is stored in the memory as a whole.

[0190] After that, the next batch of processing is carried out. For example, the above process can be carried out based on the transformed routing information 2.

[0191] Specifically, based on array 0 in the transformed routing information 2, the second tensor 4 is read from the memory to the buffer component, and the specific process can refer to the foregoing description and will not be elaborated here.

[0192] After that, based on the transformed routing information 2, the second tensor 4 is subjected to merge sorting processing, and each token result vector included in the second tensor 4 is written to the cache. The specific process can refer to the relevant descriptions of the second tensor 0 above and will not be elaborated here.

[0193] After traversing array 0 of the transformed routing information 2, based on array 1 in the transformed routing information 2, the second tensor 5 is read from the memory to the buffer component, and the second tensor 5 is subjected to merge sorting processing. The specific process refers to the above description and will not be elaborated here. The relevant processes for the second tensor 6 and the second tensor 7 are similar and will not be repeated here.

[0194] When all token result vectors in the second tensor 7 have been written to the cache, it is determined that the current second tensor 7 in the buffer component has been processed, and the tensor 1 composed of the token result vectors in the cache at this time is stored in the memory as a whole. In the memory, as Figure 9 shown, tensor 0 and tensor 1 are stored continuously, and their positional relationship is the same as that of the corresponding blocks in the input tensor.

[0195] For example, the input tensor is divided into a first block and a second block. A set of routing information corresponding to the first block is transformed to obtain the transformed routing information 1. Based on Figure 9 the above process, tensor 0 is obtained. A set of routing information corresponding to the second block is transformed to obtain the transformed routing information 2. Based on Figure 9 the above process, tensor 1 is obtained. Thus, the positional relationship between tensor 0 and tensor 1 in the memory is the same as that between the first block and the second block in the input tensor.

[0196] After that, the above process is repeated until all result tensors are traversed. Finally, what is in the memory is the processing result of the input tensor: the output tensor.

[0197] In the present disclosure, the second tensor is stored continuously in memory, and when reading the second tensor from memory to the cache, consecutive memory blocks are accessed. For example, consecutive read and write operations can be performed at the granularity of storage units, which can significantly improve bandwidth utilization. Moreover, in the second permutation operation, based on the transformed routing information, the processing results corresponding to the same first tensor in each resulting tensor are first extracted and cached in the buffer component, and then according to the transformed routing information, the token result vectors in the buffer component are written to the corresponding positions in the cache. The writing position is the same as the position of the token feature vector corresponding to the token result vector in the input vector. Then, the tensor in the cache is stored as a whole in memory, so that consecutive memory blocks are also accessed when writing to memory. For example, consecutive read and write operations can be performed at the granularity of storage units, significantly improving bandwidth utilization.

[0198] In addition, the second tensor is first read into the buffer component, and intermediate operations are all performed in the cache and the buffer component. Finally, the token result vectors written to memory form a tensor and are stored as a whole in memory. That is, when storing to memory, the token result vectors have completed permutation, that is, the token result vectors are arranged according to the positional relationship of the token feature vectors in the input tensor. For the second tensor, only a limited number of data loads from memory are required, without repeated loading, which can significantly reduce the number of memory accesses and memory access overhead compared with the current implementation method of the inverse permutation module, and can fully utilize the on-chip cache resources, reducing resource waste.

[0199] Furthermore, while the current permutation process and inverse permutation process use multiple separate operators to implement, the data processing method provided by at least one embodiment of the present disclosure can make the first permutation operation and the second permutation operation into a fused operator, further reducing data input and output and memory occupation, and significantly improving the overall performance of the permutation module.

[0200] Figure 10 It is a schematic block diagram of a data processing device provided by at least one embodiment of the present disclosure.

[0201] As Figure 10 shown, the data processing device 100 includes a reading module 101, an obtaining module 102, a conversion module 103, and a first operation module 104.

[0202] The data processing device 100 is used to sort an input tensor to aggregate multiple feature vectors included in the input tensor according to a network model, and store the sorted tensor obtained after sorting in memory.

[0203] The reading module 101 is configured to read the first tensor from the memory into the cache. The first tensor includes n token feature vectors, and each token feature vector is input into m different network models for subsequent processing. The first tensor is stored continuously in the memory. The first tensor includes at least part of the data in the input tensor, and each token feature vector is at least part of the content of a feature vector in the input tensor. m and n are positive integers.

[0204] The obtaining module 102 is configured to obtain routing information. The routing information includes n routing sub-information corresponding one-to-one to the n feature vectors, and each routing sub-information is used to indicate the m network models into which the corresponding token feature vector is input.

[0205] The conversion module 103 is configured to convert the routing information to obtain the converted routing information. The converted routing information is used to indicate, in units of network models, the linear index positions corresponding to the token feature vectors input into each network model. The linear index position is used to indicate the position of the corresponding token feature vector in the first tensor.

[0206] The first operation module 104 is configured to perform a first permutation operation based on the converted routing information to aggregate the n token feature vectors included in the first tensor according to the network models. The token feature vectors input into the same network model form a tensor and are stored as a whole in the memory.

[0207] For example, the reading module 101, the obtaining module 102, the conversion module 103, and the first operation module 104 include code and programs stored in the memory. The reading module 101, the obtaining module 102, the conversion module 103, and the first operation module 104 are implemented as a central processing unit (CPU) or other forms of processing units with data processing capabilities and / or instruction execution capabilities. The processing unit can be a general-purpose processor and can also be a single-chip microcomputer, a microprocessor, a digital signal processor, a dedicated image processing chip, or a field programmable logic array, etc. The reading module 101, the obtaining module 102, the conversion module 103, and the first operation module 104 execute the code and programs to implement some or all of the functions of the reading module 101, the obtaining module 102, the conversion module 103, and the first operation module 104 as described above. For example, the reading module 101, the obtaining module 102, the conversion module 103, and the first operation module 104 can be a circuit board or a combination of multiple circuit boards for implementing the functions described above. In the embodiments of the present application, the combination of the one or more circuit boards may include: (1) one or more processors; (2) one or more non-transitory memories connected to the processors; and (3) firmware stored in the memory that can be executed by the processors.

[0208] It should be noted that the reading module 101 can be used to implementFigure 5 Step S10 shown; the obtaining module 102 can be used to implement Figure 5 Step S20 shown; the conversion module 103 can be used to implement Figure 5 Step S30 shown; the first operation module 104 can be used to implement Figure 5 Step S40 shown. Thus, for a specific description of the functions that the reading module 101 can implement, reference can be made to the relevant description of Step S10 in the embodiments of the above data processing method. For a specific description of the functions that the obtaining module 102 can implement, reference can be made to the relevant description of Step S20 in the embodiments of the above data processing method. For a specific description of the functions that the conversion module 103 can implement, reference can be made to the relevant description of Step S30 in the embodiments of the above data processing method. For a specific description of the functions that the first operation module 104 can implement, reference can be made to the relevant description of Step S40 in the embodiments of the above data processing method. Repeated parts will not be elaborated. In addition, the data processing device 100 can achieve technical effects similar to those of the foregoing data processing method, which will not be elaborated here.

[0209] It should be noted that in at least one embodiment of the present disclosure, the data processing device 100 may include more or fewer circuits or units, and the connection relationships between the respective circuits or units are not limited and may be determined according to actual needs. The specific composition manners of the respective circuits or units are not limited and may be constituted by analog devices according to circuit principles, may be constituted by digital chips, or may be constituted in other applicable manners.

[0210] For example, the data processing device 100 can be implemented in a manner of hardware, software, or a combination of hardware and software, and the present disclosure does not make specific limitations thereto.

[0211] The data processing method and data processing device provided in at least one embodiment of the present disclosure can be applied to different systems or devices, such as being applied to Figure 12 the electronic device 300 shown. The electronic device 300 can be a terminal, such as a mobile phone terminal, a tablet computer, a laptop computer, an AR device, a VR device, a vehicle-mounted terminal, etc., or can also be a server, etc. For example, the hardware architecture of the electronic device 300 can include a graphics processor. The data processing method provided in at least one embodiment of the present disclosure can be applied to scenarios involving a CPU, high performance computing (HPC), and artificial intelligence (AI) in the electronic device 300. Of course, the present disclosure is not limited thereto, and any scenarios, devices, apparatuses, etc. involving an MOE structure or a similar structure can adopt the data processing method or data processing device provided in at least one embodiment of the present disclosure.

[0212] In some embodiments, the data processing device provided by at least one embodiment of the present disclosure may be a chip. For example, the chip is a System-on-a-Chip (SoC). The system-on-a-chip includes a processor, and the processor may be a single-core processor or a multi-core processor, a memory, and an I / O interface, etc.

[0213] Figure 11 It is a schematic structural diagram of the processor provided by at least one embodiment of the present disclosure. As Figure 11 shown, the processor 200 includes an instruction parsing unit 201 and a first execution unit 202.

[0214] For example, the instruction parsing unit 201 is used to receive and parse data processing instructions.

[0215] For example, the data processing instruction includes an input tensor and routing information as input parameters. The data processing instruction is used to sort the input tensor based on the routing information to aggregate multiple feature vectors included in the input tensor according to a network model, and store the sorted tensor obtained after sorting into the memory. The routing information includes n routing sub-information corresponding one-to-one to the n feature vectors. Each routing sub-information is used to indicate m network models into which the corresponding token feature vector is input.

[0216] For the relevant descriptions of the input tensor and the routing information, reference may be made to the relevant descriptions of the foregoing data processing method, and the repeated parts will not be elaborated.

[0217] For example, after the instruction parsing unit parses the data processing instruction, the execution unit 202 executes the data processing instruction.

[0218] For example, when the execution unit 202 executes the data processing instruction, it includes performing the following operations: reading a first tensor into a cache from the memory, where the first tensor includes n token feature vectors, each token feature vector is input into m different network models for subsequent processing, the first tensor is continuously stored in the memory, the first tensor includes at least part of the data in the input tensor, each token feature vector is at least part of the content of a feature vector in the input tensor, and m and n are positive integers; converting the routing information to obtain converted routing information, where the converted routing information is used to indicate, in units of network models, the linear index positions corresponding to the token feature vectors input into each network model, and the linear index positions are used to indicate the positions of the corresponding token feature vectors in the first tensor; based on the converted routing information, performing a first permutation operation to aggregate the n token feature vectors included in the first tensor according to network models, where the token feature vectors input into the same network model form a tensor as a whole and are stored in the memory.

[0219] Specifically, when the upper-layer software based on the processor (such as AI applications, HPC applications, and scientific computing applications, etc.) can send data processing instructions for computing and processing to the processor (such as CPU or GPU) through a uniformly encapsulated function library, the data processing instructions can carry routing information and input tensors as input parameters; when the processor receives the data processing instructions, the instruction parsing unit 201 parses the data processing instructions to obtain the routing information and input tensors as input parameters, and the processor schedules the operation unit to execute the data processing task for the input parameters. For example, after parsing the data processing instructions, the processor can store the input parameters in the data processing instructions in registers or memory, and when the execution unit 202 performs computing and processing, it can obtain the input parameters from the registers or memory.

[0220] Regarding the specific process of using the execution unit 202 to execute the data processing instructions, reference can be made to steps S20 - S40 in the data processing method described above, and the repeated parts will not be elaborated.

[0221] The processor provided in at least one embodiment of the present disclosure can achieve similar technical effects to the foregoing data processing method, and the repeated parts will not be elaborated.

[0222] Figure 12 Schematic block diagram of an electronic device provided in an embodiment of the present disclosure. As Figure 12 shown, the electronic device 300 is, for example, suitable for implementing the data processing method provided in the embodiments of the present disclosure. It should be noted that Figure 12 the components of the electronic device 300 shown are only exemplary and not restrictive. According to actual application needs, the electronic device 300 may also have other components.

[0223] As Figure 12 shown, the electronic device 300 may include a processing device 301 (such as a central processing unit, a graphics processing unit, etc.), which can perform various appropriate actions and processes according to non-transitory computer-readable instructions stored in the memory to achieve various functions.

[0224] For example, when the computer-readable instructions are run by the processing device 301, one or more steps in the data processing method according to any of the above embodiments can be executed. It should be noted that the detailed description of the processing process of the data processing method can refer to the relevant descriptions in the embodiments of the above data processing method.

[0225] For example, the memory may include any combination of one or more computer program products, and the computer program products may include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. Volatile memory may include, for example, random access memory (RAM) 303 and / or cache memory, etc. For example, computer-readable instructions may be loaded from the storage device 308 into the random access memory (RAM) 303 to run the computer-readable instructions. Non-volatile memory may include, for example, read-only memory (ROM) 302, hard disk, erasable programmable read-only memory (EPROM), portable compact disc read-only memory (CD-ROM), USB memory, flash memory, etc. Various application programs and various data may also be stored in the computer-readable storage media, such as style images, and various data used and / or generated by the application programs, etc.

[0226] For example, the processing device 301, read-only memory (ROM) 302, and random access memory (RAM) 303 are connected to each other via the bus 304. The input / output (I / O) interface 305 is also connected to the bus 304.

[0227] Generally, the following devices may be connected to the input / output (I / O) interface 305: an input device 306 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 307 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 308 including, for example, a magnetic tape, a hard disk, a flash memory, etc.; and a communication device 309. The communication device 309 may allow the electronic device 300 to communicate with other electronic devices wirelessly or wiredly to exchange data. Although Figure 12 the electronic device 300 with various devices is shown, it should be understood that it is not required to implement or have all the shown devices, and the electronic device 300 may alternatively implement or have more or fewer devices. For example, the processing device 301 may control other components in the electronic device 300 to perform desired functions. The processing device 301 may be a central processing unit (CPU), a tensor processing unit (TPU), or a graphics processing unit GPU, etc., which has data processing capabilities and / or program execution capabilities. The central processing unit (CPU) may be of X86, ARM, RISC-V architecture, etc. The GPU may be directly integrated into the SOC, directly integrated onto the motherboard, or built into the north bridge chip of the motherboard.

[0228] Figure 13 Schematic diagram of a non-transitory computer-readable storage medium provided by at least one embodiment of the present disclosure. For example, as Figure 13As shown, the storage medium 400 can be a non-transitory computer-readable storage medium, and one or more computer-readable instructions 401 can be non-temporarily stored on the storage medium 400. For example, when the computer-readable instructions 401 are executed by a processor, one or more steps of the data processing method described above can be performed.

[0229] For example, the storage medium 400 can be applied to the electronic device 300. For example, the storage medium 400 can include the storage device 308 in the electronic device 300.

[0230] For example, the storage device can include any combination of one or more computer program products. The computer program products can include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. Volatile memory can include, for example, random access memory (RAM) and / or cache memory. Non-volatile memory can include, for example, read-only memory (ROM), hard disk, erasable programmable read-only memory (EPROM), portable compact disc read-only memory (CD-ROM), USB memory, flash memory, etc. One or more computer-readable instructions can be stored on the computer-readable storage medium, and the processor can run the computer-readable instructions to implement various functions of the processor. Various application programs and various data can also be stored in the storage medium.

[0231] For example, the storage medium can include a memory card of a smart phone, a cache component of a tablet computer, a hard disk of a personal computer, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM), portable compact disc read-only memory (CD-ROM), flash memory, or any combination of the above storage media, and can also be other applicable storage media.

[0232] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram can represent a module, a program segment, or a part of code that contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks can occur in a different order than marked in the accompanying drawings. For example, two consecutive blocks shown can actually be executed substantially in parallel, and they can sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and the combination of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.

[0233] The units involved in the embodiments of the present disclosure can be implemented in software or in hardware. Among them, the name of the unit does not constitute a limitation on the unit itself in certain cases.

[0234] The functions described above herein can be performed at least in part by one or more hardware logic components. For example, without limitation, exemplary types of hardware logic components that can be used include: Field Programmable Gate Array (FPGA), Application Specific Integrated Circuit (ASIC), Application Specific Standard Product (ASSP), System on Chip (SOC), Complex Programmable Logic Device (CPLD), and so on.

[0235] The above description is only a preferred embodiment of the present disclosure and an explanation of the applied technical principles. Those skilled in the art should understand that the scope of disclosure involved in the present disclosure is not limited to the technical solutions formed by the specific combination of the above technical features, and should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above disclosure concept. For example, the technical solutions formed by mutually replacing the above features with the technical features (but not limited to) having similar functions disclosed in the present disclosure.

[0236] In addition, although the operations are depicted in a particular order, this should not be construed as requiring that the operations be performed in the particular order shown or in sequential order. In certain environments, multitasking and parallel processing may be advantageous. Similarly, although several specific implementation details are included in the above discussion, these should not be construed as limiting the scope of the present disclosure. Certain features described in the context of separate embodiments can also be implemented combinatorially in a single embodiment. Conversely, the various features described in the context of a single embodiment can also be implemented separately or in any suitable sub-combination in multiple embodiments.

[0237] Although the subject matter has been described in language specific to structural features and / or methodological logical acts, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or acts described above. Rather, the specific features and acts described above are merely example forms for implementing the claims.

[0238] For the present disclosure, the following points also need to be noted: (1) The drawings of the embodiments of the present disclosure only relate to the structures involved in the embodiments of the present disclosure, and other structures can refer to the general design.

[0239] (2) Without conflict, the embodiments of the present disclosure and the features in the embodiments can be combined with each other to obtain new embodiments.

[0240] The above are only specific embodiments of the present disclosure, but the protection scope of the present disclosure is not limited thereto. The protection scope of the present disclosure shall be subject to the protection scope of the claims.

Claims

1. A data processing method for sorting an input tensor to aggregate a plurality of feature vectors included in the input tensor according to a network model, and storing the sorted tensor obtained after sorting into a memory, The data processing method comprises: Reading a first tensor from the memory to a cache, wherein the first tensor includes n word-unit feature vectors, each word-unit feature vector is input into m different network models for subsequent processing, the first tensor is continuously stored in the memory, the first tensor includes at least part of the data in the input tensor, each word-unit feature vector is at least part of the content of a corresponding feature vector in the input tensor, and m and n are positive integers; Acquire routing information, wherein the routing information includes n routing sub-information corresponding to the n word-unit feature vectors one by one, and each routing sub-information is used to indicate m network models input by the corresponding word-unit feature vector; Converting the routing information to obtain converted routing information, wherein the converted routing information is used to indicate the linear index position corresponding to the word unit feature vector input into each network model in units of the network model, and the linear index position is used to indicate the position of the corresponding word unit feature vector in the first tensor; Based on the converted routing information, a first arrangement operation is performed to aggregate the n word-unit feature vectors included in the first tensor according to the network model, wherein the word-unit feature vectors input into the same network model form a tensor that is stored as a whole in the memory, wherein the one tensor is part of the content of the sorted tensor.

2. The data processing method according to claim 1, wherein: The routing information is converted to obtain converted routing information, including: Based on each routing sub-information, the linear index positions corresponding to the n word-unit feature vectors are respectively stored in an array corresponding to the assigned network model to obtain the converted routing information.

3. The data processing method according to claim 2, wherein: Each routing sub-information includes m model identification information, which is used to indicate the identification number of the m network models input by the corresponding word element feature vector. Based on each routing sub-information, the linear index positions corresponding to the n word-element feature vectors are stored in an array corresponding to the assigned network model, including: Traversing the n routing sub-information, for any routing sub-information, based on the index position of the word unit feature vector corresponding to the any routing sub-information in the first tensor, determining the linear index position corresponding to the word unit feature vector, wherein the linear index position corresponding to the word unit feature vector includes m linear index positions corresponding one-to-one to the m model identification information; For any one of the m linear index positions, write the any one linear index position into an array corresponding to the network model indicated by the corresponding model identification information.

4. The data processing method according to claim 1, wherein: Based on the converted routing information, a first arrangement operation is performed, including: Based on the converted routing information, for each network model, the word unit feature vector input to the network model is extracted from the cache, and the extracted word unit feature vector input to the network model is formed into a tensor and stored in the memory as a whole.

5. The data processing method according to claim 4, wherein: The converted routing information includes v arrays corresponding to v network models one by one, each array includes a plurality of valid elements for indicating the linear index position corresponding to the word element feature vector of the input corresponding network model, v is a positive integer, Based on the converted routing information, for each network model, extracting the word-unit feature vector input to the network model from the cache, and forming the extracted word-unit feature vector input to the network model into a tensor and storing it in the memory as a whole, including: For any of the v arrays: Traversing each valid element in any one of the arrays, extracting the word element feature vector indicated by each valid element from the cache and caching it in the buffer component; In response to traversing all valid elements in any one of the arrays, the word-unit feature vectors extracted from the buffer component are stored as the tensor as a whole in the memory.

6. The data processing method according to claim 5, wherein: Extracting the word element feature vector indicated by each valid element from the cache and caching it in the buffer component includes: For each valid element, based on the linear index position indicated by the valid element, determine the index position of the word element feature vector corresponding to the linear index position in the first tensor; The corresponding word-unit feature vector is extracted from the cache based on the index position and cached in the buffer component.

7. The data processing method according to any one of claims 1 to 6, further comprising: Determining a block size based on a capacity of the cache; Divide the input tensor into a plurality of blocks in units of the block size, wherein the number of word-unit feature vectors included in each block is n, and any word-unit feature vector in each block is at least part of the content of the corresponding feature vector in the input tensor; The step of reading the first tensor from the memory to a cache includes: Reading a block from the memory as the first tensor and storing it in the cache; In response to the current first tensor in the cache being processed, a next block is read into the cache until the plurality of blocks are processed.

8. The data processing method according to claim 7, wherein: Based on the block size, the plurality of routing sub-information corresponding to the plurality of feature vectors are divided into a plurality of groups, and the number of routing sub-information in each group is n. For any network model, word-unit feature vectors input into any network model extracted based on different groups of routing information are continuously stored in the memory.

9. The data processing method according to claim 1, further comprising: Based on the converted routing information, a second arrangement operation is performed on the result tensors respectively output by each network model to merge and sort the result tensors respectively output by each network model and store them in the memory, wherein the result tensors respectively output by each network model are obtained by processing the feature vectors input to each network model; Wherein, in the second arrangement operation, result feature vectors written into the cache based on the same converted routing information are formed into a tensor and stored as a whole in the memory, and are continuously stored in the memory.

10. The data processing method according to claim 9, wherein: Based on the converted routing information, a second arrangement operation is performed on the result tensors outputted by each network model, including: Based on the converted routing information, read a second tensor from the memory to a buffer component, wherein the second tensor is at least a part of a result tensor output by a network model, and the second tensor is continuously stored in the memory; Based on the converted routing information, performing a merge sort process on the second tensor, wherein in the merge sort process, each word unit result vector included in the second tensor is written into the cache based on the converted routing information, and the writing relative position relationship between each word unit result vector is the same as the relative position relationship of the word unit feature vector corresponding to each word unit result vector in the input tensor; In response to completing the merge sort process on all second tensors obtained based on the converted routing information, the tensor currently stored in the cache consisting of word unit result vectors obtained based on the merge sort process is stored as a whole in the memory.

11. The data processing method according to claim 10, wherein: The converted routing information includes v arrays corresponding to v network models one by one, the valid elements in each array indicate the linear index position corresponding to the word-unit feature vector of the network model corresponding to the input, and the number of valid elements in each array is the number of word-unit feature vectors in the network model corresponding to the input, Based on the converted routing information, reading a second tensor from the memory to a buffer component includes: Traversing the v arrays in sequence, for the jth array, according to the number x of valid elements in the jth array, reading at least part of the content of x result feature vectors from the memory as the second tensor, and caching the second tensor to the buffer component, wherein the x result feature vectors are from the result tensor output by the network model corresponding to the jth array, the second tensor includes x word unit result vectors, each word unit result vector includes at least part of the content of the corresponding result feature vector, j is a positive integer and is less than or equal to v, and x is an integer; In response to the current second tensor in the buffer component being processed, a new second tensor is read into the buffer component based on the j+1th array, wherein the new second tensor is at least part of the result tensor output by the network model corresponding to the j+1th array.

12. The data processing method according to claim 11, wherein: Based on the converted routing information, merge sorting is performed on the second tensor, including: Determine, based on valid elements in the array corresponding to the second tensor, a write position of each word unit result vector in the second tensor in the cache; According to the write position, each word unit result vector in the second tensor is written to a corresponding write position in the cache.

13. The data processing method according to claim 12, wherein: Writing each word unit result vector in the second tensor into a corresponding write position in the cache according to the write position includes: For any one of the word unit result vectors included in the second tensor, in response to data being stored in a write position corresponding to the any one of the word unit result vectors in the cache, performing a calculation operation on the any one of the word unit result vectors and the data, and writing the calculation result into the corresponding write position; In response to the fact that a write position corresponding to any word unit result vector in the cache does not store data, the any word unit result vector is written into the corresponding write position.

14. The data processing method according to claim 9, wherein: In response to the input tensor being divided into a plurality of blocks based on a block size, the routing information is divided into a plurality of groups of routing information based on the block size, The multiple groups of routing information include a first group of routing information and a second group of routing information, the first group of routing information corresponds to a first block, and the second group of routing information corresponds to a second block, Based on the converted routing information obtained by converting the first set of routing information, the second arrangement operation is performed on the result tensors respectively output by the respective network models to write the first intermediate tensor into the memory, Based on the converted routing information obtained by converting the second set of routing information, the second arrangement operation is performed on the result tensors respectively output by the respective network models to write the second intermediate tensor into the memory, When writing the first intermediate tensor and the second intermediate tensor into the memory, the storage position relationship between the first intermediate tensor and the second intermediate tensor in the memory is the same as the relative position relationship between the first block and the second block in the memory.

15. The data processing method according to any one of claims 1 to 6 and 9 to 14, wherein: The data processing method is used for data processing of a hybrid expert model structure, wherein the hybrid expert model structure includes a routing module and a plurality of expert networks as a plurality of network models. The routing information is determined by the routing module according to the input tensor of the hybrid expert model structure. Each expert network is trained to handle specific tasks and data characteristics.

16. A processor comprising an instruction parsing unit, an execution unit, a memory and a cache, wherein: The instruction parsing unit is used to receive and parse a data processing instruction, wherein the data processing instruction includes an input tensor and routing information as input parameters, and the data processing instruction is used to sort the input tensor based on the routing information to aggregate multiple feature vectors included in the input tensor according to a network model, and store the sorted tensor obtained after sorting into a memory; The execution unit executes the data processing instruction after the instruction parsing unit parses the data processing instruction. Wherein, when the execution unit executes the data processing instruction, it includes performing the following operations: Reading a first tensor from the memory to a cache, wherein the first tensor includes n word-unit feature vectors, each word-unit feature vector is input into m different network models for subsequent processing, the first tensor is continuously stored in the memory, the first tensor includes at least part of the data in the input tensor, each word-unit feature vector is at least part of the content in a feature vector in the input tensor, and m and n are positive integers; Converting the routing information to obtain converted routing information, wherein the routing information includes n routing sub-information corresponding to the n word-meta feature vectors one by one, each routing sub-information is used to indicate the m network models input by the corresponding word-meta feature vector, and the converted routing information is used to indicate the linear index position corresponding to the word-meta feature vector input to each network model in units of the network model, and the linear index position is used to indicate the position of the corresponding word-meta feature vector in the first tensor; Based on the converted routing information, a first arrangement operation is performed to aggregate the n word-unit feature vectors included in the first tensor according to a network model, wherein the word-unit feature vectors input into the same network model form a tensor that is stored as a whole in the memory, wherein the one tensor is part of the content of the sorted tensor.

17. An electronic device comprising: A memory non-transitorily stores computer executable instructions; a processor configured to execute the computer executable instructions, Wherein, when the computer executable instructions are executed by the processor, the data processing method according to any one of claims 1-15 is implemented.

18. A non-transitory computer-readable storage medium, wherein: The non-transitory computer-readable storage medium stores computer-executable instructions, When the computer executable instructions are executed by a processor, the data processing method according to any one of claims 1 to 15 is implemented.

Citation Information

Patent Citations

  • Data processing method and device, computer equipment, storage medium and program product

    CN117764116A

  • Data processing method and device, processor, electronic equipment and storage medium

    CN119312003A

  • Hybrid expert model training optimization method based on matrix routing and Token allocation

    CN119862907A

  • Mixture-of-experts model implementation method and system, electronic device, and storage medium

    WO2023201981A1

  • Data processing method and related apparatus

    WO2024067884A1