Data processing method
By configuring the weights of shared and ordinary expert models in the hybrid expert model, chunking the data blocks, optimizing data transmission and computing overlap, the problem of large communication overhead in the hybrid expert model is solved, and data processing efficiency and speed are improved.
Patent Information
- Application Number
- CN202510718111.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-30
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2045-05-30
AI Technical Summary
In the hybrid expert model, collaboration between different experts in a distributed environment requires frequent exchange of information, resulting in large communication overhead and becoming a bottleneck in data processing efficiency, especially in multi-GPU systems, which affects the speed and efficiency of data processing.
By configuring the first weight of the shared expert model and the second weight of the ordinary expert model in the target processor, the data blocks are processed in blocks, and the shared expert model is calculated first, and the ordinary expert model is calculated when needed, the processing results are integrated, and the data transmission and calculation overlap are optimized to reduce communication delay.
It significantly reduces the overall data processing time, improves the system response speed and computing efficiency, realizes efficient processing of large-scale data sets, and solves the problem of low processing efficiency caused by data transmission delay.
Smart Images

Figure CN120234286A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, and particularly to a method for processing data. Background Art
[0002] With the rapid development of artificial intelligence, the Mixture of Experts (MOE) has emerged as an efficient distributed modeling method. In the scenario of the Mixture of Experts model, a multi-GPU (Graphics Processing Unit) training method is usually adopted to accelerate the training speed through parallel computing. However, since the "experts" in the Mixture of Experts model may be distributed on different GPUs, and only some experts are selected for calculation when the Mixture of Experts model is running, frequent information exchange is required between GPUs, resulting in a large communication overhead. Especially when the number of GPUs is large, the communication overhead may become a bottleneck, resulting in low data processing efficiency of the Mixture of Experts model. Therefore, how to reduce the communication overhead and improve the large-scale data processing efficiency has become an urgent problem to be solved.
[0003] In response to the above problems, there is currently no effective solution. Summary of the Invention
[0004] This application provides a method for processing data to at least solve the problem of low data processing efficiency in related technologies.
[0005] This application provides a method for processing data, which is applied to a target processor and includes: processing each data block included in the received target data to be processed in sequence based on the first weight of the target shared expert model configured in the target processor to obtain a first processing result; when processing each received data block based on the first weight, the following operations are performed, and a second processing result is determined based on all the obtained processing results: when it is determined that there is a target data block in the currently received data block that needs to be processed based on the second weight of the target ordinary expert model, the target data block is processed based on the second weight to obtain a processing result, where the weights of the target ordinary expert model are configured in the target processor as a whole; determining the target processing result of the target data based on the first processing result and the second processing result.
[0006] The present application also provides a data processing device, which is applied to a target processor and includes: a processing module, configured to sequentially process each data block included in the target data to be processed based on the first weight of the target shared expert model configured in the target processor, and obtain a first processing result; a first determination module, configured to perform the following operations when processing each received data block based on the first weight, and determine a second processing result based on all the obtained processing results: when it is determined that there is a target data block in the currently received data block that needs to be processed based on the second weight of the target general expert model, process the target data block based on the second weight to obtain a processing result, where the weights of the target general expert model are configured in the target processor as a whole; a second determination module, configured to determine the target processing result of the target data based on the first processing result and the second processing result.
[0007] The present application also provides an electronic device, including: a memory, configured to store a computer program; a processor, configured to implement the steps of any one of the above data processing methods when executing the computer program.
[0008] The present application also provides a computer-readable storage medium, in which a computer program is stored, and when the computer program is executed by a processor, the steps of any one of the above data processing methods are implemented.
[0009] The present application also provides a computer program product, including a computer program, and when the computer program is executed by a processor, the steps of any one of the above data processing methods are implemented.
[0010] Through the present application, since during the process of processing the data blocks included in the target data based on the first weight of the shared expert model, the target data blocks that need to be processed based on the general expert model are determined, and the target data blocks are processed based on the second weight of the target general expert model, so as to finally obtain the target processing result of the target data. By first performing the calculation of the shared expert model and completing the calculation of the general expert model during the calculation of the shared expert model, the characteristics of the shared expert model and the general expert model in the MOE architecture are fully utilized, overlapping the local data calculation and data transmission, avoiding the idle time of waiting for communication to complete, enabling the aggregation of data to be completed with the minimum communication delay, significantly reducing the overall data processing time, improving the response speed and calculation efficiency of the system, and realizing the efficient processing of large-scale data sets. Therefore, the technical problem of low data processing efficiency caused by data transmission delay in the related art can be solved, and the technical effect of improving the data processing efficiency can be achieved. Description of the Drawings
[0011] To more clearly illustrate the embodiments of the present application, the following will briefly introduce the drawings required for the embodiments. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0012] Figure 1 Hardware structure block diagram of a mobile terminal for a data processing method provided by an embodiment of the present application;
[0013] Figure 2 Flowchart of a data processing method provided by an embodiment of the present application;
[0014] Figure 3 Calculation mode diagram of the first weight provided by an embodiment of the present application;
[0015] Figure 4 Ring communication flowchart of 4 GPUs provided by an embodiment of the present application;
[0016] Figure 5 Calculation mode diagram of a second weight provided by an embodiment of the present application;
[0017] Figure 6 Calculation mode diagram of each data block provided by an embodiment of the present application;
[0018] Figure 7 Typical hardware system structure diagram provided by an embodiment of the present application;
[0019] Figure 8 Structure block diagram of a data processing device provided by an embodiment of the present application. Detailed implementation manners
[0020] In recent years, with the rapid development of artificial intelligence technology, deep learning models have made breakthroughs in the fields of natural language processing (NLP), computer vision (CV), etc. However, with the continuous expansion of model scale, the limitations of traditional single-architecture large models have gradually emerged. Against this background, the Mixture of Experts (MOE) has emerged as an efficient distributed modeling method, opening up new directions for the development of large models from the following aspects:
[0021] (1) Dual drive of computing power and data:
[0022] The success of deep learning relies on powerful computing capabilities and a vast amount of data support. However, the traditional method of simply increasing the number of model parameters will lead to a sharp increase in video memory usage, even exceeding the capacity of a single hardware device. For example, although ultra-large models such as GPT-3 have excellent performance, their training costs are high and they are difficult to popularize. By decomposing tasks into multiple "expert" sub-models, MOE significantly reduces the resource requirements for a single device, thus enabling more efficient construction of large-scale models.
[0023] (2)Requirements in multi-task scenarios:
[0024] In practical applications, many tasks have diverse characteristics, and a single model often struggles to achieve the best performance for all tasks. For example, in different tasks such as machine translation, text generation, and question-and-answer systems, the model needs to possess different knowledge and skills. By introducing multiple expert networks (i.e., expert models), each of which focuses on a specific task or domain, MOE enhances the flexibility and adaptability of the model.
[0025] (3)Advantages of sparsity and dynamic routing:
[0026] Traditional dense neural networks activate the entire model during each inference, which not only wastes a large amount of computing resources but also limits the possibility of model expansion. In contrast, MOE adopts a sparse activation mechanism, that is, only a part of the experts (i.e., expert models) are called for calculation each time, thus significantly reducing the computational overhead. In addition, the dynamic routing algorithm can select the most appropriate combination of experts according to the input content, further optimizing the model performance.
[0027] Although MOE shows great potential, it still faces some challenges. For example, in a distributed environment, the collaboration between different experts requires frequent information exchange, which may lead to high communication latency; or, how to design a reasonable routing algorithm to avoid overloading some experts while other experts are idle, etc.
[0028] Based on the above problems, this application proposes a data processing method. By using block MOE calculation, it effectively overlaps the computing and remote communication (i.e., data exchange between different computing nodes or devices, which can specifically refer to the remote data exchange between GPUs (Graphics Processing Unit) or other computing devices in a multi-GPU system or a multi-node distributed computing system) workloads. At the same time, by overlapping the input / output (I / O) and computing of local single experts (i.e., expert models), it reduces the proportion of local I / O latency in computing, thus solving the bandwidth and I / O bottlenecks of MOE.
[0029] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the protection scope of the present application.
[0030] It should be noted that in the description of the present application, the terms "including", "comprising" or any other variant thereof are intended to cover a non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or device. The terms "first", "second", etc. in the present application are used to distinguish similar objects, rather than to describe a specific order or sequence.
[0031] In order to enable those skilled in the art of the present technology to better understand the solution of the present application, the present application will be further described in detail below in conjunction with the accompanying drawings and specific embodiments.
[0032] In combination with the specific application environment architecture or specific hardware architecture on which the execution of the data processing method depends, the specific application environment architecture or specific hardware architecture will be described herein.
[0033] The method embodiments provided in the embodiments of the present application can be executed on a mobile terminal, a computer terminal or a similar computing device. Taking running on a mobile terminal as an example, Figure 1 is a hardware structure block diagram of a mobile terminal for a data processing method provided in an embodiment of the present application. As Figure 1 shown, the mobile terminal may include one or more ( Figure 1 only one is shown in the figure) processors 102 (the processor 102 may include, but is not limited to, a processing device such as a microprocessor MCU or a programmable logic device FPGA) and a memory 104 for storing data. Among them, the above mobile terminal may further include a transmission device 106 for communication functions and an input / output device 108. Those of ordinary skill in the art can understand that Figure 1 the structure shown is only schematic and does not limit the structure of the above mobile terminal. For example, the mobile terminal may further include more or fewer components than Figure 1 shown in the figure, or have a different configuration from Figure 1 shown in the figure.
[0034] The memory 104 can be used to store computer programs, such as software programs and modules of application software, such as the computer program corresponding to the data processing method in the embodiments of the present application. The processor 102 executes various functional applications and data processing by running the computer program stored in the memory 104, that is, implements the above method. The memory 104 may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memories. In some instances, the memory 104 may further include a memory remotely disposed relative to the processor 102, and these remote memories may be connected to the mobile terminal through a network. Examples of the above network include but are not limited to the Internet, enterprise intranet, local area network, mobile communication network, and combinations thereof.
[0035] The transmission device 106 is used to receive or send data via a network. Specific examples of the above network may include a wireless network provided by a communication provider of the mobile terminal. In one instance, the transmission device 106 includes a network adapter (Network Interface Controller, abbreviated as NIC), which can be connected to other network devices through a base station and thus communicate with the Internet. In one instance, the transmission device 106 may be a radio frequency (RF) module, which is used to communicate with the Internet wirelessly.
[0036] Embodiments of the present application provide a data processing method. The method will be described in detail in combination with the execution process of the data processing method.
[0037] The following explains the professional terms in the embodiments of the present application:
[0038] GPGPU: General-Purpose computing on Graphics Processing Units, general-purpose image processor computing, refers to using a graphics processing unit (GPU) to perform general computing tasks other than graphics processing, and accelerating scientific computing, AI training, etc. scenarios through parallel computing.
[0039] MoE large model: Mixture of Experts, a large model architecture, mainly splits tasks for multiple "expert sub-models" to process, and only activates some parameters, which can balance performance and efficiency. For example, Google's Switch Transformer adopts this design.
[0040] In this embodiment, a data processing method is provided. Figure 2The flowchart of a data processing method provided by an embodiment of this application is as follows Figure 2 As shown, the process includes the following steps:
[0041] Step S202: Based on the first weight of the target shared expert model configured in the target processor, each data block included in the received target data to be processed is sequentially processed to obtain a first processing result;
[0042] Step S204: When processing each received data block based on the first weight, the following operations are performed, and a second processing result is determined based on all the obtained processing results: When it is determined that there is a target data block in the currently received data block that needs to be processed based on the second weight of the target general expert model, the target data block is processed based on the second weight to obtain a processing result, where the weights of the target general expert model are overall configured in the target processor;
[0043] Step S206: Determine the target processing result of the target data based on the first processing result and the second processing result.
[0044] Optionally, the execution subject of the above steps may be a background processor, or other devices with similar processing capabilities, or a machine at least integrated with an image acquisition device and a data processing device. Among them, the image acquisition device may include a graphic acquisition module such as a camera, and the data processing device may include terminals such as a computer and a mobile phone, but is not limited thereto.
[0045] Through the above steps, before processing the target data to be processed, the shared expert model and the general expert model in the target model are pre-configured in multiple processors according to a preset allocation method. Then, when processing the target data, each target processor in the multiple processors first processes each data block in the target data based on the first weight of the locally configured shared expert model to obtain a first processing result. At the same time, during the process of processing the data block, when it is found that there is a target data block that needs to be further processed based on the second weight of the locally configured target general expert model, the target data block is processed based on the second weight to obtain a second processing result, and the first processing result and the second processing result are fused together to determine the target processing result of the target data.
[0046] By first calculating the shared expert model and, during the process of calculating the shared expert model, completing the recombination (i.e., determining specific data blocks from the currently received data blocks that need to be calculated by the currently traversed general expert model) and transmission (i.e., transferring the specific data blocks from memory to the computing module) of the data blocks that require weight processing based on the general expert model, the characteristics of the shared expert model and the general expert model in the MoE architecture are fully utilized, overlapping local data calculation and data transmission. Compared with the related technology where the data blocks that each general expert model needs to calculate are sent to the processor where the general expert model is located, after calculating each general expert model, all data blocks and the processing results corresponding to each data block are synchronized, and all data blocks are recombined into a continuous matrix to perform the calculation of the shared expert model, which will generate a large amount of I / O latency. The idle time waiting for communication to complete is avoided, enabling the aggregation of data to be completed with the minimum communication delay, significantly reducing the overall data processing time, improving the system's response speed and computing efficiency, and achieving efficient processing of large-scale data sets. It solves the technical problem of low data processing efficiency caused by data transmission latency in the related technology and improves the efficiency of data processing.
[0047] In step S202, the target processor can be a hardware device that loads or configures the weights of the model in the system to calculate data through the weights of the model, that is, a hardware device used to perform model inference, such as a central processing unit (CPU), a graphics processing unit (GPU), a tensor processing unit (TPU), and a dedicated AI chip, etc. This processor usually includes a memory and a computing module. Among them, the memory is used to quickly store and access the weights of each expert model and the data blocks in the target data, and the computing module is responsible for performing computing tasks based on the weights of these expert models, such as linear transformation, etc. In the following, the graphics processing unit (GPU) is used as an example to explain the specific embodiments.
[0048] The target model can be a Mixture of Experts (MoE) model. This model architecture works together through multiple expert models (i.e., sub-models) to improve the overall performance and is mostly used in application scenarios that require efficient processing of large-scale data, such as tasks like image recognition and natural language processing. The Mixture of Experts model (i.e., the target model) mainly consists of a shared expert model and general expert models. Among them, the shared expert model can be accessed and used by multiple processors in the system and is distributed to each processor through the weights of the model to achieve parallel computing of the model. The general expert models, on the other hand, are completely resident on their respective bound processors to more precisely process specific tasks or data types.
[0049] The target shared expert model can be any one of the shared expert models included in the target model. The weights of each expert model in the target model are distributed among multiple processors according to a preset ratio. For example, the weights are evenly configured among multiple processors according to the column dimension of the weights, so that each processor has a part of the weights of the shared expert model, enabling the shared expert model to use different processor resources and more effectively utilize computing resources to process different data blocks in parallel; the first weight can specifically refer to the weights of a part of the target shared expert model configured in the target processor.
[0050] The target data can be the data input into the target model waiting to be processed. This data can be image data, text data, audio data, etc. Among them, the target data can be divided into multiple computing units (tokens), and each token generates a data block (token_output) corresponding to this token after encoding conversion, which contains the information to be processed by the expert model. The first processing result can be the processing result of the target processor for all data blocks in the target data.
[0051] Optionally, in the embodiments of the present application, the weights of the shared expert models and the weights of the ordinary expert models included in the target model are pre-configured in multiple processors in the system according to a preset allocation method. When processing the target data, each processor (i.e., the target processor) in the multiple processors processes the data blocks in the received target data in sequence based on the first weights of the shared expert models pre-configured locally to obtain the first processing result.
[0052] Optionally, taking the mixture-of-experts model architecture of deepseek v2 16B in a host system with 4 GPUs as an example, this mixture-of-experts model includes 2 shared expert models and 64 ordinary expert models. The weights of each shared expert model in this mixture-of-experts model are evenly configured in 4 GPUs according to the column dimension. For example, the weights of the shared expert model configured in GPU0 are W0, the weights of the shared expert model configured in GPU1 are W1, the weights of the shared expert model configured in GPU2 are W2, the weights of the shared expert model configured in GPU3 are W3, etc. At the same time, the weights of all ordinary expert models in this mixture-of-experts model are evenly configured in 4 GPUs. For example, the weights of the ordinary expert models configured in GPU0 are experts 0 - 15, the weights of the ordinary expert models configured in GPU1 are experts 16 - 31, the weights of the ordinary expert models configured in GPU2 are experts 32 - 47, the weights of the ordinary expert models configured in GPU3 are experts 48 - 63, etc.
[0053] Taking GPU0 as the target processor as an example, GPU0 processes each token_output (i.e., data block) in the target data based on W0 (i.e., the first weight), and obtains the first data processing result, that is, the processing result of the target data by the shared expert model configured locally on GPU0.
[0054] In step S204, the target general expert model can be a general expert model configured in the target processor and assigned a data processing task. Among them, the weights of the general expert models in the target model are configured in multiple processors according to a preset distribution method. The weight of each general expert model is completely configured in one processor (i.e., bound to one processor). For example, according to the number of general expert models, the weights of the general expert models are evenly configured in multiple GPUs, so as to disperse the computational load of all general expert models in the target model to multiple GPUs. The multiple GPUs share the workload, which can significantly improve the computational efficiency, reduce the training and inference time. At the same time, it can better utilize the available hardware resources, avoid overuse of a single processor and idleness of other processors; the second weight can specifically refer to the weight of the target general expert model configured in the target processor, where the second weight can be the weight of one general expert model or can include the weights of multiple general expert models.
[0055] The target data block can be a data block in the target data assigned to the target general expert model for processing based on the second weight. The target data block can be a single data block or a combined data block including multiple data blocks; the second processing result can be the processing result of all data blocks in the target data by the target processor.
[0056] Optionally, in the embodiment of the present application, during the process of processing the data block based on the weight of the shared expert model (i.e., processing the received data block based on the first weight), the received data block is simultaneously processed based on the weight of the general expert model. The specific processing process is as follows: when the currently received data block contains a target data block assigned to the target general expert model (i.e., needs to be processed based on the second weight), the target data block is processed based on the second weight to obtain a processing result, and then based on all the processing results configured in the target processor, the second processing result is obtained.
[0057] Optionally, taking the mixture-of-experts model architecture of DeepSeek v2 16B in a host system with 4 GPUs as an example, assuming GPU0 is the target processor, during the process of processing each token_output in the target data based on W0 on GPU0, simultaneously traverse experts 0 - 15 configured in GPU0 in sequence to determine whether there is a data block in the currently received data block that needs to be processed based on the weights of the currently traversed general expert model. For example, determine whether there is a data block in the currently received data block that needs to be processed based on the weights of expert 0. If it exists (i.e., there is a target data block), then process this data block (i.e., the target data block) based on the weights of expert 0 (i.e., the second weights) to obtain the processing result of the weights of expert 0 on the target data. Finally, based on the processing results of all target data by experts 0 - 15, determine the second processing result, that is, the processing result of GPU0 on the target data based on the weights of the general expert model configured locally.
[0058] In step S206, the target processing result can be the processing result of the target data based on the weights of all expert models (including shared expert models and general expert models) included in the target model, that is, the processing result of the target model on the target data, where the target processing result includes the processing results of all target data by multiple processors in the system.
[0059] As an optional implementation manner, based on the first weights of the target shared expert model configured in the target processor, process each data block included in the received target data to be processed in sequence to obtain the first processing result, including: repeatedly execute the following operations until all data blocks included in the target data are received: receive and store the first data block, where the first data block is the currently received data block included in the target data; process the first data block through the first weights to obtain the first data block processing result; in the case where it is determined that all data blocks included in the target data are stored in the target processor, based on the obtained first data block processing results, determine the first processing result.
[0060] Optionally, in the embodiments of the present application, the first data block may be the data block currently received by the target processor. The data block may be a single data block or a combined data block obtained by dividing the target data according to a preset rule. For example, the target data is divided according to the row dimension to obtain multiple data blocks. The target processor receives the first data block and stores the first data block in the local memory. At the same time, the calculation module processes the first data block based on the first weight to obtain the processing result of the first data block. Then, the data blocks stored in the local memory are judged. When it is determined that all the data blocks in the target data are stored in the memory (i.e., stored in the target processor), the first processing result is determined based on the obtained processing results of each first data block. Among them, the first data block may be directly allocated and received by the system or received from other processors.
[0061] Optionally, processing the first data block with the first weight to obtain the processing result of the first data block includes: performing a matrix multiplication operation on the first weight and the first data block to obtain the processing result of the first data block.
[0062] Optionally, Figure 3 is a calculation mode diagram of the first weight provided by the embodiments of the present application, as Figure 3As shown, taking GPU0 as the target processor as an example, the gating network in the system divides the target data into multiple tokens, and after encoding each token, it obtains token_output (which can be in matrix form). All token_outputs are processed according to the row dimension to obtain token_output0 (data block 0), token_output1 (data block 1), token_output2 (data block 2), and token_output3 (data block 3). The system sends token_output0 - 3 to GPU0 - 3 respectively, so that each GPU receives and stores the corresponding token_output. GPU0 first receives and stores the token_output0 allocated by the system, and performs a multiplication operation on the received token_output0 based on the first weight W0 (weight 0) to obtain the first data block processing result Z00. Then it determines whether all data blocks are stored in the local memory. If only token_output0 is stored in the current GPU0's storage, it continues to receive and store the token_output sent by other image processors, repeating the above operations until all data blocks of token_output0 - 4 are stored in the memory of GPU0. At this time, GPU0 obtains and stores all the data block processing results Z00 (result 00), Z01 (result 01), Z02 (result 02), and Z03 (result 03) of GPU0 for token_output0 - 4, and determines the first processing result based on all the data block processing results.
[0063] It should be noted that if multiple weights of the shared expert model are configured in the target processor, when calculating based on the weights of the shared expert model, the multiple weights of the shared expert model configured in the target processor can be used to process the first data block in sequence.
[0064] Through the above content, the target processor processes the received data blocks one by one, without having to wait for all data blocks to be received and then uniformly process all data blocks. Thus, it can adapt to datasets of different sizes. By adjusting the size of the data blocks and the processing strategy, it can flexibly adapt to different hardware configurations and application requirements. At the same time, multiple data blocks can be processed simultaneously, further optimizing the processing speed and improving the efficiency and speed of data processing. In addition, by processing the target data in chunks, the data does not need to load the entire dataset into the memory at once, reducing the memory requirements.
[0065] As an optional implementation, receiving and storing the first data block includes: when it is determined that only a partial data block included in the target data is stored in the target processor, receiving and storing the second data block sent by the first processor, where the first processor is a processor among the multiple processors that is ranked before the target processor in the data transmission direction, and data transmission between the multiple processors is performed through ring communication, and the second data block is the data block included in the target data except for the data block already stored in the target processor.
[0066] Optionally, in the embodiments of this application, the first processor may be a processor other than the target processor among the multiple processors included in the system, and this first processor may be a processor located before the target processor in the data transmission direction of the system; the second data block may be the data block sent by the first processor to the target processor, and this second data block may include all data blocks except for the data blocks stored in the target processor, or may be the data blocks in all the data blocks stored in the first processor except for the data block sent last time.
[0067] Optionally, Figure 4 This is the ring communication flowchart of 4 GPUs provided for the embodiments of this application. As Figure 4 shown, taking the communication form of ring all-gather as an example, the communication between GPUs (graphics processors) realizes all-gather (that is, all GPUs collect data blocks from all other GPUs) through ring communication. Assume that there are p = 4 GPUs included in the system, and all GPUs in the system are traversed and numbered, and their numbers are: rank0-3; assume that the size of the entire matrix is V, and each device initially stores a data block of size V / p. After all-gather, each device has a data block of size V. The specific implementation process of ring all-gather is as follows Figure 4 shown. It should be noted that the time required for the entire process of ring communication is (p - 1)*V / (p*B), where B is the total entry or exit bandwidth of the GPUs. If p is large enough, the required time is approximately V / B. At this time, this time is independent of the number of GPUs p.
[0068] Through the above, the transfer of data blocks is continued among multiple processors through ring communication, enabling each processor to only process and forward the corresponding data blocks, reducing redundant data transfer and processing time, and improving the efficiency and reliability of data transfer. At the same time, multiple processors work together, enabling the system to process different data blocks in parallel, thereby improving the overall data processing capacity and meeting the requirements of complex image processing tasks. In addition, since the data blocks are transferred through ring communication, the system can flexibly increase or decrease the number of processors to adapt to different application requirements without significantly affecting the overall system architecture, effectively utilizing system resources and increasing the flexibility and scalability of the system.
[0069] In addition, by combining ring communication with data processing, when the target processor receives and stores the data blocks in the target data, the process of determining the first processing result based on all the processing results is usually accompanied by communication optimization between GPUs. Through prior processing and storage, the subsequent communication can be made more efficient, enabling the system to maintain low communication latency and high bandwidth utilization even in large-scale data processing.
[0070] As an optional implementation manner, determining the first processing result based on the obtained processing results of each first data block includes: determining the position information of all the data blocks included in the target data in the target data; and splicing the obtained processing results of each first data block according to the position information to obtain the first processing result.
[0071] Optionally, in the embodiments of the present application, the position information may be the position corresponding to each data block in the target data. For example, when the target data is divided into different data blocks according to the row dimension, the position information is then the row number corresponding to each data block in the target data; according to the position information of each data block, the obtained processing results of each first data block are spliced to obtain the first processing result. For example, Figure 3 taking the target data divided into 4 tokens according to the row dimension as an example, encoding operations are performed on each token to obtain token_output0 - 3 correspondingly. According to the position of each token in the target data, the processing results Z00 - Z03 of token_output0 - 3 are spliced correspondingly to obtain the first processing result.
[0072] As an optional implementation, before determining that there is a target data block in the currently received data block that needs to be processed based on the second weight of the target general expert model, the above method further includes: determining a first sorting of the weights of multiple general expert models configured in the target processor; based on the first sorting, sequentially traversing the weights of each general expert model configured in the target processor to determine whether there is a data block in the target data that needs to be processed based on the weight of the currently traversed general expert model; in the case of determining that there is a data block in the target data that needs to be processed based on the weight of the currently traversed general expert model, determining whether there is a third data block in the currently received data block that needs to be processed based on the weight of the currently traversed general expert model; in the case of determining that there is no data block in the target data that needs to be processed based on the weight of the currently traversed general expert model, skipping the weight of the currently traversed general expert model.
[0073] Optionally, in the embodiments of the present application, the first sorting may be the sorting of the weights of the general expert models configured in the target processor. According to the first sorting, the weights of the general expert models configured in the target processor are sequentially traversed, and the data blocks in the target data are processed based on the weights of the general expert models during the traversal.
[0074] Optionally, taking the mixture-of-experts model architecture of deepseek v2 16B in a host system with 4 GPUs as an example, the mixture-of-experts model includes 2 shared expert models and 64 general expert models. The weights of the 64 general expert models in the mixture-of-experts model are evenly configured in 4 GPUs. Among them, the weights of 16 general expert models are configured in each GPU. For example, the weights of the general expert models configured in GPU0 are experts 0 - 15, the weights of the general expert models configured in GPU1 are experts 16 - 31, the weights of the general expert models configured in GPU2 are experts 32 - 47, and the weights of the general expert models configured in GPU3 are experts 48 - 63, etc. At the same time, the gating network in the system divides the target data to obtain multiple tokens, and selects 6 general expert models (i.e., Token-k, k = 6) for each token for inference, that is, processes the data block corresponding to the token (i.e., token_output) based on the weights of the selected 6 general expert models.
[0075] As shown in Table 1 below, taking GPU0 as the target processor as an example, in the process of GPU0 processing the currently received data block based on the first weight (i.e., shared expert computing), each expert configured in the GPU is traversed in sequence. Specifically: Determine whether it is necessary to process the data block based on the weight of the ordinary expert model ranked first in the sorting (i.e., expert 0), that is, whether expert 0 is assigned the tokens that need to be inferred. If not, skip expert 0 and continue to traverse the remaining experts in the same way; if so, determine the data block that needs to be inferred by expert 0 (i.e., the target data block) from the currently received data block, transfer this data block from the memory of GPU0 to the computing module (i.e., I / O), process the currently received data block based on expert 0 (i.e., the second weight), and continue to traverse the remaining experts in the same way after the processing is completed. It should be noted that only one implementation manner is described in Table 1 below. In actual operation, it may be that only the data block that needs to be inferred by expert 0 is transmitted during shared expert computing, and only the data block that needs to be inferred by expert 1 is transmitted during expert 0 computing. It may also be that during shared expert computing, multiple expert computations and the transmission of corresponding data blocks are completed. No specific restrictions are made here.
[0076] Table 1 Traversal Order
[0077]
[0078] Through the above content, by traversing the weights of the ordinary expert models in the first sorting during the process of processing the data block based on the weights of the shared expert models, the usage order and condition check of the expert model weights are optimized, enabling the I / O operation and the computing operation to be effectively overlapped, minimizing the impact of I / O latency on the computing efficiency, effectively overcoming the common I / O latency and communication bottleneck problems in large-scale computing, improving the speed of data processing, and also enabling the system to give priority to processing the most likely relevant models, avoiding redundant checks of the weights of each model, thereby saving computing resources and time. At the same time, by accurately identifying the data blocks that need to be processed with specific weights, it is ensured that only the relevant data blocks are processed, thereby improving the processing accuracy and reducing the possibility of misprocessing.
[0079] As an optional implementation manner, in the case where it is determined that there is a third data block in the currently received data block that needs to be processed based on the weight of the ordinary expert model currently traversed, the target data block is processed based on the second weight to obtain a processing result, including: determining the third data block from the currently received data block, where the target data block includes the third data block; performing a multiplication operation on the weight of the ordinary expert model currently traversed and the third data block to obtain a processing result of the third data block, where the second weight includes the weight of the ordinary expert model currently traversed, and the processing result includes the processing result of the third data block.
[0080] Optionally, in the embodiments of the present application, the third data block may be a data block that needs to be processed based on the weights of the currently traversed general expert model, that is, a data block that needs to be inferred based on the weights of the current general expert model; in the case that the currently received data block includes a data block that needs to be processed based on the weights of the currently traversed general expert model, the third data block is determined from the currently received data block, so as to perform a multiplication operation on the third data block and the currently traversed general expert model to obtain a processed result of the third data block.
[0081] As an optional implementation manner, the above method further includes: in the case that it is determined that there is no third data block in the currently received data block, performing a waiting operation until there is a third data block in the currently received data block, where the waiting operation includes: waiting for and receiving a second data block sent by a first processor, and using the second data block as the currently received data block; processing the third data block based on the weights of the currently traversed general expert model to obtain a processed result of the third data block, and the processed result includes the processed result of the third data block; where the first processor is a processor included in the multiple processors that is ranked before the target processor in the data transmission direction, and the second data block is a data block included in the target data except for the data blocks already stored in the target processor.
[0082] Optionally, in the embodiments of the present application, in the case that the currently received data block does not include a data block that needs to be processed based on the weights of the currently traversed general expert model, the processing of the data block based on the weights of the general expert model configured in the target network is temporarily suspended, waiting for the first processor to send a second data block, and when the first processor sends the second data block, receiving and storing the second data block. At this time, it is re-determined whether the currently received data block includes a data block that needs to be processed based on the weights of the currently traversed general expert model, and the above operations are repeated, where the currently received data block after storing the second data block includes the second data block.
[0083] Through the above, when it is detected that there is a third data block in the currently received data block that needs to be processed based on the weights of the general expert model, these data blocks are immediately processed based on the local weights of the general expert model (i.e., the second weights), avoiding resource idleness and maximizing the computing power of the processor. At the same time, when there is no third data block to be processed, a waiting operation is performed instead of performing invalid calculations. During the waiting operation, the processor receives the second data block from the previous processor (the first processor) and treats it as the new current data block. By ensuring that the processor continues to receive data blocks during waiting instead of idly waiting, the data stream can be smoothed, the communication idle period can be reduced, the communication efficiency of the entire system can be optimized, which helps to maintain the continuity and consistency of data transmission, reduce unnecessary data stagnation or congestion, and thus accelerate the overall computing process. In addition, the system can complete the processing of the target data by traversing the weights of the general expert model only once, saving system resources, achieving efficient management of the data stream and full utilization of computing resources, and improving the system throughput rate.
[0084] As an optional implementation, determining the second processing result based on all the obtained processing results includes: for each received data block, performing the following operations to obtain the processing result of the target processor for each data block: determining the probability value of processing the target data block based on the second weights of each target general expert model configured in advance; determining the product value of the processing result obtained by processing the target data block based on the second weights of each target general expert model and the probability value; determining the processing result of the target processor for the data block based on the product values corresponding to each target general expert model; and determining the second processing result based on the processing results of the target processor for each data block.
[0085] Optionally, in the embodiment of the present application, the probability value may be the probability configured for each general expert model by the gating network when determining 6 general expert models for each token. Among them, the higher the probability, the greater the proportion of the inference result of the weights of the general expert model for the token in the final inference result of the token.
[0086] Optionally, determining the second processing result based on the processing results of the target processor for each data block includes: performing an addition operation on the processing results of the target processor for each data block to obtain a sum value; and determining the sum value as the second processing result.
[0087] Optionally, Figure 5 This is a calculation mode diagram of the second weights provided for the embodiment of the present application, as Figure 5As shown in the figure, taking GPU0 as the target processor as an example, all data blocks (data blocks 0 - 4) in the target data are processed based on the weights of the ordinary expert models (experts 0 - 15) configured in the target processor to obtain multiple processing results (results 0 - 4), and then the second processing result, that is, the processing result of GPU0 on the target data, is determined based on results 0 - 4. Specifically: among the weights of the ordinary expert models configured in GPU0, the experts (i.e., the second weights) assigned data inference tasks respectively determine the data blocks that need to be inferred (i.e., the third data blocks), and then perform multiplication operations on the third data blocks and their corresponding experts. Finally, based on the processing results of all multiplication operations, the second processing result is determined. The second processing result can be specifically determined through the following formula:
[0088]
[0089] In the formula, is the second processing result, is the probability value of expert i (i.e., the weight of the i-th ordinary expert model) for the third data block, is the processing result of the third data block, indicates that the weight of the i-th ordinary expert model belongs to both the weights of the 6 ordinary expert models corresponding to the third data block and the weights of the ordinary expert models configured in the target processor.
[0090] As an optional implementation manner, determining the first result based on the first processing result and the second processing result includes: receiving the third processing result and the fourth processing result transmitted by other processors, where the third processing result is the result obtained by other processors processing each data block in the target data in sequence based on the weights of the shared expert models configured in other processors, and the fourth processing result is the result obtained by other processors processing each data block in the target data in sequence based on the weights of the ordinary expert models configured in other processors, where other processors are the processors other than the target processor among multiple processors; obtaining the shared processing result based on the first processing result and the third processing result; obtaining the ordinary processing result based on the second processing result and the fourth processing result; and performing an addition operation on the shared processing result and the ordinary processing result to determine the target processing result.
[0091] Optionally, in the embodiments of the present application, the third processing result may be the processing result of other processors (such as GPU1 - 3) on the target data based on the weights of the shared expert models configured in their respective processors, and the fourth processing result may be the processing result of other processors (such as GPU1 - 3) on the target data based on the weights of the ordinary expert models configured in their respective processors. Among them, the third processing result and the fourth processing result can be transmitted together with the second data block during the process of ring communication.
[0092] The shared data result can be the processing result of the target data by a system (including multiple processors) based on the weights of the shared expert model; the ordinary processing result can be the processing result of the target data by a system (including multiple processors) based on the weights of the ordinary expert model.
[0093] Optionally, Figure 6 This is a calculation mode diagram of each data block provided by an embodiment of this application. As Figure 6 shown, after each data block (data blocks 0-3) is respectively processed based on the weights of the shared expert model configured in multiple processors, the respective shared data results (results 0-3) of each data block are obtained, and the results 0-3 are concatenated to finally obtain the shared data result, that is, results 0-3 are concatenated row by row.
[0094] As an optional implementation manner, based on the first processing result and the third processing result, a shared processing result is obtained, including: concatenating the first processing result and the third processing result to obtain the shared processing result; and / or, based on the second processing result and the fourth processing result, an ordinary processing result is obtained, including: performing an addition operation on the second processing result and the fourth processing result to obtain the ordinary processing result.
[0095] As an optional implementation manner, Figure 7 This is a typical hardware system structure diagram provided by an embodiment of this application. As Figure 7 shown, this system mainly includes: CPU (Central Processing Unit Controller, central processing unit) main control, host memory, and several acceleration devices, where:
[0096] The CPU is the "brain" of the computer, responsible for executing the basic instructions in the instruction set and controlling and coordinating the operation of the entire system. The CPU main control specifically refers to the control and management role of the CPU in the computing system here. It not only executes the ordinary computing tasks in the program code, but also is responsible for scheduling and managing the computing tasks of acceleration devices (such as GPUs), and processing the data transmission with these acceleration devices.
[0097] The host memory refers to the random access memory that communicates directly with the CPU. It is the main place for storing programs and data, and the CPU can directly access the data in the host memory for reading and writing. The access speed of the host memory is much faster than that of peripherals such as hard disks, but still slower compared to GPU memory or caches. In AI computing, a large amount of data often needs to be transmitted between the CPU and the GPU, and the host memory plays a bridging role, storing the data to be processed and the results after calculation for interaction between the CPU and the acceleration devices.
[0098] Accelerating devices generally refer to dedicated hardware designed for specific types of computing tasks, such as GPUs. In the field of AI and deep learning, GPUs have become one of the most commonly used accelerating devices due to their powerful parallel processing capabilities and high-bandwidth memory. By offloading complex matrix operations and parallel computing tasks to GPUs, the computing speed can be significantly improved, especially when dealing with large-scale datasets or performing complex neural network calculations. The core of accelerating devices lies in their ability to provide a more efficient and optimized computing environment for specific tasks. For example, GPUs have thousands of computing cores dedicated to parallel computing, making them very suitable for handling large volumes of data and high-dimensional matrix operations in deep learning. In addition, accelerating devices often come with dedicated high-speed memory (i.e., GPU memory), such as HBM (High Bandwidth Memory), for storing and quickly accessing large datasets, further improving computing efficiency.
[0099] Applied to a hardware system similar to the above, taking the mixture-of-experts model architecture of deepseek v2 16B in a host system with 4 GPUs as an example, the data processing flow is specifically described as follows:
[0100] Before the first ring communication:
[0101] (1) Pre-configuration.
[0102] The mixture-of-experts model includes 2 shared expert models and 64 ordinary expert models. The weights of each shared expert model in this mixture-of-experts model are evenly configured in 4 GPUs according to the column dimension (where the target processor can be any GPU). For example, the weights of the shared expert model configured in GPU0 are W0, the weights of the shared expert model configured in GPU1 are W1, the weights of the shared expert model configured in GPU2 are W2, the weights of the shared expert model configured in GPU3 are W3, etc. At the same time, the weights of all ordinary expert models in this mixture-of-experts model are evenly configured into 4 GPUs. For example, the weights of the ordinary expert models configured in GPU0 are experts 0 - 15, the weights of the ordinary expert models configured in GPU1 are experts 16 - 31, the weights of the ordinary expert models configured in GPU2 are experts 32 - 47, the weights of the ordinary expert models configured in GPU3 are experts 48 - 63, etc.
[0103] The gating network in the system divides the target data into multiple tokens, encodes each token, and obtains token_output (which can be in matrix form). Then, all token_outputs are processed along the row dimension to obtain token_output0, token_output1, token_output2, and token_output3. After that, the system sends token_output0 - 3 to GPU0 - 3 respectively, so that each GPU receives and stores the corresponding token_output.
[0104] Meanwhile, the gating network selects the top 6 general expert models (i.e., 6 general expert models) from 64 general expert models for each token (i.e., token_output), and uses the weights of the 6 general expert models to process the corresponding token.
[0105] (2)Shared expert model calculation.
[0106] Taking GPU0 as an example, it receives and stores token_output0 (i.e., the first data block) in the memory, processes token_output0 based on W0 (i.e., the first weight), and obtains the processing result of the first data block. GPU1 - 3 are similar to GPU0.
[0107] (3)General expert model calculation.
[0108] Taking GPU0 as an example, during the calculation of the first data block based on the weights of the shared expert, it traverses experts 0 - 15 in sequence. Specifically: determine whether there is a token_output in the target data that needs to be processed based on expert 0 (i.e., the currently traversed general expert model). If not, continue to traverse experts 1 - 15 in sequence; if so, determine whether there is a data block in token_output0 that needs to be processed based on expert 0. When there is a data block in token_output0 that needs to be processed based on expert 0, at this time, expert 0 is the second weight of the target general expert model, then determine the third data block from token_output0, transfer the third data block from the memory to the calculation unit, perform a multiplication operation on expert 0 and the third data block to obtain the processing result of the third data block, and continue to traverse experts 1 - 15 in sequence; when there is a data block in token_output0 that needs to be processed based on expert 0, then wait to receive the data in token_output1 (i.e., the second data block) sent by GPU1 (i.e., the first processor).
[0109] The processing methods of GPU1 - 3 are similar to that of GPU0.
[0110] First circular communication:
[0111] Receive and store the second data block sent by the first image memory, and repeat the above steps. Specifically:
[0112] GPU0: Calculate token_output1;
[0113] GPU1: Calculate token_output2;
[0114] GPU2: Calculate token_output3;
[0115] GPU 3: Calculate token_output0.
[0116] Second circular communication:
[0117] Repeat the above steps. Specifically:
[0118] GPU0: Calculate token_output2;
[0119] GPU1: Calculate token_output3;
[0120] GPU2: Calculate token_output0;
[0121] GPU 3: Calculate token_output1.
[0122] Third circular communication:
[0123] Repeat the above steps. Specifically:
[0124] GPU0: Calculate token_output3;
[0125] GPU1: Calculate token_output2;
[0126] GPU2: Calculate token_output1;
[0127] GPU 3: Calculate token_output0.
[0128] After each GPU has calculated all data blocks (i.e., token_output0 - 3) in the target data, circular communication is performed again. After the communication is completed, each GPU stores the processing results of all data blocks in the target data by itself, as well as the processing results of all data blocks in the target data by other GPUs respectively. At this time, each GPU determines the target processing result of the target data based on all processing results (i.e., the first processing result, the second processing result, the third processing result, and the fourth processing result).
[0129] Through the description of the above embodiments, those skilled in the art can clearly understand that the method according to the above embodiments can be implemented by means of software plus a necessary general hardware platform. Of course, it can also be implemented by hardware, but in many cases, the former is a better implementation method.
[0130] The embodiments of the present application also provide a data processing device, which is used to implement the above embodiments and preferred implementation manners, and those that have been described will not be repeated. As used hereinafter, the term "module" can be a combination of software and / or hardware that can implement a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, implementation in hardware, or a combination of software and hardware is also possible and contemplated.
[0131] Figure 8 FIG. is a structural block diagram of a data processing device provided by an embodiment of the present application. As Figure 8 shown, the device is applied to a target processor and includes: a processing module 802, configured to sequentially process each data block included in the received target data based on the first weight of the target shared expert model configured in the target processor to obtain a first processing result; a first determination module 804, configured to perform the following operations when processing each received data block based on the first weight, and determine a second processing result based on all the obtained processing results: when it is determined that there is a target data block in the currently received data block that needs to be processed based on the second weight of the target general expert model, process the target data block based on the second weight to obtain a processing result, where the weights of the target general expert model are configured in the target processor as a whole; a second determination module 806, configured to determine the target processing result of the target data based on the first processing result and the second processing result.
[0132] As an optional implementation manner, the data processing device is further configured to repeatedly perform the following operations until all data blocks included in the target data are received: receive and store a first data block, where the first data block is the data block currently received and included in the target data; process the first data block through the first weight to obtain a first data block processing result; when it is determined that all data blocks included in the target data are stored in the target processor, determine the first processing result based on the obtained first data block processing results.
[0133] As an optional implementation manner, the data processing device is further configured to perform a multiplication operation on the first weight and the first data block to obtain a first data block processing result.
[0134] As an optional implementation, the data processing device is further configured to, when it is determined that only a partial data block included in the target data is stored in the target processor, receive and store a second data block sent by a first processor, where the first processor is a processor among multiple processors that is ranked before the target processor in the data transmission direction, and the multiple processors perform data transmission through ring communication, and the second data block is a data block included in the target data except for the data block already stored in the target processor.
[0135] As an optional implementation, the data processing device is further configured to determine the position information of all data blocks included in the target data in the target data; and splice the obtained respective first data block processing results according to the position information to obtain a first processing result.
[0136] As an optional implementation, the data processing device is further configured to determine a first sorting of the weights of multiple general expert models configured in the target processor; based on the first sorting, sequentially traverse the weights of each general expert model configured in the target processor to determine whether there is a data block in the target data that needs to be processed based on the weight of the currently traversed general expert model; when it is determined that there is a data block in the target data that needs to be processed based on the weight of the currently traversed general expert model, determine whether there is a third data block in the currently received data block that needs to be processed based on the weight of the currently traversed general expert model; when it is determined that there is no data block in the target data that needs to be processed based on the weight of the currently traversed general expert model, skip the weight of the currently traversed general expert model.
[0137] As an optional implementation, the data processing device is further configured to determine a third data block from the currently received data block, where the target data block includes the third data block; perform a multiplication operation on the weight of the currently traversed general expert model and the third data block to obtain a third data block processing result, where the second weight includes the weight of the currently traversed general expert model, and the processing result includes the third data block processing result.
[0138] As an optional implementation, the data processing device is further configured to perform a waiting operation until a third data block exists in the currently received data block when it is determined that the third data block does not exist in the currently received data block. The waiting operation includes: waiting for and receiving a second data block sent by a first processor, and using the second data block as the currently received data block; processing the third data block based on the weight of the ordinary expert model currently traversed to obtain a third data block processing result, where the processing result includes the third data block processing result; where the first processor is a processor among the multiple processors that is ranked before the target processor in the data transmission direction, and the second data block is a data block included in the target data other than the data blocks already stored in the target processor.
[0139] As an optional implementation, the data processing device is further configured to perform the following operations for each received data block to obtain the processing result of the target processor for each data block: determining a probability value for processing the target data block based on the second weight of each target ordinary expert model configured in advance; determining a product value of the processing result obtained by processing the target data block based on the second weight of each target ordinary expert model and the probability value; determining the processing result of the target processor for the data block based on the product values corresponding to each target ordinary expert model; and determining a second processing result based on the processing result of the target processor for each data block.
[0140] As an optional implementation, the data processing device is further configured to perform an addition operation on the processing results of the target processor for each data block to obtain a sum value; and determining the sum value as the second processing result.
[0141] As an optional implementation, the data processing device is further configured to receive a third processing result and a fourth processing result transmitted by other processors, where the third processing result is the result obtained by other processors processing each data block in the target data in sequence based on the weights of the shared expert models configured in the other processors, and the fourth processing result is the result obtained by other processors processing each data block in the target data in sequence based on the weights of the ordinary expert models configured in the other processors, where the other processors are processors other than the target processor among the multiple processors; obtaining a shared processing result based on the first processing result and the third processing result; obtaining a general processing result based on the second processing result and the fourth processing result; and performing an addition operation on the shared processing result and the general processing result to determine the target processing result.
[0142] As an optional implementation, the data processing device is further configured to splice the first processing result and the third processing result to obtain a shared processing result; and / or perform an addition operation on the second processing result and the fourth processing result to obtain a general processing result.
[0143] For the description of the features in the corresponding embodiments of the data processing device, reference can be made to the relevant descriptions in the corresponding embodiments of the data processing method, which will not be elaborated here one by one.
[0144] An embodiment of the present application further provides an electronic device, including a memory and a processor. A computer program is stored in the memory, and the processor is configured to run the computer program to execute the steps in any of the above-described embodiments of the data processing method.
[0145] An embodiment of the present application further provides a computer-readable storage medium, in which a computer program is stored. The computer program is configured to execute the steps in any of the above-described embodiments of the data processing method when running.
[0146] In an exemplary embodiment, the above computer-readable storage medium may include, but is not limited to: various media such as USB flash drives, read-only memories (ROM for short), random access memories (RAM for short), mobile hard disks, magnetic disks, or optical discs that can store computer programs.
[0147] An embodiment of the present application further provides a computer program product. The computer program product includes a computer program, and when the computer program is executed by a processor, it implements the steps in any of the above-described embodiments of the data processing method.
[0148] An embodiment of the present application further provides another computer program product, including a non-volatile computer-readable storage medium. The non-volatile computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, it implements the steps in any of the above-described embodiments of the data processing method.
[0149] Those skilled in the art can further realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present application.
[0150] The above has introduced in detail a method for processing data provided in this application. Specific examples are used in this article to elaborate on the principle and implementation manner of this application. The description of the above embodiments is only used to help understand the method of this application and its core idea. It should be noted that for those of ordinary skill in the art, without departing from the principle of this application, several improvements and modifications can be made to this application, and these improvements and modifications also fall within the protection scope of the claims of this application.
Claims
1. A method for processing data, characterized in that, Applied to a target processor, including: Based on the first weight of the target shared expert model configured in the target processor, each data block included in the received target data to be processed is sequentially processed to obtain a first processing result; When processing each of the received data blocks based on the first weight, the following operations are performed, and a second processing result is determined based on all the obtained processing results: When it is determined that there is a target data block in the currently received data block that needs to be processed based on the second weight of the target general expert model, the target data block is processed based on the second weight to obtain a processing result, where the weights of the target general expert model are configured in the target processor as a whole; Based on the first processing result and the second processing result, the target processing result of the target data is determined.
2. The data processing method according to claim 1, characterized in that Based on the first weight of the target shared expert model configured in the target processor, each data block included in the received target data to be processed is sequentially processed to obtain a first processing result, including: The following operations are repeatedly performed until all data blocks included in the target data are received: Receive and store a first data block, where the first data block is the data block included in the target data currently received; Process the first data block through the first weight to obtain a first data block processing result; When it is determined that all data blocks included in the target data are stored in the target processor, based on the obtained first data block processing results, the first processing result is determined.
3. The data processing method according to claim 2, wherein Processing the first data block through the first weight to obtain a first data block processing result, including: Performing a matrix multiplication operation on the first weight and the first data block to obtain the first data block processing result.
4. The data processing method according to claim 2, characterized in that Receiving and storing a first data block, including: When it is determined that only a part of the data blocks included in the target data are stored in the target processor, receive and store a second data block sent by a first processor, where the first processor is a processor included in multiple processors that is arranged before the target processor in the data transmission direction, and data is transmitted between the multiple processors in a circular communication manner, and the second data block is the data block included in the target data except for the data blocks already stored in the target processor.
5. The data processing method according to claim 2, wherein Based on the obtained first data block processing results, determining the first processing result, including: Determining the position information of all data blocks included in the target data in the target data; According to the position information, splicing the obtained first data block processing results to obtain the first processing result.
6. The data processing method according to claim 1, characterized in that, Before it is determined that there is a target data block in the currently received data block that needs to be processed based on the second weight of the target general expert model, the method further includes: Determining a first sorting of the weights of multiple general expert models configured in the target processor; Based on the first sorting, traverse the weights of each general expert model configured in the target processor in sequence to determine whether there is a data block in the target data that needs to be processed based on the weight of the currently traversed general expert model; When it is determined that there is a data block in the target data that needs to be processed based on the weight of the currently traversed general expert model, determine whether there is a third data block in the currently received data block that needs to be processed based on the weight of the currently traversed general expert model; When it is determined that there is no data block in the target data that needs to be processed based on the weight of the currently traversed general expert model, skip the weight of the currently traversed general expert model.
7. The data processing method according to claim 6, characterized in that, When it is determined that there is a third data block in the currently received data block that needs to be processed based on the weight of the currently traversed general expert model, process the target data block based on the second weight to obtain a processing result, including: Determine the third data block from the currently received data block, where the target data block includes the third data block; Perform a multiplication operation on the weight of the currently traversed general expert model and the third data block to obtain a processing result of the third data block, where the second weight includes the weight of the currently traversed general expert model, and the processing result includes the processing result of the third data block.
8. The method for processing data according to claim 6, wherein The method further includes: When it is determined that the third data block does not exist in the currently received data block, perform a waiting operation until the third data block exists in the currently received data block, where the waiting operation includes: waiting for and receiving a second data block sent by the first processor, and using the second data block as the currently received data block; Process the third data block based on the weight of the currently traversed general expert model to obtain the processing result of the third data block, and the processing result includes the processing result of the third data block; Wherein, the first processor is a processor included in the multiple processors that is ranked before the target processor in the data transmission direction, and the second data block is a data block included in the target data other than the data blocks already stored in the target processor.
9. The data processing method according to claim 7, wherein Determine a second processing result based on all the obtained processing results, including: For each received data block, perform the following operations to obtain the processing result of the target processor for each data block: determine the probability value of processing the target data block based on the second weight of each target general expert model; determine the product value of the processing result obtained by processing the target data block based on the second weight of each target general expert model and the probability value; determine the processing result of the target processor for the data block based on the product values corresponding to each target general expert model; Determine the second processing result based on the processing result of the target processor for each data block.
10. The data processing method according to claim 9, wherein, Determining the second processing result based on the processing result of each data block by the target processor includes: Performing an addition operation on the processing results of each data block by the target processor to obtain a sum value; Determining the sum value as the second processing result.
11. The data processing method according to claim 1, characterized in that, Determining a first result based on the first processing result and the second processing result includes: Receiving a third processing result and a fourth processing result transmitted by another processor, where the third processing result is the result obtained by the other processor after sequentially processing each data block in the target data based on the weights of a shared expert model configured in the other processor, and the fourth processing result is the result obtained by the other processor after sequentially processing each data block in the target data based on the weights of a general expert model configured in the other processor, where the other processor is a processor other than the target processor among the multiple processors; Obtaining a shared processing result based on the first processing result and the third processing result; Obtaining a general processing result based on the second processing result and the fourth processing result; Performing an addition operation on the shared processing result and the general processing result to determine the target processing result.
12. The data processing method according to claim 11, wherein Obtaining a shared processing result based on the first processing result and the third processing result includes: concatenating the first processing result and the third processing result to obtain the shared processing result; and / or Obtaining a general processing result based on the second processing result and the fourth processing result includes: performing an addition operation on the second processing result and the fourth processing result to obtain the general processing result.
13. An electronic device, characterized in that, including: A memory for storing a computer program; A processor for implementing the steps of the data processing method according to any one of claims 1 to 12 when executing the computer program.
14. A computer-readable storage medium, characterized in that, A computer program is stored in the computer-readable storage medium, wherein the computer program, when executed by a processor, implements the steps of the data processing method according to any one of claims 1 to 12.
15. A computer program product, comprising a computer program, characterized in that, The computer program, when executed by a processor, implements the steps of the data processing method according to any one of claims 1 to 12.
Citation Information
Patent Citations
Multi-task processing and model training method and device, medium and equipment
CN114282681A
Recommendation method, device and equipment based on deep learning model
CN116861092A
Customer resource determination method, model training method, electronic equipment and storage medium
CN118096229A
Text processing method and device based on hybrid expert model, equipment and medium
CN119476480A
Data processing method and device based on deep learning model, and medium
CN119514631A