Data processing method, terminal device, storage medium, chip system and computer program product

By performing preloading and first-stage processes in parallel in the end-side device, using the expert module stored in the storage unit, the problem of inefficient inference on the end-side model is solved, and more efficient model inference is achieved.

CN119831056BActive Publication Date: 2025-08-12HONOR DEVICE CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510306530.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-14
Publication Date
2025-08-12
Estimated Expiration
2045-03-14

AI Technical Summary

Technical Problem

The inference efficiency of the end-side model is low, which affects the user experience.

Method used

By executing the preloading process and the first stage flow in parallel between the first thread and the second thread, using a plurality of second expert modules stored in the first storage unit, determine and store according to the first request, parallel execution is implemented to hide I/O delay and improve the cache hit rate.

Benefits of technology

Reduces I/O delay and improves the inference efficiency of the end-side model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119831056B_ABST
    Figure CN119831056B_ABST
Patent Text Reader

Abstract

The data processing method, terminal device, storage medium, chip system and computer program product provided by the embodiments of the present application relate to the field of terminal technology. Since the multiple second expert modules in the first storage are formed by the second thread executing the preloading process based on the expert preference strategy of the first request when the first thread uses the first model to execute the prefilling stage process, the parallel execution of the preloading process and the prefilling stage process is achieved, thereby reducing the I / O delay and improving the reasoning efficiency of the first model. And because the first request has an expert preference, and the multiple second expert modules stored in the first storage are determined according to the first request, therefore, when the first thread uses the first model to execute the decoding stage process, the first thread can directly obtain the second expert module corresponding to the first expert module from the first storage, thereby improving the cache hit rate and further improving the reasoning efficiency of the first model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of terminal technology, and in particular to a data processing method, a terminal device, a storage medium, a chip system, and a computer program product. Background Art

[0002] With the advancement of device and artificial intelligence (AI) technologies, large language models (LLMs) can be deployed on edge devices. This leverages the device's local computing, low latency, low power consumption, security, reliability, availability, and privacy protection to provide users with a superior AI experience. Large models deployed on edge devices are referred to as edge models. Edge models can be applied to at least one of natural language processing (NLP) and computer vision (CV) tasks.

[0003] The inference efficiency of the client-side model is low, which affects the user experience. Summary of the Invention

[0004] Embodiments of the present application provide a data processing method, an end-side device, a storage medium, a chip system, and a computer program product, which are applied in the field of terminal technology and can improve the inference efficiency of a first model (i.e., an end-side model).

[0005] In a first aspect, embodiments of the present application provide a data processing method. The method may include: a first thread may, while executing a second phase process using a first model, retrieve a second expert module corresponding to the first expert module from the first storage if the second expert module corresponding to the first expert module exists in the first storage. The first thread may process a first word using the second expert module corresponding to the first expert module.

[0006] The first thread executing the first stage process using the first model can be executed in parallel with the second thread executing the preloading process.

[0007] The first set stored in the first storage unit can be determined by the second thread based on the first request when the first thread executes the first stage process using the first model. The first set can include multiple second expert modules. The first stage process can be used to indicate that the first request is processed using the first model to obtain the second word element. The second word element can be used to indicate the first output word element corresponding to the first request. The second stage process can be used to indicate that the first model is used to obtain an inference result based on the first request and the second word element. The first stage can also be called the pre-filling stage. The second stage can also be called the decoding stage. The pre-filling stage can refer to the stage of executing the first round of processing. The decoding stage can refer to the stage of executing the second to Pth rounds of processing. P can be an integer greater than 1. The first storage unit storing the second expert module can be understood as the first storage unit storing the model parameters of the second expert module.

[0008] The first word-element may be used to indicate an output word-element. For example, the first word-element may be the pth output word-element. The second word-element may be the 1st output word-element. The second word-element may also be referred to as the first word-element. p may be an integer greater than or equal to 2 and less than or equal to P.

[0009] The first expert module may be determined by the first thread from among multiple expert modules included in the first model based on the first word corresponding to the first request. The first expert module may have a first identifier. The expert module may have a second identifier. The first identifier may be used to indicate the first expert. The second identifier may be used to indicate the expert module.

[0010] In one implementation, the first expert module may be determined by the first thread from among the multiple expert modules included in the first model based on third weights corresponding to the respective expert modules included in the first model. The third weight may be obtained by the first thread invoking a gating module included in the first model and using the gating module based on the first word corresponding to the first request.

[0011] According to an embodiment of the present application, since the multiple second expert modules in the first storage unit are determined by the second thread based on the first request when the first thread executes the first-stage process using the first model, the parallel execution of the preloading process using the second thread and the first-stage process using the first thread is achieved. This allows I / O latency to be minimized, thereby reducing I / O latency and improving the inference efficiency of the first model. Furthermore, since the first request has an expert preference and the multiple second expert modules stored in the first storage unit can be determined based on the first request, when the first thread executes the second-stage process using the first model, the first thread can directly obtain the second expert module corresponding to the first expert module from the first storage unit, thereby improving the cache hit rate of the first thread obtaining the first expert module from the first storage unit and thereby improving the inference efficiency of the first model.

[0012] In one implementation, when a first thread executes a first-stage process using a first model, a second thread may determine, based on a first request, multiple third expert modules from the multiple expert modules included in the first model. The second thread may determine, from a second set, a second expert module corresponding to each of the multiple third expert modules. The second thread may retain the second expert modules corresponding to each of the multiple third expert modules in the second set to obtain a first set. That is, the first set stored in the first storage unit may be obtained by the second thread retaining the second expert modules corresponding to each of the multiple third expert modules in the second set when the first thread executes the first-stage process using the first model. The multiple third expert modules may be determined by the second thread from the multiple expert modules included in the first model based on the first request. The second set may be stored in the first storage unit. The second set may include the multiple second expert modules. The second set may be loaded from the second storage unit to the first storage unit by the first thread.

[0013] In the first stage (i.e., the pre-population stage), the expert module in the active state can be referred to as the second expert module. In the second stage (i.e., the decoding stage), the expert module in the active state can be referred to as the first expert module. In the pre-loading stage, the expert module predicted to be active can be referred to as the third expert module. The second expert module can have a third identifier. The third expert module can have a fourth identifier. The third identifier can be used to identify the second expert module. The fourth identifier can be used to identify the third expert module. A second expert module corresponding to a third expert module can be understood as the second expert module being the third expert module.

[0014] The second storage unit can be used to store the expert module included in the first model. The second storage unit storing the expert module can be understood as the second storage unit storing the model parameters of the expert module.

[0015] The first speed may be greater than the second speed. The first speed may be used to indicate the speed at which the first thread accesses the first storage unit. The second speed may be used to indicate the speed at which the first thread accesses the second storage unit.

[0016] For any third expert module among the multiple third expert modules corresponding to the first request, in the process of the first thread loading at least one second expert module corresponding to at least one third word from the second storage unit to the first storage unit, if the second thread determines that there is a second expert module corresponding to the third expert module in the second expert modules stored in the first storage unit, that is, if it is determined that the third expert module exists in the second expert module stored in the first storage unit, then the second expert module corresponding to the third expert module in the first storage unit can be retained.

[0017] In one implementation, the second thread may store the second expert module corresponding to the third expert module in the first storage unit into a queue.

[0018] If the second thread determines that the second set stored in the first storage unit does not contain a second expert module corresponding to the third expert module, the second thread may load the expert module corresponding to the third expert module from the second storage unit to the first storage unit. For example, the second thread may load the expert module corresponding to the third expert module from the second storage unit to the queue of the first storage unit.

[0019] Therefore, when the first thread executes the second stage process, the second expert module stored in the first storage unit corresponds to the third expert module, that is, the first storage unit stores the third expert module.

[0020] Because the first set stored in the first storage unit can be obtained by the second thread retaining the second expert modules corresponding to each of the multiple third expert modules in the second set when the first thread executes the first stage process using the first model, the parallel execution of the preloading process by the second thread and the first stage process by the first thread is achieved. This minimizes I / O latency, thereby reducing I / O latency and improving the inference efficiency of the first model. Furthermore, because the first request has an expert preference, and the third expert module can be determined by the second thread based on the first request when the first thread executes the first stage process using the first model, the first thread can directly obtain the second expert module corresponding to the first expert module from the first storage unit when the first thread executes the second stage process using the first model. This improves the cache hit rate of the first thread obtaining the first expert module from the first storage unit, thereby improving the inference efficiency of the first model.

[0021] In one implementation, the second thread determining multiple third expert modules from multiple expert modules based on the first request may include: the second thread may use a second model to determine the multiple third expert modules from the multiple expert modules included in the first model based on the first request. The second model may be trained using the second request and multiple first scores corresponding to the second request. The first scores may be used to indicate activation scores of the second expert modules.

[0022] In one implementation, the first set stored in the first storage unit may be determined by the second thread using the second model based on the first request when the first thread executes the first stage process using the first model. For example, the first set stored in the first storage unit may be obtained by the second thread inputting the first request into the second model when the first thread executes the first stage process using the first model. The second model may be trained using the second request and multiple first scores corresponding to the second request. The first score may be used to indicate the activation score of the expert module.

[0023] In one implementation, the second thread may input the first request into the second model to obtain a second score corresponding to each of the plurality of expert modules. Based on the second scores corresponding to each of the plurality of expert modules, a plurality of third expert modules corresponding to the first request are obtained. The second scores may be used to indicate a predicted activation score for the expert module. The predicted activation score may be used to indicate a probability of activation.

[0024] In one achievable embodiment, obtaining multiple third expert modules corresponding to the first request based on the second scores corresponding to each of the multiple expert modules may include: the second thread may determine U second scores from the second scores corresponding to each of the multiple expert modules. The expert module corresponding to each of the U second scores is determined as the third expert module. In the case where the probability of the expert module being accessed is agreed upon as the larger the second score, the U second scores may be the maximum U second scores. In the case where the probability of the expert module being accessed is agreed as the smaller the second score, the U second scores may be the minimum U second scores. U may be an integer greater than 1 and less than R*S. U may be determined based on the storage capacity of the first storage unit, which is not limited here.

[0025] In one implementation, the second model can be obtained by adjusting model parameters of the second model based on a loss function value. The loss function value can be obtained based on the loss function and based on first scores and second scores corresponding to each of the multiple expert modules. The second scores corresponding to each of the multiple expert modules can be obtained by inputting the second request into the second model. The second scores can be used to indicate a predicted activation score for the expert module.

[0026] In one implementation, the loss function value may be obtained based on the loss function values corresponding to the plurality of expert modules. The loss function value corresponding to the expert module may be obtained by inputting the first score and the second score corresponding to the expert module into the loss function.

[0027] Since the second model can be trained based on the expert preferences of the request, that is, the second model can be trained using the second request and multiple first scores corresponding to the second request, the first score can be used to indicate the activation score of the expert module, thereby improving the expert prediction accuracy of the second model. On this basis, the first set stored in the first storage unit can be obtained by the second thread inputting the first request into the second model when the first thread executes the first stage process using the first model. Therefore, the first thread can directly obtain the second expert module corresponding to the first expert module from the first storage unit, thereby improving the cache hit rate of the first thread obtaining the first expert module from the first storage unit, thereby improving the inference efficiency of the first model.

[0028] In one achievable embodiment, the first thread may determine at least one third word element based on the first request. The first thread may determine at least one second expert module corresponding to each of the at least one third word elements from the multiple expert modules included in the first model. The first thread may load at least one second expert module corresponding to each of the at least one third word elements from the second storage unit to the first storage unit to obtain a second set. That is, the second set may be the result of the first thread loading at least one second expert module corresponding to each of the at least one third word elements from the second storage unit to the first storage unit. The at least one second expert module corresponding to each of the at least one third word elements may be determined by the first thread from the multiple expert modules included in the first model based on the at least one third word element. The at least one third word element may be determined by the first thread based on the first request.

[0029] The first model may include at least one first layer. The first layer may include multiple expert modules. The at least one expert module corresponding to the third word may include at least one expert module corresponding to each of the at least one first layers. For any first layer in the at least one first layer, the first layer may be used as the current layer. For any third word in the at least one third word, the at least one expert module corresponding to the current layer and the third word may be determined by the first thread based on the third word from the multiple expert modules corresponding to each of the at least one first layer included in the first model.

[0030] In one implementation, if the first storage unit does not contain the second expert module corresponding to the first expert module and the first storage unit contains the second expert module corresponding to the fourth expert module, the first thread may retrieve the second expert module corresponding to the fourth expert module from the first storage unit. The first thread may use the second expert module corresponding to the fourth expert module to process the first word.

[0031] The first expert module may have a corresponding first weight. The first weight may be used to indicate the activation probability of the first expert module. The fourth expert module may have a corresponding second weight. The second weight may be used to indicate the activation probability of the fourth expert module. The absolute value of the difference between the first weight and the second weight may be less than or equal to the first threshold. Thus, the first weight may be greater than or less than the second weight, as long as the absolute value of the difference between the first weight and the second weight is less than or equal to the first threshold. The first threshold can be configured according to actual business needs and is not limited here. For example, the first threshold may be less than or equal to 0.3.

[0032] The first thread can, if there is no second expert module corresponding to the first expert module in the first storage unit, obtain the second expert module corresponding to the fourth expert module from the first storage unit if there is a fourth expert module corresponding to the first expert module among the multiple expert modules included in the first model, if there is a second expert module corresponding to the fourth expert module in the first storage unit.

[0033] Because the absolute value of the difference between the second weight corresponding to the fourth expert module and the first weight corresponding to the first expert module is less than or equal to the first threshold, the second weight can be used to indicate the activation probability of the fourth expert module, and the first weight can be used to indicate the activation probability of the first expert module. Therefore, it can be shown that the prediction accuracy of the fourth expert module is close to that of the first expert module, and thus, the fourth expert module can replace the first expert module. Based on this, if the second expert module corresponding to the first expert module does not exist in the first storage unit, and the second expert module corresponding to the fourth expert module exists in the first storage unit, the first thread can directly retrieve the second expert module corresponding to the fourth expert module from the first storage unit to process the first word-gram using the second expert module corresponding to the fourth expert module, thereby further improving the cache hit rate. Furthermore, since there is no need to first load the expert module corresponding to the first expert module from the second storage unit to the first storage unit and then retrieve the expert module corresponding to the first expert module to process the first word-gram using the expert module corresponding to the first expert module, the I / O latency between the second storage unit and the first storage unit is reduced, thereby improving the inference efficiency of the first model.

[0034] In one implementation, the first thread may call a gating module included in the first model and, using the gating module, obtain a third weight corresponding to each of the multiple expert modules based on the first word corresponding to the first request. The first thread may determine a first weight based on the third weights corresponding to each of the multiple expert modules. The first thread may determine a second weight from the third weights corresponding to each of the multiple expert modules based on the first weight and a first threshold. That is, the first weight may be determined by the first thread based on the third weights corresponding to each of the multiple expert modules included in the first model. The third weight corresponding to each of the multiple expert modules included in the first model may be obtained by the first thread calling a gating module included in the first model and, using the gating module, based on the first word corresponding to the first request. The second weight may be determined by the first thread based on the first weight and the first threshold, from the third weights corresponding to each of the multiple expert modules included in the first model.

[0035] Since the second weight corresponding to the fourth expert module can be determined from the third weights corresponding to the multiple expert modules included in the first model based on the first threshold and the first weight corresponding to the first expert module, the second weight can be used to indicate the activation probability of the fourth expert module, and the first weight can be used to indicate the activation probability of the first expert module. Therefore, the determined activation probability of the fourth expert module is closer to the activation probability of the first expert module. As a result, the fourth expert module can replace the first expert module to process the first word.

[0036] In one implementable manner, the fourth expert module may be the second expert module in the first set corresponding to the first request.

[0037] In another implementation, the first set may further include multiple second expert modules corresponding to the third request. The fourth expert module may be the second expert module in the first set corresponding to the third request. The third request may be different from the first request.

[0038] Since the fourth expert module can be the second expert module corresponding to the first request or the second expert module corresponding to the third request, the diversity of sources of the fourth expert module is improved, which helps to further reduce I / O delay.

[0039] In one implementation, the first thread may load the expert module corresponding to the first expert module from the second storage unit to the first storage unit, obtain the expert module corresponding to the first expert module from the first storage unit, and use the expert module corresponding to the first expert module to process the first word element when the first storage unit does not contain the second expert module corresponding to the first expert module and the first storage unit does not contain the second expert module corresponding to the fourth expert module.

[0040] In one implementation, if the first storage unit does not contain a second expert module corresponding to the first expert module and the first storage unit does not contain a second expert module corresponding to the fourth expert module, the first thread may, at the third thread, retrieve the expert module corresponding to the first expert module from the second storage unit and use the expert module corresponding to the first expert module to process the first word. The first thread may be used to indicate an accelerator thread, and the third thread may be used to indicate a processor thread.

[0041] Since the first thread can be used to indicate the accelerator thread, the third thread can be used to indicate the processor thread, the first storage unit can be used to indicate the accelerator memory, the second storage unit can be used to indicate the processor memory, the third speed can be used to indicate the speed at which the third thread accesses the second storage unit, and the fourth speed can be used to indicate the speed at which the third thread accesses the first storage unit, the third access speed is greater than the fourth access speed. When the first thread determines that the first storage unit does not contain a second expert module corresponding to the first expert module and does not contain a second expert module corresponding to the fourth expert module, the third thread is used to obtain the expert module corresponding to the first expert module from the second storage unit and use the expert module corresponding to the first expert module to process the first word, without the first thread first loading the expert module corresponding to the first expert module from the second storage unit to the first storage unit, obtaining the expert module corresponding to the first expert module from the first storage unit, and then using the expert module corresponding to the first expert module to process the first word. Therefore, the I / O latency between the second storage unit and the first storage unit is reduced, and the inference efficiency of the first model is improved.

[0042] In one implementation, the first thread can load the expert module corresponding to the first expert module from the second storage unit to the first storage unit when the second expert module corresponding to the first expert module does not exist in the first storage unit, obtain the expert module corresponding to the first expert module from the first storage unit, and use the expert module corresponding to the first expert module to process the first word.

[0043] In another implementation, if the first thread does not have a second expert module corresponding to the first expert module in the first storage unit, the third thread can retrieve the expert module corresponding to the first expert module from the second storage unit and use the expert module corresponding to the first expert module to process the first word. The third thread can be used to indicate a processor thread. The first thread can be used to indicate an accelerator thread.

[0044] Since the first thread can be used to indicate the accelerator thread, the third thread can be used to indicate the processor thread, the first storage unit can be used to indicate the accelerator memory, the second storage unit can be used to indicate the processor memory, the third speed can be used to indicate the speed at which the third thread accesses the second storage unit, and the fourth speed can be used to indicate the speed at which the third thread accesses the first storage unit, the third access speed is greater than the fourth access speed. When the first thread determines that the first storage unit does not contain a second expert module corresponding to the first expert module, the third thread is used to obtain the expert module corresponding to the first expert module from the second storage unit and use the expert module corresponding to the first expert module to process the first word. This does not require the first thread to first load the expert module corresponding to the first expert module from the second storage unit to the first storage unit, obtain the expert module corresponding to the first expert module from the first storage unit, and then use the expert module corresponding to the first expert module to process the first word. Therefore, the I / O delay between the second storage unit and the first storage unit is reduced, and the inference efficiency of the first model is improved.

[0045] The fourth thread may eliminate the second expert module from the first set based on the cache elimination policy when the first storage unit satisfies the first condition. The first condition may be used to trigger cache elimination. The fourth thread may be executed in parallel with the first thread and the second thread.

[0046] Because the fourth thread can eliminate the second expert module from the first set based on the cache elimination policy when the first storage unit meets the first condition, the fourth thread can be executed in parallel with the first thread and the second thread, thereby improving the cache hit rate and resource utilization when the storage capacity of the first storage unit is limited.

[0047] In one implementation, the first condition may include at least one of the following: the first capacity reaches the second threshold, the first number reaches the third threshold, the first usage rate reaches the fourth threshold, or the first hit rate is less than or equal to the fifth threshold. The first capacity can be used to indicate the total storage size of the first set. The second threshold can be determined based on the storage capacity of the first storage unit. The first number can be used to indicate the total number of second expert modules in the first set. The third threshold can be determined based on the storage capacity of the first storage unit and the storage size of the second expert module. The first usage rate can be used to indicate the ratio between the first capacity and the storage capacity of the first storage unit. The first hit rate can be used to indicate the ratio between the second number and the third number. The second number can be used to indicate the number of cache hits corresponding to the first storage unit. The third number can be used to indicate the total number of accesses corresponding to the first storage unit. And / or

[0048] The cache elimination policy may include at least one of the following: a first-in-first-out policy, a least recently used policy, a least frequently used policy, a most recently used policy, a random elimination policy, or a low interactive reference interval policy.

[0049] In one implementation, the first storage unit may be used to indicate an accelerator memory, and the second storage unit may be used to indicate a processor memory or a storage device.

[0050] In another implementation, the first storage unit may be used to indicate a processor memory, and the second storage unit may be used to indicate a storage device.

[0051] Since the second storage unit can be used to store multiple expert modules included in the first model, and the first storage unit can be used to store the expert modules of the activated experts, the cooperation between the second storage unit and the first storage unit alleviates the limited memory capacity of the end-side device, making it possible for the end-side device to deploy the first model.

[0052] In one implementation, the accelerator memory may be used to indicate a graphics processor memory, a neural network processor memory, or a tensor processor memory.

[0053] In one implementation, the third thread may be a processor thread. If the first storage unit can be used to indicate accelerator memory, the first and second threads may be accelerator threads. The fourth thread may be an accelerator thread. If the first storage unit can be used to indicate processor memory, the first, second, and fourth threads may be processor threads.

[0054] In a second aspect, an embodiment of the present application provides a data processing device, which may include: an acquisition module for, when a first thread executes a second-stage process using a first model, obtaining a second expert module corresponding to the first expert module from a first storage unit if a second expert module corresponding to the first expert module exists in the first storage unit. The first expert module is determined by the first thread from multiple expert modules included in the first model based on a first word element corresponding to a first request. A processing module is used by the first thread to process the first word element using the second expert module corresponding to the first expert module.

[0055] According to an embodiment of the present application, the first set stored in the first storage unit may be determined by the second thread based on the first request when the first thread executes the first stage process using the first model. The first set may include multiple second expert modules. The first stage process may be used to indicate the use of the first model to process the first request to obtain a second word. The second word may be used to indicate the first output word corresponding to the first request. The second stage process may be used to indicate the use of the first model to obtain an inference result based on the first request and the second word.

[0056] In a third aspect, an embodiment of the present application provides an end-side device, comprising a processor and a memory, wherein the memory is used to store computer-executable instructions, and the processor is used to run the computer-executable instructions stored in the memory to execute the method described in the first aspect or any one of the implementable embodiments of the first aspect.

[0057] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, which stores a computer program or instructions. When the computer program or instructions are run on a computer, the terminal device executes the method described in the first aspect or any one of the implementable embodiments of the first aspect.

[0058] In a fifth aspect, an embodiment of the present application provides a computer program product comprising a computer program, which, when executed, enables an end-side device to execute the method described in the first aspect or any one of the implementable embodiments of the first aspect.

[0059] In a sixth aspect, the present application provides a chip or chip system, comprising at least one processor and a communication interface, wherein the communication interface and the at least one processor are interconnected via a line, and the at least one processor is configured to execute a computer program or instruction to perform the method described in the first aspect or any one of the implementations of the first aspect. The communication interface in the chip may be an input / output interface, a pin, or a circuit.

[0060] In one possible implementation, the chip or chip system described above in this application further includes at least one memory, wherein instructions are stored in the at least one memory. The memory may be a storage unit within the chip, such as a register or cache, or a storage unit of the chip (such as a read-only memory or random access memory).

[0061] It should be understood that the second to sixth aspects of the present application correspond to the technical solutions of the first aspect of the present application, and the beneficial effects achieved by each aspect and the corresponding feasible implementation methods are similar and will not be repeated here. BRIEF DESCRIPTION OF THE DRAWINGS

[0062] Figure 1A A schematic diagram of a converter provided for related art;

[0063] Figure 1B Schematic diagram of the coding layer provided for related technologies;

[0064] Figure 1C A schematic diagram of a decoding layer provided for related art;

[0065] Figure 2A A schematic diagram of a large encoder-based model provided in an embodiment of the present application;

[0066] Figure 2B A schematic diagram of a large decoder-based model provided in an embodiment of the present application;

[0067] Figure 2C A schematic diagram of a large encoder-decoder model provided in an embodiment of the present application;

[0068] Figure 3 A schematic diagram of the application of the reasoning process of the large model provided in the embodiment of the present application;

[0069] Figure 4A A schematic diagram of a hybrid expert provided in an embodiment of the present application;

[0070] Figure 4B A schematic diagram of another hybrid expert provided in an embodiment of the present application;

[0071] Figure 4C A schematic diagram of an expert module provided in an embodiment of the present application;

[0072] Figure 5 A schematic diagram of the MoE model provided in the embodiments of the present application;

[0073] Figure 6 A schematic diagram of a first model provided in an embodiment of the present application;

[0074] Figure 7A A schematic diagram of the storage structure of a terminal device provided in an embodiment of the present application;

[0075] Figure 7B A schematic diagram of a storage method of a first model provided in an embodiment of the present application;

[0076] Figure 7C A schematic diagram of another storage method of the first model provided in an embodiment of the present application;

[0077] Figure 7D A schematic diagram of another storage method of the first model provided in an embodiment of the present application;

[0078] Figure 8A A schematic diagram illustrating the principle of a data processing method provided in an embodiment of the present application;

[0079] Figure 8B A schematic diagram illustrating another data processing method according to an embodiment of the present invention;

[0080] Figure 8C A schematic diagram of the principle of a request-based expert preference prediction method provided in an embodiment of the present application;

[0081] Figure 8D A schematic diagram of determining first scores corresponding to respective second expert modules corresponding to a third word-gram provided in an embodiment of the present application;

[0082] Figure 9A A schematic diagram of replacing a first expert module with a fourth expert module provided in an embodiment of the present application;

[0083] Figure 9B A schematic diagram of another embodiment of the present application in which a fourth expert module is used to replace a first expert module;

[0084] Figure 10 A schematic diagram of processing a first word using a third thread according to an embodiment of the present application;

[0085] Figure 11 A schematic diagram of the structure of the terminal side device provided in an embodiment of the present application;

[0086] Figure 12 A software structure diagram of the terminal device provided in the embodiment of the present application;

[0087] Figure 13 A flowchart of a data processing method provided in an embodiment of the present application;

[0088] Figure 14A A schematic diagram of an application scenario of the data processing method provided in an embodiment of the present application;

[0089] Figure 14B A schematic diagram of storing the second expert module of the first model in the second storage provided in an embodiment of the present application;

[0090] Figure 14C A schematic diagram illustrating another data processing method according to an embodiment of the present invention;

[0091] Figure 14D A schematic diagram of an embodiment of the present application showing a first thread executing a first-stage process using a first model and a second thread executing a preloading process using a second model;

[0092] Figure 14E A schematic diagram of a first thread using a first model to execute a second-stage process according to an embodiment of the present application;

[0093] Figure 15A flowchart of another data processing method provided in an embodiment of the present application;

[0094] Figure 16 A flowchart of another data processing method provided in an embodiment of the present application;

[0095] Figure 17 A flowchart of another data processing method provided in an embodiment of the present application;

[0096] Figure 18 A flowchart of another data processing method provided in an embodiment of the present application;

[0097] Figure 19 This is a structural block diagram of the data processing device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0098] The technical solutions in the embodiments of the present application will be described below with reference to the accompanying drawings.

[0099] To facilitate understanding of the embodiments of the present application, the following explanations are first made.

[0100] 1. In the embodiments of the present application, "indication" may include direct indication, indirect indication, explicit indication, or implicit indication. When describing that a certain indication information is used to indicate A, it can be understood that the indication information carries A, directly indicates A, or indirectly indicates A.

[0101] The information indicated by the indication information can be referred to as the information to be indicated. During the implementation process, there are multiple ways to indicate the information to be indicated. For example, the information to be indicated itself can be directly indicated or the index of the information to be indicated can be indicated. Optionally, the information to be indicated can also be indirectly indicated by indicating other information. There can be an association between the other information and the information to be indicated. Optionally, a part of the information to be indicated can also be indicated, while the other part of the information to be indicated can be known or otherwise agreed upon. For example, the indication of information can be achieved by using the arrangement order of multiple pieces of information that are pre-agreed (for example, agreed upon by a protocol), thereby reducing the indication overhead to a certain extent. In addition, the common parts of multiple pieces of information can be identified and indicated uniformly to reduce the indication overhead caused by indicating the same information separately.

[0102] The indication method may also be various known at least methods. For example, it may be the above indication methods and various combinations thereof. The information to be indicated may have other equivalent forms, and the technical solutions provided in the embodiments of the present application should be understood to cover various forms.

[0103] The information to be indicated can be sent as a whole or divided into multiple sub-information and sent separately, and the sending period and / or sending time of the sub-information can be the same or different, which is not limited in the embodiment of the present application. The sending period and / or sending time of the sub-information can be pre-defined.

[0104] 2. In the embodiments of this application, " / " can indicate that the associated objects are in an "or" relationship. For example, "A / B" can mean A or B. "And / or" can be used to describe the existence of three relationships between the associated objects. For example, "A and / or B" can mean that A exists alone, A and B exist simultaneously, or B exists alone. A and B can be singular or plural.

[0105] 3. In the embodiments of the present application, "at least one" may refer to one or more. "Multiple" may refer to two or more, for example, two, three, four or more. Similar expressions (for example, at least one, at least one, etc.) apply the same way. "At least one of the following", "one or more of the following" or similar expressions may refer to any combination of these items, and may include only a single item or a combination of plural items. For example, at least one of a, b, or c may mean a, b, or c; a and b; a and c; b and c; a, b, and c. Among them, a, b, and c may be singular or plural.

[0106] 4. In the embodiments of the present application, the various numerical numbers involved are only used for the convenience of description and are not used to limit the scope of protection of the embodiments of the present application. The size of the serial numbers involved in the embodiments of the present application does not mean the order of execution. The order of execution of each process should be determined by its function and internal logic. For example, the terms "first", "second", "third", "fourth" and other various terminology labels (if any) in the description, claims and drawings of the embodiments of the present application can be used to distinguish similar objects, and do not have to be used to describe a specific order or sequence. Among them, the terms used in this way can be interchangeable where appropriate, so that the embodiments described here can be implemented in an order other than what is illustrated or described here, and "first", "second", "third", "fourth" and so on do not necessarily limit them to be different.

[0107] 5. In the embodiments of this application, words such as "exemplary," "example," or "for example" are used to indicate examples, illustrations, or descriptions. Any embodiment or design described as "exemplary," "example," or "for example" should not be construed as being preferred or advantageous over other embodiments or designs. Rather, the use of words such as "exemplary," "example," or "for example" is intended to present the relevant concepts in a concrete manner to facilitate understanding.

[0108] VI. In the embodiments of this application, the term "storage" or "saving" may refer to storage in one or more memories. The one or more memories may be provided separately or integrated into an encoder or decoder, a processor, or a communication device. The one or more memories may also be partially provided separately and partially integrated into a decoder, a processor, or a communication device. The memory may be any type of storage medium and is not limited in this embodiment.

[0109] 7. In the embodiments of the present application, the terms "including", "having" and any variations thereof are intended to cover non-exclusive protection. For example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to the steps or units that are clearly listed, but may include other steps or units that are not clearly listed or are inherent to these processes, methods, systems or devices.

[0110] 8. The arrows or boxes shown by dotted lines in the schematic diagrams in the accompanying drawings of the embodiment description of the present application may represent optional steps or optional modules.

[0111] 9. Unless otherwise specified or there is no logical conflict, the terms and / or descriptions between different embodiments of the present application are consistent and can be referenced by each other. The technical features in different embodiments can be combined to form new embodiments based on their internal logical relationships.

[0112] In order to clearly describe the technical solutions of the embodiments of the present application, some of the terms involved in the embodiments of the present application are explained below.

[0113] 1. A request can be used to indicate a task to be processed by the large language model. The request can include task description data describing the task. Furthermore, the request can also include prompting data. Prompting data can be used to assist the task description data, enabling the large language model to more accurately understand the request. Prompting data can include at least one of the following: contextual data, tasks, examples, format requirements, constraints, or prompting policies.

[0114] The request can be an open request or a closed request, which is not limited in this embodiment of the present application. For example, an open request can be "Please write a story." A closed request can be "In which city is FD University located?"

[0115] Tokens are the basic units used by large language models (LLMs) to understand and process text. For example, a lemma can be a word, subword, stem, affix, single character, word group, phrase, multi-word phrase, punctuation mark, identifier, or number. Lexical elements are also called tokens or tags.

[0116] A word-element may include a word-element having a predetermined function. For example, the predetermined function may include at least one of a start function and an end function. The start function may be used to indicate the beginning of a sequence. The end function may be used to indicate the end of a sequence. A word-element used to indicate the start function may be referred to as a start word-element. A word-element used to indicate the end function may be referred to as an end word-element. As an implementation, a start word-element may be represented by [BOS]. An end word-element may be represented by [EOS].

[0117] In the embodiments of the present application, the word-gram corresponding to the input data may be referred to as the input word-gram. The word-gram corresponding to the output result (or inference result) may be referred to as the output word-gram. The input data may be used to indicate the input of the large language model. The output result may be used to indicate the output of the large language model. The input data may be used to indicate the data included in the request.

[0118] 3. A word sequence may refer to at least one word arranged in order.

[0119] 4. Tokenization can refer to the process of breaking down text into a sequence of tokens.

[0120] For example, “I like watching movies” can be lemmatized to obtain a lemma sequence formed by the following multiple lemmas [“I”, “like”, “watch”, “movie”].

[0121] 5. Detokenization can refer to the process of reassembling a sequence of lemmas into a coherent text. Detokenization is the reverse process of lemmatization.

[0122] 6. A large language model is an extremely large deep learning model that is pre-trained based on a large number of training samples. A large language model can have billions or even hundreds of billions of model parameters. A large language model can also be referred to as a large-scale pre-trained language model, a large-scale language model, a large language model, or simply a large model. For simplicity, this will be referred to as a large model below.

[0123] The large model can be applied to at least one of the following: natural language processing tasks or computer vision tasks, etc. Natural language processing can be an interdisciplinary field covering multiple disciplines such as computer science, artificial intelligence and linguistics, aiming to study and develop algorithms and technologies that enable computers to understand, analyze, generate and process human language. Natural language processing can include natural language understanding (NLU) and natural language generation (NLG). Natural language understanding can refer to enabling computers to understand natural language. Natural language generation can refer to generating text in a natural language form that humans can understand based on structured data. The embodiments of this application do not limit the application of large language models. Computer vision can refer to enabling computers to understand and analyze visual information from images or videos.

[0124] As an implementation, natural language processing tasks may include at least one of the following: intelligent question answering, intelligent dialogue, machine translation, text generation, text summarization, speech synthesis, data analysis, programming assistance, document format conversion, information extraction, information retrieval, information recommendation, syntactic analysis, part-of-speech tagging, word segmentation, sentiment analysis, text classification, or speech recognition. For example, intelligent dialogue can be applied to at least one of the following: intelligent agents, intelligent assistants, chatbots, voice assistants, virtual assistants, smart home control, medical diagnosis, educational assistance, or social media interaction.

[0125] As an implementation method, computer vision tasks may include at least one of the following: visual classification, visual retrieval, object recognition, visual segmentation, visual question answering, image description, visual object detection, 3D reconstruction, object tracking, cross-modal retrieval, multimodal matching, video summarization, emotion recognition or pedestrian re-identification, etc.

[0126] Large model architectures can include either a Transformer-based model architecture or a non-Transformer-based model architecture. A Transformer can include an encoder and a decoder. The encoder encodes input data into an encoding vector. The encoding vector indicates the semantic information of the input data. The decoder generates a complete output sequence based on the encoding vector and the generated output token sequence.

[0127] As an implementation, a Transformer-based model architecture may include at least one of the following: an encoder-only model architecture, a decoder-only model architecture, or an encoder-decoder-based model architecture. A decoder-based model architecture may include at least one of the following: a causal decoder (CD)-based model architecture or a prefix decoder (PD)-based model architecture. The large model based on the encoder-based model architecture may be referred to as an encoder-based large model. The large model based on the decoder-based model architecture may be referred to as a decoder-based large model. The large model based on the encoder-decoder model architecture may be referred to as an encoder-decoder-based large model. The decoder-based large model may utilize an autoregressive (AR) approach to gradually generate an output sequence.

[0128] It should be noted that the model architecture of the large model described in the embodiment of the present application can be a model architecture based on Transformer. The model architecture based on Transformer can be a model architecture based on a standard Transformer, or a model architecture based on a non-standard Transformer, and the embodiment of the present application does not limit this. The model architecture of the non-standard Transformer can be obtained by improving the model architecture of the standard Transformer. For example, the encoder of the standard Transformer can be improved, and the decoder of the standard Transformer can also be improved.

[0129] For ease of understanding, the following Figure 1A-Figure 1C Explain Transformer. Figure 1A-Figure 1C Based on Figure 2A-2C Explain the large model based on the Transformer model architecture.

[0130] Figure 1A Schematic diagram of a converter provided for related art.

[0131] like Figure 1A As shown, the converter 100 may include an encoder 110 and a decoder 120 .

[0132] Encoder 110 may include N cascaded encoding layers, namely, encoding layers 110_1, ..., encoding layers 110_n, ..., encoding layers 110_N. Decoder 120 may include N cascaded decoding layers, namely, decoding layers 120_1, ..., decoding layers 120_n, ..., decoding layers 120_N. Decoder 120 may generate a complete output sequence based on the generated output word sequence and the encoding vector output by encoder 120_N. N may be an integer greater than or equal to 1. n may be an integer greater than or equal to 1 and less than or equal to N.

[0133] exist Figure 1A Based on the following, the coding layer 110_n and the decoding layer 120_n are taken as examples. Figure 1B For the coding layer and Figure 1C The decoding layer is described.

[0134] Figure 1B Schematic diagram of the coding layer provided for related art.

[0135] like Figure 1B As shown, the encoding layer 110_n may include a cascaded attention layer 111_n, a layer normalization (LN) layer 112_n, a feed-forward neural network (FFN) layer 113_n, and a layer normalization layer 114_n. Furthermore, the encoding layer 110_n may also include a residual connection between the input and output of the attention layer 111_n and a residual connection between the input and output of the feed-forward neural network layer 113_n. A residual connection may refer to connecting the input of a corresponding layer to the output of the corresponding layer using a direct connection channel.

[0136] The attention layer 111_n can be used to extract global features from the input. As an implementation, the attention layer 111_n can include a multi-head self-attention (MHA) layer. The feedforward neural network layer 113_n can be used to perform nonlinear transformations on the input and / or extract higher-level feature representations. As an implementation, the feedforward neural network layer 113_n can include a cascade of linear layers 1131_n, a first nonlinear activation, and a linear layer 1132_n. The linear layer 1131_n can be used to increase the dimensionality of the input. The threading layer 1132_n can be used to reduce the dimensionality of the increased features. The first nonlinear activation function can be used to perform nonlinear transformations on the input. The first nonlinear activation function can include a rectified linear unit (ReLU) function, a Swish function, a Sigmoid function, a Gaussian error linear unit (GELU) function, a gated linear unit (GLU) function, a Swiglu function, or a GeGlu function.

[0137] Figure 1C Schematic diagram of the decoding layer provided for related art.

[0138] like Figure 1C As shown, the decoding layer 120_n may include an attention layer 121_n, a layer normalization 122_n, an attention layer 123_n, a layer normalization 124_n, a feedforward neural network layer 125_n, and a layer normalization 126_n. In addition, the decoding layer 120_n may further include a residual connection between the input and output of the attention layer 121_n, a residual connection between the input and output of the attention layer 123_n, and a residual connection between the input and output of the feedforward neural network layer 125_n.

[0139] The attention layer 121_n may include a masked multi-head attention (MMHA) layer. The attention layer 123_n may include a cross-attention (CA) layer. The input of the attention layer 123_n may include Figure 1A The output of the coding layer 110_N. As an implementation method, the model structure of the feedforward neural network layer 125_n can be the same as Figure 1B The feedforward neural network layer 113_n is similar and will not be described again here.

[0140] Figure 2A A schematic diagram of a large encoder-based model provided in an embodiment of the present application.

[0141] like Figure 2A As shown, the encoder-based large model may include cascaded coding layers 210_1, ..., coding layers 210_m, ..., coding layers 210_M. As an implementation, the model structure of the coding layer 210_m may be the same as Figure 1A The middle coding layer 110_n is similar and will not be described in detail here. M may be an integer greater than or equal to 1. m may be an integer greater than or equal to 1 and less than or equal to M.

[0142] Figure 2B A schematic diagram of a large decoder-based model provided in an embodiment of the present application.

[0143] like Figure 2B As shown, the decoder-based large model may include cascaded decoding layers 220_1, ..., decoding layers 220_m, ..., decoding layers 220_M. The decoding layers included in the decoder-based large model are described below using decoding layer 220_m as an example.

[0144] According to the configuration position of the normalization layer (NL), the model structure of the decoding layer 220_m can be implemented as follows.

[0145] The first approach is based on post-layer normalization (Post-Norm), where the normalization layer can be placed after the residual connection. The decoding layer 220_m can include a cascaded attention layer 2201_m, a normalization layer 2202_m, a feedforward neural network layer 2203_m, and a normalization layer 2204_m. In addition, the decoding layer 220_m can also include a residual connection between the input and output of the attention layer 2201_m and a residual connection between the input and output of the feedforward neural network layer 2203_m.

[0146] As an implementation manner, the normalization layer may include at least one of the following: layer normalization, root mean square layer normalization (RMSNorm), or DeepNorm, etc.

[0147] The second approach is based on pre-layer normalization (Pre-Norm), whereby a normalization layer can be placed before an attention layer or a feedforward neural network layer. The decoding layer 220_m can include a cascade of a normalization layer 2205_m, an attention layer 2206_m, a normalization layer 2207_m, and a feedforward neural network layer 2208_m. Furthermore, the decoding layer 220_m can include a residual connection between the input of the normalization layer 2205_m and the output of the attention layer 2206_m, as well as a residual connection between the input of the normalization layer 2207_m and the output of the feedforward neural network layer 2208_m.

[0148] The third approach is based on sandwich-layer normalization (Sandwich-Norm), which adds a normalization layer after the residual connection based on pre-layer normalization. The decoding layer 220_m may include a cascade of a normalization layer 2209_m, an attention layer 2210_m, a normalization layer 2211_m, a normalization layer 2212_m, a feedforward neural network layer 2213_m, and a normalization layer 2214_m. In addition, the decoding layer 220_m may also include a residual connection between the input of the normalization layer 2209_m and the output of the normalization layer 2211_m, and a residual connection between the input of the normalization layer 2212_m and the output of the normalization layer 2214_m.

[0149] Figure 2C A schematic diagram of a large encoder-decoder model provided in an embodiment of the present application.

[0150] like Figure 2C As shown, the encoder-decoder based large model may include an encoder 210 and a decoder 220. The encoder 210 may include cascaded encoding layers 210_1, ..., encoding layers 210_m, ..., encoding layers 210_M. The decoder 220 may include cascaded decoding layers 220_1, ..., decoding layers 220_m, ..., decoding layers 220_M. As an implementation, the model structure of the encoding layer 210_m may be the same as Figure 1A Similar to the coding layer 110_n, the model structure of the decoding layer 220_m can be the same as Figure 2B Similar, no further description is given here.

[0151] 7. The inference phase of a large model may refer to the phase in which the large model processes input data for P rounds to obtain an output result. P may be an integer greater than or equal to 1. P may be determined based on an end condition. The end condition may include at least one of the following: reaching a maximum sequence length, outputting an end token, or determining that a logical output result has been obtained. Thus, P may be the number of rounds corresponding to obtaining the maximum sequence length, the number of rounds corresponding to obtaining an end token, or the number of rounds corresponding to obtaining a logical output result.

[0152] As an implementation, the inference phase of a large model can include a prefill phase and a decoding phase. The prefill phase can refer to the phase in which the first round of processing is performed. During the prefill phase, the large model can be used to process the input data to obtain the first output word and the intermediate state corresponding to the input data. The first output word can refer to the first output word. The intermediate state can include a key and a value. The decoding phase can refer to the phase in which the second to Pth rounds of processing are performed. During the decoding phase, the large model can be used to process the input data, the first output word, and the intermediate state corresponding to the input data using an autoregressive approach to obtain the output result.

[0153] The pre-filling stage and the decoding stage are further described below.

[0154] In the pre-population phase, when p=1, the large model can be used to process at least one input word in parallel to obtain a first output word and an intermediate state corresponding to each of the at least one input word. The at least one first input word can be obtained by tokenizing the input data.

[0155] During the decoding phase, i.e., when 1 < p ≤ P, the large model is used to obtain the pth output word and the intermediate state corresponding to the pth word based on at least one input word, the first output word through the p-1th output word, the intermediate state corresponding to each of the at least one input word, and the intermediate state corresponding to each of the first output word through the p-1th output word. Thus, the second output word through the pth output word are obtained.

[0156] Based on the above, the output result is obtained according to the first output word-gram to the P-th output word-gram.

[0157] It should be noted that the pre-filling stage may also be referred to as the pre-processing stage, initialization stage, prompting stage, full decoding stage, full inference stage, full stage, or context stage. The decoding stage may also be referred to as the generating stage, incremental decoding stage, or incremental stage. In the embodiments of the present application, the first stage may be used to indicate the pre-filling stage. The second stage may be used to indicate the decoding stage.

[0158] In order to facilitate the understanding of the reasoning stage of the large model, the following Figure 3 Provide explanation.

[0159] Figure 3 A schematic diagram of the application of the reasoning process of the large model provided in the embodiment of the present application.

[0160] like Figure 3 As shown in the figure, the input data is "In which city is FD University located?" The large model processes the input data for P rounds to obtain the output result. P = 5.

[0161] Regarding how to use the large model to obtain output results, it can be achieved in the following ways.

[0162] In the pre-population phase, when p = 1, the question "In which city is FD University located" is tokenized, resulting in the following input tokens: ["FD"], ["university"], ["located in"], ["which city"], and ["city"]. The large model processes these multiple input tokens in parallel, resulting in the first output token ["FD"] and the intermediate state corresponding to the first output token ["FD"].

[0163] During the decoding phase, when 1 < p ≤ P, the large model is used to obtain the pth output word and the intermediate state corresponding to the pth word based on multiple input words, the first output word through the p-1th output word, the intermediate state corresponding to at least one input word, and the intermediate state corresponding to the first output word through the p-1th output word. This yields the second output word through the pth output word. The second output word is ["University"]. The third output word is ["Located"]. The fourth output word is ["SH"]. The fifth output word is ["EOS"].

[0164] Based on the above, according to the 1st output word-gram to the Pth output word-gram, the output result "FD University is located in SH" is obtained.

[0165] It should be noted that Figure 3 Intermediate states are not shown.

[0166] 8. Parameter sparsity can mean that only some model parameters are involved in task processing.

[0167] 9. A mixture of experts (MoE) is a model architecture that dynamically selects some model parameters to participate in task processing, which can improve computational efficiency and model performance.

[0168] A hybrid expert can include at least one gating module and multiple expert modules. The model structures of the multiple expert modules can be the same, but the model parameters of the multiple expert modules can be different. Optionally, the model structures of the multiple expert modules can also be different. Different expert modules can learn different feature representations. The hybrid expert can assign word units to active expert modules for processing through the gating module. An active expert module can be an expert module selected by the gating module. An inactive expert module can be an expert module not selected by the gating module. As a result, the hybrid expert has parameter sparsity.

[0169] The gating module can be used to generate weights corresponding to multiple expert modules based on word elements. The weights corresponding to the expert modules can be used to indicate the degree of contribution of the expert modules to the output results. The weights can be expressed by activation probabilities. Thus, the gating module can determine at least one target expert module from multiple expert modules based on the weights corresponding to the multiple expert modules. The target expert module can refer to an expert module in an activated state. The target expert module can be used to process word elements. The model structure of the gating module can be configured according to actual business needs, and the embodiments of the present application do not limit this. As an implementation method, the gating module may include a cascaded linear layer and a third nonlinear activation function. The gating module may also be called a routing module or a router.

[0170] The expert module can be used to perform nonlinear transformation and / or feature extraction on the input. As an implementation method, multiple expert modules can all be non-shared expert modules. Optionally, multiple expert modules can include at least one shared expert module and multiple non-shared expert modules. The non-shared expert module can refer to an expert module that is used to process word units when activated. The shared expert module can be used to cooperate with the non-shared expert module to process word units. When the multiple expert modules include at least one shared expert module and multiple non-shared expert modules, the gating module can be used to determine at least one target expert module from the multiple non-shared expert modules. The model structure of the expert module can be configured according to actual business needs, and this embodiment of the present application does not limit this. As an implementation method, the expert module can be similar to a feedforward neural network layer. The number of target expert modules can be configured according to actual business needs, and this embodiment of the present application does not limit this. For example, the number of target expert modules can be 2.

[0171] Thus, when multiple expert modules are non-shared, the gating module can determine the weights corresponding to each of the multiple expert modules based on the word element, and determine at least one target expert module from the multiple expert modules based on the weights corresponding to each of the multiple experts. The target expert module can obtain the output result corresponding to the target expert module based on the word element. The hybrid expert can obtain the output result corresponding to the word element based on the weights corresponding to each of the at least one target expert module and the output result corresponding to each of the at least one target expert module.

[0172] When the multiple expert modules include at least one shared expert module and multiple non-shared expert modules, the gating module can determine the weights corresponding to each of the multiple non-shared expert modules based on the word element, and determine at least one target expert module from the multiple non-shared expert modules based on the weights corresponding to each of the multiple non-shared expert modules. The modular expert module obtains the output result corresponding to the target expert module based on the word element. The shared expert module can obtain the output result corresponding to the shared expert module based on the word element. The hybrid expert can obtain the output result corresponding to the word element based on the weight corresponding to each of the at least one non-shared expert modules, the output result corresponding to each of the at least one non-shared expert modules, and the output result corresponding to the shared expert module.

[0173] For ease of understanding, the following Figures 4A-4C Explain the hybrid expert. Figure 4A Multiple expert modules are non-shared expert modules. Figure 4B The multiple expert modules include a shared expert module and multiple non-shared expert modules. Figure 4C Shown Figure 4A and Figure 4B Model structure of the expert module.

[0174] Figure 4A A schematic diagram of a hybrid expert provided in an embodiment of the present application.

[0175] like Figure 4A As shown, hybrid expert 400 may include gating module 410 and Q expert modules. The Q expert modules may include expert module 420_1, ..., expert module 420_q, ..., and expert module 420_Q. Q may be an integer greater than 1, for example, Q=8. q may be an integer greater than or equal to 1 and less than or equal to Q.

[0176] The gating module can determine the weights corresponding to each of the Q expert modules based on the word elements, and then determine K target expert modules from the Q expert modules based on the weights corresponding to each of the Q experts. K = 2. For example, K weights can be determined from the Q weights. The expert modules corresponding to each of the K weights are determined as target expert modules. Any of the K weights is greater than any of the Q weights other than the K weights. Since the weight corresponding to expert module 420_1 and the weight corresponding to expert module 420_q are both greater than any of the Q weights other than these two weights, expert module 420_1 and expert module 420_q are the target expert modules. q = 4.

[0177] Based on this, expert module 420_1 obtains an output result corresponding to expert module 420_1 based on the word-gram. Expert module 420_q obtains an output result corresponding to expert module 420_q based on the word-gram. Hybrid expert 400 obtains a weighted result corresponding to expert module 420_1 based on the weights and output results corresponding to expert module 420_1. Hybrid expert 400 obtains a weighted result corresponding to expert module 420_q based on the weights and output results corresponding to expert module 420_q. Hybrid expert 400 obtains an output result corresponding to the word-gram based on the weighted results corresponding to expert module 420_1 and expert module 420_q.

[0178] Figure 4B A schematic diagram of another hybrid expert provided in an embodiment of the present application.

[0179] like Figure 4B As shown, hybrid expert 400 may include a gating module 410 and Q+1 expert modules. The Q+1 expert modules may include a shared expert module 420_0 and Q non-shared expert modules. The Q non-shared expert modules may include expert modules 420_1, ..., expert modules 420_q, ..., and expert modules 420_Q. Q may be an integer greater than 1, for example, Q=8. q may be an integer greater than or equal to 1 and less than or equal to Q.

[0180] The gate control module can be used with Figure 4A In a similar manner, the expert module 420_1 and the expert module 420_q are determined as target expert modules from the Q expert modules, which will not be described in detail here.

[0181] Based on this, the shared expert module 420_0 can obtain the output result corresponding to the shared expert module 420_0 according to the word unit. The hybrid expert 400 obtains the output result corresponding to the word unit according to the weighted results corresponding to the expert module 420_1 and the expert module 420_q, as well as the output result corresponding to the shared expert module. Figure 4AIn a similar manner, weighted results corresponding to the expert module 420_1 and the expert module 420_q are determined, which will not be repeated here.

[0182] Figure 4C A schematic diagram of the expert module provided in an embodiment of the present application.

[0183] like Figure 4C As shown, the model structure of the expert module 420_q can be implemented as follows.

[0184] In a first embodiment, expert module 420_q may include a cascaded linear layer 4201_q, a first nonlinear activation function, and a linear layer 4202_q. For details about linear layer 4201_q, linear layer 4202_q, and the first nonlinear activation function, refer to the description of feedforward neural network layer 113_n above and are not further described here.

[0185] In a second manner, the expert module 420_q may include a linear layer 4203_q, a linear layer 4204_q, and a second nonlinear activation function, which may include SwiGLU.

[0186] 10. The MoE model may refer to a large model obtained by replacing at least one feedforward neural network layer in a large model with a hybrid expert layer (i.e., hybrid expert). For example, the MoE model may include at least one hybrid expert layer but not a feedforward neural network layer. Optionally, the MoE model may include at least one hybrid expert layer and at least one feedforward neural network layer, which is not limited in this embodiment of the present application. The large model may be the decoder-based large model or the encoder-decoder-based large model described above. The model structures of different hybrid expert layers may be the same or different, which is not limited in this embodiment of the present application.

[0187] The MoE model can scale up by increasing the number of expert modules, thereby improving model performance. Because the mixed expert layer has parameter sparsity, it can reduce computational overhead and improve computational efficiency (or inference efficiency) compared to dense models of the same size. Furthermore, because multiple expert modules are independent of each other, they can be run in parallel if the hardware supports it, further improving computational efficiency. Because different expert modules can focus on different features, the MoE model's performance can be improved.

[0188] For ease of understanding, the following Figure 5 Provide explanation.

[0189] Figure 5 Schematic diagram of the MoE model provided in the embodiments of the present application.

[0190] like Figure 5As shown in Figure 2, the feedforward neural network layer in the large model can be replaced by the hybrid expert layer to obtain the MoE model. The description of the hybrid expert layer can be found in the corresponding section above and will not be repeated here.

[0191] 11. End-side devices may refer to devices deployed near data sources. End-side devices are capable of processing local data without relying on cloud devices. End-side devices have at least one of the characteristics of local computing, low latency, low power consumption, reliability, availability, or privacy protection. In addition, the resources of end-side devices are limited. Resources may include computing resources and storage resources, etc. End-side devices may also be called terminal equipment (TE), user equipment (UE), edge device, mobile device, mobile terminal (MT), user terminal, terminal, access terminal, remote terminal, user agent, user device or mobile station (MS), etc.

[0192] The end-side device may include at least one of the following: a mobile phone, a tablet computer, a portable computer, a desktop computer, an end-side device in autonomous driving (AD), an end-side device in the Internet of Things (IoT), an end-side device in vehicle-to-everything (V2X), an in-vehicle terminal device, a mobile internet device (MID), a virtual reality (VR) device, an augmented reality (AR) device, a mixed reality (MR) device, a personal digital assistant (PDA), a customer premises equipment (CPE), a drone, a helicopter, an airplane, a ship, a robot, a robotic arm, an end-side device in device-to-device (D2D) communication, an end-side device in industrial control (IC), and a machine-type communication (M2C). The terminal side devices in the mobile communication (MTC), tactile terminal devices, terminal side devices in remote medical (RM), terminal side devices in smart grid (SG), terminal side devices in transportation safety (TS), terminal side devices in intelligent transportation, terminal side devices in smart city (SC), terminal side devices in smart home (SH), terminal side devices in smart office or wearable devices, etc., are not limited to the embodiments of the present application.

[0193] Wearable devices are a general term for wearable devices developed by applying wearable technology to intelligently design everyday wearables. For example, wearable devices may include at least one of the following: head-mounted displays (HMDs), glasses, gloves, watches, clothing, or shoes. A wearable device is a portable device that is worn directly on the body or integrated into the user's clothing or accessories. Wearable devices can achieve powerful functions through software support, data exchange, and cloud interaction. Broadly speaking, wearable devices include devices that are full-featured, large in size, or can achieve full or partial functions independently of a smartphone, such as smart watches or smart glasses, as well as devices that focus on a specific type of application function and require use with other devices such as smartphones, such as various smart bracelets or smart jewelry for vital sign monitoring. For example, smart glasses may include at least one of the following: VR glasses or AR glasses. Transmittable devices can also be referred to as wearable smart devices.

[0194] 12. The end-side model may refer to a large model deployed on an end-side device. The end-side model may also be called a terminal model, a mobile end model, a client model, a device-side model, a local model, an edge model, or an embedded model. The end-side model has at least one of the characteristics of low latency, bandwidth saving, privacy protection, security, reliability, availability, or personalization. The end-side model can be applied to at least one of natural language processing tasks or computer vision tasks. For the description of natural language processing tasks and computer vision tasks, please refer to the corresponding sections above and will not be repeated here. In the embodiment of the present application, the first model may refer to the end-side model.

[0195] 13. An accelerator refers to a hardware device specifically designed to accelerate a predetermined computing task. An accelerator can work in conjunction with a general-purpose processor, such as a central processing unit (CPU).

[0196] An accelerator may include at least one of the following: a graphics processing unit (GPU), a neural processing unit (NPU), a tensor processing unit (TPU), a field-programmable gate array (FPGA), or an application-specific integrated circuit (ASIC). A GPU has a large number of GPU cores and is suitable for highly parallel computations. In the embodiments of this application, an accelerator may also be referred to as a hardware accelerator.

[0197] 14. The storage structure of the end-side device may include memory and storage devices. Memory may include at least one of the following: processor memory or accelerator memory. Accelerator memory may include at least one of the following: graphics processor memory, neural network processor memory, or tensor processor memory. Storage devices may include at least one of the following: flash memory (FM), embedded multimedia card (eMMC), universal flash storage (UFS), solid state drive (SSD), or hard disk drive (HDD).

[0198] 15. Cache can refer to an area in memory used to store expert parameters. Expert parameters can be used to indicate model parameters of an expert module. In this embodiment of the present application, the model parameters of an expert module can be referred to as the expert module. Therefore, cache can also refer to an area in memory used to store expert modules.

[0199] 16. A cache hit (CH) can refer to the situation where the target expert module exists in memory. A cache miss (CM) can refer to the situation where the target expert module does not exist in memory. The target expert module can refer to the expert module in the active state.

[0200] 17. Memory access speed may refer to the time it takes a processor or accelerator to access memory. Memory access speed may be affected by at least one of memory latency, memory bandwidth, memory frequency, memory type, or memory channel. Memory frequency may refer to the memory clock frequency.

[0201] 18. I / O latency can refer to the time it takes from issuing an input / output (I / O) request to receiving a response. For example, I / O latency can include the time it takes for the accelerator memory to retrieve data from the processor memory or storage device to the expert module.

[0202] 19. An I / O bottleneck can refer to a situation where input / output operations become the performance limiting factor. I / O latency can lead to an I / O bottleneck.

[0203] 20. Parallelism can refer to the ability to execute multiple tasks at the same time.

[0204] 21. A thread is the smallest unit of execution and scheduling recognizable by the operating system. A thread can execute tasks independently. For example, a multi-core processor can execute multiple threads in parallel. Alternatively, multiple processors can execute multiple threads in parallel.

[0205] 22. Pipelining is a method of parallel computing that improves the execution efficiency of large models. For example, I / O computational pipelining can refer to a method of improving system performance by overlapping input / output operations with computation.

[0206] 23. A linked list (LL) is a linear data structure that includes at least one node. A node may include data and at least one pointer. The pointer may be used to point to a target node. A linked list may include at least one item from a unidirectional linked list or a bidirectional linked list. In the case of a unidirectional linked list, a node may include data and a pointer. The target node may point to the next node. In the case of a bidirectional linked list, a node may include data and two pointers. The target node may be the previous node or the next node.

[0207] 24. An array is a linear data structure that can include a set of continuous memory spaces for storing elements of the same type.

[0208] 25. A queue is a linear list that allows insertions at one end and removals at the other, following the first-in, first-out (FIFO) principle. The end where insertions are allowed is called the tail, and the end where removals are allowed is called the head. The data in the queue is called the queue element. Consecutive storage units can be allocated to store queue elements, and pointers are used to indicate the head and tail elements. Queues can be implemented using arrays or linked lists.

[0209] 26. A multi-core processor can be a processor that integrates multiple cores into a single chip to provide external services. A multi-core processor can also be called an on-chip multiprocessor or a single-chip multiprocessor.

[0210] A multi-core processor can include multiple computing cores, integrating multiple single-threaded computing cores or multiple multi-threaded computing cores. Deploying multiple cores on the same chip shortens the connections between cores, reduces communication latency between cores, and improves communication efficiency. Furthermore, multiple cores share on-chip resources, thereby improving on-chip resource utilization and reducing computing core power consumption.

[0211] 27. Cache eviction policy (CEP) refers to a strategy for determining which data to evict (or remove) from the cache in order to make room for new data when cache space is limited.

[0212] Cache eviction policies may include at least one of the following: first-in-first-out policy, least recently used (LRU) policy, LRU-K policy, least frequently used (LFU) policy, most recently used (MRU) policy, random replacement (RR) policy, time to live (TTL) policy, two-queue (2Q) policy, adaptive replacement cache (ARC) policy, or low inter-reference recency set (LIRS) policy.

[0213] The first-in-first-out strategy can be used to prioritize the elimination of data that entered the cache the earliest. The least recently used strategy can be used to prioritize the elimination of data that has not been accessed for the longest time. The LRU-K strategy can be used to eliminate the data that has not been accessed for the longest time from the most recent K accesses. The least frequently used strategy can be used to prioritize the elimination of data with the lowest access frequency. The most recently used strategy can be used to prioritize the elimination of data with the most recent access. The random elimination strategy can be used to randomly select data for elimination. The survival time strategy can be used to eliminate data that has exceeded the life cycle corresponding to the storage item. The secondary cache strategy can be used to determine the data to be eliminated based on the most recently used data and the frequently accessed data. The adaptive replacement cache strategy can be used to dynamically adjust the allocation of cache space. The low interactive reference interval strategy can be used to optimize cache elimination based on the low interactive reference interval.

[0214] The inventive concept of the embodiments of the present application will be described below with reference to the accompanying drawings.

[0215] With the development of terminal and artificial intelligence technologies, large models can be deployed on edge devices to provide users with a better AI experience by leveraging at least one of the following features: local computing, low latency, low power consumption, security, reliability, availability, or privacy protection. Large models deployed on edge devices are referred to as edge models. Edge models can be applied to at least one of natural language processing tasks or computer vision tasks.

[0216] Because large models have large parameters, they have high storage and computing requirements. However, due to the limited storage and computing resources of end-side devices, deploying large models on these devices is difficult. To obtain a large model that can be deployed on end-side devices (i.e., an on-end model), it is necessary to consider how to adapt the large model to the end-side devices.

[0217] Because the MoE model can reduce computing overhead and improve model performance, it can be used as an end-to-end model. However, there is a significant gap between the memory requirements of the MoE model and the memory capacity of the end-to-end device. Therefore, it is necessary to consider how to deploy the MoE model on the end-to-end device.

[0218] In order to better understand the embodiments of the present application, Figure 6 The end-side model described in the embodiments of this application is described below. In the embodiments of this application, the end-side model may be referred to as a first model. The first model may be a MoE model. The first model may include multiple cascaded first layers. The first layer may include a hybrid expert layer. The hybrid expert layer may include a gating module and multiple expert modules. For example, the first model may include R cascaded first layers. The hybrid expert layer may include S expert modules. R may be an integer greater than or equal to 1. S may be an integer greater than 1. The first layer currently being used may be referred to as the current layer.

[0219] Figure 6 A schematic diagram of the first model provided in an embodiment of the present application.

[0220] like Figure 6 As shown, the first model 600 may include R cascaded first layers, namely, first layer 600_1, ..., first layer 600_r, ..., first layer 600_R. r can be an integer greater than or equal to 1 and less than or equal to R. The model structures of different first layers may be the same or different, and the embodiment of the present application does not limit this. For example, the R first layers may all include a hybrid expert layer. Optionally, some of the R first layers may include a hybrid expert layer, and another part of the first layers may include a feedforward neural network layer. The model structures of different hybrid expert layers may be the same or different, and the embodiment of the present application does not limit this.

[0221] The first layer 600_r is taken as an example to describe the first layer 600_r, which can be understood as the current layer.

[0222] The first layer 600_r may include a hybrid expert layer 610_r. The hybrid expert layer 610_r may include a gating module 611_r and S expert modules. The S expert modules may include expert modules 612_r_1, ..., expert modules 612_r_s, ..., and expert modules 612_r_S. s may be an integer greater than or equal to 1 and less than or equal to S.

[0223] The gating module 611_r can be used to determine at least one first expert module from among the expert modules 612_r_1, ..., expert modules 612_r_s, ..., and expert modules 612_r_S. A first expert module can refer to an activated expert module. A first expert module can be used to process input (e.g., a word). The number of first expert modules can be configured based on actual business needs and is not limited here. For example, the number of first expert modules can be two. For further details on the hybrid expert layer 610_r, please refer to the above description of the hybrid expert layer and will not be repeated here.

[0224] In addition, the first layer 600_r may also include an attention layer and a normalization layer. For the description of the attention layer and the normalization layer, please refer to the corresponding parts above and will not be repeated here.

[0225] In this embodiment of the present application, the model structure in the first model other than the hybrid expert layer can be referred to as a non-hybrid expert layer. Thus, the first model can include both a non-hybrid expert layer and a hybrid expert layer. Computations involving the non-hybrid expert layer are referred to as non-expert computations. Computations involving expert modules are referred to as expert computations. The first expert module can be used to indicate the expert modules that are active in the first model during the second phase. The second expert module can be used to indicate the expert modules that are active in the first model during the first phase. The second phase can be used to indicate the decoding phase. The first phase can be used to indicate the pre-filling phase.

[0226] Next, combine Figure 6 The first model shown explains how to deploy the first model on the terminal side device.

[0227] In order to deploy the first model on the end-side device, the storage structure of the end-side device needs to be considered. Figure 7A Provide explanation.

[0228] Figure 7A A schematic diagram of the storage structure of the terminal device provided in an embodiment of the present application.

[0229] like Figure 7AAs shown, the storage structure of the edge device may include processor memory and storage devices. In addition, the edge device may also include accelerator memory. For example, the accelerator memory may include at least one of the following: graphics processor memory, neural network processor memory, or tensor processor memory.

[0230] Data transfer between processor memory and accelerator memory can be achieved through the processor and accelerator. The processor can transfer data to processor memory through the memory bus and memory controller. The accelerator can transfer data to accelerator memory through the video memory bus and video memory controller.

[0231] The processor and storage device can exchange data through the storage interface, internal bus, and storage controller. The accelerator can also exchange data with the storage device through the processor. Alternatively, the accelerator can directly exchange data with the storage device.

[0232] The storage capacity of the storage device may be greater than at least one of the storage capacity (or memory capacity) of the accelerator memory or the storage capacity (or memory capacity) of the processor memory.

[0233] In an embodiment of the present application, as one implementation, the first storage unit may be accelerator memory. The second storage unit may be processor memory or a storage device. As another implementation, the first storage unit may be processor memory. The second storage unit may be a storage device. The first storage unit may be used as a cache for the first expert module and / or the second expert module.

[0234] Since the end-side device has a multi-level storage structure, the embodiment of the present application proposes that the model parameters can be stored in different storage structures to solve the problem of limited memory capacity of the end-side device. For example, the memory capacity can be expanded based on the memory unloading method. Specifically, the model parameters that are not currently involved in the calculation can be unloaded to the second storage unit. In the case where the model parameters need to participate in the calculation, the model parameters are loaded from the second storage unit to the first storage unit, thereby reducing the memory requirements of the first model. However, since the model parameters need to be loaded from the second storage unit to the first storage unit, I / O delay will be generated.

[0235] exist Figure 7A Based on the following Figure 7B-7D The invention describes expanding memory capacity and I / O latency between a first storage unit and a second storage unit based on a memory offloading method. Figure 7B The first storage unit may be an accelerator memory, and the second storage unit may be a processor memory. Figure 7C The first storage unit may be an accelerator memory, and the second storage unit may be a storage device. Figure 7DThe first storage unit may be a processor memory. The second storage unit may be a storage device. The first storage unit may store a non-hybrid expert layer (i.e., non-hybrid expert parameters). The second storage unit may store a hybrid expert layer (i.e., hybrid expert parameters). The non-hybrid expert parameters may be used to indicate model parameters of the non-hybrid expert layer. The hybrid expert parameters may be used to indicate model parameters of the hybrid expert layer. Figure 7B-7D This is merely an exemplary description and does not constitute a limitation on the embodiments of the present application. For example, the second storage unit may store a non-hybrid expert layer and a hybrid expert layer.

[0236] Figure 7B A schematic diagram of a storage method of a first model provided in an embodiment of the present application.

[0237] like Figure 7B As shown, if an expert module in the hybrid expert layer is determined to be the first expert module, the processor and accelerator can be used to load the first expert module (i.e., first expert parameters) from the second storage unit to the first storage unit. This results in an I / O delay between the second storage unit and the first storage unit. The first expert parameters can be used to indicate the model parameters of the first expert module.

[0238] Figure 7C A schematic diagram of another storage method of the first model provided in an embodiment of the present application.

[0239] like Figure 7C As shown, Figure 7B Similarly, the first expert module can be loaded from the second storage unit to the first storage unit through the processor, accelerator and storage device, thereby generating an I / O delay between the second storage unit and the first storage unit.

[0240] Figure 7D A schematic diagram of another storage method of the first model provided in an embodiment of the present application.

[0241] like Figure 7D As shown, Figure 7B Similarly, the first expert module can be loaded from the second storage unit to the first storage unit through the processor and the storage device, thereby generating an I / O delay between the second storage unit and the first storage unit.

[0242] Since there is an I / O delay between the first storage unit and the second storage unit, the reasoning efficiency of the first model is affected. The following describes how to improve the reasoning efficiency in the embodiment of the present application with reference to the accompanying drawings.

[0243] Figure 8A A schematic diagram of the principles of a data processing method provided in an embodiment of the present application.

[0244] like Figure 8A As shown, when deploying end-side devices based on memory offloading, the data processing process involves non-expert calculations, loading the first expert module from the second storage unit to the first storage unit, and expert calculations. Because there is I / O latency when loading the first expert module from the second storage unit to the first storage unit, the inference efficiency of the first model is affected.

[0245] To improve inference efficiency, it has been discovered that preloading and computation can be used in parallel to mitigate I / O latency. For example, in the second phase of the first model (i.e., the decoding phase), while performing expert computation for the previous word, the second thread can determine a fifth expert module corresponding to the current word from the multiple expert modules included in the first model based on the expert statistical distribution and the first expert module corresponding to the previous word. The fifth expert module corresponding to the current word is preloaded into the first storage unit. The first thread determines the first expert module corresponding to the current word based on the first model. If a fifth expert module corresponding to the first expert module exists in the first storage unit, the thread retrieves the fifth expert module corresponding to the first expert module from the first storage unit. Expert computation is then performed using the fifth expert module corresponding to the first expert module. This approach is expected to enable the preloading of the fifth expert module corresponding to the current word and the execution of the expert computation for the previous word to mitigate I / O latency and reduce computing resource latency during the loading of the first expert module.

[0246] To demonstrate the effectiveness of the above implementation, this application used a Mixtral MoE-175G model, which occupied 19GB of processor memory and was processed using four processor cores, as an example for experimental verification. The Mixtral MoE-175G has 175 billion model parameters. The experimental results showed that the average time consumed to decode a word was reduced from 26.6 seconds to 21.9 seconds.

[0247] However, it was found that when it was necessary to determine whether the expert module corresponding to the first expert module participated in the calculation, the expert module was still in the loading state, which shows that the above implementation method failed to cover the I / O delay caused by loading the expert module. Further research found that this may be caused by at least one of the I / O bottleneck and preloading failure. The I / O bottleneck can refer to the expert calculation being limited by the I / O delay. The preloading duration is longer than the expert calculation duration. The preloading failure can refer to the expert module (i.e., the fifth expert module) stored in the first storage unit based on the preloading method is not the first expert module actually required. The preloading failure can be understood as a cache miss.

[0248] To improve inference efficiency, it was discovered that increasing the cache hit rate can be achieved. A cache hit means that when a first expert module is needed, it is present in the first storage unit. Therefore, in the case of a cache hit, the first expert module can be directly retrieved from the first storage unit without having to load it from the second storage unit to the first storage unit, thereby reducing I / O latency.

[0249] To improve cache hit rates, we found that storing as many expert modules as possible in the first storage unit would not only reduce I / O latency between the second storage unit and the first, but also improve cache hit rates. However, due to limited memory capacity, increasing the number of expert modules in the first storage unit is difficult to achieve.

[0250] It is further found that the use of expert modules has load balancing for the entire data set. Therefore, there may be expert preferences for requests. Therefore, the expert modules that need to be stored in the first storage unit can be determined based on the expert preference strategy of the request. Expert preference can refer to having expert modules corresponding to requests, and the expert modules corresponding to different requests may be different. For example, the expert modules corresponding to the first request A may include expert module 1 and expert module 2. The expert modules corresponding to the first request B may include expert module 3 and expert module 4. The expert modules corresponding to the first request A and the first request B are different. It can be understood that different requests may refer to the similarity between two requests being less than or equal to the sixth threshold. The sixth threshold can be configured according to actual business needs and is not limited here.

[0251] Therefore, in order to improve the cache hit rate, the embodiment of the present application proposes that an expert preference prediction process based on a request can be adopted to implement it, so that the first storage unit can store the expert module corresponding to the third expert module. The third expert module can be determined from the multiple expert modules included in the first model based on the expert preference strategy of the first request. The third expert module corresponding to the first request may be an expert module that is frequently accessed to process the first request. The third expert module determined by the request-based expert preference preloading process can be used in the decoding stage of the first model (i.e., the second stage), and can obtain the first expert module corresponding to the first request from the first storage unit as much as possible, thereby improving the cache hit rate of the decoding stage.

[0252] Next, we need to consider how to combine the request-based expert preference preloading process with the inference phase process of the first model to further reduce the I / O latency of the decoding phase.

[0253] We further discovered that, based on the above description of the inference phase of the large model, during the pre-population phase (i.e., the first phase), the first model can be used to process at least one input token in parallel. Therefore, although the first model has high parameter sparsity for each input token, since different input tokens may activate different expert modules, the parallel processing approach reduces the overall parameter sparsity of the first model. For ease of understanding, based on the above description, the expert module activated during the pre-population phase (i.e., the first phase) can be referred to as the second expert module.

[0254] Therefore, during the pre-filling phase, each expert module included in the first model may be an activated expert module, that is, the expert modules included in the first model may all be second expert modules. Therefore, it may be necessary to load all expert modules included in the first model from the second storage unit to the first storage unit, that is, it may be necessary to load all second expert modules included in the first model from the second storage unit to the first storage unit. In addition, there may be situations where an expert module is loaded into the first storage unit multiple times. This is because, if the memory capacity is limited, if the expert module is activated, the expert module (i.e., the second expert module) can be loaded from the second storage unit to the first storage unit. If the expert module is not activated, it can be removed (or eliminated) from the first storage unit.

[0255] Based on this, to further reduce I / O latency, it was discovered that the pre-population phase could be executed in parallel with the request-based expert preference preloading process. This was based on the following considerations: the request-based expert preference preloading process can include a request-based expert preference prediction process and an expert loading process. Specifically, the request-based expert preference prediction process is executed to obtain a third expert module corresponding to the first request, and then the expert loading process is executed to load the expert module corresponding to the third expert module from the second storage unit to the first storage unit. Therefore, the expert loading process will generate I / O latency. As mentioned above, during the pre-population phase of the first model, it may be necessary to load all the expert modules of the first model (i.e., the second expert modules) from the second storage unit to the first storage unit. Based on this, the request-based expert preference preloading process can take advantage of the fact that the pre-fill phase process requires loading the second expert module into the first storage unit and adjust the expert loading process to an expert retention process. That is, during the execution of the pre-fill phase process to load the second expert module from the second storage unit to the first storage unit, the expert retention process is executed to retain the second expert module corresponding to the third expert module. This can minimize the I / O delay caused by the expert loading process in the request-based expert preference preloading process. The expert module predicted to be active during the pre-load phase can be referred to as the third expert module. The first speed can be greater than the second speed. The first speed can be used to indicate the speed at which the first thread accesses the first storage unit. The second speed can be used to indicate the speed at which the first thread accesses the second storage unit.

[0256] Based on the above, continue to see Figure 8A This embodiment of the present application proposes parallel execution of the pre-filling phase and the request-based expert preference preloading process to improve cache hit rates during the decoding phase, reduce I / O latency, and thereby enhance inference efficiency. A first thread can execute the pre-filling phase, while a second thread executes the request-based expert preference preloading process. Furthermore, the first thread can also execute the decoding phase.

[0257] The following combination Figure 8B 、 8C 8D, how to implement the parallel execution of the pre-filling phase process and the request-based expert preference preloading process. For the sake of brevity, the request-based expert preference preloading process is referred to as the preloading process below.

[0258] Figure 8B A schematic diagram of the principles of another data processing method provided in an embodiment of the present application.

[0259] like Figure 8BAs shown, the first thread can be used to execute the first stage process (ie, the pre-filling stage process) and the second stage process (ie, the decoding stage process). The second thread can be used to execute the pre-loading process.

[0260] The first-stage process may include sub-processes corresponding to at least one third word-element. At least one sub-process may be executed in parallel. The sub-process may include non-expert calculation, determining a second expert module, loading the second expert module from the second storage unit to the first storage unit, and expert calculation. The third word-element may be determined by the first thread based on the first request. The third word-element may be used to indicate an input word-element.

[0261] The second stage process may include non-expert calculation, determining the first expert module, obtaining the second expert module and expert calculation corresponding to the first expert module from the first storage unit, etc.

[0262] The preloading process may include request-based expert preference prediction and retaining the second expert module corresponding to the third expert module in the first storage unit. Request-based expert preference prediction may be implemented as follows: As an implementation, multiple third expert modules corresponding to the first request may be determined based on the first request.

[0263] The first thread can load the second expert module from the second storage unit to the first storage unit. The second expert module can be determined by the first thread based on the first request using the first model. While the first thread executes the first-stage process, the second thread can retain the second expert module corresponding to the third expert module in the first storage unit. This allows the preloading process executed by the second thread and the first-stage process executed by the first thread to be executed in parallel, preserving the second expert module corresponding to the third expert module in the first storage unit and thus suppressing the I / O latency of the preloading process.

[0264] Because the third expert module is an expert module that may be accessed when the first thread executes the second stage process, the first thread can retrieve the first expert module from the multiple second expert modules in the first storage unit when executing the second stage process using the first model, and then use the first expert module to perform expert calculations. Because the first thread does not need to load the first expert module corresponding to the first expert module from the second storage unit to the first storage unit when executing the second stage using the first model, this reduces the I / O latency between the first and second storage units, thereby improving the inference efficiency of the first model.

[0265] The following combination Figure 8C and Figure 8D , describes how to predict expert preferences based on requests.

[0266] Figure 8C A schematic diagram of the principle of a request-based expert preference prediction method provided in an embodiment of the present application.

[0267] like Figure 8C As shown, a second model can be used to obtain second scores corresponding to each of the multiple expert modules based on the first request. Multiple third expert modules are determined based on the second scores corresponding to each of the multiple expert modules. The second scores can be used to indicate a predicted activation score for the expert module. The predicted activation score can be used to indicate the probability of the expert module being accessed.

[0268] The second model may be trained using at least one second request and multiple first scores corresponding to each of the at least one second request. The at least one second request may serve as a sample. For a description of the second request, please refer to the description of the request above and will not be repeated here. Furthermore, this embodiment of the present application does not limit the method for obtaining the second request. The multiple first scores corresponding to each of the at least one second request may serve as labels. The multiple first scores corresponding to the second request may include first scores corresponding to each of multiple expert modules. The first score may be used to indicate the activation score of the expert module. This activation score is the actual activation score. The second model may be obtained through offline training. The second model may be trained using at least one of an end-side device or a non-end-side device. If the second model is trained using an end-side device, the end-side device used to train the second model may be the same as or different from the end-side device used to apply the second model, and this is not limited in this embodiment of the present application. The model structure of the second model may be configured according to actual business needs and is not limited here. For example, if the second model can meet the prediction accuracy requirements, the model parameters of the second model should be kept as small as possible to reduce resource consumption of the end-side model.

[0269] Regarding how to train the second model using at least one second request and a plurality of first scores corresponding to each of the at least one second request, this can be achieved in the following manner.

[0270] As an implementation method, the second request can be input into the second model to obtain a second score corresponding to each of the multiple expert modules. Based on a loss function, a loss function value is obtained according to the first and second scores corresponding to each of the multiple expert modules. The model parameters of the second model are adjusted according to the loss function value to obtain a trained second model. The loss function can be configured according to actual business needs and is not limited here. For example, the loss function may include a cross-entropy loss function.

[0271] Regarding how to obtain the loss function value based on the loss function and the first scores and the second scores corresponding to each of the plurality of expert modules, it can be achieved in the following manner.

[0272] As an implementation method, the first score and the second score corresponding to the expert module can be input into the loss function to obtain the loss function value corresponding to the expert module. The loss function value can be obtained based on the loss function values corresponding to the multiple expert modules.

[0273] The following describes how to determine a plurality of first scores corresponding to at least one second request.

[0274] The first model may include at least one cascaded first layer. The first layer may include a gating module and multiple expert modules. The inference phase of the first model may include a pre-population phase and a decoding phase. For any second request among the at least one second request, in the pre-population phase, the second model may be used to obtain a second word (i.e., a first output word) corresponding to the second request based on at least one third word (i.e., an input word) corresponding to the second request. In the decoding phase, the second model and an autoregressive approach may be used to obtain first words (i.e., output words) one by one based on the second word and at least one third word corresponding to the second request. The second word may also be referred to as a first word.

[0275] When the pre-filling phase is performed using the second model, a second score corresponding to each of the plurality of expert modules corresponding to the at least one first layer can be determined for the third word. When the decoding phase is performed using the second model, a second score corresponding to each of the plurality of expert modules corresponding to the at least one first layer can be determined for the first word.

[0276] Based on this, the multiple expert modules corresponding to the second request may include multiple expert modules corresponding to at least one third word. Optionally, the multiple expert modules corresponding to the second request may include multiple expert modules corresponding to at least one third word and multiple expert modules corresponding to at least one first word. Optionally, the multiple expert modules corresponding to the second request may include multiple expert modules corresponding to at least one first word.

[0277] For any third word in the at least one third word-gram, a first score corresponding to any expert module in the multiple expert modules corresponding to the third word-gram can be obtained. For any first word in the at least one first word-gram, a first score corresponding to any third expert module in the multiple expert modules corresponding to the first word-gram can be obtained. Thus, the first scores corresponding to each of the multiple expert modules corresponding to the second request can be determined.

[0278] In order to facilitate understanding of how to determine the first scores corresponding to the multiple expert modules corresponding to the third word or the first scores corresponding to the multiple expert modules corresponding to the first word, the following takes a third word as an example, combined with Figure 8D Provide explanation.

[0279] Figure 8D A schematic diagram of determining first scores corresponding to multiple expert modules corresponding to a third word provided in an embodiment of the present application.

[0280] like Figure 8D As shown, the multiple expert modules corresponding to the third word may include multiple expert modules corresponding to at least one first layer respectively.

[0281] For any first layer of at least the first layer included in the first model, a gating module included in the first layer can be used to obtain, based on the third word-gram, first scores corresponding to each of the plurality of expert modules included in the first layer. Thus, first scores corresponding to each of the plurality of expert modules corresponding to the third word-gram can be determined.

[0282] In order to demonstrate the actual effect of the parallel execution of the pre-filling phase process and the pre-loading process proposed in the embodiment of the present application, an experimental verification is carried out below. When the first model can be DeepSeek MoE-32G and the first storage unit is configured to store approximately 10% (i.e. 3.1G) of the model parameters of the first model, the average cache hit rate can be increased from 0.08 to 0.25. The end-to-end decoding overhead is reduced by an average of approximately 25%. When the first model is Mixtral MoE-175G and 4 processor cores are used for calculation, if the processor memory is 19G, the average time consumed to decode a word is reduced from 26.6s to 16.2s. If the processor memory is 10G, the average time consumed to decode a word is reduced from 29.1s to 19.12s.

[0283] The above illustrates the inventive concept proposed in an embodiment of the present application of executing the pre-filling phase process executed by the first thread and the pre-loading process executed by the second thread in parallel to improve the cache hit rate, reduce I / O latency, and thereby improve the reasoning efficiency of the first model.

[0284] We further discovered that when the first thread uses the first model to execute the decoding phase, a cache miss may occur, meaning that the first storage unit does not contain the second expert module corresponding to the first expert module. Consequently, the first thread cannot retrieve the second expert module corresponding to the first expert module from the first storage unit. In this case, the first thread must first load the expert module corresponding to the first expert module from the second storage unit to the first storage unit, and then retrieve the expert module corresponding to the first expert module from the first storage unit. This results in I / O latency between the second and first storage units.

[0285] To further reduce I / O latency, it was discovered that if the first storage unit has a second expert module corresponding to the fourth expert module, and the fourth expert module can be used to replace the first expert module, I / O latency can be reduced. Therefore, the present embodiment proposes a processing method based on a cache priority strategy.

[0286] Next, we need to consider under what conditions the fourth expert module must meet in order to be able to replace the first expert module. It was found that the conditions that the fourth expert module needs to meet to replace the first expert module can be determined from the perspective of the expert module weight. For example, the first expert module can have a first weight. The fourth expert module can have a second weight. The first weight can be used to indicate the activation probability of the first expert module. The second weight can be used to indicate the activation probability of the fourth expert module. In order to enable the fourth expert module to replace the first expert module, the absolute value of the difference between the first weight and the second weight can be made as small as possible. Therefore, the embodiment of the present application proposes that the absolute value of the difference between the first weight and the second weight can be less than or equal to the first threshold. The first threshold can be configured according to actual business needs and is not limited here. For example, the first threshold can be 0.3.

[0287] Based on this, when the first thread determines the first expert module corresponding to the first request from the multiple expert modules included in the first model based on the first word corresponding to the first request, if it is determined based on the first weight and the first threshold that the multiple expert modules included in the first model include a fourth expert module corresponding to the first expert module, then if the first storage unit does not include a second expert module corresponding to the first expert module, the first expert module can be replaced with the fourth expert module. If the first storage unit includes a second expert module corresponding to the fourth expert module, the first expert module can be replaced with the fourth expert module.

[0288] The following describes how the second expert module corresponding to the fourth expert module is stored in the first storage unit.

[0289] While processing the first request, the first and second threads may also process a third request, which may be different from the first request. Therefore, the fourth expert module and the first expert module may correspond to the same request. For example, the fourth expert module and the first expert module may correspond to the first request. Therefore, the fourth expert module may be the second expert module stored in the first storage unit and corresponding to the first request. Furthermore, the fourth expert module and the first expert module may correspond to different requests. For example, the fourth expert module may correspond to the third request, while the first expert module may correspond to the first request. Therefore, the fourth expert module may be the second expert module stored in the first storage unit and corresponding to the third request.

[0290] For ease of understanding, the following Figure 9A and Figure 9B It indicates that if the first thread does not have a second expert module corresponding to the first expert module in the first storage unit and the first storage unit has a second expert module corresponding to the fourth expert module, the first expert module is replaced by the fourth expert module. Figure 9A This is for the case where the fourth expert module and the first expert module correspond to the same request. Figure 9B This is the case where the fourth expert module and the first expert module correspond to different requests.

[0291] Figure 9A A schematic diagram of replacing a first expert module with a fourth expert module is provided in an embodiment of the present application.

[0292] like Figure 9A As shown, the first model may include a first layer 900. The first layer 900 may include a gating module 910 and four expert modules. The four expert modules may include an expert module 921, an expert module 922, an expert module 923, and an expert module 924.

[0293] The first thread can use gating module 910 to determine, based on the first word corresponding to the first request, expert modules 921 and 923 from among expert modules 921, 922, 923, and 924 as the first expert module. The absolute value of the difference between the third weight corresponding to expert module 922 and the third weight corresponding to expert module 923 is less than or equal to the first threshold. Therefore, expert module 922 can be designated as the fourth expert module corresponding to expert module 923 (i.e., the first expert module). The third weight corresponding to expert module 922 can be referred to as the second weight. The third weights corresponding to expert modules 921 and 923 can each be referred to as the first weight.

[0294] The first thread determines that the first storage unit has a second expert module corresponding to the expert module 921 (ie, the first expert module), and thus, the second expert module corresponding to the first expert module can be obtained from the first storage unit.

[0295] The first thread determines that the first storage unit does not contain a second expert module corresponding to expert module 923 (i.e., the first expert module). However, the first storage unit does contain a second expert module corresponding to the fourth expert module (i.e., expert module 922). Expert module 922 is the fourth expert module corresponding to expert module 923 (i.e., the first expert module). Therefore, the fourth expert module (i.e., expert module 922) can replace the first expert module (i.e., expert module 923). The first thread can then retrieve the second expert module corresponding to the fourth expert module from the first storage unit, improving cache hit rates. Furthermore, since there is no need to first load the expert module corresponding to the first expert module from the second storage unit to the first storage unit and then retrieve the expert module corresponding to the first expert module from the first storage unit, I / O latency between the second and first storage units is reduced.

[0296] Based on this, the first thread may utilize the first expert module (ie, expert module 921 ) and the fourth expert module (ie, expert module 922 ) to process the first word.

[0297] Figure 9B Another schematic diagram of replacing the first expert module with the fourth expert module is provided in an embodiment of the present application.

[0298] like Figure 9B As shown, Figure 9A The difference is that the fourth expert module (i.e., expert module 922) corresponding to the first expert module (i.e., expert module 923) is the second expert module corresponding to the third request stored in the first storage unit. The third request is different from the first request. The other parts are the same as Figure 9A Similar, no further description is given here.

[0299] To demonstrate the effectiveness of the cache-first processing approach proposed in this embodiment, an experimental verification was conducted. If the first model is a DeepSeek MoE-32G and the first storage unit is configured to store approximately 10% (i.e., 3.1G) of the first model's parameters, and the pre-population and pre-loading processes are executed in parallel, the average cache hit rate can be increased from 0.25 to 0.60. This reduces the end-to-end decoding overhead by an average of approximately 41.6%.

[0300] It was further discovered that, in the event of a cache miss, it is also possible to consider using a third thread to retrieve the expert module corresponding to the first expert module from the second storage unit, and using the expert module corresponding to the first expert module to process the first word. The third speed can be greater than the fourth speed. The third speed can be used to indicate the speed at which the third thread accesses the second storage unit. The fourth speed can be used to indicate the speed at which the third thread accesses the first storage unit. For example, the third thread can be a processor thread. The first thread can be an accelerator thread. This is because it was found that the average time it takes to process the first word using the third thread is less than the average time it takes for the first thread to load the expert module corresponding to the first expert module from the second storage unit to the first storage unit, retrieve the expert module corresponding to the first expert module from the first storage unit, and use the expert module corresponding to the first expert module to process the first word.

[0301] In order to understand how to use the third thread to process the first word, the following Figure 10 Provide explanation.

[0302] Figure 10 A schematic diagram of processing a first word using a third thread according to an embodiment of the present application.

[0303] like Figure 10 As shown, the first model may include a first layer 1000. The first layer 1000 may include a gating module 1010 and four expert modules. The four expert modules may include an expert module 1021, an expert module 1022, an expert module 1023, and an expert module 1024.

[0304] The first thread may utilize the gating module 1010 to determine, based on the first word corresponding to the first request, expert module 1021 and expert module 1023 as the first expert module from among expert modules 1021, 1022, 1023, and 1024.

[0305] The first thread determines that the first storage unit has a second expert module corresponding to the expert module 1021 (ie, the first expert module), and thus, the second expert module corresponding to the first expert module can be obtained from the first storage unit.

[0306] The first thread determines that the first storage unit does not contain a second expert module corresponding to the expert module 1023 (ie, the first expert module). The third thread can obtain the expert module corresponding to the expert module 1023 (ie, the first expert module) from the second storage unit.

[0307] Based on this, the first thread can use the second expert module corresponding to the first expert module (ie, expert module 1021) to process the first word-gram. The third thread can use the expert module corresponding to the first expert module (ie, expert module 1023) to process the first word-gram.

[0308] Because the memory capacity of the first storage unit is limited, during the execution of the data processing method based on the first thread and the second thread, the first storage unit may still meet the first condition, thereby affecting the cache hit rate and resource utilization. The first condition can be used to trigger cache eviction. As an implementation method, the first condition can be used to indicate that the first storage unit is full.

[0309] In order to improve the cache hit rate and resource utilization, an embodiment of the present application proposes that if the first storage unit meets the first condition, the fourth thread can be used to execute the cache elimination process based on the cache elimination strategy. In this way, the second expert module among the multiple second expert modules stored in the first storage unit can be eliminated, thereby improving the cache hit rate and resource utilization.

[0310] For example, the fourth thread may eliminate the second expert module from the plurality of second expert modules stored in the first storage unit based on the cache elimination policy when the first storage unit meets the first condition. The fourth thread may be executed in parallel with the first thread and the second thread.

[0311] Based on the above, the embodiment of the present application builds a first model based on MoE, completes the scheduling of model parameters from the second storage unit to the first storage unit, and can configure a predetermined number of expert modules to be stored in the first storage unit. Furthermore, pipeline parallel processing is implemented for calculating and storing the model parameters of the expert modules in the first storage unit. A request-based expert preference strategy is utilized to improve the cache hit rate of the first storage unit. Furthermore, a cache priority strategy can be utilized to further improve the cache hit rate.

[0312] The above describes the inventive concept of the embodiment of the present application. The data processing method provided by the embodiment of the present application will be described in detail below with reference to the accompanying drawings.

[0313] First, combine Figure 11 and Figure 12 , the end-side device to which the data processing method is applicable is explained. Figure 11 and Figure 12 Taking the end-side device shown in the figure as an example, in conjunction with the accompanying drawings (for example, Figure 13 、 Figure 14A 、 Figure 14B 、 Figure 14C 、 Figure 14D 、 Figure 14E 、 Figure 15 、 Figure 16 、 Figure 17 、 Figure 18 ) and application scenarios, the data processing method provided in the embodiment of the present application is specifically described.

[0314] It should be noted that the embodiments of the present application can be implemented independently or in conjunction with each other, and the same or similar concepts or processes will not be described in detail in some embodiments.

[0315] Figure 11 A schematic diagram of the structure of the terminal side device provided in an embodiment of the present application.

[0316] The end-side device 1100 may include a processor 1110, an external memory interface 1120, an internal memory 1121, a universal serial bus (USB) interface 1130, a charging management module 1140, a power management module 1141, a battery 1142, an antenna 1, an antenna 2, a mobile communication module 1150, a wireless communication module 1160, an audio module 1170, a speaker 1170A, a receiver 1170B, a microphone 1170C, an earphone interface 1170D, a sensor module 1180, a button 1190, a motor 1191, an indicator 1192, a camera 1193, a display screen 1194, and a subscriber identification module (SIM) card interface 1195, etc. Among them, the sensor module 1180 can include a pressure sensor 1180A, a gyroscope sensor 1180B, an air pressure sensor 1180C, a magnetic sensor 1180D, an acceleration sensor 1180E, a distance sensor 1180F, a proximity light sensor 1180G, a fingerprint sensor 1180H, a temperature sensor 1180J, a touch sensor 1180K, an ambient light sensor 1180L, a bone conduction sensor 1180M, etc.

[0317] It should be understood that the structure illustrated in the embodiments of this application does not constitute a specific limitation on the end-side device 1100. In other embodiments of this application, the end-side device 1100 may include more or fewer components than shown, or may combine or separate certain components, or arrange the components differently. The illustrated components may be implemented in hardware, software, or a combination of software and hardware.

[0318] Processor 1110 may include one or more processing units. For example, processor 1110 may include an application processor (AP). The application processor may include a central processing unit, a hardware accelerator, a memory controller, a video codec, and a baseband processor (BP). The hardware accelerator may include at least one of the following: a graphics processor, a neural network processor, a tensor processor, an image signal processor (ISP), or a digital signal processor (DSP). Among them, different processing units may be independent devices or integrated into one or more processors. For example, the application processor may be a system on chip. In the embodiment of the present application, the hardware accelerator may be referred to as an accelerator.

[0319] Processor 1110 may also include a memory for storing instructions and data. In some embodiments, the memory in processor 1110 is a cache memory. This memory can store instructions or data that have just been used or are being recycled by processor 1110. If processor 1110 needs to use the instruction or data again, it can directly access the memory. This avoids duplicate accesses, reduces processor 1110 latency, and thus improves system efficiency.

[0320] In some instances, the memory of the embodiment of the present application may store instructions and data for implementing the data processing method described in the embodiment of the present application, and when the processor is running, the data processing method described in the embodiment of the present application may be implemented.

[0321] The internal memory 1121 can be used to store computer-executable program code, including instructions. The internal memory 1121 may include a program storage area and a data storage area. The program storage area may store, for example, an operating system or applications required for at least one function (e.g., AI applications). The data storage area may store, for example, data created during use of the end-side device 700. Furthermore, the internal memory 1121 may include high-speed random access memory and non-volatile memory, such as at least one disk storage device, flash memory device, or universal flash storage (UFS). The processor 1110 executes various functional applications and data processing of the end-side device 700 by executing instructions stored in the internal memory 1121 and / or instructions stored in a memory provided within the processor 1110. In embodiments of the present application, the internal memory 1121 may be used to implement the instructions and data for the data processing methods described in embodiments of the present application.

[0322] The software system of the terminal device 1100 can adopt a layered architecture, an event-driven architecture, a micro-kernel architecture, a microservice architecture, or a cloud architecture. In this embodiment of the application, the Android system with a layered architecture is used as an example to illustrate the software structure of the terminal device 1100.

[0323] Figure 12 This is a block diagram of the software structure of the terminal device provided in the embodiment of the present application.

[0324] A layered architecture divides software into multiple layers, each with distinct roles and responsibilities. Layers can communicate with each other through software interfaces. In some embodiments, the Android system is divided into four layers: from top to bottom: the application layer, the application framework layer (i.e., framework layer), the hardware abstraction layer (HAL), the driver layer (i.e., driver layer), and the hardware layer.

[0325] The application layer may include an application package. In an embodiment of the present application, the application package may include an AI application. As an implementation method, when the end-side device runs an AI application, the AI application may execute the data processing method of the embodiment of the present application.

[0326] The application framework layer provides an application programming interface (API) and programming framework for applications in the application layer. The application framework layer may include predefined functions. In embodiments of the present application, the application framework layer may include a neural network API (NNAPI). The NN API is a system-level API. The NN API may be a framework provided by Android for running end-side models on end-side devices. The NN API may support at least one of loading end-side models obtained using a machine learning framework, compiling end-side models into a hardware executable format, or determining the hardware used to perform inference tasks. The machine learning framework may be a framework that supports running machine learning models on end-side devices. For example, the hardware may include an application processor. The application processor may include a central processing unit (CPU) and a hardware accelerator. The hardware accelerator may include at least one of the following: a graphics processor, a neural network processor, or a tensor processor.

[0327] The hardware abstraction layer is an interface layer located between the application framework layer and the driver layer. It can be an encapsulation of the hardware driver and provide a unified interface for calling upper-layer applications. In an embodiment of the present application, the hardware abstraction layer may include an AI hardware abstraction layer. The AI hardware abstraction layer can be used to transmit AI application requests to the hardware. The AI hardware abstraction layer can also be used to compile the end-side model into a hardware executable format or assign inference tasks to the hardware. The AI hardware abstraction layer may include an AI hardware abstraction layer interface. The AI hardware abstraction layer interface can be used to transmit AI application requests to the hardware of the driver layer. The AI hardware abstraction layer interface may include a first interface, a second interface, and a third interface. The first interface can be used to provide basic information about the hardware. For example, the information may include at least one of the following: operation type or performance indicator. The second interface can be used to load and compile the end-side model. The third interface can be used to perform inference tasks.

[0328] The neural network application programming interface can call the AI hardware abstraction layer interface to compile the end-side model into a hardware executable format. The neural network application programming interface can determine the hardware used to perform the inference task based on the basic information of the hardware provided by the AI hardware abstraction layer. For example, the neural network application programming interface can determine that the hardware used to perform the inference task is the central processing unit when there is no hardware accelerator on the end-side device. Optionally, the neural network application program interface can determine that the hardware used to perform the inference task is the hardware accelerator when there is a hardware accelerator on the end-side device. Optionally, the neural network application program interface can determine that the hardware used to perform the inference task is the central processing unit when there is a hardware accelerator and a predetermined operation. The predetermined operation may refer to an operation that needs to be performed by the central processing unit. The AI hardware abstraction layer can transmit inference tasks to the hardware and manage inference tasks.

[0329] The driver layer can be a layer between hardware and software that can be used to drive the hardware. In an embodiment of the present application, the driver layer may include the Android kernel. The Android kernel can be used to manage hardware and provide system services to upper layers (e.g., the hardware abstraction layer or the application framework layer). System services may include at least one of the following: process scheduling, memory management, or a file system. Hardware drivers can be written for central processing units and hardware accelerators to support the end-side model. The hardware driver is integrated into the Android kernel.

[0330] The application layer can indirectly utilize services provided by the Android kernel through the application framework layer. The AI hardware abstraction layer can access hardware through the Android kernel. The application framework layer can access the Android kernel through system calls.

[0331] The hardware layer can include a central processing unit, a graphics processing unit, a neural network processor, a tensor processor, etc., to perform inference tasks.

[0332] It should be noted that the embodiments of the present application are only illustrated using the Android system as an example. In other operating systems, if the functions implemented by each functional module are similar to those in the embodiments of the present application, the data processing method described in the embodiments of the present application can also be implemented.

[0333] The following combination Figure 13 、 Figure 14A 、 Figure 14B 、 Figure 14C 、 Figure 14D 、 Figure 14E 、 Figure 15 、 Figure 16 and Figure 17 The data processing method based on the first thread, the second thread, the third thread, the fourth thread, the first storage unit, the second storage unit, the first model and the second model provided in the embodiment of the present application is described. Figure 13 、 Figure 15 、 Figure 16 and Figure 17 The difference lies in that different implementation methods are provided for the case where the first storage unit does not have a second expert module corresponding to the first expert module. Figures 14A-14E Shown for Figure 13 Specific examples.

[0334] For ease of understanding, the first model, the first storage unit, the second storage unit and the third thread are supplementarily explained below in combination with the above, as well as the first set and the second set involved below.

[0335] The inference phase of the first model may include a first phase and a second phase. The first phase may also be referred to as a pre-population phase. The second phase may also be referred to as a decoding phase. As described above, the pre-population phase may refer to the phase in which the first round of processing is performed. The decoding phase may refer to the phase in which rounds 2 through P of processing are performed.

[0336] The first model may include multiple cascaded first layers. The first layer may include a mixed expert layer. The mixed expert layer may include a gating module and multiple expert modules. The gating module may be used to determine an activated expert module from the multiple expert modules. The activated expert module may be used to perform nonlinear transformation and feature extraction on the input. For example, the first model may include R cascaded first layers, namely, the 1st first layer, ..., the rth first layer, ..., the Rth first layer. The rth first layer may include the rth mixed expert layer. The rth mixed expert layer may include S expert modules, namely, the r_1th expert module, ..., the r_sth expert module, ..., the r_Sth expert module. R may be an integer greater than or equal to 1. r may be an integer greater than or equal to 1 and less than or equal to R. S may be an integer greater than 1. s may be an integer greater than or equal to 1 and less than or equal to S.

[0337] For example, the first model could be Figure 6 The model shown in Figure 1. The first layer of the rth layer can refer to Figure 6 The first layer 600_r. The rth mixed expert layer can be Figure 6 The r_s expert module may refer to Figure 6 Expert module 612_r_s.

[0338] It should be noted that the first layer currently being used can be referred to as the current layer. For example, the rth first layer can be used to indicate the current layer. In the first phase (i.e., the pre-population phase), the activated expert module can be referred to as the second expert module. In the second phase (i.e., the decoding phase), the activated expert module can be referred to as the first expert module. The first expert module can have a first identifier. The expert module can have a second identifier. The second expert module can have a third identifier. The first identifier can be used to indicate the first expert module. The second identifier can be used to indicate the expert module. The third identifier can be used to indicate the second expert module.

[0339] The second model can be used to determine the third expert module that needs to be preloaded into the first storage unit. The second model can be trained using the method described above. During the preloading phase, the expert module predicted to be active can be referred to as the third expert module. The third expert module can have a fourth identifier. The fourth identifier can be used to indicate the third expert module.

[0340] The first storage unit can be used to store a second expert module corresponding to the third expert module. The second expert module corresponding to the third expert module can be understood as the second expert module being the third expert module. The first storage unit storing the second expert module can be understood as the first storage unit storing model parameters of the second expert module. The model parameters of the second expert module can be associated with a third identifier. The number of second expert modules stored in the first storage unit can be configured based on actual business needs and is not limited herein.

[0341] The second storage unit can be used to store the expert module included in the first model. The second storage unit storing the expert module can be understood as the second storage unit storing the model parameters of the expert module.

[0342] The second set may include multiple second expert modules.The second set may be obtained by the first thread loading the expert module corresponding to the second expert module from the second storage unit to the first storage unit.

[0343] The first set may include the second expert module corresponding to the third expert module.The first set may be obtained by the second thread retaining the second expert module corresponding to the third expert module in the second set.

[0344] The third thread can be used for the first thread to obtain the expert module corresponding to the first expert module from the second storage unit when it is determined that the second expert module corresponding to the first expert module does not exist in the first storage unit, or when it is determined that both the second expert module corresponding to the first expert module and the second expert module corresponding to the fourth expert module do not exist in the first storage unit, and to execute the second stage process using the expert module corresponding to the first expert module.

[0345] The data processing method based on the first thread, the second thread, the third thread, the fourth thread, the first storage unit, the second storage unit, the first model and the second model provided in an embodiment of the present application is described below with reference to the accompanying drawings.

[0346] Figure 13 This is a flow chart of a data processing method provided in an embodiment of the present application. The method can be applied to a first thread and a second thread.

[0347] like Figure 13 As shown, the method includes S1301-S1316. The first thread may use the first model to execute the first phase process (i.e., the pre-filling phase process) and the second phase process (i.e., the decoding phase process). The second thread may use the second model to execute the pre-loading process. The first thread's execution of the first phase process using the first model may be performed in parallel with the second thread's execution of the pre-loading process using the second model.

[0348] The pre-filling phase process (i.e., the first phase process) may include S1301 - S1303 , the decoding phase process (i.e., the second phase process) may include S1306 - S1316 , and the pre-loading process may include S1304 - S1305 .

[0349] According to the above, the pre-filling stage can refer to the stage of performing the first round of processing. The decoding stage can refer to the stage of performing the second to P rounds of processing. Figure 13 The decoding phase process may refer to the decoding phase process corresponding to the p-th round. p may be an integer greater than or equal to 2 and less than or equal to P. P may be an integer greater than 1. The first thread may repeatedly execute S1306-S1315 for rounds 2 to P based on an autoregressive method until an inference result (i.e., an output result) is obtained. The input corresponding to the p-th round may include at least one third word element corresponding to the first request, and the first word element corresponding to each of rounds 1 to p-1. The output corresponding to the p-th round may include the first word element corresponding to the p-th round. The third word element may be used to indicate an input word element. The first word element may be used to indicate an output word element. The first word element corresponding to the 1st round may also be referred to as a second word element.

[0350] The first thread can execute the first phase process (i.e., the pre-filling phase process) using the first model in parallel with the second thread executing the pre-loading process using the second model. The following describes the pre-filling phase process, the pre-loading phase process, and the decoding phase process.

[0351] How the first thread uses the first model to execute the first phase process (ie, the pre-filling phase process) can be implemented through the following S1301 - S1303 .

[0352] In S1301 , the first thread determines, based on at least one third word corresponding to the first request, at least one second expert module corresponding to each of the at least one third word from a plurality of expert modules included in the first model, to obtain a second set.

[0353] According to an embodiment of the present application, the first request may be used to indicate a task to be processed by the first model. For an explanation of the first request, please refer to the above explanation of the request, which will not be repeated here.

[0354] The third word-gram may be used to indicate an input word-gram. The at least one third word-gram may be obtained by the first thread tokenizing the first request.

[0355] For example, the first request may be “in which city is FD University located”. The at least one third word-gram corresponding to the first request may include a third word-gram [“FD”], a third word-gram [“university”], a third word-gram [“located in”], a third word-gram [“which city”], and a third word-gram [“city”].

[0356] According to an embodiment of the present application, as described above, the multiple expert modules included in the first model may include multiple expert modules corresponding to at least one first layer, that is, the first model may include at least one first layer. Each first layer may include multiple expert modules. The at least one second expert module corresponding to the third word may include at least one second expert module corresponding to at least one first layer, that is, each first layer may have at least one second expert module corresponding to the first layer.

[0357] According to an embodiment of the present application, the second set may include at least one second expert module corresponding to each of the at least one third word-grams.

[0358] How to determine at least one second expert module corresponding to at least one third word-gram can be achieved in the following manner.

[0359] As an implementation, the first thread may determine, based on the at least one third word corresponding to the first request, at least one second expert module corresponding to each of the at least one third word-grams from a plurality of expert modules included in the first model and corresponding to each of the at least one first layer. The steps of determining the at least one second expert module corresponding to each of the at least one third word-grams may be performed in parallel.

[0360] How the first thread determines at least one second expert module corresponding to at least one third word from a plurality of expert modules corresponding to at least one first layer can be implemented in the following manner.

[0361] As an implementation method, the first layer may include a gating module. Thus, the first thread may utilize the gating module corresponding to each of the at least one first layer to determine, based on the at least one third word corresponding to the first request, at least one second expert module corresponding to each of the at least one third word from a plurality of expert modules corresponding to each of the at least one first layer. For example, for any third word in the at least one third word, any first layer in the at least one first layer, with the first layer as the current layer, the first thread may utilize the gating module corresponding to the current layer to determine, based on the third word, at least one second expert module corresponding to the third word and the current layer from a plurality of expert modules corresponding to the current layer.

[0362] How to determine at least one second expert module corresponding to the third word-gram and the current layer can be achieved in the following manner.

[0363] As an implementation, the first thread may utilize the gating module corresponding to the current layer to obtain, based on the third word-gram, the fourth weights corresponding to each of the multiple expert modules corresponding to the current layer. Based on the fourth weights corresponding to each of the multiple expert modules corresponding to the current layer, the first thread may determine, from the multiple expert modules corresponding to the current layer, at least one second expert module corresponding to the third word-gram.

[0364] How the first thread determines at least one second expert module corresponding to the third word from the multiple expert modules corresponding to the current layer according to the fourth weights corresponding to the multiple expert modules corresponding to the current layer can be achieved in the following manner.

[0365] The first method is based on the Top-T method. For example, the first thread can determine T fourth weights from the fourth weights corresponding to the multiple expert modules corresponding to the current layer. The expert modules corresponding to the T fourth weights are determined to be the second expert modules corresponding to the third word and the current layer. When the larger the fourth weight is, the greater the probability of the expert module being activated is, the T fourth weights can be the maximum T fourth weights. When the smaller the fourth weight is, the greater the probability of the expert module being activated is, the T fourth weights can be the minimum T fourth weights. T can be an integer greater than or equal to 1 and less than or equal to S. T can be configured according to actual business needs and is not limited here. For example, S=16. T=2.

[0366] The second method is based on the threshold method. For example, when the fourth weight is larger, the probability of the expert module being activated is larger. The first thread can determine a fourth weight greater than or equal to the sixth threshold from the fourth weights corresponding to the multiple expert modules corresponding to the current layer. The expert module corresponding to the fourth weight greater than or equal to the seventh threshold is determined as the second expert module corresponding to the third word and the current layer. When the fourth weight is smaller, the probability of the expert module being activated is larger. The first thread can determine a fourth weight less than or equal to the eighth threshold from the fourth weights corresponding to the multiple expert modules corresponding to the current layer. The expert module corresponding to the fourth weight less than or equal to the eighth threshold is determined as the second expert module corresponding to the third word and the current layer. The seventh threshold can be greater than the eighth threshold. The seventh and eighth thresholds can be configured according to actual business needs and are not limited here.

[0367] As described above, the first model may include R cascaded first layers, namely the 1st first layer, ..., the rth first layer, ..., the Rth first layer. The rth first layer may include the rth mixed expert layer. The rth mixed expert layer may include the rth gating module and S rth expert modules. The S rth expert modules may include the r_1th expert module, ..., the r_sth expert module, ..., the r_sth expert module. For the description of the first model, please refer to the corresponding part above and will not be repeated here. The at least one second expert module corresponding to the third word may include at least one second expert module corresponding to each of the R first layers. The current layer mentioned above may refer to the rth first layer. The gating module corresponding to the current layer may refer to the rth gating module. The expert module corresponding to the current layer may refer to the r_sth expert module.

[0368] When r=1, the first thread can input the first first feature into the first gating module to obtain the fourth weight corresponding to each of the S first expert modules, that is, the fourth weight corresponding to each of the 1_1 expert module, ..., the 1_s expert module, ..., and the 1_S expert module. The fourth weight corresponding to the 1_s expert module can be used to indicate the activation probability of the 1_s expert module. The first thread can determine at least one second expert module corresponding to the third word element from the S first expert modules based on the fourth weights corresponding to each of the S first expert modules. The first first feature can be obtained by the first thread based on the third word element using the first first layer.

[0369] When 1<r≤R, the first thread can input the r first features into the rth gating module to obtain the fourth weights corresponding to the S rth expert modules, that is, the fourth weights corresponding to the r_1th expert module, ..., the r_sth expert module, ..., and the r_Sth expert module. The fourth weight corresponding to the r_sth expert module can be used to indicate the activation probability of the r_sth expert module. The first thread can determine at least one second expert module corresponding to the third word element from the S rth expert modules based on the fourth weights corresponding to the S rth expert modules. The rth first feature can be obtained by the first thread based on the third word element using the first first layer.

[0370] In S1302 , the first thread loads at least one second expert module corresponding to each of the at least one third word-grams from the second storage unit to the first storage unit to obtain a second set.

[0371] According to an embodiment of the present application, the second set may include at least one second expert module corresponding to each of the at least one third word-grams. The at least one second expert module corresponding to the third word-grams may include at least one second expert module corresponding to each of the at least one first layer.

[0372] For any one of the at least one third word-gram and any one of the at least one first layer, with the first layer being the current layer, the first thread may load at least one second expert module corresponding to the third word-gram and the current layer from the second storage unit to the first storage unit. The steps of loading the at least one second expert module corresponding to each of the at least one third word-grams may be performed in parallel.

[0373] At 1303 , the first thread uses the second set and obtains a second word-gram corresponding to the first request according to at least one third word-gram.

[0374] According to an embodiment of the present application, the first thread can use at least one second expert module included in the first model and corresponding to at least one third word to execute the first stage process to obtain the second word corresponding to the first request. The second word can be used to indicate the first output word corresponding to the first request. For example, the second word can be Figure 3 The first output word ["FD"].

[0375] How the second thread uses the second model to execute the preloading process can be implemented through the following S1304-S1305.

[0376] In S1304 , the second thread uses the second model to obtain a plurality of third expert modules corresponding to the first request according to the first request.

[0377] According to an embodiment of the present application, the second thread can input the first request into the second model to obtain multiple third expert modules corresponding to the first request. The third expert modules may be expert modules that need to be preloaded into the first storage unit so that the first thread can retrieve the first expert modules from the first storage unit when executing the second phase process. The second model is described in the corresponding section above and will not be repeated here.

[0378] Regarding how the second thread utilizes the second model to obtain multiple third expert modules corresponding to the first request according to the first request, this can be achieved in the following manner.

[0379] As an implementation, the second thread may input the first request into the second model to obtain multiple third expert modules corresponding to the first request. For example, the first thread may input the first request into the second model to obtain second scores corresponding to each of the multiple expert modules. Based on the second scores corresponding to each of the multiple expert modules, multiple third expert modules corresponding to the first request are obtained. The second scores can be used to indicate the predicted activation score of the expert module. The predicted activation score can be used to indicate the probability of activation.

[0380] How the second thread obtains multiple third expert modules corresponding to the first request based on the second scores corresponding to the multiple expert modules can be achieved in the following manner.

[0381] As an implementation method, the second thread can determine U second scores from the second scores corresponding to each of the multiple expert modules. The expert modules corresponding to the U second scores are determined as third expert modules. When the larger the second score is, the higher the probability of the expert module being accessed is, the U second scores can be the maximum U second scores. When the smaller the second score is, the greater the probability of the expert module being accessed is, the U second scores can be the minimum U second scores. U can be an integer greater than 1 and less than R*S. U can be determined based on the storage capacity of the first storage unit and is not limited here. For the description of R and S, please refer to the corresponding parts above and will not be repeated here.

[0382] In S1305 , the second thread retains the second expert modules corresponding to the plurality of third expert modules in the second set to obtain a first set.

[0383] According to an embodiment of the present application, the first set may include second expert modules in the second set corresponding to each of the plurality of third expert modules.

[0384] For any third expert module among the multiple third expert modules corresponding to the first request, in the process of the first thread loading at least one second expert module corresponding to at least one third word from the second storage unit to the first storage unit, if the second thread determines that there is a second expert module corresponding to the third expert module in the second set stored in the first storage unit, that is, if it is determined that the third expert module exists in the second set stored in the first storage unit, then the second expert module corresponding to the third expert module in the first storage unit can be retained.

[0385] How the second thread retains the second expert module corresponding to the third expert module in the second set can be achieved in the following manner.

[0386] As an implementation manner, the second thread may store the second expert module in the second set corresponding to the third expert module in a queue.

[0387] Furthermore, if the second thread determines that the second set stored in the first storage unit does not contain a second expert module corresponding to the third expert module, the second thread may load the expert module corresponding to the third expert module in the second storage unit from the second storage unit to the first storage unit. For example, the second thread may load the expert module corresponding to the third expert module in the second storage unit from the second storage unit to the queue of the first storage unit.

[0388] Therefore, when the first thread executes the second stage process, the second expert module stored in the first storage unit corresponds to the third expert module, that is, the first storage unit stores the third expert module.

[0389] How the first thread uses the first model to execute the second phase process can be implemented through the following S1306-S1316.

[0390] At S1306 , the first thread uses the gating module included in the first model to determine, based on the first word-gram corresponding to the first request, at least one first expert module corresponding to the first word-gram from a plurality of expert modules included in the first model.

[0391] According to an embodiment of the present application, the first word-element may be used to indicate an output word-element. For example, the first word-element may be the p-th output word-element described above. p may be an integer greater than or equal to 2 and less than or equal to P.

[0392] As an implementation, the first thread may utilize the gating module included in the first model to determine, based on the first word element corresponding to the first request, the third weight corresponding to each of the multiple expert modules included in the first model. For example, the first thread may utilize the gating module included in the first model to determine, based on the first word element corresponding to the first request and at least one third word element corresponding to the first request, the third weight corresponding to each of the multiple expert modules included in the first model. Based on the third weights corresponding to each of the multiple expert modules included in the first model, at least one first expert module corresponding to the first word element is determined from the multiple expert modules. The third weight may be used to indicate the activation probability of the expert module.

[0393] According to an embodiment of the present application, as described above, the first model may include at least one first layer. The first layer may include a gating module and multiple expert modules. The at least one first expert module corresponding to the first word may include at least one first expert module corresponding to each of the at least one first layers.

[0394] Regarding how the first thread utilizes the gating module included in the first model to determine the third weights corresponding to the plurality of expert modules included in the first model according to the first word corresponding to the first request, this can be achieved in the following manner.

[0395] As an implementation, for any first layer in at least one first layer, the first layer may be used as the current layer, and the first thread may utilize the gating module corresponding to the current layer to determine, based on the first word element corresponding to the first request, the third weights corresponding to the multiple expert modules corresponding to the current layer. For example, the third weights corresponding to the multiple expert modules corresponding to the current layer may be determined based on the first word element corresponding to the first request and the second word element corresponding to the first request.

[0396] How the first thread determines at least one first expert module corresponding to the first word-gram from the multiple expert modules according to the third weights corresponding to the multiple expert modules included in the first model can be implemented in the following manner.

[0397] As an implementation method, the first thread can determine at least one first expert module corresponding to the first word and the current layer from the multiple expert modules corresponding to the current layer based on the third weights corresponding to the multiple expert modules corresponding to the current layer.

[0398] How the first thread determines at least one first expert module corresponding to the first word-gram and the current layer from the multiple expert modules corresponding to the current layer can be implemented in the following manner.

[0399] The first method is based on the Top-T method. The first thread can determine T third weights from the third weights corresponding to the multiple expert modules corresponding to the current layer. The expert module corresponding to each of the T third weights is determined to be the first expert module corresponding to the third word and the current layer. When the larger the third weight is, the greater the probability of the expert module being activated is, the T third weights can be the maximum T third weights. When the smaller the third weight is, the greater the probability of the expert module being activated is, the T third weights can be the minimum T third weights. T can be an integer greater than or equal to 1 and less than or equal to S. T can be configured according to actual business needs and is not limited here. For example, S=16. T=2.

[0400] The second method is based on the threshold method. For example, when the third weight is larger, the probability of the expert module being activated is larger. The first thread can determine a third weight greater than or equal to the sixth threshold from the third weights corresponding to the multiple expert modules corresponding to the current layer. The expert module corresponding to the third weight greater than or equal to the seventh threshold is determined as the first expert module corresponding to the third word and the current layer. When the third weight is smaller, the probability of the expert module being activated is larger. The first thread can determine a third weight less than or equal to the eighth threshold from the third weights corresponding to the multiple expert modules corresponding to the current layer. The expert module corresponding to the third weight less than or equal to the eighth threshold is determined as the first expert module corresponding to the third word and the current layer. The seventh threshold can be greater than the eighth threshold. The seventh and eighth thresholds can be configured according to actual business needs and are not limited here.

[0401] In S1307 , the first thread determines whether there is a second expert module corresponding to the first expert module in the first set stored in the first storage unit. If so, S1308 - S1309 are executed; if not, S1310 is executed.

[0402] According to an embodiment of the present application, the second expert module may have a third identifier. The first expert module may have a first identifier. The first thread may determine whether a third identifier corresponding to the first identifier exists among the multiple third identifiers corresponding to the first set. If it is determined that the third identifier corresponding to the first identifier exists, it may be determined that a second expert module corresponding to the first expert module exists among the multiple second expert modules. If it is determined that the third identifier corresponding to the first identifier does not exist, it may be determined that a second expert module corresponding to the first expert module does not exist among the multiple second expert modules.

[0403] According to an embodiment of the present application, the second expert module corresponding to the first expert module can be understood as the second expert module being the first expert module. The third identifier corresponding to the first identifier can be understood as the third identifier being the first identifier.

[0404] At S1308 , the first thread obtains a second expert module corresponding to the first expert module from the first storage unit.

[0405] According to an embodiment of the present application, the first expert module may have a first identifier, and the second expert module may have a third identifier.

[0406] As an implementation, the first thread may determine a third identifier corresponding to the first identifier from the third identifiers corresponding to the plurality of second expert modules. The third expert module corresponding to the third identifier is determined as the second expert module corresponding to the first expert module. The first thread may obtain model parameters corresponding to the third identifier from the first storage unit, thereby obtaining the second expert module corresponding to the first expert module.

[0407] In S1309 , the first thread processes the first word-gram using a second expert module corresponding to the first expert module.

[0408] According to an embodiment of the present application, upon determining that a second expert module corresponding to the first expert module exists among the plurality of second expert modules stored in the first storage unit, the first thread can retrieve the second expert module corresponding to the first expert module from the first storage unit. Thus, the first thread can use the second expert module corresponding to the first expert module to process the first word, and then the first thread executes the second stage process using the first model to obtain an inference result.

[0409] In S1310 , the first thread determines whether there is a fourth expert module corresponding to the first expert module among the multiple expert modules included in the first model. If so, operation S1311 is executed; if not, S1314 - S1316 are executed.

[0410] At S1311 , the first thread determines whether there is a second expert module corresponding to the fourth expert module in the first set stored in the first storage unit. If so, S1312 - S1313 are executed; if not, S1314 - S1316 are executed.

[0411] At S1312 , the first thread obtains a second expert module corresponding to the fourth expert module from the first storage unit.

[0412] According to an embodiment of the present application, the fourth expert module may have a fifth identifier. The second expert module may have a third identifier.

[0413] As an implementation, the first thread may determine a third identifier corresponding to the fifth identifier from the third identifiers corresponding to each of the plurality of second expert modules corresponding to the first set. The second expert module corresponding to the third identifier is determined as the second expert module corresponding to the fourth expert module. The first thread may obtain model parameters corresponding to the third identifier from the first storage unit, thereby obtaining the second expert module corresponding to the fourth expert module.

[0414] At S1313 , the first thread processes the first word-gram using the second expert module corresponding to the fourth expert module.

[0415] At S1314 , the first thread loads the expert module corresponding to the first expert module from the second storage unit to the first storage unit.

[0416] According to an embodiment of the present application, the expert module may have a second identifier. The first expert module may have a first identifier.

[0417] As an implementation, the first thread may determine a second identifier corresponding to the first identifier from the second identifiers corresponding to the plurality of expert modules stored in the second storage unit. The expert module corresponding to the second identifier is determined as the expert module corresponding to the first expert module. The first thread may obtain model parameters corresponding to the second identifier from the second storage unit, and load the model parameters corresponding to the second identifier from the second storage unit to the first storage unit. Thus, the first thread loads the expert module corresponding to the first expert module from the second storage unit to the first storage unit.

[0418] At S1315 , the first thread obtains an expert module corresponding to the first expert module from the first storage unit.

[0419] At S1316 , the first thread processes the first word-gram using an expert module corresponding to the first expert module.

[0420] According to an embodiment of the present application, the fourth expert module can be used to replace the first expert module to process the first word. Figure 9A and Figure 9B Chinese expert module 922.

[0421] The following describes how to determine whether there is a fourth expert module corresponding to the first expert module among the multiple expert modules included in the first model.

[0422] The first expert module may have a corresponding first weight. The fourth expert module may have a corresponding second weight. The first weight may be used to indicate the activation probability of the first expert module. The second weight may be used to indicate the activation probability of the fourth expert module. The second weight needs to meet the following condition: the absolute value of the difference between the second weight and the first weight is less than or equal to the first threshold. Thus, the first weight may be greater than the second weight or less than the second weight, as long as the absolute value of the difference between the first weight and the second weight is less than or equal to the first threshold. The first threshold can be configured according to actual business needs and is not limited here. For example, the first threshold may be less than or equal to 0.3.

[0423] The above has explained how the first thread determines at least one first expert module from the multiple expert modules included in the first model according to the first request. Based on this, the following describes how to determine a fourth expert module corresponding to the first expert module.

[0424] The first approach is based on the Top-T approach. For example, the first thread can determine T third weights from the third weights corresponding to the multiple expert modules corresponding to the current layer. These T third weights can be determined as T first weights. For any one of the T first weights, the first thread can determine a second weight from the other third weights based on the first weight and a first threshold. For example, the first thread can determine a first range based on the first weight and the first threshold. The first thread can determine a third weight within the first range from the other third weights and determine the third weight within the first range as the second weight. Alternatively, the first thread can determine a third weight from the other third weights whose absolute difference with the first weight is less than or equal to the first threshold. The third weight whose absolute difference with the first weight is less than or equal to the first threshold is determined as the second weight. The other third weights can refer to third weights from the multiple third weights other than the first weight. Thus, the expert module corresponding to the second weight can be determined as the fourth expert module corresponding to the first expert module. If a larger third weight indicates a greater probability of expert module activation, the T third weights can be the maximum T third weights. In the case where the smaller the third weight, the greater the probability of the expert module being activated, the T third weights can be the minimum T third weights. T can be an integer greater than or equal to 1 and less than or equal to S. T can be configured based on actual business needs and is not limited here. For example, S = 16. T = 2.

[0425] It can be understood that if there is no third weight within the first range or a third weight whose absolute difference with the first weight is less than or equal to the first threshold among the other third weights, then the first expert module does not have a corresponding fourth expert module. In addition, if there are multiple third weights within the first range among the other third weights or there are multiple third weights whose absolute difference with the first weight is less than or equal to the first threshold among the other third weights, then the first expert module has multiple corresponding fourth expert modules.

[0426] The second approach is based on a threshold. For example, if a larger third weight indicates a greater probability of expert module activation, the first thread may determine a third weight greater than or equal to a seventh threshold from the third weights corresponding to the multiple expert modules corresponding to the current layer. The third weight greater than or equal to the seventh threshold may be determined as the first weight. If a smaller third weight indicates a greater probability of expert module activation, the first thread may determine a third weight less than or equal to an eighth threshold from the third weights corresponding to the multiple expert modules corresponding to the current layer. The third weight less than or equal to the eighth threshold may be determined as the first weight. The second weight may be determined from other third weights based on the first weight and the first threshold. For example, the first thread may determine a first range based on the first weight and the first threshold. The first thread may determine a third weight within the first range from other third weights. The third weight within the first range may be determined as the second weight. Alternatively, the first thread may determine a third weight from other third weights whose absolute difference with the first weight is less than or equal to the first threshold. The third weight whose absolute difference with the first weight is less than or equal to the first threshold may be determined as the second weight. The other third weights may refer to third weights other than the first weight among the multiple third weights. Thus, the second expert module corresponding to the second weight can be determined as the fourth expert module corresponding to the first expert module. The seventh threshold can be greater than the eighth threshold. The seventh threshold and the eighth threshold can be configured according to actual business needs and are not limited here.

[0427] The following describes how the fourth expert module is stored in the first storage unit.

[0428] As described above, the fourth expert module and the first expert module may correspond to the same request. For example, the fourth expert module and the first expert module may correspond to the first request. Alternatively, the fourth expert module and the first expert module may correspond to different requests. For example, the fourth expert module may correspond to the third request, while the first expert module may correspond to the first request. The following describes how to implement storing the fourth expert module in the first storage unit for each of the above two scenarios.

[0429] As an implementation method, if the fourth expert module and the first expert module correspond to the first request, the fourth expert module can be the third expert module determined by the second thread according to the first request using the second model when the first thread uses the first model to execute the first stage process corresponding to the first request, that is, the fourth expert module is the third expert module. For example, the fourth expert module can be Figure 9A In the expert module 922. How the second thread uses the second model to determine the third expert module according to the first request can be referred to the corresponding part above and will not be repeated here. In the above case, the fourth expert module can be stored in the first storage unit in the following manner.

[0430] For example, when a first thread executes a first-stage process corresponding to a first request using a first model, the first thread may determine, based on at least one third word corresponding to the first request, at least one second expert module corresponding to each of the at least one third word elements from a plurality of expert modules included in the first model, and load the at least one second expert module corresponding to each of the at least one third word elements from the second storage unit to the first storage unit to obtain a second set. When the first thread executes the first-stage process corresponding to the first request, the second thread may retain the second expert module corresponding to the third expert module in the first storage unit if it is determined that a second expert module corresponding to the third expert module exists in the second set stored in the first storage unit, that is, if it is determined that the third expert module exists in the second set stored in the first storage unit.

[0431] As another implementation, if the fourth expert module corresponds to the third request and the first expert module corresponds to the first request, then the fourth expert module can be stored in the first storage unit in the following manner, as the processing stages of the first request and the third request are different. For example, the fourth expert module can be Figure 9B Chinese expert module 922.

[0432] As an implementation, the fourth expert module can be the third expert module determined by the second thread when executing the preload process corresponding to the third request. For example, the fourth expert module can be the third expert module determined by the second thread based on the third request using the second model, while the first thread executes the first-stage process corresponding to the third request using the first model. The second thread can use a similar method to process the first request to determine the third expert module based on the third request using the second model, and will not be further described here. In this case, the fourth expert module can be stored in the first storage unit in the following manner.

[0433] For example, when the first thread uses the first model to execute the first stage process corresponding to the third request, the first thread can determine, based on the at least one third word corresponding to the third request, at least one second expert module corresponding to each of the at least one third word from the multiple expert modules included in the first model, and load the at least one second expert module corresponding to each of the at least one third word from the second storage unit to the first storage unit. The second set can include at least one second expert module corresponding to the third request. When the first thread executes the first stage process corresponding to the third request, the second thread can retain the second expert module corresponding to the third expert module in the first storage unit if it is determined that the second expert module corresponding to the third expert module exists in the second set stored in the first storage unit, that is, if it is determined that the third expert module exists in the second set stored in the first storage unit. Thus, the first set can include at least one second expert module corresponding to the third request.

[0434] As another implementation, the fourth expert module may be the second expert module determined by the first thread using the first model to execute the first stage process corresponding to the third request. If the second expert module does not exist in the first storage unit, the first thread may load the expert module corresponding to the second expert module from the second storage unit to the first storage unit.

[0435] As another implementation, the fourth expert module may be the first expert module determined by the first thread using the first model to execute the second stage process corresponding to the third request. If the first storage unit does not contain a second expert module corresponding to the first expert module, the first thread may load the expert module corresponding to the first expert module from the second storage unit to the first storage unit.

[0436] The following describes how the first thread processes the first word.

[0437] The fourth expert module may have a fifth identification. The second expert module may have a third identification.

[0438] When the first thread determines that a fourth expert module corresponding to the first expert module exists among the multiple expert modules, the first thread may determine whether a third identifier corresponding to the fifth identifier exists among the multiple third identifiers. When the first thread determines that a third identifier corresponding to the fifth identifier exists, the first thread may obtain a second expert module corresponding to the fourth expert module from the first storage unit, thereby processing the first word element using the second expert module corresponding to the fourth expert module. When the first thread determines that a third identifier corresponding to the fifth identifier does not exist, the first thread may load the expert module corresponding to the first expert module from the second storage unit to the first storage unit, obtain the expert module corresponding to the first expert module from the first storage unit, and process the first word element using the expert module corresponding to the first expert module.

[0439] When it is determined that there is no fourth expert module corresponding to the first expert module among multiple expert modules, the first thread may load the expert module corresponding to the first expert module from the second storage unit to the first storage unit, obtain the expert module corresponding to the first expert module from the first storage unit, and use the expert module corresponding to the first expert module to process the first word.

[0440] Based on the above, the first thread may repeatedly execute S1306 - S1316 until an inference result is obtained.

[0441] Because the storage capacity (or memory capacity) of the first storage unit is limited, during the execution of the data processing method based on the first and second threads, the first storage unit may still meet the first condition, thereby affecting the cache hit rate and resource utilization. The first condition can be used to trigger cache eviction. As an implementation, the first condition can be used to indicate that the first storage unit is full.

[0442] To improve cache hit rate and resource utilization, the fourth thread can eliminate the second expert module from the first set stored in the first storage unit based on the cache elimination policy when the first storage unit meets the first condition. The fourth thread can be executed in parallel with the first thread and the second thread.

[0443] The following describes how the fourth thread determines whether the first storage unit meets the first condition.

[0444] The first condition may include at least one of the following: the first capacity reaches the second threshold, the first number reaches the third threshold, the first usage rate reaches the fourth threshold, the first hit rate is less than or equal to the fifth threshold, or the second hit rate is greater than or equal to the ninth threshold.

[0445] The first capacity may be used to indicate the total storage size of the first set stored in the first storage unit. The second threshold may be determined based on the storage capacity of the first storage unit. For example, the second threshold may be the storage capacity of the first storage unit. Alternatively, the second threshold may be the product of the first coefficient and the storage capacity of the first storage unit. The first coefficient may be greater than 0 and less than 1. The first coefficient may be configured based on actual business needs and is not limited here.

[0446] The first number may be used to indicate the total number of second expert modules in the first set. The third threshold may be determined based on the storage capacity of the first storage unit and the storage size of the second expert modules. For example, the third threshold may be a ratio between the storage capacity of the first storage unit and the storage size of the second expert modules.

[0447] The first usage rate may be used to indicate the ratio between the first capacity and the storage capacity. The fourth threshold may be configured according to actual business needs and is not limited here. For example, the fourth threshold may be 0.9.

[0448] The first hit rate may be used to indicate a ratio between the second number and the third number. The second number may be used to indicate the number of cache hits corresponding to the first storage unit. The third number may be used to indicate the total number of accesses corresponding to the first storage unit. The number of cache hits may be used to indicate the number of times a first condition occurs. The first condition may be used to indicate the presence of a first expert module in the first storage unit.

[0449] The second hit ratio may be used to indicate a ratio between the fourth number and the third number. The fourth number may be used to indicate a number of cache misses corresponding to the first storage unit. The number of cache misses may be used to indicate a number of times a second situation occurs. The second situation may be used to indicate that the first expert module is not present in the first storage unit.

[0450] The fifth and ninth thresholds can be configured according to actual business needs and are not limited here. For example, the fifth threshold can be 0.9 and the ninth threshold can be 0.3.

[0451] The cache elimination strategy may include at least one of the following: first-in-first-out strategy, least recently used strategy, LRU-K strategy, least frequently used strategy, most recently used strategy, random elimination strategy, survival time strategy, second-level cache strategy, adaptive replacement cache strategy or low interactive reference interval strategy, etc.

[0452] The following describes how the fourth thread eliminates the second expert module from the first set stored in the first storage unit based on the cache elimination policy. The eliminated second expert module can be referred to as the sixth expert module.

[0453] As an implementation manner, the plurality of second expert modules in the first set have respective corresponding first moments, which can be used to indicate the moment when the second expert modules are stored in the first storage unit.

[0454] The fourth thread can eliminate the sixth expert module from the first set stored in the first storage unit based on the first moments corresponding to each of the multiple second expert modules, if the first storage unit satisfies the first condition. For example, the fourth thread can determine a second moment from the first moments corresponding to each of the multiple second expert modules. The second expert module corresponding to the second moment is determined as the sixth expert module, and the sixth expert module is eliminated. The second moment can be used to indicate the earliest first moment among the first moments corresponding to each of the multiple second expert modules. That is, the fourth thread can eliminate the sixth expert module from the multiple second expert modules if the first storage unit satisfies the first condition. The second moment is earlier than the first moment. The second moment can be used to indicate the storage moment of the sixth expert module in the first storage unit. The first moment can be used to indicate the moment corresponding to any second expert module among the multiple second expert modules except the sixth expert module.

[0455] As an implementation, the first set stored in the first storage unit can be stored in a queue, i.e., the first storage unit stores a queue. The queue can include multiple expert modules in the first set. A fourth thread can eliminate the second expert module at the head of the queue if the first storage unit meets the first condition. The second expert module at the head of the queue can be the sixth expert module. The queue can be implemented using an array or a linked list.

[0456] As another implementation, the plurality of second expert modules in the first set have respective corresponding third times, which may be used to indicate the time when the second expert module was most recently accessed.

[0457] If the first storage unit satisfies the first condition, the fourth thread may eliminate the sixth expert module from the first set stored in the first storage unit based on the third times corresponding to each of the multiple second expert modules. For example, the fourth thread may determine a fourth time from the third times corresponding to each of the multiple second expert modules. The second expert module corresponding to the fourth time is determined as the sixth expert module. The fourth time may indicate the earliest third time among the third times corresponding to each of the multiple second expert modules, or the fourth time may indicate the latest third time among the third times corresponding to each of the multiple second expert modules. In other words, if the first storage unit satisfies the first condition, the fourth thread may eliminate the sixth expert module from the multiple second expert modules. The fourth time may indicate the time when the sixth expert module was most recently accessed. The fourth time may be earlier than the third time, or the fourth time may be later than the third time. The third time may indicate the time corresponding to any second expert module other than the sixth expert module.

[0458] As another implementation, the plurality of second expert modules in the first set have respective corresponding first frequencies. The first frequencies may be used to indicate the frequency with which the second expert modules are accessed.

[0459] If the first storage unit satisfies the first condition, the fourth thread may eliminate the sixth expert module from the first set stored in the first storage unit based on the first frequencies corresponding to each of the multiple second expert modules. For example, the fourth thread may determine a second frequency from the first frequencies corresponding to each of the multiple second expert modules. The second expert module corresponding to the second frequency is determined as the sixth expert module, and the sixth expert module is eliminated. The second frequency may indicate the smallest first frequency among the first frequencies corresponding to each of the multiple second expert modules. In other words, the fourth thread may eliminate the sixth expert module from the multiple second expert modules if the first storage unit satisfies the first condition. The second frequency is less than the first frequency. The second frequency may indicate the frequency with which the sixth expert module is accessed. The first frequency may indicate the frequency corresponding to any second expert module in the multiple second expert modules other than the sixth expert module.

[0460] When the first thread executes the first stage process, if the first storage unit meets the first condition, the fourth thread can eliminate the second expert module from the first set stored in the first storage unit based on the cache elimination strategy, so that the first thread can load the expert module corresponding to the second expert module from the second storage unit to the first storage unit.

[0461] When the first thread executes the second stage process, if the first storage unit meets the first condition, the fourth thread can eliminate the second expert module from the first set stored in the first storage unit based on the cache elimination strategy, so that the first thread can load the expert module corresponding to the first expert module from the second storage unit to the first storage unit or the third thread can load the expert module corresponding to the first expert module from the second storage unit to the first storage unit.

[0462] For ease of understanding, the following takes the first model deployed on the intelligent assistant of the terminal device as an example. Figures 14A-14E right Figure 13 The embodiments described are described.

[0463] Figure 14A Schematic diagram of an application scenario of the data processing method provided in an embodiment of the present application.

[0464] like Figure 14A As shown, the display interface of the end-side device displays a dialog box for communicating with the intelligent assistant. The first request "Which city is FD University located in" is entered in the dialog box. The intelligent assistant can use Figure 13 The data processing method processes the first request and obtains the inference result "FD University is located in SH".

[0465] In addition, there are multiple ways to start the smart assistant. For example, the smart assistant can be started by long pressing the navigation bar at the bottom of the terminal device. The embodiments of this application do not limit the way to start the smart assistant.

[0466] The following combination Figures 14B-14E right Figure 14A How to use intelligent assistants Figure 13 The data processing method is described.

[0467] Figure 14B A schematic diagram of storing the expert module of the first model in the second storage unit provided in an embodiment of the present application.

[0468] like Figure 14B As shown, the first model 1410 may include a first layer 1411 and a first layer 1412. The first layer 1411 may include a gating module 1411_1 and 16 expert modules. The 16 expert modules may include expert modules 1411_2_1, ..., expert modules 1411_2_s, ..., and expert modules 1411_2_16. The first layer 1412 may include a gating module 1412_1 and 16 expert modules. The 16 expert modules may include expert modules 1412_2_1, ..., expert modules 1412_2_s, ..., and expert modules 1412_2_16. s may be an integer greater than or equal to 1 and less than or equal to 16. As an implementation, Figure 14B The first model 1410 can be Figure 6 The first model 600 is obtained when R = 2 and S = 16. In addition, the first model 1410 and the first model 600 may be different.

[0469] The 32 expert modules of the first model can be stored in the second storage unit. In addition, the second storage unit can also store a non-mixed expert layer.

[0470] exist Figure 14B On the basis of Figure 14C A schematic diagram of the principles of another data processing method provided in an embodiment of the present application.

[0471] like Figure 14C As shown, the first thread can tokenize the first request "In which city is FD University located?" to obtain multiple third tokens corresponding to the first request, namely, the third token ["FD"], the third token ["University"], the third token ["located in"], the third token ["which city"], and the third token ["city"]. The first thread can use the first model 1410 to perform the first phase process (i.e., the pre-filling phase process) to obtain the second token (i.e., the first output token). Specifically, the first thread can use the first model to obtain the second token based on the multiple third tokens corresponding to the first request. The first thread can use the first model 1410 to perform the second phase process (i.e., the decoding phase process). Specifically, the first thread can use the first model to obtain an inference result based on the first request and the second token. The inference result can be determined based on the multiple first tokens. The third token can be used to indicate an input token. The first token can be used to indicate an output token. The second token can also be referred to as a first token.

[0472] When the first thread executes the first phase process using the first model 1410 , the second thread may execute the preloading process using the second model 1420 .

[0473] exist Figure 14C Based on the following Figure 14D , it is described how the first thread uses the first model 1410 to execute the first stage process and how the second thread uses the second model 1420 to execute the preloading process.

[0474] Figure 14D A schematic diagram of an embodiment of the present application showing a first thread using a first model to execute a first-stage process and a second thread using a second model to execute a preloading process.

[0475] like Figure 14DAs shown, the first thread can implement the execution of the first-phase process using the first model in the following manner. For example, using the first model 1410, according to the first request, multiple second expert modules corresponding to multiple third tokens are obtained. For any third token among the multiple third tokens, the first thread can load the multiple second expert modules corresponding to the third token from the second storage unit to the first storage unit. Thus, a second set is obtained.

[0476] When the first thread is executing the first-phase process using the first model 1410, the second thread can use the second model 1420 to obtain multiple third expert modules according to the first request. For any one of the multiple third expert modules corresponding to the first request, during the process where the first thread loads the multiple second expert modules corresponding to the multiple third tokens from the second storage unit to the first storage unit, if the second thread determines that there is a second expert module corresponding to the third expert module in the second set stored in the first storage unit, that is, if it determines that the third expert module exists in the second set stored in the first storage unit, it can retain the second expert module corresponding to the third expert module in the first storage unit. Thus, the first storage unit stores a first set. The first set can include the second expert modules corresponding to the third expert modules. The second expert modules in the first set can be used to obtain the second expert module corresponding to the first expert module from the multiple second expert modules when the first thread uses the first model to execute the second-phase process.

[0477] Based on Figures 14B-14D this, taking the first token ["located at"] as an example, in combination with Figure 14E it is explained how the first thread uses the first model to execute the second-phase process.

[0478] Figure 14E This is a schematic diagram of the first thread using the first model to execute the second-phase process provided by an embodiment of this application.

[0479] As Figure 14EAs shown, the first thread uses gating module 1411_1 to determine, based on the first word ["located"], the first word ["FD"], the first word ["university"], and multiple third words corresponding to the first request, multiple first expert modules corresponding to the first word ["located"] and the first layer 1411 from the 16 expert modules included in the first layer 1411. The multiple first expert modules corresponding to the first word ["located"] and the first layer 1411 may include expert module 1411_2_1 and expert module 1411_2_16. The first thread determines that a second expert module corresponding to the first expert module (i.e., expert module 1411_2_1) and a second expert module corresponding to the first expert module (i.e., expert module 1411_2_16) exist in the first set stored in the first storage unit. Therefore, the first thread can retrieve these second expert modules from the first storage unit and use them to process the first word.

[0480] The first thread can use gating module 1412_1 to determine, based on the second characteristic, multiple first expert modules corresponding to the first word ["located"] and first layer 1412 from the 16 expert modules included in first layer 1412. The multiple first expert modules corresponding to the first word ["located"] and first layer 1412 may include expert module 1412_2_1 and expert module 1412_2_16. The second characteristic can be obtained by the first thread using first layer 1411 based on the first word ["located"], the first word ["FD"], the first word ["university"], and multiple third words corresponding to the first request. The first thread can determine that the first set stored in the first storage unit contains a second expert module corresponding to the first expert module (i.e., expert module 1412_2_1) and a second expert module corresponding to the first expert module (i.e., expert module 1412_2_16). Therefore, the first thread can retrieve the second expert modules from the first storage unit and use them to process the first word.

[0481] Figure 15 This is a flowchart of another data processing method provided in an embodiment of the present application. This method can be applied to the first thread, the second thread, and the third thread.

[0482] like Figure 15 As shown, the method includes S1501-S1515.

[0483] In S1501 , the first thread determines, based on at least one third word-gram corresponding to the first request, at least one second expert module corresponding to each of the at least one third word-grams from a plurality of expert modules included in the first model.

[0484] At S1502 , the first thread loads at least one second expert module corresponding to each of the at least one third word-grams from the second storage unit to the first storage unit to obtain a second set.

[0485] At S1503 , the first thread uses the second set and obtains a second word-gram corresponding to the first request according to at least one third word-gram.

[0486] At S1504 , the second thread uses the second model to obtain a plurality of third expert modules corresponding to the first request according to the first request.

[0487] At S1505 , the second thread retains the second expert module corresponding to the third expert module in the second set to obtain a first set.

[0488] At S1506 , the first thread uses the gating module included in the first model to determine, based on the first word-gram corresponding to the first request, at least one first expert module corresponding to the first word-gram from a plurality of expert modules included in the first model.

[0489] At S1507 , the first thread determines whether there is a second expert module corresponding to the first expert module in the first set stored in the first storage unit. If so, S1508 - S1509 are executed; if not, S1510 is executed.

[0490] At S1508 , the first thread obtains a second expert module corresponding to the first expert module from the first storage unit.

[0491] At S1509 , the first thread processes the first word-gram using a second expert module corresponding to the first expert module.

[0492] In S1510 , the first thread determines whether there is a fourth expert module corresponding to the first expert module among the multiple expert modules included in the first model. If so, S1511 is executed; if not, S1514 - S1515 are executed.

[0493] At S1511 , the first thread determines whether there is a second expert module corresponding to the fourth expert module in the first set. If so, S1512 - S1513 are executed; if not, S1514 - S1515 are executed.

[0494] At S1512 , the first thread obtains a second expert module corresponding to the fourth expert module from the first storage unit.

[0495] In S1513, the first thread processes the first word using the second expert module corresponding to the fourth expert module.

[0496] Since the absolute value of the difference between the second weight corresponding to the fourth expert module and the first weight corresponding to the first expert module is less than or equal to the first threshold, the second weight can be used to indicate the activation probability of the fourth expert module, and the first weight can be used to indicate the activation probability of the first expert module. Therefore, it can be explained that the prediction accuracy of the fourth expert module is relatively close to the prediction accuracy of the first expert module, and thus, the fourth expert module can replace the first expert module.

[0497] Based on this, if the second expert module corresponding to the first expert module does not exist in the first storage unit and the second expert module corresponding to the fourth expert module exists in the first storage unit, the first thread can directly obtain the second expert module corresponding to the fourth expert module from the first storage unit to use the second expert module corresponding to the fourth expert module to process the first word, thereby further improving the cache hit rate. In addition, because there is no need to first load the expert module corresponding to the first expert module from the second storage unit to the first storage unit, and then obtain the expert module corresponding to the first expert module and use the expert module corresponding to the first expert module to process the first word, the I / O latency between the second storage unit and the first storage unit is reduced, thereby improving the inference efficiency of the first model.

[0498] At S1514 , the third thread obtains an expert module corresponding to the first expert module from the second storage unit.

[0499] At S1515 , the third thread processes the first word-gram using an expert module corresponding to the first expert module.

[0500] According to the embodiment of the present application, the description of S1501-S1512 can be found in the above description of Figure 13 The description will not be repeated here.

[0501] and Figure 13 The difference between the illustrated embodiments is that, in this embodiment, if the first thread determines that the second expert module corresponding to the fourth expert module does not exist in the first set stored in the first storage unit, it uses the third thread to process the first word-unit. That is, the third thread can obtain the expert module corresponding to the first expert module from the second storage unit and use the expert module corresponding to the first expert module to process the first word-unit. The third speed can be greater than the fourth speed. The third speed can be used to indicate the speed at which the third thread accesses the second storage unit. The fourth speed can be used to indicate the speed at which the third thread accesses the first storage unit.

[0502] Since the first thread can be used to indicate the accelerator thread, the third thread can be used to indicate the processor thread, the first storage unit can be used to indicate the accelerator memory, the second storage unit can be used to indicate the processor memory, the third speed can be used to indicate the speed at which the third thread accesses the second storage unit, and the fourth speed can be used to indicate the speed at which the third thread accesses the first storage unit, the third access speed is greater than the fourth access speed. When the first thread determines that the first set stored in the first storage unit does not contain a second expert module corresponding to the fourth expert module, the third thread is used to obtain the expert module corresponding to the first expert module from the second storage unit and use the expert module corresponding to the first expert module to process the first word. This eliminates the need for the first thread to first load the expert module corresponding to the first expert module from the second storage unit to the first storage unit, obtain the expert module corresponding to the first expert module from the first storage unit, and then use the expert module corresponding to the first expert module to process the first word. Therefore, the I / O latency between the second storage unit and the first storage unit is reduced, and the inference efficiency of the first model is improved.

[0503] Figure 16 This is a flowchart of another data processing method provided in an embodiment of the present application. This method can be applied to the first thread and the second thread.

[0504] like Figure 16 As shown, the method includes S1601-S1612.

[0505] In S1601 , the first thread determines, based on at least one third word-gram corresponding to the first request, at least one second expert module corresponding to each of the at least one third word-grams from a plurality of expert modules included in the first model.

[0506] At S1602 , the first thread loads at least one second expert module corresponding to each of the at least one third word-grams from the second storage unit to the first storage unit to obtain a second set.

[0507] At S1603 , the first thread uses the second set and obtains a second word-gram corresponding to the first request according to at least one third word-gram.

[0508] At S1604 , the second thread uses the second model to obtain a plurality of third expert modules corresponding to the first request according to the first request.

[0509] At S1605 , the second thread retains the second expert module corresponding to the third expert module in the second set to obtain a first set.

[0510] At S1606 , the first thread uses the gating module included in the first model to determine, based on the first word-gram corresponding to the first request, at least one first expert module corresponding to the first word-gram from a plurality of expert modules included in the first model.

[0511] In S1607 , the first thread determines whether there is a second expert module corresponding to the first expert module in the first set stored in the first storage unit. If so, S1608 - S1609 are executed; if not, S1610 - S1612 are executed.

[0512] At S1608 , the first thread obtains a second expert module corresponding to the first expert module from the first storage unit.

[0513] At S1609 , the first thread processes the first word-gram using a second expert module corresponding to the first expert module.

[0514] At S1610 , the first thread loads an expert module corresponding to the first expert module from the second storage unit to the first storage unit.

[0515] At S1611 , the first thread obtains an expert module corresponding to the first expert module from the first storage unit.

[0516] At S1612 , the first thread processes the first word-gram using an expert module corresponding to the first expert module.

[0517] According to the embodiment of the present application, the description of S1601-S1609 can be found in the above description of Figure 13 The description will not be repeated here.

[0518] and Figure 13 The difference between the embodiments shown is that, in this embodiment, when the first thread determines that there is no second expert module corresponding to the first expert module in the first set stored in the first storage unit, S1610-S1612 is executed. Figure 13 S1314-S1316 shown.

[0519] Figure 17 This is a flowchart of another data processing method provided in an embodiment of the present application. This method can be applied to the first thread, the second thread, and the third thread.

[0520] like Figure 17 As shown, the method includes S1701-S1711.

[0521] In S1701 , the first thread determines, based on at least one third word-gram corresponding to the first request, at least one second expert module corresponding to each of the at least one third word-grams from a plurality of expert modules included in the first model.

[0522] At S1702 , the first thread loads at least one second expert module corresponding to each of the at least one third word-grams from the second storage unit to the first storage unit to obtain a second set.

[0523] At S1703 , the first thread uses the second set and obtains a second word-gram corresponding to the first request according to at least one third word-gram.

[0524] At S1704 , the second thread uses the second model to obtain a plurality of third expert modules corresponding to the first request according to the first request.

[0525] At S1705 , the second thread retains the second expert module corresponding to the third expert module in the second set to obtain a first set.

[0526] At S1706 , the first thread uses the gating module included in the first model to determine, based on the first word-gram corresponding to the first request, at least one first expert module corresponding to the first word-gram from a plurality of expert modules included in the first model.

[0527] In S1707 , the first thread determines whether there is a second expert module corresponding to the first expert module in the first set stored in the first storage unit. If so, S1708 - S1709 are executed; if not, S1710 - S1711 are executed.

[0528] At S1708 , the first thread obtains a second expert module corresponding to the first expert module from the first storage unit.

[0529] At S1709 , the first thread processes the first word-gram using a second expert module corresponding to the first expert module.

[0530] At S1710 , the third thread obtains an expert module corresponding to the first expert module from the second storage unit.

[0531] At S1711 , the third thread processes the first word-gram using an expert module corresponding to the first expert module.

[0532] According to the embodiment of the present application, the description of S1701-S1709 can be found in the above description of Figure 13 The description will not be repeated here.

[0533] and Figure 13 The difference between the illustrated embodiments is that, in this embodiment, when the first thread determines that there is no second expert module corresponding to the first expert module in the first set stored in the first storage unit, it uses the third thread to obtain the expert module corresponding to the first expert module from the second storage unit, and uses the expert module corresponding to the first expert module to process the first word.

[0534] Since the first thread can be used to indicate the accelerator thread, the third thread can be used to indicate the processor thread, the first storage unit can be used to indicate the accelerator memory, the second storage unit can be used to indicate the processor memory, the third speed can be used to indicate the speed at which the third thread accesses the second storage unit, and the fourth speed can be used to indicate the speed at which the third thread accesses the first storage unit, the third access speed is greater than the fourth access speed. When the first thread determines that the second expert module corresponding to the first expert module does not exist in the first set, the third thread is used to obtain the expert module corresponding to the first expert module from the second storage unit, and the expert module corresponding to the first expert module is used to process the first word. This does not require the first thread to first load the expert module corresponding to the first expert module from the second storage unit to the first storage unit, obtain the expert module corresponding to the first expert module from the first storage unit, and then use the expert module corresponding to the first expert module to process the first word. Therefore, the I / O delay between the second storage unit and the first storage unit is reduced, and the inference efficiency of the first model is improved.

[0535] It should be noted that the first layer 600_r, the first layer 900 and the first layer 1000 may be different from each other, may be partially the same, or may be completely the same.

[0536] The first layer 1411, the first layer 900 and the first layer 1000 may be different from each other, or may be partially the same, or may be completely the same. The first layer 1412, the first layer 900 and the first layer 1000 may be different from each other, or may be partially the same, or may be completely the same.

[0537] Expert modules 420_q, 612_r_s, 921, and 1021 can be different, partially identical, or completely identical. Expert modules 420_q, 612_r_s, 922, and 1022 can be different, partially identical, or completely identical. Expert modules 420_q, 612_r_s, 923, and 1023 can be different, partially identical, or completely identical. Expert modules 420_q, 612_r_s, 924, and 1024 can be different, partially identical, or completely identical. Gating modules 611_r, 910, and 1010 can be different, partially identical, or completely identical. q∈{1,…,Q}. r∈{1,…,R}. s∈{1,…,S}.

[0538] Expert module 1411_2_s, expert module 420_q, expert module 921, and expert module 1021 can be different from each other, or partially identical, or completely identical. Expert module 1411_2_s, expert module 420_q, expert module 922, and expert module 1022 can be different from each other, or partially identical, or completely identical. Expert module 1411_2_s, expert module 420_q, expert module 923, and expert module 1023 can be different from each other, or partially identical, or completely identical. Expert module 1411_2_s, expert module 420_q, expert module 924, and expert module 1024 can be different from each other, or partially identical, or completely identical. s∈{1,…,16}.

[0539] Expert module 1412_2_s, expert module 420_q, expert module 921, and expert module 1021 may be different from each other, partially identical, or completely identical. Expert module 1412_2_s, expert module 420_q, expert module 922, and expert module 1022 may be different from each other, partially identical, or completely identical. Expert module 1412_2_s, expert module 420_q, expert module 923, and expert module 1023 may be different from each other, partially identical, or completely identical. Expert module 1412_2_s, expert module 420_q, expert module 924, and expert module 1024 may be different from each other, partially identical, or completely identical.

[0540] Figure 18 A flowchart of another data processing method provided in an embodiment of the present application.

[0541] like Figure 18 As shown, the method includes S1810-S1820.

[0542] At S1810, when the first thread executes the second stage process using the first model, if the first storage unit contains a second expert module corresponding to the first expert module, the first thread retrieves the second expert module corresponding to the first expert module from the first storage unit. The first expert module is determined by the first thread from among multiple expert modules included in the first model based on the first word corresponding to the first request.

[0543] According to an embodiment of the present application, the first expert module may be determined by the first thread from among the multiple expert modules included in the first model based on third weights corresponding to the respective expert modules included in the first model. The third weight may be obtained by the first thread invoking a gating module included in the first model and using the gating module based on the first word corresponding to the first request.

[0544] It should be noted that for the description of how the first thread determines the first expert module, how the first thread determines whether the first expert module exists in the first storage unit, and how the first thread obtains the first expert module from the first storage unit, please refer to the corresponding parts above and will not be repeated here.

[0545] At S1820 , the first thread processes the first word-gram using a second expert module corresponding to the first expert module.

[0546] According to an embodiment of the present application, the first set stored in the first storage unit may be determined by the second thread based on the first request when the first thread executes the first stage process using the first model. The first set may include multiple second expert modules. The first stage process may be used to indicate the use of the first model to process the first request to obtain a second word. The second word may be used to indicate the first output word corresponding to the first request. The second stage process may be used to indicate the use of the first model to obtain an inference result based on the first request and the second word.

[0547] It should be noted that, for the description of how the first storage unit stores the plurality of second expert modules, please refer to the corresponding part above, which will not be repeated here.

[0548] According to an embodiment of the present application, since the multiple second expert modules of the first storage unit are determined by the second thread according to the first request when the first thread executes the first stage process using the first model, the parallel execution of the preloading process using the second thread and the first stage process using the first thread is achieved, thereby hiding the I / O delay as much as possible, thereby reducing the I / O delay and improving the reasoning efficiency of the first model. And because the first request has an expert preference, and the multiple second expert modules stored in the first storage unit can be determined according to the first request, therefore, when the first thread executes the second stage process using the first model, the first thread can directly obtain the second expert module corresponding to the first expert module from the first storage unit, thereby improving the cache hit rate of the first thread obtaining the first expert module from the first storage unit, thereby improving the reasoning efficiency of the first model. How the first storage unit stores the first set can be achieved in the following manner.

[0549] As an implementation, when a first thread executes a first-stage process using a first model, a second thread may determine, based on a first request, multiple third expert modules from the multiple expert modules included in the first model. The second thread may determine, from a second set, a second expert module corresponding to each of the multiple third expert modules. The second thread may retain, from the second set, the second expert modules corresponding to each of the multiple third expert modules, to obtain a first set. That is, the first set stored in the first storage unit may be obtained by the second thread retaining, when the first thread executes the first-stage process using the first model, the second expert modules corresponding to each of the multiple third expert modules in the second set. The multiple third expert modules may be determined by the second thread from the multiple expert modules included in the first model based on the first request.

[0550] The second set may be stored in the first storage unit, may include a plurality of second expert modules, and may be loaded from the second storage unit to the first storage unit by the first thread.

[0551] It should be noted that, for the description of the third expert module, the first set and the second set, please refer to the corresponding parts above and will not be repeated here.

[0552] Because the first set stored in the first storage unit can be obtained by the second thread retaining the second expert modules corresponding to each of the multiple third expert modules in the second set when the first thread executes the first stage process using the first model, the parallel execution of the preloading process by the second thread and the first stage process by the first thread is achieved. This minimizes I / O latency, thereby reducing I / O latency and improving the inference efficiency of the first model. Furthermore, because the first request has an expert preference, and the third expert module can be determined by the second thread based on the first request when the first thread executes the first stage process using the first model, the first thread can directly obtain the second expert module corresponding to the first expert module from the first storage unit when the first thread executes the second stage process using the first model. This improves the cache hit rate of the first thread obtaining the first expert module from the first storage unit, thereby improving the inference efficiency of the first model.

[0553] How the second thread determines multiple third expert modules from multiple expert modules according to the first request can be implemented in the following manner.

[0554] As an implementation, the second thread may utilize the second model to determine, based on the first request, multiple third expert modules from the multiple expert modules included in the first model. The second model may be trained using the second request and multiple first scores corresponding to the second request. The first scores may be used to indicate activation scores of the second expert modules.

[0555] It should be noted that, for the description of the second model, the second request and the first score, please refer to the corresponding parts above and will not be repeated here.

[0556] How the first thread stores the second set can be implemented in the following manner.

[0557] As an implementation method, the first thread may determine at least one third word element based on the first request. The first thread may determine at least one second expert module corresponding to each of the at least one third word elements from the multiple expert modules included in the first model. The first thread may load at least one second expert module corresponding to each of the at least one third word elements from the second storage unit to the first storage unit to obtain a second set. That is, the second set may be the result of the first thread loading at least one second expert module corresponding to each of the at least one third word elements from the second storage unit to the first storage unit. The at least one second expert module corresponding to each of the at least one third word elements may be determined by the first thread from the multiple expert modules included in the first model based on the at least one third word element. The at least one third word element may be determined by the first thread based on the first request.

[0558] It should be noted that, for the description of the second expert module, the third word unit and the second set, please refer to the corresponding parts above and will not be repeated here.

[0559] How the second thread determines the third expert module according to the first request can be implemented in the following manner.

[0560] As an implementation, the first set stored in the first storage unit can be determined by the second thread based on the first request using the second model when the first thread executes the first stage process using the first model. That is, the second thread can determine the first set stored in the first storage unit based on the first request using the second model when the first thread executes the first stage process using the first model. The second model can be trained using the second request and multiple first scores corresponding to the second request. The first score can be used to indicate the activation score of the expert module.

[0561] How the second model uses the second request and the multiple second scores corresponding to the second request to train the second model can be implemented in the following manner.

[0562] As an implementation, the second model can be obtained by adjusting model parameters of the second model based on a loss function value. The loss function value can be obtained based on the loss function and based on first scores and second scores corresponding to each of the multiple expert modules. The second scores corresponding to each of the multiple expert modules can be obtained by inputting the second request into the second model. The second scores can be used to indicate the predicted activation scores of the expert modules.

[0563] Regarding how to obtain the loss function value based on the loss function and the first scores and the second scores corresponding to each of the plurality of expert modules, it can be achieved in the following manner.

[0564] As an implementation, the loss function value may be obtained based on the loss function values corresponding to the plurality of expert modules. The loss function value corresponding to the expert module may be obtained by inputting the first score and the second score corresponding to the expert module into the loss function.

[0565] Since the second model can be trained based on the expert preferences of the request, that is, the second model can be trained using the second request and multiple first scores corresponding to the second request, the first score can be used to indicate the activation score of the expert module, thereby improving the expert prediction accuracy of the second model. On this basis, the first set stored in the first storage unit can be obtained by the second thread inputting the first request into the second model when the first thread executes the first stage process using the first model. Therefore, the first thread can directly obtain the second expert module corresponding to the first expert module from the first storage unit, thereby improving the cache hit rate of the first thread obtaining the first expert module from the first storage unit, thereby improving the inference efficiency of the first model.

[0566] How the first thread implements data processing when the second expert module corresponding to the first expert module does not exist in the first storage unit can be implemented in the following manner.

[0567] As an implementation, if the first storage unit does not contain the second expert module corresponding to the first expert module and the first storage unit contains the second expert module corresponding to the fourth expert module, the first thread may retrieve the second expert module corresponding to the fourth expert module from the first storage unit. The first thread may use the second expert module corresponding to the fourth expert module to process the first word.

[0568] According to an embodiment of the present application, the first expert module may have a corresponding first weight. The first weight may be used to indicate the activation probability of the first expert module. The fourth expert module may have a corresponding second weight. The second weight may be used to indicate the activation probability of the fourth expert module. The absolute value of the difference between the first weight and the second weight may be less than or equal to a first threshold.

[0569] Because the absolute value of the difference between the second weight corresponding to the fourth expert module and the first weight corresponding to the first expert module is less than or equal to the first threshold, the second weight can be used to indicate the activation probability of the fourth expert module, and the first weight can be used to indicate the activation probability of the first expert module. Therefore, it can be shown that the prediction accuracy of the fourth expert module is close to that of the first expert module, and thus, the fourth expert module can replace the first expert module. Based on this, if the second expert module corresponding to the first expert module does not exist in the first storage unit, and the second expert module corresponding to the fourth expert module exists in the first storage unit, the first thread can directly retrieve the second expert module corresponding to the fourth expert module from the first storage unit to process the first word-gram using the second expert module corresponding to the fourth expert module, thereby improving the cache hit rate. Furthermore, because there is no need to first load the expert module corresponding to the first expert module from the second storage unit to the first storage unit and then retrieve the expert module corresponding to the first expert module and use it to process the first word-gram, the I / O latency between the second storage unit and the first storage unit is reduced, thereby improving the inference efficiency of the first model.

[0570] As an implementation manner, the fourth expert module may be the second expert module in the first set corresponding to the first request.

[0571] As another implementation, the first set may further include multiple second expert modules corresponding to the third request. The fourth expert module may be the second expert module in the first set corresponding to the third request. The third request may be different from the first request.

[0572] It should be noted that, for the description of how the fourth expert module is stored in the first storage unit, please refer to the corresponding part above, which will not be repeated here.

[0573] Since the fourth expert module can be the second expert module corresponding to the first request or the second expert module corresponding to the third request, the diversity of sources of the fourth expert module is improved, which helps to further reduce I / O delay.

[0574] How the first thread determines the first expert module and the fourth expert module can be implemented in the following manner.

[0575] As an implementation, the first thread may call a gating module included in the first model and, using the gating module, obtain a third weight corresponding to each of the multiple expert modules based on the first word corresponding to the first request. The first thread may determine the first weight based on the third weight corresponding to each of the multiple expert modules. The first thread may determine the second weight from the third weights corresponding to each of the multiple expert modules based on the first weight and a first threshold. That is, the first weight may be determined by the first thread based on the third weight corresponding to each of the multiple expert modules included in the first model. The third weight corresponding to each of the multiple expert modules included in the first model may be obtained by the first thread calling a gating module included in the first model and, using the gating module, based on the first word corresponding to the first request. The second weight may be determined by the first thread based on the first weight and a first threshold, from the third weights corresponding to each of the multiple expert modules included in the first model.

[0576] It should be noted that for the description of how the first thread determines the first weight based on the third weights corresponding to each of the multiple expert modules included in the first model, and how the first thread determines the second weight from the third weights corresponding to each of the multiple expert modules included in the first model based on the first weight and the second weight, please refer to the corresponding part above and will not be repeated here.

[0577] Since the second weight corresponding to the fourth expert module can be determined from the third weights corresponding to the multiple expert modules included in the first model based on the first threshold and the first weight corresponding to the first expert module, the second weight can be used to indicate the activation probability of the fourth expert module, and the first weight can be used to indicate the activation probability of the first expert module. Therefore, the determined activation probability of the fourth expert module is closer to the activation probability of the first expert module. As a result, the fourth expert module can replace the first expert module to process the first word.

[0578] As another implementation method, the first thread can load the first expert module from the second storage unit to the first storage unit, obtain the second expert module corresponding to the first expert module from the first storage unit, and use the second expert module corresponding to the first expert module to process the first word when there is no second expert module corresponding to the first expert module in the first storage unit and there is no fourth expert module corresponding to the first expert module in the first storage unit.

[0579] As another implementation, if the first expert module is not present in the first storage unit and the second expert module corresponding to the fourth expert module is not present in the first storage unit, the third thread may retrieve the expert module corresponding to the first expert module from the second storage unit and use the expert module corresponding to the first expert module to process the first word. The third thread may be used to indicate a processor thread. The first thread may be used to indicate an accelerator thread.

[0580] Since the first thread can be used to indicate the accelerator thread, the third thread can be used to indicate the processor thread, the first storage unit can be used to indicate the accelerator memory, the second storage unit can be used to indicate the processor memory, the third speed can be used to indicate the speed at which the third thread accesses the second storage unit, and the fourth speed can be used to indicate the speed at which the third thread accesses the first storage unit, the third access speed is greater than the fourth access speed. When the first thread determines that the first storage unit does not contain a second expert module corresponding to the first expert module and does not contain a second expert module corresponding to the fourth expert module, the third thread is used to obtain the expert module corresponding to the first expert module from the second storage unit and use the expert module corresponding to the first expert module to process the first word, without the first thread first loading the expert module corresponding to the first expert module from the second storage unit to the first storage unit, obtaining the expert module corresponding to the first expert module from the first storage unit, and then using the expert module corresponding to the first expert module to process the first word. Therefore, the I / O latency between the second storage unit and the first storage unit is reduced, and the inference efficiency of the first model is improved.

[0581] As another implementation method, the first thread can load the expert module corresponding to the first expert module from the second storage unit to the first storage unit when the second expert module corresponding to the first expert module does not exist in the first storage unit, obtain the expert module corresponding to the first expert module from the first storage unit, and use the expert module corresponding to the first expert module to process the first word.

[0582] As another implementation, if the first thread does not have a second expert module corresponding to the first expert module in the first storage unit, the third thread can retrieve the expert module corresponding to the first expert module from the second storage unit and use the expert module corresponding to the first expert module to process the first word. The third thread can be used to indicate a processor thread. The first thread can be used to indicate an accelerator thread.

[0583] Since the first thread can be used to indicate the accelerator thread, the third thread can be used to indicate the processor thread, the first storage unit can be used to indicate the accelerator memory, the second storage unit can be used to indicate the processor memory, the third speed can be used to indicate the speed at which the third thread accesses the second storage unit, and the fourth speed can be used to indicate the speed at which the third thread accesses the first storage unit, the third access speed is greater than the fourth access speed. When the first thread determines that the first storage unit does not contain a second expert module corresponding to the first expert module, the third thread is used to obtain the expert module corresponding to the first expert module from the second storage unit and use the expert module corresponding to the first expert module to process the first word. This does not require the first thread to first load the expert module corresponding to the first expert module from the second storage unit to the first storage unit, obtain the expert module corresponding to the first expert module from the first storage unit, and then use the expert module corresponding to the first expert module to process the first word. Therefore, the I / O delay between the second storage unit and the first storage unit is reduced, and the inference efficiency of the first model is improved.

[0584] Regarding how to adjust the second expert module stored in the first storage unit when the first storage unit meets the first condition, it can be achieved in the following manner.

[0585] As an implementation, the fourth thread may eliminate the second expert module from the first set stored in the first storage unit based on a cache elimination policy when the first storage unit satisfies a first condition. The first condition may be used to trigger cache elimination. The fourth thread may be executed in parallel with the first and second threads.

[0586] Because the fourth thread can eliminate the second expert module from the first set stored in the first storage unit based on the cache elimination policy when the first storage unit meets the first condition, the fourth thread can be executed in parallel with the first thread and the second thread, thereby improving the cache hit rate and resource utilization when the storage capacity of the first storage unit is limited.

[0587] The first condition may include at least one of the following: the first capacity reaches the second threshold, the first number reaches the third threshold, the first usage rate reaches the fourth threshold, or the first hit rate is less than or equal to the fifth threshold. The first capacity can be used to indicate the total storage size of the first set. The second threshold can be determined based on the storage capacity of the first storage unit. The first number can be used to indicate the total number of second expert modules in the first set. The third threshold can be determined based on the storage capacity of the first storage unit and the storage size of the second expert module. The first usage rate can be used to indicate the ratio between the first capacity and the storage capacity of the first storage unit. The first hit rate can be used to indicate the ratio between the second number and the third number. The second number can be used to indicate the number of cache hits corresponding to the first storage unit. The third number can be used to indicate the total number of accesses corresponding to the first storage unit. And / or

[0588] The cache elimination policy may include at least one of the following: a first-in-first-out policy, a least recently used policy, a least frequently used policy, a most recently used policy, a random elimination policy, or a low interactive reference interval policy.

[0589] It should be noted that, for the description of how the fourth thread eliminates the second expert module from the first set stored in the first storage unit based on the cache elimination strategy, please refer to the corresponding part above and will not be repeated here.

[0590] As an implementation, the first storage unit may be used to indicate an accelerator memory, and the second storage unit may be used to indicate a processor memory or a storage device.

[0591] As another implementation, the first storage unit may be used to indicate a processor memory, and the second storage unit may be used to indicate a storage device.

[0592] Since the second storage unit can be used to store multiple expert modules included in the first model, and the first storage unit can be used to store the expert modules of the activated experts, the cooperation between the second storage unit and the first storage unit alleviates the limited memory capacity of the end-side device, making it possible for the end-side device to deploy the first model.

[0593] As an implementation, accelerator memory may be used to refer to graphics processor memory, neural network processor memory, or tensor processor memory.

[0594] As an implementation, the third thread may be a processor thread. If the first storage unit can be used to indicate accelerator memory, the first and second threads may be accelerator threads. The fourth thread may be an accelerator thread. If the first storage unit can be used to indicate processor memory, the first, second, and fourth threads may be processor threads.

[0595] It should be noted that the module names involved in the embodiments of the present application can be defined as other names as long as the functions of each module can be achieved, and there is no specific restriction on the names of the modules.

[0596] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in the embodiments of this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with the relevant laws, regulations and standards of relevant countries and regions, and provide corresponding operation entrances for users to choose to authorize or refuse.

[0597] The data processing method of the embodiment of the present application has been described above. The device for performing the above method provided in the embodiment of the present application is described below. Those skilled in the art will understand that the method and device can be combined and referenced with each other, and the relevant device provided in the embodiment of the present application can perform the steps in the above-mentioned list sorting method.

[0598] Figure 19 This is a structural block diagram of the data processing device provided in an embodiment of the present application.

[0599] like Figure 19 As shown, the data processing device 1900 may include an acquisition module 1910 and a processing module 1920 .

[0600] Acquisition module 1910 is configured to, when the first thread executes the second phase of the process using the first model, acquire, from the first storage unit, a second expert module corresponding to the first expert module, if the first thread has a second expert module corresponding to the first expert module in the first storage unit. The first expert module is determined by the first thread from among multiple expert modules included in the first model based on the first word corresponding to the first request.

[0601] The processing module 1920 is configured for the first thread to process the first word using the second expert module corresponding to the first expert module.

[0602] According to an embodiment of the present application, the first set stored in the first storage unit may be determined by the second thread based on the first request when the first thread executes the first stage process using the first model. The first set may include multiple second expert modules. The first stage process may be used to indicate the use of the first model to process the first request to obtain a second word. The second word may be used to indicate the first output word corresponding to the first request. The second stage process may be used to indicate the use of the first model to obtain an inference result based on the first request and the second word.

[0603] The data processing method provided in the embodiment of the present application can be applied to a terminal device with a communication function. The specific device form of the terminal device can refer to the above related description and will not be repeated here.

[0604] An embodiment of the present application provides an end-side device, which may include a processor and a memory. The memory stores computer-executable instructions. The processor executes the computer-executable instructions stored in the memory, causing the end-side device to perform the data processing method in the above-mentioned embodiment. The implementation principles and technical effects are similar to those of the above-mentioned related embodiments and will not be repeated here. The processor may include at least one of a central processing unit, a graphics processing unit, a neural network processor, or a tensor processor.

[0605] An embodiment of the present application provides a chip or chip system. The chip or chip system includes at least one processor and a communication interface. The communication interface and the at least one processor are interconnected via a line. The at least one processor is configured to execute a computer program or instruction to perform the data processing method in the above embodiment. The implementation principles and technical effects are similar to those of the above-mentioned related embodiments and are not further described here. The at least one processor may include at least one of a central processing unit, a graphics processing unit, a neural network processor, or a tensor processor.

[0606] The embodiments of the present application also provide a computer-readable storage medium. The computer-readable storage medium stores a computer program. When the computer program is executed by a processor, the above-mentioned method is implemented. The methods described in the above embodiments can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. If implemented in software, the functions can be stored as one or more instructions or codes on a computer-readable medium or transmitted on a computer-readable medium. Computer-readable media can include computer storage media and communication media, and can also include any medium that can transfer a computer program from one place to another. The storage medium can be any target medium that can be accessed by a computer.

[0607] In one possible implementation, a computer-readable medium may include random access memory (RAM), read-only memory (ROM), compact disc read-only memory (CD-ROM) or other optical disk storage, magnetic disk storage or other magnetic storage devices, or any other medium designed to carry or store the desired program code in the form of instructions or data structures and accessible by a computer. Furthermore, any connection is appropriately termed a computer-readable medium. For example, if software is transmitted from a website, server, or other remote source using a coaxial cable, fiber optic cable, twisted pair, digital subscriber line (DSL), or wireless technology (e.g., infrared, radio, or microwave), the coaxial cable, fiber optic cable, twisted pair, DSL, or wireless technology such as infrared, radio, or microwave is included in the definition of medium. Disk and disc, as used herein, include optical disc, laser disc, digital versatile disc (DVD), floppy disk, or Blu-ray disc. Disks typically reproduce data magnetically, while discs reproduce data optically using lasers. Combinations of the above should also be included within the scope of computer-readable media.

[0608] An embodiment of the present application provides a computer program product, which includes a computer program. When the computer program is executed, the computer executes the above method.

[0609] The embodiments of the present application are described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processing unit of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable device to generate a machine, so that the instructions executed by the processing unit of the computer or other programmable data processing device generate a device for implementing the functions specified in one or more processes in the flowchart and / or one or more blocks in the block diagram.

[0610] The above specific implementation methods further explain in detail the purpose, technical solutions and beneficial effects of the embodiments of the present application. It should be understood that the above are only specific implementation methods of the embodiments of the present application and are not intended to limit the scope of protection of the embodiments of the present application. Any modifications, equivalent replacements, improvements, etc. made on the basis of the technical solutions of the embodiments of the present application should be included in the scope of protection of the embodiments of the present application.

Claims

1. A data processing method, characterized in that: include: When the first thread executes the second stage process using the first model, The first thread obtains the second expert module corresponding to the first expert module from the first storage unit if the second expert module corresponding to the first expert module exists in the first storage unit, wherein the first expert module is determined by the first thread from a plurality of expert modules included in the first model according to the first word corresponding to the first request; The first thread processes the first word using a second expert module corresponding to the first expert module; Among them, the first set stored in the first storage unit is determined by the second thread according to the first request when the first thread uses the first model to execute the first stage process. The first set includes multiple second expert modules. The first stage process is used to use the first model to process the first request to obtain a second word element. The second word element is used to indicate the first output word element corresponding to the first request. The second stage process is used to use the first model to obtain an inference result based on the first request and the second word element.

2. The method according to claim 1, characterized in that Also includes: When the first thread executes the first stage process using the first model, The second thread determines a plurality of third expert modules from the plurality of expert modules according to the first request; The second thread determines, from a second set, a second expert module corresponding to each of the plurality of third expert modules, wherein the second set is stored in the first storage unit, the second set includes a plurality of second expert modules, and the second set is loaded from the second storage unit to the first storage unit by the first thread; The second thread retains the second expert modules in the second set corresponding to each of the plurality of third expert modules to obtain the first set.

3. The method according to claim 2, characterized in that Also includes: The first thread determines a third word according to the first request; The first thread determines at least one second expert module corresponding to the third word-gram from the plurality of expert modules; The first thread loads at least one second expert module corresponding to the third word-gram from the second storage unit to the first storage unit to obtain the second set.

4. The method according to claim 2 or 3, characterized in that The second thread determines, according to the first request, a plurality of third expert modules from the plurality of expert modules, including: The second thread uses a second model to determine the multiple third expert modules from the multiple expert modules according to the first request, wherein the second model is trained using the second request and multiple first scores corresponding to the second request, and the first scores are used to indicate the activation scores of the expert modules.

5. The method according to any one of claims 1 to 3, characterized in that Also includes: The first thread obtains the second expert module corresponding to the fourth expert module from the first storage unit when the first storage unit does not contain the second expert module corresponding to the first expert module and the first storage unit contains the second expert module corresponding to the fourth expert module; The first thread processes the first word using a second expert module corresponding to the fourth expert module; The first expert module has a corresponding first weight, which is used to indicate the activation probability of the first expert module; the fourth expert module has a corresponding second weight, which is used to indicate the activation probability of the fourth expert module; and the absolute value of the difference between the first weight and the second weight is less than or equal to the first threshold.

6. The method according to claim 5, characterized in that Also includes: The first thread calls a gating module included in the first model, and uses the gating module to obtain third weights corresponding to the plurality of expert modules according to a first word corresponding to the first request; The first thread determines the first weight according to third weights corresponding to each of the plurality of expert modules; The first thread determines the second weight from third weights corresponding to each of the plurality of expert modules according to the first weight and the first threshold.

7. The method according to claim 5, characterized in that The fourth expert module is the second expert module in the first set corresponding to the first request; or The first set further includes a plurality of second expert modules corresponding to a third request, and the fourth expert module is a second expert module in the first set corresponding to the third request, wherein the third request is different from the first request.

8. The method according to claim 5, characterized in that Also includes: The first thread, if the first storage unit does not contain a second expert module corresponding to the first expert module and the first storage unit does not contain a second expert module corresponding to the fourth expert module, loads the expert module corresponding to the first expert module from the second storage unit to the first storage unit, obtains the expert module corresponding to the first expert module from the first storage unit, and processes the first word using the expert module corresponding to the first expert module; or When the first storage unit does not contain a second expert module corresponding to the first expert module and the first storage unit does not contain a second expert module corresponding to the fourth expert module, the third thread obtains the expert module corresponding to the first expert module from the second storage unit and uses the expert module corresponding to the first expert module to process the first word element, wherein the third thread is used to indicate a processor thread, and the first thread is used to indicate an accelerator thread.

9. The method according to any one of claims 1 to 3, characterized in that Also includes: If the first storage unit does not contain a second expert module corresponding to the first expert module, the first thread loads the expert module corresponding to the first expert module from the second storage unit to the first storage unit, obtains the expert module corresponding to the first expert module from the first storage unit, and processes the first word using the expert module corresponding to the first expert module; or When the first storage unit does not contain a second expert module corresponding to the first expert module, the third thread obtains the expert module corresponding to the first expert module from the second storage unit, and uses the expert module corresponding to the first expert module to process the first word, wherein the third thread is used to indicate a processor thread, and the first thread is used to indicate an accelerator thread.

10. The method according to any one of claims 1 to 3, characterized in that Also includes: The fourth thread eliminates the second expert module from the first set based on a cache elimination policy when the first storage meets a first condition, wherein the first condition is used to trigger cache elimination, and the fourth thread is executed in parallel with the first thread and the second thread.

11. The method according to claim 10, characterized in that The first condition includes at least one of the following: the first capacity reaches a second threshold, the first number reaches a third threshold, the first usage rate reaches a fourth threshold, or the first hit rate is less than or equal to a fifth threshold, the first capacity is used to indicate the total storage size of the first set, the second threshold is determined according to the storage capacity of the first storage unit, the first number is used to indicate the total number of second expert modules in the first set, the third threshold is determined according to the storage capacity of the first storage unit and the storage size of the second expert modules, the first usage rate is used to indicate the ratio between the first capacity and the storage capacity of the first storage unit, the first hit rate is used to indicate the ratio between the second number and the third number, the second number is used to indicate the number of cache hits corresponding to the first storage unit, and the third number is used to indicate the total number of accesses corresponding to the first storage unit; and / or The cache elimination strategy includes at least one of the following: a first-in-first-out strategy, a least recently used strategy, a least frequently used strategy, a most recently used strategy, a random elimination strategy, or a low interactive reference interval strategy.

12. The method according to any one of claims 1 to 3, characterized in that The first storage unit is used to indicate the accelerator memory, and the second storage unit is used to indicate the processor memory or storage device; or The first storage unit is used to indicate the processor memory, and the second storage unit is used to indicate the storage device.

13. The method according to claim 12, characterized in that The accelerator memory is used to indicate a graphics processor memory, a neural network processor memory, or a tensor processor memory.

14. A terminal side device, characterized in that: include: processor and memory; The memory stores computer-executable instructions; The processor executes the computer-executable instructions stored in the memory, so that the end-side device executes the method according to any one of claims 1 to 13.

15. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 13 is implemented.

16. A chip system, characterized in that: The system comprises at least one processor and a communication interface, wherein the communication interface and the at least one processor are interconnected via a line, and the at least one processor is configured to run a computer program or instruction to execute the method according to any one of claims 1 to 13.

17. A computer program product, characterized in that The method comprises a computer program, which, when executed, causes a terminal-side device to execute the method according to any one of claims 1 to 13.

Citation Information

Patent Citations

  • Model training method and device and data processing method and device

    CN112651511A

  • Memory management method and device for large language model

    CN119576805A