Model processing method, electronic device, storage medium, and program product

By acquiring the activity level and current deployment devices of expert submodules, the deployment location and resource allocation of expert submodules are dynamically adjusted, solving the problem of resource allocation mismatch in hybrid expert models and improving model processing efficiency and resource utilization.

CN122285303APending Publication Date: 2026-06-26INSPUR SUZHOU INTELLIGENT TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
INSPUR SUZHOU INTELLIGENT TECH CO LTD
Filing Date
2026-05-28
Publication Date
2026-06-26

AI Technical Summary

Technical Problem

In existing technologies, the processing efficiency of models based on hybrid expert models is relatively low, mainly due to insufficient resource supply for expert sub-modules activated at high frequencies, while expert sub-modules activated at low frequencies occupy redundant resources, resulting in a mismatch in resource allocation.

Method used

By acquiring the activity level and current deployed devices of expert submodules, the deployment location and resource allocation of expert submodules are dynamically adjusted, including migration and compression processing, to ensure that high-frequency modules are matched with high-computing-power devices and that low-frequency modules have their resource usage compressed, thereby achieving adaptive optimization of resources.

Benefits of technology

It improves model processing efficiency and resource utilization, avoids resource mismatch problems caused by static configuration, and realizes dynamic optimization and flexible deployment in heterogeneous environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122285303A_ABST
    Figure CN122285303A_ABST
Patent Text Reader

Abstract

This application discloses a model processing method, electronic device, storage medium, and program product, relating to the field of computer technology. It includes the ability to acquire the activity level of each expert sub-module in a first hybrid expert model and its currently deployed first execution device, and dynamically determine the first expert sub-module that needs to be migrated based on the matching relationship between activity thresholds and device types. The method then performs migration processing on its execution device to obtain a second hybrid expert model. Therefore, it effectively avoids the problem in related technologies where static deployment makes it difficult to adapt to dynamic changes in the frequency of expert sub-module calls. It also solves the technical problem of low processing efficiency caused by the mismatch between hardware resource allocation and the real-time activity level of expert sub-modules during model operation, achieving the technical effect of dynamically optimizing the deployment devices of expert sub-modules and improving the overall processing efficiency and resource utilization of the model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, and in particular to a model processing method, electronic device, storage medium, and program product. Background Technology

[0002] Models based on the Mixture-of-Experts (MoE) architecture are an important form of model deployment. To ensure the inference efficiency of MoE models on heterogeneous hardware, each expert sub-module of the model needs to be processed.

[0003] In related technologies, static configuration is typically used for model processing. This involves adjusting the parameters of each expert submodule and deploying them on the corresponding components before model execution. However, during model execution, the activation status of each expert submodule changes dynamically. This can lead to insufficient resource supply for frequently activated expert submodules, while infrequently activated submodules may consume redundant resources, resulting in low model processing efficiency. Summary of the Invention

[0004] This application provides a model processing method, electronic device, storage medium, and program product to at least solve the problem of low efficiency in model processing in related technologies.

[0005] This application provides a model processing method, including:

[0006] The activity level of each expert submodule in the first hybrid expert model and the first execution device for deploying each expert submodule are obtained. The activity level is used to indicate the frequency of the expert submodule calls.

[0007] Based on the activity levels of multiple expert sub-modules and the first execution device, in each expert sub-module of the first hybrid expert model, a first expert sub-module is determined. The activity level of the first expert sub-module is greater than or equal to a first threshold, and the first execution device is a first preset device; or, the activity level of the first expert sub-module is less than a second threshold, and the first execution device is a second preset device.

[0008] The first expert submodule is transferred to obtain the second hybrid expert model.

[0009] This application also provides a model processing apparatus, comprising: an acquisition module, a determination module, and a processing module, wherein:

[0010] The acquisition module is used to acquire the activity level of each expert sub-module in the first hybrid expert model and the first execution device for deploying each expert sub-module. The activity level is used to indicate the call frequency of the expert sub-module.

[0011] The determination module is used to determine the first expert submodule among the expert submodules of the first hybrid expert model based on the activity levels of the multiple expert submodules and the first execution device. The activity level of the first expert submodule is greater than or equal to a first threshold and the first execution device is a first preset device; or the activity level of the first expert submodule is less than a second threshold and the first execution device is a second preset device.

[0012] The processing module is used to perform transfer processing on the first expert submodule to obtain the second hybrid expert model.

[0013] This application also provides an electronic device, including: a memory for storing a computer program; and a processor for implementing the steps of any of the above-described model processing methods when executing the computer program.

[0014] This application also provides a computer-readable storage medium storing a computer program, wherein the computer program, when executed by a processor, implements the steps of any of the above-described model processing methods.

[0015] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the steps of any of the above-described model processing methods.

[0016] The model processing method, electronic device, storage medium, and program product provided in this application can obtain the activity level of each expert sub-module in the first hybrid expert model and its currently deployed first execution device. Based on the matching relationship between the activity threshold and the device type, the first expert sub-module that needs to be migrated is dynamically determined, and its execution device is migrated to obtain the second hybrid expert model. Therefore, it can effectively avoid the problem in related technologies where the static deployment method is difficult to adapt to the dynamic changes in the calling frequency of expert sub-modules. It also solves the technical problem of low processing efficiency caused by the mismatch between hardware resource allocation and the real-time activity level of expert sub-modules during model operation. This achieves the technical effect of dynamically optimizing the deployment device of expert sub-modules and improving the overall processing efficiency and resource utilization of the model. Attached Figure Description

[0017] To more clearly illustrate the embodiments of this application, the accompanying drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0018] Figure 1 This is a schematic diagram of the system architecture provided for an embodiment of this application;

[0019] Figure 2A schematic flowchart illustrating a model processing method provided in an embodiment of this application;

[0020] Figure 3 A schematic diagram illustrating a method for determining a second hybrid expert model provided in an embodiment of this application;

[0021] Figure 4 A schematic diagram illustrating a method for determining a third hybrid expert model provided in an embodiment of this application;

[0022] Figure 5 A flowchart illustrating another model processing method provided in an embodiment of this application;

[0023] Figure 6 This is a schematic diagram of the structure of a model processing device provided in an embodiment of this application;

[0024] Figure 7 A schematic diagram of the structure of the electronic device provided in this application. Detailed Implementation

[0025] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, other embodiments obtained by those of ordinary skill in the art without creative effort are all within the protection scope of this application.

[0026] It should be noted that, in the description of this application, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. The terms "first," "second," etc., in this application are used to distinguish similar objects and are not used to describe a specific order or sequence.

[0027] In related technologies, static configuration is typically used for model processing. This involves adjusting the parameters of each expert submodule and deploying them on the corresponding components before model execution. However, during model execution, the activation status of each expert submodule changes dynamically. This can lead to insufficient resource supply for frequently activated expert submodules, while infrequently activated submodules may consume redundant resources, resulting in low model processing efficiency.

[0028] To address the aforementioned issues, in this embodiment, when model processing is required, the activity level of each expert submodule in the first hybrid expert model and the first execution device deploying each expert submodule are obtained. The activity level is used to indicate the call frequency of the expert submodule. Based on the activity levels and first execution devices corresponding to multiple expert submodules, a first expert submodule is determined among the expert submodules in the first hybrid expert model. The activity level of the first expert submodule is greater than or equal to a first threshold, and the first execution device is a first preset device; or, the activity level of the first expert submodule is less than a second threshold, and the first execution device is a second preset device. The first expert submodule is then subjected to migration processing to obtain the second hybrid expert model. In this way, the optimal deployment location and compression status of each expert submodule can be dynamically determined based on the actual activity level of each submodule during runtime and the real-time status of device resources. This achieves the matching of computing resource supply with the real-time load demand of the modules, and can adaptively adjust the resource allocation between high-frequency and low-frequency calling modules. This effectively avoids the resource mismatch problem caused by the rigidity of traditional static configuration methods, and enables dynamic optimization of the model in heterogeneous environments based on load and resource changes. This improves resource utilization and deployment flexibility while enhancing model processing efficiency.

[0029] To enable those skilled in the art to better understand the present application, the present application will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0030] This section describes the specific application environment architecture or hardware architecture upon which the model processing method depends. (References) Figure 1 , Figure 1 This is a schematic diagram of the system architecture provided for an embodiment of this application. Please refer to [link / reference]. Figure 1 This includes electronic devices and terminal devices. Electronic devices can be any device with on-device computing capabilities, such as servers. Each electronic device includes multiple expert sub-modules. Terminal devices can send tasks to be executed to the electronic device. After receiving the task, the electronic device can execute it. During task execution, it processes each expert sub-module in the model used to execute the task (e.g., device migration) to obtain an updated model, and then continues executing the task to improve model processing efficiency and overall resource utilization.

[0031] Figure 2 This is a flowchart illustrating a model processing method provided in an embodiment of this application, as shown below. Figure 2 As shown, embodiments of this application provide a model processing method, which is described in detail below:

[0032] S201. Obtain the activity level of each expert submodule in the first hybrid expert model and the first execution device for deploying each expert submodule.

[0033] The execution subject of this application embodiment can be an electronic device or a model processing device installed in an electronic device. The model processing device can be implemented by software or by a combination of software and hardware.

[0034] The activity level is used to indicate the frequency of calls to expert submodules. In other words, the activity level refers to how frequently an expert submodule is called, and its value directly reflects the importance or popularity of that expert submodule under the current workload.

[0035] For example, the higher the activity level of an expert submodule, the more frequently it is called and the higher its contribution to the model inference results. Conversely, the lower the activity level of an expert submodule, the sparser it is called and the lower its contribution to the model inference results.

[0036] The first hybrid expert model can be the MoE model, which contains multiple independent expert sub-modules and selects some expert sub-modules to perform corresponding tasks through a routing mechanism.

[0037] Each expert submodule has independent inference and computation capabilities, and can receive routing instructions and execute specific computational tasks (e.g., attention computation, feedforward layer operations, etc.) independently.

[0038] The first execution device can be a physical or logical computing device that currently deploys and runs multiple expert sub-modules; that is, the first execution device can be a heterogeneous computing device that deploys each expert sub-module in the first hybrid expert model.

[0039] For example, the first execution device may be a graphics processing unit (GPU), a central processing unit (CPU), or a specific memory region thereon, wherein the GPU is a preferred high-performance computing device.

[0040] In some embodiments, the electronic device may obtain the activity level of each expert sub-module in the first hybrid expert model based on the following implementation: obtaining the number of task executions, the number of calls to each expert sub-module, the initial activity level, and the adjustment coefficient of the first hybrid expert model within a historical period; and for any expert sub-module, weighting the number of task executions, the number of calls, and the initial activity level according to the adjustment coefficient to obtain the activity level.

[0041] The historical period can be the most recent continuous running period of the first hybrid expert model. In other words, the historical period can refer to the time window that summarizes the call status of each expert submodule. The historical period can be a configurable fixed period.

[0042] For example, the historical period can be the most recent 1 hour, the most recent 24 hours, etc., and its length can be flexibly set according to the frequency of model inference (for example, a shorter historical period can be selected for high-frequency inference scenarios, and a longer historical period can be selected for low-frequency inference scenarios).

[0043] The number of task executions can be the total number of task batches processed by the first hybrid expert model within a historical period. For example, if the first hybrid expert model processed a total of 50 inference batches within a historical period, then the number of task executions is 50.

[0044] The number of calls can be the number of times a specific expert submodule is selected by the routing mechanism of the first hybrid expert model and actually participates in the execution of tasks within a historical period. For example, if expert submodule A is called 60 times in 100 inference batches within a historical period (the most recent 5 hours), then its number of calls is 60.

[0045] The initial activity level can be the initial setting value of the activity level of the expert sub-modules before the start of the historical period. It can be set according to the initialization configuration of the first hybrid expert model. For example, the initial activity level of each expert sub-module can be uniformly set to 0.5. Alternatively, it can be obtained by summarizing the calling frequency of the expert sub-modules during the model pre-training stage. For example, the initial activity level of the expert sub-modules called frequently during pre-training can be set to 0.8, and the initial activity level of the expert sub-modules called infrequently can be set to 0.2.

[0046] The adjustment coefficient can be a preset weight parameter used to balance the influence of historical activity and recent call data when calculating new activity. The smaller the adjustment coefficient, the higher the weight of recent call data, and the more sensitive the activity is to changes in call frequency.

[0047] In this embodiment, the adjustment coefficient can be understood as an attenuation coefficient. The adjustment coefficient can be preset to a fixed value, for example, the adjustment coefficient is 0.8. Alternatively, the coefficient can be dynamically adjusted according to actual needs, so that the adjustment coefficient can reflect the reduction of the historical proportion and the enhancement of the proportion of new data collection.

[0048] In some embodiments, the electronic device can obtain the number of times the first hybrid expert model has executed tasks, the number of times each expert submodule has been called, the initial activity level, and the adjustment coefficient through the log recording module within a historical period. Specifically, when the first hybrid expert model executes a task, the log recording module will record the completion status of each task, the inference batch number, and the call status of each expert submodule (whether it has been selected for execution by routing) in real time.

[0049] When it is necessary to obtain data for a historical period, the electronic device filters the task logs within that historical period from the log recording module, obtains the initial activity level and adjustment coefficient, calculates the total number of task batches in the log records to obtain the number of task executions, and for each expert submodule, it iterates through the log records to determine the total number of records marked as having been called and executed, thus obtaining the number of calls to each expert submodule.

[0050] In some embodiments, the electronic device may employ an exponentially decaying weighted average algorithm, by adjusting the coefficients. To balance initial activity (historical data) and recent call percentage (current data), the relevant formula for weighted processing can be expressed as:

[0051]

[0052] in, This represents the activity level of the i-th expert submodule after the update. This represents the adjustment result of the initial activity level of the i-th expert submodule. This represents the number of times the i-th expert submodule is called, and N represents the number of task executions. This represents the average number of times the i-th expert submodule was triggered by the route and actually executed in the most recent N batches. This represents the adjustment factor, and its value is (0, 1).

[0053] In this way, the weighted processing retains historical activity information while introducing recent call data, making activity a dynamic indicator that smoothly reflects the recent call trend of expert submodules, thus avoiding interference from instantaneous fluctuations.

[0054] S202. Based on the activity levels of multiple expert sub-modules and the first execution device, determine the first expert sub-module among the expert sub-modules of the first hybrid expert model.

[0055] Wherein, the activity level of the first expert submodule is greater than or equal to the first threshold and the first execution device is a first preset device, or the activity level of the first expert submodule is less than the second threshold and the first execution device is a second preset device.

[0056] In some embodiments, when the activity level of the first expert submodule is greater than or equal to the first threshold and the first execution device is the second preset device, or when the activity level of the first expert submodule is less than the second threshold and the first execution device is the first preset device, it is not necessary to switch the first execution device.

[0057] The first expert submodule can be any expert submodule to be processed in the first hybrid expert model. In the actual execution process, the electronic device will traverse each expert submodule and process them one by one as the first expert submodule to ensure that the deployment method of each first expert submodule is adapted to its activity level.

[0058] The first threshold can be a pre-set high-activity threshold, used to define whether the expert submodule belongs to the high-frequency call category. This indicates that the value range is (0, 1). The first threshold can be configured according to business needs or model characteristics (e.g., 0.7), and expert sub-modules with an activity level greater than or equal to this threshold are determined to be high-frequency call modules and should be prioritized for matching with high computing power resources.

[0059] The first preset device can be a general-purpose computing device adapted to low-activity expert sub-modules, that is, the first preset component can be a CPU. Its core advantages are flexible resource scheduling, low deployment cost, and efficient handling of inference tasks of low-frequency call modules, avoiding the occupation of valuable GPU resources.

[0060] The second threshold can be a pre-set low-activity threshold, used to define whether the expert submodule belongs to the low-frequency call category, and the second threshold is less than the first threshold. This indicates that the value range is (0, 1). Among them, expert submodules with an activity level less than the second threshold (e.g., 0.3) are determined to be low-frequency call modules, which can be deployed on general computing resources or compressed to save high computing power resources.

[0061] The second preset device can be a high-performance heterogeneous device adapted to the highly active expert sub-module, that is, the second preset device is a GPU, whose core advantages are strong parallel computing capabilities and low inference latency, which can meet the computing power requirements of high-frequency calling modules.

[0062] In some embodiments, the electronic device obtains the current activity level of the expert submodule. The first execution device (GPU or CPU) currently deployed, and a preset first threshold is invoked. Second threshold Then, condition judgment is performed to determine the expert submodules that do not require migration processing.

[0063] Specifically, the first scenario is that the activity level of the expert submodule... (High activity level), and the currently deployed first execution device is the second preset device, GPU. In this case, the highly active expert submodule needs to be matched with a high-computing-power device. The deployment status in this case is highly compatible with the module requirements, which can give full play to the parallel computing advantages of GPU and avoid the increase in inference latency caused by deployment on CPU. Therefore, there is no need to switch the first execution device.

[0064] The second scenario is that the activity level of the expert submodule... (Low activity), and the currently deployed first execution device is the first preset device CPU. At this time, the low-activity expert submodule is deployed on a general-purpose computing device, which will not occupy the GPU's valuable video memory and computing power resources, and can meet its low-frequency inference needs. The resource configuration is reasonable, so there is no need to switch the first execution device.

[0065] In some embodiments, the electronic device obtains the current activity level of the expert submodule. The first execution device (GPU or CPU) currently deployed, and a preset first threshold is invoked. Second threshold Then, condition judgment is performed to determine the first expert submodule that needs to be migrated.

[0066] Specifically, the first scenario is that the activity level of the expert submodule... (Highly active module), but the first execution device currently deployed is the first preset device CPU. At this time, the highly active expert submodule occupies CPU resources for a long time, and its high-frequency calls will lead to insufficient CPU computing power and significantly increased inference latency. Meanwhile, the high parallel computing capability of the GPU is not fully utilized, and the deployment status does not match the module requirements. Therefore, this expert submodule can be identified as the first expert submodule.

[0067] The second scenario: If the activity level of the expert submodule... (Low-activity module), but the first execution device currently deployed is the second preset device GPU. At this time, the low-activity module occupies the GPU's video memory and computing power resources, which may cause high-frequency calling modules to be unable to execute in time due to insufficient GPU resources, resulting in resource waste and scheduling congestion. The deployment status does not match the resource optimization requirements. Therefore, this expert submodule can be identified as the first expert submodule.

[0068] S203. Perform migration processing on the first expert submodule to obtain the second hybrid expert model.

[0069] The second hybrid expert model can be a version of the model generated by adjusting the deployment positions of each first expert submodule in the first hybrid expert model.

[0070] In some embodiments, the electronic device can migrate the model data of the corresponding first expert submodule from the current device to the device to be switched. After the migration operation is completed, each expert submodule constitutes a second hybrid expert model in the new deployment location.

[0071] exist Figure 2In the illustrated embodiment, when model processing is required, the activity level of each expert submodule in the first hybrid expert model and the first execution device for deploying each expert submodule are obtained. The activity level indicates the calling frequency of the expert submodule. Based on the activity level and the first execution device corresponding to multiple expert submodules, a first expert submodule is determined among the expert submodules in the first hybrid expert model. The activity level of the first expert submodule is greater than or equal to a first threshold, and the first execution device is a first preset device; or, the activity level of the first expert submodule is less than a second threshold, and the first execution device is a second preset device. The first expert submodule is then subjected to migration processing to obtain the second hybrid expert model. In this way, the optimal deployment location and compression state can be dynamically determined based on the actual activity level of each expert submodule during runtime and the real-time status of device resources. This achieves the matching of computing resource supply and the real-time load demand of modules, and can adaptively adjust the resource allocation between high-frequency calling modules and low-frequency calling modules. This effectively avoids the resource mismatch problem caused by the rigidity of traditional static configuration methods, and realizes dynamic optimization of the model in a heterogeneous environment based on load and resource changes. This improves resource utilization and deployment flexibility while enhancing model processing efficiency.

[0072] Based on any of the above embodiments, the method by which the electronic device performs migration processing on the first expert submodule to obtain the second hybrid expert model will be described in detail.

[0073] Figure 3 This is a schematic diagram illustrating a method for determining a second hybrid expert model provided in an embodiment of this application. Please refer to... Figure 3 ,include:

[0074] S301. Determine the first constraint condition.

[0075] The first constraint is used to determine whether migration processing of each first expert submodule is feasible.

[0076] In some embodiments, the electronic device may determine the first constraint condition based on the following implementation: obtaining a first time list of migration processing for each first expert submodule in the first hybrid expert model within a historical period; determining a first duration of migration processing for each first expert submodule, and determining a second duration in the first time list; and determining the first duration as less than the second duration as the first constraint condition.

[0077] The first time list is used to store historical time records of the actual time consumed by performing migration and other processing operations on each first expert submodule in the first hybrid expert model within a historical period.

[0078] The first duration can be defined as the total time required to perform the corresponding migration process for the first expert submodule. The relevant formula can be expressed as:

[0079]

[0080] in, Indicates the first duration. This indicates that the expert submodule will be used. Data transfer time required to migrate from the current device (first execution device) to the preset device.

[0081] The second duration can be the estimated time saved after performing the migration operation. In other words, the second duration can be understood as the expected latency saving time after the migration. For example, the latency reduction due to memory release, bandwidth reduction, and faster activation can be expressed as... .

[0082] Only when The equipment migration operation is only actually carried out when the company is established.

[0083] S302. Based on the first constraint, perform transfer processing on each first expert sub-module in the first hybrid expert model to obtain the second hybrid expert model.

[0084] In some embodiments, the electronic device may perform migration processing on each first expert sub-module in the first hybrid expert model according to a first constraint condition to obtain a second hybrid expert model based on the following implementation: determining a second execution device among a plurality of first execution devices; determining the device type of the second execution device, wherein the device type is a first preset device or a second preset device; and performing migration processing on each first expert sub-module in the first hybrid expert model according to the device type and the first constraint condition to obtain a second hybrid expert model.

[0085] The second execution device is the execution device where the first expert submodule is located.

[0086] In some embodiments, after determining at least one first expert submodule based on activity level and a first threshold, and a first execution device and a preset device, the execution device corresponding to these first expert submodules can be determined as a second execution device. Furthermore, under the constraint of migration time, the second execution device can be switched to an execution device that is compatible with the current first expert submodule. In this way, the first expert submodule can be loaded and perform inference tasks on the compatible execution device, thereby improving the deployment efficiency of model processing.

[0087] In some embodiments, the electronic device may perform migration processing on each first expert sub-module in the first hybrid expert model according to the device type and the first constraint condition to obtain a second hybrid expert model: when the device type of the second execution device is a first preset device and the first duration is less than the second duration, the first expert sub-module is migrated to the second preset device to obtain the second hybrid expert model; when the device type of the second execution device is a second preset device and the first duration is less than the second duration, the first expert sub-module is migrated to the first preset device to obtain the second hybrid expert model.

[0088] In some embodiments, the electronic device first determines whether the constraints are met. If the constraints are met, a device migration operation is performed. Specifically, necessary storage resources are allocated on the device to be switched. Then, the model weights and current running context of the first expert submodule are asynchronously transmitted from its currently deployed first execution device to the device to be switched via a system bus (e.g., a peripheral component high-speed interconnect bus) or network. Before the data transmission and loading verification are completed, requests are still processed by the original module on the first execution device. After the new module is ready, the routing table is updated through atomic operations to direct subsequent requests to the new instance on the device to be switched. After the migration of such modules is completed, the model runs under the new device layout, thus obtaining the second hybrid expert model.

[0089] In some embodiments, if the device type of the second execution device is CPU and the migration processing time meets the first constraint condition, the second execution device can be switched, that is, the corresponding first expert submodule can be migrated to GPU. If the device type of the second execution device is GPU and the migration processing time meets the first constraint condition, the second execution device can be switched, that is, the corresponding first expert submodule can be migrated to CPU.

[0090] Optionally, this solution's device migration scenarios are not limited to switching between a single GPU and a single CPU. It can also support flexible migration configurations for multi-heterogeneous device clusters (including multiple GPUs and multiple CPUs), and a device priority mechanism is introduced during the migration process to achieve optimal resource allocation. When a device switch is required for a specific first-expert submodule, an optimal device can be dynamically selected as the migration target.

[0091] Specifically, electronic devices can pre-configure priority weights for available heterogeneous devices (e.g., GPU1, GPU2, CPU1, CPU2, etc.). When determining the type of device to be switched, the electronic device filters the suitable device type based on the activity of the first expert sub-module (high-activity modules are preferentially matched with GPUs, and low-activity modules can be matched with CPUs). Then, it selects the device with the highest priority and that meets the resource constraints from the suitable type of devices as the device to be switched.

[0092] For example, if the highly active first expert submodule is currently deployed on GPU2 (lower priority), then all GPUs are screened first, and GPU1 with the highest priority and sufficient available video memory is selected as the migration target. If a GPU resource is scarce, the less active first expert submodule can select CPU2, the highest priority CPU, from among multiple CPUs for migration. This ensures that the first expert submodule can be matched with the optimal hardware resources, further improving the resource utilization of the heterogeneous device cluster and the overall inference efficiency of the model.

[0093] exist Figure 3 In the illustrated embodiment, since a corresponding migration processing strategy can be generated based on the activity level of each expert submodule and the current device type, and its feasibility is determined by comparing the first estimated migration duration with the second duration determined based on historical data, the migration operation is executed under the premise of meeting the constraints. Therefore, the resource waste and performance fluctuation caused by static configuration and blind execution of optimization operations can be effectively avoided. This solves the technical problem in related technologies where the lack of time estimation for model processing operations leads to low execution efficiency and affects the stability of online services. While realizing the dynamic determination of model processing strategies, it significantly improves the reliability and execution efficiency of the optimization operation itself, providing a reliable technical guarantee for the efficient and stable deployment and continuous optimization of models in heterogeneous environments.

[0094] Based on any of the above embodiments, the method by which the electronic device compresses each expert sub-module to obtain a third hybrid expert model in the above processing method will be described in detail.

[0095] Figure 4 This is a schematic diagram illustrating a method for determining a third hybrid expert model provided in an embodiment of this application. Please refer to... Figure 4 ,include:

[0096] S401. Obtain resource indicator data for each expert submodule in the second hybrid expert model.

[0097] Resource metrics data are used to reflect the real-time resource usage and availability status of the first execution device. These resource metrics data include GPU memory usage, CPU memory usage, and bandwidth utilization of the Peripheral Component Interconnect Express (PCIe) / Compute Express Link (CXL) bus.

[0098] In some embodiments, for the GPU as the first execution device, its video memory usage and total video memory can be obtained by querying the driver, and the video memory utilization rate can be calculated; for the CPU, its memory usage and total memory can be obtained, and the memory utilization rate can be calculated; for the PCIe or CXL bus of the connected device, its current data transfer rate and bandwidth are monitored, and the bandwidth utilization rate B is calculated. These real-time sampled and calculated utilization rates constitute resource indicator data, providing hardware environment status input for subsequent scheduling decisions.

[0099] Among them, GPU memory usage CPU memory utilization The relevant formula for the bandwidth utilization B of the PCIe / CXL bus can be expressed as:

[0100]

[0101]

[0102]

[0103] In some embodiments, after acquiring resource index data of each expert submodule in the second hybrid expert model, the electronic device must meet preset resource constraints when determining whether to perform compression processing based on these data. This ensures that the storage usage of the expert submodule after deployment and compression does not exceed the available memory of the corresponding component, avoiding inference interruption issues caused by memory overflow or insufficient resources. The relevant formula for this resource constraint can be expressed as:

[0104]

[0105]

[0106] Among them, Represents the i-th expert submodule In quantization bit width Pruning ratio Storage usage below; This indicates the available memory of the GPU. This indicates the available memory for the CPU.

[0107] S402. Based on resource indicator data and activity level, determine the second expert submodule among the expert submodules of the second hybrid expert model.

[0108] Specifically, the activity level of the second expert submodule is less than the second threshold, or the resource indicator data of the second expert submodule is greater than the third threshold.

[0109] The third threshold can be a pre-defined resource utilization threshold used to determine whether the current hardware resources are under strain or overloaded. For example, the third threshold could represent the upper limit of GPU memory utilization (e.g., 90%), CPU memory utilization, or system bus bandwidth utilization. When resource metrics exceed this threshold, it indicates that the system resources are facing a bottleneck.

[0110] In some embodiments, the electronic device obtains the current activity level of each expert submodule in the second hybrid expert model. Resource metrics for each primary execution device (GPU memory utilization) (Bandwidth utilization B), and a preset second threshold. The third threshold corresponding to various resources (e.g., The third threshold The third threshold of B Then, based on OR logic, conditional judgments are performed to determine the second expert submodule that needs to be compressed.

[0111] Specifically, the first scenario is that the activity level of the expert submodule... (Low activity). Since the low activity of expert submodules is called infrequently, it has little impact on the overall inference performance of the model. Even if it is compressed (e.g., quantized to a low bit width, pruning some redundant weights), it will not significantly reduce the effect. At the same time, it can effectively reduce the storage resources occupied by them, freeing up more hardware resources for expert submodules that are called frequently. Therefore, the corresponding expert submodule can be identified as the second expert submodule.

[0112] The second scenario is that if any resource metric exceeds the corresponding third threshold, it indicates that the resources of the current first execution device (GPU / CPU) are already at a bottleneck. Continuing to maintain the resource usage of the existing expert submodule will lead to increased inference latency, stuttering, or even task interruption. In this case, regardless of the activity level of the expert submodule, its storage usage and resource consumption need to be reduced through compression to alleviate resource pressure. Therefore, the electronic device can designate the corresponding expert submodule as the second expert submodule. The relevant formula for the judgment condition can be expressed as:

[0113]

[0114] in, This represents a critical value for GPU memory utilization (e.g., 90%). This represents the critical value for bandwidth utilization. The low threshold representing activity level.

[0115] In some embodiments, when the activity level of the expert submodule is greater than or equal to a first threshold and the first execution device is a second preset device, or when the activity level of the expert submodule is not less than a second threshold and the resource indicator data is less than a third threshold, it is not necessary to compress the expert submodule.

[0116] In some embodiments, the electronic device obtains the activity level of the expert submodule. The currently deployed first execution device type, resource indicator data, and preset first threshold. Second threshold And the third threshold, and based on the AND logic, conditional judgment is performed to determine that the expert submodule does not need to be compressed.

[0117] Specifically, the first scenario is that the activity level of the expert submodule... (High activity level), and the first execution device currently deployed is the second preset device, the GPU. The high-activity module is the core of model inference and plays a key role in inference performance. The GPU has sufficient high computing power resources to handle its original resource usage without sacrificing accuracy or speed through compression. If forced compression is performed, it may lead to a decrease in module calculation accuracy and a reverse increase in inference latency, which will affect the overall performance. Therefore, the electronic device can determine that it is not necessary to compress the expert sub-module.

[0118] The second scenario is that the activity level of the expert submodule... (Low activity) and all resource metrics are below the corresponding third threshold, indicating that current hardware resources are sufficient and there is no resource bottleneck pressure. At this time, the resource consumption of the low-activity module will not affect the normal operation of other modules, and the compression operation itself consumes computing resources and increases system overhead. Therefore, no additional compression processing is required, and the electronic device can determine that it does not need to compress the expert submodule.

[0119] S403. Compress the second expert submodule to obtain the third hybrid expert model.

[0120] The third hybrid expert model can be a version of the model generated by adjusting the internal structure of each expert sub-module in the second hybrid expert model.

[0121] In some embodiments, the electronic device may compress the second expert submodule to obtain a third hybrid expert model based on the following implementation: obtaining a second time list of compression processing of each second expert submodule in the second hybrid expert model within a historical period; determining a third duration for compression processing of each second expert submodule, and determining a fourth duration in the second time list; determining a second constraint condition if the third duration is less than the fourth duration; and compressing the second expert submodule according to the second constraint condition to obtain the third hybrid expert model.

[0122] The second time list is used to store historical time records of the actual time consumed by performing compression and other processing operations on each expert submodule in the second hybrid expert model within a historical period.

[0123] The third duration can be the total time required to perform the corresponding compression processing on the second expert submodule. The relevant formula can be expressed as:

[0124]

[0125] in, Indicates the third duration. This indicates the quantization bit width of the expert submodule before compression. This indicates the quantization bit width of the expert submodule after compression. This indicates the pruning ratio of the expert submodule before compression. This indicates the pruning ratio of the expert submodule after compression. This indicates the time required to perform a quantization / pruning switch (including loading new quantization weights, updating the execution engine state, etc.).

[0126] The fourth duration can be the estimated time saved after performing this compression process. In other words, the fourth duration can be understood as the estimated latency saving time after compression. For example, the latency reduction due to memory release, bandwidth reduction, and faster activation can be expressed as... .

[0127] Only when Only when it is established does the quantization / pruning controller actually perform bit width / pruning state switching.

[0128] In some embodiments, the electronic device may compress the second expert submodule according to the second constraint condition to obtain a third hybrid expert model based on the following implementation: determining a first compression parameter and a second compression parameter for the second expert submodule, wherein the first compression parameter is used to compress the bit width of each weight in the second expert submodule, and the second compression parameter is used to clear multiple weights of the second expert submodule; when the third duration is less than the fourth duration, compressing the second expert submodule according to the first compression parameter and the second compression parameter to obtain a third hybrid expert model.

[0129] The first compression parameter is used to compress the bit width of each weight in the second expert submodule. The first compression parameter can be the precision bit width used when performing quantization compression on the second expert submodule, for example, converting the weights from 16-bit floating-point numbers (FP16) to 8-bit integers (INT8).

[0130] The second compression parameter is used to remove multiple weights from the second expert submodule. This second compression parameter can be the weight removal ratio used when performing pruning and compression on the second expert submodule; for example, removing 50% of the weights with the smallest absolute value. It determines the sparsity of the model structure.

[0131] In this embodiment of the application, the formula related to compression can be expressed as:

[0132] ,

[0133] in, Indicates the high bit width (e.g., 8 bits or 4 bits). Indicates the low bit width (e.g., 2 bits). This indicates a high pruning percentage (e.g., 50%~70%). This indicates a low pruning percentage (e.g., 0%~10%). For example, quantization (i.e., the first compression parameter) ) and pruning parameters (i.e., the second compression parameter) Typical values ​​for ) can be: =8 (or 4 digits) =2; =0%~10%, =50%~70%.

[0134] In some embodiments, the electronic device acquires a first compression parameter and a second compression parameter, and determines whether to perform a compression operation based on a second constraint condition. Quantization processing is performed, that is, based on the first compression parameter, the module's weight parameters are converted into quantized data of the corresponding bit width. During the conversion, core weight information is preserved through linear mapping. Structured pruning processing is also performed, that is, based on the second compression parameter, weights with low contribution in the second expert submodule are selected, redundant weights are removed proportionally, and the module parameter structure is reconstructed. This ensures that the logic of the second expert submodule is not affected after pruning. Furthermore, the pruning ratio can be appropriately adjusted during the compression process to meet accuracy requirements.

[0135] After completing the above operations, save the compressed parameter data and replace the original parameters of the second expert submodule. Once all the second expert submodules requiring compression have completed the above operations, the electronic device integrates the compressed model parameters to form a resource-optimized third hybrid expert model.

[0136] Optionally, if it is necessary to simultaneously migrate and compress each expert module of the first hybrid expert model, that is, to migrate the expert sub-modules to the device to be switched and to compress the model according to the first compression parameter and the second compression parameter, the electronic device can perform a combined optimization operation.

[0137] At this point, it's necessary not only to determine the device to be migrated, but also to determine the first compression parameter (quantization bit width) and the second compression parameter (pruning ratio). Specifically, if the migration is triggered due to low activity of the expert submodule (below the first threshold) or any resource metric data exceeding the third threshold, and it is currently deployed on a GPU, then the device to be migrated is determined to be a CPU, and the first compression parameter can be determined from... (For example, downsizing from 8-bit to 2-bit), the second compression parameter can be derived from... (For example, increasing pruning from 10% to 50%).

[0138] This requires determining whether the migration and compression operations meet the duration constraints. The relevant formula can be expressed as:

[0139]

[0140] in, This represents the total time spent switching components between each expert submodule and compressing each expert submodule.

[0141] Only when Only upon establishment is the actual device migration and bit-width / pruning state switching performed. The estimated time saved by performing this migration and compression process can be understood as the estimated time saved by delay.

[0142] In some embodiments, the electronic device performs constraint condition judgment on the processing of each expert submodule one by one. If the constraint condition is met, the corresponding operation is executed (if only the component is switched, the module asynchronous migration is completed; if only compression is performed, quantization and pruning are performed; if switching and compression are performed, migration and compression are performed); if the preset constraint condition is not met (the processing time is higher than the delay saving), the processing operation is not executed for the time being, and the current state of the module is maintained.

[0143] After each expert submodule completes its condition judgment and corresponding processing, the electronic device integrates the optimized states of each expert submodule (including the final deployment device and compressed parameters), updates the model's deployment configuration and parameter storage structure, and forms a hybrid expert model with more reasonable resource consumption and better processing efficiency.

[0144] Optionally, a compressed version of the module can be generated at the source or management node based on compression parameters (significantly reducing the data volume). This smaller, compressed model data is then transmitted to the second component. This not only reduces the data transfer volume during the migration itself but also ensures that the expert submodules obtained on the migrated device are in a compressed state. After the compressed module is loaded and verified on the migrated device, the routing is updated. By integrating compression as part of the migration process, rather than executing two separate operations sequentially, the total time consumption can be effectively controlled, and constraints can be met. After each expert submodule completes its combination operation, the model runs in the compressed form under the new device layout, resulting in the optimized hybrid expert model.

[0145] In some embodiments, once the switch is complete, the scheduler updates the deployment status of the expert submodule, and the execution engine loads and runs according to the new settings. Subsequently, execution data (actual latency, activation count, device load) is fed back to the monitoring module, forming a closed loop. Through the above mechanism, this system can dynamically switch the submodule to a more resource-efficient execution mode or device when the activity of the expert submodule decreases or device resources are scarce, avoiding problems such as GPU memory overflow, bandwidth congestion, and increased latency.

[0146] exist Figure 4 In the illustrated embodiment, since the model compression can be dynamically performed based on the resource index data of each expert submodule in the second hybrid expert model and the constraint that the third duration is less than the fourth duration, the resource waste caused by the long execution time of traditional model optimization methods can be effectively avoided. This solves the technical problem of affecting the real-time performance and stability of model services due to blind execution. While achieving dynamic model optimization, the efficiency and controllability of the optimization process itself are ensured. While ensuring the feasibility and effectiveness of processing operations, the model deployment efficiency is significantly improved, providing a reliable technical guarantee for the flexible adaptation and efficient operation of the model on heterogeneous hardware.

[0147] Figure 5 A flowchart illustrating another model processing method provided in this application embodiment is shown below. Figure 5 As shown, embodiments of this application provide a model processing method, which is described in detail below:

[0148] This system mainly consists of five sub-modules: activity monitoring module, resource status awareness module, device-sub-module scheduler, quantization / pruning controller, and execution engine. Among them, the activity monitoring module is responsible for collecting and maintaining the activation data of each expert sub-module 1-N (first expert sub-module 1, second expert sub-module 2, ..., Nth expert sub-module N) in real time, that is, the number of activations in the most recent N batches (to calculate activity).

[0149] The resource status awareness module is responsible for monitoring the resource status of each device (the first execution device currently deployed) in the current heterogeneous computing environment, i.e., resource indicator data, specifically including GPU memory utilization, CPU memory utilization, and PCIe / CXL bandwidth utilization.

[0150] The device-submodule scheduler is used to determine the deployment device (GPU or CPU), quantization bit width, and pruning ratio for each expert submodule based on the input data (activity and resource index data) from the activity monitoring module and the resource status awareness module.

[0151] The quantization / pruning controller is used to execute the quantization downgrading and pruning scheme of the weights of each expert submodule according to the scheduler's decision, and is responsible for detecting and triggering state switching during operation (e.g., downgrading from 8-bit to 2-bit, or migrating from GPU to CPU).

[0152] The execution engine is responsible for loading expert sub-modules on the specified device, executing subsequent tasks according to the specified quantization and pruning patterns, and feeding back the running status to the activity monitoring and resource status awareness modules.

[0153] like Figure 5The entire system architecture, from left to right (or counter-clockwise), presents five major modules: activity monitoring, resource awareness, scheduling decision-making, quantization / pruning control, and the execution engine, forming a closed-loop control process. First, the activity monitoring module continuously receives the activation frequency (number of calls) of each expert sub-module from the execution engine and generates activity metrics for each sub-module based on this. Simultaneously, the resource awareness module acquires resource metrics such as current GPU memory usage, CPU memory usage, PCIe / CXL bandwidth, and device load. Then, the scheduler combines activity and resource conditions to make device allocation (GPU or CPU) and compression strategy (quantization bit width, pruning ratio) decisions. This decision is executed by the quantization / pruning controller, which performs bit width changes, prunes, or migrates devices for the expert sub-modules and instructs the execution engine to load the corresponding version model on the specified device. The execution engine runs tasks according to the scheduling results and simultaneously feeds back the activation status and resource usage during operation to the activity monitoring and resource awareness modules, forming a closed loop of monitoring, decision-making, execution, and feedback. Through this structure, each expert submodule can dynamically switch quantization / pruning strategies or deploy devices based on its own activity level and changes in the resource environment, so as to achieve efficient collaboration and inference optimization on heterogeneous platforms.

[0154] Below, in conjunction with Figure 5 The process of model processing is explained in detail.

[0155] The specific process is as follows: Step one is to activate data collection, that is, at the beginning of startup or inference, the execution engine module loads each expert submodule and starts the inference job. The activity monitoring module collects the activation status of each expert submodule in real time, including the number of times it is called in this batch, and stores this data in the activity cache.

[0156] Step two involves resource status monitoring. Simultaneously, the resource status awareness module monitors the resource status of each device in the system, including GPU memory usage, available CPU memory, PCIe / CXL bandwidth utilization, and device load rate. This module samples the current resource status at fixed intervals (e.g., every 10 seconds) and provides the scheduler with the latest resource overview.

[0157] Step three involves calculating activity metrics. The activity monitoring module updates the activity levels of each expert sub-module based on the latest collected data and historical cache, according to a pre-set decay rule (adjustment coefficient). For example, it reduces the proportion of historical data and increases the proportion of newly collected data to generate a set of metrics reflecting recent activity trends.

[0158] Step four involves generating scheduling decisions. The scheduler module receives activity data from the monitoring module and resource status (GPU memory, RAM, bandwidth) from the status awareness module. Based on preset activity thresholds (first and second thresholds) and resource constraints, the scheduler selects the execution device (GPU or CPU), quantization bit width, and pruning ratio for each expert submodule, forming a scheduling decision list (processing strategy).

[0159] Step five involves the execution of quantization and pruning control. After receiving the scheduling decision, the quantization / pruning controller module performs the following operations on the designated expert submodule: if migration to a new device is required, model weight migration is performed; the quantization bit width and pruning degree are modified according to the decision; and the required model version is prepared. The controller ensures that the model version matches the device environment.

[0160] Step six involves model loading and execution. The execution engine loads the quantized / pruned versions of each expert submodule onto a specified device (e.g., GPU1, GPU2, or CPU) according to the scheduling results. Then, it executes the inference task, routing the input data to each expert submodule and returning the output results. Simultaneously, the execution engine records a summary of the execution process for each expert submodule, including latency, activation counts, and device load changes.

[0161] Step seven is the feedback data return, whereby the execution engine returns the activation count and actual device resource usage (video memory usage, bandwidth transmission status) collected during the operation to the activity monitoring module and the resource status perception module, and the two modules update their parameters and resource status data respectively.

[0162] Step eight involves closed-loop iteration, where the entire process forms a closed-loop mechanism of monitoring, decision-making, execution, feedback, and re-monitoring. As inference batches iterate continuously, each batch updates its activity and resource status, the scheduler module regenerates the schedule, the quantization / pruning controller may switch schemes, and the execution engine continues to execute and provide feedback. In this way, the deployment and compression strategies of sub-modules can be dynamically adjusted in heterogeneous device environments to improve overall deployment efficiency.

[0163] exist Figure 5In the illustrated embodiment, a device-submodule allocation scheduler and quantization / pruning state switching mechanism based on the activity and resource status of expert submodules are constructed. This dynamically determines whether each expert submodule should be deployed on a GPU or CPU based on its activity and current hardware resources, and selects appropriate quantization bit widths and pruning ratios. Furthermore, during operation, if a resource bottleneck or a decrease in expert activity is detected, quantization degradation or device migration is triggered. By integrating structured / semi-structured quantization, pruning, and device collaboration in a heterogeneous device environment, this method achieves reduced memory load, lower bandwidth overhead, and significant optimization of throughput and latency. In this approach, in a heterogeneous device environment, the execution device is dynamically selected and its quantization / pruning scheme is adjusted based on the activity frequency / importance index of the expert submodule and combined with device resources, thereby improving the deployment efficiency and resource utilization of the MoE model from a system-model collaboration perspective.

[0164] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods according to the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method.

[0165] Figure 6 This is a schematic diagram of a model processing device provided in an embodiment of this application. Figure 6 As shown, embodiments of this application also provide a model processing apparatus 60, including: an acquisition module 61, a determination module 62, and a processing module 63, wherein:

[0166] The acquisition module 61 is used to acquire the activity level of each expert sub-module in the first hybrid expert model and the first execution device for deploying each expert sub-module. The activity level is used to indicate the calling frequency of the expert sub-module.

[0167] The determining module 62 is used to determine the first expert submodule among the expert submodules of the first hybrid expert model based on the activity levels of the multiple expert submodules and the first execution device. The activity level of the first expert submodule is greater than or equal to a first threshold, and the first execution device is a first preset device; or, the activity level of the first expert submodule is less than a second threshold, and the first execution device is a second preset device.

[0168] Processing module 63 is used to perform migration processing on the first expert submodule to obtain the second hybrid expert model.

[0169] For a description of the features in the embodiment corresponding to the model processing device, please refer to the relevant description in the embodiment corresponding to the model processing method, which will not be repeated here.

[0170] In one possible implementation, the processing module 63 is specifically used for:

[0171] Determine the first constraint condition, which is used to determine whether it is feasible to perform migration processing on each first expert submodule.

[0172] Based on the first constraint, the first expert sub-modules in the first hybrid expert model are transferred to obtain the second hybrid expert model.

[0173] In one possible implementation, the processing module 63 is specifically used for:

[0174] Obtain the first time list of migration processing performed on each first expert submodule in the first hybrid expert model within a historical period;

[0175] Determine the first duration for migration processing of each first expert submodule, and determine the second duration from the first duration list;

[0176] The first duration being less than the second duration is defined as the first constraint condition.

[0177] In one possible implementation, the processing module 63 is specifically used for:

[0178] Among multiple first execution devices, a second execution device is determined, which is the execution device where the first expert submodule is located;

[0179] Determine the device type of the second execution device, which is either a first preset device or a second preset device;

[0180] Based on the equipment type and the first constraint, the first expert sub-modules in the first hybrid expert model are transferred to obtain the second hybrid expert model.

[0181] In one possible implementation, the processing module 63 is specifically used for:

[0182] When the device type of the second execution device is the first preset device and the first duration is less than the second duration, the first expert submodule is migrated to the second preset device to obtain the second hybrid expert model;

[0183] When the device type of the second execution device is the second preset device and the first duration is less than the second duration, the first expert submodule is migrated to the first preset device to obtain the second hybrid expert model.

[0184] In one possible implementation, the acquisition module 61 is specifically used for:

[0185] Obtain the number of task executions, the number of calls to each expert submodule, the initial activity level, and the adjustment coefficient of the first hybrid expert model within a historical period.

[0186] For any expert submodule, the activity level is obtained by weighting the number of task executions, the number of calls, and the initial activity level according to the adjustment coefficient.

[0187] In one possible implementation, the device further includes a compression module, which is specifically used for:

[0188] Obtain resource indicator data for each expert submodule in the second hybrid expert model;

[0189] Based on resource indicator data and activity level, a second expert submodule is determined among the expert submodules of the second hybrid expert model. The activity level of the second expert submodule is less than the second threshold, or the resource indicator data of the second expert submodule is greater than the third threshold.

[0190] The second expert submodule is compressed to obtain the third hybrid expert model.

[0191] In one possible implementation, the compression module is specifically used for:

[0192] Obtain a second time list of each second expert sub-module in the second hybrid expert model within a historical period, after compression processing.

[0193] Determine the third duration for compression processing of each second expert submodule, and determine the fourth duration from the second time list;

[0194] The third duration being less than the fourth duration is defined as the second constraint condition.

[0195] Based on the second constraint, the second expert submodule is compressed to obtain the third hybrid expert model.

[0196] In one possible implementation, the compression module is specifically used for:

[0197] Determine the first compression parameter and the second compression parameter of the second expert submodule. The first compression parameter is used to compress the bit width of each weight in the second expert submodule, and the second compression parameter is used to clear multiple weights in the second expert submodule.

[0198] When the third duration is less than the fourth duration, the second expert submodule is compressed according to the first compression parameter and the second compression parameter to obtain the third hybrid expert model.

[0199] Figure 7 A schematic diagram of the structure of the electronic device provided in this application. Figure 7 As shown, the electronic device 70 provided in this embodiment includes at least one processor 71 and a memory 72. Optionally, the electronic device 70 further includes a communication component 73. The processor 71, memory 72, and communication component 73 are connected via a bus.

[0200] In a specific implementation, at least one processor 71 executes computer execution instructions stored in memory 72, causing at least one processor 71 to execute the above-described model processing method embodiment.

[0201] The specific implementation process of processor 71 can be found in the above method embodiments, and its implementation principle and technical effect are similar. It will not be repeated here.

[0202] In the above embodiments, it should be understood that the processor can be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), etc. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the method disclosed in the application can be directly manifested as being executed by a hardware processor, or executed by a combination of hardware and software modules within the processor.

[0203] The memory may include random access memory (RAM) and may also include non-volatile memory (NVM), such as at least one disk storage device.

[0204] The bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus, etc. Buses can be categorized as address buses, data buses, control buses, etc. For ease of illustration, the buses shown in the accompanying drawings are not limited to a single bus or a single type of bus.

[0205] Embodiments of this application also provide a computer-readable storage medium storing a computer program, wherein the computer program is configured to execute the steps in any of the above-described model processing method embodiments at runtime.

[0206] In one exemplary embodiment, the aforementioned computer-readable storage medium may include, but is not limited to, various media capable of storing computer programs, such as a USB flash drive, read-only memory (ROM), random access memory (RAM), portable hard disk, magnetic disk, or optical disk.

[0207] Embodiments of this application also provide a computer program product, including a computer program that, when executed by a processor, implements the steps in any of the above-described model processing method embodiments.

[0208] Embodiments of this application also provide another computer program product, including a non-volatile computer-readable storage medium storing a computer program, which, when executed by a processor, implements the steps in any of the above-described model processing method embodiments.

[0209] Any of the components, modules, units, parts, methods, and operations described herein can be implemented using software, firmware, hardware (e.g., fixed logic circuitry), manual processing, or any combination thereof. Alternatively or additionally, any functionality described herein can be executed at least in part by one or more hardware logic components, such as, but not limited to, a central processing unit (CPU), a field-programmable gate array (FPGA), an application-specific integrated circuit (ASIC), an application-specific standard product (ASSP), a system-on-a-chip (SoC), a complex programmable logic device (CPLD), a microprocessor (MCU), etc. The terms "system," "computing device," or "apparatus" as used herein encompass various means, devices, and machines for processing data, including, for example, one or more programmable processors, computers, SoCs, or combinations thereof. The apparatus may also include code that creates an execution environment for the computer program in question, such as code constituting processor firmware, a protocol stack, a database management system, an operating system, a cross-platform runtime environment, a virtual machine, or one or more combinations thereof. The aforementioned computer program (also known as a program, software, software application, app, script, or code) can be written in any form of programming language, including compiled or interpreted languages, declarative or procedural languages, and can be deployed in any form, including as a standalone program or as a module, component, subroutine, object, or other unit suitable for a computing environment.

[0210] Those skilled in the art will further recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for specific applications, but such implementations should not be considered beyond the scope of this application.

[0211] The above provides a detailed description of a model processing method provided in this application. Specific examples have been used to illustrate the principles and implementation methods of this application. The descriptions of the embodiments above are merely for the purpose of helping to understand the method and its core ideas. It should be noted that those skilled in the art can make various improvements and modifications to this application without departing from its principles, and these improvements and modifications also fall within the protection scope of this application.

Claims

1. A model processing method characterized by comprising: The method comprises the following steps: obtaining the activity of each expert sub-module in the first mixed expert model and the first execution device deploying the expert sub-module, wherein the activity is used to indicate the calling frequency of the expert sub-module; determining a first expert sub-module from each expert sub-module of the first mixed expert model according to the activity corresponding to each expert sub-module and the first execution device, wherein the activity of the first expert sub-module is greater than or equal to a first threshold value, and the first execution device is a first preset device, or the activity of the first expert sub-module is less than a second threshold value, and the first execution device is a second preset device; performing migration processing on the first expert sub-module to obtain a second mixed expert model.

2. The method of claim 1, wherein, The migration processing on the first expert sub-module to obtain a second mixed expert model comprises the following steps: determining a first constraint condition for judging whether the migration processing on each first expert sub-module is feasible; performing migration processing on each first expert sub-module of the first mixed expert model according to the first constraint condition to obtain the second mixed expert model.

3. The method of claim 2, wherein, The determination of the first constraint condition comprises the following steps: obtaining a first time list of the migration processing on each first expert sub-module of the first mixed expert model in a historical period; determining a first time length of the migration processing on each first expert sub-module and a second time length in the first time list; determining the first constraint condition when the first time length is less than the second time length.

4. The method of claim 3, wherein, The migration processing on each first expert sub-module of the first mixed expert model according to the first constraint condition to obtain the second mixed expert model comprises the following steps: determining a second execution device from a plurality of first execution devices, wherein the second execution device is the execution device where the first expert sub-module is located; determining the device type of the second execution device, wherein the device type is a first preset device or a second preset device; performing migration processing on each first expert sub-module of the first mixed expert model according to the device type and the first constraint condition to obtain the second mixed expert model.

5. The method of claim 4, wherein, The migration processing on each first expert sub-module of the first mixed expert model according to the device type and the first constraint condition to obtain the second mixed expert model comprises the following steps: when the device type of the second execution device is the first preset device and the first time length is less than the second time length, migrating the first expert sub-module to the second preset device to obtain the second mixed expert model; when the device type of the second execution device is the second preset device and the first time length is less than the second time length, migrating the first expert sub-module to the first preset device to obtain the second mixed expert model.

6. The method according to any one of claims 1 to 5, characterized in that, The obtaining of the activity of each expert sub-module in the first mixed expert model comprises the following steps: obtaining the task execution times of the first mixed expert model, the calling times of each expert sub-module, the initial activity and the adjustment coefficient in a historical period; For any expert submodule, the number of task executions, the number of calls, and the initial activity level are weighted according to the adjustment coefficient to obtain the activity level.

7. The method of claim 1, wherein, The method further includes: Obtain resource indicator data for each expert submodule in the second hybrid expert model; Based on the resource indicator data and the activity level, a second expert submodule is determined in each expert submodule of the second hybrid expert model. The activity level of the second expert submodule is less than the second threshold, or the resource indicator data of the second expert submodule is greater than the third threshold. The second expert submodule is compressed to obtain the third hybrid expert model.

8. The method of claim 7, wherein, The second expert submodule is compressed to obtain the third hybrid expert model, which includes: Obtain a second time list of each second expert submodule in the second hybrid expert model that has been compressed within a historical period; A third duration for compressing each of the second expert submodules is determined, and a fourth duration is determined from the second time list; The third duration being less than the fourth duration is determined as the second constraint condition; Based on the second constraint, the second expert submodule is compressed to obtain the third hybrid expert model.

9. The method of claim 8, wherein, Based on the second constraint, the second expert submodule is compressed to obtain the third hybrid expert model, which includes: A first compression parameter and a second compression parameter are determined for the second expert submodule. The first compression parameter is used to compress the bit width of each weight in the second expert submodule, and the second compression parameter is used to clear multiple weights in the second expert submodule. When the third duration is less than the fourth duration, the second expert submodule is compressed according to the first compression parameter and the second compression parameter to obtain the third hybrid expert model.

10. An electronic device, comprising: include: Memory, used to store computer programs; A processor for executing the computer program to implement the steps of the model processing method as described in any one of claims 1 to 9.