Mass inference and data flow optimization system for MOE architecture-oriented large model

By employing a collaborative architecture comprising a request access module, an environment awareness module, an expert routing engine, a resource scheduling module, and a dynamic optimization control module, the problems of inference latency fluctuations, uneven resource utilization, and limited throughput in large-scale MOE models in online education platforms are resolved, achieving stable response and resource optimization in high-concurrency scenarios.

CN120849141BActive Publication Date: 2025-11-21VIRTAI TECH BEIJING CO LTD

Patent Information

Application Number
CN202511365194.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-09-23
Publication Date
2025-11-21
Estimated Expiration
2045-09-23

AI Technical Summary

Technical Problem

In online education platforms with high concurrency and low latency requirements, when deploying large-scale MOE models for real-time batch inference, traditional inference frameworks are inefficient in managing the dynamic expert routing, large-scale weight loading and unloading, and high-throughput data streams unique to the MOE architecture. This results in large fluctuations in inference latency, uneven resource utilization, limited throughput, and poor adaptability to dynamic load.

Method used

The system adopts a collaborative architecture consisting of a request access module, an environment awareness module, an expert routing engine, a resource scheduling module, and a dynamic optimization control module. The dynamic optimization control module intelligently triggers optimization strategies based on routing conflict factors, and combines batch processing scale adjustment, resource instance scaling, and concurrency control optimization operations to achieve adaptive optimization of the system.

Benefits of technology

It achieves stable low-latency response and resource collaborative optimization in high-concurrency scenarios, adaptively balances throughput and latency, intelligently responds to traffic fluctuations, ensures service quality and system stability, and improves the responsiveness and resource utilization efficiency of online education platforms.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120849141B_ABST
    Figure CN120849141B_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of project management, in particular to a large model batch inference and data flow optimization system for MOE architecture. The present application sets a collaborative architecture of a request access module, an environment perception module, an expert routing engine, a resource scheduling module and a dynamic optimization control module, uses the request access module to extract the text length and the subject type to provide the basis for accurate routing, collects the GPU display memory, the I / O bandwidth and the request queue depth in real time through the environment perception module, comprehensively monitors the system load, activates the expert sub-network according to the request characteristics through the expert routing engine to avoid invalid calculation, manages the weight loading and resource allocation through the resource scheduling module to reduce the I / O bottleneck, and finally intelligently triggers the optimization strategy based on the routing conflict factor through the dynamic optimization control module to solve the problems of large inference delay fluctuation and unbalanced resource utilization mentioned in the background technology, and realize stable low-delay response and resource collaborative optimization in a high-concurrency scenario.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of project management technology, and more specifically, to a large-model batch inference and data flow optimization system for MOE architecture. Background Technology

[0002] With the widespread application of Large Language Models (LLMs) and Mixture-of-Experts (MOE) architectures in fields such as Natural Language Processing, efficient batch inference services have become crucial for their implementation. However, deploying large-scale MOE models for real-time batch inference in online education platforms that require high concurrency and low latency presents significant challenges.

[0003] When the platform needs to process tens of thousands of math, physics, chemistry, and other subject problems submitted by students simultaneously during peak periods, traditional inference frameworks are inefficient in managing the dynamic expert routing, large-scale weight loading and unloading, and high-throughput data streams unique to the MOE architecture, such as the frequent data exchange between GPU memory and CPU / SSD / NVMe.

[0004] Existing inference optimization techniques, such as static batch processing, simple pipeline parallelism, or expert parallelism, often fail to fully utilize the potential of the MOE architecture when handling batch inference tasks unique to K12 online education platforms, which involve short texts, high concurrency, strong real-time requirements, and relatively concentrated distribution of requested content. They may also have efficiency bottlenecks in data flow scheduling, resulting in overall service costs and response latency that are difficult to meet the requirements of large-scale commercial applications. Summary of the Invention

[0005] The purpose of this invention is to provide a large-scale model batch inference and data flow optimization system for MOE architecture, so as to solve the problems of large inference latency fluctuation, uneven resource utilization, limited throughput and poor dynamic load adaptability mentioned in the background art.

[0006] To achieve the above objectives, the present invention aims to provide a large-scale model batch inference and data flow optimization system for MOE architecture, the system comprising a request access module, an environment awareness module, an expert routing engine, a resource scheduling module, and a dynamic optimization control module;

[0007] The request access module is used to receive user request streams and extract request features, including at least text length and subject type;

[0008] The environment perception module is used to collect inference environment parameters in real time, including GPU memory usage, storage I / O bandwidth, and request queue depth.

[0009] The expert routing engine is used to dynamically activate the expert sub-network of the MOE model based on the request characteristics;

[0010] The resource scheduling module is used to manage the loading / unloading of weight data and the allocation of computing resources;

[0011] The dynamic optimization control module receives data from the environment perception module and the access request module, and analyzes and determines whether an optimization strategy needs to be activated. Specifically:

[0012] Calculate the routing conflict factor based on the GPU memory limit and average text length;

[0013] A preset routing conflict factor threshold is set, and the routing conflict factor is compared with the routing conflict factor threshold. If the routing conflict factor exceeds the threshold, a dynamic optimization operation set is activated, which includes batch size adjustment operation, resource instance scaling operation, and concurrency control optimization operation.

[0014] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0015] 1. This invention achieves dynamic optimization of the entire inference chain of the MOE model by setting up a collaborative architecture of a request access module, an environment perception module, an expert routing engine, a resource scheduling module, and a dynamic optimization control module. The request access module extracts text length and subject type to provide a basis for accurate routing. The environment perception module collects GPU memory, I / O bandwidth, and request queue depth in real time to comprehensively monitor system load. At the same time, the expert routing engine activates expert sub-networks according to request characteristics to avoid invalid calculations. The resource scheduling module manages weight loading and resource allocation to reduce I / O bottlenecks. Finally, the dynamic optimization control module intelligently triggers optimization strategies based on routing conflict factors, solving the problems of large fluctuations in inference latency and uneven resource utilization mentioned in the background technology, and achieving stable low-latency response and resource collaborative optimization in high-concurrency scenarios.

[0016] 2. This invention achieves an adaptive balance between throughput and latency by setting up a batch processing scale adjustment operation in a dynamic optimization control module. It monitors inference latency in real time, calculates the batch processing adjustment coefficient by combining the target latency and the maximum I / O bandwidth, and dynamically scales the batch processing scale based on the comparison between the batch processing adjustment coefficient and a preset threshold. This solves the throughput limitation problem mentioned in the background technology and maximizes system throughput efficiency while ensuring real-time performance.

[0017] 3. This invention achieves intelligent response to traffic fluctuations by setting resource instance scaling operations. It counts the number of requests in time windows and calculates the load fluctuation factor of adjacent windows. When the fluctuation factor exceeds the threshold, it elastically scales up and down the number of instances proportionally. This solves the problem of poor dynamic load adaptability mentioned in the background technology and significantly improves the system's response capability to drastic changes in request volume during breaks or evening peaks in online education scenarios.

[0018] 4. This invention optimizes operations by setting concurrency control, achieving a dual guarantee of service quality and system stability. It counts the number of requests and successful responses within a time window, calculates the dynamic value of service quality, and compares it with the service quality threshold. If the value is higher than the threshold, the concurrency limit is increased to handle more requests; if the value is lower than the threshold, the concurrency limit is decreased to avoid system overload. This solves the problem of unstable response caused by high concurrency and ensures the success rate and service quality of student requests during peak periods.

[0019] 5. This invention achieves efficient decision-making under resource conflicts by setting a priority sorting mechanism for multiple optimization operations. When batch processing adjustment, resource scaling, and concurrency control need to be executed simultaneously, the priority of each operation is calculated, sorted from highest to lowest priority, and executed in sequence to ensure that core optimizations take effect first. This solves the pain point of difficult resource collaborative optimization, and scientifically schedules optimization actions in complex load scenarios to avoid performance degradation caused by resource idleness or competition. Attached Figure Description

[0020] Figure 1 This is a schematic diagram of the overall system operation process of the present invention.

[0021] Figure 2 This is a schematic diagram of the overall dynamic optimization operation process of the present invention.

[0022] Figure 3 This is a schematic diagram of the batch processing scale adjustment operation process of the present invention.

[0023] Figure 4 This is a schematic diagram of the resource instance scaling operation process of the present invention.

[0024] Figure 5 This is a schematic diagram of the concurrent control optimization operation process of the present invention. Detailed Implementation

[0025] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0026] In one specific embodiment, such as Figures 1-5 As shown, the large-model batch inference and data flow optimization system for the MOE architecture includes a request access module, an environment awareness module, an expert routing engine, a resource scheduling module, a dynamic optimization control module, a fault-tolerant migration module, and a performance monitoring and logging module. The functions of each module are as follows:

[0027] Request access module: Used to receive user request streams and extract request features (text length L, subject type T).

[0028] Environment Awareness Module: Used to collect inference environment parameters in real time (GPU memory usage G, storage I / O bandwidth B, request queue depth Q).

[0029] Expert Routing Engine: Used to dynamically activate expert sub-networks of the MOE model based on request characteristics.

[0030] Resource scheduling module: Used to manage the loading / unloading of weight data and the allocation of computing resources.

[0031] Dynamic optimization control module: Used to intelligently trigger optimization strategies based on environmental perception and access request data.

[0032] Fault-tolerant migration module: Used to monitor the status of GPU nodes and automatically migrate requests and weight data in case of failure.

[0033] Performance monitoring and logging module: Used to monitor system performance in real time and record operation logs.

[0034] In practical applications, the system first receives user request streams through the request access module and extracts request features, including text length L and subject type T. These extracted features provide crucial information for subsequent precise routing and dynamic optimization, helping to improve inference efficiency and accuracy. Simultaneously, the environment awareness module collects inference environment parameters in real time, including GPU memory utilization G, storage I / O bandwidth B, and request queue depth Q. Real-time monitoring of these environmental parameters provides real-time data support for resource scheduling and dynamic optimization, facilitating the rational allocation and optimized utilization of resources.

[0035] Then, the expert routing engine dynamically activates the expert sub-network of the MOE model based on the request characteristics. Dynamically activating the expert sub-network avoids unnecessary computation, improves inference efficiency and accuracy, and reduces resource consumption.

[0036] The expert routing engine establishes a mapping relationship between the subject type T and the expert sub-network based on the text length L and subject type T extracted by the request access module, and selects the most suitable expert sub-network for inference based on the text length L, prioritizing high-performance experts for long texts.

[0037] Meanwhile, the expert routing engine dynamically adjusts the number of activated expert sub-networks based on GPU memory utilization G and request queue depth Q. Specifically:

[0038] If GPU memory utilization G > 70% or request queue depth Q > 100, reduce the number of active experts.

[0039] If GPU memory usage G < 30% or request queue depth Q < 30, then activate more experts.

[0040] Next, the loading / unloading of weight data and the allocation of computational resources are managed through the resource scheduling module. By dynamically adjusting the loading strategy of weight data and the allocation of computational resources, I / O bottlenecks are reduced, resource utilization is improved, and the smooth progress of the inference process is ensured.

[0041] The resource scheduling module dynamically adjusts the loading strategy of weight data based on the GPU memory utilization rate G and storage I / O bandwidth B collected in real time by the environment perception module. Specifically, if the GPU memory utilization rate G > 80%, idle expert weights are unloaded to the CPU or SSD. If the storage I / O bandwidth B is idle, high-frequency subject weights are preloaded to the GPU.

[0042] Meanwhile, in high-concurrency request scenarios, the resource scheduling module predicts future request volumes and preloads the expert weights that may be needed. Specifically, it defines a time window w and calculates the request percentage W for each subject within that time window. i The percentage of filtering requests is W i Disciplines with a percentage greater than 15% are marked as high-frequency disciplines. This is based on a sliding window model and the formula... Predict the number of requests in the next window , where q t Let α be the request queue depth for the current time window, and α be the smoothing coefficient used to adjust the smoothness of the prediction results; it is typically a constant between 0 and 1. This formula predicts future request volumes based on a sliding window model. The request queue depth reflects the current system load and potential demand, while the smoothing coefficient adjusts the accuracy and stability of the prediction results. By predicting the request volume for the next window, the system can preload potentially needed resources, such as expert weights, to cope with upcoming load spikes. This helps reduce resource loading latency and overhead, improving system response speed and overall performance.

[0043] Furthermore, the dynamic optimization control module receives data from the environment perception module and the access request module, and analyzes and determines whether an optimization strategy needs to be activated, such as... Figure 2 As shown, specifically:

[0044] According to the formula Calculate the routing conflict factor Rc, where G max L is the upper limit of GPU video memory. avgThe formula, representing the average text length, aims to evaluate the relationship between GPU memory resources and the amount of data in the current processing task. GPU memory resources are critical during the MOE model inference process, and text length directly affects memory usage. This formula can determine whether the current task will put significant pressure on GPU memory resources, thus serving as a basis for triggering dynamic optimization strategies. When the routing conflict factor exceeds a preset threshold, it indicates that GPU memory resources are relatively strained, requiring the activation of optimization strategies to alleviate the pressure.

[0045] Preset routing conflict factor threshold Rc th And the routing conflict factor Rc and the routing conflict factor threshold Rc th Compare them, if Rc > Rc th Then, the dynamic optimization operation set S={O1,O2,O3} is activated, where O1 represents batch size adjustment operation, O2 represents resource instance scaling operation, and O3 represents concurrency control optimization operation.

[0046] Among them, such as Figure 3 As shown, the specific steps for performing batch size adjustment are as follows:

[0047] Monitoring real-time average latency D a ;

[0048] According to the formula Calculate the batch processing adjustment factor B s D t arg et For the target delay, B max This formula represents the maximum I / O bandwidth, used to measure the amount of data (expressed in terms of I / O bandwidth) that the system can process under a given target latency. Batch size directly impacts system throughput and latency; therefore, it needs to be adjusted based on the target latency and I / O bandwidth.

[0049] Preset batch processing adjustment coefficient threshold B th And the batch adjustment factor Bs and the batch adjustment factor threshold B are compared. th Compare them, if B s <B th If the batch size is too large, the batch size is reduced; otherwise, it is increased. This adaptive balance between throughput and latency ensures system performance in high-concurrency scenarios. This is achieved by calculating the batch adjustment coefficient B. s This is compared with a preset threshold, allowing for dynamic adjustment of the batch size. When the batch adjustment coefficient B... s If the value is less than the threshold, it means that the current batch size may cause the latency to exceed the target value, so the batch size needs to be reduced; conversely, the size can be increased to improve throughput.

[0050] Among them, such as Figure 4As shown, the specific steps for performing resource instance scaling operations are as follows:

[0051] The number of requests N is counted in units of time window M. r ;

[0052] According to the formula Calculate the load fluctuation factor L f , where N r-prev This formula represents the request volume in the previous time window and is used to assess changes in system load. Load fluctuations reflect dynamic changes in system request volume and are an important reference for system resource scheduling.

[0053] Preset load fluctuation factor threshold and load fluctuation factor With load fluctuation factor threshold To make a comparison, if > The number of instances is calculated proportionally. This intelligently handles traffic fluctuations, improving system flexibility and responsiveness. By comparing the request volume in the current time window with that in the previous time window, the trend of system load changes can be determined. When the load fluctuation factor L_f exceeds a preset threshold, it indicates a significant change in system load, and the number of resource instances may need to be adjusted to cope with load fluctuations.

[0054] Among them, such as Figure 5 As shown, the specific steps for performing concurrency control optimization operations are as follows:

[0055] The number of requests N within the statistical time window M r and the number of successful responses N s ;

[0056] According to the formula Calculate the dynamic value of service quality Q d This formula is used to evaluate the quality of service (QoS) of a system under current load. QoS is a crucial metric for measuring system performance and is closely related to the success rate and efficiency of request processing.

[0057] Preset dynamic service quality threshold And the dynamic value of service quality Q d With service quality dynamic threshold For comparison, if Q d > If the concurrency limit is high, then the concurrency limit is increased; otherwise, it is decreased. This ensures service quality and system stability, avoiding system overload and unstable response issues. This is achieved by calculating the dynamic service quality value Q. d By comparing the dynamic service quality value Q with a preset threshold, it can be determined whether the current service quality of the system meets the requirements. dWhen the threshold is exceeded, it indicates that the system may face the risk of overload or a decline in service quality, and the concurrency limit needs to be adjusted to improve system stability and service quality.

[0058] Furthermore, after the dynamic optimization control module activates the dynamic optimization operation set S, if it is necessary to execute any two or all of O1, O2, and O3, the following operations will be performed:

[0059] According to the formula , and Calculate the priority of batch size adjustment operation P1, resource instance scaling operation P2, and concurrency control optimization operation P3 respectively;

[0060] Sort P1, P2, and P3 in order of size, and then sort O1, O2, and O3 in that order, while simultaneously performing the corresponding dynamic optimization operations in that order.

[0061] At the same time, after the dynamic optimization control module activates the dynamic optimization operation set S, if the independent triggering conditions of operations O1, O2 and O3 are not met, the dynamic optimization operation set S is forcibly executed in the default priority order, which is O1 > O2 > O3.

[0062] Furthermore, the fault-tolerant migration module monitors the GPU node's operating status in real time. When a node failure is detected, pending requests are automatically migrated to healthy nodes, and relevant weight data is transferred simultaneously. This improves the system's reliability and fault tolerance, ensuring the continuity and stability of the inference service. Specifically:

[0063] Preset GPU node failure rate threshold ;

[0064] Define time window T E Count the number of fault events F within the window, where fault events include node unresponsiveness and memory errors.

[0065] Through formula Calculate GPU node failure rate ,in For the total number of GPU nodes, when > At that time, the faulty node is marked, and the pending requests are routed to the healthy node, while the weight data of the faulty node is migrated to the backup node.

[0066] The performance monitoring and logging module monitors key system performance metrics in real time, including GPU memory utilization (G), CPU utilization, storage I / O bandwidth (B), request processing latency, and throughput. It also records all important operation logs, including the execution status and results of request inbound, expert routing decisions, resource scheduling, and dynamic optimization operations, and provides historical data analysis and visualization reports. This provides data support for system performance evaluation and optimization, facilitates troubleshooting and fault analysis, and contributes to continuous system improvement and optimization.

[0067] In summary, the system first receives user request streams through the request access module and extracts request features. Simultaneously, the environment awareness module collects inference environment parameters in real time. The expert routing engine dynamically activates the expert sub-network of the MOE model based on request features, while the resource scheduling module manages the loading / unloading of weight data and the allocation of computational resources. The dynamic optimization control module determines whether to activate optimization strategies based on real-time data to cope with different load conditions. Furthermore, the fault tolerance migration module and the performance monitoring and logging module are responsible for improving the system's reliability and fault tolerance, as well as monitoring and recording system performance, respectively.

[0068] Through the collaborative work of the above modules, this invention realizes batch inference and data flow optimization for large models in the MOE architecture, and solves the problems of large inference latency fluctuations, uneven resource utilization, limited throughput and poor dynamic load adaptability, providing stable and efficient inference services for online education platforms with high concurrency and low latency requirements.

[0069] The foregoing has shown and described the basic principles, main features, and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited to the above embodiments. The embodiments and descriptions in the specification are merely preferred examples and are not intended to limit the invention. Various changes and modifications can be made to the invention without departing from its spirit and scope, and all such changes and modifications fall within the scope of the present invention as claimed. The scope of protection of the present invention is defined by the appended claims and their equivalents.

Claims

1. A system for batch inference and data flow optimization of large models for MOE architecture, characterized in that, The system includes a request access module, an environment perception module, an expert routing engine, a resource scheduling module, and a dynamic optimization control module; The request access module is used to receive user request streams and extract request features, including at least text length and subject type; The environment perception module is used to collect inference environment parameters in real time, including GPU memory usage, storage I / O bandwidth, and request queue depth. The expert routing engine is used to dynamically activate the expert sub-network of the MOE model based on the request characteristics; The resource scheduling module is used to manage the loading / unloading of weight data and the allocation of computing resources; The dynamic optimization control module receives data from the environment perception module and the access request module, and analyzes and determines whether an optimization strategy needs to be activated. Specifically: Calculate the routing conflict factor based on the GPU memory limit and average text length; A preset routing conflict factor threshold is set, and the routing conflict factor is compared with the routing conflict factor threshold. If the routing conflict factor exceeds the threshold, the dynamic optimization operation set is activated, which includes batch size adjustment operation, resource instance scaling operation, and concurrency control optimization operation. When dynamically activating the expert sub-network of the MOE model, the expert routing engine adopts the following strategy: Based on the text length and subject type extracted by the request access module, a mapping relationship between subject type and expert sub-network is established, and the most suitable expert sub-network is selected for inference based on the text length, with high-performance experts given priority for long texts. The number of activated expert subnetworks is dynamically adjusted based on GPU memory utilization and request queue depth. Specifically: if GPU memory utilization is higher than a first set value or request queue depth is higher than a second set value, the number of activated experts is reduced; if GPU memory utilization is lower than a third set value or request queue depth is lower than a fourth set value, more experts are activated.

2. The large-model batch inference and data flow optimization system for MOE architecture according to claim 1, characterized in that, After the dynamic optimization control module activates the dynamic optimization operation set, the specific steps for performing the batch size adjustment operation are as follows: Monitor real-time average latency; Calculate batch processing adjustment coefficients based on target latency and maximum I / O bandwidth; A preset batch processing adjustment coefficient threshold is set, and the batch processing adjustment coefficient is compared with the batch processing adjustment coefficient threshold. If the batch processing adjustment coefficient is less than the threshold, the batch processing size is reduced; otherwise, the size is increased.

3. The large-model batch inference and data flow optimization system for MOE architecture according to claim 2, characterized in that, After the dynamic optimization control module activates the dynamic optimization operation set, the specific steps for performing resource instance scaling operations are as follows: Request volume is counted in units of time windows; Calculate the load fluctuation factor based on the request volume of the current time window and the request volume of the previous time window; A preset load fluctuation factor threshold is set, and the load fluctuation factor is compared with the load fluctuation factor threshold. If the load fluctuation factor exceeds the threshold, the number of instances is calculated proportionally.

4. The large-model batch inference and data flow optimization system for MOE architecture according to claim 3, characterized in that, After the dynamic optimization control module activates the dynamic optimization operation set, the specific steps for performing concurrent control optimization operations are as follows: The number of requests and successful responses within the statistical time window; Calculate dynamic service quality values ​​based on request volume and successful response count; A preset dynamic threshold for service quality is set, and the dynamic value of service quality is compared with the dynamic threshold. If the dynamic value of service quality exceeds the threshold, the concurrency limit is increased; otherwise, the concurrency limit is decreased.

5. The large-model batch inference and data flow optimization system for MOE architecture according to claim 4, characterized in that, After the dynamic optimization control module activates the dynamic optimization operation set, if it needs to execute any two or all of the batch size adjustment operation, resource instance scaling operation, and concurrency control optimization operation, the following operations will be performed: Calculate the priority of batch size adjustment operation, resource instance scaling operation, and concurrency control optimization operation respectively; Sort by priority from highest to lowest, and execute the corresponding dynamic optimization operations in sequence.

6. The system for large-scale model batch inference and data flow optimization for MOE architecture according to claim 5, characterized in that, After the dynamic optimization control module activates the dynamic optimization operation set, if the independent triggering conditions for batch size adjustment operation, resource instance scaling operation, and concurrency control optimization operation are not met, the dynamic optimization operation set will be forcibly executed in the default priority order, whereby batch size adjustment operation takes precedence over resource instance scaling operation, and resource instance scaling operation takes precedence over concurrency control optimization operation.

7. The system for large-scale model batch inference and data flow optimization for MOE architecture according to claim 1, characterized in that, When performing the loading / unloading of weighted data and the allocation of computational resources, the resource scheduling module specifically implements the following functions: Based on the GPU memory usage and storage I / O bandwidth collected in real time by the environmental perception module, the loading strategy of weight data is dynamically adjusted. Specifically: if the GPU memory usage is higher than the set threshold, idle expert weights are unloaded to the CPU or SSD; if the storage I / O bandwidth is idle, high-frequency subject weights are preloaded to the GPU. In high-concurrency request scenarios, predict future request volume and preload the expert weights that may be needed. Specifically, define a time window, count the request ratio of each subject within that time window, filter subjects whose request ratio is higher than the ratio threshold and mark them as high-frequency subjects, and predict the request volume of the next window based on a sliding window model.

8. The system for large-scale model batch inference and data flow optimization for MOE architecture according to claim 1, characterized in that, The system also includes a fault-tolerant migration module, used to monitor the running status of GPU nodes in real time. When a node failure is detected, it automatically migrates pending requests to healthy nodes and synchronously transfers relevant weight data. Specifically: Preset GPU node failure rate threshold; Define a time window and count the number of fault events within the window. The fault events include node unresponsiveness and memory errors. Calculate the GPU node failure rate. When the failure rate exceeds a preset threshold, mark the faulty node and route the pending requests to the healthy node. At the same time, migrate the weight data of the faulty node to the backup node.

9. The system for large-scale model batch inference and data flow optimization for MOE architecture according to claim 1, characterized in that, The system also includes a performance monitoring and logging module, which is used to monitor key performance indicators of the system in real time, including GPU memory utilization, CPU utilization, storage I / O bandwidth, request processing latency and throughput. It is also used to record all important operation logs, including the execution status and results of request access, expert routing decisions, resource scheduling and dynamic optimization operations, and to provide historical data analysis and visualization reports.

Citation Information

Patent Citations

  • Hybrid expert model reasoning method based on cooperation of CPU and GPU

    CN120235253A

  • Multi-model collaborative operation method based on large model efficient training

    CN120295784A

Cited By

  • A deep learning inference method for large model local deployment

    CN122529073A