Multi-tenant large model reasoning service architecture and operation method thereof

Through multi-level cache management, dynamic routing strategies and load balancing mechanism, the problems of low utilization of memory resources and insufficient model load switching efficiency are solved, efficient resource utilization and low latency services are realized, and on-demand loading and parameter integration of personalized models of large-scale tenants are supported.

CN120295727APending Publication Date: 2025-07-11和创(北京)科技股份有限公司

Patent Information

Application Number
CN202510348112.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-24
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

The utilization rate of video memory resources in the existing technology is low, and it is impossible to support the coexistence of personalized models of large-scale tenants. The model loading and switching efficiency is insufficient, and the lack of dynamic routing strategies leads to increased request delays.

Method used

Multi-level cache management, dynamic routing strategy and load balancing mechanism are adopted to manage the storage of LoRA models in video memory, memory, disk and remote storage according to priority through the multi-level cache module. Combined with the routing registration module, the server's model cache location and load status are recorded. The dynamic routing module allocates requests based on the tenant identity and server status, and superimposes the LoRA model and the basic model to generate a personalized inference model through the parameter fusion module.

Benefits of technology

It significantly improves resource utilization, reduces video memory usage by more than 90%, supports tenants at the magnitude of tens of times, shortens the average response time by 40%, and can automatically deal with traffic peaks to avoid single-point overload.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120295727A_ABST
    Figure CN120295727A_ABST
Patent Text Reader

Abstract

The invention provides a multi-tenant large model reasoning service architecture and an operation method thereof, and the architecture comprises a multi-level cache module which is used for managing the storage of a LoRA model in a video memory, an internal memory, a disk and a remote memory according to the priority; the route registration module is used for recording the model cache position and the load state of each server; the dynamic routing module is used for allocating requests according to the tenant identifiers and the server states; and the parameter fusion module is used for superposing the LoRA model and the basic model to generate a personalized inference model. According to the method, efficient resource utilization and low-delay service are realized through multi-stage cache management, a dynamic routing strategy and a load balancing mechanism. The framework supports on-demand loading and parameter fusion of the tenant personalized LoRA model, and the reasoning service efficiency in a multi-tenant scene is remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of artificial intelligence model services, and particularly to a multi-tenant large model inference service architecture and its operation method. Background Art

[0002] With the wide application of large models (such as GPT, LLaMA, etc.), how to provide personalized inference services for different tenants has become a technical challenge. Traditional solutions usually adopt two methods: independent deployment and static model switching. When using the independent deployment method, each tenant separately deploys a complete large model instance, resulting in waste of video memory and computing resources. When using static model switching, a shared base model is used and the fine-tuning parameters of the tenant (such as LoRA) are dynamically loaded, but the video memory capacity is limited and cannot support real-time switching of a large number of tenants.

[0003] The main problems of the prior art include: low utilization rate of video memory resources, inability to support the coexistence of personalized models for a large number of tenants. Insufficient efficiency of model loading and switching, resulting in increased request latency. Lack of dynamic routing strategies, unable to optimize request allocation according to server load and model cache status. Summary of the Invention

[0004] The purpose of the present invention is to provide a multi-tenant large model inference service architecture and its operation method, which combines multi-level cache management, dynamic routing strategies and load balancing mechanisms to achieve efficient resource utilization and low-latency services for large model inference services.

[0005] According to an object of the present invention, the present invention provides a multi-tenant large model inference service architecture, including:

[0006] A multi-level cache module for managing the storage of LoRA models in video memory, memory, disk and remote storage according to priority;

[0007] A routing registration module for recording the model cache locations and load statuses of each server;

[0008] A dynamic routing module for allocating requests according to tenant identification and server status;

[0009] A parameter fusion module for superimposing the LoRA model and the base model to generate a personalized inference model.

[0010] Further, the multi-level cache module has a custom cache eviction policy, and synchronously updates the routing registration information when evicting models.

[0011] Further, the dynamic routing module preferentially selects a server where the model exists in the video memory and the load is lower than the threshold.

[0012] Further, the multi-level cache module includes a video memory space, a memory cache, and a disk storage. The video memory space stores the base model and serves as a high-priority cache. The memory cache is used to quickly load active micro-models, and the disk storage is used to store cold models and support expansion to external storage.

[0013] Further, the dynamic routing module includes a priority routing policy that preferentially routes requests to the server with the corresponding model present in the video memory and the lowest load. If there are no available nodes, it selects the server with the next-level storage and distributes them to the most idle node through global load balancing.

[0014] Further, the parameter fusion module superimposes the incremental parameters of the LoRA model on the original parameters of the base model to generate a personalized inference model.

[0015] Further, the cache eviction policy includes a cache eviction policy based on LRU or LFU, which is used to dynamically adjust the storage location of the LoRA model in the video memory, memory, disk, and remote storage according to the model access frequency, and ensure the maximum utilization of the video memory space.

[0016] According to another object of the present invention, the present invention provides an operating method for the above multi-tenant large model inference service architecture, including the following steps:

[0017] Step 1: Model training and cache initialization

[0018] The tenant submits private data, and the platform trains to generate the corresponding LoRA model and initially stores all LoRA models in remote storage;

[0019] Step 2: Request routing and load balancing

[0020] Receive a user request and query the routing registry according to the routing service, and allocate servers according to the priority. Among them, priority 1 is that there is a model in the video memory and the concurrency is lower than the threshold, priority 2 is that there is a model in the memory / disk and the load is lower, and priority 3 is the server with the lowest global concurrency;

[0021] Step 3: Dynamic loading and fusion inference

[0022] After the target server receives the request, it checks whether there is a LoRA model in the video memory. If it exists, it directly performs parameter fusion inference; if not, it retrieves and loads it level by level in the order of video memory → memory → disk → remote storage;

[0023] Step 4: Status synchronization and eviction management

[0024] When the video memory is insufficient, the cache eviction policy is used to move the low-frequency model from the video memory to the memory and update the routing registry information to ensure the reasonable use of the video memory space.

[0025] Furthermore, the routing registry records the storage locations and load status of LoRA models on each server in real time. When a model in the video memory is eliminated, it actively notifies the routing service to update the registry to avoid incorrect routing.

[0026] Furthermore, the routing decision-making process performs load balancing through a dynamic load evaluation system, and the threshold is dynamically adjusted. Based on the moving window average calculation of historical load data, abnormal traffic automatically triggers an increase in the threshold, thereby effectively coping with burst traffic and optimizing request routing.

[0027] The technical solution of the present invention realizes efficient resource utilization and low-latency services through multi-level cache management, dynamic routing policies, and load balancing mechanisms. This architecture supports the on-demand loading and parameter fusion of tenant-specific LoRA models, significantly improving the inference service efficiency in multi-tenant scenarios. BRIEF DESCRIPTION OF THE DRAWINGS

[0028] In order to more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the following will briefly introduce the drawings required for the description of the specific embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0029] Figure 1 It is the system architecture diagram of the embodiment of the present invention;

[0030] Figure 2 It is the flowchart of LoRA model loading and elimination of the embodiment of the present invention;

[0031] Figure 3 It is the routing decision flowchart of the embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0032] The following will clearly and completely describe the technical solutions of the present invention in conjunction with the embodiments. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0033] In the description of the present invention, it should be understood that the orientation or positional relationship indicated by the terms "center", "longitudinal", "transverse", "length", "width", "thickness", "upper", "lower", "front", "rear", "left", "right", "vertical", "horizontal", "top", "bottom", "inner", "outer", "clockwise", "counterclockwise", etc. is based on the orientation or positional relationship shown in the drawings. It is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and thus should not be construed as a limitation to the present invention.

[0034] In addition, the terms "first" and "second" are only used for descriptive purposes and should not be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include one or more of the said features. In the description of the present invention, "a plurality" means two or more unless otherwise specifically defined. In addition, the terms "mounted", "connected", "connected to" should be understood in a broad sense. For example, it may be a fixed connection, a detachable connection, or an integral connection; it may be a mechanical connection or an electrical connection; it may be directly connected or indirectly connected through an intermediate medium, and it may be the communication inside two elements. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to specific situations.

[0035] Embodiment 1

[0036] A multi-tenant large model inference service architecture, comprising:

[0037] A multi-level cache module for managing the storage of the LoRA model in video memory, memory, disk and remote storage according to priorities;

[0038] A routing registration module for recording the model cache locations and load statuses of each server;

[0039] A dynamic routing module for allocating requests according to tenant identifiers and server statuses;

[0040] A parameter fusion module for superimposing the LoRA model and the base model to generate a personalized inference model.

[0041] Among them, the multi-level cache module customizes a cache eviction policy and synchronously updates the routing registration information when evicting models.

[0042] The dynamic routing module preferentially selects a server where the model exists in the video memory and the load is lower than the threshold.

[0043] Such as Figure 1As shown in the figure, it is the system architecture diagram of the embodiment of the present invention, specifically showing the interaction process of the basic model, multi-level cache, routing service, and tenant requests.

[0044] The system hierarchical architecture includes a user access layer, a routing and governance layer, and a model execution layer, where: The user access layer has the function of multi-tenant support (users A / B / C) and submits inference requests through independent channels.

[0045] The routing and governance layer specifically includes a model routing service and a model governance service, where:

[0046] The model routing service has the characteristics of dynamic routing decision (based on node load / model status) and real-time synchronization of micro-model node information with the governance service.

[0047] The model governance service has the characteristics of node registration and status monitoring and maintaining a global service directory (including the deployment locations of the basic model and micro-models).

[0048] The inference execution layer includes a cluster of model inference service nodes, which can process basic model (resident in video memory) and micro-model requests in parallel, and supports a three-level storage architecture of video memory - memory - disk.

[0049] The service architecture of the present invention realizes hierarchical storage optimization:

[0050] Video memory space: The basic large model resides in video memory permanently (high priority); micro-models reside dynamically according to popularity (cooperating with the cache eviction policy).

[0051] Memory cache: Fast loading layer for active micro-models; cache policy based on LRU or LFU.

[0052] Disk storage: Persistent storage space for micro-models; cold model archiving (possibly docking with external storage expansion).

[0053] The service architecture of the present invention realizes dynamic resource scheduling:

[0054] The cache eviction policy takes effect bidirectionally

[0055] Load balancing mechanism: Node load information is fed back to the routing service in real time; elastic scaling based on QPS / video memory occupancy / response latency.

[0056] The service architecture of the present invention realizes model life cycle management:

[0057] Fast response for hot models: Reside in video memory → immediate inference;

[0058] On-demand loading for warm models: Pre-load in memory → dynamic replacement in video memory;

[0059] Cold model archiving mechanism: Disk storage → external storage (long-term preservation).

[0060] The service architecture of the present invention realizes the maximization of video memory resources through a three - level storage system, and differentially processes the basic model (resident) and the micro - model (dynamically loaded). The governance service ensures automatic elimination of node failures, and the caching strategy prevents OOM caused by video memory overflow. It supports horizontal expansion of inference nodes and realizes the management of a large number of micro - models through external storage docking. The present invention has high practical value in the scenario of model service - orientation, and is particularly suitable for service scenarios that need to support both basic models and a large number of long - tail micro - models at the same time.

[0061] A running method for a multi - tenant large - model inference service architecture includes the following steps:

[0062] Step 1: Model training and cache initialization

[0063] The tenant submits private data, and the platform trains to generate the corresponding LoRA model (only the incremental parameters are saved);

[0064] In the initial state, all LoRA models are stored in the remote storage;

[0065] Step 2: Request routing and load balancing

[0066] 1. The user request is sent to the routing service with the tenant identifier;

[0067] 2. The routing service queries the registry and allocates servers according to the following priorities:

[0068] Priority 1: There is a model in the video memory and the concurrency number < threshold.

[0069] Priority 2: There is a model in the memory / disk and the load is low.

[0070] Priority 3: The server with the lowest global concurrency number (triggering remote loading).

[0071] Step 3: Model loading and inference

[0072] 1. After receiving the request, the target server checks whether the corresponding LoRA model exists in the video memory:

[0073] If it exists, directly stack the parameters onto the basic model for inference.

[0074] If it does not exist, retrieve and load it step by step in the order of video memory → memory → disk → remote storage.

[0075] 2. When loading a new model, if the video memory is insufficient, eliminate the least recently used model to the memory and update the routing registry.

[0076] Step 4: State synchronization and elimination management

[0077] The server regularly reports the load and model cache status to the routing service.

[0078] When the model in the video memory is eliminated, the routing service is notified to remove the corresponding registration record.

[0079] The core innovation points of the present invention are as follows:

[0080] 1. Model management with multi-level caching and custom cache eviction policies

[0081] Storage classification: Cache the LoRA model in the video memory, memory, local disk, and remote storage according to priority, and dynamically adjust the storage location according to the access frequency.

[0082] Gradual eviction mechanism: Custom cache eviction policies (such as an improved LRU policy) can be defined at different levels to manage the cache. When the video memory is insufficient, low-frequency models are evicted to the next-level storage, and the routing registration information is updated.

[0083] 2. Dynamic routing and load balancing mechanism

[0084] Routing registration table: Real-time record the storage location (video memory / memory / disk) of the LoRA model in each server and the server load status.

[0085] Priority routing strategy: First, route the request to the server with the corresponding model in the video memory and the lowest load; if there are no available nodes, select the server with the next-level storage; finally, allocate it to the most idle node through global load balancing.

[0086] 3. Model dynamic loading and parameter fusion

[0087] Load on demand: Retrieve the model level by level during request processing. If the cache is not hit, trigger remote loading and update the cache.

[0088] Parameter fusion: The LoRA model parameters in the video memory are superimposed on the base model in real time to generate a personalized inference model.

[0089] 4. Real-time status synchronization

[0090] When the model is evicted from the video memory, actively notify the routing service to update the registration table to avoid incorrect routing.

[0091] The server periodically reports the load status (such as the number of concurrent requests, video memory occupancy rate) for the routing service to make decisions.

[0092] The inference service architecture of the multi-tenant large model of the present invention and its operation method significantly improve resource utilization. Through multi-level caching, the video memory occupancy is reduced by more than 90%, and it supports serving tenants in a multiple order of magnitude simultaneously. Low-latency guarantee, the dynamic routing strategy enables 90% of the requests to hit the model in the video memory, and the average response time is shortened by 40%. Elastic expansion, the load balancing mechanism can automatically handle traffic peaks and avoid single-point overload.

[0093] As Figure 2 shown, it is the flowchart of LoRA model loading and elimination in the embodiment of the present invention, which describes the process of loading the model from remote storage to the video memory and the cache elimination process.

[0094] Specifically, the core processing flow of LoRA model loading and elimination mainly includes:

[0095] 1. Request acceptance and model location:

[0096] After receiving the user's inference request, the system first detects the storage location of the target LoRA model;

[0097] Video memory (GPU video memory): If it is already resident, directly enter the parameter fusion inference stage;

[0098] Memory / disk: If it is not resident in the video memory but exists in the local storage, trigger video memory loading;

[0099] Remote storage: When there is no local copy, it is necessary to remotely download the model file;

[0100] 2. Dynamic management of video memory space:

[0101] When the video memory is sufficient, directly load the target model and update the routing table;

[0102] When the video memory is insufficient, trigger a three-level elimination mechanism:

[0103] 3. Parameter fusion inference execution

[0104] After the model is loaded, dynamically fuse the LoRA adapter parameters with the base model;

[0105] Output the inference result and keep the base model resident in the video memory (to avoid repeated loading).

[0106] The LoRA model hierarchical storage strategy of the present invention mainly includes:

[0107] Video memory hot layer: Frequently accessed LoRA models are resident in the GPU video memory;

[0108] Memory warm layer: Active but not frequently accessed models are resident in the memory (fast loading channel);

[0109] Disk cold layer: Infrequently accessed models are persistently stored (supporting fast reading);

[0110] Remote repository: Model version archiving and disaster recovery storage.

[0111] The intelligent elimination algorithm of the present invention: Keep memory / disk copies during model migration to avoid repeated downloads.

[0112] The real-time synchronization mechanism of the routing table of the present invention: After the model location changes, achieve second-level state synchronization through a distributed routing table; support multi-node collaborative updates to prevent requests from being routed to invalid nodes.

[0113] The technical solution of the present invention is particularly applicable to large-scale personalized model service scenarios (such as AIGC applications with thousands of personalized faces). It realizes the reuse of the base model through LoRA lightweight adapters, and combines a hierarchical storage mechanism to balance resource utilization and service quality. It is a typical optimization solution for the implementation of current large model services.

[0114] Such as Figure 3 As shown, it is the routing decision flow chart of the embodiment of the present invention, showing the routing logic based on tenant identification, model cache location, and load status.

[0115] The routing decision process adopts a multi-level cache routing strategy:

[0116] Priority routing for the hot layer of video memory. The node needs to meet two conditions: the model resides in the video memory and the real-time load is lower than the preset threshold (such as GPU utilization < 80%); a millisecond-level response channel (to avoid disk I / O bottlenecks).

[0117] Fallback routing for local cache. The memory / disk cache node needs to meet the load factor ≤ 0.7 (dynamically adjust the threshold through the PID algorithm); loading time control: memory cache < 200ms, disk cache < 2s (SSD environment).

[0118] The routing decision process adopts a dynamic load evaluation system:

[0119] Threshold dynamic adjustment, calculated based on the moving window average of historical load data (window period

[0120] = 5min); automatic trigger of threshold increase for abnormal traffic (such as during burst traffic, the threshold of the video memory node changes from 80% to 90%).

[0121] Cross-node load synchronization, realize the state synchronization of the routing registry through etcd (update delay < 500ms).

[0122] The routing decision process adopts a remote loading acceleration scheme: Model sharding transmission (based on HTTP / 3 protocol multiplexing); an edge loading and preheating mechanism: After the first 50% of the model parameters are loaded, the inference request queue is opened, and the background continues to load the remaining parameters (cooperated by two threads).

[0123] The technical solution of the present invention realizes the deep combination of hierarchical routing decision-making and elastic resource scheduling, and is particularly applicable to scenarios with a large number of long-tail model requests (such as personalized recommendation systems for thousands of people). Through the four-level resource scheduling of video memory - memory - disk - remote, the hardware utilization rate is increased by more than 40% while ensuring the SLA.

[0124] The multi-tenant large model inference service architecture and its operation method of the present invention realize efficient resource utilization and low-latency services through multi-level cache management, dynamic routing strategies, and load balancing mechanisms. This architecture supports the on-demand loading and parameter fusion of tenant-specific LoRA models, significantly improving the inference service efficiency in multi-tenant scenarios.

[0125] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some or all of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. A multi-tenant large model inference service architecture, characterized in that, Including: A multi-level cache module for managing the storage of the LoRA model in video memory, memory, disk, and remote storage according to priority; A routing registration module for recording the model cache locations and load statuses of each server; A dynamic routing module for allocating requests based on tenant identification and server status; A parameter fusion module for superimposing the LoRA model and the base model to generate a personalized inference model.

2. The multi-tenant large model inference service architecture according to claim 1, wherein The multi-level cache module has a custom cache eviction policy, and synchronously updates the routing registration information when evicting a model.

3. The multi-tenant large model inference service architecture according to claim 1, wherein The dynamic routing module preferentially selects a server where the model exists in the video memory and the load is lower than the threshold.

4. The multi-tenant large model inference service architecture according to claim 1, wherein The multi-level cache module includes a video memory space, a memory cache, and a disk storage. The video memory space stores the base model and is a high-priority cache. The memory cache is used to quickly load active micro-models. The disk storage is used to store cold models and supports expansion to external storage.

5. The multi-tenant large model inference service architecture according to claim 1, characterized in that, The dynamic routing module includes a priority routing policy, which preferentially routes requests to the server where the corresponding model exists in the video memory and has the lowest load. If there are no available nodes, it selects the server with the next-level storage and distributes it to the most idle node through global load balancing.

6. The multi-tenant large model inference service architecture according to claim 1, characterized in that, The parameter fusion module superimposes the incremental parameters of the LoRA model and the original parameters of the base model to generate a personalized inference model.

7. The multi-tenant large model inference service architecture according to claim 2, wherein The cache eviction policy includes a cache eviction policy based on LRU or LFU, which is used to dynamically adjust the storage location of the LoRA model in video memory, memory, disk, and remote storage according to the model access frequency, and ensure the maximum utilization of the video memory space.

8. The operating method of the multi-tenant large model inference service architecture according to any one of claims 1-6, characterized in that, Including the following steps: Step 1, Model training and cache initialization The tenant submits private data, the platform trains to generate the corresponding LoRA model, and initially stores all LoRA models in remote storage; Step 2, Request routing and load balancing Receive a user request and query the routing registration table according to the routing service, and allocate servers according to priority. Among them, priority 1 is that the model exists in the video memory and the concurrency is lower than the threshold, priority 2 is that the model exists in the memory / disk and the load is lower, and priority 3 is the server with the lowest global concurrency; Step 3, Dynamic loading and fusion inference After the target server receives the request, it checks whether the LoRA model exists in the video memory. If it exists, it directly performs parameter fusion inference; if it does not exist, it retrieves and loads it level by level in the order of video memory, memory, disk, and remote storage; Step 4, Status synchronization and eviction management When the video memory is insufficient, the low-frequency model is moved from the video memory to the memory through the cache eviction policy, and the routing registration table information is updated to ensure the reasonable use of the video memory space.

9. The operating method of the multi-tenant large model inference service architecture according to claim 8, characterized in that, The routing registration table records the LoRA model storage locations and load statuses of each server in real time, and when the model in the video memory is evicted, it actively notifies the routing service to update the registration table to avoid incorrect routing.

10. The operating method of the multi-tenant large model inference service architecture according to claim 8, wherein The routing decision process performs load balancing through a dynamic load evaluation system, and the threshold is dynamically adjusted based on the moving average of historical load data. Abnormal traffic automatically triggers the threshold to float, so as to effectively handle burst traffic and optimize request routing.

Citation Information

Patent Citations

  • KServe model reasoning acceleration method and device based on distributed cache and medium

    CN118153694A

  • Drawing service method and system based on adaptive scheduling LoRA model and storage medium

    CN118447116A

  • Inference system, inference method, inference device, storage medium, and program product

    CN119129755A

  • Load balancing scheduling method and device for model reasoning and electronic equipment

    CN119440847A

  • Large model online reasoning service implementation method based on no-service calculation

    CN119512744A

Cited By

  • Method and system for realizing resource use and concurrent reasoning of large-model all-in-one machine

    CN120930806A

  • Large language model reasoning method and device, equipment and storage medium

    CN121189489A

  • Video flaw tracing and fixed-point regeneration method and system

    CN122493372A

  • Video flaw traceability and fixed-point regeneration method and system

    CN122493372B

  • Multi-tenant ai inference service platform lora hot switching method and system, medium

    CN122661349A