Multi-LoRA large language model deployment system based on cloud computing platform
By combining LoRA technology and container orchestration on the cloud computing platform, the video memory sharing and resource optimization of large language models are achieved, and the high cost and resource waste problems in traditional deployment methods are solved, and the training efficiency and resource utilization rate in multiple business scenarios are improved.
Patent Information
- Application Number
- CN202510365835.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-26
- Publication Date
- 2025-07-25
AI Technical Summary
The traditional large language model deployment methods have problems such as high cost, waste of video memory resources, fragmented R&D processes, lack of unified management, waste of resources, complex management, high training costs, insufficient resource flexibility and scheduling capabilities, and lack of priority management. In particular, the cloud computing platform has failed to effectively combine LoRA technology with cloud native resource scheduling.
The multi-LoRA large language model deployment system based on the cloud computing platform is adopted, including the cloud computing AI platform layer, multi-LoRA dynamic loading layer and resource scheduling optimization layer. Through the deep integration of LoRA adapter and container orchestration, it realizes video memory sharing, dynamic loading and optimizes resource scheduling, and supports unified management and automated iteration of multiple business scenarios.
Significantly reduce the memory usage and deployment costs, improve training efficiency and resource utilization, realize efficient training and deployment of large language models in multiple business scenarios, and improve development efficiency and resource utilization.
Smart Images

Figure CN120371320A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to artificial intelligence and cloud computing technologies, and particularly to a method for deploying a large language model based on a cloud computing platform. Background Art
[0002] A large language model is a large model specifically designed for processing natural language tasks. It is usually based on the Transformer architecture and has extremely strong text understanding and generation capabilities.
[0003] With the popularization of AI technology, enterprises need to train dedicated large language models for different business scenarios. The traditional deployment methods have the following problems:
[0004] 1. The deployment cost of large language models is too high:
[0005] In the traditional mode, a fine-tuned large language model needs to be independently deployed for each business scenario. Taking a 70B parameter model as an example, a single deployment requires 4 A100 GPUs (about 80GB of video memory), and the total video memory requirement for 10 scenarios reaches 800GB, and the hardware cost reaches several million yuan.
[0006] Repeated occupation of video memory resources: The basic model parameters of different businesses are exactly the same, only some parameters are fine-tuned, but the full amount needs to be loaded, resulting in a video memory waste rate of more than 90%.
[0007] 2. Fragmentation and inefficiency of the R & D process:
[0008] The model training, evaluation, and deployment links are scattered on different platforms (such as PAI training and ECS deployment), and the data flow depends on manual operations, and the iteration cycle is as long as 1-2 weeks.
[0009] Lack of unified management: There is no associated metadata among model versions, LoRA adapters, and data sets, making it difficult to trace the corresponding relationship between model lineage and business scenarios.
[0010] 3. Resource waste: Each scenario needs to independently deploy a complete large model, occupying a large amount of GPU video memory.
[0011] 4. Complex management: Model versions are iterated frequently, and traditional containerization solutions are difficult to achieve dynamic decoupling of models and inference services.
[0012] 5. High training cost: Fine-tuning a large model with all parameters requires consuming a huge amount of computing resources, which is difficult for small and medium-sized enterprises to bear.
[0013] 6. Insufficient resource elasticity and scheduling capabilities: Traditional resource allocation adopts the static ECS mode, which cannot cope with sudden traffic (such as a 10-fold increase in QPS during a large promotion) or large-scale training tasks (such as distributed training of a model with hundreds of billions of parameters).
[0014] 7. Lack of priority management: High-priority online inference tasks compete with low-priority offline training tasks for resources, resulting in service latency fluctuations.
[0015] The LoRA (Low-Rank Adaptation) technique is a technique for fine-tuning large-scale pre-trained models. It performs low-rank decomposition on the weight matrix of the pre-trained model, introduces low-rank matrices to achieve fine-tuning, reduces the number of parameters, and avoids directly modifying the original weights, thereby reducing computational and storage overhead. Only a small number of parameters need to be fine-tuned to adapt to new tasks, avoiding the high cost of full-model fine-tuning. Low-rank matrix operations are more efficient than full-rank matrices, reducing the demand for computing resources and being suitable for resource-constrained environments.
[0016] The multi-LoRA technique is an extension of LoRA that allows multiple low-rank adaptation matrices to be used simultaneously to fine-tune the model to handle more complex tasks or multi-task learning scenarios. Multi-LoRA achieves independent adaptation to different tasks or features by introducing multiple low-rank matrices into the weight matrix of the pre-trained model. Each low-rank matrix can capture different task-specific information, thereby enhancing the flexibility and performance of the model.
[0017] LoRA adapters are usually used to fine-tune pre-trained language models by introducing low-rank matrices to reduce the number of parameters.
[0018] The applicant has found through research that the current problems of cloud computing platforms when using the multi-LoRA technique are as follows:
[0019] Multi-LoRA technique: Existing solutions require manual parameter merging or independent deployment, cannot dynamically switch adapters, and are not combined with cloud-native resource scheduling, making it difficult to achieve video memory sharing and elastic scaling.
[0020] AI platform: Although it realizes containerized AI task management, it does not solve the problem of sharing models in multi-business scenarios and lacks in-depth support for the LoRA technique.
[0021] Therefore, although cloud computing platforms can improve resource utilization through resource adjustment and scheduling, they lack in-depth optimization for the AI model lifecycle; while the multi-LoRA technique can reduce training costs, its deployment still relies on specific frameworks, which are more inclined to traditional monolithic application architectures compared to the microservices architecture required by cloud-native, making it difficult to seamlessly integrate with the cloud-native ecosystem. Summary of the Invention
[0022] The technical problem to be solved by the present invention is to provide a system solution that combines cloud-native elastic resource management, multi-LoRA dynamic loading, and automated model iteration.
[0023] The technical solution adopted by the present invention to solve the above technical problems is a multi-LoRA large language model deployment system based on a cloud computing platform, including a cloud computing AI platform layer, a multi-LoRA dynamic loading layer, and a resource scheduling optimization layer;
[0024] The cloud computing AI platform layer is used to receive text task data, train a LoRA adapter adapted to the basic large language model for the inference process, and send the resource requirements of the LoRA model to the multi-LoRA dynamic loading layer;
[0025] The multi-LoRA dynamic loading layer is used to store LoRA adapter files for each business scenario; send a resource scheduling application to the resource scheduling optimization layer according to the received resource requirements of the LoRA model; then dynamically switch the LoRA adapter to be mounted according to the request parameters, perform GPU optimization configuration according to the business priority in the request parameters, load the basic large language model, mount several LoRA adapters superimposed in the form of a sparse matrix, and complete the deployment of the LoRA model for this inference; afterwards, the LoRA model outputs a task response for the text task data; among them, when performing GPU optimization configuration, based on the service quality mechanism of the container orchestration platform, exclusive GPU resources are allocated to high-priority services, and medium and low-priority tasks share the resource pool;
[0026] In response to the resource scheduling application, the resource scheduling optimization layer outputs request parameters optimized through priority management and hardware awareness to the LoRA dynamic loading layer based on the container orchestration platform.
[0027] The present invention deeply integrates LoRA with container orchestration to realize an automated pipeline for training and inference, improve the efficiency of large model deployment, and enable users who do not understand technical details to quickly deploy a multi-LoRA large language model. The video memory sharing mechanism allows multiple business scenarios to share the same basic large language model and only load different LoRA adapters, greatly reducing video memory occupancy.
[0028] Specifically, the cloud computing AI platform layer includes a dataset management module, a model training module, a model management module, and a model service management module;
[0029] The dataset management module is used to support the unified access of multiple data sources;
[0030] The model training module is used to generate a LoRA adapter adapted to the basic large language model through an AI-Pipeline engine that supports distributed training tasks;
[0031] The model management module is used to manage existing basic large language models;
[0032] The model service management module is used to support the dynamic binding of LoRA adapters and basic large language models.
[0033] Furthermore, the multi-LoRA dynamic loading layer adopts a dynamic scheduling strategy to hierarchically cache LoRA adapters according to usage frequency. High-frequency adapters are resident in video memory, and low-frequency adapters are loaded on demand, achieving the effect of improving the response speed of large language models.
[0034] Specifically, the multi-LoRA dynamic loading layer includes a LoRA repository, a video memory optimization module, and a Vllm inference engine.
[0035] The LoRA repository is used to store LoRA adapter files for each business scenario, record versions, and associated metadata. The associated metadata includes the base model version and business scenario tags.
[0036] The video memory optimization module is used to adopt a hierarchical caching mechanism based on the Quality of Service (QoS) mechanism of the container orchestration platform. By monitoring the call frequency of LoRA adapters, LoRA adapters with high-frequency access are retained in the GPU.
[0037] LoRA adapters with medium and low-frequency access are persisted to the Object Storage Service (OSS).
[0038] The Vllm inference engine is used to support the simultaneous loading of a base large language model and several LoRA adapters in the GPU to form a LoRA model through the integrated VLLM inference framework, and dynamically switch adapters through request parameters, and output task responses for text task data inferred by the LoRA model.
[0039] Furthermore, the resource scheduling optimization layer designs an elastic resource pool and time-aware training / inference task scheduling through resource optimization management, improving GPU utilization, increasing resource utilization, and thus reducing costs.
[0040] Specifically, the resource scheduling optimization layer includes a priority management module, an elastic computing power resource allocation module, and a hardware-aware optimization module.
[0041] The priority management module is used to schedule tasks based on a priority queue, dividing the business task priorities into three levels: high, medium, and low, and adopting the Gang Scheduling strategy to ensure the full resource allocation of distributed training tasks.
[0042] The elastic computing power resource allocation module is used to automatically select a resource pool according to the task type. The task types include training tasks or inference tasks. Training tasks are preferentially scheduled to the elastic resource pool, and inference tasks are bound to the online resource pool.
[0043] The hardware-aware optimization module is used to dynamically bind tasks to appropriate CPU or GPU cores through dynamic core binding technology.
[0044] The beneficial effects of the present invention are as follows: the basic large language model is shared among different business scenarios, and multiple LoRAs are dynamically loaded, resulting in a significant reduction in the video memory occupancy rate and the model deployment cost; the containerized resource scheduling and unified model management technology greatly improve the training efficiency and resource utilization rate of the deployed large language model in multiple business scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0045] Figure 1 is a system architecture diagram;
[0046] Figure 2 is an AI-Pipeline flow chart;
[0047] Figure 3 is a schematic diagram of the video memory sharing mechanism;
[0048] Figure 4 is a resource scheduling diagram. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0049] The embodiments will be described in detail below with reference to the accompanying drawings.
[0050] Figure 1 The multi-LoRA large language model system of the cloud computing platform in the embodiments is shown. The system is divided into four layers: the user interaction layer, the cloud computing AI platform layer, the multi-LoRA dynamic loading layer, and the resource scheduling optimization layer.
[0051] Kubernetes is used as an open-source container orchestration platform for automating the deployment, scaling, and management of containerized applications. It can run on various infrastructures, including local data centers, public clouds, private clouds, and hybrid cloud environments.
[0052] The user interaction layer is used to receive text task data and display the task response for the text task data; it mainly includes a Web interface and an API gateway. On the Web interface, users submit training tasks, select data sets, and configure LoRA parameters through the cloud computing platform, while the API gateway receives HTTP requests and routes them to the corresponding LoRA adapters. Among them, LoRA parameters such as the rank R and the scaling factor α; HTTP requests such as X-LoRA-Id:scen1. When the client request is routed, X-LoRA-Id:scen1 is specified in the HTTP Header, and the server dynamically switches the adapter and dynamically merges the parameters during calculation.
[0053] The cloud computing AI platform layer is used to train LoRA adapters suitable for the basic large language model and send the resource requirements of the LoRA adapters to the multi-LoRA dynamic loading layer; it includes four modules, namely, the data set management module, the model training module, the model management module, and the model service management module.
[0054] The dataset management module is used to support the unified access of multiple data sources and accelerate data caching through the Fluid engine; the multiple data sources include the Object Storage Service (OSS), Network Attached Storage (NAS), and Open Data Processing Service (ODPS). Among them, OSS is suitable for storing unstructured data and is suitable for static resources, backups, and archiving. NAS is suitable for file-level shared storage and is suitable for file sharing and persistent storage. ODPS is suitable for big data analysis and processing and is suitable for building data warehouses and machine learning. The Fluid engine is an open-source cloud-native data orchestration and acceleration system that focuses on efficiently managing, accessing, and accelerating data-intensive applications in a Kubernetes environment.
[0055] The model training module is used to generate a LoRA adapter adapted to the base large language model through the AI-Pipeline engine that supports distributed training tasks; the AI-Pipeline engine is used to implement and manage the automated execution of the AI-Pipeline, and the AI-Pipeline is a technology that organizes multiple steps of AI tasks into an automated workflow. More specifically, the model training module of the embodiment provides a Notebook interactive development environment, supports distributed training tasks such as TFJob and PyTorchJob, and integrates the LoRA fine-tuning framework;
[0056] The training of the LoRA adapter is automatically orchestrated by the AI-Pipeline, and the pipeline includes stages of data preprocessing, distributed training, model verification, and service release. Figure 2 The AI-Pipeline flowchart for the embodiment is to orchestrate multi-stage tasks based on ArgoWorkflow. First, in the data preprocessing stage, the source data is output to the dedicated data storage server NAS for storage after data cleaning and format conversion. Then, it enters the LoRA fine-tuning stage. After the user configures the rank and the number of training epochs, the PytorchJob that supports the Torch distributed training task is started to execute the multi-GPU training task to generate the LoRA parameter file. Finally, in the model training stage, the evaluation script is executed to obtain the accuracy of the model and the video memory occupancy of the model for model verification. The models that do not meet the requirements are retrained, and for the models that meet the accuracy and video memory requirements, the VLLM engine loads the model and the LoRA adapter to provide the communication interface REST API between the client and the server, thus completing the release of the service.
[0057] The model management module is used to manage the completed base large language model, such as Deepseek v3.
[0058] The model service management module is used to support the dynamic binding of the LoRA adapter and the base large language model and provide a hot update interface.
[0059] The present invention deeply integrates cloud computing with multiple LoRAs: embed a LoRA fine-tuning template in the platform, and the user selects a base large language model and LoRA parameters (such as rank R≤64) through configuration. The platform automatically generates a training task and outputs an adapter file. During deployment, the model service declares the list of LoRA adapters it depends on through custom resources, and the platform dynamically mounts the adapters to the video memory when the service starts.
[0060] The multi-LoRA dynamic loading layer is used to implement the storage, version management, and video memory sharing of LoRA adapters; store LoRA adapter files for each business scenario, record the version of the base large language model and associated metadata, and the associated metadata is the business scenario label; for example, LoRA parameter files stored by business label classification, gray release and rollback of versions; send a resource scheduling application to the resource scheduling optimization layer according to the resource requirements of the received LoRA adapter; then dynamically switch the LoRA adapter to be mounted according to the request parameters, perform GPU optimization configuration according to the business priority in the request parameters, load the base large language model, and mount several LoRA adapters superimposed in the form of a sparse matrix to complete the deployment of the LoRA model for inference this time; afterwards, the LoRA model outputs a task response for the text task data; among them, when performing GPU optimization configuration, based on the quality of service mechanism of the container orchestration platform, exclusive GPU resources are allocated to high-priority services, and medium- and low-priority tasks share the resource pool.
[0061] The multi-LoRA dynamic loading layer includes a LoRA repository, a video memory optimization module, and a Vllm inference engine.
[0062] The video memory optimization module is used to adopt a hierarchical caching mechanism based on the QoS mechanism of Kubernetes to monitor the call frequency of LoRA adapters. The LoRA adapters with high-frequency access are retained in the GPU, and the LoRA adapters with medium- and low-frequency access are persisted to the object storage service OSS; specifically, in the hierarchical caching mechanism, the LoRA adapters with low-frequency access are loaded on demand; only the hot models are retained in the video memory to reduce the GPU memory pressure; the LoRA adapters with medium-frequency access are stored in the RAM memory, and the LoRA adapters with low-frequency access are stored in the disk Disk; in addition, a video memory recycling strategy is executed: for LoRA adapters without requests within 24 hours, they are automatically unloaded and stored in the RAM memory to release the video memory resources;
[0063] Figure 3The video memory sharing mechanism of the embodiment is shown, which consists of two parts: the basic large language model parameters resident in the video memory and the dynamic superposition of LoRA adapters. The basic large language model parameter matrix is usually determined by the size of the base model. The LoRA adapters are divided into three layers: hot, warm, and cold according to the call frequency and are stored in different storage spaces and loaded into the video memory at different times. The basic large language model is only loaded once, and each LoRA adapter is superimposed in the form of a sparse matrix. Model registration: The LoRA file is associated with the basic large language model (such as DeepSeek-R1 70B), and the metadata is stored in the model repository. Each LoRA adapter is superimposed in the form of a sparse matrix, and parameter updates are achieved through matrix addition. The formula is expressed as:
[0064] W′ = W + α·(A × B)
[0065] Among them, W is the basic large language model parameter, A and B are LoRA low-rank matrices, and α is the scaling coefficient.
[0066] The Vllm inference engine is used to support the simultaneous loading of a basic large language model and several LoRA adapters in the GPU through the integrated VLLM inference framework to form a LoRA model, and dynamically switch adapters through request parameters; output the task response of the LoRA model inferred for text task data.
[0067] The resource scheduling optimization layer responds to the resource scheduling application, and based on the intelligent scheduling engine of Kubernetes, ensures the resource utilization rate and the task service level agreement SLA, and outputs the request parameters after priority management and hardware-aware optimization to the LoRA dynamic loading layer.
[0068] The resource scheduling optimization layer performs differential scheduling according to the task type. For tasks of the training type, elastic resource pools (such as Spot instances) are preferentially allocated, and the CPU utilization rate is improved by using the core binding strategy. For tasks of the inference type, based on the Horizontal Pod Autoscaler (HPA), it automatically scales in and out according to the queries per second (QPS), and combines GPU sharing technologies such as GPU virtualization technology MIG to improve resource density.
[0069] The resource scheduling optimization layer includes a priority management module, an elastic computing power resource allocation module, and a hardware-aware optimization module.
[0070] The priority management module is used to schedule tasks based on the priority queue, divide the business task priorities into three levels: high, medium, and low, and adopt the Gang Scheduling strategy to ensure the full resource allocation of distributed training tasks;
[0071] The elastic computing power resource allocation module is used to automatically select a resource pool according to the task type. The task types include training tasks or inference tasks. Training tasks are preferentially scheduled to the elastic resource pool, and inference tasks are bound to the online resource pool;
[0072] Furthermore, the elastic computing power resource allocation module can also perform intelligent resource scheduling. For example, dynamic resource quota: Based on historical load prediction, such as the time series algorithm Prophet, the training resources are automatically expanded at night, and online inference is preferentially guaranteed during the day; or, cost-aware scheduling: The training task automatically selects Spot instances and triggers task migration when the price fluctuates.
[0073] The hardware-aware optimization module is used to dynamically bind tasks to appropriate CPU or GPU cores through dynamic core binding technology. For example, CPU / GPU core binding is enabled for large-scale training tasks (such as the 70B model) to reduce the communication overhead of accessing non-uniform memory access (NUMA) nodes across different cores.
[0074] Figure 4 The resource scheduling diagram of the present invention is shown. Tasks are divided into online inference tasks and offline training tasks. Online inference tasks require fixed GPU resources, while offline tasks automatically expand instances at night to improve the overall resource utilization rate.
[0075] The embodiment reduces the video memory occupancy and deployment cost by dynamically loading multiple LoRA adapters and reusing the same basic large language model, realizing the sharing of the same basic large language model in multiple business scenarios. Integrating LoRA fine-tuning, parameter management, and service publishing in the cloud computing platform, the training and deployment processes are integrated, improving the development efficiency. Based on the elastic resource pool and priority policy of Kubernetes, efficient scheduling of mixed tasks (training / inference) is achieved, and the optimization of dynamic resource scheduling is completed.
[0076] Through experimental verification, the dynamic deployment of the multi-LoRA shared large language model of the present invention compared with the system solution that only uses the large language model for deployment:
[0077] 1. Deployment cost: For 10 business scenarios sharing the basic large language model, the video memory occupancy is reduced from 220GB (10x22GB) to 22.3GB, a reduction of 89.8%.
[0078] 2. Training efficiency: The AI-Pipeline shortens the model iteration cycle from 7 days to 1.5 days, an increase of 78.5%.
[0079] 3. Resource utilization rate: The GPU utilization rate is increased from 32% to 82%, and the task throughput is increased by 4.2 times.
[0081] Example 1: LoRA fine-tuning and deployment in multiple business scenarios
[0082] Step 1: Model Training
[0083] 1. The user logs in to the AI platform, creates a LoRA fine-tuning project, and selects a base large language model
[0084] (DeepSeek-R1-1.5B) and a dataset (such as CV-Product-Images).
[0085] 2. Configure LoRA parameters: rank R = 32, α = 64, training epochs = 50, trigger the distributed training task (4x4090 DGPU)
[0086] 3. After the task is completed, the platform automatically uploads the LoRA adapter file to the repository and records the metadata
[0087] Business label "CV-Product Recognition_v1"
[0088] Step 2: Model Management
[0089] 1. In the model service management interface, create a new inference service and associate the base large language model with the LoRA adapter.
[0090] 2. The platform generates a VLLM startup command and injects the adapter path:
[0091] vllm serve / models / deepseek-r1-1.5b--enable-LoRA--LoRA-modules
[0092] cv-brand= / LoRA / cv-hats-v1
[0093] 3. After the service is started, Deepseek-r1-1.5B (4GB) and the LoRA adapter (0.3GB) are loaded into the video memory, with a total occupancy of 4.3GB.
[0094] Step 3: Request Invocation
[0095] When the client sends a request, it specifies the LoRA identifier:
[0096] POST / infer HTTP / 1.1
[0097] X-LoRA-Id:cv-hats_v1
[0098] Content-Type:application / json
[0099] {"prompt":"What are the characteristics of this type of product?"}
[0100] Example 2: Elastic Resource Scheduling and Cost Optimization
[0101] Scenario: Coexistence of daytime online inference (QPS = 200) and nighttime batch training tasks.
[0102] Scheduling strategy:
[0103] 1. Online service: Fixed allocation of 2x GPUs (22GB video memory) to ensure P99 latency < 200ms.
[0104] 2. Training tasks:
[0105] Daytime: Use Spot instances (1x GPU) in the elastic resource pool with medium priority.
[0106] Nighttime: Expand to 4x GPUs with high priority to accelerate task completion.
[0107] 3. Resource recycling: Automatically release Spot instances after training is completed, reducing costs by 70%.
Claims
1. A multi-LoRA large language model deployment system based on a cloud computing platform, including a cloud computing AI platform layer, a multi-LoRA dynamic loading layer, and a resource scheduling optimization layer; The cloud computing AI platform layer is used to receive text task data, train a LoRA adapter suitable for the large language model for the inference process, and send the resource requirements of the LoRA adapter to the multi-LoRA dynamic loading layer; The multi-LoRA dynamic loading layer is used to store LoRA adapter files for each business scenario; send a resource scheduling application to the resource scheduling optimization layer according to the received resource requirements of the LoRA adapter; then dynamically switch the LoRA adapter to be mounted according to the request parameters, perform GPU optimization configuration according to the business priority in the request parameters, load the basic large language model, and mount several LoRA adapters superimposed in the form of a sparse matrix to complete the deployment of the LoRA model for this inference; After that, the LoRA model outputs a task response to the text task data; among them, when performing GPU optimization configuration, based on the service quality mechanism of the container orchestration platform, exclusive GPU resources are allocated to high-priority services, and medium- and low-priority tasks share the resource pool; The resource scheduling optimization layer responds to the resource scheduling application and outputs request parameters optimized by priority management and hardware awareness to the LoRA dynamic loading layer based on the container orchestration platform.
2. The system according to claim 1, wherein A hierarchical caching mechanism is adopted for LoRA adapter mounting. Frequently accessed LoRA adapters are stored in the GPU, and medium- and low-frequency accessed LoRA adapters are loaded on demand; medium-frequency accessed LoRA adapters are stored in the RAM memory, and low-frequency accessed LoRA adapters are stored on the disk.
3. The system according to claim 1, wherein, The cloud computing AI platform layer includes a dataset management module, a model training module, a model management module, and a model service management module; The dataset management module is used to support the unified access of multiple data sources; The model training module is used to generate a LoRA adapter suitable for the basic large language model through the AI-Pipeline engine that supports distributed training tasks; The model management module is used to manage the existing basic large language models; The model service management module is used to support the dynamic binding of the LoRA adapter and the basic large language model.
4. The system according to claim 3, wherein The dataset management module accelerates the caching of multiple data sources through the Fluid engine.
5. The system according to claim 3, wherein The multiple data sources include object storage service OSS data, network attached storage NAS data, and open data processing service ODPS data.
6. The system according to claim 1, wherein The container orchestration platform is Kubernetes.
7. The system according to claim 6, wherein The multi-LoRA dynamic loading layer includes a LoRA warehouse, a video memory optimization module, and a Vllm inference engine; The LoRA warehouse is used to store LoRA adapter files for each business scenario, record the basic large language model version and associated metadata; the associated metadata is a business scenario label; The video memory optimization module is used to adopt a hierarchical caching mechanism based on the Kubernetes service quality QoS mechanism. By monitoring the LoRA adapter call frequency, frequently accessed LoRA adapters are retained in the GPU, and medium- and low-frequency accessed LoRA adapters are persisted to the object storage service OSS; The Vllm inference engine is used to support the simultaneous loading of a base large language model and several LoRA adapters in the GPU through the integrated VLLM inference framework to form a LoRA model, and dynamically switch adapters through request parameters; output the task responses for text task data inferred by the LoRA model.
8. The system according to claim 1, wherein The resource scheduling optimization layer includes a priority management module, an elastic computing power resource allocation module, and a hardware-aware optimization module; The priority management module is used to schedule tasks based on a priority queue, divide the priorities of business tasks into three levels: high, medium, and low, and adopt the Gang Scheduling strategy to ensure the full resource allocation of distributed training tasks; The elastic computing power resource allocation module is used to automatically select a resource pool according to the task type. The task types include training tasks or inference tasks. Training tasks are preferentially scheduled to the elastic resource pool, and inference tasks are bound to the online resource pool; The hardware-aware optimization module is used to dynamically bind tasks to appropriate GPU cores through dynamic core binding technology.
Citation Information
Cited By
Large language model processing system and method based on prompt adaptive tuning
CN121457604A
AI platform blue-green deployment method and system based on dynamic routing and resource multiplexing
CN121478499A
LoRA fine-tuning computing power resource dynamic allocation method
CN121560567A