The invention provides a multi-LoRA large
language model deployment
system based on a
cloud computing platform. A
cloud computing AI platform layer trains a LoRA adapter matched with a basic large
language model for a reasoning process; the multi-LoRA
dynamic loading layer dynamically switches LoRA adapters needing to be mounted according to the request parameters, GPU optimization configuration is carried out according to service priorities in the request parameters, a basic large
language model is loaded, and a plurality of LoRA adapters stacked in a
sparse matrix form are mounted; and the
resource scheduling optimization layer responds to the
resource scheduling application, and outputs the request parameters subjected to priority management and hardware sensing optimization to the LoRA
dynamic loading layer based on the container arrangement platform. According to the method, LoRA and container arrangement are deeply fused, and an automatic
assembly line of training and reasoning is achieved. A
video memory sharing mechanism enables a plurality of service scenes to share the same basic large language model, a LoRA adapter is loaded as required, and
video memory occupation is greatly reduced. Full-life-cycle management from
data preparation to model service is supported, and the
large model deployment cost is remarkably reduced.