The invention discloses a fast and quality-aware cloud collaborative large
language model reasoning method and
system (CEC-LM), and belongs to the field of computer networks and
artificial intelligence, and the method comprises the steps: employing a task execution selector based on a lightweight model, evaluating the task complexity in real time according to a prompt word requested by a user, and carrying out the real-time evaluation of the task complexity according to the prompt word; the complex tasks are routed to a cloud LLM to ensure response quality, and meanwhile the simple tasks are distributed to a local small
language model to be executed. And when sensing that the local SLM is high in load, dynamically unloading the complexity prediction model from the GPU to the CPU, and realizing zero-copy parameter transmission by utilizing a unified
memory architecture. According to the method, a task monitoring and
priority scheduling mechanism during operation is introduced, the task routing is adaptively adjusted by continuously monitoring the network condition and the
system load, and the emergency task is preferentially processed, so that the achievement rate of the service level target is ensured, the reasoning quality,
delay and cost are effectively balanced, and the
system overhead is remarkably reduced.