The invention discloses a
large model reasoning optimization method for SLO
perception in an edge heterogeneous computing power network, aiming at an edge heterogeneous computing power computing cluster scene providing
large model reasoning service, requests are strategically scheduled to make full use of heterogeneous computing power of a cluster, so that SLO heterogeneous requests are met to the maximum extent; the method comprises the following steps: firstly, separating a pre-filling stage and a decoding stage of a reasoning process to different nodes, and respectively adopting a data parallel deployment strategy and an
assembly line parallel deployment strategy to ensure that the computing power of equipment is fully utilized; then, according to heterogeneous computing power characteristics of decoding nodes, a distributed
assembly line non-uniform deployment strategy is realized to improve the model reasoning
throughput; finally, round rewards are collected in offline exploration, a
large model is used for assisting in encoding of potential rewards, a decoder is trained to decode agency rewards in each step, a
reinforcement learning PPO
algorithm is assisted in training a scheduler, effective scheduling requests are achieved, the
resource utilization rate is increased, meanwhile, cluster
energy consumption is reduced, and the
throughput of large model reasoning service is increased.