The invention provides a big
language model collaborative reasoning method and
system for an edge ad hoc network, and the method comprises the steps: S1, finding at least one available
inference engine in a
local area network through a point-to-point protocol, and obtaining the performance parameters of the
inference engine, the performance parameters comprising the calculation capability and an available memory; s2, according to the number of the
inference engines and the performance parameters, a parallel strategy is dynamically selected to allocate calculation tasks of the large
language model, and the parallel strategy comprises a hierarchical parallel strategy or a hierarchical collaborative strategy; s3, based on the parallel strategy, driving the
inference engine to cooperatively execute the calculation task so as to generate an inference result; and S4, managing an intermediate result of the calculation task through a distributed key value cache, and monitoring the state of the
inference engine so as to dynamically adjust task distribution. According to the method, the reasoning efficiency of the large
language model in the edge network is remarkably improved, the single-
machine load pressure is reduced, the robustness and expandability of the
system are enhanced, and the method is suitable for dynamically changing edge calculation scenes.