The invention discloses a sparse large
language model weight refinement method,
system and device and a storage medium, which are corresponding schemes, and the scheme realizes training-free and plug-and-play refinement of sparse weight after one-time
pruning, significantly reduces the
confusion degree, improves the zero sample task performance, and especially has an outstanding effect in a high sparse rate scene. The space is optimized in a unified mode through global soft constraint, heterogeneous Hessian inversion is avoided,
pruning prior is effectively utilized in combination with historical
momentum, Neumann series extrapolation solution is adopted, iteration time can be remarkably shortened, and expenses of computing resources and memory resources are reduced. The method is oriented to online and operation and maintenance processes of
inference services such as
text generation, question and answer, classification and retrieval enhanced question and answer and the like, and accesses the current network through hot update or gray release after completing intra-layer weight refinement; the existing word segmentation coding,
service gateway and monitoring and
rollback mechanisms are reused in the whole process, and it is ensured that the method can be implemented under real service distribution and resource budget.