The invention relates to the technical field of big
language model dialogue
system optimization, in particular to a big
language model dynamic dialogue history
compression method and
system.The method comprises the steps that the maximum length of a
context window matched with a target big
language model, the maximum number of newly-generated lexical elements and the size of a safety buffer area are set, and then dialogue history is loaded; initial compression and
verification are carried out through a keyword and TF-IDF mixed
scoring system, multiple times of dynamic compression are carried out according to gradients if the conditions are not met, and finally, parameters are adjusted to adapt to the
residual space when a model is called to generate response. The
system comprises a dialogue history loading module, a parameter configuration module, a dynamic compression engine module and a large language
model integration module. Through a multi-stage compression
verification mechanism and a progressive multi-stage dynamic compression strategy, super-long texts such as
engineering technology documents can be processed, service interruption is reduced, the compression efficiency is improved on the premise that key
semantics are reserved, and multi-language dynamic compression is supported. The problems of system token overrun and service
instability in the prior art are solved.