The invention discloses a long-time-history
large model processing method based on asynchronous
reinforcement learning, and relates to the technical field of large language models.The method comprises the steps that the problem complexity is improved through dual-operation iteration, high-quality training data is obtained through triple
verification, a training end and a track generation end are decoupled through a two-stage track
parallel generation mechanism, and the problem is solved; after a training end collects a preset batch track, long-time-history
large model updating is started, then historical data is managed by depending on a basic tool set sub-model, end-to-end asynchronous
reinforcement learning optimization is executed on the whole process of an
intelligent agent, finally, the problem of long-time-history reward sparseness is solved based on a GRPO
algorithm, and a model strategy is optimized. According to the scheme, the core bottlenecks of weak long-range exploration, low training efficiency and insufficient
data quality of an open-source search intelligent body are solved, the search intelligence of
fuzzy query analysis, key
information extraction by
noise, cross-document reasoning
verification and super-long-range tool calling is unlocked, the limitation of a traditional search round is broken through, and a long-time-range task is supported.