The invention discloses a multi-level cache and self-feedback driven retrieval enhancement generation reasoning acceleration
system. The
system comprises an input
processing module, a multi-level
cache management module, a retrieval enhancement generation module and a self-feedback optimization module. Wherein the multi-level
cache management module adopts a three-level cache structure and is used for pre-filling and caching completely matched historical interaction data, similar matched associated reasoning results and hotspot knowledge of a
knowledge base, and redundant calculation is reduced by
multiplexing cache data in the reasoning stage; the retrieval enhancement generation module has dual functions: on one hand, a cache or a
knowledge base is retrieved based on a query vector generated based on a user request, a
retrieval result and
user input are fused to generate an enhancement cue word, and a large
language model is driven to perform reasoning so as to generate targeted output; on the other hand, when there is no user request, the hot knowledge of the
knowledge base is actively recognized, knowledge arrangement type cue words are generated through a preset template, and a large
language model is driven to generate preprocessing results (such as a core abstract and a preset question and answer pair) and store the preprocessing results in a three-level cache. And the self-feedback optimization module dynamically adjusts retrieval parameters, generation parameters and a cache strategy by evaluating the accuracy, correlation and timeliness of a generated result to form closed-
loop optimization, so that invalid calculation and redundant information generation are reduced, the reasoning speed and the output quality are improved, and the collaborative efficiency of a large
language model is remarkably improved.