The invention discloses a large
language model output
safety control system and a control method thereof. The
system comprises a
data acquisition module, a safety
processing subsystem, an output control module and a
reproduction and tracing data module. The
data acquisition module is used for intercepting, receiving or mirroring data interacted between a user side and the main
language model, and the interacted data at least comprises candidate output of
user input and / or main
language model output; the security
processing subsystem is used for receiving the interactive data and generating a
control signal and / or transformed content data based on at least one security
processing strategy, and the transformed content data
is security output content obtained by filtering,
rewriting, replacing or abstracting candidate output; the output control module responds to the
control signal and / or the transformed content data to perform at least one operation of terminating transmission, replacing / masking, outputting the transformed content data, or sending a protocol instruction including withdrawing, covering or rendering a failure mark to output a final response; and the
reproduction and tracing data module records
reproduction information for playback reproduction, strategy
rollback or scheduling parameter updating. Preferably, the candidate outputs include streaming token window segments or interlayer vector representations, and the
system supports sliding
window detection and dynamic review paths to defend
black box antagonism detection and achieve load balancing.