This invention relates to a method, apparatus, and
system for security evaluation of large language models based on dynamic custom logic encoding. It constructs a temporary
logic mapping protocol for encoding; after acquiring malicious semantic text to be tested, it encodes the malicious semantic text into a sequence of logical instructions according to the temporary
logic mapping protocol, constructs test prompts, inputs them into the large
language model under test, and determines whether the large
language model under test has been jailbroken based on the output. The apparatus includes an acquisition module, a protocol construction module, an encoding module, a construction module, an output acquisition module, and a judgment module; the
system is implemented based on the method. This invention can bypass security protection mechanisms based on blacklists and fixed
pattern recognition, achieving a high avoidance rate, breaking through static defense mechanisms, touching security vulnerabilities at the
logical reasoning level, possessing high adaptability and
scalability, ensuring the long-term effectiveness of the evaluation, accurately locating weak points in the model's security strategy, and performing targeted and refined reinforcement.