The invention discloses a large
language model security testing method based on evolutionary dynamic adversarial attacks, and relates to the technical field of
security testing. Comprising the following steps: 1, creating a dynamic
attack generation engine, constructing an adversarial evolution architecture by utilizing the dynamic
attack generation engine, and generating a
test sample based on the adversarial evolution architecture for a large
language model security test; the method comprises the following steps: 1, establishing a large-scale
language model, 2, calculating investigation parameters of harmlessness, honesty and helpfulness based on a Constancy AI principle, calculating a security alignment gap index SAGI by using the investigation parameters, and intelligently judging a value drift condition of the large-scale language model according to the security alignment gap index SAGI, 3, establishing a multi-
modal joint defense engine, and carrying out intelligent judgment on the value drift condition of the large-scale language model according to the value drift condition of the large-scale language model.
Steganalysis,
syntax tree analysis and audio
anomaly detection functions are integrated, and all-around
threat detection coverage of texts, codes, images and voices is carried out on a detected large language model.