This invention belongs to the field of
artificial intelligence security technology, specifically relating to a corpus security detection method,
system, and computer-readable storage medium. The corpus security detection method includes: S1, preprocessing the original corpus to obtain
paragraph-level corpus text; S2, inputting the
paragraph-level corpus text into a five-dimensional
detector to obtain a five-dimensional risk
score; wherein, the five-dimensional
detector includes a harmful information classifier, a sensitive
entity identifier, a low-quality scorer, a false statement verifier, and a confidential
fingerprint comparer; S3, weighting and fusing the five-dimensional risk scores to obtain a comprehensive risk
score; S4, determining whether the comprehensive risk
score is greater than a blocking threshold; if so, blocking is performed. This invention can perform integrated real-time detection of harmful, sensitive, low-quality, false, and confidential information in large-scale pre-training corpora before training, achieving real-time blocking before streaming into the
database, and is applicable to corpus cleaning at various stages such as
large model pre-training, fine-tuning, and continuous pre-training.