The disclosure relates to a method and
system for dynamically mitigating threats of generative
Artificial Intelligence (AI) models. Conventional systems often suffer from inefficiencies due to sequentially applying
threat detection checks leading to unnecessary preprocessing and increased computational demands. Additionally, such systems typically focus only on input data, neglecting potential threats in outputs. The disclosed
system and method addresses these drawbacks by employing a hierarchical structure of
macro and nano classifiers. The
system utilizes
macro classifiers for broad initial
threat categorization followed by specialized nano classifiers for detailed analysis of specific
threat subtypes, thereby optimizing
processing time and computational resources. The system operates in real time, applying predefined moderation rules to both input and output data to ensure comprehensive
threat mitigation. Additionally, continuous
telemetry data updates refine nano classifiers and threat identification mechanisms, maintaining high accuracy and adaptability. The disclosed method enhances safety efficiency and reliability of generative AI models.