This invention discloses a zero-shot adversarial robust model, method, and
computer device based on text information enhancement. First, text descriptions and
random text are generated to enhance semantic expression diversity. Features of images and different texts are extracted using a visual
language model. Adversarial examples are iteratively generated using image features and text description features, and used for adversarial training. Cross-entropy loss is calculated using the image features and text description features of the adversarial examples to
train the adversarial model. The
cosine similarity between
random text features, adversarial example features, and clean example features is calculated separately, and the two are aligned using a loss. To maintain zero-shot performance on clean examples, the
cosine similarity between
random text features and clean example features in the original and target models is calculated, and the two are aligned using a loss. The three losses are integrated to obtain a zero-shot adversarial robust model based on text information enhancement. This method can improve the adversarial robustness and generalization performance of the model under zero-shot conditions.