The application discloses a performance evaluation and
adaptive optimization method and
system for a visual-
language model, constructs an open set
metadata set, for any
data set, extracts image and text embedding by using the visual and text encoders of the visual-
language model, and obtains class-level visual prototypes and text prototypes based on cross-
modal similarity aggregation, to form the cross-
modal structure identifier of the
data set; a double-level accuracy estimator containing a representation
bottleneck and a structure
bottleneck is constructed, and the structure identifier is mapped to a predicted accuracy; on unlabeled data in a target domain, the predicted accuracy is used as a differentiable global supervision
signal, and only the affine transformation parameters of the normalization layer in the visual-
language model are optimized through gradient back propagation, to realize efficient
adaptive optimization of parameters. The application deeply fuses the accuracy
estimation and model optimization, forms a
closed loop of "
performance estimation-model optimization", and improves the generalization ability and reliability of the visual-language model in an open environment.