The invention discloses a generative zero sample learning method based on two-state collaborative decoupling and semantic refining, and belongs to the field of generative zero sample learning. According to the method, static features such as a background, a structure and details of an image are decoupled, and dynamic common features are extracted by a cross-
modal label generation module, so that complementary expression of the static and dynamic features is realized; the feature focus is dynamically adjusted according to different
confusion types, the feature discrimination is enhanced, and cross-class interference is relieved; on the semantic level, by constructing a vision-semantic
mirror image cross attention mechanism, bidirectional alignment between semantic features and visual features is achieved, and the multi-
granularity capability and adaptability of
semantic representation are further improved. According to the method,
feature structure decoupling,
confusion adaptive regulation and control and dynamic
semantic alignment are taken as the core, the cross-category generalization ability and the generated
sample quality are effectively improved, the performance
bottleneck of traditional generative zero sample learning is broken through, and the method has high theoretical value and wide application prospects.