The invention discloses a
remote sensing image open vocabulary segmentation method and
system based on vision-language interaction. According to the method, the advantages of a multi-
modal large
language model and a semantic segmentation network based on a
visual basic model are cooperatively utilized, language-pixel two-way mapping is taken as a core, five types of marks of images, texts, categories, objects and segmentation are introduced as carriers of cross-
modal information, and three types of cross-
modal fine-grained information interaction modules are taken as bridges, so that the cross-modal information interaction is realized. Bidirectional mapping and alignment of fine-grained information of the multi-modal large
language model and the semantic segmentation network are realized, the open vocabulary segmentation capability of the semantic segmentation model is improved, and the method can adapt to different
remote sensing scenes and category definitions. The method has the following advantages: the method has high performance, strong generalization ability and good expansibility, can provide any category of semantic segmentation maps for unlabeled target domain images based on instructions, and has high application value in the aspects of
urban planning,
map making, disaster response and the like.