The invention discloses a
remote sensing image indication segmentation method based on multi-
modal large
language model enhancement
fine tuning, and aims to realize high-precision and high-generalization pixel-level
remote sensing target segmentation of any
natural language instruction. According to the method, a two-stage enhanced
fine tuning framework of reasoning and segmentation decoupling is constructed, and a
remote sensing task-oriented semantic positioning alignment
hybrid verifiable reward function is combined, so that the deep reasoning and space understanding capabilities of a multi-
modal large
language model are effectively excited. The method comprises the following specific steps: 1, reading a remote sensing image and corresponding
natural language instruction data; 2, performing supervised
fine tuning (SFT) based on a pre-trained multi-
modal large
language model to obtain a preliminary alignment capability; 3, designing an
inference model and segmentation model collaborative architecture, generating a structured thinking chain and a position prompt by the
inference model and segmentation model collaborative architecture, and outputting a high-precision
mask based on an SAM architecture by the segmentation model collaborative architecture; 4, a group relative strategy optimization (GRPO)
algorithm is introduced, and enhanced fine tuning is carried out in combination with triple rewards of format, positioning accuracy and remote sensing
semantic consistency; and 5, executing end-to-end indication segmentation, and outputting a pixel-level binary
mask. The method breaks through a traditional closed set segmentation normal form, has the advantages of open vocabulary understanding, few sample generalization, interpretable reasoning and the like, and remarkably improves the
automation level and
adaptive capacity of remote sensing intelligent interpretation.