The application relates to a
remote sensing image description generation method based on external
knowledge retrieval enhancement, which comprises the following steps: step 1, visual
feature coding; step 2, external
knowledge retrieval enhancement; step 3, cross-
modal generation decoder; and step 4, multi-task joint optimization. The application has the beneficial effects that: the retrieval enhancement generation paradigm is used; through manifold alignment based on
principal component analysis and a diversity
perception reordering strategy, high-confidence semantic knowledge is efficiently extracted and fused from an external
knowledge base, and the limitation of a parameterized model in fine-grained semantic
cognition is effectively made up. Meanwhile, a cross-
modal attention entropy regularization mechanism is adopted, the cross-
modal alignment quality in the generation stage is improved, the attention collapse phenomenon is inhibited, and the model is encouraged to balance attention to
global information of an image. In addition, a joint
loss function is used in combination with a focal loss and length normalization, and the long-
tail vocabulary prediction and short
sentence generation problems are systematically alleviated.