Remote sensing target detection sample generation method and system based on multi-modal large model

By combining a multimodal large language model with a diffusion model, a remote sensing image target detection sample generation method is proposed. This method addresses the issues of sample scarcity and semantic logic deficiency in remote sensing target detection, generating high-quality samples that conform to physical laws and improving detection performance.

CN122391907APending Publication Date: 2026-07-14WUHAN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610529879.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-04-21
Publication Date
2026-07-14

AI Technical Summary

Technical Problem

Existing remote sensing target detection technologies suffer from problems such as sample scarcity, missing semantic logic, chaotic perspective relationships, and harsh background fusion when acquiring scarce samples, resulting in poor detection performance in fields such as natural disaster emergency rescue.

Method used

A remote sensing image target detection sample generation method based on multimodal large language model (MLLM) is adopted. Combined with diffusion model, realistic target detection samples are generated through multi-task learning framework and dedicated multimodal inference model, including placement center prediction, scale prediction and angle prediction, and boundary fusion and illumination coordination are performed.

Benefits of technology

The generated samples conform to the semantic and physical laws of complex scenes, which improves the detection accuracy of scarce categories, solves the problem of sample scarcity, and improves the performance of the detection model in complex scenes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122391907A_ABST
    Figure CN122391907A_ABST
Patent Text Reader

Abstract

The application discloses a remote sensing target detection sample generation method and system based on a multi-modal large model, first, instruction fine-tuning data sets suitable for multi-modal large language model training are acquired; and a special multi-modal inference model facing remote sensing target placement is constructed, and fine-tuning training is performed on the special multi-modal inference model; a pure background image to be enhanced and a foreground reference image of a target category are input into the fine-tuned special multi-modal inference model, a placement strategy set containing a coordinate position, a suggested size and a corresponding reference foreground index is output through multi-round inference; finally, a diffusion generation model is constructed, the output placement strategy set is taken as a geometric control condition, and a reference foreground image set is taken to provide appearance guidance, target generation, boundary fusion, local light coordination and shadow compensation are performed in a specified area of the pure background image, and a final synthetic sample is generated.
Need to check novelty before this filing date? Find Prior Art