Remote sensing target detection sample generation method and system based on multi-modal large model
By combining a multimodal large language model with a diffusion model, a remote sensing image target detection sample generation method is proposed. This method addresses the issues of sample scarcity and semantic logic deficiency in remote sensing target detection, generating high-quality samples that conform to physical laws and improving detection performance.
Patent Information
- Application Number
- CN202610529879.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-04-21
- Publication Date
- 2026-07-14
AI Technical Summary
Existing remote sensing target detection technologies suffer from problems such as sample scarcity, missing semantic logic, chaotic perspective relationships, and harsh background fusion when acquiring scarce samples, resulting in poor detection performance in fields such as natural disaster emergency rescue.
A remote sensing image target detection sample generation method based on multimodal large language model (MLLM) is adopted. Combined with diffusion model, realistic target detection samples are generated through multi-task learning framework and dedicated multimodal inference model, including placement center prediction, scale prediction and angle prediction, and boundary fusion and illumination coordination are performed.
The generated samples conform to the semantic and physical laws of complex scenes, which improves the detection accuracy of scarce categories, solves the problem of sample scarcity, and improves the performance of the detection model in complex scenes.
Smart Images

Figure CN122391907A_ABST