This invention discloses a unified five-
modal target detection method, relating to the fields of
computer vision and multimodal target detection technology. This method targets visible light,
infrared,
synthetic aperture radar, multispectral, and hyperspectral images. First, it performs unified
annotation and three-channel input unification
processing on different
modal inputs. Then, it constructs a unified target detection network comprising five modality-specific input adapters, a shared
backbone network, a shared
feature fusion network, and five modality-
specific detection heads. Based on the modality identifier of the input sample, the corresponding input adapter and detection head are selected to achieve modality-specific input correction, shared
feature extraction, and modality-
specific detection output. During the
training phase, a round-robin single-batch single-modality joint training mechanism is adopted, and the
loss function of the detection head corresponding to the current modality is calculated only. During the
inference phase, the target detection result for the corresponding modality is output. This invention can achieve target detection of five heterogeneous
modal images within a unified framework, possessing good versatility and
engineering application value.