The invention discloses a cross-
modal target detection method based on learnable
Fourier transform, and mainly solves the problem of insufficient fusion of a visible light image and an
infrared image in a complex scene due to inter-domain difference in the prior art. According to the implementation scheme, the method comprises the steps that bimodal features are extracted through a double-flow CSPDarknet53 network; a target position guiding module is utilized to enhance target area representation and suppress background interference; the features are converted to a
frequency domain, and amplitude texture information of the visible light image and phase contour information of the
infrared image are adaptively enhanced through a learnable
frequency domain feature enhancement module; suppressing
noise through global filtering and then inversely transforming back to a
spatial domain; and finally, outputting a target detection result of the multi-
modal image by the detection head. According to the method,
frequency domain physical characteristics are fully utilized, full
complementation and adaptive fusion of cross-
modal features are realized, the detection precision and robustness of vehicles, pedestrians and other targets under low-illumination and complex backgrounds are remarkably improved, meanwhile, high calculation efficiency is kept, and the method can be applied to the fields of automatic driving, intelligent monitoring and the like.