The invention discloses a cross-
modal target detection
system and method, and relates to the technical field of
computer vision, and the method comprises the steps: collecting a visible light image and a
point cloud, carrying out the calibration of an internal reference and an
external reference, projecting the
point cloud to a pixel grid to generate a
depth map, calculating the gradient of the visible light image, the gradient of the
depth map and the
point density of the
point cloud, and constructing a geometric reference domain. Bidirectional mapping is obtained; according to a geometric reference domain and bidirectional mapping, forming a general bottom layer feature, obtaining a two-mode feature graph through lightweight
adaptation, and generating a pixel-level credibility graph; determining a dominant mode according to the pixel-level credibility graph, constructing an initial fusion field, constraining a propagation boundary through a geometric reference domain, executing edge preserving linear fusion and
anisotropic diffusion, and forming a fusion feature
pyramid; and generating anchor points on the fusion feature
pyramid according to the fusion features, the geometric reference domain and the pixel-level credibility map, generating a target candidate frame, performing boundary correction, and outputting the target position and category. And multi-scale positioning stability is realized.