The application discloses a multi-
modal three-dimensional target detection method based on structural feature and
semantic feature fusion, and belongs to the technical field of target detection.The application solves the problems of low precision, poor real-time performance and insufficient robustness of the existing method.The application constructs an explicit matching
fusion mechanism of low-layer structure guided initialization, high-layer semantic optimization and
time sequence expansion, projects 3D boundary boxes of a
laser radar and a camera to multi-view image planes, calculates
geometric similarity and category consistency constraints of 2D projection regions, and constructs an explicit matching graph in combination with low-layer structure information and high-layer semantic features.The application guides weighted aggregation of cross-
modal target features with high confidence through a sparse matching graph, aligns historical frame targets to a current frame coordinate
system to construct a space-time matching graph, aggregates historical information through expansion to a
time sequence dimension, and realizes efficient and accurate three-dimensional target detection of the
laser radar and the visual multi-
view camera.The application method can be applied to multi-
modal three-dimensional target detection.