The invention discloses a roadside multi-mode fusion sensing method. According to the method, an initial 3D view cone
tensor with depth information is generated through a depth
estimation network, features of pixel points are projected to a coordinate
system in combination with internal and external parameter matrixes of a camera, then a three-dimensional space point coordinate set is generated according to projected feature position information, a
Gaussian function is adopted to distribute weights of neighborhood grids, and a three-dimensional space point coordinate set is obtained. And finally, inputting the image features and the features into a fusion unit, respectively carrying out uncertainty modeling on the two types of features, automatically generating a fusion weight, and finally, carrying out normalization weighted aggregation on the features in the neighborhood so as to reduce the position
approximation error and obtain the calibrated image features, and finally, inputting the image features and the features into the fusion unit, and respectively carrying out uncertainty modeling on the two types of features and automatically generating a fusion weight. And dynamic optimization is carried out according to a self-supervised
loss function in the fusion process, so that the stability of multi-
modal feature fusion is ensured, the detection accuracy is high, the performance is relatively good, and the problem of global feature
dislocation caused by external parameter deviation and depth error can be avoided.