The invention relates to an end-to-end automatic driving method based on Transform multi-
modal feature fusion, and the method comprises the steps: collecting an
RGB image and an original depth image, and converting the original depth image into an HHA image; inputting the
RGB image and the HHA image into a sensing module for fusion to obtain a fusion feature map, and performing global average
pooling and flattening operation to obtain environment features; acquiring the real-time speed of a vehicle, an advanced navigation command and a target position, connecting in series to form measurement input, and
processing based on an MLP measurement
encoder to obtain measurement characteristics; adding the environment features and the measurement features
element by element to obtain combined features, performing downsampling step by step, and connecting a ReLU
activation function behind each layer to obtain track features; and a loss estimator is constructed, training loss is dynamically predicted in real
time based on the track features, and the weight of each
feature parameter is dynamically adjusted according to the loss, so that the vehicle is dynamically controlled and adjusted in real time. According to the method, the multi-
modal features can be well fused, so that a more accurate end-to-end automatic driving method is realized.