The present application relates to the technical field of automatic driving end-to-end
perception, in particular to an end-to-end automatic driving long-
tail recognition method based on contrast learning pre-training, first, a synthetic image data with long-
tail distribution characteristics is generated through a conditional
diffusion model, then a fine-grained scene classifier is used to systematically organize and semantically
label the generated samples, and a structured multi-
modal graph-
text alignment dataset is constructed; finally, the enhanced dataset and the original
training set are fused, the visual-linguistic joint embedding space is optimized through a multi-task contrast
loss function, and the parameter update of the pre-training model is realized. The method innovatively establishes a closed-
loop optimization mechanism of generative data enhancement and contrast learning framework, effectively alleviates the data scarcity problem under the long-
tail distribution scene, and significantly improves the cross-
modal representation ability and downstream task generalization performance of the model on low-resource classes.