The application discloses a 3D multi-target detection method based on semantic driving
single image for ATS, which processes input RGB images, extracts 3D bounding boxes of each object in the images, generates all potential 2D projections of 3D objects in a scene, extracts keywords, phrases and
semantic information in the description to form feature information P representing the language description t ; fuses 2D image information and 3D geometric information of the object to obtain complete object representation a ; associates language features extracted from the language description with the detected 3D object, captures the semantic correspondence between the text and the visual
modal; filters the target according to the generated matching
score to obtain all targets meeting the
natural language description. The application improves the accuracy and efficiency of retrieval and recognition, realizes higher recognition accuracy, significantly reduces the computational complexity, can improve the accuracy and speed of cross-
modal retrieval, and greatly improves the accurate recognition and positioning ability of traffic events.