The application discloses a three-dimensional target tracking method,
system, device and medium based on space-time enhancement, and relates to the technical field of automatic driving. The method comprises the following steps: inputting three-dimensional
point cloud sequence data into a trained SMTrack network for
processing to obtain a final bounding box prediction result. The SMTrack network comprises a target-specific
encoder, an STFM module and an STT module connected in sequence. The target-specific
encoder is used for extracting target-specific features from the
point cloud sequence. The STFM module is used for modeling appearance information and motion information in stages according to the target-specific features to generate a preliminary bounding box prediction result. The STT module is used for optimizing the preliminary bounding box prediction result. The application adopts a twin structure, extracts features from historical frame and current frame sequences, and introduces part
perception motion modeling and coarse-to-fine bounding box regression mechanisms, so that the target positioning accuracy can be greatly improved.