This invention provides a
multimodal data annotation method and apparatus that combines frame-overlapping and single-frame
annotation, applicable to the field of
data processing technology. This application first acquires core
source data such as 2D images, 3D / 4D point clouds, target
object motion states, and dual coordinate systems. Relying on a real-time dual coordinate
system transformation algorithm, combined with displacement thresholds, it identifies the dynamic and static attributes of the target object, generating a classification-adapted
annotation base dataset. Then, it constructs multi-dimensional annotation association nodes according to dynamic and static classifications, generates a bidirectional mapping association graph through 2D-3D / 4D bidirectional projection, and obtains annotation space
verification feature vectors by mining
spatial consistency through a multi-coordinate
system verification model. These feature vectors are then imported into an optimization engine, combining static cross-frame propagation, dynamic trajectory calibration, and manual correction mechanisms to construct a multimodal annotation base model. Finally, it integrates real-time motion data and deviation correction signals for dynamic annotation, balancing accuracy and efficiency, to generate multimodal annotation information combining frame-overlapping and single-frame annotations adapted to complex scenes.