The invention provides a multi-
modal data annotation method and device combining overlapped frame
annotation and
single frame annotation, and is applied to the technical field of
data processing. According to the method, core
source data such as a 2D image, a 3D / 4D
point cloud, a target
object motion state and a double-coordinate
system are firstly obtained, dynamic and static attributes of a target object are recognized by means of a double-coordinate
system real-time conversion
algorithm in combination with a displacement threshold value, and a labeling basic
data set matched in a classified mode is generated. The method comprises the following steps: firstly, constructing a multi-dimensional
label association node by pressing static classification, generating a bidirectional mapping association graph through 2D-3D / 4D bidirectional projection, and mining space consistency through a multi-coordinate
system verification model to obtain a
label space
verification feature vector; the method comprises the following steps of: selecting a multi-
modal labeling model, importing the multi-
modal labeling model into an optimization engine, constructing a multi-modal labeling basic model in combination with static cross-frame propagation, dynamic track calibration and a manual correction mechanism, and finally fusing real-time motion data and deviation correction
signal dynamic labeling, balancing precision and efficiency to generate multi-modal labeling information which is matched with stacked frame and
single frame combination of a complex scene.