The invention provides a training teaching video intelligent analysis and knowledge point automatic marking method based on multi-
modal fusion, and the method comprises the steps: collecting multi-
modal data in a training teaching process, and carrying out the
time alignment processing; extracting feature information of the
modal data, fusing the feature information through a cross-modal fusion architecture, and generating a unified teaching behavior representation vector; based on the teaching behavior representation vector, identifying operation steps in the practical teaching process through a
sequence labeling model, and determining the category and the starting and ending time boundary of each operation step; matching the identified operation steps with a preset skill
knowledge base, and generating a standardized knowledge point
label containing knowledge point content, starting and ending timestamps and confidence information; and storing the standardized knowledge point labels into a
database, and constructing a retrieval index. According to the invention, the unstructured teaching video is converted into a searchable, navigable and analyzable
knowledge unit, and automatic identification and structured marking of practical teaching operation are realized.