The invention relates to the technical field of
video processing, particularly discloses a fine-grained multi-mode AI video semantic
data processing method, and belongs to the technical field of
video processing and
content management. The method comprises the following steps: carrying out multi-
modal content analysis on an original video, and respectively extracting an
image frame, an audio Mel
spectrogram and a text transcription with a
timestamp; respectively extracting high-dimensional feature vectors of vision 768 dimension, audio 128 dimension and text 384 dimension by using a pre-trained
deep learning model; unifying the multi-
modal features to second-level time
granularity through
feature dimension standardization and time axis alignment, and splicing to generate a 1280-dimensional comprehensive
feature matrix; constructing an AI model comprising a bidirectional LSTM
time sequence feature extraction component and a fine-grained classifier, and training and classifying fusion features; and finally outputting a video file with structured classification information to realize second-level fine-grained semantic recognition. The method solves the problems of low multi-
modal information fusion efficiency, inconsistent
time sequence, high manual analysis cost and the like, and is suitable for safety supervision and
quality control in industrial scenes such as
building construction,
chemical engineering and the like.