The invention discloses a long video content
information acquisition method based on an OCR and voice recognition technology. The method comprises the following steps: S1, carrying out preprocessing n on input long video data to extract an
image frame sequence and an audio
stream; s2, inputting the
image frame sequence into an OCR recognition module, inputting the audio
stream into an ASR recognition module, and obtaining a preliminary recognition result; s3, constructing a multi-target
fitness function, and optimizing an OCR and ASR parameter combination by using a Kanglizard optimization
algorithm; s4, respectively applying the optimal parameter group to an OCR identification module and an ASR identification module to obtain an optimized identification result; s5, constructing a fusion
factor graph, executing edge
message passing by adopting a
belief propagation algorithm, and generating a multi-
modal semantic block set; and S6,
processing the multi-
modal semantic block set to generate a unified multi-
modal content information set. According to the invention, through fusion of the horny
lizard optimization
algorithm and the
belief propagation mechanism, high-precision recognition and multi-modal
semantic consistency extraction of the image text and the voice information in the long video are realized.