The application discloses a script-style
record generation method and
system based on cross-
modal semantic reasoning, comprising the following steps: S1, synchronously collecting and preprocessing audio
stream, video
stream and auxiliary text data in an interrogation scene, and outputting audio
stream A'(t), video
frame sequence V'(x, y, t) and
formatted text T(n); S2, extracting audio stream A'(t) and video
frame sequence V'(x, y, t) features in parallel, and outputting text sequence Text_Seq and visual feature sequence; S3, performing semantic conversion on the visual feature sequence to generate visual description sequence Vis_Seq described in
natural language; and S4, performing
time sequence alignment and semantic fusion on the text sequence Text_Seq and the visual description sequence Vis_Seq to generate a comprehensive
record sequence T_comprehensive. The application generates script-style records with narrative logic and easy-to-understand, breaks the stereotyped mode of template generation, can clearly show the strategy adjustment of interrogation, the emotional change of the parties and the key turning point of the case, and makes the records more in line with the reading and cognitive habits of human beings.