This invention discloses a method and
system for multimodal video acquisition and tagging indexing in fire scenes. The method includes: S1: Deploying a multimodal
camera array and
sensor array, and synchronously acquiring video,
smoke concentration, and temperature data through a hardware synchronization trigger; S2:
Spatial registration of the video using a pre-calibrated
homography matrix; S3: Adaptively adjusting weights based on
smoke concentration to generate a fused video; S4: Automatically generating labels for fire stage,
fuel type, and key moments based on temperature data and video features; S5: Storing the original video, fused video, temperature data, and labels in a three-
level structure in a
database, establishing a
time index and a
label index. This invention can achieve
microsecond-level synchronous acquisition of multimodal videos such as visible light, near-
infrared, and
thermal infrared, while automatically identifying fire stage,
fuel type, and key events, generating structured labels without manual intervention; providing a high-quality, multi-dimensional, and searchable experimental data foundation for fire
mechanism analysis, firefighter case teaching, and fire identification.