This invention provides a cross-
modal spatiotemporal event prediction method,
system, computer, and storage medium. The method includes the following steps: receiving a multimodal high-level structured
semantic feature sequence of
monitoring data output from multiple edge nodes; mapping each
semantic feature in the feature sequence to a unified embedding space; summing the weighted results of each modality according to a gating
fusion mechanism to obtain joint representation parameters; performing average
pooling operation on the joint representation parameters corresponding to the
time step within the current
sliding time window, and inputting the result into a lightweight prediction head to output decision parameters for the probability of event occurrence. By implementing spatiotemporal modeling through a
sliding time window and a memory mechanism, the robustness and adaptability of the
system are enhanced, detection sensitivity is improved, and false alarms are reduced. This solution, with its low
power consumption, low latency, and high accuracy, solves the core bottlenecks of high computational cost, large latency, and poor
scalability in existing technologies, providing a sustainable solution for
smart city monitoring.