The invention provides a spatial intelligent visual physical process
inference method based on an implicit physical
large model, and belongs to the field of spatial intelligent and
artificial intelligence modeling calculation. The problems of low prediction precision and lack of physical consistency for complex physical scenes in the prior art are solved. The method comprises the following specific steps: acquiring multi-
modal data of an environment, and representing the multi-
modal data in a unified coordinate
system; preprocessing the multi-
modal data, designing a geometric coding model, a
dynamic prediction model and an
energy conservation constraint, introducing a space-time attention mechanism, predicting the motion state of an object according to the multi-
modal data, and obtaining prediction data of the motion state of the object; acquiring real
observation data of
object motion, comparing the difference between the prediction data and the real
observation data, and performing correction and weight updating on the prediction data of the model through a self-adaptive residual term; according to the method, visual information and implicit
physical law modeling are fused, and the self-supervised physical consistency constraint is utilized, so that
automatic learning and prediction of a potential physical process are realized.