A
system and method for enhancing static photographs to simulate motion using
machine learning models. The
system comprises one or more processors and a non-transitory computer readable medium storing instructions that, when executed, cause the
system to receive a static photograph, apply a
machine learning model to generate a sequence of modified images creating an illusion of motion, and display the sequence to simulate a moving video. The
machine learning model, trained on a dataset of static photographs and corresponding video sequences, extracts features using a
convolutional neural network, processes the features with a
recurrent neural network to generate motion vectors, and applies the vectors to create the modified images. The simulated motion may include tilting, vibrating, shaking, zooming, panning, and rotating. A
user interface allows specifying the desired type or intensity of motion. The method enables creating video-like effects from static images, enhancing expressiveness and engagement of
visual media.