An automated
video editing system facilitates the creation of non-linear editing (NLE) timelines using a single prompt from a user. The
system automatically ingests digital media, including video, audio, text, and images, and processes them to generate a proxy version with extracted features such as
speech transcription, shot detection, facial recognition, and
text recognition. A prompt-driven editing engine interprets
user input and generates an edit
decision list (EDL) using a large
language model, which guides the
assembly of an edited video timeline. The
system also applies advanced editing features-such as captioning, animated title cards,
font and color styling, sound effects, and transitions-based on learned user preferences. Additionally, it enables contextual overlays, chapter cards, and hierarchical timelines, while continuously learning user preferences to personalize editing results. The system may be integrated with
social media platforms and existing media libraries for content sourcing, customization, and automated publishing.