A
system may include a content repository that stores multi-
modal content comprising text, visual content, and / or audio content. The
system may include a processor programmed to: access a prompt comprising a
text query, a
visual query comprising text that describes a visual to be found, and / or an audio query comprising text that describes audio to be found, execute a
language model based on the prompt to identify content from the content repository, receive, from the
language model, a request for a
callback function that seeks additional information to satisfy the multi-
modal query, execute the
callback function to obtain the additional information and provide the additional information to the
language model in response to the request for the
callback function, re-execute the language model based on the multi-
modal query and the additional information, obtain, from the language model, content responsive to the prompt based on the additional information.