Natural-language queries are matched to LLM-generated video captions to detect relevant vehicle events without retraining static dashcam models.
Text prompts are converted into filtering objects that guide video detectors, improving object identification without manual review.