视频生成方法、装置、设备、计算机可读介质和程序产品
By generating a candidate video set through a pre-trained multimodal large language model and selecting a subset of videos with high target accuracy, and combining text modal feature information, the problems of high computational complexity and insufficient accuracy in existing technologies are solved, achieving efficient and accurate video localization.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- JD DIGITS HAIYI INFORMATION TECHNOLOGY CO LTD
- Filing Date
- 2025-09-29
- Publication Date
- 2026-07-17
AI Technical Summary
Existing technologies suffer from high computational complexity and insufficient accuracy in video localization, making it difficult to effectively utilize multimodal large language models for accurate video localization.
A candidate event location video set is generated by a pre-trained multimodal large language model, and a subset of candidate event location videos is selected by using the target hit accuracy. The event location videos are further filtered by combining video description text and event description text, which reduces the computational load of the model and improves the accuracy.
While reducing computational complexity, it improves the accuracy and efficiency of video localization, ensuring that the event localization video that best matches the event description text is selected.
Smart Images

Figure CN121334458B_ABST