一种基于知识增强的迭代式自优化文生视频方法及系统

By constructing a static physical knowledge base and a dynamic constraint memory, and combining static appearance verification with dynamic physical verification, the problem of implicit understanding of physical rules in the T2V model is solved, and explicit injection of physical rules and reuse of historical experience are realized, thereby improving the physical fidelity and generalization ability of video generation.

CN122138025BActive Publication Date: 2026-07-17SHANDONG UNIV

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
SHANDONG UNIV
Filing Date
2026-05-06
Publication Date
2026-07-17

AI Technical Summary

Technical Problem

Existing T2V models lack an explicit understanding of physical rules, resulting in physical illusions in generated videos. Furthermore, they lack dynamic memory and experience reuse mechanisms across sessions, failing to meet the physical fidelity requirements of professional scenarios.

Method used

We construct an iterative self-optimizing text-based video system based on knowledge enhancement. Through a static physical knowledge base and a dynamic constraint memory, we achieve explicit injection of physical rules and cross-use reuse of historical experience. By combining static appearance verification and dynamic physical verification, we form a closed loop of knowledge retrieval, prompt generation, verification scoring, and memory accumulation.

Benefits of technology

It significantly improves the physical rule compliance and temporal coherence of video generation, enhances the generalization ability of out-of-distribution scenes, and achieves continuous iterative improvement in the quality of system generation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122138025B_ABST
    Figure CN122138025B_ABST
Patent Text Reader

Abstract

本发明提供了一种基于知识增强的迭代式自优化文生视频方法及系统,涉及视频生成技术领域,包括:获取用户原始文本输入,分别从静态物理知识库中匹配物理规则集,从动态约束记忆库中召回历史约束集,执行物理规则约束、记忆融合,生成知识增强后的视频生成提示词;将生成的视频生成提示词输入到预训练的扩散式T2V模型中,生成目标视频;对目标视频执行静态外观校验与动态物理校验,并计算语义一致性评分与物理常识评分;基于校验结果和评分结果,判断是否满足迭代终止条件,若满足则输出最终视频及优化提示词;本发明形成“知识检索‑规划生成‑校验反馈‑记忆沉淀”的完整闭环,显著提升分布外场景的泛化能力,实现系统生成质量的持续迭代提升。
Need to check novelty before this filing date? Find Prior Art