A knowledge-enhanced iterative self-optimizing text-to-video method and system
By constructing a static physical knowledge base and a dynamic constraint memory, and combining static appearance with dynamic physical verification, the problem of implicit understanding of physical rules in the T2V model is solved, and explicit injection of physical rules and cross-session reuse of errors are realized, thereby improving the physical fidelity and generalization ability of video generation.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SHANDONG UNIV
- Filing Date
- 2026-05-06
- Publication Date
- 2026-06-02
AI Technical Summary
Existing T2V models lack an explicit understanding of physical rules, often resulting in physical illusions that violate gravity, fluid dynamics, and optical principles when generating videos. Furthermore, they lack dynamic memory and experience reuse mechanisms across sessions, leading to low physical fidelity, weak generalization ability, and unsustainable optimization.
We construct an iterative self-optimizing text-based video system based on knowledge enhancement. Through a static physical knowledge base and a dynamic constraint memory, we achieve explicit injection of physical rules and cross-use reuse of historical experience. By combining static appearance verification and dynamic physical verification, we form a closed loop of 'knowledge retrieval - prompt generation - verification scoring - memory accumulation' to accurately locate and correct physical errors.
It significantly improves the physical rule compliance and temporal coherence of video generation, enhances the generalization ability of out-of-distribution scenes, and achieves continuous iterative improvement of the system's generation quality without modifying the model architecture or adding training data.
Smart Images

Figure CN122138025A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of video generation technology, specifically to an iterative self-optimizing text-to-video generation method and system based on knowledge enhancement. Background Technology
[0002] With the rapid development of generative artificial intelligence, diffusion-based text-to-video (T2V) generation models based on the Transformer architecture can now generate highly realistic video content, with single-frame image quality approaching the level of real-world shooting, demonstrating enormous application potential in the entertainment industry. However, existing T2V models only learn about physical principles through implicit feature fitting in the training data, lacking explicit understanding and input of physical rules. The generated videos often exhibit physical illusions that violate gravity, fluid dynamics, optical principles, and causal logic, failing to meet the core requirements of realistic physical simulation in professional scenarios.
[0003] To address the aforementioned issues, existing optimization solutions have significant shortcomings: First, they rely on large-scale physical scene datasets for retraining or fine-tuning, resulting in high computational costs and poor generalization; or they depend on implicit reasoning based on the built-in common sense of large language models, lacking support from a structured physical knowledge base, which easily leads to rule illusions and makes the process untraceable. Second, they lack dynamic memory and experience reuse mechanisms across sessions, causing similar physical errors to recur and preventing the system from continuously improving with use; at the same time, existing verification relies solely on black-box scoring from textual implication models, lacking frame-level visual language verification, making it difficult to accurately locate temporal physical errors. Third, existing prompt word optimization methods only optimize visual descriptions, failing to address the issue of physical fidelity, lacking explicit physical rule embedding, closed-loop feedback, and reliable reasoning mechanisms, resulting in unstable optimization effects.
[0004] Therefore, existing technologies suffer from problems such as untraceable physical knowledge injection, lack of dynamic memory reuse, and lack of frame-level verification loop, resulting in low physical fidelity, weak generalization ability, and unsustainable optimization. Summary of the Invention
[0005] To address the aforementioned issues, this invention proposes an iterative self-optimizing text-generated video method and system based on knowledge enhancement. This system forms a complete closed loop of "knowledge retrieval - prompt generation - verification and scoring - memory retention," significantly improving the generalization ability across distributed scenarios and enabling continuous iterative improvement in the quality of system-generated content.
[0006] According to some embodiments, the present invention adopts the following technical solution: A knowledge-enhanced iterative self-optimizing text-generated video method includes: Obtain the user's original text input and load the static physics knowledge base and dynamic constraint memory. The system performs vectorized retrieval on the user's original text input, matches the physical rule set from the static physical knowledge base, recalls the historical constraint set from the dynamic constraint memory base, and performs physical rule constraints and memory fusion based on the physical rule set and historical constraint set to generate knowledge-enhanced video generation prompts. The generated video prompts are input into a pre-trained diffusion-based T2V model to generate the target video; Perform static appearance verification and dynamic physical verification on the target video, and calculate semantic consistency score and physical common sense score; Based on the verification and scoring results, determine whether the iteration termination condition is met. If it is met, output the final video and optimized prompt words; otherwise, update the dynamic constraint memory and re-optimize the video generation prompt words to perform iterative generation of the target video.
[0007] According to some embodiments, the present invention adopts the following technical solution: A knowledge-enhanced iterative self-optimizing text-generated video system, comprising: The text acquisition module is configured to: acquire the user's original text input and load the static physical knowledge base and dynamic constraint memory. The prompt generation module is configured to: perform vectorized retrieval on the user's original text input, match the physical rule set from the static physical knowledge base, recall the historical constraint set from the dynamic constraint memory base, and perform physical rule constraints and memory fusion based on the physical rule set and the historical constraint set to generate knowledge-enhanced video prompt words; The video generation module is configured to input the generated video generation prompts into a pre-trained diffusion-based T2V model to generate the target video; The inspection and scoring module is configured to perform static appearance verification and dynamic physical verification on the target video, and calculate semantic consistency score and physical common sense score. The termination judgment module is configured to: determine whether the iteration termination condition is met based on the verification result and the scoring result; if it is met, output the final video and optimized prompt words; otherwise, update the dynamic constraint memory and re-optimize the video generation prompt words to perform iterative generation of the target video.
[0008] According to some embodiments, the present invention adopts the following technical solution: A computer program product includes a computer program that, when executed by a processor, implements the aforementioned knowledge-enhanced iterative self-optimizing text-generated video method.
[0009] According to some embodiments, the present invention adopts the following technical solution: A non-transitory computer-readable storage medium is provided for storing computer instructions, which, when executed by a processor, implement the aforementioned knowledge-enhanced iterative self-optimizing text-to-video generation method.
[0010] According to some embodiments, the present invention adopts the following technical solution: An electronic device includes a processor, a memory, and a computer program; wherein the processor is connected to the memory, the computer program is stored in the memory, and when the electronic device is running, the processor executes the computer program stored in the memory to enable the electronic device to implement the aforementioned knowledge-enhanced iterative self-optimizing text-generated video method.
[0011] Compared with the prior art, the beneficial effects of the present invention are as follows: This invention achieves explicit introduction of physical knowledge and cross-use reuse of historical experience by acquiring the user's original text input and loading a static physical knowledge base and a dynamic constraint memory base, avoiding the physical illusions that can easily arise from implicit reasoning relying on large language models. By performing vectorized retrieval on the user's original text input to match the physical rule set and recall the historical constraint set, and then performing physical rule constraint and memory fusion to generate knowledge-enhanced prompt words, the comprehensiveness and accuracy of physical constraints in the prompt words are significantly improved. Explicit embedding of physical rules can be achieved without modifying the T2V model architecture or adding new training data. The generated prompt words are input into a diffusion-based T2V model to generate a target video, and static appearance verification and dynamic physical correction are performed on the target video. The system simultaneously calculates semantic consistency scores and physics common sense scores, achieving an upgrade from black-box quantization scoring to interpretable frame-level fine-grained feedback, enabling precise location of physical error types and positions in videos. By determining the iteration termination condition based on the verification and scoring results, and updating the dynamic constraint memory when the condition is not met, the system re-optimizes the prompt words and iteratively generates videos, forming a complete closed loop of "knowledge retrieval - prompt generation - verification scoring - memory accumulation". This allows the system to continuously reuse historical failure experience and avoid the recurrence of similar physical errors. Thus, without modifying the model architecture or increasing training data, the system efficiently and cost-effectively improves the physical rule compliance, temporal coherence, and generalization ability of T2V generated videos in out-of-distribution scenarios. Attached Figure Description
[0012] The accompanying drawings, which form part of this invention, are used to provide a further understanding of the invention. The illustrative embodiments of the invention and their descriptions are used to explain the invention and do not constitute an improper limitation of the invention.
[0013] Figure 1 This is a flowchart of the method in Example 1.
[0014] Figure 2This is a detailed flowchart of the dual-database semantic retrieval and physical prior acquisition in Example 1.
[0015] Figure 3 This is a flowchart of the two-stage frame-level video verification and evaluation process for the verification agent in Example 1. Detailed Implementation
[0016] The present invention will be further described below with reference to the accompanying drawings and embodiments.
[0017] It should be noted that the following detailed descriptions are exemplary and intended to provide further illustration of the invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains.
[0018] It should be noted that the terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the scope of exemplary embodiments according to the invention. As used herein, the singular form is intended to include the plural form as well, unless the context clearly indicates otherwise. Furthermore, it should be understood that when the terms "comprising" and / or "including" are used in this specification, they indicate the presence of features, steps, operations, devices, components, and / or combinations thereof.
[0019] Example 1 One embodiment of the present invention provides an iterative self-optimizing text-generated video method based on knowledge enhancement, comprising: Step 1: Obtain the user's original text input and load the static physics knowledge base and dynamic constraint memory base; Step 2: Perform vectorized retrieval on the user's original text input, match the physical rule set from the static physical knowledge base, recall the historical constraint set from the dynamic constraint memory base, and perform physical rule constraints and memory fusion based on the physical rule set and historical constraint set to generate knowledge-enhanced video generation prompts; Step 3: Input the generated video prompts into the pre-trained diffusion-based T2V model to generate the target video; Step 4: Perform static appearance verification and dynamic physical verification on the target video, and calculate semantic consistency score and physical common sense score; Step 5: Based on the verification results and scoring results, determine whether the iteration termination condition is met. If it is met, output the final video and optimized prompt words; otherwise, update the dynamic constraint memory and re-optimize the video generation prompt words to perform iterative generation of the target video.
[0020] This embodiment relates to an iterative self-optimizing text-to-video generation method based on knowledge enhancement, which aims to solve the problems in existing T2V generation technologies that conform to physical laws, such as insufficient user professional physical prior knowledge leading to incomplete physical constraints and easy hallucinations, as well as the lack of semantic information and closed-loop feedback mechanisms leading to poor interpretability, weak generalization ability, and recurrence of similar errors.
[0021] In T2V generation tasks, traditional methods typically rely on user-specified physical constraints or implicit reasoning based on built-in common sense in large language models. However, for complex and specialized physical scenarios in the real world, this approach suffers from limited coverage, insufficient accuracy, and susceptibility to illusions. Especially when facing unknown scenarios where users lack prior knowledge, it is difficult to accurately provide comprehensive and reasonable physical rule constraints, thus affecting the physical fidelity and scene adaptability of the final generated video.
[0022] Therefore, the first technical problem this embodiment aims to solve is: how to automatically mine and construct a more comprehensive and applicable set of physical constraints when users lack sufficient prior knowledge of physical knowledge. To address this problem, this embodiment proposes constructing a structured static physical knowledge base, breaking down textbook-level and domain-level physical knowledge into standardized concept records containing visual descriptions, counterfactual scenarios, and verification templates; automatically matching the physical rules corresponding to user queries through semantic retrieval to achieve explicit injection of physical knowledge; and further compressing abstract physical rules into visual descriptions recognizable by the T2V model, solving the problem of poor physical compliance in existing T2V technologies, and enabling the coverage of necessary physical constraints even when users lack prior knowledge.
[0023] Furthermore, existing physical fidelity T2V optimization methods suffer from serious deficiencies in feedback mechanisms and experience reuse. Most methods only address single-prompt optimization or single-round iteration, lacking cross-session experience reuse mechanisms, leading to the recurrence of similar physical errors. Simultaneously, the verification process relies solely on black-box quantitative scoring, lacking structured localization and interpretability analysis of error details, failing to form a complete closed-loop optimization system, resulting in insufficient generalization capability in out-of-distribution scenarios. Therefore, the second technical problem this embodiment aims to solve is: how to construct an interpretable closed-loop iterative optimization mechanism to achieve cross-session reuse of historical failure experience, while simultaneously locating physical errors in the video, providing fine-grained feedback for optimization, and improving the generalization capability in out-of-distribution scenarios and the continuous iterative quality of system generation.
[0024] To address the aforementioned issues, this embodiment designs a multi-agent collaborative architecture that decouples planning, generation, and verification responsibilities. It performs a two-stage frame-level verification process—static appearance and dynamic physics—through a visual language model, accurately locating error positions, types, and correction directions, upgrading black-box scoring to interpretable, fine-grained feedback. Simultaneously, a dynamic constraint memory is constructed to persistently store historical failure cases, learning constraints, and successful correction patches. New requests can directly reuse historical experience, preventing the recurrence of similar errors. Ultimately, this forms a complete closed loop of "knowledge retrieval - prompt generation - verification scoring - memory accumulation," significantly improving the generalization ability in distributed scenarios and achieving continuous iterative improvement in the system's generation quality.
[0025] This embodiment provides an iterative self-optimizing text-generated video method based on knowledge enhancement, such as... Figure 1 As shown, the specific steps are as follows: S101, Initial input acquisition.
[0026] Obtain the user's original text input, complete the initialization of system parameters, load and verify the static physical knowledge base and dynamic constraint memory, and set the maximum number of iterations, the verification threshold, the convergence threshold, and the prompt word length limit parameters.
[0027] Specifically, the system acquires the user's original text input q, such as "In the space station, a glass of water is slowly poured out, and the liquid is released into the surrounding area," as well as the user-selected model configuration and iteration parameters, to complete the system environment initialization. The set iteration parameters include: the maximum number of iterations T, and the SA / PC pass / fail threshold. Convergence threshold Physical concept retrieval volume Number of similar failure cases searched Initialize the iteration round counter .
[0028] S102, Dual-database semantic retrieval and physical prior acquisition.
[0029] The user's original text input is vectorized, and relevant physical concept records are retrieved from the static physical knowledge base, while semantically similar historical failure records are retrieved from the dynamic constraint memory base. This provides physical priors and historical experience to support subsequent video generation. Figure 2 As shown, specifically: S1021. Static physics knowledge base retrieval to obtain relevant physics rule sets. .
[0030] A pre-trained text embedding model is used to vectorize the user's original text input q to generate a query embedding vector. The cosine similarity search is performed in the vector database of the static physics knowledge base, returning the Top-K records of the most similar physical concepts. Complete physical concept information is then matched from a JSON structured file using unique concept identifiers to form a set of physical rules. .
[0031] In this embodiment, each physical concept record in the static physical knowledge base includes a unique concept identifier (id), concept name (name), domain classification (domain), standard definition (definition), related concepts (related_concepts), formulas (formulas), visual description (visual_grounding), counterfactuals (counterfactuals), and verification templates (verification_templates). The visual description includes a list of material appearance descriptions (material_appearance) and a list of motion dynamics descriptions (motion_dynamics), used to transform abstract physical laws into visual descriptions recognizable by the T2V model. These will be explained in detail below: (1) Concept unique identifier (id), used to uniquely identify a physical concept record, which facilitates indexing, retrieval and association, such as “concept_other_free_fall_0137”; (2) Concept name (name), which represents the standard name of the current physical concept, such as "free fall"; (3) Domain classification, used to identify the discipline or business area to which the physical concept belongs, such as "solid mechanics"; (4) Standard definition: The standard definition used to describe physical concepts is the basis for subsequent physical rule reasoning and verification. For example, "The motion of an object falling from rest under the action of gravity alone, with an initial velocity of zero and an acceleration equal to the acceleration due to gravity." (5) Related concepts are used to expand related knowledge retrieval and assist reasoning, such as "["gravity", "acceleration", "inertia", "velocity"]".
[0032] (6) Formulas are used to store the formula expressions, meanings and variable descriptions related to the physical concept. For example, “[{"expression": "v = gt", "meaning": "free fall velocity changes with time", "variables": ["v: velocity (m / s)", "g: gravitational acceleration (9.8 m / s²)", "t: time (s)"]}, {"expression": "h = 1 / 2 gt²", "meaning": "fall height is proportional to the square of time", "variables": ["h: fall height (m)"]}, {"expression": "v² = 2gh", "meaning": "final velocity square is proportional to height"}].
[0033] (7) Visual description (visual_grounding), including material appearance and motion dynamics, specifically: The visual_grounding.material_appearance list is used to transform abstract physical laws into static visual descriptions that the T2V model can recognize, such as "["solid object like a stone or metal piece", "opaque and dense appearance"]".
[0034] The list of motion dynamics descriptions (visual_grounding.motion_dynamics) is used to provide dynamic motion representations that conform to physical laws, such as "["object moving vertically downwards", "velocity increasing uniformly over time"]".
[0035] (8) Counterfactual descriptions are used to describe erroneous situations that do not conform to the physical law, so as to avoid them during dynamic verification. For example, "["The object falls up and down after being released", "The object falls at a constant speed without air resistance", "The object is suspended in the air"]".
[0036] (9) Verification question templates (verification_templates) are used to guide the visual language model to verify the video content from the perspective of predefined questions. For example, "[{"question": "Does the object accelerate when it falls?", "focus": "The existence of gravitational acceleration", "expected": "Yes"},{"question": "Is the falling trajectory of the object a vertical straight line?", "focus": "Direction of the trajectory", "expected": "Yes"}]".
[0037] In this embodiment, a general physics knowledge base is constructed using junior high school, high school, and university physics textbooks, and a storage method combining structured files and vector databases is adopted. It can be replaced with professional physics knowledge bases for different business scenarios such as industrial simulation, film and television special effects, and aerospace, and supplemented with industry-specific physical rules and visual constraints. In terms of storage method, the vector database can be replaced with other similar vector databases, and the structured files can be replaced with relational databases. The ultimate goal is to achieve structured storage and efficient semantic retrieval of physics knowledge.
[0038] S1022. Dynamic constraint memory retrieval to obtain the historical constraint set M(q).
[0039] Using the same text embedding model, the user's original text input q is vectorized, and cosine similarity retrieval is performed in the vector database of the dynamic constraint memory. The top-N semantically similar historical success / failure records are returned. Metadata filtering is supported by physical concept identifiers and error types. The learned constraint list, successful correction patch, and error description are extracted from the records to form a historical constraint set M(q).
[0040] In this embodiment, each historical record in the dynamic constraint memory includes the following fields: unique record identifier (record_id), experiment session identifier (experiment_id), iteration round number (iteration), list of associated physical concept identifiers (concept_ids), user's original text input (user_query), prompt words used to generate the agent (agent_prompt), error type classification (error_type), error summary (error_summary), list of learned constraints (learned_constraints), successful patch (successful_patch), static appearance verification result (static_check_result), dynamic physics verification result (dynamic_check_result), list of problem frame numbers (problem_frames), video file storage path (video_path), list of keyframe paths (keyframe_paths), semantic consistency score (sa_score), physics common sense score (pc_score), overall score (overall_score), whether the verification passed (is_pass), and last update timestamp (last_updated).
[0041] The error types are categorized into three types: visual artifacts, physical violations, and description ambiguities. Successfully recorded error types are listed in an empty list, which will be explained in detail below: (1) Record a unique identifier (record_id) to uniquely identify the historical records in a certain iteration process, such as “exp_20260307_152348_6afce0_iter1”; (2) Experiment session identifier (experiment_id), multiple iterations requested by the same user share the same experiment_id, for example "exp_20260307_152348_6afce0"; (3) Iteration round number, starting from 0 and incrementing, indicates which optimization round it was generated, for example, "1"; (4) Associated physical concept identifiers (concept_ids), used to identify the physical concept associated with the record, supporting filtering by concept and searching for similar cases, such as "["concept_fluid_microgravity_behavior", "concept_surface_tension_liquid_shape"]"; (5) User raw input (user_query) is used to store the user's raw generation requirements without optimization, such as "In the space station, a glass of water is slowly poured out, and the liquid is released into the surrounding area".
[0042] (6) Generate prompt words (agent_prompt) to save the text conditions that are actually input to the T2V model in this round, such as "Inside the space station, a cup of water is slowly released, and the liquid forms spherical droplets that float in mid-air. The droplets reflect light, showing the fluid behavior under microgravity. The instruments inside the space station serve as the background."
[0043] (7) Error type classification (error_type), used to identify the error type. Optional values: which one or more of the following: Visual Artifact, Physics Violation, Description Ambiguity. Successful records are recorded as an empty array [], for example, "['Visual Artifact']", which in this example is "[]". (8) Error summary (error_summary) is used to compress and summarize the main problems in this round. If the iteration is successful, it is recorded as an empty string, which is convenient for the planning agent to call quickly. For example, "The liquid looks like a continuous downward flow (Frams1-8) instead of forming floating spherical droplets, which violates the behavior of microgravity fluid." In this example, it is an empty string "".
[0044] (9) The list of learned constraints (learned_constraints), see Table 1 below, is used to store reusable constraint rules summarized from failed cases and reusable constraints extracted from successful records. If the current iteration is a failure, the record is an empty list []. (10) Successful patch (successful_patch), see Table 2 below, is used to store the prompt word optimization strategy that has been verified to be effective in subsequent iterations. If this iteration fails, this part of the record will be null; (11) Static appearance verification result (static_check_result), which saves the pass status and explanation of static verification, such as "{"pass":true, "details":"Object, background, color and material are all correct. Liquid forms the expected spherical droplet."}"; (12) Dynamic physical check result (dynamic_check_result), which saves the pass status and explanation of the dynamic physical check, such as "{"pass":true, "details":"Motion, surface tension, and microgravity behavior are all correct."}"; (13) Problem frame number (problem_frames) refers to the sample frame number that is determined to have a major problem after sampling and numbering in step S1051. If the current success is achieved, it is recorded as an empty list [], for example, "[1, 2, 3, 4, 5, 6, 7, 8]"; in this example, it is an empty list []. (14) Video file path (video_path), used to locate the corresponding generated video file, for example " / mnt / nfs / ... / output / exp_20260307_152348_6afce0_iter1.mp4"; (15) Keyframe path list (keyframe_paths), used to store key sampling frame paths for easy review, for example "["data / frames / exp_20260307_152348_6afce0_iter1_frame_00.jpg", ..., "data / frames / exp_20260307_152348_6afce0_iter1_frame_07.jpg"]"; (16) Semantic consistency score (scores.sa_score), which represents the degree of semantic alignment between the video content and the text conditions, for example, "0.77734375"; (17) Physics knowledge score (scores.pc_score), which indicates the degree to which the video content conforms to the laws of physics, for example, "0.6796875"; (18) Overall score (scores.overall_score), calculated by weighting sa_score and pc_score, for example "0.728515625"; (19) Whether the validation is passed (is_pass), true / false indicates whether the static and dynamic validations are passed and the score is up to standard, for example, "true"; (20) Timestamp (last_updated) is used to record the update time of the history, such as "2026-03-07T15:43:36.076273" or "2026-03-07T15:37:08.472619".
[0045] Table 1 Expanded list of learned constraints (learned_constraints)
[0046] Table 2. Expanded table of successful patches (successful_patch)
[0047] S103, Planning intelligent agent video generation text condition generation.
[0048] The planning agent uses the user's original text input q and the retrieved physical rule set as input. Historical constraints It executes physical rule constraints and memory fusion to generate text that meets the conditions for video generation.
[0049] The physical rule constraints mentioned above are: from the physical rule set Extract a preset number of visual descriptions and counterfactuals from each concept record, concatenate them into a short text summary, prioritize the use of visual descriptions in the output content, write professional physics terms as specific visual descriptions, that is, transform abstract physical laws into observable visual descriptions, and disable unnecessary professional physics terms; The memory fusion is defined as: integrating the set of historical constraints. The learned constraint list (learned_constraints) and successful patch (successful_patch) are internalized into prompt word optimization logic to avoid the recurrence of similar errors during the generation process. Specifically, constraint rules summarized from previous iterations are extracted from the learned constraint list (learned_constraints) (such as "when a liquid is released in a microgravity environment, it should form spherical or near-spherical droplets to float, rather than flowing continuously downwards"), and verified effective optimization prompt words (such as "add "form floating spherical droplets") are extracted from the successful patch (successful_patch) and added as positive examples to the prompt generation process.
[0050] Specifically, a standardized prompt template is built to display the user's original text input. Physical rule set Historical constraints Enter the template and generate optimized prompt words using a large language model. The core logic can be formally expressed as:
[0051] In the formula, To generate the prompt function for the planning agent, For the first The text-based video prompts generated through iterative cycles.
[0052] Here is a standardized prompt engineering template, the content of which is shown in Table 3: Table 3. Project Template Tips
[0053] In the template above, the placeholders {iteration_round}, {user_query}, {physics_rules_text}, and {memory_hints_text} are replaced with actual content, as shown in Table 4. Table 4. Explanation of Placeholders for Planning Intelligent Agent Templates
[0054] In this embodiment, a large language model is used to transform physical rules and optimize prompt words. The ultimate goal is to generate optimized prompt words that are more in line with physical constraints and adapted to the T2V model.
[0055] S104, Generating intelligent agent video.
[0056] The pre-trained T2V diffusion model corresponding to the generating agent takes the prompt words generated by the planning agent as input, generates the target video file according to the preset generation parameters, and completes persistent storage.
[0057] Load the user-selected pre-trained diffusion-based T2V generative model, and configure the corresponding scheduler, inference steps, bootstrapping scale, number of frames, frame rate, resolution, and random seed parameters. Wheel text conditions Use the input to generate the target video. Formalized expression:
[0058] In the formula, For generating video functions that generate intelligent agents, For the first The target video generated through rounds of iteration.
[0059] In this embodiment, the T2V generation model uses CogVideoX-5B, and VAE slicing and VAE block optimization are enabled to reduce video memory usage. The generation parameters are set as follows: 50 inference steps, 6.0 guidance scale, 49 frames per second, 8fps frame rate, 720×480 resolution, and 42+t random seed to ensure the comparability of generation results between rounds. The generated video files are saved to the preset output directory according to the rounds and recorded synchronously in the experimental log.
[0060] S105, Verification of the agent's two-stage video verification.
[0061] Temporal frame sampling is performed on the generated target video. Static appearance verification and dynamic physical verification are performed using a visual language model, outputting a two-stage pass result and feedback information. Simultaneously, the VideoCon-Physics automatic evaluator is used to calculate the semantic consistency (SA) score and the physical commonsense (PC) score, completing a quantitative assessment of video quality. Figure 3 As shown, specifically: S1051, Timing frame sampling preprocessing.
[0062] From the generated target video The system samples a preset number of video frames evenly over time. In this embodiment, the number of frames sampled is 8. Each frame is labeled with a time sequence number from 1 to 8 to generate a multi-image list arranged in chronological order, which is adapted to the input requirements of the visual language model.
[0063] S1052, Static appearance verification, output the first pass result and the first feedback information.
[0064] Input the user's original text , generated text conditions The standard definitions of relevant physical concepts, related concepts, and multi-image lists obtained from the static knowledge base are input into the visual language model. In this embodiment, GPT-4o is selected. Frame-level consistency checks are performed on four dimensions: integrity of core elements, consistency of color and material, background fit, and object accuracy. The first pass result (Boolean value PASS / FAIL) and the first feedback information are output. The first feedback information includes fine-grained error descriptions and is accompanied by a frame number label.
[0065] S1053, Dynamic physical verification, outputs the second pass result and the second feedback information.
[0066] Based on the retrieved physical rule set The system defines definitions, motion dynamics, counterfactuals, and verification templates. It inputs the physical knowledge context and a list of multiple sampled frames into the visual language model (GPT-4o) to perform cross-frame temporal tracking. It performs dynamic physical verification on five dimensions: counterfactual scene avoidance, causal logic rationality, material-motion consistency, compliance with physical laws, and motion trajectory and velocity changes. It outputs a second pass / fail result and a second feedback message. The second feedback message includes fine-grained error descriptions and the corresponding problem frame number.
[0067] This embodiment employs a two-stage visual verification mechanism of "static appearance verification + dynamic physical verification," combined with quantitative scoring. Custom verification dimensions such as temporal coherence, image stability, and lighting consistency can be added according to business needs. At the verification logic level, single-round question-and-answer verification can be replaced with multi-round question-and-answer verification. Multiple sets of subdivided questions are generated based on a physical concept-based verification template to verify the video content one by one, further enhancing the rigor of the verification. At the quantitative scoring level, other video quality assessment models can be used, adding more assessment dimensions. The ultimate goal is to achieve accurate verification of the physical compliance and semantic consistency of the generated video. Specifically, when using a visual language model (GPT-4o in this example) for static appearance verification and dynamic physical verification evaluation, the evaluation template shown in Table 5 is adopted: Table 5 Evaluation Template
[0068] In the template above, the placeholders {user_query}, {agent_prompt}, and {physics_context} are replaced with actual content, as shown in Table 6. Table 6 Explanation of Placeholders for Validating Smart Agent Templates
[0069] The generated target video is not passed through text placeholders, but is appended to the user message as multimodal image parameters (image_url list) when the visual language model is called. Therefore, no corresponding text placeholders are needed in the template.
[0070] S1054, SA / PC evaluation index quantitative scoring calculation.
[0071] The VideoCon-Physics automatic evaluator, based on the VideoPhy protocol, is used to evaluate the generated video. A discrimination test is performed. This process involves inputting video frames and specific physical verification queries into a multimodal model, extracting the model's output probabilities for affirmative (Yes) and negative (No) terms, and calculating the semantic consistency (SA) score and the physical common sense (PC) score, respectively. The SA score measures the semantic alignment between the video content and the text conditions, while the PC score measures the degree to which the video content conforms to physical rules. Both scores range from [0,1].
[0072] S106, Iteration termination judgment.
[0073] Determine whether the current iteration meets the preset termination condition. If the termination condition is met, execute step S108 and output the final target video and optimized text conditions. If the termination condition is not met, execute the dynamic memory update step S107.
[0074] The preset termination condition includes any one of the following; the iteration will terminate if any one of them is met: Eligible termination condition: Both the first and second passing results of the current round are PASS, and , ,in In this embodiment, the preset qualified threshold is used. ; Convergence termination condition: The improvement in the overall score after two consecutive iterations is less than the preset convergence threshold. ,Right now and In this embodiment ;
[0075] Forced termination condition: current iteration round In this embodiment, the maximum number of iterations T is exceeded. .
[0076] S107, Dynamic Memory Update and Closed-Loop Iteration.
[0077] The generated data, verification results, and quantitative scores of the current round are synchronized to the dynamic constraint memory in real time. Through experience accumulation and patch generation mechanisms, this provides a reference for subsequent iterations or cross-session requests. The specific steps are as follows: Structured extraction of verification results: The system parses the structured JSON data output by the verification agent in step S105, extracts key information, and automatically classifies errors, directly obtaining the error_type field to determine whether the problem in this round is a visual artifact, a physical violation, or a description ambiguity; precise problem localization: extracting the problem frame number list (problem_frame) and error summary (error_summary) for accurate error-prone reminders in the next round of planning; state synchronization: obtaining the is_pass boolean value as the core metadata for updating the memory record.
[0078] Experience accumulation and patch generation: This part is triggered only in successful rounds. If the current round satis_pass=true, the experience backfilling logic based on the large language model is activated. The system will input the original user text. Using the prompts for failed rounds, successful rounds, and validation feedback as context, the following content is automatically generated: Learned Constraints: Summarizes the visual constraints that must be followed in this scenario. Successful Patch: Records the prompt editing rules and the final optimized text; if this round is a failure (is_pass = false), the above fields are recorded as null or an empty list.
[0079] Persistent writing to the memory: The extracted and generated structured data is encapsulated into a complete historical record and synchronously written to the database.
[0080] Iterative state update: After the data writing is complete, the iteration round counter will be updated. Updated to :like If the maximum number of iterations is reached, return to step S103 with the updated dynamic memory context to start the next round of prompt word optimization and video generation. If the termination condition has been triggered, stop jumping and proceed to step S108.
[0081] S108, Final result output.
[0082] When the iteration meets the termination condition, the final optimized text prompt, target video file, and score data for each iteration are output.
[0083] Through the above solution, this embodiment solves the following two problems: 1) Difficulty in obtaining physical constraints: Existing T2V models cannot understand professional physical terms and rely on user-provided or LLM implicit reasoning. However, users lack prior knowledge and LLM is prone to illusions, resulting in incomplete coverage of physical rules and poor accuracy; 2) Lack of closed-loop optimization mechanism: Existing solutions lack closed-loop feedback and experience reuse capabilities, leading to repeated occurrences of similar errors, preventing continuous system iteration, and severely lacking generalization ability when facing off-distribution physical scenarios; The following improvements have been achieved: (1) This embodiment achieves explicit injection of physical knowledge by constructing a structured static physical knowledge base, reducing the illusion of physical rules in large language models and significantly improving the comprehensiveness and accuracy of physical constraints. Addressing the problems of existing technologies relying on built-in common sense in large language models, the susceptibility of implicit reasoning of physical rules to illusion, and incomplete coverage of physical constraints when user prior knowledge is insufficient, this embodiment constructs a domain-specific physical knowledge base. It decomposes abstract physical laws into standardized records of visual descriptions, counterfactual constraints, and verification templates. Through semantic retrieval, it automatically matches the physical rules required for the scenario, achieving explicit and traceable injection of physical knowledge. Compared to the shortcomings of traditional methods such as implicit fitting and susceptibility to illusion, this embodiment significantly reduces the reliance on user professional physical prior knowledge, greatly improving the accuracy of physical rules and scenario coverage, and is adaptable to the stringent physical requirements of professional scenarios such as science and education and industry.
[0084] (2) This embodiment achieves cross-session reuse of historical failure experience through a dynamic constraint memory, constructing a continuous iteration mechanism that becomes more accurate with use, thus solving the industry pain point of repeated occurrences of similar physical errors. Addressing the problems of independent processing of single requests, lack of experience reuse mechanisms, and recurrence of similar errors in existing technologies, this embodiment designs a cross-session dynamic constraint memory. This memory can persistently reuse historical failure cases, error types, learning constraints, and successful correction patches. New requests can directly reuse historical experience through semantic retrieval, avoiding similar errors at their source. This mechanism breaks the limitations of single iteration in existing technologies, enabling the system's generation quality to be continuously optimized with increasing usage, forming a positive-loop capability iteration system.
[0085] (3) This embodiment upgrades the existing black-box quantitative scoring to white-box interpretable fine-grained feedback through a two-stage frame-level visual verification mechanism, significantly improving the accuracy and optimization efficiency of error correction. Addressing the problems of poor interpretability, inability to locate error locations and causes, and ambiguous feedback signals in existing verification schemes, this embodiment employs a visual language model to perform static appearance + dynamic physical two-stage cross-frame verification. This accurately locates the frame location, temporal interval, error type, and correction direction of the error, providing precise feedback signals for prompt word optimization. Compared to traditional black-box verification that relies solely on SA / PC scoring, the verification mechanism in this embodiment offers stronger interpretability.
[0086] (4) This embodiment constructs a complete physical alignment closed loop through a multi-agent collaborative architecture with decoupled responsibilities, achieving a breakthrough in the generalization ability of physical scenes outside the distribution, while possessing strong versatility and scalability. This embodiment addresses the problems of high module coupling, inflexible replacement, and weak generalization ability of existing technologies in physical scenes outside the distribution. It designs a multi-agent architecture with decoupled planning, generation, and verification responsibilities. T2V models, LLM, VLM, and vector databases can be flexibly replaced without modifying the model architecture or retraining and fine-tuning. It is plug-and-play and adaptable to all mainstream T2V models. At the same time, through the complete closed loop of "knowledge retrieval - planning and generation - verification feedback - memory accumulation", it can still stably generate video content that conforms to physical rules even when facing physical scenes outside the distribution that are not covered by training data.
[0087] Example 2 One embodiment of the present invention provides an iterative self-optimizing text-generated video system based on knowledge enhancement, comprising: The text acquisition module is configured to: acquire the user's original text input and load the static physical knowledge base and dynamic constraint memory. The prompt generation module is configured to: perform vectorized retrieval on the user's original text input, match the physical rule set from the static physical knowledge base, recall the historical constraint set from the dynamic constraint memory base, and perform physical rule constraints and memory fusion based on the physical rule set and the historical constraint set to generate knowledge-enhanced video prompt words; The video generation module is configured to input the generated video generation prompts into a pre-trained diffusion-based T2V model to generate the target video; The inspection and scoring module is configured to perform static appearance verification and dynamic physical verification on the target video, and calculate semantic consistency score and physical common sense score. The termination judgment module is configured to: determine whether the iteration termination condition is met based on the verification result and the scoring result; if it is met, output the final video and optimized prompt words; otherwise, update the dynamic constraint memory and re-optimize the video generation prompt words to perform iterative generation of the target video.
[0088] Example 3 One embodiment of the present invention provides a computer program product, including a computer program that, when executed by a processor, implements the aforementioned knowledge-enhanced iterative self-optimizing text-generated video method.
[0089] Example 4 In one embodiment of the present invention, a non-transitory computer-readable storage medium is provided for storing computer instructions, which, when executed by a processor, implement the knowledge-enhanced iterative self-optimizing text-generated video method described above.
[0090] Example 5 One embodiment of the present invention provides an electronic device, including: a processor, a memory, and a computer program; wherein the processor is connected to the memory, the computer program is stored in the memory, and when the electronic device is running, the processor executes the computer program stored in the memory to enable the electronic device to implement the knowledge-enhanced iterative self-optimizing text-generated video method described above.
[0091] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0092] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0093] While the specific embodiments of the present invention have been described above in conjunction with the accompanying drawings, this is not intended to limit the scope of protection of the present invention. Those skilled in the art should understand that various modifications or variations that can be made by those skilled in the art without creative effort based on the technical solutions of the present invention are still within the scope of protection of the present invention.
Claims
1. A knowledge-enhanced iterative self-optimizing text-generated video method, characterized in that, include: Obtain the user's original text input and load the static physics knowledge base and dynamic constraint memory. The system performs vectorized retrieval on the user's original text input, matches the physical rule set from the static physical knowledge base, recalls the historical constraint set from the dynamic constraint memory base, and performs physical rule constraint and memory fusion based on the physical rule set and historical constraint set to generate knowledge-enhanced video generation prompts. The generated video prompts are input into a pre-trained diffusion-based T2V model to generate the target video; Perform static appearance verification and dynamic physical verification on the target video, and calculate semantic consistency score and physical common sense score; Based on the verification and scoring results, determine whether the iteration termination condition is met. If it is met, output the final video and optimized prompt words; otherwise, update the dynamic constraint memory and re-optimize the video generation prompt words to perform iterative generation of the target video.
2. The iterative self-optimizing text-generated video method based on knowledge enhancement as described in claim 1, characterized in that, Each physical concept record in the static physical knowledge base includes a unique concept identifier, domain classification, standard definition, visual description, counterfactual description, and verification question template; The visualization description includes a list of material appearance descriptions and a list of motion dynamic descriptions, which are used to transform abstract physical laws into visual descriptions that can be recognized by the T2V model.
3. The iterative self-optimizing text-generated video method based on knowledge enhancement as described in claim 1, characterized in that, Each historical record in the dynamic constraint memory includes user query, generated prompt words, error type classification, natural language error description, learned constraint list, and successful correction patch; The error types are categorized into three types: visual problems, violations of physical laws, and descriptive ambiguities.
4. The iterative self-optimizing text-generated video method based on knowledge enhancement as described in claim 1, characterized in that, The physical rule constraints specifically include: extracting motion dynamics descriptions and counterfactual prohibited scenario descriptions from each concept record in the physical rule set, transforming abstract physical laws into observable visual descriptions, and disabling unnecessary technical physical terms. The memory fusion specifically includes: internalizing the learning constraints and successful correction patches from the historical constraint set into prompt word optimization logic.
5. The iterative self-optimizing text-generated video method based on knowledge enhancement as described in claim 1, characterized in that, The static appearance verification and dynamic physical verification of the target video specifically include: uniformly sampling a preset number of video frames of the target video over time to generate a multi-image list arranged in chronological order; performing static appearance verification on the multi-image list for the integrity of core elements, consistency of color and material, background fit, and accuracy of objects, as well as dynamic physical verification for counterfactual scene avoidance, causal logic rationality, material-motion consistency, compliance with physical laws, and changes in motion trajectory and speed.
6. The iterative self-optimizing text-generated video method based on knowledge enhancement as described in claim 1, characterized in that, The calculation of semantic consistency score and physical common sense score is as follows: Calculate a semantic consistency score to measure the degree of semantic alignment between video content and text conditions; A physics common sense score is calculated to measure how well video content adheres to physical rules.
7. A knowledge-enhanced iterative self-optimizing text-generated video system, characterized in that, include: The text acquisition module is configured to: acquire the user's original text input and load the static physical knowledge base and dynamic constraint memory. The prompt generation module is configured to: perform vectorized retrieval on the user's original text input, match the physical rule set from the static physical knowledge base, recall the historical constraint set from the dynamic constraint memory base, and perform physical rule constraints and memory fusion based on the physical rule set and the historical constraint set to generate knowledge-enhanced video prompt words; The video generation module is configured to input the generated video generation prompts into a pre-trained diffusion-based T2V model to generate the target video; The inspection and scoring module is configured to perform static appearance verification and dynamic physical verification on the target video, and calculate semantic consistency score and physical common sense score. The termination judgment module is configured to: determine whether the iteration termination condition is met based on the verification result and the scoring result; if it is met, output the final video and optimized prompt words; otherwise, update the dynamic constraint memory and re-optimize the video generation prompt words to perform iterative generation of the target video.
8. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the knowledge-enhanced iterative self-optimizing text-generated video method according to any one of claims 1-6.
9. A non-transitory computer-readable storage medium, characterized in that, The non-transitory computer-readable storage medium is used to store computer instructions, which, when executed by a processor, implement the knowledge-enhanced iterative self-optimizing text-generated video method as described in any one of claims 1-6.
10. An electronic device, characterized in that, include: The device includes a processor, a memory, and a computer program; wherein the processor is connected to the memory, the computer program is stored in the memory, and when the electronic device is running, the processor executes the computer program stored in the memory to enable the electronic device to perform an iterative self-optimizing text-generated video method based on knowledge enhancement as described in any one of claims 1-6.