Optimization method, system and equipment of video generation prompt model and storage medium

By building a supervised fine-tuning data set and performing preference optimization training, optimizing the video generation prompt model, the problem of insufficient security and accuracy of prompt optimization in the existing technology is solved, and the video generation quality is improved.

CN120179859APending Publication Date: 2025-06-20BEIJING KNOWLEDGE ATLAS TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510243714.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-03
Publication Date
2025-06-20

AI Technical Summary

Technical Problem

The existing video generation model has problems that prompts that the optimization of the cue is not safe and accurate enough when training and inference, resulting in poor quality of the generated video.

Method used

Optimize video generation prompt models by building a safe, accurate and useful supervised fine-tuning dataset, combining text-level and video-level feedback signals, and perform preference optimization training.

Benefits of technology

Improve the security, accuracy and usefulness of the video generation model, ensuring that the generated video quality is higher and conforms to human values ​​and preferences.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120179859A_ABST
    Figure CN120179859A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of artificial intelligence, and relates to an optimization method and system of a video generation prompt model, equipment and a storage medium. The method comprises the following steps: 1) supervised fine-tuning training: constructing a safe, accurate and useful supervised fine-tuning data set, and performing supervised fine-tuning training on a video generation prompt model by using the supervised fine-tuning data set; 2) preference optimization training: performing preference optimization training; and respectively constructing a text level preference data set comprising the positive example data and the negative example data and a video level preference data set comprising the positive example data and the negative example data, and carrying out preference optimization training on the video generation prompt model after supervision fine tuning training by utilizing the text level preference data set and the video level preference data set. By optimizing the video generation prompt model, the difference between training and reasoning of the video generation model is made up, and the video generation quality is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of artificial intelligence, and relates to an optimization method, system, device and storage medium for a video generation prompt model. Background Art

[0002] In recent years, video generation models have made remarkable progress, especially in generating high-quality video content, such as Sora, CogVideoX, etc. However, these models usually require a large amount of high-quality video and text annotation data for training, and the video labels used in training depict the video content in detail. However, in the inference stage, the user's input is usually short, unstructured, and even unclear in intention, which makes it crucial to optimize and rewrite the user input when generating videos.

[0003] Most existing video generation models rely on large language models (LLMs) for prompt optimization, usually rewriting the user input through in-context learning. However, such methods often rely only on the understanding ability of the large language model itself during the rewriting process, without taking into account the accuracy and security of the rewritten text. For example, the large language model may change the user's original intention, omit key information, or fail to identify potential security risks during the optimization process, resulting in content deviation or security hazards in the rewritten user input. In addition, these methods do not consider the effect of the finally generated video when optimizing the user input. Even if the rewritten text is more detailed semantically, it cannot guarantee that it can effectively guide the video generation model, thus possibly resulting in poor-quality generated videos. Although these problems are crucial in model optimization, there is currently no systematic solution for how to train a harmless model that can accurately optimize user input and promote high-quality video generation.

[0004] Prompt optimization technology has always been an important research issue for large language models and generative models. Early automatic prompt optimization work can be traced back to AutoPrompt. With the rapid development of LLMs, there are more and more works on using LLMs for automatic prompt optimization. Existing video prompt optimization methods also rely on LLMs to rewrite the user input, usually through in-context learning. However, such methods only rely on the capabilities of the LLMs themselves, without considering the accuracy and security of the rewriting. At the same time, since the effect of the finally generated video is not considered, the quality of the finally generated video cannot be guaranteed either. Prompt-A-Video introduced an image and video reward model when optimizing the model to improve the quality of the generated video. However, they ignored the more important text-level alignment - the security and accuracy of the rewriting. An ideal prompt model needs to express the user's intention safely and accurately, without omitting information in the user input or generating harmful expressions, and needs to help the video generation model generate higher-quality videos.

[0005] Therefore, in view of the deficiencies existing in the above-mentioned prior art, it is necessary to develop a new optimization method, system, device and storage medium for a video generation prompt model. Summary of the Invention

[0006] In order to overcome the deficiencies of the prior art, the present invention proposes an optimization method, system, device and storage medium for a video generation prompt model, which makes up for the gap in training and inference of the video generation model by optimizing the video generation prompt model, and improves the quality of video generation.

[0007] To achieve the above object, the present invention provides the following technical solutions:

[0008] An optimization method for a video generation prompt model, characterized by comprising the following steps:

[0009] 1) Supervised fine-tuning training: Construct a safe, accurate and useful supervised fine-tuning data set and use the supervised fine-tuning data set to perform supervised fine-tuning training on the video generation prompt model;

[0010] 2) Preference optimization training: Respectively construct a text-level preference data set including positive example data and negative example data and a video-level preference data set including positive example data and negative example data, and use the text-level preference data set and the video-level preference data set to perform preference optimization training on the video generation prompt model after supervised fine-tuning training.

[0011] Preferably, the construction of the safe, accurate and useful supervised fine-tuning data set in step 1) specifically includes:

[0012] 11) Data collection and collation: Collect various user inputs related to video generation from open sources, perform diversity and rule-based quality screening on the user inputs, and use a safety classifier to label the user inputs with safety problems among them;

[0013] 12) Preliminary prompt generation: Based on the collected and collated user inputs, use a large language model for in-context learning to generate preliminary prompts;

[0014] 13) Criticism and problem identification: Based on pre-established principles, use a large language model to criticize and identify problems in the preliminary prompts;

[0015] 14) Rule-guided improvement: Use a large language model to modify the preliminary prompts with identified problems to obtain optimized prompts, thereby generating the final supervised fine-tuning data set.

[0016] Preferably, in step 13), performing criticism and problem identification includes identifying potential safety hazards, inaccurately expressed parts, and parts missing key information in the preliminary prompts.

[0017] Preferably, the specific steps of separately constructing a text-level preference dataset including positive example data and negative example data and a video-level preference dataset including positive example data and negative example data in step 2) are as follows:

[0018] 21) Data sampling: Use the video generation prompt model after supervised fine-tuning training to generate multiple prompts for each user input.

[0019] 22) Construction of text-level preference dataset: Based on pre-established principles, use a large language model to respectively criticize multiple prompts of each user input to check whether there are problems. If there are problems, use the large language model to modify the problematic prompts. Among them, the problematic prompts are used as negative example data, and the modified prompts are used as positive example data.

[0020] 23) Construction of video-level preference dataset: For the prompts without problems, use the video generation model to generate videos and use the reward model to evaluate the quality of the generated videos. The prompt with the lowest video score is used as negative example data, and the prompt with the highest score is used as positive example data.

[0021] Preferably, the check for whether there are problems in step 22) includes checking whether there are potential safety hazards, whether there are problems with inaccurate expressions, and whether there are problems with missing key information.

[0022] In addition, the present invention also provides an optimization system for a video generation prompt model, which is characterized by including:

[0023] A supervised fine-tuning training module, which is used to construct a safe, accurate, and useful supervised fine-tuning dataset and use the supervised fine-tuning dataset to perform supervised fine-tuning training on the video generation prompt model.

[0024] A preference optimization training module, which is used to separately construct a text-level preference dataset including positive example data and negative example data and a video-level preference dataset including positive example data and negative example data, and use the text-level preference dataset and the video-level preference dataset to perform preference optimization training on the video generation prompt model after supervised fine-tuning training.

[0025] Moreover, the present invention also provides an optimization device for a video generation prompt model, which is characterized by including:

[0026] One or more processors;

[0027] A memory for storing one or more programs;

[0028] When the one or more programs are executed by the one or more processors, the one or more processors are caused to implement the optimization method of the video generation prompt model as described above.

[0029] Finally, the present invention provides a computer-readable storage medium, on which a computer program is stored, characterized in that when the program is executed by a processor, the steps of the optimization method of the video generation prompt model as described above are implemented.

[0030] Compared with the prior art, the optimization method, system, device and storage medium of the video generation prompt model of the present invention have one or more of the following beneficial technical effects:

[0031] 1. The present invention combines principle-based supervised fine-tuning and preference optimization based on multiple feedbacks to optimize the video generation prompt model, and can construct a video generation prompt model that conforms to human values, enabling it to focus on the safety, accuracy and usefulness of the generated video, thereby promoting the generation of high-quality videos.

[0032] 2. In the principle-based supervised fine-tuning stage, the present invention uses the context learning ability of the LLM to construct supervised fine-tuning data, and then uses the LLM to criticize and further improve the supervised fine-tuning data based on principles such as harmlessness and alignment, which can ensure the safety, accuracy and usefulness of the supervised fine-tuning data.

[0033] 3. The present invention systematically optimizes the prompts input by users by combining text-level alignment feedback and feedback signals of video quality, ensuring that the optimized prompts are harmless, accurate and helpful for generating high-quality videos.

[0034] 4. The present invention is significantly superior to other video prompt technologies and can generate high-quality videos that are more in line with human preferences. Description of the Drawings

[0035] Figure 1 is a flowchart of the optimization method of the video generation prompt model of the present invention.

[0036] Figure 2 is a flowchart of constructing a safe, accurate and useful supervised fine-tuning data set in the present invention.

[0037] Figure 3 is a flowchart of constructing a text-level preference data set and a video-level preference data set in the present invention.

[0038] Figure 4 is a schematic diagram of the composition of the optimization system of the video generation prompt model of the present invention. Detailed Embodiments

[0039] Before describing any embodiments of the present invention in detail, it should be understood that the present invention is not limited in its application to the details of the construction and arrangement of components set forth in the following description or illustrated in the following drawings. The present invention is capable of other embodiments and of being practiced or carried out in various ways. Additionally, it should be understood that the language and terminology used herein are for the purpose of description and should not be regarded as limiting. As used herein, the terms "comprising" or "having" and their variants are intended to cover the listed items and their equivalents as well as additional items. Unless otherwise specified or limited, the terms "mounted," "connected," "supported," and "coupled" and their variants are used broadly and cover both direct and indirect mounting, connection, support, and coupling. Further, "connected" and "coupled" are not limited to physical or mechanical connection or coupling.

[0040] And, in the first aspect, in the disclosure of the present invention, the orientation or positional relationship indicated by terms such as "longitudinal," "transverse," "upper," "lower," "front," "rear," "left," "right," "vertical," "horizontal," "top," "bottom," "inner," "outer," etc. is based on the orientation or positional relationship shown in the drawings. It is only for the convenience of describing the present invention and simplifying the description, and does not indicate or imply that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation. Therefore, the above terms should not be construed as limiting the present invention; in the second aspect, the term "one" should be understood as "at least one" or "one or more." That is, in one embodiment, the number of an element can be one, while in other embodiments, the number of the element can be multiple. The term "one" should not be construed as limiting the quantity.

[0041] In view of the problems existing in the prior art, the present invention proposes an optimization method for a new video generation prompt model, which mainly includes principle-based supervised fine-tuning and preference optimization combining multiple feedbacks, aiming to make up for the gap in training and inference of the video generation model and improve the quality of video generation. When optimizing, three dimensions are mainly concerned:

[0042] 1. Security: When optimizing the prompt, harmless treatment should be carried out on user inputs with security problems to avoid harmful elements therein.

[0043] 2. Accuracy: Except for security scenarios, when optimizing the prompt, it should be ensured that it is precisely aligned with the user input, and the user's intention cannot be tampered with or important details in the user input be lost.

[0044] 3. Usefulness: The optimized prompt should be as friendly as possible to the video generation model to assist in generating high-quality videos.

[0045] Figure 1 The flowchart of the optimization method for the video generation prompt model of the present invention is shown. As Figure 1As shown in the figure, the optimization method of the video generation prompt model of the present invention includes the following steps:

[0046] I. Supervised fine-tuning training.

[0047] In the supervised fine-tuning stage, the present invention optimizes the prompt corresponding to the user input through a series of carefully designed steps, that is, the Prompt, aiming to construct a safe, accurate and useful supervised fine-tuning data set and use the supervised fine-tuning data set to perform supervised fine-tuning training on the video generation prompt model.

[0048] As Figure 2 shown in the figure, constructing a safe, accurate and useful supervised fine-tuning data set specifically includes:

[0049] 1. Data collection and collation.

[0050] First, collect diverse open-source user inputs for video generation, which cover various video generation scenarios, so as to ensure that the optimized prompts can adapt to a wide range of application requirements.

[0051] Since some of the user inputs involve potential security issues, it is necessary to perform diversity and rule-based quality screening on the user inputs and use existing security classifiers (such as ShieldGemma) to label the user inputs with security issues.

[0052] Among them, when performing diversity and rule-based quality screening, the rule can be that the user input cannot include unsafe hidden dangers, for example, "there cannot be bloody violence", etc. In this way, through diversity and rule-based quality screening, some user inputs involving potential security issues can be removed.

[0053] 2. Preliminary prompt generation.

[0054] Based on the collected and collated user inputs, use the large language model for in-context learning to generate preliminary prompts.

[0055] That is, rewrite the collected and collated user inputs through the large language model (LLM) plus in-context learning (ICL), that is, give the large language model some artificial rewriting examples and let the large language model imitate to generate preliminary prompts. This is a conventional technique in the field of large language models and will not be described in detail here.

[0056] 3. Criticism and problem identification.

[0057] Since there may be problems with the generated preliminary prompts in terms of security, accuracy, and usefulness, in the present invention, based on pre-established principles, a large language model is used to criticize and identify problems with the preliminary prompts.

[0058] Specifically, some principles can be pre-established manually. These principles are related to security, accuracy, and usefulness, such as "Do not omit the detailed requirements in the user input" and "The rewritten prompt should not have security risks, such as blood and violence". Then, the large language model is used to determine whether the generated preliminary prompt violates these principles and point out how it violates the principles specifically.

[0059] In this way, through criticism and problem identification, potential security risks, inaccurately expressed parts, and parts missing key information in the preliminary prompt can be identified.

[0060] 4. Improvement guided by rules.

[0061] Through criticism and problem identification, the problems existing in the generated preliminary prompt are found. Then, through the improvement ability of the large language model, combined with rules and critical feedback, the preliminary prompt can be further optimized. That is, some guiding rules are set manually, such as "Supplement the details omitted in the user input" and "Delete the security risks in the prompt, such as blood and violence". The large language model is used to modify the preliminary prompt with identified problems based on the guiding rules to obtain an optimized prompt, thereby generating the final supervised fine-tuning dataset.

[0062] Through the improvement guided by rules, the security and accuracy of the prompt are further optimized, thereby generating the final supervised fine-tuning dataset.

[0063] Through the above steps, the present invention constructs a supervised fine-tuning dataset covering different types of user inputs, and then the video generation prompt model can be standardly supervised and fine-tuned using the supervised fine-tuning dataset to train the video generation prompt model after supervised fine-tuning training.

[0064] II. Preference optimization training.

[0065] The preference optimization stage aims to further optimize the video generation prompt model after supervised fine-tuning training according to the feedback signals at the text level and the video level, so as to generate more harmless and high-quality videos. The feedback at the text level mainly focuses on whether there are security issues in the generated prompts and whether the original input of the user can be accurately expressed. The present invention uses a large language model to make judgments based on the established principles; the feedback at the video level mainly focuses on the quality of the generated videos, and the present invention uses a reward model for evaluation. That is, a text-level preference dataset including positive example data and negative example data and a video-level preference dataset including positive example data and negative example data are respectively constructed, and the video generation prompt model after supervised fine-tuning training is subjected to preference optimization training by using the text-level preference dataset and the video-level preference dataset.

[0066] As Figure 3 shown, the construction of the text-level preference dataset and the video-level preference dataset specifically includes:

[0067] 1. Data sampling.

[0068] When performing data sampling, the present invention uses the video generation prompt model after supervised fine-tuning training to generate multiple prompts for each user input respectively. Preferably, four prompts are generated for each user input respectively.

[0069] By generating multiple prompts for each user input, it can be used to construct preference optimization data pairs, improve the diversity of preference optimization data, and help obtain higher-quality preference optimization data.

[0070] 2. Construction of the text-level preference dataset.

[0071] In the preference optimization stage, the present invention gives priority to the feedback at the text level to ensure that the generated prompts have no security issues and can accurately express the original intention of the user. For this purpose, based on the pre-established principles, a large language model is used to respectively criticize multiple prompts for each user input to check whether there are problems.

[0072] When respectively criticizing multiple prompts for each user input, some principles can be formulated manually in advance. These principles are some principles related to security, accuracy, and usefulness, such as "the detailed requirements in the user input cannot be omitted", "the generated prompts cannot have security hazards, such as bloody violence, etc.", and then let the large language model judge whether the prompts violate these principles and point out how specifically they violate the principles. In this way, through criticism, the security hazards, inaccurately expressed parts, and parts missing key information in the prompts can be identified.

[0073] If there are problems, the large language model is used to modify the prompts with problems. When modifying the prompts with problems, the improvement ability of the large language model can be used to further optimize the prompts in combination with rules and critical feedback. That is, some guiding rules are set manually, such as "supplement the details missing in the user input" and "delete the potential safety hazards in the prompts, such as blood and violence", and the large language model is used to modify the prompts with problems based on the guiding rules to obtain the modified prompts.

[0074] Finally, the prompts with problems are used as negative example data, and the modified prompts are used as positive example data, thus forming a text-level preference dataset including positive example data and negative example data.

[0075] 3. Construction of video-level preference dataset.

[0076] For the prompts without problems, the present invention further introduces video-level feedback signals. That is, based on the prompts without problems, an existing video generation model is used to generate videos and an existing reward model is used to evaluate the quality of the generated videos. That is, the prompts without problems are input into the existing video generation model to generate videos, and then the generated videos are input into the existing reward model to obtain the quality scores of the videos. The prompts with the lowest video scores are used as negative example data, and the prompts with the highest scores are used as positive example data, thus forming a video-level preference dataset including positive example data and negative example data.

[0077] Finally, the text-level preference dataset including positive example data and negative example data and the video-level preference dataset including positive example data and negative example data collected through the above steps will be used for direct preference optimization (DPO) training of the video generation prompt model after supervised fine-tuning training. This process optimizes the video generation prompt model after supervised fine-tuning training through reinforcement learning methods, enabling it to generate prompts that are both harmless, accurate and can support high-quality video generation.

[0078] Figure 4 The schematic diagram of the composition of the optimization system of the video generation prompt model of the present invention is shown. As Figure 4 shown, the optimization system of the video generation prompt model of the present invention includes:

[0079] 1. Supervised fine-tuning training module.

[0080] The supervised fine-tuning training module is used to construct a safe, accurate and useful supervised fine-tuning dataset and use the supervised fine-tuning dataset to perform supervised fine-tuning training on the video generation prompt model.

[0081] 2. Preference optimization training module.

[0082] The preference optimization training module is used to respectively construct a text-level preference dataset including positive example data and negative example data and a video-level preference dataset including positive example data and negative example data, and use the text-level preference dataset and the video-level preference dataset to perform preference optimization training on the video generation prompt model after supervised fine-tuning training.

[0083] In addition, the present invention also provides an optimization device for a video generation prompt model, which includes: one or more processors; a memory for storing one or more programs; when the one or more programs are executed by the one or more processors, the one or more processors are caused to implement the optimization method of the video generation prompt model as described above.

[0084] Finally, the present invention also provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, the steps of the optimization method of the video generation prompt model as described above are implemented.

[0085] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than limiting the protection scope of the present invention. Those skilled in the art, based on the idea of the present invention, can modify or equivalently replace the technical solutions of the present invention without departing from the essence and scope of the technical solutions of the present invention.

Claims

1. A method for optimizing a video generation prompt model, characterized in that: The following steps are involved: 1) Supervised fine-tuning training: Construct a safe, accurate and useful supervised fine-tuning dataset and use the supervised fine-tuning dataset to perform supervised fine-tuning training on the video generation prompt model; 2) Preference optimization training: construct a text-level preference dataset including positive data and negative data and a video-level preference dataset including positive data and negative data respectively, and use the text-level preference dataset and the video-level preference dataset to perform preference optimization training on the video generation prompt model after supervised fine-tuning training.

2. The optimization method of the video generation prompt model according to claim 1, characterized in that: The step 1) of constructing a safe, accurate and useful supervised fine-tuning dataset specifically includes: 11) Data collection and organization: Collect various open-source user inputs for video generation, perform diversity and rule-based quality screening on the user inputs, and use security classifiers to mark user inputs with security issues; 12) Preliminary prompt generation: Based on the collected and organized user input, a large language model is used for context learning to generate preliminary prompts; 13) Criticism and problem identification: Based on the pre-established principles, the large language model is used to criticize and identify problems with the preliminary prompts; 14) Rule-guided refinement: The initial prompts of the identified problems are modified using a large language model to obtain optimized prompts, thereby generating the final supervised fine-tuning dataset.

3. The optimization method of the video generation prompt model according to claim 2, characterized in that: In the step 13), the criticism and problem identification includes identifying safety hazards, inaccurate expressions and parts that omit key information in the preliminary prompt.

4. The method for optimizing the video generation prompt model according to claim 3, characterized in that: The step 2) of constructing a text-level preference dataset including positive data and negative data and a video-level preference dataset including positive data and negative data respectively specifically includes: 21) Data sampling: Generate multiple prompts for each user input using the video prompt generation model trained by supervised fine-tuning; 22) Construction of text-level preference dataset: Based on the pre-established principles, a large language model is used to criticize multiple prompts input by each user to check whether there are any problems. If there are any problems, the large language model is used to modify the problematic prompts, where the problematic prompts are used as negative example data and the modified prompts are used as positive example data; 23) Construction of video-level preference dataset: For prompts that do not have problems, a video generation model is used to generate videos and a reward model is used to evaluate the quality of the generated videos. The prompt with the lowest video score is used as the negative data, and the prompt with the highest score is used as the positive data.

5. The method for optimizing the video generation prompt model according to claim 4, characterized in that: Checking whether there are problems in step 22) includes checking whether there are potential safety hazards, whether there are problems of inaccurate expression, and whether there are problems of missing key information.

6. A video generation prompt model optimization system, characterized in that: include: A supervised fine-tuning training module, which is used to construct a safe, accurate and useful supervised fine-tuning dataset and use the supervised fine-tuning dataset to perform supervised fine-tuning training on the video generation prompt model; A preference optimization training module is used to construct a text-level preference dataset including positive data and negative data and a video-level preference dataset including positive data and negative data, respectively, and use the text-level preference dataset and the video-level preference dataset to perform preference optimization training on a video generation prompt model after supervised fine-tuning training.

7. An optimization device for a video generation prompt model, characterized in that: include: one or more processors; A memory for storing one or more programs; When the one or more programs are executed by the one or more processors, the one or more processors implement the method for optimizing the video generation prompt model according to any one of claims 1 to 5.

8. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the steps of the method for optimizing the video generation prompt model as described in any one of claims 1 to 5 are implemented.