Model training method and device and video generation method and device

Through joint training of the AIGC model and the watermark extraction model, a video generation model is generated, which solves the problem of low watermark traceability accuracy in AIGC video, and realizes embedding recognizable watermark information in AIGC video and improves traceability accuracy.

CN120378559APending Publication Date: 2025-07-25ALIPAY (HANGZHOU) INFORMATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510506047.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-22
Publication Date
2025-07-25

AI Technical Summary

Technical Problem

When existing watermarking technology embeds watermarks in AIGC videos, the traceability accuracy is low, making it difficult to effectively track the source of the video.

Method used

By jointly training the AIGC model and the watermark extraction model, a video generation model is generated to ensure that the generated video contains watermark information and can be recognized by the extracted model, and avoid being extracted by the text extraction model.

Benefits of technology

The traceability accuracy of AIGC videos is improved, ensuring that the watermark information can be identified and not extracted by others, and effective traceability of the videos is achieved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120378559A_ABST
    Figure CN120378559A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a model training method and device and a video generation method and device, an AIGC model and a watermark extraction model can be jointly trained according to a sample text, sample watermark information and a text extraction model, the trained AIGC model is determined as a video generation model, and when a watermark is embedded in an AIGC video, the video generation efficiency is improved. A first text used for generating an AIGC video and first watermark information needing to be embedded can be input into a video generation model, and the video generation model can output a target video containing the first watermark information.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This document relates to the field of artificial intelligence technology, and in particular, to a method and device for model training and video generation. Background Art

[0002] Artificial Intelligence Generated Content (AIGC) is a technology that uses AI technology to automatically generate content. With the development of AIGC technology, more and more users use AIGC models (or AIGC large models) to generate content, such as generating videos (which can be represented as AIGC videos), etc.

[0003] In order to facilitate the traceability of AIGC videos, watermarks can be embedded in AIGC videos. However, current watermarking technologies are usually for real videos captured or recorded, while AIGC videos are randomly generated, and some video features are different from real videos. Therefore, when embedding watermarks in AIGC videos based on current watermarking technologies, problems such as performance degradation will occur, resulting in low traceability accuracy. Summary of the Invention

[0004] Embodiments of this specification provide a method and device for model training and video generation, which are used to solve the problem of low traceability accuracy when embedding watermarks in AIGC videos based on current watermarking technologies.

[0005] To solve the above technical problems, the embodiments of this specification are implemented as follows: In a first aspect, a model training method is proposed, including: Obtain sample text, sample watermark information, an Artificial Intelligence Generated Content (AIGC) model, a watermark extraction model, and a text extraction model. The AIGC model is used to generate videos, the watermark extraction model is used to extract watermark information from videos, and the text extraction model is used to extract the text based on which the videos are generated; Jointly train the AIGC model and the watermark extraction model according to the sample text, the sample watermark information, and the text extraction model, and determine the trained AIGC model as a video generation model. The video generation model is used to generate videos containing watermark information.

[0006] In a second aspect, a video generation method is proposed, including: Obtain first text and first watermark information; Input the first text and the first watermark information into a pre-trained video generation model to obtain a target video; Among them, the video generation model is obtained by training an AIGC model. The AIGC model is jointly trained with a watermark extraction model according to sample text, sample watermark information, and a text extraction model. The AIGC model is used to generate videos, the watermark extraction model is used to extract watermark information in the videos, the text extraction model is used to extract the text on which the generated videos are based, and the target video contains the first watermark information.

[0007] In a third aspect, a model training device is proposed, including: An acquisition module that acquires sample text, sample watermark information, an artificial intelligence generated content (AIGC) model, a watermark extraction model, and a text extraction model. The AIGC model is used to generate videos, the watermark extraction model is used to extract watermark information in the videos, and the text extraction model is used to extract the text on which the generated videos are based; A training module that jointly trains the AIGC model and the watermark extraction model according to the sample text, the sample watermark information, and the text extraction model, and determines the trained AIGC model as a video generation model. The video generation model is used to generate videos containing watermark information.

[0008] In a fourth aspect, a video generation device is proposed, including: An acquisition module that acquires first text and first watermark information; A video generation module that inputs the first text and the first watermark information into a pre-trained video generation model to obtain a target video; Among them, the video generation model is obtained by training an AIGC model. The AIGC model is jointly trained with a watermark extraction model according to sample text, sample watermark information, and a text extraction model. The AIGC model is used to generate videos, the watermark extraction model is used to extract watermark information in the videos, the text extraction model is used to extract the text on which the generated videos are based, and the target video contains the first watermark information.

[0009] In a fifth aspect, an electronic device is proposed, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the computer program is executed by the processor, it implements the steps of the method described in the first aspect, or implements the steps of the method described in the second aspect.

[0010] In a sixth aspect, a computer-readable storage medium is proposed. A computer program is stored on the computer-readable storage medium. When the computer program is executed by a processor, it implements the steps of the method described in the first aspect, or implements the steps of the method described in the second aspect.

[0011] In a seventh aspect, a computer program product is provided. The computer program product includes a non-transitory computer-readable storage medium storing a computer program, and the computer program is operable to cause a computer to execute some or all of the steps of the method described in the first aspect, or to execute some or all of the steps of the method described in the second aspect.

[0012] The embodiments of this specification provide a watermarking technology applicable to AIGC videos. According to sample text, sample watermark information, and a text extraction model, the AIGC model and the watermark extraction model can be jointly trained. The trained AIGC model is determined as a video generation model. When generating an AIGC video, the first text used to generate the AIGC video and the first watermark information to be embedded can be input into the video generation model, and the video generation model can output a target video containing the first watermark information, thereby achieving the purpose of embedding watermark information in the AIGC video. Among them, since the text extraction model is used for supervised training when training the video generation model, it can be ensured that the target video generated based on the video generation model can conform to the text description and the watermark information in it cannot be extracted by the text extraction model. In addition, since the watermark extraction model is jointly trained when training the video generation model, it can be ensured that the trained watermark extraction model can extract the watermark information in the target video generated based on the video generation model, thereby facilitating the traceability of the target video and improving the traceability accuracy. BRIEF DESCRIPTION OF THE DRAWINGS

[0013] To more clearly illustrate the technical solutions in the embodiments of this specification or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings described below are only some embodiments recorded in the embodiments of this specification. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0014] Figure 1 is a flowchart of a model training method according to an embodiment of this specification; Figure 2 is a flowchart of a video generation method according to an embodiment of this specification; Figure 3 is a structural diagram of an electronic device according to an embodiment of this specification; Figure 4 is a structural diagram of a model training device according to an embodiment of this specification Figure 5 is a structural diagram of a video generation device according to an embodiment of this specification. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0015] In related technologies, to facilitate the traceability of videos, watermark information can be embedded (or added) to videos. For example, personal information of the video owner can be embedded in the video. However, current watermark technologies are usually applicable to real videos obtained by shooting or recording. For AIGC videos, due to the strong randomness in the generation of AIGC videos and the significant differences in some video features (such as video continuity) from real videos, applying current watermark technologies to AIGC videos will reduce the traceability accuracy.

[0016] The embodiments of this specification provide a watermark technology applicable to AIGC videos. According to sample texts, sample watermark information, and a text extraction model, the AIGC model and the watermark extraction model can be jointly trained. The trained AIGC model is determined as the video generation model. When generating an AIGC video, the first text used to generate the AIGC video and the first watermark information to be embedded can be input into the video generation model, and the video generation model can output a target video containing the first watermark information, thereby achieving the purpose of embedding watermark information in the AIGC video. Among them, since the text extraction model is used for supervised training when training the video generation model, it can be ensured that the target video generated based on the video generation model conforms to the text description and the watermark information in it cannot be extracted by the text extraction model. In addition, since the watermark extraction model is jointly trained when training the video generation model, it can be ensured that the trained watermark extraction model can extract the watermark information in the target video generated based on the video generation model, thus facilitating the traceability of the target video and improving the traceability accuracy.

[0017] To enable those skilled in the art to better understand the technical solutions in the embodiments of this specification, the technical solutions in the embodiments of this specification will be clearly and completely described below in conjunction with the accompanying drawings in one or more embodiments of this specification. Obviously, the described embodiments are only a part of the embodiments of this specification, rather than all the embodiments. Based on the embodiments in this specification, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of this document.

[0018] It should be noted that the watermark technology provided in the embodiments of this specification is not only applicable to embedding watermarks in AIGC videos but also applicable to embedding watermarks in AIGC images (images generated according to the AIGC model). The embodiments of this specification are described by taking AIGC videos as an example.

[0019] The following will detail the technical solutions provided in each embodiment of this specification in conjunction with the accompanying drawings.

[0020] Figure 1 It is a flowchart of the model training method in an embodiment of this specification.Figure 1 The model training method shown can include the following steps.

[0021] S102: Obtain sample text, sample watermark information, an AIGC model, a watermark extraction model, and a text extraction model. The AIGC model is used to generate videos, the watermark extraction model is used to extract watermark information from videos, and the text extraction model is used to extract the text based on which the video is generated.

[0022] In the scenario of embedding watermark information in an AIGC video, a video generation model can be trained to generate a video containing watermark information according to this video generation model. Among them, when training the video generation model, sample text, sample watermark information, an AIGC model, a watermark extraction model, and a text extraction model can be obtained first. The AIGC model is a model that can use AI technology to generate videos (or AIGC videos), and specifically can be a currently publicly available model, such as the StyleGAN model, the stablediffusion model, etc., which are not specifically limited here. The sample text is the text used to generate a video (or AIGC video) during model training. That is, after inputting the sample text into the AIGC model, the AIGC model can output a video that conforms to the description of the sample text. Optionally, the sample text can be expressed as a prompt. The sample watermark information is the watermark information that needs to be embedded in the video during model training, and this sample watermark information can specifically be text information. The watermark extraction model is used to extract watermark information from videos, and specifically can be a watermark extractor, etc., which are not specifically limited here. The text extraction model is used to extract the text based on (or used) when generating the video. For example, the text extraction model can perform content analysis on the video generated by the AIGC model and extract the text based on which the AIGC model generates the video according to the analysis result. This text extraction model can be a currently publicly available model that summarizes and extracts the prompt for AIGC content.

[0023] S104: Jointly train the AIGC model and the watermark extraction model according to the sample text, sample watermark information, and the text extraction model, and determine the trained AIGC model as the video generation model. The video generation model is used to generate a video containing watermark information.

[0024] After obtaining the text, sample watermark information, AIGC model, watermark extraction model, and text extraction model, during model training, the AIGC model and the watermark extraction model can be jointly trained according to the sample text, sample watermark information, and the text extraction model.

[0025] During the model training process, the sample text and the AIGC model are used to generate videos without watermark information. The sample text, the sample watermark information, and the AIGC model are used to generate videos with watermark information. The text extraction model plays a supervisory role. Specifically, it can supervise whether the videos without watermark information generated by the AIGC model conform to the description of the sample text, and whether the watermark information in the videos with watermark information generated by the AIGC model can be extracted by the text extraction model. Thus, it can be ensured that after the video generation model is trained, the videos with watermark information generated based on the video generation model can conform to the text description, and the watermark information in the videos cannot be extracted by others. The watermark extraction model is used to extract the watermark information in the videos generated by the AIGC model, so that after the video generation model is trained, the watermark information in the videos generated based on the video generation model can be extracted for video traceability.

[0026] After the model training is completed, the trained AIGC model can be determined as the video generation model. This video generation model can be used to generate videos with watermark information. That is, by inputting the text for video generation and the watermark information to be embedded into the video generation model, the video generation model can output videos with watermark information.

[0027] It should be noted that the video generation model is also a type of AIGC model. Different from the AIGC models in the related technologies, the video generation model can generate videos with watermark information.

[0028] In some embodiments, the joint training of the AIGC model and the watermark extraction model according to the sample text, the sample watermark information, and the text extraction model may include: Input the sample text and the sample watermark information into the AIGC model to obtain the first video; Input the sample text into the AIGC model to obtain the second video; Based on the predicted values of the first video and the second video by the watermark extraction model and the predicted values of the first video and the second video by the text extraction model, jointly train the AIGC model and the watermark extraction model. The predicted values include predicted losses or predicted probabilities.

[0029] For the first video, we hope it conforms to the description of the sample text and contains the sample watermark information, and at the same time, the sample watermark information in it cannot be extracted by others using the text extraction model. For the second video, we hope the second video conforms to the description of the sample text and does not contain the sample watermark information. Therefore, after obtaining the first video and the second video, when performing model training, the watermark extraction model and the text extraction model can be used to process the first video and the second video to obtain the predicted values of the watermark extraction model for the first video and the second video and the predicted values of the text extraction model for the first video and the second video, and then the AIGC model and the watermark extraction model can be jointly trained based on these predicted values. Among them, the predicted value can be the prediction loss or prediction probability between the extraction result of the model for the video and the real sample. Optionally, the predicted value of the watermark extraction model for the first video and the second video can be the prediction loss, and the predicted value of the text extraction model for the first video and the second video can be the prediction probability.

[0030] In some embodiments, the predicted value of the watermark extraction model for the first video may include a first loss. The first loss can characterize the accuracy of the watermark extraction model for watermark extraction from the first video. The smaller the first loss, the higher the accuracy, and the larger the first loss, the lower the accuracy. Among them, the first loss can be determined according to the watermark extraction result of the watermark extraction model for the first video and the sample watermark information.

[0031] In some embodiments, the predicted value of the watermark extraction model for the second video may include a second loss. The second loss can characterize the false recall rate of the watermark extraction model for watermark extraction from the second video. The smaller the second loss, the lower the false recall rate, and the larger the second loss, the higher the false recall rate. Among them, the second loss can be determined according to the watermark extraction result of the watermark extraction model for the second video and the empty watermark information.

[0032] In some embodiments, the predicted values of the text extraction model for the first video may include a first probability and a second probability. The first probability can characterize the probability that the text extraction model extracts the sample text from the first video. The higher the probability, the more the first video conforms to the description of the sample text. The second probability can characterize the probability that the text extraction model extracts the sample watermark information from the first video. The lower the probability, the lower the possibility that the watermark information in the first video is extracted by others using the text extraction model. Among them, the first probability can be determined according to the text extraction result of the text extraction model for the first video and the sample text, and the second probability can be determined according to the text extraction result of the text extraction model for the first video and the sample watermark information.

[0033] In some embodiments, the predicted values of the text extraction model for the second video may include a third probability and a fourth probability. The third probability may represent the probability that the text extraction model extracts the sample text from the second video. The higher the probability, the more the second video conforms to the description of the sample text. The fourth probability may represent the probability that the text extraction model extracts the sample watermark information from the second video. The lower the probability, the lower the possibility that the text extraction model extracts the sample watermark information from the second video. Among them, the third probability may be determined according to the text extraction result of the text extraction model for the second video and the sample text, and the fourth probability may be determined according to the text extraction result of the text extraction model for the second video and the empty watermark information.

[0034] It should be noted that the above-mentioned third probability can play an anchoring role for the first probability to ensure that the text extraction results for the video are consistent when the sample watermark information is embedded and not embedded, that is, to ensure that the generated video still conforms to the description of the sample text after the sample watermark information is embedded. In addition, the above-mentioned fourth probability can play an anchoring role for the second probability to ensure that the sample watermark information cannot be extracted by the text extraction model whether the watermark information is embedded or not, so as to ensure that the embedded watermark information in the AIGC video cannot be extracted by others using the text extraction model. Based on these two anchoring effects, the text extraction model can play a supervisory role in the model training process, so as to train a model that meets the requirements.

[0035] In the case where the predicted values of the watermark extraction model for the first video and the second video include the above-mentioned first loss and second loss, and the predicted values of the text extraction model for the first video and the second video include the above-mentioned first probability, second probability, third probability and fourth probability, in some embodiments, when jointly training the AIGC video and the watermark extraction model according to these predicted values, it may include: Transmitting the first loss and the second loss to the AIGC model and the watermark extraction model, and transmitting the first probability, the second probability, the third probability and the fourth probability to the AIGC model to jointly train the AIGC model and the watermark extraction model; During the joint training process, optimize the first loss, the second loss, the first probability, the second probability, the third probability and the fourth probability.

[0036] Since the first loss can represent the accuracy of the watermark extraction model in extracting watermarks from the first video, and the second loss can represent the false recall rate of the watermark extraction model in extracting watermarks from the second video, these two losses will affect the performance of the watermark extraction model and also have a certain impact on the performance of the AIGC model. Thus, during model training, the first loss and the second loss can be passed to the AIGC model and the watermark extraction model to train the AIGC model and the watermark extraction model based on the first loss and the second loss.

[0037] Since the first probability can represent the probability that the text extraction model extracts sample text from the first video, the second probability can represent the probability that the text extraction model extracts sample watermark information from the first video, the third probability can represent the probability that the text extraction model extracts sample text from the second video, and the fourth probability can represent the probability that the text extraction model extracts sample watermark information from the second video, the first probability, the second probability, the third probability, and the fourth probability will affect the performance of the AIGC model. Thus, during model training, the first probability, the second probability, the third probability, and the fourth probability can be passed to the AIGC model to train the AIGC model based on the first probability, the second probability, the third probability, and the fourth probability.

[0038] After passing the first loss and the second loss to the AIGC model and the watermark extraction model, and passing the first probability, the second probability, the third probability, and the fourth probability to the AIGC model, during the process of jointly training the AIGC model and the watermark extraction model, the first loss, the second loss, the first probability, the second probability, the third probability, and the fourth probability can be continuously optimized to obtain a model that meets the requirements.

[0039] In some embodiments, when optimizing the first loss, the second loss, the first probability, the second probability, the third probability, and the fourth probability, it may include: Minimize the first loss; Minimize the second loss; Minimize the relative entropy (i.e., KL divergence) between the first probability and the third probability; Minimize the relative entropy (i.e., KL divergence) between the second probability and the fourth probability.

[0040] Since the first loss can represent the accuracy of the watermark extraction model in extracting watermarks from the first video, and the smaller the first loss, the higher the accuracy, thus, when optimizing, the first loss can be minimized.

[0041] Since the second loss can represent the false recall rate of the watermark extraction model in extracting watermarks from the second video, and the smaller the second loss, the lower the false recall rate, thus, when optimizing, the second loss can be minimized.

[0042] Since the first probability can represent the probability that the text extraction model extracts the sample text from the first video, and the third probability can represent the probability that the text extraction model extracts the sample text from the second video, the third probability can anchor the first probability, that is, it is expected that after embedding the watermark information in the video, the video still conforms to the text description. Therefore, when optimizing, the relative entropy between the first probability and the third probability can be minimized.

[0043] Since the second probability can represent the probability that the text extraction model extracts the sample watermark information from the first video, and the fourth probability can represent the probability that the text extraction model extracts the sample watermark information from the second video, the fourth probability can anchor the second probability, that is, it is expected that after embedding the watermark in the video, the watermark cannot be extracted by the text extraction model. Therefore, when optimizing, the relative entropy between the second probability and the fourth probability can be minimized.

[0044] After optimizing the first loss, the second loss, the first probability, the second probability, the third probability, and the fourth probability, when the optimization conditions are met (or other conditions for ending the training are met), the training can be ended, and the trained AIGC model can be determined as the video generation model. This video generation model can be used to generate videos containing watermark information, and the watermark information in the video can be extracted by the trained watermark extraction model but cannot be extracted by the text extraction model (that is, cannot be extracted by others).

[0045] To facilitate the understanding of the training method of the video generation model in the embodiments of this specification, the following will take the AIGC model as the currently publicly available stable-diffusion model (abbreviated as SD) as an example for illustration.

[0046] Before model training, the following four components or models can be defined first: SD-text encoder: Encodes text information; SD-diffusion: Denoises to generate videos; Extractor: Extracts watermark results; Open-Revertor: A publicly available model that summarizes AIGC content and extracts prompts.

[0047] Among them, SD-text encoder and SD-diffusion are two components in the SD model. SD-textencoder can understand the semantic information of the input text, and SD-diffusion can generate AIGC videos (or generate images, here taking video generation as an example) according to the semantic information understood by SD-text encoder.

[0048] When performing model training, for a sample prompt used to generate a video, it can be concatenated with the sample watermark text. For example, if the sample prompt is origin_prompt and the sample watermark text is build by antgroup, after concatenating the two, the concatenated result is: <water>build by antgroup< / water> origin_prompt. After obtaining the concatenated result, the concatenated result can be input into the SD-text encoder to generate Video 1. When generating Video 1, the Extractor can be used to extract the watermark text in Video 1, and the watermark extraction result can be compared with build by antgroup to determine the <water>xxx< / water> prediction loss L11 for this segment. L11 represents the accuracy of the Extractor in extracting the watermark from Video 1. The smaller L11 is, the higher the accuracy; the larger L11 is, the lower the accuracy. In addition, the Open-Revertor can also be used to extract or restore the text based on which Video 1 is generated. The text extraction result is compared with origin_prompt to determine the prediction probability P21 of the Open-Revertor for origin_prompt, and the text extraction result is compared with build by antgroup to determine the <water>The prediction probability P31 of the segment. P21 can represent the probability that Open-Revertor extracts the origin_prompt from Video 1. The higher the probability, the more Video 1 conforms to the description of the origin_prompt. P31 can represent the probability that Open-Revertor extracts the sample watermark information from Video 1. The lower the probability, the lower the probability that the watermark information in Video 1 is extracted by others.

[0049] After that, the origin_prompt without watermark can be input into the SD-text encoder to generate Video 2. When generating Video 2, the Extractor can be used to extract the watermark text in Video 2 and compare the watermark extraction result with the empty watermark to determine the <water>< / water> prediction loss L12 of the Extractor. L12 represents the false positive recall rate of the Extractor for watermark extraction in Video 2. The smaller L12 is, the lower the false positive recall rate. The larger L12 is, the higher the false positive recall rate. In addition, the Open-Revertor can also be used to extract or restore the text based on which Video 2 is generated, compare the text extraction result with the origin_prompt to determine the prediction probability P22 of the Open-Revertor for the origin_prompt, and compare the text extraction result with the empty watermark to determine the Open-Revertor's <water>The predicted probability P32 of the segment. P22 can represent the probability that Open-Revertor extracts origin_prompt from Video 2. The higher the probability, the more Video 2 conforms to the description of origin_prompt. P32 can represent the probability that Open-Revertor extracts sample watermark information from Video 2. The lower the probability, the lower the possibility of extracting watermark information from Video 2.

[0050] After that, freeze open-revertor (i.e., do not change the model parameters of open-revertor), and optimize L11, L12, P21, P22, P31, and P32. Specifically, the optimization function can be as follows: min(L11); min(L12); min(KL(P21||P22) + KL(P31||P32)).

[0051] That is, make Extractor able to extract the watermark as much as possible without affecting the extraction result of Open-Revertor.

[0052] After that, new noise can be introduced in each round of denoising process, and min(L11) can be optimized to ensure the extraction stability of the watermark. After iteratively optimizing L11, L12, P21, P22, P31, and P32, when the optimization conditions or other conditions for ending training are met, the trained SD-text encoder and SD-diffusion can be determined as the video generation model. This video generation model is used to generate videos containing watermarks. The watermark information in this video can be extracted by the trained Extractor, but cannot be extracted by Open-Revertor, and the text extraction result of Open-Revertor for this video is consistent (or basically consistent) with the prompt based on which the video is generated.

[0053] Compared with invisible watermarks, there is no strong consistency in the generated content itself. Even if visible watermarks are added during the generation process, they will have an impact during subsequent denoising, and finally invisible but "perceptible" watermarks will be generated in the generated content. Therefore, stronger robustness can be obtained, and it performs better on AIGC images / videos. In addition, through the contrast loss of Open-Revertor, the risk of the watermark being detected by the model is reduced.

[0054] Based on the video generation model generated by the model training method provided in the embodiments of this specification, the embodiments of this specification also provide a video generation method, which can generate videos containing watermark information. Please refer to Figure 2 .

[0055] Figure 2 It is a schematic flowchart of a video generation method according to an embodiment of this specification. Figure 2 The video generation method shown may include the following steps.

[0056] S202: Obtain the first text and the first watermark information.

[0057] When generating a target video containing watermark information, the first text and the first watermark information can be obtained first. Among them, the first text is the text based on which the target video is generated, and the generated target video needs to conform to the description of the first text. Optionally, the first text can be expressed as a prompt. The first watermark information is the watermark information to be embedded in the target video, and this watermark information can be text information.

[0058] S204: Input the first text and the first watermark information into a pre-trained video generation model to obtain a target video. Among them, the video generation model is obtained by training an AIGC model. The AIGC model is jointly trained with a watermark extraction model according to sample text, sample watermark information, and a text extraction model. The AIGC model is used to generate videos, the watermark extraction model is used to extract the watermark information in the videos, the text extraction model is used to extract the text based on which the videos are generated, and the target video contains the first watermark information.

[0059] After obtaining the first text and the first watermark information, the two can be spliced, and then the splicing result is input into the pre-trained video generation model, and the video generation model can output the target video. Among them, for the training method of the video generation model, reference can be made to Figure 1 the embodiment shown, and it will not be repeated here. The generated target video conforms to the description of the first text and contains the first watermark information.

[0060] In some embodiments, after generating the target video, in the scenario of tracing the target video, it may include: Extract the first watermark information in the target video according to the trained watermark extraction model; Trace the target video according to the first watermark information.

[0061] Since the watermark extraction model is jointly trained when training the video generation model, it can be ensured that the trained watermark extraction model can extract the watermark information in the video generated based on the video generation model. In this way, after generating the target video according to the video generation model, when tracing the target video, the trained watermark extraction model can be used to extract the first watermark information in the target video, and the target video can be traced according to the first watermark information, so that the tracing of the target video can be realized and the tracing accuracy can be improved.

[0062] The embodiments of this specification provide a watermarking technology applicable to AIGC videos. According to sample texts, sample watermark information, and text extraction models, the AIGC model and the watermark extraction model can be jointly trained. The trained AIGC model is determined as the video generation model. When generating an AIGC video, the first text used to generate the AIGC video and the first watermark information to be embedded can be input into the video generation model, and the video generation model can output a target video containing the first watermark information, thereby achieving the purpose of embedding watermark information in the AIGC video. Among them, since the text extraction model is used for supervised training when training the video generation model, it can be ensured that the watermark information in the target video generated by the video generation model cannot be extracted by the text extraction model, thus avoiding the extraction of the watermark information in the target video by others. In addition, since the watermark extraction model is jointly trained when training the video generation model, it can be ensured that the trained watermark extraction model can extract the watermark information in the target video generated by the video generation model, thereby facilitating the traceability of the target video and improving the traceability accuracy.

[0063] The above describes specific embodiments of this specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be performed in a different order than in the embodiments and still achieve the desired results. Additionally, the processes depicted in the figures do not necessarily require the specific order or sequential order shown to achieve the desired results. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0064] Figure 3 is a schematic structural diagram of an electronic device according to an embodiment of this specification. Please refer to Figure 3 , at the hardware level, the electronic device includes a processor, and optionally also includes an internal bus, a network interface, and a memory. Among them, the memory may include internal memory, such as high-speed random access memory (Random-Access Memory, RAM), and may also include non-volatile memory, such as at least one disk memory, etc. Of course, the electronic device may also include other hardware required for other services.

[0065] The processor, network interface, and memory can be interconnected through an internal bus, which can be an ISA (Industry Standard Architecture) bus, a PCI (Peripheral Component Interconnect) bus, an EISA (Extended Industry Standard Architecture) bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 3 only a bidirectional arrow is used in

[0066] Memory, for storing programs. Specifically, the program can include program code, and the program code includes computer operation instructions. The memory can include a memory and a non-volatile memory, and provide instructions and data to the processor.

[0067] The processor reads the corresponding computer program from the non-volatile memory into the memory and then runs it, forming a model training device at the logical level. The processor executes the program stored in the memory and is specifically used to perform the following operations: Obtain a sample text, sample watermark information, an artificial intelligence generated content (AIGC) model, a watermark extraction model, and a text extraction model, where the AIGC model is used to generate a video, the watermark extraction model is used to extract watermark information from the video, and the text extraction model is used to extract the text based on which the video is generated; According to the sample text, the sample watermark information, and the text extraction model, jointly train the AIGC model and the watermark extraction model, and determine the trained AIGC model as a video generation model, where the video generation model is used to generate a video containing watermark information.

[0068] It should be understood that the model training device in this embodiment can be used as Figure 1 the execution subject of the Figure 1 shown method, and thus can implement the steps and functions of the

[0069] shown method, which will not be elaborated here. Obtain a first text and first watermark information; Input the first text and the first watermark information into a pre-trained video generation model to obtain a target video; Among them, the video generation model is obtained by training an AIGC model. The AIGC model is jointly trained with a watermark extraction model according to sample text, sample watermark information, and a text extraction model. The AIGC model is used to generate videos, the watermark extraction model is used to extract watermark information from videos, the text extraction model is used to extract the text on which the generated videos are based, and the target video contains the first watermark information.

[0070] It should be understood that the video generation device in this embodiment can be used as Figure 2 the execution subject of the method shown, and thus can implement Figure 2 the steps and functions of the method shown, which will not be elaborated here.

[0071] As described above in this specification Figure 1 or Figure 2 The method disclosed in the embodiments shown can be applied to or implemented by a processor. The processor may be an integrated circuit chip with signal processing capabilities. During implementation, the steps of the above method can be completed by the integrated logic circuit in the hardware of the processor or by instructions in software form. The above processor may be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it may also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. It can implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of this specification. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. The steps of the method disclosed in combination with the embodiments of this specification can be directly implemented by the execution of the hardware decoding processor or by the combination of the hardware and software modules in the decoding processor. The software module may be located in a mature storage medium in the art such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory, or an electrically erasable programmable memory, a register, etc. This storage medium is located in the memory, and the processor reads the information in the memory and combines its hardware to complete the steps of the above method.

[0072] Of course, in addition to the software implementation, the electronic devices in the embodiments of this specification do not exclude other implementation manners, such as logic devices or a combination of software and hardware, etc. That is to say, the execution subject of the following processing flow is not limited to each logic unit, and can also be hardware or logic devices.

[0073] The embodiments of this specification also propose a computer-readable storage medium that stores one or more programs. The one or more programs include instructions that, when executed by a portable electronic device including multiple application programs, can cause the portable electronic device to execute Figure 1 the method of the illustrated embodiment, and specifically used to perform the following operations: Obtain a sample text, sample watermark information, an artificial intelligence generated content (AIGC) model, a watermark extraction model, and a text extraction model. The AIGC model is used to generate a video, the watermark extraction model is used to extract the watermark information in the video, and the text extraction model is used to extract the text on which the video is generated; According to the sample text, the sample watermark information, and the text extraction model, jointly train the AIGC model and the watermark extraction model, and determine the trained AIGC model as a video generation model. The video generation model is used to generate a video containing watermark information.

[0074] Alternatively, when the instruction is executed by a portable electronic device including multiple application programs, it can cause the portable electronic device to execute Figure 2 the method of the illustrated embodiment, and specifically used to perform the following operations: Obtain a first text and first watermark information; Input the first text and the first watermark information into a pre-trained video generation model to obtain a target video; Wherein, the video generation model is obtained by training an AIGC model. The AIGC model is jointly trained with the watermark extraction model according to the sample text, the sample watermark information, and the text extraction model. The AIGC model is used to generate a video, the watermark extraction model is used to extract the watermark information in the video, the text extraction model is used to extract the text on which the video is generated, and the target video contains the first watermark information.

[0075] In addition, the embodiments of the present disclosure also propose a computer program product, which includes a non-transitory computer-readable storage medium storing a computer program. The computer program is operable to cause a computer to execute: Obtain sample text, sample watermark information, an artificial intelligence-generated content (AIGC) model, a watermark extraction model, and a text extraction model. The AIGC model is used to generate videos. The watermark extraction model is used to extract watermark information from videos. The text extraction model is used to extract the text based on which the videos are generated. According to the sample text, the sample watermark information, and the text extraction model, jointly train the AIGC model and the watermark extraction model, and determine the trained AIGC model as a video generation model. The video generation model is used to generate videos containing watermark information.

[0076] Or cause a computer to execute: Obtain first text and first watermark information; Input the first text and the first watermark information into a pre-trained video generation model to obtain a target video; Wherein, the video generation model is obtained by training an AIGC model. The AIGC model is jointly trained with the watermark extraction model according to sample text, sample watermark information, and the text extraction model. The AIGC model is used to generate videos. The watermark extraction model is used to extract watermark information from videos. The text extraction model is used to extract the text based on which the videos are generated. The target video contains the first watermark information.

[0077] Figure 4 It is a schematic structural diagram of a model training device 40 according to an embodiment of this specification. Please refer to Figure 4 In a software implementation manner, the model training device 40 may include: An acquisition module 41, which acquires sample text, sample watermark information, an artificial intelligence-generated content (AIGC) model, a watermark extraction model, and a text extraction model. The AIGC model is used to generate videos. The watermark extraction model is used to extract watermark information from videos. The text extraction model is used to extract the text based on which the videos are generated. A training module 42, which jointly trains the AIGC model and the watermark extraction model according to the sample text, the sample watermark information, and the text extraction model, and determines the trained AIGC model as a video generation model. The video generation model is used to generate videos containing watermark information.

[0078] In some implementation manners, the training module 42 jointly trains the AIGC model and the watermark extraction model according to the sample text, the sample watermark information, and the text extraction model, including: Input the sample text and the sample watermark information into the AIGC model to obtain a first video; Input the sample text into the AIGC model to obtain a second video; Jointly train the AIGC model and the watermark extraction model according to the prediction values of the watermark extraction model for the first video and the second video and the prediction values of the text extraction model for the first video and the second video, where the prediction values include prediction losses or prediction probabilities.

[0079] In some embodiments, the prediction value of the watermark extraction model for the first video includes a first loss, and the prediction value for the second video includes a second loss. The first loss represents the accuracy of the watermark extraction model for extracting watermarks from the first video, and the second loss represents the false recall rate of the watermark extraction model for extracting watermarks from the second video; The prediction values of the text extraction model for the first video include a first probability and a second probability, and the prediction values for the second video include a third probability and a fourth probability. The first probability represents the probability that the text extraction model extracts the sample text from the first video, the second probability represents the probability that the text extraction model extracts the sample watermark information from the first video, the third probability represents the probability that the text extraction model extracts the sample text from the second video, and the fourth probability represents the probability that the text extraction model extracts the sample watermark information from the second video.

[0080] In some embodiments, the training module 42 jointly trains the AIGC model and the watermark extraction model according to the prediction values of the watermark extraction model for the first video and the second video and the prediction values of the text extraction model for the first video and the second video, including: Transmit the first loss and the second loss to the AIGC model and the watermark extraction model, and transmit the first probability, the second probability, the third probability, and the fourth probability to the AIGC model to jointly train the AIGC model and the watermark extraction model; During the joint training process, optimize the first loss, the second loss, the first probability, the second probability, the third probability, and the fourth probability.

[0081] In some embodiments, the training module 42 optimizes the first loss, the second loss, the first probability, the second probability, the third probability, and the fourth probability, including: Minimize the first loss; Minimize the second loss; Minimize the relative entropy of the first probability and the third probability; Minimize the relative entropy between the second probability and the fourth probability.

[0082] The above modules in the model training device provided by the embodiments of this specification can also implement the method steps provided by the above embodiments of the model training method. Alternatively, the model training device provided by the embodiments of this specification may further include other modules in addition to the above modules to implement the method steps provided by the above embodiments of the model training method. Moreover, the model training device provided by the embodiments of this specification can achieve the technical effects that the above embodiments of the model training method can achieve, which will not be elaborated here.

[0083] Figure 5 It is a schematic structural diagram of a video generation device 50 according to an embodiment of this specification. Please refer to Figure 5 , in a software implementation manner, the video generation device 50 may include: An acquisition module 51 that acquires a first text and first watermark information; A video generation module 52 that inputs the first text and the first watermark information into a pre-trained video generation model to obtain a target video; Among them, the video generation model is obtained by training an AIGC model. The AIGC model is jointly trained with a watermark extraction model according to a sample text, sample watermark information, and a text extraction model. The AIGC model is used to generate videos, the watermark extraction model is used to extract watermark information in videos, the text extraction model is used to extract the text on which the generated videos are based, and the target video includes the first watermark information.

[0084] In some embodiments, the video generation device 50 further includes a tracing module, and the tracing module: Extracts the first watermark information in the target video according to the trained watermark extraction model; Traces the target video according to the first watermark information.

[0085] The above modules in the video generation device provided by the embodiments of this specification can also implement the method steps provided by the above embodiments of the video generation method. Alternatively, the video generation device provided by the embodiments of this specification may further include other modules in addition to the above modules to implement the method steps provided by the above embodiments of the video generation method. Moreover, the video generation device provided by the embodiments of this specification can achieve the technical effects that the above embodiments of the video generation method can achieve, which will not be elaborated here.

[0086] In summary, the above description is only a preferred embodiment of this specification and is not intended to limit the protection scope of this document. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of one or more embodiments of this specification shall be included within the protection scope of this document.

[0087] The systems, devices, modules, or units illustrated in the above embodiments can be specifically implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a computer. Specifically, the computer can be, for example, a personal computer, a laptop computer, a cellular phone, a camera phone, a smart phone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or any combination of these devices.

[0088] Computer-readable media includes both permanent and non-permanent, removable and non-removable media and can be implemented by any method or technology for storing information. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile discs (DVD) or other optical storage, magnetic cassettes, magnetic tape magnetic disk storage or other magnetic storage devices, or any other non-transitory medium that can be used to store information accessible by a computing device. As defined herein, computer-readable media does not include transitory computer-readable media such as modulated data signals and carrier waves.

[0089] It should also be noted that the term "comprising", "including" or any other variation thereof is intended to cover a non-exclusive inclusion, such that a process, method, commodity or device comprising a series of elements not only includes those elements but also includes other elements not expressly listed, or further includes elements inherent to such process, method, commodity or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the existence of additional identical elements in the process, method, commodity or device comprising the said element.

[0090] Each embodiment in this specification is described in a progressive manner. For the same or similar parts among the embodiments, reference can be made to each other, and the differences between each embodiment and other embodiments are emphasized. In particular, for the system embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and for the relevant parts, reference can be made to the partial description of the method embodiments.< / water> < / water>

Claims

1. A model training method, comprising: Obtaining sample text, sample watermark information, an artificial intelligence generated content (AIGC) model, a watermark extraction model, and a text extraction model, where the AIGC model is used to generate videos, the watermark extraction model is used to extract watermark information from videos, and the text extraction model is used to extract the text based on which the videos are generated; Jointly training the AIGC model and the watermark extraction model according to the sample text, the sample watermark information, and the text extraction model, and determining the trained AIGC model as a video generation model, where the video generation model is used to generate videos containing watermark information.

2. The method according to claim 1, where the jointly training the AIGC model and the watermark extraction model according to the sample text, the sample watermark information, and the text extraction model comprises: Inputting the sample text and the sample watermark information into the AIGC model to obtain a first video; Inputting the sample text into the AIGC model to obtain a second video; Jointly training the AIGC model and the watermark extraction model according to the predicted values of the watermark extraction model for the first video and the second video and the predicted values of the text extraction model for the first video and the second video, where the predicted values include predicted losses or predicted probabilities.

3. The method according to claim 2, where the predicted value of the watermark extraction model for the first video includes a first loss, and the predicted value of the watermark extraction model for the second video includes a second loss, where the first loss represents the accuracy of the watermark extraction model in extracting watermarks from the first video, and the second loss represents the false recall rate of the watermark extraction model in extracting watermarks from the second video; The predicted values of the text extraction model for the first video include a first probability and a second probability, and the predicted values of the text extraction model for the second video include a third probability and a fourth probability, where the first probability represents the probability that the text extraction model extracts the sample text from the first video, the second probability represents the probability that the text extraction model extracts the sample watermark information from the first video, the third probability represents the probability that the text extraction model extracts the sample text from the second video, and the fourth probability represents the probability that the text extraction model extracts the sample watermark information from the second video.

4. The method according to claim 3, where the jointly training the AIGC model and the watermark extraction model according to the predicted values of the watermark extraction model for the first video and the second video and the predicted values of the text extraction model for the first video and the second video comprises: Transmitting the first loss and the second loss to the AIGC model and the watermark extraction model, and transmitting the first probability, the second probability, the third probability, and the fourth probability to the AIGC model to jointly train the AIGC model and the watermark extraction model; During the joint training process, optimize the first loss, the second loss, the first probability, the second probability, the third probability, and the fourth probability.

5. The method according to claim 4, wherein optimizing the first loss, the second loss, the first probability, the second probability, the third probability, and the fourth probability includes: Minimize the first loss; Minimize the second loss; Minimize the relative entropy of the first probability and the third probability; Minimize the relative entropy of the second probability and the fourth probability.

6. A video generation method, comprising: Obtain a first text and first watermark information; Input the first text and the first watermark information into a pre-trained video generation model to obtain a target video; Wherein, the video generation model is obtained by training an AIGC model, and the AIGC model is jointly trained with a watermark extraction model according to sample texts, sample watermark information, and a text extraction model. The AIGC model is used to generate videos, the watermark extraction model is used to extract watermark information in the videos, the text extraction model is used to extract the text based on which the videos are generated, and the target video contains the first watermark information.

7. The method according to claim 6, further comprising: Extract the first watermark information in the target video according to the trained watermark extraction model; Trace the source of the target video according to the first watermark information.

8. A model training device, comprising: An acquisition module, which acquires sample texts, sample watermark information, an artificial intelligence generated content (AIGC) model, a watermark extraction model, and a text extraction model. The AIGC model is used to generate videos, the watermark extraction model is used to extract watermark information in the videos, and the text extraction model is used to extract the text based on which the videos are generated; A training module, which jointly trains the AIGC model and the watermark extraction model according to the sample texts, the sample watermark information, and the text extraction model, and determines the trained AIGC model as a video generation model. The video generation model is used to generate videos containing watermark information.

9. A video generation device, comprising: An acquisition module, which acquires a first text and first watermark information; A video generation module, which inputs the first text and the first watermark information into a pre-trained video generation model to obtain a target video; Wherein, the video generation model is obtained by training an AIGC model, and the AIGC model is jointly trained with a watermark extraction model according to sample texts, sample watermark information, and a text extraction model. The AIGC model is used to generate videos, the watermark extraction model is used to extract watermark information in the videos, the text extraction model is used to extract the text based on which the videos are generated, and the target video contains the first watermark information.

10. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein when the computer program is executed by the processor, it implements the steps of the method according to any one of claims 1 to 5, or implements the steps of the method according to any one of claims 6 to 7.

11. A computer-readable storage medium, having a computer program stored thereon, wherein when the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 5, or implements the steps of the method according to any one of claims 6 to 7.

12. A computer program product, comprising a non-transitory computer-readable storage medium storing a computer program, wherein the computer program is operable to cause a computer to execute some or all of the steps of the method according to any one of claims 1 to 5, or execute some or all of the steps of the method according to any one of claims 6 to 7.