Watermark embedding method and apparatus for video diffusion model
By embedding watermarks through training a video diffusion model, the problem of insufficient protection in existing technologies is solved, efficient generation of watermarked videos is achieved, and the security and flexibility of the model are improved.
Patent Information
- Application Number
- PCT/CN2024/137506
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-04-29
- Filing Date
- 2024-12-06
- Publication Date
- 2025-09-25
AI Technical Summary
The existing technology has a limited protection effect on the diffusion model when protecting the generated video. The model needs to be obtained and disassembled to identify the watermark information, which reduces the security of the model.
The video diffusion model is trained using watermarked videos and prompt words through computing devices, so that the trained model can generate videos containing target watermarks, and the model parameters and watermark-related parameters are not limited to specific layers, avoiding watermark attacks.
It improves the efficiency of video generation, enhances the protection effect of the model, reduces the possibility of watermark attacks, and meets the diverse needs of users.
Smart Images

Figure CN2024137506_25092025_PF_FP_ABST
Abstract
Description
A video diffusion model watermark embedding method and device
[0001] This application claims priority to the Chinese patent application filed with the State Intellectual Property Office on March 18, 2024, with application number 202410316370.X and application name “A model watermark generation and embedding method and device”. This application also claims priority to the Chinese patent application filed with the State Intellectual Property Office on April 29, 2024, with application number 202410546990.2 and invention name “A video diffusion model watermark embedding method and device”. The entire contents of the above applications are incorporated into this application by reference. Technical Field
[0002] The present application relates to the field of artificial intelligence technology, and in particular to a video diffusion model watermark embedding method and device. Background Art
[0003] Video generation refers to the process by which a diffusion model generates a video based on first input information. To protect the diffusion model used to generate the video, the processing layer or network layer specified in the diffusion model can be modified to a watermark layer that carries watermark information. However, this method requires obtaining and disassembling the diffusion model to obtain the watermark information carried in the diffusion model in order to identify the source of the diffusion model, resulting in limited protection of the diffusion model. Summary of the Invention
[0004] The present application provides a video diffusion model watermark embedding method and device, which are used to solve the technical problem that the diffusion model of the generated video cannot be effectively protected.
[0005] In a first aspect, the present application provides a video diffusion model watermark embedding method. The method can be executed by a computing device, a chip (such as a processor) in a computing device, or a computing device cluster consisting of multiple computing devices. The embedding method includes: the computing device obtains a first video diffusion model and training data. The first video diffusion model is used to generate a first video based on first input information. The first input information includes a target prompt word. The training data includes: a prompt word and multiple groups of watermarked videos. Each group of watermarked videos includes: a training video and a watermark of the training video. The computing device uses the training data to train the first video diffusion model to obtain a second video diffusion model. The second video diffusion model is used to generate a second video based on the first input information. The second video includes a target watermark that matches the target prompt word.
[0006] Compared to the need to obtain and disassemble the model to obtain the watermark, and then identify the source of the model based on the watermark. In this application, the computing device uses the watermarked video and the prompt word to train the video diffusion model, so that the trained video diffusion model can generate a video containing a target watermark corresponding to the target prompt word based on the target prompt word in the first input information, without the need for post-processing or secondary processing of the video, so that the video has the target watermark, thereby improving the efficiency of video generation. Moreover, because the training data of the video diffusion model contains the target watermark and the target prompt word, the model parameters related to the target prompt word and the target watermark in the video diffusion model are not limited to a specific layer or layers, but all model parameters are related to the watermark, avoiding the problem of model attackers launching watermark attacks on the video diffusion model, resulting in reduced model security, and improving the protection effect of the video diffusion model.
[0007] In one possible implementation, the method further includes: the computing device displays a first window. In response to the user's first input operation, the computing device uses the first window to display the first information input by the user. The first information includes one or more of the following: one or more prompt words, one or more watermarks. In the case where the first information includes a prompt word, the prompt word included in the first information is a target prompt word. In the case where the first information includes multiple prompt words, the multiple prompt words include the target prompt word. In the case where the first information includes a watermark, the watermark included in the first information is a target watermark. In the case where the first information includes multiple watermarks, the multiple watermarks include the target watermark. The computing device uses the first information to obtain training data. In this way, the user can flexibly set one or more prompt words and watermarks according to the needs of actual applications, meet the diverse needs of users, and improve friendliness.
[0008] In another possible implementation, if the first information includes multiple prompt words and multiple watermarks, the computing device uses the first information to obtain training data, including: the computing device displays a second window. In response to a second user input operation, the computing device displays the second information input by the user in the second window. The second information includes a correspondence between the target prompt word and the target watermark. The computing device obtains training data based on the first and second information. In this way, if the video diffusion model has multiple sources, the user can set the correspondence to reflect the multiple sources of the video diffusion model, meeting the diverse needs of users and improving user friendliness.
[0009] In another possible implementation, the computing device uses the second video diffusion model to simultaneously generate the video data and the target watermark included in the second video according to the first input information.
[0010] In another possible implementation, the computing device can store the video data and the target watermark in a single data stream. This makes it easier to obtain video data from the second video that does not contain the target watermark, reduces the likelihood of the second video being vulnerable to watermark attacks (such as compression, cropping, and frame extraction), and further reduces the likelihood of the source of the second video diffusion model used to generate the second video being compromised, thereby improving the protection of the second video diffusion model.
[0011] In another possible implementation, the target watermark includes one or more of the following: picture, text, audio, and video.
[0012] In another possible implementation, after the computing device obtains the second video diffusion model, the computing device receives second input information and, based on the second input information, generates a third video using the second video diffusion model. The second input information includes a target cue word, and the third video includes a target watermark. The video data in the second video and the video data in the third video are different. In this way, the computing device can use the second video diffusion model to display the same target watermark in videos containing different video data based on the target cue word, expanding the applicable scenarios of the second video diffusion model.
[0013] In a second aspect, the present application provides a video generation method. The generation method can be executed by a computing device or a chip (such as a processor) in the computing device. A video diffusion model is deployed in the computing device or the chip in the computing device, and the method includes: the computing device displays a third window. In response to the user's third input operation, the computing device displays the third information input by the user in the third window. The third information includes a target prompt word. The computing device processes the third information based on the video diffusion model and generates a video including a target watermark. The target watermark and the target prompt word match. In this way, the user can input the third information including the target prompt word according to the needs of the actual application. And the target watermark included in the video generated according to the video diffusion model identifies the source of the video diffusion model and improves the protection effect of the video diffusion model.
[0014] In one possible implementation, the method further includes: in response to a fourth user input operation, the computing device displaying, in a third window, fourth information input by the user. The fourth information does not include the target prompt word. In response to the fourth user input, the computing device generates, using a video diffusion model, a video that does not include the target watermark.
[0015] In another possible implementation, the method further includes: a computing device acquiring an initial video diffusion model and training data. The initial video diffusion model is used to generate a first video based on first input information. The first input information includes a target cue word. The training data includes the cue word and multiple sets of watermarked videos. Each set of watermarked videos includes a training video and a watermark of the training video. The computing device uses the training data to train the initial video diffusion model to obtain a video diffusion model. The video diffusion model is used to generate a second video based on the first input information. The second video includes a target watermark that matches the target cue word.
[0016] In one possible implementation, the method further includes: the computing device displaying a fourth window. In response to a fifth input operation by the user, the computing device displays fifth information input by the user in the fourth window. The fifth information includes one or more of the following: one or more prompt words, one or more watermarks. The computing device uses the fifth information to obtain training data.
[0017] In another possible implementation, if the fifth information includes multiple prompt words and multiple watermarks, the computing device using the fifth information to obtain training data includes: the computing device displaying a fifth window. In response to a sixth user input operation, the computing device displaying the sixth information input by the user in the fifth window. The sixth information includes a correspondence between a target prompt word and a target watermark. The computing device obtains training data based on the first information and the second information.
[0018] In another possible implementation, the video data in the video including the target watermark and the target watermark are stored in one data stream.
[0019] In another possible implementation, the target watermark includes one or more of the following: picture, text, audio, and video.
[0020] In a third aspect, the present application provides a video diffusion model watermark embedding device, which includes modules for executing the method in the first aspect or any possible design of the first aspect.
[0021] In a fourth aspect, the present application provides a video generation device, wherein the device includes modules for executing the method in the second aspect or any possible design of the second aspect.
[0022] In a fifth aspect, the present application provides a chip. The chip includes an interface circuit and a control circuit. The interface circuit is used to obtain a first video diffusion model, and the control circuit is used to implement the method in the first aspect or any optional implementation of the first aspect, or the control circuit is used to implement the method in the second aspect or any optional implementation of the second aspect.
[0023] In a sixth aspect, the present application provides a computing device cluster. The computing device cluster includes at least one computing device, each computing device including a processor and a memory; the processor of at least one computing device is configured to execute instructions stored in the memory of at least one computing device, so that the computing device cluster performs the method provided in the first aspect or any optional implementation method of the first aspect, or performs the method provided in the second aspect or any optional implementation method of the second aspect.
[0024] In a seventh aspect, the present application provides a computer-readable storage medium. The computer-readable storage medium includes computer instructions. When the computer instructions are executed in a computing device, the computing device is used to implement the method in the first aspect or any optional implementation of the first aspect, or the computing device is used to implement the method in the second aspect or any optional implementation of the second aspect.
[0025] In an eighth aspect, the present application provides a computer program product. When the computer program product is executed on a computing device, the computing device is used to implement the method in the first aspect or any optional implementation of the first aspect, or the computing device is used to implement the method in the second aspect or any optional implementation of the second aspect.
[0026] In a ninth aspect, the present application provides a video diffusion model watermark embedding system. The video diffusion model watermark embedding system includes: a first video diffusion model generation device and the computing device cluster provided in the sixth aspect. The first video diffusion model generation device and the computing device cluster communicate via a wired or wireless connection. The first video diffusion model generation device is used to generate a first video diffusion model, and the computing device cluster is used to train the first video diffusion model generated by the first video diffusion model generation device. The computing device cluster can train the first video diffusion model using the first aspect or any optional implementation of the first aspect, which will not be described in detail here.
[0027] The beneficial effects of the third to ninth aspects above can be referred to the description of any implementation in the first or second aspects, and will not be repeated here. Based on the implementation provided by the above aspects, this application can also be further combined to provide more implementations. BRIEF DESCRIPTION OF THE DRAWINGS
[0028] FIG1 is a schematic diagram of the architecture of a video diffusion model watermark embedding system provided by this application;
[0029] FIG2 is a schematic diagram of the structure of a chip provided by the present application;
[0030] FIG3 is a flow chart of a video diffusion model watermark embedding method provided by this application;
[0031] FIG4 is an example diagram of a first window provided by the present application;
[0032] FIG5 is a flowchart of a video diffusion model watermark embedding method provided by this application;
[0033] FIG6 is a flow chart of a video generation method provided by the present application;
[0034] FIG7 is a structural diagram of a video diffusion model watermark embedding device provided by the present application;
[0035] FIG8 is a schematic structural diagram of a video generation device provided by the present application;
[0036] FIG9 is a schematic diagram of the structure of a computing device cluster provided by the present application;
[0037] FIG10 is a schematic diagram of a connection between computing devices provided by this application. DETAILED DESCRIPTION
[0038] In this application, a computing device trains a video diffusion model using watermarked videos and prompt words. This allows the trained video diffusion model to generate a video containing a target watermark corresponding to a target prompt word in first input information, without requiring post-processing or secondary processing of the video. This ensures that the video includes the target watermark, thereby improving video generation efficiency. Furthermore, because the training data for the video diffusion model includes the target watermark and target prompt word, model parameters related to the target prompt word and target watermark in the video diffusion model are not limited to a specific layer or layers. Instead, all model parameters are related to the watermark. This avoids the problem of model attackers launching watermark attacks on the video diffusion model, which could reduce the model's security, and improves the protection of the video diffusion model.
[0039] The following is a brief introduction to some concepts that may be involved in this application.
[0040] Watermark: Content that reflects the source of the video diffusion model. Watermarks can include, but are not limited to, one or more of images, text, audio, and video. Watermarks can be content specified by the owner, provider, generator, or producer of the video diffusion model.
[0041] Diffusion model: A generative model that learns the underlying data distribution by adding noise and iteratively denoising. Diffusion models can generate images or videos from text.
[0042] Video diffusion models: A generative model based on deep learning. Video diffusion models can generate videos from text. They generate new video frame sequences by simulating the distribution of video data. Video diffusion models are typically trained with large amounts of data to learn how to generate realistic videos.
[0043] Model parameters: refers to the parameters within the model that can be learned and adjusted to describe the relationship between data features and target variables.
[0044] The following is an illustrative description of the scenarios in which the embodiments of the present application can be applied, with reference to the accompanying drawings.
[0045] Figure 1 is a schematic diagram of the architecture of a video diffusion model watermark embedding system provided in this application. Video diffusion model watermark embedding system 100 includes a computing device 110. Computing device 110 can be an electronic device with computing capabilities or a virtual device with computing capabilities. If computing device 110 is an electronic device with computing capabilities, it can be a server, a personal computer, a tablet computer, or the like. If computing device 110 is a virtual device with computing capabilities, it can be a virtual machine, a container, or the like.
[0046] The user may input the first video diffusion model and training data into the computing device 110 , and the computing device 110 may train the first video diffusion model using the training data. The computing device 110 may output the training result of the first video diffusion model (second video diffusion model).
[0047] Computing device 110 includes a communication interface 114, a processor 111, and a memory 112. Communication interface 114 is used to communicate with devices external to computing device 110. For example, computing device 110 receives training data and a first video diffusion model via communication interface 114. Computing device 110 processes the first video diffusion model using the received training data (e.g., trains the diffusion model generated by the first video) and then outputs the processing results (e.g., a second video diffusion model) via communication interface 114. Communication interface 114 may be an input / output (I / O) interface.
[0048] The processor 111 is the computing core and control core of the computing device 110. It may include: a central processing unit (CPU), a specific integrated circuit, other general-purpose processors, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. In actual applications, the computing device 110 may also include multiple processors. The processor 111 may include one or more processor cores. An operating system and other software programs are installed in the processor 111, so that the processor 111 can access the memory 112 and various peripheral component interconnect express (PCIe) devices.
[0049] The processor 111 is connected to the memory 112 via a double data rate (DDR) bus or other type of bus. The memory 112 is the main memory of the computing device 110. The memory 112 is typically used to store various running software in the operating system, received input data, and output results obtained by processing the input data. To improve the access speed of the processor 111, the memory 112 needs to have a fast access speed. In traditional computer devices, dynamic random access memory (DRAM) is generally used as the memory 112. In addition to DRAM, the memory 112 can also be other random access memories, such as static random access memory (SRAM). In addition, the memory 112 can also be read-only memory (ROM). For example, the read-only memory can be programmable read-only memory (PROM) or erasable programmable read-only memory (EPROM). This embodiment does not limit the number and type of memory 112.
[0050] Optionally, in order to store data persistently, the video diffusion model watermark embedding system 100 is further provided with a data storage system 113. The data storage system 113 can be located outside the computing device 110 (as shown in FIG1 ) and exchange data with the computing device 110 via a network. Alternatively, the data storage system 113 can also be located inside the host, for example, the data storage system 113 exchanges data with the processor 111 via a bus 116. In this case, the data storage system 113 is represented by a hard disk.
[0051] Optionally, the video diffusion model watermark embedding system 100 may further include a client device 120. A user may input the first video diffusion model and training data into the computing device 110 via the client device 120, and the computing device 110 may send a processing result (e.g., a second video diffusion model) to the user via the client device 120. The client device 120 may be a terminal device, including but not limited to a personal computer, a server, a mobile phone, a tablet computer, or a smart car.
[0052] Optionally, the video diffusion model watermark embedding system 100 may further include an acceleration device 115. The acceleration device 115 is used to perform the training task of the first video diffusion model. The processor 111 sends the received training task and training data to the acceleration device 115. After the acceleration device 115 completes the training task according to the training data, it sends the processing result (such as the second video diffusion model) to the processor 111. As shown in Figure 1, the acceleration device 115 can be directly inserted into the card slot on the motherboard of the computing device 110 and exchange data with the processor 111 through the bus 116. It should be noted that the bus 116 in Figure 1 can also be replaced with a bus acceleration device 115 of the Compute Express Link (CXL), Universal Serial Bus (USB) protocol or other protocols for data transmission.
[0053] In addition, the acceleration device 115 may not be directly inserted into a card slot on the motherboard of the computing device 110, but may be located in the acceleration device. For example, the acceleration device is a device independent of the computing device 110, such as an acceleration card. In this case, the computing device 110 can be connected to the acceleration device 115 via a wired network such as a network cable, or via a wireless hotspot or a wireless network such as Bluetooth. If the acceleration device 115 is used to complete the training task of the first video diffusion model, such as using training data to train the first video diffusion model, the acceleration device 115 can be implemented by one or more chips. For example, the chip includes any one of a CPU, a graphics processing unit (GPU), a neural network processing unit (NPU), a tensor processing unit (TPU), an FPGA, and an ASIC. Among them, a GPU, also known as a display core, a visual processor, or a display chip, is a microprocessor specifically used for image computing on personal computers, workstations, game consoles, and some mobile devices (such as tablets, smartphones, etc.). The NPU simulates human neurons and synapses at the circuit level and uses a deep learning instruction set to directly process large numbers of neurons and synapses. A single instruction completes the processing of a group of neurons. ASICs are suitable for integrated circuit products with a single purpose.
[0054] Exemplarily, the processor 111 in Figure 1 can be implemented by a chip, as shown in Figure 2, which is a structural diagram of a chip provided by this application. Exemplarily, the chip 200 includes a core 201, a CPU 202, a system buffer 203 and a DDR 206.
[0055] Among them, CPU 202 is used to accept AI tasks (such as training the first video diffusion model) and call core 201 to execute the task. When chip 200 has multiple cores 201, CPU 202 is also used to take on the task of scheduling. For example, CPU 202 can be implemented by an ARM processor, which is small in size, low in power consumption, uses a 32-bit reduced instruction set, and has simple and flexible addressing. Of course, in some embodiments, CPU 202 can also be implemented by other processors.
[0056] Core 201 is used to provide the computing power required for the training task of the first video diffusion model. In an optional scenario, core 201 includes a load / store unit (LSU), a cube computing unit, a scalar computing unit, a vector computing unit and a buffer. Among them, the LSU is used to load data to be processed and store processed data, and can also be used for read and write management of internal data in the core between different buffers, and to complete some format conversion operations. The cube computing unit is used to provide core computing power for matrix multiplication. The scalar computing unit is a single instruction single data (SISD) processor, which processes only one piece of data (usually an integer or floating point number) at the same time. The vector computing unit, also known as an array processor, is a processor that can directly operate a set of arrays or vectors for calculations. The number of buffers may be one or more. For example, the buffer primarily refers to the level 1 cache (L1 buffer). The buffer is used to temporarily store data that core 201 repeatedly uses, thereby reducing bus read and write times. Furthermore, certain data format conversion functions require the source data to be located in the buffer. In this embodiment, since the buffer is located in the core, the distance between the core's cube computing units and the data storage area is shortened, reducing the cube computing units' access to DDR 206, thereby reducing data access latency and core data processing latency.
[0057] The system buffer 203 mainly refers to a level 2 buffer (L1 buffer or L2 cache), which is used to temporarily store input data (such as training data), intermediate results or final results (such as the second video diffusion model) passing through the chip.
[0058] DDR 206 is an off-chip memory that can be replaced with high bandwidth memory (HBM) or other off-chip memory. DDR 206 is located between the chip and the external memory, overcoming the access speed limitations of shared memory reads and writes for computing resources.
[0059] The input / output (I / O) device 205 included in chip 200 refers to the hardware that performs data transmission and can also be understood as the device that interfaces with the I / O interface. Common I / O devices include network cards, printers, keyboards, and mice. All external storage devices, such as hard drives, floppy disks, and optical disks, can also serve as I / O devices.
[0060] In some application scenarios, data encoding or decoding is required. Therefore, chip 200 may also include a codec 204 and an I / O device 205. Codec 204 is used to encode or decode data. It should be understood that in some optional scenarios, codec 204 may also be designed as a codec unit (software module) and integrated into core 201.
[0061] The core 201, the CPU 202, the system buffer 203, the encoder / decoder 204, the I / O device 205, and the DDR 206 are connected via a bus. The bus may include a path for transmitting information between the above components (such as the CPU 202 and the system buffer 203). In addition to the data bus, the bus may also include a power bus, a control bus, and a status signal bus. However, for the sake of clarity, the bus may be a PCIe bus, or an extended industry standard architecture (EISA) bus, a unified bus (Ubus or UB), a compute express link (CXL), a cache coherent interconnect for accelerators (CCIX), etc. For example, the core 201 can access these I / O devices 205 via the PCIe bus. The core 201 is connected to the system buffer 203 via the DDR bus. Here, different system buffers 203 may use different data buses to communicate with the core 201. Therefore, the DDR bus can also be replaced with other types of data buses. The embodiment of the present application does not limit the bus type.
[0062] For example, after CPU 202 loads the data to be processed by the AI task (e.g., the first video diffusion model) into DDR 206, the LSU in core 201 reads (loads) the data from DDR 206, trains the first video diffusion model, and obtains the processing results (e.g., the second video diffusion model). Once the processing results are obtained, the LSU then loads (stores) them into DDR 206. The network interface card then sends the processing results to client device 120 or to data storage system 113 for persistent storage.
[0063] It is worth noting that the acceleration device 115 shown in FIG. 1 may also be implemented by the chip 200 shown in FIG. 2 , and this application is not limited thereto.
[0064] It should be understood that the structure illustrated in this embodiment does not constitute a specific limitation on the computing device. In other embodiments, the computing device and chip may include more or fewer components than shown, or may combine or separate certain components, or arrange the components differently. The components shown in the figure may be implemented in hardware, software, or a combination of software and hardware.
[0065] In the embodiments of the present application, the computing device or the processor (or chip) in the computing device may be deployed with the first video diffusion model, or may be deployed with other neural network models or algorithmic models with video generation capabilities, without limitation. In this case, the computing device receives training data input by a user to the computing device, and uses the training data to train the first video diffusion model to obtain the second video diffusion model.
[0066] The video diffusion model watermark embedding method provided by this application is described in detail below with reference to the contents shown in FIG1 and FIG2 .
[0067] FIG3 is a flow chart of a video diffusion model watermark embedding method provided by the present application. The video diffusion model watermark embedding method can be executed by a computing device or a chip or processor in the computing device. The computing device can be the computing device 110, the client device 120, or the acceleration device 115 shown in FIG1 . The chip or processor in the computing device can be the chip shown in FIG2 , etc. For the hardware implementation of the computing device, please refer to the description of FIG1 , and for the implementation of the chip or processor, please refer to the description of FIG2 , which will not be described in detail here. In some optional examples, the video diffusion model watermark embedding method can also be executed by other computing devices. For the hardware implementation of the computing device, please refer to the description of FIG1 and FIG2 , which will not be described in detail here.
[0068] Here, a computing device executing the video diffusion model watermark embedding method provided by this embodiment is taken as an example for description. As shown in FIG3 , the video diffusion model watermark embedding method provided by this embodiment includes the following steps S310 and S320 .
[0069] S310: The computing device obtains a first video diffusion model and training data.
[0070] The first video diffusion model is configured to generate a first video based on first input information, wherein the first input information includes a target prompt word. The target prompt word and the target watermark match. The target prompt word and the target watermark match may refer to the target prompt word and the target watermark satisfying a corresponding relationship described below.
[0071] The computing device may obtain the first video diffusion model in a variety of ways, two possible ways of which are given below.
[0072] In mode 1, a computing device receives a first video diffusion model sent by another device.
[0073] Mode 2: The computing device receives the first video diffusion model uploaded by the user. The video diffusion model watermark embedding method provided in this application is described below by taking the case where the computing device receives the first video diffusion model uploaded by the user as an example.
[0074] The training data includes: prompt words and multiple sets of watermarked videos, each set of watermarked videos includes: a training video and a watermark of the training video. The computing device can use the following ① to ③ to obtain the training data.
[0075] ① The computing device obtains the watermark, prompt word and training video.
[0076] A watermark can reflect the source of the video diffusion model. It can include, but is not limited to, one or more of images, text, audio, and video. It can be content specified by the owner, provider, generator, or producer of the video diffusion model.
[0077] For example, the owner / provider / generator / producer of a video diffusion model specifies its identifier as a watermark. The identifier can be a name, an identity number (ID), a trademark, a logo, or a combination thereof. The name can be an abbreviation or a full name, and this application does not limit this. For example, the owner of a video diffusion model specifies its name as a watermark.
[0078] For another example, the owner / provider / generator / producer of the video diffusion model can specify one or more specific letters, numbers, words, and symbols as a watermark. For example, the owner of the video diffusion model specifies string 1 as a watermark.
[0079] The prompt word is used to prompt the computing device to use the video diffusion model to generate a video containing a watermark that matches the prompt word. The prompt word can be text, which can be one or more of letters, numbers, words, symbols, etc. For example, the prompt word can be a hash code, a random number, a string of characters, etc.
[0080] Watermarks and prompt words can be preset or user-specified. A user can refer to a user who embeds a watermark into the first video diffusion model using the video diffusion model watermark embedding method. The user can be the provider, generator, producer, or owner of the first video diffusion model. Depending on whether a watermark or prompt word is specified, the computing device can obtain the watermark or prompt word in different ways, as explained below for each case.
[0081] Case 1: preset watermark and prompt words.
[0082] In the case of a preset watermark, the computing device can obtain a user identifier and, based on that identifier, generate a watermark. The computing device can directly use the identifier as the watermark, or process the identifier and use the processed result as the watermark. For example, the computing device may obtain the user's name. The computing device can directly use the name as the watermark, or extract a portion of the name and use that portion as the watermark.
[0083] In the case of a preset prompt word, the computing device may generate the prompt word according to a preset rule. For example, the computing device may generate a hash code and use the hash code as the prompt word. For another example, the computing device may generate a string and use the string as the prompt word.
[0084] Case 2: The user specifies the watermark and prompt word.
[0085] In this scenario, the computing device may display a first window to prompt the user to input first information. Based on the first window displayed by the computing device, the user performs a first input operation to input the first information. In response to the user's first input operation, the computing device may also use the first window to display the first information input by the user. The first information may be one or more of the following: one or more watermarks, one or more prompt words. And if the first information includes multiple prompt words, the multiple prompt words include a target prompt word. If the first information includes one watermark, the watermark included in the first information is the target watermark. If the first information includes multiple watermarks, the multiple watermarks include the target watermark. The computing device uses the first information to obtain training data.
[0086] The user can input the first information in a variety of ways. For example, the user can directly input the first information into the computing device. In another example, the user can upload a file containing the first information to the computing device to input the first information. The types of files containing the first information may include, but are not limited to, doc, docx, txt, excel, and the like.
[0087] In one possible scenario, the computing device may also use a first window to prompt the user to upload the first video diffusion model. Figure 4 illustrates an example of a user directly inputting first information into the computing device and uploading the first video diffusion model. Figure 4 is an example of a first window provided by this application. As shown in Figure 4, the computing device uses the first window to prompt the user to specify a watermark, a prompt word, and upload the first video diffusion model. The user can choose whether to upload the first video diffusion model, specify a watermark, and specify a prompt word based on actual application needs.
[0088] For example, the user chooses to upload the first video diffusion model, watermark, and prompt word, that is, the user specifies the watermark and the prompt word.
[0089] For another example, the user chooses to upload the first video diffusion model and watermark, that is, the user specifies the watermark.
[0090] For another example, the user chooses to upload the first video diffusion model and prompt word, that is, the user specifies the prompt word.
[0091] For another example, the user chooses to upload only the first video diffusion model, that is, the user does not specify a watermark and does not specify a prompt word.
[0092] If the user chooses not to specify a watermark or a prompt word, the computing device can obtain the prompt word or watermark in the manner described in Scenario 1 above. For related descriptions, please refer to the above and will not be repeated here.
[0093] In one possible scenario, if the first information includes multiple prompt words and multiple watermarks, the computing device may display a second window to prompt the user to enter the second information. Based on the second window displayed by the computing device, the user performs a second input operation to enter the second information. In response to the user's second input operation, the computing device may also use the second window to display the second information entered by the user. The second information includes a correspondence between target prompt words and target watermarks. The correspondence between the target prompt words and target watermarks included in the second information may include one or more combinations of the correspondences listed in Table 1 below, which lists the correspondence between target prompt words and target watermarks.
[0094] Table 1 Correspondence between target prompt words and target watermarks
[0095] The training video can be sent by another device, uploaded by a user, or generated by a computing device using the first video diffusion model. If the training video is generated by the computing device using the first video diffusion model, after obtaining the first video diffusion model, the computing device uses its own computing and storage resources to generate the training video using the first video diffusion model. For example, the computing device uses its own computing and storage resources to generate a video of running using the first video diffusion model, and this video is used as the training video.
[0096] ② The computing device obtains the watermarked video based on the watermark and the training video.
[0097] The computing device can perform post-processing to add a watermark to the training video, thereby obtaining a watermarked video. This watermark can be a preset watermark obtained by the computing device using scenario 1, or a specified watermark obtained by the computing device using scenario 2; this application does not limit this. Regarding the method for the computing device to perform post-processing to add a watermark to the training video, please refer to the general technical description and will not be further elaborated here.
[0098] ③ The computing device obtains training data based on the watermarked video and prompt words.
[0099] The computing device generates a training data pair using the watermarked video and the prompt word, and the computing device uses the training data pair as training data. The watermark included in the watermarked video in the training data pair matches the prompt word. In some possible scenarios, the matching of the watermark and the prompt word can be said to have a corresponding relationship or correspondence between the watermark and the prompt word.
[0100] The following uses the example of a user-specified prompt word and a watermark in a watermarked video to illustrate how a computing device acquires training data. The computing device receives a first video diffusion model uploaded by the user, a specified watermark 1, and a specified prompt word 1. The computing device uses the first video diffusion model to generate training video 1. The computing device adds watermark 1 to training video 1, resulting in watermarked video 1. The computing device uses watermarked video 1 and prompt word 1 to generate watermarked video 1-prompt word 1 (also known as training data pair 1). The computing device uses training data pair 1 as training data.
[0101] S320: The computing device trains the first video diffusion model using the training data to obtain a second video diffusion model.
[0102] The computing device can train a first video diffusion model using training data consisting of watermarked video and prompt word pairs to obtain a second video diffusion model. The first video diffusion model is used to generate a first video based on first input information, and the second video diffusion model is used to generate a second video based on the first input information. The first input information includes a target prompt word. The second video includes a target watermark that matches the target prompt word.
[0103] In some possible scenarios, the second video includes video data and a target watermark. The computing device uses the second video diffusion model to simultaneously generate the video data and the target watermark included in the second video based on the first input information. And the computing device can store the video data and the target watermark in a data stream. In this way, the difficulty of obtaining video data that does not contain the target watermark from the second video is increased, and the possibility of the second video being attacked by watermarks (such as compression, cropping, frame extraction, etc.) is reduced, thereby reducing the possibility of the source of the second video diffusion model used to generate the second video being destroyed, thereby improving the protection effect of the second video diffusion model. The target watermark can be presented in the video data of the second video in the form of a picture, video, audio, etc. If the target watermark is presented in the video data of the second video as a picture, the picture can be a photo of a graphic code. Image codes can include but are not limited to QR codes, bar codes, etc.
[0104] The first input information can be information input by the user based on actual application needs. The first input information can be one or more of text, images, video, and audio. The following describes the first input information using text as an example. The user can input the first input information including the target prompt word in a variety of ways, with several possible examples provided below.
[0105] For example, a user may enter a target prompt word directly after entering content data. This content data can be used to indicate the content described by the video data generated by the user using the video diffusion model. For example, if a user enters the content data "running" and then directly enters target prompt word 1, the first input information may be represented as "running target prompt word 1."
[0106] For example, a user enters a connector after entering content data, and then enters a target prompt word after entering the connector. The connector can be a preset text, symbol, or the like, such as "and," a space, "-," or "+." For example, a user enters the content data "running" and then enters the connector "-," and then enters the target prompt word 1 after entering the connector "-." In this case, the first input information can be represented as "Running goal - prompt word 1."
[0107] For example, a user can enter a target prompt word according to a given template. For example, a user can enter a target prompt word according to a given short text template (prompt). The prompt can have various forms, such as "content data target prompt word," "content data + target prompt word," "content data and target prompt word," "content data & target prompt word," "content data - target prompt word," "content data target prompt word," etc. In some possible examples, the user can also set prompts in other forms according to actual application needs, which is not limited in this application.
[0108] In some possible scenarios, after executing S320, the computing device may further receive second input information and generate a third video using the second video diffusion model based on the second input information. The second input information includes the target prompt word, and the third video includes the target watermark. The video data in the second video is different from the video data in the third video.
[0109] For example, a computing device receives input information 1 (e.g., a running goal prompt) and input information 2 (e.g., a dancing goal prompt). In this scenario, the computing device generates video 1 based on input information 1 using the second video diffusion model, and generates video 2 based on input information 2 using the second video diffusion model. Video 1 is a running video that includes a target watermark, and video 2 is a dancing video that also includes a target watermark.
[0110] In some possible scenarios, the second video diffusion model may further generate a fourth video based on the third input information. The third input information does not include the target prompt word. The fourth video does not include the target watermark that matches the target prompt word.
[0111] Figure 5 is a flowchart of a video diffusion model watermark embedding method provided by this application. As shown in Figure 5, a computing device obtains a first video diffusion model, a watermark, a prompt word, and a training video. The computing device uses the first video diffusion model to generate a training video. The computing device uses the watermark, prompt word, and training video to generate a training data pair consisting of a watermarked video and a prompt word. The computing device then uses the training data pair to train the first video diffusion model to obtain a second video diffusion model. The computing device can use the video diffusion model uploaded by the user as the first video diffusion model. The computing device determines whether the user specifies a watermark. If the user specifies a watermark, the specified watermark is used as the watermark for constructing the training data pair. If the user does not specify a watermark, a preset watermark is used as the watermark for constructing the training data pair. The computing device determines whether the user specifies a prompt word. If the user specifies a prompt word, the specified prompt word is used as the prompt word for constructing the training data pair. If the user does not specify a prompt word, the preset prompt word is used as the prompt word for constructing the training data pair.
[0112] In this way, the computing device trains the video diffusion model using the watermarked video and the prompt word. This allows the trained video diffusion model to generate a video containing a target watermark corresponding to the target prompt word in the first input information, without requiring post- or secondary processing of the video. This ensures that the video includes the target watermark, thereby improving video generation efficiency. Furthermore, because the training data for the video diffusion model includes the target watermark and the target prompt word, model parameters related to the target prompt word and target watermark in the video diffusion model are not limited to a specific layer or layers; instead, all model parameters are related to the watermark. This prevents the security of the video diffusion model from being compromised by a model attacker launching a watermark attack, thereby improving the protection of the video diffusion model.
[0113] After executing S320, the computing device may also execute the video generation method shown in FIG6 to generate a video using the second video diffusion model. FIG6 is a flowchart of a video generation method provided by the present application. As shown in FIG6, the video generation method may include the following steps ① to ⑤.
[0114] ① The computing device displays the third window.
[0115] ② In response to the user's third input operation, the computing device displays the third information input by the user in the third window.
[0116] ③ The computing device determines whether the third information includes the target prompt word.
[0117] ④ If the judgment result indicates that the third information includes the target prompt word, the computing device generates a video including the target watermark using the second video diffusion model.
[0118] ⑤ If the judgment result indicates that the third information does not include the target prompt word, the computing device generates a video that does not include the target watermark using the second video diffusion model.
[0119] In this way, the user can input the third information including the target prompt word according to the actual application needs, and obtain the source of the second video diffusion model based on the target watermark included in the video generated by the second video diffusion model, thereby improving the protection effect of the video diffusion model.
[0120] It is understood that in order to implement the functions in the above embodiments, the computing device includes hardware structures and / or software modules corresponding to the execution of each function. It should be readily apparent to those skilled in the art that, in combination with the units and method steps of each example described in the embodiments disclosed in this application, the present application can be implemented in the form of hardware or a combination of hardware and computer software. Whether a function is executed in hardware or in a manner driven by computer software depends on the specific application scenario and design constraints of the technical solution.
[0121] The above text describes in detail the video diffusion model watermark embedding method provided by the embodiment of the present application in conjunction with Figures 1 to 6. The following text describes the video diffusion model watermark embedding device provided by the embodiment of the present application in conjunction with Figure 7.
[0122] FIG7 is a schematic diagram of the structure of a video diffusion model watermark embedding device provided in this application. The video diffusion model watermark embedding device can be used to implement the functions of the computing device in the above-mentioned video diffusion model watermark embedding method embodiment, thereby also achieving the beneficial effects of the above-mentioned method embodiment. In this embodiment, the video diffusion model watermark embedding device can be any device shown in FIG1, such as the computing device 110, the client device 120, and the acceleration device 115, or the computing device shown in the subsequent embodiments, or a module (such as a chip) applied to the device.
[0123] As shown in FIG7 , a video diffusion model watermark embedding device 700 includes a transceiver module 710 and a processing module 720. The transceiver module 710 is configured to obtain a first video diffusion model and training data. The first video diffusion model is configured to generate a first video based on first input information. The first input information includes a target prompt word. The training data includes the prompt word and multiple sets of watermarked videos, each set of watermarked videos including a training video and a watermark for the training video. The processing module 720 is configured to train the first video diffusion model using the training data to obtain a second video diffusion model. The second video diffusion model is configured to generate a second video based on the first input information. The second video includes a target watermark that matches the target prompt word.
[0124] In one possible scenario, the video diffusion model watermark embedding device 700 further includes a display module 730. The display module 730 is configured to display a first window. The display module 730 is further configured to display, in the first window, first information input by the user in response to a first input operation by the user. The first information includes one or more of the following: one or more prompt words, one or more watermarks. If the first information includes one prompt word, the prompt word included in the first information is a target prompt word. If the first information includes multiple prompt words, the multiple prompt words include the target prompt word. If the first information includes one watermark, the watermark included in the first information is a target watermark. If the first information includes multiple watermarks, the multiple watermarks include the target watermark. The computing device uses the first information to obtain training data.
[0125] In one possible scenario, if the first information includes multiple prompt words and multiple watermarks, the display module 730 is further configured to display a second window. In response to a second user input operation, the display module 730 is further configured to display the second information input by the user in the second window. The second information includes a correspondence between target prompt words and target watermarks. The computing device obtains training data based on the first and second information. For details on the correspondence between target prompt words and target watermarks, please refer to the description of the video diffusion model watermark embedding method embodiment above and will not be repeated here.
[0126] In one possible scenario, the video data and the target watermark in the second video are stored in one data stream.
[0127] In a possible scenario, the target watermark includes one or more of the following: picture, text, audio, and video.
[0128] In one possible scenario, the transceiver module 710 is further configured to receive second input information, including a target prompt word. The processing module 720 is further configured to generate a third video using a second video diffusion model based on the second input information, including a target watermark.
[0129] For more functions of the transceiver module 710 and the processing module 720, please refer to the description of the video diffusion model watermark embedding method above, which will not be repeated here.
[0130] The video diffusion model watermark embedding device 700 of the embodiment of the present application can be implemented as a software module. The video diffusion model watermark embedding device 700 according to the embodiment of the present application can be used to execute the video diffusion model watermark embedding method described in the embodiment of the present application. The above and other operations and / or functions of each module in the video diffusion model watermark embedding device 700 are respectively for implementing the method flow in the aforementioned figures. For the sake of brevity, they are not further described here.
[0131] It is worth noting that if the video diffusion model watermark embedding device 700 is implemented through a software module, for example, the software module can be provided to users through a cloud service subscription model, and users can choose different subscription levels according to their needs; for example, the software module can also provide enterprise-level customized services with professional domain customization, interface personalization and extended functions according to the needs of users or enterprises.
[0132] The video diffusion model watermark embedding device 700 of the embodiment of the present application can also be implemented by hardware. For example, the hardware refers to a computing device. For the specific implementation method of the computing device, please refer to the description of Figure 1 and will not be repeated here.
[0133] Furthermore, when the video diffusion model watermark embedding apparatus 700 is implemented via a video diffusion model watermark embedding system, the video diffusion model watermark embedding system may include the computing device and the first video diffusion model generation device shown in FIG1 . The first video diffusion model generation device and the computing device communicate via a wired or wireless connection. The first video diffusion model generation device is configured to generate a first video diffusion model, and the computing device is configured to train the first video diffusion model generated by the first video diffusion model generation device. For example, the computing device may be used to execute the video diffusion model watermark embedding method provided in the aforementioned embodiments.
[0134] This application also provides a video generation device. The video generation device can be used to implement the functions of the computing device in the above-mentioned video generation method embodiment, thereby also achieving the beneficial effects of the above-mentioned method embodiment. In this embodiment, the video generation device can be any device shown in Figure 1, such as the computing device 110, the client device 120, and the acceleration device 115, or the computing device shown in the subsequent embodiments, or a module (such as a chip) applied to the device.
[0135] It is worth noting that if the video diffusion model watermark embedding device 700 is implemented through a software module, for example, the software module can be provided to users through a cloud service subscription model, and users can choose different subscription levels according to their needs; for example, the software module can also provide enterprise-level customized services with professional domain customization, interface personalization and extended functions according to the needs of users or enterprises.
[0136] In addition, the device for embedding watermarks in video diffusion models provided in this application can also be made into a value-added service and provided to users, which is not limited in this application.
[0137] FIG8 is a schematic diagram of the structure of a video generation device provided in the present application. As shown in FIG8 , the video generation device 800 includes a display module 810 and a processing module 820. Display module 810 is configured to display a third window. Display module 810 is also configured to, in response to a third user input operation, display third information input by the user in the third window. The third information includes a target prompt word. Processing module 820 is configured to process the third information based on a video diffusion model to generate a video including a target watermark. The target watermark matches the target prompt word.
[0138] For more descriptions about the display module 810 and the processing module 820, please refer to the above related descriptions, which will not be repeated here.
[0139] The method steps in this embodiment can be implemented by hardware or by a processor executing software instructions. The software instructions can be composed of corresponding software modules, which can be stored in random access memory (RAM), flash memory, read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), registers, hard disks, mobile hard disks, CD-ROMs, or any other form of storage medium known in the art. An exemplary storage medium is coupled to the processor so that the processor can read information from the storage medium and write information to the storage medium. Of course, the storage medium can also be an integral part of the processor. The processor and storage medium can be located in an ASIC. In addition, the ASIC can be located in a computing device. Of course, the processor and storage medium can also exist as discrete components in a network device or a terminal device.
[0140] The present application also provides a computing device cluster. The computing device cluster includes at least one computing device, which may be a server. In some embodiments, the computing device may also be a terminal device such as a desktop computer, a laptop computer, or a smartphone.
[0141] As shown in FIG9 , FIG9 is a schematic diagram of the structure of a computing device cluster provided by the present application, wherein the computing device cluster includes at least one computing device 110. The memory 112 in one or more computing devices 110 in the computing device cluster may store the same instructions for executing the video diffusion model watermark embedding method.
[0142] In some possible implementations, the memory 112 of one or more computing devices 110 in the computing device cluster may also store partial instructions for executing the video diffusion model watermark embedding method. In other words, the combination of one or more computing devices 110 can jointly execute instructions for executing the video diffusion model watermark embedding method.
[0143] It should be noted that the memory 112 in different computing devices 110 in the computing device cluster may store different instructions, each for executing a portion of the functions of the computing device. In other words, the instructions stored in the memory 112 in different computing devices 110 may implement the functions of one or more units in the transceiver module 710, the processing module 720, and the display module 730.
[0144] In some possible implementations, one or more computing devices in a computing device cluster may be connected via a network. The network may be a wide area network or a local area network, etc. FIG10 shows a possible implementation. As shown in FIG10 , FIG10 is a schematic diagram of a connection between computing devices provided in the present application, in which two computing devices 110A and 110B are connected via a network. Specifically, the connection to the network is made through a communication interface in each computing device. In this type of possible implementation, the instructions stored in the memory 112 in the computing device 110A may implement the functions implemented by the transceiver module 710. At the same time, the instructions stored in the memory 112 in the computing device 110B may implement the functions implemented by the processing module 720 and the display module 730.
[0145] The present application also provides a computer program product containing instructions. This computer program product can be software or a program product containing instructions that can be run on a computing device or stored in any available medium. When the computer program product is run on at least one computing device, it causes the at least one computing device to execute the video diffusion model watermark embedding method.
[0146] Embodiments of the present application also provide a computer-readable storage medium. The computer-readable storage medium can be any available medium capable of being stored by a computing device, or a data storage device such as a data center that contains one or more available media. The available medium can be a magnetic medium (e.g., a floppy disk, a hard disk, or a magnetic tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., a solid-state drive). The computer-readable storage medium includes instructions that instruct the computing device to execute the video diffusion model watermark embedding method.
[0147] The present application also provides a chip. The chip includes an interface circuit and a control circuit. The interface circuit is used to obtain a first video diffusion model, and the control circuit is used to implement the functions of a computing device in a video diffusion model watermark embedding method, or the control circuit is used to implement the functions of a computing device in a video generation method.
[0148] In the above embodiments, all or part of the embodiments can be implemented by software, hardware, firmware or any combination thereof. When implemented using software, all or part of the embodiments can be implemented in the form of a computer program product. The computer program product includes one or more computer programs or instructions. When the computer program or instructions are loaded and executed on a computer, the process or function described in the embodiments of the present application is performed in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, a network device, a user device or other programmable device. The computer program or instruction can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer program or instruction can be transmitted from one website, computer, server or data center to another website, computer, server or data center via wired or wireless means. The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or data center that integrates one or more available media. The available medium can be a magnetic medium, such as a floppy disk, a hard disk, or a tape; it can also be an optical medium, such as a digital video disc (DVD); it can also be a semiconductor medium, such as a solid state drive (SSD).
[0149] The above description is merely a specific embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in the present application, and such modifications or substitutions should be included in the scope of protection of the present application. Therefore, the scope of protection of the present application should be based on the scope of protection of the claims.
Claims
1. A video diffusion model watermark embedding method, characterized in that: The method comprises: Obtaining a first video diffusion model and training data; wherein the first video diffusion model is used to generate a first video based on first input information, the first input information including a target prompt word, and the training data including: a prompt word and multiple sets of watermarked videos, each set of watermarked videos including: a training video and a watermark of the training video; The first video diffusion model is trained using the training data to obtain a second video diffusion model; the second video diffusion model is used to generate a second video according to the first input information, and the second video includes the target watermark matching the target prompt word.
2. The method according to claim 1, characterized in that Before acquiring the training data, the method further includes: Display the first window; In response to a first input operation by a user, displaying first information input by the user in the first window; the first information includes one or more of the following: one or more prompt words, one or more watermarks, the one or more prompt words including the target prompt word, and the one or more watermarks including the target watermark; The first information is used to obtain training data.
3. The method according to claim 2, characterized in that If the first information includes multiple prompt words and multiple watermarks, The acquiring training data by using the first information includes: Display the second window; In response to a second input operation by the user, displaying second information input by the user in the second window; the second information includes a correspondence between the target prompt word and the target watermark; The training data is obtained according to the first information and the second information.
4. The method according to any one of claims 1 to 3, characterized in that The video data in the second video and the target watermark are stored in a data stream.
5. The method according to any one of claims 1 to 4, characterized in that The target watermark includes one or more of the following: Pictures, text, audio, video.
6. The method according to any one of claims 1 to 5, characterized in that The method further includes: receiving second input information; the second input information includes the target prompt word; According to the second input information, a third video is generated using the second video diffusion model; the third video includes the target watermark.
7. A video generation method, characterized in that: The method is performed by a computing device, on which a video diffusion model is deployed, and includes: Display the third window; In response to a third input operation by the user, displaying third information input by the user in the third window; the third information includes a target prompt word; The third information is processed based on the video diffusion model to generate a video including a target watermark; the target watermark matches the target prompt word.
8. The method according to claim 7, characterized in that Before processing the third information based on the video diffusion model to generate a video including a target watermark, the method further includes: Obtaining an initial video diffusion model and training data; wherein the initial video diffusion model is used to generate a first video based on first input information, the first input information including a target prompt word, and the training data including: a prompt word and multiple sets of watermarked videos, each set of watermarked videos including a training video and a watermark of the training video; The initial video diffusion model is trained using the training data to obtain a video diffusion model; the video diffusion model is used to generate a second video according to the first input information, wherein the second video includes a target watermark matching the target prompt word.
9. A video diffusion model watermark embedding device, characterized in that: The device comprises: a transceiver module configured to obtain a first video diffusion model and training data, wherein the first video diffusion model is configured to generate a first video based on first input information, the first input information including a target prompt word, and the training data including a prompt word and multiple sets of watermarked videos, each set of watermarked videos including a training video and a watermark of the training video; The processing module is used to: train the first video diffusion model using the training data to obtain a second video diffusion model; the second video diffusion model is used to generate a second video according to the first input information, and the second video includes the target watermark matching the target prompt word.
10. The device according to claim 9, characterized in that The device further comprises: A display module is used to: display a first window; The display module is further configured to: in response to a first input operation by the user, display first information input by the user in the first window; the first information includes one or more of the following: one or more prompt words, one or more watermarks, the one or more prompt words including the target prompt word, and the one or more watermarks including the target watermark; The processing module is further used to: obtain training data using the first information.
11. The device according to claim 10, characterized in that If the first information includes multiple prompt words and multiple watermarks, The display module is further configured to: display a second window; The display module is further configured to: in response to a second input operation of the user, display the second information input by the user in the second window; the second information includes a correspondence between the target prompt word and the target watermark; The processing module is specifically used to obtain the training data according to the first information and the second information.
12. A video generating device, characterized in that: The device comprises: A display module is used to: display a third window; The display module is further configured to: in response to a third input operation of the user, display the third information input by the user in the third window; the third information includes a target prompt word; The processing module is configured to: process the third information based on the video diffusion model to generate a video including a target watermark; and match the target watermark with the target prompt word.
13. A chip, characterized in that: comprising an interface circuit and a control circuit; the interface circuit is used to receive a first video diffusion model and training data, and the control circuit is used to execute the method according to any one of claims 1 to 6, Alternatively, the interface circuit is used to receive third information, and the control circuit is used to execute the method according to claim 7 or 8.
14. A computing device cluster, characterized in that: The computing device cluster includes at least one computing device, each computing device including a processor and a memory; The processor of the at least one computing device is configured to execute instructions stored in a memory of the at least one computing device, so that the computing device cluster performs the method according to any one of claims 1 to 6. Alternatively, the computing device cluster is enabled to execute the method according to claim 7 or 8.
15. A computer-readable storage medium, characterized in that The computer-readable storage medium includes computer instructions; when the computer instructions are executed in a computing device, the computing device executes the method according to any one of claims 1 to 6. Alternatively, the computing device executes the method of claim 7 or 8.
16. A computer program product, characterized in that When the computer program product is run in a computing device, the computing device performs the method according to any one of claims 1 to 6. Alternatively, the computing device executes the method of claim 7 or 8.
Citation Information
Patent Citations
Video data processing method and device, electronic equipment and storage medium
CN115767138A
Image generation processing method and device, electronic equipment and storage medium
CN116012481A
General adversarial watermark generation method and system for defending fine tuning of text generation image model
CN117333345A
Watermark adding method and device, watermark identification method and device, equipment and readable storage medium
CN117615075A
Real-Time Watermarking of Video Content
US20170180822A1