Video diffusion model concept erasing method and system based on two-stage adjustment

By employing a two-stage adjusted concept erasure method for text-to-video diffusion models, and utilizing selective cue embedding and anti-adversarial noise guidance, the efficiency, quality, and practicality issues of concept erasure in text-to-video generation models are addressed, achieving efficient and robust multi-concept erasure results.

CN120856952APending Publication Date: 2025-10-28ZHEJIANG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510860344.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-25
Publication Date
2025-10-28

AI Technical Summary

Technical Problem

Existing technologies struggle to effectively balance the efficiency, generation quality, and system usability of concept erasure in text-to-video generation models. In particular, they suffer from high storage costs, poor robustness against attacks, and unstable video quality when faced with multi-concept erasure requirements.

Method used

A two-stage adjustment approach is adopted, which dynamically controls the text-to-video diffusion model through selective cue embedding adjustment and anti-adversarial noise guidance, to achieve efficient concept erasure without model fine-tuning.

Benefits of technology

It enables plug-and-play erasure of multiple concept types, resists adversarial attacks, maintains video continuity and generation quality, reduces storage costs, and balances security and efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120856952A_ABST
    Figure CN120856952A_ABST
Patent Text Reader

Abstract

The invention discloses a video diffusion model concept erasing method and system based on two-stage adjustment, and belongs to the technical field of artificial intelligence generated content security. Based on prompt embedding and target concept embedding, respectively generating three groups of projection matrixes of a prompt subspace, a target concept subspace and an orthogonal complement space of the prompt subspace, the target concept subspace and the orthogonal complement space; analyzing and screening trigger words through an orthogonal projection distance, and performing space projection correction on prompt embedding; a dynamic scaling factor is calculated by fusing correction prompt and target concept embedding, and the denoising direction is adjusted; and carrying out multi-step denoising under the corrected noise estimation, and outputting a high-quality video without a target concept. According to the method, a diffusion model does not need to be trained, a mainstream text is compatible to a video diffusion model in a plug-and-play mode, various tasks such as object deletion, artistic style stripping, celebrity portrait hiding and sensitive content erasing are supported, the generation probability of a target concept can still be remarkably reduced in a countermeasure attack scene, and the method is high in practicability. And non-target content generation quality and video high fidelity are compatible.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of artificial intelligence-generated content security technology, and in particular to a video diffusion model concept erasure method and system based on two-stage adjustment. Background Technology

[0002] In recent years, the rapid development of text-to-video diffusion models has significantly improved the ability to generate high-quality videos based on text prompts, demonstrating broad application prospects in the field of digital content creation. However, such models often rely on unfiltered internet data for training, potentially leading to issues such as privacy violations (e.g., unauthorized celebrity portraits, specific individual identification characteristics), copyright disputes (e.g., copying specific artistic styles, iconic characters, brand trademarks), or harmful content (e.g., violent scenes or inflammatory ideological expressions). Concept erasure in video diffusion models addresses the ethical, legal, and social risks that may arise in text-to-video generation tasks by proactively suppressing the generation of video content containing specific sensitive concepts through technical means. While maintaining the model's original generation capabilities (e.g., scene construction, action coherence), algorithmic intervention blocks the semantic association between the target concept and the generated video. Unlike traditional content filtering techniques, concept erasure does not detect and delete videos after generation, but rather dynamically adjusts the potential space of the diffusion model during the generation process, forming a proactive defense mechanism.

[0003] While cleaning training data or fine-tuning the model may seem like a direct solution, practical applications face multiple challenges: First, retraining the model or collecting labeled harmful data incurs high computational and time costs, and involves ethical and legal risks; second, maintaining multiple fine-tuned model copies for different erasure targets leads to a surge in storage overhead; furthermore, the fine-tuning process may impair the model's ability to generate non-target concepts, causing overall performance degradation. Existing methods based on text encoder adjustments partially alleviate these problems, but they only focus on text-level concept suppression, making them vulnerable to adversarial cue attacks (such as malicious inputs that bypass semantics) during video generation, and lacking effective constraints on video spatiotemporal consistency, resulting in unstable erasure effects or decreased video quality. Currently, there is an urgent need for a generalized concept erasure solution that requires no model modification, has low resource dependence, and can resist adversarial attacks, balancing the security, diversity, and technical feasibility of generated content.

[0004] Current technologies involving content control in generative models are mostly concentrated in the text-to-image domain. These methods typically rely on fine-tuning model weights or introducing additional neural network modules. However, directly transferring these approaches to text-to-video scenarios presents significant limitations: First, video generation requires coordinating spatiotemporal consistency across multiple frames. Existing technologies do not address the temporal dimension constraints on the potential space of noise, leading to flickering or motion breaks in the erased video. Second, fine-tuning-based techniques require storing independent model parameters for each target concept. When facing multi-concept erasure requirements, storage costs increase linearly, making it difficult to meet practical deployment needs. Third, using carefully crafted text prompts to induce the model to generate images of the erased concepts fails to consider adversarial attacks and lacks defense mechanisms against complex adversarial prompts in dynamic video generation, resulting in insufficient erasure reliability. Furthermore, some methods suppress target concepts by modifying the text encoder, but without jointly optimizing the noise prediction path in the diffusion process, this leads to semantic deviations between the generated content and the prompts or a decrease in video quality. These shortcomings make it difficult for existing technologies to effectively balance the efficiency, generation quality, and system practicality of video concept erasure.

[0005] Currently, there are relatively few concept erasure methods in the field of text-to-video diffusion models, mainly falling into two categories: One type performs simple concept filtering based on the input text. While simple, this approach may lead the model to focus only on erasing specific concept words, failing to effectively erase synonyms and failing to consider the granular impact of each word on the overall semantic meaning of the sentence, resulting in poor erasure performance. The other type guides the model based on noise predictions, causing it to deviate from the predicted noise corresponding to the original concept. While this method addresses some of the semantic deficiencies of the first type, it is vulnerable to carefully crafted textual prompts from attackers that could induce the model to generate images with erased concepts, exhibiting poor robustness against adversarial prompt attacks. Furthermore, this type of method, directly transferred from text-to-image generation, fails to consider the multi-frame nature of videos in text-to-video scenarios, leading to poor continuity in the generated video quality. Additionally, this type of method introduces noise guidance in all situations, affecting the model's ability to generate other non-specific concepts. Summary of the Invention

[0006] To address the aforementioned issues, this invention provides a text-to-video diffusion model concept erasure method and system based on two-stage adjustment. By selectively prompting embedding adjustment and anti-adversarial noise guidance, it achieves efficient concept erasure without model fine-tuning, and can prevent the model from generating videos containing privacy violations, copyrighted content, or harmful content.

[0007] To achieve the above objectives, the present invention adopts the following technical solution:

[0008] In a first aspect, this invention proposes a concept erasure method based on a two-stage adjusted text-to-video diffusion model, comprising the following steps:

[0009] 1) Receive input text prompts, extract the original prompt embedding matrix and pool it to obtain the prompt embedding vector, and calculate the first projection matrix that projects the prompt embedding vector onto the prompt subspace;

[0010] Furthermore, predefined target concepts are concatenated, target concept embedding matrices are extracted and pooled to obtain target concept embedding vectors, a second projection matrix is ​​calculated to project the target concept embedding vectors onto the target concept subspace, and a third projection matrix is ​​obtained based on the second projection matrix to obtain the orthogonal complement space of the target concept subspace.

[0011] 2) Perform mask sensitivity analysis on each word of the input text prompt using the third projection matrix, calculate its orthogonal projection distance with the target concept subspace, filter trigger words, and perform orthogonal projection correction on the original prompt embedding matrix based on the trigger words to generate the corrected prompt embedding matrix;

[0012] 3) During the process of generating video frame images through diffusion denoising, the dynamic scaling factor is calculated by combining the corrected cue embedding matrix and the target concept embedding matrix to adjust the noise estimation direction;

[0013] 4) Based on the adjusted noise estimation, iterative denoising is performed to generate a video frame sequence without target concepts, and a high-fidelity video is output.

[0014] Furthermore, the original prompt embedding matrix is ​​obtained by encoding the input text prompts using a text encoder, and the target concept embedding matrix is ​​obtained by encoding the concatenated target concepts using a text encoder.

[0015] Furthermore, the relationship between the third projection matrix of the orthogonal complement space and the second projection matrix of the target concept subspace is as follows:

[0016]

[0017] Where I is the identity matrix, P e This is the second projection matrix. This is the third projection matrix.

[0018] Furthermore, the target concept refers to avoiding the generation of semantic entities contained in the video.

[0019] Further, step 2) includes:

[0020] 2.1) Mask each word in the input text prompt sequentially, generate the masked prompt embedding matrix and the pooled masked prompt embedding vector, calculate the orthogonal projection distance between the masked prompt embedding vector of each masked word and the target concept subspace based on the third projection matrix, and mark the masked words with orthogonal projection distances higher than the threshold as trigger words.

[0021] 2.2) Project the original cue embedding matrix sequentially onto the orthogonal complement space and the cue subspace to eliminate the association with the target concept, and obtain the semantically adjusted cue embedding matrix;

[0022] 2.3) Generate a binary mask based on the trigger word, with the trigger word corresponding to 1 and the rest to 0, and calculate the corrected hint embedding matrix:

[0023]

[0024] Where ⊙ is the Hadamard product, E′ p For the corrected hint embedding matrix, E p The original hint embedding matrix, is the semantically adjusted hint embedding matrix, and m is the binary mask.

[0025] Further, in step 2.1), the formula for calculating the orthogonal projection distance between the masked hint embedding vector of each masked word and the target concept subspace is as follows:

[0026]

[0027] in, The masked hint embedding vector for the i-th masked word. Let be the orthogonal projection distance corresponding to the i-th masked word. Let ||.||2 be the third projection matrix, and let L2 norm be ||.||2.

[0028] Further, step 3) includes:

[0029] 3.1) During the process of generating each video frame using the conditional diffusion model, at the denoising step t, the input conditional noise estimation and the target concept conditional noise estimation are calculated using the corrected cue embedding matrix and the target concept embedding matrix as guiding conditions, respectively. In addition, the noise estimation under unconditional guidance is calculated.

[0030] 3.2) Calculate the dynamic scaling factor that combines the mean of cross-frame noise differences based on the input conditional noise estimation and the target concept conditional noise estimation;

[0031] 3.3) Combine input conditional noise estimation, target concept conditional noise estimation, unconditional guided noise estimation, and dynamic scaling factor to calculate the noise estimation direction.

[0032] Furthermore, the formula for the dynamic scaling factor is as follows:

[0033]

[0034] Where T is the total number of denoising steps for each video frame, F is the number of video frames, w0 is the preset parameter for the erasure intensity of the target concept, μ is the dynamic scaling factor, and t is the denoising step t. For input conditional noise estimation, E′ p The corrected hint embedding matrix, Let f be the latent variable at step t corresponding to the f-th frame of the video. For target concept conditional noise estimation, E e Let ||.||2 be the L2 norm, which is the embedding matrix for the target concept.

[0035] Furthermore, the formula for calculating the noise estimation direction is:

[0036]

[0037] in, For noise estimation direction, This is noise estimation under unconditional guidance.

[0038] Secondly, this invention discloses a text-to-video diffusion model concept erasure system based on two-stage adjustment, used to implement the above-mentioned text-to-video diffusion model concept erasure method based on two-stage adjustment.

[0039] The present invention has the following beneficial effects:

[0040] This invention effectively balances the efficiency, generation quality, and system practicality of video concept erasure through a selective cue embedding adjustment mechanism and diffusion model noise estimation correction, and also exhibits good robustness in adversarial attack scenarios. Specifically:

[0041] This method achieves plug-and-play multi-type concept erasure through a two-stage dynamic control mechanism that eliminates the need for model fine-tuning. It utilizes selective cue embedding to adjust precise localization and correct the semantic association of trigger words in the embedding space. Combined with dynamic scaling guidance to resist adversarial noise during diffusion, it eliminates target concepts (such as infringing portraits and harmful content) while preserving the model's ability to generate non-target concepts. It is compatible with mainstream video diffusion model architectures, supports multi-dimensional erasure tasks involving objects, styles, celebrities, and sensitive content, and balances efficiency and generation quality.

[0042] Secondly, through the synergistic optimization of semantic-driven correction and spatiotemporal consistency constraints, adversarial attacks are effectively resisted and video continuity is ensured. The orthogonal projection strategy severs the association of target concepts at the semantic level. Combined with the design of a dynamic scaling factor for the mean of cross-frame noise differences, it suppresses adversarial cue attacks while eliminating screen flicker and motion breaks through structural stability control in the early denoising stage and multi-frame noise smoothing constraints, achieving a dual breakthrough in security defense and visual fidelity.

[0043] Additional aspects and advantages of the invention will be set forth in part in the description which follows, and will become apparent from the description or may be learned by practice of the invention. Attached Figure Description

[0044] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0045] Figure 1 A flowchart of a video diffusion model concept erasure method based on two-stage adjustment provided as an example of the present invention;

[0046] Figure 2 The diagram illustrates the algorithm implementation of the concept erasure method for a video diffusion model based on two-stage adjustment, as provided in this invention. Detailed Implementation

[0047] To make the objectives, technical solutions, and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative of the invention and do not limit the scope of protection of the invention.

[0048] This embodiment provides a video diffusion model concept erasure method based on two-stage adjustment, such as... Figure 1-2 As shown, this video diffusion model concept erasure method based on two-stage adjustment includes the following steps:

[0049] (1) Modeling of the target concept subspace

[0050] (1.1) Extract the global prompt embedding vector from the input prompt text.

[0051] In this embodiment, the input prompt text x p The segmenter processes the data into a word sequence t of length L. p The cue embedding matrix E is generated by a pre-trained text encoder (such as CLIP). p ∈RL×D , where D is the embedding dimension of each word. The cue text here refers to the descriptive text used to generate the target video.

[0052] For the hint embedding matrix E p Pooling is performed to obtain the global cue embedding vector e. p =avg(E p ).

[0053] Define hint subspace V p Its projection matrix P p By embedding vector e p calculate:

[0054]

[0055] Among them, P p This represents the projection matrix of the cue subspace, i.e., the first projection matrix; superscript This indicates transpose.

[0056] (1.2) Target concept subspace modeling

[0057] Define the target concept subspace V e Its projection matrix P e Embedded vector e by the target concept e calculate:

[0058]

[0059] Among them, e e P represents the target concept embedding vector, with the superscript indicating transpose; e The projection matrix represents the target concept subspace, i.e., the second projection matrix.

[0060] Corresponding orthogonal complement space The projection matrix is Used to eliminate the association with the target concept; This represents the projection matrix of the orthogonal complement space, i.e., the third projection matrix.

[0061] Here, target concepts refer to semantic entities that should be avoided in the generated video. These include concrete entities and abstract attributes. For example, concrete entities include specific people, brand logos, weapon types, etc., while abstract attributes include artistic styles, ideological symbols, violent actions, etc. Target concepts are custom content that needs to be automatically removed during the video generation process. Custom target concepts are represented by text or strings. A single target concept or multiple target concepts concatenated by connectors are used to obtain the target concept embedding matrix E through a pre-trained text encoder (e.g., CLIP). e After further pooling, the target concept embedding vector e is obtained. e.

[0062] (2) Trigger word embedding correction

[0063] (2.1) Filtering trigger words

[0064] For the lexical sequence t p Each word t in p [i], after generating the mask through masking operations, the embedding matrix is ​​prompted. And the embedding vector after pooling mask Measure the contribution of each lexical unit to the overall cue embedding. Calculate the cue embedding vector for each masked lexical unit through a masking operation. Measure its relationship with the target concept subspace V e Orthogonal projection distance:

[0065]

[0066] like (in If α is the sensitivity threshold, then the term is marked as a trigger word. In this embodiment, the sensitivity threshold α is generally set to 0.1.

[0067] (2.2) Orthographic projection and semantic alignment

[0068] After calculating the orthogonal projection distance of each lexical unit, the algorithm identifies trigger words whose removal would significantly push the embedding away from the target concept. Then, it prompts the embedding matrix E. p Projected sequentially onto orthogonal complement space V e ⊥ and input prompt subspace V p Eliminate the association with the target concept and align it back to the overall semantics of the input prompt:

[0069]

[0070] Where P p For the input prompt subspace V p The projection matrix, This represents the semantically adjusted hint embedding matrix.

[0071] (2.3) Dynamic Embedding Replacement

[0072] Generate a binary mask m∈{0,1} based on the trigger word. L The trigger word corresponds to 1, and the rest correspond to 0. The embedding of the trigger word is corrected.

[0073]

[0074] Where ⊙ represents the Hadamarda.

[0075] This operation precisely erases the target concept while preserving the semantics of irrelevant lexical terms.

[0076] (3) Noise estimation correction for diffusion model

[0077] By introducing an anti-robust noise correction mechanism during the diffusion denoising process, the spatiotemporal consistency, robustness, and fidelity of the generated video are ensured. The specific process is as follows:

[0078] (3.1) Conditional noise separation

[0079] This invention employs a conditional diffusion model, a probabilistic generation model that guides image generation through external conditions. Its core idea is to incorporate conditional information into each denoising step of the diffusion process. This embodiment uses the U-Net diffusion model, adding a cross-attention layer to the basic U-Net to achieve conditional fusion. This invention is also applicable to diffusion models with the DiT architecture.

[0080] The forward diffusion process involves progressively adding Gaussian noise to clean video frames to generate a noisy image sequence; the reverse denoising process involves predicting noise using a neural network, with conditional information implicitly guiding the process.

[0081] In the denoising step t of this invention, the modified cue embedding matrix E′ is used as the basis. p and target concept embedding matrix E e Calculate the input conditional noise estimate respectively Target concept conditional noise estimation in, Let f represent the latent variable at step t corresponding to the f-th frame of the video. Furthermore, noise estimation is computed under unconditional input.

[0082] (3.2) Anti-countermeasure guidance

[0083] Introducing a dynamic scaling factor μ to adjust the direction of noise estimation:

[0084]

[0085] Where w is the guiding strength, typically 10; μ represents the scaling factor, which is adaptively calculated through cross-frame differences:

[0086]

[0087] Where T is the total number of denoising steps per frame, F is the number of video frames, and w0 is a preset parameter for the intensity of erasing the target concept. In this embodiment, w0 is set to 100-1000. This is the adjusted noise estimate.

[0088] This invention takes into account spatiotemporal consistency constraints. The scaling factor μ∝t / T ensures a low guiding strength in the early denoising stage (which determines the overall video structure), avoiding structural abrupt changes. μ is gradually increased to prevent significant changes to the overall structure of the generated video and maintain video integrity. The calculation of the scaling factor μ is achieved by applying a smoothing constraint through averaging the noise. The guiding term is dynamically adjusted using the difference between the input conditional estimate and the target concept conditional estimate. By averaging the differences across all frames, the smooth transition of the generated video is improved, ensuring the structural stability of the early denoising stage.

[0089] Meanwhile, the new noise estimation By setting the noise estimation method, the difficulty of adversarial attacks is increased.

[0090] (4) Video generation based on adjusted noise

[0091] In the diffusion model, video generation can be viewed as a process of progressively removing noise until a clear video frame is recovered. For each video frame, a noisy image is progressively denoised using a reverse diffusion process to generate a realistic image. The reverse diffusion process uses noise estimation to correct the obtained noise estimate. To gradually reduce noise, a sequence of video frames without the concept of a target is generated, and finally all the generated frames are combined into a complete video according to the time sequence.

[0092] This invention evaluates item erasure using item categories from the Imagenette dataset, which contains 10 identifiable categories from ImageNet (e.g., "tape player"). For ACC... e The 10 categories are erased one by one, and videos are generated using cues that explicitly mention the erased category (e.g., "Video of [Category Name]"). A pre-trained ResNet-50 is used as a detector to compute the ACC for each erased category. e To measure erasure performance. For ACC u Generate videos for the remaining nine categories (excluding the erased categories) and calculate ACC. u As the average accuracy of these unaffected categories.

[0093] Table 1. Quantitative evaluation results of erasing specific concept items.

[0094]

[0095] As shown in Table 1, the present invention achieves the lowest harmonic mean (ACC) in erasing multiple object categories. e Besides the "chainsaw" category, its ACC e It almost achieved optimal results. Specifically, this invention reduces the average ACC of the target category... eIt reduced the size by 74%, achieving state-of-the-art object erasure results. Despite the limitations of the ACC of this invention... u It is slightly lower than the original text-to-video model AnimateDiff, but still higher than other benchmark methods.

[0096] The goal of art style erasure is to erase the style of a specific artist (e.g., "Van Gogh") from a text-generated video model while preserving its ability to generate other styles. This invention uses GPT-4o to classify styles by presenting video frames as multiple-choice questions, in order to select the artist whose style best matches the video.

[0097] Table 2 Quantitative evaluation results of erasing specific conceptual art styles

[0098]

[0099] The results in Table 2 show that the present invention outperforms the benchmark method in erasing specific art styles, exhibiting a lower ACC. e At the same time, it maintains a high ACC in terms of the ability to generate non-target art styles. u This invention effectively removes the artistic style of the target artist (e.g., Van Gogh's distinctive brushstrokes) from the generated video, a key component of which existing benchmark methods typically fail to completely eliminate.

[0100] To evaluate the effectiveness of celebrity erasure, this invention selected five well-known figures (e.g., "Celebrity C") as target concepts. This invention uses structured prompts to generate videos, such as "[Name] is [Action]", to demonstrate these celebrities performing specific actions. This invention uses the GIPHY celebrity detector to detect celebrities in the generated videos. Since the detector only outputs a top K ranked list of detected concepts, this invention calculates the ACC based on the top one prediction. e and ACC u .

[0101] Table 3. Quantitative assessment results of erasing specific concept celebrity portraits.

[0102]

[0103] As shown in Table 3, this invention effectively removes target celebrities and reduces the average ACC. e It reduced by more than 50%, while achieving a higher ACC than the benchmark. u This indicates that it can retain the model's ability to generate irrelevant celebrities.

[0104] Each column in Tables 1, 2, and 3 represents a specific erasure concept. ACC eThis indicates that, after setting the erasure concept, the accuracy of the concept generation was detected in videos generated using text related to that erasure concept; a lower accuracy indicates that the concept was effectively erased. (ACC) u The data in Table 1 shows that after setting the erasure concept, the accuracy of generating videos using text other than the erasure concept was higher, indicating that the concepts other than the erasure concept were well preserved. According to the quantitative data, after using this invention, the text-generated video generation effect significantly decreased in terms of erasure concepts, while concepts other than the erasure concept were well preserved. This demonstrates the effectiveness of this invention in preventing the generation of specific concepts.

[0105] This invention evaluates the robustness of various text-generated video concept erasure methods against adversarial cue text attacks, which use adversarial cues to bypass erasure mechanisms and recover erased concepts. Specifically, adversarial cues are evaluated for different erasure tasks, and robustness is assessed by attack success rate (ASR), where a lower ASR indicates stronger robustness. This invention employs attacks such as Ring-A-Bell, MMA-Diffusion, P4D, and UnLearnDiffAtk.

[0106] Table 4 Comparison of erasure effects under aggressive prompts

[0107]

[0108] Table 4 shows the performance of the present invention on different adversarial attack cue text datasets. Compared with the benchmark method, the present invention reduces ASR by more than 40% on average. It greatly reduces the generation of specific concepts, indicating that the present invention has strong robustness against different types of adversarial attack cue texts and can resist potential adversarial attacks in the real world. This robustness is attributed to: (1) selective cue infiltration adjustment effectively detects and replaces lexical units related to the target concept according to semantics, and (2) diffusion model noise correction resists adversarial cueing.

[0109] Based on the same inventive concept, this embodiment also provides a text-to-video diffusion model concept erasure system based on two-stage adjustment, including:

[0110] The input text embedding extraction module is used to receive input text prompts, extract the original prompt embedding matrix and pool it to obtain the prompt embedding vector, and calculate the first projection matrix that projects the prompt embedding vector onto the prompt subspace;

[0111] The target concept subspace modeling module is used to concatenate predefined target concepts, extract the target concept embedding matrix and pool it to obtain the target concept embedding vector, calculate the second projection matrix that projects the target concept embedding vector onto the target concept subspace, and obtain the third projection matrix of the orthogonal complement space of the target concept subspace based on the second projection matrix.

[0112] The prompt embedding dynamic correction module is used to perform mask sensitivity analysis on each word of the input text prompt using the third projection matrix, calculate its orthogonal projection distance with the target concept subspace, filter trigger words, and perform orthogonal projection correction on the original prompt embedding matrix based on the trigger words to generate the corrected prompt embedding matrix.

[0113] The anti-adversarial guidance module is used to calculate a dynamic scaling factor to adjust the noise estimation direction by combining the corrected cue embedding matrix and the target concept embedding matrix during the process of generating video frame images through diffusion denoising.

[0114] The video generation module is used to iteratively denoise based on the adjusted noise estimation, generate a video frame sequence without the target concept, and output a high-fidelity video.

[0115] For the system embodiments, since they basically correspond to the method embodiments, relevant details can be found in the descriptions of the method embodiments; the implementation methods of the remaining modules will not be repeated here. The system embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of the present invention according to actual needs. Those skilled in the art can understand and implement this without any creative effort.

[0116] The system embodiments of the present invention can be applied to any device with data processing capabilities, such as a computer or other similar device. The system embodiments can be implemented in software, hardware, or a combination of both. Taking software implementation as an example, as a logical device, it is formed by the processor of any data processing device loading the corresponding computer program instructions from non-volatile memory into memory for execution.

[0117] The above-described embodiments are merely illustrative of several implementations of the present invention, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of the invention. Those skilled in the art can make various modifications and improvements without departing from the concept of the present invention, and these modifications and improvements all fall within the scope of protection of the present invention.

Claims

1. A concept erasure method based on a two-stage adjusted text-to-video diffusion model, characterized in that, Includes the following steps: 1) Receive input text prompts, extract the original prompt embedding matrix and pool it to obtain the prompt embedding vector, and calculate the first projection matrix that projects the prompt embedding vector onto the prompt subspace; Furthermore, predefined target concepts are concatenated, target concept embedding matrices are extracted and pooled to obtain target concept embedding vectors, a second projection matrix is ​​calculated to project the target concept embedding vectors onto the target concept subspace, and a third projection matrix is ​​obtained based on the second projection matrix to obtain the orthogonal complement space of the target concept subspace. 2) Perform mask sensitivity analysis on each word of the input text prompt using the third projection matrix, calculate its orthogonal projection distance with the target concept subspace, filter trigger words, and perform orthogonal projection correction on the original prompt embedding matrix based on the trigger words to generate the corrected prompt embedding matrix; 3) During the process of generating video frame images through diffusion denoising, the dynamic scaling factor is calculated by combining the corrected cue embedding matrix and the target concept embedding matrix to adjust the noise estimation direction; 4) Based on the adjusted noise estimation, iterative denoising is performed to generate a video frame sequence without target concepts, and a high-fidelity video is output.

2. The concept erasure method based on a two-stage adjustment text-to-video diffusion model according to claim 1, characterized in that, The original prompt embedding matrix is ​​obtained by encoding the input text prompts using a text encoder, and the target concept embedding matrix is ​​obtained by encoding the concatenated target concepts using a text encoder.

3. The concept erasure method based on a two-stage adjustment text-to-video diffusion model according to claim 1, characterized in that, The relationship between the third projection matrix of the orthogonal complement space and the second projection matrix of the target concept subspace is as follows: Where I is the identity matrix, P e This is the second projection matrix. This is the third projection matrix.

4. The concept erasure method based on a two-stage adjustment text-to-video diffusion model according to claim 1, characterized in that, The target concept refers to avoiding the generation of semantic entities contained in the video.

5. The concept erasure method based on a two-stage adjustment text-to-video diffusion model according to claim 1, characterized in that, Step 2) includes: 2.1) Mask each word in the input text prompt sequentially, generate the masked prompt embedding matrix and the pooled masked prompt embedding vector, calculate the orthogonal projection distance between the masked prompt embedding vector of each masked word and the target concept subspace based on the third projection matrix, and mark the masked words with orthogonal projection distances higher than the threshold as trigger words. 2.2) Project the original cue embedding matrix sequentially onto the orthogonal complement space and the cue subspace to eliminate the association with the target concept, and obtain the semantically adjusted cue embedding matrix; 2.3) Generate a binary mask based on the trigger word, with the trigger word corresponding to 1 and the rest to 0, and calculate the corrected hint embedding matrix: Where ⊙ represents the Hadamard product, E p ′ For the corrected hint embedding matrix, E p The original hint embedding matrix, is the semantically adjusted hint embedding matrix, and m is the binary mask.

6. The concept erasure method based on a two-stage adjustment text-to-video diffusion model according to claim 5, characterized in that, In step 2.1), the formula for calculating the orthogonal projection distance between the masked hint embedding vector of each masked word and the target concept subspace is as follows: in, The masked hint embedding vector for the i-th masked word. Let be the orthogonal projection distance corresponding to the i-th masked word. Let ||.||2 be the third projection matrix, and let L2 norm be ||.||2.

7. The concept erasure method based on a two-stage adjustment text-to-video diffusion model according to claim 1, characterized in that, Step 3) includes: 3.1) During the process of generating each video frame using the conditional diffusion model, at the denoising step t, the input conditional noise estimation and the target concept conditional noise estimation are calculated using the corrected cue embedding matrix and the target concept embedding matrix as guiding conditions, respectively. In addition, the noise estimation under unconditional guidance is calculated. 3.2) Calculate the dynamic scaling factor that combines the mean of cross-frame noise differences based on the input conditional noise estimation and the target concept conditional noise estimation; 3.3) Combine input conditional noise estimation, target concept conditional noise estimation, unconditional guided noise estimation, and dynamic scaling factor to calculate the noise estimation direction.

8. The concept erasure method based on a two-stage adjustment text-to-video diffusion model according to claim 7, characterized in that, The formula for the dynamic scaling factor is as follows: Where T is the total number of denoising steps for each video frame, F is the number of video frames, w0 is the preset parameter for the erasure intensity of the target concept, μ is the dynamic scaling factor, and t is the denoising step t. For input conditional noise estimation, E p ′ The corrected hint embedding matrix, Let f be the latent variable at step t corresponding to the f-th frame of the video. For target concept conditional noise estimation, E e Let ||.||2 be the L2 norm, which is the embedding matrix for the target concept.

9. The concept erasure method based on a two-stage adjustment text-to-video diffusion model according to claim 8, characterized in that, The formula for calculating the noise estimation direction is: in, For noise estimation direction, This is noise estimation under unconditional guidance.

10. A concept erasure system based on a two-stage adjusted text-to-video diffusion model, used to implement the concept erasure method of claim 1, characterized in that, The system includes: The input text embedding extraction module is used to receive input text prompts, extract the original prompt embedding matrix and pool it to obtain the prompt embedding vector, and calculate the first projection matrix that projects the prompt embedding vector onto the prompt subspace; The target concept subspace modeling module is used to concatenate predefined target concepts, extract the target concept embedding matrix and pool it to obtain the target concept embedding vector, calculate the second projection matrix that projects the target concept embedding vector onto the target concept subspace, and obtain the third projection matrix of the orthogonal complement space of the target concept subspace based on the second projection matrix. The prompt embedding dynamic correction module is used to perform mask sensitivity analysis on each word of the input text prompt using the third projection matrix, calculate its orthogonal projection distance with the target concept subspace, filter trigger words, and perform orthogonal projection correction on the original prompt embedding matrix based on the trigger words to generate the corrected prompt embedding matrix. The anti-adversarial guidance module is used to calculate a dynamic scaling factor to adjust the noise estimation direction by combining the corrected cue embedding matrix and the target concept embedding matrix during the process of generating video frame images through diffusion denoising. The video generation module is used to iteratively denoise based on the adjusted noise estimation, generate a video frame sequence without the target concept, and output a high-fidelity video.