A secure controllable video generation method and system

By employing singular value decomposition and weight adjustment methods, the problem of removing harmful information in diffusion model video generation is solved, achieving a balance between security and generation quality. This approach is applicable to various model architectures and reduces the security risks of video generation.

CN120897107BActive Publication Date: 2025-12-16UNIV OF SCI & TECH OF CHINA
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511430517.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-10-09
Publication Date
2025-12-16
Estimated Expiration
2045-10-09

AI Technical Summary

Technical Problem

Existing video generation technologies based on diffusion models struggle to effectively remove harmful information while generating realistic videos, and they also fail to improve security and applicability while preserving normal semantics. Existing methods often result in decreased generation quality or limited generalization ability.

Method used

Harmful concept features are decomposed into multiple independent semantic directions through singular value decomposition, projection coefficients and projection variances are calculated, semantic distribution and concept structure weights are constructed, text features are dynamically adjusted, and a noise estimator and a 3D decoder are combined to generate safe videos.

Benefits of technology

Without compromising the quality of the generated video, it significantly reduces security risks in video generation, ensuring the security and consistency of the generated video, and is applicable to various video generation models.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120897107B_ABST
    Figure CN120897107B_ABST
Patent Text Reader

Abstract

The application relates to the technical field of layered semantic adjustment, and discloses a safe and controllable video generation method and system; the method comprises the following steps: encoding a text prompt input by a user and a harmful concept text respectively to obtain text prompt features and harmful concept features; decomposing the harmful concept features into multiple independent semantic directions, and calculating projection coefficients and projection variances of the text prompt features in different semantic directions; constructing and fusing semantic distribution weights and concept structure weights, scaling and adjusting the text prompt features to obtain safe text features after removing harmful semantics; taking the safe text features as conditions, performing multi-step iterative denoising in a video latent space, predicting and removing noises, and obtaining denoised video latent space representation; and mapping the denoised video latent space representation to an original video space to generate a final safe video. The application can significantly reduce the safety risk in video generation.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of hierarchical semantic adjustment, and particularly relates to a safe and controllable video generation method and system. BACKGROUND

[0002] In recent years, diffusion models have achieved remarkable success in generating realistic and imaginative videos based on textual descriptions. Therefore, video generation technology based on diffusion models is widely used in various fields, including film production, virtual reality, and social media.

[0003] Video generation technology based on diffusion models mainly includes forward and reverse processes. In the forward process, a pre-trained 3D encoder compresses the training video into features in the latent space, and noise satisfying the Gaussian distribution is gradually added to the video features at each time step. In the reverse process, a deep neural network is usually used as a noise estimator, which predicts the noise added at each time step using text features as conditions. By minimizing the estimated noise and the real noise, the noise estimator is trained. In the inference process, the noise estimator gradually predicts the noise from the Gaussian distribution of noise according to the text features as conditions, and performs denoising, and finally obtains the video features. After passing through the 3D decoder, the final video is obtained.

[0004] However, the rapid development of video generation technology based on diffusion models also brings new security challenges. In particular, they can generate videos carrying harmful information, posing a huge social risk and severely limiting their reliability and practical applicability. Since video generation models have great application potential, it is particularly important to ensure their security. Compared with images, videos introduce a time dimension and contain more information, which can fully exhibit various semantic information of text prompts. Therefore, an important challenge to effectively improve the security of video generation content is to remove harmful semantics in the text prompt while retaining normal semantic information to ensure that the generated video meets user needs.

[0005] Although existing video generation techniques have made significant progress in terms of quality and diversity, they still face serious challenges in terms of content safety. Existing methods mainly use fine-tuning or modifying model parameters to eliminate unsafe elements, which has the following main shortcomings: 1) limited generalization ability of the model. Methods based on fine-tuning model parameters of a certain architecture are difficult to apply to other architectures. This architecture dependency makes it difficult for these methods to effectively control safety when facing new model architectures or complex scenarios. 2) Difficulty in balancing safety and semantic preservation. Existing methods often weaken normal semantic expression when eliminating unsafe content, resulting in a decline in the quality of generated content. For example, overly strict filtering can result in generated videos lacking details or not meeting user needs, and even introduce new errors or unnatural expressions. SUMMARY

[0006] To solve the above technical problems, the present application provides a safe and controllable video generation method and system. The present application operates on text features in a text feature space, removes unsafe content in prompt features while preserving core semantics, decomposes the harmful concept subspace into multiple independent semantic directions, and quantifies the semantic distribution of text features on harmful concepts by calculating the projection variance of harmful concepts in these directions, thereby ensuring the hierarchical removal of unsafe semantics in key directions while ensuring that the generated safe content is consistent with user needs.

[0007] To solve the above technical problems, the present application adopts the following technical solutions:

[0008] In a first aspect, the present application provides a safe and controllable video generation method, comprising:

[0009] Encoding the user input text prompt and the harmful concept text predefined based on the large language model to obtain text prompt features and harmful concept features;

[0010] Decomposing the harmful concept features into multiple independent semantic directions through singular value decomposition, and calculating the projection coefficients and projection variances of the text prompt features in different semantic directions to quantify the correlation strength between the text prompt and the harmful concept text; constructing and fusing semantic distribution weights and concept structure weights, and scaling and adjusting the text prompt features according to the generated maximum weight to obtain safe text features with harmful semantics removed;

[0011] Using a noise estimator to perform multi-step iterative denoising in the video latent space conditioned on the safe text features, predict and remove noise, and obtain denoised video latent space representation;

[0012] Using a 3D decoder to map the denoised video latent space representation to the original video space to generate the final safe video.

[0013] In one embodiment, the step of decomposing harmful concept features into multiple independent semantic directions through singular value decomposition specifically includes:

[0014] Harmful Concept Features Through Singular Value Decomposition Decompose:

[0015] ;

[0016] Among them, the right singular vector matrix It constitutes The entire semantic space, The number of right singular vectors. Indicates the first A right singular vector, Represents a diagonal matrix. Describes a left singular vector matrix. This indicates transpose.

[0017] In one embodiment, calculating the projection coefficients and projection variances of the text prompt features in different semantic directions to quantify the association strength between the text prompt and the harmful concept text specifically includes:

[0018] Calculate text prompt features Projection of harmful conceptual characteristics: ,

[0019] in, This represents the overall projection matrix of textual prompt features in the harmful concept space; for The The element in row j represents the first element in the text prompt. The text prompt word is in the j-th semantic direction Projection coefficients on;

[0020] Calculate the projection variance :

[0021] ,

[0022] in, Indicates the length of the text prompt.

[0023] In one embodiment, the construction and fusion of semantic distribution weights and conceptual structure weights, and the scaling and adjustment of the text prompt features based on the generated final weights to obtain safe text features after removing harmful semantics, specifically includes:

[0024] ;

[0025] The weights represent the normalized semantic distribution. Let represent the projection variance in the j-th semantic direction;

[0026] Based on the characteristics of the concept of harmfulness Singular value construction of conceptual structure weights : ; Indicates harmful concept characteristics The j-th singular value obtained after decomposition;

[0027] Through hyperparameters Dynamically fuse semantic distribution weights and conceptual structure weights to generate the final weights. ;

[0028] Construct a diagonal scaling matrix Furthermore, hierarchical semantic adjustment is applied to the text prompt features to obtain secure text features. ; This represents the operation of converting a vector into a diagonal matrix.

[0029] In one embodiment, the step of using a noise estimator, conditioned on the secure text features, to perform multi-step iterative denoising in the video latent space, predicting and removing noise to obtain a denoised video latent space representation, specifically includes:

[0030] Noise estimator Utilizing secure text features To predict the noise at each time step Perform multiple rounds of noise reduction: ,in, Indicates time step The following video features For time step index, and All are control time steps The parameters of noise level; after After rounds of denoising, the final denoised video latent space representation is obtained. .

[0031] Secondly, the present invention provides a secure and controllable video generation system, comprising:

[0032] The encoding module encodes the text prompts input by the user and the harmful concept texts predefined based on the large language model, respectively, to obtain text prompt features and harmful concept features;

[0033] The hierarchical semantic adjustment module decomposes harmful concept features into multiple independent semantic directions through singular value decomposition, and calculates the projection coefficients and projection variances of text prompt features in different semantic directions to quantify the association strength between text prompts and harmful concept text; it constructs and merges semantic distribution weights and concept structure weights, and scales and adjusts the text prompt features according to the generated final weights to obtain safe text features after removing harmful semantics.

[0034] The noise estimation module uses a noise estimator to perform multi-step iterative denoising in the video latent space, based on the secure text features, to predict and remove noise, and obtain the denoised video latent space representation.

[0035] The decoding module uses a 3D decoder to map the denoised video latent space representation to the original video space, generating the final secure video.

[0036] In one embodiment, the step of decomposing harmful concept features into multiple independent semantic directions through singular value decomposition specifically includes:

[0037] Harmful Concept Features Through Singular Value Decomposition Decompose:

[0038] ;

[0039] Among them, the right singular vector matrix It constitutes The entire semantic space, The number of right singular vectors. Indicates the first One right singular vector; Represents a diagonal matrix. Describes a left singular vector matrix. This indicates transpose.

[0040] In one embodiment, calculating the projection coefficients and projection variances of the text prompt features in different semantic directions to quantify the association strength between the text prompt and the harmful concept text specifically includes:

[0041] Calculate text prompt features Projection of harmful conceptual characteristics: ,

[0042] in, This represents the overall projection matrix of textual prompt features in the harmful concept space; for The The element in row j represents the first element in the text prompt. The text prompt word is in the j-th semantic direction projection coefficients on the text prompt;

[0043] computing projection variance :

[0044] ,

[0045] wherein, denotes the text length of the text prompt.

[0046] In one of the embodiments, the construction and fusion of semantic distribution weight and concept structure weight are performed, and the generated maximum weight is used to scale and adjust the text prompt features to obtain safe text features after removing harmful semantics, which specifically includes:

[0047] ;

[0048] denotes the normalized semantic distribution weight, denotes the projection variance in the jth semantic direction;

[0049] based on the singular value of harmful concept features to construct the concept structure weight : ; denotes the jth singular value obtained after decomposition of harmful concept features ;

[0050] dynamically fuse the semantic distribution weight and the concept structure weight through the hyperparameter to generate the maximum weight ;

[0051] construct a diagonal scaling matrix , and apply hierarchical semantic adjustment to the text prompt features to obtain safe text features ; denotes the conversion of a vector into a diagonal matrix.

[0052] In one of the embodiments, the noise estimator is used to perform multi-step iterative denoising in the video latent space conditioned on the safe text features, predict and remove noise to obtain the denoised video latent space representation, which specifically includes:

[0053] noise estimator use the safe text features to predict the noise of each time step , and perform multi-round denoising: wherein, denotes the video feature at time step , is the index of time step , and All are control time steps The parameters of noise level; after After rounds of denoising, the final denoised video latent space representation is obtained. .

[0054] The system and method in this invention correspond to each other; the specific technical solutions applicable to the method are also applicable to the system.

[0055] Compared with the prior art, the beneficial technical effects of the present invention are:

[0056] 1. This invention provides an end-to-end solution to the problem of insecure content in videos generated by current methods. By constructing a unified framework that includes a text encoder, a hierarchical semantic adjustment module, a noise estimator, and a 3D decoder, this invention significantly reduces security risks in video generation and provides strong support for the secure application of video generation technology.

[0057] 2. This invention addresses the problem of balancing the security and quality of generated videos in current methods by proposing a hierarchical semantic adjustment method that does not require retraining the model or adjusting model parameters. By dynamically adjusting semantic weights in the text prompt feature space, this invention can effectively remove harmful content while preserving normal semantics, thereby significantly improving the security of video generation without reducing the quality of the generated video.

[0058] 3. This invention addresses the problem that most current methods are based on specific model architectures and are difficult to generalize to other architectures. It proposes a general, training-independent framework. This framework seamlessly integrates the hierarchical semantic adjustment module into existing diffusion models, making it applicable to various video generation models and greatly improving the method's versatility and applicability. Attached Figure Description

[0059] Figure 1 This is a flowchart of the method in an embodiment of the present invention.

[0060] Figure 2 This is a schematic diagram of the frame structure in an embodiment of the present invention. Detailed Implementation

[0061] A preferred embodiment of the present invention will now be described in detail with reference to the accompanying drawings.

[0062] To better express the safe controllable video generation method proposed in the application, the following will be combined with the drawings, taking violence as an unsafe element, the user input text prompt as "The man is stabbing a cat with a knife, and the cat is bleeding.", and the harmful concept text as "Blood, Aggression, War, Abuse, Self-Harm", to further illustrate the application.

[0063] The safe controllable video generation method proposed in the application seamlessly integrates hierarchical semantic adjustment into the diffusion model without the need for retraining the model or adjusting the model parameters, forming a unified framework for safe generation. The framework consists of four parts: a text encoder, a hierarchical semantic adjustment module, a noise estimator, and a 3D decoder. Initially, the text encoder encodes the user input text prompt and the harmful concept text respectively to obtain the corresponding features. Subsequently, the hierarchical semantic adjustment module measures the distance between the features of each word in the text prompt and the harmful concept features in the feature space, and quantifies the association strength of each word with the harmful concept by calculating the projection variance. Based on these association strengths, the hierarchical semantic adjustment module dynamically adjusts the semantic weights in the text prompt features, and the stronger the association strength, the stronger the adjustment strength of the text prompt. The safe text features obtained by processing the hierarchical semantic adjustment module are passed to the noise estimator, which uses the safe text features to guide the denoising step in the video generation process. Since the safe text features have been adjusted by the hierarchical semantic adjustment, the denoising process can more effectively avoid generating video content related to harmful concepts, thereby ensuring the safety of the generated video. Finally, the 3D decoder maps the denoised video latent space representation back to the original video space to generate the final safe video.

[0064] As shown in Figure 1 , a safe controllable video generation method in the application includes the following steps:

[0065] S1, encoding the user input text prompt and the harmful concept text based on the pre-defined large language model to obtain text prompt features and harmful concept features;

[0066] S2, decomposing the harmful concept features into multiple independent semantic directions by singular value decomposition, and calculating the projection coefficients and projection variances of the text prompt features in different semantic directions to quantify the association strength between the text prompt and the harmful concept text; constructing and fusing semantic distribution weights and concept structure weights, and scaling the text prompt features according to the generated maximum weight to obtain safe text features with harmful semantics removed;

[0067] S3. Using a noise estimator, with the secure text features as conditions, perform multi-step iterative denoising in the video latent space, predict and remove noise, and obtain the denoised video latent space representation.

[0068] S4, using a 3D decoder to map the denoised video latent space representation to the original video space, generating the final secure video.

[0069] In a preferred embodiment, the predefined harmful concept texts can be obtained by first defining different harmful concept categories according to OpenAI's usage strategy, and aligning them with real-world deployment standards. Then, the present invention uses the GPT-4o model to generate diverse harmful concept texts covering both explicit and implicit expressions.

[0070] like Figure 2 As shown, in one embodiment, step S1 specifically includes:

[0071] For user-inputted text prompts and harmful conceptual text, a pre-trained text encoder is used for encoding. The text encoder can be a Transformer-based model, such as the CLIP model or the T5 model. For text prompts... and harmful conceptual texts Each through a text encoder Obtain text prompt features and harmful concept characteristics ,in, and These represent the lengths of the text prompt and the harmful concept text, respectively. yes and The feature dimension. In this embodiment, It is 16. It is 5. The value is 4096.

[0072] like Figure 2 As shown, in one embodiment, step S2 specifically includes:

[0073] First, we use singular value decomposition to analyze the harmful concept characteristics. Decompose: The right singular vector matrix It constitutes The entire semantic space, in which for and The minimum value is then calculated. Next, the text prompt features are calculated. The formula for projecting onto harmful conceptual features is: ,in, Indicates the first The text prompt word is in the j-th semantic direction The projection coefficients on the surface. This represents a diagonal matrix whose diagonal elements are singular values.

[0074] To quantify the specific distribution intensity, the projection variance is calculated. , ,in, The weights representing the normalized semantic distribution reflect... exist For each direction, the degree of dependence is considered; larger values ​​indicate a stronger association with harmful concepts and should be removed preferentially. Similarly, based on... Singular value construction of conceptual structure weights Among these, higher singular values ​​correspond to the core semantic directions of the concept subspace, and these directions are preferentially removed. This is achieved through hyperparameters. The two weights are dynamically combined to generate the final weight. In this embodiment, The value is 0.5. Construct a diagonal scaling matrix. Furthermore, hierarchical semantic adjustment is applied to the text prompt features to obtain secure text features. Safe text features shape and same.

[0075] In one embodiment, step S3 specifically includes:

[0076] Noise estimator Predicting noise at each time step using secure text features ,in Indicates time step Based on the video features, perform multiple rounds of noise reduction: ,in and All are control time steps The parameter for noise level. After... After rounds of denoising, the final denoised video latent space representation is obtained. .

[0077] In a preferred embodiment, the noise estimator may employ a deep learning network, such as a Unet network or a Transformer network.

[0078] In one embodiment, step S4 specifically includes:

[0079] via 3D decoder Representing the latent space of the denoised video mapping back to the original video space to generate the final safety video .

[0080] In a preferred embodiment, the 3D decoder can employ a deep learning network, generally composed of multiple layers of 3D convolutional layers and upsampling operations interleaved.

[0081] The present application can significantly reduce the time risk in video generation without reducing the generation quality, ensuring the consistency and safety of the generated video in the time dimension.

[0082] The present application designs a safety video generation framework without retraining the model or adjusting the model parameters, which includes a text encoder, a hierarchical semantic adjustment module, a noise estimator and a 3D decoder. In the hierarchical semantic adjustment module, the present application calculates the distance between the features of each word in the text prompt and the harmful concept features in the feature space, quantifies the association strength of each word with the harmful concept, and dynamically adjusts the semantic weight in the text prompt features based on the association strength, thereby removing the harmful semantics while retaining the normal semantics. The entire framework of the present application can significantly reduce the time risk in video generation without reducing the generation quality, ensuring the consistency and safety of the generated video in the time dimension.

[0083] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to be limiting of the present application. As used herein, the term "includes" and tautological derivatives thereof means that the named features, steps, operations and / or components are present, but does not exclude the presence or addition of one or more other features, steps, operations or components.

[0084] It should be understood that although the steps in the flowcharts of the specification drawings are shown in sequence according to the direction of the arrows, these steps are not necessarily executed in sequence according to the direction of the arrows. Unless otherwise specified herein, the execution of these steps is not strictly limited in sequence, and these steps can be executed in other sequences. Moreover, at least part of the steps in the flowcharts of the specification drawings can include multiple steps or multiple stages, which are not necessarily executed at the same time, but can be executed at different times, and the execution sequence of these steps or stages is not necessarily sequential, but can be executed in rotation or alternation with other steps or steps or stages in other steps.

[0085] Based on the description of the method embodiments, the present application further provides a system. The system can be a software (application), module, component, server, client, etc. using the method described in the embodiments of the present application and combined with the necessary implementation hardware. Based on the same innovative concept, the system in one or more embodiments provided by the embodiments of the present application is as described in the following embodiments. Since the implementation scheme of the system solving the problem is similar to the method, the implementation of the specific system in the embodiments of the present application can be referred to the implementation of the foregoing method, and the repeated parts will not be described herein. The term "module" or "module" used below is a combination of software and / or hardware that can realize a predetermined function. Although the system described in the following embodiments is preferably implemented in software, hardware, or a combination of software and hardware is also possible and is conceived.

[0086] A safe controllable video generation system comprises:

[0087] An encoding module encodes a user input text prompt and a harmful concept text predefined based on a large language model to obtain text prompt features and harmful concept features, respectively.

[0088] A hierarchical semantic adjustment module decomposes the harmful concept features into multiple independent semantic directions through singular value decomposition, calculates the projection coefficients and projection variances of the text prompt features in different semantic directions to quantify the correlation strength between the text prompt and the harmful concept text, constructs and fuses semantic distribution weights and concept structure weights, and scales and adjusts the text prompt features according to the generated final weights to obtain safe text features after removing harmful semantics.

[0089] A noise estimation module uses a noise estimator to perform multi-step iterative denoising in a video latent space conditioned on the safe text features, predict and remove noise, and obtain a denoised video latent space representation.

[0090] A decoding module uses a 3D decoder to map the denoised video latent space representation to an original video space to generate a final safe video.

[0091] The technical features of the above embodiments can be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described, however, as long as the combination of the technical features does not exist contradictory, it should be considered as the scope of the present application.

[0092] It will be obvious to a person skilled in the art that the application is not limited to the details of the foregoing exemplary embodiments and can be implemented in other concrete forms without departing from the spirit or essential characteristics of the application. The embodiments are therefore to be considered in all respects as illustrative and not restrictive, the scope of the application being indicated by the appended claims rather than by the foregoing description, and all changes which come within the meaning and range of equivalency of the claims are therefore intended to be embraced therein and no

[0093] Furthermore, it should be understood that although the description is made on the basis of embodiments, not every embodiment contains only one independent technical solution, and the description is made in this way only for the sake of clarity, and a person skilled in the art should consider the description as a whole, and the technical solutions in each embodiment can also be appropriately combined to form other embodiments that can be understood by a person skilled in the art.

Claims

1. A method for secure controllable video generation, characterized in that, The method comprises the following steps: encoding the user input text prompt and the harmful concept text predefined based on the large language model to obtain text prompt features and harmful concept features; decompose the harmful concept features into multiple independent semantic directions through singular value decomposition, and calculate the projection coefficients and projection variances of the text prompt features in different semantic directions to quantify the correlation strength between the text prompt and the harmful concept text; construct and fuse the semantic distribution weight and the concept structure weight, and scale and adjust the text prompt features according to the generated maximum weight to obtain safe text features after removing harmful semantics; using a noise estimator, the safe text features are used as conditions to perform multi-step iterative denoising in the video latent space, predict and remove noise, and obtain denoised video latent space representation; using a 3D decoder, the denoised video latent space representation is mapped to the original video space to generate a final safe video.

2. The method of claim 1, wherein, The singular value decomposition of the harmful concept features into multiple independent semantic directions comprises: Decomposing harmful concept features by singular value decomposition Decomposition: ; where the right singular vector matrix constitutes the entire semantic space of , where is the number of right singular vectors, is the th right singular vector, is a diagonal matrix, denotes the transpose.

3. The method of claim 2, wherein, The calculation of the projection coefficients and projection variances of the text prompt features in different semantic directions to quantify the correlation strength between the text prompt and the harmful concept text comprises: Computing text cues features In projection of harmful concept features: , wherein, represents the overall projection matrix of the textual cue feature in the harmful concept space; is the first row jthcolumn element of the matrix represents the projection coefficient of the jthsemantic direction of the thtextual cue word in the textual cue on the jthsemantic direction. Computing the projection variance : , wherein, represents the text length of the text prompt.

4. The method of claim 3, wherein, The construction and fusion of the semantic distribution weight and the concept structure weight, and the scaling and adjustment of the text prompt features according to the generated maximum weight to obtain safe text features after removing harmful semantics comprises: ; denotes the normalized semantic distribution weight, denotes the projection variance in the j-th semantic direction; Constructing concept structure weights based on singular values of harmful concept features : : ; representing harmful concept features the jth singular value obtained after decomposition; By hyperparameters Dynamically fusing semantic distribution weights and concept structure weights to generate the most right weights ; Constructing a diagonal scaling matrix , and applying hierarchical semantic adjustment to the text prompt features to obtain secure text features ; represents an operation of converting a vector into a diagonal matrix.

5. The method of claim 1, wherein, The use of a noise estimator, the safe text features are used as conditions to perform multi-step iterative denoising in the video latent space, predict and remove noise, and obtain denoised video latent space representation comprises: Noise estimator Utilizing secure text features To predict noise at each time step , performing multiple rounds of denoising: where, represents the video feature at time step , is the index of time step , and are parameters that control the degree of noise at time step ; after rounds of denoising, the final denoised video latent space representation is obtained.

6. A secure controllable video generation system, characterized by The method comprises the following steps: The encoding module encodes the user input text prompt and the harmful concept text predefined based on the large language model to obtain text prompt features and harmful concept features; The hierarchical semantic adjustment module decomposes the harmful concept features into multiple independent semantic directions through singular value decomposition, and calculates the projection coefficients and projection variances of the text prompt features in different semantic directions to quantify the correlation strength between the text prompt and the harmful concept text; construct and fuse the semantic distribution weight and the concept structure weight, and scale and adjust the text prompt features according to the generated maximum weight to obtain safe text features after removing harmful semantics; The noise estimation module uses a noise estimator to perform multi-step iterative denoising in the video latent space based on the safe text features, predicts and removes noise, and obtains denoised video latent space representation; The decoding module uses a 3D decoder to map the denoised video latent space representation to the original video space to generate a final safe video.

7. A secure controllable video generation system according to claim 6, wherein, The singular value decomposition of the harmful concept features into multiple independent semantic directions comprises: Decomposing harmful concept features by singular value decomposition Decomposition: ; wherein the right singular vector matrix constitutes the entire semantic space of the number of right singular vectors denotes the th right singular vector; denotes a diagonal matrix, denotes a left singular vector matrix, denotes the transpose.

8. A secure controllable video generation system according to claim 7, wherein, The calculation of the projection coefficients and projection variances of the text prompt features in different semantic directions to quantify the correlation strength between the text prompt and the harmful concept text comprises: Computing text cues features In projection of harmful concept features: , wherein, represents the overall projection matrix of the textual cue feature in the harmful concept space; is the first row, jthcolumn element of the projection coefficient of the thtextual cue word in the textual cue in the jthsemantic direction Computing the projection variance : , wherein, represents the text length of the text prompt.

9. A secure controllable video generation system according to claim 8, wherein, The construction and fusion of the semantic distribution weight and the concept structure weight, and the scaling and adjustment of the text prompt features according to the generated maximum weight to obtain safe text features after removing harmful semantics comprises: ; denotes the normalized semantic distribution weight, denotes the projection variance in the j-th semantic direction; Constructing concept structure weights based on singular values of harmful concept features : : ; representing harmful concept features the jth singular value obtained after decomposition By hyperparameters Dynamically fusing semantic distribution weights and concept structure weights to generate the most right weights ; Constructing a diagonal scaling matrix , and applying hierarchical semantic adjustment to the text prompt features to obtain secure text features ; denotes converting a vector to a diagonal matrix.

10. The secure controllable video generation system of claim 6, wherein, The noise estimator is used to iteratively denoise in the video latent space conditioned on the secure text features, predict and remove noise, and obtain a denoised video latent space representation, specifically including: Noise estimator Utilizing safety text features To predict noise at each time step , performing multiple rounds of denoising: where, represents the video feature at time step , is the index of time step , and are parameters that control the degree of noise at time step ; after rounds of denoising, the final denoised video latent space representation is obtained.

Citation Information

Patent Citations

  • Cooperative control text generation method and system based on potential semantic space transformation

    CN120524997A

  • Security test method based on concept decomposition and recombination

    CN120597276A