An Adaptive Content Eliminator and Elimination Method for Image Generation Large Models
By constructing the elimination space and the retention space, projecting token by token and correcting the direction according to the correlation degree, the precise elimination of the target content in the image generation model and the protection of irrelevant content are achieved, and the problem of difficult control of the elimination effect and negative impact on the generation process in the prior art is solved.
Patent Information
- Application Number
- CN202510213518.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-26
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2045-02-26
AI Technical Summary
When the prior art removes target content of unsafe or copyright issues in the image generation model, it is difficult to accurately control the elimination effect, and may have a negative impact on the generation process of irrelevant content, resulting in ethical, legal and security challenges.
By constructing elimination spaces and retention spaces, project the text prompts entered by the user token by token to these spaces, calculate the correlation degree and reconstruct the retention space based on the results, correcting the direction of the retention component to achieve accurate elimination of the target content and protection of irrelevant content.
Accurate removal of single or multiple target content is achieved, and the model is fine-tuned without spending a lot of computing resources. No additional parameters are introduced, which significantly improves the content elimination performance of the generated model and maximizes the generation quality of irrelevant content.
Smart Images

Figure CN119719631B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer vision technology, and more specifically to an adaptive content eliminator and elimination method for a large image generation model. Background Art
[0002] In recent years, the rapid development of text-to-image (T2I) diffusion models has enabled users to easily generate high-quality images with simple text prompts. With their powerful generation capabilities, these models have been widely used in content creation, art design, and other fields. However, the training of such models usually relies on large-scale datasets crawled from the Internet, which are of varying quality and inevitably contain unsafe content. This results in the model accidentally presenting offensive content or content involving copyright disputes when generating images, which brings ethical, legal, and security challenges to practical applications. To address these issues, an intuitive solution is to retrain the generative model after strictly screening the dataset. However, this approach has significant limitations. First, it is often necessary to design specific detectors to screen data for different types of target content, which not only increases the complexity of development, but also cannot cope with unknown or newly emerging target content. Second, retraining large generative models from scratch is extremely costly, time-consuming, and resource-intensive. In addition, even if the model is retrained after screening, it may still introduce unexpected bad tendencies due to data bias. Another simple solution is to use text filters or safety checkers. For example, text filters can directly block text prompts containing target content, and security checkers can filter results containing unsafe content by detecting generated images. However, these methods have limited protection capabilities and are easily bypassed by cleverly designed malicious prompts, and cannot fundamentally solve the problem. Therefore, it is particularly important to develop a low-cost, efficient, and accurate technology to remove unwanted semantics in generated images (i.e., content removal).
[0003] Existing content elimination techniques are mainly divided into two categories. First, there are methods based on fine-tuning. These methods fine-tune the model parameters to specifically remap the semantics of the target content and the replacement content, or weaken the expression of the target content. Although these methods are relatively excellent in the effect of content elimination, their practical applications face two major bottlenecks: one is that separate training is required for each target content, resulting in the consumption of time and resources; the other is that it is difficult to simultaneously take into account the precise elimination of the target content and the protection of the generation process of irrelevant content during this fine-tuning process, which is likely to have a negative impact on the generation quality. This also directly leads to the poor performance of these methods in scenarios where multiple target contents need to be eliminated. The second category of methods is those that do not require additional training. The methods without training achieve content elimination by directly intervening in the image generation process, such as the Safe Latent Diffusion (SLD) and Negative Prompt. These methods can significantly reduce the computational cost and do not require time for fine-tuning. However, their disadvantages are also very prominent: the elimination effect is usually difficult to precisely control, especially in terms of protecting the generation quality of irrelevant content. Summary of the Invention
[0004] In this embodiment, an adaptive content eliminator, an elimination method, an electronic device, and a storage medium for an image generation large model are provided to accurately remove unsafe or copyright-infringing target content during the image generation process while ensuring that there is no negative impact on the generation process of other irrelevant content.
[0005] In a first aspect, an embodiment of the present invention provides an adaptive content elimination method for an image generation large model. The adaptive content elimination method for the image generation large model includes:
[0006] Construct an elimination space and a retention space based on the target content. The elimination space corresponds to the semantic information that needs to be removed, and the retention space corresponds to the semantic information that needs to be retained;
[0007] Convert the text prompt input by the user into a semantic representation, including tokenizing the text prompt, using the generated words or phrases as tokens, and projecting them token by token into the elimination space and the retention space to obtain the retention components in the retention space;
[0008] Measure the semantic distance between the semantic representation of the current text prompt and the semantic representation of the target content, and calculate the correlation degree;
[0009] According to the correlation degree result, reconstruct the retention space and correct the direction of the retention components.
[0010] In an optional embodiment, an elimination space and a retention space are constructed according to the target content. The elimination space corresponds to the semantic information to be removed, and the retention space corresponds to the semantic information to be retained, including:
[0011] Define the target content and obtain the semantic representation of the target content;
[0012] For a single target content, use the semantic representation of the target content as a basis vector to span the elimination space;
[0013] For multiple target contents, process the semantic representations of each target content through the Schmidt orthogonalization method to obtain a set of orthonormal bases, and use the orthonormal bases to span the elimination space;
[0014] Take the space that is the orthogonal complement of the elimination space as the retention space.
[0015] In an optional embodiment, the text prompt input by the user is converted into a semantic representation, including tokenizing the text prompt, using the generated words or phrases as tokens, and projecting them token by token into the elimination space and the retention space to obtain the retention components in the retention space, including:
[0016] Tokenize the text input by the user to generate independent words or phrases. Use the generated words or phrases as tokens, and encode the tokens using a pre-trained language model to generate a semantic representation;
[0017] For a single target content, calculate the projections onto the elimination space and the retention space to obtain the elimination components of the token in the elimination space and the retention components in the retention space;
[0018] For multiple target contents, perform projection calculations on each target content separately to obtain the total elimination components and retention components;
[0019] Replace the original semantic representation with the calculated retention components to participate in the subsequent generation process.
[0020] In an optional embodiment, the tokens include special tokens, and the special tokens are not projected.
[0021] In an optional embodiment, measure the semantic distance between the semantic representation of the current text prompt and the semantic representation of the target content, and calculate the correlation degree, including:
[0022] Obtain the semantic representation of the current text prompt and the semantic representation of the target content;
[0023] Use cosine similarity to calculate the correlation degree between the semantic representation of the text prompt and the semantic representation of the target content.
[0024] In an alternative embodiment, according to the relevance result, reconstructing the retention space and correcting the direction of the retention component includes:
[0025] When the relevance is greater than a preset threshold, it indicates a strong correlation between the semantic representation of the text prompt and the semantic representation of the target content. Then, strengthen the orthogonal constraint between the retention space and the elimination space;
[0026] When the relevance is less than the preset threshold, it indicates a weak or no correlation between the semantic representation of the text prompt and the semantic representation of the target content. Then, relax the orthogonal constraint between the retention space and the elimination space to reduce the negative impact of target content elimination on the generation of irrelevant content.
[0027] In an alternative embodiment, according to the relevance result, reconstructing the retention space and correcting the direction of the retention component further includes:
[0028] For a single target content, directly correct the direction of the retention component according to the cosine similarity between the semantic representation of the target content and the semantic representation corresponding to the current text prompt;
[0029] For multiple target contents, scale the modulus of the projection of the current input text prompt along the direction corresponding to each target content according to the similarity, and correct the direction of the retention component along each target content.
[0030] Compared with the prior art, the beneficial effects of the adaptive content elimination method for the large model for image generation of the present invention are as follows:
[0031] This method can effectively remove single or multiple target contents from the generated image, without consuming a large amount of computing resources for fine-tuning the model, and without introducing additional parameters. While successfully removing single or multiple target contents, the impact on the generation of irrelevant content is negligible, significantly improving the content elimination performance of the generation model.
[0032] In a second aspect, an embodiment of the present invention provides an adaptive content eliminator for a large model for image generation, including:
[0033] A projection module, the projection module includes a semantic space construction unit and a per-token projection unit. The semantic space construction unit is used to construct an elimination space and a retention space according to the target content. The elimination space corresponds to the semantic information to be removed, and the retention space corresponds to the semantic information to be retained;
[0034] The per-token projection unit is used to convert the text prompt input by the user into a semantic representation, including tokenizing the text prompt, using the generated words or phrases as tokens, and projecting them per-token into the elimination space and the retention space to obtain the retention component of the retention space;
[0035] An adaptive direction correction module, the adaptive direction correction module includes a correlation capture unit and a direction correction unit, the correlation capture unit is used to measure the semantic distance between the semantic representation of the current text prompt and the semantic representation of the target content, and calculate the correlation;
[0036] The direction correction unit is used to reconstruct the reserved space and correct the direction of the reserved components according to the correlation result.
[0037] In a third aspect, an embodiment of the present invention provides an electronic device, including a processor, a communication interface, a memory, and a bus. Among them, the processor, the communication interface, and the memory complete communication with each other through the bus, and the processor can call the logical instructions in the memory to execute the steps of the method provided in the first aspect.
[0038] In a fourth aspect, an embodiment of the present invention provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the steps of the adaptive content elimination method for an image generation large model as described in the first aspect.
[0039] Compared with the prior art, the beneficial effects of the adaptive content eliminator, electronic device, and storage medium for an image generation large model of the present invention are the same as those of the adaptive content elimination method for an image generation large model described in the first aspect, so they will not be elaborated here. Description of the Drawings
[0040] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0041] Figure 1 It is a flowchart of the adaptive content elimination method for an image generation large model in an embodiment of the present invention;
[0042] Figure 2 It is a schematic diagram of the application of the adaptive content elimination method for an image generation large model in an embodiment of the present invention to an image generation large model;
[0043] Figure 3 It is a structural block diagram of the adaptive content eliminator for an image generation large model in an embodiment of the present invention;
[0044] Figure 4 It is a structural block diagram of the electronic device in an embodiment of the present invention. Detailed Embodiments
[0045] To better understand the purpose, technical solution, and advantages of this application, the following describes and explains this application in conjunction with the accompanying drawings and embodiments.
[0046] Unless otherwise defined, the technical terms or scientific terms involved in this application shall have the general meaning understood by those with ordinary skills in the technical field to which this application belongs. In this application, words such as "a", "one", "a kind of", "the", "these", etc. do not indicate a limitation in quantity, and they can be singular or plural. The terms "include", "comprise", "have" and any variants thereof involved in this application are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or device that includes a series of steps or modules (units) is not limited to the listed steps or modules (units), but may include unlisted steps or modules (units), or may include other steps or modules (units) inherent in these processes, methods, products, or devices. The terms "connected", "coupled", etc. involved in this application do not limit to physical or mechanical connections, but may include electrical connections, whether directly connected or indirectly connected. The "multiple" involved in this application means two or more. "And / or" describes the association relationship of associated objects, indicating that three relationships can exist. For example, "A and / or B" can represent: A exists alone, A and B exist simultaneously, and B exists alone. Usually, the character " / " indicates that the objects associated before and after are in an "or" relationship. The terms "first", "second", "third", etc. involved in this application are only used to distinguish similar objects and do not represent a specific sorting of the objects.
[0047] In an embodiment of the present invention, an adaptive content elimination method for an image generation large model is provided. Figure 1 It is a flowchart of the adaptive content elimination method for the image generation large model of the present invention, as Figure 1 shown, and this process includes the following steps:
[0048] S100. Construct an elimination space and a retention space according to the target content, where the elimination space corresponds to the semantic information to be removed, and the retention space corresponds to the semantic information to be retained;
[0049] First, it should be noted that common text-image generation models usually include a compression network for compressing an image into a latent space, and a UNet that gradually denoises from randomly sampled Gaussian noise in the latent space. UNet is the core component of text-image generation, and the cross-attention mechanism therein is used to interact between text and image representations, so that text information can guide the visual content of the generated image to be consistent with the text prompt.
[0050] In this embodiment, in the first step, a semantic subspace (elimination space and retention space) is constructed. According to the semantic information of the target content, two subspaces are constructed in the semantic space of the generation model: an elimination space representing the semantic information related to the target content to be removed and a retention space representing the semantic information unrelated to the target content. Ensure that the two subspaces are completely independent and complementary, and can cover the semantic range in the current space.
[0051] In the cross-attention mechanism, the Value space determines some content information in the generated image. Therefore, we select the Value space as the semantic space for layer-by-layer projection. Denote the overall semantic space as , and the corresponding retention space and elimination space are denoted as: and .
[0052] Specifically, an elimination space and a retention space are constructed according to the target content. The elimination space corresponds to the semantic information to be removed, and the retention space corresponds to the semantic information to be retained, including:
[0053] Define the target content and obtain the semantic representation of the target content; clarify which content needs to be removed from the generated image. For example, if the goal is to remove copyrighted artworks, the features of these works will be defined as the target content.
[0054] For a single target content, use the semantic representation of the target content as the basis vector to span the elimination space; this means that all semantic information related to this target content will be projected into this space.
[0055] For multiple target contents, process the semantic representations of each target content through the Schmidt orthogonalization method to obtain a set of orthonormal bases, and use the orthonormal bases to span the elimination space; the purpose of doing this is to ensure that the semantic information between different target contents does not interfere with each other and is convenient for subsequent projection calculations.
[0056] Take the space that is the orthogonal complement of the elimination space as the retention space. It contains all semantic information that does not belong to the target content. That is, it is the orthogonal complement space of the elimination space, ensuring that the entire original semantic space can be completely covered by these two subspaces and do not overlap with each other.
[0057] Once the elimination space and the retention space are established, the next step is to project the semantic representations obtained from the input text prompt at different layers and time steps of the generation model token by token into these two subspaces, as follows:
[0058] S200. Convert the text prompt input by the user into a semantic representation, including tokenizing the text prompt, using the generated words or phrases as tokens, and projecting them token by token into the elimination space and the retention space to obtain the retention components of the retention space;
[0059] Specifically, the text prompt input by the user is transformed into a semantic representation, including tokenizing the text prompt, using the generated words or phrases as tokens, and projecting each token into the elimination space and the retention space to obtain the retention components in the retention space, including:
[0060] Tokenize the text input by the user to generate independent words or phrases (tokens), use the generated words or phrases (tokens) as tokens, and encode the tokens using a pre-trained language model to generate a semantic representation; First, tokenize the text prompt input by the user, break it down into individual words or phrases (i.e., tokens), and use a pre-trained language model (such as CLIP, BERT, etc.) to encode these tokens to generate a high-dimensional vector representation corresponding to each token, which is their semantic representation.
[0061] It should be noted that the tokens include special tokens, and the special tokens are not projected. Exemplarily, for some language models, the first token may be a special symbol (such as [CLS]) used to capture the information of the entire sentence, and the last few tokens may also be special padding symbols. Decide whether to project these special tokens according to needs. For example, in actual operations, the prefix tokens do not contain discriminative semantic information, so they are not projected; at the same time, since the text of the target content is short, its corresponding text representation contains a large number of suffix tokens, and the information contained in the suffix tokens is often redundant, so the last valid token of the target content is used to copy and overwrite the suffix tokens for subsequent calculations.
[0062] For a single target content, calculate the projections on the elimination space and the retention space to obtain the elimination components of the tokens in the elimination space and the retention components in the retention space; that is, through the orthogonal projection formula in linear algebra, the purpose is to find the components (elimination components and retention components) of the token in the two subspaces.
[0063] For multiple target contents, perform projection calculations on each target content separately to obtain the total elimination components and retention components;
[0064] Replace the calculated retention components with the original semantic representation to participate in the subsequent generation process. While removing the unnecessary target content, the integrity of other non-target content is maintained.
[0065] Specifically, in combination with Figure 2 as shown, the representation corresponding to the current text prompt in this semantic space , where l represents the number of tokens (tokens) included in the semantic representation, represents the feature dimension of each token. By using the semantic representation of the text prompt Projected separately onto the retention space and the elimination space respectively, and by utilizing the complementary and independent properties of the two, we can obtain: , where is the retention component corresponding to in this semantic representation, is the elimination component corresponding to in this semantic representation, and there is . To achieve more flexible and precise elimination, perform token-by-token projection on , denoted as , , are respectively the j-th tokens of , , , then there are:
[0066] ;
[0067] ;
[0068] ;
[0069] where = represents the semantic representation corresponding to the target content in the semantic space when eliminating a single target content. In actual operation, the first token is a prefix token and does not contain discriminative semantic information, so no projection is performed. In addition, since the text of the target content is short, its corresponding text representation contains a large number of suffix tokens, and the information contained in the suffix tokens is often redundant. To avoid redundancy, we use the last token corresponding to the text representation of the target content to copy and overwrite the suffix tokens for the calculation of the semantic representation, and then perform subsequent projection, that is, use to replace to participate in the calculation, that is, the semantic representation calculated from is used as the basis for the spanned elimination space to ensure that the semantic information related to the target content is effectively eliminated at all positions. Finally, the retention component = is used to replace = to participate in subsequent calculations. When eliminating multiple target contents, the representation of each target content in the semantic space can be denoted as { }, then the corresponding elimination space is composed of the subspace spanned by the tokens of the target content at the corresponding positions, and the retention space is its complementary independent space. That is:
[0070] ;
[0071] where:
[0072] ;
[0073] We can obtain:
[0074] ;
[0075] S300. Measure the semantic distance between the semantic representation of the current text prompt and the semantic representation of the target content, and calculate the correlation degree;
[0076] Measuring the semantic distance between the semantic representation of the current text prompt and the semantic representation of the target content, and calculating the correlation degree, includes:
[0077] Obtain the semantic representation of the current text prompt and the semantic representation of the target content; convert the text prompt input by the user into a representation form in a high-dimensional vector space through a pre-trained language model (such as CLIP, BERT, etc.). Similarly, use the same or compatible language model to encode all target content that needs to be removed to obtain their respective semantic representations.
[0078] It should be noted that, in order to ensure the consistency and comparability of the calculation results, these semantic representations are usually normalized so that the norm of each vector is equal to 1. This can avoid the deviation of the distance metric caused by the difference in vector lengths.
[0079] Use cosine similarity to calculate the correlation degree between the semantic representation of the text prompt and the semantic representation of the target content.
[0080] S400. According to the correlation degree result, reconstruct the retention space and correct the direction of the retention components.
[0081] First, set one or more thresholds to distinguish between high and low correlation degrees. Based on the calculated correlation degree, judge the strength of the relationship between the current text prompt and the target content.
[0082] Specifically, according to the correlation degree result, reconstruct the retention space and correct the direction of the retention components, including:
[0083] When the correlation degree is greater than the preset threshold, it indicates a strong correlation between the semantic representation of the text prompt and the semantic representation of the target content, then strengthen the orthogonal constraint between the retention space and the elimination space; to ensure more accurate removal of the target content. This means to more strictly maintain the independence and complementarity of the two subspaces.
[0084] When the correlation degree is less than the preset threshold, it indicates a weak or no correlation between the semantic representation of the text prompt and the semantic representation of the target content, then relax the orthogonal constraint between the retention space and the elimination space to reduce the negative impact of target content elimination on the generation of irrelevant content.
[0085] According to the relevance results, reconstruct the reserved space and correct the direction of the reserved components, which also includes:
[0086] For a single target content, directly correct the direction of the reserved component according to the cosine similarity between the semantic representation of the target content and the semantic representation corresponding to the current text prompt;
[0087] For multiple target contents, according to the similarity, scale the modulus of the projection of the current input text prompt along the direction corresponding to each target content, and correct the direction of the reserved component along each target content.
[0088] According to the results of the relevance, dynamically adjust the basis vectors used to span the reserved space. For high relevance, more orthogonalization processing may be required to ensure that the new basis vectors better represent the semantic information of non-target contents; for low relevance, some restrictions can be considered to be relaxed, allowing some degree of overlap, but not affecting the overall effect.
[0089] Specifically, when eliminating a single target content, directly perform direction correction according to the cosine similarity between the semantic representation of the target content and the semantic representation corresponding to the current text prompt, that is:
[0090] ;
[0091] ;
[0092] Among them, represents the cosine similarity between the target content and the j-th token of the current text prompt, is the hyperparameter of the direction correction unit, and directly corrects the direction of the reserved component.
[0093] When eliminating multiple target contents, similar to the case of erasing a single target content, it is necessary to scale the modulus of the projection of the current input text prompt along the direction corresponding to each target content according to the similarity, that is, correct the direction of the reserved component along each target content, that is:
[0094] ;
[0095] ;
[0096] Through the above steps, the reconstruction of the reserved space and the correction of the direction of the reserved components are achieved according to the relevance results. This method not only ensures the accurate elimination of target contents, but also maximally protects the generation quality of irrelevant contents from being disturbed.
[0097] To illustrate the effectiveness of the present invention, the following experiments were carried out for verification.
[0098] Apply this method to the Or-Eliminator embedded in Stable Diffusion V1.4 to explore its performance in eliminating single or multiple IP characters. At the same time, use irrelevant IP to verify that the adaptive content eliminator has less impact on the generation of irrelevant content:
[0099]
[0100] Table 1 Performance comparison when eliminating IP characters
[0101] As can be seen from Table 1, for the target content to be eliminated, whether it is single or multiple, the adaptive content elimination method for the image generation large model shows content elimination performance. At the same time, the generated images of irrelevant content are hardly affected, and the FID is much lower than other methods, less than half of other methods, showing outstanding prior protection ability, and this protection ability for the generation of irrelevant content is not affected by the increase in the eliminated content.
[0102] Apply this method to Stable Diffusion V1.4 to explore its performance in eliminating painting styles as follows:
[0103]
[0104] Table 2 Performance comparison when eliminating painting styles
[0105] As can be seen from Table 2, when eliminating a certain painting style, NP shows the best elimination result, but NP has a great impact on irrelevant content while eliminating the target content. And our adaptive content elimination method for the image generation large model shows strong accuracy in the process of content elimination, and well solves the impact of content elimination on the generation of irrelevant content with almost no loss of elimination performance, achieving the best trade-off between target content elimination and irrelevant content protection.
[0106] Apply this method to Stable Diffusion V1.4 to explore its performance in eliminating unsafe content (nudity) as follows:
[0107]
[0108] Table 3 Performance comparison when eliminating unsafe content
[0109] As can be seen from Table 3, for unsafe content, our adaptive content elimination method for the image generation large model shows the best content elimination performance, and can eliminate unsafe content to the greatest extent and prevent the generation of unsafe visual information.
[0110] An embodiment of the present invention also provides an adaptive content eliminator for an image generation large model. This eliminator is used to implement the above method embodiment, and the parts that have been described will not be repeated here. The following terms such as "module", "unit", "sub-unit", etc. can be a combination of software and / or hardware that can achieve a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, implementation in hardware or a combination of software and hardware is also possible and contemplated.
[0111] As Figure 3 shown, Figure 3 is a structural block diagram of the adaptive content eliminator for an image generation large model in the present invention. The eliminator includes:
[0112] A projection module, the projection module includes a semantic space construction unit and a token-by-token projection unit. The semantic space construction unit is used to construct an elimination space and a retention space according to the target content. The elimination space corresponds to the semantic information that needs to be removed, and the retention space corresponds to the semantic information that needs to be retained;
[0113] The token-by-token projection unit is used to convert the text prompt input by the user into a semantic representation, including tokenizing the text prompt, using the generated words or phrases as tokens, and projecting them token by token into the elimination space and the retention space to obtain the retention components in the retention space;
[0114] An adaptive direction correction module, the adaptive direction correction module includes a correlation capture unit and a direction correction unit. The correlation capture unit is used to measure the semantic distance between the semantic representation of the current text prompt and the semantic representation of the target content and calculate the correlation;
[0115] The direction correction unit is used to reconstruct the retention space and correct the direction of the retention components according to the correlation result.
[0116] Figure 4 is a structural block diagram of the electronic device provided by the embodiment of the present invention. As Figure 4 shown, the electronic device may include: a processor 610, a communication interface 620, a memory 630, and a communication bus 640. Among them, the processor 610, the communication interface 620, and the memory 630 communicate with each other through the communication bus 640. The processor 610 can call the logical instructions in the memory 630 to execute the following methods:
[0117] Construct an elimination space and a retention space according to the target content. The elimination space corresponds to the semantic information that needs to be removed, and the retention space corresponds to the semantic information that needs to be retained;
[0118] The text prompt input by the user is converted into a semantic representation, including tokenizing the text prompt, using the generated words or phrases as tokens, and projecting them token by token into the elimination space and the retention space to obtain the retention components in the retention space;
[0119] Measure the semantic distance between the semantic representation of the current text prompt and the semantic representation of the target content, and calculate the correlation degree;
[0120] According to the correlation degree result, reconstruct the retention space and correct the direction of the retention components.
[0121] In addition, when the logical instructions in the above-mentioned memory 630 are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs that can store program codes.
[0122] The embodiments of the present invention further provide a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it is configured to execute the methods provided in the above-mentioned various embodiments, for example, including:
[0123] Construct an elimination space and a retention space according to the target content, where the elimination space corresponds to the semantic information to be removed, and the retention space corresponds to the semantic information to be retained;
[0124] The text prompt input by the user is converted into a semantic representation, including tokenizing the text prompt, using the generated words or phrases as tokens, and projecting them token by token into the elimination space and the retention space to obtain the retention components in the retention space;
[0125] Measure the semantic distance between the semantic representation of the current text prompt and the semantic representation of the target content, and calculate the correlation degree;
[0126] According to the correlation degree result, reconstruct the retention space and correct the direction of the retention components.
[0127] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on such an understanding, the essence of the above technical solution, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to enable a computer device (which can be a personal computer, server, or network device, etc.) to execute the methods of each embodiment or some parts of the embodiments.
[0128] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of each embodiment of the present invention.
Claims
1. An adaptive content elimination method for image generation large model, characterized in that: The elimination method comprises: Constructing an elimination space and a retention space according to the target content, wherein the elimination space corresponds to the semantic information to be removed, and the retention space corresponds to the semantic information to be retained; The text prompts input by the user are converted into semantic representations, including word segmentation of the text prompts, the generated words or phrases are used as tokens, and they are projected token by token into the elimination space and the retention space to obtain the retention component of the retention space; Measuring the semantic distance between the semantic representation of the current text prompt and the semantic representation of the target content, and calculating the relevance; According to the correlation results, the retention space is reconstructed and the direction of the retained component is corrected; According to the correlation results, the reserved space is reconstructed and the direction of the reserved component is corrected, including: When the correlation is greater than a preset threshold, the orthogonal constraints of the retained space and the eliminated space are strengthened; When the correlation is less than the preset threshold, relax the orthogonal constraints of the retained space and the eliminated space; According to the correlation results, the reserved space is reconstructed and the direction of the reserved component is corrected, which also includes: For a single target content, the direction of the retained component is directly corrected according to the cosine similarity between the semantic representation of the target content and the semantic representation corresponding to the current text prompt; For multiple target contents, the projection of the currently input text prompt along the direction corresponding to each target content is scaled according to the similarity, and the direction of the retained component is corrected along each target content.
2. The adaptive content elimination method for image generation large model according to claim 1, characterized in that: Constructing an elimination space and a retention space according to the target content, wherein the elimination space corresponds to the semantic information to be removed, and the retention space corresponds to the semantic information to be retained, including: Define the target content and obtain the semantic representation of the target content; For a single target content, the semantic representation of the target content is used as the basis vector to span the elimination space; For multiple target contents, the semantic representation of each target content is processed by Schmidt orthogonalization method to obtain a set of standard orthogonal bases, and the standard orthogonal bases are used to span the elimination space; The space that is orthogonal to the elimination space is taken as the retention space.
3. The adaptive content elimination method for image generation large model according to claim 1, characterized in that: The text prompts input by the user are converted into semantic representations, including word segmentation of the text prompts, and the generated words or phrases are used as tokens, and are projected token by token into the elimination space and the retention space to obtain the retention components of the retention space, including: Perform word segmentation on the text input by the user to generate independent words or phrases, use the generated words or phrases as tokens, encode the tokens using the pre-trained language model, and generate semantic representations; For a single target content, calculating the projection on the elimination space and the retention space to obtain the elimination component of the token in the elimination space and the retention component of the retention space; For multiple target contents, projection calculation is performed on each target content to obtain a total eliminated component and a retained component; The calculated retained component replaces the original semantic representation and participates in the subsequent generation process.
4. The adaptive content elimination method for image generation large model according to claim 3, characterized in that: The tokens include special tokens, and the special tokens are not projected.
5. The adaptive content elimination method for image generation large model according to claim 1, characterized in that: Measure the semantic distance between the semantic representation of the current text prompt and the semantic representation of the target content, and calculate the relevance, including: Obtaining the semantic representation of the current text prompt and the semantic representation of the target content; The cosine similarity is used to calculate the correlation between the semantic representation of the text prompt and the semantic representation of the target content.
6. An adaptive content remover for large image generation models, characterized in that include: A projection module, the projection module comprising a semantic space construction unit and a token-by-token projection unit, the semantic space construction unit is used to construct an elimination space and a retention space according to the target content, the elimination space corresponds to the semantic information to be removed, and the retention space corresponds to the semantic information to be retained; The token-by-token projection unit is used to convert the text prompt input by the user into a semantic representation, including word segmentation processing of the text prompt, using the generated words or phrases as tokens, and projecting them token by token into the elimination space and the retention space to obtain the retention component of the retention space; An adaptive direction correction module, the adaptive direction correction module comprising a relevance capture unit and a direction correction unit, the relevance capture unit is used to measure the semantic distance between the semantic representation of the current text prompt and the semantic representation of the target content, and calculate the relevance; The direction correction unit is used to reconstruct the reserved space and correct the direction of the reserved component according to the correlation result; According to the correlation results, the reserved space is reconstructed and the direction of the reserved component is corrected, including: When the correlation is greater than a preset threshold, the orthogonal constraints of the retained space and the eliminated space are strengthened; When the correlation is less than the preset threshold, relax the orthogonal constraints of the retained space and the eliminated space; According to the correlation results, the reserved space is reconstructed and the direction of the reserved component is corrected, which also includes: For a single target content, the direction of the retained component is directly corrected according to the cosine similarity between the semantic representation of the target content and the semantic representation corresponding to the current text prompt; For multiple target contents, the projection of the currently input text prompt along the direction corresponding to each target content is scaled according to the similarity, and the direction of the retained component is corrected along each target content.
7. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the program, the adaptive content elimination method for image generation large model according to any one of claims 1 to 6 is implemented.
8. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the adaptive content elimination method for image generation large model as described in any one of claims 1 to 6 are implemented.
Citation Information
Patent Citations
Image generation content suppression method and system based on text graph diffusion model
CN117251589A