Content generation method and device, electronic equipment, storage medium and program product
Through selective suppression of joint attention matrix and collaborative double-layer optimization, combined with reverse contrast learning and low-rank parameterized adaptation, the problem of precise suppression and high-quality generation of sensitive concepts by the content generation model under the implicit cross attention mechanism is solved, and the security and generation efficiency are improved.
Patent Information
- Application Number
- CN202510272202.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-07
- Publication Date
- 2025-07-25
AI Technical Summary
The existing content generation model is prone to learning content with improper or sensitive attributes during training, which leads to ineffective control when generating content, poses security risks, and it is difficult for existing methods to achieve accurate suppression and high-quality generation under the implicit cross attention mechanism.
By selectively suppressing the target concept related column vectors in the joint attention matrix, combining collaborative double-layer optimization and reverse contrast learning, precise suppression of target attributes and maintenance of non-target attributes are achieved, and low-rank parameterized adaptation is used to fine-tune the existing model, which is suitable for content generation models of different architectures.
Accurate suppression of sensitive concepts under the implicit cross attention mechanism, while maintaining the model's ability to generate high-quality compliant content, improving the robustness and generation efficiency of the model.
Smart Images

Figure CN120372303A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of artificial intelligence technology, and in particular, to a content generation method, device, electronic device, storage medium, and program product. Background Art
[0002] With the rapid development of artificial intelligence, it is possible to guide a model to generate content by inputting text.
[0003] However, these models are usually trained based on large-scale, undifferentiated data sets, resulting in the possibility that they may learn content with improper or sensitive attributes. At this time, regardless of the text input by the user, content will be generated according to the input text, which is likely to cause security problems. Summary of the Invention
[0004] In view of this, the purpose of the present application is to provide a content generation method, device, electronic device, storage medium, and program product, which can suppress text with target attributes (also known as target concepts or negative concepts), and maintain text with non-target attributes to output content, thereby being able to purify the generated content.
[0005] In a first aspect, an embodiment of the present application provides a content generation method, the method includes: obtaining a target text, the target text including a first sub-text with a target attribute; inputting the target text into a pre-trained target content generation model, so as to suppress the first sub-text with the target attribute through the target content generation model, and maintain text other than the target attribute, and output a target content, the target content not including content related to the first sub-text.
[0006] In a possible implementation, the target text further includes a second sub-text with a non-target attribute, and the target content includes content related to the second sub-text.
[0007] In a possible implementation, the target content includes a target image, and the image content of the target image is irrelevant to the first sub-text and relevant to the second sub-text.
[0008] In a possible implementation, the target content generation model is trained in the following manner: obtaining a first training sample and a second training sample, the first training sample including a first text sample with a target attribute, and the second training sample including a second text sample with a non-target attribute; obtaining an initial content generation model, the initial content generation model being used to output content based on the input text: training the initial content generation model using the first training sample and the second training sample until the end condition for the training of the initial content generation model to end is met, to obtain the target content generation model.
[0009] In a possible implementation, the end conditions include a first end condition and a second end condition. Training the initial content generation model using the first training sample and the second training sample until the end conditions for the end of the training of the initial content generation model are satisfied includes: training the ability of the initial content generation model to suppress text with target attributes using the first training sample until the first end condition is satisfied; training the ability of the initial content generation model to maintain text with non-target attributes using the second training sample until the second end condition is satisfied.
[0010] In a possible implementation, the method further includes: during the process of training the initial content generation model using the first training sample and the second training sample, alternately training the ability of the initial content generation model to suppress text with target attributes and the ability to maintain text with non-target attributes.
[0011] In a possible implementation, alternately training the ability of the initial content generation model to suppress text with target attributes and the ability to maintain text with non-target attributes includes: training the initial content generation model using a part of the samples in the first training sample, and when the first end condition is not satisfied, updating the model parameters of the initial content generation model to the first model parameters; training the initial content generation model using a part of the samples in the second training sample, and when the second end condition is not satisfied, updating the model parameters of the initial content generation model from the first model parameters to the second model parameters; training the initial content generation model using another part of the samples in the first training sample, and when the first end condition is not satisfied, updating the model parameters of the initial content generation model from the second model parameters to the third model parameters; training the initial content generation model using another part of the samples in the second training sample, and when the second end condition is not satisfied, updating the model parameters of the initial content generation model from the third model parameters to the fourth model parameters.
[0012] In a possible implementation, training the ability of the initial content generation model to suppress text with target attributes using the first training sample until the first end condition is satisfied includes: inputting the first text sample into the initial content generation model to obtain the first content output by the initial content generation model; determining a first target loss based on the first content; if it is determined that the first end condition is satisfied based on the first target loss, stopping the training of the ability of the initial content generation model to suppress text with target attributes; if it is determined that the first end condition is not satisfied based on the first target loss, updating the model parameters of the initial content generation model and continuing to train the initial content generation model.
[0013] In a possible implementation, the first target loss includes a first loss, a second loss, or a third loss, and the third loss is determined based on the first loss and the second loss; the first loss is expressed as the difference between a first output, a second output, and a third output, where the first output includes the output generated by the initial content generation model for the first text sample under the current model parameters, the second output includes the output generated by the initial content generation model for the first text sample under the model parameters before the current model parameters, and the third output includes the output generated by the initial content generation model for an empty text under the current model parameters; the second loss is expressed as the attention distribution of the attention matrices generated by each attention head among multiple attention heads in the initial content generation model for the target attribute.
[0014] In a possible implementation, the ability of the initial content generation model to maintain text with non-target attributes is trained using a second training sample until a second end condition is met, including: inputting the second text sample into the initial content generation model to obtain the second content output by the initial content generation model; determining a second target loss based on the second content; if it is determined that the second end condition is met based on the second target loss, stop training the ability of the initial content generation model to maintain text with non-target attributes; if it is determined that the second end condition is not met based on the second target loss, update the model parameters of the initial content generation model and continue training the initial content generation model.
[0015] In a possible implementation, the second target loss includes a fourth loss, a fifth loss, or a sixth loss, and the sixth loss is determined based on the fourth loss and the fifth loss; the fourth loss is expressed as the difference between a fourth output and a fifth output, where the fourth output includes the output generated by the initial content generation model for the second text sample under the current model parameters, and the fifth output includes the output generated by the initial content generation model for the second text sample under the original model parameters; the fifth loss is expressed as the ratio between a first product and a second product, where the first product includes the product between the feature vectors of the first text sample and the feature vectors of each second text sample, and the second product includes the product between the feature vector of the first text sample and the feature vector of a synonym or variant of the first text sample.
[0016] In a possible implementation, the initial content generation model includes a text encoding module, an image encoding module, a feature fusion module, a multi-head self-attention module, and an output module; the text encoding module is used to extract text vectors from the input text of the initial content generation model to obtain text feature vectors; the image encoding module is used to convert the input image latent variable of the initial content generation model into an image feature vector; the feature fusion module is used to fuse the text feature vector and the image feature vector to obtain a fused feature vector; the multi-head self-attention module is used to perform self-attention calculation on the fused feature vector to obtain a self-attention calculation result; the output module is used to output an output result based on the self-attention calculation result.
[0017] In a second aspect, an embodiment of the present application provides a content generation device, which includes: an acquisition module, configured to acquire a target text, where the target text includes a first sub-text of a target attribute; a content generation module, configured to input the target text into a pre-trained target content generation model, so as to suppress the first sub-text of the target attribute through the target content generation model, and maintain the text other than the target attribute, and output a target content, where the target content does not include content related to the first sub-text.
[0018] In a third aspect, an embodiment of the present application provides an electronic device, including a processor and a memory, where the memory stores machine-executable instructions that can be executed by the processor, and the processor executes the machine-executable instructions to implement the above method.
[0019] In a fourth aspect, an embodiment of the present application provides a machine-readable storage medium, which stores machine-executable instructions. When the machine-executable instructions are called and executed by a processor, the machine-executable instructions cause the processor to implement the above method.
[0020] In a fifth aspect, an embodiment of the present application provides a computer program product, which includes a computer program stored on a computer-readable storage medium. The computer program includes program instructions. When the program instructions are executed by a computer device, the computer device is caused to execute the above method.
[0021] The embodiments of the present application bring the following beneficial effects: By acquiring a target text, where the target text includes a first sub-text of a target attribute; inputting the target text into a pre-trained target content generation model, so as to suppress the first sub-text of the target attribute through the target content generation model, and maintain the text other than the target attribute, and output a target content, where the target content does not include content related to the first sub-text, in this way, when generating content from the input text, it is possible to suppress a part of the text with the target attribute in the input text, so as to purify the generated content.
[0022] Other features and advantages of the present application will be set forth in the following description, and in part will be obvious from the description, or can be learned by practice of the present application. The objectives and other advantages of the present application are realized and attained by the structure particularly pointed out in the description, claims and drawings.
[0023] In order to make the above objectives, features and advantages of the present application more obvious and understandable, the following specific preferred embodiments are given, and in conjunction with the accompanying drawings, the detailed description is as follows. Description of the Drawings
[0024] In order to more clearly illustrate the specific embodiments of the present application or the technical solutions in the prior art, the following will briefly introduce the drawings required for the description of the specific embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present application. For those skilled in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0025] Figure 1 A flowchart of a content generation method provided by an embodiment of the present application; Figure 2 A schematic illustration of a target content generation model obtaining an output image based on an input target text provided by an embodiment of the present application; Figure 3 A schematic system architecture diagram provided by an embodiment of the present application; Figure 4 A schematic illustration of alternating training provided by an embodiment of the present application; Figure 5 A schematic structural diagram of an initial content generation model provided by an embodiment of the present application; Figure 6 A schematic structural diagram of a content generation device provided by an embodiment of the present application; Figure 7 A schematic structural diagram of an electronic device provided by an embodiment of the present application. Detailed Embodiments
[0026] In order to make the objectives, technical solutions and advantages of the embodiments of the present application clearer, the following will clearly and completely describe the technical solutions of the present application in conjunction with the drawings. Obviously, the described embodiments are some, but not all, of the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative efforts fall within the scope of protection of the present application.
[0027] Currently, you can guide the model to generate content by inputting text. However, these models are usually trained based on large-scale, undifferentiated data sets, which may cause them to learn inappropriate or sensitive content. In this case, no matter what the user inputs, the content will be generated according to the input text, which can easily cause security issues. For example, the user inputs text such as "weapons", "violent content", "discriminatory content", etc. to guide the model to generate content. At this time, if the content is directly generated according to the text entered by the user, some security issues will arise.
[0028] In the related technologies, there are but are not limited to the following solutions to reduce security issues: 1. Input filtering: Relying on mechanisms such as keyword blacklists or sensitive word matching to intercept input, or deleting relevant training samples to prevent the model from learning negative concepts. However, these methods are easily bypassed, such as evading detection through homographs, variants, and inflected words, or causing the model to significantly reduce the quality of related concepts. In addition, modifications to training data will also lead to a decrease in the generalization performance of the model.
[0029] 2. Output Detection: Relying on post-processing detectors to identify and block non-compliant output content, but when the detector is bypassed or fails, it still cannot prevent the generation of harmful content. This method cannot eliminate the emergence of bad content at the model generation level.
[0030] 3. Explicit Attention Suppression: Some studies have tried to fine-tune the model internally to reduce the model's attention to sensitive concepts. For models with explicit cross-attention structures (such as models based on the "U-Net+Cross-Attention" structure), the cross-attention weights can be modified directly. However, when the model adopts an implicit cross-attention mechanism to embed text and image features into a unified transformer stream, this explicit method is no longer applicable and difficult to directly migrate. In addition, explicit attention suppression methods usually target specific layers in the model and are not universal.
[0031] Therefore, how to achieve accurate suppression of sensitive concepts under the implicit cross-attention mechanism while maintaining the model's ability to generate high-quality compliant content has become a key challenge that needs to be solved urgently. In addition, how to make it more universally applicable to different model architectures is still a problem.
[0032] Based on this, the embodiments of the present application provide a content generation method, device, electronic device, storage medium and program product, which can suppress text of target attributes (also called target concepts) and maintain text of non-target attributes to output content, thereby purifying the generated content.
[0033] It should be noted that the technical points involved in this embodiment include but are not limited to the following: 1. Selective Suppression of Joint Attention Matrix: In the absence of an explicit cross-attention layer, a joint attention matrix is generated after the fusion of text and image features. This application achieves precise control of negative concepts by selectively suppressing the column vectors related to the target concept in the joint attention matrix. Unlike traditional methods, this method directly acts on the column vectors of the attention matrix instead of the output of a certain layer, so it has stronger versatility.
[0034] 2. Collaborative Dual-Level Optimization: Pure negative suppression may lead to a decline in the quality of model generation, while pure positive retention cannot eradicate negative concepts. This application introduces a dual-layer optimization mechanism, focusing on the suppression of negative concepts at the bottom layer, and striving to retain the model's capabilities in other aspects at the upper layer, and through collaborative training, the model achieves the overall optimal performance. Unlike traditional staged training methods, traditional methods only have bottom-level suppression or top-level retention. Among them, only bottom-level suppression: while suppressing negative concepts, it will lead to a decline in the performance of common concepts, thereby damaging the overall generation ability of the model. Only top-level retention: sensitive concepts may not be completely removed, and variants of confusing negative concepts are also difficult to completely eradicate. This method introduces a dual-layer optimization mechanism, focusing on the suppression of negative concepts at the bottom layer, and striving to retain the model's capabilities in other aspects at the upper layer, and through collaborative training, the model can converge faster and achieve better results.
[0035] 3. Reverse Contrastive Learning Based on Synonyms and Variants: To prevent users from bypassing keyword detection by using synonyms, variants, or misspellings, this application introduces a reverse contrastive learning method based on synonym variants at the top level, treating various variants of negative concepts as negative samples and distinguishing them from normal content, thereby effectively avoiding bypassing. This method uses a special sampling strategy to make the model training process more robust, thereby avoiding the overfitting problem that is prone to occur in previous technologies.
[0036] 4. Generic Model Fine-Tuning via Low-Rank Parameterization Adaptation: Using the low-rank parameterization adaptation layer ( ) to fine-tune an existing model without modifying the original model backbone. This method is universal, thus supporting content purification for content generation models of different architectures (with or without cross-attention modules), and is easy to deploy and scale. This method uses a more efficient training scheme, can use fewer training resources, and can complete model fine-tuning faster.
[0037] In one embodiment of the present application, the content generation method can be run on a local terminal device or a server. When the content generation method is run on a server, the method can be implemented and executed based on a cloud interaction system, wherein the cloud interaction system includes a server and a client device (also called a local terminal device or a terminal device).
[0038] In an optional embodiment, various cloud applications and models can be run under the cloud interaction system, and the model can be deployed on the client device or on the server. Cloud applications can be, for example: cloud games. Taking cloud games as an example, cloud games refer to a game method based on cloud computing. In the operation mode of cloud games, the running body of the game program and the game screen presentation body are separated, the storage and operation of the image editing method are completed on the cloud game server, and the role of the client device is used for receiving and sending data and presenting the game screen. For example, the client device can be a display device with data transmission function close to the user side, such as a mobile terminal, a TV, a computer, a handheld computer, etc.; but the cloud game server in the cloud is used for information processing. When playing the game, the player operates the client device to send an operation instruction to the cloud game server. The cloud game server runs the game according to the operation instruction, encodes and compresses the game screen and other data, and returns it to the client device through the network. Finally, the client device decodes and outputs the game screen.
[0039] In an optional embodiment, taking a game as an example, the local terminal device stores a game program and is used to present game pictures. The local terminal device is used to interact with players through a graphical user interface, that is, conventionally, the game program is downloaded and installed through an electronic device and then run. The ways for the local terminal device to provide the graphical user interface to players can include various methods. For example, it can be rendered and displayed on the display screen of the terminal, or provided to players through holographic projection. For example, the local terminal device can include a display screen and a processor. The display screen is used to present the graphical user interface, and the graphical user interface includes game pictures. The processor is used to run the game, generate the graphical user interface, and control the display of the graphical user interface on the display screen.
[0040] To facilitate the understanding of this embodiment, first, an exemplary description of a content generation method disclosed in the embodiments of the present application will be given. The content generation method of this embodiment can be executed by a server, or by a client device, or jointly executed by a server and a client device. Please refer to Figure 1 , Figure 1 which is a schematic flowchart of a content generation method provided by the embodiments of the present application. As Figure 1 shown, the content generation method may include: S110. Obtain a target text, where the target text includes a first sub-text of a target attribute.
[0041] Among them, the target text may be text input by a user. In this embodiment, the target text is used to guide the model to generate content. The target attribute may be an attribute that needs to be suppressed, such as a violent attribute, a discriminatory attribute, etc., which can be set as needed. The target text includes a first sub-text of the target attribute, that is, this first sub-text needs to be suppressed, or it can be understood that content related to the first sub-text is not output as much as possible.
[0042] S120. Input the target text into a pre-trained target content generation model to suppress the first sub-text of the target attribute through the target content generation model, and maintain the text other than the target attribute, and output a target content that does not include content related to the first sub-text.
[0043] Among them, the target content generation model can be a pre-trained model for outputting content based on the input text. In this embodiment, suppressing the first sub-text of the target attribute can be understood as minimizing the output of content related to the first sub-text. Maintaining the text other than the target attribute can be understood as outputting content related to the text other than the target attribute. Among them, the target content can refer to the content output by the target content generation model, serving as a carrier for carrying information. In this embodiment, the target content can include but is not limited to text, images, or videos. That is to say, the target content generation model can output text based on text (for example, the input text guides the model to generate a novel), output an image based on text, or output a video based on text, which is not limited here. In this embodiment, the target content does not include content related to the first sub-text, which can be understood as the first sub-text with the target attribute being suppressed in the target content generation model.
[0044] The technical solution of this embodiment obtains a target text, where the target text includes the first sub-text of the target attribute; inputs the target text into a pre-trained target content generation model to suppress the first sub-text of the target attribute through the target content generation model, and maintain the text other than the target attribute, and outputs the target content, where the target content does not include content related to the first sub-text. In this way, when generating content from the input text, it is possible to suppress part of the text with the target attribute in the input text, thus achieving the purification of the generated content. Moreover, the target content generation model of this embodiment not only has the ability to suppress the text of the target attribute, but also has the ability to maintain (also known as retain) the text of non-target attributes, so that content can be normally generated while suppressing.
[0045] In a possible implementation, the target text further includes a second sub-text of non-target attribute, and the target content includes content related to the second sub-text.
[0046] Among them, the non-target attribute (also known as the compliance concept or non-target concept) can be understood as an attribute other than the target attribute. In this embodiment, if the target text further includes a second sub-text of non-target attribute, then the target content includes content related to the second sub-text. That is to say, the second sub-text in the target text can be maintained, so as to output content related to the second sub-text, such as outputting text, image, or video related to the second sub-text.
[0047] In a possible implementation, the target content includes a target image, and the image content of the target image has nothing to do with the first sub-text and is related to the second sub-text.
[0048] Exemplarily, assuming the target text input by the user is: "XXX (violent content) under the sun", the output image can be, for example, a scene under the sun without involving violent content.
[0049] Please refer to Figure 2 , Figure 2 for an exemplary illustration using the output content as an image. Figure 2 This is a schematic diagram of a target content generation model provided by an embodiment of the present application to obtain an output image based on an input target text. As Figure 2 shown, the target content generation model includes a text encoding module (Text Encoding Module) and a content generation module (Content Generation Module). Among them, the text encoding module extracts feature vectors from the input text to obtain a sequence of feature vectors (including word vectors and position information). Then, the content generation module converts the text feature vectors into an output image or latent variables (Output Image or Latent Variables).
[0050] It should be noted that a transducer or a diffusion model can be used to convert the text feature vectors into an output image or latent variables. The transducer or the diffusion model includes core parameters θ_base, trainable adaptation parameters Δθ, and intermediate variables such as an internal computational attention matrix.
[0051] Next, an explanation of how to train the target content generation model will be given.
[0052] In one possible implementation, the target content generation model is trained in the following manner: Obtain a first training sample and a second training sample. The first training sample includes a first text sample of a target attribute, and the second training sample includes a second text sample of a non-target attribute; obtain an initial content generation model, and the initial content generation model is used to output content based on the input text: Use the first training sample and the second training sample to train the initial content generation model until the end condition for the end of the training of the initial content generation model is met, and obtain the target content generation model.
[0053] Exemplarily, the first label and the second label can be, for example, text, image, or video, which are related to the content output by the target content generation model. For example, if the target content generation model is used to output text, the first and second labels can be, for example, text; if the target content generation model is used to output an image, the first and second labels can be, for example, an image; if the target content generation model is used to output a video, the first and second labels can be, for example, a video. In this embodiment, the initial content generation model can be a pre-trained model.
[0054] In this embodiment, since the first training sample includes the first text sample of the target attribute and the second training sample includes the second text sample of the non-target attribute, when training the initial content generation model using the first training sample and the second training sample, the initial content generation model can be made to suppress the content of the target attribute and maintain the content of the target attribute.
[0055] The technical solution of this embodiment trains the initial content generation model by using the first training sample and the second training sample, so that the trained target content generation model not only has the ability to suppress the text of the target attribute, but also has the ability to maintain (which can also be called retain) the text of the non-target attribute, so that content can be normally generated while suppressing.
[0056] In a possible implementation, training the initial content generation model using the first training sample and the second training sample until the end condition for the end of the training of the initial content generation model is satisfied includes: Training the ability of the initial content generation model to suppress the text of the target attribute using the first training sample until the first end condition is satisfied; Training the ability of the initial content generation model to maintain the text of the non-target attribute using the second training sample until the second end condition is satisfied; The end condition includes the first end condition and the second end condition.
[0057] Among them, the first end condition can refer to the condition for judging whether the training of the ability of the initial content generation model to suppress the text of the target attribute converges, or it can be understood as the condition for judging whether the ability of the initial content generation model to suppress the text of the target attribute meets the requirements. The second end condition can be the condition for judging whether the training of the ability of the initial content generation model to maintain the text of the non-target attribute converges, or it can be understood as the condition for judging whether the ability of the initial content generation model to maintain the text of the non-target attribute meets the requirements.
[0058] In this embodiment, by setting the first end condition and the second end condition to judge whether the training of the initial content generation model is completed, the ability of the target content generation model to suppress the text of the target attribute and the ability to maintain the text of the non-target attribute can both meet the requirements, improving the ability of the target content generation model to generate content.
[0059] In another possible implementation, the end condition can also include one of the first end condition or the second end condition, which can improve the training efficiency of the initial content generation model.
[0060] In a possible implementation, the method further includes: In the process of training the initial content generation model using the first training sample and the second training sample, alternately train the ability of the initial content generation model to suppress the text of the target attribute and the ability to maintain the text of the non-target attribute.
[0061] In this embodiment, alternately training the ability of the initial content generation model to suppress the text of the target attribute and the ability to maintain the text of the non-target attribute can be that after training the ability of the initial content generation model to suppress the text of the target attribute to update the model parameters, then training the ability to maintain the text of the non-target attribute to update the model parameters, and then continuing to train the ability of the initial content generation model to suppress the text of the target attribute to update the model parameters, and then training the ability to maintain the text of the non-target attribute to update the model parameters. By alternately training in this way until the first end condition and the second end condition are met, in this way, the training efficiency of the initial content generation model can be improved.
[0062] In another possible implementation, it can also be that the ability of the initial content generation model to suppress the text of the target attribute and the ability to maintain the text of the non-target attribute are trained synchronously. For example, training the ability of the initial content generation model to suppress the text of the target attribute and the ability to maintain the text of the non-target attribute, and then updating the model parameters, which can improve the relevance between the ability to suppress the text of the target attribute and the ability to maintain the text of the non-target attribute.
[0063] In one possible implementation, alternately training the ability of the initial content generation model to suppress the text of the target attribute and the ability to maintain the text of the non-target attribute includes: Using a part of the samples in the first training sample to train the initial content generation model, and when the first end condition is not met, updating the model parameters of the initial content generation model to the first model parameters; using a part of the samples in the second training sample to train the initial content generation model, and when the second end condition is not met, updating the model parameters of the initial content generation model from the first model parameters to the second model parameters; using another part of the samples in the first training sample to train the initial content generation model, and when the first end condition is not met, updating the model parameters of the initial content generation model from the second model parameters to the third model parameters; using another part of the samples in the second training sample to train the initial content generation model, and when the second end condition is not met, updating the model parameters of the initial content generation model from the third model parameters to the fourth model parameters.
[0064] Among them, model parameters are adjustable values inside the model, which are used to capture the patterns in the data. For example, in a neural network, the model parameters are weights and biases, which affect how the input data is transformed into the output result. In this embodiment, the model parameters affect how to suppress the text of the target attribute and how to maintain the text of the non-target attribute to generate content. In this embodiment, any two of the first model parameter, the second model parameter, the third model parameter, and the fourth model parameter are different model parameters. For example, the weights and / or biases of any two of the first model parameter, the second model parameter, the third model parameter, and the fourth model parameter can be different.
[0065] It should be noted that the model training in this embodiment can be, for example, full-scale training or fine-tuning training. Specifically, full-scale training can, for example, refer to the complete training of all the parameters of the model, including training from scratch (random initialization) or full-parameter adjustment based on a pre-trained model. Fine-tuning training can be based on a pre-trained model and adjust some or all of the parameters for a specific task. According to the adjustment range, it can be divided into: full-scale fine-tuning and partial fine-tuning. Among them, full-scale fine-tuning can, for example, be: unfreezing all parameters for training. Partial fine-tuning can, for example, be: only adjusting the top classification layer or adding a new adapter (such as LoRA), and freezing the underlying parameters to reduce the computational amount.
[0066] Next, the process of model update will be exemplarily described with full-scale training and fine-tuning training respectively.
[0067] First, the process of model update with full-scale training will be exemplarily described.
[0068] Assume that the initial parameters of the model are θ0. After updating the model parameters, they can be updated to the first model parameter θ1. Updating the first model parameter θ1 can obtain the second model parameter θ2. Updating the second model parameter θ2 can obtain the third model parameter θ3. Updating the third model parameter θ3 can obtain the fourth model parameter θ4. Then, if the end condition is met during the training of the fourth model parameter θ4, the fourth model parameter θ4 is used as the model parameter during model inference.
[0069] Then, the process of model update with fine-tuning training will be exemplarily described.
[0070] Assume that the initial parameters of the model are the base parameters θ0. After updating the model parameters, they can be updated to the first model parameter θ 0+ Δθ1. Updating the first model parameter θ 0+ Δθ1 can obtain the second model parameter θ 0+ Δθ 1+ Δθ2. Updating the second model parameter θ 0+ Δθ1+ The update of Δθ2 can obtain the third model parameter θ 0+ Δθ 1+ Δθ 2+ Δθ3, for the third model parameter θ 0+ Δθ 1+ Δθ 2+ The update of Δθ3 can obtain the fourth model parameter θ 0+ Δθ 1+ Δθ 2+ Δθ 3+ Δθ4. Then, if during the training of the fourth model parameter θ 0+ Δθ 1+ Δθ 2+ Δθ 3+ the end condition is satisfied during the training of Δθ4, then the fourth model parameter θ 0+ Δθ 1+ Δθ 2+ Δθ 3+ Δθ4 is used as the model parameter during model inference.
[0071] It should be noted that generally, a model includes one or more modules. Then, the update of model parameters can be to update the parameters of at least one module among one or more modules. The specific module for which the parameters are updated is related to the structure of the model. This will be further described in the following embodiments. In this embodiment, no specific limitation is imposed on the module for which the parameters are updated.
[0072] In this embodiment, if the first end condition is satisfied, the training of the initial content generation model using the first training sample is stopped. If the second end condition is satisfied, the training of the initial content generation model using the second training sample is stopped.
[0073] In a possible implementation, training the ability of the initial content generation model to suppress text with target attributes using the first training sample until the first end condition is satisfied includes: Inputting the first text sample into the initial content generation model to obtain the first content output by the initial content generation model; determining the first target loss based on the first content; if it is determined that the first end condition is satisfied based on the first target loss, then stop training the ability of the initial content generation model to suppress text with target attributes; if it is determined that the first end condition is not satisfied based on the first target loss, then update the model parameters of the initial content generation model and continue training the initial content generation model.
[0074] In this embodiment, the first target loss is used to determine whether the first end condition is satisfied. If it is determined that the first end condition is satisfied based on the first target loss, the training of the ability of the initial content generation model to suppress the text of the target attribute is stopped; if it is determined that the first end condition is not satisfied based on the first target loss, the model parameters of the initial content generation model are updated, and the training of the initial content generation model is continued.
[0075] It should be noted that the first end condition can be that the first target loss is less than the first loss threshold, or the first target loss is the smallest first target loss during the training process, which can be set as needed and is not limited here.
[0076] In a possible implementation manner, the first target loss includes a first loss, a second loss, or a third loss, and the third loss is determined based on the first loss and the second loss; The first loss is expressed as: the difference between the first output, the second output, and the third output. The first output includes the output generated by the initial content generation model for the first text sample under the current model parameters, the second output includes the output generated by the initial content generation model for the first text sample under the model parameters before the current model parameters, and the third output includes the output generated by the initial content generation model for the empty text under the current model parameters.
[0077] Exemplarily, the first loss is expressed as: ; where represents the expected value, represents under the current model parameters of the initial content generation model for the first text sample when the initial content generation model, at time step t, generates an output through the latent variable ; represents the first text sample of the target attribute to be suppressed; represents the negative concept suppression coefficient, which is used to adjust the intensity of negative concept suppression, and the value can be set as needed, and the default value range is [0.1, 1.5].
[0078] represents the output generated by the initial content generation model when inputting the first text sample through the latent variable ; represents the output generated by the initial content generation model at time step t when inputting the empty text through the latent variable , which is equivalent to unconditional generation. represents the square of the Euclidean norm of the vector; The second loss is expressed as: the attention distribution of the attention matrix generated by each attention head among multiple attention heads in the initial content generation model to the target attribute.
[0079] Exemplarily, the second loss is expressed as: The second loss is expressed as: ; where is the number of attention heads in the initial content generation model; is the attention matrix corresponding to the h-th attention head; represents the i-th column vector of the attention matrix in the h-th attention head. The column vector represents the attention distribution of the image features of the image to the target attribute, represents the 1-norm of the vector, and the word vector position index is where and are the start position index and end position index of the negative concept word in the sequence in the text.
[0080] In this embodiment, determining the third loss based on the first loss and the second loss may be to perform weighted calculation on the first loss and the second loss, and the result of the weighted calculation is used as the third loss.
[0081] In a possible implementation manner, using the second training sample to train the ability of the initial content generation model to maintain text with non-target attributes until the second end condition is met includes: Inputting the second text sample into the initial content generation model to obtain the second content output by the initial content generation model; determining the second target loss based on the second content; if it is determined that the second end condition is met based on the second target loss, stop training the ability of the initial content generation model to maintain text with non-target attributes; if it is determined that the second end condition is not met based on the second target loss, update the model parameters of the initial content generation model and continue to train the initial content generation model.
[0082] In a possible implementation manner, the second target loss includes a fourth loss, a fifth loss, or a sixth loss, and the sixth loss is determined based on the fourth loss and the fifth loss; The fourth loss is expressed as: the difference between the fourth output and the fifth output. The fourth output includes the output generated by the initial content generation model for the second text sample under the current model parameters, and the fifth output includes the output generated by the initial content generation model for the second text sample under the original model parameters.
[0083] Exemplarily, the fourth loss is expressed as: ; Represents the current model parameters of the initial content generation model Generated, using the second text sample c as the guiding text, at time step t, through the latent variable The generated output; Represents the feature vector of the image generated under the original parameters of the initial content generation model, with the same random seed and input text guidance.
[0084] The fifth loss is expressed as: the ratio between the first product and the second product. The first product includes the product between the feature vector of the first text sample and the feature vectors of each second text sample, and the second product includes the product between the feature vector of the first text sample and the feature vector of a synonym or variant of the first text sample.
[0085] Exemplarily, the fifth loss is expressed as: ; Represents the feature vector of the text of the target attribute (such as "weapon"); Represents the feature vector of a synonym or variant of the target attribute (such as "weap0n", "gun"), Represents the feature vector of the text of the i-th non-target attribute (such as "cat", "flower"); K represents the number of texts of non-target attributes (the recommended value range is [3, 10]), Is an adjustment parameter used to adjust the intensity of contrast learning, and the recommended value range is [0.05, 0.2].
[0086] For ease of understanding, the co-training of the embodiments of the present application will be described below. Please refer to Figure 3 , Figure 3 Is a schematic diagram of a system architecture provided by the embodiments of the present application. As Figure 3 Shown, taking the input text to guide the model to output an image as an example for illustration.
[0087] As Figure 3As shown in the figure, first, the text can be input into the text-guided image synthesis model, that is, the content generation model, and then an image can be output. Then, perform bottom-layer optimization (Suppression Level), that is, the optimization of the suppression ability, including but not limited to optimization through the negative guidance loss and the selective suppression loss of the column vector. The negative guidance loss can refer to the description of the first loss, and the selective suppression loss of the column vector can refer to the description of the second loss. Then, top-layer optimization (that is, the optimization of the maintenance ability) can also be performed, including but not limited to the content fidelity loss and the reverse contrast loss based on the synonym variants. Then, the optimal adaptation parameter Δθ (Generate Optimal Adaptation Parameter Δθ*) can be generated, and the adaptation parameter Δθ* is fused with the original model parameter θ_base for image synthesis (Fuse Adaptation Parameter Δθ* with Original Model Parameter θ_base for Image Synthesis).
[0088] In this embodiment, the goal of the bottom-layer optimization: through the negative guidance loss and the attention suppression loss, when the model inputs text containing negative concepts, its output result approaches the image generated unconditionally or guided by empty text, thereby weakening the model's expression ability for negative concepts.
[0089] Attention Suppression Loss: For a model using the implicit cross-attention mechanism, after the text features and image features are fused, a joint attention matrix will be obtained through the multi-head attention mechanism , where represents the sequence length of the joint features. If there are negative words in the input text, the word vector position index is , where and are the start position index and end position index of the negative concept word in the sequence in the text, then the column vector corresponding to this attention matrix can be selectively suppressed: .
[0090] The attention matrix here can be one or more attention matrices selected from multiple attention modules inside the model, and specific selection needs to be based on the actual model. In actual use, it is recommended to select the attention matrix closer to the output side.
[0091] The underlying optimization objective is expressed as:
[0092] Among them, is the weight for adjusting the attention suppression loss, and the default value range is [0.1, 1].
[0093] In this embodiment, the objective of the top-level optimization (Preservation Level) is to ensure that while suppressing negative concepts, the model's generation ability for normal content (such as "cat", "flower") is not affected through the Content Fidelity Loss and the Reverse Contrastive Loss.
[0094] The top-level optimization objective is expressed as:
[0095] Among them, is the weight for adjusting the content fidelity loss, and the default value range is [0.5, 2].
[0096] Please refer to Figure 4 , Figure 4 which is a schematic diagram of an alternative training provided by the embodiment of the present application. As Figure 4 shown, first, perform several iterations on the underlying optimization, and then perform several iterations on the top-level optimization, alternating. This iterative method can make the model converge faster and have better results.
[0097] Next, the model architecture of one of the initial content generation models will be described.
[0098] Please refer to Figure 5 , Figure 5 which is a schematic diagram of the structure of an initial content generation model provided by the embodiment of the present application. In one possible implementation, the initial content generation model includes a text encoding module, an image encoding module, a feature fusion module, a multi-head self-attention module, and an output module.
[0099] The text encoding module is used to extract text vectors from the input text of the initial content generation model to obtain text feature vectors. The image encoding module is used to convert the input image latent variables of the initial content generation model into image feature vectors. The feature fusion module is used to fuse the text feature vectors and the image feature vectors to obtain the fused feature vectors. The multi-head self-attention module is used to perform self-attention calculations on the fused feature vectors to obtain the self-attention calculation results. The output module is used to output the output results based on the self-attention calculation results.
[0100] Specifically, Text Input: The text prompt input by the user, such as "a photo of a weapon". Text Encoder: Converts the text prompt into a series of token embeddings. Pretrained text encoders such as BERT, GPT, T5, etc. can be used. Input: Text prompt; Output: Sequence of text feature vectors.
[0101] Specifically, Image Latents: The latent representation of the image, which is effective for the Encoder. The Decoder may not require this input. Image Encoder: Converts the image latents into image feature vectors, which is an optional component. Input: Image latents; Output: Sequence of image feature vectors.
[0102] Specifically, Feature Fusion: Fuses the text feature vectors and the image feature vectors together. Common fusion methods include: Concatenation: Directly concatenates the sequence of text feature vectors and the sequence of image feature vectors. Addition: Element-wise adds the text feature vectors and the image feature vectors. Input: Text features, image features (optional); Output: Sequence of fused feature vectors.
[0103] Specifically, Multi-Head Self-Attention: Performs self-attention calculation on the sequence of fused feature vectors. Q, K, V: Each attention head maps the sequence of fused feature vectors into Query (Q), Key (K), and Value (V) vectors respectively. Linear: A linear transformation layer used to map the input vectors into Q, K, V vectors. A1, ..., AH: Attention matrices calculated by each attention head. Input: Sequence of fused feature vectors; Output: Sequence of feature vectors after self-attention calculation.
[0104] Specifically, Residual Connection & Layer Normalization: Performs residual connection and layer normalization on the output of each attention head to enhance the stability and training efficiency of the model.
[0105] Specifically, Feed-Forward Network: Performs a non-linear transformation on the output of each attention head to enhance the expressive power of the model.
[0106] Specifically, Residual Connection & Layer Normalization: Perform residual connection and layer normalization on the output of the feedforward network.
[0107] Specifically, Output Module: If it is a diffusion model, according to the model type, it can be: noise prediction, denoised image prediction, velocity field prediction. If it is other generative models, the output may be different according to the model type. Input: The feature vector processed by the multi-layer self-attention module, Output: The generated image or latent variable.
[0108] It should be noted that during training, the image latent variable input to the image encoding module can be the latent variable of the first image (also known as the third content) including the relevant content of the first text sample or the second image (also known as the fourth content) including the relevant content of the second text sample.
[0109] During model inference, the image latent variable input to the image encoding module can be the latent variable of the intermediate image (also known as the fifth content) generated based on the target text and including the relevant content of all sub-texts of the target text. For example, the intermediate image includes the image content related to the first sub-text with the target attribute in the target text. The output module outputs the target image that does not include the image content related to the first sub-text.
[0110] It should be noted that in this embodiment, when training the initial content generation model, the parameters of the multi-head self-attention module can be updated. Among them, the way to update the parameters can refer to the description of updating the model parameters in the above embodiment, such as referring to the description of the first model parameter, the second model parameter, the third model parameter, and the fourth model parameter, which will not be elaborated here.
[0111] Generally speaking, this application proposes a method and system for purifying the output content of a text-guided image synthesis model through a novel dual-level collaborative optimization strategy, suppressing or removing preset negative concepts (such as "weapons", "violent content", "discriminatory content", etc.). This method is particularly applicable to single-stream or multi-stream transformer-based architectures that adopt an implicit cross-attention mechanism, achieving precise suppression of specified concepts while maintaining the model's generation ability for other compliant concepts, and effectively improving the quality and safety of the output content. This method obtains a purification effect similar to traditional methods in models without an explicit cross-attention layer, and even outperforms traditional methods in some metrics, by selectively suppressing column vectors of the attention matrix and combining with reverse contrast learning.
[0112] Specifically, it involves the following processes: Training and inference processes: 1. Variant collection: In the model training preparation stage, synonyms and common variants of the target negative concept will be collected, such as "weap0n", "weapon-1", "wepon", etc. These variants will be trained together with the negative concept to improve the model's recognition ability and robustness for the negative concept. Word variants can be obtained through edit distance calculation or generated using large language models.
[0113] 2. Multi-target suppression: If multiple target concepts need to be suppressed simultaneously, such as "weapons", "violent content", etc., the method of this application can be achieved by superimposing the corresponding loss functions. In the underlying optimization stage, the negative guidance loss and attention suppression loss of all target concepts to be suppressed can be calculated separately, and these losses are weighted and summed.
[0114] 3. Hyperparameter setting: The hyperparameters in the method of this application can be adjusted according to the actual situation, and the default value range of the parameters is: Negative guidance coefficient : [0.1, 1.5]; Attention suppression loss weight : [0.1, 1]; Content fidelity loss weight : [0.5, 2]; Reverse contrast loss temperature coefficient : [0.05, 0.2]; These parameters can be adjusted using hyperparameter search methods such as Grid Search to achieve better results.
[0115] 4. Inference stage: In the inference stage, only the core parameters and adaptation parameters need to be loaded. When the text input by the user contains negative concepts (such as "weapon" and its variants), the images generated by the model will be significantly weakened or directly removed.
[0116] Generally speaking, the technical effects and advantages of this embodiment include but are not limited to the following: 1. Balancing negative inhibition and normal content retention: The method proposed in this application can effectively inhibit the model from generating negative content while retaining the ability of the model to generate normal content, avoiding over-inhibition.
[0117] 2. Capable of preventing various bypass strategies: This method uses the reverse contrast learning strategy, which can use negative concepts and their synonyms and variants as negative samples for training, thus preventing the model from being bypassed by users using synonyms or deformed words.
[0118] 3. Wide applicability: By introducing a low-rank parameterized adaptation module, this method can be applied to various text-guided image generation models without modifying the structure of the model.
[0119] 4. Applicable to implicit cross-attention mechanism: The method of this application can be directly applied to models using the implicit cross-attention mechanism without relying on an explicit cross-attention layer.
[0120] 5. Higher training efficiency: This application adopts collaborative training, which can make the model converge faster and can use fewer training resources.
[0121] 6. Stronger practical application value: This application can optimize the model to improve the quality and security of the output images, with high practical application value.
[0122] Please refer to Figure 6 , Figure 6 which is the structural schematic diagram of a content generation device provided by an embodiment of this application. As Figure 6 shown, the content generation device includes: An acquisition module 610 for acquiring a target text, where the target text includes a first sub-text of a target attribute; a content generation module 620 for inputting the target text into a pre-trained target content generation model, so as to suppress the first sub-text of the target attribute through the target content generation model, and maintain the text other than the target attribute, and output a target content, where the target content does not include content related to the first sub-text.
[0123] For the device in this embodiment, reference may be made to the description of the above method embodiment, which will not be elaborated here.
[0124] This embodiment also provides an electronic device, including a processor and a memory. The memory stores machine-executable instructions that can be executed by the processor, and the processor executes the machine-executable instructions to implement the above image editing method. The electronic device may be a server or a terminal device.
[0125] See Figure 7 As shown, the electronic device includes a processor 100 and a memory 101. The memory 101 stores machine-executable instructions that can be executed by the processor 100, and the processor 100 executes the machine-executable instructions to implement the steps of the above method.
[0126] Furthermore, Figure 7 the electronic device shown further includes a bus 102 and a communication interface 103. The processor 100, the communication interface 103, and the memory 101 are connected through the bus 102.
[0127] Among them, the memory 101 may include a high-speed random access memory (RAM), and may also include a non-volatile memory, such as at least one disk memory. Through at least one communication interface 103 (which may be wired or wireless), a communication connection is established between this system network element and at least one other network element, and the Internet, wide area network, local area network, metropolitan area network, etc. can be used. The bus 102 may be an ISA bus, a PCI bus, or an EISA bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of simplicity of representation, Figure 7 only a bidirectional arrow is used in the figure, but it does not mean that there is only one bus or one type of bus.
[0128] The processor 100 may be an integrated circuit chip with signal processing capabilities. In the implementation process, each step of the above method can be completed by the integrated logic circuit of the hardware in the processor 100 or the instructions in the form of software. The above-mentioned processor 100 may be a general-purpose processor, including a central processing unit (CPU for short), a network processor (NP for short), etc.; it may also be a digital signal processor (DSP for short), an application specific integrated circuit (ASIC for short), a field-programmable gate array (FPGA for short), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. It can implement or execute the various methods, steps and logic block diagrams disclosed in the embodiments of the present application. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. The steps of the method disclosed in combination with the embodiments of the present application can be directly embodied as being executed and completed by the hardware decoding processor, or executed and completed by the combination of the hardware and software modules in the decoding processor. The software module may be located in a mature storage medium in the art such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory, or an electrically erasable programmable memory, a register, etc. This storage medium is located in the memory 101, and the processor 100 reads the information in the memory 101 and combines its hardware to complete the steps of the method in the foregoing embodiments.
[0129] This embodiment also provides a machine-readable storage medium. The machine-readable storage medium stores machine-executable instructions. When the machine-executable instructions are called and executed by the processor, the machine-executable instructions cause the processor to implement the steps of the above method.
[0130] This embodiment also provides a computer program product. The computer program product includes a computer program stored on a computer-readable storage medium. The computer program includes program instructions. When the program instructions are executed by a computer device, the computer device is caused to execute the steps of the above method.
[0131] Those skilled in the art can clearly understand that for the convenience and simplicity of description, the specific working processes of the systems and devices described above can refer to the corresponding processes in the foregoing method embodiments, and will not be elaborated herein.
[0132] In addition, in the description of the embodiments of the present application, unless otherwise clearly specified and limited, the terms "installation", "connection", and "coupling" should be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be directly connected or indirectly connected through an intermediate medium, and it can be the communication inside two components. For those skilled in the art, the specific meanings of the above terms in the present application can be understood according to specific situations.
[0133] If a function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art or a part of the technical solution can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods of the various embodiments of the present application. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs that can store program codes.
[0134] In the description of the present application, it should be noted that the orientation or positional relationship indicated by the terms "center", "upper", "lower", "left", "right", "vertical", "horizontal", "inner", "outer", etc. is based on the orientation or positional relationship shown in the drawings. It is only for the convenience of describing the present application and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and thus should not be construed as a limitation of the present application. In addition, the terms "first", "second", and "third" are only used for descriptive purposes and cannot be understood as indicating or implying relative importance.
[0135] Finally, it should be noted that the above embodiments are only specific embodiments of the present application, used to illustrate the technical solutions of the present application, rather than limiting them. The protection scope of the present application is not limited thereto. Although the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that: any person skilled in the art within the technical scope disclosed by the present application can still modify the technical solutions recorded in the foregoing embodiments, or can easily think of changes, or perform equivalent replacements on some of the technical features; and these modifications, changes, or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application, and should all be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.
Claims
1. A content generation method, characterized in that, The method includes: Obtaining a target text, where the target text includes a first sub-text of a target attribute; Inputting the target text into a pre-trained target content generation model to suppress the first sub-text of the target attribute through the target content generation model, and maintain the text other than the target attribute, and outputting a target content, where the target content does not include content related to the first sub-text.
2. The method according to claim 1, characterized in that, The target text further includes a second sub-text of a non-target attribute, and the target content includes content related to the second sub-text.
3. The method according to claim 2, wherein The target content includes a target image, and the image content of the target image is irrelevant to the first sub-text and relevant to the second sub-text.
4. The method according to any one of claims 1-3, characterized in that, The target content generation model is trained in the following manner: Obtaining a first training sample and a second training sample, where the first training sample includes a first text sample of a target attribute, and the second training sample includes a second text sample of a non-target attribute; Obtaining an initial content generation model, where the initial content generation model is used to output content based on the input text: Training the initial content generation model using the first training sample and the second training sample until an end condition for the end of the training of the initial content generation model is satisfied, to obtain the target content generation model.
5. The method according to claim 4, characterized in that, Wherein, The end condition includes a first end condition and a second end condition. Training the initial content generation model using the first training sample and the second training sample until the end condition for the end of the training of the initial content generation model is satisfied includes: Training the ability of the initial content generation model to suppress the text of the target attribute using the first training sample until the first end condition is satisfied; Training the ability of the initial content generation model to maintain the text of the non-target attribute using the second training sample until the second end condition is satisfied.
6. The method according to claim 5, characterized in that, The method further includes: During the process of training the initial content generation model using the first training sample and the second training sample, alternately training the ability of the initial content generation model to suppress the text of the target attribute and the ability to maintain the text of the non-target attribute.
7. The method according to claim 6, wherein The alternately training the ability of the initial content generation model to suppress the text of the target attribute and the ability to maintain the text of the non-target attribute includes: Training the initial content generation model using a part of the samples in the first training sample, and updating the model parameters of the initial content generation model to first model parameters when the first end condition is not satisfied; Training the initial content generation model using a part of the samples in the second training sample, and updating the model parameters of the initial content generation model from the first model parameters to second model parameters when the second end condition is not satisfied; Training the initial content generation model using another part of the samples in the first training sample, and updating the model parameters of the initial content generation model from the second model parameters to third model parameters when the first end condition is not satisfied; Use another part of the samples in the second training sample to train the initial content generation model, and update the model parameters of the initial content generation model from the third model parameters to the fourth model parameters when the second end condition is not satisfied.
8. The method according to any one of claims 5-7, characterized in that, The training of the ability of the initial content generation model to suppress text with target attributes using the first training sample until the first end condition is met includes: Input the first text sample into the initial content generation model to obtain the first content output by the initial content generation model; Determine the first target loss based on the first content; If it is determined that the first end condition is met based on the first target loss, stop training the ability of the initial content generation model to suppress text with target attributes; If it is determined that the first end condition is not met based on the first target loss, update the model parameters of the initial content generation model and continue to train the initial content generation model.
9. The method according to claim 8, wherein The first target loss includes the first loss, the second loss, or the third loss, and the third loss is determined based on the first loss and the second loss; The first loss is expressed as the difference between the first output, the second output, and the third output. The first output includes the output generated by the initial content generation model for the first text sample under the current model parameters, the second output includes the output generated by the initial content generation model for the first text sample under the model parameters before the current model parameters, and the third output includes the output generated by the initial content generation model for an empty text under the current model parameters; The second loss is expressed as the attention distribution of the attention matrix generated by each attention head among multiple attention heads in the initial content generation model for the target attribute.
10. The method according to any one of claims 5-7, characterized in that, The training of the ability of the initial content generation model to maintain text with non-target attributes using the second training sample until the second end condition is met includes: Input the second text sample into the initial content generation model to obtain the second content output by the initial content generation model; Determine the second target loss based on the second content; If it is determined that the second end condition is met based on the second target loss, stop training the ability of the initial content generation model to maintain text with non-target attributes; If it is determined that the second end condition is not met based on the second target loss, update the model parameters of the initial content generation model and continue to train the initial content generation model.
11. The method according to claim 10, wherein The second target loss includes the fourth loss, the fifth loss, or the sixth loss, and the sixth loss is determined based on the fourth loss and the fifth loss; The fourth loss is expressed as the difference between the fourth output and the fifth output. The fourth output includes the output generated by the initial content generation model for the second text sample under the current model parameters, and the fifth output includes the output generated by the initial content generation model for the second text sample under the original model parameters; The fifth loss is expressed as the ratio between a first product and a second product. The first product includes the products between the feature vectors of the first text sample and the feature vectors of each second text sample. The second product includes the product between the feature vector of the first text sample and the feature vector of a synonym or variant of the first text sample.
12. The method according to claim 9 or 11, characterized in that, The initial content generation model includes a text encoding module, an image encoding module, a feature fusion module, a multi-head self-attention module, and an output module; The text encoding module is configured to extract text vectors from the text input to the initial content generation model to obtain text feature vectors; The image encoding module is configured to convert the image latent variable input to the initial content generation model into an image feature vector; The feature fusion module is configured to fuse the text feature vector and the image feature vector to obtain a fused feature vector; The multi-head self-attention module is configured to perform self-attention calculation on the fused feature vector to obtain a self-attention calculation result; The output module is configured to output an output result based on the self-attention calculation result.
13. A content generation device, characterized in that, The device includes: An acquisition module, configured to acquire a target text, where the target text includes a first sub-text of a target attribute; A content generation module, configured to input the target text into a pre-trained target content generation model, so as to suppress the first sub-text of the target attribute through the target content generation model, maintain the text other than the target attribute, and output a target content that does not include content related to the first sub-text.
14. An electronic device, characterized in that, It includes a processor and a memory. The memory stores machine-executable instructions that can be executed by the processor. The processor executes the machine-executable instructions to implement the content generation method according to any one of claims 1-12.
15. A machine-readable storage medium, characterized in that, The machine-readable storage medium stores machine-executable instructions. When the machine-executable instructions are called and executed by a processor, the machine-executable instructions cause the processor to implement the content generation method according to any one of claims 1-12.
16. A computer program product, characterized in that, The computer program product includes a computer program stored on a computer-readable storage medium. The computer program includes program instructions. When the program instructions are executed by a computer device, the computer device is caused to execute the content generation method according to any one of claims 1-12.
Citation Information
Patent Citations
Generating objects of mixed concepts using text-to-image diffusion models
US20240144544A1