A text-based image editing method and an electronic device

By introducing a sampling and encoding module in ManiGAN, the intermediate editing results are output to the user to control the editing process, the problem that the ManiGAN output results do not meet user requirements is solved, and an image editing effect that is more in line with user needs is achieved.

CN114092758BActive Publication Date: 2025-07-01SHENZHEN INST OF ADVANCED TECH CHINESE ACAD OF SCI
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111188395.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-10-12
Publication Date
2025-07-01
Estimated Expiration
2041-10-12

AI Technical Summary

Technical Problem

When existing ManiGANs edit source images based on input text, the output image editing results often do not meet user requirements.

Method used

The sampling and encoding module is introduced in the existing ManiGAN to form an improved ManiGAN, which outputs the intermediate editing results to the user so that the user can judge and control the editing process, thereby eliminating intermediate results that do not meet the requirements.

Benefits of technology

By introducing the sampling and encoding module, the improved ManiGAN can better control the intermediate editing results, ensuring that the final output image editing results are more in line with user requirements.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114092758B_ABST
    Figure CN114092758B_ABST
Patent Text Reader

Abstract

The present application relates to the field of image processing, and provides a text-based image editing method and an electronic device. The method includes: obtaining the overall image feature and the local image feature of a target source image, as well as the overall sentence feature and the word feature of a target text; editing the target source image based on the overall image feature, the local image feature, the overall sentence feature and the word feature by means of an image editing model to obtain a target edited image; wherein, the image editing model includes: a sampling and encoding module and at least one cascaded generation module. The encoding module can output an intermediate editing result in the image editing process. If the intermediate editing result does not meet the user's requirements, the image editing model can adjust the intermediate editing result and input the adjusted intermediate result into at least one cascaded generation module, thereby solving the problem that the image editing result output by ManiGAN does not meet the user's requirements.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of image processing, and in particular, to a text-based image editing method and an electronic device. Background Art

[0002] As is well known, text-based image editing is a technology for editing a source image according to a given text, which is a research hotspot in the multimedia field and has important application value. The existing Manipulating Attention Generative Adversarial Network (ManiGAN) is used to perform image editing on a source image to be edited according to the content of a text description. However, when ManiGAN edits a source image according to an input text, the output image editing result often does not meet the user's requirements.

[0003] Therefore, how to make the image editing result output by ManiGAN meet the user's requirements is an urgent problem to be solved currently. Summary of the Invention

[0004] The present application provides a text-based image editing method and an electronic device, which can solve the problem that the image editing result output by ManiGAN does not meet the user's requirements.

[0005] In a first aspect, a text-based image editing method is provided, including: obtaining an overall image feature and a local image feature of a target source image, and an overall sentence feature and a sentence word feature of a target text; editing the target source image based on an image editing model according to the overall image feature, the local image feature, the overall sentence feature, and the sentence word feature to obtain a target edited image; where the image editing model includes: a sampling encoding module and at least one cascaded generation module; the process of the image editing model for processing the target source image includes: using the sampling encoding module to perform sampling encoding processing on the overall image feature, the overall sentence feature, and the local image feature to obtain a first edited image, and outputting the first edited image; in response to a user instruction, inputting the first edited image, the local image feature, and the sentence word feature into the at least one cascaded generation module to perform high-dimensional visual feature extraction to obtain the target edited image, or inputting the target source image, the local image feature, and the sentence word feature into the at least one cascaded generation module to perform high-dimensional visual feature extraction to obtain the target edited image.

[0006] The above method can be executed by a chip on an electronic device. Compared with the existing ManiGAN that edits the source image to be edited according to the text and directly outputs an editing result that may not meet the user's requirements, the present application introduces a sampling encoding module into the existing ManiGAN to form an improved ManiGAN; the sampling encoding module will output the intermediate editing result (i.e., the first edited image) to facilitate the user to judge whether the intermediate editing result meets the requirements. If it meets the requirements, the intermediate editing result will be passed to at least one cascaded generation module; if it does not meet the requirements, the intermediate result will not be passed to at least one cascaded generation module, but the target source image will be used to replace the intermediate editing result and continue to be passed to at least one cascaded generation module. It can be seen that when the improved ManiGAN edits the target source image according to the target text, it can control the intermediate editing result and timely eliminate the intermediate editing results that do not meet the requirements, so as to prevent the inaccurate output results of the previous level from affecting the accuracy of the output results of the subsequent level, thereby editing a target edited image that better meets the requirements for the user.

[0007] Optionally, the generation module includes: a first auto-decoder, a first self-attention module, a second upsampling module, and a second auto-encoder. The first auto-decoder is used to restore the high-dimensional visual features of the input information to obtain a first high-dimensional feature image, and the input information is the first edited image or the output information of the previous layer generation module; the first self-attention module is used to fuse and process the first high-dimensional feature image and the sentence word features to obtain sentence semantic information features; the second upsampling module is used to perform feature fusion and upsampling processing on the sentence semantic information features to obtain a second upsampling result; the second auto-encoder is used to extract high-dimensional visual features from the second upsampling result to obtain output information. When the generation module is the last-level generation module in the at least one cascaded generation module, the output information is the target edited image.

[0008] Optionally, the first self-attention module includes: a self-attention layer and a first noise-affine combination module. The self-attention layer is used to fuse the first high-dimensional feature image and the sentence word features; the first noise-affine combination module is used to fuse the concatenation result of the output result of the self-attention layer and the first high-dimensional feature image, and the local image features.

[0009] By introducing the first noise-affine combination module into the above first self-attention module, the first noise-affine combination module can enhance the reliability of the generated image of the generation module by introducing Gaussian noise, thus avoiding the situation where the reliability of the editing result is affected by random noise in the image in the generation module.

[0010] Optionally, the sampling and encoding module includes: a first upsampling module and a first autoencoder. The first upsampling module is configured to perform upsampling processing on the overall image feature, the overall sentence feature, and the local image feature to obtain a first upsampling result. The first autoencoder is configured to generate a first edited image according to the first upsampling result.

[0011] Optionally, the first upsampling module includes: a plurality of identical upsampling layers, a second noise-affine combination module, and a third noise-affine combination module. The inputs of the first upsampling module are the overall sentence feature, the overall image feature, and the local image feature. In the plurality of identical upsampling layers, for two adjacent upsampling layers, the input of the latter upsampling layer is the output of the former upsampling layer. The second noise-affine combination module is located between any two upsampling layers in the plurality of identical upsampling layers and is configured to perform feature fusion on the result output by the former upsampling layer and the local image feature among the any two upsampling layers. The third noise-affine combination module is configured to perform feature fusion on the output result of the last upsampling layer in the plurality of identical upsampling layers and the local image feature.

[0012] Introducing the second noise-affine combination module and the third noise-affine combination module in the first upsampling module can further enhance the visual features of the output results of different upsampling layers in the first upsampling module.

[0013] Optionally, the image editing model further includes: a detail correction model for performing detail modification on the target edited image. The detail correction model is configured to process the local image feature, the sentence word feature, and the target edited image to obtain a target corrected image. The detail correction model includes: a first detail correction module, a second detail correction module, a fusion module, and a generator. The first detail correction module is configured to perform detail modification on the local image feature, the first random noise, and the sentence word feature to obtain a first detail feature. The second detail correction module is configured to perform detail modification on the local image feature corresponding to the target edited image, the second random noise, and the sentence word feature to obtain a second detail feature. The fusion module is configured to perform feature fusion on the first detail feature and the second detail feature. The generator is configured to generate the target corrected image according to the output result of the fusion module.

[0014] Adding the detail correction model to the above image editing model can further modify and enhance the details of the target edited image output by the image editing model, so as to obtain a high-resolution target corrected image.

[0015] Optionally, the first detail correction module includes a fourth noise-affine combination module, a fifth noise-affine combination module, a sixth noise-affine combination module, a second self-attention module, a first residual network, and a first linear network; the fourth noise-affine combination module is used to perform feature fusion on the first random noise and the local image features to obtain first fusion features; the second self-attention module is used to perform feature fusion on the first fusion features and the sentence word features; the fifth noise-affine combination module performs feature fusion on the concatenation result of the output result of the second self-attention module and the first random noise, and the local image features; the first residual network is used to extract visual features from the output result of the fifth noise-affine combination module; the first linear network is used to perform a linear transformation on the local image features; the sixth noise-affine combination module is used to perform feature fusion on the output result of the first residual network and the output result of the first linear network.

[0016] Adding multiple noise-affine combination modules to the above first detail correction module can enhance the reliability of the detail correction model.

[0017] Optionally, the second detail correction module includes a seventh noise-affine combination module, an eighth noise-affine combination module, a ninth noise-affine combination module, a third self-attention module, a second residual network, and a second linear network; the seventh noise-affine combination module is used to perform feature fusion on the second random noise and the local image features corresponding to the target edited image to obtain first fusion features; the third self-attention module is used to perform feature fusion on the first fusion features and the sentence word features; the eighth noise-affine combination module performs feature fusion on the concatenation result of the output result of the third self-attention module and the second random noise, and the local image features corresponding to the target edited image; the second residual network is used to extract visual features from the output result of the eighth noise-affine combination module; the second linear network is used to perform a linear transformation on the local image features corresponding to the target edited image; the ninth noise-affine combination module is used to perform feature fusion on the output result of the second residual network and the output result of the second linear network.

[0018] Adding multiple noise-affine combination modules to the above second detail correction module can enhance the reliability of the detail correction model.

[0019] Optionally, the training method of the detail correction model includes: training the generator of the detail correction model according to a conditional generator loss function, an unconditional generator loss function, and a semantic contrast function; training the discriminator of the detail correction model according to a conditional discriminator loss function and an unconditional discriminator loss function.

[0020] Training the generator of the detailed correction model according to the conditional generator loss function, unconditional generator loss function, and semantic contrast function can make the image editing result (i.e., the target edited image) generated by the generator more conform to the content described in the target text and the user's requirements. Training the discriminator of the detailed correction model according to the conditional discriminator loss function and unconditional discriminator loss function can make the recognition result of the discriminator more accurate.

[0021] Optionally, the training method of the image editing model includes: training an initial model using a preset loss function and a training set to obtain the image editing model; wherein, the preset loss function includes sub-functions corresponding to N sub-networks respectively and loss functions of N-1 autoencoders, the initial model includes N sub-networks, and the N sub-networks are initial models corresponding to the sampling encoding module and at least one generation module respectively; during the training process, when the output image of the i-th sub-network does not meet the preset conditions, the sub-functions corresponding to the i+1-th to N-th sub-networks and the loss functions of the i-th to i+1-th autoencoders are used to train the initial model, where 0 ≤ i < N.

[0022] During the training of the above image editing model, training the N sub-networks is to skip the previous-stage sub-networks whose output results do not meet the requirements through the autoencoder and give priority to training the subsequent-stage sub-networks. Since the target source image rather than the intermediate editing result (such as the first edited image) is used during the skipping, it is possible to avoid the propagation of incorrect results output by the potential previous-stage sub-networks to the subsequent-stage sub-networks. Prioritizing the training of the subsequent-stage sub-networks can bring better update gradients to the previous-stage sub-networks, thereby making the convergence effect of the previous-stage sub-networks better.

[0023] In a second aspect, an electronic device is provided, including a module for executing any one of the methods in the first aspect.

[0024] In a third aspect, a computer-readable storage medium is provided. The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the processor is caused to execute the method described in any one of the first aspect. BRIEF DESCRIPTION OF THE DRAWINGS

[0025] To more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for use in the embodiments or the description of the prior art. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0026] Figure 1Schematic flowchart of an image editing method based on text in an embodiment of the present invention;

[0027] Figure 2 Schematic structural diagram of an image editing model in an embodiment of the present invention;

[0028] Figure 3 Schematic structural diagram of a detail correction model in an embodiment of the present invention;

[0029] Figure 4 Schematic diagram of the processing process of an image editing model for a target source image in an embodiment of the present invention;

[0030] Figure 5 Schematic structural diagram of an electronic device in an embodiment of the present invention. Detailed implementation manners

[0031] In the following description, for the purpose of illustration rather than limitation, specific details such as specific system structures and technologies are proposed to thoroughly understand the embodiments of the present application. However, those skilled in the art should clearly understand that the present application can also be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, and methods are omitted to avoid unnecessary details from interfering with the description of the present application.

[0032] It should be understood that when used in the specification of the present application and the appended claims, the term "comprising" indicates the presence of the described features, wholes, steps, operations, elements, and / or components, but does not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components, and / or their combinations.

[0033] It should also be understood that the term "and / or" as used in the specification of the present application and the appended claims refers to any combination and all possible combinations of one or more of the associated listed items, and includes these combinations.

[0034] In addition, in the description of the specification of the present application and the appended claims, the terms "first", "second", "third", etc. are only used for distinguishing descriptions and cannot be understood as indicating or implying relative importance.

[0035] References to "one embodiment" or "some embodiments" etc. described in the specification of this application mean that specific features, structures, or characteristics described in connection with that embodiment are included in one or more embodiments of this application. Thus, statements such as "in one embodiment", "in some embodiments", "in some other embodiments", "in still some other embodiments", etc. that appear in different places in this specification do not necessarily all refer to the same embodiment, but mean "one or more but not all embodiments", unless otherwise specifically emphasized. The terms "comprising", "including", "having" and their variants all mean "including but not limited to", unless otherwise specifically emphasized.

[0036] Text-based image editing is a research hotspot in the multimedia field and has important application value. ManiGAN is used to perform image editing on a source image to be edited according to the content of a text description. However, existing ManiGAN cannot process the intermediate results of image editing when editing the source image according to the text description content, often resulting in the image editing results output by ManiGAN not meeting the user's requirements. In order to make the image editing results output by existing ManiGAN more in line with the user's requirements, an autoencoder is introduced into the multi-level generative adversarial network of existing ManiGAN. This autoencoder can output the intermediate editing results to the user to facilitate the user's direct control of the editing results output in the middle of ManiGAN, so as to obtain a target edited image that more meets the user's requirements.

[0037] The following further elaborates on this application in detail with reference to the accompanying drawings and specific embodiments.

[0038] To solve the problem that the image editing results output by ManiGAN do not meet the user's requirements, this application proposes a text-based image editing method, as Figure 1 shown. This method is executed by an electronic device and includes:

[0039] S101, obtaining the overall image feature and local image feature of a target source image, as well as the overall sentence feature and word feature of a target text.

[0040] Exemplarily, the electronic device acquires the overall image features and local image features of the target source image, as well as the overall sentence features and sentence word features of the target text. The target source image is from the MS-COCO (i.e., Microsoft Common Objects in Context) dataset and the CUB200 dataset; the above-mentioned target text refers to the text information recording the user's editing of the target source image. For example, if the target source image is a bird and the user wants to dye the feathers of this bird red and the head yellow, these editing requirements can be recorded in the target text in text form, that is, the specific content of this target text is to dye the feathers of this bird red and the head yellow.

[0041] The above-mentioned overall image features (i.e., global image features) refer to the features that can represent the entire image and are used to describe the overall features such as image color and shape. For example, color features, texture features, and shape features; the above-mentioned local image features (i.e., local image features) are the local expressions of image features, which reflect the local characteristics of the image.

[0042] Usually, the visual geometry group (VGG) network can be used to extract the overall image features and local image features of the target source image; a special recurrent neural network (RNN), namely the long short-term memory (LSTM) network, can be used to extract the overall sentence features and sentence word features of the target text.

[0043] S102, based on the overall image features, local image features, overall sentence features, and sentence word features, the target source image is edited based on the image editing model to obtain the target edited image; wherein, the image editing model includes: a sampling encoding module and at least one cascaded generation module; the processing process of the image editing model for the target source image includes: using the sampling encoding module to perform sampling encoding processing on the overall image features, overall sentence features, and local image features to obtain a first edited image and output the first edited image; in response to the user instruction, inputting the first edited image, local image features, and sentence word features into at least one cascaded generation module for high-dimensional visual feature extraction to obtain the target edited image, or inputting the target source image, local image features, and sentence word features into at least one cascaded generation module for high-dimensional visual feature extraction to obtain the target edited image.

[0044] Exemplarily, such as Figure 2As shown, the target source image I is a picture of a bird with a size of 128×128 and 3 channels. In this bird picture, the feathers on the bird's abdomen are white, the beak is gray, and the feathers on the head and neck are both gray and white; the VGG network extracts features from the target source image I of 128×128 and obtains the overall image feature c corresponding to the target source image I I and the local image feature M I , where the overall image feature c I is a column vector of 128×1; the local image feature M I has a size of 128×128 and 128 channels; for another example, the specific content of the target text T is to change the feathers on the bird's abdomen to yellow, the beak to yellow, and the feathers on the head and neck to gray and yellow; the LSTM network extracts features from the target text T and obtains the overall sentence feature c corresponding to the target text T T and the word feature M of the sentence T , where the overall sentence feature c T and the word feature M of the sentence T are both column vectors of 128×1.

[0045] As Figure 2 shown, the image editing model 201 includes: a sampling encoding module G 00 and at least one cascaded generation module (for example, G 01 , G 02 ), where the sampling encoding module G 00 includes a first upsampling module F0 and a first autoencoder G0. The first upsampling module F0 performs upsampling on the overall image feature c I and the overall sentence feature c T to obtain a first upsampling result; the first autoencoder G0 encodes the first upsampling result and generates a first edited image

[0046] Optionally, as Figure 2 shown, the first upsampling module F0 includes: multiple identical upsampling layers, a second noisy affine combination module 2011a, and a third noisy affine combination module 2011b. The input of the first upsampling module F0 is the overall sentence feature c T , the overall image feature c I and the local image feature M I . In two adjacent upsampling layers among multiple identical upsampling layers, the input of the latter upsampling layer is the output of the former upsampling layer; the second noisy affine combination module 2011a is located between any two upsampling layers among multiple identical upsampling layers and is used to combine the result output by the former upsampling layer and the local image feature M IPerform feature fusion; the third noise-affine combination module 2011b is used to perform feature fusion on the output result of the last upsampling layer among multiple identical upsampling layers and the local image feature M I Perform feature fusion.

[0047] The above-mentioned multiple identical upsampling layers can be three upsampling layers or four upsampling layers. This application does not make any restrictions on this, and users can set the specific number of upsampling layers according to actual needs. This application only takes Figure 2 the four upsampling layers in the first upsampling module F0 as an example to illustrate the process of feature extraction and feature fusion of the four upsampling layers combined with the second noise-affine combination module 2011a for the overall sentence feature c T and the overall image feature c I and the local image feature M I Perform the process of feature extraction and feature fusion.

[0048] The above-mentioned second noise-affine combination module 2011a can be located between any two of the four upsampling layers in the first upsampling module F0. For example, Figure 2 in the four upsampling layers of the first upsampling module F0, the first three upsampling layers perform upsampling processing on the input 128×1 overall image feature c I and the 128×1 overall sentence feature c T to output an internal feature with a size of 32×32 and 64 channels (i.e., an internal feature of 32×32×64); the second noise-affine combination module 2011a is located between the third upsampling layer and the fourth upsampling layer in the first upsampling module F0; the second noise-affine combination module 2011a performs visual feature enhancement on the upsampling result output by the third upsampling layer (i.e., the 32×32×64 internal feature output by the first three upsampling layers) according to the local image feature M I to obtain a first enhanced upsampling result of 32×32×64; the first enhanced upsampling result first undergoes upsampling processing by the fourth upsampling layer and outputs a visual feature of 64×64×32, and the 64×64×32 visual feature is further enhanced by the third noise-affine combination module 2011b and outputs an enhanced visual feature of 64×64×32; the third noise-affine combination module 2011b performs visual feature enhancement on the upsampling result output by the fourth upsampling layer (i.e., the output result of the last upsampling layer among multiple identical upsampling layers) according to the local image feature M I Perform visual feature enhancement.

[0049] The above-mentioned first autoencoder G0 performs feature extraction and encoding processing on the enhanced visual features of 64×64×32 output by the third noisy affine combination module 2011b, and outputs a first edited image with a size of 64×64 and 3 channels (i.e., the first edited image of 64×64×3 ). The user can directly observe the first edited image of 64×64×3 to determine whether it meets the requirements; if the first edited image of 64×64×3 meets the user's requirements, the first edited image of 64×64×3 is input into at least one cascaded generation module for subsequent processing, such as Figure 2 described in (I) above; if the first edited image of 64×64×3 does not meet the user's requirements, the first edited image of 64×64×3 is discarded and replaced by the target source image I of 128×128×3 and input into the generation module G 01 for subsequent editing processing, such as Figure 2 described in (II) above. Among them, since the first edited image 00 generated by the sampling encoding module G is discarded, the network structure included in the sampling encoding module G 00 is represented by a dotted line box to illustrate that because the first edited image does not meet the user's requirements, the step of generating the first edited image 00 by the sampling encoding module G is skipped, and the target source image I of 128×128×3 is directly input into the generation module G 01 for image editing. Thus, because the first edited image 00 generated by the sampling encoding module G does not meet the requirements, the target source image I of 128×128×3 is directly used to replace the first edited image and input into the generation module G 01 for image editing, thereby avoiding the situation where the first edited image (i.e., the incorrect first edited image ) continues to propagate backward

[0050] Exemplarily, if the user determines that the first edited image of 64×64×3 meets the requirements, the user will send a confirmation instruction to the electronic device, and the electronic device will, according to the received confirmation instruction from the user, use the first edited image of 64×64×3 Performing high-dimensional visual feature extraction in at least one cascaded generation module to obtain a target edited image, as Figure 2 described in (I) therein; if the user determines that the first edited image of 64×64×3 does not meet the requirements, the user will send a rejection instruction to the electronic device, and the electronic device will discard the first edited image of 64×64×3 according to the received rejection instruction from the user and input the target source image I of 128×128×3 into at least one cascaded generation module (for example, generation module G ) to perform high-dimensional visual feature extraction to obtain a target edited image, as 01 described in (II) therein. Figure 2

[0051] The above at least one cascaded generation module can be 1 generation module or two generation modules, and the present application does not make any limitation thereto. The user can set the number of generation modules according to actual needs. The present application only takes Figure 2 the image editing model in which there are 2 generation modules as an example to illustrate the process of further image editing of the intermediate editing result (for example, the first edited image) by the two generation modules in cooperation with the sampling encoding module to generate a target edited image.

[0052] Exemplarily, as Figure 2 shown, the above generation module G 01 includes: a first auto-decoder E0, a first self-attention module 2012c, a second upsampling module F1, and a second auto-encoder G1. The first auto-decoder E0 is used to restore the high-dimensional visual features of the first edited image (i.e., the input information) to obtain a first high-dimensional feature image; the first self-attention module 2012c is used to perform fusion and splicing processing on the first high-dimensional feature image and the sentence word feature M T to obtain sentence semantic information features; the second upsampling module F1 is used to perform feature fusion and upsampling processing on the sentence semantic information features to obtain a second upsampling result; the second auto-encoder G1 is used to perform high-dimensional visual feature extraction on the second upsampling result to obtain output information.

[0053] Figure 2 The processing process of the generation module G 01 for the first edited image (for example, the above first edited image of 64×64×3 ) is as follows: The first auto-decoder E0 processes the above first edited image of 64×64×3 ​Perform high-dimensional visual feature restoration to obtain a first high-dimensional feature image with a size of 64×64 and 32 channels (i.e., 64×64×32); then, the first self-attention module 2012c performs feature fusion and splicing processing on the first high-dimensional feature image of 64×64×32, the sentence word feature M of 128×1 T and the image local feature M of 128×128×128 I to generate a high-dimensional visual feature with fine-grained sentence semantic information (i.e., sentence semantic information feature); the size of this sentence semantic information feature is 64×64 and the number of channels is 32 (i.e., 64×64×32); then, the second upsampling module F1 performs feature fusion and upsampling processing on the sentence semantic information feature of 64×64×32 to generate a visual feature image with a size of 128×128 and 32 channels (i.e., the second upsampling result); the second upsampling module F1 includes two residual networks and an upsampling layer, where the two residual networks are used to fuse the sentence semantic information feature, and the upsampling layer is used to increase the spatial resolution of the image; finally, the second autoencoder G1 performs high-dimensional visual feature extraction and encoding processing on the second upsampling result to generate a second edited image with a size of 128×128 and 3 channels (i.e., 128×128×3)

[0054] Optionally, the second upsampling result can first pass through the noise-affine combination module 2012b, which fuses the second upsampling result and the image local feature M of 128×128×128 I to generate a visual feature image with a size of 128×128 and 32 channels; then, the second autoencoder G1 performs high-dimensional visual feature extraction and encoding processing on the visual feature image of 128×128×32 and outputs a second edited image with a size of 128×128 and 3 channels (i.e., the second edited image of 128×128×3 ).

[0055] The above first self-attention module 2012c includes: a self-attention layer F Atten and a first noise-affine combination module 2012a. The self-attention layer F Atten is used to fuse the first high-dimensional feature image and the sentence word feature; the first noise-affine combination module 2012a is used to fuse the concatenation result of the output of the self-attention layer F Atten and the first high-dimensional feature image, and the image local feature M I For example, the self-attention layer F AttenPerform feature fusion on the above-mentioned first high-dimensional feature image of 64×64×32 and the sentence word feature of 128×1, and splice the fusion result with the first high-dimensional feature image; the first noisy affine combination module 2012a performs feature fusion on the splicing result and the image local feature M of 128×128×128 I Perform feature fusion.

[0056] The second edited image of 128×128×3 output by the second autoencoder G1 The user can directly observe the second edited image of 128×128×3 to determine whether it meets the requirements.

[0057] Exemplarily, if the user determines that the second edited image of 128×128×3 meets the requirements, the user will send a confirmation instruction to the electronic device, and the electronic device will input the second edited image of 128×128×3 according to the received confirmation instruction sent by the user into the next cascaded generation module for high-dimensional visual feature extraction to obtain the target edited image; if the user determines that the second edited image of 128×128×3 does not meet the requirements, the user will send a rejection instruction to the electronic device, and the electronic device will discard the generation module G 01 that generates the second edited image of 128×128×3 (that is, does not consider the generation module G 00 that generates the first edited image ) and input the target source image of 128×128×3 into the next cascaded generation module (for example, the generation module G 02 ) for high-dimensional visual feature extraction to obtain the target edited image, thus avoiding the situation where the second edited image (that is, the incorrect second edited image ) continues to propagate backward.

[0058] Exemplarily, as Figure 2 shown, the next cascaded generation module G 01 after the above-mentioned generation module G 02 includes: a second auto-decoder E1, a second self-attention module 2013c, a third upsampling module F2, and a generator G2. The second auto-decoder E1 is used to restore the high-dimensional visual features of the second edited image (that is, the output information of the previous generation module G 01 ) to obtain the second high-dimensional feature image; the second self-attention module 2013c is used to fuse and splice the second high-dimensional feature image and the sentence word feature M T and output the second edited image The corresponding sentence semantic information features; the third upsampling module F2 is used to perform feature fusion and upsampling processing on the second edited image of the corresponding sentence semantic information features to obtain a third upsampling result; the generator G2 is used to perform high-dimensional visual feature extraction on the third upsampling result.

[0059] Figure 2 in the generation module G 02 for the second edited image (For example, the above-mentioned second edited image of 128×128×3 ) The processing process is as follows: the second auto-decoder E1 performs high-dimensional visual feature restoration on the above-mentioned second edited image of 128×128×3 to obtain a second high-dimensional feature image with a size of 128×128 and 32 channels (i.e., 128×128×32); then, the second self-attention module 2013c performs feature fusion and splicing processing on the second high-dimensional feature image of 128×128×32, the sentence word feature M of 128×1 T and the image local feature M of 128×128×128 I to generate high-dimensional visual features with fine-grained sentence semantic information (i.e., the corresponding sentence semantic information features of the second edited image ); the size of the corresponding sentence semantic information features of this second edited image is 128×128× and 32 channels (i.e., 128×128××32); then, the third upsampling module F2 performs feature fusion and upsampling processing on the sentence semantic information features of 128×128××32 to generate a visual feature image with a size of 256×256 and 32 channels (i.e., the third upsampling result); the third upsampling module F2 includes two residual networks and an upsampling layer, where the two residual networks are used to fuse the sentence semantic information features, and the upsampling layer is used to improve the spatial resolution of the image; finally, the generator G2 performs high-dimensional visual feature extraction and encoding processing on the third upsampling result to generate a third edited image with a size of 256×256 and 3 channels (i.e., 256×256×3)

[0060] Optionally, the third upsampling result of 256×256×32 can first pass through the noise-affine combination module 2013b, and the noise-affine combination module 2013b combines the third upsampling result of 256×256×32 and the image local feature M of 128×128×128 Iare fused to generate a visual feature image with a size of 256×256 and 32 channels (i.e., 256×256×32); then, the generator G2 performs high-dimensional visual feature extraction and encoding processing on the visual feature image of 256×256×32 to obtain a third edited image with a size of 256×256 and 3 channels (i.e., the third edited image of 256×256×3 ).

[0061] The second self-attention module 2013c includes: a self-attention layer F Atten and a noisy affine combination module 2012a. The self-attention layer F Atten is used to perform feature fusion on the second high-dimensional feature image and the sentence word features; the noisy affine combination module 2012a is used to perform feature fusion on the concatenation result of the output result of the self-attention layer F Atten and the second high-dimensional feature image, and the image local feature M I . For example, the self-attention layer F Atten performs feature fusion on the second high-dimensional feature image of 128×128×32 and the sentence word features of 128×1, and concatenates the fusion result with the second high-dimensional feature image; the noisy affine combination module 2013a performs feature fusion on the concatenation result and the image local feature M of 128×128×128 I .

[0062] The third edited image of 256×256×3 output by the generator G2 The user can directly observe the third edited image of 256×256×3 to determine whether it meets the requirements. If the user determines that the third edited image of 256×256×3 meets the requirements, the user will send a confirmation instruction to the electronic device. The electronic device will input the third edited image of 256×256×3 received from the user into the next-level network for further processing or directly output it to the user (i.e., since the generation module G 02 is the last generation module in the two cascaded generation modules, therefore, the output information of the generator G2 in the generation module G 02 is the target edited image (i.e., the third edited image )); if the user determines that the third edited image of 256×256×3 does not meet the requirements, the user will send a rejection instruction to the electronic device. The electronic device will input the third edited image of 256×256×3 Discard and reuse the image editing model, and repeatedly execute the foregoing processing process on the target source image of 128×128×3.

[0063] The core algorithm of the above first noisy affine combination module 2012a is as follows:

[0064] F NAC (h, M) = h⊙W rnd (M) + b rnd (M), (1)

[0065] Where, F NAC (h, M) represents the core algorithm function of the first noisy affine combination module 2012a, W rnd (M) = W2(W1(M) + noise), b rnd (M) = b2(b1(M) + noise), through W rnd (M) and b rnd (M) calculate the weights and biases of the region feature M of the target source image (for example, the aforementioned local image feature M I ), h represents the visual feature (for example, the concatenation result of the output result of the self-attention layer F Atten ) and the first high-dimensional feature image), noise is Gaussian noise, and ⊙ is the Hadamard element-wise product. The first noisy affine combination module 2012a can enhance the reliability of the sampling encoding module and the generation module for image editing by introducing Gaussian noise, so that the sampling encoding module and the generation module will not affect the reliability of the editing result due to the random noise existing in the image.

[0066] It should be noted that, except for the first noisy affine combination module 2012a, the core algorithms of other noisy affine combination modules (such as the second noisy affine combination module, the third noisy affine combination module, etc.) mentioned in this application are the same as those of the first noisy affine combination module 2012a, and this application will not elaborate on this.

[0067] Exemplarily, the training method of the image editing model includes: training the initial model using a preset loss function and a training set to obtain the image editing model; wherein, the preset loss function includes sub-functions corresponding to N sub-networks and N - 1 autoencoder loss functions, the initial model includes N sub-networks, and these N sub-networks are the initial models corresponding to the sampling encoding module and at least one generation module respectively; during the training process, when the output image of the i-th sub-network does not meet the preset conditions, the sub-functions corresponding to the i + 1-th to N-th sub-networks and the i-th to i + 1-th autoencoder loss functions are used to train the initial model, where 0 ≤ i < N.

[0068] For example, as Figure 2As shown in the figure, the training method of the image editing model 201 includes: using the MS-COCO and CUB200 datasets and a preset loss function, and with the help of discriminators (such as discriminator D0, discriminator D1, and discriminator D2), training is carried out in an adversarial form to obtain the image editing model 201; the above initial model refers to the network model of the image editing model 201 before training; this initial model includes N = 3 sub-networks and two (N - 1 = 3 - 1) autoencoders. These 3 sub-networks are respectively the sampling encoding module G 00 , the generation module G 01 , and the generation module G 02 ; the two autoencoders are respectively the autoencoder G0 - the decoder E0 and the autoencoder G1 - the decoder E1; the above preset loss function is the loss functions of the 3 sub-networks and the loss functions of the two autoencoders. During the training process, for example, when N = 3, when the output image of the 1st (i.e., i = 1) sub-network (such as the sampling encoding module G 00 ) does not meet the preset conditions, the sub-functions corresponding to the 2nd (i.e., i + 1 = 1 + 1) to the 3rd sub-networks (i.e., the loss functions corresponding to the generation module G 01 and the loss function corresponding to the generation module G 02 ) and the loss functions of the 1st to the 2nd autoencoders (i.e., the loss function of the autoencoder G0 - the decoder E0 and the loss function of the autoencoder G1 - the decoder E1) are used to train the initial model.

[0069] When training the initial model, the following three preset loss functions will be used for training. These three preset loss functions are specifically as follows:

[0070] The first preset loss function:

[0071] The second preset loss function:

[0072] The third preset loss function: In the above formula, b represents the number of skipped sub-networks during training, i = 0, 1, 2; L G,i represents the loss function of the corresponding layer in the ManiGAN network, L G,0 is the loss function corresponding to the layer where the sampling encoding module G 00 is located, L G,1 is the loss function corresponding to the layer where the generation module G 01 is located, L G,2 is the loss function corresponding to the layer where the generation module G 02 is located; in the above formula, I′ represents the target edited image, I′~P G,i(I,T)It means that the target edited image I′ is generated by the initial model given the training image I and the randomly selected target text T.

[0073] The loss function of the above autoencoder is while where is the loss function of the autoencoder G0 - decoder E0, is the loss function of the autoencoder G1 - decoder E1; G i consists of a 3×3 convolutional layer and a tanh activation function, and E i includes an atanh function, a 3×3 convolutional layer, a LeakyReLU layer, and an Instance Normalization layer. The autoencoder composed of G i and E i feeds the intermediate editing result back to the user, and the user can directly control the intermediate results generated by different sub - networks to prevent the inaccurate intermediate editing results from affecting the accuracy of the output result of the entire image editing model 201.

[0074] Exemplarily, taking the training of the initial model with the first preset loss function as an example, when training the initial model with the first preset loss function, no sub - network needs to be skipped. That is, the overall image feature c I of the training image and the local image feature M I , as well as the overall sentence feature c T of the training text, are input into the sampling encoding module G 00 . The sampling encoding module G 00 outputs the first training edited image. If the user determines that the first training edited image meets the user's requirements, the first training edited image is input into the generation module G 01 . The generation module G 01 edits the first training edited image and the local image feature M I of the training image and generates the second training edited image. If the user determines that the second training edited image meets the user's requirements, the second training edited image is input into the generation module G 02 . The generation module G 02 edits the second training edited image and the local image feature M I of the training image and generates the third training edited image. If the user determines that the third training edited image meets the user's requirements, the training of the initial model according to the first preset loss function is completed. If the user determines that the third training edited image does not meet the user's requirements, the above training process is repeated. Since the autoencoder G0 - decoder E0 and the autoencoder G1 - decoder E1 are distributed in the sampling encoding module G 00 , the generation module G01 and the generation module G 02 therebetween (see Figure 2 ), thus, when training the initial model according to the first preset loss function, the autoencoder G0-auto decoder E0 and the autoencoder G1-auto decoder E1 are also trained.

[0075] Exemplarily, taking the training of the initial model with the second preset loss function as an example, when training the initial model with the second preset loss function, it is necessary to skip the sampling encoding module G 00 to process the overall image feature c of the training image I and the local image feature M I , as well as the overall sentence feature c of the training text T . Instead, the training image is directly input into the generation module G 01 ; the generation module G 01 edits the training image and the local image feature M of the training image I and generates a fourth training edited image; if the user determines that the fourth training edited image meets the user's requirements, the fourth training edited image is input into the generation module G 02 ; the generation module G 02 edits the fourth training edited image and the local image feature M of the training image I and generates a fifth training edited image; if the user determines that the fifth training edited image meets the user's requirements, the training of the initial model according to the second preset loss function is completed; if the user determines that the fifth training edited image does not meet the user's requirements, the above training process is repeated. Since there is only the autoencoder G1-auto decoder E1 between the generation module G 01 and the generation module G 02 (see Figure 2 ), thus, when training the initial model according to the second preset loss function, the autoencoder G1-auto decoder E1 is also trained.

[0076] Exemplarily, taking the training of the initial model with the third preset loss function as an example, when training the initial model with the third preset loss function, it is necessary to skip the process of the sampling encoding module G 00 processing the overall image feature c of the training image I and the local image feature M I , as well as the overall sentence feature c of the training text T , and skip the process of the generation module G 01 editing the editing result output by the sampling encoding module G 00 and the local image feature M of the training image I . Instead, the training image is directly input into the generation module G 02 ; the generation module G02 For the training image and the image local feature M of the training image I perform editing processing and generate the sixth training edited image; if the user determines that the sixth training edited image meets the user's requirements, the training of the initial model according to the third preset loss function is completed; if the user determines that the sixth training edited image does not meet the user's requirements, repeat the above training process. Since the generation module G 02 does not have an autoencoder, therefore, when training the initial model according to the third preset loss function, there is no need to train the autoencoder.

[0077] Due to the characteristics that the multi-layer adversarial network (i.e., the above N sub-networks) is difficult to train and converge, therefore, when training the initial model, an autoencoder is used to randomly skip the lower-layer adversarial network (for example, skip the sampling encoding module G 00 ) and preferentially train the higher-layer network (for example, the generation module G 01 , or, the generation module G 02 ), for example, when b = 1, skip the sampling encoding module G 00 and preferentially train the generation module G 01 and the generation module G 02 ; when b = 2, skip the sampling encoding module G 00 and the generation module G 01 . The above training method has the following advantages: a) Since when skipping the lower-layer sub-network, the input of the next-level sub-network uses the original training image instead of the incorrect editing result (i.e., the wrong editing result) output by the lower-layer sub-network that does not meet the user's requirements, it is possible to avoid the propagation of the incorrect editing result generated by the potential lower-layer sub-network to the higher-layer sub-network; b) Preferentially training the higher-layer sub-network can bring better update gradients to the lower-layer sub-network, so that the convergence effect of the lower-layer sub-network can be better.

[0078] Exemplarily, the image editing model further includes: a detail correction model (Symmetrical Detail Correction Module, SCDM) for modifying the details of the target editing image; the detail correction model is used to process the image local feature, the sentence word feature and the target editing image to obtain the target corrected image; the detail correction model includes: a first detail correction module, a second detail correction module, a fusion module and a generator, wherein the first detail correction module is used to modify the details of the image local feature, the first random noise and the sentence word feature to obtain the first detail feature; the second detail correction module is used to modify the details of the image local feature corresponding to the target editing image, the second random noise and the sentence word feature to obtain the second detail feature; the fusion module is used to perform feature fusion on the first detail feature and the second detail feature; the generator is used to generate the target corrected image according to the output result of the fusion module.

[0079] As shown Figure 2 in the figure, the image editing model 201 further includes a detail correction model 2014, which corrects the details of the target edited image generated by the generation module G I according to the local image feature M of the target source image I T and the sentence word feature M of the text description T 02 to obtain the target corrected image The detail correction model 2014 includes: a first detail correction module 301, a second detail correction module 302, a fusion module F fuse and a generator G 0S , where the first detail correction module 301 corrects the details of the local image feature M I , the first random noise noise1 and the sentence word feature M T to obtain the first detail feature, and inputs the first detail feature into the fusion module F fuse ; the second detail correction module 302 corrects the details of the local image feature corresponding to the target edited image (for example, the third edited image output by the generation module G 02 ) , the second random noise noise2 and the sentence word feature M to obtain the second detail feature, and inputs the second detail feature into the fusion module F T ; the fusion module F fuse performs feature fusion on the first detail feature and the second detail feature, and inputs the fusion result into the generator G fuse ; the generator G 0S encodes the output result of the fusion module F 0S and generates the target corrected image fuse

[0080] Exemplarily, as Figure 3 shown in the figure, the first detail correction module 301 includes a fourth noise-affine combination module 3011, a fifth noise-affine combination module 3013, a sixth noise-affine combination module 3015, a second self-attention module 3012, a first residual network 3014 and a first linear network 3016; where the fourth noise-affine combination module 3011 fuses the first random noise noise1 and the local image feature M I to obtain the first fusion feature, and inputs the first fusion feature into the second self-attention module 3012; the second self-attention module 3012 performs self-attention on the first fusion feature and the sentence word feature M T and the sentence word feature MPerform feature fusion, splice the fusion result with the first random noise noise1 to obtain a splicing result; input the splicing result into the fifth noisy affine combination module 3013; the fifth noisy affine combination module 3013 performs feature fusion on the splicing result and the local image feature M I Perform feature fusion, and input the fusion result into the first residual network 3014; the first residual network 3014 extracts visual features from the fusion result output by the fifth noisy affine combination module 3013, and inputs the visual feature extraction result into the sixth noisy affine combination module 3015; the first linear network 3016 performs a linear transformation on the local image feature M I and inputs the linear transformation result into the sixth noisy affine combination module 3015; the sixth noisy affine combination module 3015 performs feature fusion on the visual feature extraction result output by the first residual network 3014 and the linear transformation result output by the first linear network 3016 to obtain the first detailed feature x I .

[0081] Exemplarily, as Figure 3 shown, the second detail correction module 302 includes a seventh noisy affine combination module 3021, an eighth noisy affine combination module 3023, a ninth noisy affine combination module 3025, a third self-attention module 3022, a second residual network 3024, and a second linear network 3026; among them, the seventh noisy affine combination module 3021 performs feature fusion on the second random noise noise2 and the local image feature corresponding to the target edited image (for example, the third edited image 02 output by the generation module G corresponding local image feature ) to obtain a first fusion feature, and input the first fusion feature into the third self-attention module 3022; the third self-attention module 3022 performs feature fusion on the first fusion feature and the sentence word feature M T and splices the feature fusion result with the second random noise noise2, and finally inputs the splicing result into the eighth noisy affine combination module 3023; the eighth noisy affine combination module 3023 performs feature fusion on the splicing result and the local image feature corresponding to the target edited image (for example, the third edited image 02 output by the generation module G corresponding local image feature ) and inputs the fusion result into the second residual network 3024; the second residual network 3024 extracts visual features from the fusion result output by the eighth noisy affine combination module and inputs the visual feature extraction result into the ninth noisy affine combination module 3025; the second linear network 3026 performs a linear transformation on the local image feature corresponding to the target edited image (for example, the generation module G02 The third edited image output The corresponding local image features ) Perform a linear transformation and input the result of the linear transformation into the ninth noisy affine combination module 3025; the ninth noisy affine combination module 3025 performs feature fusion on the visual feature extraction result output by the second residual network 3024 and the result of the linear transformation output by the second linear network 3026 to obtain the second detailed feature The above-mentioned local image features are obtained by using the VGG network to extract features from the third edited image

[0082] The above-mentioned fusion module F fuse performs feature fusion on the first detailed feature x output by the sixth noisy affine combination module 3015 I and the second detailed feature output by the ninth noisy affine combination module 3025 and inputs the fusion result into the generator G 0S ; the generator G 0S encodes the output fusion result of the fusion module F fuse and generates the target corrected image

[0083] Optionally, the inputs of the first detailed correction module 301 and the second detailed correction module 302 can be exchanged, that is, the input of the first detailed correction module 301 can be the target edited image (for example, the third edited image output by the generation module G 02 ) corresponding local image features the second random noise noise2 and the sentence word feature M ; the input of the second detailed correction module 302 can be the local image feature M T , the first random noise noise1 and the sentence word feature M I . Correspondingly, the output of the first detailed correction module 301 is the second detailed feature T . The output of the first detailed correction module 301 is the first detailed feature x I .

[0084] Figure 3 In, the core algorithm of the fusion module F fuse is as follows:

[0085]

[0086] Among them, F residual is a residual network, and β1 and β2 are linear network layers for the input first detailed feature x I and the second detailed feature​​ A pair of attention weights obtained through calculation. The fusion module F fuse can adaptively select whether to modify the image features based on the target source image or the target edited image generated by the image editing model, so as to enhance the detailed features of the modified image. The above-mentioned generator G 0S converts the output result of the fusion module F fuse into the final target corrected image

[0087] Exemplarily, the training method of the detail correction model 2014 includes: training the generator of the detail correction model according to the conditional generator loss function, the unconditional generator loss function, and the semantic contrast function; training the discriminator of the detail correction model according to the conditional discriminator loss function and the unconditional discriminator loss function; wherein, the training data set is the MS-COCO data set and the CUB200 data set; the conditional generator loss function is as follows:

[0088]

[0089] wherein, L Gs,0 is the conditional generator loss function, and L Ds,0 is the conditional discriminator loss function; D S is the discriminator of the detail correction model in the image editing model; refers to the target corrected image is generated by the detail correction model (SDCM) given the training image I and the description text T of the training image I I ; I~P data means that the training image I is sampled from real data; T is a randomly selected target text, is the semantic contrast function, and the purpose is to make the target corrected image closer to the description text T of the training image I than the randomly selected target text T I . The above function is defined as:

[0090]

[0091] wherein, ρ c is the contrast control threshold; L corre is the correlation function in ControlGAN, which is used to describe the matching degree between the training text and the target corrected image.

[0092] The above unconditional generator loss function is as follows:

[0093]

[0094] Among them, L Gs,1 is the unconditional generator loss function, and L Ds,1 is the unconditional discriminator loss function;; D S is the discriminator of the detail correction module in the image editing model; refers to the target corrected image is generated by the detail correction model (SDCM) given the training image I and the randomly selected target text T. The total loss functions of the generator and discriminator in the detail correction model 2014 are respectively:

[0095]

[0096] Among them, L Gs is the total loss function of the generator, and L Ds is the total loss function of the discriminator; L ControlGAN is the multi-modal loss function of text and image, which is used to improve the matching degree between the image editing result and the target text; The above image local features are obtained by using the VGG network to extract features from the target corrected image , and the image local feature M I is obtained by using the VGG network to extract features from the target source image I. The function L DAMSM is the text-image similarity function defined in ManiGAN; L reg The function is the regularization term defined in the image editing model, which is used to strengthen the modification effect, During the training of the detail correction model, the training image I and the target corrected image are randomly replaced with each other to accelerate the training process of the detail correction model.

[0097] For the sake of easy understanding, the overall process of the text-based image editing method provided by this application will be exemplarily described below in combination with Figure 4

[0098] As Figure 4 shown, the configurable editing part means that the user can control the intermediate editing results output by the sampling encoding module G 00 , the generation module G 01 and the generation module G 02 in the image editing model, and selectively skip some modules by judging the intermediate editing results. Replace the intermediate editing results that do not meet the user's requirements with the target source image, so as to realize the editing process of the target source image according to the target text. The following uses the sub-flowcharts of (a) to (h) in Figure 4 to illustrate the user's operations on the sampling encoding module G 00 , the generation module G​01 and the generation module G 02 The process of randomly controlling the intermediate editing results output.

[0099] Figure 4 Among them, (a) complete process: The user inputs the overall image features and local image features of the target source image, as well as the overall sentence features and sentence word features of the target text into the image editing model. The sampling encoding module G in the image editing model 00 The first edited image output meets the user's requirements; The sampling encoding module G 00 Inputs the first edited image into the generation module G 01 In, the generation module G 01 Processes the first edited image and the second edited image output also meets the user's requirements; Inputs the second edited image into the generation module G 02 ; The generation module G 02 Processes the second edited image and the third edited image generated also meets the user's requirements. The generation module G 02 Inputs the third edited image (i.e., the target edited image) into the detail correction model (SCDM). This detail correction model (SCDM) performs detail correction on the third edited image and obtains the target corrected image.

[0100] Figure 4 Among them, (b) skipping G 00 : The user inputs the overall image features and local image features of the target source image, as well as the overall sentence features and sentence word features of the target text into the image editing model. The sampling encoding module G in the image editing model 00 The first edited image output does not meet the user's requirements; The first edited image output by this sampling encoding module G 00 Is discarded, and the target source image is used to replace the first edited image and input into the generation module G 01 In, the generation module G 01 Processes the target source image and the second edited image output meets the user's requirements; Inputs the second edited image into the generation module G 02 ; The generation module G 02 Processes the second edited image and the third edited image generated also meets the user's requirements. The generation module G 02 Inputs the third edited image (i.e., the target edited image) into the detail correction model (SCDM). This detail correction model (SCDM) performs detail correction on the third edited image and obtains the target corrected image.

[0101] Figure 4 Among them, (c) skipping G 00 and G 01: The user inputs the overall image features and local image features of the target source image, as well as the overall sentence features and sentence word features of the target text into the image editing model. The sampling encoding module G in the image editing model 00 The first edited image output does not meet the user's requirements; this sampling encoding module G 00 Discards the first edited image output, and uses the target source image to replace the first edited image and inputs it into the generation module G 01 In, the generation module G 01 The second edited image processed and output by the generation module G for the target source image also does not meet the user's requirements; this generation module G 01 Discards the second edited image, and uses the target source image to replace the second edited image and inputs it into the generation module G 02 ; The generation module G 02 The third edited image processed and generated by the generation module G for the target source image meets the user's requirements. The generation module G 02 Inputs the third edited image (i.e., the target edited image) into the detail correction model (SCDM). This detail correction model (SCDM) performs detail correction on the third edited image and obtains the target corrected image.

[0102] Figure 4 In, (d) only uses SCDM (i.e., skips G 00 、G 01 And G 02 ): The user inputs the overall image features and local image features of the target source image, as well as the overall sentence features and sentence word features of the target text into the image editing model. The sampling encoding module G in the image editing model 00 The first edited image output does not meet the user's requirements; this sampling encoding module G 00 Discards the first edited image output, and uses the target source image to replace the first edited image and inputs it into the generation module G 01 In, the generation module G 01 The second edited image processed and output by the generation module G for the target source image also does not meet the user's requirements; this generation module G 01 Discards the second edited image, and uses the target source image to replace the second edited image and inputs it into the generation module G 02 ; The generation module G 02 The third edited image processed and generated by the generation module G for the target source image also does not meet the user's requirements. This generation module G 02 Discards the third edited image, and uses the target source image to replace the third edited image and inputs it into the detail correction model (SCDM). This detail correction model (SCDM) performs detail correction on the target source image according to the local image features of the target source image and the sentence word features of the target text and obtains the target corrected image.

[0103] Figure 4Among them, (e) Skip SCDM: The user inputs the overall image features and local image features of the target source image, as well as the overall sentence features and sentence word features of the target text into the image editing model. The sampling and encoding module G in the image editing model 00 The first edited image output meets the user's requirements; The sampling and encoding module G 00 Inputs the first edited image into the generation module G 01 Among them, the generation module G 01 Processes the first edited image and the second edited image output also meets the user's requirements; Inputs the second edited image into the generation module G 02 ; The generation module G 02 Processes the second edited image and the third edited image generated also meets the user's requirements. This third edited image is the target edited image.

[0104] Figure 4 Among them, (f) Repeat G 01 : The user inputs the overall image features and local image features of the target source image, as well as the overall sentence features and sentence word features of the target text into the image editing model. The sampling and encoding module G in the image editing model 00 The first edited image output meets the user's requirements; The sampling and encoding module G 00 Inputs the first edited image into the generation module G 01 Among them, the generation module G 01 Processes the first edited image and the second edited image output does not meet the user's requirements; This generation module G 01 Can be reused until the second edited image that meets the user's requirements is output; For example, the generation module G 01 Re-processes according to the first editing result and generates a new second edited image that meets the requirements, and inputs this new second edited image into the generation module G 02 ; The generation module G 02 Processes the new second edited image and the third edited image generated also meets the user's requirements. The generation module G 02 Inputs the third edited image (i.e., the target edited image) into the detail correction model (SCDM). This detail correction model (SCDM) corrects the details of the third edited image and obtains the target corrected image.

[0105] Figure 4 Among them, (g) Repeat G 02 : The user inputs the overall image features and local image features of the target source image, as well as the overall sentence features and sentence word features of the target text into the image editing model. The sampling and encoding module G in the image editing model 00 The first edited image output meets the user's requirements; The sampling and encoding module G 00Input the first edited image into generation module G 01 In the generation module G 01 The second edited image processed and output from the first edited image meets the user requirements; the generation module G 02 The third edited image processed and generated from the second edited image does not meet the user requirements. This generation module G 02 Can be reused until the third edited image that meets the user requirements (i.e., the target edited image) is output; for example, the generation module G 02 Re-processes according to the second editing result and generates a new third edited image that meets the requirements, and inputs this new third edited image into the generation module G 02 ; the generation module G 02 Inputs the new third edited image (i.e., the target edited image) into the detail correction model (SCDM). This detail correction model (SCDM) corrects the details of the new third edited image and obtains the target corrected image.

[0106] Figure 4 In, (h) Repeat SCDM: The user inputs the overall image features and local image features of the target source image, as well as the overall sentence features and sentence word features of the target text into the image editing model. The sampling encoding module G in the image editing model 00 The output first edited image meets the user requirements; the sampling encoding module G 00 Inputs the first edited image into the generation module G 01 In the generation module G 01 The second edited image processed and output from the first edited image meets the user requirements; the generation module G 02 The third edited image processed and generated from the second edited image meets the user requirements. The generation module G 02 Inputs the third edited image (i.e., the target edited image) into the detail correction model (SCDM). This detail correction model (SCDM) corrects the details of the third edited image and obtains a fourth edited image that does not meet the user requirements. This detail correction model (SCDM) can be reused until the fourth edited image that meets the user requirements (i.e., the target corrected image) is output; for example, the detail correction model (SCDM) re-processes according to the third editing result and generates a new fourth edited image that meets the requirements, and this new fourth edited image is the target corrected image.

[0107] The beneficial effects of the above technical solutions proposed in this application compared with the prior art are as follows:

[0108] Compared with the existing ManiGAN that edits the source image to be edited according to the text and directly outputs an editing result that may not meet the user's requirements, the present application introduces a sampling encoding module into the existing ManiGAN to form an improved ManiGAN; the sampling encoding module will output an intermediate editing result (i.e., the first edited image) to facilitate the user to judge whether the intermediate editing result meets the requirements. If it meets the requirements, the intermediate editing result will be passed on to at least one cascaded generation module; if it does not meet the requirements, the intermediate result will not be passed on to at least one cascaded generation module, but the target source image will be used to replace the intermediate editing result and continue to be passed on to at least one cascaded generation module. It can be seen that when the improved ManiGAN edits the target source image according to the target text, it can control the intermediate editing result and promptly eliminate the intermediate editing results that do not meet the requirements, so as to prevent the inaccurate results output by the previous stage from affecting the accuracy of the results output by the subsequent stage, thereby editing a target edited image that better meets the requirements for the user.

[0109] A first noise-affected affine combination module is introduced into the above-mentioned first self-attention module. By introducing Gaussian noise, the first noise-affected affine combination module can enhance the reliability of the image edited by the generation module, thus avoiding the situation where the reliability of the editing result is affected by random noise in the image in the generation module.

[0110] Introducing a second noise-affected affine combination module and a third noise-affected affine combination module into the first upsampling module can further enhance the visual features of the output results of different upsampling layers in the first upsampling module.

[0111] Adding a detail correction model to the above-mentioned image editing model can further modify and enhance the details of the target edited image output by the image editing model, so as to obtain a high-resolution target corrected image.

[0112] Adding multiple noise-affected affine combination modules to the above-mentioned first detail correction module can enhance the reliability of the detail correction model.

[0113] Adding multiple noise-affected affine combination modules to the above-mentioned second detail correction module can enhance the reliability of the detail correction model.

[0114] Training the generator of the detail correction model according to the conditional generator loss function, unconditional generator loss function and semantic contrast function can make the image editing result (i.e., the target edited image) generated by the generator more in line with the content described in the target text and the user's requirements. Training the discriminator of the detail correction model according to the conditional discriminator loss function and unconditional discriminator loss function can make the recognition result of the discriminator more accurate.

[0115] During the training of the above image editing model, the training of N sub-networks is carried out by an autoencoder to skip the previous-stage sub-networks whose output results do not meet the requirements and preferentially train the subsequent-stage sub-networks. Since the target source image rather than the intermediate editing result (e.g., the first edited image) is used during the skipping, potential error results output by the previous-stage sub-networks can be avoided from propagating to the subsequent-stage sub-networks. The preferential training of the subsequent-stage sub-networks can bring better update gradients to the previous-stage sub-networks, thereby making the convergence effect of the previous-stage sub-networks better.

[0116] Figure 5 FIG. shows a schematic structural diagram of an electronic device provided by the present application. Figure 5 The dotted line in indicates that the unit or module is optional. The electronic device 500 can be used to implement the method described in the above method embodiments. The electronic device 500 can be a terminal device, a server, or a chip.

[0117] The electronic device 500 includes one or more processors 501, and the one or more processors 501 can support the electronic device 500 to implement Figure 1 the method in the corresponding method embodiment. The processor 501 can be a general-purpose processor or a dedicated processor. For example, the processor 501 can be a central processing unit (CPU). The CPU can be used to control the electronic device 500, execute software programs, and process the data of software programs. The electronic device 500 can also include a communication unit 505 for implementing signal input (reception) and output (transmission).

[0118] For example, the electronic device 500 can be a chip, and the communication unit 505 can be the input and / or output circuit of the chip, or the communication unit 505 can be the communication interface of the chip, and the chip can be a component of the terminal device.

[0119] Again, for example, the electronic device 500 can be a terminal device, and the communication unit 505 can be the transceiver of the terminal device, or the communication unit 505 can be the transceiver circuit of the terminal device.

[0120] The electronic device 500 may include one or more memories 502, on which there is a program 504. The program 504 can be run by the processor 501 to generate instructions 503, so that the processor 501 executes the method described in the above method embodiment according to the instructions 503. Optionally, data (such as the ID of the chip to be tested) can also be stored in the memory 502. Optionally, the processor 501 can also read the data stored in the memory 502. The data can be stored at the same storage address as the program 504, or the data can be stored at a different storage address from the program 504.

[0121] The processor 501 and the memory 502 can be separately provided or integrated together. For example, they can be integrated on a system on chip (SOC) of a terminal device.

[0122] For the specific manner in which the processor 501 executes the method for starting the aging test, reference can be made to the relevant descriptions in the method embodiments.

[0123] It should be understood that the steps of the above method embodiments can be completed by a logic circuit in hardware form or an instruction in software form in the processor 501. The processor 501 can be a CPU, a digital signal processor (DSP), a field programmable gate array (FPGA), or other programmable logic devices. For example, discrete gates, transistor logic devices, or discrete hardware components.

[0124] The present application also provides a computer program product. When the computer program product is executed by the processor 501, it implements the method described in any of the method embodiments of the present application.

[0125] The computer program product can be stored in the memory 502. For example, it is the program 504. After processes such as preprocessing, compilation, assembly, and linking, the program 504 is finally converted into an executable target file that can be executed by the processor 501.

[0126] The present application also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a computer, it implements the method described in any of the method embodiments of the present application. The computer program can be a high-level language program or an executable target program.

[0127] The computer-readable storage medium is, for example, the memory 502. The memory 502 may be a volatile memory or a non-volatile memory, or the memory 502 may include both a volatile memory and a non-volatile memory. Among them, the non-volatile memory may be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory may be a random access memory (RAM), which is used as an external cache. By way of example but not limitation, many forms of RAM are available, such as static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchlink dynamic random access memory (SLDRAM), and direct rambus random access memory (DRRAM).

[0128] Those skilled in the art can clearly understand that, for the convenience and conciseness of description, the specific working processes and the technical effects generated by the above-described devices and apparatuses can refer to the corresponding processes and technical effects in the foregoing method embodiments, and will not be elaborated herein again.

[0129] In several embodiments provided in the present application, the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, some features of the above-described method embodiments can be ignored or not executed. The above-described apparatus embodiments are merely illustrative. The division of units is only a logical function division, and there may be other division methods in actual implementation. Multiple units or components can be combined or integrated into another system. In addition, the coupling between units or the coupling between components can be a direct coupling or an indirect coupling. The above couplings include electrical, mechanical, or other forms of connection.

[0130] The above-described embodiments are only used to illustrate the technical solutions of the present application and are not intended to limit the same. Although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features. These modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the various embodiments of the present application, and should all be included within the protection scope of the present application.

Claims

1. A text-based image editing method, characterized in that The method includes: obtaining the overall image feature and local image feature of the target source image, as well as the overall sentence feature and word feature of the target text; editing the target source image based on the overall image feature, the local image feature, the overall sentence feature, and the word feature according to an image editing model to obtain a target edited image; wherein, the image editing model includes: a sampling encoding module and at least one cascaded generation module; the processing process of the image editing model for the target source image includes: using the sampling encoding module to perform sampling encoding processing on the overall image feature, the overall sentence feature, and the local image feature to obtain a first edited image and output the first edited image; in response to a user instruction, inputting the first edited image, the local image feature, and the word feature into the at least one cascaded generation module to perform high-dimensional visual feature extraction to obtain the target edited image, or inputting the target source image, the local image feature, and the word feature into the at least one cascaded generation module to perform high-dimensional visual feature extraction to obtain the target edited image, and the generation module includes: a first auto-decoder, a first self-attention module, a second upsampling module, and a second auto-encoder, the first auto-decoder is used to restore the high-dimensional visual feature of the input information to obtain a first high-dimensional feature image, and the input information is the first edited image or the output information of the previous generation module; the first self-attention module is used to fuse and splice the first high-dimensional feature image and the word feature to obtain a sentence semantic information feature; the second upsampling module is used to perform feature fusion and upsampling processing on the sentence semantic information feature to obtain a second upsampling result; the second auto-encoder is used to perform high-dimensional visual feature extraction on the second upsampling result to obtain output information, and when the generation module is the last generation module in the at least one cascaded generation module, the output information is the target edited image.

2. The method according to claim 1, wherein The first self-attention module includes: a self-attention layer and a first noise-affine combination module; the self-attention layer is used to fuse the first high-dimensional feature image and the word feature; the first noise-affine combination module is used to fuse the concatenation result of the output result of the self-attention layer and the first high-dimensional feature image, and the local image feature.

3. The method according to any one of claims 1 to 2, characterized in that, The sampling encoding module includes: a first upsampling module and a first auto-encoder; the first upsampling module is used to perform upsampling processing on the overall image feature, the overall sentence feature, and the local image feature to obtain a first upsampling result; the first auto-encoder is used to generate a first edited image according to the first upsampling result.

4. The method according to claim 3, characterized in that, The first upsampling module includes: a plurality of identical upsampling layers, a second noise-affine combination module, and a third noise-affine combination module. The input of the first upsampling module is the overall sentence feature, the overall image feature, and the local image feature. Among two adjacent upsampling layers in the multiple identical upsampling layers, the input of the latter upsampling layer is the output of the former upsampling layer; The second noisy affine combination module is located between any two upsampling layers in the multiple identical upsampling layers, and is used to perform feature fusion on the result output by the former upsampling layer and the local image feature among the any two upsampling layers; The third noisy affine combination module is used to perform feature fusion on the output result of the last upsampling layer in the multiple identical upsampling layers and the local image feature.

5. The method according to any one of claims 1 to 2, characterized in that The image editing model further includes: a detail correction model for modifying the details of the target edited image; The detail correction model is used to process the local image feature, the sentence word feature, and the target edited image to obtain a target corrected image; The detail correction model includes: a first detail correction module, a second detail correction module, a fusion module, and a generator. Among them, the first detail correction module is used to modify the details of the local image feature, the first random noise, and the sentence word feature to obtain a first detail feature; The second detail correction module is used to modify the details of the local image feature corresponding to the target edited image, the second random noise, and the sentence word feature to obtain a second detail feature; The fusion module is used to perform feature fusion on the first detail feature and the second detail feature; The generator is used to generate the target corrected image according to the output result of the fusion module.

6. The method according to claim 5, wherein The first detail correction module includes a fourth noisy affine combination module, a fifth noisy affine combination module, a sixth noisy affine combination module, a second self-attention module, a first residual network, and a first linear network; The fourth noisy affine combination module is used to perform feature fusion on the first random noise and the local image feature to obtain a first fusion feature; The second self-attention module is used to perform feature fusion on the first fusion feature and the sentence word feature; The fifth noisy affine combination module performs feature fusion on the concatenation result of the output result of the second self-attention module and the first random noise, and the local image feature; The first residual network is used to extract visual features from the output result of the fifth noisy affine combination module; The first linear network is used to perform a linear transformation on the local image feature; The sixth noisy affine combination module is used to perform feature fusion on the output result of the first residual network and the output result of the first linear network.

7. The method according to claim 5, wherein The second detail correction module includes a seventh noisy affine combination module, an eighth noisy affine combination module, a ninth noisy affine combination module, a third self-attention module, a second residual network, and a second linear network; The seventh noisy affine combination module is used to perform feature fusion on the second random noise and the local image feature corresponding to the target edited image to obtain a first fusion feature; The third self-attention module is used to perform feature fusion on the first fusion feature and the sentence word feature; The eighth noisy affine combination module performs feature fusion on the concatenation result of the output result of the third self-attention module and the second random noise, and the image local feature corresponding to the target edited image; The second residual network is used to extract visual features from the output result of the eighth noisy affine combination module; The second linear network is used to perform a linear transformation on the image local feature corresponding to the target edited image; The ninth noisy affine combination module is used to perform feature fusion on the output result of the second residual network and the output result of the second linear network.

8. The method according to claim 5, characterized in that, The training method of the detail correction model includes: Training the generator of the detail correction model according to a conditional generator loss function, an unconditional generator loss function, and a semantic contrast function; Training the discriminator of the detail correction model according to a conditional discriminator loss function and an unconditional discriminator loss function.

9. The method according to any one of claims 1 to 2, characterized in that, The training method of the image editing model includes: Training an initial model using a preset loss function and a training set to obtain the image editing model; Wherein, the preset loss function includes sub-functions respectively corresponding to N sub-networks and loss functions of N-1 autoencoders, the initial model includes N sub-networks, and the N sub-networks are initial models respectively corresponding to the sampling encoding module and at least one generation module; During the training process, when the output image of the i-th sub-network does not meet the preset condition, the sub-functions corresponding to the (i + 1)-th to N-th sub-networks and the loss functions of the i-th to (i + 1)-th autoencoders are used to train the initial model, where 0 ≤ i < N.

10. An electronic device, characterized in that, The device includes a processor and a memory, the memory is used to store a computer program, and the processor is used to call and run the computer program from the memory, so that the device executes the method according to any one of claims 1 to 9.

11. A computer-readable storage medium, characterized in that, A computer program is stored in the computer-readable storage medium, and when the computer program is executed by a processor, the processor executes the method according to any one of claims 1 to 9.

Citation Information

Patent Citations

  • Method for editing picture according to text based on generative adversarial network and dynamic editing module

    CN112818646A

  • Interactive image editing method and device, readable storage medium and electronic equipment

    CN113448477A