Image editing method and device
Through the autoregressive image generation model, the causal relationship between the first reference image and the second reference image is understood, and the target image of the source image is generated, which solves the problem of specifying the editing area in the prior art, improves the image editing efficiency and user experience, and the generated images are more in line with the causal relationship.
Patent Information
- Application Number
- CN202510615646.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-13
- Publication Date
- 2025-08-15
AI Technical Summary
The existing image editing methods have shortcomings in efficiency and user experience, especially the need for users to specify the area to be edited in the source image, resulting in inefficiency and poor user experience.
The autoregressive image generation model is adopted to obtain the causal relationship between the first reference image and the second reference image, and the target image corresponding to the source image is generated. The encoding module, the fusion module, the causal understanding module and the decoding module are used for image editing, including a multi-layer Transformer module and the self-attention module, to realize cross-modal attention interaction and progressive feature synthesis.
There is no need to specify the editing area in the source image in advance, which improves image editing efficiency and achieves a more intelligent editing effect, breaks through the limitations of the traditional image generation field, and the generated images are more realistic and in line with causal relationships.
Smart Images

Figure CN120495432A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of artificial intelligence technology, and in particular to an image editing method and device. Background Art
[0002] The rapid development and widespread application of deep learning technology has provided a powerful impetus for image editing. Users can input relevant instructions and source images, invoking deep learning models to edit the source images accordingly. However, current image editing methods leave much to be desired, both in terms of efficiency and user experience. Summary of the Invention
[0003] The present application provides an image editing method and device to improve image editing efficiency and user experience.
[0004] This application provides the following solutions:
[0005] According to a first aspect, there is provided an image editing method, the method comprising:
[0006] Acquire a first reference image, a second reference image, and a source image, wherein there is a causal relationship between the first reference image and the second reference image, and the causal relationship is used to represent an editing process performed when the first reference image is converted into the second reference image;
[0007] An autoregressive image generation model is used to generate a target image corresponding to the source image based on the first reference image and the second reference image, and the causal relationship exists between the source image and the target image.
[0008] According to an achievable method in an embodiment of the present application, the autoregressive image generation model includes an encoding module, a fusion module, a causal relationship understanding module, and a decoding module;
[0009] The step of generating a target image corresponding to the source image based on the first reference image and the second reference image by using an autoregressive image generation model includes:
[0010] Encoding the first reference image, the second reference image, and the source image using the encoding module to obtain an encoding result of the first reference image, an encoding result of the second reference image, and an encoding result of the source image, respectively;
[0011] Using the fusion module, fusing the encoding result of the first reference image, the encoding result of the second reference image, and the encoding result of the source image to obtain a fused feature representation;
[0012] Using the causal relationship understanding module, based on the fused feature representation, an autoregressive approach is used to predict the target image feature representation;
[0013] The target image feature representation is decoded using the decoding module to obtain the target image.
[0014] According to an achievable method in an embodiment of the present application, the causal relationship understanding module includes a multi-layer Transformer module and a prediction module;
[0015] The Transformer module performs attention processing on the input intermediate feature representation, the encoding result of the first reference image, and the encoding result of the second reference image in each round of prediction to obtain the intermediate feature representation output by the Transformer module;
[0016] The intermediate feature representation input to the first-layer Transformer module includes the fused feature representation and the feature representation of the predicted elements of the target image. The intermediate feature representation input to other Transformer modules is the intermediate feature representation output by the previous-layer Transformer module. The last-layer Transformer module outputs the intermediate feature representation to the prediction module.
[0017] The prediction module predicts the feature representation of the next element of the target image using the input intermediate feature representation;
[0018] The feature representation of the next element is added to the feature representation of the element predicted for the target image to provide it to the first-layer Transformer module to perform the next round of prediction until all elements of the target image are predicted.
[0019] According to an implementable method in an embodiment of the present application, the Transformer module includes a first self-attention module, a gated self-attention module, and a second self-attention module;
[0020] The first self-attention module performs a first self-attention process on the intermediate feature representation input to the Transformer module, and outputs the intermediate feature representation obtained after the first self-attention process to the gated self-attention module;
[0021] The gated self-attention module splices the encoding result of the first reference image, the encoding result of the second reference image, and the intermediate feature representation input to the gated self-attention module to obtain a spliced feature representation; performs a second self-attention process on the spliced feature representation, and obtains the intermediate feature representation after the second self-attention process; performs weighted processing on the intermediate feature representation input to the gated self-attention module and the intermediate feature representation after the second self-attention process, and outputs the intermediate feature representation obtained after the weighted processing to the second self-attention module; wherein the weight corresponding to the intermediate feature representation after the second self-attention process is determined by a gating parameter, and the gating parameter is learned when training the autoregressive image generation model;
[0022] The second self-attention module performs a third self-attention process on the input intermediate feature representation to obtain the intermediate feature representation output by the Transformer module.
[0023] According to an achievable manner in an embodiment of the present application, the method further includes: obtaining type information of the editing task; encoding the type information of the editing task using the encoding module to obtain an encoding result of the type information of the editing task;
[0024] Using the fusion module, fusing the encoding result of the first reference image, the encoding result of the second reference image, and the encoding result of the source image to obtain a fused feature representation includes:
[0025] The fusion module is used to fuse the encoding result of the first reference image, the encoding result of the second reference image, the encoding result of the source image, and the encoding result of the type information of the editing task to obtain a fused feature representation.
[0026] According to a second aspect, a method for training an autoregressive image generation model is provided, the method comprising:
[0027] Obtaining a training data set, the training data set including a plurality of training samples, each training sample including two image sample pairs having the same causal relationship, using the two images included in one of the image sample pairs as a first reference image sample and a second reference image sample, and using the two images included in the other image sample pair as a source image sample and a target image sample;
[0028] The first reference image sample, the second reference image sample and the source image sample are used as inputs of the autoregressive image generation model, and the target image sample is used as an output target of the autoregressive image generation model to train the autoregressive image generation model.
[0029] According to an achievable method in an embodiment of the present application, the two images included in the image sample pair are generated in the following manner:
[0030] Inputting a first image and a first editing instruction into an image editing model, obtaining a second image generated by the image editing model based on the first image and the first editing instruction, wherein the first image and the second image constitute a first image sample pair;
[0031] Inputting a third image and a second editing instruction into the image editing model, obtaining a fourth image generated by the image editing model based on the third image and the second editing instruction, wherein the third image and the fourth image constitute a second image sample pair;
[0032] The first image sample pair and the second image sample pair constitute a training sample;
[0033] The first editing instruction and the second editing instruction correspond to the same editing process, and the image editing model is implemented based on a diffusion model.
[0034] According to an achievable method in an embodiment of the present application, the autoregressive image generation model includes an encoding module, a fusion module, a causal relationship understanding module, and a decoding module;
[0035] Encoding the first reference image sample, the second reference image sample, and the source image sample using the encoding module to obtain encoding results of the first reference image sample, encoding results of the second reference image sample, and encoding results of the source image sample, respectively;
[0036] Using the fusion module, fusing the encoding result of the first reference image sample, the encoding result of the second reference image sample, and the encoding result of the source image sample to obtain a fused feature representation;
[0037] Using the causal relationship understanding module, based on the fused feature representation, an autoregressive approach is used to predict the target image feature representation;
[0038] The target image feature representation is decoded using the decoding module to obtain the target image.
[0039] According to an achievable method in an embodiment of the present application, the causal relationship understanding module includes a multi-layer Transformer module and a prediction module;
[0040] The Transformer module performs attention processing on the input intermediate feature representation, the encoding result of the first reference image sample, and the encoding result of the second reference image sample in each round of prediction to obtain the intermediate feature representation output by the Transformer module;
[0041] The intermediate feature representation input to the first-layer Transformer module includes the fused feature representation and the feature representation of the predicted elements of the target image. The intermediate feature representation input to other Transformer modules is the intermediate feature representation output by the previous-layer Transformer module. The last-layer Transformer module outputs the intermediate feature representation to the prediction module.
[0042] The prediction module predicts the feature representation of the next element of the target image using the input intermediate feature representation;
[0043] The feature representation of the next element is added to the feature representation of the element predicted for the target image to provide it to the first-layer Transformer module to perform the next round of prediction until all elements of the target image are predicted.
[0044] According to an implementable method in an embodiment of the present application, the Transformer module includes a first self-attention module, a gated self-attention module, and a second self-attention module;
[0045] The first self-attention module performs a first self-attention process on the intermediate feature representation input to the Transformer module, and outputs the intermediate feature representation obtained after the first self-attention process to the gated self-attention module;
[0046] The gated self-attention module splices the encoding result of the first reference image sample, the encoding result of the second reference image sample, and the intermediate feature representation input to the gated self-attention module to obtain a spliced feature representation; performs a second self-attention process on the spliced feature representation, and obtains the intermediate feature representation after the second self-attention process; performs weighted processing on the intermediate feature representation input to the gated self-attention module and the intermediate feature representation after the second self-attention process, and outputs the intermediate feature representation obtained after the weighted processing to the second self-attention module; wherein the weight corresponding to the intermediate feature representation after the second self-attention process is determined by a gating parameter; and the gating parameter is updated during the training process;
[0047] The second self-attention module performs a third self-attention process on the input intermediate feature representation to obtain the intermediate feature representation output by the Transformer module.
[0048] According to a third aspect, there is provided an image editing apparatus, the apparatus comprising:
[0049] an acquisition unit configured to acquire a first reference image, a second reference image, and a source image, wherein a causal relationship exists between the first reference image and the second reference image, and the causal relationship is used to represent an editing process performed when the first reference image is converted into the second reference image;
[0050] A generating unit is configured to generate a target image corresponding to the source image based on the first reference image and the second reference image using an autoregressive image generation model, wherein the causal relationship exists between the source image and the target image.
[0051] According to a fourth aspect, a device for training an autoregressive image generation model is provided, the device comprising:
[0052] a data acquisition unit configured to acquire a training data set, the training data set including a plurality of training samples, each training sample including two image sample pairs having the same causal relationship, the two images included in one of the image sample pairs being used as a first reference image sample and a second reference image sample, and the two images included in the other image sample pair being used as a source image sample and a target image sample;
[0053] The model training unit is configured to use the first reference image sample, the second reference image sample and the source image sample as inputs of the autoregressive image generation model, and use the target image sample as an output target of the autoregressive image generation model to train the autoregressive image generation model.
[0054] According to a fifth aspect, a computer-readable storage medium is provided, on which a computer program is stored. When the program is executed by a processor, the steps of the method described in any one of the first and second aspects are implemented.
[0055] According to a sixth aspect, an electronic device is provided, comprising:
[0056] one or more processors; and
[0057] A memory associated with the one or more processors, the memory being used to store program instructions, wherein the program instructions, when read and executed by the one or more processors, execute the steps of the method described in any one of the first and second aspects above.
[0058] According to a seventh aspect, a computer program product is provided, comprising a computer program, which, when executed by a processor, implements the steps of the method described in any one of the first and second aspects.
[0059] According to the specific embodiments provided in this application, this application discloses the following technical effects:
[0060] 1) In the solution provided by this application, the autoregressive image generation model is able to understand the causal relationship between a first reference image and a second reference image, generating a target image corresponding to the source image, thereby achieving the same editing effect as the aforementioned causal relationship on the source image. This approach not only eliminates the need to pre-define the editing area in the source image, but also improves image editing efficiency. Furthermore, causal-based editing is more intelligent, enhancing the user experience. Furthermore, it overcomes the limitation of traditional image generation, which only allows image editing based on a single reference image.
[0061] 2) The autoregressive image generation model of the present application includes an encoding module, a fusion module, a causal relationship understanding module, and a decoding module. The present application can first encode the first reference image, the second reference image, and the source image, then use an autoregressive method to predict the target image feature representation, and further decode the target image feature representation to obtain the target image. This method enables the causal relationship understanding module to more accurately learn the causal relationship between the first reference image and the second reference image, and then accurately edit the source image.
[0062] 3) The causal relationship understanding module in this application includes a multi-layer Transformer module and a prediction module. The multi-layer Transformer module implements cross-modal attention interaction and progressive feature synthesis, fully understanding the relationship between the first reference image, the second reference image, and the source image to generate a more accurate target image. Furthermore, based on the feature representation of the predicted element of the target image, the feature representation of the next element of the target image is predicted, resulting in a more realistic predicted target image.
[0063] 4) The Transformer module in this application may include a first self-attention module, a gated self-attention module, and a second self-attention module. The gated self-attention module can dynamically adjust the flow of information according to the input, determine which information in the first reference image and the second reference image should be paid attention to or suppressed, and thus achieve an understanding of the causal relationship between the first reference image and the second reference image, so that the autoregressive image generation model can decouple information in the first reference image and the second reference image that is not related to the "causal relationship".
[0064] 5) This application can obtain the type information of the editing task, encode the type information of the editing task, and then add the encoding result of the type information of the editing task to the fusion feature representation. In this way, the autoregressive image generation model can perform more accurate image editing on the source image based on the type information of the editing task.
[0065] 6) The present application provides a method for training an autoregressive image generation model. The autoregressive image generation model trained by this method can support the input of a first reference image and a second reference image that have a causal relationship. The autoregressive image generation model edits the source image after understanding the causal relationship. And based on the trained autoregressive image generation model, more delicate image editing tasks can be achieved, such as: the migration of characteristic scenes (such as meteorological scene migration), virtual makeup trials with reference to pictures of models before and after makeup, etc. The inference mode of the autoregressive image generation model designed in the present application can explicitly model causal variables (such as weather / material changes), breaking through the limitation of traditional methods that rely on single reference image conditional generation.
[0066] 7) The two images in the image sample pair can be obtained using the image editing model. This method can quickly and conveniently obtain a large training data set, laying a good foundation for training the autoregressive image generation model.
[0067] Of course, any product implementing the present application does not necessarily need to achieve all of the advantages described above at the same time. BRIEF DESCRIPTION OF THE DRAWINGS
[0068] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0069] Figure 1 is a system architecture diagram applicable to the embodiments of the present application;
[0070] Figure 2 A flowchart of the image editing method provided in an embodiment of the present application;
[0071] Figure 3 A schematic diagram of a first reference image and a second reference image provided in an embodiment of the present application;
[0072] Figure 4a A schematic diagram of one of the architectures of the autoregressive image generation model provided in an embodiment of the present application;
[0073] Figure 4b A schematic diagram of another architecture of the autoregressive image generation model provided in an embodiment of the present application;
[0074] Figure 5 A schematic diagram of the architecture of the causal relationship understanding model provided in an embodiment of the present application;
[0075] Figure 6A schematic diagram of the architecture of the Transformer module provided in an embodiment of the present application;
[0076] Figure 7 A flowchart of a method for training an autoregressive image generation model provided in an embodiment of the present application;
[0077] Figure 8 A schematic block diagram of an image editing device provided in an embodiment of the present application;
[0078] Figure 9 A schematic block diagram of an apparatus for training an autoregressive image generation model provided in an embodiment of the present application;
[0079] Figure 10 A schematic block diagram of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0080] The following will be combined with the accompanying drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field are within the scope of protection of this application.
[0081] The terms used in the embodiments of the present invention are only for the purpose of describing specific embodiments and are not intended to limit the present invention. The singular forms "a", "an", "the" and "the" used in the embodiments of the present invention and the appended claims are also intended to include plural forms unless the context clearly indicates otherwise.
[0082] It should be understood that the term "and / or" as used herein is merely a description of the relationship between associated objects, indicating that three possible relationships exist. For example, "A and / or B" can represent: A exists alone, A and B exist simultaneously, or B exists alone. Furthermore, the character " / " in this document generally indicates that the associated objects are in an "or" relationship.
[0083] The word "if," as used herein, may be interpreted as "at the time of" or "when" or "in response to determining" or "in response to detecting," depending on the context. Similarly, the phrases "if it is determined" or "if (stated condition or event) is detected" may be interpreted as "when it is determined" or "in response to the determination" or "when detecting (stated condition or event)" or "in response to detecting (stated condition or event)," depending on the context.
[0084] Currently, some image editing technologies already exist. For example, users input a source image and a single reference image, and specify the area to be edited in the source image. Through a diffusion model, the reference image is used as a guide to edit the specified editing area in the source image. For example, for a source image containing a person's image, the user can specify that the area to be edited in the source image is the hair of the person's image. The given single reference image can be for a certain hairstyle. Then, based on the given single reference image, the model replaces the hairstyle of the person's image in the area to be edited in the source image with the hairstyle in the reference image. However, this method often requires the user to specify the area to be edited in the source image, which reduces image editing efficiency and user experience.
[0085] In view of this, the present application provides a new approach. To facilitate understanding of the present application, the system architecture on which the present application is based is first described. Figure 1 An exemplary system architecture to which the embodiments of the present application can be applied is shown. Figure 1 As shown in , the system architecture may include: a user device, an image editing device located at the image editing application server end, and a device for training an autoregressive image generation model.
[0086] The user equipment and the server can communicate with each other. The user equipment and the server can be directly or indirectly connected through wired or wireless communication, which is not limited in this application.
[0087] User devices may include, but are not limited to, smart mobile terminals, smart home devices, wearable devices, and personal computers (PCs). Smart mobile devices may include mobile phones, tablets, laptops, PDAs (Personal Digital Assistants), and internet-connected cars. Smart home devices may include smart TVs and smart refrigerators. Wearable devices may include smart watches, smart glasses, virtual reality devices, augmented reality devices, and mixed reality devices (i.e., devices that support both virtual reality and augmented reality).
[0088] Image editing application servers can be standalone servers, server clusters, or even cloud servers. Cloud servers, also known as cloud computing servers or cloud hosts, are a type of hosting product within the cloud computing service ecosystem. They address the management difficulties and limited scalability of traditional physical hosting and virtual private servers (VPS).
[0089] Before performing an image editing task, the autoregressive image generation model used by the image editing device may be trained by a device for training an autoregressive image generation model to obtain a trained autoregressive image generation model.
[0090] As one possible implementation, a user can input a source image, a first reference image, and a second reference image through a user device. The user device then carries the source image, the first reference image, and the second reference image in an image editing request and sends it to an image editing device on a server via a network. The image editing device generates a target image corresponding to the source image using the method provided in an embodiment of the present application and returns the target image to the user device via the network. The user device then displays the received target image to the user.
[0091] As another feasible method, a user can input the source image, information about the first reference image, and information about the second reference image through a user device. The user device then carries the source image, information about the first reference image, and information about the second reference image in an image editing request and sends the request to the image editing device on the server side via the network. The image editing device obtains the corresponding first reference image and second reference image based on the information about the first reference image and the second reference image, and generates a target image corresponding to the source image based on the first reference image and the second reference image using the method provided in the embodiment of the present application, and returns the target image to the user device via the network. The user device then displays the received target image to the user.
[0092] Apart from Figure 1 In addition to the illustrated architecture, a computer terminal with strong computing power may also adopt the method provided in the embodiment of the present application to obtain a source image, a first reference image, and a second reference image, and then generate a target image corresponding to the source image.
[0093] It should be understood that Figure 1 The user equipment and server end in the figure are only illustrative. According to the implementation requirements, there can be any number of user equipment and server ends.
[0094] It should be noted that the limitations such as "first" and "second" involved in this disclosure do not have restrictions on size, order and quantity, but are only used to distinguish them in name. For example, "first reference image" and "second reference image" are used to distinguish two reference images.
[0095] Figure 2 This is a flow chart of the image editing method provided in the embodiment of the present application. The method can be performed by Figure 1 The image editing device in the system shown is executed. Figure 2 As shown in , the method may include the following steps:
[0096] Step 201: Acquire a first reference image, a second reference image, and a source image. There is a causal relationship between the first reference image and the second reference image, and the causal relationship is used to characterize the editing process of converting the first reference image into the second reference image.
[0097] Step 202 : Generate a target image corresponding to a source image based on a first reference image and a second reference image using an autoregressive image generation model, wherein a causal relationship exists between the source image and the target image.
[0098] As can be seen from the above process, the autoregressive image generation model in the solution provided by this application is able to understand the causal relationship between the first and second reference images and generate a target image corresponding to the source image, thereby achieving the same editing effect as the above causal relationship on the source image. On the one hand, this method does not require the editing area to be specified in the source image in advance, which improves image editing efficiency. On the other hand, editing based on causal relationships is more intelligent, improving the user experience. And finally, it breaks through the limitation of traditional image generation that can only perform image editing based on a single reference image.
[0099] The following describes in detail the steps in the above process and the effects that can be further produced in conjunction with the embodiments.
[0100] First, the above step 201, namely “obtaining the first reference image, the second reference image and the source image”, is described in detail with reference to the embodiment.
[0101] In an embodiment of the present application, a first reference image, a second reference image, and a source image are obtained. A causal relationship exists between the first reference image and the second reference image. The causal relationship is used to characterize the editing process that the first reference image undergoes to be converted into the second reference image. The editing process here can be adding a certain attribute, modifying a certain attribute, or deleting a certain attribute, etc. The source image can be an image that is completely different from the first reference image and the second reference image.
[0102] The source image can be a Figure 1 The source image may be uploaded by the user device, that is, the server receives the source image uploaded by the user device; or the source image information (such as the source image identification, number, etc.) may be input by the user, and after the server receives the source image information sent by the user, the server obtains the corresponding image as the source image from the database of the server according to the source image information.
[0103] The first reference image and the second reference image can be obtained by the user through Figure 1 The first reference image and the second reference image uploaded by the user device in the server side are received by the user device. Figure 1The server receives the information of the reference image pair selected by the user (such as the identification and number of the reference image pair), and obtains the corresponding images from the database on the server as the first reference image and the second reference image based on the information of the reference image pair. Alternatively, the first reference image and the second reference image are determined based on the type of editing task selected by the user, that is, the server receives the type information of the editing task selected by the user (such as the identification and number of the type of editing task), and obtains the corresponding images from the database on the server as the first reference image and the second reference image based on the type information of the editing task. For example, there can be multiple types of editing tasks, such as adding, deleting, modifying, etc. When the user selects the type of editing task as adding, the first reference image and the second reference image corresponding to the adding type are determined, such as Figure 3 As shown, the conversion from the first reference image to the second reference image requires the addition of a "castle".
[0104] In addition, users can also input the type of editing task information through their device. The server receives the type of editing task information, and the autoregressive image generation model can then further perform image editing based on the type of editing task information. The type of image editing information is usually in the form of text.
[0105] Next, the above step 202, i.e., “generating a target image corresponding to the source image based on the first reference image and the second reference image by using the autoregressive image generation model”, is described in detail with reference to an embodiment.
[0106] In an embodiment of the present application, an autoregressive image generation model is used to generate a target image corresponding to a source image based on a first reference image and a second reference image. The causal relationship between the source image and the target image is the same as the causal relationship between the first reference image and the second reference image. For example, if the editing process of converting the first reference image to the second reference image is to add a "castle", then the editing process of converting the source image to the target image is also to add a "castle".
[0107] Autoregressive Image Generation Models (AIGM) are a type of image generation model based on the autoregressive mechanism. The core idea is to generate each element of the image (which can be pixels, tiles, etc.) one by one in sequence. Each element depends on the information of the previously generated elements, and image generation is achieved by modeling the conditional probability distribution between tiles.
[0108] As one of the feasible solutions, the autoregressive image generation model in this application may include an encoding module, a fusion module, a causal relationship understanding module and a decoding module, such as Figure 4aThe core idea is to encode the causal relationship between the first reference image and the second reference image as a transferable latent variable, and to leverage the multimodal understanding and generation capabilities of the autoregressive image generation model to achieve high-fidelity, causally constrained image editing.
[0109] When generating a target image using an autoregressive image generation model, the encoding module is first used to encode the first reference image, the second reference image, and the source image, and the encoding results of the first reference image, the second reference image, and the source image are obtained respectively. Among them, the encoding module may include an image encoder. The present application can use the image encoder to encode the first reference image, the second reference image, and the source image. Depending on the usage scenario, the image encoder can adopt networks with different architectures such as VQ-VAE (Vector Quantized Variational Autoencoder), VQGAN (Vector Quantized Generative Adversarial Network), 1D Tokenizer (one-dimensional word segmenter), etc.
[0110] The encoding process can be specifically described by first constructing a codebook containing predefined discrete vectors that represent different features in the image. The image to be encoded is then divided into fixed-size patches. Each patch is mapped to a high-dimensional vector representation using an image encoder (such as a convolutional neural network). These high-dimensional vectors are then quantized into discrete tokens in the codebook, with each discrete token corresponding to a discrete vector in the codebook. After quantization, the image to be encoded is converted into a series of discrete tokens. This token sequence serves as input to the fusion module.
[0111] Next, the fusion module is used to fuse the encoding results of the first reference image, the encoding results of the second reference image, and the encoding results of the source image to obtain a fused feature representation. It should be noted that the fusion processing here can be directly splicing the encoding results of the first reference image, the encoding results of the second reference image, and the encoding results of the source image to obtain a fused feature representation; or it can be fused using a cross-attention mechanism to obtain a fused feature representation.
[0112] Furthermore, the causal relationship understanding module is utilized to predict the target image feature representation based on the fused feature representation using an autoregressive approach. Specifically, the causal relationship understanding module of the present application can predict the feature representation of the first element of the target image in the first round of prediction based on the fused feature representation, and predict the feature representation of the next element of the target image based on the fused feature representation and the feature representation of the predicted elements of the target image in subsequent rounds of prediction, until the feature representations of all elements of the target image, i.e., the target image feature representation, are predicted.
[0113] Finally, the decoding module is used to decode the target image feature representation to obtain the target image, wherein the target image feature representation can be a series of discrete tokens, and the decoding module can decode these discrete token sequences back to the original image data form.
[0114] In order to enable the autoregressive image generation model to better understand the causal relationship between the first reference image and the second reference image to generate a more accurate target image, the causal relationship understanding module of the present application may include a multi-layer Transformer module and a prediction module, such as Figure 5 As shown in the figure, the Transformer module is used to perform attention processing on the intermediate feature representations, while the prediction module is used to predict the feature representation of the next element in the target image based on the input intermediate feature representations. Attention processing is a neural network mechanism for processing sequential data, widely used in natural language processing (NLP) and computer vision. Its core idea is to enable the causal relationship understanding module to dynamically pay attention to other elements in the sequence when processing a certain element, thereby capturing the dependencies between elements.
[0115] Specifically, in each round of prediction, the Transformer module uses the input intermediate feature representation, the encoding result of the first reference image, and the encoding result of the second reference image to perform attention processing to obtain the intermediate feature representation output by the Transformer module. The intermediate feature representation input to the first-layer Transformer module includes the fused feature representation and the feature representation of the elements predicted for the target image, while the intermediate feature representation input to the Transformer modules of other layers is the intermediate feature representation output by the previous layer Transformer module, and the intermediate feature representation output by the last layer Transformer module is sent to the prediction module.
[0116] That is, the first-layer Transformer module is input with the fused feature representation and the feature representation of the elements predicted for the target image (for the first round of prediction, since the target image has not yet been predicted, the feature representation of the predicted elements is empty), and after performing processing such as self-attention, the intermediate feature representation is output to the next-layer Transformer module.
[0117] The Transformer module in the middle layer is input with the fused feature representation output by the Transformer module in the previous layer, and after performing processing such as self-attention, the intermediate feature representation is output to the Transformer module in the next layer.
[0118] The last layer of Transformer module is fed with the fused feature representation output by the previous layer of Transformer module, and after processing such as self-attention, it outputs the intermediate feature representation to the prediction module.
[0119] Furthermore, the prediction module uses the input intermediate feature representation to predict the feature representation of the next element of the target image, and adds the feature representation of the next element to the feature representation of the elements predicted for the target image to provide it to the first-layer Transformer module to perform the next round of prediction until all elements of the target image are predicted.
[0120] In essence, the above intermediate features are represented as latent vectors obtained during the model processing process, which can eliminate redundant information and retain key semantics.
[0121] In other words, for each prediction round, the causal understanding module uses multiple layers of Transformer modules and the prediction module to output a feature representation for an element in the target image. After completing all prediction rounds, the feature representations of all elements in the target image are fully predicted. It should be noted that, except for the feature representation of the first element in the target image, the feature representations of all other elements are combined with the feature representations of the predicted elements.
[0122] The Transformer module in the above process can include one or more self-attention modules. A self-attention module can capture the long-distance dependencies of the input intermediate feature representations, but may be limited by its processing power. By stacking multiple self-attention modules, more complex dependencies can be captured. Each self-attention module can further refine the attention distribution based on the previous self-attention module, thereby more accurately capturing the long-distance dependencies of the input intermediate feature representations, and then gradually extracting and abstracting the features in the input sequence, allowing the causal understanding module to learn more advanced and abstract representations. In addition, each self-attention module can include a residual connection, allowing the causal understanding module to learn residual mapping, which helps alleviate the gradient vanishing problem and stabilize the training process.
[0123] In order to further extract the causal relationship between the first reference image and the second reference image, the Transformer module of the present application may include a gated self-attention module. As one feasible solution, the Transformer module may include two self-attention modules (hereinafter referred to as the first self-attention module and the second self-attention module) and a gated self-attention module, such as Figure 6 shown.
[0124] In a preferred embodiment of the present application, the intermediate feature representation input to the Transformer module can be first subjected to a first self-attention processing through a first self-attention module, and the intermediate feature representation obtained after the first self-attention processing can be output to a gated self-attention module.
[0125] The gated self-attention module concatenates the encoding results of the first reference image, the encoding results of the second reference image, and the intermediate feature representation input to the gated self-attention module to obtain a concatenated feature representation. The concatenated feature representation is then subjected to a second self-attention process to obtain an intermediate feature representation after the second self-attention process. The intermediate feature representation input to the gated self-attention module and the intermediate feature representation after the second self-attention process are weighted, and the weighted intermediate feature representation is output to the second self-attention module. The weights corresponding to the intermediate feature representation after the second self-attention process are determined by gating parameters, which are learned during training of the autoregressive image generation model.
[0126] The process of obtaining weighted intermediate feature representation using the gated self-attention module can be expressed as follows:
[0127]
[0128] Among them, ν on the left side of the formula represents the intermediate feature representation after weighted processing, ν on the right side of the formula represents the intermediate feature representation input to the gated self-attention module, γ is the gating parameter (learned when training the autoregressive image generation model), and β is the hyperparameter of the injection strength (usually a fixed value). Represents the splicing feature representation, SelfAttn The concatenated feature represents the intermediate feature representation after the second self-attention processing. Here, the intermediate feature representation after the second self-attention processing only retains the part of the intermediate feature representation input to the gated self-attention module.
[0129] Furthermore, the intermediate feature representation obtained after weighted processing of the gated self-attention module output is input into the second self-attention module, and the second self-attention module performs a third self-attention process on the input intermediate feature representation to obtain the intermediate feature representation output by the Transformer module.
[0130] It should be noted that the number of gated self-attention modules and self-attention modules included in the Transformer module in the above embodiment is not fixed. In actual applications, the number and order of gated self-attention modules and self-attention modules can be adjusted according to actual needs.
[0131] It should also be noted that when the first reference image and the second reference image are input into the autoregressive image generation model, since there is a causal relationship between the first reference image and the second reference image, the present application can input the first reference image and the second reference image into the autoregressive image generation model in sequence, or manually annotate the sequence of the first reference image and the second reference image, so that the causal relationship learned by the autoregressive image generation model is from the first reference image to the second reference image. For example, for Figure 3 For the two reference images in the image, if the order is from the first reference image to the second reference image, the causal relationship learned by the autoregressive image generation model is the editing process of adding the "castle". If the order is from the second reference image to the first reference image, the causal relationship learned by the autoregressive image generation model is the editing process of deleting the "castle".
[0132] Furthermore, the present application may also obtain type information of the editing task, wherein the type of the editing task may be, for example, adding a certain attribute, modifying a certain attribute, or deleting a certain attribute.
[0133] In one achievable embodiment, a user uploads a source image to a preset location on a display page and selects the type of editing task they want. Based on the user's selection, information about the editing task type is obtained. Based on the editing task type information, a first reference image and a second reference image corresponding to the editing task type are matched from a preset storage space.
[0134] As another feasible embodiment, the user selects the type of required editing task on the display page, and uploads the source image, the first reference image, and the second reference image at a preset location.
[0135] Furthermore, if Figure 4b As shown in , the encoding module is used to encode the type information of the editing task to obtain an encoding result of the type information of the editing task. In an embodiment of the present application, in order to accurately encode information of two different modalities, image and text, as a preferred implementation scheme, the encoding module includes an image encoder and a text encoder. The image encoder is used to encode the first reference image, the second reference image, and the source image to obtain an encoding result of the first reference image, an encoding result of the second reference image, and an encoding result of the source image. The text encoder is used to encode the type information of the editing task to obtain an encoding result of the type information of the editing task.
[0136] When fusing the encoding, a fusion module is used to fuse the encoding result of the first reference image, the encoding result of the second reference image, the encoding result of the source image, and the encoding result of the type information of the editing task to obtain a fusion feature representation.
[0137] Regarding the process of encoding the type information of the editing task using a text encoder, the present application can convert the type information of the editing task into a unified format, for example, converting the text to lowercase and removing punctuation or special characters. The text corresponding to the type information of the editing task is then segmented into basic semantic units (tokens), such as words, subwords, or characters, and each token is assigned a unique integer index for subsequent processing.
[0138] The autoregressive image generation model in the above process can be trained in advance. Figure 7 This is a flowchart of a method for training an autoregressive image generation model provided in an embodiment of the present application. The method can be performed by Figure 1 The apparatus for training the autoregressive image generation model in the system shown is executed. Figure 7 As shown in , the method may include the following steps:
[0139] S701: Obtain a training data set, where the training data set includes multiple training samples, each training sample includes two image sample pairs with the same causal relationship, and the two images included in one of the image sample pairs are used as the first reference image sample and the second reference image sample, respectively, and the two images included in the other image sample pair are used as the source image sample and the target image sample, respectively.
[0140] S702: Using the first reference image sample, the second reference image sample, and the source image sample as inputs of an autoregressive image generation model, using the target image sample as an output target of the autoregressive image generation model, and training the autoregressive image generation model.
[0141] As can be seen from the above process, this application provides a method for training an autoregressive image generation model. The autoregressive image generation model trained by this method can support the input of a first reference image and a second reference image with a causal relationship. After the autoregressive image generation model understands the causal relationship, it edits the source image. Moreover, based on the trained autoregressive image generation model, more sophisticated image editing tasks can be achieved, such as: the migration of characteristic scenes (such as meteorological scene migration), virtual makeup trials with reference to pictures of models before and after makeup, etc.
[0142] The following describes in detail the steps in the above process and the effects that can be further produced in conjunction with the embodiments.
[0143] First, the above step 701, namely “obtaining a training data set”, is described in detail with reference to an embodiment.
[0144] An embodiment of the present application can obtain a training data set, which includes multiple training samples in the form of quads, each training sample including a source image sample, a target image sample, a first reference image sample, and a second reference image sample, wherein the causal relationship between the source image sample and the target image sample is the same as the causal relationship between the first reference image sample and the second reference image sample. For example, if the editing process of converting the first reference image sample into the second reference image sample is "adding rainy day elements", then the editing process of converting the source image sample into the target image sample is also "adding rainy day elements".
[0145] That is to say, the present application can use two image sample pairs with the same causal relationship as a training sample, that is, the two images included in one of the image sample pairs are used as the first reference image sample and the second reference image sample, and the two images included in the other image sample pair are used as the source image sample and the target image sample.
[0146] In order to quickly and easily obtain a large amount of training data sets, the present application can use the image editing model in the prior art to generate image sample pairs. The image editing model can be implemented based on the diffusion model, such as paint-by-example, magicbrush, etc., among which Paint-by-Example is an example-based image editing project. Paint-by-Example uses diffusion models to implement image editing. It decouples and reorganizes the source image from the reference image through self-supervised training, thereby achieving high-fidelity image editing; MagicBrush is also built on an advanced deep learning framework and uses the diffusion model as one of its core technologies. It combines natural language processing and computer vision technology to improve the performance of the image editing model by providing detailed editing instructions and target images.
[0147] The two images included in the image sample pair are generated in the following manner: a first image and a first editing instruction are input into the image editing model, a second image generated by the image editing model based on the first image and the first editing instruction is obtained, and a first image sample pair is formed based on the first image and the second image. A third image and a second editing instruction are input into the image editing model, a fourth image generated by the image editing model based on the third image and the second editing instruction is obtained, and a second image sample pair is formed based on the third image and the fourth image. Then, the first image sample pair and the second image sample pair are used as a training sample. Among them, one image sample pair provides a cross-sample causal reasoning basis for the autoregressive image generation model by explicitly representing causal variables (such as weather conditions, object attributes or scene structure), and the other image sample pair requires that the image editing process strictly follows the causal law derived from the reference image pair.
[0148] It should be noted that the first editing instruction and the second editing instruction correspond to the same editing process. For example, the editing process from the first image to the second image can be "adding rainy elements", then the editing process from the third image to the fourth image should also be "adding rainy elements".
[0149] It should also be noted that in addition to being obtained using image editing models, training samples can also be manually selected from an image library, or by inputting question descriptions (Prompt) into a large language model so that the large language model outputs image sample pairs with a causal relationship.
[0150] Next, in conjunction with the embodiment, the above-mentioned step 702, namely, "using the first reference image sample, the second reference image sample and the source image sample as the input of the autoregressive image generation model, using the target image sample as the output target of the autoregressive image generation model, and training the autoregressive image generation model" is described in detail.
[0151] In an embodiment of the present application, the first reference image sample, the second reference image sample, and the source image sample are used as inputs of the autoregressive image generation model, and the target image sample is used as the output target of the autoregressive image generation model. That is, the source image, the first reference image, and the second reference image in the training sample are used as inputs of the autoregressive image generation model to obtain the target image predicted by the autoregressive image generation model. In each round of iteration, the difference between the predicted target image and the target image sample in the training sample is used to obtain the loss function, and the value of the loss function is used to update the parameters of the autoregressive image generation model using a method such as gradient descent until the training end condition is reached. The training end condition can be, for example, that the loss function is less than or equal to a preset loss function threshold, or that the number of iterations reaches a preset round number threshold.
[0152] based on Figure 7 The training method is used to train the autoregressive image generation model, and the trained autoregressive image generation model is used to implement Figure 2 The method shown enables the autoregressive image generation model to learn the causal relationship between the first reference image and the second reference image, perform image editing on the source image, and obtain the target image.
[0153] The autoregressive image generation model may include an encoding module, a fusion module, a causal relationship understanding module, and a decoding module. During the training process, the encoding module is used to encode the first reference image sample, the second reference image sample, and the source image sample, respectively obtaining the encoding results of the first reference image sample, the second reference image sample, and the source image sample. The fusion module is then used to fuse the encoding results of the first reference image sample, the second reference image sample, and the source image sample to obtain a fused feature representation. The causal relationship understanding module is then used to predict the target image feature representation using an autoregressive approach based on the fused feature representation. Finally, the decoding module is used to decode the target image feature representation to obtain the target image.
[0154] Furthermore, the causal relationship understanding module includes a multi-layer Transformer module and a prediction module.
[0155] In each round of prediction, the Transformer module uses the input intermediate feature representation, the encoding results of the first reference image samples, and the encoding results of the second reference image samples to perform attention processing to obtain the intermediate feature representation output by the Transformer module. The intermediate feature representation input to the first-layer Transformer module includes the fused feature representation and the feature representation of the predicted elements of the target image. The intermediate feature representation input to other Transformer modules is the intermediate feature representation output by the previous layer Transformer module. The last layer Transformer module outputs the intermediate feature representation to the prediction module.
[0156] The prediction module uses the input intermediate feature representation to predict the feature representation of the next element of the target image, and adds the feature representation of the next element to the feature representation of the elements predicted for the target image to provide it to the first-layer Transformer module for the next round of prediction until all elements of the target image are predicted.
[0157] Furthermore, the Transformer module includes a first self-attention module, a gated self-attention module and a second self-attention module. The first self-attention module performs a first self-attention processing on the intermediate feature representation input to the Transformer module, and outputs the intermediate feature representation obtained after the first self-attention processing to the gated self-attention module.
[0158] The gated self-attention module concatenates the encoding results of the first reference image sample, the encoding results of the second reference image sample, and the intermediate feature representation input to the gated self-attention module to obtain a concatenated feature representation. The concatenated feature representation undergoes a second self-attention process, and an intermediate feature representation after the second self-attention process is obtained. The intermediate feature representation input to the gated self-attention module and the intermediate feature representation after the second self-attention process are weighted, and the resulting intermediate feature representation is output to the second self-attention module. The weights corresponding to the intermediate feature representations after the second self-attention process are determined by gating parameters, which are updated during training.
[0159] The second self-attention module performs the third self-attention processing on the intermediate feature representation input to obtain the intermediate feature representation output by the Transformer module.
[0160] The specific structure of the autoregressive image generation model mentioned above has been described in detail in step 202 and will not be repeated here.
[0161] It should be noted that before training the model, to help the autoregressive image generation model understand causal relationships, simple text annotations can be performed on the training samples, namely, the type of editing task corresponding to each training sample can be labeled. When training the model, the first reference image sample, the second reference image sample, the source image sample, and the corresponding editing task type information can be used as the input of the autoregressive image generation model, and the target image sample can be used as the output target of the autoregressive image generation model to train the autoregressive image generation model.
[0162] In this case, the encoding module of the autoregressive image generation model can include two encoders: a text encoder and an image encoder. The text encoder is used to encode the type information of the editing task, and the image encoder is used to encode the first reference image sample, the second reference image sample, and the source image sample.
[0163] The above method provided in the embodiment of the present application can be applied to a variety of application scenarios, including but not limited to: for virtual makeup trial, by referring to the pictures of the model before and after makeup, the effects of different makeup on the user's face can be simulated to realize the virtual makeup trial function, which not only improves the consumer's shopping experience, but also helps beauty brands better demonstrate the product effects; for clothing matching, by referring to the images before and after changing clothes, different matching effects can be displayed to help users make more appropriate purchasing decisions; for scene migration, elements in the featured scene can be synthesized into a new scene to create a unique visual effect. For example, by referring to images containing and excluding the starry sky, the city night scene can be synthesized with the starry sky.
[0164] The foregoing description of this specification describes specific embodiments. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be performed in an order different from that described in the embodiments and still achieve the desired results. Furthermore, the processes depicted in the accompanying drawings do not necessarily require the specific order shown or the sequential order to achieve the desired results. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0165] Figure 8 A schematic block diagram of an image editing device according to an embodiment is shown. Figure 1 The server side of the architecture shown in Figure 1. Figure 8 As shown, the apparatus 800 includes:
[0166] The acquisition unit 801 is configured to acquire a first reference image, a second reference image and a source image. There is a causal relationship between the first reference image and the second reference image, and the causal relationship is used to represent the editing process of converting the first reference image into the second reference image.
[0167] The generating unit 802 is configured to generate a target image corresponding to a source image based on a first reference image and a second reference image using an autoregressive image generating model, where a causal relationship exists between the source image and the target image.
[0168] As one of the possible implementation methods, the autoregressive image generation model includes an encoding module, a fusion module, a causal relationship understanding module and a decoding module.
[0169] The generation unit 802, when using the autoregressive image generation model to generate a target image corresponding to the source image based on the first reference image and the second reference image, can be configured as follows: using the encoding module to encode the first reference image, the second reference image and the source image to obtain the encoding results of the first reference image, the encoding results of the second reference image and the encoding results of the source image respectively; using the fusion module to fuse the encoding results of the first reference image, the encoding results of the second reference image and the encoding results of the source image to obtain a fused feature representation; using the causal relationship understanding module to predict the target image feature representation based on the fused feature representation using an autoregressive method; using the decoding module to decode the target image feature representation to obtain the target image.
[0170] As one of the possible implementation methods, the causal relationship understanding module includes multiple layers of Transformer modules and prediction modules; in each round of prediction, the Transformer module uses the input intermediate feature representation, the encoding result of the first reference image, and the encoding result of the second reference image to perform attention processing to obtain the intermediate feature representation output by the Transformer module; wherein, the intermediate feature representation input to the first-layer Transformer module includes the fused feature representation and the feature representation of the elements predicted for the target image, and the intermediate feature representation input to other Transformer modules is the intermediate feature representation output by the previous layer Transformer module, and the last layer Transformer module outputs the intermediate feature representation to the prediction module; the prediction module uses the input intermediate feature representation to predict the feature representation of the next element of the target image; the feature representation of the next element is added to the feature representation of the elements predicted for the target image, so as to provide it to the first-layer Transformer module to perform the next round of prediction until all elements of the target image are predicted.
[0171] As one of the possible implementation methods, the Transformer module includes a first self-attention module, a gated self-attention module and a second self-attention module; the first self-attention module performs a first self-attention process on the intermediate feature representation input to the Transformer module, and outputs the intermediate feature representation obtained after the first self-attention process to the gated self-attention module; the gated self-attention module splices the encoding result of the first reference image, the encoding result of the second reference image and the intermediate feature representation input to the gated self-attention module to obtain a spliced feature representation; the spliced feature representation performs a second self-attention process and obtains the intermediate feature representation after the second self-attention process; the intermediate feature representation input to the gated self-attention module and the intermediate feature representation after the second self-attention process are weighted, and the intermediate feature representation obtained after the weighted process is output to the second self-attention module; wherein the weight corresponding to the intermediate feature representation after the second self-attention process is determined by the gating parameter, which is learned when training the autoregressive image generation model; the second self-attention module performs a third self-attention process on the intermediate feature representation input to obtain the intermediate feature representation output by the Transformer module.
[0172] As one possible implementation method, the acquisition module 801 may also be configured to: acquire type information of the editing task.
[0173] The generating module 802 may also be configured to: utilize the encoding module to encode the type information of the editing task to obtain an encoding result of the type information of the editing task.
[0174] The generation unit 802, when using the fusion module to fuse the encoding result of the first reference image, the encoding result of the second reference image and the encoding result of the source image to obtain the fused feature representation, can be configured to: use the fusion module to fuse the encoding result of the first reference image, the encoding result of the second reference image, the encoding result of the source image and the encoding result of the type information of the editing task to obtain the fused feature representation.
[0175] According to another embodiment, an apparatus for training an autoregressive image generation model is provided. Figure 9 A schematic block diagram of a device for training an autoregressive image generation model according to an embodiment is shown, wherein the device is arranged at Figure 1 The server side of the architecture shown in Figure 1. Figure 9 As shown, the apparatus 900 includes:
[0176] The data acquisition unit 901 is configured to acquire a training data set, where the training data set includes multiple training samples, each training sample includes two image sample pairs with the same causal relationship, the two images included in one of the image sample pairs are used as the first reference image sample and the second reference image sample, and the two images included in the other image sample pair are used as the source image sample and the target image sample.
[0177] The model training unit 902 is configured to use the first reference image sample, the second reference image sample and the source image sample as inputs of the autoregressive image generation model, and use the target image sample as the output target of the autoregressive image generation model to train the autoregressive image generation model.
[0178] As one of the feasible methods, when generating the two images included in the image sample pair, the data acquisition unit 901 can be configured as follows: inputting the first image and the first editing instruction into the image editing model, obtaining the second image generated by the image editing model based on the first image and the first editing instruction, and the first image and the second image constitute a first image sample pair; inputting the third image and the second editing instruction into the image editing model, obtaining the fourth image generated by the image editing model based on the third image and the second editing instruction, and the third image and the fourth image constitute a second image sample pair; the first image sample pair and the second image sample pair constitute a training sample; wherein, the first editing instruction and the second editing instruction correspond to the same editing process, and the image editing model is implemented based on a diffusion model.
[0179] As one of the feasible methods, the autoregressive image generation model includes an encoding module, a fusion module, a causal relationship understanding module and a decoding module.
[0180] The model training unit 902 can be configured to: use the encoding module to encode the first reference image sample, the second reference image sample and the source image sample to obtain the encoding results of the first reference image sample, the encoding results of the second reference image sample and the encoding results of the source image sample, respectively; use the fusion module to fuse the encoding results of the first reference image sample, the encoding results of the second reference image sample and the encoding results of the source image sample to obtain a fused feature representation; use the causal relationship understanding module to predict the target image feature representation based on the fused feature representation using an autoregressive method; use the decoding module to decode the target image feature representation to obtain the target image.
[0181] As one of the feasible methods, the causal relationship understanding module includes a multi-layer Transformer module and a prediction module; in each round of prediction, the Transformer module uses the input intermediate feature representation, the encoding result of the first reference image sample and the encoding result of the second reference image sample to perform attention processing to obtain the intermediate feature representation output by the Transformer module; wherein, the intermediate feature representation input to the first-layer Transformer module includes the fused feature representation and the feature representation of the elements predicted for the target image, the intermediate feature representation input to other Transformer modules is the intermediate feature representation output by the previous layer Transformer module, and the last layer Transformer module outputs the intermediate feature representation to the prediction module; the prediction module uses the input intermediate feature representation to predict the feature representation of the next element of the target image; the feature representation of the next element is added to the feature representation of the elements predicted for the target image, so as to provide it to the first-layer Transformer module to perform the next round of prediction until all elements of the target image are predicted.
[0182] As one of the feasible methods, the Transformer module includes a first self-attention module, a gated self-attention module and a second self-attention module; the first self-attention module performs a first self-attention process on the intermediate feature representation input to the Transformer module, and outputs the intermediate feature representation obtained after the first self-attention process to the gated self-attention module; the gated self-attention module splices the encoding result of the first reference image sample, the encoding result of the second reference image sample and the intermediate feature representation input to the gated self-attention module to obtain a spliced feature representation; the spliced feature representation performs a second self-attention process and obtains the intermediate feature representation after the second self-attention process; the intermediate feature representation input to the gated self-attention module and the intermediate feature representation after the second self-attention process are weighted, and the intermediate feature representation obtained after the weighted process is output to the second self-attention module; wherein the weight corresponding to the intermediate feature representation after the second self-attention process is determined by the gating parameter; the gating parameter is updated during the training process; the second self-attention module performs a third self-attention process on the intermediate feature representation input to obtain the intermediate feature representation output by the Transformer module.
[0183] Each embodiment in this specification is described in a progressive manner. The same or similar parts between the embodiments can be referred to each other. Each embodiment focuses on the differences from other embodiments. In particular, for the device embodiment, since it is basically similar to the method embodiment, the description is relatively simple. For the relevant parts, refer to the partial description of the method embodiment. The device embodiment described above is merely illustrative, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. A person of ordinary skill in the art can understand and implement it without expending creative work.
[0184] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with the relevant laws, regulations and standards of relevant countries and regions, and provide corresponding operation entrances for users to choose to authorize or refuse.
[0185] In addition, an embodiment of the present application further provides a computer-readable storage medium on which a computer program is stored. When the program is executed by a processor, the steps of any one of the methods in the aforementioned method embodiments are implemented.
[0186] And an electronic device comprising:
[0187] one or more processors; and
[0188] A memory associated with one or more processors, the memory being used to store program instructions, which, when read and executed by one or more processors, execute the steps of any one of the method embodiments described above.
[0189] The present application also provides a computer program product, comprising a computer program, which implements the steps of any one of the method embodiments described above when executed by a processor.
[0190] in, Figure 10The electronic device architecture is shown as an example, and may include a processor 1010, a video display adapter 1011, a disk drive 1012, an input / output interface 1013, a network interface 1014, and a memory 1020. The processor 1010, the video display adapter 1011, the disk drive 1012, the input / output interface 1013, the network interface 1014, and the memory 1020 may be communicatively connected via a communication bus 1030.
[0191] Among them, the processor 1010 can be implemented by a general-purpose CPU, a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, etc., to execute relevant programs to implement the technical solutions provided in this application.
[0192] The memory 1020 can be implemented in the form of ROM (Read Only Memory), RAM (Random Access Memory), static storage device, dynamic storage device, etc. The memory 1020 can store an operating system 1021 for controlling the operation of the electronic device 1000, and a basic input and output system (BIOS) 1022 for controlling the low-level operations of the electronic device 1000. In addition, a web browser 1023, a data storage management system 1024, and an image editing device 800 or a device for training an autoregressive image generation model 900, etc. can also be stored. The above-mentioned image editing device 800 or the device for training an autoregressive image generation model 900 can be an application program that specifically implements the operations of the aforementioned steps in the embodiment of the present application. In short, when the technical solution provided by the present application is implemented by software or firmware, the relevant program code is stored in the memory 1020 and is called and executed by the processor 1010.
[0193] The input / output interface 1013 is used to connect input / output modules to implement information input and output. The input / output modules can be configured as components within the device (not shown in the figure) or can be externally connected to the device to provide corresponding functions. Input devices may include a keyboard, mouse, touch screen, microphone, various sensors, etc., and output devices may include a display, speaker, vibrator, indicator light, etc.
[0194] The network interface 1014 is used to connect to a communication module (not shown) to enable communication between the device and other devices. The communication module can communicate via a wired method (such as USB, network cable, etc.) or a wireless method (such as mobile network, WIFI, Bluetooth, etc.).
[0195] The bus 1030 comprises a pathway for transmitting information between the various components of the device (eg, the processor 1010 , the video display adapter 1011 , the disk drive 1012 , the input / output interface 1013 , the network interface 1014 , and the memory 1020 ).
[0196] It should be noted that although the above device only shows the processor 1010, the video display adapter 1011, the disk drive 1012, the input / output interface 1013, the network interface 1014, the memory 1020, the bus 1030, etc., in the specific implementation process, the device may also include other components necessary for normal operation. In addition, it will be understood by those skilled in the art that the above device may also include only the components necessary to implement the solution of the present application, and does not necessarily include all the components shown in the figure.
[0197] Through the description of the above embodiments, it can be seen that those skilled in the art can clearly understand that the present application can be implemented by means of software plus a necessary general hardware platform. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, can be embodied in the form of a computer program product, which can be stored in a storage medium such as ROM / RAM, a magnetic disk, an optical disk, etc., and includes a number of instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in various embodiments of the present application or certain parts of the embodiments.
[0198] The above is a detailed introduction to the technical solutions provided by this application. Specific examples are used herein to illustrate the principles and implementation methods of this application. The description of the above embodiments is only intended to help understand the method and core concept of this application. At the same time, for those skilled in the art, based on the concept of this application, there may be changes in the specific implementation methods and application scope. In summary, the contents of this specification should not be understood as limiting this application.
Claims
1. An image editing method, characterized in that: The method comprises: Acquire a first reference image, a second reference image, and a source image, wherein there is a causal relationship between the first reference image and the second reference image, and the causal relationship is used to represent an editing process performed when the first reference image is converted into the second reference image; An autoregressive image generation model is used to generate a target image corresponding to the source image based on the first reference image and the second reference image, and the causal relationship exists between the source image and the target image.
2. The method according to claim 1, characterized in that The autoregressive image generation model includes an encoding module, a fusion module, a causal relationship understanding module and a decoding module; The step of generating a target image corresponding to the source image based on the first reference image and the second reference image by using an autoregressive image generation model includes: Encoding the first reference image, the second reference image, and the source image using the encoding module to obtain an encoding result of the first reference image, an encoding result of the second reference image, and an encoding result of the source image, respectively; Using the fusion module, fusing the encoding result of the first reference image, the encoding result of the second reference image, and the encoding result of the source image to obtain a fused feature representation; Using the causal relationship understanding module, based on the fused feature representation, an autoregressive approach is used to predict the target image feature representation; The target image feature representation is decoded using the decoding module to obtain the target image.
3. The method according to claim 2, characterized in that The causal relationship understanding module includes a multi-layer Transformer module and a prediction module; The Transformer module performs attention processing on the input intermediate feature representation, the encoding result of the first reference image, and the encoding result of the second reference image in each round of prediction to obtain the intermediate feature representation output by the Transformer module; The intermediate feature representation input to the first-layer Transformer module includes the fused feature representation and the feature representation of the predicted elements of the target image. The intermediate feature representation input to other Transformer modules is the intermediate feature representation output by the previous-layer Transformer module. The last-layer Transformer module outputs the intermediate feature representation to the prediction module. The prediction module predicts the feature representation of the next element of the target image using the input intermediate feature representation; The feature representation of the next element is added to the feature representation of the element predicted for the target image to provide it to the first-layer Transformer module to perform the next round of prediction until all elements of the target image are predicted.
4. The method according to claim 3, characterized in that The Transformer module includes a first self-attention module, a gated self-attention module and a second self-attention module; The first self-attention module performs a first self-attention process on the intermediate feature representation input to the Transformer module, and outputs the intermediate feature representation obtained after the first self-attention process to the gated self-attention module; The gated self-attention module splices the encoding result of the first reference image, the encoding result of the second reference image, and the intermediate feature representation input to the gated self-attention module to obtain a spliced feature representation; performs a second self-attention process on the spliced feature representation, and obtains the intermediate feature representation after the second self-attention process; performs weighted processing on the intermediate feature representation input to the gated self-attention module and the intermediate feature representation after the second self-attention process, and outputs the intermediate feature representation obtained after the weighted processing to the second self-attention module; wherein the weight corresponding to the intermediate feature representation after the second self-attention process is determined by a gating parameter, and the gating parameter is learned when training the autoregressive image generation model; The second self-attention module performs a third self-attention process on the input intermediate feature representation to obtain the intermediate feature representation output by the Transformer module.
5. The method according to any one of claims 2 to 4, characterized in that The method further includes: obtaining type information of the editing task; and encoding the type information of the editing task using the encoding module to obtain an encoding result of the type information of the editing task; Using the fusion module, fusing the encoding result of the first reference image, the encoding result of the second reference image, and the encoding result of the source image to obtain a fused feature representation includes: The fusion module is used to fuse the encoding result of the first reference image, the encoding result of the second reference image, the encoding result of the source image, and the encoding result of the type information of the editing task to obtain a fused feature representation.
6. A method for training an autoregressive image generation model, characterized in that The method comprises: Obtaining a training data set, the training data set including a plurality of training samples, each training sample including two image sample pairs having the same causal relationship, using the two images included in one of the image sample pairs as a first reference image sample and a second reference image sample, and using the two images included in the other image sample pair as a source image sample and a target image sample; The first reference image sample, the second reference image sample and the source image sample are used as inputs of the autoregressive image generation model, and the target image sample is used as an output target of the autoregressive image generation model to train the autoregressive image generation model.
7. The method according to claim 6, characterized in that The two images included in the image sample pair are generated in the following manner: Inputting a first image and a first editing instruction into an image editing model, obtaining a second image generated by the image editing model based on the first image and the editing instruction, wherein the first image and the second image constitute a first image sample pair; Inputting a third image and a second editing instruction into the image editing model, obtaining a fourth image generated by the image editing model based on the third image and the two editing instructions, wherein the third image and the fourth image constitute a second image sample pair; The first image sample pair and the second image sample pair constitute a training sample; The first editing instruction and the second editing instruction correspond to the same editing process, and the image editing model is implemented based on a diffusion model.
8. The method according to claim 6, characterized in that The autoregressive image generation model includes an encoding module, a fusion module, a causal relationship understanding module and a decoding module; Encoding the first reference image sample, the second reference image sample, and the source image sample using the encoding module to obtain encoding results of the first reference image sample, encoding results of the second reference image sample, and encoding results of the source image sample, respectively; Using the fusion module, fusing the encoding result of the first reference image sample, the encoding result of the second reference image sample, and the encoding result of the source image sample to obtain a fused feature representation; Using the causal relationship understanding module, based on the fused feature representation, an autoregressive approach is used to predict the target image feature representation; The target image feature representation is decoded using the decoding module to obtain the target image.
9. The method according to claim 8, characterized in that The causal relationship understanding module includes a multi-layer Transformer module and a prediction module; The Transformer module performs attention processing on the input intermediate feature representation, the encoding result of the first reference image sample, and the encoding result of the second reference image sample in each round of prediction to obtain the intermediate feature representation output by the Transformer module; The intermediate feature representation input to the first-layer Transformer module includes the fused feature representation and the feature representation of the predicted elements of the target image. The intermediate feature representation input to other Transformer modules is the intermediate feature representation output by the previous-layer Transformer module. The last-layer Transformer module outputs the intermediate feature representation to the prediction module. The prediction module predicts the feature representation of the next element of the target image using the input intermediate feature representation; The feature representation of the next element is added to the feature representation of the element predicted for the target image to provide it to the first-layer Transformer module to perform the next round of prediction until all elements of the target image are predicted.
10. The method according to claim 9, characterized in that The Transformer module includes a first self-attention module, a gated self-attention module and a second self-attention module; The first self-attention module performs a first self-attention process on the intermediate feature representation input to the Transformer module, and outputs the intermediate feature representation obtained after the first self-attention process to the gated self-attention module; The gated self-attention module splices the encoding result of the first reference image sample, the encoding result of the second reference image sample, and the intermediate feature representation input to the gated self-attention module to obtain a spliced feature representation; performs a second self-attention process on the spliced feature representation, and obtains the intermediate feature representation after the second self-attention process; performs weighted processing on the intermediate feature representation input to the gated self-attention module and the intermediate feature representation after the second self-attention process, and outputs the intermediate feature representation obtained after the weighted processing to the second self-attention module; wherein the weight corresponding to the intermediate feature representation after the second self-attention process is determined by a gating parameter; and the gating parameter is updated during the training process; The second self-attention module performs a third self-attention process on the input intermediate feature representation to obtain the intermediate feature representation output by the Transformer module.
11. An image editing device, characterized in that: The device comprises: an acquisition unit configured to acquire a first reference image, a second reference image, and a source image, wherein a causal relationship exists between the first reference image and the second reference image, and the causal relationship is used to represent an editing process performed when the first reference image is converted into the second reference image; A generating unit is configured to generate a target image corresponding to the source image based on the first reference image and the second reference image using an autoregressive image generation model, wherein the causal relationship exists between the source image and the target image.
12. A device for training an autoregressive image generation model, the device comprising: a data acquisition unit configured to acquire a training data set, the training data set including a plurality of training samples, each training sample including two image sample pairs having the same causal relationship, the two images included in one of the image sample pairs being used as a first reference image sample and a second reference image sample, and the two images included in the other image sample pair being used as a source image sample and a target image sample; The model training unit is configured to use the first reference image sample, the second reference image sample and the source image sample as inputs of the autoregressive image generation model, and use the target image sample as an output target of the autoregressive image generation model to train the autoregressive image generation model.
13. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the steps of the method according to any one of claims 1 to 10 are implemented.
14. An electronic device, characterized in that: include: one or more processors; as well as A memory associated with the one or more processors, the memory being used to store program instructions, wherein the program instructions, when read and executed by the one or more processors, perform the steps of the method according to any one of claims 1 to 10.
15. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 10 are implemented.