Migration image editing method and system based on text graph large model

By constructing a context learning data set and fine-tuning the literary-generated image model using LoRA technology, the problem of inconsistent generation of complex spatial relationship images in the existing technology is solved, and high-quality image editing effect is achieved.

CN120355819APending Publication Date: 2025-07-22COMMUNICATION UNIVERSITY OF CHINA
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510304994.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-14
Publication Date
2025-07-22

AI Technical Summary

Technical Problem

When the existing literary and artistic image model generates non-rigid moving images of complex spatial relationships, it is difficult to accurately describe the specific location and relationship of the limb parts, resulting in inconsistent generation effects.

Method used

Construct a context learning data set, supervised and fine-tuned training of the literary image big model through LoRA technology, enhance the model's visual information capture ability in complex spatial relationship image editing, and use the fine-tuned model to generate the same edited image as the edited image example.

Benefits of technology

The consistency of the generation effect in the generation of complex spatial relationship images is achieved, the quality and efficiency of image editing are improved, and the problem of the inability to generate consistent high-quality images in the prior art is solved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120355819A_ABST
    Figure CN120355819A_ABST
Patent Text Reader

Abstract

The invention provides a migration image editing method and system based on a text graph large model, and the method comprises the steps: constructing a plurality of image editing types of example image pairs as a context learning data set, and carrying out the supervision and fine tuning training of the text graph large model based on the context learning data set. An image editing migration task is executed in the reasoning process through the fine-adjusted text graph large model, a user gives an editing image pair example and a source image, the fine-adjusted text graph large model can apply the image editing effect in the editing image pair example to the source image, and an editing image corresponding to the source image is generated. According to the invention, the technical effect of solving the problem of image generation of non-rigid motion in a complex spatial relationship can be achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and more specifically, to a method and system for migrating image editing based on a text-to-image large model. Background Art

[0002] Existing text-to-image large models can generate an edited image that fits the target text description based on the source image and target text description given by the user. However, this technology has limitations. Specifically, although text has advantages in expressing abstract semantic information, for non-rigid motions involving complex spatial relationships, such as a jumping action scene with complex limb changes, it is difficult for text to accurately describe details such as the specific positions, heights, and mutual relationships of each limb part, thus unable to provide sufficiently precise guidance for image editing, resulting in the model being difficult to generate a high-quality image that is exactly the same as the target text.

[0003] In the prior art, images with rich visual information are used to make up for the deficiency of text expression ability; however, the above method can only achieve the effect of improving the appearance texture style of low-dimensional representations, and still cannot solve the problem of generating images of non-rigid motions with complex spatial relationships.

[0004] Therefore, there is an urgent need for a method that can achieve image generation with non-rigid transformation. Summary of the Invention

[0005] In view of the above problems, the purpose of the present invention is to provide a method and system for migrating image editing based on a text-to-image large model to solve at least one problem existing in the prior art.

[0006] According to one aspect of the present invention, there is provided a method for migrating image editing based on a text-to-image large model, which is applied to an electronic device and includes:

[0007] Constructing a context learning data set based on an image array of different image editing types; wherein, the image array at least includes example image pairs and query image pairs with the same image editing type; both the example image pairs and query image pairs include a source image and an edited image corresponding to the source image;

[0008] Supervising and fine-tuning the text-to-image large model based on the context learning data set to obtain a fine-tuned text-to-image large model.

[0009] In addition, an optional technical solution is that it further includes generating, based on at least one example of an edited image pair, an edited image for the source image to be processed using the fine-tuned text-to-image large model, with the same image editing effect as the example of the edited image pair.

[0010] In addition, an alternative technical solution is that, in the process of generating, for a source image to be processed, an edited image with the same image editing effect as the example of the edited image pair by using the fine-tuned text-to-image large model based on at least one example of the edited image pair, the input of the fine-tuned text-to-image large model is the example of the edited image pair and the source image to be processed; the output of the fine-tuned text-to-image large model is a four-grid image, and the four-grid image includes the example of the edited image pair, the source image to be processed, and the edited image corresponding to the generated source image.

[0011] In addition, an alternative technical solution is that the method for generating, for a source image to be processed, an edited image with the same image editing effect as the example of the edited image pair by using the fine-tuned text-to-image large model based on at least one example of the edited image pair includes encoding the example of the edited image pair and the source image to be processed through a variational autoencoder to generate a latent image token sequence; splicing the latent image token sequence and a randomly sampled latent noise image token sequence; in a denoising network, capturing the global features of the image through a self-attention mechanism; and decoding to obtain the edited image corresponding to the source image.

[0012] In addition, an alternative technical solution is that the method for generating, for a source image to be processed, an edited image with the same image editing effect as the example of the edited image pair by using the fine-tuned text-to-image large model based on at least one example of the edited image pair includes inputting the example of the edited image pair, the source image to be processed, and a text description into the fine-tuned text-to-image large model; converting the example of the image pair and the source image to be processed through a variational autoencoder into a latent image token sequence and splicing it with a randomly sampled latent noise sequence; converting the text description through a text encoder into a latent text token sequence; capturing the global features of the image based on the latent image token sequence and the latent text token sequence through a self-attention mechanism; and decoding to obtain a four-grid image, and the four-grid image includes the example of the edited image pair, the source image to be processed, and the edited image corresponding to the generated source image.

[0013] In addition, an alternative technical solution is that the method for supervising and fine-tuning the training of the text-to-image large model based on the context learning dataset includes encoding an image array including example image pairs and query image pairs through a variational autoencoder; obtaining a latent image token sequence based on the query source images in the example image pairs and the query image pairs; obtaining a latent noise image token based on the query edited images in the query image pairs; freezing the latent image token sequence, adding noise to the latent noise image token, so as to realize the supervised fine-tuning training of the text-to-image large model based on the LoRA technology.

[0014] In addition, an alternative technical solution is that the image array is a four-grid image composed of example image pairs and query image pairs with the same image editing type.

[0015] On the other hand, the present invention also provides a transfer image editing system based on a text-to-image large model, which performs image generation by using the transfer image editing method based on the text-to-image large model as described above. The system includes:

[0016] A dataset construction unit for constructing a context learning dataset based on image arrays of different editing types; wherein, the image array at least includes example image pairs and query image pairs with the same image editing type; both the example image pairs and the query image pairs include a source image and an edited image corresponding to the source image.

[0017] A model fine-tuning unit for performing supervised fine-tuning training on a pre-trained text-to-image large model based on the context learning dataset to obtain a fine-tuned text-to-image large model.

[0018] The above-mentioned transfer image editing method and system based on a text-to-image large model construct example image pairs of several editing types as a context learning dataset, and perform supervised fine-tuning training on the text-to-image large model based on the context learning dataset. The fine-tuned text-to-image large model can execute an image editing transfer task during the inference process, that is, when the user gives an example of an edited image pair and a source image, the fine-tuned text-to-image large model can apply the image editing effect in the example of the edited image pair to the source image to generate an edited image corresponding to the source image.

[0019] To achieve the above and related purposes, one or more aspects of the present invention include features that will be described in detail later. The following description and the accompanying drawings illustrate certain exemplary aspects of the present invention in detail. However, these aspects indicate only some of the various ways in which the principles of the present invention can be used. In addition, the present invention is intended to include all these aspects and their equivalents. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] By referring to the following description in conjunction with the accompanying drawings, and with a more comprehensive understanding of the present invention, other objects and results of the present invention will become more apparent and easier to understand. In the drawings:

[0021] Figure 1 It is a schematic flowchart of a transfer image editing method based on a text-to-image large model according to an embodiment of the present invention;

[0022] Figure 2 It is a schematic diagram of the input and output of a text-to-image large model according to an embodiment of the present invention;

[0023] Figure 3Schematic diagram of the principle of the migration image editing method based on the text-to-image large model according to an embodiment of the present invention;

[0024] Figure 4 Schematic diagram for comparing the effects of the migration image editing method based on the text-to-image large model according to an embodiment of the present invention and other existing image editing solutions;

[0025] Figure 5 Example diagram of the effect of image editing using the migration image editing method based on the text-to-image large model according to an embodiment of the present invention;

[0026] Figure 6 Module schematic diagram of the migration image editing system based on the text-to-image large model provided by an embodiment of the present invention;

[0027] Figure 7 Internal structure schematic diagram of an electronic device for implementing the migration image editing method based on the text-to-image large model provided by an embodiment of the present invention.

[0028] In all the drawings, the same reference numerals indicate similar or corresponding features or functions. Detailed implementation manners

[0029] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Apparently, the described embodiments are some but not all of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0030] The technical solutions in the embodiments of the present application will be clearly and elaborately described below with reference to the accompanying drawings. Among them, in the description of the embodiments of the present application, unless otherwise specified, "and / or" in the text is only an association relationship describing associated objects, indicating that three relationships may exist. For example, A and / or B may represent: A exists alone, A and B exist simultaneously, and B exists alone. These three situations.

[0031] Hereinafter, the terms "first" and "second" are only used for descriptive purposes and cannot be construed as implying or suggesting relative importance or implicitly indicating the quantity of the indicated technical features. Thus, features defined with "first" and "second" may explicitly or implicitly include one or more of such features. Additionally, in the description of the embodiments of the present application, "a plurality" means two or more than two.

[0032] References to "one embodiment" or "some embodiments" or the like described in this specification mean that a particular feature, structure, or characteristic described in connection with the embodiment is included in one or more embodiments of the present application. Thus, statements such as "in one embodiment", "in some embodiments", "in other some embodiments", "in still other embodiments", etc., which appear in different places in this specification, are not necessarily all referring to the same embodiment, but mean "one or more but not all embodiments", unless otherwise specifically emphasized. The terms "comprising", "including", "having" and their variants all mean "including but not limited to", unless otherwise specifically emphasized.

[0033] To describe in detail the method and system for migrating image editing based on the text-to-image large model of the present invention, the specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings.

[0034] A text-to-image model is a model that generates corresponding images through simple text descriptions (in English). Existing text-to-image models include the text-to-image large model FLUX and Stable Diffusion v3, etc. FLUX is a generative large model in the field of computer vision and can perform image generation tasks such as text-to-image (txt2img) and image-to-image (img2img).

[0035] DiT (Diffusion Transformer) is a generative model that applies the Transformer architecture to the diffusion model, aiming to improve the performance and quality of image generation tasks. In DiT, the input image is divided into several fixed-size image patches, and each image patch is linearly embedded into a vector to form a series of token sequences. These token sequences are input into the Transformer encoder, and the global features of the image are captured through the self-attention mechanism. This design enables the model to more effectively process complex image generation tasks. Corresponding to the present invention, the denoising network in Flux is composed of several DiT modules.

[0036] The variational autoencoder (VAE) is a generative model for learning the data distribution. It consists of an encoder and a decoder. The encoder maps the input data to the latent space, and the decoder reconstructs the input data from the latent space. In the FLUX model, the role of the VAE is to encode the input text information into a latent token sequence, and after the denoising process, then generate the corresponding image through the decoder. This method enables the model to capture the complex distribution of the data during the generation process and improve the quality and diversity of the generated images.

[0037] LoRA (Low-Rank Adaptation) is an efficient parameter fine-tuning technique designed to optimize the adaptation ability of large pre-trained models for specific tasks. Its core idea is to insert trainable low-rank matrices into key parts of the model while freezing the weights of the pre-trained model, thereby achieving adaptation to specific tasks. This method significantly reduces the number of parameters and computational resources required for fine-tuning while maintaining or improving model performance. Compared with traditional full-parameter fine-tuning, LoRA only needs to train the newly added low-rank matrix part, avoiding re-training of the entire model's parameters. This not only reduces computational and storage costs but also reduces the risk of overfitting, enabling the model to more efficiently adapt to new tasks.

[0038] Specifically, the present invention proposes an image editing transfer task and provides an image generation method and system based on a text-to-image large model for this task. The definition of the image editing transfer task is as follows: Given an example image pair including an example source image and an example edited image, the text-to-image large model learns the transformation relationship from the example source image to the example edited image in this pair and applies this transformation relationship to the query source image to generate the corresponding query edited image of the query source image. The present invention obtains a text-to-image large model that can apply the transfer image editing method based on the text-to-image large model through constructing a context learning dataset and few-shot context fine-tuning.

[0039] Figure 1 The flowchart of the transfer image editing method based on the text-to-image large model according to an embodiment of the present invention is shown.

[0040] As Figure 1 shown, the transfer image editing method based on the text-to-image large model provided in this embodiment, which is applied to an electronic device, mainly includes the following steps:

[0041] S110: Construct a context learning dataset based on an image array of different image editing types; wherein, the image array at least includes an example image pair and a query image pair with the same image editing type; both the example image pair and the query image pair include a source image and the edited image corresponding to the source image.

[0042] It should be noted that the image editing types may include but are not limited to non-rigid motions with complex spatial relationships, which may include daily actions such as laughing, covering the face, squatting, hugging, etc.; they may also include actions in figure skating such as triple toe loop; actions in gymnastics, like the single-ring support backswing rotation of 360 degrees followed by a single-ring flyover on the pommel horse; and the ribbon throwing and catching with jumping and turning in rhythmic gymnastics. In the specific implementation process, the image type editing is not limited to human figures only, but can also be animals or robots, etc., and no specific limitation is made here. Also, the example image pairs, query image pairs, and the example editing image pairs in the subsequent steps all include a source image and the editing image corresponding to the source image.

[0043] Among them, constructing the context learning dataset includes collecting a number of example image pairs of image editing types, and splicing two example image pairs with consistent image editing type effects into a four-grid image as a training data. In an embodiment of the present invention, the above context learning dataset contains 21 image editing types, each image editing type contains two samples, and a total of 42 training data are included. That is to say, the present invention can achieve the migration of complex image editing tasks through a small-scale dataset.

[0044] Exemplarily, a text-to-image large model can be used to generate an image array of different editing types. In this embodiment, the image array includes an example image pair and a query image pair with the same image editing type; in the specific implementation process, it can also be an example image pair and multiple query image pairs with the same image editing type, as well as multiple example image pairs and a query image pair with the same image editing type; as long as it can be used to perform supervised fine-tuning training on the text-to-image large model based on the context learning dataset to obtain a fine-tuned text-to-image large model, no specific limitation is made here. The input method of the image array can be a 2×2 grid layout, that is, a four-grid image, or a 1×4 horizontal linear layout or a 4×1 vertical linear layout.

[0045] Taking the input method of the image array being a 2×2 grid layout as an example, two example image pairs with similar image editing types (i.e., the example image pair and the query image pair) are constructed into a large image I * , as a training data. Specifically, the image pair in the first row (I s , I t ) is regarded as the example image pair, and the image pair in the second row is regarded as the query image pair. That is, I * is a four-grid image, including I s (upper left corner), I t (upper right corner), (lower left corner), (lower right corner).

[0046] S120: Perform supervised fine-tuning training on the large model of the text graph based on the context learning data set to obtain a fine-tuned large model of the text graph.

[0047] Few-sample context fine-tuning refers to fine-tuning the large model of the text graph based on this small-scale dataset using LoRA technology to enhance the ability of the large model of the text graph to capture visual information in context. As an example, the large model of the text graph based on the present invention can adopt the large model provided by FLUX, such as Flux.1-dev.

[0048] Based on at least one edited image pair example, the fine-tuned Vincent graph model is used to generate an edited image with the same image editing effect as the edited image pair example for the source image to be processed. That is, after the Vincent graph model is fine-tuned using the LoRA technology, the fine-tuned Vincent graph model can perform image editing migration tasks during the inference process. Specifically, a trainable low-rank matrix (such as Rank=16) is inserted into the attention layer or feedforward layer of DiT, the original model weights are frozen, and only the newly added parameters are optimized. That is, by decomposing the low rank into a low rank matrix, the amount of fine-tuning parameters is significantly reduced, and the few sample scenarios are adapted. In the specific implementation process, the input of the fine-tuned Vincent graph model is the edited image pair example and the source image to be processed; the output of the fine-tuned Vincent graph model is a four-square grid image, which includes the edited image pair example, the source image to be processed, and the generated edited image corresponding to the source image.

[0049] Figure 2 Schematic diagram of input and output of the large model of Wensheng graph according to an embodiment of the present invention; Figure 2 As shown in Figure 1, given an edited image pair example and a source image, the fine-tuned Wenshengtu model can apply the image editing effect in the edited image pair example to the source image to generate an edited image corresponding to the source image. s , I t ), the source image to be processed is the query source image Input the fine-tuned Wensheng graph model; wherein, I s represents an example source image, I t represents the edited image of the example, represents the query source image, Indicates querying the edited image (i.e., the edited image). Figure 2 In the last two groups of four-grid images, the fine-tuned Wensheng graph model is the query source image Generate and edit image pairs example (I s , I t )

[0050] Figure 3 Schematic diagram of the principle of the migration image editing method based on the text-to-image large model according to an embodiment of the present invention; for the text-to-image large model of the present invention, the backbone of the network structure is a text-to-image large model based on DiT (Diffusion Transformer) and a variational autoencoder (VAE). As Figure 3 shown, taking the FLUX model as an example of the text-to-image large model, the whole process is divided into three main stages: data exemplification, fine-tuning stage, and inference stage. For the first stage of data exemplification, two pairs of example and query images are shown in the figure, where one pair is training data and the other pair is test data Among them, the training data is used to fine-tune the text-to-image large model; the test data is used to evaluate the text-to-image large model. The text encoder converts the text description (P) into a sequence of text tokens (C T ), and these sequences are used to guide image generation. The text token sequence (C T ) and the image token sequence (C I ) generate query (Q), key (K), and value (V) matrices through layer normalization and linear transformation. The query (Q), key (K), and value (V) matrices are fed into the multi-modal attention (MMA) for processing to extract features in different sub-images and the relationships between them.

[0051] For the second stage: fine-tuning stage, the method for supervised fine-tuning training of the text-to-image large model based on the context learning dataset includes: S121, encoding the image array including the example image pair and the query image pair using a variational autoencoder; S122, obtaining a latent image token sequence based on the query source image in the example image pair and the query image pair; obtaining a latent noise image token based on the query editing image in the query image pair; S123, freezing the latent image token sequence and adding noise to the latent noise image token to implement supervised fine-tuning training of the text-to-image large model based on the LoRA technology. Specifically, the training data I * is encoded into a latent image token sequence, and the sub-image is converted into a latent conditional image token sequence c I through the VAE encoder. The latent noise image token encoded for the sub-image is z0; z0 is added noise to z t , c IRemain unchanged. The noisy latent image token sequence is fed into the DiT module for processing, and image features are extracted through multi-modal attention (MMA) and the forward process. By comparing the generated image features and the target image features, the loss function (L CFM ) is calculated to guide the update of the model parameters.

[0052] That is to say, during the fine-tuning process using the LoRA technique, the model repeatedly performs noise addition and denoising operations, which are only applied to z0. The loss function used is as follows:

[0053]

[0054] where v θ (z, t, c T , c I ) is the velocity field predicted by the model, t is the time step, and u t (z|∈) is the target vector field. The LoRA technique enables rapid task adaptation without changing the backbone architecture, avoiding the resource overhead of full-parameter fine-tuning.

[0055] For the third stage: Inference Stage, the method for obtaining the edited image corresponding to the source image includes: S131, encoding the edited image pair example and the source image to be processed through a variational autoencoder to generate a latent image token sequence; S132, concatenating the latent image token sequence and a randomly sampled latent noise image token sequence; S133, in the denoising network, capturing the global features of the image through the self-attention mechanism; and decoding to obtain the edited image corresponding to the source image.

[0056] Specifically, the test data image sub-image is converted into a latent conditional image token sequence c I through the VAE encoder, and is concatenated with a randomly sampled latent noise image token sequence z T , and then fed into the DiT module for processing, and image features are extracted through multi-modal attention (MMA) and the forward process. The latent representation processed by the DiT module is decoded into the generated image In general, the dataset construction is to leverage the contextual learning of the FLUX model to generate images. The fine-tuning process is to further enhance the FLUX model's ability to extract visual contextual relationships for the image editing transfer task on the constructed dataset. The inference process is to use the fine-tuned model to perform image transfer tasks directly on the user-provided data. In the inference stage, the latent conditional tag sequence (the encoded example edited pair and the query source image) is concatenated with the random noise tag as the input of the diffusion model. The diffusion process is performed in the latent space to reduce the computational complexity. In the denoising process (such as Euler method sampling), the contextual information of the latent conditional image tag sequence and the latent noisy image tag sequence are fused through multimodal attention to guide the edited image pair example (I s , I t ) By splicing potential tags and injecting context, the generated image is ensured to be consistent with the editing effect of the example editing pair, while retaining the content characteristics of the query source image. This paper takes into account the generation quality, computational efficiency and task generalization, and provides a scalable technical framework for image migration based on the large model of text graphs.

[0057] In a specific embodiment, the image array is a four-square grid image composed of example image pairs and query image pairs of the same image editing type. The image array is a large grid image composed of example image pairs and query image pairs of the same image editing type. * , as a training data. The first row of edit pairs (I s ,I t ) as an example edit pair, and the edit pair in the second row Treated as a query edit pair. * is a four-grid image, including I s (upper left corner), I t (upper right corner), (lower left corner), That is to say, in the contextual learning dataset, the large image formed by the four-grid sub-image is used as a training data; then the input of the edit transfer image generation process corresponding to the fine-tuned Wensheng graph large model is three images, including the example edit pair input by the user (I s ,I t ), query source image After encoding these three images into a potential conditional image edit sequence using a variational autoencoder, they are concatenated with the potential noise image sequence. After the denoising process is completed, the potential noise image sequence corresponds to the generated edited image That is, the output of the migration image generation process of the fine-tuned Wenshengtu large model is a four-square image.

[0058] Specifically, the method for generating an edited image with the same image editing effect as the example of the edited image pair for the source image to be processed by using the fine-tuned text-to-image large model based on at least one example of the edited image pair includes: inputting the example of the edited image pair and the source image to be processed into the fine-tuned text-to-image large model, converting them into a sequence of latent image tokens through a variational autoencoder, and concatenating them with a randomly sampled sequence of latent noise; capturing the global features of the image through a self-attention mechanism based on the sequence of latent image tokens; and decoding to obtain a four-grid image, where the four-grid image includes the example of the edited image pair, the source image to be processed, and the edited image corresponding to the source image to be processed.

[0059] In a specific embodiment, as Figure 3 shown, the input of the fine-tuned text-to-image large model further includes a text description;

[0060] In addition, an example of the processing process of the text description during the image editing migration process of the text-to-image large model is as follows: The method for generating an edited image with the same image editing effect as the example of the edited image pair for the source image to be processed by using the fine-tuned text-to-image large model based on at least one example of the edited image pair includes: inputting the example of the edited image pair, the source image to be processed, and the text description into the fine-tuned text-to-image large model; converting the example of the image pair and the source image to be processed into a sequence of latent image tokens through a variational autoencoder and concatenating them with a randomly sampled sequence of latent noise; converting the text description into a sequence of latent text tokens through a text encoder; capturing the global features of the image through a self-attention mechanism based on the sequence of latent image tokens and the sequence of latent text tokens; and decoding to obtain a four-grid image, where the four-grid image includes the example of the edited image pair, the source image to be processed, and the edited image corresponding to the generated source image.

[0061] In a specific embodiment, the context learning dataset contains 21 editing types, each editing type contains two samples, and there are a total of 42 training data. Among them, the rank of LoRA is 16, and the number of training iterations is 6000 times. During the inference process, the number of time steps for denoising processing is 35.

[0062] Figures 4 to 5 shows the effect of applying the migration image editing method based on the text-to-image large model of the present invention. Among them, Figure 4 is a schematic diagram of the effect comparison between the migration image editing method based on the text-to-image large model of the embodiment of the present invention and other existing image editing schemes; Figure 5 is an example diagram of the effect of image editing by applying the migration image editing method based on the text-to-image large model of the embodiment of the present invention.

[0063] AsFigure 4 As shown, the first and second columns of images are example editing pairs, the third column of images is the query source image input by the user, the fourth column of images "ours" represents the result image obtained by editing the image using the migration image editing method based on the text-to-image large model of the present invention, and the fifth column to the last column are the result images of P2P, RF-Solver-Edit, and MimicBrush. For the comparison scheme of the prior art, refer to P2P: the paper Prompt-to-prompt image editing with cross attention control; RF-Solver-Edit: the paper Taming Rectified Flow for Inversion and Editing; MimicBrush: the paper Zero-shot Image Editing with Reference Imitation. By observing Figure 4 , it is obvious that the editing image corresponding to the source image to be processed generated by the migration image editing method based on the text-to-image large model of the present invention s , I t ) has the highest similarity in image editing type with the image editing type of the example editing pair (I

[0064] Among them, in Figures 4 - 5 the four-grid image I * contains the example editing pair I s (upper left corner), I t (upper right corner), the new source image (lower left corner), and the result generated by the model (lower right corner). By observing Figure 5 , it is obvious that the migration image editing method based on the text-to-image large model of the present invention is proven to have good context generation ability and can generate a large image containing several sub-images with consistent styles.

[0065] Such as Figure 6As shown in the figure, the present invention provides a migration image editing system based on a text-to-image large model, which performs image editing migration by using the migration image editing method based on the text-to-image large model as described above. According to the functions implemented, the migration image editing system 600 based on the text-to-image large model may include a dataset construction unit 610 and a model fine-tuning unit 620. The units of the present invention may also be referred to as modules, which refer to a series of computer program segments that can be executed by the processor of an electronic device and can complete fixed functions, and are stored in the memory of the electronic device.

[0066] In this embodiment, the functions of each module / unit are as follows:

[0067] Perform image generation by using the migration image editing method based on the text-to-image large model as described above; including:

[0068] The dataset construction unit 610 is used to construct a context learning dataset based on an image array of different image editing types; wherein, the image array at least includes an example image pair and a query image pair with the same image editing type; both the example image pair and the query image pair include a source image and an edited image corresponding to the source image;

[0069] The model fine-tuning unit 620 is used to perform supervised fine-tuning training on the text-to-image large model based on the context learning dataset to obtain a fine-tuned text-to-image large model.

[0070] The migration image editing system based on the text-to-image large model of the present invention constructs an example image pair of several editing types as a context learning dataset, and performs supervised fine-tuning training on the text-to-image large model based on the context learning dataset. The fine-tuned text-to-image large model can execute the image editing migration task during the inference process, that is, when the user gives an example of an edited image pair and a source image, the fine-tuned text-to-image large model can apply the image editing effect in the example of the edited image pair to the source image to generate an edited image corresponding to the source image; achieving the technical effect of solving the problem of image generation for non-rigid motion with complex spatial relationships.

[0071] For more specific implementation manners of the above-mentioned migration image editing system based on the text-to-image large model, reference can be made to the description of the embodiments of the migration image editing method based on the text-to-image large model as described above, and details will not be repeated here.

[0072] As Figure 7 shown in the figure, the present invention also correspondingly provides an electronic device 1 for the migration image editing method based on the text-to-image large model.

[0073] The electronic device 1 may include a processor 10, a memory 11, and a bus. It may also include a computer program stored in the memory 11 and executable on the processor 10, such as a migration image editing program 12 based on a text-to-image large model. The memory 11 may also include both an internal storage unit of the migration image editing system based on the text-to-image large model and an external storage device. The memory 11 can be used not only to store installed application software and various types of data, such as the code of the migration image editing program based on the text-to-image large model, but also to temporarily store data that has been output or will be output.

[0074] The electronic device 1 may include a processor 10, a memory 11, and a bus. It may also include a computer program stored in the memory 11 and executable on the processor 10, such as a migration image editing program 12 based on a text-to-image large model. The memory 11 may also include both an internal storage unit of the migration image editing system based on the text-to-image large model and an external storage device. The memory 11 can be used not only to store installed application software and various types of data, such as the code of the migration image editing program based on the text-to-image large model, but also to temporarily store data that has been output or will be output.

[0075] Among them, the memory 11 includes at least one type of readable storage medium, and the readable storage medium includes flash memory, mobile hard disks, multimedia cards, card-type memories (such as SD or DX memories, etc.), magnetic memories, magnetic disks, optical discs, etc. In some embodiments, the memory 11 may be an internal storage unit of the electronic device 1, such as the mobile hard disk of the electronic device 1. In other embodiments, the memory 11 may also be an external storage device of the electronic device 1, such as a plug-in mobile hard disk, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, etc., equipped on the electronic device 1. Further, the memory 11 may also include both an internal storage unit and an external storage device of the electronic device 1. The memory 11 can be used not only to store installed application software and various types of data on the electronic device 1, such as the code of the migration image editing program based on the text-to-image large model, but also to temporarily store data that has been output or will be output.

[0076] In some embodiments, the processor 10 may be composed of an integrated circuit. For example, it may be composed of a single packaged integrated circuit, or may be composed of multiple packaged integrated circuits with the same or different functions, including one or more central processing units (CPUs), microprocessors, digital processing chips, graphics processors, and combinations of various control chips, etc. The processor 10 is the control core (Control Unit) of the electronic device, connecting various components of the entire electronic device through various interfaces and lines, and by running or executing programs or modules stored in the memory 11 (such as the migration image editing program based on the text-to-image large model, etc.), and calling the data stored in the memory 11, to execute various functions of the electronic device 1 and process data.

[0077] The bus may be a peripheral component interconnect (PCI) bus or an extended industry standard architecture (EISA) bus, etc. This bus can be divided into an address bus, a data bus, a control bus, etc. The bus is configured to enable connection communication between the memory 11 and at least one processor 10, etc.

[0078] Figure 7 Only the electronic device with components is shown. Those skilled in the art can understand that, Figure 7 the shown structure does not constitute a limitation on the electronic device 1, and it may include fewer or more components than shown, or combine certain components, or have different component arrangements.

[0079] For example, although not shown, the electronic device 1 may further include a power supply (such as a battery) for powering each component. Preferably, the power supply can be logically connected to the at least one processor 10 through a power management system, so as to implement functions such as charge management, discharge management, and power consumption management through the power management system. The power supply may also include any components such as one or more DC or AC power supplies, a recharge system, a power failure detection circuit, a power converter or inverter, a power status indicator, etc. The electronic device 1 may also include various sensors, a Bluetooth module, a Wi-Fi module, etc., which will not be elaborated here.

[0080] Furthermore, the electronic device 1 may further include a network interface. Optionally, the network interface may include a wired interface and / or a wireless interface (such as a WI-FI interface, a Bluetooth interface, etc.), which is generally used to establish a communication connection between the electronic device 1 and other electronic devices.

[0081] Optionally, the electronic device 1 may further include a user interface, which may be a display, an input unit (such as a keyboard), and optionally, the user interface may also be a standard wired interface or a wireless interface. Optionally, in some embodiments, the display may be an LED display, a liquid crystal display, a touch liquid crystal display, and an OLED (Organic Light-Emitting Diode) toucher, etc. Among them, the display may also be appropriately referred to as a display screen or a display unit, which is used to display the information processed in the electronic device 1 and to display a visual user interface.

[0082] It should be understood that the above embodiments are only for illustration purposes and are not limited by this structure in the scope of the patent application.

[0083] The migration image editing program 12 based on the text-to-image large model stored in the memory 11 of the electronic device 1 is a combination of multiple instructions. When running in the processor 10, it can achieve: constructing a context learning data set based on an image array of different image editing types; wherein, the image array at least includes example image pairs and query image pairs with the same image editing type; both the example image pairs and the query image pairs include a source image and an edited image corresponding to the source image; performing supervised fine-tuning training on the text-to-image large model based on the context learning data set to obtain a fine-tuned text-to-image large model. Based on at least one example of an edited image pair, using the fine-tuned text-to-image large model, generating an edited image for the source image to be processed, which has the same image editing effect as the edited image pair example.

[0084] Specifically, for the specific implementation method of the above instructions by the processor 10, reference may be made to Figure 1 the description of the relevant steps in the corresponding embodiments, which will not be elaborated here. Further, if the modules / units integrated in the electronic device 1 are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. The computer-readable medium may include: any entity or system that can carry the computer program code, a recording medium, a USB flash drive, a mobile hard disk, a magnetic disk, an optical disk, a computer memory, a read-only memory (ROM, Read-Only Memory).

[0085] An embodiment of the present invention further provides a computer-readable storage medium, which may be non-volatile or volatile. The storage medium stores a computer program, and when the computer program is executed by a processor, it realizes: constructing a context learning data set based on an image array of different image editing types; wherein, the image array at least includes example image pairs and query image pairs with the same image editing type; both the example image pairs and the query image pairs include a source image and an edited image corresponding to the source image; performing supervised fine-tuning training on the text-to-image large model based on the context learning data set to obtain a fine-tuned text-to-image large model. Based on at least one example of an edited image pair, using the fine-tuned text-to-image large model, an edited image with the same image editing effect as the example of the edited image pair is generated for the source image to be processed.

[0086] Specifically, the specific implementation method when the computer program is executed by the processor can refer to the description of the relevant steps in the embodiment of the wearing detection method, which will not be elaborated here.

[0087] In several embodiments provided by the present invention, it should be understood that the disclosed devices, systems, and methods can be implemented in other ways. For example, the system embodiments described above are merely illustrative. For example, the division of the modules is only a logical function division, and there may be other division methods in actual implementation.

[0088] The modules described as separate components may or may not be physically separated, and the components shown as modules may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0089] In addition, in each embodiment of the present invention, the functional modules can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above integrated unit can be implemented in the form of hardware or in the form of a hardware plus a software functional module.

[0090] For those skilled in the art, it is obvious that the present invention is not limited to the details of the above exemplary embodiments, and without departing from the spirit or basic characteristics of the present invention, the present invention can be implemented in other specific forms.

[0091] Therefore, in any case, the embodiments should be regarded as exemplary and non-limiting. The scope of the present invention is defined by the appended claims rather than the above description. Therefore, all changes falling within the meaning and scope of the equivalent elements of the claims are intended to be embraced by the present invention. Any reference signs in the claims should not be construed as limiting the claims concerned.

[0092] In addition, it is obvious that the word "comprising" does not exclude other units or steps, and the singular does not exclude the plural. A plurality of units or systems stated in the system claims can also be implemented by one unit or system through software or hardware.

[0093] However, those skilled in the art should understand that various improvements can be made to the above-described method for migrating image editing based on the text-to-image large model and the system for migrating image editing based on the text-to-image large model without departing from the content of the present invention. Therefore, the protection scope of the present invention should be determined by the content of the appended claims.

Claims

1. A method for migrating image editing based on a text-to-image large model, applied to an electronic device, characterized in that, Including: Construct a context learning dataset based on image arrays of different image editing types; wherein, the image array includes at least example image pairs and query image pairs with the same image editing type; both the example image pairs and query image pairs include a source image and an edited image corresponding to the source image; Supervisedly fine-tune and train a text-to-image large model based on the context learning dataset to obtain a fine-tuned text-to-image large model.

2. The migration image editing method based on the text-to-image large model according to claim 1, wherein Based on at least one example of an edited image pair, use the fine-tuned text-to-image large model to generate an edited image for the source image to be processed, which has the same image editing effect as the example of the edited image pair.

3. The method for migrating image editing based on a text-to-image large model according to claim 2, wherein, During the process of using the fine-tuned text-to-image large model based on at least one example of an edited image pair to generate an edited image for the source image to be processed, which has the same image editing effect as the example of the edited image pair, The input of the fine-tuned text-to-image large model is the example of the edited image pair and the source image to be processed; the output of the fine-tuned text-to-image large model is A four-grid image, which includes the example of the edited image pair, the source image to be processed, and the generated edited image corresponding to the source image.

4. The migration image editing and generation method based on the text-to-image large model according to claim 2, wherein The method of using the fine-tuned text-to-image large model based on at least one example of an edited image pair to generate an edited image for the source image to be processed, which has the same image editing effect as the example of the edited image pair, includes Encoding the example of the edited image pair and the source image to be processed through a variational autoencoder to generate a latent image token sequence; Concatenating the latent image token sequence and a randomly sampled latent noise image token sequence; In a denoising network, capturing the global features of the image through a self-attention mechanism; Decoding to obtain the edited image corresponding to the source image.

5. The migration image editing method based on the text-to-image large model according to claim 2, wherein The method of using the fine-tuned text-to-image large model based on at least one example of an edited image pair to generate an edited image for the source image to be processed, which has the same image editing effect as the example of the edited image pair, includes Inputting the example of the edited image pair, the source image to be processed, and a text description into the fine-tuned text-to-image large model; The example of the image pair and the source image to be processed are converted into a latent image token sequence through a variational autoencoder and concatenated with a randomly sampled latent noise sequence; the text description is converted into a latent text token sequence through a text encoder; Based on the latent image token sequence and the latent text token sequence, capturing the global features of the image through a self-attention mechanism; Decoding to obtain a four-grid image, which includes the example of the edited image pair, the source image to be processed, and the generated edited image corresponding to the source image.

6. The migration image editing method based on the text-to-image large model according to claim 1, wherein The method of supervisedly fine-tuning and training a text-to-image large model based on the context learning dataset includes Encoding an image array including example image pairs and query image pairs through a variational autoencoder; Based on the query source images in the example image pairs and query image pairs, obtaining a latent image token sequence; Obtaining a latent noise image token based on the query edited images in the query image pairs; Freeze the sequence of the latent image tokens, and add noise to the latent noise image tokens to achieve supervised fine-tuning training of the text-to-image large model based on the LoRA technology.

7. The method for migrating image editing based on the text-to-image large model according to claim 1, wherein The image array is a four-grid image composed of example image pairs and query image pairs with the same image editing type.

8. A transfer image editing system based on a text-to-image large model, which uses the transfer image editing method based on the text-to-image large model according to any one of claims 1-7 to generate images; comprising: A dataset construction unit for constructing a context learning dataset for image arrays of different image editing types; wherein, the image array at least includes example image pairs and query image pairs with the same image editing type; both the example image pairs and the query image pairs include a source image and an edited image corresponding to the source image; A model fine-tuning unit for performing supervised fine-tuning training on the text-to-image large model based on the context learning dataset to obtain a fine-tuned text-to-image large model.

9. An electronic device, characterized in that The electronic device includes a memory, a processor, and a transfer image editing program based on the text-to-image large model stored in the memory and executable on the processor. When the transfer image editing program based on the text-to-image large model is executed by the processor, it implements the steps of the transfer image editing method based on the text-to-image large model according to any one of claims 1 to 7.

10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the transfer image editing method based on the text-to-image large model according to any one of claims 1 to 7.