Drag-based image editing apparatus and method

The drag-based image editing system combines GAN and diffusion models to achieve fast, memory-efficient, and high-quality image editing by processing drag inputs, addressing resource-intensive challenges in conventional methods.

JP2026025822APending Publication Date: 2026-02-16SEOUL NATIONAL UNIVERSITY R&DB FOUNDATION
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
JP2024203228
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-08-01
Filing Date
2024-11-21
Publication Date
2026-02-16

AI Technical Summary

Technical Problem

Conventional drag-based image editing methods require significant computational resources and memory, are cumbersome for users, and lack efficient techniques for real-time high-quality editing without additional input.

Method used

A drag-based image editing system utilizing a combination of a GAN-based optical flow generation model (FlowGen) and a diffusion model-based image generation model (FlowDiffusion) to process drag inputs, minimizing memory usage and optimizing computational efficiency.

Benefits of technology

Enables high-quality image editing at speeds 10 to 100 times faster and with up to 5 times less GPU memory than existing technologies, allowing real-time interaction and natural movement reflection using video datasets.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026025822000001_ABST
    Figure 2026025822000001_ABST
Patent Text Reader

Abstract

To provide an image editing device and an image editing method for editing an image based on dragging.SOLUTION: The image editing apparatus 100 includes an input / output unit configured to acquire a drag input command and an image, and a controller configured to acquire an opticalflow based on the drag input command and the image using a first artificial intelligence model trained to receive the drag input command and the image and output the opticalflow, acquire an edited image as an output of a second artificial intelligence model different from the first artificial intelligence model by inputting the opticalflow and the image to the second artificial intelligence model, and provide the edited image.SELECTED DRAWING: Figure 2
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The embodiments disclosed herein relate to a drag-based image editing device and method, and more particularly to a device and method for processing image editing requests based on drag input commands using artificial intelligence.

[0002] This research was conducted as a result of the "Artificial Intelligence Graduate School Program (Seoul National University)" project (IITP-2021-0-01343) of the Ministry of Science and ICT and the Institute for Information and Communications Technology Planning (IITP)'s ICT Broadcasting Innovation Talent Development Project. [Background technology]

[0003] Recently, the market for image editing programs has been growing steadily, and the demand for image editing techniques based on artificial intelligence (AI) is increasing rapidly. In particular, the demand for drag-based image editing techniques is increasing. The drag-based editing technique is a generative AI-based image editing technique in which, when a user drags a specific part of an image with a mouse to move it to the desired position, the system automatically adjusts the remaining parts naturally, taking into account realistic movements.

[0004] The drag-based image editing method is highly intuitive and allows precise adjustment, making it a technique that can be widely used by both experts and general users, and its "dragging" feature makes it highly suitable for touch-based mobile environments. However, conventional drag-based image editing methods must take into account realistic motions present in the image, which requires additional optimization and learning for each individual image, resulting in large amounts of calculation and memory. In addition, it is cumbersome for users to have to input text associated with the image and masks indicating areas that can be moved.

[0005] Meanwhile, prior art document Korean Patent Publication No. 10-2014-0120628 (published October 14, 2014) is an invention relating to an image editing method and apparatus, and includes the steps of determining a main image, receiving an input to determine an individual, displaying the individual on a mobile device screen, receiving a touch input for the individual, displaying a menu for the individual on the screen, receiving a first drag input to change the touch point, displaying the main image in a first area if the touch point is located on the menu, overlaying an icon on the main image at the touch point, receiving a second drag input to position the icon at a point on the main image, and generating an individual insertion image with the individual inserted when the touch input ends. While this prior art proposes a technique for editing images through drag input, it does not propose how to improve the resource-intensive nature of image editing, nor does it propose how to implement this technique when editing images through drag input on an image, since it only inserts one image into another image. Therefore, there is a need for a technique that allows for editing images to a satisfactory quality with real-time interaction while requiring a small amount of calculation and memory.

[0006] On the other hand, the above-mentioned background art is technical information that the inventor possessed in order to derive the present invention or that he acquired in the process of deriving the present invention, and it cannot necessarily be said to be publicly known art that was disclosed to the general public prior to the filing of the present invention. [Prior art documents] [Patent documents]

[0007] [Patent Document 1] Korean Patent Publication No. 10-2014-0120628 Summary of the Invention [Problem to be solved by the invention]

[0008] The disclosed embodiments aim to provide an apparatus and method for drag-based image editing.

[0009] The disclosed embodiments aim to provide a real-time drag-based image editing device and method that allows a user to naturally move and edit an object in an image in response to a drag input.

[0010] The disclosed embodiments aim to provide an image editing apparatus and method that is memory-efficient and allows for real-time level interaction.

[0011] The disclosed embodiments aim to provide an image editing apparatus and method that allows high-quality editing without the need for additional information and allows real-time interaction.

[0012] The disclosed embodiments aim to provide an image editing apparatus and method that performs high-quality editing at high speed without optimization, unlike existing optimization-based methods.

[0013] The disclosed embodiments aim to provide an image editing device and method that can perform high-speed editing while reflecting detailed movements by learning natural movements that occur in everyday life using video datasets.

[0014] The disclosed embodiment aims to provide a fast yet user-friendly image editing device and method by proposing and training a new specialized model specifically for drag-based editing, rather than optimizing a pre-trained huge model at the time of editing. [Means for solving the problem]

[0015] According to one embodiment, a technical means for achieving the above-mentioned technical object may include an input / output unit that acquires a drag input command and an image, and a control unit that uses a first artificial intelligence model that has been trained to receive a drag input command and an image and output an optical flow, acquires an optical flow based on the drag input command and the image, inputs the optical flow and the image to a second artificial intelligence model different from the first artificial intelligence model, acquires an edited image as an output of the second artificial intelligence model, and provides the edited image.

[0016] According to another embodiment, an image editing method performed by an image editing device may include the steps of: acquiring a drag input command and an image; acquiring an optical flow based on the drag input command and the image using a first artificial intelligence model trained to receive the drag input command and the image and output an optical flow; inputting the optical flow and the image into a second artificial intelligence model different from the first artificial intelligence model to acquire an edited image as an output of the second artificial intelligence model; and providing the edited image.

[0017] According to yet another embodiment, there is provided a computer-readable recording medium having a program for performing an image editing method recorded thereon, the image editing method including the steps of: acquiring a drag input command and an image; acquiring an optical flow based on the drag input command and the image using a first artificial intelligence model trained to output an optical flow in response to input of the drag input command and the image; inputting the optical flow and the image into a second artificial intelligence model different from the first artificial intelligence model to acquire an edited image as an output of the second artificial intelligence model; and providing the edited image.

[0018] According to yet another embodiment, there is provided a computer program executed by an image editing device and stored on a medium for performing an image editing method, the image editing method including the steps of: acquiring a drag input command and an image; acquiring an optical flow based on the drag input command and the image using a first artificial intelligence model trained to receive the drag input command and the image and output an optical flow; inputting the optical flow and the image into a second artificial intelligence model different from the first artificial intelligence model to acquire an edited image as an output of the second artificial intelligence model; and providing the edited image. [Effects of the Invention]

[0019] According to any one of the above-mentioned solutions, an apparatus and method for editing an image based on dragging can be provided.

[0020] According to any one of the above-mentioned solutions, it is possible to provide an image editing device and method based on real-time dragging, which allows a user to naturally move and edit an object in an image through drag input.

[0021] According to any one of the above-described solutions, it is possible to provide an image editing device and method that requires little memory and allows real-time interaction, i.e., it is possible to provide an image editing device and method that is capable of real-time processing, requires low computing resources, and can be implemented on various hardware.

[0022] According to any one of the above-mentioned solutions, it is possible to provide an image editing apparatus and method that allows high-quality editing without additional information and allows real-time interaction.

[0023] According to any one of the above-mentioned solutions, it is possible to provide an image editing device and method that performs high-quality editing at high speed without optimization, unlike existing optimization-based methods. By enabling sophisticated drag editing using only simple drag commands and images without masks or additional text information, and by omitting the optimization process that requires many calculations for each individual image, it is possible to provide an image editing device and method that is 10 to 100 times faster and uses up to 5 times less GPU memory than existing technologies.

[0024] According to any one of the above-mentioned solutions, it is possible to provide an image editing device and method that can perform high-speed editing while reflecting detailed movements by learning natural movements that occur in everyday life using a video dataset.

[0025] According to any one of the above-mentioned means for solving the problem, a fast yet user-friendly image editing device and method can be presented by proposing and learning a new specialized model specialized for drag editing, rather than optimizing a huge pre-trained model at the time of editing.

[0026] The effects obtained from the disclosed embodiments are not limited to those mentioned above, and other effects not mentioned will be clearly understood by those skilled in the art to which the embodiments disclosed below pertain. [Brief explanation of the drawings]

[0027] [Figure 1] 1 is a block diagram illustrating an image editing device according to an embodiment; [Figure 2] 1 is a block diagram illustrating an image editing device according to an embodiment. [Figure 3] 1 is an exemplary diagram illustrating an image editing apparatus according to an embodiment; [Figure 4] 1 is an exemplary diagram illustrating an image editing apparatus according to an embodiment; [Figure 5]1 is an exemplary diagram illustrating an image editing apparatus according to an embodiment; [Figure 6] 1 is an exemplary diagram illustrating an image editing apparatus according to an embodiment; [Figure 7] 1 is an exemplary diagram illustrating an image editing apparatus according to an embodiment; [Figure 8] 1 is an exemplary diagram illustrating an image editing apparatus according to an embodiment; [Figure 9] 1 is an exemplary diagram illustrating an image editing apparatus according to an embodiment; [Figure 10] 1 is a flowchart illustrating an image editing method according to an embodiment. [Figure 11] 1 is an exemplary diagram illustrating an image editing method according to an embodiment; [Figure 12] 1 is an exemplary diagram illustrating an image editing method according to an embodiment; DETAILED DESCRIPTION OF THE INVENTION

[0028] Various embodiments will be described in detail below with reference to the accompanying drawings. The embodiments described below may be implemented in various modified forms. In order to more clearly describe the features of the embodiments, detailed descriptions of matters that are well known to those skilled in the art to which the following embodiments pertain will be omitted. In addition, parts of the drawings that are not relevant to the description of the embodiments will be omitted, and similar parts will be designated by similar reference numerals throughout the specification.

[0029] Throughout the specification, when a certain component is said to be "connected" to another component, this includes not only "directly connected" but also "connected via another component in between." Furthermore, when a certain component is said to "include" another component, this does not exclude the other component, but means that the other component may also be included, unless otherwise specified.

[0030] Hereinafter, the embodiments will be described in detail with reference to the accompanying drawings.

[0031] Fig. 1 is a configuration diagram for explaining an image editing device according to an embodiment, Fig. 2 is a block diagram for explaining an image editing device according to an embodiment, Fig. 3 to Fig. 9 are exemplary diagrams for explaining an image editing device according to an embodiment, and Fig. 3 to Fig. 9 will be referred to when explaining the image editing device.

[0032] The image editing device 100 is a device for editing an image in response to a drag command, that is, the image editing device 100 may be a drag-based image editing artificial intelligence device capable of real-time interaction.

[0033] The image editing device 100 can be implemented as a user terminal, a server, or a server-client system.

[0034] According to one embodiment, when the image editing device 100 is implemented as a user terminal, all operations of the image editing device 100 described below can be performed on the user terminal, and when the image editing device 100 is implemented as a server, all operations of the image editing device 100 described below can be performed on the server.

[0035] According to another embodiment, the image editing device 100 may be implemented as a server-client system, which includes a user terminal 10 and a server 20, as shown in FIG. 1, and communicates with each other via a network.

[0036] Here, the user terminal 10 may be implemented as a computer, a portable terminal, a television, a wearable device, etc. that can be connected to a remote server via a network or can be connected to other terminals and servers. Here, the computer may include, for example, a notebook PC, a desktop PC, or a laptop PC equipped with a web browser, and the portable terminal may include, for example, a wireless communication device that ensures portability and mobility, such as a Personal Communication System (PCS), a Personal Digital Cellular (PDC), a Personal Handyphone System (PHS), a Personal Digital Assistant (PDA), a Global System for Mobile communications (GSM), an International Mobile Telecommunication (IMT)-2000, a Code Division Multiple Access (CDMA)-2000, a W-Code Division Multiple Access (W-CDMA), a Wireless Broadband Internet (Wibro), a smartphone, a Mobile Worldwide Interoperability for Microwave Access (Mobile WiMAX), etc. The television may include Internet Protocol Television (IPTV), Internet Television, terrestrial TV, cable TV, etc. A wearable device is a type of information processing device that can be worn directly on the human body, such as a watch, glasses, accessory, clothing, or footwear, and can be connected to a remote server or other terminal via a network, either directly or through another information processing device.

[0037] The server 20 may be implemented as a computer that can communicate with the user terminal 10 via a network, or as a cloud computer server. When the image editing device 100 is implemented as a server-client, the components that make up the image editing device 100 may be implemented on the user terminal 10 and the server 20.

[0038] Meanwhile, referring to FIG. 2, the image editing apparatus 100 may include an input / output unit 110, a control unit 120, a communication unit 130, and a memory 140.

[0039] The input / output unit 110 may include an input unit for receiving input from a user and an output unit for displaying information such as a result of execution of a task or a status of the image editing device 100. For example, the input / output unit 110 may include an operation panel for receiving user input and a display panel for displaying a screen.

[0040] Specifically, the input unit may include a device capable of receiving various types of user input, such as a keyboard, physical buttons, a touch screen, a camera, a microphone, etc. The output unit may include a display panel, a speaker, etc. However, without being limited thereto, the input / output unit 110 may include a configuration supporting various inputs and outputs.

[0041] According to the embodiment, the input / output unit 110 can acquire a drag input command and an image, for example, by outputting an image, can acquire a user's drag input command for the output image, and can transmit the output image and the drag input command for the image to the control unit 120.

[0042] The control unit 120 controls the overall operation of the image editing device 100 and may include a processor such as a CPU, a GPU, etc. The control unit 120 may control other components included in the image editing device 100 to perform an operation corresponding to a user input received via the input / output unit 110.

[0043] For example, the control unit 120 may execute a program stored in the memory 140 , read a file stored in the memory 140 , or store a new file in the memory 140 .

[0044] The control unit 120 is described in more detail below.

[0045] The communication unit 130 may perform wired or wireless communication with other devices or networks. To this end, the communication unit 130 may include a communication module that supports at least one of various wired or wireless communication methods. For example, the communication module may be implemented in the form of a chipset.

[0046] The wireless communication supported by the communication unit 130 may be, for example, Wireless Fidelity (Wi-Fi), Wi-Fi Direct, Bluetooth (registered trademark), Ultra Wide Band (UWB), Near Field Communication (NFC), etc. The wired communication supported by the communication unit 130 may be, for example, USB or High Definition Multimedia Interface (HDMI), etc.

[0047] Various types of data, such as files, applications, and programs, may be embedded and stored in the memory 140. The control unit 120 may access and use data stored in the memory 140 or store new data in the memory 140. The control unit 120 may also execute programs embedded in the memory 140. A program for executing an image editing method may be embedded in the memory 140.

[0048] According to an embodiment, the memory 140 can store models for the optical flow generation module and the diffusion model-based image generation module, respectively. Here, the optical flow generation module and the diffusion model-based image generation module are each trained based on video image data present in the real world, and can generalize the learned motion in the process to reflect the motion in a user's input image.

[0049] According to one embodiment, when a drag command input for an image is received from a user through the input / output unit 110, the control unit 120 can execute a program stored in the memory 140 to output an edited image reflecting the drag on the image.

[0050] Meanwhile, the control unit 120 according to the embodiment may generate an edited image by editing the input image based on the drag command and the image.

[0051] The image editing device 100, which learns from real-world motion present in video data, may include an optical flow generation model (FlowGen) based on a generative adversarial network (GAN) and an image generation model (FlowDiffusion) based on a diffusion model. Hereinafter, the terms "optical flow generation module," "optical flow generation model," and "FlowGen" may be used interchangeably and are interpreted as having the same meaning. Furthermore, the terms "diffusion model-based image generation module," "diffusion model-based image generation model," and "FlowDiffusion" may be used interchangeably and are interpreted as having the same meaning.

[0052] FlowGen generates appropriate movement based on the user's input image and drag command, and FlowDiffusion plays a role in generating a new actual image based on the generated movement. That is, using FlowGen, which has been trained to receive a drag input command and an image and output an optical flow, the control unit 120 can obtain an optical flow based on the drag input command and the image, and can obtain an edited image by inputting the image and optical flow to FlowDiffusion, which is different from FlowGen.

[0053] This drag-based image editing system is divided into two well-designed modules. By learning natural movements from video data, it accelerates inference speed, minimizes memory usage, and enables the synthesis of images that reflect natural movements without requiring additional user input (e.g., text prompts) other than the drag command. It effectively combines FlowGen, a GAN-based system that is fast but has difficulty generating high-quality, complex data, with FlowDiffusion, a diffusion model-based system that is slow but can generate high-quality data, leveraging the strengths of each network. In this way, two different models can be used to generate edited images that reflect drag commands on an image, allowing image editing to be performed through simple network forwarding operations without optimization. Therefore, even with a typical consumer GPU (RTX3090), image editing can be performed at high speeds of around one second, requiring approximately 3GB of memory during calculation, which is approximately 75 times faster and 3.5 times less memory than the prior art DragDiffusion method.

[0054] More specifically, as shown in Figure 3, when a user inputs an image I and a drag instruction, an optical flow generation model (FlowGen) estimates dense optical flow, and a diffusion model-based image generation model (FlowDiffusion) edits the original image using flow guidance. According to the embodiment disclosed herein, since auxiliary input such as text or foreground masks is not required, and inversion and optimization are not required, an edited image I' can be provided in a short time (about 1 second).

[0055] 4, the control unit 120 may include an optical flow generation module (FlowGen) 450 and a diffusion model-based image generation module (FlowDiffusion) 470. Therefore, as shown in FIG. 4, when an input image 410 to be edited and a drag command 411 are obtained, the input image 410 and the drag command 411 are input to the optical flow generation module 450, and the result is input to the diffusion model-based image generation module 470 to obtain an edited image 490.

[0056] Here, the optical flow generation module (FlowGen) 450 can be a GAN (Generative Adversarial Networks)-based network for motion generation. FlowGen is a newer approach to generate a motion vector field based on an optical flow representation from an image.

[0057] Here, the control unit 120 can train FlowGen, which is composed of a generator and a discriminator of a Generative Adversarial Network (GAN), by executing a process in which the generator generates a fake synthetic optical flow based on an input image and a conditional drag input, and the discriminator distinguishes between the fake optical flow and the true optical flow.

[0058] 5, FlowGen includes a user input processor 510 that acquires a drag command and an input image and converts the drag command into appropriate sparse optical flow vectors, an optical flow generator 520 that generates an optical flow based on a generative adversarial network (GAN), and a normalizer 530 that normalizes the generated optical flow. Compared to FlowDiffusion, which will be described later, FlowGen learns normalized data using a different method, so FlowGen can newly process normalization at the final stage. The optical flow output through the normalizer 530 is input to FlowDiffusion, so the normalizer 530 can perform fixed-size normalization on the optical flow.

[0059] FlowGen generates an appropriate optical flow based on the input image and the user's drag input. It is based on the model structure proposed in the GAN architecture Pix2Pix (Isola et al. 2017. Image-To-Image Translation With Conditional Adversarial Networks. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR)). The generator and classifier utilize the latest research findings, using group normalization instead of instance normalization, and increasing model size by adding deeper layers. To accelerate learning and increase model capacity for large datasets, group normalization is used instead of instance normalization, and a deeper architecture is used for the PatchGAN-based classifier.

[0060] JPEG2026025822000002.jpg93169

[0061] JPEG2026025822000003.jpg29169

[0062]

number

[0063] JPEG2026025822000005.jpg41169

[0064] JPEG2026025822000006.jpg39169

[0065] Meanwhile, according to an embodiment, the diffusion model based image generation module (FlowDiffusion) 470 may be a diffusion-based network for motion-conditioned image generation.

[0066] 7 and 8, FlowDiffusion receives an input image and a generated optical flow as input, and generates an image that reflects the motion conditioned on the input image and optical flow through a denoising process 720 based on the input image and optical flow. For efficiency, this entire process is performed in the latent space of the image, and a pre-trained variational autoencoder (VAE) encoder 710 and decoder 730 are used to project the image into the latent space.

[0067] Compared to Instruct-Pix2Pix (Tim Brooks et al. 2023. Instructpix2pix: Learning to follow image editing instructions. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR)), the key difference with FlowDiffusion is that editing signals are encoded in an additional channel for the flow dimension, not the text prompt. While Instruct-Pix2Pix must reflect signals from both the text and image domains, FlowDiffusion must maintain consistency except in dragged regions and only reflect the dense flow, not the text. FlowDiffusion has the advantage of reducing the computational cost associated with text encoders.

[0068] FlowDiffusion receives the optical flow generated by FlowGen and the input image to generate the final edited image. It is based on a diffusion model and uses a U-Net structure. To perform calculations in a more efficient latent image space, it can include a pre-trained variational autoencoder encoder and decoder, which project and restore a 3xhxw image into a 4x(h / / 8)x(w / / 8) latent space. FlowDiffusion trains the diffusion model using data defined in this latent space. FlowDiffusion's U-Net receives 10 input channels, consisting of 4 channels of latent noise, 4 channels of latent image, and 2 channels of optical flow. FlowDiffusion is based on Instruct-Pix2Pix, but its main difference is that it uses optical flow conditions instead of text input. FlowDiffusion's loss function is defined as follows:

[0069]

number

[0070] FlowDiffusion uses the theorem of the noise estimation method of the diffusion model used in DDPM (Jonathan Ho et al. 2020. Denoising diffusion probabilistic models. In Conference on Neural Information Processing Systems (NeurIPS)). The control unit 120 can train FlowDiffusion to gradually remove noise required to restore an image from random noise based on the image and optical flow. That is, FlowDiffusion operates in a manner that gradually removes noise required to restore an appropriate image from random noise, taking into account the input image and optical flow indicating motion.

[0071] JPEG2026025822000008.jpg46169

[0072]

number

[0073] JPEG2026025822000010.jpg93169

[0074] Meanwhile, according to the embodiment, the control unit 120 can perform pre-processing on the video training data.

[0075] That is, the control unit 120 preprocesses the video data by acquiring a training dataset including multiple samples consisting of two images, two masks, and optical flow based on any video data, and can train FlowGen or FlowDiffusion using the preprocessed video data.

[0076] The biggest problem when training a drag-specific model (e.g., output = model(input, cdrag)) is the lack of a curated dataset consisting of three pairs: an input image, an output image, and a drag condition. Therefore, in this embodiment, a video dataset is utilized to learn natural movements. A large-scale facial video dataset, CelebV-Text (Yu et al. 2023. Celebv-text: A large-scale facial text-video dataset. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR)), is used, and the following preprocessing steps are performed: 1) The control unit 120 extracts multiple frames from the video, for example, at 10 fps. 2) The control unit 120 samples image pairs spaced a maximum of 8 frames apart using a sliding window technique. 3) The control unit 120 extracts optical flow between image pairs using FlowFormer (Huang et al. 2022. Flowformer: A transformer architecture for optical flow. In European Conference on Computer Vision (ECCV)). 4) The control unit 120 generates an object mask. According to an embodiment, the object mask can be quickly generated using YOLO (Redmon et al. 2016. You only look once: Unified, real-time object detection. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR)).

[0077] JPEG2026025822000011.jpg77167

[0078] Meanwhile, according to the embodiment, the control unit 120 can extract pseudo drag instructions for learning.

[0079] For training FlowGen, we generate pseudo-drag instructions that take into account various drag input scenarios.

[0080] JPEG2026025822000012.jpg99168

[0081] Incidentally, the control unit 120 may use some dlib key points as special points for face moving images, and may apply grid-based sampling when fine-tuning general scenes.

[0082] To provide flexibility, various GAN configurations can be used depending on the user's intentions, for example, a GAN configuration can be selected for single-point dragging, fine-motion editing based on key points, or highly detailed editing such as hair.

[0083] User inputs can vary from a single point to a large set of points, and the above allows for a wide range of variations to be handled by generating plausible motion.

[0084] Meanwhile, according to an embodiment, the control unit 120 can maintain the background based on a mask operation. For FlowDiffusion, background consistency can be applied, and the generated image can be used for FlowDiffusion training, as described below.

[0085] The difference between video generation and drag-based image editing models lies in the background. In video generation models, natural movement requires both the background and the object to change, but drag-based image editing techniques require only the specific object the user drags to move. Unlike video generation, which allows any plausible movement, drag-based image editing requires a consistent background and allows movement only in the dragged object. Therefore, for drag-based image editing, a consistent background must be maintained.

[0086] JPEG2026025822000013.jpg114168

[0087]

number

[0088] As shown in Figure 9, image 930 can be generated by extracting an object from image 910, extracting a background from another image 920, and unifying the background. Here, solid line 931 in Figure 9 represents the mask of image 920, and dashed line 932 represents the expanded mask. When training an actual diffusion model, the model receives an image that has been calculated based on the mask as input, and is trained to output images 920 and 930. That is, data is input so that only the object moves while the background remains unchanged. Here, the mask is only necessary in the training stage, not in the inference stage.

[0089] JPEG2026025822000015.jpg113170

[0090] Meanwhile, according to an embodiment, the control unit 120 can perform optical flow normalization.

[0091] For FlowGen and FlowDiffusion to effectively conditionally train on optical flow, the optical flow must be normalized to an appropriate scale. There are two normalization methods. The first is fixed-size normalization, which divides the spatial dimension and normalizes the optical flow by the image size. The second is sample-wise normalization, which divides each channel by the absolute maximum value of the flow vector for that channel. Fixed-size normalization maintains the actual size and scale, resulting in numbers directly proportional to the dimension. However, sample-wise normalization does not represent the actual flow size because all samples are densely packed around 0, resulting in a very narrow distribution. Here, sample-wise normalization is more effective for FlowGen, while fixed-size normalization is more effective for FlowDiffusion during training. This is because in FlowGen, the loss is calculated directly for the flow, while in FlowDiffusion, information on the actual scale is important. During inference, the rare flow per sample is normalized, the size is rescaled using the maximum absolute magnitude of the input rare drag command, and a fixed-size normalization is applied before inputting it into the diffusion model. This normalization process allows the diffusion model to generate reasonable results for optical flow without a specialized encoder network for optimization.

[0092] FIG. 10 is a flowchart illustrating an image editing method according to an embodiment.

[0093] The image editing method according to the embodiment shown in Figure 10 includes steps that are processed in time series by the image editing device 100 shown in Figures 1 to 9. Therefore, even if omitted below, the contents described above regarding the image editing device 100 shown in Figures 1 to 9 can also be applied to the image editing method according to the embodiment shown in Figure 10. Figure 10 will be described below with reference to Figures 11 and 12. Figures 11 and 12 are example diagrams for explaining the image editing method according to one embodiment.

[0094] As shown in FIG. 10, when the image editing device 100 acquires a drag input command and an image (S1010), it can acquire an optical flow based on the drag input command and the image using a first artificial intelligence model trained to output an optical flow in response to the input of the drag input command and the image (S1020). Here, the first artificial intelligence model is the same as the above-described FlowGen model. Therefore, to perform step S1020, the image editing device 100 can train a first artificial intelligence model including a GAN generator and a classifier, and can train the first artificial intelligence model before the inference process of image editing. According to an embodiment, the image editing device 100 can train the first artificial intelligence model by executing a process in which the generator generates a false composite optical flow based on the input image and the conditional drag input, and the classifier distinguishes between the false optical flow and the true optical flow. In addition, training data used to train the first artificial intelligence model can be generated based on any video data. The image editing device 100 can preprocess any video data. That is, based on any video data, a training data set including two images, two masks, and optical flow can be obtained. The preprocessed video data can also be used to train the second artificial intelligence model described below.

[0095] Once the image editing device 100 acquires the optical flow, it can input the optical flow and the image into a second artificial intelligence model to acquire an edited image. Here, the edited image refers to an image edited in response to a drag input command for an original image with a drag command. The optical flow is acquired as the output of the first artificial intelligence model, and the image editing device 100 can perform fixed-size normalization on the optical flow output from the first artificial intelligence model and input the result to the second artificial intelligence model. The second artificial intelligence model has the same object as FlowDiffusion. The image editing device 100 can train the second artificial intelligence model before inferring the edited image using the second artificial intelligence model. The second artificial intelligence model can be trained to progressively remove noise necessary to restore the image from random noise based on the image and optical flow. According to one embodiment, the image editing device 100 can perform both a step of training the first artificial intelligence model and a step of training the second artificial intelligence model prior to the inference process.

[0096] The image editing device 100 may provide the edited image acquired from the second artificial intelligence model. Here, providing may mean providing the edited image via a screen of a user terminal or saving it in a user account so that the user who input the drag can check it.

[0097] As shown in FIG. 11 , the image editing device 100 or the image editing method according to the embodiment disclosed in this specification performs drag input commands (shaded points (drag start points), It can be seen that when a round dot (drag end point) and an arrow (drag direction) are given, images 1111, 1116, 1121, and 1131 reflecting the drag edit are output. For example, if a drag input command is performed on input image 1110, image 1111 reflecting the command can be output. If another drag input command is performed on image 1111, as shown in image 1115, which is the same as image 1111, image 1116 can be output. Also, if a drag input command is performed on input images 1120 and 1130, images 1121 and 1131 can be output, allowing the user to naturally and quickly confirm the completed drag-edited image. Here, optical flows 1112, 1117, 1122, and 1132 when resultant images 1111, 1116, 1121, and 1131 are output from input images 1110, 1115, 1120, and 1130, respectively, are shown in FIG. 11.

[0098] It can be seen that the image editing device 100 or image editing method according to the embodiment disclosed in this specification outputs an image 1220 of the best quality when a drag input command (a diagonal dot (start point of drag), a round dot (end point of drag), and an arrow (direction of drag); Input+Drag Instr.) is performed on an input image 1210, as shown in FIG. 12 .

[0099] Such an image editing device 100 or image editing method can perform high-quality editing at high speed.

[0100] Incidentally, Table 1 below shows the evaluation results of implementing each drug-based editing technique.

[0101] [Table 1]

[0102] Image editing techniques that may be compared to the embodiments disclosed herein include DragDiffusion (Yujun Shi et al. 2023. DragDiffusion: Harnessing Diffusion Models for Interactive Point-based Image Editing. arXiv preprint arXiv:2306.14435 (2023)), DragonDiffusion (Chong Mou et al. 2023. DragonDiffusion: Enabling Drag-style Manipulation on Diffusion Models. In International Conference on Learning Representations (ICLR)), SDE-Drag (Shen Nie et al. 2023. The Blessing of Randomness: SDE Beats ODE in General Diffusion-based Image Editing. In International Conference on Learning Representations (ICLR)), and Readout Guidance (Grace Luo et al. 2024. Readout Guidance: Learning Control from Diffusion Features. IEEE Conference on Computer Vision and Pattern Recognition (CVPR)). The evaluation was performed on human face frames extracted from TalkingHead-1KH video data. 68 facial key points, such as eyes, nose, and mouth, were extracted from two different frames, and one was used as the input image and the other as the ground truth editing result. The drag input was created using the 68 key points. The type of input required for editing, time, and memory, as well as how well it reflected the original image, and the similarity to the ground truth frame were measured.Evaluation indices such as PSNR, SSIM, LPIPS, and CLIP Similarity Score measure the structural similarity and perceptual distance between two images. The higher the PSNR, SSIM, and CLIP Similarity Score, and the lower the LPIPS, the better the indices. O represents the comparison result between the input image and the edited image, and E represents the comparison result between the GT ground truth frame and the edited image. In particular, the better the indices in column E, the more realistic the editing appears to be. Here, as shown in Table 1, the embodiment (InstantDrag) disclosed herein is several tens of times faster than other techniques and uses up to five times less memory.

[0103] The term "module" used in the above embodiments refers to software or hardware components such as FPGAs (field programmable gate arrays) or ASICs, and the "module" performs a certain function. However, the term "module" is not limited to software or hardware. A "module" may be configured to reside on an addressable storage medium or to execute one or more processors. Thus, by way of example, "module" includes components such as software components, object-oriented software components, class components, and task components, as well as processes, functions, attributes, procedures, subroutines, segments of program code, drivers, firmware, microcode, circuits, data, databases, data structures, tables, arrays, and variables.

[0104] The functionality provided within components and units may be combined into fewer components and units or separated into additional components and units.

[0105] Furthermore, the components and "units" may be implemented to implement one or more CPUs within a device or a secure multimedia card.

[0106] The image editing method according to the embodiment may also be embodied in the form of a computer-readable medium storing computer-executable instructions and data. Here, the instructions and data may be stored in the form of program code, which, when executed by a processor, generates a predetermined program module and performs a predetermined operation. Furthermore, the computer-readable medium may be any available medium accessible by a computer, including both volatile and nonvolatile media, and both separable and non-separable media. The computer-readable medium may also be a computer recording medium. The computer recording medium may include both volatile and non-volatile, separable and non-separable media embodied in any method or technology for storing information such as computer-readable instructions, data structures, program modules, or other data. For example, the computer recording medium may be a magnetic storage medium such as a hard disk drive (HDD) or solid-state drive (SSD), an optical storage medium such as a CD, DVD, or Blu-ray disc, or a memory included in a server accessible via a network.

[0107] The image editing method according to the embodiment may also be implemented as a computer program (or computer program product) including computer-executable instructions. The computer program includes programmable machine instructions to be processed by a processor, and may be implemented in a high-level programming language, an object-oriented programming language, an assembly language, a machine language, etc. The computer program may also be recorded on a computer-readable recording medium (e.g., a memory, a hard disk, a magnetic / optical medium, or a solid-state drive (SSD)).

[0108] Therefore, the image editing method according to the embodiment can be implemented by executing the above-described computer program on a computing device. The computing device can include a processor, a memory, a storage device, a high-speed interface connected to the memory and a high-speed expansion port, and at least a portion of a low-speed interface connected to a low-speed bus and the storage device. Each of these components is connected to each other using various buses and can be mounted on a common motherboard or in other suitable manners.

[0109] Here, the processor may process instructions within a computing device. Such instructions may include instructions stored in a memory or storage device for displaying graphic information to provide a GUI (Graphical User Interface) on an external input and output device, such as a display connected to a high-speed interface. In other embodiments, multiple processors and / or multiple buses may be used, along with multiple memories and memory types, as appropriate. The processor may also be implemented as a chipset, which may include multiple independent analog and / or digital processors.

[0110] Also, memory stores information within a computing device. As an example, memory may be comprised of a volatile memory unit or collection thereof. As another example, memory may be comprised of a non-volatile memory unit or collection thereof. Memory may also be in other forms of computer-readable media, such as magnetic or optical disks.

[0111] The storage device can provide a large amount of storage space to a computing device. The storage device may be a computer-readable medium or a configuration that includes such a medium, such as a device in a Storage Area Network (SAN) or other configuration, and may be a floppy disk drive, hard disk drive, optical disk drive, tape drive, flash memory, or other similar semiconductor memory device or device array.

[0112] The above-described embodiments are merely illustrative, and those skilled in the art will understand that the above-described embodiments may be easily modified into other specific forms without changing the technical ideas or essential features of the above-described embodiments. Therefore, it should be understood that the above-described embodiments are illustrative in all respects and are not limiting. For example, each component described as a single component may be implemented in a distributed form, and similarly, each component described as a distributed component may be implemented in a combined form.

[0113] The scope of protection sought by this specification is determined by the claims below rather than the above detailed description, and should be construed to include all modifications or variations derived from the meaning and scope of the claims and their equivalents. [Explanation of symbols]

[0114] 100 Image editing equipment 110 Input / output section 120 control section 130 Communications Department 140 memory

Claims

1. an input / output unit for acquiring drag input commands and images; an image editing device comprising: a control unit that uses a first artificial intelligence model trained to receive a drag input command and an image input and output an optical flow, acquires an optical flow based on the drag input command and the image, inputs the optical flow and the image to a second artificial intelligence model different from the first artificial intelligence model, acquires an edited image as an output of the second artificial intelligence model, and provides the edited image.

2. 2. The image editing device of claim 1, wherein the control unit trains a first artificial intelligence model including a generator and a discriminator of a Generative Adversarial Network (GAN), the generator generating a fake synthetic optical flow based on an input image and a conditional drag input, and the discriminator distinguishing between the fake optical flow and a true optical flow, thereby training the first artificial intelligence model.

3. 3. The image editing device of claim 2, wherein the control unit preprocesses the video data by acquiring a training dataset including a plurality of samples consisting of two images, two masks, and optical flow based on arbitrary video data, and trains the first artificial intelligence model using the preprocessed video data.

4. 4. The image editing device of claim 3, wherein the control unit generates an image in which a background of a first image, one of the two images, is filled with a second image, the other image, and trains the second artificial intelligence model using a training dataset including a plurality of samples constructed by replacing the first image with the generated image.

5. 2. The image editing device of claim 1, wherein the control unit trains the second artificial intelligence model based on a diffusion model and trains the second artificial intelligence model to progressively remove noise necessary to restore an image from random noise based on an image and optical flow.

6. The image editing device of claim 1 , wherein the control unit performs fixed size normalization on the optical flow from the first artificial intelligence model and inputs the normalized optical flow to the second artificial intelligence model.

7. The control unit is configured to generate a rare flow initialized to an arbitrary value sampled at U(0, 1). [Equation 1] 2. The image editing device according to claim 1, wherein the pseudo drag input command is generated based on the pseudo drag input command, and the generated pseudo drag input command is used to train the first artificial intelligence model.

8. 2. The image editing device of claim 1, wherein the control unit performs sample-wise normalization on the optical flow when training the first artificial intelligence model, and performs fixed size normalization on the optical flow when training the second artificial intelligence model.

9. An image editing method executed by an image editing device, comprising: obtaining a drag input command and an image; acquiring an optical flow based on the drag input command and the image using a first artificial intelligence model trained to receive a drag input command and an image and output an optical flow; inputting the optical flow and the image into a second artificial intelligence model different from the first artificial intelligence model to obtain an edited image as an output of the second artificial intelligence model; and providing the edited image.

10. 10. The image editing method of claim 9, further comprising: training a first artificial intelligence model including a generator and a discriminator of a Generative Adversarial Network (GAN); the generator generating a fake synthetic optical flow based on an input image and a conditional drag input; and the discriminator distinguishing between the fake optical flow and a true optical flow, thereby training the first artificial intelligence model.

11. 11. The image editing method of claim 10, wherein the step of training the first artificial intelligence model includes the steps of preprocessing the video data by obtaining a training dataset including a plurality of samples consisting of two images, two masks, and optical flow based on arbitrary video data, and training the first artificial intelligence model using the preprocessed video data.

12. 10. The image editing method of claim 9, further comprising: training the second artificial intelligence model based on a diffusion model; and training the second artificial intelligence model to progressively remove noise necessary to restore the image from random noise based on the image and optical flow.

13. The image editing method of claim 9 , wherein the obtaining of the optical flow comprises performing fixed size normalization on the optical flow from the first artificial intelligence model.

14. A computer-readable recording medium having a program recorded thereon for executing the method of claim 9.

15. A computer program stored on a medium for executing by an image editing device and for carrying out the method of claim 9.

Citation Information

Patent Citations

  • Medical image segmentation method and device, computer device and readable storage medium

    EP3923238A1

  • Machine learning based controllable animation of still images

    US20240005587A1

  • Method and Apparatus for Generating and Editing a Detailed Image

    KR1020140120628A