Device and method for generating lora-based drawings and automating spatial structures

The LoRA-based drawing generation method addresses inefficiencies in existing methods by automating high-quality drawing production from text inputs, enhancing precision and reducing costs through iterative model fine-tuning and data augmentation.

KR1020260113495APending Publication Date: 2026-07-21JIK TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
KR · KR
Patent Type
Applications
Current Assignee / Owner
JIK TECH CO LTD
Filing Date
2025-01-13
Publication Date
2026-07-21

AI Technical Summary

Technical Problem

Existing drawing generation methods are inefficient and costly, lacking the ability to automatically produce high-quality, structured drawings from user-specified text inputs, which hinders design efficiency and increases project time and costs.

Method used

A LoRA-based drawing generation method utilizing a fine-tuned image generation model, enhanced with labeled drawing data and data augmentation, iteratively improved through user feedback, to generate high-quality drawings with customizable detail and accuracy.

Benefits of technology

Efficiently generates high-quality, structured drawings from user text inputs, reducing time and costs by automating the design process and ensuring precise spatial structure implementation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure PAT00001_ABST
    Figure PAT00001_ABST
Patent Text Reader

Abstract

According to various embodiments of the present invention, a LoRA-based automatic drawing generation method may include: a step of selecting a pre-trained image generation model, preparing labeled drawing data for drawing-specific learning of the model, and fine-tuning the model using a LoRA technique; and a step of generating a drawing image that meets the user's requirements through the fine-tuned image generation model based on a text prompt and parameter variables entered by the user.
Need to check novelty before this filing date? Find Prior Art

Description

Technology Field

[0001] The present invention relates to a LoRA-based drawing generation automation device and method. Background Technology

[0002] The content described in this section merely provides background information regarding the present embodiment and does not constitute prior art.

[0003] In architecture, drawings are a core tool for design and construction, serving to visually represent the building's ideas and specific plans. They convey details regarding the building's form, structure, materials, and spatial configuration, and enable smooth communication between the designer and the contractor. Through this, they accurately reflect design intent and minimize errors or misunderstandings, contributing to improved project efficiency and quality.

[0004] Furthermore, drawings play a crucial role in the legal approval process and construction management. They serve as a standard for verifying whether a building complies with local laws and regulations, and during the construction phase, they function as a guideline providing specific instructions on the sequence and methods of work. Therefore, the precision and accuracy of drawings are directly linked to the success of a construction project, enabling the realization of a high-quality building. Prior art literature

[0005] Republic of Korea Published Patent Application No. 10-2024-0171916 A (December 9, 2024) The problem to be solved

[0006] The objective of the present invention is to provide a LoRA-based drawing generation automation device and method that can efficiently generate high-quality drawing images by applying the LoRA technique to an image generation model to perform learning specialized for drawings, thereby maximizing the efficiency of design work by automatically generating data with structured drawings and labeled objects based solely on user-specified text input, and reducing time and costs by quickly and accurately implementing spatial structures or design requirements desired by the user.

[0007] Other unspecified objects of the present invention may be further considered to the extent that they can be easily inferred from the following detailed description and effects. means of solving the problem

[0008] A method performed in a device comprising a memory storing one or more programs for automatically generating drawings based on LoRA according to an embodiment of the present invention for achieving the above-described purpose, and one or more processors performing operations according to said one or more programs, may include: a step of selecting a pre-trained image generation model and preparing labeled drawing data for drawing-specific learning of said model, and fine-tuning the model using a LoRA technique; and a step of generating a drawing image that meets the user's requirements through the fine-tuned image generation model based on a text prompt and parameter variables entered by the user.

[0009] The step of preparing the above-mentioned labeled drawing data includes data in which objects and spatial structures are labeled by color, and is characterized by applying a data augmentation technique to improve the drawing generation performance of the model.

[0010] The text prompt entered by the user includes specific structures, object placements, or design features of the drawing, and the level of detail and quality of the generated drawing are customizable through the adjustment of the guidance scale and the number of inference steps.

[0011] The above-described fine-tuned image generation model is characterized by being iteratively improved by verifying the generated drawing and readjusting the model with additional training data if it contains an incorrect structure or object placement.

[0012] A computer program according to one embodiment of the present invention for achieving the above-described purpose is stored on a computer-readable recording medium and executes any one of the above-described LoRA-based drawing generation automation methods on a computer. Effects of the invention

[0013] As described above, according to one embodiment of the present invention, by applying a LoRA-based drawing generation automation device and method, high-quality drawing images can be efficiently generated by applying LoRA techniques to an image generation model and performing learning specialized for drawings. Through this, structured drawings and data labeled with objects can be automatically generated with only user-specified text input, thereby maximizing the efficiency of design work and reducing time and costs by quickly and accurately implementing the spatial structure or design requirements desired by the user.

[0014] Even if an effect is not explicitly mentioned herein, the effects and potential effects described in the following specification expected by the technical features of the present invention are treated as described in the specification of the present invention. Brief explanation of the drawing

[0015] FIG. 1 is a flowchart illustrating a LoRA-based drawing generation automation method according to one embodiment of the present invention. FIG. 2 is a diagram for explaining the architecture of a stable Diffusion model used in a LoRA-based drawing generation automation device and method according to an embodiment of the present invention, a diffusion process operating in latent space, a denoising U-Net structure, and an image generation process through conditions such as text or images. FIG. 3 is a diagram for explaining the training of a stable Diffusion model or a similar text-image based generation model and the generation of high-resolution images used in a LoRA-based drawing generation automation device and method according to an embodiment of the present invention. FIG. 4 is a diagram showing the configuration of a LoRA-based drawing generation automation device according to one embodiment of the present invention. Specific details for implementing the invention

[0016] Hereinafter, embodiments of the present invention will be described in detail with reference to the attached drawings. The advantages and features of the present invention, and the methods for achieving them, will become clear by referring to the embodiments described below in detail together with the attached drawings. However, the present invention is not limited to the embodiments disclosed below but may be implemented in various different forms. These embodiments are provided merely to ensure that the disclosure of the present invention is complete and to fully inform those skilled in the art of the scope of the invention, and the present invention is defined only by the scope of the claims. Unless otherwise defined, all terms used in this specification (including technical and scientific terms) may be used in a meaning that is commonly understood by those skilled in the art to which the present invention belongs. Furthermore, terms defined in commonly used dictionaries are not to be interpreted ideally or excessively unless explicitly and specifically defined otherwise.

[0017] The terms used in this application are used merely to describe specific embodiments and are not intended to limit the invention. Singular expressions include plural expressions unless the context clearly indicates otherwise. In this application, terms such as “have,” “may have,” “include,” or “may include” are intended to indicate the existence of the features, numbers, steps, actions, components, parts, or combinations thereof described in the specification, and should be understood as not precluding the existence or addition of one or more other features, numbers, steps, actions, components, parts, or combinations thereof. Terms including ordinal numbers, such as “second,” “first,” etc., may be used to describe various components, but said components are not limited by said terms.

[0018] The above terms are used solely for the purpose of distinguishing one component from another. For example, without departing from the scope of the present invention, the second component may be named the first component, and similarly, the first component may be named the second component. The term "and / or" includes a combination of a plurality of related described items or any of a plurality of related described items.

[0019] In this specification, identification symbols (e.g., a, b, c, etc.) for each step are used for convenience of explanation and do not indicate the order of the steps; the steps may occur differently from the specified order unless the context clearly indicates a specific order. That is, the steps may occur in the same order as specified, may be performed substantially simultaneously, or may be performed in the reverse order.

[0020] Various embodiments of the LoRA-based drawing generation automation apparatus and method according to the present invention will be described in detail below with reference to the attached drawings.

[0021] FIG. 1 is a flowchart illustrating a LoRA-based drawing generation automation method according to one embodiment of the present invention.

[0022] A LoRA-based drawing generation automation method may be performed in a device comprising a memory storing one or more programs for automatically generating drawings based on LoRA and one or more processors performing operations according to one or more programs. For example, a LoRA-based drawing generation automation method may be performed by a LoRA-based drawing generation automation device described through FIG. 4.

[0023] In step S100, a pre-trained image generation model may be selected, labeled drawing data may be prepared for drawing-specific training of the model, and the model may be fine-tuned using the LoRA technique.

[0024] In step S200, a step of generating a drawing image that meets the user's requirements through a finely tuned image generation model based on text prompts and parameter variables entered by the user may be performed.

[0025] The step of preparing labeled drawing data includes data in which objects and spatial structures are labeled by color, and data augmentation techniques may be applied to improve the drawing generation performance of the model.

[0026] The text prompt entered by the user includes specific structures, object placements, or design features of the drawing, and the level of detail and quality of the generated drawing can be customized by adjusting the guidance scale and the number of inference steps.

[0027] A fine-tuned image generation model can be iteratively improved by verifying the generated drawings and readjusting the model with additional training data if incorrect structures or object placements are included.

[0028] Image generation model

[0029] It uses Stable Diffusion and latent diffusion models. This model adopts a method of generating images in a low-dimensional latent space instead of a high-dimensional image space.

[0030] The core structure of the model consists of an autoencoder, U-Net, and a text encoder. The autoencoder is responsible for encoding and decoding images, while U-Net removes noise and progressively generates images. The text encoder performs the important function of converting user-input prompts into latent vectors.

[0031] The image generation process begins with random noise and completes the final image through an iterative denoising process, removing noise and progressively adding detailed features of the image at each stage.

[0032] The input elements include a text prompt containing an image description, a seed controlling randomness, a guidance scale determining how much to follow the text prompt, and the number of inference steps determining image quality.

[0033] The output image can generally be generated with a resolution of 512x512, 768x768, or 1024x1024 pixels.

[0034] LoRA Application Mechanism

[0035] Low-Rank Adaptation is a method for efficiently fine-tuning the weights of existing large-scale pre-trained models. With the primary model weights fixed, the model is selectively adapted to a specific domain through low-dimensional adapter matrices. This LoRA is used in drawing learning. Labeled drawing data is used for additional training with LoRA, based on pretrained weights such as the previously described image generation model, Stable Diffusion. A model tuned in this way performs the functions of an image generation model while creating a generative model specialized for drawings.

[0036] Drawing data generation.

[0037] Select the image generation model and additionally select the LoRA model trained on the drawing. Enter "Generate drawing image" at the prompt. Generate the drawing while adjusting the parameter variables. The generated drawing is similar to the drawing trained using LoRA. The generated drawing is data that is already color-labeled according to structure and objects.

[0038] Although FIG. 1 describes each process as being executed sequentially, this is merely an illustrative description, and a person skilled in the art may modify and adapt the process in various ways without departing from the essential characteristics of the embodiment of the present invention, such as changing the order described in FIG. 1, executing one or more processes in parallel, or adding other processes.

[0039] FIG. 2 is a diagram for explaining the architecture of a stable Diffusion model used in a LoRA-based drawing generation automation device and method according to an embodiment of the present invention, a diffusion process operating in latent space, a denoising U-Net structure, and an image generation process through conditions such as text or images.

[0040] Stable Diffusion is an approach for generating high-resolution images that aims to overcome the limitations of existing diffusion models. Instead of directly predicting pixel values, this model utilizes an autoencoder to compress images in a latent space and performs training based on that compressed representation. This reduces computational load and allows for greater focus on the semantic aspects of the images.

[0041] The overall structure of Stable Diffusion consists of three main parts. First, there is the 'Perceptual Image Compression' stage, which compresses images into a latent space using an autoencoder. Second, the 'Latent Diffusion Model' is at the center, learning the compressed latent representations. Third, there is a 'Conditioning Mechanism' to apply conditional information, such as text or layout, to the model. In particular, conditional information is applied internally within the model through Cross Attention operations.

[0042] Through this structure, Stable Diffusion can perform image generation based on various conditions. For example, it is possible to generate images based on text descriptions or create images tailored to specific layouts. Additionally, it can be utilized for 'Super Resolution' tasks to convert low-resolution images to high resolution, or for 'Inpainting' tasks to fill specific parts of an image.

[0043] Another feature of Stable Diffusion is that it trains the autoencoder using perceptual loss during the 'Perceptual Image Compression' process. This ensures that the compressed latent representation retains important features of the original image. Furthermore, the model's conditional inputs are effectively applied through a cross-attention mechanism, enabling the generation of images under various conditions.

[0044] Experimental results showed that Stable Diffusion successfully generated high-quality images under various conditions. It demonstrated excellent performance in diverse tasks, including image generation based on text descriptions, image generation according to specific layouts, conversion of low-resolution images to high resolution, and filling of specific parts of images.

[0045] In conclusion, Stable Diffusion presents a novel approach for efficiently generating high-resolution images by combining image compression via an autoencoder with a latent diffusion model. This reduces the computational complexity of existing models and demonstrates superior performance in image generation tasks under various conditions.

[0046] FIG. 3 is a diagram for explaining the training of a stable Diffusion model or a similar text-image based generation model and the generation of high-resolution images used in a LoRA-based drawing generation automation device and method according to an embodiment of the present invention.

[0047] Fine-tuning methods for Stable Diffusion models include Textual Inversion, DreamBooth, and full model fine-tuning. Textual Inversion involves training a text encoder with a dataset containing new words, while DreamBooth is a method for fine-tuning UNets. Full model fine-tuning is also possible if a sufficient dataset is available.

[0048] Huggingface's Accelerate library simplifies complex configurations, handling setups for Multi-GPU, TPU, and fp16 with just a few lines of code. This makes fine-tuning Stable Diffusion models easier.

[0049] Fine-tuning of the entire Stable Diffusion model can be performed using Accelerate and Diffusers. For example, you can fine-tune the model by specifying the training directory using the CompVis / stable-diffusion-v1-4 model and running the train_text_to_image.py script with mixed precision (fp16) settings. While training is possible with a 24GB GPU, a GPU of 30GB or more is recommended to increase the batch size.

[0050] DreamBooth is a text-to-image model released by Google Research that aims to be a personalized text-to-image model. It allows for fine-tuning with a small number of images by using unique identifiers ([V]) and class names, and can generate images in new contexts while preserving the semantic knowledge of existing models. However, it may have drawbacks such as overfitting when training data is scarce or when the prompt and the actual image are similar.

[0051] You can apply DreamBooth to the Stable Diffusion model using Diffusers. You can configure the pretrained model, instance data directory, and output directory, and fine-tune the model by running the train_dreambooth.py script. Afterward, you can generate images from the trained model using the StableDiffusionPipeline.

[0052] Textual Inversion is a technique that learns new concepts using a small number of images by adding new words to a text encoder. It can be easily used via Diffusers and Accelerate; training is performed by setting up a pretrained model and a training data directory, and running the textual_inversion.py script. The trained model can then generate images with prompts containing new words.

[0053] Low-Rank Adaptation (LoRA) is a method for efficiently fine-tuning the weights of large-scale pre-trained models, and it can also be applied to the cross-attention layer of the UNet in the Stable Diffusion model. This offers advantages such as faster training speeds, reduced memory requirements, and lighter trained file sizes. Fine-tuning with LoRA can be performed using Diffusers, and training can be performed using the `train_text_to_image_lora.py` script. The trained model can generate images by loading the LoRA layer into the UNet of the StableDiffusionPipeline.

[0054] Since Riffusion does not use the StableDiffusionPipeline, the RiffusionPipeline must be modified directly.

[0055] FIG. 4 is a diagram showing the configuration of a LoRA-based drawing generation automation device according to one embodiment of the present invention.

[0056] A LoRA-based drawing generation automation device (100) includes at least one processor (110), a computer-readable storage medium (120), and a communication bus (150).

[0057] The processor (110) can be controlled to operate as a LoRA-based drawing generation automation device (100). For example, the processor (110) can execute one or more programs (121) stored in a computer-readable storage medium (120). One or more programs (121) may include one or more computer-executable instructions, and the computer-executable instructions may be configured to cause the LoRA-based drawing generation automation device (100) to perform operations according to exemplary embodiments when executed by the processor (110).

[0058] A computer-readable storage medium (120) is configured to store computer-executable instructions or program code, program data and / or other suitable forms of information. Computer-executable instructions or program code, program data and / or other suitable forms of information may also be provided through an input / output interface (130) or a communication interface (140). A program (121) stored in the computer-readable storage medium (120) includes a set of instructions executable by a processor (110). In one embodiment, the computer-readable storage medium (120) may be memory (volatile memory such as random access memory, non-volatile memory, or a suitable combination thereof), one or more magnetic disk storage devices, optical disk storage devices, flash memory devices, or other forms of storage media that are accessed by the LoRA-based drawing generation automation device (100) and capable of storing desired information, or a suitable combination thereof.

[0059] The communication bus (150) interconnects various other components of the LoRA-based drawing generation automation device (100), including the processor (110) and the computer-readable storage medium (120).

[0060] The LoRA-based drawing generation automation device (100) may also include one or more input / output interfaces (130) and one or more communication interfaces (140) that provide interfaces for one or more input / output devices. The input / output interface (130) and the communication interface (140) are connected to a communication bus (150). An input / output device (not shown) may be connected to other components of the LoRA-based drawing generation automation device (100) through the input / output interface (130).

[0061] The processor (110) can perform operations according to the contents described through FIGS. 1 to 3.

[0062] Among the various components illustrated exemplarily in FIG. 4, some components may be omitted or additional components may be included.

[0063] The present application also provides a computer storage medium. Program instructions are stored in the computer storage medium, and when the program instructions are executed by a processor, the above-described LoRA-based automatic drawing generation method is realized.

[0064] A computer storage medium according to one embodiment of the present invention may be a U disk, SD card, PD optical drive, mobile hard disk, large capacity floppy drive, flash memory, multimedia memory card, server, etc., but is not necessarily limited thereto.

[0065] Although it is described that all components constituting the embodiments of the present invention described above are combined or operate in combination, the present invention is not necessarily limited to such embodiments. That is, within the scope of the purpose of the present invention, all such components may be selectively combined in one or more ways to operate. Furthermore, while all such components may each be implemented as a single independent piece of hardware, they may also be implemented as a computer program having a program module that performs some or all of the combined functions in one or more pieces of hardware by selectively combining some or all of the components. Additionally, such a computer program may be stored on a computer-readable media such as a USB memory, CD disk, or flash memory, and read and executed by a computer to implement the embodiments of the present invention. Magnetic recording media, optical recording media, etc., may be included as recording media for the computer program.

[0066] The foregoing description is merely an illustrative explanation of the technical concept of the present invention, and those skilled in the art to which the present invention pertains will be able to make various modifications, changes, and substitutions within the scope of the essential characteristics of the present invention. Accordingly, the embodiments disclosed in the present invention and the accompanying drawings are intended to explain, not limit, the technical concept of the present invention, and the scope of the technical concept of the present invention is not limited by such embodiments and accompanying drawings. The scope of protection of the present invention shall be interpreted by the claims below, and all technical concepts within an equivalent scope shall be interpreted as being included within the scope of rights of the present invention.

Claims

Claim 1 A method for automatically generating drawings based on LoRA, performed in a device comprising a memory storing one or more programs for automatically generating drawings based on LoRA and one or more processors performing operations according to said one or more programs, comprising: a step of selecting a pre-trained image generation model and preparing labeled drawing data for drawing-specific learning of said model, and fine-tuning the model using a LoRA technique; and a step of generating a drawing image that meets the user's requirements through the fine-tuned image generation model based on a text prompt and parameter variables entered by the user. Claim 2 A LoRA-based automatic drawing generation method according to claim 1, wherein the step of preparing the labeled drawing data includes data in which objects and spatial structures are labeled by color, and applies a data augmentation technique to improve the drawing generation performance of the model. Claim 3 A LoRA-based automatic drawing generation method according to claim 1, wherein the text prompt entered by the user includes a specific structure, object placement, or design feature of the drawing, and the level of detail and quality of the generated drawing can be customized through the adjustment of the guidance scale and the number of inference steps. Claim 4 A LoRA-based automatic drawing generation method according to claim 1, wherein the fine-tuned image generation model is iteratively improved by verifying the generated drawing and readjusting the model through additional training data if it contains an incorrect structure or object placement. Claim 5 A computer program stored on a computer-readable recording medium for executing the LoRA-based drawing automatic generation method described in any one of paragraphs 1 to 4 on a computer.