User prompt-based image generation method and system therefor

WO2026205681A1PCT designated stage Publication Date: 2026-10-01NDOTLIGHT CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/KR2025/019330
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2025-03-28
Filing Date
2025-11-20
Publication Date
2026-10-01

Smart Images

  • Figure KR2025019330_01102026_PF_FP_ABST
    Figure KR2025019330_01102026_PF_FP_ABST
Patent Text Reader

Abstract

Provided are a user prompt-based image generation method and a system therefor. The image generation method according to some embodiments may comprise the steps of: acquiring a user prompt including a description of a target image; generating an enhanced prompt by supplementing scene context of the user prompt by using a generative language model; and generating the target image from the enhanced prompt by using a generative image model. According to this method, the difficulty of writing a user prompt can be reduced, inconvenience can be reduced, and as a specified prompt is provided to the generative image model, the accuracy and consistency of target image generation can also be improved.
Need to check novelty before this filing date? Find Prior Art

Description

User prompt-based image generation method and system

[0001] The present disclosure relates to a technique for generating images based on a user prompt.

[0002] Recently, technology for generating images based on user prompts has been rapidly advancing. This technology utilizes deep learning-based generative models to interpret text entered by the user and automatically generate a corresponding image.

[0003] Prompt-based image generation technology is widely utilized in various fields such as design, content creation, advertising, and research, and is receiving significant attention for its ability to generate customized images tailored to user needs. However, existing technologies have the following limitations.

[0004] The first limitation is that creating prompts to obtain the desired results is difficult and cumbersome. Since the image generation outcome depends heavily on the input prompts, simply entering keywords is insufficient to achieve the desired result; instead, the target image must be described in detail to align with the requirements of the generative model. For instance, the desired style, object placement, and lighting environment must be clearly described. However, drafting prompts that consider all these specific details is a challenging and difficult task for the average user, and the inconvenience is significant as it requires multiple revisions and iterative attempts to achieve the desired outcome.

[0005] The second limitation is the difficulty of consistently generating sophisticated target images that match the prompt. Due to the limitations of generative models, the generated images may vary even when the same prompt is input, and issues such as a lack of specific details or distortion can frequently occur.

[0006] Various technical problems to be solved through some embodiments of the present disclosure relate to a method and system for accurately generating a target image based on a user prompt.

[0007] Specifically, the technical problem to be solved through some embodiments of the present disclosure is to provide a method and a system that can lower the difficulty of creating user prompts and reduce the inconvenience.

[0008] In addition, another technical problem to be solved by some embodiments of the present disclosure is to provide a method and a system capable of consistently generating a sophisticated target image according to an input prompt.

[0009] The technical problems of the present disclosure are not limited to those mentioned above, and other unmentioned technical problems will be clearly understood by a person skilled in the art of the present disclosure from the description below.

[0010] A user prompt-based image generation method according to some embodiments of the present disclosure for solving the above-described technical problem may include a method performed by at least one processor, comprising: a step of obtaining a user prompt including a description of a target image; a step of generating an enhanced prompt by supplementing the scene context of the user prompt through a generative language model; and a step of generating the target image from the enhanced prompt through a generative image model.

[0011] In some embodiments, the scene context may include placement information of foreground objects and background information.

[0012] In some embodiments, the scene context may include at least one of camera view information and lighting information.

[0013] In some embodiments, the generative language model may be fine-tuned using a design-related text dataset.

[0014] In some embodiments, the step of generating the reinforcement prompt may include the step of preparing a reinforcement prompt template—the reinforcement prompt template includes one or more incomplete fields related to the scene context—; and the step of supplementing the values ​​of the one or more incomplete fields based on the user prompt through the generative language model.

[0015] In some embodiments, the generative language model is a vision-language model further equipped with image understanding capabilities, and the step of supplementing the value of one or more incomplete fields may include the step of obtaining a source image associated with the user prompt; and the step of supplementing the value of one or more incomplete fields based further on the source image.

[0016] In some embodiments, the step of generating the reinforcement prompt may include: generating a first reinforcement prompt through the generative language model; deriving self-feedback for the first reinforcement prompt through the generative language model; and generating a second reinforcement prompt by updating the first reinforcement prompt by reflecting the self-feedback through the generative language model.

[0017] In some embodiments, the description includes text regarding the design concept of the target image, and the step of generating the reinforcement prompt may include the step of generating the reinforcement prompt by supplementing the scene context to the design concept through the generative language model.

[0018] In some embodiments, the step of generating the reinforcement prompt may include: generating an initial reinforcement prompt through the generative language model and providing it to the user; obtaining the user's modification prompt for the initial reinforcement prompt; and updating the modification prompt through the generative language model to generate the reinforcement prompt.

[0019] In some embodiments, the scene context includes at least one of camera view information and lighting information related to a foreground object, and the step of generating the enhancement prompt may include: obtaining the at least one of the information from a 3D model file or 3D modeling tool of the foreground object; supplementing the values ​​of at least some of the incomplete fields of the enhancement prompt template by reflecting the obtained information; and further supplementing the values ​​of the incomplete fields through the generative language model to generate the enhancement prompt.

[0020] In some embodiments, the step of generating the target image may include: acquiring a source image associated with the user prompt; estimating a depth map of the source image; generating depth guidance information based on the depth map; and conditioning the depth guidance information to the generative image model to generate the target image.

[0021] In some embodiments, the step of generating the depth guidance information may include: adjusting the noise level of an area corresponding to the background of the source image in the depth map so that some noise remains; and generating the depth guidance information based on the adjusted depth map.

[0022] In some embodiments, the step of generating the target image may include: acquiring a source image associated with the user prompt—the source image includes a foreground object—; generating a mask that distinguishes the foreground object and the background in the source image; and generating the target image by conditioning the mask to the generative image model.

[0023] In some embodiments, the step of generating the target image may include: obtaining a source image associated with the user prompt—the source image includes a foreground object—; obtaining an output image of the generative image model; and compositing the foreground object onto the output image to generate the target image.

[0024] In some embodiments, the generative image model is a diffusion-based model, and the step of generating the target image may include: encoding the enhancement prompt to generate prompt guidance information; preparing a noise image; and conditioning the prompt guidance information to the generative image model and performing a denoising process on the noise image to generate the target image.

[0025] An image generation system according to some embodiments of the present disclosure for solving the technical problem described above comprises: one or more processors; and a memory for storing a computer program executed by the one or more processors, wherein the computer program may include instructions for: obtaining a user prompt containing a description of a target image; generating an enhanced prompt by supplementing the scene context of the user prompt through a generative language model; and generating the target image from the enhanced prompt through a generative image model.

[0026] A non-transitory computer-readable recording medium according to some embodiments of the present disclosure for solving the above-described technical problem is a non-transitory computer-readable recording medium in which instructions are stored that, when executed by at least one processor, cause said at least one processor to perform a user prompt-based image generation method, wherein the image generation method may include: obtaining a user prompt containing a description of a target image; generating an enhanced prompt by supplementing the scene context of said user prompt through a generative language model; and generating said target image from said enhanced prompt through a generative image model.

[0027] According to some embodiments of the present disclosure, a high-quality enhanced prompt can be generated by sufficiently supplementing missing elements of the scene context in the user prompt through a generative language model. In this case, the difficulty of writing the user prompt is lowered and the hassle can be reduced, and the accuracy and consistency of target image generation can also be improved as the specified prompt is provided to the generative image model.

[0028] In addition, user prompts can be enhanced through a generative language model fine-tuned using a design-related text dataset. In this case, high-quality enhanced prompts can be effectively generated.

[0029] In addition, a reinforcement prompt template containing incomplete fields regarding various elements of the scene context can be prepared, and reinforcement prompts can be generated by supplementing the values ​​of these incomplete fields through a generative language model. In this case, even if the user enters only a few keywords related to the target image, a professional, high-quality prompt (i.e., reinforcement prompt) can be generated, and as a result, the accuracy and consistency of target image generation can be further improved.

[0030] In addition, enhancement prompts can be generated by further considering information contained in the source image (e.g., placement, type, shape, design, etc. of foreground objects) through a vision-image model. In such cases, high-quality enhancement prompts can be generated more effectively.

[0031] In addition, by utilizing masks and / or depth maps obtained from source images as guidance information for a generative image model, sophisticated target images that correspond to enhancement prompts can be consistently generated (e.g., consistency in the placement of foreground objects, backgrounds, etc., and consistency in image generation according to the description of enhancement prompts can be maintained at a high level).

[0032] Additionally, the noise level can be adjusted so that some noise is present in the background area of ​​the depth map estimated from the source image. In this case, background generation errors that may occur due to excessive noise (e.g., the generation of unintended background objects) are prevented, while the residual noise acts as a hint during the background generation process, thereby further improving the accuracy and consistency of target image generation.

[0033] The effects according to the technical concept of the present disclosure are not limited to those mentioned above, and other unmentioned effects will be clearly understood by a person skilled in the art from the description below.

[0034] FIGS. 1 and 2 are exemplary drawings for explaining the operation of an image generation system according to some embodiments of the present disclosure at the system level.

[0035] FIGS. 3 and 4 are exemplary drawings for further explaining the operation of an image generation system according to some embodiments of the present disclosure.

[0036] FIG. 5 is an exemplary flowchart schematically illustrating a user prompt-based image generation method according to some embodiments of the present disclosure.

[0037] FIG. 6 is an exemplary flowchart illustrating a method for generating an enhancement prompt according to some embodiments of the present disclosure.

[0038] FIG. 7 illustrates a reinforcement prompt template that can be used in some embodiments of the present disclosure.

[0039] FIGS. 8 and 9 are exemplary drawings for further illustrating a method for generating an enhanced prompt according to some embodiments of the present disclosure.

[0040] Figures 10a and 10b show actual examples of generating an enhancement prompt.

[0041] FIG. 11 is an exemplary drawing for illustrating a method for generating an enhancement prompt according to some other embodiments of the present disclosure.

[0042] FIG. 12 is an exemplary flowchart illustrating a method for generating a target image according to some embodiments of the present disclosure.

[0043] FIG. 13 is an exemplary drawing for further explaining a target image generation method according to some embodiments of the present disclosure.

[0044] FIG. 14 and FIG. 15a to 15c are exemplary drawings for illustrating a depth map noise adjustment method according to some embodiments of the present disclosure.

[0045] FIG. 16 is an exemplary drawing for illustrating a method for generating a target image using a diffusion-based generative image model according to some embodiments of the present disclosure.

[0046] FIG. 17 illustrates an exemplary computing device capable of implementing an image generation system according to some embodiments of the present disclosure.

[0047] Hereinafter, various embodiments of the present disclosure will be described in detail with reference to the attached drawings. The advantages and features of the present disclosure and the methods for achieving them will become clear by referring to the embodiments described below in detail together with the attached drawings. However, the technical concept of the present disclosure is not limited to the following embodiments but can be implemented in various different forms. The following embodiments are provided merely to complete the technical concept of the present disclosure and to fully inform those skilled in the art of the scope of the present disclosure, and the technical concept of the present disclosure is defined only by the scope of the claims.

[0048] In describing the various embodiments of the present disclosure, if it is determined that a detailed description of related known configurations or functions could obscure the essence of the present disclosure, such detailed description is omitted.

[0049] Unless otherwise defined, terms used in the following embodiments (including technical and scientific terms) may be used in a meaning commonly understood by those skilled in the art to which this disclosure pertains, but this may vary depending on the intent of those skilled in the art, case law, the emergence of new technology, etc. The terms used in this disclosure are for describing the embodiments and are not intended to limit the scope of this disclosure.

[0050] In the following embodiments, singular expressions include plural concepts unless the context clearly specifies them as singular. Additionally, plural expressions include singular concepts unless the context clearly specifies them as plural.

[0051] In addition, terms such as first, second, A, B, (a), (b), etc. used in the following embodiments are used merely to distinguish one component from another, and the essence, order, or sequence of the said component is not limited by such terms.

[0052] The components described by reference to terms such as part or unit, module, block, ~or, ~er, etc. used in the following embodiments, and the functional blocks illustrated in the drawings may be implemented in the form of software, hardware, or a combination thereof. Software may be, for example, machine code, firmware, embedded code, and application software. Additionally, hardware may include, for example, electrical circuits, electronic circuits, processors, computers, integrated circuits, integrated circuit cores, passive components, or a combination thereof.

[0053] Hereinafter, various embodiments of the present disclosure will be described in detail with reference to the attached drawings.

[0054] FIG. 1 is an exemplary drawing for explaining the operation of an image generation system (10) according to some embodiments of the present disclosure at the system level.

[0055] As illustrated in FIG. 1, the image generation system (10) is a computing device / system equipped with the ability to generate images based on a user prompt (12). For example, the image generation system (10) can generate a consistently sophisticated target image (13) according to the user prompt (12) using a generative image model (11). Alternatively, as illustrated in FIG. 2, the image generation system (10) can generate a consistently sophisticated target image (23) based on a source image (22) containing a user prompt (21) and a foreground object.

[0056] The generative image model (11) is a deep learning model equipped with image generation capabilities, and may be a model trained (or configured) to generate a target image (e.g., 13) using a user prompt (e.g., 12) as a key input / condition, guidance information, etc. Examples of the generative image model (11) may include a diffusion-based model and a GAN (Generative Adversarial Network)-based model, but the scope of the present disclosure is not limited thereto. In addition, examples of diffusion-based models may include DDPM (Denoising Diffusion Probabilistic Model), DDIM (Denoising Diffusion Implicit Model), LDM (Latent Diffusion Model), and a score-based generative model, but the scope of the present disclosure is not limited thereto. The generative image model (11) may be located inside the image generation system (10) or may operate externally (e.g., at a remote location). The generative image model (11) may be named, depending on the case, as an 'image generation model', 'generative AI (Artificial Intelligence) model', 'generative AI', 'generative deep learning model', 'image generator', 'text-to-image model', etc.

[0057] A user prompt (e.g., 12) is information that induces the generative image model (11) to generate a desired result, and may include a description of a target image (e.g., 13). The description of the target image (e.g., 13) may include, for example, information (or text / description) about at least some of the various elements of the scene context, and examples of such elements include, but are not limited to, the type, placement and shape of foreground objects, the relationship between foreground objects, the background (e.g., type, placement and shape of background objects, the relationship between foreground objects and background objects, etc.), lighting, camera view, style, mood, design concept, etc.

[0058] The target image (e.g., 13) may refer to an image desired by the user (i.e., an image corresponding to the user prompt (e.g., 12)) or an image generated through the generative image model (11). The target image (e.g., 13) may be an output image of the generative image model (11), or an image to which post-processing has been applied to the output image. Depending on the case, the target image (e.g., 13) may be named as 'target image', 'synthetic image', 'result image', 'generated image', etc.

[0059] The source image (22) refers to an image that serves as the basis for the creation of the target image (23) or is used as a reference. The source image (22) may be, for example, an image containing a foreground object (e.g., a bag), but the scope of the present disclosure is not limited thereto. For instance, the source image (22) may be an image related to a specific background or style. The source image (22) can be utilized for various purposes, which will be described later. Depending on the method of use, the source image (22) may be named as an 'object (or background) image', 'reference image', 'base image', etc.

[0060] Below, the operation of the image generation system (10) described above will be explained in more detail with reference to FIG. 3.

[0061] FIG. 3 is an exemplary drawing for further explaining the operation of an image generation system (10) according to some embodiments of the present disclosure. FIG. 3 assumes a case where a user prompt (33) and a source image (34) are used together.

[0062] As illustrated in FIG. 3, the image generation system (10) can acquire a prompt (33) written by a user (32) and enhance the user prompt (33) through a generative language model (31). For example, if the user prompt (33) consists of only a few keywords (e.g., bag, table), the image generation system (10) can understand the content of the user prompt (33) through the generative language model (31) and, based on the result, generate an enhanced prompt (35) by sufficiently supplementing the missing elements of the scene context. By doing so, the difficulty of writing the user prompt (33) is lowered, the hassle is reduced, and the effect of consistently generating a sophisticated target image (36) can also be achieved.

[0063] For reference, the generative language model (31) refers to a deep learning model equipped with the ability to understand and generate language (or text) (e.g., a large-scale language model such as LLaMA (Large Language Model Meta AI)), and in some cases, it may be a multi-modal model equipped with the ability to understand images (e.g., a vision-language model).

[0064] Next, the image generation system (10) can generate a target image (36) from an enhancement prompt (35) through a generative image model (11). For example, the image generation system (10) can consistently generate a sophisticated target image (36) by using the enhancement prompt (35) and the source image (34) as guidance information for the generative image model (11). The detailed operations of the image generation system (10) will be described in detail with reference to the drawings from Fig. 5 onwards.

[0065] Meanwhile, in some embodiments, as illustrated in FIG. 4, an image generation system (10) may provide image generation services to a number of users (e.g., 42). For instance, the image generation system (10) may receive a user prompt (43) from a terminal (41) of a specific user (42) and generate and provide a target image (44) based on it. The image generation system (10) may provide such services through a web or app-based interface, but the scope of the present disclosure is not limited thereto. The network illustrated in FIG. 4 may be implemented as any type of wired or wireless network, including a Local Area Network (LAN), a Wide Area Network (WAN), a mobile radio communication network, and Wibro (Wireless Broadband Internet).

[0066] The image generation system (10) described above may be implemented with at least one computing device. For example, all functions of the image generation system (10) may be implemented in a single computing device, or the first function of the image generation system (10) may be implemented in a first computing device and the second function may be implemented in a second computing device. Alternatively, specific functions of the image generation system (10) may be implemented in multiple computing devices.

[0067] A computing device may encompass any device equipped with computing (or processing) functions, and for an example of such a device, refer to FIG. 17. Since a computing device is a collection of various components (e.g., memory, processor, etc.) that interact with each other, it may be referred to as a 'computing system' depending on the case. Of course, the term computing system may also encompass the concept of a collection of multiple computing devices that interact with each other.

[0068] Up to now, an image generation system (10) according to several embodiments of the present disclosure has been described with reference to FIGS. 1 to 4. Hereinafter, various methods (i.e., detailed operations) that can be performed in the image generation system (10) described above will be described with reference to FIGS. 5 and subsequent drawings.

[0069] For the sake of convenience of understanding, the following description will continue under the assumption that all steps / operations of the methods described below are performed in the image generation system (10, e.g., at least one processor) described above. Therefore, if the subject of a specific step / operation is omitted, it can be understood that the step / operation is performed in the image generation system (10). However, in an actual environment, some steps / operations of the methods described below may be performed in other computing devices.

[0070] For convenience of explanation, the image generation system (10) will be abbreviated as 'system' below.

[0071] FIG. 5 is an exemplary flowchart schematically illustrating a user prompt-based image generation method according to some embodiments of the present disclosure. However, this is merely an exemplary embodiment for achieving the purpose of the present disclosure, and it is understood that some steps may be added or deleted as necessary.

[0072] As illustrated in FIG. 5, the image generation method according to the embodiments may begin at step S51, which involves obtaining a user prompt containing a description of a target image. For instance, the system (10) may obtain a user prompt composed of keywords describing some elements of the scene context of the target image (e.g., type of foreground object, placement, background, style, design concept, etc.). At this time, the system (10) may further obtain a source image containing a foreground object.

[0073] The specific method for obtaining the user prompt and / or source image may vary depending on the embodiment.

[0074] In some embodiments, the system (10) may receive a prompt and / or source image from a user. For example, the system (10) may receive a user prompt consisting of several keywords and a source image regarding a foreground object from a user (e.g., received from a user terminal (41), etc.).

[0075] In some other embodiments, the system (10) may acquire a prompt created by a user and select example images from among pre-prepared example images that correspond to the scene context and / or foreground object of the prompt and provide them to the user. Then, the system (10) may use the example image selected by the user as a source image.

[0076] In some other embodiments, the system (10) may acquire a prompt written by a user and automatically select an example image from among pre-prepared example images that corresponds to the scene context and / or foreground object of the prompt as a source image.

[0077] In some other embodiments, the system (10) may provide the user with a prompt-image collection consisting of pairs of example prompts and corresponding example images. For example, the system (10) may obtain keywords related to a target image from the user (e.g., keywords regarding foreground objects, design concepts, styles, etc., or an initial user prompt containing a brief description) and select and provide the user with a prompt-image collection associated with the keywords (e.g., composed of multiple prompt-image pairs). Here, the example image may be an image generated using the example prompt, or an image that outlines the result of generating an image according to the example prompt. Then, the system (10) may use the example prompt selected by the user as a user prompt for generating the target image, or use the example prompt selected and modified by the user as a user prompt. At this time, the example image corresponding to the selected example prompt may be used as a source image as needed. For reference, if the example prompt contains sufficient detailed information regarding the scene context, step S52 described below may be omitted.

[0078] In step S52, an enhancement prompt is generated by supplementing the scene context of the user prompt through the generative language model (31). For example, the system (10) can generate a meta prompt for enhancing the prompt based on the user prompt and a pre-prepared system prompt (e.g., an instruction), and input this into the generative language model (31) to generate the enhancement prompt. However, the specific method may vary depending on the embodiment.

[0079] Hereinafter, various embodiments of a method for generating an enhanced prompt will be described with reference to FIGS. 6 to 11.

[0080] First, a method for generating an enhanced prompt according to some embodiments of the present disclosure will be described with reference to FIGS. 6 to 10b.

[0081] FIG. 6 is an exemplary flowchart illustrating a method for generating an enhancement prompt according to some embodiments of the present disclosure. However, this is merely an exemplary embodiment for achieving the purpose of the present disclosure, and it is understood that some steps may be added or deleted as necessary.

[0082] As illustrated in FIG. 6, the embodiments may begin with step S61 of preparing an enhancement prompt template. Here, the prompt template refers to a prompt of a fixed structure designed to dynamically fill in the values ​​of one or more incomplete fields, and the enhancement prompt template refers to a prompt template having one or more incomplete fields regarding a scene context. Additionally, the fields may be named as 'item', 'variable', 'attribute', 'placeholder', 'slot', 'element', etc. Depending on the case, the fields may be named as 'item', 'variable', 'attribute', 'placeholder', 'slot', 'element', etc. The enhancement prompt template may be designed in a structured format (see FIG. 7) or in a plain text format.

[0083] An example of a reinforcement prompt template is illustrated in FIG. 7. FIG. 7 illustrates an example where the reinforcement prompt template (71) is designed as a structure in the format of JSON (JavaScript Object Notation), and for ease of understanding, the field values ​​are filled in.

[0084] As illustrated in FIG. 7, the reinforcement prompt template (71) may be designed (or defined) to include a scene field (72), a lighting field (73), a camera view field (74), and / or an atmosphere field (75). However, the scope of the present disclosure is not limited thereto, and the details of the scene context included in the reinforcement prompt template (71) may vary depending on the embodiment.

[0085] The scene field (72) is a field in which various information about the main elements that constitute the scene (e.g., foreground object, background, etc.) is recorded. The scene field (72) may include, for example, the type of foreground object (see 'product', e.g., type of product, etc.), placement (see 'placement', e.g., standing state, lying state, location, information about the structure where the object is located, etc.) and / or the surrounding environment (see 'environment', e.g., surrounding structure that highlights the foreground object, description of the out-of-focus background, etc.), but is not limited thereto.

[0086] The lighting field (73) is a field in which various information about the lighting (i.e., lighting setting information) is recorded. The lighting field (73) includes detailed fields (or information) regarding, for example, primary light (see 'primary_light'), accent light (see 'accent_light', e.g., lighting that emphasizes a specific part, etc.) and / or fill light (see 'fill_light', e.g., lighting that softly fills shadows, etc.), and each detailed field may include subfields such as type, direction, intensity and / or purpose, but is not limited thereto.

[0087] The camera view field (74) is a field in which various information about the camera (i.e., camera setting information) is recorded. The camera view field (74) may include, for example, detailed fields (or information) regarding the lens (see 'lens'), aperture (see 'aperture'), focus (see 'focus'), and / or image effect (see 'effect'), but is not limited thereto.

[0088] Finally, the atmosphere field (75) is a field where information about the overall atmosphere, feeling, design concept, etc. of the scene is recorded.

[0089] Referring again to Fig. 6, the explanation will be provided.

[0090] In the above-described step S61, before inputting the reinforcement prompt template into the generative language model (31), the system (10) can supplement (or fill in) the values ​​of some incomplete fields (e.g., empty fields) of the template.

[0091] For example, the system (10) can supplement the values ​​of incomplete fields based on user prompts. For instance, the system (10) can extract information about foreground objects, backgrounds, etc. from user prompts and supplement the values ​​of related incomplete fields (e.g., product, placement fields, etc. in FIG. 7) based on this.

[0092] As another example, the system (10) can obtain lighting information and / or camera view information from the metadata of the source image's 3D model file and / or 3D modeling tool, and based on this, supplement the values ​​of the related incomplete fields (e.g., lighting, camera_settings fields in FIG. 7, etc.).

[0093] As another example, the system (10) may supplement the values ​​of incomplete fields based on various combinations of the examples described above.

[0094] In step S62, the value of an incomplete field of a reinforcement prompt template is supplemented based on user prompts through a generative language model (31). Here, the generative language model (31) may be fine-tuned using a design-related text dataset, but the scope of the present disclosure is not limited thereto. A generative language model (31) trained on such a text dataset can supplement the value of an incomplete field with higher accuracy (or completeness) by utilizing design knowledge.

[0095] For example, as illustrated in FIG. 8, the system (10) can generate a meta prompt (83) based on a reinforcement prompt template (84) and a system prompt (85, e.g., a directive requesting the creation of a reinforcement prompt, a directive requesting the supplementation of incomplete field values, etc.), and input this into a generative language model (31) to supplement the values ​​of incomplete fields of the reinforcement prompt template (84), thereby generating a reinforcement prompt (86) (i.e., the description of the prompt (82) written by the user (81) is reinforced). At this time, the user prompt (82) may be reflected in the reinforcement prompt template (84) or individually included in the meta prompt (83), and both methods may be applied simultaneously.

[0096] FIG. 9 illustrates the process of supplementing incomplete field values ​​of a reinforcement prompt template (92). For clarity of the present disclosure, 'Field A', 'Field B', and 'Field C' will be referred to as 'First Field', 'Second Field', and 'Third Field', respectively. FIG. 9 illustrates, as an example, a case where the value of the first field (93, e.g., the product field in FIG. 7) of the reinforcement prompt template (92) is supplemented (or filled) by a user prompt (91).

[0097] As illustrated in FIG. 9, the system (10) can input a reinforcement prompt template (92) into a generative language model (31) to appropriately supplement the values ​​of incomplete fields (94, 95). The generative language model (31) can understand the content of the reinforcement prompt template (92) by utilizing knowledge obtained from various text data, and based on the results, supplement the values ​​of incomplete fields (94, 95) with high accuracy (or completeness). If the generative language model (31) is fine-tuned using a design-related text dataset, the incomplete fields (94, 95) can be supplemented with even higher accuracy as design knowledge is utilized.

[0098] For reference, FIG. 9 illustrates that only the reinforcement prompt template (92) is input into the generative language model (31), but this is for convenience of understanding only, and as described above, the reinforcement prompt template (92) can be included in the meta prompt (e.g., 83) and input into the generative language model (31).

[0099] FIGS. 10a and FIGS. 10b illustrate actual examples of generating reinforcement prompts. FIGS. 10a and FIGS. 10b illustrate reinforcement prompts (102, 104) generated based on different user prompts (101, 103), respectively, and said reinforcement prompts (102, 104) were generated in the manner shown in FIG. 8.

[0100] As shown in FIGS. 10a and 10b, it can be seen that when the description of the user prompt (101, 103) is changed, the content of the reinforcement prompt (102, 104) is also changed accordingly. For example, it can be seen that the values ​​of lighting, environment, and mood fields are changed according to the description of the user prompt (101, 103), and through this, it can be seen that the generative language model (31) can effectively reinforce the user prompt.

[0101] Hereinafter, a method for generating an enhanced prompt according to several other embodiments of the present disclosure will be described with reference to FIG. 11.

[0102] FIG. 11 is an exemplary drawing for illustrating a method for generating an enhancement prompt according to some other embodiments of the present disclosure.

[0103] As illustrated in FIG. 11, the embodiments relate to a method for generating an enhancement prompt (117) when the generative language model (31) is a Vision-Language Model (VLM) that further possesses image understanding capabilities.

[0104] Specifically, the system (10) can generate a meta prompt (114) including a reinforcement prompt template (115), a system prompt (116), etc., and input this meta prompt (114) together with a source image (112) into a generative language model (31) to generate a reinforcement prompt (117). At this time, the reinforcement prompt template (115) may be a form in which some incomplete field values ​​are supplemented based on a prompt (113) written by a user (111).

[0105] In the above case, the generative language model (31) further considers the information contained in the source image (112) (e.g., placement, type, shape, design, etc. of foreground objects) to supplement the reinforcement prompt template (115), thereby generating a more sophisticated reinforcement prompt (117), and as a result, the quality of the target image can be further improved.

[0106] For reference, FIG. 11 illustrates an example in which a single source image is input into a generative language model (31), but the scope of the present disclosure is not limited thereto. For instance, the system (10) may extract information (e.g., style, design concept, background object, etc.) of other elements of the scene context (i.e., elements other than the foreground object) from a user prompt (113) and further input a reference image associated with the extracted elements into the generative language model (31).

[0107] Meanwhile, in some other embodiments of the present disclosure, the system (10) may progressively improve the quality of a reinforcement prompt based on self-feedback of the generative language model (31). In this case, the self-feedback may include, for example, a self-evaluation score for the generated reinforcement prompt, parts requiring improvement and / or reasons therefor (e.g., reasons why the evaluation score was calculated as such, reasons requiring improvement, etc.), but is not limited thereto. Specifically, the system (10) may generate a first reinforcement prompt through the generative language model (31) and derive self-feedback by evaluating the quality of the first reinforcement prompt. Then, the system (10) may generate a second reinforcement prompt by updating the first reinforcement prompt through the generative language model (31) to reflect the self-feedback. This process may be repeated until a preset condition (e.g., the self-evaluation score exceeding a threshold) is satisfied.

[0108] The technical concept of the preceding embodiments may also be applied to the fine-tuning (or training) process of the generative language model (31). For example, the system (10) may tune (or update / adjust) the parameters of the generative language model (31) through a process of updating input prompt samples so that the self-feedback of the generative language model (31) leads to a more positive direction (e.g., a direction in which the self-evaluation score increases). Any specific tuning method may be used.

[0109] In some other embodiments, the system (10) can generate an enhancement prompt that matches the design concept desired by the user through a generative language model (31). For example, let us assume that there is text regarding the design concept of a target image in the user prompt. In this case, the system (10) can supplement the scene context of the user prompt through the generative language model (31) so that the design concept is reflected (e.g., supplementing incomplete field values ​​of the enhancement prompt template to match the design concept). At this time, the design concept may be reflected in the enhancement prompt template and input into the generative language model (31), or it may be individually included in other parts of the meta prompt. Additionally, the generative language model (31) may be fine-tuned with a design-related text dataset, but the scope of the present disclosure is not limited thereto. According to these embodiments, since a concrete prompt is automatically generated when the user inputs only keywords and design concepts regarding the foreground object, the difficulty of creating the prompt is significantly reduced, and the usability of the generative image model (11) in design-related fields can be greatly increased.

[0110] In some other embodiments, the system (10) may automatically update the reinforcement prompt to reflect the user's modifications. Specifically, let us assume that the system (10) has generated an initial reinforcement prompt through a generative language model (31) and provided it to the user, and has obtained a modification prompt for it from the user. In this case, the system (10) may update the modification prompt through the generative language model (31) to generate a reinforcement prompt to be used for generating a target image. This update process can be understood as a refining process (i.e., a process of generating a user-customized prompt) so that the user's modifications are reflected throughout the scene context within the prompt.

[0111] Various embodiments of a method for generating an enhancement prompt have been described so far with reference to FIGS. 6 to 11.

[0112] Referring again to Fig. 5, the explanation will be provided.

[0113] In step S53, a target image is generated from an enhancement prompt through a generative image model (11). For example, the system (10) can generate a target image by conditioning (or inputting) an enhancement prompt to the generative image model (11).

[0114] The specific method for generating the target image in step S53 may vary depending on the embodiment.

[0115] In some embodiments, the system (10) may generate a target image by using a reinforcement prompt as guidance information for a generative image model (11). For example, the system (10) may generate prompt guidance information (e.g., prompt embedding vector) by encoding the reinforcement prompt through a text encoder and condition it to the generative image model (11), for which further reference is made to the description in FIG. 13, etc.

[0116] In some other embodiments, the system (10) may extract a mask (e.g., an inverted mask in binary format or a regular mask) that distinguishes foreground objects and backgrounds from a source image and further condition this onto a generative image model (11) to generate a target image (i.e., the mask is further conditioned in addition to prompt guidance information). By doing so, the placement of foreground objects, backgrounds, etc., can be generated while maintaining consistency, and further reference should be made to the description in FIG. 13, etc. For reference, the mask may or may not be conditioned in an encoded state. Also, the mask may be understood as a type of guidance information.

[0117] In some other embodiments, the system (10) may estimate a depth map of a source image and use it further as guidance information for a generative image model (11) to generate a target image (i.e., the depth map is further conditioned in addition to the prompt guidance information). By doing so, the placement of foreground objects, backgrounds, etc., can be generated while maintaining consistency, and further reference should be made to the description in FIG. 13, etc.

[0118] In some other embodiments, the system (10) may generate a target image by further conditioning a mask and a depth map to the generative image model (11) in addition to the reinforcement prompt. By doing so, a sophisticated target image that fits well with the reinforcement prompt can be consistently generated, and these embodiments will be described in detail below with reference to FIGS. 12 to 16.

[0119] FIG. 12 is an exemplary flowchart illustrating a method for generating a target image according to some embodiments of the present disclosure. However, this is merely an exemplary embodiment for achieving the purpose of the present disclosure, and it is understood that some steps may be added or deleted as necessary.

[0120] As illustrated in FIG. 12, the embodiments may begin with step S121, which estimates a depth map of a source image containing a foreground object. For instance, as illustrated in FIG. 13, the system (10) may estimate a depth map (135) of a source image (133) through a depth estimation model (132). The depth map (135) thus generated may contain noise (e.g., depth values ​​of background objects not present in the source image (133)) due to scene context bias of the depth estimation model (132) formed during the training process (e.g., the model learns specific object placements or object-background relationships that appear repeatedly in the training data). This noise may serve as a hint for generating the background of the target image (139) (e.g., when noise is generated in an appropriate location where a background object mentioned in the reinforcement prompt (137) can be placed), or conversely, it may cause a background generation error.

[0121] For reference, in FIG. 13, the text encoder (131), depth estimation model (132), generative image model (11), etc. may be trained models, and in some cases, additional fine-tuning may be performed during the process of generating a target image (e.g., 139). However, the scope of the present disclosure is not limited thereto.

[0122] In some embodiments, the noise level of the depth map may be adjusted to prevent background generation errors in the target image while simultaneously increasing background generation accuracy. These embodiments will be further described below with reference to FIGS. 14 to 15c.

[0123] FIG. 14 is an exemplary drawing for illustrating a depth map noise adjustment method according to some embodiments of the present disclosure.

[0124] As illustrated in FIG. 14, the system (10) can adjust the noise level of the background area of ​​the estimated depth map (141) (i.e., the area corresponding to the background of the source image and located outside the foreground object (144)). FIG. 14 illustrates the process of gradually adjusting the noise level of the depth map (141) (see 142, 143).

[0125] Specifically, the system (10) can adjust the noise level of the background area of ​​the depth map (141) so that some noise remains. For example, the system (10) can remove the noise in the background area until the noise level reaches a threshold (e.g., about 20%). In this case, the boundaries of the foreground object become more distinct, and the possibility of background generation errors can be significantly reduced. Here, it can be understood that the reason for leaving some noise is that the noise can act as a hint during the background generation process.

[0126] Hereinafter, with reference to FIGS. 15a to 15c, the effects according to the noise adjustment method described above will be further explained.

[0127] FIG. 15a illustrates an image (151) generated based on a depth map (141) before noise level adjustment, and FIG. 15b and FIG. 15c illustrate images (155, 156) generated based on each of the adjusted depth maps (142, 143).

[0128] As illustrated in FIG. 15a, noise in the lower region (154) of the foreground object (144) acts as a hint for the intended background object (see 'table' in the right image (151)) (assuming the table is mentioned in the reinforcement prompt), whereas noise in the right region (153) can cause an error in generating unintended background objects (152, i.e., background objects not mentioned in the reinforcement prompt).

[0129] In contrast, in the depth maps (142, 143) of FIG. 15b and FIG. 15c, as the noise level is partially reduced, no background generation error occurs, and it can be seen that the noise remaining in the area below the foreground object (144) (e.g., 154) acts only as a hint for the intended background object.

[0130] Up to now, a depth map noise adjustment method according to some embodiments of the present disclosure has been described with reference to FIGS. 13 to 15c.

[0131] Referring again to Fig. 12, the explanation will be provided.

[0132] In step S122, a mask is generated to distinguish between the foreground object and the background in the source image. For example, referring again to FIG. 13, the system (10) can generate an inversion mask (134) to distinguish between the foreground object and the background in the source image (133).

[0133] In step S123, a depth map and a mask are conditioned on the generative image model (11), and a target image is generated based on an enhancement prompt. For example, the system (10) can generate a target image using the depth map, mask, and enhancement prompt as guidance information for the generative image model (11). In this case, a sophisticated target image corresponding to the enhancement prompt can be consistently generated (e.g., consistency in placement of foreground objects, background, etc., and consistency in image generation according to the description of the enhancement prompt can be maintained at a high level).

[0134] Referring again to FIG. 13, the system (10) can generate prompt guidance information (138, e.g., prompt embedding vector) by encoding a reinforcement prompt (137) through a text encoder (131), and generate depth guidance information (136) based on an estimated depth map (135). The depth guidance information (136) may also be generated by encoding the depth map (135) in an appropriate manner, but the scope of the present disclosure is not limited thereto.

[0135] Next, the system (10) can generate a target image (139) by conditioning (or inputting) guidance information (136, 138) and an inversion mask (134) to a generative image model (11). The inversion mask (134) may be encoded in an appropriate manner and conditioned to the generative image model (11), or it may be conditioned as is.

[0136] Below, for the sake of greater understanding, a method for generating a target image (166) when the generative image model (11) is a diffusion-based model will be explained with reference to FIG. 16.

[0137] As illustrated in FIG. 16, the system (10) can prepare a noise image (162) and generate depth guidance information (164) and prompt guidance information (165) in the manner described above.

[0138] Next, the system (10) can generate a target image (166) by conditioning guidance information (164, 165) and an inversion mask (161) to a generative image model (11) and performing a denoising process on a noise image (162). For example, the system (10) can input guidance information (164, 165) and / or an inversion mask (161) into a noise prediction neural network (or score prediction neural network) of the generative image model (11) to predict the noise (or score) of a specific time step and perform a denoising process on the noise image (162) based on the predicted noise. The system (10) can generate a target image (166) by repeatedly performing this process. As those skilled in the art are likely already familiar with the operating principles of such denoising processes, a detailed explanation thereof will be omitted.

[0139] Meanwhile, the system (10) may perform various post-processing steps on the output image or target image of the generative image model (11).

[0140] For example, the system (10) can generate a target image by compositing the foreground object of the source image onto the output image of the generative image model (11). In this case, a more sophisticated and higher quality target image can be generated.

[0141] As another example, the system (10) can perform post-processing to transform the style of the target image. For instance, the system (10) can perform this post-processing in conjunction with a deep learning model equipped with style transfer capabilities.

[0142] As another example, the system (10) can perform post-processing to relight the background of the target image (e.g., changing shadows and / or lighting). For instance, the system (10) can perform this post-processing in conjunction with a deep learning model equipped with background relighting capabilities.

[0143] As another example, the system (10) may perform post-processing based on various combinations of the examples described above.

[0144] The aforementioned post-processing steps may be performed according to a user's request (or command) or based on the system's own judgment. For example, the system (10) may perform style conversion on a target image according to a user's request, or it may perform style conversion based on the judgment that the style of the output image of the generative image model does not match the style in the reinforcement prompt. Alternatively, the system (10) may composite the foreground object of the source image with the output image based on the judgment that the difference between the foreground object of the output image and the foreground object of the source image is greater than a threshold.

[0145] Additionally, the system (10) may regenerate the target image if a generation error is detected.

[0146] For example, the system (10) can generate a caption for a target image through an image captioning model (e.g., a vision-language model, etc.) and compare it with a reinforcement prompt. Specifically, the system (10) can detect generation errors by comparing the caption with the reinforcement prompt using a generative language model (31) (or another language model). Then, the system (10) can update the existing reinforcement prompt through the generative language model (31) in a way that prevents generation errors (e.g., reinforcement of descriptions of ungenerated background objects, or addition of descriptions to prevent the generation of unintended background objects). Then, the system (10) can regenerate the target image from the updated reinforcement prompt through a generative image model (11). In some cases, the system (10) may further remove residual noise from the depth map to prevent the generation of unintended background objects, regenerate depth guidance information using this, and then regenerate the target image using the depth guidance information.

[0147] As another example, the system (10) can detect generation errors by comparing a target image and a reinforcement prompt through a vision-language model, and update the existing reinforcement prompt in a way that prevents generation errors through a generative language model (31). Then, the system (10) can regenerate the target image from the updated reinforcement prompt through a generative image model (11).

[0148] As another example, the system (10) may regenerate the target image based on various combinations of the examples described above.

[0149] Up to this point, a method for generating user prompt-based images according to several embodiments of the present disclosure has been described with reference to FIGS. 5 to 16. According to the method described above, a high-quality enhanced prompt can be generated by sufficiently supplementing the elements of the scene context that are lacking in the user prompt through a generative language model (31). In this case, the difficulty of creating the user prompt is lowered, the hassle can be reduced, and the accuracy and consistency of target image generation can be improved as the enhanced prompt is provided to the generative image model (11).

[0150] Additionally, user prompts can be enhanced through a generative language model (31) fine-tuned using a design-related text dataset. In this case, high-quality enhanced prompts can be effectively generated.

[0151] Additionally, an enhancement prompt template containing incomplete fields regarding various elements of the scene context can be prepared, and the enhancement prompt can be generated by supplementing the values ​​of these incomplete fields through a generative language model (31). In this case, even if the user enters only a few keywords related to the target image, a professional, high-quality prompt (i.e., an enhancement prompt) can be generated, and as a result, the accuracy and consistency of the target image generation can be further improved.

[0152] In addition, an enhancement prompt can be generated by further considering the information contained in the source image (e.g., placement, type, shape, design, etc. of foreground objects) through the vision-image model (31). In this case, a high-quality enhancement prompt can be generated more effectively.

[0153] In addition, by utilizing the mask and / or depth map obtained from the source image as guidance information for the generative image model (11), a sophisticated target image corresponding to the reinforcement prompt can be consistently generated (e.g., consistency in placement of foreground objects, background, etc., and consistency in image generation according to the description of the reinforcement prompt can be maintained at a high level).

[0154] Additionally, the noise level can be adjusted so that some noise is present in the background area of ​​the depth map estimated from the source image. In this case, background generation errors that may occur due to excessive noise (e.g., the generation of unintended background objects) are prevented, while the residual noise acts as a hint during the background generation process, thereby further improving the accuracy and consistency of target image generation.

[0155] Hereinafter, with reference to FIG. 17, an exemplary computing device capable of implementing the system (10) described above will be described.

[0156] FIG. 17 is an exemplary hardware configuration diagram showing a computing device (170).

[0157] As illustrated in FIG. 17, a computing device (170) may include one or more processors (171), a bus (173), a communication interface (174), a memory (172) for loading a computer program (176) executed by the processor (171), and a storage (175) for storing the computer program (176). However, FIG. 17 illustrates only the components related to the embodiments of the present disclosure. Therefore, a person skilled in the art to which the present disclosure belongs will understand that other general-purpose components (e.g., input devices such as a keyboard and mouse, output devices such as a speaker and display) may be included in addition to the components (171 to 176) illustrated in FIG. 17. That is, the computing device (170) may include various additional components in addition to the components (171 to 176) illustrated in FIG. 17. Additionally, depending on the case, the computing device (170) may be configured with some of the components (171 to 176) shown in FIG. 17 omitted. Each component of the computing device (170) will be described below.

[0158] The processor (171) can control the overall operation of each component of the computing device (170). The processor (171) may be configured to include at least one of a CPU (Central Processing Unit), MPU (Micro Processor Unit), MCU (Micro Controller Unit), GPU (Graphic Processing Unit), NPU (Neural Processing Unit), TPU (Tensor Processing Unit), VPU (Vision Processing Unit), APU (Accelerated Processing Unit), or any other type of processor well known in the art of the present disclosure. Additionally, the processor (171) may perform operations for at least one application or program for executing specific operations / steps / methods. The computing device (170) may have one or more processors.

[0159] Next, the memory (172) may store various data, commands and / or information. The memory (172) may load a computer program (176) from storage (175) to execute specific operations / steps / methods. The memory (172) may be implemented as volatile memory such as RAM, but the technical scope of the present disclosure is not limited thereto.

[0160] Next, the bus (173) can provide communication functions between components of the computing device (170). The bus (173) can be implemented as various types of buses, such as an address bus, a data bus, and a control bus.

[0161] Next, the communication interface (174) may support wired and wireless internet communication of the computing device (170). Additionally, the communication interface (174) may support various communication methods other than internet communication. To this end, the communication interface (174) may be configured to include a communication module well known in the art of the present disclosure.

[0162] Next, the storage (175) may store one or more computer programs (176) non-temporarily. The storage (175) may be configured to include non-volatile memory such as ROM (Read Only Memory), EPROM (Erasable Programmable ROM), EEPROM (Electrically Erasable Programmable ROM), flash memory, a hard disk, a removable disk, or any form of computer-readable recording medium well known in the art to which this disclosure belongs.

[0163] Next, the computer program (176) may include instructions that cause the processor (171) to perform specific operations / steps / methods when loaded into memory (172). That is, the processor (171) can perform specific operations / steps / methods by executing the loaded instructions.

[0164] For example, a computer program (176) may include instructions to perform the action of obtaining a user prompt containing a description of a target image, the action of generating an enhancement prompt by supplementing the scene context of the user prompt through a generative language model (31), and the action of generating a target image from the enhancement prompt through a generative image model (11).

[0165] As another example, a computer program (176) may include instructions to perform at least some of the methods / steps / actions described with reference to FIGS. 1 through 16.

[0166] As illustrated, a system (10) according to some embodiments of the present disclosure can be implemented through a computing device (170).

[0167] Meanwhile, in some embodiments, the computing device (170) illustrated in FIG. 17 may refer to a virtual machine implemented based on cloud technology. For example, the computing device (170) may be a virtual machine running on one or more physical servers included in a server farm. In this case, at least some of the processor (171), memory (172), and storage (175) illustrated in FIG. 17 may be virtual hardware, and the communication interface (174) may also be implemented as a virtualized networking element such as a virtual switch.

[0168] Up to now, with reference to FIG. 17, an exemplary computing device (170) capable of implementing a system (10) according to some embodiments of the present disclosure has been described.

[0169] Various embodiments of the present disclosure and effects according to those embodiments have been described with reference to FIGS. 1 to 17. The effects according to the technical concept of the present disclosure are not limited to those described above, and other unmentioned effects will be clearly understood by a person skilled in the art from the description below.

[0170] Furthermore, just because the above embodiments describe a plurality of components being combined into one or operating in combination, the technical concept of the present disclosure is not necessarily limited to these embodiments. That is, within the scope of the purpose of the technical concept of the present disclosure, all such components may be selectively combined into one or more combinations to operate.

[0171] The technical concept of the present disclosure described above may be implemented as computer-readable code on a computer-readable recording medium. A computer program recorded on a computer-readable recording medium may be transmitted to another computing device via a network such as the Internet and installed on said computing device, thereby being used on said computing device.

[0172] Although operations are depicted in a specific order in the drawings, it should not be understood that the operations must necessarily be executed in the specific order depicted or in a sequential order, or that all depicted operations must be executed to obtain the desired result. In certain situations, multitasking and parallel processing may be advantageous. Although various embodiments of the present disclosure have been described above with reference to the attached drawings, those skilled in the art will understand that the technical concept of the present disclosure may be implemented in other specific forms without altering the technical concept or essential features thereof. Therefore, the embodiments described above should be understood as illustrative in all respects and not restrictive. The scope of protection of the present disclosure shall be interpreted by the claims below, and all technical concepts within the equivalent scope shall be interpreted as being included within the scope of rights of the technical concept defined by the present disclosure.

Claims

1. A method performed by at least one processor, A step of obtaining a user prompt containing a description of a target image; A step of generating an enhanced prompt by supplementing the scene context of the user prompt through a generative language model; and A step comprising generating the target image from the enhancement prompt through a generative image model, User prompt-based image generation method.

2. In Paragraph 1, The above scene context includes placement information of foreground objects and background information, User prompt-based image generation method.

3. In Paragraph 1, The above scene context includes at least one of camera view information and lighting information, User prompt-based image generation method.

4. In Paragraph 1, The above generative language model is fine-tuned using a design-related text dataset, User prompt-based image generation method.

5. In Paragraph 1, The step of generating the above reinforcement prompt is, Step of preparing an enhancement prompt template - the enhancement prompt template includes one or more incomplete fields related to the scene context -; and A step comprising supplementing the value of one or more incomplete fields based on the user prompt through the generative language model, User prompt-based image generation method.

6. In Paragraph 5, The above generative language model is a vision-language model further equipped with image understanding capabilities, and The step of supplementing the values ​​of one or more of the above-mentioned incomplete fields is, A step of obtaining a source image associated with the above user prompt; and A step comprising supplementing the value of one or more incomplete fields based further on the source image above, User prompt-based image generation method.

7. In Paragraph 1, The step of generating the above reinforcement prompt is, A step of generating a first reinforcement prompt through the above generative language model; A step of deriving self-feedback for the first reinforcement prompt through the generative language model; and A step comprising generating a second reinforcement prompt by updating the first reinforcement prompt by reflecting the self-feedback through the generative language model, User prompt-based image generation method.

8. In Paragraph 1, The above description includes text regarding the design concept of the target image, and The step of generating the above reinforcement prompt is, A step comprising generating the reinforcement prompt by supplementing the scene context to match the design concept through the generative language model, User prompt-based image generation method.

9. In Paragraph 1, The step of generating the above reinforcement prompt is, A step of generating an initial reinforcement prompt through the above generative language model and providing it to the user; A step of obtaining the user's modification prompt for the initial reinforcement prompt; and A step comprising updating the modification prompt through the generative language model to generate the reinforcement prompt, User prompt-based image generation method.

10. In Paragraph 1, The above scene context includes at least one of camera view information and lighting information related to the foreground object, and The step of generating the above reinforcement prompt is, A step of obtaining at least one piece of information from a 3D model file or 3D modeling tool of the foreground object; A step of supplementing the values ​​of at least some of the incomplete fields of the reinforcement prompt template by reflecting the information obtained above; and A step comprising generating the reinforcement prompt by further supplementing the values ​​of the incomplete fields through the generative language model. User prompt-based image generation method.

11. In Paragraph 1, The step of generating the above target image is, A step of obtaining a source image associated with the above user prompt; A step of estimating the depth map of the above source image; A step of generating depth guidance information based on the depth map above; and A method comprising the step of generating the target image by conditioning the depth guidance information to the generative image model. User prompt-based image generation method.

12. In Paragraph 11, The step of generating the above depth guidance information is, A step of adjusting the noise level of an area corresponding to the background of the source image in the depth map so that some noise remains; and A method comprising the step of generating depth guidance information based on the adjusted depth map. User prompt-based image generation method.

13. In Paragraph 1, The step of generating the above target image is, A step of obtaining a source image associated with the above user prompt - the source image includes a foreground object -; A step of generating a mask that distinguishes the foreground object and the background in the source image; and A method comprising the step of generating the target image by conditioning the mask to the generative image model. User prompt-based image generation method.

14. In Paragraph 1, The step of generating the above target image is, A step of obtaining a source image associated with the above user prompt - the source image includes a foreground object -; A step of obtaining an output image of the above-mentioned generative image model; and A method comprising the step of generating the target image by compositing the foreground object onto the output image. User prompt-based image generation method.

15. In Paragraph 1, The above generative image model is a diffusion-based model, and The step of generating the above target image is, A step of generating prompt guidance information by encoding the above-mentioned reinforcement prompt; Step of preparing a noise image; and A method comprising the step of conditioning the prompt guidance information to the generative image model and performing a denoising process on the noise image to generate the target image. User prompt-based image generation method.

16. One or more processors; and It includes memory for storing computer programs executed by one or more of the above processors, and The above computer program is: An action of obtaining a user prompt containing a description of the target image; An operation to generate an enhanced prompt by supplementing the scene context of the user prompt through a generative language model; and Instructions for an operation to generate the target image from the enhancement prompt through a generative image model, Image generation system.

17. In Paragraph 16, The above scene context is, Includes placement information of foreground objects and background information, including at least one of camera view information and lighting information, Image generation system.

18. In Paragraph 16, The operation of generating the above reinforcement prompt is, Action of preparing a reinforcement prompt template - the reinforcement prompt template includes one or more incomplete fields related to the scene context -; and A method comprising supplementing the value of one or more incomplete fields based on the user prompt through the generative language model. Image generation system.

19. A non-transitory computer-readable recording medium storing instructions that, when executed by at least one processor, cause said at least one processor to perform a user prompt-based image generation method, The above image generation method is: A step of obtaining a user prompt containing a description of a target image; A step of generating an enhanced prompt by supplementing the scene context of the user prompt through a generative language model; and A step comprising generating the target image from the enhancement prompt through a generative image model, Non-transient computer-readable recording medium.