Method for correcting scale and composition during artificial intelligence-based space image generation, and electronic device for performing same

By adding human figures and iteratively refining AI-generated spatial images using multiple AI models, the method corrects scale and composition errors, ensuring accurate and efficient furniture arrangement in large spaces.

WO2026111323A1PCT designated stage Publication Date: 2026-05-28PLANBY TECHNOLOGIES INC
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-11-14
Publication Date
2026-05-28

AI Technical Summary

Technical Problem

Current AI-based image generation technologies for interior design often fail to accurately reflect the physical characteristics of large spaces, leading to visual errors in furniture arrangement and composition that can result in inefficient use of space and incorrect atmosphere.

Method used

A method involving the use of AI models to add human figures to spatial images, process them through preprocessing and guide information, and iteratively refine the scale and composition using multiple AI models to correct visual errors.

Benefits of technology

Minimizes visual errors in AI-generated spatial images by accurately adjusting scale and composition, ensuring furniture is arranged efficiently and harmoniously within the space.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure KR2025018836_28052026_PF_FP_ABST
    Figure KR2025018836_28052026_PF_FP_ABST
Patent Text Reader

Abstract

The present disclosure relates to a method for correcting scale and composition during artificial intelligence-based space image generation, and an electronic device for performing same. According to an embodiment of the present disclosure, the method may comprise the steps of: acquiring a first space image and an input prompt; acquiring a first contour image and a first depth image on the basis of the first space image; acquiring first guidance information by using a first artificial intelligence model taking, as inputs, the first contour image and the first depth image; acquiring a second space image by adding a plurality of human-shaped figures to the first space image; acquiring a second contour image and a second depth image on the basis of the second space image; acquiring second guidance information by using the first artificial intelligence model taking, as inputs, the second contour image and the second depth image; and acquiring a final space image in which at least one object is added to the first space image, by using a second artificial intelligence model taking, as inputs, the first space image, the input prompt, the first guidance information, and the second guidance information.
Need to check novelty before this filing date? Find Prior Art

Description

Method for correcting scale and composition when generating AI-based spatial images and electronic device for performing the same

[0001] The present disclosure relates to a method for correcting scale and composition when generating an artificial intelligence-based spatial image and an electronic device for performing the same. More specifically, it relates to a method for generating a spatial image desired by a user by correcting the scale and composition of the space using an image in which a human figure has been added.

[0002] With the advancement of artificial intelligence and computer vision technologies, active research is being conducted on image generation techniques utilizing AI. In particular, there is a growing number of attempts to apply these technologies to the interior design field. AI-powered image generation technology offers the advantage of proposing various interior designs tailored to the style or theme desired by the user. Through this, users can more easily visualize their desired spaces and compare various design options.

[0003] However, various forms of visual errors can occur in the current generation of images for interior design or spatial composition. For example, when arranging furniture in large spaces based on images such as photographs, models, or sketches, the scale or composition frequently does not match naturally. This is because images generated by AI fail to perfectly reflect the physical characteristics of the actual space. Such visual errors can confuse users when selecting a final design and may cause unexpected problems during actual spatial configuration.

[0004] This problem is particularly pronounced in large spaces such as offices and lobbies. In such large areas, the arrangement and scale of furniture significantly impact the overall atmosphere and functionality, meaning even minor visual errors can lead to major issues. For instance, if the size or placement of furniture does not match the actual space, it may appear cramped or be used inefficiently. This highlights the need for further advancements in AI-powered image generation technology within the field of interior design.

[0005] The objective of the present disclosure is to provide a method for correcting scale and composition when generating an AI-based spatial image, which generates a spatial image desired by a user by correcting the scale and composition of the space using an image in which a human figure is added to the spatial image, and an electronic device for performing the same.

[0006] In one embodiment of the present disclosure, a method for correcting scale and composition when generating an AI-based spatial image may be provided. The method may include the steps of: acquiring a first spatial image and an input prompt; acquiring a first contour image and a first depth image based on the first spatial image; acquiring first guide information using a first AI model that takes the first contour image and the first depth image as inputs; acquiring a second spatial image by adding a plurality of human figures to the first spatial image; acquiring a second contour image and a second depth image based on the second spatial image; acquiring second guide information using the first AI model that takes the second contour image and the second depth image as inputs; and acquiring a final spatial image in which at least one object is added to the first spatial image using a second AI model that takes the first spatial image, the input prompt, the first guide information, and the second guide information as inputs.

[0007] In one embodiment of the present disclosure, the method may include the step of acquiring the second spatial image, the step of adding a first human figure located at a first depth on the spatial image, and the step of adding a second human figure located at a second depth corresponding to a distance greater than the first depth on the spatial image.

[0008] In one embodiment of the present disclosure, the size of the first person shape may be larger than the size of the second person shape.

[0009] In one embodiment of the present disclosure, the first artificial intelligence model may be a model pre-trained to output a feature map corresponding to an input image, and the second artificial intelligence model may be a model pre-trained to output final noise by removing input noise a predetermined number of times.

[0010] In one embodiment of the present disclosure, the step of acquiring the final spatial image may include: a step of repeatedly removing the noise using the second artificial intelligence model that takes the first spatial image, the input prompt, and the second guide information as inputs for a first number of times among the previously defined number of times; a step of repeatedly removing the noise using the second artificial intelligence model that takes the first spatial image, the input prompt, and the first guide information as inputs for a second number of times among the previously defined number of times; and a step of acquiring the final spatial image based on the final noise from which the noise has been repeatedly removed.

[0011] In one embodiment of the present disclosure, the first number may be less than the second number.

[0012] In one embodiment of the present disclosure, the first number corresponds to 30% of the previously defined number, and the second number may correspond to 70% of the previously defined number.

[0013] In one embodiment of the present disclosure, the first weight for the first guide information and the second weight for the second guide information in the second artificial intelligence model may be different from each other.

[0014] In one embodiment of the present disclosure, the plurality of human figures may be two to three.

[0015] In one embodiment of the present disclosure, an electronic device may be provided. The electronic device may include at least one processor comprising a processing circuit and at least one memory storing at least one instruction. By the at least one processor executing the at least one instruction, the electronic device may acquire a first spatial image and an input prompt, acquire a first contour image and a first depth image based on the first spatial image, acquire first guide information using a first artificial intelligence model that takes the first contour image and the first depth image as inputs, acquire a second spatial image by adding a plurality of human figures to the first spatial image, acquire a second contour image and a second depth image based on the second spatial image, acquire second guide information using the first artificial intelligence model that takes the second contour image and the second depth image as inputs, and acquire a final spatial image in which at least one object is added to the first spatial image using a second artificial intelligence model that takes the first spatial image, the input prompt, the first guide information, and the second guide information as inputs.

[0016] According to one embodiment of the present disclosure, the scale and composition of an output image can be corrected by adding a plurality of human figures to a spatial image and providing them to an artificial intelligence model as a guide for scale and composition. According to one embodiment of the present disclosure, visual errors in the output image can be minimized by utilizing guide information for the spatial image and guide information for the image with added human figures together.

[0017] FIG. 1 is a block diagram showing an electronic device according to one embodiment of the present disclosure.

[0018] FIG. 2 is a block diagram showing an image preprocessing module according to one embodiment of the present disclosure.

[0019] FIG. 3 is a block diagram showing a human shape addition module and an image preprocessing module according to one embodiment of the present disclosure.

[0020] FIG. 4 is a block diagram showing a first artificial intelligence model according to one embodiment of the present disclosure.

[0021] FIG. 5 is a block diagram showing a first artificial intelligence model according to one embodiment of the present disclosure.

[0022] FIG. 6 is a block diagram showing a second artificial intelligence model according to one embodiment of the present disclosure.

[0023] FIG. 7 is a conceptual diagram showing a noise removal step of a second artificial intelligence model according to one embodiment of the present disclosure.

[0024] FIG. 8 is a conceptual diagram showing an electronic device according to one embodiment of the present disclosure.

[0025] FIG. 9 is a flowchart showing a method for correcting scale and composition when generating an artificial intelligence-based spatial image according to one embodiment of the present disclosure.

[0026] FIG. 10 is a diagram illustrating the operation of a generative model including a first artificial intelligence model and a second artificial intelligence model according to one embodiment of the present disclosure.

[0027] Hereinafter, embodiments of the present disclosure will be described in detail with reference to the accompanying drawings. The embodiments will be described clearly and in detail so that a person skilled in the art can easily practice the present disclosure. However, the scope of the rights is not limited or restricted by these embodiments. Identical or similar reference numerals are used for similar components in each drawing, and redundant descriptions of identical or similar components are omitted.

[0028] The terms used in the following description have been selected as common and universal in the relevant technical field, but other terms may exist depending on technological development and / or changes, conventions, preferences of the skilled technician, etc. Therefore, the terms used in the following description should not be understood as limiting the technical concept, but as illustrative terms to explain the embodiments.

[0029] In addition, there are terms arbitrarily selected by the applicant in specific cases, and their detailed meanings will be described in the relevant explanatory section. Therefore, the terms used in the description below must be understood not merely as their names, but based on their meanings and the content throughout the specification.

[0030] Singular expressions may include plural expressions unless the context clearly indicates otherwise. Terms used herein, including technical or scientific terms, may have the same meaning as generally understood by those skilled in the art as described in this specification. Additionally, terms including ordinal numbers, such as "first" or "second," used in this specification may be used to describe various components, but said components should not be limited by said terms. Such terms are used solely for the purpose of distinguishing one component from another.

[0031] When a part of a specification is described as "comprising" a certain component, this means that, unless specifically stated otherwise, it does not exclude other components but may include additional components. Furthermore, terms such as "part" or "module" as used in the specification refer to a unit that processes at least one function or operation, and this may be implemented in hardware or software, or as a combination of hardware and software.

[0032] Embodiments of the present disclosure are described below with reference to the attached drawings so that those skilled in the art can easily implement them. However, the present disclosure may be embodied in various different forms and is not limited to the embodiments described herein. Furthermore, in order to clearly explain the present disclosure in the drawings, parts unrelated to the explanation have been omitted, and similar parts throughout the specification have been given similar reference numerals. Also, the reference numerals used in each drawing are for the purpose of explaining each drawing, and different reference numerals used in different drawings are not intended to indicate different elements. The present disclosure will be described in detail below with reference to the attached drawings.

[0033] In the present disclosure, an artificial intelligence model may be composed of a plurality of neural network layers. Each of the plurality of neural network layers has a plurality of weight values ​​and performs neural network operations through operations between the results of operations of a previous layer and the plurality of weights. The plurality of weights possessed by the plurality of neural network layers may be optimized by the learning results of the deep neural network model. For example, the plurality of weights may be updated so that the loss value or cost value obtained from the deep neural network model during the learning process is reduced or minimized. For example, the deep neural network model may include, but is not limited to, a CNN (Convolutional Neural Network), DNN (Deep Neural Network), RNN (Recurrent Neural Network), RBM (Restricted Boltzmann Machine), DBN (Deep Belief Network), BRDNN (Bidirectional Recurrent Deep Neural Network), or Deep Q-Networks.

[0034] In the present disclosure, functions related to 'Artificial Intelligence' are operated through a processor and memory. The processor may be composed of one or more processors. In this case, the one or more processors may be general-purpose processors such as CPUs, APs, and DSPs (Digital Signal Processors), graphics-dedicated processors such as GPUs and VPUs (Vision Processing Units), or AI-dedicated processors such as NPUs. The one or more processors control the processing of input data according to predefined operation rules or AI models stored in memory. Alternatively, if the one or more processors are AI-dedicated processors, the AI-dedicated processors may be designed with a hardware structure specialized for processing a specific AI model.

[0035] The predefined rules of operation or artificial intelligence models are characterized by being created through learning. Here, being created through learning means that a basic artificial intelligence model is trained using a number of training data by a learning algorithm, thereby creating predefined rules of operation or artificial intelligence models configured to perform desired characteristics (or objectives). Such learning may be performed on the device itself where the artificial intelligence according to the present disclosure is executed, or it may be performed through a separate server and / or system. Examples of learning algorithms include supervised learning, unsupervised learning, semi-supervised learning, or reinforcement learning, but are not limited to the examples described above.

[0036] 'Scale' may refer to the size or ratio of objects within an image. In one embodiment of the present disclosure, scale may refer to the size or ratio of objects within a final spatial image output by an artificial intelligence model.

[0037] "Composition" may refer to the way objects are arranged in an image. In one embodiment of the present disclosure, composition may refer to the way objects are arranged within a final spatial image output by an artificial intelligence model.

[0038] "Input prompt" may refer to input text that instructs an artificial intelligence model to perform a specific task. However, the present disclosure is not limited thereto, and the input prompt may include data of any format, such as text, audio, image, or video. In one embodiment of the present disclosure, the input prompt is text provided to the artificial intelligence model and may serve to provide information necessary for generating a spatial image.

[0039] 'Outline image' can refer to an image reconstructed by emphasizing the outlines of objects within the image.

[0040] A 'depth image' can refer to an image that includes depth information for each pixel within the image.

[0041] A 'feature map' is a map that emphasizes the characteristics of an image or data, and can refer to data output by an encoder.

[0042] 'Noise' can refer to the data that serves as input to the diffusion model. The initial noise input to the diffusion model can take the form of Gaussian noise.

[0043] FIG. 1 is a block diagram showing an electronic device according to one embodiment of the present disclosure.

[0044] Referring to FIG. 1, a first spatial image (10) can be processed by an electronic device (100) and used to generate a final spatial image (20). The first spatial image (10) can be used with an input prompt (15) that can reflect the characteristics of the spatial image desired by the user. The first spatial image (10) may include various forms of spatial images. For example, the first spatial image (10) may be an indoor image, an outdoor image, an architectural image, a nature image, a city image, an art image, or an abstract image. However, the present disclosure is not limited thereto.

[0045] The first spatial image (10) can contribute to generating a final spatial image containing the style and objects desired by the user by adding a human figure to the spatial image. The first spatial image (10) can be processed by an electronic device (100) together with an input prompt (15) to generate a final spatial image (20) with scale and composition corrected. The first spatial image (10) can be preprocessed by an image preprocessing module (120) and converted into a form that can be processed by a first artificial intelligence model (130) and a second artificial intelligence model (140).

[0046] The input prompt (15) may indicate the characteristics of the spatial image desired by the user. The input prompt (15) may be provided via text, image, voice, gesture, or other user interface. However, the present disclosure is not limited thereto.

[0047] The input prompt (15) can be processed by the electronic device (100) together with the first spatial image (10) to generate the final spatial image (20). The input prompt (15) can specify the style, color, composition, scale, placement of objects, lighting conditions, or other visual elements of the spatial image desired by the user. The input prompt (15) can play an important role in the process in which the electronic device (100) generates the final spatial image (20) based on the first spatial image (10).

[0048] In one embodiment of the present disclosure, the input prompt (15) may be utilized in the processing of the second artificial intelligence model (140). The input prompt (15) may be provided to the second artificial intelligence model (140) along with guide information generated by the first artificial intelligence model (130). The input prompt (15) may help to more accurately reflect the characteristics of the final spatial image (20) desired by the user.

[0049] The electronic device (100) is a device capable of generating a final spatial image (20) by processing a first spatial image (10) and an input prompt (15). The electronic device (100) may include a human figure addition module (110), an image preprocessing module (120), a first artificial intelligence model (130), and a second artificial intelligence model (140). For example, the electronic device (100) may be a computer, a smartphone, a tablet, a server, a workstation, an embedded system, a cloud computing platform, or network equipment. However, the present disclosure is not limited thereto.

[0050] The electronic device (100) can generate a final spatial image (20) based on a first spatial image (10) and an input prompt (15). The electronic device (100) can add a human figure to the spatial image through a human figure addition module (110). The electronic device (100) can preprocess the input image through an image preprocessing module (120) to convert it into a form suitable for processing by an artificial intelligence model. The electronic device (100) can generate guide information based on the preprocessed image through a first artificial intelligence model (130). The electronic device (100) can generate a final spatial image based on the first spatial image (10), the input prompt (15), and the guide information through a second artificial intelligence model (140).

[0051] The human figure addition module (110) can perform the function of adding a human figure to a spatial image. The human figure addition module (110) can generate a second spatial image by adding a human figure to the first spatial image (10). For example, the human figure addition module (110) may be a full-body human figure. However, the present disclosure is not limited thereto, and may be a half-body, face, hands, feet, silhouette, or shadow figure of a human figure.

[0052] The human figure addition module (110) can transmit a second spatial image, in which human figures have been added to the first spatial image, to the image preprocessing module (120). The human figure addition module (110) can contribute to correcting the scale and composition of the spatial image. According to one embodiment of the present disclosure, by adding a plurality of human figures to the first spatial image (10), the human figures can be used as a measure of the scale and composition of the image generation when utilizing an artificial intelligence model trained on humans.

[0053] In one embodiment of the present disclosure, the human shape addition module (110) can store various human shape data and select and add a human shape suitable for a spatial image based thereon. The human shape addition module (110) may include a function to adjust the size, position, direction, etc. of the human shape according to user input. Through this function, a user-customized spatial image can be generated.

[0054] The image preprocessing module (120) can preprocess an input image and convert it into a form suitable for processing by an artificial intelligence model. For example, the image preprocessing module (120) can generate an outline image, a depth image, a color correction image, a noise removal image, a resolution adjustment image, a brightness adjustment image, a contrast adjustment image, etc. However, the present disclosure is not limited thereto.

[0055] The image preprocessing module (120) can generate a contour image and a depth image from the first spatial image (10). The image preprocessing module (120) can provide these preprocessed images to the first artificial intelligence model (130). The first artificial intelligence model (130) can generate guide information based on the preprocessed images. The guide information can be used by the second artificial intelligence model (140) to generate the final spatial image (20).

[0056] In one embodiment of the present disclosure, the image preprocessing module (120) can improve the quality of the input image by applying various preprocessing techniques. For example, the image preprocessing module (120) can perform operations such as increasing the resolution of the image, adjusting the color balance, or removing noise. Such preprocessing processes can contribute to improving the quality of the final spatial image (20).

[0057] The first artificial intelligence model (130) can generate guide information based on an image preprocessed within the electronic device (100). The first artificial intelligence model (130) can play an important role in an artificial intelligence-based image processing system. For example, the first artificial intelligence model (130) may be a ControlNet model, a Diffusion model, a Convolutional Neural Network (CNN), a Recurrent Neural Network (RNN), a Generative Adversarial Network (GAN), a Transformer, an Autoencoder, a Reinforcement Learning Model, a Deep Neural Network (DNN), etc. However, the present disclosure is not limited thereto.

[0058] The first artificial intelligence model (130) can generate guide information by receiving the contour image and depth image of the preprocessed image as input. The first artificial intelligence model (130) can provide the generated guide information to the second artificial intelligence model (140) so that the second artificial intelligence model (140) can generate a final spatial image (20) with high accuracy in scale and composition.

[0059] A second artificial intelligence model (140) can be used to generate a final spatial image within an electronic device (100). The second artificial intelligence model (140) can generate a final spatial image with clear scale and composition based on the first spatial image (10), input prompt (15), and guide information. The second artificial intelligence model (140) may be a diffusion model. However, the present disclosure is not limited thereto, and for example, the second artificial intelligence model (140) may be a Generative Adversarial Network (GAN), a Transformer, a Convolutional Neural Network (CNN), a Recurrent Neural Network (RNN), an Autoencoder, a Reinforcement Learning Model, or a Deep Neural Network (DNN).

[0060] The second artificial intelligence model (140) can generate a final spatial image (20) by utilizing guide information provided by the first artificial intelligence model (130). The second artificial intelligence model (140) can generate a final spatial image containing the style and objects desired by the user based on the first spatial image (10) and the input prompt (15). The second artificial intelligence model (140) can correct the scale and composition of the spatial image according to the input information.

[0061] In one embodiment of the present disclosure, the second artificial intelligence model (140) may generate a final spatial image (20) by accepting additional user input data in addition to the first spatial image (10) and the input prompt (15).

[0062] The final space image (20) may be a space image in which the first space image (10) has been modified to include the style and objects desired by the user. The final space image (20) may be generated based on the first space image (10), input prompt (15), and guide information.

[0063] In one embodiment of the present disclosure, the final spatial image (20) may be obtained in a form in which at least one object is added to the first spatial image (10). Alternatively, the final spatial image (20) may be obtained in a form in which at least one object of the first spatial image (10) is modified and / or in a form in which at least one object is added to the first spatial image (10). For example, at least one object may be furniture such as a desk or chair, or a building such as a window or a column.

[0064] FIG. 2 is a block diagram showing an image preprocessing module according to an embodiment of the present disclosure. Regarding the first spatial image (10) and the image preprocessing module (120), details that overlap with those described in FIG. 1 will be omitted. Hereinafter, the same reference numerals will be assigned to configurations identical to those described in FIG. 1, and redundant descriptions will be omitted.

[0065] Referring to FIG. 2 along with FIG. 1, the image preprocessing module (120) can receive a first spatial image (10) as input and generate a first contour image (11) and a first depth image (12).

[0066] The first contour image (11) may represent a contour extracted from the first spatial image (10). The first contour image (11) may represent the visual information of the first spatial image (10) in the form of a contour by simplifying it. For example, the first contour image (11) may include contours of buildings, furniture, people, etc. However, the present disclosure is not limited thereto.

[0067] The first depth image (12) may represent depth information extracted from the first spatial image (10). The first depth image (12) may serve to visually represent the depth information of the spatial image. For example, the first depth image (12) may be distance, height, depth, density, transparency, color, and brightness. However, the present disclosure is not limited thereto.

[0068] The first depth image (12) can be generated based on depth information extracted from the first spatial image (10). The first depth image (12) can be used as input to the first artificial intelligence model together with the first contour image (11). The first artificial intelligence model can obtain first guide information using the first depth image (12) and the first contour image (11).

[0069] FIG. 3 is a block diagram showing a human shape addition module and an image preprocessing module according to an embodiment of the present disclosure. Regarding the first spatial image (10), the human shape addition module (110), and the image preprocessing module (120), details that overlap with those described in FIG. 1 and FIG. 2 will be omitted. Hereinafter, the same reference numerals will be assigned to configurations identical to those described in FIG. 1 and FIG. 2, and duplicate descriptions will be omitted.

[0070] Referring to FIG. 3 together with FIG. 1, the human figure addition module (110) can generate a second spatial image (30) by adding a plurality of human figures to a first spatial image (10). The human figure addition module (110) can add a plurality of human figures to the first spatial image (10) based on user input. The human figure addition module (110) can transmit the second spatial image (30) to an image preprocessing module (120).

[0071] In one embodiment of the present disclosure, a human shape addition module (110) may add a first human shape and a second human shape to a first spatial image (10). The first human shape and the second human shape may have different sizes and ratios. For example, the first human shape may be placed at a position corresponding to the foreground in the first spatial image (10), and the second human shape may be placed at a position corresponding to the background in the first spatial image (10). In this case, the size of the first human shape may be larger than the size of the second human shape.

[0072] In one embodiment of the present disclosure, a first person shape may be positioned at a location corresponding to the left side of the center of the first spatial image (10), and a second person shape may be positioned at a location corresponding to the right side of the center of the first spatial image (10). In this case, the size of the first person shape may be the same as the size of the second person shape.

[0073] In one embodiment of the present disclosure, each of the plurality of human figures may be cropped from different human figure images. However, the present disclosure is not limited thereto, and each of the plurality of human figures may be cropped from the same human figure image.

[0074] In one embodiment of the present disclosure, a human figure addition module (110) may be placed on a first spatial image (10) based on user input. For example, the user input may correspond to coordinate values ​​of the first spatial image (10).

[0075] The image preprocessing module (120) can process the second spatial image (30) to generate a second contour image (31) and a second depth image (32). The second contour image (31) may represent a contour extracted from the second spatial image (30). The second depth image (32) may represent depth information extracted from the second spatial image (30).

[0076] FIG. 4 is a block diagram showing a first artificial intelligence model according to an embodiment of the present disclosure. Regarding the second contour image (31), the second depth image (32), and the first artificial intelligence model (130), details that overlap with those described in FIG. 3 and FIG. 1 will be omitted. Hereinafter, the same reference numerals will be assigned to configurations identical to those described in FIG. 1 to FIG. 3, and redundant descriptions will be omitted.

[0077] Referring to FIG. 4 together with FIG. 1, the first artificial intelligence model (130) can receive a second contour image (31) and a second depth image (32) as inputs and output second guide information that contributes to correcting the scale and composition of the space. The first artificial intelligence model (130) can generate second guide information (40). The first artificial intelligence model (130) may be a model that has been pre-trained to output a feature map corresponding to the input image.

[0078] The second guide information (40) can provide guidance necessary for generating a spatial image. The second guide information (40) can be generated from the first artificial intelligence model (130). The second guide information (40) can be generated using the second contour image (31) and the second depth image (32) as inputs. The second guide information (40) can be input a first number of times, which is an initial, predefined number of times, during the noise removal step of the second artificial intelligence model (140).

[0079] The second guide information (40) can contribute to correcting the scale and composition of the space through the first artificial intelligence model (130). The second guide information (40) can reflect the scale and composition desired by the user during the process of generating the space image.

[0080] FIG. 5 is a block diagram showing a first artificial intelligence model according to an embodiment of the present disclosure. Regarding the first contour image (11), the first depth image (12), and the first artificial intelligence model (130), details that overlap with those described in FIG. 2, FIG. 1, and FIG. 4 will be omitted. Hereinafter, the same reference numerals will be assigned to configurations identical to those described in FIG. 1 to FIG. 4, and redundant descriptions will be omitted.

[0081] Referring to FIG. 5 together with FIG. 1, the first artificial intelligence model (130) can receive a first contour image (11) and a first depth image (12) as inputs and generate first guide information (50). The first guide information (50) can provide guidance necessary for generating a spatial image. The first guide information (50) can be generated by the first artificial intelligence model (130). The first guide information (50) can be input a predefined second time after the first time during the noise removal step of the second artificial intelligence model (140).

[0082] Unlike the second guide information (40), the first guide information (50) may be guide information that does not contain information about human figures in the space. The first guide information (50) can perform the function of controlling the second artificial intelligence model (140) so that objects and depths in the space can be clearly determined after the scale and composition of the space image are corrected through the second guide information (40).

[0083] In one embodiment of the present disclosure, the first guide information (50) can serve as guidance during the noise removal process of the second artificial intelligence model (140). The first guide information (50) can be used together with the second guide information to improve the quality of the final spatial image. According to one embodiment of the present disclosure, in the process where noise containing a human shape is generated by the second guide information (40), the first guide information (50) performs the role of controlling the generation basis of the second artificial intelligence model (140), thereby enabling the removal of distortion caused by the human shape generated by the second guide information (40).

[0084] FIG. 6 is a block diagram showing a second artificial intelligence model according to an embodiment of the present disclosure. Regarding the first spatial image (10), input prompt (15), second guide information (40), first guide information (50), second artificial intelligence model (140), and final spatial image (20), details that overlap with those described in FIG. 1, FIG. 2, FIG. 3, FIG. 4, and FIG. 5 will be omitted. Hereinafter, the same reference numerals will be assigned to configurations identical to those described in FIG. 1 to FIG. 5, and redundant descriptions will be omitted.

[0085] Referring to FIG. 6 together with FIG. 1 to 5, a first spatial image (10) can be input to a second artificial intelligence model (140) along with an input prompt (15). A first guide information (50) and a second guide information (40) can be input to the second artificial intelligence model (140). The second artificial intelligence model (140) can output a final spatial image (20) by taking the first spatial image, the input prompt, the first guide information (50), and the second guide information (40) as inputs.

[0086] In one embodiment of the present disclosure, the time at which the first guide information (50) is input to the second artificial intelligence model (140) and the time at which the second guide information (40) is input to the second artificial intelligence model (140) may be different from each other. For example, the first guide information (50) may be input during a first time interval corresponding to an initial order of the denoising step (or noise removal step) of the second artificial intelligence model (140), and the second guide information (40) may be input during a second time interval following the first time interval of the denoising step (or noise removal step) of the second artificial intelligence model (140). However, the present disclosure is not limited thereto, and the order of the denoising step in which the first guide information (50) and the second guide information (40) are input may overlap or alternate with each other.

[0087] FIG. 7 is a conceptual diagram showing a noise removal step of a second artificial intelligence model according to one embodiment of the present disclosure. Hereinafter, the same reference numerals are assigned to configurations identical to those described in FIG. 1 to 6, and redundant descriptions are omitted.

[0088] Referring to FIG. 7 along with FIG. 1, the second artificial intelligence model (140) may be a diffusion model. The second artificial intelligence model (140) may be pre-trained to output noise-removed data by repeating for a number of times (T) defined as a noise removal step.

[0089] The predefined number of times (T) or total denoising order (T) may represent the total number of times or order in which the current noise is input to the second artificial intelligence model (140) and the next noise is output, from the initial noise to the final noise. The predefined number of times (T) may be composed of the sum of the first number of times (Tn) and the second number of times (n).

[0090] The first number of times (Tn) can be used in the initial step of the noise removal step. The first number of times (Tn) can be composed of the remaining number of times excluding the second number of times (n) from the predefined number of times (T). During the first number of times (Tn), first guide information extracted from the first contour image (31) and the first depth image (32) corresponding to the image with the added human shape can be input into the second artificial intelligence model (140).

[0091] The second number of times (n) can be used in the mid-to-late stage of the noise removal step. The second number of times (n) can be composed of the remaining number of times excluding the first number of times (Tn) from the predefined number of times (T). During the second number of times (n), second guide information extracted from the second contour image (11) and the second depth image (12) corresponding to the existing image without a human shape can be input into the second artificial intelligence model (140).

[0092] The reason for using an image without a human figure in the initial step of the noise removal step is that the human figure acts as a scale and compositional measure of the spatial image. Since the second artificial intelligence model (140) has been pre-trained with training data including a human figure, the human figure can perceive the scale and composition of the spatial image more clearly than the figures of other objects.

[0093] The reason for using an image without a human figure in the mid-to-late stages of the noise removal step is that the tendency to generate objects around the human figure is reduced as the human figure is referenced. Since the first guide information extracted from the first contour image (31) and the first depth image (32) corresponding to the image with the human figure added is continuously input into the second artificial intelligence model (140), the generation of furniture around the human figure may be limited or distorted, it can be understood that the existing image without a human figure, or an image from which the human figure has been removed from the image with the human figure added, is input into the second artificial intelligence model (140).

[0094] FIG. 8 is a conceptual diagram showing an electronic device according to one embodiment of the present disclosure. Regarding the human shape addition module (110), image preprocessing module (120), first artificial intelligence model (130), and second artificial intelligence model (140), details that overlap with those described in FIG. 1, FIG. 3, FIG. 2, FIG. 4, FIG. 5, and FIG. 6 will be omitted. Hereinafter, the same reference numerals will be assigned to configurations identical to those described in FIG. 1 to FIG. 7, and redundant descriptions will be omitted.

[0095] Referring to FIG. 8 together with FIG. 1, the electronic device (200) can generate a spatial image desired by the user by adding a human figure to the spatial image and correcting the scale and composition. The electronic device (200) may include a processor (210), storage (220), a communication interface (230), and memory (240).

[0096] The processor (210) is a device capable of performing various operations within the electronic device (200). The processor (210) can operate in cooperation with other components of the electronic device (200). The processor (210) can provide a method for generating a spatial image desired by the user by utilizing an image in which a human figure is added to the spatial image, thereby correcting the scale and composition of the space. The processor (210) can acquire a first spatial image and an input prompt. Based on the first spatial image, the processor (210) can acquire a first contour image and a first depth image. The processor (210) can acquire first guide information by using a first artificial intelligence model (130) that takes the first contour image and the first depth image as input. The processor (210) can acquire a second spatial image by adding a plurality of human figures to the first spatial image. Based on the second spatial image, the processor (210) can acquire a second contour image and a second depth image. The processor (210) can obtain second guide information by using a first artificial intelligence model (130) that takes a second contour image and a second depth image as inputs. The processor (210) can obtain a final spatial image in which at least one object is added to the first spatial image by using a second artificial intelligence model (140) that takes a first spatial image, an input prompt, first guide information, and second guide information as inputs.

[0097] Storage (220) is a device capable of storing and managing data. For example, storage (220) may be a hard disk drive (HDD), a solid state drive (SSD), a flash memory, a network attached storage (NAS), cloud storage, an optical disk, or a magnetic tape. However, the present disclosure is not limited thereto.

[0098] Storage (220) may include a database that stores multiple images, such as spatial images, human images, etc. A processor (210) may load data from storage (220) to perform a method according to the present disclosure.

[0099] The communication interface (230) may be a component that enables data transmission and reception between the electronic device (200) and an external device. For example, the communication interface (230) may be wired communication, wireless communication, short-range communication, long-range communication, Internet Protocol (IP) communication, Bluetooth communication, or Near Field Communication (NFC). However, the present disclosure is not limited thereto.

[0100] The communication interface (230) can support a method for generating a space image desired by the user by utilizing an image in which a human figure is added to the space image and correcting the scale and composition of the space. The communication interface (230) can receive a first space image and an input prompt from an external device. The communication interface (230) can transmit data necessary to acquire a first contour image and a first depth image based on the first space image. The communication interface (230) can transmit the final space image generated through the second artificial intelligence model (140) to an external device such as a user terminal.

[0101] The memory (240) can store and manage data within the electronic device (200). The memory (240) can store at least one instruction corresponding to the human figure addition module (110), the image preprocessing module (120), the first artificial intelligence model (130), and the second artificial intelligence model (140). For example, the memory (240) may be a volatile memory, a non-volatile memory, a flash memory, a dynamic random access memory (DRAM), a static random access memory (SRAM), a read-only memory (ROM), or an electrically eraseable programmable read-only memory (EEPROM). However, the present disclosure is not limited thereto.

[0102] The memory (240) can provide a method for generating a space image desired by the user by utilizing an image in which a human figure is added to the space image and correcting the scale and composition of the space. The memory (240) can store instructions required to acquire a first space image and an input prompt, and to acquire a first contour image and a first depth image. Additionally, it can store instructions required to acquire first guide information using a first artificial intelligence model (130) and to acquire a second space image by adding a plurality of human figures to the first space image.

[0103] In one embodiment of the present disclosure, the memory (240) may store instructions necessary to acquire a second contour image and a second depth image based on a second spatial image, and to acquire a final spatial image using a second artificial intelligence model (140). The memory (240) may store instructions necessary to generate a final spatial image in which at least one object is added to a first spatial image, using a first spatial image, an input prompt, first guide information, and second guide information as inputs.

[0104] The human figure addition module (110) can add a human figure to a spatial image by executing an instruction stored in the memory (240) of the electronic device (200). The human figure addition module (110) can obtain a second spatial image by adding a plurality of human figures to the first spatial image. The second spatial image is a form in which the scale and composition are corrected compared to the first spatial image, and can generate a spatial image desired by the user.

[0105] In one embodiment of the present disclosure, the human figure addition module (110) may add a first human figure located at a first depth and a second human figure located at a second depth corresponding to a distant view from the first depth. The size of the first human figure may be larger than the size of the second human figure. In this way, the human figure addition module (110) can more accurately correct the depth and composition of the spatial image.

[0106] The image preprocessing module (120) can execute instructions stored in the memory (240) of the electronic device (200) to obtain a first spatial image and an input prompt, and obtain a first contour image and a first depth image based on the first spatial image.

[0107] In one embodiment of the present disclosure, the image preprocessing module (120) can obtain a second contour image and a second depth image based on a second spatial image. The second contour image and the second depth image can be used as inputs to a first artificial intelligence model (130) to obtain second guide information. The first spatial image, input prompt, first guide information, and second guide information can be used as inputs to a second artificial intelligence model (140) to obtain a final spatial image.

[0108] The first artificial intelligence model (130) can receive a first contour image and a first depth image as inputs to obtain first guide information. The first artificial intelligence model (130) may be a model that has been pre-trained to output a feature map corresponding to the input image. The first artificial intelligence model (130) can receive a second contour image and a second depth image as inputs to obtain second guide information. The first artificial intelligence model (130) can operate according to instructions stored in the memory (240) of the electronic device (200).

[0109] The second artificial intelligence model (140) can output a final noise by removing input noise a predetermined number of times. The second artificial intelligence model (140) can generate a final spatial image by receiving a first spatial image, an input prompt, first guide information, and second guide information as inputs. The second artificial intelligence model (140) can generate a final spatial image by applying different weights to the first guide information and the second guide information.

[0110] In one embodiment of the present disclosure, the second artificial intelligence model (140) can repeatedly remove noise by receiving the first spatial image, input prompt, and second guide information as input during a first number of times. The second artificial intelligence model (140) can repeatedly remove noise by receiving the first spatial image, input prompt, and first guide information as input during a second number of times. The second artificial intelligence model (140) can obtain a final spatial image based on the final noise. The first number of times corresponds to 30% of a predefined number of times, and the second number of times corresponds to 70% of a predefined number of times. According to one embodiment of the present disclosure, if the first number of times exceeds 30%, the space around the human shape may be distorted by the human shape.

[0111] FIG. 9 is a flowchart illustrating a method for correcting scale and composition when generating an AI-based spatial image according to one embodiment of the present disclosure. Hereinafter, the same reference numerals are assigned to configurations identical to those described in FIG. 1 to 8, and redundant descriptions are omitted.

[0112] Referring to FIG. 9 together with FIG. 1, at step S910, the electronic device (100) may acquire a first spatial image and an input prompt. The first spatial image may be an image taken in various environments. The input prompt may be provided in various forms, such as text, voice, or gestures. For example, the input prompt may be text describing a specific element of the image desired by the user. However, the present disclosure is not limited thereto.

[0113] In one embodiment of the present disclosure, the step of acquiring a first spatial image and an input prompt may include the step of collecting data through various sensors or input devices.

[0114] In step S920, the electronic device (100) may acquire a first contour image and a first depth image based on a first spatial image. The first contour image may represent the boundaries of an object within the image. The first depth image may include depth information for each pixel.

[0115] In one embodiment of the present disclosure, the step of acquiring a first contour image and a first depth image based on a first spatial image may include the step of extracting detailed information of the image using various image processing techniques. For example, a contour may be extracted using an edge detection algorithm.

[0116] In step S930, the electronic device (100) can obtain first guide information by using a first artificial intelligence model that takes a first contour image and a first depth image as inputs. The first guide information may include information such as the location and size of an object within the image.

[0117] In step S940, the electronic device (100) can obtain a second spatial image by adding a plurality of human figures to the first spatial image. The human figures can be added in various poses and sizes.

[0118] In one embodiment of the present disclosure, the step of obtaining a second spatial image by adding a plurality of human figures to a first spatial image may include the step of generating human figures using 3D modeling technology and integrating them into the image. For example, natural human movement may be realized using motion capture data.

[0119] In step S950, the electronic device (100) can acquire a second contour image and a second depth image based on the second spatial image. The second contour image can additionally provide depth information of the human shape relative to the first contour data. The second depth image can additionally provide depth information of the human shape relative to the first depth data.

[0120] In one embodiment of the present disclosure, the step of acquiring a second contour image and a second depth image based on a second spatial image may utilize a pre-trained image segmentation model to more accurately identify objects within the image.

[0121] In step S960, the electronic device (100) can obtain second guide information by using a second artificial intelligence model that takes a second contour image and a second depth image as inputs.

[0122] In step S970, the electronic device (100) can obtain a final spatial image in which at least one object is added to the first spatial image by using a second artificial intelligence model that takes the first spatial image, input prompt, first guide information, and second guide information as inputs. The final spatial image may be an image adjusted to suit user requirements. The second artificial intelligence model can generate an optimal result by integrating various input data.

[0123] In one embodiment of the present disclosure, the step of acquiring a final spatial image using a second artificial intelligence model that takes a first spatial image, an input prompt, first guide information, and second guide information as inputs may include a step of improving the quality of the image by utilizing various artificial intelligence technologies. For example, a Generative Adversarial Network (GAN) may be additionally used to enhance the details of the image.

[0124] FIG. 10 is a drawing for explaining the operation of a generative model including a first artificial intelligence model and a second artificial intelligence model according to an embodiment of the present disclosure. Hereinafter, the same reference numerals are assigned to configurations identical to those described in FIG. 1 to 9, and redundant descriptions are omitted.

[0125] Referring to FIG. 10 together with FIG. 1 to 5, an input prompt and a first spatial image can be encoded through a first encoder (900). At this time, the first encoder (900) can encode the input prompt (200) into a suitable feature vector to provide to a second generative model (122) depending on the type of input prompt. For example, the input prompt may include at least one of text, image, video, and audio. In one embodiment of the present disclosure, when the input prompt (200) is text, the first encoder (900) may use Contrastive Language-Image Pretraining (CLIP).

[0126] In one embodiment of the present disclosure, the second artificial intelligence model (140) may be provided with an encoded input prompt, a first spatial image, a time vector (1010) encoded with time information, and noise data (1020) as inputs. In this case, the time vector (1010) may be encoded with information for providing order information of a noise removal step or a denoising step performed in the second generation model (122).

[0127] In one embodiment of the present disclosure, the second artificial intelligence model (140) may perform a denoising operation on noise data (1020) provided as input repeatedly as a number of times T. Herein, T may be referred to as the denoising number, which may be a predefined number of times.

[0128] In one embodiment of the present disclosure, noise data provided as input to the second artificial intelligence model (140) at any one order t included in T may be referred to as Zt. The second artificial intelligence model (140) may obtain differential noise data at t-1 as output based on an encoded input prompt provided as input at t, a time vector (910) in which time information is encoded, and Zt, which is noise data.

[0129] In one embodiment of the present disclosure, the second artificial intelligence model (140) performs denoising so that the order of T becomes smaller, and removes differential noise data obtained at each order from the input data so that Z0, which does not contain noise components at t=0, can be generated.

[0130] In one embodiment of the present disclosure, the noise data (1020) provided to the second artificial intelligence model (140) may be a latent vector encoded through an encoder. In this case, the final spatial image (20) may be generated by decoding Z0, which is an output component obtained through the second artificial intelligence model (140), through a decoder (940).

[0131] However, the present disclosure is not limited thereto, and the final spatial image (20) may be an image obtained as an output through the second artificial intelligence model (140).

[0132] In one embodiment of the present disclosure, an electronic device (100) can preprocess a second spatial image (30) to which a plurality of human figures have been added from a first spatial image (10) through an image preprocessing module (120). The preprocessed image may include a second contour image (31) and a second depth image (32). The electronic device (100) can provide the second contour image (31) and the second depth image (32) to a first artificial intelligence model (130).

[0133] In one embodiment of the present disclosure, the first artificial intelligence model (130) can output second guide information by taking the second contour image (31) and the second depth image (32) as inputs. The second guide information can be input to the second artificial intelligence model (140) from the order corresponding to T to Tn during the denoising step of the second artificial intelligence model (140). By utilizing the second guide information in the initial step of the denoising step, the intermediate noise Z, which has the spatial scale and composition on the spatial image, is formed. t-n-1 can be printed.

[0134] In one embodiment of the present disclosure, the electronic device (100) may preprocess the acquired first spatial image (10) through an image preprocessing module (120). The preprocessed image may include a first contour image (11) and a first depth image (12). The electronic device (100) may provide the first contour image (11) and the first depth image (12) to a first artificial intelligence model (130).

[0135] In one embodiment of the present disclosure, the first artificial intelligence model (130) can output second guide information by taking the second contour image (31) and the second depth image (32) as inputs. The second guide information can be input to the second artificial intelligence model (140) from the order corresponding to Tn-1 to 1 during the denoising step of the second artificial intelligence model (140). By utilizing the first guide information during the mid-to-late step of the denoising step, distortion caused by the human shape on the spatial image can be eliminated.

[0136] In one embodiment of the present disclosure, the electronic device (100) can determine the accuracy of the scale and composition while performing the denoising step. If the accuracy is less than a predefined threshold value, the first number of steps can be increased and the process can be carried out.

[0137] In one embodiment of the present disclosure, the method may further include the step of encoding a preprocessed image (contour image and depth image). The electronic device (100) may encode the preprocessed image (210) through a second encoder (1030). At this time, the second encoder (1030) may encode the preprocessed image such that the size of the input image matches the size of the noise data (1020) provided to the second artificial intelligence model (140). However, the present disclosure is not limited thereto, and it is understood that the first encoder (1030) may be omitted depending on the configuration and operation of the first artificial intelligence model (130) and the second artificial intelligence model (140).

[0138] In one embodiment of the present disclosure, the electronic device (100) may provide an encoded image as an input to a first artificial intelligence model (130). In one embodiment of the present disclosure, the first artificial intelligence model (130) may acquire noise data (1020) as an input along with the encoded image.

[0139] In one embodiment of the present disclosure, based on the encoded image and noise data (1020) provided as input to the first artificial intelligence model (130), guide information (first guide information or second guide information) may be provided to the second artificial intelligence model (140).

[0140] In one embodiment of the present disclosure, the electronic device (100) may place a first weight on guide information provided to the second artificial intelligence model (140) and a second weight on the input prompt. The greater the second weight is than the first weight, the more the second artificial intelligence model (140) can reflect the information to be modified through the input prompt (e.g., text) to generate the final spatial image (20).

[0141] In one embodiment of the present disclosure, the greater the first weight is than the second weight, the more the second artificial intelligence model (140) can generate a final spatial image (20) by reflecting the guide information. In this case, the accuracy of the scale and composition of the final spatial image (20) can be increased.

[0142] Hereinafter, embodiments of the present disclosure will be described in detail with reference to the accompanying drawings. The embodiments will be described clearly and in detail so that a person skilled in the art can easily practice the present disclosure. However, the scope of the rights is not limited or restricted by these embodiments. Identical or similar reference numerals are used for similar components in each drawing, and redundant descriptions of identical or similar components are omitted.

[0143] The terms used in the following description have been selected as common and universal in the relevant technical field, but other terms may exist depending on technological development and / or changes, conventions, preferences of the skilled technician, etc. Therefore, the terms used in the following description should not be understood as limiting the technical concept, but as illustrative terms to explain the embodiments.

[0144] In addition, there are terms arbitrarily selected by the applicant in specific cases, and their detailed meanings will be described in the relevant explanatory section. Therefore, the terms used in the description below must be understood not merely as their names, but based on their meanings and the content throughout the specification.

[0145] Singular expressions may include plural expressions unless the context clearly indicates otherwise. Terms used herein, including technical or scientific terms, may have the same meaning as generally understood by those skilled in the art as described in this specification. Additionally, terms including ordinal numbers, such as "first" or "second," used in this specification may be used to describe various components, but said components should not be limited by said terms. Such terms are used solely for the purpose of distinguishing one component from another.

[0146] When a part of a specification is described as "comprising" a certain component, this means that, unless specifically stated otherwise, it does not exclude other components but may include additional components. Furthermore, terms such as "part" or "module" as used in the specification refer to a unit that processes at least one function or operation, and this may be implemented in hardware or software, or as a combination of hardware and software.

[0147] Embodiments of the present disclosure are described below with reference to the attached drawings so that those skilled in the art can easily implement them. However, the present disclosure may be embodied in various different forms and is not limited to the embodiments described herein. Furthermore, in order to clearly explain the present disclosure in the drawings, parts unrelated to the explanation have been omitted, and similar parts throughout the specification have been given similar reference numerals. Also, the reference numerals used in each drawing are for the purpose of explaining each drawing, and different reference numerals used in different drawings are not intended to indicate different elements. The present disclosure will be described in detail below with reference to the attached drawings.

Claims

1. In a method for correcting scale and composition when generating AI-based spatial images, Step of acquiring a first spatial image and an input prompt; A step of acquiring a first contour image and a first depth image based on the first spatial image above; A step of obtaining first guide information using a first artificial intelligence model that takes the first contour image and the first depth image as inputs; A step of obtaining a second spatial image by adding a plurality of human figures to the first spatial image; A step of acquiring a second contour image and a second depth image based on the second spatial image above; A step of obtaining second guide information using the first artificial intelligence model that takes the second contour image and the second depth image as inputs; and A method comprising the step of obtaining a final spatial image in which at least one object is added to the first spatial image by using a second artificial intelligence model that takes as input the first spatial image, the input prompt, the first guide information, and the second guide information.

2. In Paragraph 1, The step of acquiring the second spatial image above is: A step of adding a first human figure located at a first depth on the spatial image above; A method comprising the step of adding a second human figure located at a second depth corresponding to a distant view from the first depth on the spatial image.

3. In Paragraph 2, A method in which the size of the first human figure is larger than the size of the second human figure.

4. In Paragraph 1, The above-mentioned first artificial intelligence model is a model pre-trained to output a feature map corresponding to an input image, and The above second artificial intelligence model is a method in which the model is pre-trained to output final noise by removing input noise a predetermined number of times.

5. In Paragraph 4, The step of acquiring the final spatial image above is: A step of repeatedly removing the noise using the second artificial intelligence model, which takes the first spatial image, the input prompt, and the second guide information as inputs, during a first number of times among the previously defined number of times; A step of repeatedly removing the noise using the second artificial intelligence model, which takes the first spatial image, the input prompt, and the first guide information as inputs, during a second number of times among the previously defined number of times; and A method comprising the step of obtaining the final spatial image based on the final noise from which the noise has been repeatedly removed.

6. In Paragraph 5, A method in which the first number is less than the second number.

7. In Paragraph 6, The above first number corresponds to 30% of the above previously defined number, and The above second number corresponds to 70% of the above previously defined number, a method.

8. In Paragraph 5, A method in which, in the above-mentioned second artificial intelligence model, the first weight for the first guide information and the second weight for the second guide information are different from each other.

9. In Paragraph 1, The above plurality of human figures is 2 to 3, a method.

10. In an electronic device, At least one processor including a processing circuit; and It includes at least one memory that stores at least one instruction, By the above at least one processor executing the above at least one instruction, the electronic device: Acquire the first spatial image and input prompt, Based on the first spatial image above, a first contour image and a first depth image are obtained, and Using a first artificial intelligence model that takes the first contour image and the first depth image as input, first guide information is obtained, and By adding a plurality of human figures to the first spatial image, a second spatial image is obtained, and Based on the second spatial image above, a second contour image and a second depth image are obtained, and Using the first artificial intelligence model that takes the second contour image and the second depth image as input, second guide information is obtained, and An electronic device that obtains a final spatial image in which at least one object is added to the first spatial image by using a second artificial intelligence model that takes the first spatial image, the input prompt, the first guide information, and the second guide information as inputs.