Image processing apparatus, image processing method, and image processing program
The image processing apparatus addresses the challenge of generating room images without unnecessary furniture and creating interior-coordinated images by using feature detection images to maintain the floor plan and shape of the area, resulting in efficient and automated image processing.
Patent Information
- Application Number
- JP2023211414
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-12-14
- Publication Date
- 2025-06-26
- Estimated Expiration
- 2043-12-14
AI Technical Summary
Existing technologies face challenges in efficiently generating room images without unnecessary furniture while maintaining the floor plan, and in creating interior-coordinated room images through virtual home staging without manual intervention.
An image processing apparatus and method that generates a feature detection image for object arrangement based on geometric features in the target image, and uses this image to create an object arrangement image where objects are added or removed while maintaining the shape of the area represented by the image.
Enables easy generation of new images with objects erased or added while preserving the area's shape, facilitating efficient image processing for room imaging applications.
Smart Images

Figure 2025095422000001_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to an image processing apparatus, an image processing method, and an image processing program for processing images.
Background Art
[0002] On property introduction sites, real estate companies disclose room images taken by them. When the room image is an image as follows, the room image is effective for attracting customers. A room image without unnecessary furniture in the room is effective for attracting customers because the floor plan can be understood. A room image coordinated with interior design is effective for attracting customers because an image of living can be created.
[0003] In the case of a property where people are living, it is not realistic to move out furniture in order to take a room image without unnecessary furniture. It is not realistic to perform home staging to coordinate the interior design for the actual room in order to take a room image coordinated with interior design because it is necessary to clean and move in furniture after moving out the furniture.
[0004] In the case of a property where no one is living, although it is possible to take a room image without unnecessary furniture, it is not realistic to perform home staging to take a room image coordinated with interior design because it is necessary to move in furniture.
Prior Art Documents
Patent Documents
[0005]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0006] Regarding a room image without unnecessary furniture, due to the above circumstances, it is efficient to obtain a room image without unnecessary furniture based on an image with furniture arranged. When obtaining a room image without unnecessary furniture by manual image editing, it takes time. Therefore, generating a room image without unnecessary furniture using an image generation model called image generation AI is being considered.
[0007] When generating a room image without unnecessary furniture in a room based on a room image with furniture arranged using an image generation model, there is a problem that the floor plan of the room is not maintained. For example, the floor plan of a room is the shape of an area formed by geometric features. For example, geometric features are boundary lines in an image of a plane forming a floor, wall, ceiling, etc., and figures in an image of a door, window, etc.
[0008] Regarding an interior-coordinated room image, due to the above circumstances, it is efficient to obtain an interior-coordinated room image based on a room image without unnecessary furniture by virtual home staging that performs interior coordination on the room image. When obtaining an interior-coordinated room image from virtual home staging in manual image editing, not only does it take time, but a sense of interior coordination is also required. Therefore, generating an interior-coordinated room image using an image generation model is being considered.
[0009] When generating an interior-coordinated room image based on a room image without unnecessary furniture by virtual home staging using an image generation model, there is a problem that the floor plan of the room is not maintained. Furthermore, there is also a problem that colors and patterns of the floor, wall, ceiling, etc. in the original room image, detailed shapes of windows, etc., and other facilities that should not be erased are not maintained.
[0010] Patent Document 1 discloses generating an image in which CG objects such as furniture and equipment are arranged at predetermined positions in a virtual area corresponding to a target area. However, Patent Document 1 does not generate an image by an image generation model based on the input image.
[0011] On one aspect, an object of the present invention is to provide a technology capable of easily generating a new image in which the objects arranged in this image are erased while maintaining the shape of the area represented by the image.
[0012] On another aspect, an object of the present invention is to provide a technology capable of generating a new image in which a new object is arranged in this area while maintaining the shape of the area represented by the image.
Means for Solving the Problems
[0013] The image processing apparatus according to the embodiment generates a feature detection image for object arrangement based on the detection of geometric features in the arrangement target image representing an area, which is an image of the target for arranging an object, and generates an object arrangement image representing the area where the objects not arranged in the arrangement target image are arranged based on the feature detection image for object arrangement, and includes an arrangement processing unit.
Effects of the Invention
[0014] According to one aspect of the embodiment, it is possible to easily generate a new image in which the objects arranged in the image are erased while maintaining the shape of the area represented by the image.
[0015] According to another aspect of the embodiment, it is possible to generate a new image in which a new object is arranged in this area while maintaining the shape of the area represented by the image.
Brief Description of the Drawings
[0016]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Figure 12
Figure 13
Figure 14
Figure 15
Figure 16
Figure 17
DETAILED DESCRIPTION OF THE INVENTION
[0017] [Embodiment] Hereinafter, several embodiments will be described with reference to the drawings. Note that the scales of the respective drawings used in the following description of the embodiments may be changed as appropriate. In addition, the respective drawings used in the following description of the embodiments may show the configuration with some components omitted for the purpose of explanation.
[0018] (Configuration example) FIG. 1 is a block diagram showing a configuration example of the image processing system S. The image processing system S is a system that generates an image using a machine learning model.
[0019] The machine learning model is a program module generated by performing machine learning on learning data. The machine learning model generates data based on the input data and outputs the generated data.
[0020] An image is an image drawn on a two-dimensional plane represented by a coordinate system of the x-axis and the y-axis. The notation "image" includes the meaning of image data representing an image unless explicitly stated to be an image displayed on a display device. Therefore, the notation "image" can be read as image data representing an image.
[0021] The image processing system S includes a server 1 and a terminal 2. The server 1 and the terminal 2 are connected to each other so as to be communicable via a network NW. The network NW is composed of one or more of various networks such as the Internet, a mobile communication network, and a LAN (Local Area Network). The one or more networks may include a wireless network or a wired network. In FIG. 1, one terminal 2 is illustrated, but the image processing system S can include a plurality of terminals 2.
[0022] Server 1 is a device that processes an image representing a region. The region is a three-dimensional space. The space may be a space closed by walls or the like, or may also be an open space. Here, a region inside a property will be described as an example. The property may be a residential property in various buildings such as condominiums, apartments, and detached houses, or a non-residential property such as an office. The residential property may be a rental property or a property for sale. In the case of a residential property, the region inside the property is a region such as a room or a living area, but is not limited thereto. Server 1 may be a server on the cloud. A configuration example of Server 1 will be described later. Server 1 is an example of an image processing device that processes images.
[0023] Terminal 2 is a device having an input function, a display function, a communication function, and the like. For example, Terminal 2 is a PC (Personal Computer), a smartphone, a tablet terminal, or the like, but is not limited thereto. For example, Terminal 2 is a terminal of a user of a real estate company, but may be a terminal of any user.
[0024] A configuration example of Server 1 will be described. Server 1 includes a processing circuit 11, a main memory 12, an auxiliary storage device 13, and a communication interface 14. The processing circuit 11, the main memory 12, the auxiliary storage device 13, and the communication interface 14 are connected to each other via a bus or the like so that signals can be input and output. In FIG. 1, the interface is described as "I / F".
[0025] The processing circuit 11 corresponds to the central part of the server 1. The processing circuit 11 constitutes the computer of the server 1. The processing circuit 11 includes one or more circuits that execute a plurality of processes by a plurality of functions. For example, the circuit is a processor, an ASIC (Application Specific Integrated Circuit), or an FPGA (Field-Programmable Gate Array), but is not limited thereto. For example, the processor is a CPU (Central Processing Unit) or a GPU (Graphics Processing Unit), but is not limited thereto. The processing circuit 11 expands the image processing program stored in advance in the main memory 12 or the auxiliary storage device 13 into the main memory 12. The image processing program is a program that enables the processing circuit 11 to execute the processing by each part described later. The processing circuit 11 can execute various processes by executing the image processing program expanded in the main memory 12.
[0026] The main memory 12 corresponds to the main storage part of the server 1. The main memory 12 includes a non-volatile memory area and a volatile memory area. The main memory 12 stores an operating system or a program in the non-volatile memory area. The main memory 12 uses the volatile memory area as a work area where data is appropriately rewritten by the processing circuit 11. For example, the main memory 12 includes a ROM (Read Only Memory) as the non-volatile memory area. For example, the main memory 12 includes a RAM (Random Access Memory) as the volatile memory area. The main memory 12 can store the image processing program.
[0027] The auxiliary storage device 13 corresponds to the auxiliary storage part of the server 1. The auxiliary storage device 13 includes one or more storage devices. The storage device is, but not limited to, EEPROM (registered trademark) (Electric Erasable Programmable Read-Only Memory), HDD (Hard Disc Drive), SSD (Solid State Drive), flash memory, etc. The auxiliary storage device 13 stores the above-mentioned image processing program, data used by the processing circuit 11 for performing various processes, and data generated by the processes in the processing circuit 11.
[0028] The auxiliary storage device 13 includes an image storage area 131. The image storage area 131 stores a plurality of images. The image storage area 131 is an example of a storage unit that stores images.
[0029] The image storage area 131 can store an image to be erased. The image to be erased is an image of an object to be erased. The image to be erased is an image representing an area where an object is placed. The image to be erased is an image of a photograph taken of an area where an object is placed. For example, the area where an object is placed is an area within a property where a person is living. The image to be erased is an image transmitted from the terminal 2 to the server 1.
[0030] An object has a shape. For example, an object is movable from an area. Movable objects include not only those not fixed to a wall or the like forming an area, but also those fixed to a wall or the like forming an area but removable. In this example, an object is, but not limited to, furniture, electrical appliances, lighting fixtures, curtains, or articles, etc. An object may form an area. In this example, an object is, but not limited to, a window or a floor, etc. Hereinafter, the notation "area where an object is placed" is intended to mean an area where a movable object such as furniture is placed. The notation "area where no object is placed" is intended to mean an area where no movable object such as furniture is placed. A movable object is a different object from an object forming an area.
[0031] The image storage area 131 can store an image for setting an object to be erased. The image for setting an object to be erased is an image obtained by drawing a mask on one or more objects to be erased in the image of the object to be erased. The object to be erased is an object to be erased in the image of the object to be erased. The object to be erased is an unnecessary object in the image of the object to be erased. Erasing an object includes the meaning of leaving no object or making the object invisible. The type of the object to be erased may be set via the terminal 2. Generating the image for setting an object to be erased is an example of setting one or more objects to be erased in the image of the object to be erased. The image for setting an object to be erased is an example of setting the object to be erased.
[0032] The image storage area 131 can store an object-erased image. The object-erased image is an image obtained by erasing the object to be erased from the image of the object to be erased.
[0033] The image storage area 131 can store a feature detection image for object erasing. The feature detection image for object erasing is an image showing geometric features detected from the object-erased image. Geometric features are geometric features detectable from an image. For example, geometric features are boundary lines in an image forming planes such as a floor, a wall, and a ceiling, and line segments of figures in an image such as a door and a window, but are not limited thereto. The line segment may include a straight line or a curve. Since regions in a property include many straight lines, geometric features preferably include straight lines. Here, a straight line is taken as an example of geometric features for explanation. Since the feature detection image for object erasing shows the shape of a region formed by geometric features, it is used to generate an object-erased image while maintaining the shape of the region represented by the image of the object to be erased. For example, the shape of the region formed by geometric features is a floor plan of a room, but is not limited thereto.
[0034] The image storage area 131 can store an image for object placement. The image for object placement is an image of an object placement target. The image for object placement is an image representing a region. For example, the image for object placement is an object-erased image.
[0035] Note that the image to be placed is not limited to the object-eliminated image. In the case of an unoccupied property, the server 1 does not need to generate an object-eliminated image based on the image to be eliminated. Therefore, the image to be placed may be an image of a photograph taken of the area. For example, the area is an area in an unoccupied property where furniture or the like is not placed. In this case, the image to be placed is an image transmitted from the terminal 2 to the server 1.
[0036] The image storage area 131 can store a feature detection image for object placement. The feature detection image for object placement is an image showing geometric features detected from the image to be placed. The geometric features are as described above. Since the feature detection image for object placement shows the shape of the area formed by the geometric features, it is used to generate an object placement image while maintaining the shape of the area represented by the image to be placed.
[0037] The image storage area 131 can store a reference image. The reference image is an image representing an area where an object is placed. The reference image is an image used as a reference for generating an object placement image described later. The reference image may be an image transmitted from the terminal 2 to the server 1. The image transmitted from the terminal 2 to the server 1 may be an image of a photograph taken or a generated image. Hereinafter, the reference image transmitted from the terminal 2 to the server 1 is also referred to as a first reference image. The reference image may be an image generated using an image generation model for the reference image described later. Hereinafter, the reference image generated using the image generation model for the reference image is also referred to as a second reference image.
[0038] The image storage area 131 can store the image to be held. The image to be held is an image in which a mask is drawn on one or more objects to be held in the image to be arranged. The object to be held is the object to be held in the image to be arranged. The object to be held is an object that is not to be redrawn in the image to be arranged. Holding an object includes the meaning of leaving the object and not redrawing the object. The type of the object to be held may be set via the terminal 2. Generating the image to be held is an example of setting one or more objects to be held in the image to be arranged. The image to be held is an example of setting the object to be held.
[0039] The image storage area 131 can store the object arrangement image. The object arrangement image is an image representing an area where one or more objects not arranged in the image to be arranged are arranged. The object arrangement image is an image in which objects suitable for this area are arranged while maintaining the shape of the area represented by the image to be arranged.
[0040] The image storage area 131 can store the image of the object to be synthesized. The image of the object to be synthesized is an image in which a mask is drawn on one or more objects to be synthesized in the object arrangement image. The object to be synthesized is the object to be synthesized in the image to be arranged. The object to be synthesized is an object newly arranged in the object arrangement image. That is, the object to be synthesized is an object that is not arranged in the image to be arranged among the objects arranged in the area represented by the object arrangement image. Generating the image of the object to be synthesized is an example of setting one or more objects to be synthesized in the object arrangement image. The image of the object to be synthesized is an example of setting the object to be synthesized.
[0041] The image storage area 131 can store a composite image. The composite image is an image obtained by synthesizing an object onto an image to be arranged. The composite image is an image representing an area where one or more objects not arranged in the image to be arranged are arranged. The object synthesized onto the image to be arranged to generate the composite image is an object not arranged in the image to be arranged. The synthesized object is an object to be synthesized with a mask drawn thereon in the image of the object to be synthesized. The composite image is an image in which an object suitable for the area is arranged while maintaining the shape of the area represented by the image to be arranged.
[0042] The image storage area 131 can store an edge detection image. The edge detection image is an image showing edges detected from the composite image. An edge includes the meaning of a contour line. For example, the edges detected from the composite image are the edges of each object. Since the edge detection image shows the shape of an object with edges, it is used to generate an image that fits while maintaining the shape of each object arranged in the area represented by the composite image.
[0043] The image storage area 131 can store an outer edge setting image. The outer edge setting image is an image with a mask drawn on the outer edges of one or more objects synthesized onto the image to be arranged to generate the composite image. The outer edge of an object is the boundary between the object and the image to be arranged onto which the object is synthesized. Generating the outer edge setting image is an example of setting the outer edges of one or more objects synthesized onto the image to be arranged to generate the composite image. The outer edge setting image is an example of the setting of the outer edge.
[0044] The image storage area 131 can store an image that fits. The image that fits is an image representing an area where one or more objects not arranged in the image to be arranged are arranged. The image that fits is an image with the outer edges made to fit in the composite image. Making the outer edges fit includes making the color of the boundary between the object and the image to be arranged onto which the object is synthesized closer to a natural state. The image that fits is an image in which an object suitable for the area is arranged while maintaining the shape of the area represented by the image to be arranged.
[0045] Here, an example has been described in which the image storage area 131 stores each of the above-described images, but the present invention is not limited thereto. Each image may be stored in a different image storage area for each type. In this example, each image storage area included in the auxiliary storage device 13 is an example of a storage unit. A part of each image may be stored in the main memory 12 instead of the auxiliary storage device 13. In this example, the main memory 12 is an example of a storage unit.
[0046] The auxiliary storage device 13 includes a machine learning model storage area 132. The machine learning model storage area 132 stores a plurality of machine learning models. The machine learning model storage area 132 is an example of a storage unit that stores machine learning models.
[0047] The machine learning model storage area 132 can store an object detection model that detects an object in an image. The object detection model is a machine learning model generated by performing machine learning on learning data. The object detection model generates a detection result of an object detected in the input image based on the input image, and outputs the detection result. For example, the detection result includes the object name of the detected object. The object detection model may be a method based on zero-shot object detection. The object detection model may be a method based on instance segmentation. For example, the object detection model is Detic (Detector with Image Classes), but the present invention is not limited thereto.
[0048] The machine learning model storage area 132 can store a geometric feature detection model that detects geometric features in an image. The geometric feature detection model is a machine learning model generated by performing machine learning on learning data. The geometric feature detection model generates an image indicating geometric features detected from the input image based on the input image, and outputs the generated image. When the geometric feature is a straight line, for example, the geometric feature detection model is MLSD (Mobile Line Segment Detection), but the present invention is not limited thereto.
[0049] The machine learning model storage area 132 can store an edge detection model for detecting edges in an image. The edge feature detection model is a machine learning model generated by performing machine learning on learning data. The edge feature detection model generates an image indicating the detected edges from the input image based on the input image, and outputs the generated image. For example, the edge detection model is Canny, but is not limited thereto.
[0050] The machine learning model storage area 132 can store an image generation model for object elimination that generates an object-eliminated image. The image generation model for object elimination is a program module generated by performing machine learning on learning data. The learning data can include a plurality of images representing areas where no objects are placed. That is, the image generation model for object elimination is generated by performing machine learning on an image whose area shape is known, such as an image representing an area where no objects are placed. The diffusion model (Stable Diffusion) will be described as an example of the image generation model for object elimination, but the image generation model for object elimination is not limited to the diffusion model. For example, the diffusion model is Stable Diffusion v1.5_inpainting. The image generation model for object elimination generates an object-eliminated image by i2i (image to image) using an inpainting function based on the input image, and outputs the object-eliminated image. As will be described below, the image generation model for object elimination can incorporate, as an extended model, a combination of zero or more machine learning models for fine-tuning and zero or more machine learning models for control.
[0051] For the diffusion model for object elimination, as an extended model, a machine learning model for fine-tuning the diffusion model can be incorporated. The machine learning model for fine-tuning is a program module generated by performing machine learning on learning data. The learning data can include a plurality of images representing areas where no objects are placed. That is, the machine learning model for fine-tuning is generated by performing machine learning on an image whose area shape is known, such as an image representing an area where no objects are placed. The machine learning model for fine-tuning attempts to draw an area where no objects are placed regardless of the input. Since the machine learning model for fine-tuning attempts to generate an image representing an area where no objects are placed from just noise, it can also be used in t2i. Since the machine learning model for fine-tuning attempts to generate an image representing an area where no objects are placed from the noise applied on the area represented by the input image rather than just noise, it can also be used in i2i. Therefore, the machine learning model for fine-tuning is a model that outputs an image representing an area where no objects are placed based on at least one of the input text data and the image representing the area where the object is placed. In this way, the diffusion model for object elimination incorporates a machine learning model for fine-tuning that is more proficient in object elimination. The machine learning model for fine-tuning is used to draw the target area. For example, the machine learning model for fine-tuning is LoRA (Low-Rank Adaptation) or LoCon (LoRA for Convolution network), but is not limited thereto.
[0052] In the diffusion model for object elimination, as an extended model, a control machine learning model for controlling the generation of an image by the diffusion model can be incorporated. For example, the control machine learning model is, but not limited to, ControlNet. The control machine learning model has a function of geometric feature detection such as MLSD. The control machine learning model is a program module generated by machine learning of learning data. The learning data can include a plurality of images showing geometric features of areas where no objects are placed and a plurality of images representing areas where no objects are placed. The image showing the geometric features of the area where no objects are placed only needs to show geometric features such that the shape of the area such as the floor plan can be understood. Therefore, the image showing the geometric features of the area where no objects are placed can be easily generated by a geometric feature detection model such as MLSD based on the image representing the area where no objects are placed, but is not limited to this. The image showing the geometric features of the area where no objects are placed may be one in which geometric features are drawn freehand. For example, the learning data may include 100 pairs of images, each pair consisting of an image showing the geometric features of the area where no objects are placed and an image representing the area where no objects are placed. By machine learning two types of images in pairs, the method of drawing an area from geometric features is machine-learned using a neural network. The control machine learning model is a model that outputs an image that retains geometric features and represents an area where no objects are placed, based on the input image showing the geometric features. In this way, in the diffusion model for object elimination, a model having a function of geometric feature detection that is good at retaining the shape of the area is incorporated as the control machine learning model. In this way, when an image showing geometric features is input to the model having a function of geometric feature detection as the control machine learning model, an image representing an area retaining geometric features is more likely to be generated. Furthermore, by combining the above-mentioned machine learning model for fine-tuning with the model having a function of geometric feature detection as the control machine learning model, object elimination is more likely to occur.
[0053] The machine learning model memory area 132 can store an image generation model for object placement that generates an object placement image. The image generation model for object placement is a program module generated by machine learning of learning data. The learning data can include a plurality of images representing areas where objects are placed. That is, the image generation model for object placement is generated by machine learning an image in which an object of the correct size is placed in an area, such as an image representing an area where an object is placed. A diffusion model will be used as an example of the image generation model for object placement, but the image generation model for object placement is not limited to the diffusion model. For example, the diffusion model is Stable Diffusion v1.5. As will be described below, the image generation model for object placement can incorporate, as an extended model, a combination of zero or more machine learning models for fine-tuning and zero or more machine learning models for control.
[0054] In the first placement mode, the image generation model for object placement refers to the input feature detection image for object placement, generates an object placement image by t2i (text to image) based on the input text data, and outputs the object placement image. For example, the text data includes data of a text about the object placement image. In the second placement mode, the image generation model for object placement refers to the input feature detection image for object placement, generates an object placement image by i2i based on the input reference image, and outputs the object placement image. In the third placement mode, the image generation model for object placement refers to the input feature detection image for object placement, generates an object placement image by i2i using an inpainting function based on the input retention target setting image, and outputs the object placement image. The image generation model for object placement inpaints portions of the retention target setting image other than the object to be retained.
[0055] For the diffusion model for object placement, as an extended model, a machine learning model for fine-tuning the diffusion model can be incorporated. The machine learning model for fine-tuning is a program module generated by performing machine learning on learning data. The learning data can include a plurality of images representing the area where the object is placed. That is, the machine learning model for fine-tuning is generated by performing machine learning on an image in which an object of the correct size is placed in the area, such as an image representing the area where the object is placed. The machine learning model for fine-tuning attempts to draw the area where the object is placed regardless of the input. Since the machine learning model for fine-tuning attempts to generate an image representing the area where the object is placed from just noise, it can also be used in t2i. Since the machine learning model for fine-tuning attempts to generate an image representing the area where the object is placed from the noise applied on the area represented by the input image rather than just noise, it can also be used in i2i. Therefore, the machine learning model for fine-tuning is a model that outputs an image representing the area where the object is placed based on at least one of the input text data and the image representing the area where the object is not placed. In this way, the diffusion model for object placement incorporates a machine learning model for fine-tuning that is more proficient in object placement. The machine learning model for fine-tuning is used to draw the target area. For example, the machine learning model for fine-tuning is LoRA or LoCon, but is not limited thereto.
[0056] For the diffusion model for object placement, as an extended model, a control machine learning model for controlling the generation of images by the diffusion model can be incorporated. For example, the control machine learning model can be, but is not limited to, ControlNet. In one example, the control machine learning model has a function of geometric feature detection such as MLSD. In this example, the control machine learning model is a program module generated by machine learning of learning data. The learning data can include a plurality of images showing geometric features of areas where no objects are placed and a plurality of images representing areas where no objects are placed, similar to the control machine learning model in the diffusion model for object elimination described above. The control machine learning model is a model that outputs an image that retains geometric features and represents an area where no objects are placed, based on the input image showing geometric features. In this way, the diffusion model for object placement incorporates, as the control machine learning model, a model with a function of geometric feature detection that is good at retaining the shape of the area. In this way, when an image showing geometric features is input to a model with a function of geometric feature detection as the control machine learning model, an image representing an area that retains geometric features is more likely to be generated. Furthermore, by combining the above-described machine learning model for fine-tuning with a model having a function of geometric feature detection as the control machine learning model, object placement is more likely to occur. In another example, the control machine learning model has a function of converting an area where no objects are placed to an area where objects are placed. In this example, the control machine learning model is a program module generated by machine learning of learning data. The learning data can include a plurality of images representing areas where objects are placed. The control machine learning model outputs an image representing an area where objects are placed, based on the input image representing an area where no objects are placed. In this way, the diffusion model for object placement incorporates, as the control machine learning model, a model with a function of converting an area where no objects are placed to an area where objects are placed.
[0057] The machine learning model memory area 132 can store an image generation model for a reference image. The image generation model for a reference image is a program module generated by performing machine learning on learning data. The learning data can include a plurality of images representing regions where objects are arranged. Although the diffusion model will be described as an example of the image generation model for a reference image, the image generation model for a reference image is not limited to the diffusion model. For example, the diffusion model is Stable Diffusion v1.5. The image generation model for a reference image generates a second reference image by t2i based on the input text data and outputs the second reference image. For example, the text data includes data of a text about the reference image. As will be described below, the image generation model for a reference image can incorporate zero or more machine learning models for fine-tuning as an extended model.
[0058] For the diffusion model for reference images, as an extended model, a machine learning model for fine-tuning the diffusion model can be incorporated. The machine learning model for fine-tuning is a program module generated by performing machine learning on learning data. The learning data can include a plurality of images representing regions where objects are placed. That is, the machine learning model for fine-tuning is generated by performing machine learning on images with objects of the correct size placed in regions, such as images representing regions where objects are placed. The machine learning model for fine-tuning attempts to draw a region where an object is placed regardless of the input. Since the machine learning model for fine-tuning attempts to generate an image representing a region where an object is placed from just noise, it can also be used in text-to-image (t2i). Since the machine learning model for fine-tuning attempts to generate an image representing a region where an object is placed from noise applied on the region represented by the input image rather than just noise, it can also be used in image-to-image (i2i). Therefore, the machine learning model for fine-tuning is a model that outputs an image representing a region where an object is placed based on at least one of the input text data and an image representing a region where no object is placed. Thus, the diffusion model for reference images incorporates a machine learning model for fine-tuning that is more proficient in object placement. The machine learning model for fine-tuning is used to draw a target region. For example, the machine learning model for fine-tuning is LoRA or LoCon, but is not limited thereto.
[0059] The machine learning model memory area 132 can store an image generation model for synthesis. The image generation model for synthesis is a program module generated by performing machine learning on learning data. The learning data can include a plurality of images representing regions where objects are placed. Although a diffusion model will be described as an example of the image generation model for synthesis, the image generation model for synthesis is not limited to the diffusion model. For example, the diffusion model is Stable Diffusion v1.5_inpainting. The image generation model for synthesis refers to the input edge detection image and generates and outputs an adapted image by i2i using an inpainting function based on the input outer edge setting image. As will be described below, the image generation model for synthesis can incorporate, as an extended model, a combination of zero or more machine learning models for fine-tuning and zero or more control machine learning models.
[0060] The diffusion model for synthesis can incorporate, as an extended model, a machine learning model for fine-tuning for fine-tuning the diffusion model. The machine learning model for fine-tuning is a program module generated by performing machine learning on learning data. The learning data can include a plurality of images representing regions where objects are placed. That is, the machine learning model for fine-tuning is generated by performing machine learning on an image in which an object of the correct size is placed in a region, such as an image representing a region where an object is placed. The machine learning model for fine-tuning is a model that outputs an image representing a region where an object is placed based on an input image representing a region where no object is placed. In this way, the diffusion model for synthesis incorporates a machine learning model for fine-tuning that is more proficient in object placement. The machine learning model for fine-tuning is used to draw the target region. For example, the machine learning model for fine-tuning is LoRA or LoCon, but is not limited thereto.
[0061] A machine learning model for control for controlling the generation of an image by the diffusion model can be incorporated as an extended model into the diffusion model for synthesis. For example, the machine learning model for control is ControlNet, but is not limited to this. In one example, the machine learning model for control has an edge detection function such as Canny. In this example, the machine learning model for control is a program module generated by machine learning the learning data. The learning data can include a plurality of images showing the edges of an area where an object is not placed and a plurality of images representing an area where an object is not placed. The image showing the edges of an area where an object is not placed may be one that allows the edges of the area to be recognized. Therefore, the image showing the edges of an area where an object is not placed can be easily generated by an edge detection model such as Canny based on the image representing the area where an object is not placed, but is not limited to this. The image showing the edges of an area where an object is not placed may be one in which the edges are drawn freehand. For example, the learning data may include 100 images each of an image showing the edges of an area where an object is not placed and an image representing an area where an object is not placed, which are paired. By machine learning two types of images in pairs, a method of drawing an area from an edge is machine learned using a neural network. The machine learning model for control is a model that outputs an image that retains edges and represents an area where an object is not placed, based on an image showing an input edge. In this way, a model with an edge detection function that is good at retaining objects is incorporated as a machine learning model for control in the diffusion model for synthesis. The reason for using a model with an edge detection function as a machine learning model for control is to distinguish between the contour of an object and the lines of incongruity caused by image synthesis. The incongruity caused by image synthesis does not appear as a clear line. Therefore, by generating an image of the edge of an object, which is a clear line, as the machine learning model for control, it is possible to eliminate the incongruity caused by image synthesis. For example, ControlNet has a parameter called strength, and if it is set to a threshold value or higher that makes the edge of the object visible, the incongruity is not taken into account and is drawn as the background of the area.In another example, the machine learning model for control has a function of converting an area where an object is not placed into an area where the object is placed. In this example, the machine learning model for control is a program module generated by machine learning of learning data. The learning data can include a plurality of images representing an area where an object is placed. The machine learning model for control outputs an image representing an area where an object is placed based on an input image representing an area where an object is not placed. In this way, a model having a function of converting an area where an object is not placed into an area where the object is placed is incorporated as the machine learning model for control into the diffusion model for synthesis.
[0062] Here, an example in which the machine learning model storage area 132 stores each of the above-described machine learning models has been described, but the present invention is not limited to this. Each machine learning model may be stored in a different machine learning model storage area for each type. In this example, each machine learning model storage area included in the auxiliary storage device 13 is an example of a storage unit.
[0063] The communication interface 14 includes various interfaces that connect the server 1 communicably to other devices using a communication protocol defined by the network NW.
[0064] Note that the hardware configuration of the server 1 is not limited to the above-described configuration. The server 1 can appropriately omit and change the above-described components and add new components.
[0065] Each part realized by the processing circuit 11 will be described. The processing circuit 11 realizes an erasing processing unit 111, a placement processing unit 112, and a synthesis processing unit 113. Each part realized by the processing circuit 11 can also be referred to as each function. Each part realized by the processing circuit 11 can also be said to be realized by a control unit including the processing circuit 11 and the main memory 12. The erasing processing unit 111 executes erasing processing described later. The placement processing unit 112 executes placement processing described later. The synthesis processing unit 113 executes the synthesis processing described later.
[0066] (Operation example) Next, an operation example of the server 1 configured as described above will be described. Note that the processing procedures described below are merely examples, and each process may be changed as much as possible. Also, regarding the processing procedures described below, depending on the embodiment, steps can be omitted, replaced, and added as appropriate.
[0067] FIG. 2 is a flowchart showing an example of processing by the server 1.
[0068] The erasing processing unit 111 executes erasing processing (step S1). The erasing processing includes generating an object-erased image using an image generation model for object erasing based on the image to be erased. A typical example of the erasing processing will be described later.
[0069] The placement processing unit 112 executes placement processing (step S2). The placement processing includes generating an object-placement image using an image generation model for object placement based on the image to be placed. When the processing circuit 11 executes the placement processing after executing the erasing processing, the image to be placed is the object-erased image. The processing circuit 11 may omit the erasing processing. In this case, the image to be placed is not the object-erased image but the image of the photo taken of the area. A typical example of the placement processing will be described later.
[0070] The synthesis processing unit 113 executes synthesis processing (step S3). The synthesis processing includes generating a synthesized image based on the image to be placed and the object-placement image. The synthesis processing includes generating an adapted image using an image generation model for synthesis based on the synthesized image. A typical example of the synthesis processing will be described later.
[0071] FIG. 2 shows an example in which the processing circuit 11 executes the deletion process, the arrangement process, and the synthesis process, but is not limited thereto. The processing circuit 11 executes the deletion process, but may not execute the arrangement process and the synthesis process. The processing circuit 11 may execute the arrangement process and the synthesis process without executing the deletion process. The processing circuit 11 executes at least the arrangement process, but may not execute the synthesis process. In the synthesis process, the processing circuit 11 executes the generation of the synthesized image, but may not execute the generation of the adapted image.
[0072] The processing circuit 11 transmits the output image generated by the server 1 to the terminal 2 via the communication interface. For example, the output image is an image representing an area where one or more objects not arranged in the object deletion image or the placement target image are arranged, but may be an image generated by the server 1 other than these. The image representing the area where one or more objects not arranged in the placement target image are arranged is an object placement image, a synthesized image, or an adapted image. The terminal 2 receives the output image from the server 1. The terminal 2 causes the display device to display an image based on the output image from the server 1.
[0073] FIG. 3 is a flowchart showing an example of the deletion process by the server 1.
[0074] The deletion processing unit 111 acquires the deletion target image from the image storage area 131 (step S101).
[0075] The erasing processing unit 111 detects an object in the image to be erased (step S102). In step S102, for example, the erasing processing unit 111 detects an object in the image to be erased using an object detection model. Here, the erasing processing unit 111 inputs the image to be erased to the object detection model. The object detection model generates a detection result of the object detected in the image to be erased based on the image to be erased and outputs the detection result. The erasing processing unit 111 acquires the detection result output from the object detection model. The erasing processing unit 111 detecting an object in the image to be erased includes the erasing processing unit 111 inputting the image to be erased to the object detection model. The erasing processing unit 111 detecting an object in the image to be erased includes the erasing processing unit 111 acquiring the detection result output from the object detection model.
[0076] The erasing processing unit 111 generates an image to be erased setting image based on the object detection in the image to be erased (step S103). In step S103, for example, the erasing processing unit 111 draws a mask on one or more objects to be erased in the image to be erased based on the detection result in the image to be erased. The erasing processing unit 111 can draw a mask on one or more objects to be erased in the image to be erased based on the type of the object to be erased set via the terminal 2. The erasing processing unit 111 generates an image to be erased setting image based on the drawing of the mask.
[0077] The erasure processing unit 111 generates an object erasure image based on the erasure target image and the erasure target setting image (step S104). In step S104, for example, the erasure processing unit 111 generates an object erasure image using an image generation model for object erasure. Here, the erasure processing unit 111 inputs the erasure target image and the erasure target setting image into the image generation model for object erasure. The image generation model for object erasure refers to the erasure target setting image, generates an object erasure image based on the erasure target image, and outputs the object erasure image. For example, the image generation model for object erasure inpaints the part of the object to be erased in the erasure target image to generate an object erasure image. As described above, the image generation model for object erasure incorporates a machine learning model for fine-tuning that is more proficient in object erasure. By using the machine learning model for fine-tuning, the image generation model for object erasure can obtain the effect of maintaining the shape of the region. This is because the machine learning model for fine-tuning is generated by machine learning an image in which the shape of the region is known. By using the machine learning model for fine-tuning, the image generation model for object erasure can generate an object erasure image using information such as what the shape of the region is. The image generation model for object erasure redraws only the part of the object to be erased by looking at the entire image, but by using the machine learning model for fine-tuning, it can draw a correct straight line for a part that is hidden by an extra object but should originally be a straight line. The erasure processing unit 111 acquires the object erasure image output from the image generation model for object erasure. The erasure processing unit 111 generating the object erasure image includes the erasure processing unit 111 inputting the erasure target image and the erasure target setting image into the image generation model for object erasure. The erasure processing unit 111 generating the object erasure image includes the erasure processing unit 111 acquiring the object erasure image output from the image generation model for object erasure. The erasure processing unit 111 stores the object erasure image in the image storage area 131.
[0078] The elimination processing unit 111 determines whether to finalize the object-eliminated image (step S105). In step S105, for example, if the object to be eliminated has been eliminated in the object-eliminated image, the elimination processing unit 111 determines that the object-eliminated image is finalized. The degree of elimination can be set as appropriate.
[0079] When the elimination processing unit 111 determines to finalize the object-eliminated image (step S105, YES), the process ends. When the elimination processing unit 111 determines not to finalize the object-eliminated image (step S105, NO), the process transitions from step S105 to step S106.
[0080] The elimination processing unit 111 generates a feature detection image for object elimination based on the detection of geometric features in the object-eliminated image (step S106). In step S106, for example, the elimination processing unit 111 uses a geometric feature detection model to generate a feature detection image for object elimination. Here, the elimination processing unit 111 inputs the object-eliminated image to the geometric feature detection model. The geometric feature detection model generates a feature detection image for object elimination based on the object-eliminated image and outputs the feature detection image for object elimination. The elimination processing unit 111 acquires the feature detection image for object elimination output from the geometric feature detection model. The elimination processing unit 111 generating the feature detection image for object elimination includes the elimination processing unit 111 inputting the object-eliminated image to the geometric feature detection model. The elimination processing unit 111 generating the feature detection image for object elimination includes the elimination processing unit 111 acquiring the feature detection image for object elimination output from the geometric feature detection model.
[0081] The elimination processing unit 111 generates a new object elimination image based on the feature detection image for object elimination, the original object elimination image which is the source of the feature detection image for object elimination, and the elimination target setting image (step S107). Hereinafter, the original object elimination image which is the source of the feature detection image for object elimination is also referred to as the original object elimination image. The newly generated object elimination image is also referred to as the new object elimination image. In step S107, for example, the elimination processing unit 111 uses an image generation model for object elimination to generate an object elimination image. Here, the elimination processing unit 111 inputs the feature detection image for object elimination and the original object elimination image into the image generation model for object elimination. The image generation model for object elimination refers to the feature detection image for object elimination and the elimination target setting image, generates a new object elimination image based on the original object elimination image, and outputs the new object elimination image. For example, the image generation model for object elimination inpaints the part of the object to be eliminated in the original object elimination image to generate an object elimination image. As described above, the image generation model for object elimination incorporates a machine learning model for fine-tuning which is better at object elimination. As described above, the image generation model for object elimination can obtain the effect of maintaining the shape of the region by using the machine learning model for fine-tuning. As described above, the image generation model for object elimination incorporates a model having a function of geometric feature detection as a control machine learning model. By using this control machine learning model, the image generation model for object elimination can control the geometric features by the feature detection image for object elimination and maintain the shape of the region. By combining the machine learning model for fine-tuning and this control machine learning model, the image generation model for object elimination can draw a correct boundary line for the part that should originally be a straight line by the machine learning model for fine-tuning and the part that is controlled to be a straight line by the control machine learning model. The image generation model for object elimination may also refer to the elimination target setting image and generate a new object elimination image based on the original object elimination image. The elimination processing unit 111 acquires the new object elimination image output from the image generation model for object elimination.The generation of a new object deletion image by the deletion processing unit 111 includes the deletion processing unit 111 inputting a feature detection image for object deletion and an original object deletion image into an image generation model for object deletion. The generation of an object deletion image by the deletion processing unit 111 includes the deletion processing unit 111 obtaining a new object deletion image output from the image generation model for object deletion. The deletion processing unit 111 stores the newly generated object deletion image in the image storage area 131. Note that the deletion processing unit 111 may generate a new object deletion image based on the feature detection image for object deletion and the original object deletion image without using the deletion target setting image. In this example, the image generation model for object deletion may refer to the feature detection image for object deletion and generate a new object arrangement image by i2i based on the original object deletion image.
[0082] As described above, based on the object detection in the deletion target image, the deletion processing unit 111 can generate a deletion target setting image and generate an object deletion image based on the deletion target image and the deletion target setting image. Through such processing, the deletion processing unit 111 can automatically generate an object deletion image without the need for the user to specify the object to be deleted in the deletion target image.
[0083] As described above, based on the geometric feature detection in the object deletion image, the deletion processing unit 111 can generate a feature detection image for object deletion and generate a new object deletion image based on at least the feature detection image for object deletion and the original object deletion image. In a typical example, the deletion processing unit 111 can generate a new object deletion image based on the feature detection image for object deletion, the original object deletion image, and the deletion target setting image. By using the feature detection image for object deletion, the deletion processing unit 111 can generate an object deletion image in which necessary furniture is deleted or unnecessary furniture does not remain while maintaining the shape of the area represented by the deletion target image.
[0084] As described above, the erasure processing unit 111 can generate an object erasure image using an image generation model for object erasure in which one or both of a machine learning model for fine-tuning and a machine learning model for control are incorporated. By using such an image generation model for object erasure, the erasure processing unit 111 can generate an object erasure image in which necessary furniture is erased or unnecessary furniture does not remain while maintaining the shape of the region represented by the erasure target image.
[0085] FIG. 4 is a flowchart showing an example of the arrangement processing by the server 1.
[0086] The arrangement processing unit 112 acquires the arrangement target image from the image storage area 131 (step S201).
[0087] The arrangement processing unit 112 generates a feature detection image for object arrangement based on the detection of geometric features in the arrangement target image (step S202). In step S202, for example, the arrangement processing unit 112 generates a feature detection image for object arrangement using a geometric feature detection model. Here, the arrangement processing unit 112 inputs the arrangement target image to the geometric feature detection model. The geometric feature detection model generates a feature detection image for object arrangement based on the arrangement target image and outputs the feature detection image for object arrangement. The arrangement processing unit 112 acquires the feature detection image for object arrangement output from the geometric feature detection model. The arrangement processing unit 112 generating a feature detection image for object arrangement includes the arrangement processing unit 112 inputting the arrangement target image to the geometric feature detection model. The arrangement processing unit 112 generating a feature detection image for object arrangement includes the arrangement processing unit 112 acquiring the feature detection image for object arrangement output from the geometric feature detection model.
[0088] The arrangement processing unit 112 determines whether the type of the object to be held is set (step S203). If the type of the object to be held is set (step S203, YES), the process transitions from step S203 to step S204. If the type of the object to be held is not set (step S203, NO), the process transitions from step S203 to step S206.
[0089] The arrangement processing unit 112 detects an object in the image to be arranged (step S204). In step S204, for example, the arrangement processing unit 112 uses an object detection model to detect an object in the image to be arranged. Here, the arrangement processing unit 112 inputs the image to be arranged into the object detection model. The object detection model generates a detection result of the object detected in the image to be arranged based on the image to be arranged and outputs the detection result. The arrangement processing unit 112 obtains the detection result output from the object detection model. The arrangement processing unit 112 detecting an object in the image to be arranged includes the arrangement processing unit 112 inputting the image to be arranged into the object detection model. The arrangement processing unit 112 detecting an object in the image to be arranged includes the arrangement processing unit 112 obtaining the detection result output from the object detection model.
[0090] The arrangement processing unit 112 generates a holding target setting image based on the object detection in the image to be arranged (step S205). In step S205, for example, the arrangement processing unit 112 draws a mask on one or more objects to be held in the image to be arranged based on the detection result in the image to be arranged. The arrangement processing unit 112 can draw a mask on one or more objects to be held in the image to be arranged based on the type of the object to be held set via the terminal 2.
[0091] The arrangement processing unit 112 determines whether to use a reference image for generating the object arrangement image (step S206). Whether to use a reference image may be set via the terminal 2. When the arrangement processing unit 112 determines to use a reference image (step S206, YES), the process transitions from step S206 to step S207. When the arrangement processing unit 112 determines not to use a reference image (step S206, NO), the process transitions from step S206 to step S209.
[0092] The arrangement processing unit 112 determines whether there is a first reference image (step S207). When there is a first reference image (step S207, YES), the process transitions from step S207 to step S209. When there is no first reference image (step S207, NO), the process transitions from step S207 to step S208.
[0093] The arrangement processing unit 112 generates a second reference image (step S208). In step S208, for example, the arrangement processing unit 112 generates a reference image using an image generation model for the reference image. Here, the arrangement processing unit 112 inputs text data to the image generation model for the reference image. The text data may be data based on an input operation of the terminal 2. The image generation model for the reference image generates a second reference image based on the text data and outputs the second reference image. For example, the image generation model for the reference image generates a second reference image by t2i based on the text data. As described above, the image generation model for the reference image incorporates a machine learning model for fine-tuning that is more proficient in object arrangement. The machine learning model for fine-tuning is generated by machine learning an image in which an object of the correct size is arranged in a region. Therefore, the image generation model for the reference image can generate a second reference image in which objects of the correct ratio are arranged in a region by using the machine learning model for fine-tuning. The arrangement processing unit 112 acquires the second reference image output from the image generation model for the reference image. The arrangement processing unit 112 generating the second reference image includes the arrangement processing unit 112 inputting text data to the image generation model for the reference image. The arrangement processing unit 112 generating the second reference image includes the arrangement processing unit 112 acquiring the second reference image output from the image generation model for the reference image. The arrangement processing unit 112 stores the second reference image in the image storage area 131.
[0094] The arrangement processing unit 112 generates an object arrangement image based on the feature detection image for object arrangement (step S209). In step S209, for example, the arrangement processing unit 112 generates an object arrangement image using an image generation model for object arrangement. Here, the arrangement processing unit 112 inputs the feature detection image for object arrangement into the image generation model for object arrangement. The image generation model for object arrangement refers to the feature detection image for object arrangement, generates an object arrangement image, and outputs the object arrangement image. The arrangement processing unit 112 acquires the object arrangement image output from the image generation model for object arrangement. The arrangement processing unit 112 generating the object arrangement image includes the arrangement processing unit 112 inputting the feature detection image for object arrangement into the image generation model for object arrangement. The arrangement processing unit 112 generating the object arrangement image includes the arrangement processing unit 112 acquiring the object arrangement image output from the image generation model for object arrangement. The arrangement processing unit 112 stores the object arrangement image in the image storage area 131.
[0095] The first arrangement mode is a mode that does not use the holding target setting image and the reference image. The arrangement processing unit 112 generates an object arrangement image based on the feature detection image for object arrangement and the text data. In the first arrangement mode, the arrangement processing unit 112 inputs the feature detection image for object arrangement and the text data into the image generation model for object arrangement. The text data may be data based on the input operation of the terminal 2. The image generation model for object arrangement refers to the feature detection image for object arrangement, generates an object arrangement image based on the text data, and outputs the object arrangement image. For example, the image generation model for object arrangement generates an object arrangement image by t2i based on the text data. The arrangement processing unit 112 acquires the object arrangement image output from the image generation model for object arrangement. The arrangement processing unit 112 generating the object arrangement image includes the arrangement processing unit 112 inputting the feature detection image for object arrangement and the text data into the image generation model for object arrangement. The arrangement processing unit 112 generating the object arrangement image includes the arrangement processing unit 112 acquiring the object arrangement image output from the image generation model for object arrangement.
[0096] The second arrangement mode does not use the image to be held as a setting image, but uses a reference image. The arrangement processing unit 112 generates an object arrangement image based on the feature detection image for object arrangement and the reference image. In the second arrangement mode, the arrangement processing unit 112 inputs the feature detection image for object arrangement and the reference image to the image generation model for object arrangement. The reference image may be the first reference image or the second reference image. The image generation model for object arrangement refers to the feature detection image for object arrangement, generates an object arrangement image based on the reference image, and outputs the object arrangement image. For example, the image generation model for object arrangement generates an object arrangement image by i2i based on the reference image. The arrangement processing unit 112 acquires the object arrangement image output from the image generation model for object arrangement. The arrangement processing unit 112 generating the object arrangement image includes the arrangement processing unit 112 inputting the feature detection image for object arrangement and the reference image to the image generation model for object arrangement. The arrangement processing unit 112 generating the object arrangement image includes the arrangement processing unit 112 acquiring the object arrangement image output from the image generation model for object arrangement. Note that the arrangement processing unit 112 may generate an object arrangement image based on text data in addition to the feature detection image for object arrangement and the reference image.
[0097] The third arrangement mode is a mode that uses the image to be held as a setting image but does not use a reference image. The arrangement processing unit 112 generates an object arrangement image based on the feature detection image for object arrangement and the image to be held as a setting image. In the third arrangement mode, the arrangement processing unit 112 inputs the feature detection image for object arrangement and the image to be held as a setting image to the image generation model for object arrangement. The image generation model for object arrangement refers to the feature detection image for object arrangement, generates an object arrangement image based on the image to be held as a setting image, and outputs the object arrangement image. For example, the image generation model for object arrangement performs inpainting on parts other than the object to be held in the image to be held as a setting image to generate an object arrangement image. The area represented by the object arrangement image is the area where the object to be held is arranged. The arrangement processing unit 112 acquires the object arrangement image output from the image generation model for object arrangement. The arrangement processing unit 112 generating the object arrangement image includes the arrangement processing unit 112 inputting the feature detection image for object arrangement and the image to be held as a setting image to the image generation model for object arrangement. The arrangement processing unit 112 generating the object arrangement image includes the arrangement processing unit 112 acquiring the object arrangement image output from the image generation model for object arrangement. Note that the arrangement processing unit 112 may generate an object arrangement image based on text data in addition to the feature detection image for object arrangement and the image to be held as a setting image.
[0098] As described above, the image generation model for object placement incorporates a machine learning model for fine-tuning that is more proficient in object placement. The machine learning model for fine-tuning is generated by machine learning an image in which an object of the correct size is placed in a region. Therefore, the image generation model for object placement can generate an object placement image in which objects of the correct ratio are placed in a region by using the machine learning model for fine-tuning. As described above, the image generation model for object placement incorporates, as a control machine learning model, a model having a function of geometric feature detection. By using this control machine learning model, the image generation model for object placement can control the geometric features in the feature detection image for object placement and maintain the shape of the region. As described above, the image generation model for object placement incorporates, as a control machine learning model, a model having a function of converting a region where no object is placed into a region where an object is placed. By using this control machine learning model, the image generation model for object placement becomes more likely to place an object.
[0099] As described above, the placement processing unit 112 can generate an object placement image based on the feature detection image for object placement. By using the feature detection image for object placement, the placement processing unit 112 can generate an object placement image in which a new object is placed in the placement target region while maintaining the shape of the region represented by the placement target image.
[0100] As described above, the placement processing unit 112 can generate an object placement image by using the image generation model for object placement that incorporates a machine learning model for fine-tuning that is more proficient in object placement. By using such an image generation model for object placement, the placement processing unit 112 can generate an object placement image in which objects of the correct ratio are placed in a region. For example, the placement processing unit 112 can prevent the generation of an image in which furniture outside common sense, such as furniture that is too large for the region, is placed.
[0101] As described above, the placement processing unit 112 can generate an object placement image using an image generation model for object placement in which a model having a function of detecting geometric features is incorporated as a machine learning model for control. By using such an image generation model for object placement, the placement processing unit 112 can control the geometric features based on the feature detection image for object placement and maintain the shape of the region.
[0102] As described above, the placement processing unit 112 can generate an object placement image using an image generation model for object placement in which a model having a function of converting a region where an object is not placed into a region where an object is placed is incorporated as a machine learning model for control. By using such an image generation model for object placement, the placement processing unit 112 makes it easier to place an object.
[0103] As described above, the placement processing unit 112 can generate an object placement image based on the feature detection image for object placement and text data. By using the text data, the placement processing unit 112 can generate an object placement image with high degrees of freedom and many variations of object placement.
[0104] As described above, the placement processing unit 112 can generate an object placement image based on the feature detection image for object placement and a reference image. By using the reference image, the placement processing unit 112 can easily generate an object placement image in which an object reflecting the preference of the reference image is placed.
[0105] As described above, the placement processing unit 112 can generate an object placement image based on the feature detection image for object placement and a retention target setting image. By using the retention target setting image, the placement processing unit 112 can generate an object placement image in which an object to be retained in the placement target image is retained.
[0106] FIG. 5 is a flowchart showing an example of the synthesis process by the server 1.
[0107] The synthesis processing unit 113 acquires the object arrangement image from the image storage area 131 (step S301).
[0108] The synthesis processing unit 113 detects the objects in the object arrangement image (step S302). In step S302, the synthesis processing unit 113 uses the object detection model to detect the objects in the object arrangement image. Here, the synthesis processing unit 113 inputs the object arrangement image to the object detection model. The object detection model generates a detection result of the objects detected in the object arrangement image based on the object arrangement image and outputs the detection result. The synthesis processing unit 113 acquires the detection result output from the object detection model. The synthesis processing unit 113 detecting the objects in the object arrangement image includes the synthesis processing unit 113 inputting the object arrangement image to the object detection model. The synthesis processing unit 113 detecting the objects in the object arrangement image includes the synthesis processing unit 113 acquiring the detection result output from the object detection model.
[0109] The synthesis processing unit 113 generates a synthesis target setting image based on the object detection in the object arrangement image (step S303). In step S303, for example, the synthesis processing unit 113 draws a mask on one or more objects to be synthesized in the object arrangement image based on the detection result in the object arrangement image. The synthesis processing unit 113 generates a synthesis target setting image based on the drawing of the mask.
[0110] The synthesis processing unit 113 generates a synthesized image based on the placement target image and the synthesis target setting image (step S304). In step S304, for example, the synthesis processing unit 113 synthesizes the objects to be synthesized with masks drawn in the synthesis target setting image into the placement target image. The synthesis processing unit 113 generates a synthesized image based on the synthesis. The synthesis processing unit 113 stores the synthesized image in the image storage area 131.
[0111] The synthesis processing unit 113 generates an edge detection image based on edge detection in the synthesized image (step S305). In step S305, for example, the synthesis processing unit 113 uses an edge detection model to generate an edge detection image. Here, the synthesis processing unit 113 inputs the synthesized image into the edge detection model. The edge detection model generates an edge detection image based on the synthesized image and outputs the edge detection image. The synthesis processing unit 113 acquires the edge detection image output from the edge detection model. The synthesis processing unit 113 generating the edge detection image includes the synthesis processing unit 113 inputting the synthesized image into the edge detection model. The synthesis processing unit 113 generating the edge detection image includes the synthesis processing unit 113 acquiring the edge detection image output from the edge detection model.
[0112] The synthesis processing unit 113 detects an object in the synthesized image (step S306). In step S306, the synthesis processing unit 113 uses an object detection model to detect an object in the synthesized image. Here, the synthesis processing unit 113 inputs the synthesized image into the object detection model. The object detection model generates a detection result of the object detected in the synthesized image based on the synthesized image and outputs the detection result. The synthesis processing unit 113 acquires the detection result output from the object detection model. The synthesis processing unit 113 detecting an object in the synthesized image includes the synthesis processing unit 113 inputting the synthesized image into the object detection model. The synthesis processing unit 113 detecting an object in the synthesized image includes the synthesis processing unit 113 acquiring the detection result output from the object detection model. The synthesis processing unit 113 generates an outer edge setting image based on object detection in the synthesized image (step S307). In step S307, for example, the synthesis processing unit 113 draws a mask on the object synthesized in the synthesized image based on the detection result in the synthesized image. The synthesis processing unit 113 generates an outer edge setting image based on the drawing of the mask.
[0113] The synthesis processing unit 113 generates an adaptation image based on the edge detection image and the outer edge setting image (step S308). In step S308, for example, the synthesis processing unit 113 generates an adaptation image using an image generation model for synthesis. Here, the synthesis processing unit 113 inputs the edge detection image and the outer edge setting image into the image generation model for synthesis. The image generation model for synthesis refers to the edge detection image, generates an adaptation image based on the outer edge setting image, and outputs the adaptation image. For example, the image generation model for synthesis paints in the portion of the outer edge of the object where the mask is drawn in the outer edge setting image to generate an adaptation image. As described above, the image generation model for synthesis incorporates a machine learning model for fine-tuning that is more proficient in object placement. The machine learning model for fine-tuning is generated by machine learning an image in which an object of the correct size is placed in a region. Therefore, the image generation model for synthesis can generate an adaptation image in which an object of the correct ratio is placed in a region by using the machine learning model for fine-tuning. As described above, a model having an edge detection function is incorporated as a control machine learning model in the image generation model for synthesis. The image generation model for synthesis can control the edges in the edge detection image and maintain the shape of the object by using this control machine learning model. As described above, a model having a function of converting from a region where no object is placed to a region where an object is placed is incorporated as a control machine learning model in the image generation model for synthesis. The image generation model for synthesis becomes more likely to place an object by using this control machine learning model. The synthesis processing unit 113 acquires the adaptation image output from the image generation model for synthesis. The synthesis processing unit 113 generating an adaptation image includes the synthesis processing unit 113 inputting the edge detection image and the outer edge setting image into the image generation model for synthesis. The synthesis processing unit 113 generating an adaptation image includes the synthesis processing unit 113 acquiring the adaptation image output from the image generation model for synthesis. The synthesis processing unit 113 stores the adaptation image in the image storage area 131.
[0114] As described above, the composition processing unit 113 can generate a composite image based on the image to be arranged and the composition target setting image. Through such processing, the composition processing unit 113 can generate a composite image that arranges objects not arranged in the image to be arranged while maintaining the color and pattern of the floor, walls, ceiling, etc., the detailed shape of windows, etc., and other facilities that should not be erased in the image to be arranged.
[0115] As described above, the composition processing unit 113 can generate a blending image based on the edge detection image and the outer edge setting image. By using the edge detection image, the composition processing unit 113 can generate a blending image with smoothed boundaries while maintaining the shape of the object represented by the edge detection image.
[0116] As described above, the composition processing unit 113 can generate a blending image by using an image generation model for composition in which a machine learning model for fine-tuning, which is more proficient in object placement, is incorporated. By using such an image generation model for composition, the composition processing unit 113 can generate a blending image in which objects are arranged in the correct ratio for the region.
[0117] As described above, the composition processing unit 113 can generate a blending image by using an image generation model for composition in which a model having an edge detection function is incorporated as a machine learning model for control. By using such an image generation model for composition, the composition processing unit 113 can control the edges in the edge detection image and maintain the shape of the object.
[0118] As described above, the composition processing unit 113 can generate a blending image by using an image generation model for composition in which a model having a function of converting from a region where no object is arranged to a region where an object is arranged is incorporated as a machine learning model for control. By using such an image generation model for synthesis, the synthesis processing unit 113 makes it easier to place objects.
[0119] (Image example) An example of an image displayed on the display device of the terminal 2 will be described.
[0120] FIG. 6 is a diagram showing an example of an image to be erased. The image to be erased is a photographic image representing an area where an object such as furniture is placed.
[0121] FIG. 7 is a diagram showing an example of an image for setting an object to be erased. In the image for setting an object to be erased, a mask is drawn on one or more objects to be erased.
[0122] FIG. 8 is a diagram showing an example of an object-erased image. In the object-erased image, the object to be erased has been erased. Here, the object-erased image is assumed to correspond to the image for placement target.
[0123] FIG. 9 is a diagram showing an example of a feature detection image for object placement. In the feature detection image for object placement, the straight lines detected from the image for placement target are shown.
[0124] Although not shown, the feature detection image for object erasure is an image in which the straight lines detected from the object-erased image are shown, similar to the feature detection image for object placement.
[0125] FIG. 10 is a diagram showing an example of a reference image. The reference image is an image representing an area where an object such as furniture different from the image to be erased is placed.
[0126] FIG. 11 is a diagram showing an example of an image for setting an object to be retained. In the image for setting an object to be retained, a mask is drawn on one or more objects to be retained.
[0127] FIG. 12 is a diagram showing an example of an object arrangement image. In the object arrangement image, one or more objects that are not arranged in the arrangement target image such as a sofa are arranged.
[0128] FIG. 13 is a diagram showing an example of an image to be synthesized. In the image to be synthesized, masks are drawn on one or more objects to be synthesized. The objects with masks drawn are objects such as sofas newly arranged in the object arrangement image.
[0129] FIG. 14 is a diagram showing an example of a synthesized image. The image to be synthesized is an image obtained by synthesizing the object to be synthesized into the arrangement target image. Therefore, in the synthesized image, objects such as lighting fixtures, floors, and windows other than the objects to be synthesized are the same as those in the arrangement target image.
[0130] FIG. 15 is a diagram showing an example of an edge detection image. In the edge detection image, edges detected from the synthesized image are shown.
[0131] FIG. 16 is a diagram showing an example of an outer edge setting image. In the outer edge setting image, masks are drawn on the outer edges of one or more objects such as sofas synthesized into the arrangement target image to generate the synthesized image.
[0132] FIG. 17 is a diagram showing an example of an image for adaptation. In the image for adaptation, the color of the boundary between one or more objects such as sofas synthesized into the arrangement target image to generate the synthesized image and other objects is adapted to a natural state.
[0133] According to the embodiment, the server 1 can easily generate a new image obtained by eliminating the objects arranged in the elimination target image while maintaining the shape of the region represented by the elimination target image, such as the object elimination image. According to the embodiment, the server 1 can generate a new image in which an object suitable for a region represented by an arrangement target image is arranged while maintaining the shape of the region represented by the arrangement target image, such as an object arrangement image, a composite image, or a familiarization image.
[0134] [Other Embodiments] In the above-described embodiment, an example in which the auxiliary storage device 13 of the server 1 stores each image has been described, but the present invention is not limited to this. A server different from the server 1 may store each image instead of the server 1. A plurality of servers different from the server 1 may store each image in a distributed manner instead of the server 1.
[0135] In the above-described embodiment, an example in which the auxiliary storage device 13 of the server 1 stores each machine learning model has been described, but the present invention is not limited to this. A server different from the server 1 may store each machine learning model instead of the server 1. A plurality of servers different from the server 1 may store each machine learning model in a distributed manner instead of the server 1.
[0136] The image processing apparatus has been described by taking the server 1 as an example, but the present invention is not limited to this. The image processing apparatus may be realized by a device having the same functions as the server 1. The device may be a terminal such as a PC (Personal Computer), a smartphone, or a tablet terminal.
[0137] The image processing apparatus may be realized by one device such as the server 1 described in the above-described embodiment, or may be realized by a plurality of devices in which functions are distributed.
[0138] The above-described embodiment may be applied not only to the apparatus but also to a method executed by the apparatus. The above-described embodiment may be applied to a program that enables each function to be executed by a computer of the apparatus. A program that enables each function to be executed by a computer of the apparatus is a program that enables the processing of each part provided in the apparatus to be executed by the computer of the apparatus. The above-described embodiment may be applied to a recording medium that stores the program.
[0139] Each of the one or more circuits constituting the processing circuit executes one or more of a plurality of processes. When the processing circuit is constituted by a single circuit, the single circuit executes all of the plurality of processes. When the processing circuit is constituted by a plurality of circuits, each of the plurality of circuits executes a part of the plurality of processes. A part of the plurality of processes may be one of the plurality of processes or two or more of the plurality of processes. When the processing circuit is constituted by a plurality of circuits, the plurality of circuits may be included in one device or may be distributed among a plurality of devices.
[0140] The program may be transferred in a state stored in the device according to the embodiment, or may be transferred in a state not stored in the device. In the latter case, the program may be transferred via a network or may be transferred in a state recorded on a recording medium. The recording medium is a non-transitory tangible medium. The recording medium is a computer-readable medium. The recording medium may be any medium that can store a program and is readable by a computer, such as a CD-ROM or a memory card, regardless of its form.
[0141] In short, the present invention is not limited to the present embodiment as it is, and at the implementation stage, the components can be modified and embodied without departing from the gist thereof. Also, various inventions can be formed by appropriately combining a plurality of components disclosed in the present embodiment. For example, some components may be deleted from all the components shown in the present embodiment. Furthermore, components from different embodiments may be appropriately combined.
[0142] Some of the above embodiments may be expressed as follows. [1] An image of an object to be arranged, generating a feature detection image for object arrangement based on detection of geometric features in the arrangement target image representing a region, generating an object arrangement image representing a region where an object not arranged in the arrangement target image is arranged, based on the feature detection image for object arrangement, An image processing apparatus including an arrangement processing unit. [2] The image processing apparatus according to [1], wherein the arrangement processing unit generates the object arrangement image based on the feature detection image for object arrangement and the text data. [3] The image processing apparatus according to [1], wherein the arrangement processing unit generates the object arrangement image based on the feature detection image for object arrangement and a reference image representing an area where an object is arranged. [4] The arrangement processing unit sets an object to be held in the arrangement target image based on object detection in the arrangement target image, generates the object arrangement image based on the feature detection image for object arrangement and the setting of the object to be held, and an area represented by the object arrangement image is an area where the object to be held is arranged. The image processing apparatus according to [1]. [5] Based on object detection in the object arrangement image, an object to be synthesized is set in the object arrangement image, and a composite image in which the object to be synthesized is synthesized into the arrangement target image is generated based on the arrangement target image and the setting of the object to be synthesized. The image processing apparatus according to [1], further comprising a composite processing unit. [6] The composite processing unit generates an edge detection image based on edge detection in the composite image, sets an outer edge of the object synthesized in the composite image, and generates an image in which the outer edge is adapted in the composite image based on the edge detection image and the setting of the outer edge. The image processing apparatus according to [5]. [7] The arrangement processing unit uses an image generation model for object arrangement incorporated with a machine learning model for fine-tuning generated by machine learning of learning data including a plurality of images representing an area where an object is arranged, to generate the object arrangement image. The image processing apparatus according to [1]. [8] The placement processing unit uses an image generation model for object placement incorporated with a machine learning model for control that outputs an image representing the area where the object is placed based on an image representing the area where the input object is not placed, to generate the object placement image. The image processing apparatus according to [1]. [9] For an image of an object to be erased, based on object detection in an erasure target image representing the area where the object is placed, the object to be erased in the erasure target image is set. Based on the erasure target image and the setting of the object to be erased, an object erasure image is generated by erasing the object to be erased from the erasure target image. Further comprising an erasure processing unit. The placement target image is the object erasure image generated by the erasure processing unit. The image processing apparatus according to any one of [1] to [8].
[10] The erasure processing unit generates a feature detection image for object erasure based on detection of geometric features in the object erasure image. Based on the feature detection image for object erasure and the object erasure image that is the source of the generation of the feature detection image for object erasure, a new object erasure image is generated. The image processing apparatus according to [9].
[11] The erasure processing unit uses an image generation model for object erasure incorporated with a machine learning model for fine-tuning generated by machine learning of learning data including a plurality of images representing the area where the object is not placed, to generate the object erasure image. The image processing apparatus according to [9].
[12] For an image of an object to be erased, based on object detection in an erasure target image representing the area where the object is placed, the object to be erased in the erasure target image is set. Based on the erasure target image and the setting of the object to be erased, an object erasure image is generated by erasing the object to be erased from the erasure target image. Based on detection of geometric features in the generated object erasure image, a feature detection image for object erasure is generated. Generate a new object elimination image based on the feature detection image for object elimination and the object elimination image that is the source of the feature detection image for object elimination. An image processing apparatus including an elimination processing unit.
[13] An image processing method executed by an image processing apparatus, For an image of an object placement target, generate a feature detection image for object placement based on geometric feature detection in the placement target image representing a region, Based on the feature detection image for object placement, generate an object placement image representing a region where an object not placed in the placement target image is placed. An image processing method including the above.
[14] An image processing method executed by an image processing apparatus, For an image of an object elimination target, set an object to be eliminated in the elimination target image based on object detection in the elimination target image representing a region where an object is placed. Based on the elimination target image and the setting of the object to be eliminated, generate an object elimination image obtained by eliminating the object to be eliminated from the elimination target image. Generate a feature detection image for object elimination based on geometric feature detection in the generated object elimination image. Generate a new object elimination image based on the feature detection image for object elimination and the object elimination image that is the source of the feature detection image for object elimination. An image processing method including the above.
[15] An image processing program capable of causing a computer to execute the processing of each part included in the image processing apparatus according to any one of [1] to
[12] .
Explanation of Signs
[0143] 1... Server, 2... Terminal, 11... Processing circuit, 12... Main memory, 13... Auxiliary storage device, 14... Communication interface, 111... Elimination processing unit, 112... Placement processing unit, 113... Composition processing unit, 131... Image storage area, 132... Machine learning model storage area, S... Image processing system.
Claims
1. An image of an object placement target, which generates a feature detection image for object placement based on the detection of geometric features in the placement target image representing a region, and generates an object placement image representing a region where an object not placed in the placement target image is placed, based on the feature detection image for object placement. An image processing apparatus comprising a placement processing unit.
2. The placement processing unit generates the object placement image based on the feature detection image for object placement and text data. The image processing apparatus according to Claim 1.
3. The placement processing unit generates the object placement image based on the feature detection image for object placement and a reference image representing a region where an object is placed. The image processing apparatus according to Claim 1.
4. The placement processing unit sets an object to be held in the placement target image based on object detection in the placement target image, generates the object placement image based on the feature detection image for object placement and the setting of the object to be held, and a region represented by the object placement image is a region where the object to be held is placed. The image processing apparatus according to Claim 1.
5. Based on object detection in the object placement image, an object to be synthesized in the object placement image is set, and a composite image in which the object to be synthesized is synthesized into the placement target image is generated based on the placement target image and the setting of the object to be synthesized. The image processing apparatus according to Claim 1, further comprising a composite processing unit.
6. The composite processing unit generates an edge detection image based on edge detection in the composite image, sets an outer edge of an object synthesized in the composite image, and generates an image in which the outer edge is blended in the composite image based on the edge detection image and the setting of the outer edge. The image processing apparatus according to Claim 5.
7. The placement processing unit uses an object placement image generation model incorporated with a machine learning model for fine-tuning generated by machine learning of learning data including a plurality of images representing regions where objects are placed, to generate the object placement image. The image processing apparatus according to Claim 1.
8. The placement processing unit uses an object placement image generation model incorporated with a machine learning model for control that outputs an image representing a region where an object is placed based on an input image representing a region where an object is not placed, to generate the object placement image. The image processing apparatus according to claim 1.
9. An image of an object to be erased, based on object detection in an erasure target image representing an area where the object is placed, an object to be erased in the erasure target image is set, Based on the erasure target image and the setting of the object to be erased, an object erasure image obtained by erasing the object to be erased from the erasure target image is generated. Further comprising an erasure processing unit, The placement target image is the object erasure image generated by the erasure processing unit. The image processing apparatus according to claim 1.
10. The erasure processing unit Based on detection of geometric features in the generated object erasure image, a feature detection image for object erasure is generated, Based on the feature detection image for object erasure and the object erasure image that is the source of the generation of the feature detection image for object erasure, a new object erasure image is generated. The image processing apparatus according to claim 9.
11. The erasure processing unit uses an image generation model for object erasure incorporated with a machine learning model for fine-tuning generated by machine learning of learning data including a plurality of images representing areas where no object is placed, to generate the object erasure image. The image processing apparatus according to claim 9.
12. An image processing method executed by an image processing apparatus, Generating a feature detection image for object placement based on detection of geometric features in a placement target image that is an image of an object to be placed and represents an area; Based on the feature detection image for object placement, generating an object placement image representing an area where an object not placed in the placement target image is placed; An image processing method comprising the above.
13. An image processing program capable of causing a computer to execute the processing of each part provided in the image processing apparatus according to any one of claims 1 to 11.
Citation Information
Patent Citations
Method for specifying coordinate axis in three dimensional space and method for specifying plane
JP2021051660A
Display device, method, and program
WO2020054203A1
Image processing system, image processing method and program
JP2021149679A