Image processing device, image processing method and program
The image processing device uses AI to transform still images from smartphones into virtual indoor spaces, addressing the limitations of existing systems by enabling flexible object replacement and efficient home staging for real estate promotion.
Patent Information
- Application Number
- JP2024013533
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-01-31
- Publication Date
- 2025-08-13
AI Technical Summary
Existing image processing systems for home staging require imaging devices that capture images of indoor spaces in all directions and fail to accurately incorporate virtual objects into areas with existing furniture, especially in occupied properties.
An image processing device that utilizes a widely available imaging device like a smartphone to capture still images of indoor spaces with or without objects, employing AI models to generate masked images and replace or interpolate objects with virtual ones based on user input and space attributes.
Enables effective home staging by quickly generating virtual indoor space images that match desired designs, allowing real estate sales without displacing tenants, and providing flexible object replacement options.
Smart Images

Figure 2025118295000001_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to an image processing device, an image processing method, and an image processing program. [Background technology]
[0002] In recent years, image processing techniques for home staging have been proposed as a technique for promoting the sale of real estate. For example, Patent Document 1 discloses an image processing system including a structure estimation means, an area estimation means, and an image processing means. The structure estimation means estimates the structure of a space from a background image showing the interior space of a structure in all directions. The area estimation means estimates an area in the space where a virtual object can be placed based on the estimated structure. The image processing means composites the virtual object into the estimated area of the background image. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Publication No. 2022-77148 Summary of the Invention [Problem to be solved by the invention]
[0004] However, the invention described in Patent Document 1 requires an imaging device that captures images of the space inside a structure, i.e., the indoor space, in all directions. Also, in the case of a property that currently has tenants, for example, objects such as furniture are arranged in the indoor space, and the image generated by the invention described in Patent Document 1 only has virtual objects (virtual targets) arranged in areas where no objects are arranged.
[0005] Therefore, the present invention aims to provide an image processing device, an image processing method, and an image processing program that perform home staging using two-dimensional still images captured using widely available imaging devices such as those installed in smartphones, and still images captured of indoor spaces in which objects are installed. [Means for solving the problem]
[0006] An image processing device of the present invention includes an acquisition unit, a masking processing unit, and a designated image generation unit. The acquisition unit acquires, as a first target image, a still image captured of an indoor space in which an object is installed. The masking processing unit generates a masked image by masking either an installed object, which is a subject image of the object, or an interior object, which is a subject image of the interior of the indoor space, based on a display request requesting display of the selected one of image objects, which are subject images of the first target image. When the designated image generation unit acquires designation information about the type and attributes of the indoor space and an instruction to generate an image, the designated image generation unit generates, as a first designated image, an image in which a first virtual object corresponding to the selected one is provided based on the designation information, using a first learning model that has learned from still images of the indoor space in which the object is installed.
[0007] When a masking processing unit receives a request to select at least one image object from the plurality of masked image objects and to unmask the image object, the masking processing unit preferably generates a masked image in which the masking of the selected image object is unmasked.
[0008] It is preferable that the designated image generating section generates a first designated image in which the first virtual object is provided in the masked area of the masking image.
[0009] It is preferable that the designated image generating section interpolates a non-formed area of the masked masking area where no first virtual object is provided with an image object to generate the first designated image.
[0010] The acquisition unit may acquire, as the second target image, a still image of an indoor space in which no object is installed. In this case, when acquiring the designation information and the image generation instruction, the designated image generation unit estimates the type and size of the indoor space shown in the second target image using a second learning model that has learned learning data including a still image of the indoor space and a content label that indicates the content of the subject image of the still image, and generates, as the second designated image, an image in which the second virtual object is provided based on the designation information.
[0011] Preferably, the indoor space is a room, the type is a classification according to use, and the attribute is an external appearance feature.
[0012] It is preferable that the acquisition unit acquires a first object image having a floor, a ceiling, and at least two adjacent walls as subject images, and it is preferable that the first learning model is one that has learned from a still image having a floor, a ceiling, and at least two adjacent walls as subject images.
[0013] The image processing method of the present invention includes an acquisition step, a masking processing step, and a designated image generation step. The acquisition step acquires a still image of an indoor space in which an object is installed as a first target image. The masking processing step generates a masked image by masking an installed object, which is a subject image of the object, or an interior object, which is a subject image of the interior of the indoor space, based on a display request for selecting and displaying one of the image objects, which are subject images of the first target image. The designated image generation step, when receiving designation information about the type and attributes of the indoor space and an instruction to generate an image, generates, as the first designated image, an image in which a first virtual object corresponding to the selected one is provided based on the designation information, using a first learning model that has learned from a still image of the indoor space in which the object is installed.
[0014] The image processing program of the present invention causes a computer to execute the above steps. [Effects of the Invention]
[0015] According to the present invention, home staging can be performed using still images of an indoor space in which an object is installed, captured using a widely available imaging device such as one installed in a smartphone. [Brief explanation of the drawings]
[0016] [Figure 1] 1 is a schematic diagram of an image processing system according to an embodiment. [Figure 2] FIG. 10 is an explanatory diagram of an image obtained using the image processing device. [Figure 3] FIG. 10 is an explanatory diagram of an image obtained using the image processing device. [Figure 4] FIG. 10 is an explanatory diagram of an image obtained using the image processing device. [Figure 5] FIG. 10 is an explanatory diagram of an image obtained using the image processing device. [Figure 6] FIG. 1 is a configuration diagram of an image processing device. [Figure 7] FIG. 10 is an explanatory diagram of a display image displayed on the client terminal. [Figure 8] FIG. 10 is an explanatory diagram of a display image displayed on the client terminal. [Figure 9] FIG. 10 is an explanatory diagram of a display image displayed on the client terminal. [Figure 10] FIG. 10 is an explanatory diagram of a display image displayed on the client terminal. [Figure 11] FIG. 10 is an explanatory diagram of a display image displayed on the client terminal. [Figure 12] FIG. 10 is an explanatory diagram of a display image displayed on the client terminal. DETAILED DESCRIPTION OF THE INVENTION
[0017] The image processing system 10 shown in FIG. 1 performs image processing on still images captured of an indoor space of a real estate property, generates an image showing a virtual indoor space, and displays the image on a display unit (display) of a predetermined terminal. The still images are captured using a commonly used imaging device such as that installed in a smartphone or digital camera. The image processing system 10 includes a client terminal 11, a management terminal 12, and an image processing device 13. The client terminal 11 and the image processing device 13 communicate with each other via a communication network CN, i.e., transmit and receive various data. For example, the client terminal 11 transmits a still image captured of an indoor space of a real estate property to the image processing device 13, and the image processing device 13 performs predetermined processing on the received still image to generate an image showing a virtual indoor space and transmits the image to the client terminal 11. As a result, the generated image is displayed on the display unit of the client terminal 11, allowing a client using the client terminal 11 to grasp the virtual indoor space through the image.
[0018] The client terminal 11 is an example of the terminal used by a client that causes the image processing device 13 to perform various processes. The client is, for example, a real estate sales company that sells real estate or a brokerage company that mediates sales and purchases, and the client terminal 11 is used by employees of the company. Other examples of clients include, but are not limited to, real estate management companies and real estate lessors (landlords). The client terminal 11 is a computer that includes the display unit and an input unit (keyboard, mouse, touch panel display, etc.) for performing various input operations such as a transmission operation for transmitting the still image and a display request operation for requesting the image processing device 13 to display a predetermined image, and may be any of a personal computer, a mobile terminal, a smartphone, etc., and is not particularly limited.
[0019] The client terminal 11 may operate on a browser by receiving a program that runs on the browser from the image processing device 13, or may be embedded with predetermined application software and operate by executing the program of this application software. Although only one client terminal 11 is depicted in Fig. 1, there may be multiple client terminals 11 so that multiple clients each use one terminal.
[0020] The management terminal 12 is a terminal device that manages various settings and / or data of the image processing device 13. Like the client terminal 11, the management terminal 12 is a computer that includes an input unit for input operations and a display unit.
[0021] The image processing device 13 performs predetermined processing in response to input of various pieces of information from the client terminal 11 and the management terminal 12, respectively.
[0022] The image processing device 13 is configured with a computer. A predetermined program is installed in the image processing device 13, and by executing this program, the image processing device 13 functions as each unit described below and performs predetermined processing. The program causes the computer to execute an acquisition step, a masking processing step, and a designated image generation step. The acquisition step acquires a still image of an indoor space in which an object is installed as a first target image. The masking processing step generates a masked image by masking one of the image objects that are subject images of the first target image, based on a display request requesting display of either an installation object that is a subject image of the object in the first target image or an interior object that is a subject image of the interior of the indoor space. The designated image generation step, when receiving specification information on the type and attributes of the indoor space and an image generation instruction, generates, as the first designated image, an image in which a first virtual object corresponding to the selected object is provided based on the specification information, using a first learning model that has learned a still image of the indoor space in which the object is installed. Note that the program may be installed in the client terminal 11, causing the client terminal 11 to function as the image processing device 13.
[0023] In the image processing system 10, a client uses the client terminal 11 to transmit a still image G1, captured of an indoor space with furniture, home appliances, and other objects installed, as the first target image described above, to the image processing device 13. The example still image G1 shown in FIG. 2 includes subject images SI1 and SI2 of a table and two chairs as furniture. The table in subject image SI1 is made of wood, and the chair in subject image SI2 has a curved shape. In order to change the design of the indoor space, the client transmits a generation instruction to the image processing device 13 to generate an image of a virtual indoor space in which the table and the two chairs have been changed to match the changed design of the indoor space. In response to this generation instruction, the image processing device 13 generates an image G2, as the first specified image described above, in which the subject images of the table and chairs have been changed to virtual objects of a different table and chair, as shown in FIG. 3. In the example of image G2 shown in FIG. 3, a virtual object VO1 of a glass table and a virtual object VO2 of a chair with a linear shape are provided. The image processing device 13 transmits the generated image G2 to the client terminal 11, and the client visually recognizes the generated image G2 on the display unit of the client terminal 11 and understands the virtual indoor space. Similarly, instead of the above-mentioned objects such as home appliances and furniture, it is also possible to generate virtual images of interior decorations such as walls, ceilings, and floors of the indoor space that match the design of the changed indoor space.
[0024] Furthermore, in the image processing system 10, a client can use the client terminal 11 to transmit a still image G3, as shown in FIG. 4 , captured of an indoor space where no objects such as furniture or home appliances are installed, as a second target image to the image processing device 13. The client transmits a generation instruction to the image processing device 13 to generate an image of a virtual indoor space that matches the desired design of the indoor space. In response to this generation instruction, the image processing device 13 generates an image G4 with virtual objects, as shown in FIG. 5. In the example of image G4 shown in FIG. 5, a virtual curtain object VO3, a virtual sofa object VO4, a virtual cushion object VO5, a virtual table object VO6, a virtual plant object VO7, a virtual television object VO8, and a virtual cupboard object VO9 are provided. The image processing device 13 transmits the generated image G4 to the client terminal 11 as a second specified image, and the client visually recognizes the generated image on the display unit of the client terminal 11 to grasp the virtual indoor space.
[0025] The image processing device 13 will be described with reference to Fig. 6. The image processing device 13 includes a control unit 20, an acquisition unit 21, a masking processing unit 22, and a designated image generation unit 23. The control unit 20 comprehensively controls each unit in the image processing device 13, such as the acquisition unit 21 and the masking processing unit 22.
[0026] The image processing device 13 includes a first AI (Artificial Intelligence) to a third AI, and incorporates a first learning model M1 and a second learning model M2. The first learning model M1 is composed of an installed object learning model M1a and an interior object learning model M1b. The first AI constitutes a masking processing unit 22 and a designated image generating unit 23, the second AI similarly constitutes a masking processing unit 22 and a designated image generating unit 23, and the third AI constitutes the designated image generating unit 23. The masking processing unit 22 and the designated image generating unit 23 perform each process using an AI corresponding to the process to be performed.
[0027] When the control unit 20 receives a first target image to be processed from the client terminal 11, it sends it to the acquisition unit 21, and when it receives a second target image, it sends it to the designated image generation unit 23. A menu screen is displayed on the display unit of the client terminal 11 when the client logs in to the image processing device 13, and this menu screen has a selection button for selecting whether the still image to be sent to the image processing device 13 is the first target image or the second target image. Based on an input operation on this selection button, the client terminal 11 sends discrimination information indicating whether the image is the first target image or the second target image to the control unit 20, and then sends a still image corresponding to the selected one. In this way, the control unit 20 sends the discriminated first target image and second target image to the acquisition unit 21.
[0028] As described above, the acquisition unit 21 acquires a still image of an indoor space in which objects such as furniture and home appliances are installed as a first target image to be subjected to image processing, and outputs this first target image to the masking processing unit 22. It is preferable that the first target image include subject images of the floor and ceiling of the indoor space and at least two adjacent walls, from the viewpoint of more appropriate processing in the masking processing unit 22 and the designated image generation unit 23. The masking processing unit 22 more accurately detects each of the subject images of objects installed in the indoor space, such as furniture and home appliances, and the interior objects of the indoor space, such as the ceiling, floor, walls, and air conditioner, and masks them as described below. Furthermore, the designated image generation unit 23 generates an image in which each of these detected subject images is more accurately changed to a desired virtual object.
[0029] Acquisition unit 21 further acquires a still image of an indoor space without any objects such as furniture or home appliances as a second target image to be processed, and outputs this second target image to designated image generation unit 23. It is preferable that the second target image also include subject images of the floor and ceiling of the indoor space and at least two adjacent walls, from the viewpoint of more appropriate processing in masking processing unit 22 and designated image generation unit 23. Masking processing unit 22 more accurately detects subject images of the interior of the indoor space, such as the ceiling, floor, walls, and air conditioner, individually, and detects the size and shape of the indoor space more accurately. Furthermore, designated image generation unit 23 generates an image in which the desired virtual object is more accurately provided for each of these detected subject images.
[0030] The size (data capacity) of the first target image and the second target image is preferably at most 10 MB (megabytes), compared to a size exceeding 10 MB, from the viewpoint of further shortening the time required to transmit an image with a virtual object provided to the client terminal 11. The size (data capacity) of the first target image and the second target image is preferably at least 500 KB (kilobytes), compared to a size smaller than 500 KB, from the viewpoint of image recognition accuracy.
[0031] When a first target image is input, and a display request is input from client terminal 11 via control unit 20 requesting the selection and display of either an installation object or an interior object among the image objects that are subject images of the first target image, masking processing unit 22 generates a masked image by masking the selected object based on the display request and sends the masked image to client terminal 11. Installation objects are subject images of objects installed in an indoor space, such as home appliances, furniture, curtains, and ornaments, and interior objects are subject images of interior objects, which are indoor decorations and equipment, such as floors, ceilings, walls, air conditioners, ventilation fans, and vents in an indoor space. The initial setting (default) may be a state in which either an installation object or an interior object is selected, and a masked image of the selected object may be sent to client terminal 11. In this case, when a switching request from client terminal 11 is input to masking processing unit 22 via control unit 20, the switching request is treated as a display request requesting the selection and display of the other object, and a masked image of the other object is generated and sent to client terminal 11. Furthermore, when a request to switch from the other side to the one side is input, the masking processing unit 22 similarly treats this request as a display request and performs the same processing. In this example, the initial setting is a state in which the installed object is selected. The initial setting may be incorporated into the aforementioned program.
[0032] When an installation object is selected, the masking processing unit 22 performs a masking process using the installation object learning model M1a with the first AI. The installation object learning model M1a is a trained model that is pre-trained to input a still image containing an installation object, detect the installation object from the input still image, identify the type of the installation object, and mask it, as well as identify the type and attributes of the indoor space. That is, the model is pre-trained using a still image containing an installation object as input parameters, and the installation objects in the input still image and their respective types, an image in a masked state, and the type and attributes of the indoor space as output parameters. The type of the installation object can be identified and masked by trimming the area and training it in association with the type of the installation object. The first AI uses the installation object learning model M1a to identify each installation object in the input first target image and perform a masking process. In this example, the first AI uses TENSORFLOW (registered trademark) or keras.
[0033] When an interior object is selected, the masking processing unit 22 performs a masking process using the interior object learning model M1b with the second AI. The interior object learning model M1b is a trained model that is pre-trained to input a still image containing an installed object, detect the interior object from the input still image, identify the type of the interior object, and mask it, as well as identify the type and attributes of the indoor space. As with the case of an installed object, the type of the interior object can be identified and masked by trimming the area and learning it in association with the type of the interior object. The second AI uses the interior object learning model M1b to identify each interior object in the input first target image and perform a masking process. The above group information is input using the management terminal 12 (see Figure 1). In this example, the second AI uses TENSORFLOW or keras.
[0034] The still images used as input information for generating the installation object learning model M1a and the interior object learning model M1b, like the first and second target images, are images captured by a commonly used imaging device such as one installed in a smartphone or digital camera. Similarly to the sizes of the first and second target images, the size (data capacity) of the still images is preferably at most 10 MB (megabytes), compared to a size exceeding 10 MB, from the viewpoint of shortening the time required to transmit an image with a virtual object to the client terminal 11. Furthermore, a size of at least 500 KB (kilobytes) is preferable from the viewpoint of image recognition accuracy, compared to a size less than 500 KB. Furthermore, it is preferable that the still images used as input information include subject images of the floor and ceiling of an indoor space and at least two adjacent walls, from the viewpoint of more appropriate processing by the masking processing unit 22 and the designated image generating unit 23.
[0035] Masking is a process of making the area inside the outline of the subject image a single color as a target area so as to hide the subject image of the target. The color of the masking is not limited, and in this example it is white.
[0036] The masking processing unit 22 further specifies the type of masked installation object when masking an installation object, and specifies the type of masked interior object when masking an interior object. The control unit 20 generates a display image showing the installation object or interior object of the type specified by the masking processing unit 22 together with a masking image, and displays the image on the display unit of the client terminal 11. In this way, the control unit 20 also functions as a display control unit that generates and displays various display images that display information from each unit of the image processing device 13 on the display unit of the client terminal 11.
[0037] When performing masking processing using either the first AI or the second AI, the masking processing unit 22 further infers the type and attributes of the indoor space based on the first AI or the second AI used, and causes the control unit 20 to display the inferred type and attributes on the client terminal 11. In this example, the indoor space is a living room, but it may also be a connecting space connecting living rooms, such as a hallway or staircase. A living room is a room used continuously for living, office work, work, circulating, entertainment, or other similar purposes. The type of indoor space is classified according to its use. In this example, if the indoor space is a living room, it is classified and stored as a bar, bathroom, bedroom / ensuite, dining room, foyer, game area / run bathroom, hobby / craft room, children's room, kitchen, laundry room, living room / family room / lounge, media room, nursery, pantry, single room studio / unit, study room, and sunroom, but this is not limited to this example.
[0038] The attributes of an indoor space are exterior features, and in this example, they are a "design theme" that indicates the type of design, a "color preference" that indicates the color, and a "design idea" that indicates the exterior features of the design that are not expressed by the type of design or color. In this example, the "design themes" are categorized and stored as bohemian, coastal, contemporary, farmhouse, French country, glam, industrial, Japandi, mid-century modern, minimalist, modern, rustic, Scandinavian, traditional, transitional, etc., but are not limited to these examples.
[0039] When a plurality of image objects are masked and a release request is received from the client terminal 11 via the control unit 20 to select at least one image object from the plurality of masked image objects and release the masking, the masking processing unit 22 generates a masked image in which the masking of the selected image object is released. The masked image thus generated is also sent to the client terminal 11 by the control unit 20 and displayed.
[0040] When the designated image generation unit 23 receives designation information about the type and attributes of the indoor space and an instruction to generate an image from the client terminal 11 via the control unit 20, the designated image generation unit 23 uses the first learning model M1 to generate, as a first designated image, an image in which a first virtual object corresponding to the selected one is provided based on the designation information. The designated image generation unit 23 in this example includes a first generation unit 31 and a second generation unit 32, and the first generation unit 31 includes an installation object generation subunit 31a and an interior object generation subunit 31b.
[0041] When the masked image in which the installation object is masked is input from the masking processing unit 22 and the client terminal 11 receives designation information and an instruction to generate an image via the control unit 20, the installation object generation subunit 31a uses the installation object learning model M1a to generate a first designation image showing an indoor space in which the masked installation object is replaced with a first virtual object based on the designation information. In this way, the first virtual object is placed in the masking area of the masked installation object. In order for the installation object generation subunit 31a to perform this process, in addition to the aforementioned learning, the installation object learning model M1a is pre-trained to input designation information about the type and attributes of the indoor space and the installation object, and to generate a still image corresponding to the input designation information and the installation object. That is, the installation object generation subunit 31a is pre-trained to input designation information about the type and attributes of the indoor space and the installation object as input parameters, and the input designation information and a still image corresponding to the installation object as output parameters.
[0042] If the area occupied by the first virtual object does not include the masking area of the installed object, the first AI estimates the part of the masking area outside the outline of the first virtual object from the image object that is the subject image in the background of the installed object, and interpolates with the interior object to create the first specified image.
[0043] When the masked image in which the interior objects are masked is input from the masking processing unit 22 and the client terminal 11 receives specification information and an instruction to generate an image via the control unit 20, the interior object generation subunit 31b uses the interior object learning model M1b to generate a first specified image showing an indoor space in which the masked interior objects are replaced with first virtual objects based on the specification information, using the second AI. In this case, the first virtual object is also placed in the masking area of the masked interior object. In order for the interior object generation subunit 31b to perform this processing, in addition to the above-mentioned learning, the interior object learning model M1b is pre-trained to input specification information about the type and attributes of the indoor space and interior objects and generate a still image corresponding to the input specification information and interior objects.
[0044] When the second generation unit 32 of the designated image generation unit 23 receives a second target image from the acquisition unit 21 and acquires the designation information and an image generation instruction from the client terminal 11 via the control unit 20, the second generation unit 32 uses the second learning model M2 to estimate the type and size of the indoor space shown in the second target image and generates an image in which the second virtual object is provided based on the designation information as a second designated image. When the second target image is received, the third AI identifies fixtures such as doors and window frames and equipment such as air conditioners (hereinafter, fixtures and equipment are collectively referred to as fixtures, etc.) in the second target image and estimates the size and shape of the indoor space captured in the second target image from the size of the identified fixtures, etc. The second learning model M2 is generated by learning from learning data including still images of indoor spaces and content labels indicating the content of each subject image in the still images. Specifically, the second learning model M2 is obtained by inputting a still image without an object, specification information about the type and size of fixtures and fittings shown in the still image, the type and attributes of the indoor space, and the content label, and is trained in advance to generate a still image containing an image object corresponding to the input specification information and content label. That is, the still image without an object, specification information about the type and size of fixtures and fittings shown in the still image, the type and attributes of the indoor space, and the content label are used as input parameters, and the still image containing an image object corresponding to the input specification information and content label is used as an output parameter. In this example, the third AI uses TENSORFLOW or keras.
[0045] The size (data capacity) of the still image used as input information to generate the second learning model M2, like the sizes of the first and second target images, is preferably at most 10 MB (megabytes). This is preferable from the perspective of shortening the time required to transmit an image with a virtual object to the client terminal 11 compared to a size exceeding 10 MB. Furthermore, a size of at least 500 KB (kilobytes) is preferable from the perspective of image recognition accuracy compared to a size smaller than 500 KB. Furthermore, it is preferable for the still image used as input information to include subject images of the floor and ceiling of an indoor space and at least two adjacent walls, from the perspective of more appropriate processing by the second generation unit 32. The third AI uses the second learning model M2 to estimate the type and size of the indoor space from the input second target image and generate a second designated image. The above group information is input using the management terminal 12 (see FIG. 1). The second generation unit 32 outputs the generated second designated image to the control unit 20, which then displays the second designated image on the client terminal 11.
[0046] The operation of the above configuration will be described. When the acquisition unit 21 acquires a first target image obtained by capturing an image of an indoor space in which an object is installed, the acquisition unit 21 sends the first target image to the masking processing unit 22. When the acquisition unit 21 acquires a second target image obtained by capturing an image of an indoor space in which no object is installed, the acquisition unit 21 sends the second target image to the specified image generation unit 23.
[0047] When a first target image is input, the masking processing unit 22 regards this input as a display request input with the installation objects selected, as shown in FIG. 7 in this example, and generates a masked image in which all of the installation objects are masked. In FIG. 7, the state in which the installation objects are selected is indicated by dot hatching with the "furniture and home appliances" tag colored. When the above-mentioned still image G1 (see FIG. 2) is input, the masking processing unit 22 in this example detects (estimates) all of the installation objects in the still image G1 and performs masking processing to generate a masked image G5a with the masked region as shown in FIG. 7, and also generates an extracted image G5b in which the type of the detected installation object is identified and displayed. The masked image G5a and extracted image G5b generated in this way are displayed as display image G5 on the display unit of the client terminal 11 by the control unit 20.
[0048] The extracted image G5b of the display image G5 has a check box for each installed object, and all check boxes are marked with a check. For example, in the example shown in Fig. 7, a window glass, a curtain, a door, a television, a cupboard, a table, and a chair are shown with a check mark. Note that, in the example shown in Fig. 7, the door is classified as an installed object, but as described above, it may also be classified as an interior object.
[0049] The masking processing unit 22 generates a masked image G5a using the first AI, and also infers the type and attributes of the indoor space to generate an image G5c showing these. The image G5c has an "Image Type" column D1, a "Design Theme" column D2, a "Color Preference" column D3, and a "Design Idea" column D4. Column D1 indicates the type of indoor space, and columns D2 to D4 indicate the attributes of the indoor space. In the example shown in FIG. 7, the type of indoor space is displayed as "Living Room / Family Room / Lounge" and the attributes are displayed as "Minimalist" and "Blue, Beige" as the inference results of the first AI. These columns D1 to D4 are also used for performing input operations (including selection operations) on the client terminal 11.
[0050] The client performs a cancellation operation to cancel the masking by deleting the check marks in the check boxes in extraction image G5b on client terminal 11. For example, by deleting the check marks for the window glass, curtains, doors, television, and cupboards, masking processing unit 22 generates extraction image G6b in which these masking marks are cancelled, as in display image G6 shown in Fig. 8, and also generates masked image G6a in which the table and chairs are masked, and these are displayed on client terminal 11. Note that the check marks can be deleted by clicking with a mouse as an input unit or by touching on the touch panel display.
[0051] The client performs input operations in fields D1 to D4 on the client terminal 11. This input operation is the input of the aforementioned specified information. D1 is a pull-down selection field, and is selected from the aforementioned room types, bar, bathroom, bedroom / ensuite, etc. In the example shown in Figure 8, the prediction result of the first AI is retained and remains "living room / family room / lounge."
[0052] Similarly, field D2 is a pull-down selection field, and in this example, as described above, the selection is made from the stored categories of bohemian, coastal, contemporary, etc. In the example shown in Figure 8, the first AI's prediction result of "minimal" (see Figure 7) is changed to "modern" by a selection operation on client terminal 11.
[0053] 8, the first AI's prediction result of "blue, beige" (see FIG. 7) is changed to "white, gray, black" by inputting the color on the client terminal 11. Column D4 is also a text input column like column D3, and in the example shown in FIG. 8, "linear, simple" is input by inputting the color on the client terminal 11, meaning an indoor space with straight lines and a simple design.
[0054] Display image G6 has a "Generate New Design" button B for the client to send an instruction to generate an image. When this button B is operated on client terminal 11, designation information and an instruction to generate an image are sent from client terminal 11 to installed object generation subunit 31a of designated image generation unit 23. In response to this input, installed object generation subunit 31a uses a first AI to infer and generate first virtual objects representing a table and chairs based on the designation information. Then, as shown in FIG. 9, image G7a in which these first virtual objects are arranged is generated as a first designated image, and control unit 20 generates display image G7 including image G7a and displays it on client terminal 11.
[0055] Next, when the masking release or masking specification operation on the extracted image G6b, the input operation of the specification information in the fields D1 to D4, and the operation on the button B are performed on the client terminal 11, the masking processing unit 22 and the installation object generation subunit 31a perform the same processing again, and a new display image is displayed on the client terminal 11.
[0056] When a first target image is input and a display request is input from the client terminal 11 in which an interior object is selected, the masking processing unit 22 generates a masked image in which all of the interior objects are masked using the second AI. In FIG. 10, the state in which an interior object is selected is indicated by dot hatching with the “interior” tag colored. In this example, when a switching operation to “interior object” is performed on the client terminal 11, this switching information is input to the control unit 20, and the control unit outputs this input to the masking processing unit 22 as a display request in which the interior object is selected. In this example, the masking processing unit 22 detects (estimates) all of the interior objects in the still image G1 and generates an image G8a in which the interior objects are masked by performing a masking process, as shown in FIG. 10. Note that, together with generating the image G8a, an extracted image (not shown) may be generated in which the type of the detected interior object is identified and displayed. The image G8a generated in this manner, together with image G8b, is arranged in the display image G8 by the control unit 20 and displayed on the display unit of the client terminal 11.
[0057] Image G8b has a "Color Preference" column D5 and a "Design Idea" column D6. Columns D5 and D6 indicate the attributes of indoor spaces. In the example shown in FIG. 10, the attributes are displayed as "Blue, Beige" in column D5 as the result of estimation by the second AI. These columns D5 and D6 are also used for performing input operations (including selection operations) on client terminal 11, and these input operations are input operations for the above-mentioned specified information. When the client performs a text input operation in columns D5 and D6 and an operation on button B on client terminal 11, second generation unit 32 uses the second AI to generate a first virtual object of the interior according to the specified information, and a display image including this first virtual object is transmitted from control unit 20 to client terminal 11 and displayed.
[0058] According to the above-described configuration, home staging generates a first designated image in which an object is masked and turned into a virtual object for an indoor space. Therefore, even for a currently occupied property, a first designated image with a different design for the indoor space can be displayed on client terminal 11 without the tenant having to move out. Therefore, a client using client terminal 11 can provide customers with a design image that matches their preferences, even when the property is occupied, contributing to promoting real estate sales. Since the image object to be masked can be either a set of installed objects or interior objects, customers can choose to change only the interior or only the installed objects, such as furniture and appliances, providing a variety of home staging options and promoting sales. Furthermore, the still images used by the first target image, first AI, and second AI for training are all images, for example, between 500 KB and 10 MB, captured by commonly used imaging devices such as smartphones and digital cameras. Therefore, the process of generating the first designated image is extremely short. In this example, it has been confirmed that the first designated image is displayed on client terminal 11 in a time ranging from 10 to 20 seconds. As a result, the first designated image can be provided to the customer quickly, and a large number of first designated images can be provided in a short period of time. Furthermore, since it is possible to select objects to be masked and replaced from among the installed objects, home staging is possible in which certain currently installed objects are retained as they are while other specific objects are replaced with virtual objects. Therefore, for example, home staging can be performed while referring to currently installed objects, allowing the customer to grasp a more realistic design. Thus, with the above configuration, useful home staging can be performed using still images captured of the indoor space in which the objects are arranged.
[0059] When the second target image is input from the acquisition unit 21, the second generation unit 32 of the specified image generation unit 23 generates an image G9a representing the second target image as shown in FIG. 11. The control unit 20 then generates a display image G9 including the image G9a and an image G9b for inputting specified information, and displays the display image G9 on the client terminal 11. Image G9b is similar to the aforementioned images G5c and G8b, and displays the result of the estimation made by the third AI and functions as a means for inputting specified information via the client terminal 11. For example, in the example shown in FIG. 11, the third AI estimates the interior color as "blue, beige," and displays the result in field D5. Assume that, for example, "white, gray" is input into field D5 and, for example, "simple" is input into field D6 on the client terminal 11, and text such as "modern," "white, gray, ivory," and "simple" is input into fields D2 to D4 of the "furniture and home appliances" tag similar to the example shown in FIG. 8, and button B is then operated. In this case, the second generating unit 32 is sent designation information indicating these and an instruction to generate the image.
[0060] Based on the designation information, the second generation unit 32 generates a second designated image G10a in which a second virtual object is provided, as shown in Fig. 12. The control unit 20 generates a display image G10 including the second designated image G10a and displays it on the client terminal 11. In this manner, with this configuration, home staging is similarly performed on the second target image captured of an indoor space in which no object is installed. [Explanation of symbols]
[0061] 10 Image Processing System 11 Client terminal 13 Image processing device 21 Acquisition Department 22 Masking processing section 23 Specified image generation unit
Claims
1. an acquisition unit that acquires, as a first target image, a still image captured of an indoor space in which an object is installed; a masking processing unit that, based on a display request for selecting and displaying one of an installation object that is a subject image of the object and an interior object that is a subject image of an interior of the indoor space, among image objects that are subject images of the first target image, masks the selected one to generate the masked image; a designated image generation unit that, when acquiring designation information about the type and attributes of the indoor space and an instruction to generate an image, generates, as a first designated image, an image in which a first virtual object corresponding to the selected one of the indoor spaces is provided based on the designation information, using a first learning model that has learned a still image of the indoor space in which an object is installed; An image processing device comprising:
2. 2. The image processing device according to claim 1, wherein, when a request to select at least one image object from among the plurality of masked image objects and to unmask the image object is received while the plurality of image objects are masked, the masking processing unit generates the masked image in which the masking of the selected image object is unmasked.
3. The designated image generation unit The image processing device according to claim 1 or 2, wherein the first designated image is generated by providing the first virtual object in a masked area of the masking image.
4. The designated image generation unit 3. The image processing device according to claim 1, wherein a non-provided area of the masked masked area where the first virtual object is not provided is interpolated with the image object to form the first designated image.
5. the acquisition unit acquires, as a second target image, a still image captured of the indoor space in which the object is not installed; The designated image generation unit 3. The image processing device according to claim 1, wherein, when the specified information and an instruction to generate an image are acquired, the type and size of the indoor space shown in the second target image are estimated using a second learning model that has learned learning data including a still image of an indoor space and a content label indicating the content of the subject image of the still image, and an image in which a second virtual object is provided based on the specified information is generated as a second specified image.
6. The indoor space is a living room, The types are classified according to their uses, 3. The image processing device according to claim 1, wherein the attribute is an appearance feature.
7. The image processing device according to claim 1 , wherein the acquisition unit acquires the first object image having, as subject images, a floor, a ceiling, and at least two adjacent walls.
8. The image processing device according to claim 1 or 2, wherein the first learning model is obtained by learning the still image having a floor, a ceiling, and at least two adjacent walls as subject images.
9. an acquisition step of acquiring a still image of an indoor space in which an object is installed as a first target image; a masking processing step of selecting and displaying either an installation object that is a subject image of the object or an interior object that is a subject image of the interior of the indoor space, among image objects that are subject images of the first target image, based on a display request for requesting display, and masking the selected one to generate the masked image; a designated image generating step of generating, when designation information on the type and attributes of the indoor space and an instruction to generate an image, as a first designated image, an image in which a first virtual object corresponding to the selected one of the indoor spaces is provided based on the designation information, using a first learning model that has learned from a still image of the indoor space in which an object is installed; An image processing method comprising:
10. an acquisition step of acquiring a still image of an indoor space in which an object is installed as a first target image; a masking processing step of selecting and displaying either an installation object that is a subject image of the object or an interior object that is a subject image of the interior of the indoor space, among image objects that are subject images of the first target image, based on a display request for requesting display, and masking the selected one to generate the masked image; a designated image generating step of generating, when designation information on the type and attributes of the indoor space and an instruction to generate an image, as a first designated image, an image in which a first virtual object corresponding to the selected one of the indoor spaces is provided based on the designation information, using a first learning model that has learned from a still image of the indoor space in which an object is installed; An image processing program that causes a computer to execute the following.
Citation Information
Patent Citations
Image processing method, program, and image processing system
JP2022077148A