Architectural image output device, architectural image creation method, and computer program

The architectural image output device uses generative AI to transform perspective images and spatial data based on customer requests, addressing discrepancies in building design images and enhancing customer satisfaction.

JP2026078920APending Publication Date: 2026-05-15FINE
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
FINE
Filing Date
2024-10-29
Publication Date
2026-05-15

AI Technical Summary

Technical Problem

Existing systems struggle to create perspective images that accurately reflect customer preferences and variations in building designs, leading to discrepancies between customers and designers.

Method used

An architectural image output device that utilizes a trained model to transform perspective images and spatial data based on customer requests, using generative AI models to convert images and spatial data to match desired atmospheres or styles.

Benefits of technology

Enables the creation of images that meet specific customer requirements, improving satisfaction by ensuring alignment between customer and designer visions without pre-prepared texture combinations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026078920000001_ABST
    Figure 2026078920000001_ABST
Patent Text Reader

Abstract

This invention provides an architectural image output device, an architectural image creation method, and a computer program capable of creating images of buildings that meet specific requirements. [Solution] The architectural image output device includes a processing unit that acquires a perspective image from a specific viewpoint inside or outside a building, or spatial data indicating the distance to equipment or furniture shown in the perspective image, accepts reference data corresponding to the requirements for the building, provides the acquired perspective image or spatial data to a trained model that transforms the perspective image or spatial data using the features of the reference data, and outputs the image transformed by the trained model as a transformed perspective image.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a building image output device, a building image creation method, and a computer program capable of creating an image of a building according to a request.

Background Art

[0002] In a scene where a building is designed or an interior is selected, a customer who is an owner of a building and a designer need to share images with each other. Generally, the customer and the designer share an image of an already built building or interior, or a word that evokes an image.

[0003] In order to share an image between a customer and a designer, a device that automatically creates an image from a three-dimensional model has been proposed in Patent Document 1. The image creation device of Patent Document 1 associates different texture information or colors such as various wall colors, floor materials, and furniture materials with the three-dimensional model, and creates a perspective image for each different texture. In advance, an interior coordinator or the like creates various perspective images using textures selected from an interior catalog, so that the customer can grasp the image while switching textures and share requests with the designer.

Prior Art Documents

Patent Documents

[0004]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0005] By pre-associating texture information or color with a three-dimensional model and creating a perspective drawing from the three-dimensional model and the texture information or color, it is possible to provide aesthetically pleasing perspective drawings. To minimize discrepancies in image between the customer and the designer, it is desirable to be able to create perspective images corresponding to texture information or combinations that are not pre-prepared, in order to accommodate a wider variety of variations.

[0006] The present invention aims to provide an architectural image output device, an architectural image creation method, and a computer program capable of creating images of buildings that meet specific requirements. [Means for solving the problem]

[0007] An architectural image output device according to one embodiment of the present disclosure includes a processing unit that acquires a perspective image from a specific viewpoint inside or outside a building, or spatial data indicating the distance to equipment or furniture shown in the perspective image, receives reference data corresponding to a request for the building, provides the acquired perspective image or spatial data to a trained model that transforms the perspective image or spatial data using the features of the reference data, and outputs the image transformed by the trained model as a transformed perspective image. [Effects of the Invention]

[0008] According to this disclosure, it is possible to create images of buildings that meet the requirements. [Brief explanation of the drawing]

[0009] [Figure 1] This is a schematic diagram of the architectural image provision system in the first embodiment. [Figure 2] This is a block diagram showing the configuration of an architectural image output device. [Figure 3] This is an overview diagram of the learning model. [Figure 4] This is a block diagram showing the configuration of a terminal device. [Figure 5]It is a flowchart showing an example of image processing by the building image output device in the first embodiment. [Figure 6] It is a flowchart showing an example of image processing by the building image output device in the first embodiment. [Figure 7] A display example for sharing the customer's image is shown. [Figure 8] A display example for sharing the customer's image is shown. [Figure 9] A display example for sharing the customer's image is shown. [Figure 10] It is a schematic diagram of the learning model used in the building image output device of the second embodiment. [Figure 11] It is a flowchart showing an example of image processing by the building image output device in the second embodiment. [Figure 12] It is a flowchart showing an example of image processing by the building image output device in the second embodiment. [Figure 13] It is a schematic diagram of the building image providing system of the third embodiment. [Figure 14] It is a block diagram showing the configuration of the building image output device of the third embodiment. [Figure 15] It is a schematic diagram of the learning model. [Figure 16] It is a flowchart showing an example of image processing by the building image output device in the third embodiment. [Figure 17] It is a schematic diagram of the building image providing system of the fourth embodiment. [Figure 18] It is a flowchart showing an example of image processing by the building image output device in the fourth embodiment. [Embodiments for Carrying Out the Invention]

[0010] The present disclosure will be specifically described with reference to the drawings showing its embodiments. In the following embodiments, a building image providing system including the building image output device of the present disclosure will be described.

[0011] [First Embodiment] FIG. 1 is a schematic diagram of a building image providing system 100 according to the first embodiment. The building image providing system 100 of the first embodiment includes a building image output device 1, a parse image creation device 2, a three-dimensional model database 3, and a terminal device 4. The building image output device 1, the parse image creation device 2, and the three-dimensional model database 3 can input and output data via a network N1 of a service provider that provides a service for providing parse images. The building image output device 1 and the terminal device 4 used by a customer or a salesperson of a construction company who shows images to the customer can transmit and receive data via a network N2 which is a public communication network.

[0012] The parse image creation device 2 can automatically set an appropriate viewpoint for the data of the three-dimensional model read from the three-dimensional model database 3 or the data of the newly input three-dimensional model, and create a parse image from the set viewpoint. In the building image providing system of the present disclosure, the parse image creation device 2 is set to output the spatial data corresponding to the created parse image together with the parse image.

[0013] The building image output device 1 acquires the parse image (or spatial data combined) created by the parse image creation device 2, and creates a new parse image corresponding to the reference data close to the customer's request such as an image or text given via the terminal device 4 from the acquired image, and displays it on the terminal device 4. Details of the configuration and processing of the building image output device 1 will be described later.

[0014] The three-dimensional model database 3 stores data of three-dimensional models created in advance. The data of the three-dimensional model may be tagged with keywords such as gross floor area, number of rooms, presence or absence of a children's room, presence or absence of a void, number of floors, and floor number where the living room is located. In addition, the data of the three-dimensional model only needs to be stored in the three-dimensional model database 3 so that it can be searched. Texture is associated with the data of the three-dimensional model, and a keyword explaining the texture is tagged. The data of the three-dimensional model may be tagged with a suitable family composition and presence or absence of pets.

[0015] In the architectural image provision system 100 of the first embodiment, a customer of a building or a sales representative of a construction company showing images to a customer connects to the architectural image output device 1 using a terminal device 4. The customer or sales representative uses the terminal device 4 to select a perspective image of a building that meets the customer's requirements and obtains the result of converting the selected perspective image into a perspective image with a texture according to the customer's requirements.

[0016] The configuration and processing of each device for realizing such an architectural image provision system 100 will be explained in detail. Figure 2 is a block diagram showing the configuration of the architectural image output device 1. The architectural image output device 1 uses a server computer. In the following description, the architectural image output device 1 is assumed to be a single server computer, but it is not limited to this, and a configuration in which multiple server computers are connected by communication to distribute processing and storage area may be used. The architectural image output device 1 comprises a processing unit 10, a storage unit 11, and a communication unit 12.

[0017] The processing unit 10 includes one or more processors such as a CPU (Central Processing Unit), MPU (Micro-Processing Unit), or GPU (Graphics Processing Unit). The processing unit 10 also includes memory, which is a temporary storage medium such as SRAM (Static Random Access Memory) or DRAM (Dynamic Random Access Memory). The processing unit 10 may be configured as a single hardware (SoC: System On a Chip) integrating the processor, memory, storage unit 11, and communication unit 12. The processing unit 10 reads the architectural image output program P1 stored in the storage unit 11 into memory and executes it, thereby causing a general-purpose computer to perform various processes described later, and to function as the architectural image output device 1 of this disclosure.

[0018] The storage unit 11 is a relatively large-capacity non-temporary storage medium such as a hard disk, flash memory, or SSD (Solid State Drive). The storage unit 11 stores the program (program product) necessary for the processing unit 10 to execute processing. The program product includes an architectural image output program P1, a web server program P2 for performing web server functions, a learning model M1, and images and their designs (arrangement) to be included in the screen to be displayed on the terminal device 4.

[0019] The architectural image output program P1 and the learning model M1 stored in the memory unit 11 may be downloaded by the processing unit 10 from a download server via the communication unit 12 and stored in the memory unit 11, or the architectural image output program P9 and the learning model M9 stored in a non-temporary storage medium 9 readable from a computer may be read by the processing unit 10 and stored in the memory unit 11.

[0020] The learning model M1 may be executed in part or in whole on an external server accessible via network N.

[0021] The communication unit 12 enables communication with the three-dimensional model database 3 and the perspective image creation device 2 via the network N2, and communication with the terminal device 4 via the network N1. The communication unit 12 is a wired communication device such as Ethernet® or a wireless communication device for WiFi.

[0022] The architectural image output device 1 makes perspective images created based on three-dimensional models stored in the three-dimensional model database 3, or perspective images created from scratch, accessible from the terminal device 4 via a web page. The architectural image output device 1 accepts requests from the terminal device 4 to change the atmosphere of the interior, change the texture of the equipment, or change the type of equipment in the perspective image by inputting reference data, and converts the perspective image according to the request before outputting it to the terminal device 4.

[0023] The learning model M1 used for this purpose is a so-called generative AI model that uses a language model. Figure 3 is an overview diagram of the learning model M1. The learning model M1 employs a multimodal model and accepts both text and image inputs. The learning model M1 is optimized, for example, to convert the input image data into image data that corresponds to the features of the input data used as reference data. The learning model M1 utilizes so-called generative AI using an architecture such as Stable Diffusion, and includes an image encoder and a text encoder in a VAE (Variational Auto-Encoder), a decoder in the VAE that converts the features output from each encoder back into an image, and an extension network in between.

[0024] The learning model M1 is trained to output an image obtained by converting a given perspective image to the color tone of another perspective image represented by text. Alternatively, the learning model M1 may be trained to output an image obtained by converting a given perspective image and spatial data, or a depth image corresponding to the spatial data, to a perspective image that matches the atmosphere of another image. By using the learning model M1, the processing unit 10 can output a perspective image that meets the requirements.

[0025] The learning model M1 may be a model that uses an autoencoder for image conversion to convert the input perspective image data and a perspective image of another building into a perspective image with the same color tones as the equipment or furniture in the perspective image of the other building. The processing unit 10 can also obtain perspective images from different images without having to recreate a three-dimensional model with different texture information by providing the learning model M1 with the perspective image and a perspective image of another building or a real-world image.

[0026] A multimodal model including a language model provided by an external server may be used as the learning model M1. The learning model M1 may also be, for example, a model that transforms images using a generative AI employing a different architecture. The learning model M1 may be a language model such as a large language model (LLM) or a small language model (SLM).

[0027] Figure 4 is a block diagram showing the configuration of the terminal device 4. The terminal device 4 comprises a processing unit 40, a storage unit 41, a communication unit 42, a display unit 43, and an operation unit 44.

[0028] The processing unit 40 includes one or more arithmetic processing units such as CPUs, MPUs, and GPUs. The processing unit 40 also includes a temporary storage medium such as SRAM or DRAM. The processing unit 40 reads the web browser program stored in the storage unit 41 into the temporary storage medium and executes it, thereby causing a general-purpose computer to perform various processes described later and display a perspective image.

[0029] The storage unit 41 is a relatively large-capacity non-volatile storage area such as an SSD or hard disk. The storage unit 41 stores the programs (program products) necessary for the processing unit 40 to execute processing, as well as configuration data for reference. The program products include a web browser program.

[0030] The communication unit 42 enables communication with the architectural image output device 1 via the network N2. The communication unit 42 is a wired communication device such as Ethernet® or a wireless communication device for WiFi. The communication unit 42 may also be a wireless communication module connected to a carrier network.

[0031] The display unit 43 is a display such as a liquid crystal display or an organic EL (Electro Luminescence) display. The display unit 43 may be, for example, a touch panel-integrated display.

[0032] The operation unit 44 is a user interface that can input and output to the processing unit 40. The operation unit 44 may be, for example, a keyboard and mouse. The operation unit 44 may be a touch panel built into the display unit 43, or it may be configured to include physical buttons, switches, and physical dials. The operation unit 44 may also be configured to accept voice operation using a microphone and a voice recognition processing unit.

[0033] Figures 5 and 6 are flowcharts illustrating an example of image processing by the architectural image output device 1 in the first embodiment. When a customer or sales representative accesses a web page provided by the architectural image output device 1 using a terminal device 4, the following processing steps are executed. The following processing in the terminal device 4 may be automatically detected on the web server side of the architectural image output device 1 by a script included in the web page.

[0034] The processing unit 40 of the terminal device 4 receives search keywords for buildings that meet the request on a web page (step S401). In step S401, the processing unit 40 receives keywords such as family structure, room conditions, total floor area, and construction method. The processing unit 40 may also receive these keywords via voice recognition. The processing unit 40 sends a search request for a three-dimensional model specifying the received search keywords to the building image output device 1 (step S402).

[0035] The architectural image output device 1 receives a request to search for a three-dimensional model (step S101). The processing unit 10 searches for a three-dimensional model from the three-dimensional model database 3 using the search keywords specified in the received search request (step S102). Each three-dimensional model is associated with the total floor area of ​​the building and tags, as described above, so it can be searched using these. Alternatively, the search may be based on the size of the building, records of actual residents, the strength of the correlation between the content of the interview items when creating the housing design plan, such as the customer's family structure, hobbies, presence of pets, preferred colors, preferred materials, etc., and the actual building. The processing unit 10 extracts a predetermined number of three-dimensional models from the search results (step S103). In step S103, the processing unit 10 extracts, for example, five three-dimensional models with the highest degree of matching with the search keywords or the content of the customer's interview items from the search results.

[0036] The processing unit 10 provides each extracted three-dimensional model to the perspective image creation device 2 to create a perspective image and corresponding spatial data (step S104). In step S104, the processing unit 10 creates a perspective image using the texture most frequently used for each three-dimensional model, or using the texture with the highest degree of matching to the search request. Here, the texture of the perspective image is used to select the perspective image, so any texture can be used. The spatial data may be a depth image showing the distance to each object in the three-dimensional model from a set viewpoint in pixel values, or it may be a matrix of data showing the distance to the object corresponding to each pixel in the image.

[0037] The processing unit 10 creates a web page that includes a screen listing the created perspective images for selection (step S105) and sends it to the terminal device 4 (step S106).

[0038] The processing unit 40 receives a web page containing a list of perspective images (step S403), displays the web page (step S404), and accepts the selection of a perspective image (step S405).

[0039] The processing unit 40 receives reference data corresponding to the customer's request for the selected perspective image (step S406). The reference data may be other perspective images, keywords indicating a style such as "Nordic style" or "modern," or the name of a chain store to evoke a standardized interior in a chain store. The processing unit 40 transmits the identification data of the selected perspective image and the received reference data to the architectural image output device 1 (step S407).

[0040] The architectural image output device 1 receives identification data of the selected perspective image and reference data received by the terminal device 4 (step S107). The processing unit 10 of the architectural image output device 1 creates a prompt for the learning model M1 to convert the perspective image to a perspective image that evokes an impression from the reference data, based on the selected perspective image, the spatial data corresponding to the perspective image, and the received reference data (step S108). In step S108, the processing unit 10 creates a natural language prompt, for example, "Please convert the attached perspective image to a perspective image that has the atmosphere of the attached reference data." The processing unit 10 may also create a prompt that instructs the processing unit 10 to vectorize the reference data by providing it to another language model, convert it to vectorized features, and then convert the perspective image to another perspective image. The processing unit 10 may also create a prompt that instructs the processing unit 10 to convert the spatial data (depth information) itself and the reference data to a perspective image.

[0041] The processing unit 10 provides the created prompt and the selected perspective image to the learning model M1 (step S109), and obtains the converted image output from the learning model M1 (step S110). In step S109, the processing unit 10 may provide spatial data together with the perspective image, or it may provide spatial data alone to the learning model M1, depending on how the learning model M1 has been trained. The processing unit 10 transmits the converted image to the terminal device 4 as a perspective image according to the customer's request (step S111). In step S111, the processing unit 10 may also transmit floor plan data output from the three-dimensional model corresponding to the selected perspective image to the terminal device 4 together with the perspective image.

[0042] The terminal device 4 receives the converted perspective image (step S408), and the processing unit 40 displays the web page containing the received perspective image on the display unit 43 (step S409). The processing unit 40 then decides whether or not to convert it to a perspective image with a different impression (step S410). In step S410, the processing unit 40 makes this decision based on whether or not a switching interface (button, icon, etc.) for switching the impression displayed on the web page has been selected.

[0043] If the processing unit 40 determines that it will convert to a perspective image with a different impression (S410: YES), it returns to step S406 and accepts other reference data (S406). If the processing unit 40 determines that it will not convert to a perspective image with a different impression (S410: NO), it terminates the process.

[0044] Figures 7-9 show examples of displays for sharing customer images. Figure 7 shows an example of a search screen 430 for searching for a three-dimensional model of a building suitable for the customer's requirements. The search screen 430 is an interface provided by the Web server program P2 of the building image output device 1. The search screen 430 includes an input field 431 for search keywords and a search button 432. In the example in Figure 7, "12 tatami mat LD" is entered in the input field 431. This searches for buildings designed to include a large living and dining room. When the search button 432 is selected, the processing unit 40 of the terminal device 4 retrieves the keyword entered in the input field 431 and sends a search request specifying the retrieved keyword to the building image output device 1. In the search screen 430 shown in Figure 7, the search is configured to use text, but it may also be configured to search from images in a plan collection provided by a building provider, for example.

[0045] Figure 8 shows an example of a conversion screen 433 for converting search results for a three-dimensional model. The conversion screen 433 includes a display area 434 for a perspective image created from the three-dimensional model of the building selected from the search results. The conversion screen 433 includes a switching interface 435 for switching the impression of the perspective image. The switching interface 435 is icon-shaped and can be converted to other perspective images with impressions corresponding to the text or marks, colors, etc., within the icon. As shown in Figure 8, the switching interface 435 includes not only text indicating atmospheres such as "Nordic style" and "modern," but also "ABC style," which includes the name of a restaurant chain "ABC" that shares common features in its interior and exterior. When a customer or sales representative selects the switching interface 435 on the display screen, the processing unit 40 of the terminal device 4 provides the learning model M1 with the spatial data corresponding to the original perspective image, along with the image data of the exterior or interior associated with the text "Nordic style" or "modern," or "ABC style," as reference data. The processing unit 40 acquires the converted perspective image output from the learning model M1 and displays it on the display unit 43.

[0046] The conversion screen 433 includes an input field 436 for text to be entered as reference data and a drag-and-drop field 437 for image data. The conversion screen 433 also includes a conversion button 438. When the conversion button 438 is selected, the processing unit 40 provides the text or image entered in the input field 436 or the drag-and-drop field 437 to the learning model M1 as reference data along with the perspective image. The processing unit 40 acquires the converted perspective image output from the learning model M1 and displays it on the display unit 43.

[0047] Figure 9 shows another example of the conversion screen 433. Figure 9 shows an example of what is displayed when, in the conversion screen 433 shown in Figure 8, image data showing dark-colored interiors is selected as reference data for the perspective image displayed in the display area 434, and the convert button 438 is selected. The perspective image 439 shown in Figure 9 has darker furniture colors compared to the perspective image shown in Figure 8. The conversion screen 433 shown in Figure 9 may include not only a switching interface 435 for further conversion as described above, but also an interface that allows the floor plan data output from the three-dimensional model corresponding to the converted perspective image 439 to be downloaded and displayed on the terminal device 4, alongside the converted perspective image 439.

[0048] In this way, by providing the original perspective image along with reference data to the learning model M1, it is possible to create a perspective image that gives a common impression with the impression corresponding to the reference data. Compared to conventional methods of selecting and converting textures to be combined with a three-dimensional model, this method allows for obtaining a perspective image converted by applying textures that are not pre-prepared. It also becomes possible for customers to easily and repeatedly create trial perspective images with altered impressions by inputting natural language as reference data themselves using the terminal device 4 that they use directly. This makes it possible to improve customer satisfaction by ensuring that the image is shared between the customer and the designer or interior coordinator.

[0049] In the architectural image provision system 100 described above, a three-dimensional model pre-stored in the three-dimensional model database 3 was used. However, in step S104, the processing unit 10 of the architectural image output device 1 may, instead of having the perspective image creation device 2 create a perspective image and spatial data from the extracted three-dimensional model, provide the perspective image creation device 2 with a three-dimensional model output from the CAD system to create the perspective image and spatial data.

[0050] [Second Embodiment] In the second embodiment, the processing unit 10 of the architectural image output device 1 uses data of objects that appear (are projected) in the perspective image (hereinafter referred to as subject data) to suppress the conversion of parts of the image to incorrect objects when the perspective image is converted by the learning model. The subject data may be a segmentation image that shows which object each pixel in the perspective image corresponds to, or it may be data that describes the objects appearing in the perspective image in text, corresponding to data that indicates their position in the perspective image.

[0051] Figure 10 is a schematic diagram of the learning model M2 used in the architectural image output device 1 of the second embodiment. In the second embodiment, the learning model M2, like the learning model M1 of the first embodiment, uses so-called generative AI and includes an encoder and decoder for images in a VAE (Variational Auto-Encoder) and an extended network. The learning model M2 accepts subject data input to the decoder along with reference data. As described above, if the subject data is segmentation data, it is an image in which areas corresponding to outdoor equipment (walls, doors, windows) or indoor equipment (walls, doors, windows), furniture, or home appliances of a building shown in the perspective image are colored with a color or pattern set for each identified piece of equipment, furniture, or home appliance. The subject data may be output from the three-dimensional model by the perspective image creation device 2, or it may be included in the three-dimensional model.

[0052] The learning model M2 is optimized to convert perspective images (or perspective images and spatial data) into image data that corresponds to the subject data input as reference data and the text or image features.

[0053] Similar to learning model M1, the learning model M2 converts the input perspective image (or perspective image and spatial data) into image data that corresponds to the input image or text used as reference data. By using subject data as reference data, the distinction between walls and floors, and between walls and windows, becomes clear, preventing erroneous conversions where window areas are mistakenly colored as walls during color tone conversion.

[0054] In the second embodiment, the processing unit 10 may compare the converted perspective image obtained using the learning model M1 of the first embodiment with the subject data, and modify the converted perspective image based on the comparison result.

[0055] Figures 11 and 12 are flowcharts illustrating an example of image processing by the architectural image output device 1 in the second embodiment. For the processing steps shown in Figures 11 and 12 that are common to the processing steps shown in Figures 5 and 6 of the first embodiment, the same step numbers are used, and detailed explanations are omitted.

[0056] In the second embodiment, the processing unit 10 of the architectural image output device 1 provides each of the three-dimensional models extracted from the search results to the perspective image creation device 2 to create a perspective image and corresponding spatial data and subject data (step S124). In step S124, the processing unit 10 may acquire the spatial data and subject data using means other than the perspective image creation device 2, for example, an external image creation service.

[0057] In the second embodiment, when the processing unit 10 receives the perspective image selected by the terminal device 4 and the received reference data (S107), it creates a prompt using the subject data (step S128). In step S128, the processing unit 10 creates a prompt to convert the selected perspective image into a perspective image that evokes an impression from the reference data, using information about the objects that should be depicted as defined in the subject data.

[0058] The transformation by the learning model M2 using the subject data of the second embodiment makes it clear to distinguish between walls and floors, and between walls and windows, thus preventing erroneous transformations where the window area is mistakenly colored as a wall when converting the color tone.

[0059] [Third Embodiment] In the third embodiment of the architectural image provision system 100, the system outputs a building plan that meets the customer's requirements without requiring the customer or sales representative to select a perspective image. The customer or sales representative only needs to input the results of the interview to obtain a floor plan and a perspective image as the building plan.

[0060] Figure 13 is a schematic diagram of the architectural image provision system 100 of the third embodiment. The architectural image provision system 100 of the third embodiment includes an architectural image output device 1, a perspective image creation device 2, a three-dimensional model database 3, and a terminal device 4, similar to the first embodiment, and the connection configuration with networks N1 and N2 is also the same as that of the first embodiment. For components of the architectural image provision system 100 of the third embodiment that are common with the first embodiment, the same reference numerals are used and detailed explanations are omitted.

[0061] In the third embodiment, the architectural image output device 1 is connected to by a customer of a building or a sales representative of a building company who shows images to a customer, using a terminal device 4. In the third embodiment, the customer or sales representative uses the terminal device 4 to input the results of a hearing regarding the customer's requirements for a building, such as a custom-built house, and the architectural image output device 1 can then provide floor plan data corresponding to the customer's requirements and a perspective image of one room in that floor plan data.

[0062] Figure 14 is a block diagram showing the configuration of the architectural image output device 1 of the third embodiment. The architectural image output device 1 of the third embodiment stores learning model M1 and learning model M3 in the storage unit 11. Learning model M3 also uses a language model and is trained to output floor plan data from given text without going through a three-dimensional model. Except for the fact that learning model M3 is stored, the configuration of the architectural image output device 1 is the same as that of the architectural image output device 1 of the first embodiment, so a detailed explanation is omitted.

[0063] Figure 15 is an overview diagram of the learning model M3. As mentioned above, the learning model M3 includes a language model. The learning model M3 employs, for example, a Transformer and includes a text encoder and a decoder that uses the features output from the text encoder to output descriptive data of the floor plan data (a structure described in, for example, JSON). The learning model M3 is a model that translates the word group resulting from the hearing into descriptive data of the floor plan. The learning model M3 may also employ a multimodal model and use a text encoder and a VAE decoder that, given the features output from the encoder, converts those features into an image to output the floor plan data as a design drawing (two-dimensional CAD data composed of vector data, etc.). Part or all of the learning model M3 may be provided by an external server via the network N2.

[0064] Preferably, the learning model M3 is trained to output more accurate floor plan data by using the relationship between tags in the three-dimensional model database 3 and floor plan data based on the three-dimensional model as knowledge.

[0065] The text of the interview results includes, for example, the conditions of the room the customer wants, the customer's family structure, preferred interior style, preferred furniture (color scheme, furniture manufacturer), preferred colors, preferred materials, preferred color scheme of accessories, hobbies, points that are important to the customer, and the image of the preferred space, as answered by the customer in a questionnaire. The learning model M3 accepts this text of interview results as input. The learning model M3 is trained to use data from past construction projects and the initial interview results text and the data of the final designed floor plan (descriptive data) as training data, so that the final floor plan data is output.

[0066] Figure 16 is a flowchart showing an example of image processing by the architectural image output device 1 in the third embodiment. When a customer or sales representative accesses a web page provided by the architectural image output device 1 using a terminal device 4, the following processing steps are executed. The following processing in the terminal device 4 may be automatically detected on the web server side of the architectural image output device 1 by a script included in the web page.

[0067] The processing unit 40 of the terminal device 4 receives the hearing results on the web page (step S421). In step S421, the processing unit 40 may receive the hearing results by accepting selections on the web page using input fields, radio buttons, and check buttons. The processing unit 40 may also receive the hearing results by speech recognition. The processing unit 40 transmits the received hearing results to the architectural image output device 1 (step S422).

[0068] The processing unit 10 of the architectural image output device 1 receives the interview results (step S131), inputs the interview results to the learning model M3 (step S132), and acquires the floor plan data (design drawing) output from the learning model M3 (step S133). The processing unit 10 provides the floor plan data to the perspective image creation device 2 to create a perspective image and spatial data (step S134). In step S134, the processing unit 10 may output subject data in addition to the perspective image and spatial data.

[0069] The processing unit 10 creates reference data from the hearing results received in step S131 (step S135). The reference data consists of the keywords related to preferred interior styles included in the hearing results, or image data that is highly correlated with the keywords. The reference data may also be graph data (matrix) that includes nodes corresponding to each text and directed edges indicating the strength of the relationship between the texts.

[0070] The processing unit 10 uses the perspective image and spatial data created in step S134 and the created reference data to create a prompt for the learning model M1 to convert the perspective image to another perspective image that evokes an impression from the reference data (step S136). In step S136, the processing unit 10 creates a natural language prompt such as, "Please convert the attached perspective image and spatial data to a perspective image that captures the atmosphere of the attached reference data." The processing unit 10 may also create a prompt that instructs the reference data to be given to another language model to be vectorized, converted into vectorized features, and then converted into a new perspective image.

[0071] The processing unit 10 provides the created prompt, the created perspective image, and the spatial data to the learning model M1 (step S137), and acquires the image output from the learning model M1 (step S138). The processing unit 10 sends the converted image to the terminal device 4 as a perspective image according to the customer's request (step S139). The processing unit 10 sends the floor plan data acquired in step S133 to the terminal device 4 (step S140).

[0072] Terminal device 4 receives the perspective image and floor plan data (step S423). The processing unit 40 of terminal device 4 displays a web page on the display unit 43 that can display the received perspective image and floor plan data (step S424), and then terminates the process.

[0073] The processing unit 10 of the architectural image output device 1 may acquire floor plan data from the learning model M3 multiple times based on the hearing results received in step S131, and create multiple perspective images. In the architectural image provision system 100 of the third embodiment, by inputting the customer's hearing results, floor plan data and perspective images that make it easy to understand the building corresponding to that floor plan can be created without actually creating a three-dimensional model. This makes it possible to share an image at the time of the initial plan presentation between the customer and the building provider.

[0074] In the architectural image provision system 100 of the third embodiment, the architectural image output device 1 can apply the method of the first embodiment, which extracts and uses a three-dimensional model that matches the input interview results from the three-dimensional model database 3. Alternatively, a three-dimensional model created from the interview results in a CAD system may be used.

[0075] [Fourth Embodiment] Figure 17 is a schematic diagram of the architectural image provision system 100 in the fourth embodiment. The architectural image provision system 100 in the fourth embodiment includes an architectural image output device 1, a perspective image creation device 2, a three-dimensional model database 3, a terminal device 4, a network N1, and a network N2, similar to the first and third embodiments. In the fourth embodiment, similar to the third embodiment, the architectural image output device 1 automatically creates floor plan data. In the fourth embodiment, the architectural image output device 1 outputs floor plan data that accurately reflects the customer's requests. For components of the architectural image provision system 100 in the fourth embodiment that are common with the first or third embodiment, the same reference numerals are used and detailed descriptions are omitted.

[0076] In the fourth embodiment, the architectural image output device 1 uses a language model M4 in addition to the learning models M1 and M3. In the fourth embodiment, instead of directly inputting the hearing results into the learning model M3 shown in the third embodiment, the language model M4 is used to process the data so that it can output floor plan data with high accuracy. The language model M4 may be a large language model or a small language model. The language model M4 may be provided from an external server via the network N2.

[0077] Therefore, the language model M4 is a model that outputs a response to an input query by outputting a set of texts related to the context represented by the input set of texts. The architectural image output device 1 of the fourth embodiment uses the language model M4 to classify the text of the hearing results into text used to output floor plan data and text used to create data to be used as reference data.

[0078] Figure 18 is a flowchart showing an example of image processing by the architectural image output device 1 in the fourth embodiment. When a customer or sales representative accesses a web page provided by the architectural image output device 1 using a terminal device 4, the following processing steps are executed. The following processing in the terminal device 4 may be automatically detected on the web server side of the architectural image output device 1 by a script included in the web page. Detailed explanations of the processing steps shown in Figure 18 that are common with the processing steps shown in Figure 16 of the third embodiment are omitted.

[0079] In the fourth embodiment, the processing unit 10 of the architectural image output device 1 creates prompts to create data for creating floor plan data and reference data from the hearing results received in step S131, i.e., the word group (step S141). In step S141, the processing unit 10 creates prompts to cause the language model M4 to output data corresponding to the conditions of the building or specifications, in order to input it into the learning model M3 that creates floor plan data. The processing unit 10 creates prompts to cause the language model M4 to output requests for floor plans or lifestyles, and requests for interior styles, according to the hearing results, in order to create characters or images that will become reference data. The processing unit 10 provides the created prompts to the language model M4 (step S142). The processing unit 10 obtains the data for the learning model M3 and the reference data for the learning model M1 from the language model M4 (step S143). The processing in steps S141-S143 is not limited to processing using language model M4; data for learning model M3 and reference data that have been pre-associated with the answer choices for each item of the hearing may also be read.

[0080] The processing unit 10 inputs the data for the building or specifications intended for the learning model M3 from the data acquired in step S143 into the learning model M3 (step S144). The data for the building or specifications includes the total floor area, number of floors, number of rooms, whether the living room is an LDK or LD type, the floor on which the living room is located, the floor on which the bathroom is located, etc. The processing unit 10 acquires the floor plan data (design drawing) output from the learning model M3 (step S145).

[0081] The processing unit 10 provides the floor plan data to the perspective image creation device 2, causing it to create a perspective image and spatial data (step S146). In step S146, the processing unit 10 may output subject data in addition to the perspective image and spatial data, similar to the second embodiment.

[0082] The processing unit 10 uses the perspective image and spatial data created in step S146 and the reference data for the learning model M1 acquired in step S143 to create a prompt for the learning model M1 to convert the perspective image and spatial data into a perspective image that evokes an impression from the reference data (step S147). The processing in step S147 is the same as the processing in step S136 of the third embodiment. In step S147, a prompt may be created to generate another perspective image from the spatial data, which is a depth image, without using the original perspective image.

[0083] The processing unit 10 provides the created prompt, the created perspective image, and the spatial data to the learning model M1 and acquires the perspective image (S138).

[0084] In the fourth embodiment, the language model M4, learning model M3, and learning model M1 are used separately. In order to elicit customer requests for buildings, it becomes possible to accurately separate and generate information about the structure, which can be directly obtained as customer requests for the floor plan during the hearing, from information about decorations such as interiors or exteriors, which can be obtained as customer requests during the perspective drawing creation stage.

[0085] The embodiments disclosed above are illustrative in all respects and not restrictive. The scope of the present invention is defined by the claims, and all modifications within the meaning and scope equivalent to the claims are included. [Explanation of Symbols]

[0086] 1. Architectural image output device 10 Processing Unit 11 Storage section P1 Architectural Image Output Program (Computer Program) M1, M3 Learning Models 2. Perspective image creation device 3. Three-dimensional model database 4 Terminal devices 43 Display section

Claims

1. Obtain a perspective image of a building from a specific viewpoint, either indoors or outdoors, or spatial data indicating the distance to equipment or furniture shown in the perspective image. We accept reference data corresponding to the requirements for the aforementioned building. The acquired perspective image or spatial data is given to a trained model that transforms perspective images or spatial data using the features of the reference data. The image transformed by the aforementioned trained model is output as the transformed perspective image. A building image output device equipped with a processing unit.

2. The aforementioned reference data is a perspective image or photograph of another building corresponding to the aforementioned request, or text corresponding to the aforementioned request. The architectural image output device according to claim 1.

3. The aforementioned processing unit, For the perspective image or spatial data, subject data indicating which object, including building equipment or furniture, each pixel in the perspective image corresponds to is provided to the trained model as reference data. The architectural image output device according to claim 1.

4. Computers Obtain a perspective image of a building from a specific viewpoint, either indoors or outdoors, or spatial data indicating the distance to equipment or furniture shown in the perspective image. We accept reference data to reflect the requirements for the aforementioned building. The acquired perspective image or spatial data is given to a trained model that transforms perspective images or spatial data using the features of the reference data. The image transformed by the aforementioned trained model is output as the transformed perspective image. How to create architectural images.

5. On the computer, Obtain a perspective image of a building from a specific viewpoint, either indoors or outdoors, or spatial data indicating the distance to equipment or furniture shown in the perspective image. We accept reference data to reflect the requirements for the aforementioned building. The acquired perspective image or spatial data is given to a trained model that transforms perspective images or spatial data using the features of the reference data. The image transformed by the aforementioned trained model is output as the transformed perspective image. A computer program that executes a process.