Building structure recognition system and building structure recognition method

The system addresses labor-intensive and accuracy-dependent building recognition by using machine learning to adapt to site conditions, removing noise, and providing real-time scanning feedback, enhancing structure recognition accuracy and efficiency.

JP7862801B2Active Publication Date: 2026-05-20SAI CORPORATION +1
View PDF 7 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
SAI CORPORATION
Filing Date
2022-08-18
Publication Date
2026-05-20

Smart Images

  • Figure 0007862801000001
    Figure 0007862801000001
  • Figure 0007862801000002
    Figure 0007862801000002
  • Figure 0007862801000003
    Figure 0007862801000003
Patent Text Reader

Abstract

This building interior structure recognition system comprises: a first machine learning model generation unit that uses BIM data and a first site image to generate a first machine-learned model; a second machine learning model generation unit that uses a second site image containing an image of a noise component to carry out relearning with respect to the first machine-learned model; a scanning unit that scans the interior of a building while determining whether scanning of a structure inside the building is successful, and acquires a third site image and 3D point cloud data of a structure inside the building; a noise component removal unit that extracts an image of a noise component; a third machine learning model generation unit uses the image of the noise component to carry out relearning with respect to the second machine-learned model; a building interior structure recognition unit that uses the third machine-learned model to recognize a structure inside the building; and a point group data output unit that extracts and outputs point cloud data of the structure inside the building.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an in-building structure recognition system and an in-building structure recognition method, and particularly to an in-building structure recognition system, an in-building structure recognition method, and a program for recognizing structures arranged inside a building such as a building using deep learning by a neural network.

Background Art

[0002] Conventionally, as a method for confirming the construction status of a building such as a building under construction, a human directly measures and confirms at the construction site using a two-dimensional construction drawing or the like, or compares it with a BIM (Building Information Modeling) model using a remote sensing technology capable of measuring distance using reflected light such as LiDER (Light Detection and Ranging).

[0003] However, when measuring with LiDER or the like, it is necessary to measure multiple locations at the construction site according to the on-site situation based on experience, and there is a problem that the accuracy of the data obtained depends on the proficiency of the measurer. In addition, there are problems such as the labor required for registering the obtained point cloud data and the labor required to manually identify structures inside the building such as pipes and measure their positions and sizes. There are also problems with the accuracy of the captured point cloud data and the processed data, and the difficulty of reusing the data.

[0004] Emphasizing data accuracy and measuring all locations at the construction site is difficult to adopt in reality because the amount of information becomes enormous. When the measurer has a high level of proficiency, it is possible to measure only the necessary locations based on their own experience, but due to variations due to proficiency and the need to improve the efficiency of measurement, automation of measurement is required.

[0005] To compare the construction status of a building site with the completed structure, it is expected that a pre-trained model using deep learning with a neural network will be used to automate the identification of the areas of structures placed at the construction site and the recognition of what those structures are.

[0006] To create a trained model for automating the recognition of structures in images, a sufficient number of construction site images are needed as input data for training. Furthermore, the ground truth data for training requires annotations of the structures contained in those images—that is, the recognition results of identifying what each part of the image represents. However, collecting a large number of actual construction site photographs suitable for training and annotating a vast number of structures to use as ground truth data is difficult.

[0007] Alternatively, instead of using actual photographs of the construction site, it is conceivable to perform machine learning using rendered images—3D models of the completed construction site rendered to closely resemble the actual appearance—to create a trained model. However, rendered images are primarily created for marketing purposes of buildings, and their production costs are high, making it difficult to prepare a sufficient number of rendered images for training. Furthermore, the annotation work required for the structures included in the rendered images is enormous and time-consuming if done manually.

[0008] Therefore, it is necessary to prepare a sufficient number of training images of construction sites for learning purposes, and to automate the annotation of structures included in those training images. Furthermore, it is necessary that the trained model created in this way can recognize structures with high accuracy.

[0009] Furthermore, in actual construction sites, various items are present in addition to the structure being measured, such as wire mesh, protective netting or sheets, temporarily installed iron fences or poles, garbage, and materials. Also, people may unintentionally appear in the scan during the site. These noise components interfere with the recognition of the structure being measured, affecting the accuracy of structure recognition.

[0010] Therefore, a system is needed that can accurately recognize site images that include noise components such as wire mesh, protective netting and sheets, temporarily installed iron fences and poles, debris, and materials present at the construction site. Furthermore, in actual construction sites, the presence of such noise components changes moment by moment as they are moved or added according to the construction progress, so it is desirable to use a model that matches the latest site conditions. In this regard, it is desirable to minimize the time and cost required to regenerate a model that matches the site.

[0011] Non-patent document 1 points out that in as-built modeling, which creates 3D models based on 3D measurements of existing large-scale equipment, the amount of point cloud data becomes enormous, and that "it is important to note that the measurement principles of the measurement devices used for as-built modeling of large-scale equipment differ from those of point cloud measurement devices for small parts. In point cloud measurement of small parts, triangulation is generally performed using a laser output device and a CCD camera, but with this method, the device also becomes larger as the size of the object increases. In addition, in the measurement of small parts, the number of point clouds measured is often only a few million at most, but in the case of large-scale equipment, a large amount of point cloud data is required for modeling."

[0012] For example, Patent Document 1 discloses a building production system comprising: an existing part survey means that converts digitized data of existing parts of a building obtained from existing drawings into 3D CAD data and stores it together with various on-site survey data including point cloud data obtained by a 3D laser scanner and a 3D polygon model created from the point cloud data; a construction member design means that places newly constructed member objects selected from member objects stored in a member library in advance onto the 3D polygon model; a member construction position output means that searches for and outputs member objects corresponding to a unique ID of the member object obtained by reading an electronic tag attached to a member that has been pre-cut at a member factory according to the member objects placed by the construction member design means, along with its construction position information, from the 3D CAD model designed by the construction member design means; and an automatic position indicating device that indicates the construction position of the member in the existing part based on the construction position information of the member object output by the member construction position output means of the CPU.

[0013] Furthermore, Patent Document 2 discloses an image processing apparatus comprising: an image acquisition unit that acquires an input image generated by imaging a real space using an imaging device; a recognition unit that recognizes the relative position and orientation between the real space and the imaging device based on the positions of one or more feature points reflected in the input image; an application unit that provides an augmented reality application using the recognized relative position and orientation; and a display control unit that superimposes guidance objects onto the input image according to the distribution of the feature points to guide a user operating the imaging device, so as to stabilize the recognition processing performed by the recognition unit.

[0014] However, while Patent Documents 1 and 2 both disclose technologies for grasping three-dimensional space or objects within three-dimensional space, they do not solve the problem of the enormous amount of data, such as three-dimensional point cloud data, that occurs in large-scale facilities such as buildings and factories, and are not suitable for automating the recognition of structures in images in order to quickly grasp the situation of a construction site in progress.

[0015] Furthermore, neither Patent Documents 1 nor 2 took into consideration the ability to recognize target structures with high accuracy even in site images that include noise components such as wire mesh, protective netting or sheets, temporarily installed iron fences or poles, garbage, and materials present at the construction site, nor did they consider the retraining of the model to match the latest site conditions. [Prior art documents] [Patent Documents]

[0016] [Patent Document 1] Japanese Patent Publication No. 2013-149119 [Patent Document 2] Japanese Patent Publication No. 2013-225245 [Non-patent literature]

[0017] [Non-Patent Document 1] Hiroshi Masuda, "Digitalization Technologies for Large-Scale Environments and Their Problems," Proceedings of the Japan Society for Precision Engineering Annual Conference (Proceedings of the Japan Society for Precision Engineering Symposium), Autumn 2007, pp. 81-84, September 3, 2007. [Overview of the project] [Problems that the invention aims to solve]

[0018] Therefore, the present invention solves the above problems and provides a building structure recognition system and method that can recognize target structures with high accuracy even in site images that include noise components such as wire mesh, protective netting or sheets, temporarily installed iron fences or poles, garbage, and materials present at the construction site, and that enables high-precision structural recognition that is adapted to the latest site conditions.

[0019] Furthermore, the present invention provides a program for causing a computer to execute each step of a method for recognizing structures inside a building. [Means for solving the problem]

[0020] To solve the above problems, the present invention provides a building structure recognition system for recognizing structures inside a building using a machine learning model, comprising: a first machine learning model generation unit that performs machine learning using BIM (Building Information Modeling) data and a first site image to generate a first machine learning model; a second machine learning model generation unit that retrains the first machine learning model using a second site image that includes images of noisy structures that do not have BIM data to generate a second machine learning model; a scanning unit that scans the inside of a building while determining the success or failure of scanning the structures inside the building, and acquires 3D point cloud data and a third site image inside the building; and an image of noisy structures from the third site image acquired by the scanning unit. The present invention provides a building structure recognition system comprising: a noise removal unit that removes images; a third machine learning model generation unit that retrains a second machine learning model using the image from which noise components have been removed by the noise removal unit to generate a third machine learning model; a building structure recognition unit that recognizes building structures contained in a third field image using the third machine learning model; and a point cloud data output unit that extracts and outputs point cloud data of building structures recognized by the building structure recognition unit from 3D point cloud data acquired by the scanning unit.

[0021] In a building interior structure recognition system according to an aspect of the present invention, a first machine learning model generation unit performs machine learning using, as correct data, an image generated from BIM data, and using, as observation data, an image obtained by processing an image generated by rendering BIM data with information obtained from a first on-site image, to generate a first trained machine learning model.

[0022] In a building interior structure recognition system according to an aspect of the present invention, a second machine learning model generation unit inputs a set of correct data and observation data for the second on-site image including an image of a noise component having no BIM data into the first trained machine learning model for retraining, to generate a second trained machine learning model.

[0023] In a building interior structure recognition system according to an aspect of the present invention, a scanning unit acquires an image inside a building and determines whether there is at least one corresponding reference point or a reference structure between consecutive frames, and when there is no at least one corresponding reference point or reference structure, notifies an alert prompting rescan.

[0024] In a building interior structure recognition system according to an aspect of the present invention, a noise component removal unit extracts an image of a noise component from a third on-site image acquired by the scanning unit by stereo matching, generates a mask image of the noise component, reconstructs the image so as to interpolate a portion of the mask image, and generates an image from which the noise component has been removed.

[0025] In a building interior structure recognition system according to an aspect of the present invention, a third machine learning model generation unit inputs a set of correct data and observation data for an image from which a noise component has been removed from the third on-site image obtained by the noise component removal unit into the second trained machine learning model for retraining, to generate a third trained machine learning model.

[0026] In a building structure recognition system according to one aspect of the present invention, the building structure recognition unit is characterized by inputting a third site image into a third machine learning model and recognizing structures inside the building included in the third site image.

[0027] Furthermore, the present invention provides a method for recognizing structures inside a building using a machine learning model, comprising the steps of: performing machine learning using BIM (Building Information Modeling) data and a first site image to generate a first machine-learned model; retraining the first machine-learned model using a second site image that includes images of noisy components that do not have BIM data to generate a second machine-learned model; scanning the inside of a building while determining the success or failure of scanning the structures inside the building to acquire 3D point cloud data of the structures inside the building and images of the inside of the building; removing images of noisy components from the images of the inside of the building acquired by the scanning unit; retraining the second machine-learned model using the images from which the noisy components obtained in the removal step have been removed to generate a third machine-learned model; recognizing structures inside the building using the third machine-learned model; and extracting and outputting point cloud data of structures inside the building recognized by the building structure recognition unit from the 3D point cloud data acquired by the scanning unit.

[0028] In a method for recognizing structures inside a building according to one aspect of the present invention, the scanning step is characterized by acquiring an image of the inside of the building, determining whether or not at least one corresponding reference point or reference structure exists, and notifying an alert prompting a rescan if at least one corresponding reference point or reference structure does not exist.

[0029] Furthermore, the present invention provides a program characterized by causing a computer to execute each step of the above-described method for recognizing structures inside a building.

[0030] In this invention, "BIM (Building Information Modeling) data" refers to data of a three-dimensional model of a building reproduced on a computer. [Effects of the Invention]

[0031] According to the present invention, it is possible to recognize on-site images that include noise components such as wire mesh, protective netting and sheets, temporarily installed iron fences and poles, garbage, and materials present at the construction site with high accuracy.

[0032] Furthermore, by retraining the model to match the latest on-site conditions in response to the constantly changing noise components, the accuracy of structural recognition can be improved.

[0033] Furthermore, the system can notify the user of the success or failure of the scan during the scanning process, and prompt them to perform a rescan on the spot if the scan is unsuccessful, thus avoiding the need to redo the scanning work at a later date. Other objects, features, and advantages of the present invention will become apparent from the following description of embodiments of the invention with respect to the accompanying drawings. [Brief explanation of the drawing]

[0034] [Figure 1] Figure 1 is a schematic diagram showing the overall structure recognition system inside a building according to the present invention. [Figure 2] Figure 2 shows the flow of each process in the building structure recognition system according to the present invention. [Figure 3] Figure 3 is a schematic diagram showing the first machine learning model generation unit of the present invention. [Figure 4] Figure 4 is a schematic diagram showing the second machine learning model generation unit of the present invention. [Figure 5] Figure 5 is a schematic diagram showing the scanning unit of the present invention. [Figure 6] Figure 6 is a schematic diagram showing the noise component removal unit of the present invention. [Figure 7]Figure 7 is a schematic diagram showing the third machine learning model generation unit of the present invention. [Figure 8] Figure 8 is a schematic diagram showing the building structure recognition unit of the present invention. [Figure 9] Figure 9 shows the overall flow of the building structure recognition method according to the present invention. [Modes for carrying out the invention] [Examples]

[0035] Figure 1 is a schematic diagram showing the overall structure recognition system 1 inside a building according to the present invention. The building structure recognition system 1 according to the present invention comprises: a first machine learning model generation unit 11 that performs machine learning using BIM (Building Information Modeling) data and a first site image to generate a first machine learning model; a second machine learning model generation unit 12 that retrains the first machine learning model M1 using a second site image that includes images of noisy structures that do not have BIM data to generate a second machine learning model M2; a scanning unit 20 that scans the inside of a building while determining the success or failure of scanning the structures inside the building and acquires 3D point cloud data and a third site image inside the building; and the scanning unit 20 acquires the inside of the building The system includes: a noise structure removal unit 30 that removes images of noise structures from the image; a third machine learning model generation unit 13 that retrains the second machine learning model using the image from which noise structures have been removed from the third field image obtained by the noise structure removal unit to generate a third machine learning model M3; a building structure recognition unit 40 that recognizes structures inside the building using the third machine learning model M3; and a point cloud data output unit 50 that extracts and outputs point cloud data of structures inside the building recognized by the building structure recognition unit from the 3D point cloud data acquired by the scanning unit.

[0036] The first machine learning model generation unit 11 performs machine learning using BIM (Building Information Modeling) data and the first site image to generate the first machine-trained model. Specifically, the first machine learning model generation unit 11 uses the image generated from the BIM data as ground truth data, and uses the image obtained by processing the image generated by rendering the BIM data with information obtained from the first site image as observation data to perform machine learning and generate the first machine-trained model.

[0037] The second machine learning model generation unit 12 retrains the first machine learning model M1 using a second field image that includes images of noisy structures for which BIM data is not available, and generates a second machine learning model M2. Specifically, the second machine learning model generation unit 12 inputs a set of ground truth data and observation data for the second field image that includes images of noisy structures for which BIM data is not available into the first machine learning model and retrains it, thereby generating a second machine learning model.

[0038] The scanning unit 20 scans the inside of the building while determining whether the scanning of structures inside the building is successful, and acquires 3D point cloud data and a third site image of the inside of the building. The scanning unit 20 acquires the image of the inside of the building and also determines whether or not there is at least one corresponding reference point or reference structure, and if at least one corresponding reference point or reference structure does not exist, it notifies an alert prompting a rescan.

[0039] The noise removal unit 30 extracts an image from which noise components have been removed from the image of the building acquired by the scanning unit 20. The noise removal unit 30 extracts images of noise components from the third on-site image acquired by the scanning unit 20 by stereo matching, generates a mask image of the noise components, and reconstructs the image by interpolating the portion of the mask image to generate an image from which the noise components have been removed.

[0040] The third machine learning model generation unit 13 uses the image from which noise components extracted by the noise component removal unit has been removed to retrain the second machine learning model and generate the third machine learning model M3. Specifically, the third machine learning model generation unit 13 inputs the set of ground truth data and observation data for the image from which noise components have been removed from the third field image, obtained by the noise component removal unit 30, into the second machine learning model and retrains it to generate the third machine learning model.

[0041] The building structure recognition unit 40 recognizes structures inside the building using a third machine learning model M3. The building structure recognition unit 40 inputs a third site image into the third machine learning model and recognizes structures inside the building included in the third site image.

[0042] The point cloud data output unit 50 extracts and outputs point cloud data of structures inside the building recognized by the building structure recognition unit from the 3D point cloud data acquired by the scanning unit.

[0043] Figure 2 shows the flow of each process in the building structure recognition system according to the present invention. Figure 2 shows the relationships between the following processes performed in the building structure recognition system: machine learning model generation, data acquisition / scanning, noise component removal, and building structure recognition. The machine learning model generation is performed by the first machine learning model generation unit 11, the second machine learning model generation unit 12, or the third machine learning model generation unit 13. The data acquisition / scanning is performed by the scanning unit 20 or the point cloud data output unit 50. Pre-processing before scanning may be performed by an external imaging device or scanning device not shown. The noise component removal is performed by the noise component removal unit 30. The building structure recognition is performed by the building structure recognition unit 40.

[0044] The overall processing flow can be broadly divided into pre-processing before scanning and processing after scanning. First, we will explain the pre-processing before performing the actual scanning at the site. Pre-processing before scanning includes generating a first machine learning model and retraining the first machine learning model to generate a second machine learning model. First, BIM data and first site images are acquired to generate the first machine learning model. The first machine learning model is generated using the acquired BIM data and first site images. The first machine learning model is created assuming an ideal site and can be used universally for various sites. However, in actual sites, there may be items that are not usually included in the BIM data (wire mesh, protective netting or sheets, temporarily installed iron fences or poles, garbage, materials, etc.). In addition to such items, unintended people may also be captured in the images. Therefore, a second site image containing such noise components that would be noise in recognizing structures inside the building is acquired, and the first machine learning model is retrained to learn that these noise components are noise. This retraining process yields a second machine-learned model that is tailored to the actual field conditions. The acquisition of this second field image is typically done during site inspections before the actual scanning takes place.

[0045] Next, we will explain the processing after the actual scanning at the site. The processing after scanning includes performing the scan, removing noise components, generating a third machine learning model, recognizing structures inside the building, and acquiring point cloud data. First, the inside of the building, which is the site, is scanned to acquire a third site image and 3D point cloud data. During this process, the success or failure of the scan is determined while scanning is in progress, and if a rescan is necessary, an alert prompting a rescan is sent. If an alert prompting a rescan is sent, a rescan is performed. This is repeated until all scan targets inside the building are scanned. Next, noise components are removed from the scanned third site image. Noise components present at the site often change moment by moment, and the third site image may contain new noise components that were not present in the second site image acquired in the pre-processing before scanning. Also, noise components that were present in the second site image may have moved to other locations. Unintended appearances of people, etc., can also become noise components. Therefore, noise components are removed from the scanned images, and the second machine learning model is retrained to further train the model on the images from which the noise components have been removed. Through this retraining, a third machine learning model is obtained that is adapted to the latest conditions of the actual site. Next, the third machine learning model is used to recognize structures inside the building in the third site image. Finally, 3D point cloud data corresponding to the recognized structures inside the building is obtained.

[0046] Figure 3 is a schematic diagram showing the first machine learning model generation unit of the present invention. The first machine learning model generation unit 11 uses images generated from BIM data as ground truth data (ground truth images) and images obtained by processing images generated by rendering BIM data using information obtained from the first site image as observation data (observation images) to perform machine learning and generate the first machine-learned model. The ground truth images generated from BIM data are those that distinguish structures in the image from the background. The ground truth images may be manually generated, for example, by manually filling in the parts of the image containing structures. The ground truth images may also be binarized images that can distinguish between parts of the structure and the background. The observation images are also generated from BIM data. First, a rendering image is generated by rendering the BIM data. Next, textures and other information extracted from the first site image are added to the rendering image to generate an image that is closer to an actual photograph of the site, and this is used as the observation image. The machine learning model generation process is performed using this set of ground truth images and observation images to generate the first machine-learned model M1. The first machine-learned model M1 can be used as a general-purpose machine-learned model for recognizing structures inside a building. In particular, when scanning a building in an ideal environment free of noise components and recognizing structures within the building, the first machine learning model M1 can be used.

[0047] Figure 4 is a schematic diagram showing the second machine learning model generation unit of the present invention. The second machine learning model generation unit 12 inputs a set of ground truth data and observation data for noisy structure images that do not have BIM data, which are included in the second field image, into the first machine learning model M1 and retrains it to generate the second machine learning model M2. When generating the first machine learning model M1, ground truth images and observation images are generated based on BIM data and used for machine learning, so the first machine learning model M1 is suitable for scanning an ideal building environment where no noisy structures exist. On the other hand, in actual field sites, noisy structures that do not have BIM data often exist. Therefore, by further training the first machine learning model M1 with the second field image which includes noisy structures, a second machine learning model M2 is generated that can accurately recognize structures even in a building environment which includes noisy structures. The set of ground truth images and observation images used for retraining the first machine learning model M1 is generated from the second field image. The ground truth image of the second field image is shown in a way that distinguishes the structures in the image from the background. The ground truth image of the second field image may be manually generated, for example, by manually filling in the structural parts within the image. The ground truth image of the second field image may also be a binarized image that can distinguish between structural parts and background parts. The observed image of the second field image may be the second field image including noise components as is. Alternatively, the observed image of the second field image may be an image that has been preprocessed as necessary from the second field image including noise components. Using such a set of the ground truth image and observed image of the second field image, the first machine learning model M1 is retrained to generate the second machine learning model M2.

[0048] Figure 5 is a schematic diagram showing the scanning unit of the present invention. The scanning unit 20 acquires images of the building interior and determines whether at least one corresponding reference point or reference structure exists. If at least one corresponding reference point or reference structure does not exist, it notifies an alert prompting a rescan. The scanning unit 20 may also have a scanning success / failure determination unit 203, an alert notification unit 204, and a rescan processing unit 205 for each function.

[0049] The scanning success / failure determination unit 203 determines whether at least one corresponding reference point or reference structure exists. Here, a "reference point" is a point that serves as a reference for matching continuous frames, and for example, a marker or the like may be attached to a structure inside a building to serve as a reference point. A "reference structure" is a structure that serves as a reference for matching continuous frames, and for example, a structure having a straight part such as the corner of a column may be used as the reference structure.

[0050] The alert notification unit 204 issues an alert prompting a rescan if at least one corresponding reference point or reference structure is not present. The alert is intended to notify the user that a rescan is necessary and may be displayed as an icon or message on the display screen of a scanning device such as a distance measuring scanner 201 or imaging device 202, or as a warning sound.

[0051] The rescan processing unit 205 receives a rescan instruction from the user and performs a rescan.

[0052] The scanning success / failure determination unit 203, the alert notification unit 204, and the rescan processing unit 205 repeat the processing until all scanning of the target is completed. Once scanning is complete, a third field image and 3D point cloud data are acquired. The third field image and 3D point cloud data acquired by the scanning unit 20 may contain information about noise components.

[0053] Figure 6 is a schematic diagram showing the noise component removal unit of the present invention. The noise removal unit 30 extracts images of noise components from the third field image acquired by the scanning unit 20 by stereo matching, generates a mask image of the noise components, reconstructs the image by interpolating the portion of the mask image, and generates an image from which the noise components have been removed. The noise removal unit 30 may also have a stereo matching unit 301, a mask image generation unit 302, and an image reconstruction unit 303 for each function.

[0054] The stereo matching unit 301 inputs two images (typically a right image and a left image) of the same object captured from different viewpoints from the third field image acquired by the scanning unit 20 into an existing machine learning model for stereo matching, estimates the three-dimensional depth for each pixel, and generates a distance image in which the estimated depth for each pixel is represented by a gradient. The existing machine learning model for stereo matching may be, for example, a machine learning model trained using a convolutional neural network (CNN) to obtain a mapping function that represents the parallax between two images (typically a right image and a left image) of the same object captured from different viewpoints. Alternatively, the existing machine learning model for stereo matching may use a hierarchical network equipped with recursive refinement that updates the parallax from coarse to fine, or a stacked cascade architecture for inference.

[0055] In other embodiments, the stereo matching unit 301 may use an existing stereo matching method that does not use a machine learning model, instead of the above-described method that uses an existing machine learning model for stereo matching. In this case, for example, two images (typically the right image and the left image) captured from two locations may be used to estimate the three-dimensional depth for each pixel, and a distance image may be generated in which the estimated depth for each pixel is represented by a gradient.

[0056] The mask image generation unit 302 performs thresholding on the distance image generated by the stereo matching unit 301 to generate a mask image of the noise components. For example, if a wire mesh, which is a noise component, is present in front of the structure to be recognized, a mask image of the wire mesh portion, which is the noise component, is generated.

[0057] The image reconstruction unit 303 removes the mask portion of the mask image generated by the mask image generation unit 302 from the original image and obtains a reconstructed image that interpolates the portion of the mask image. The image reconstruction unit 303 may also input the third field image containing noise components and the mask image of the noise components generated by the mask image generation unit 302 into an existing machine learning model for image reconstruction, and obtain a reconstructed image that removes the area of ​​the mask image and interpolates the portion of the mask image. The existing machine learning model for image reconstruction may be an existing machine learning model generated by deep learning using a neural network. As a result, the image reconstruction unit 303 obtains an image from which the noise components have been removed from the third field image.

[0058] Figure 7 is a schematic diagram showing the third machine learning model generation unit of the present invention. The third machine learning model generation unit 13 inputs the set of ground truth data and observation data for images from which noise components have been removed, extracted from images of the building interior extracted by the noise component removal unit 30, into the second machine learning model and retrains it to generate the third machine learning model. When generating the second machine learning model M2, the ground truth image and observation image of the second field image including noise components, acquired in the pre-processing before scanning, are used for retraining. However, in order to also handle new noise components that did not exist at the time of pre-processing before scanning, the third machine learning model M3 is generated, which is capable of recognizing structures that are even more appropriate to the target site, by additionally training the model with images from which noise components have been removed, acquired during the actual scanning, which include noise components. The set of ground truth images and observation images used for retraining the second machine learning model M2 is generated from images from which noise components have been removed, obtained by the noise component removal unit 30. The ground truth image of the image from which noise components have been removed is one in which structures in the image are distinguished from the background. The ground truth image of the image from which noise components have been removed may be manually generated, for example, by manually filling in the parts of the image containing structures. The ground truth image of the image from which noise components have been removed may also be a binarized image that can distinguish between the parts of the structure and the background. The observed image of the image from which noise components have been removed may be the image from which the noise components have been removed from a third field image containing noise components. Alternatively, the observed image of the image from which noise components have been removed may be an image that has been preprocessed as needed from the image from which the noise components have been removed from the third field image containing noise components. Using such a set of the ground truth image and the observed image of the image from which noise components have been removed, the second machine learning model M2 is retrained to generate the third machine learning model M3.

[0059] Figure 8 is a schematic diagram showing the building structure recognition unit and point cloud data output unit of the present invention. The building structure recognition unit 40 inputs the third field image to the third machine learning model M3 and recognizes the structures inside the building contained in the third field image. The output from the third machine learning model M3 is shown in a way that distinguishes the structures inside the building from the background in the image. The output from the third machine learning model M3 may be a binarized image that can distinguish between the parts of the building structure and the background. The third field image may contain noise components that were present when the site was scanned. The third machine learning model M3 has been further trained with images from which the noise components have been removed from the third field image containing noise components, and can accurately recognize structures inside the building even when the third field image contains noise components.

[0060] The point cloud data output unit 50 extracts 3D point cloud data corresponding to the recognized structures inside the building from the 3D point cloud data scanned by the scanning unit 20, based on the image information of the structures inside the building recognized by the building structure recognition unit 40, and outputs it as point cloud data. This provides 3D point cloud data of the structures inside the building, and by rendering or performing other operations on the obtained 3D point cloud data, it can be used for purposes such as generating 3D CAD data of the inside of the building including the structures.

[0061] The following describes a method for recognizing structures inside a building using a machine learning model. Figure 9 shows the overall flow of the building structure recognition method according to the present invention. The building structure recognition method according to the present invention includes: step S901, which involves machine learning using BIM (Building Information Modeling) data and a first site image to generate a first machine-learned model; step S902, which involves retraining the first machine-learned model M1 using a second site image that includes an image of a noisy structure that does not have BIM data to generate a second machine-learned model M2; a scanning step S903, which involves scanning the inside of the building while determining the success or failure of scanning the structures inside the building to acquire 3D point cloud data of the structures inside the building and an image of the inside of the building; and a scanning step S903, which involves scanning the inside of the building while determining the success or failure of scanning the structures inside the building to acquire 3D point cloud data of the structures inside the building and an image of the inside of the building. The present invention provides a method for recognizing structures inside a building, which includes the steps of: S904 removing images of noise components from indoor images; S905 retraining a second machine learning model M2 using the image from which the noise components obtained in the removal step have been removed to generate a third machine learning model M3; S906 recognizing structures inside the building using the third machine learning model M3; and S907 extracting and outputting point cloud data of structures inside the building recognized by the building structure recognition unit from 3D point cloud data acquired by the scanning unit.

[0062] In a method for recognizing structures inside a building according to one aspect of the present invention, the scanning step S903 acquires an image of the inside of the building and determines whether or not at least one corresponding reference point or reference structure exists. If at least one corresponding reference point or reference structure does not exist, an alert prompting a rescan is issued.

[0063] Furthermore, the present invention provides a program characterized by causing a computer to execute each step of the above-described method for recognizing structures inside a building. [Examples]

[0064] Below, as Example 2, we will describe an example in which the building structure recognition system 1 described in Example 1 does not have a second machine learning model generation unit 12, and an example in which the building recognition method described in Example 1 does not have a step of generating a second machine learning model. Points not specifically described below are the same as in the building structure recognition system 1 of Example 1.

[0065] The building structure recognition system for recognizing structures inside a building using the machine learning model in Example 2 is a Building Information Modeling (BIM) system. The system is characterized by comprising: a first machine learning model generation unit 11 that performs machine learning using modeling data and a first on-site image to generate a first machine learning model M1; a scanning unit 20 that scans the inside of a building while determining the success or failure of scanning the structures inside the building and acquires 3D point cloud data inside the building and a third on-site image; a noise structure removal unit 30 that removes images of noise structures from the images inside the building acquired by the scanning unit 20; a third machine learning model generation unit 13 that retrains the first machine learning model M1 using the images from which noise structures have been removed obtained by the noise structure removal unit 30 to generate a third machine learning model M3; an indoor structure recognition unit 40 that recognizes structures inside the building using the third machine learning model M3; and a point cloud data output unit 50 that extracts and outputs point cloud data of the structures inside the building recognized by the indoor structure recognition unit 40 from the 3D point cloud data acquired by the scanning unit 20.

[0066] The difference between the building structure recognition system 1 of Example 1 and Example 2 is that Example 2 does not have a second machine learning model generation unit 12 that retrains the first machine learning model using a second field image containing images of noisy components that do not have BIM data, thereby generating a second machine learning model. Furthermore, in Example 2, the third machine learning model generation unit 13 retrains the first machine learning model M1, not the second machine learning model M2, to generate a third machine learning model M3. That is, in Example 2, the third machine learning model generation unit 13 inputs a set of ground truth data and observation data for images from which noise components have been removed, extracted from images of the building interior extracted by the noise component removal unit 30, into the first machine learning model M1, retrains it, and generates a third machine learning model M3.

[0067] The set of ground truth images and observed images to be used for retraining the first machine learning model M1 is generated from the third field image from which noise components have been removed, obtained by the noise component removal unit 30. The ground truth image of the image from which noise components have been removed is shown in a way that distinguishes structures in the image from the background. The ground truth image of the image from which noise components have been removed may be manually generated, for example, by manually filling in the parts of the image containing structures. The ground truth image of the image from which noise components have been removed may also be a binarized image that can distinguish between the parts of the structure and the background. The observed image of the image from which noise components have been removed may be the image from which noise components have been removed from the third field image containing noise components as is. Alternatively, the observed image of the image from which noise components have been removed may be an image that has been preprocessed as needed from the image from which noise components have been removed from the third field image containing noise components. Using a set of ground truth images and observed images from which such noise components have been removed, the first machine learning model M1 is retrained to generate a third machine learning model M3.

[0068] The Building Structure Recognition Method for Recognizing Structures Inside a Building Using a Machine Learning Model in Example 2 is a building structure recognition method for recognizing structures inside a building using a machine learning model, and includes the steps of: S901, which involves performing machine learning using BIM (Building Information Modeling) data and a first site image to generate a first machine-learned model; S903, which involves scanning the inside of a building while determining the success or failure of scanning the structures inside the building to acquire 3D point cloud data of the structures inside the building and an image of the inside of the building; S904, which involves removing images of noise components from the image of the inside of the building acquired by the scanning unit; S905, which involves retraining the first machine-learned model using the image from which the noise components obtained in the removal step have been removed to generate a third machine-learned model; S906, which involves recognizing structures inside the building using the third machine-learned model; and S907, which involves extracting and outputting point cloud data of structures inside the building recognized by the building structure recognition unit from the 3D point cloud data acquired by the scanning unit.

[0069] The difference between the building structure recognition method of Example 1 and Example 2 is that it does not include step S902, which involves retraining the first machine-learned model using a second field image containing images of noisy components that do not have BIM data, in order to generate a second machine-learned model. Furthermore, in Example 2, step S905, which generates a third machine-learned model, involves retraining the first machine-learned model M1, not the second machine-learned model M2, in order to generate a third machine-learned model M3. That is, in Example 2, the set of ground truth data and observation data for the image from which the noise components have been removed from the third field image containing the noise components, obtained by the noise component removal unit 30, is input to the first machine-learned model M1 for retraining to generate a third machine-learned model M3.

[0070] As described above, the building structure recognition system and method according to the present invention enable the measurement of shape and position by focusing on components of interest at a construction site, thereby improving accuracy and speed. Furthermore, it reduces the amount of components that need to be managed at a construction site, significantly reducing the amount of data handled by the construction site component management system. Additionally, the building structure recognition system and method according to the present invention can recognize site images containing noise components such as wire mesh, protective netting and sheets, temporarily installed iron fences and poles, debris, and materials with high accuracy. Moreover, by retraining the model to match the latest site conditions in response to constantly changing noise components, the accuracy of structure recognition can be improved. Furthermore, the success or failure of scanning can be notified to the user during the scanning process, and if scanning is unsuccessful, a rescan can be initiated immediately, avoiding the need to redo the scanning work at a later date. Thus, according to the present invention, it is possible to minimize the time and cost required to generate a model tailored to various site-specific situations at actual construction sites. Although the above description concerns examples, it will be apparent to those skilled in the art that the present invention is not limited thereto, and various changes and modifications can be made within the scope of the principles of the present invention and the appended claims. [Explanation of Symbols]

[0071] 1. Building Structure Recognition System 11. First Machine Learning Model Generation Unit 12. Second Machine Learning Model Generation Unit 13. Third Machine Learning Model Generation Unit 20 Scanning Section 30 Noise component removal section 40 Building structure recognition unit 50-point cloud data output unit

Claims

1. A building structure recognition system for recognizing structures inside a building using a machine learning model, A first machine learning model generation unit performs machine learning using BIM (Building Information Modeling) data and first on-site images to generate a first machine learning model, A second machine learning model generation unit generates a second machine learning model by retraining the first machine learning model using a second field image that includes an image of a noisy structure without BIM data, A scanning unit scans the inside of a building while determining the success or failure of scanning the structures inside the building, and acquires 3D point cloud data and a third site image of the inside of the building. A noise component removal unit removes images of noise components from the third field image acquired by the scanning unit, A third machine learning model generation unit generates a third machine learning model by using the image from which the noise components have been removed, obtained by the noise component removal unit, to retrain the second machine learning model. A building structure recognition unit recognizes structures inside a building using the aforementioned third machine learning model, A point cloud data output unit extracts and outputs point cloud data of structures inside the building recognized by the building structure recognition unit from the three-dimensional point cloud data acquired by the scanning unit. A building structure recognition system characterized by comprising the following features.

2. A building structure recognition system for recognizing structures inside a building using a machine learning model, A first machine learning model generation unit performs machine learning using BIM (Building Information Modeling) data and first on-site images to generate a first machine learning model, A scanning unit scans the inside of a building while determining the success or failure of scanning the structures inside the building, and acquires 3D point cloud data and a third site image of the inside of the building. A noise component removal unit removes images of noise components from the third field image acquired by the scanning unit, A third machine learning model generation unit generates a third machine learning model by using the image from which the noise components have been removed, obtained by the noise component removal unit, to retrain the first machine learning model. A building structure recognition unit recognizes structures inside a building using the aforementioned third machine learning model, A point cloud data output unit extracts and outputs point cloud data of structures inside the building recognized by the building structure recognition unit from the three-dimensional point cloud data acquired by the scanning unit. A building structure recognition system characterized by comprising the following features.

3. The building structure recognition system according to claim 1 or 2, characterized in that the first machine learning model generation unit uses an image generated from the BIM data as ground truth data, processes the image generated by rendering the BIM data using information obtained from the first on-site image, performs machine learning on the resulting image as observation data, and generates a first machine learning model.

4. The building structure recognition system according to claim 1, characterized in that the second machine learning model generation unit inputs a set of ground truth data and observation data for the second field image, which includes an image of a noisy structure that does not have BIM data, into the first machine learning model to retrain it and generate the second machine learning model.

5. The building structure recognition system according to claim 1 or 2, characterized in that the scanning unit acquires an image of the inside of the building, determines whether or not there is at least one corresponding reference point or reference structure between consecutive frames, and notifies an alert prompting a rescan if there is at least one corresponding reference point or reference structure.

6. The building structure recognition system according to claim 1 or 2, characterized in that the noise component removal unit extracts images of the noise components from the third on-site image acquired by the scanning unit by stereo matching, generates a mask image of the noise components, reconstructs the image so as to interpolate the portion of the mask image, and generates an image from which the noise components have been removed.

7. The building structure recognition system according to claim 1, characterized in that the third machine learning model generation unit inputs the set of ground truth data and observation data for the image from which the noise components have been removed from the third field image, obtained by the noise component removal unit, into the second machine learning model to retrain it and generate a third machine learning model.

8. The building structure recognition system according to claim 2, characterized in that the third machine learning model generation unit inputs a set of ground truth data and observation data for an image from which the noise components have been removed by the noise component removal unit into the first machine learning model to retrain it and generate a third machine learning model.

9. The building structure recognition system according to claim 1 or 2, characterized in that the building structure recognition unit inputs the third site image to the third machine learning model and recognizes the building structure included in the third site image.

10. A method for recognizing structures inside a building using a machine learning model, The process involves a step of performing machine learning using BIM (Building Information Modeling) data and a first on-site image to generate a first machine-learned model, and The first machine learning model is retrained using a second field image that includes an image of a noisy structure that does not have BIM data, thereby generating a second machine learning model. The scanning step involves scanning the inside of the building while determining the success or failure of scanning the structures inside the building, and acquiring 3D point cloud data of the structures inside the building and a third site image. The steps include removing images of noise components from the images of the building interior acquired in the scanning step, The steps include: using the image from which the noise components obtained in the removal step have been removed, retraining the second machine learning model to generate a third machine learning model; The third step involves using the machine learning model described above to recognize structures inside the building, A step to extract and output point cloud data of the structures inside the building recognized in the recognition step from the three-dimensional point cloud data acquired in the scanning step. A method for recognizing structures inside a building, including the above.

11. The method for recognizing structures inside a building according to claim 10, characterized in that the scanning step acquires an image of the inside of the building, determines whether or not there is at least one corresponding reference point or reference structure between consecutive frames, and notifies an alert prompting a rescan if there is at least one corresponding reference point or reference structure.

12. A method for recognizing structures inside a building using a machine learning model, The process involves a step of performing machine learning using BIM (Building Information Modeling) data and a first on-site image to generate a first machine-learned model, and The scanning step involves scanning the inside of the building while determining the success or failure of scanning the structures inside the building, and acquiring 3D point cloud data of the structures inside the building and images of the inside of the building. The steps include removing images of noise components from the images of the building interior acquired in the scanning step, The steps include: using the image from which the noise components obtained in the removal step have been removed, retraining the first machine learning model to generate a third machine learning model; The third step involves using the machine learning model described above to recognize structures inside the building, A step to extract and output point cloud data of the structures inside the building recognized in the recognition step from the three-dimensional point cloud data acquired in the scanning step. A method for recognizing structures inside a building, including the above.

13. The method for recognizing structures inside a building according to claim 12, characterized in that the scanning step acquires an image of the inside of the building, determines whether there is at least one corresponding reference point or reference structure between consecutive frames, and notifies an alert prompting a rescan if there is no at least one corresponding reference point or reference structure.

14. A program characterized by causing a computer to perform each step of the method according to any one of claims 10 to 13.