Building detection system, building detection method, and control program

The system improves building detection accuracy by using multiple trained models tailored to geometric features, integrating their outputs to enhance detection precision and facilitate better analysis of building changes.

JP2025127839APending Publication Date: 2025-09-02PASCO CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024024772
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-02-21
Publication Date
2025-09-02

AI Technical Summary

Technical Problem

Conventional building detection systems struggle to accurately identify buildings due to variations in building characteristics, such as size and shape, leading to suboptimal detection accuracy.

Method used

A building detection system that utilizes multiple trained models, each tailored to specific geometric features like aspect ratio, degree of depression, and circularity, integrating their outputs to enhance detection accuracy.

Benefits of technology

The system achieves higher accuracy in detecting buildings by classifying and training models based on geometric features, reducing missed detections and enabling better analysis of building changes and attributes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025127839000001_ABST
    Figure 2025127839000001_ABST
Patent Text Reader

Abstract

To provide a building detection system, a building detection method, and a control program capable of detecting a building more accurately.SOLUTION: A building detection system includes: acquisition means which acquires an input image obtained by imaging a building; integration means which inputs the input image to a plurality of building detectors corresponding to multiple groups, respectively, categorized based on geometric features including at least shape features of the building, to integrate output data output from the building detectors; and output means which outputs information related to the integrated output data. Each of the building detectors is a trained model which has been trained using learning images including buildings having corresponding geometric features and ground-truth data corresponding to the buildings.SELECTED DRAWING: Figure 6
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to a building detection system, a building detection method, and a control program. [Background technology]

[0002] BACKGROUND ART Conventionally, research has been conducted into techniques for detecting buildings from images captured from the air, such as aerial photographs or satellite images, for purposes such as interpreting building changes, creating maps, and creating city models.

[0003] For example, Patent Document 1 discloses a building extraction system in which building detectors for each of a plurality of area ranges are used to detect building regions from data to be processed. That is, in the building extraction system described in Patent Document 1, buildings in a training image are classified by area, and a building detector is trained for each group of classified buildings, thereby preparing a building detector suited to the area characteristics of the building. Then, building regions are detected by inputting the image, which is data to be processed, into the building detectors for the plurality of groups and integrating the outputs obtained. [Prior art documents] [Patent documents]

[0004] [Patent Document 1] Japanese Patent Application Publication No. 2019-175139 Summary of the Invention [Problem to be solved by the invention]

[0005] However, because buildings can be small and narrow, or large and narrow, there can be variations in the characteristics of buildings classified by area alone. Therefore, there is room for improvement in the conventional technology to detect buildings with higher accuracy.

[0006] An object of the present invention is to provide a building detection system, a building detection method, and a control program that are capable of detecting buildings with higher accuracy. [Means for solving the problem]

[0007] The building detection system of the present invention comprises an acquisition means for acquiring an input image of a building, an integration means for inputting the input image to a plurality of building detectors corresponding to each of a plurality of groups classified based on geometric features including at least the shape features of the building, and integrating the output data output from each of the plurality of building detectors, and an output means for outputting information related to the integrated output data, wherein each of the plurality of building detectors is a trained model trained using a training image including a building having the corresponding geometric features and correct answer data corresponding to that building.

[0008] In the building detection system according to the present invention, the geometric features preferably include a plurality of types of shape features of the building, or a shape feature and an area of ​​the building.

[0009] In the building detection system according to the present invention, the shape feature is preferably the aspect ratio, degree of depression, or degree of circularity of the building.

[0010] Furthermore, in the building detection system of the present invention, it is preferable that each of the multiple building detectors is a trained model that has been trained to minimize the error between the training output data that is output when a training image containing a building with corresponding geometric features is input, and the correct answer data that indicates the area of ​​the building with that geometric feature within the training image.

[0011] Furthermore, it is preferable that the building detection system of the present invention further comprises a classification means for classifying multiple groups based on the geometric features of the buildings, and a learning means for training the multiple building detectors so as to minimize the error between the learning output data output when a learning image containing a building with corresponding geometric features is input and the correct answer data indicating the area of ​​the building with those shape features.

[0012] In addition, the building detection method of the present invention includes obtaining an input image of an object, inputting the input image to a plurality of building detectors corresponding to each of a plurality of groups classified based on geometric features including at least the shape features of the building, integrating the output data output from each of the plurality of building detectors, and outputting information related to the integrated output data, wherein each of the plurality of building detectors is a trained model trained using a training image including a building having the corresponding geometric features and ground truth data corresponding to that building.

[0013] In addition, the control program of the present invention is a control program for a building detection device, and causes the building detection device to acquire an input image of a building, input the input image to a plurality of building detectors corresponding to each of a plurality of groups classified based on geometric features including at least the shape features of the building, integrate the output data output from each of the plurality of building detectors, and output information related to the integrated output data, and each of the plurality of building detectors is a trained model trained using a training image including a building with corresponding geometric features and correct answer data corresponding to that building. [Effects of the Invention]

[0014] The building detection system, building detection method, and control program monitoring system according to the present invention are capable of detecting buildings with higher accuracy. [Brief explanation of the drawings]

[0015] [Figure 1] 1 is a configuration diagram of an example of a building detection system 1. FIG. [Figure 2] 10 is a flowchart illustrating an example of the flow of a learning process. [Figure 3] FIG. 10 is a schematic diagram for explaining shape features. [Figure 4] FIG. 2 is a schematic diagram for explaining a learning detector 112. [Figure 5] FIG. 10 is a schematic diagram for explaining a scale. [Figure 6]10 is a flowchart illustrating an example of the flow of a detection process. DETAILED DESCRIPTION OF THE INVENTION

[0016] Various embodiments of the present invention will be described below with reference to the drawings. It should be noted that the technical scope of the present invention is not limited to these embodiments, but encompasses the inventions set forth in the claims and their equivalents.

[0017] FIG. 1 is a diagram showing an example of a building detection system 1 according to the present invention.

[0018] The building detection system 1 includes a building detection device 100 and a server device 200. The building detection device 100 and the server device 200 are communicatively connected to each other via a network N. The network N is an intranet, the Internet, or the like.

[0019] The building detection device 100 is a personal computer, a notebook personal computer, a server, etc. The building detection device 100 includes an operation device 101, a display device 102, a communication device 103, a storage device 110, a processing circuit 120, etc.

[0020] The operation device 101 has input devices such as a keyboard and a mouse, and an interface circuit that acquires signals from the input devices, accepts operations by a user, and outputs to the processing circuit 120 a signal according to the user's input.

[0021] The display device 102 is an example of an output unit. The display device 102 has a display configured with a liquid crystal display, an organic electroluminescence display, or the like, and an interface circuit that outputs image data to the display, and displays image data on the display in accordance with instructions from the processing circuit 120.

[0022] The communication device 103 is an example of an output unit. The communication device 103 includes a wired or wireless communication interface circuit and connects the building detection device 100 to a communication network. The communication device 103 performs wired communication in accordance with a communication protocol such as TCP / IP (Transmission Control Protocol / Internet Protocol). The communication device 103 may also perform wireless communication using a wireless communication method conforming to the IEEE (Institute of Electrical and Electronics Engineers) 802.11 standard. The communication device 103 transmits information supplied from the processing circuit 120 to an external device. The communication device 103 also supplies information received from an external device to the processing circuit 120.

[0023] The storage device 110 includes, for example, semiconductor memory such as RAM (Random Access Memory) and ROM (Read Only Memory), a fixed disk device such as a hard disk, or a portable storage device such as an optical disk. The storage device 110 stores computer programs, data, and the like used for processing by the processing circuit 120. The computer programs are installed in the storage device 110 from a server (not shown) via the communication device 103. Note that the computer programs may be installed in the storage device 110 from a computer-readable portable recording medium using a known setup program or the like. The portable recording medium is, for example, a CD-ROM, a DVD-ROM, or the like. The computer programs may be distributed from a server or the like and installed in the storage device 140.

[0024] Furthermore, the storage device 110 stores a first building detector 111-1, a second building detector 111-2, ..., an N-th building detector 111-N, and a training detector 112. The first building detector 111-1, the second building detector 111-2, ..., the N-th building detector 111-N are examples of multiple building detectors, and detect buildings from images in which buildings are captured. Hereinafter, the first building detector 111-1, the second building detector 111-2, ..., the N-th building detector 111-N may be collectively referred to as the building detector 111. N is a preset integer of 2 or greater, and the building detector 111 includes the same number of building detectors as the number N of groups, which will be described later. The training detector 112 is used to generate the building detector 111.

[0025] The processing circuit 120 is, for example, a CPU (Central Processing Unit). The processing circuit 120 may be an LSI (Large Scale Integration), an ASIC (Application Specific Integrated Circuit), a DSP (Digital Signal Processor), an FPGA (Field Programmable Gate Array), or the like. The processing circuit 120 is connected to the operation device 101, the display device 102, the communication device 103, the storage device 110, and the like, and controls each of these components. The processing circuit 120 reads a program stored in the storage device 110 and operates in accordance with the read program, thereby functioning as a classification means 121, a learning means 122, an acquisition means 123, an integration means 124, and an output control means 125. The processing circuit 120 uses a building detector 111 to detect buildings from images of buildings.

[0026] FIG. 2 is a flowchart showing an example of the flow of the learning process executed by the building detection device 100.

[0027] An example of the operation of the learning process of the building detection device 100 will be described below with reference to the flowchart shown in Fig. 2. The flow of the operation described below is executed mainly by the processing circuit 120 in cooperation with each element of the building detection device 100 based on a program stored in advance in the storage device 110.

[0028] First, the classification means 121 acquires a training dataset including a set of multiple training overall images and correct answer data (step S101). The classification means 121 acquires the training dataset by receiving it from the server device 200 via the communication device 103. Note that the training dataset may be pre-stored in the storage device 110, and the classification means 121 may acquire the training dataset by reading it from the storage device 110. Each training overall image is an aerial photograph, satellite image (or an orthoimage based on the aerial photograph or satellite image), or the like, of the ground surface, which is the target area for the building detection process, and includes buildings captured from the sky. The training dataset also includes information on the scale of each training overall image. The correct answer data indicates the shape of the buildings included in the training overall image. The correct answer data is, for example, image data including the same number of pixels as the number of pixels included in the training overall image, and each pixel indicating whether the corresponding pixel in the training overall image includes a building. In addition, the correct answer data includes polygon data (vector data with vertex coordinates arranged in order) representing the area of ​​each building and a unique building ID that identifies the building, making it possible to determine the correspondence between each building, the area of ​​the building (building area), and the pixels representing the building.

[0029] Each learning overall image includes one or more buildings, each having one or more types of geometric features. The geometric features of the buildings include at least the shape features of the buildings. The geometric features of the buildings may further include the area of ​​the buildings. The geometric features are, for example, a feature vector whose components are the parameters to be used from among the shape features and area. The shape features are the aspect ratio, degree of depression, or circularity of the building, etc. The parameters to be used for classification from among the aspect ratio, degree of depression, degree of circularity, and area are set in advance.

[0030] Fig. 3 is a schematic diagram for explaining shape characteristics, showing a building C having a hook-like shape as seen from above.

[0031] The aspect ratio is a parameter that indicates the length and narrowness of a building's shape. For example, the aspect ratio is the value obtained by dividing the short side E1 of the smallest rectangle (circumscribing rectangle) D1 that contains building C by the long side E2, as shown in the following formula (1). In other words, the closer the aspect ratio is to a square, the closer it is to 1, and the longer and narrower it is, the closer it is to 0. (Aspect ratio) = (short side of the smallest rectangle that can contain the building) / (long side of the smallest rectangle that can contain the building) (1)

[0032] The degree of depression is a parameter that indicates the size and / or number of depressions in the shape of a building. A depression is a part formed by an angle whose interior angle is greater than 180°. The degree of depression is, for example, the value obtained by dividing the area of ​​building C by the area of ​​the convex hull D2 of building C, as shown in the following equation (2). A convex hull is a shape that contains an object and has no depressions. In other words, the degree of depression is 1 when the shape of a building has no depressions, and approaches 0 as the size of the depressions in the shape of the building increases or the number of depressions increases. (Degree of depression) = (area of ​​building) / (area of ​​convex hull of building) (2)

[0033] The circularity is a parameter that indicates the degree of approximation of the building shape to a circle. For example, as shown in the following formula (3), the circularity is calculated by multiplying the area of ​​building C by 4π, taking the square root of the product, and dividing the result by the perimeter of building C (the sum of F1 to F6). In other words, the closer the building shape is to a circle, the closer the circularity is to 1, and the further the building shape is from a circle, the closer the circularity is to 0. (Circularity) = {4π × (area of ​​building)} 1 / 2 / (building perimeter) (3) The circularity may be calculated by multiplying the area of ​​building C by 4π and dividing the result by the square of the perimeter of building C (the sum of F1 to F6), as in the following formula (4). (Circularity) = 4π × (area of ​​building) / (perimeter of building) 2 (4)

[0034] By including the aspect ratio, degree of depression, or degree of circularity of a building in the geometric features, the building detection device 100 can appropriately classify buildings based on the geometric features and detect them with high accuracy.

[0035] Furthermore, the geometric features of a building preferably include multiple types of shape features of the building, or the shape features and area of ​​the building. By including multiple types of elements in the geometric features, the building detection device 100 can appropriately classify buildings based on the geometric features and detect them with higher accuracy.

[0036] Next, the classification means 121 classifies the buildings included in each learning overall image into multiple groups based on the geometric features of the buildings (step S102). The classification means 121 uses the ground truth data to identify building regions corresponding to each building in each learning overall image. The classification means 121 calculates a feature vector, whose components are parameters used among the aspect ratio, degree of depression, degree of circularity, and area of ​​each identified building region, as the geometric features of each building. The classification means 121 classifies the buildings included in each learning image into a predetermined number (N) of groups based on the calculated geometric features using a clustering method such as the k-means method. The classification means 121 may also classify the buildings using other methods such as DBSCAN, Gaussian Mixture Model, and Mean Shift. The classification means 121 associates the building ID of each building with the group number (1 to N) into which the building is classified, and stores the associations in the storage device 110.

[0037] Furthermore, the classification means 121 stores group identification criteria for identifying the group to which each building belongs in association with the group number in the storage device 110. For example, the classification means 121 stores the center of gravity (cluster center) of each group as the group identification criteria in the storage device 110. For example, when k-means is used, the center of gravity of each group is the average vector of the geometric features (feature vectors) of the buildings belonging to each group.

[0038] Next, the learning means 122 causes the learning detector 112 to learn (step S103).

[0039] FIG. 4 is a schematic diagram for explaining the learning detector 112. As shown in FIG.

[0040] The training detector 112 is a model using a neural network. As shown in FIG. 4, the training detector 112 includes a common unit 113 and an individual unit 114. The common unit 113 is trained commonly regardless of the geometric features of the building, while the individual unit 114 is trained according to the geometric features of the building. The individual unit 114 includes a first type unit 114-1, a second type unit 114-2, ..., an Mth type unit 114-M. Each type unit is set for each combination of a model type and a scale variation of an image input to the training detector 112. For example, if there are two model types and three scale variations and they are combined in a round-robin manner, the number M of type units 114 is six. The model type indicates the type of neural network in each type unit. The model type includes a model that combines a convolutional layer and a pooling layer in a CNN (Convolutional Neural Network), a model that uses a convolutional layer that performs dilated convolution operations, and the like. The model type may include models other than CNN, such as a Transformer model or a CRF (Conditional Random Field). A reference scale is determined in advance, and scale variations are determined in advance to be 1.0, 0.5, 2.0 times the reference scale, etc. The learning means 122 cuts out multiple window images from the entire training image as training images and inputs them to the training detector 112. At this time, the learning means 122 adjusts the scale of the entire training image as necessary before cutting out the training images.

[0041] FIG. 5 is a schematic diagram for explaining the scale.

[0042] FIG. 5 shows training images R1, R2, and R3, which are window images including buildings C1, C2, and C3. In this example, the scale of the training image matches the reference scale, and three scale variations are set: 1.0, 0.5, and 2.0 times the reference scale. Training image R1 is an image extracted from the training image acquired by the classification means 121 without changing the scale. Training image R2 is an image extracted from the training image acquired by the classification means 121 after reducing it by thinning out pixels so that the number of pixels in the width and height directions is halved. Training image R3 is an image extracted from the training image acquired by the classification means 121 after enlarging it by linear interpolation or the like so that the number of pixels in the width and height directions is doubled. Training image R2 is an image with a scale 0.5 times that of training image R1, and training image R3 is an image with a scale 2.0 times that of training image R1. 5, the number of pixels in each training image is the same (Px × Py), and the larger the scale of each training image, the larger the physical size of the object contained in one pixel. Px and / or Py are, for example, 16, 32, or 64, and are determined in advance when modeling the training detector 112.

[0043] Returning to Figure 4, each of the various classification sections includes a first feature section 115-1, a second feature section 115-2, ..., an Nth feature section 115-N, etc. Each feature section corresponds to a group classified based on the geometric features of the building. That is, the number N of feature sections included in each classification section is the same as the number of groups.

[0044] The common unit 113 includes the input layer and part of the intermediate layer of the learning detector 112, and the individual unit 114 includes part of the intermediate layer and the output layer of the learning detector 112. A learning image cut out from the entire learning image is input to the common unit 113, and information output from the common unit 113 is input to all feature units of all classification units. Each feature unit outputs learning output data related to buildings included in the learning image. The learning output data is, for example, image data including the same number of pixels as the number of pixels included in the learning image, with each pixel indicating the probability of a building being present at the position of each pixel.

[0045] The learning means 122 first enlarges or reduces the learning overall image as needed to match the scale of each separate part. Note that a plurality of learning overall images corresponding to the scale of each separate part may be prepared in advance, and in step S101, the classification means 121 may acquire a plurality of learning overall images corresponding to the scale of each separate part.

[0046] Next, the learning means 122 extracts and acquires a plurality of window images as learning images from each of the scale-adjusted learning whole images. The learning means 122 randomly selects a plurality of positions from one learning whole image, and extracts learning images from each of the selected positions. Note that the learning means 122 may extract learning images so that buildings are always included, or so that the total number of buildings per group included in all learning images is equal.

[0047] Next, the learning means 122 inputs each training image into the input layer of the common unit 113 and compares the training output data output from each feature unit of the individual unit 114 with the correct answer data, thereby training the training detector 112. When each training image is input, the learning means 122 calculates the error between the training output data output from each feature unit (in each individual unit) corresponding to a group of geometric features of the building included in each training image and the correct answer data corresponding to each training image. Based on the calculated error, the learning means 122 changes the values ​​of parameters such as weights in each feature unit by error backpropagation or the like. Furthermore, the learning means 122 accumulates the errors to be propagated from the top layer of each feature unit to the bottom layer of the common unit 113, and based on the accumulated errors, changes the values ​​of parameters such as weights in the common unit 113 by error backpropagation or the like.

[0048] That is, when each training image is input, the learning means 122 trains each feature portion so that the error between the training output data output from the feature portion corresponding to the group of geometric features of the buildings included in each training image and the correct answer data corresponding to each training image is minimized. On the other hand, the learning means 122 does not take into account the training output data output from the feature portion that does not correspond to the group of geometric features of the buildings included in each training image. In this way, each feature portion is trained to be suitable for detecting buildings having the corresponding geometric features.

[0049] Next, the learning means 122 acquires an evaluation dataset including a set of multiple evaluation overall images and correct answer data (step S104). The learning means 122 acquires the evaluation dataset by receiving it from the server device 200 via the communication device 103. The evaluation dataset may be pre-stored in the storage device 110, and the classification means 121 may acquire the evaluation dataset by reading it from the storage device 110. Each evaluation overall image is an aerial photograph, satellite image (or an orthoimage based on the aerial photograph or satellite image), or the like, of the ground surface, which is the target area for the building detection process, and includes buildings captured from the sky. The correct answer data indicates the shape of the buildings included in the evaluation overall image. The correct answer data is, for example, image data that includes the same number of pixels as the number of pixels included in the evaluation overall image, and each pixel indicates whether the corresponding pixel in the evaluation overall image includes a building. Each evaluation overall image includes one or more buildings, each having one or more types of geometric features. The correct answer data also includes polygon data representing the area of ​​each building and a unique building ID that identifies the building. The evaluation dataset also includes information on the scale of each evaluation full image. The evaluation full images may be the same as the learning full images. Some or all of the data pairs in the evaluation dataset (pairs of the entire training image and correct answer data or pairs of the training image and correct answer data) may be the same as the data pairs in the training dataset, or they may be entirely different data pairs.

[0050] Next, the learning means 122 inputs each evaluation whole image to the learning detector 112 learned in step S103, and acquires evaluation output data output from the learning detector 112 (step S105).

[0051] The learning means 122 first enlarges or reduces the evaluation overall image to match the scale of the various separate parts. Note that a plurality of evaluation overall images corresponding to the scale of the various separate parts may be prepared in advance, and the learning means 122 may acquire a plurality of evaluation overall images corresponding to the scale of the various separate parts in step S104.

[0052] Next, the learning means 122 cuts out a plurality of window images from each of the scale-adjusted evaluation whole images to obtain them as evaluation images. The learning means 122 cuts out the evaluation images from the evaluation whole images while shifting the cut-out positions so that the evaluation images do not overlap each other and all of the areas included in the evaluation whole images are included in one of the evaluation images. Note that the evaluation images may be cut out so that they overlap each other.

[0053] Next, the learning means 122 inputs each evaluation image to the input layer of the common part 113 and acquires evaluation output data output from each feature part of the individual part 114.

[0054] Next, the learning means 122 generates binary evaluation data for each of the multiple overall evaluation images and for each feature portion based on the corresponding evaluation output data (step S106). The binary evaluation data is a binary image with the same number of pixels as the overall evaluation image. If the existence probability indicated for each pixel in the evaluation output data is equal to or greater than a predetermined threshold, the learning means 122 sets a value indicating that a building is included as the pixel value of the corresponding pixel in the binary evaluation data. On the other hand, if the existence probability indicated for each pixel in the evaluation output data is less than the predetermined threshold, the learning means 122 sets a value indicating that a building is not included as the pixel value of the corresponding pixel in the binary evaluation data. Note that if the evaluation images are cropped so as to overlap with each other, the learning means 122 sets pixel values ​​for the overlapping positions by comparing the maximum existence probability indicated for corresponding pixels in each evaluation output data with a predetermined threshold. Instead of the maximum value, an average value or a median value may be used, and which value is to be used is determined in advance. In this way, by the process of step S106, M×N sets of binary evaluation data corresponding to each feature portion are generated.

[0055] Next, the learning means 122 identifies the feature part with the highest accuracy for each group classified based on the geometric features of the buildings (step S107). The learning means 122 first uses each correct answer data to identify a building area corresponding to each building in each evaluation entire image. Next, the learning means 122 calculates the geometric features of each identified building.

[0056] Next, the learning means 122 identifies the group to which each building belongs based on the group identification criteria stored in step S102. For example, the learning means 122 identifies the group whose center of gravity stored in step S102 is closest to the geometric feature of each building as the group to which each building belongs.

[0057] Next, the learning means 122 calculates a matching rate for each feature part belonging to each group for each group. To this end, the learning means 122 first calculates, as a matching rate for each feature part, the proportion of pixels in the evaluation binary data based on the evaluation output data output from each feature part belonging to each group for each evaluation overall image that indicate that a building belonging to each group is included among the pixels in the correct answer data for all evaluation overall images that indicate that a building is included. This calculates M × N matching rates corresponding to each evaluation binary data, i.e., each feature part. The learning means 122 then identifies the feature part with the highest matching rate in each group as the feature part with the highest accuracy. This identifies N features.

[0058] Next, the learning means 122 generates a building detector 111 by combining the common part 113 with each feature part identified as the feature part with the highest accuracy (step S108), and then ends the series of steps. To do this, the learning means 122 first identifies the type part with the highest matching rate for each of the first feature part 115-1, the second feature part 115-2, ..., the Nth feature part 115-N. Then, the learning means 122 generates each building detector by combining the common part 113 with the feature part included in the identified type part for each of the first feature part 115-1, the second feature part 115-2, ..., the Nth feature part 115-N. That is, the learning means 122 generates the first building detector 111-1 by combining the common part 113 with the first feature part 115-1 of the type part with the highest matching rate among the first feature parts 115-1 included in each type of feature part. Similarly, the learning means 122 generates the Nth building detector 111-N by combining the common part 113 with the Nth feature part 115-N of the type part with the highest degree of match among the Nth feature parts 115-N included in the various type parts.

[0059] In the learning detector 112, the common unit 113 may be an extraction unit that extracts image features without learning and outputs them to each feature unit. In this case, the image features extracted by the extraction unit may be, for example, HOG (Histograms of Oriented Gradients) features, SIFT (Scaled Invariance Feature Transform) features, etc. Furthermore, the common part 113 may be omitted from the learning detector 112. In this case, the learning means 122 generates a building detector so that each of the first feature part 115-1, the second feature part 115-2, ..., the Nth feature part 115-N includes only feature parts included in the identified type part. When the common part 113 is omitted, a learning image (or an input image) may be directly input to each feature part.

[0060] In this way, the learning means 122 trains each building detector 111 using a training image containing a building having the corresponding geometric feature and the correct answer data corresponding to that building. In particular, the learning means 122 trains each building detector 111 so as to minimize the error between the training output data output when the training image is input and the correct answer data indicating the area of ​​the building having the geometric feature within the training image.

[0061] In other words, each building detector 111 is a trained model trained using training images containing buildings with corresponding geometric features and correct answer data corresponding to those buildings. In particular, each building detector 111 is a trained model trained to minimize the error between training output data output when that training image is input and correct answer data indicating the area of ​​the building with that geometric feature within that training image.

[0062] That is, each building detector 111 is generated corresponding to each of a plurality of groups classified based on the geometric features of buildings. The building detection device 100 generates a building detector for each of a plurality of types of geometric features that can detect buildings having each geometric feature with high accuracy, and can detect buildings having each geometric feature with high accuracy.

[0063] FIG. 6 is a flowchart showing an example of the flow of the detection process executed by the building detection device 100.

[0064] An example of the operation of the detection process of the building detection device 100 will be described below with reference to the flowchart shown in Fig. 6. The flow of the operation described below is executed mainly by the processing circuit 120 in cooperation with each element of the building detection device 100 based on a program stored in advance in the storage device 110.

[0065] First, the acquisition means 123 acquires an input overall image in which one or more buildings are captured (step S201). The acquisition means 123 acquires the input overall image by receiving it from the server device 200 via the communication device 103. The input overall image is an aerial photograph or satellite image (or an orthoimage based on the aerial photograph or satellite image) of the ground surface, which is the target area for the building detection process, and includes buildings captured from the sky. The acquired information includes the scale of the input overall image.

[0066] Next, the acquisition means 123 acquires an input image in which one or more buildings are captured (step S202). The acquisition means 123 enlarges or reduces the input overall image to match the scale of each building detector 111, i.e., the scale of the classification section that was the basis for each building detector 111. Next, the acquisition means 123 cuts out and acquires multiple window images from the scale-adjusted input overall image as input images in which buildings are captured. The acquisition means 123 cuts out input images from the input overall image while shifting the cut-out positions so that the input images do not overlap each other and so that all areas included in the input overall image are included in one of the input images. Note that the input images may be cut out so that they overlap each other.

[0067] Next, the integration means 124 inputs each input image to each building detector 111 (inputs each to all of the building detectors 111), and acquires output data output from each building detector 111 (step S203).

[0068] Next, the integration means 124 generates overall data by combining the output data from each building detector, for each of the first building detector 111-1, the second building detector 111-2, ..., the Nth building detector 111-N (step S204). The overall data is a multi-valued image whose number of pixels is the same as the number of pixels in the input overall image scaled to each building detector 111. The integration means 124 sets the existence probability indicated by the pixel in the corresponding output data as the pixel value of each pixel in the overall data. Note that if the input images are cropped so that they overlap with each other, the integration means 124 sets the maximum value of the existence probability indicated by the pixel in each corresponding output data as the pixel value of the pixel in the overall data corresponding to the overlapping position. The average or median may be used instead of the maximum value. However, which value is used is predetermined to be the same as that used when generating evaluation binary data based on the evaluation output data.

[0069] Next, the integration means 124 integrates the generated overall data to generate integrated data (step S205). The integration means 124 first enlarges or reduces each overall data to match the scale of the input overall image. The integration means 124 identifies the maximum pixel value for each corresponding pixel in each overall data, and generates integrated data such that each identified maximum value becomes the pixel value of the corresponding pixel.

[0070] In this way, the integrating means 124 integrates the output data output from each building detector 111 to generate integrated data. The integrated data is an example of integrated output data. In other words, "integrating output data" is not limited to simply merging pixel values ​​indicated in multiple output data, but also means combining information included in multiple output data.

[0071] As described above, each building detector 111 corresponds to a model type and / or scale that can detect buildings with corresponding geometric features with high accuracy. Therefore, buildings with corresponding geometric features are detected with high accuracy in the output data output from each building detector 111. The integration means 124 integrates the output data output from each building detector 111, thereby generating data in which buildings with various types of geometric features are collectively detected with high accuracy.

[0072] Next, the integration means 124 binarizes the generated integrated data to generate overall binary data (step S206). The overall binary data is a binary image with the same number of pixels as the number of pixels in the integrated data. If the existence probability (maximum value) indicated for each pixel in the integrated data is equal to or greater than a predetermined threshold, the integration means 124 sets a value indicating that a building is included as the pixel value of the corresponding pixel in the overall binary data. Furthermore, if the existence probability indicated for each pixel in the integrated data is less than the predetermined threshold, the integration means 124 sets a value indicating that a building is not included as the pixel value of the corresponding pixel in the overall binary data.

[0073] Next, the output control means 125 outputs information about the generated integrated data by transmitting it to an external device via the communication device 103 or by displaying it on the display device 102 (step S207), thereby completing the series of steps. The information about the integrated data is, for example, entire binary data. The information about the integrated data may be the integrated data itself. The information about the integrated data may be information indicating the center of gravity or outer edge position of a group of pixels in the integrated data that each indicate that a building is included.

[0074] As described above, the building detection system 1 inputs an input image to multiple building detectors corresponding to groups classified based on geometric features, and integrates the output data output from the multiple building detectors. This enables the building detection system 1 to detect buildings with various geometric features with high accuracy, making it possible to detect buildings with even higher accuracy.

[0075] Furthermore, the building detection system 1 can reduce the chances that users will miss buildings. Furthermore, users of the building detection system 1 can grasp the construction and destruction of buildings, and can obtain basic statistical information on house movements. Furthermore, users of the building detection system 1 can more easily grasp the changes in individual buildings over time. Furthermore, users of the building detection system 1 can easily determine the detailed attributes of buildings (for example, the type of building, such as a detached house, condominium, or factory) from the size and shape of the detected building area. Furthermore, by automating the process of extracting information about buildings from images, users of the building detection system 1 can quickly and inexpensively extract information from a wide area of ​​the ground.

[0076] Although preferred embodiments have been described above, the embodiments are not limited to these. For example, in step S105 of FIG. 2, the learning means 122 may acquire evaluation output data only from feature portions corresponding to geometric features of buildings included in the evaluation images, rather than acquiring evaluation output data from all feature portions. In this case, the learning means 122 identifies each building included in each evaluation image and identifies the group to which each identified building belongs, similar to the processing in step S107. The learning means 122 acquires evaluation output data only from feature portions corresponding to the identified group. This allows the building detection device 100 to reduce the processing load and processing time required to acquire each output data.

[0077] Furthermore, the learning method of each building detector 111 is not limited to minimizing the error between learning output data output when a learning image is input and ground truth data that indicates the area of ​​a building having that geometric feature in the learning image. For example, each building detector 111 may be trained to output whether or not a building exists in the learning image when the learning image is input, or the number of buildings that exist, etc.

[0078] Also, instead of the building detection device 100, the server device 200 may have the classification means 121 and / or the learning means 122, and may execute the learning process of FIG.

[0079] Alternatively, the server device 200 may store the building detector 111 instead of the building detection device 100. In that case, in step S203 of FIG. 6, the integrating means 124 transmits each input image to the server device 200 via the communication device 103. The server device 200 inputs each input image received from the building detection device 100 to each building detector 111 and transmits output data output from each building detector 111 to the building detection device 100. The integrating means 124 acquires the output data by receiving it from the server device 200 via the communication device 103. In that case, the server device 200 may have an acquiring means 123 instead of the building detection device 100 and execute the processing of steps S201 and / or S202 of FIG. 6 to acquire the entire input image and / or the input image.

[0080] It should be understood by those skilled in the art that various changes, substitutions, and alterations can be made to the present invention without departing from the spirit and scope of the present invention. For example, the above-described embodiments and modifications may be implemented in appropriate combinations within the scope of the present invention. [Explanation of symbols]

[0081] 1 Building detection system, 100 Building detection device, 102 Display device, 103 Communication device, 111-1 First building detector, 111-2 Second building detector, 111-N Nth building detector, 121 Classification means, 122 Learning means, 123 Acquisition means, 124 Integration means

Claims

1. an acquisition means for acquiring an input image of a building; an integration means for inputting the input image to a plurality of building detectors corresponding to a plurality of groups classified based on geometric features including at least shape features of buildings, and integrating output data output from each of the plurality of building detectors; an output means for outputting information relating to the integrated output data, Each of the plurality of building detectors is a trained model trained using a training image including a building having a corresponding geometric feature and ground truth data corresponding to the building. A building detection system comprising:

2. The building detection system according to claim 1 , wherein the geometric features include a plurality of types of shape features of a building, or a shape feature and an area of ​​a building.

3. The building detection system according to claim 1 or 2, wherein the shape characteristics are an aspect ratio, a degree of concavity, or a degree of circularity of the building.

4. 3. The building detection system of claim 1, wherein each of the plurality of building detectors is a trained model that is trained to minimize an error between training output data output when a training image including a building having a corresponding geometric feature is input and correct answer data indicating an area of ​​a building having the geometric feature within the training image.

5. a classification means for classifying the plurality of groups based on the geometric features of the buildings; The building detection system of claim 1 or 2, further comprising a learning means for training the plurality of building detectors so as to minimize an error between learning output data output when a learning image including a building having corresponding geometric features is input and correct data indicating the area of ​​the building having the shape feature.

6. Obtain an input image of a building, inputting the input image to a plurality of building detectors corresponding to a plurality of groups classified based on geometric features including at least a shape feature of a building, and integrating output data output from each of the plurality of building detectors; outputting information about the integrated output data; Each of the plurality of building detectors is a trained model trained using a training image including a building having a corresponding geometric feature and ground truth data corresponding to the building. A building detection method comprising:

7. A control program for a building detection device, Obtain an input image of a building, inputting the input image to a plurality of building detectors corresponding to a plurality of groups classified based on geometric features including at least a shape feature of a building, and integrating output data output from each of the plurality of building detectors; outputting information about the integrated output data; Each of the plurality of building detectors is a trained model trained using a training image including a building having a corresponding geometric feature and ground truth data corresponding to the building. A control program comprising:

Citation Information

Patent Citations

  • Architectural structure extraction system

    JP2019175139A