Method and device for identifying buildings in hollow village, electronic equipment and storage medium

By employing a building classification and recognition model in hollow villages, and combining semantic segmentation and attention mechanism for feature extraction, the problems of time-consuming, labor-intensive, and low-accuracy recognition in traditional methods are solved, achieving efficient and accurate building classification.

CN115457384BActive Publication Date: 2026-02-03BEIJING UNIV OF CIVIL ENG & ARCHITECTURE
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202211020560.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-08-24
Publication Date
2026-02-03
Estimated Expiration
2042-08-24

AI Technical Summary

Technical Problem

Traditional methods for identifying buildings in hollow villages rely on manual operation, which is time-consuming and labor-intensive. Furthermore, the insufficient clarity of high-resolution satellite imagery makes identification difficult and results in a high false positive rate. Existing machine learning methods are not very accurate, shallow features are easily affected by the environment, and there are serious problems with missed detections at multiple scales.

Method used

A building classification and recognition model is adopted, which combines a semantic segmentation layer and a feature extraction layer. The attention mechanism is used to segment and extract features from building areas. Building classification information is obtained through machine learning training to enhance the depth and accuracy of feature extraction.

Benefits of technology

It improves the accuracy and efficiency of building identification in hollow villages, reduces the false positive rate, and achieves accurate classification of building categories.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115457384B_ABST
    Figure CN115457384B_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of image processing, and provides a method and device for identifying buildings in a hollow village, electronic equipment and a storage medium, comprising: determining an image of an identification area where a target village is located; inputting the image into a building classification identification model to obtain building classification information output by the building classification identification model, the building classification information being used to represent the categories of buildings in the target village; and the building classification identification model comprising a semantic segmentation layer and a feature extraction layer. The buildings in the image are classified by the building classification identification model, which comprises a semantic segmentation layer for segmenting building areas and a feature extraction layer for extracting features of each building area based on an attention mechanism. Through the operation process of area segmentation and feature extraction, more accurate image features can be extracted, and the accuracy of the classification of each building in the image is provided.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of image processing technology, and in particular to a method, apparatus, electronic device, and storage medium for identifying buildings in hollow villages. Background Technology

[0002] The revitalization of hollow villages is an important strategy for revitalizing abandoned land, optimizing land use, and improving the rural living environment. To this end, it is necessary to use technical means to identify and detect various types of buildings in hollow villages.

[0003] Traditional detection methods rely heavily on the professionalism of operators, are time-consuming and labor-intensive, and cannot achieve large-scale automated measurement; secondly, the types of measurement data are limited, making it impossible to quickly obtain the geomorphological features of the entire area.

[0004] The high resolution of high-resolution satellite imagery poses certain limitations to the feature recognition of buildings in hollowed-out villages. Detailed features of buildings in high-resolution satellite imagery (such as roof vegetation and roof damage) are small and not clearly visible, leading to greater difficulty in recognition, a higher false positive rate, and problems such as feature vanishing in deep network structures during model training.

[0005] Traditional machine learning extraction methods rely heavily on shallow features of remote sensing images, including spectral and morphological indices. While these methods have enabled automatic building extraction, they suffer from low accuracy, are highly susceptible to environmental influences, and are prone to multi-scale missed detections.

[0006] This shows that the current identification of buildings in hollow villages is not accurate enough. Summary of the Invention

[0007] In view of the problems existing in the prior art, the present invention provides a method, device, electronic device and storage medium for identifying buildings in hollow villages.

[0008] In a first aspect, the present invention provides a method for identifying buildings in hollow villages, comprising:

[0009] Image of the identification area where the target village is located;

[0010] The image is input into the building classification and recognition model to obtain building classification information output by the building classification and recognition model. The building classification information is used to characterize the category of each building in the target village.

[0011] The building classification and recognition model is a model trained by machine learning, which takes the image features of each building in the sample image and the category identifier of each building in the sample image as input, and is used to classify each building in the image.

[0012] The building classification and recognition model includes a semantic segmentation layer and a feature extraction layer. The semantic segmentation layer is used to segment each building region in the image; the feature extraction layer is used to extract features from each building region based on an attention mechanism.

[0013] In one embodiment, inputting the image into a building classification and recognition model to obtain building classification information output by the building classification and recognition model includes:

[0014] The image is input into the semantic segmentation layer, and the semantic segmentation layer outputs each building region.

[0015] Each building region is input into the feature extraction layer, which increases the extraction depth based on an attention mechanism. Based on the extraction depth, features are extracted from each building region, and the image features of each building region are output.

[0016] The image features of each building area are input into the classification layer of the building classification and recognition model, and the classification layer outputs the building classification information of each building area.

[0017] In one embodiment, the feature extraction layer includes a first operational unit and a second operational unit. Accordingly, each building region is input to the feature extraction layer, which increases the extraction depth based on an attention mechanism. Features are extracted from each building region according to the extraction depth, and image features of each building region are output, including:

[0018] Each building region is input into the first processing unit, which then generates a weight matrix based on the extraction depth and the pixel values ​​of each building region.

[0019] Each building area is input into the second processing unit, which determines the initial image features based on the pixel values ​​of each building area, and then determines the image features based on the initial image features and the weight matrix.

[0020] In one embodiment, the image of the identification area where the target village is located includes:

[0021] The original image of the identification area where the target village is located is acquired, and the original image is orthorectified to determine the image.

[0022] Secondly, the present invention provides a device for identifying buildings in hollow villages, comprising:

[0023] The determination module is used to determine the image of the recognition area where the target village is located.

[0024] The recognition module is used to input the image into the building classification recognition model to obtain building classification information output by the building classification recognition model. The building classification information is used to characterize the category of each building in the target village.

[0025] The building classification and recognition model is a model trained by machine learning, which takes the image features of each building in the sample image and the category identifier of each building in the sample image as input, and is used to classify each building in the image.

[0026] The building classification and recognition model includes a semantic segmentation layer and a feature extraction layer. The semantic segmentation layer is used to segment each building region in the image; the feature extraction layer is used to extract features from each building region based on an attention mechanism.

[0027] In one embodiment, the identification module is specifically used for:

[0028] The image is input into the semantic segmentation layer, and the semantic segmentation layer outputs each building region.

[0029] Each building region is input into the feature extraction layer, which increases the extraction depth based on an attention mechanism. Based on the extraction depth, features are extracted from each building region, and the image features of each building region are output.

[0030] The image features of each building area are input into the classification layer of the building classification and recognition model, and the classification layer outputs the building classification information of each building area.

[0031] In one embodiment, the feature extraction layer includes a first operational unit and a second operational unit. Correspondingly, in the process of the recognition module inputting each building region into the feature extraction layer, the feature extraction layer increasing the extraction depth based on an attention mechanism, extracting features from each building region according to the extraction depth, and outputting the image features of each building region, the specific processing is as follows:

[0032] Each building region is input into the first processing unit, which then generates a weight matrix based on the extraction depth and the pixel values ​​of each building region.

[0033] Each building area is input into the second processing unit, which determines the initial image features based on the extraction depth and the pixel values ​​of each building area, and then determines the image features based on the initial image features and the weight matrix.

[0034] In one embodiment, the determining module is specifically used for:

[0035] The original image of the identification area where the target village is located is acquired, and the image is determined by orthorectification based on the original image.

[0036] Thirdly, the present invention provides an electronic device, including a memory and a memory storing a computer program, wherein the processor executes the program to implement the steps of the method for identifying buildings in hollow villages as described in the first aspect.

[0037] Fourthly, the present invention provides a processor-readable storage medium storing a computer program for causing the processor to perform the steps of the method for identifying buildings in a hollow village as described in the first aspect.

[0038] The present invention provides a method, apparatus, electronic device, and storage medium for identifying buildings in hollow villages. The method classifies buildings in images using a building classification and recognition model. The building classification and recognition model includes a semantic segmentation layer for segmenting building regions and a feature extraction layer for extracting features from each building region based on an attention mechanism. Through the segmentation of regions and the feature extraction process, relatively accurate image features can be extracted, providing accuracy in classifying each building in the image. Attached Figure Description

[0039] To more clearly illustrate the technical solutions in this invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.

[0040] Figure 1 This is a flowchart illustrating the method for identifying buildings in hollow villages provided by the present invention;

[0041] Figure 2 This is a schematic diagram of the structure of the identification device for buildings in hollow villages provided by the present invention;

[0042] Figure 3 This is a schematic diagram of the structure of the electronic device provided by the present invention. Detailed Implementation

[0043] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.

[0044] Figure 1 This diagram illustrates a flowchart of a method for identifying buildings in a hollow village, as provided by the present invention. (See attached diagram.) Figure 1 The method includes:

[0045] 11. Determine the image of the recognition area where the target village is located;

[0046] 12. Input the image into the building classification and recognition model to obtain building classification information output by the model. The building classification information is used to characterize the category of each building in the target village. The building classification and recognition model is a model trained by machine learning based on the image features and category labels of each building in the sample image as input. It is used to classify each building in the image. The building classification and recognition model includes a semantic segmentation layer and a feature extraction layer. The semantic segmentation layer is used to segment each building region in the image. The feature extraction layer is used to extract features from each building region based on an attention mechanism.

[0047] Regarding steps 11 and 12, it should be noted that in this invention, the target village is a hollow village where building identification is to be performed. Hollow villages contain village houses with varying degrees of newness and damage. Therefore, it is necessary to classify the various types of buildings in the hollow village to obtain the category of each building. This category can be divided into gray roofs, red tile roofs, corrugated steel roofs, terracotta tile roofs, blue steel roofs, lightweight roofs, damaged roofs, slightly overgrown thatched roofs, and severely overgrown thatched roofs, etc. In this invention, different identification information is used to mark the location area of ​​each building in the image to distinguish the category of each building.

[0048] The identification area covers the region of hollowed-out villages where image acquisition will take place. Images of this area will be captured. Drones can be used to capture low-altitude images of the identification area.

[0049] After image acquisition, the image is input into the building classification and recognition model to obtain building classification information output by the building classification and recognition model. The building classification information is used to characterize the category of each building in the target village.

[0050] This building classification and recognition model takes as input the image features of each building extracted from sample images and the category of each building in the sample images, and takes the building classification information representing each building in the sample images as the expected output. It is trained on a pre-built initial model. The initial model includes a semantic segmentation layer that divides the building regions in the image into regions and a feature extraction layer that uses an attention mechanism for deep feature extraction.

[0051] In this invention, the semantic segmentation layer is used to simultaneously extract the location and boundaries of each building region in the image, which can improve the accuracy of building recognition and facilitate the subsequent accurate extraction of image features of the building regions.

[0052] The feature extraction layer enhances the extraction depth by incorporating the SE attention mechanism, extracts image features from each segmented building region, improves the feature selection capability of image pixel channels, and reduces the possibility of misjudgment in model detection.

[0053] The method for identifying buildings in hollow villages provided by this invention classifies buildings in images using a building classification and recognition model. This model includes a semantic segmentation layer for segmenting building regions and a feature extraction layer for extracting features from each building region based on an attention mechanism. Through the segmentation and feature extraction processes, more accurate image features can be extracted, providing accuracy in classifying each building in the image.

[0054] A further method described above primarily explains the process of inputting an image into a building classification and recognition model to obtain building classification information output by the model, as detailed below:

[0055] The image is input into the semantic segmentation layer, which then outputs the building regions.

[0056] Each building region is input into the feature extraction layer, which increases the extraction depth based on the attention mechanism. Based on the extraction depth, features are extracted from each building region, and the image features of each building region are output.

[0057] The image features of each building area are input into the classification layer of the building classification and recognition model, and the classification layer outputs the building classification information of each building area.

[0058] It should be noted that in this invention, the semantic segmentation layer simultaneously extracts the location and boundaries of each building region in the image. By determining the location and defining the boundaries, the region where the building is located can be divided. The semantic segmentation layer outputs an image in which the divided building regions are marked.

[0059] An image with predefined building regions is input into the feature extraction layer, which is equivalent to inputting each building region into the feature extraction layer. The feature extraction layer increases the extraction depth based on an attention mechanism, and extracts features from each building region according to the extraction depth, outputting the image features corresponding to each building region.

[0060] In this invention, the building classification and recognition model also includes a classification layer. The classification layer can perform similarity matching between image features and classification information based on the image features of each building area, and use the successfully matched classification information as the classification information of the corresponding building area.

[0061] In a further step of the present invention, through the process of region segmentation and feature extraction, more accurate image features can be extracted, providing accuracy in classifying the various buildings in the image.

[0062] In a further step of the above method, the feature extraction layer includes a first operational unit and a second operational unit. Accordingly, this further method mainly explains the process of inputting each building region into the feature extraction layer, increasing the extraction depth based on an attention mechanism, extracting features from each building region according to the extraction depth, and outputting the image features of each building region. The details are as follows:

[0063] Each building region is input into the first processing unit, which generates a weight matrix based on the extraction depth and the pixel values ​​of each building region.

[0064] Each building area is input into the second processing unit, which determines the initial image features based on the pixel values ​​of each building area, and then determines the image features based on the initial image features and the weight matrix.

[0065] It should be noted that in this invention, a deep residual network model is established based on the attention mechanism to understand the dependencies between image pixel channels. This increases the extraction depth, allowing for adaptive adjustment of the feature weights of each channel based on the extraction depth. Therefore, each building region is input into the first computation unit, which generates a weight matrix based on the pixel values ​​of each building region according to the extraction depth.

[0066] Each building region is input into the second computational unit, which determines initial image features based on the pixel values ​​of each region, and then determines the final image features based on these initial features and a weight matrix. The initial image features are adjusted using the weight matrix to obtain image features at a preset depth. This allows gradients with attention effects to flow into deeper networks, promoting better allocation of computational resources and ultimately improving the accuracy of target detection.

[0067] In a further aspect of the present invention, by performing calculations on the initial image features based on adjustable weight values ​​under the attention mechanism during the feature extraction process, more accurate image features can be extracted, providing accuracy for classifying each building in the image.

[0068] In a further step of the above method, the processing of the image for determining the identification area where the target village is located is explained in detail below:

[0069] The original image of the identification area where the target village is located is acquired, and the original image is orthorectified to determine the image.

[0070] It should be noted that in this invention, low-altitude images collected by UAVs are used as the original images, and high-resolution orthophotos are synthesized using Pix4D software as the final images. Compared with high-resolution remote sensing images, this orthophoto data improves the ability to represent the subtle features of buildings.

[0071] In a further step of the above method, the acquisition of the building classification and recognition model needs to be explained, as follows:

[0072] Several abandoned villages were selected as the study area. After defining the study area, drones equipped with cameras were used to collect low-altitude data, resulting in a large number of image samples. The collected image samples were then manually screened to obtain qualified raw image samples.

[0073] Based on the distribution of village houses in hollow villages, manual screening was conducted, and the two-dimensional features of village houses in remote sensing images were carefully observed and extracted to determine the classification criteria for hollow village buildings. After comparative analysis, features such as vegetation growing on the roof and obvious roof damage were used as positive samples (abandoned buildings), while other features such as intact roofs and bright gloss, as well as undamaged roofs and backgrounds, were used as negative samples.

[0074] Orthophotos were generated by mosaicking and preprocessing the acquired image samples using Pix4D software. Next, the TIFF format remote sensing images were cropped to 1024*1024 pixels using Global Mapper, while selecting appropriate overlapping areas. An instance segmentation model was introduced for building recognition in the images, performing pixel-level semantic segmentation simultaneously with target detection. Finally, based on the characteristics of positive and negative samples, the LabelMe professional annotation software was used to label the categories of each building in the image samples.

[0075] By employing various data augmentation techniques such as flip and scale, Gaussian noise, and gamut mapping, the number and diversity of training samples (noisy data) are increased, reducing the model's dependence on inherent properties and improving the model's generalization ability and robustness during training.

[0076] During the training process, a transfer learning strategy is adopted to transfer knowledge and share knowledge structures from a large amount of labeled data in the auxiliary domain, thereby promoting the learning tasks in this domain.

[0077] To address the issues of dense target distribution and information interference such as occlusion, an SE channel attention module is introduced into the residual structure of the Identity layer of the ResNet residual network. This enhances the feature filtering capability of channel information and reduces the possibility of misjudgment in model detection.

[0078] To address the issue of low feature abstraction in shallow neural networks, the depth of the backbone feature extraction network is increased. For example, the residual network in the 4-layer network is increased to 23 layers, which solves the gradient vanishing problem while increasing the network depth.

[0079] After training the model with a large number of image samples, a building classification and recognition model was obtained.

[0080] The following describes the identification device for buildings in hollow villages provided by the present invention. The identification device for buildings in hollow villages described below can be referred to in correspondence with the identification method for buildings in hollow villages described above.

[0081] Figure 2 This diagram illustrates a process flow chart of a building identification device for hollow villages provided by the present invention. (See attached diagram.) Figure 2 The device includes a determining module 21 and an identifying module 22, wherein:

[0082] Module 21 is used to determine the image of the recognition area where the target village is located;

[0083] The recognition module 22 is used to input the image into the building classification recognition model to obtain the building classification information output by the building classification recognition model. The building classification information is used to characterize the category of each building in the target village.

[0084] Among them, the building classification and recognition model is a model that is trained by machine learning based on the image features and category labels of each building in the sample image as input, and is used to classify each building in the image.

[0085] The building classification and recognition model includes a semantic segmentation layer and a feature extraction layer. The semantic segmentation layer is used to segment each building region in the image; the feature extraction layer is used to extract features from each building region based on an attention mechanism.

[0086] In a further embodiment of the above-described apparatus, the identification module is specifically used for:

[0087] The image is input into the semantic segmentation layer, which then outputs the building regions.

[0088] Each building region is input into the feature extraction layer, which increases the extraction depth based on the attention mechanism. Based on the extraction depth, features are extracted from each building region, and the image features of each building region are output.

[0089] The image features of each building area are input into the classification layer of the building classification and recognition model, and the classification layer outputs the building classification information of each building area.

[0090] In a further embodiment of the above-mentioned device, the feature extraction layer includes a first operational unit and a second operational unit. Correspondingly, in the process of inputting each building region into the feature extraction layer, the feature extraction layer increasing the extraction depth based on an attention mechanism, extracting features from each building region according to the extraction depth, and outputting the image features of each building region, the recognition module is specifically used for:

[0091] Each building region is input into the first processing unit, which generates a weight matrix based on the extraction depth and the pixel values ​​of each building region.

[0092] Each building area is input into the second processing unit, which determines the initial image features based on the extraction depth and the pixel values ​​of each building area, and then determines the image features based on the initial image features and the weight matrix.

[0093] In a further embodiment of the above apparatus, the determining module is specifically used for:

[0094] The original image of the identification area where the target village is located is acquired, and the image is determined by orthorectification based on the original image.

[0095] Since the device described in this embodiment of the invention is based on the same principle as the method described in the above embodiments, more detailed explanations will not be repeated here.

[0096] It should be noted that, in the embodiments of the present invention, the relevant functional modules can be implemented by a hardware processor.

[0097] The present invention provides a building identification device for hollow villages. The device classifies buildings in images using a building classification and identification model. The building classification and identification model includes a semantic segmentation layer for segmenting building regions and a feature extraction layer for extracting features from each building region based on an attention mechanism. Through the segmentation of regions and the feature extraction process, relatively accurate image features can be extracted, providing accuracy in classifying each building in the image.

[0098] Figure 3 An example is a schematic diagram of the physical structure of an electronic device, such as... Figure 3As shown, the electronic device may include a processor 31, a communication interface 32, a memory 33, and a communication bus 34. The processor 31, communication interface 32, and memory 33 communicate with each other via the communication bus 34. The processor 31 can call a computer program stored in the memory 33 to execute steps of a method for identifying buildings in a hollow village, such as: determining an image of the identification area where the target village is located; inputting the image into a building classification and recognition model to obtain building classification information output by the model, the building classification information being used to characterize the category of each building in the target village;

[0099] Among them, the building classification and recognition model is a model that is trained by machine learning based on the image features and category labels of each building in the sample image as input, and is used to classify each building in the image.

[0100] The building classification and recognition model includes a semantic segmentation layer and a feature extraction layer. The semantic segmentation layer is used to segment each building region in the image; the feature extraction layer is used to extract features from each building region based on an attention mechanism.

[0101] Furthermore, the logical instructions in the aforementioned memory 33 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0102] On the other hand, the present invention also provides a computer program product, the computer program product including a computer program stored on a non-transitory computer-readable storage medium, the computer program including program instructions, when the program instructions are executed by a computer, the computer is able to perform the steps of a method for identifying buildings in a hollow village, such as: determining an image of the identification area where the target village is located; inputting the image into a building classification and identification model to obtain building classification information output by the building classification and identification model, the building classification information being used to characterize the category of each building in the target village;

[0103] Among them, the building classification and recognition model is a model that is trained by machine learning based on the image features and category labels of each building in the sample image as input, and is used to classify each building in the image.

[0104] The building classification and recognition model includes a semantic segmentation layer and a feature extraction layer. The semantic segmentation layer is used to segment each building region in the image; the feature extraction layer is used to extract features from each building region based on an attention mechanism.

[0105] On the other hand, embodiments of the present invention also provide a processor-readable storage medium storing a computer program for causing the processor to execute steps of a method for identifying buildings in a hollow village, such as: determining an image of the identification area where the target village is located; inputting the image into a building classification and identification model to obtain building classification information output by the building classification and identification model, wherein the building classification information is used to characterize the category of each building in the target village;

[0106] Among them, the building classification and recognition model is a model that is trained by machine learning based on the image features and category labels of each building in the sample image as input, and is used to classify each building in the image.

[0107] The building classification and recognition model includes a semantic segmentation layer and a feature extraction layer. The semantic segmentation layer is used to segment each building region in the image; the feature extraction layer is used to extract features from each building region based on an attention mechanism.

[0108] The processor-readable storage medium can be any available medium or data storage device that the processor can access, including but not limited to magnetic memory (e.g., floppy disk, hard disk, magnetic tape, magneto-optical disk (MO)), optical memory (e.g., CD, DVD, BD, HVD), and semiconductor memory (e.g., ROM, EPROM, EEPROM, non-volatile memory (NAND FLASH), solid-state drive (SSD)).

[0109] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.

[0110] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.

[0111] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for identifying buildings in hollow villages, characterized in that, include: Image of the identification area where the target village is located; The hollow village contains village houses of varying degrees of newness and dilapidation; The image is input into the building classification and recognition model to obtain building classification information output by the building classification and recognition model. The building classification information is used to characterize the category of each building in the target village. The building classification and recognition model is a model trained by machine learning, which takes the image features of each building in the sample image and the category identifier of each building in the sample image as input, and is used to classify each building in the image. The building classification and recognition model includes a semantic segmentation layer and a feature extraction layer. The semantic segmentation layer is used to segment each building region in the image; the feature extraction layer is used to extract features from each building region based on an attention mechanism.

2. The method for identifying buildings in hollow villages according to claim 1, characterized in that, The step of inputting the image into the building classification and recognition model to obtain building classification information output by the building classification and recognition model includes: The image is input into the semantic segmentation layer, and the semantic segmentation layer outputs each building region. Each building region is input into the feature extraction layer, which increases the extraction depth based on an attention mechanism. Based on the extraction depth, features are extracted from each building region, and the image features of each building region are output. The image features of each building area are input into the classification layer of the building classification and recognition model, and the classification layer outputs the building classification information of each building area.

3. The method for identifying buildings in hollow villages according to claim 2, characterized in that, The feature extraction layer includes a first operation unit and a second operation unit. Correspondingly, each building region is input into the feature extraction layer, which increases the extraction depth based on an attention mechanism. Features are extracted from each building region according to the extraction depth, and the image features of each building region are output, including: Each building region is input into the first processing unit, which then generates a weight matrix based on the extraction depth and the pixel values ​​of each building region. Each building area is input into the second processing unit, which determines the initial image features based on the pixel values ​​of each building area, and then determines the image features based on the initial image features and the weight matrix.

4. The method for identifying buildings in hollow villages according to claim 1, characterized in that, The image of the identification area where the target village is located includes: The original image of the identification area where the target village is located is acquired, and the original image is orthorectified to determine the image.

5. A device for identifying buildings in a hollow village, characterized in that, include: The determination module is used to determine the image of the recognition area where the target village is located. The hollow village contains village houses of varying degrees of newness and dilapidation; The recognition module is used to input the image into the building classification recognition model to obtain building classification information output by the building classification recognition model. The building classification information is used to characterize the category of each building in the target village. The building classification and recognition model is a model trained by machine learning, which takes the image features of each building in the sample image and the category identifier of each building in the sample image as input, and is used to classify each building in the image. The building classification and recognition model includes a semantic segmentation layer and a feature extraction layer. The semantic segmentation layer is used to segment each building region in the image; the feature extraction layer is used to extract features from each building region based on an attention mechanism.

6. The identification device for buildings in hollow villages according to claim 5, characterized in that, The identification module is specifically used for: The image is input into the semantic segmentation layer, and the semantic segmentation layer outputs each building region. Each building region is input into the feature extraction layer, which increases the extraction depth based on an attention mechanism. Based on the extraction depth, features are extracted from each building region, and the image features of each building region are output. The image features of each building area are input into the classification layer of the building classification and recognition model, and the classification layer outputs the building classification information of each building area.

7. The identification device for buildings in hollow villages according to claim 6, characterized in that, The feature extraction layer includes a first operation unit and a second operation unit. Correspondingly, in the process of the recognition module inputting each building region into the feature extraction layer, the feature extraction layer increasing the extraction depth based on an attention mechanism, extracting features from each building region according to the extraction depth, and outputting the image features of each building region, the specific processing is as follows: Each building region is input into the first processing unit, which then generates a weight matrix based on the extraction depth and the pixel values ​​of each building region. Each building area is input into the second processing unit, which determines the initial image features based on the extraction depth and the pixel values ​​of each building area, and then determines the image features based on the initial image features and the weight matrix.

8. The identification device for buildings in hollow villages according to claim 6, characterized in that, The determining module is specifically used for: The original image of the identification area where the target village is located is acquired, and the image is determined by orthorectification based on the original image.

9. An electronic device comprising a processor and a memory storing a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method for identifying buildings in hollow villages as described in any one of claims 1 to 4.

10. A processor-readable storage medium, characterized in that, The processor-readable storage medium stores a computer program for causing the processor to perform the steps of the method for identifying buildings in a hollow village as described in any one of claims 1 to 4.

Citation Information

Patent Citations

  • Target identification method and device

    CN111539456A

  • Method and device for automatically classifying village and town architectural properties based on three-dimensional remote sensing image

    CN113792667A

  • Double-branch network remote sensing image building semantic segmentation method fusing rich scale features

    CN114821069A