A high-resolution remote sensing image residential area fine-grained classification method and device

By using a multi-task learning model to perform fine-grained classification of residential features on high-resolution remote sensing imagery, the problem of difficulty in fine classification in existing technologies is solved, and efficient fine classification of residential features is achieved, supporting topographic map updates and urban planning.

CN115601658BActive Publication Date: 2026-07-24BEIJING AEROSPACE HONGTU INFORMATION TECH
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
BEIJING AEROSPACE HONGTU INFORMATION TECH
Filing Date
2022-10-20
Publication Date
2026-07-24

AI Technical Summary

Technical Problem

Existing technologies are insufficient for fine-classifying residential features, failing to meet the needs of topographic map updates and remote sensing applications.

Method used

A multi-task learning model is used to perform fine-grained classification of residential features in high-resolution remote sensing images. This includes multi-class labeling, segmentation, and training and validation set partitioning of sample high-resolution remote sensing images. Semantic segmentation and object detection are performed on the images to be classified using ConvNeXt-Base and UNet structures. The fusion process is then used to obtain the fine classification results.

Benefits of technology

It enables fine-grained classification of settlement elements, provides more detailed information on settlement distribution, and supports topographic map data updates and urban planning.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115601658B_ABST
    Figure CN115601658B_ABST
Patent Text Reader

Abstract

The application provides a high-resolution remote sensing image resident area fine-grained classification method and device, relates to the technical field of remote sensing image processing, and comprises the following steps: obtaining a sample high-resolution remote sensing image, performing multi-category labeling on resident area elements in the sample high-resolution remote sensing image, and obtaining a target sample high-resolution remote sensing image; the target sample high-resolution remote sensing image is segmented according to a preset size, a sample data set is obtained, and the sample data set is divided into a training set and a verification set according to a preset proportion; the multi-task learning model is trained by using the training set and the verification set, and a target multi-task learning model is obtained; after obtaining a high-resolution remote sensing image to be classified, the target multi-task learning model is used for fine-grained classification of resident area elements of the high-resolution remote sensing image to be classified, and a fine classification result of the resident area elements of the high-resolution remote sensing image to be classified is obtained, so that the technical problem that the prior art cannot finely classify resident area elements is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the technical field of remote sensing image processing, and in particular to a method and apparatus for fine-grained classification of residential areas in high-resolution remote sensing images. Background Technology

[0002] Settlements are places where humans gather and live according to their needs for production and daily life. In topographic maps of various scales, settlement elements are among the most basic geographic information and also among the most important and fastest-changing elements in geospatial databases. With the development of socio-economic levels, users have increasingly higher requirements for the timeliness of topographic maps.

[0003] The need to update existing topographic map data is becoming increasingly urgent. Currently, methods for extracting settlements based on remote sensing imagery are constantly emerging. However, most of these methods treat all settlement features as a single type of feature for identification and extraction, failing to obtain fine-grained classifications within settlement features. This makes it difficult to meet the needs of topographic map updates and other remote sensing applications that require detailed classification of settlement features.

[0004] No effective solutions have yet been proposed to address the above problems. Summary of the Invention

[0005] In view of this, the purpose of the present invention is to provide a method and apparatus for fine-grained classification of residential areas in high-resolution remote sensing images, so as to alleviate the technical problem that the prior art is difficult to classify residential area elements in detail.

[0006] In a first aspect, embodiments of the present invention provide a method for fine-grained classification of residential areas in high-resolution remote sensing imagery, comprising: acquiring sample high-resolution remote sensing imagery and performing multi-class labeling on residential area elements in the sample high-resolution remote sensing imagery to obtain a target sample high-resolution remote sensing imagery; segmenting the target sample high-resolution remote sensing imagery according to a preset size to obtain a sample dataset, and dividing the sample datasetry into a training set and a validation set according to a preset ratio; training a multi-task learning model using the training set and the validation set to obtain a target multi-task learning model; and after acquiring the high-resolution remote sensing imagery to be classified, using the target multi-task learning model to perform fine-grained classification of residential area elements in the high-resolution remote sensing imagery to be classified, thereby obtaining the fine-grained classification result of residential area elements in the high-resolution remote sensing imagery to be classified.

[0007] Furthermore, the residential features include: detached houses, blocks, playgrounds, and high-rise buildings. Multi-category labeling is performed on the residential features in the sample high-resolution remote sensing image to obtain the target high-resolution remote sensing image. This includes: adding polygon labels to the detached houses, blocks, and playgrounds in the sample high-resolution remote sensing image to obtain the intermediate sample high-resolution remote sensing image; and adding bounding box labels to the high-rise buildings in the intermediate sample high-resolution remote sensing image to obtain the target sample high-resolution remote sensing image.

[0008] Furthermore, using the target multi-task learning model, fine-grained classification of residential features is performed on the high-resolution remote sensing image to be classified, obtaining the fine-classification results of residential features in the high-resolution remote sensing image to be classified. This includes: inputting the high-resolution remote sensing image to be classified into the target multi-task learning model to obtain a semantic segmentation prediction map and a high-rise building result map, wherein the semantic segmentation prediction map includes detached houses, blocks, and playgrounds in the high-resolution remote sensing image to be classified, and the high-rise building result map includes high-rise buildings in the high-resolution remote sensing image to be classified; fusing the semantic segmentation prediction map and the high-rise building result map to obtain the fine-classification results of residential features in the high-resolution remote sensing image to be classified.

[0009] Further, the semantic segmentation prediction map and the high-rise building result map are fused to obtain the detailed classification result of the residential area features in the high-resolution remote sensing image to be classified. This includes: performing hole filling and fragment removal processing on the semantic segmentation prediction map to obtain a target semantic segmentation prediction map; determining the intermediate high-rise buildings in the high-rise building result map based on the residential area feature patches contained in the target semantic segmentation prediction map, wherein the intermediate high-rise buildings are high-rise buildings whose overlapping area with the residential area feature patches contained in the target semantic segmentation prediction map is greater than a preset threshold; determining the target high-rise buildings among the intermediate high-rise buildings, wherein the target high-rise buildings are intermediate high-rise buildings with an area greater than a preset area and independent intermediate high-rise buildings with an area less than a preset area; fusing the target high-rise buildings with the target semantic segmentation prediction map to obtain a fused image; simplifying and smoothing the contours of the residential area features in the fused image, and converting the classification raster in the fused image into vectors to obtain the detailed classification result of the residential area features.

[0010] Furthermore, the backbone network of the multi-task learning model is ConvNeXt-Base, the semantic segmentation classification head is a UNet structure, and the object detection localization classification head is YOLOv5s.

[0011] Secondly, embodiments of the present invention also provide a fine-grained classification device for residential areas in high-resolution remote sensing imagery, comprising: an acquisition unit, a construction unit, a training unit, and a classification unit. The acquisition unit is used to acquire sample high-resolution remote sensing images and perform multi-category labeling on residential area elements in the sample high-resolution remote sensing images to obtain target sample high-resolution remote sensing images. The construction unit is used to segment the target sample high-resolution remote sensing images according to a preset size to obtain a sample dataset, and divide the sample dataset into a training set and a validation set according to a preset ratio. The training unit is used to train a multi-task learning model using the training set and the validation set to obtain a target multi-task learning model. The classification unit is used, after acquiring the high-resolution remote sensing images to be classified, to perform fine-grained classification of residential area elements in the high-resolution remote sensing images to be classified using the target multi-task learning model to obtain the fine-grained classification results of the residential area elements in the high-resolution remote sensing images to be classified.

[0012] Furthermore, the residential elements include: detached houses, blocks, playgrounds, and high-rise buildings. The acquisition unit is used to: add polygon labels to the detached houses, blocks, and playgrounds in the sample high-resolution remote sensing image to obtain intermediate sample high-resolution remote sensing image; and add bounding box labels to the high-rise buildings in the intermediate sample high-resolution remote sensing image to obtain the target sample high-resolution remote sensing image.

[0013] Further, the classification unit is used to: input the high-resolution remote sensing image to be classified into the target multi-task learning model to obtain a semantic segmentation prediction map and a high-rise building result map, wherein the semantic segmentation prediction map includes detached houses, blocks, and playgrounds in the high-resolution remote sensing image to be classified, and the high-rise building result map includes high-rise buildings in the high-resolution remote sensing image to be classified; and fuse the semantic segmentation prediction map and the high-rise building result map to obtain the detailed classification result of the residential land features of the high-resolution remote sensing image to be classified.

[0014] Thirdly, embodiments of the present invention also provide an electronic device, including a memory and a processor, wherein the memory is used to store a program that supports the processor in executing the method described in the first aspect above, and the processor is configured to execute the program stored in the memory.

[0015] Fourthly, embodiments of the present invention also provide a computer-readable storage medium on which a computer program is stored.

[0016] In this embodiment of the invention, a target sample high-resolution remote sensing image is obtained by acquiring sample high-resolution remote sensing images and performing multi-category labeling on the residential features in the sample high-resolution remote sensing images. The target sample high-resolution remote sensing image is then segmented according to a preset size to obtain a sample dataset, and the sample dataset is divided into a training set and a validation set according to a preset ratio. A multi-task learning model is trained using the training set and the validation set to obtain a target multi-task learning model. After acquiring the high-resolution remote sensing image to be classified, the target multi-task learning model is used to perform fine-grained classification of residential features in the high-resolution remote sensing image to be classified, thereby obtaining the fine classification results of residential features in the high-resolution remote sensing image to be classified. This achieves the purpose of fine classification of residential features, thus solving the technical problem that it is difficult to perform fine classification of residential features in the prior art, and thus realizing the technical effect of providing convenience for updating topographic map data.

[0017] Other features and advantages of the invention will be set forth in the description which follows, and will be apparent in part from the description, or may be learned by practicing the invention. The objects and other advantages of the invention are realized and obtained in accordance with the structures particularly pointed out in the description, claims and drawings.

[0018] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, preferred embodiments are described below in detail with reference to the accompanying drawings. Attached Figure Description

[0019] To more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the drawings used in the description of the specific embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.

[0020] Figure 1 A flowchart illustrating a fine-grained classification method for residential areas in high-resolution remote sensing imagery provided in this embodiment of the invention;

[0021] Figure 2 This is a schematic diagram of a fine-grained classification device for residential areas in high-resolution remote sensing imagery provided in an embodiment of the present invention;

[0022] Figure 3 This is a schematic diagram of an electronic device provided in an embodiment of the present invention. Detailed Implementation

[0023] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0024] Example 1:

[0025] According to an embodiment of the present invention, an embodiment of a fine-grained classification method for residential areas in high-resolution remote sensing imagery is provided. It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. Furthermore, although a logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order than that shown here.

[0026] Figure 1 This is a flowchart of a high-resolution remote sensing imagery fine-grained classification method for residential areas according to an embodiment of the present invention, such as... Figure 1 As shown, the method includes the following steps:

[0027] Step S102: Obtain sample high-resolution remote sensing images and perform multi-category annotation on the residential features in the sample high-resolution remote sensing images to obtain target sample high-resolution remote sensing images.

[0028] Specifically, the high-resolution remote sensing images of the acquired samples were visually interpreted to label residential features in multiple categories, obtaining area labels for three types of features: detached houses, blocks, and playgrounds. At the same time, bounding boxes were added to high-rise buildings to obtain labels for them.

[0029] Step S104: Segment the target sample high-resolution remote sensing image according to a preset size to obtain a sample dataset, and divide the sample dataset into a training set and a validation set according to a preset ratio;

[0030] Specifically, the high-resolution remote sensing images of the target samples are converted and cropped to obtain sample data of sizes such as 512 or 1024 pixels. A sample dataset is constructed based on the sample data, and the sample dataset is divided into a training set and a validation set according to a preset ratio.

[0031] Step S106: Train the multi-task learning model using the training set and the validation set to obtain the target multi-task learning model;

[0032] It should be noted that the multi-task learning model selects a UNet-like network and inserts an object detection and localization classification head into the model structure. In this embodiment of the invention, the model backbone network is ConvNeXt-Base, the semantic segmentation classification head is a UNet structure, and the object detection and localization classification head is YOLOv5s, which together form a multi-task learning network model.

[0033] The semantic segmentation samples of detached houses, blocks, and playgrounds and the target detection samples of high-rise buildings in the training set are input into the multi-task learning network model for model training and optimization. The optimal model is selected by the validation set index to obtain the target multi-task learning network model.

[0034] In this embodiment of the invention, the semantic segmentation loss function is the cross-entropy loss l. ce and Dice loss dice The combined loss function, the object detection loss function is divided into classification loss l cls Positioning loss box and confidence loss l obj The loss function consists of three parts, with weighting coefficients of 0.6 and 0.4 for semantic segmentation and object detection, respectively. The specific loss function expressions are as follows:

[0035] L = 0.6 × (l ce +l dice )+0.4×(l cls +l box +l obj ).

[0036] Step S108: After acquiring the high-resolution remote sensing image to be classified, the target multi-task learning model is used to perform fine-grained classification of residential features on the high-resolution remote sensing image to be classified, so as to obtain the fine-grained classification results of residential features of the high-resolution remote sensing image to be classified.

[0037] In this embodiment of the invention, a target sample high-resolution remote sensing image is obtained by acquiring sample high-resolution remote sensing images and performing multi-category labeling on the residential features in the sample high-resolution remote sensing images. The target sample high-resolution remote sensing image is then segmented according to a preset size to obtain a sample dataset, and the sample dataset is divided into a training set and a validation set according to a preset ratio. A multi-task learning model is trained using the training set and the validation set to obtain a target multi-task learning model. After acquiring the high-resolution remote sensing image to be classified, the target multi-task learning model is used to perform fine-grained classification of residential features in the high-resolution remote sensing image to be classified, thereby obtaining the fine classification results of residential features in the high-resolution remote sensing image to be classified. This achieves the purpose of fine classification of residential features, thus solving the technical problem that it is difficult to perform fine classification of residential features in the prior art, and thus realizing the technical effect of providing convenience for updating topographic map data.

[0038] In this embodiment of the invention, step S108 includes the following steps:

[0039] Step S11: Input the high-resolution remote sensing image to be classified into the target multi-task learning model to obtain a semantic segmentation prediction map and a high-rise building result map. The semantic segmentation prediction map includes detached houses, blocks and playgrounds in the high-resolution remote sensing image to be classified, and the high-rise building result map includes high-rise buildings in the high-resolution remote sensing image to be classified.

[0040] Step S12: The semantic segmentation prediction map and the high-rise building result map are fused to obtain the detailed classification result of the residential land features of the high-resolution remote sensing image to be classified.

[0041] Specifically, the semantic segmentation prediction map and the high-rise building result map are fused to obtain the detailed classification results of residential features in the high-resolution remote sensing image to be classified, including:

[0042] The semantic segmentation prediction map is subjected to hole filling and fragment removal processing to obtain the target semantic segmentation prediction map;

[0043] Based on the residential land feature patches contained in the target semantic segmentation prediction map, the intermediate high-rise buildings in the high-rise building result map are determined, wherein the intermediate high-rise buildings are high-rise buildings whose overlapping area with the residential land feature patches contained in the target semantic segmentation prediction map is greater than a preset threshold.

[0044] Identify the target high-rise buildings among the intermediate high-rise buildings, wherein the target high-rise buildings are intermediate high-rise buildings with an area greater than a preset area and independent intermediate high-rise buildings with an area less than a preset area.

[0045] The target high-rise building is fused with the target semantic segmentation prediction map to obtain a fused image;

[0046] The contours of the settlement features in the fused image are simplified and smoothed, and the classification raster in the fused image is converted into vectors to obtain the fine classification results of the settlement features.

[0047] In this embodiment of the invention, after inputting the high-resolution remote sensing image to be classified into the target multi-task learning model, a semantic segmentation prediction map S and a high-rise building result map D are obtained.

[0048] Hole filling and fragment removal are performed on the semantic segmentation prediction map S to obtain the semantic segmentation result Si. post .

[0049] It should be noted that, in this embodiment of the invention, morphological closing operations are used to fill small holes, and morphological opening operations are used to remove small fragmented patches.

[0050] Results D for high-rise buildings and semantic segmentation results S post The images are then fused to obtain detailed classification results of residential features from the high-resolution remote sensing imagery to be classified. The specific fusion steps are as follows:

[0051] The bounding box and semantic segmentation result S in the high-rise building result D are compared. post Calculate the area of ​​intersecting patches sequentially, and mark the patches with an area greater than 0 (i.e., the preset threshold) as suspected high-rise building patches B (i.e., middle high-rise buildings).

[0052] Based on the suspected high-rise building image patches, an adaptive selection strategy is adopted to obtain the target high-rise buildings, as detailed below:

[0053] If the area of ​​the suspected high-rise building is less than T s (i.e., preset area), then determine whether the suspected high-rise building patch is an independent patch. If it is an independent patch, mark it as the target high-rise building; if it is not an independent patch, mark it as the original category.

[0054] If the area of ​​the suspected high-rise building is greater than or equal to T s If the area is an independent area, then the entire area containing that area will be marked as the target high-rise building.

[0055] Finally, the outlines of the extracted residential features for the four categories are simplified and smoothed, and the classification raster is converted into a vector to obtain the detailed classification results of the residential features.

[0056] Compared to existing residential area extraction methods based on high-resolution remote sensing imagery that treat all buildings as a single category and fail to yield finer-grained classification results, this invention employs a multi-task learning-based fine-grained classification method for residential area features in remote sensing imagery. This method can extract four types of residential area features—detached houses, street blocks, high-rise buildings, and playgrounds—end-to-end, providing more detailed residential area distribution information for remote sensing applications such as topographic map updates and urban planning.

[0057] Example 2:

[0058] This invention also provides a fine-grained classification device for high-resolution remote sensing image residential areas. This device is used to perform the fine-grained classification method for high-resolution remote sensing image residential areas provided in the above-described embodiments of this invention. The following is a detailed description of the device provided in this invention.

[0059] like Figure 2 As shown, Figure 2The above-mentioned high-resolution remote sensing image residential area fine-grained classification device includes: acquisition unit 10, construction unit 20, training unit 30 and classification unit 40.

[0060] The acquisition unit is used to acquire sample high-resolution remote sensing images and perform multi-category annotation on the residential features in the sample high-resolution remote sensing images to obtain target sample high-resolution remote sensing images.

[0061] The construction unit is used to segment the target sample high-resolution remote sensing image according to a preset size to obtain a sample dataset, and to divide the sample dataset into a training set and a validation set according to a preset ratio.

[0062] The training unit is used to train the multi-task learning model using the training set and the validation set to obtain the target multi-task learning model.

[0063] The classification unit is used to perform fine-grained classification of residential features on the high-resolution remote sensing image to be classified after acquiring the high-resolution remote sensing image to be classified, using the target multi-task learning model, to obtain the fine-grained classification results of the residential features of the high-resolution remote sensing image to be classified.

[0064] In this embodiment of the invention, a target sample high-resolution remote sensing image is obtained by acquiring sample high-resolution remote sensing images and performing multi-category labeling on the residential features in the sample high-resolution remote sensing images. The target sample high-resolution remote sensing image is then segmented according to a preset size to obtain a sample dataset, and the sample dataset is divided into a training set and a validation set according to a preset ratio. A multi-task learning model is trained using the training set and the validation set to obtain a target multi-task learning model. After acquiring the high-resolution remote sensing image to be classified, the target multi-task learning model is used to perform fine-grained classification of residential features in the high-resolution remote sensing image to be classified, thereby obtaining the fine classification results of residential features in the high-resolution remote sensing image to be classified. This achieves the purpose of fine classification of residential features, thus solving the technical problem that it is difficult to perform fine classification of residential features in the prior art, and thus realizing the technical effect of providing convenience for updating topographic map data.

[0065] Example 3:

[0066] This invention also provides an electronic device, including a memory and a processor. The memory is used to store a program that supports the processor in executing the method described in Embodiment 1 above, and the processor is configured to execute the program stored in the memory.

[0067] See Figure 3The present invention also provides an electronic device 100, including: a processor 50, a memory 51, a bus 52 and a communication interface 53, wherein the processor 50, the communication interface 53 and the memory 51 are connected through the bus 52; the processor 50 is used to execute executable modules, such as computer programs, stored in the memory 51.

[0068] The memory 51 may include high-speed random access memory (RAM) or non-volatile memory, such as at least one disk storage device. Communication between this system network element and at least one other network element is achieved through at least one communication interface 53 (which can be wired or wireless), such as the Internet, wide area network, local area network, metropolitan area network, etc.

[0069] Bus 52 can be an ISA bus, PCI bus, or EISA bus, etc. The bus can be divided into address bus, data bus, control bus, etc. For ease of representation, Figure 3 The symbol is represented by a single double-headed arrow, but this does not mean that there is only one bus or one type of bus.

[0070] The memory 51 is used to store programs. After receiving an execution instruction, the processor 50 executes the programs. The method executed by the device for defining the flow process disclosed in any of the foregoing embodiments of the present invention can be applied to the processor 50 or implemented by the processor 50.

[0071] Processor 50 may be an integrated circuit chip with signal processing capabilities. In implementation, each step of the above method can be completed by the integrated logic circuitry in the hardware of processor 50 or by instructions in software form. Processor 50 can be a general-purpose processor, including a Central Processing Unit (CPU), a Network Processor (NP), etc.; it can also be a Digital Signal Processor (DSP), an Application Specific Integrated Circuit (ASIC), a Field-Programmable Gate Array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. It can implement or execute the methods, steps, and logic block diagrams disclosed in the embodiments of this invention. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the methods disclosed in the embodiments of this invention can be directly embodied in the execution of a hardware decoding processor, or executed by a combination of hardware and software modules in the decoding processor. The software modules can reside in random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, or other mature storage media in the art. The storage medium is located in memory 51. The processor 50 reads the information in memory 51 and, in conjunction with its hardware, completes the steps of the above method.

[0072] Example 4:

[0073] This invention also provides a computer-readable storage medium storing a computer program, which, when executed by a processor, performs the steps of the method described in Embodiment 1 above.

[0074] Furthermore, in the description of the embodiments of the present invention, unless otherwise explicitly specified and limited, the terms "installation," "connection," and "linking" should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral connection; they can refer to a mechanical connection or an electrical connection; they can refer to a direct connection or an indirect connection through an intermediate medium; and they can refer to the internal connection of two components. Those skilled in the art can understand the specific meaning of the above terms in the present invention based on the specific circumstances.

[0075] In the description of this invention, it should be noted that the terms "center," "upper," "lower," "left," "right," "vertical," "horizontal," "inner," and "outer," etc., indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings. They are used only for the convenience of describing the invention and for simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, they should not be construed as limitations on the invention. Furthermore, the terms "first," "second," and "third" are used for descriptive purposes only and should not be construed as indicating or implying relative importance.

[0076] In the several embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. The apparatus embodiments described above are merely illustrative. For example, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. Furthermore, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Additionally, the shown or discussed mutual couplings, direct couplings, or communication connections may be through some communication interfaces; indirect couplings or communication connections between devices or units may be electrical, mechanical, or other forms.

[0077] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0078] In addition, the functional units in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.

[0079] Finally, it should be noted that the above-described embodiments are merely specific implementations of the present invention, used to illustrate the technical solutions of the present invention, and not to limit it. The scope of protection of the present invention is not limited thereto. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that any person skilled in the art can still modify or easily conceive of changes to the technical solutions described in the foregoing embodiments within the technical scope disclosed in the present invention, or make equivalent substitutions for some of the technical features; and these modifications, changes, or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be covered within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.

Claims

1. A fine-grained classification method for residential areas in high-resolution remote sensing imagery, characterized in that, include: A high-resolution remote sensing image of a sample is acquired, and the residential features in the sample high-resolution remote sensing image are labeled with multiple categories to obtain a target sample high-resolution remote sensing image. The residential features include: detached houses, blocks, playgrounds, and high-rise buildings. The multi-category labeling of the residential features in the sample high-resolution remote sensing image to obtain the target high-resolution remote sensing image includes: adding polygon labels to detached houses, blocks, and playgrounds in the sample high-resolution remote sensing image to obtain an intermediate sample high-resolution remote sensing image; adding bounding box labels to high-rise buildings in the intermediate sample high-resolution remote sensing image to obtain the target sample high-resolution remote sensing image. The target sample high-resolution remote sensing image is segmented according to a preset size to obtain a sample dataset, and the sample dataset is divided into a training set and a validation set according to a preset ratio. The multi-task learning model is trained using the training set and the validation set to obtain the target multi-task learning model; After acquiring the high-resolution remote sensing image to be classified, the target multi-task learning model is used to perform fine-grained classification of residential features on the high-resolution remote sensing image to be classified, obtaining the fine-grained classification results of residential features in the high-resolution remote sensing image to be classified, including: The high-resolution remote sensing image to be classified is input into the target multi-task learning model to obtain a semantic segmentation prediction map and a high-rise building result map. The semantic segmentation prediction map includes detached houses, blocks and playgrounds in the high-resolution remote sensing image to be classified, and the high-rise building result map includes high-rise buildings in the high-resolution remote sensing image to be classified. The semantic segmentation prediction map and the high-rise building result map are fused to obtain the detailed classification results of residential features in the high-resolution remote sensing image to be classified, including: The semantic segmentation prediction map is subjected to hole filling and fragment removal processing to obtain the target semantic segmentation prediction map; Based on the residential land feature patches contained in the target semantic segmentation prediction map, the intermediate high-rise buildings in the high-rise building result map are identified, wherein the intermediate high-rise buildings are high-rise buildings whose overlapping area with the residential land feature patches contained in the target semantic segmentation prediction map is greater than a preset threshold; target high-rise buildings among the intermediate high-rise buildings are identified, wherein the target high-rise buildings are intermediate high-rise buildings with an area greater than a preset area and independent intermediate high-rise buildings with an area smaller than a preset area; the target high-rise buildings are fused with the target semantic segmentation prediction map to obtain a fused image; the contours of the residential land features in the fused image are simplified and smoothed, and the classification raster in the fused image is converted into vectors to obtain the fine classification results of the residential land features.

2. The method according to claim 1, characterized in that, The backbone network of the multi-task learning model is ConvNeXt-Base, the semantic segmentation classification head is a UNet structure, and the object detection and localization classification head is YOLOv5s.

3. A fine-grained classification device for residential areas in high-resolution remote sensing imagery, characterized in that, include: Acquisition unit, construction unit, training unit and classification unit, where, The acquisition unit is used to acquire sample high-resolution remote sensing images and perform multi-category labeling on residential features in the sample high-resolution remote sensing images to obtain target sample high-resolution remote sensing images. The residential features include: detached houses, blocks, playgrounds, and high-rise buildings. Performing multi-category labeling on the residential features in the sample high-resolution remote sensing images to obtain target high-resolution remote sensing images includes: adding polygon labels to detached houses, blocks, and playgrounds in the sample high-resolution remote sensing images to obtain intermediate sample high-resolution remote sensing images; adding bounding box labels to high-rise buildings in the intermediate sample high-resolution remote sensing images to obtain target sample high-resolution remote sensing images. The construction unit is used to segment the target sample high-resolution remote sensing image according to a preset size to obtain a sample dataset, and to divide the sample dataset into a training set and a validation set according to a preset ratio. The training unit is used to train the multi-task learning model using the training set and the validation set to obtain the target multi-task learning model. The classification unit is used, after acquiring the high-resolution remote sensing image to be classified, to perform fine-grained classification of residential features on the high-resolution remote sensing image to be classified using the target multi-task learning model, to obtain the fine-grained classification results of residential features in the high-resolution remote sensing image to be classified, including: The high-resolution remote sensing image to be classified is input into the target multi-task learning model to obtain a semantic segmentation prediction map and a high-rise building result map. The semantic segmentation prediction map includes detached houses, blocks and playgrounds in the high-resolution remote sensing image to be classified, and the high-rise building result map includes high-rise buildings in the high-resolution remote sensing image to be classified. The semantic segmentation prediction map and the high-rise building result map are fused to obtain the detailed classification results of residential features in the high-resolution remote sensing image to be classified, including: The semantic segmentation prediction map is subjected to hole filling and fragment removal processing to obtain the target semantic segmentation prediction map; Based on the residential land feature patches contained in the target semantic segmentation prediction map, the intermediate high-rise buildings in the high-rise building result map are identified, wherein the intermediate high-rise buildings are high-rise buildings whose overlapping area with the residential land feature patches contained in the target semantic segmentation prediction map is greater than a preset threshold; target high-rise buildings among the intermediate high-rise buildings are identified, wherein the target high-rise buildings are intermediate high-rise buildings with an area greater than a preset area and independent intermediate high-rise buildings with an area smaller than a preset area; the target high-rise buildings are fused with the target semantic segmentation prediction map to obtain a fused image; the contours of the residential land features in the fused image are simplified and smoothed, and the classification raster in the fused image is converted into vectors to obtain the fine classification results of the residential land features.

4. An electronic device, characterized in that, The device includes a memory and a processor, the memory being used to store a program that enables the processor to execute the method of any one of claims 1 to 2, and the processor being configured to execute the program stored in the memory.

5. A computer-readable storage medium storing a computer program thereon, characterized in that, When a computer program is run by a processor, it performs the steps of the method described in any one of claims 1 to 2.