A training sample determination method for high-resolution remote sensing image intelligent classification
Patent Information
- Application Number
- CN202610595369.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-04-30
- Publication Date
- 2026-10-09
- Estimated Expiration
- 2046-04-30
AI Technical Summary
然而,现有方法在实际应用中仍存在一定局限:部分研究依赖外部土地覆盖产品作为先验信息,样本选择受制于参考数据的可获取性与时效性;另一些方法则基于已标注的语义信息计算复杂性指标,其流程本质上是"先标注、后选样",无法从源头上减少冗余标注
[0016]The training sample determination scheme provided in this application involves acquiring at least one remote sensing image, cropping each remote sensing image into at least two image blocks, and forming an image block set based on each image block corresponding to each remote sensing image; determining the spectral feature complexity and spatial structure complexity of each image block in the image block set; determining at least two spectral-spatial two-dimensional complexity levels and a stratification criterion corresponding to each spectral-spatial two-dimensional complexity level; for each spectral-spatial two-dimensional complexity level, determining an image block subset composed of image blocks belonging to the spectral-spatial two-dimensional complexity level from the image block set based on the stratification criterion, the spectral feature complexity, and the spatial structure complexity corresponding to the spectral-spatial two-dimensional complexity level; selecting a preset number of image block samples from each image block subset, and forming an image block training sample set for a semantic segmentation model based on the image block samples selected from each image block subset; wherein, the semantic segmentation model is a deep learning model used to perform remote sensing image classification tasks. This approach quantifies the scene complexity of image patches from two dimensions: spectral features and spatial structure. Based on this, it performs joint hierarchical and balanced sampling to construct a training sample set with high representativeness for the semantic segmentation model, effectively improving the accuracy and generalization ability of the semantic segmentation model in remote sensing image interpretation tasks.
Smart Images

Figure CN122176558B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of remote sensing image processing technology, and in particular to a method for determining training samples for intelligent classification of high-resolution remote sensing images. Background Technology
[0002] High spatial resolution remote sensing imagery, with its rich texture details, complete spatial geometry, and clear feature boundaries, can more realistically depict the morphology and structural features of urban landscapes, providing crucial data support for fine-scale feature mapping and information extraction. However, due to factors such as diverse feature types, complex scene structures, and significant scale differences, the spatial heterogeneity of high-resolution remote sensing imagery is significantly enhanced, which places higher demands on the accuracy and generalization ability of interpretation models.
[0003] Traditional remote sensing classification methods, such as exponential feature extraction, maximum likelihood estimation, and support vector machines, generally rely on manually designed features. With the continuous improvement of remote sensing image resolution and the increasing complexity of urban environments, the limitations of traditional methods in terms of computational efficiency, classification accuracy, and model generalization ability have become increasingly apparent, making it difficult to meet the practical needs of refined urban feature extraction. In recent years, deep learning models, with their powerful feature learning capabilities, have rapidly become the mainstream research direction in this field and have demonstrated superior performance in remote sensing image classification and feature extraction tasks.
[0004] While deep learning methods offer significant advantages in extraction accuracy, they are inherently data-driven, meaning model performance heavily relies on large amounts of high-quality labeled data. Theoretically, larger and more diverse training datasets enable models to learn richer knowledge, thereby improving their generalization ability to unseen data. However, in practical applications of large-scale remote sensing mapping, obtaining sufficient labeled samples often requires substantial investment of manpower and time. This contradiction raises a key technical question: how to construct a small but highly representative training sample set to achieve an effective balance between model accuracy and labeling costs.
[0005] Achieving these goals largely depends on a scientifically sound sampling design. The core of sampling is selecting samples from the study area that can adequately represent the distribution characteristics of ground features, providing comprehensive information support for the model and reducing sampling bias during training. Common sampling methods include simple random sampling, systematic sampling, and stratified random sampling. Research shows that, compared to the former two, stratified random sampling is more effective in improving the overall representativeness of the samples.
[0006] However, most existing sampling studies focus on pixel-level remote sensing classification tasks, using individual pixels as the basic sampling unit and performing statistical analysis and selection based on pixel category labels. But in deep learning-based semantic segmentation tasks, the sampling unit is typically an image patch containing multiple land cover categories. Taking urban green space scenes as an example, a single image patch often simultaneously encompasses multiple land cover categories such as green space, buildings, and roads. This category heterogeneity makes pixel-level sampling strategies difficult to directly adapt to the sample selection requirements at the image patch level, limiting their applicability in deep learning-based remote sensing image semantic segmentation scenarios.
[0007] In current remote sensing semantic segmentation tasks, image patch-level sample selection still primarily employs simple random sampling. This method assumes that all image patches have an equal probability of being selected, essentially assuming that the land cover composition within an image patch is relatively uniform and that different image patches contribute similarly to model training. However, in high-resolution remote sensing imagery, due to the diversity of land cover types and complex spatial patterns, the complexity within different image patches varies significantly. Taking urban green space as an example, some image patches may contain only a single land cover type (such as large continuous green areas), while others may simultaneously encompass multiple land cover types such as green spaces, buildings, and roads, presenting a complex intermingling pattern of land cover features. This difference in intra-patch heterogeneity directly determines the informational value of image patches for model training: complex image patches contain richer boundary information and land cover co-occurrence patterns, helping the model learn more discriminative features; while simple image patches provide relatively limited supervisory information, and repeatedly learning from such samples has a lower marginal contribution to improving model performance.
[0008] Against this backdrop, the limitations of simple random sampling become increasingly apparent: because it does not consider the complexity differences between image patches, this method struggles to ensure that complex image patches are effectively included in the training set. When the study area is small and the variability of ground features is limited, the complexity differences between image patches are not significant, and the bias introduced by random sampling is relatively controllable. However, as the study area expands to a large scale, the heterogeneity of the spatial pattern of ground features increases significantly, and the complexity of image patches exhibits a wider distribution range. At this point, the samples selected by simple random sampling often consist mainly of a large number of simple image patches, resulting in insufficient coverage of complex geographical scenes by the training set, which in turn restricts the generalization performance of the model in areas with complex ground features.
[0009] To address the aforementioned issues, existing research has attempted to introduce complexity-related indicators to stratify training samples, thereby improving the sample distribution's coverage of complex geographical scenarios. However, existing methods still have certain limitations in practical applications: some studies rely on external land cover products as prior information, and sample selection is constrained by the availability and timeliness of reference data; other methods calculate complexity indicators based on labeled semantic information, and their process is essentially "label first, then select samples," which cannot reduce redundant labeling from the source. Summary of the Invention
[0010] This application provides a method for determining training samples for intelligent classification of high-resolution remote sensing images. It can construct a training sample set with high representativeness of the semantic segmentation model, effectively improving the accuracy and generalization ability of the semantic segmentation model in remote sensing image interpretation tasks.
[0011] According to one aspect of this application, a method for determining training samples for intelligent classification of high-resolution remote sensing images is provided, the method comprising: Acquire at least one remote sensing image, crop each remote sensing image into at least two image blocks, and construct an image block set based on each image block corresponding to each remote sensing image; The spectral feature complexity and spatial structure complexity of each image block in the image block set are determined respectively; Determine at least two spectral-spatial two-dimensional complexity levels and the stratification criteria corresponding to each spectral-spatial two-dimensional complexity level; For each of the spectral-spatial two-dimensional complexity levels, based on the hierarchical determination criteria, the spectral feature complexity, and the spatial structure complexity corresponding to the spectral-spatial two-dimensional complexity level, a subset of image blocks belonging to the spectral-spatial two-dimensional complexity level is determined from the image block set. A predetermined number of image block samples are selected from each of the image block subsets, and an image block training sample set for the semantic segmentation model is constructed based on the image block samples selected from each of the image block subsets; wherein, the semantic segmentation model is a deep learning model used to perform remote sensing image classification tasks.
[0012] According to one aspect of this application, a training sample determination device for intelligent classification of high-resolution remote sensing images is provided, the device comprising: The image block set determination module is used to acquire at least one remote sensing image, crop each remote sensing image into at least two image blocks, and form an image block set based on each image block corresponding to each remote sensing image. The complexity determination module is used to determine the spectral feature complexity and spatial structure complexity of each image block in the image block set, respectively. The stratification determination criterion module is used to determine at least two spectral-spatial two-dimensional complexity levels and the stratification determination criterion corresponding to each of the spectral-spatial two-dimensional complexity levels. The image block subset determination module is used to determine, for each of the spectral-spatial two-dimensional complexity levels, an image block subset consisting of image blocks belonging to the spectral-spatial two-dimensional complexity level from the image block set, based on the hierarchical determination criteria, the spectral feature complexity, and the spatial structure complexity corresponding to the spectral-spatial two-dimensional complexity level. The image patch training sample set determination module is used to select a preset number of image patch samples from each of the image patch subsets, and to construct an image patch training sample set for the semantic segmentation model based on the image patch samples selected from each of the image patch subsets; wherein, the semantic segmentation model is a deep learning model used to perform remote sensing image classification tasks.
[0013] According to another aspect of this application, an electronic device is provided, the electronic device comprising: At least one processor; and A memory that is communicatively connected to at least one processor; wherein, The memory stores a computer program that can be executed by at least one processor, such that the at least one processor is able to perform the training sample determination method for intelligent classification of high-resolution remote sensing images according to any embodiment of the present application.
[0014] According to another aspect of this application, a computer-readable storage medium is provided, which stores computer instructions for causing a processor to execute and implement the training sample determination method for intelligent classification of high-resolution remote sensing images according to any embodiment of this application.
[0015] According to another aspect of this application, a computer program product is provided, comprising a computer program that, when executed by a processor, implements the training sample determination method for intelligent classification of high-resolution remote sensing images according to any embodiment of this application.
[0016] The training sample determination scheme provided in this application involves acquiring at least one remote sensing image, cropping each remote sensing image into at least two image blocks, and forming an image block set based on each image block corresponding to each remote sensing image; determining the spectral feature complexity and spatial structure complexity of each image block in the image block set; determining at least two spectral-spatial two-dimensional complexity levels and a stratification criterion corresponding to each spectral-spatial two-dimensional complexity level; for each spectral-spatial two-dimensional complexity level, determining an image block subset composed of image blocks belonging to the spectral-spatial two-dimensional complexity level from the image block set based on the stratification criterion, the spectral feature complexity, and the spatial structure complexity corresponding to the spectral-spatial two-dimensional complexity level; selecting a preset number of image block samples from each image block subset, and forming an image block training sample set for a semantic segmentation model based on the image block samples selected from each image block subset; wherein, the semantic segmentation model is a deep learning model used to perform remote sensing image classification tasks. This approach quantifies the scene complexity of image patches from two dimensions: spectral features and spatial structure. Based on this, it performs joint hierarchical and balanced sampling to construct a training sample set with high representativeness for the semantic segmentation model, effectively improving the accuracy and generalization ability of the semantic segmentation model in remote sensing image interpretation tasks.
[0017] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this application, nor is it intended to limit the scope of this application. Other features of this application will become readily apparent from the following description. Attached Figure Description
[0018] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0019] Figure 1 A flowchart illustrating a method for determining training samples for intelligent classification of high-resolution remote sensing images, provided as an embodiment of this application; Figure 2 This application provides a schematic diagram illustrating the conversion between an image block and its corresponding grayscale image. Figure 3 A schematic diagram illustrating the conversion of a grayscale image into a binary edge image, provided as an embodiment of this application; Figure 4 A schematic diagram illustrating the distribution of a spectral-spatial two-dimensional complexity hierarchy provided for an embodiment of this application; Figure 5This is a schematic diagram illustrating the training and evaluation process of a semantic segmentation model provided in an embodiment of this application. Figure 6 Performance comparison results of stratified sampling and simple random sampling under different sampling ratios on three models: U-Net, DeepLabV3, and MFFTNet, provided in the embodiments of this application; Figure 7 A visual comparison chart of urban green space interpretation results under different sampling methods provided in the embodiments of this application; Figure 8 A schematic diagram of a training sample determination device for intelligent classification of high-resolution remote sensing images provided in an embodiment of this application; Figure 9 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation
[0020] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present application, and not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative effort should fall within the scope of protection of the present application.
[0021] It should be noted that the terms "first," "second," "third," "fourth," "actual," "preset," etc., used in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0022] Figure 1This document provides a flowchart of a method for determining training samples for intelligent classification of high-resolution remote sensing images, applicable to situations where training samples for a semantic segmentation model performing a remote sensing image classification task are determined. This method can be executed by a training sample determination device for intelligent classification of high-resolution remote sensing images, which can be implemented in hardware and / or software and can be configured in an electronic device. Figure 1 As shown, the method includes: S110. Acquire at least one remote sensing image, crop each remote sensing image into at least two image blocks, and form an image block set based on each image block corresponding to each remote sensing image.
[0023] In this embodiment, at least one remote sensing image is acquired, wherein the high-sensitivity image can be a high-resolution remote sensing image with a resolution greater than a preset resolution threshold. Each remote sensing image is cropped into at least two image blocks, and an image block set is constructed based on the image blocks corresponding to each remote sensing image. The image blocks cropped from each remote sensing image can be image blocks of the same size and non-overlapping, such as image blocks of a fixed size of 256×256 pixels.
[0024] S120. Determine the spectral feature complexity and spatial structure complexity of each image block in the image block set.
[0025] In this embodiment, for each image block in the image block set, the spectral feature complexity and spatial structure complexity of the image block are determined respectively. The spectral feature complexity reflects the complexity of the spectral information of the image block, and the spatial structure complexity reflects the complexity of the spatial structure of the image block. The spectral feature complexity and spatial structure complexity together reflect the overall complexity of the image block. For example, the grayscale entropy of an image block can be used to characterize its spectral feature complexity, and the edge complexity of the image block can be used to characterize its spatial structure complexity.
[0026] Optionally, determining the spectral feature complexity of each image block in the image block set includes: for each image block in the image block set, converting the image block into a grayscale image, and determining a grayscale histogram corresponding to the image block based on the grayscale image; determining the probability of occurrence of each grayscale value in the image block based on the grayscale histogram; wherein the probability of occurrence is the proportion of pixels with the grayscale value in the image block; determining the grayscale entropy of the image block according to the probability of occurrence of each grayscale value, and using the grayscale entropy as the spectral feature complexity of the image block. Since remote sensing images are usually RGB images, for each image block in the image block set, converting the image block from an RGB image (i.e., RGB Image) to a grayscale image (i.e., Grayscale Image), for example, Figure 2 This document illustrates the conversion between an image block and its corresponding grayscale image, as provided in an embodiment of this application. The number of pixels corresponding to each grayscale value in the grayscale image is statistically analyzed, and a grayscale histogram corresponding to the image block is determined based on the statistical results. The probability of occurrence of each grayscale value in the image block is determined based on the grayscale histogram, where the probability of occurrence of a certain grayscale value is the proportion of pixels with that grayscale value in the image block, i.e., the ratio of the number of pixels corresponding to that grayscale value in the image block to the total number of pixels in the image block. The grayscale entropy of the image block is calculated according to the Shannon entropy formula based on the probability of occurrence of each grayscale value. For example, the grayscale entropy of the image block can be determined using the following formula. : in, Indicates grayscale value The probability of its occurrence. Gray entropy. This provides a quantitative basis for assessing the complexity of image blocks from the perspective of spectral statistical characteristics.
[0027] Optionally, determining the spatial structure complexity of each image block in the image block set includes: for each image block in the image block set, converting the image block into a grayscale image, performing edge detection on the grayscale image to determine a binary edge image; determining the number of edge pixels in the image block based on the binary edge image; determining the edge complexity of the image block based on the number of edge pixels, and using the edge complexity as the spatial structure complexity of the image block. For example, for each image block in the image block set, edge detection is performed on the grayscale image corresponding to the image block to determine a binary edge image (i.e., a binary edge image). Figure 3This is a schematic diagram illustrating the conversion of a grayscale image to a binary edge image, provided in an embodiment of this application. The number of edge pixels in an image block is determined based on the binary edge image, and the edge complexity of the image block is determined based on the number of edge pixels. For example, the edge complexity of the image block can be calculated using the following formula: in, Indicates the edge complexity of an image block. This represents the number of edge pixels in the image block. This represents the total number of pixels in the image block. This is an adjustable scaling factor. In the embodiments within this community, edge complexity is used as the spatial structural complexity of image blocks. Edge complexity provides a supplementary quantitative basis for evaluating the complexity of image blocks from the perspective of spatial structure, and together with grayscale entropy, it constitutes a comprehensive measurement framework for characterizing the complexity of image blocks from both spectral and spatial dimensions.
[0028] S130. Determine at least two spectral-spatial two-dimensional complexity levels and the stratification criteria corresponding to each of the spectral-spatial two-dimensional complexity levels.
[0029] In this embodiment, at least two pre-defined spectral-spatial two-dimensional complexity levels are determined. Each spectral-spatial two-dimensional complexity level is a two-dimensional complexity level formed by a two-dimensional cross-combination of spectral complexity levels and spatial complexity levels. Pre-defined hierarchical determination criteria are obtained for each spectral-spatial two-dimensional complexity level. Each hierarchical determination criterion for each spectral-spatial two-dimensional complexity level is a combination of the determination criteria for the corresponding spectral complexity level and the determination criteria for the spatial complexity level.
[0030] Optionally, determining at least two spectral-spatial two-dimensional complexity levels and a stratification criterion corresponding to each spectral-spatial two-dimensional complexity level includes: determining a pre-defined spectral complexity level and a spatial complexity level; wherein the number of spectral complexity levels is at least two, and / or the number of spatial complexity levels is at least two; determining a spectral stratification threshold range corresponding to each spectral complexity level and a spatial stratification threshold range corresponding to each spatial complexity level; determining at least two spectral-spatial two-dimensional complexity levels based on the spectral complexity levels and the spatial complexity levels, and determining a stratification criterion corresponding to each spectral-spatial two-dimensional complexity level based on the spectral stratification threshold range corresponding to each spectral complexity level and the spatial stratification threshold range corresponding to each spatial complexity level.
[0031] In this embodiment, pre-defined spectral complexity levels and spatial complexity levels are determined, and the spectral stratification threshold range corresponding to each spectral complexity level and the spatial stratification threshold range corresponding to each spatial complexity level are also determined. The number of spectral complexity levels is at least two, and / or the number of spatial complexity levels is at least two. The spectral stratification threshold range corresponding to each spectral complexity level and the spatial stratification threshold range corresponding to each spatial complexity level can be pre-defined based on actual conditions. Then, a spectral-spatial two-dimensional complexity level is determined based on a two-dimensional cross-combination of the spectral complexity levels and spatial complexity levels. For example, if the pre-defined spectral complexity levels include three levels (low, medium, and high), and the pre-defined spatial complexity levels include two levels (low and high), then 3 × 2 = 6 spectral-spatial two-dimensional complexity levels can be determined based on the two-dimensional cross-combination of the spectral complexity levels and spatial complexity levels. Based on the spectral stratification threshold range corresponding to each spectral complexity level and the spatial stratification threshold range corresponding to each spatial complexity level, the stratification determination criterion corresponding to each spectral-spatial two-dimensional complexity level is determined. For example, the spectral stratification threshold ranges corresponding to the low, medium, and high spectral complexity levels are [E1, E2), [E2, E3), and [E3, E4], respectively; the spatial stratification threshold ranges corresponding to the low and high spatial complexity levels are [C1, C2) and [C2, C3], respectively. Therefore, the stratification determination criterion corresponding to the low spectral-low spatial two-dimensional complexity level is: and The stratification criteria corresponding to the low-spectral-high spatial two-dimensional complexity hierarchy are as follows: and The stratification criteria corresponding to the mid-spectral-low spatial two-dimensional complexity level are as follows: and The stratification criteria corresponding to the mid-spectral-high spatial two-dimensional complexity level are as follows: and The stratification criteria corresponding to the hyperspectral-low spatial two-dimensional complexity hierarchy are as follows: and The stratification criteria corresponding to the hyperspectral-high spatial two-dimensional complexity hierarchy are as follows: and .
[0032] Optionally, determining the spectral stratification threshold range corresponding to each spectral complexity level and the spatial stratification threshold range corresponding to each spatial complexity level includes: taking the spectral feature complexity and the spatial structure complexity as target complexities respectively, determining the effective value range of the complexity index based on the target complexity of each image block in the image block set; determining the stratification step size of each complexity index level based on the number of complexity index levels corresponding to the target complexity and the effective value range of the complexity index, and determining the complexity index stratification threshold range corresponding to the complexity index level based on the stratification step size and the effective value range of the complexity index. The advantage of this approach is that it effectively improves the rationality of determining the spectral stratification threshold range corresponding to each spectral complexity level and the spatial stratification threshold range corresponding to each spatial complexity level, thereby improving the rationality of determining the stratification judgment criteria corresponding to each spectral-spatial two-dimensional complexity level.
[0033] In this embodiment, spectral feature complexity and spatial structure complexity are used as target complexities, respectively. The effective value range of the complexity index is determined based on the target complexity of each image block in the image block set. For example, the maximum and minimum values of the target complexity of all image blocks in the image block set are determined, and the effective value range of the complexity index is determined based on these values. The upper limit of the effective value range of the complexity index is the maximum value of the target complexity, and the lower limit is the minimum value of the target complexity. Optionally, the 5th percentile of the target complexity of all image blocks in the image block set is determined. ) and the 95th percentile ( ), the 5th percentile of the objective complexity ( ) is used as the lower limit of the effective range of the complexity index, and the 95th percentile of the target complexity is ( This serves as the upper limit of the effective range for the complexity index, in order to exclude the influence of extreme values on the stratification results.
[0034] Based on the number of complexity index levels corresponding to the target complexity and the effective value range of the complexity index, the level step size of each complexity index level is determined, where the level step size is the ratio of the length of the effective value range of the complexity index to the number of complexity index levels.
[0035] Then, based on the hierarchical step size and the effective value range of the complexity index, the stratification threshold range corresponding to the complexity index level is determined. For example, if the number of spectral complexity levels is the same as the number of spatial complexity levels and both are 3, then the hierarchical step size for each complexity index level can be determined according to the following formula: in, This represents the step size of the complexity index level. This indicates the complexity of spectral features respectively. and spatial structure complexity As the target complexity.
[0036] For example, if the spectral complexity hierarchy includes three levels: Low, Medium, and High, and the spatial complexity hierarchy also includes three levels: Low, Medium, and High, then the threshold range for each complexity index level is as follows: It should be noted that percentiles are only used to determine a reasonable range for the stratification threshold; in the actual stratification stage, the spectral feature complexity and spatial structure complexity of all image patches in the image patch set (including those located in...) are considered. The following or All the above samples are mapped to the corresponding complexity level according to the above-mentioned hierarchical judgment criteria, ensuring that all image blocks in the image block set are included in the hierarchical system.
[0037] It is understandable that, based on the two-dimensional cross combination of spectral complexity level and spatial complexity level, the entire sample space can be divided into 3×3=9 spectral-spatial two-dimensional complexity levels. Figure 4 This is a schematic diagram of the distribution of a spectral-spatial two-dimensional complexity hierarchy provided for an embodiment of this application.
[0038] S140. For each of the spectral-spatial two-dimensional complexity levels, based on the hierarchical determination criteria, the spectral feature complexity, and the spatial structure complexity corresponding to the spectral-spatial two-dimensional complexity level, determine a subset of image blocks from the image block set that consists of image blocks belonging to the spectral-spatial two-dimensional complexity level.
[0039] In this embodiment, for each image block in the image block set, the spectral-spatial two-dimensional complexity level to which the image block belongs is determined based on the spectral feature complexity, spatial structure complexity, and the hierarchical determination criteria corresponding to each spectral-spatial two-dimensional complexity level. It is understood that the spectral-spatial two-dimensional complexity level to which each image block in the image block set belongs can be determined in the above manner. Then, the set of all image blocks belonging to the same spectral-spatial two-dimensional complexity level in the image block set is called the image block subset corresponding to that spectral-spatial two-dimensional complexity level. It is understood that the image block subset belonging to each spectral-spatial two-dimensional complexity level can be determined from the image block set in the above manner.
[0040] S150. Select a preset number of image block samples from each of the image block subsets, and construct an image block training sample set for the semantic segmentation model based on the image block samples selected from each of the image block subsets; wherein, the semantic segmentation model is a deep learning model used to perform remote sensing image classification tasks.
[0041] In this embodiment, for each spectral-spatial two-dimensional complexity level, a predetermined number of image patch samples are selected from the image patch subset corresponding to that level. It is understood that the number of image patch samples selected from each spectral-spatial two-dimensional complexity level subset is the same. Then, an image patch training sample set for the semantic segmentation model is constructed based on the image patch samples selected from each image patch subset. That is, based on the image patch training sample set constructed from the image patch samples selected from each image patch subset, a predetermined deep learning model is trained to generate a semantic segmentation model for performing remote sensing image classification tasks. It is understood that a balanced sampling strategy is used to construct the image patch training sample set. Given a total training sample size, the same number of image patch samples are extracted from each of the nine spectral-spatial two-dimensional complexity levels to ensure that different complexity levels contribute evenly during model training.
[0042] The training sample determination scheme provided in this application involves acquiring at least one remote sensing image, cropping each remote sensing image into at least two image blocks, and forming an image block set based on each image block corresponding to each remote sensing image; determining the spectral feature complexity and spatial structure complexity of each image block in the image block set; determining at least two spectral-spatial two-dimensional complexity levels and a stratification criterion corresponding to each spectral-spatial two-dimensional complexity level; for each spectral-spatial two-dimensional complexity level, determining an image block subset composed of image blocks belonging to the spectral-spatial two-dimensional complexity level from the image block set based on the stratification criterion, the spectral feature complexity, and the spatial structure complexity corresponding to the spectral-spatial two-dimensional complexity level; selecting a preset number of image block samples from each image block subset, and forming an image block training sample set for a semantic segmentation model based on the image block samples selected from each image block subset; wherein, the semantic segmentation model is a deep learning model used to perform remote sensing image classification tasks. This approach quantifies the scene complexity of image patches from two dimensions: spectral features and spatial structure. Based on this, it performs joint hierarchical and balanced sampling to construct a training sample set with high representativeness for the semantic segmentation model, effectively improving the accuracy and generalization ability of the semantic segmentation model in remote sensing image interpretation tasks.
[0043] In some embodiments, after constructing an image block training sample set for a semantic segmentation model based on image block samples selected from each of the image block subsets, the method further includes: determining a remote sensing image classification label corresponding to each image block in the image block training sample set; wherein the remote sensing image classification label includes the urban land cover type corresponding to each pixel in the image block; and training a preset deep learning model based on the image block training sample set and the remote sensing image classification label corresponding to each image block in the image block training sample set to generate a semantic segmentation model.
[0044] In this embodiment, for each image block in the image block training sample set, the urban feature type corresponding to each pixel in the image block is determined, and the urban feature type corresponding to each pixel in the image block is used as the remote sensing image classification label for that image block. The urban feature type can include various features such as green space, buildings, and roads. Based on the image block training sample set and the remote sensing image classification label corresponding to each image block in the image block training sample set, a preset deep learning model is iteratively trained until preset training conditions are met, generating a semantic segmentation model, thereby enabling intelligent extraction of urban features based on the semantic segmentation model. The preset training conditions can include the number of iterations reaching a preset threshold, and can also include the loss value calculated based on the loss function corresponding to the preset deep learning model being less than a preset loss threshold.
[0045] Optionally, after generating the semantic segmentation model, the method further includes: obtaining an image patch test sample set; determining a segmentation performance evaluation index for the semantic segmentation model based on the image patch test sample set, and evaluating the segmentation performance of the semantic segmentation model based on the segmentation performance evaluation index; wherein the segmentation performance evaluation index includes pixel accuracy and average intersection-over-union ratio (IoU); the pixel accuracy is the ratio of the number of pixels correctly classified by the semantic segmentation model in the image patch test sample set to the total number of pixels in the image patch test sample set; the average IoU is the average of the IoU values of the number of pixels in all categories classified by the semantic segmentation model in the image patch test sample set.
[0046] In this embodiment, an image patch test sample set is obtained, wherein the image patch test sample set contains at least one image patch test sample. Optionally, after obtaining at least one remote sensing image and cropping it into image patches of a fixed size (256×256 pixels) that do not overlap, all image patches can be used as basic sampling units. Then, simple random sampling is used to extract approximately 20% of the image patches from the overall basic sampling units (i.e., all 256×256 image patches) as an independent image patch test sample set. The remaining 80% of the image patches constitute the original training sample pool for selecting the image patch training sample set. Based on the image patch test sample set, the segmentation performance evaluation index of the semantic segmentation model is determined, wherein the segmentation performance evaluation index includes pixel accuracy and average intersection-over-union ratio (IoU). Specifically, the image patch test sample set is input into the semantic segmentation model, and based on the output of the semantic segmentation model, the predicted urban land cover type corresponding to each pixel in each image patch test sample in the image patch test sample set is determined. The predicted urban feature type for each pixel in each image patch test sample set is compared with the corresponding real urban feature type to determine the total number of correctly classified pixels in the image patch test sample set. Correctly classified pixels are those whose predicted urban feature type matches their corresponding real urban feature type. Then, the pixel accuracy PA is calculated using the following formula: in, The number of classification categories for urban land cover types. Indicates the first The number of pixels correctly classified in the class (i.e., the number of real urban features) And the urban land cover type prediction results are (number of pixels) Indicates the true category is But it was predicted to be The number of pixels.
[0047] In this embodiment of the application, the average crossover ratio (MIoU) is calculated according to the following formula: in, For the number of categories, Indicates the first The number of pixels in the class that are correctly classified (i.e., the true class is 0) And it was predicted to be (number of pixels) Indicates the true category is But it was predicted to be The number of pixels, Indicates the true category is But it was predicted to be The number of pixels. It is understandable that the average Intersection over Union (MIoU) is the average of the intersection over union ratios for all classes.
[0048] In this embodiment, the segmentation performance of the semantic segmentation model is evaluated based on pixel precision and average intersection-union ratio (IU). A higher pixel precision indicates better segmentation performance of the semantic segmentation model, and a higher IU also indicates better segmentation performance of the semantic segmentation model.
[0049] In this embodiment of the application, the Urban Green Space Dataset (UGSet), constructed based on GF-2 satellite data, is used as an example to illustrate the specific implementation of this embodiment. Figure 5 This is a schematic diagram illustrating the training and evaluation process of a semantic segmentation model provided in an embodiment of this application.
[0050] Step 1: Data Preparation and Preprocessing. The UGSet dataset contains 4544 images, each 512×512 pixels in size, with a spatial resolution of approximately 1m. Using Python image processing libraries, the original RGB images and their corresponding labels were cropped into 256×256 pixel non-overlapping image blocks, resulting in 18176 sample units. Subsequently, approximately 20% of these samples were randomly selected as the independent test set, with the remaining 80% forming the original training sample pool. For subsequent feature analysis, the rgb2gray function from the skimage library was used to convert all RGB image blocks into grayscale images.
[0051] Step 2: Image Patch Complexity Measurement. For each image patch in the original training sample pool, two core metrics are calculated based on its grayscale image: grayscale entropy and edge complexity. First, a histogram of grayscale values from 0 to 255 levels is plotted, and grayscale entropy is calculated using the Shannon entropy formula to quantify the complexity of the grayscale value distribution. Second, the grayscale image is normalized, and edge features are extracted using the Canny edge detection algorithm with sigma=2.0. The proportion of edge pixels to total pixels is calculated and multiplied by a scaling factor of 10 to obtain the edge complexity. Simultaneously, the binary image after edge detection is output. During batch processing, the filenames, grayscale entropy, and edge complexity results of all image patches are summarized and saved as a CSV file, and the corresponding edge images are uniformly stored in a designated folder to provide a quantitative basis for subsequent stratified sampling.
[0052] Step 3: Construct a highly representative training sample set based on joint stratification using gray-level entropy and marginal complexity. First, calculate the 5th and 95th percentiles of the distributions of gray-level entropy and marginal complexity in the original sample pool to determine the effective value ranges for each indicator, thus eliminating the interference of extreme values on subsequent stratification. Based on this, divide the effective range length of each indicator into three equal parts, classifying gray-level entropy and marginal complexity into low, medium, and high levels respectively. Samples outside the effective value range (below the 5th percentile) are classified as low-level, and those above the 95th percentile are classified as high-level. After completing the single-indicator stratification, the levels of gray-level entropy and marginal complexity are combined in a two-dimensional cross-combination, thereby dividing the sample space into 3×3=9 different complexity levels. Finally, given the total training sample size, a balanced sampling strategy is adopted to extract the same number of image block samples from each of the nine levels. If the number of samples in a certain level is less than the number that should be extracted from that level, all samples from that level are extracted, and the remaining required number of samples are evenly distributed to other levels to ensure that the total sample size remains unchanged.
[0053] Step 4: Model Training and Accuracy Evaluation. Based on the training sample set constructed in Step 3, three representative semantic segmentation networks—U-Net, DeepLabV3, and MFFTNet—were used for model training, and their performance was compared with that of a model trained using samples selected through simple random sampling. To comprehensively evaluate the effectiveness of different sampling strategies, the training sample size was set to 1% to 10% of the total original training sample pool, with an interval of 1%. At each sampling ratio, stratified sampling and simple random sampling were used to extract samples from the original training sample pool according to the corresponding proportion for model training. Pixel accuracy (PA) and mean intersection-over-union ratio (MIoU) were used to evaluate the segmentation performance of the model. Figure 6 This document presents a performance comparison of stratified sampling and simple random sampling under different sampling ratios on three models: U-Net, DeepLabV3, and MFFTNet, as provided in the embodiments of this application. Figure 6The red line represents the stratified sampling result, and the blue line represents the random sampling result; the dashed lines indicate the accuracy values and corresponding sampling ratios when the two sampling methods reach their highest accuracy. Subfigures (a)–(f) show the performance comparison of different models under the two metrics: Subfigure (a) is the curve of PA value of U-Net model changing with sampling ratio; Subfigure (b) is the curve of MIoU value of U-Net model changing with sampling ratio; Subfigure (c) is the curve of PA value of DeepLabV3 model changing with sampling ratio; Subfigure (d) is the curve of MIoU value of DeepLabV3 model changing with sampling ratio; Subfigure (e) is the curve of PA value of MFFTNet model changing with sampling ratio; Subfigure (f) is the curve of MIoU value of MFFTNet model changing with sampling ratio. Table 1 shows the PA comparison results of different deep learning models provided in the embodiments of this application using random and stratified sampling on UGSet, and Table 2 shows the MIoU comparison results of different deep learning models provided in the embodiments of this application using random and stratified sampling on UGSet. Step 5: Model Prediction and Visualization. Based on the trained model, predict the test set and visualize the extraction results of typical regions. For example, Figure 7 A visual comparison of urban green space interpretation results under different sampling methods provided in the embodiments of this application (taking the DeepLabV3 model with a 10% sampling ratio as an example).
[0054] In the urban green space extraction experiment in the example, the stratified sampling method significantly outperformed traditional simple random sampling in both PA and MIoU under the same sample size, especially in the low sample proportion stage of 1%–4%, where the advantage was more prominent, with the highest MIoU improvement reaching 7.5%. Under the condition of achieving the same accuracy, the training sample size required by the stratified sampling method is much smaller than that of random sampling, and the proportion of training samples can be reduced from 10% to 4%, showing the potential to significantly reduce the cost of manual annotation. Figure 7 shows the prediction results of a typical area. As can be seen from Figure 7, the segmentation results of the present invention have significantly higher accuracy in boundary restoration, and the loss of detailed information and misjudgment are significantly reduced. This indicates that the embodiments of this application use gray-level entropy and edge complexity as core indicators to quantify the scene complexity of image blocks from two dimensions: spectral information distribution and spatial structure features. Based on this, a two-dimensional joint stratification system is constructed, and balanced sampling is used to achieve balanced coverage of complex geographical scenes by training samples, which can effectively improve the representativeness of training samples and the generalization ability of the model.
[0055] Compared with existing deep learning-based semantic segmentation training sample selection methods, the main advantage of the technical solution provided in this application is that the entire sample selection process does not rely on any external prior knowledge or pre-annotated information, but is completed solely based on the image's own features, realizing a paradigm shift from "relying on priors" to "data-driven". This application uses grayscale entropy and edge complexity as core indicators to quantify the scene complexity of image patches from two dimensions: spectral information distribution and spatial structure features. Based on this, a two-dimensional joint hierarchical system is constructed. Balanced sampling achieves balanced coverage of complex geographical scenes by training samples, achieving higher urban feature interpretation accuracy than simple random sampling with the same training sample size. Under the condition of achieving the same extraction accuracy, the required number of training samples can be significantly reduced, thereby effectively reducing manual annotation costs.
[0056] Figure 8 This is a schematic diagram of a training sample determination device for intelligent classification of high-resolution remote sensing images, provided as an embodiment of this application. This device can execute the training sample determination method for intelligent classification of high-resolution remote sensing images provided in any embodiment of this application, and possesses the corresponding functional modules and beneficial effects of the method. For example... Figure 8 As shown, the device includes: The image block set determination module 810 is used to acquire at least one remote sensing image, crop each remote sensing image into at least two image blocks, and form an image block set based on each image block corresponding to each remote sensing image. The complexity determination module 820 is used to determine the spectral feature complexity and spatial structure complexity of each image block in the image block set, respectively. The stratification determination criterion module 830 is used to determine at least two spectral-spatial two-dimensional complexity levels and the stratification determination criterion corresponding to each spectral-spatial two-dimensional complexity level. The image block subset determination module 840 is used to determine, for each of the spectral-spatial two-dimensional complexity levels, an image block subset consisting of image blocks belonging to the spectral-spatial two-dimensional complexity level from the image block set, based on the hierarchical determination criteria, the spectral feature complexity, and the spatial structure complexity corresponding to the spectral-spatial two-dimensional complexity level. The image patch training sample set determination module 850 is used to select a preset number of image patch samples from each of the image patch subsets, and to construct an image patch training sample set for the semantic segmentation model based on the image patch samples selected from each of the image patch subsets; wherein, the semantic segmentation model is a deep learning model for performing remote sensing image classification tasks.
[0057] Optional, complexity determination module, used for: For each image block in the image block set, the image block is converted into a grayscale image, and a grayscale histogram corresponding to the image block is determined based on the grayscale image. The probability of occurrence of each gray value in the image block is determined based on the gray-level histogram; wherein, the probability of occurrence is the proportion of pixels with the gray value in the image block; The gray entropy of the image block is determined based on the probability of occurrence of each gray value, and the gray entropy is used as the spectral feature complexity of the image block.
[0058] Optional, complexity determination module, used for: For each image block in the image block set, the image block is converted into a grayscale image, and edge detection is performed on the grayscale image to determine a binary edge image; The number of edge pixels in the image block is determined based on the binary edge image; The edge complexity of the image block is determined based on the number of edge pixels, and the edge complexity is used as the spatial structure complexity of the image block.
[0059] Optionally, the stratification determination criterion module includes: A complexity level determination unit is used to determine a pre-defined spectral complexity level and a spatial complexity level; wherein the number of spectral complexity levels is at least two, and / or the number of spatial complexity levels is at least two. The stratification threshold range determination unit is used to determine the spectral stratification threshold range corresponding to each spectral complexity level and the spatial stratification threshold range corresponding to each spatial complexity level. The stratification determination criterion unit is used to determine at least two spectral-spatial two-dimensional complexity levels based on the spectral complexity level and the spatial complexity level, and to determine the stratification determination criterion corresponding to each spectral-spatial two-dimensional complexity level based on the spectral stratification threshold range corresponding to each spectral complexity level and the spatial stratification threshold range corresponding to each spatial complexity level.
[0060] Optionally, a stratified threshold range determination unit is used for: The spectral feature complexity and the spatial structure complexity are respectively used as target complexity, and the effective value range of the complexity index is determined according to the target complexity of each image block in the image block set. Based on the number of complexity index levels corresponding to the target complexity and the effective value range of the complexity index, the level step size of each complexity index level is determined, and the layer threshold range of the complexity index corresponding to the complexity index level is determined based on the level step size and the effective value range of the complexity index.
[0061] Optional, also includes: The remote sensing image classification label determination module is used to determine the remote sensing image classification label corresponding to each image block in the image block training sample set after the image block samples selected from each of the image block subsets constitute a semantic segmentation model; wherein, the remote sensing image classification label includes the urban land cover type corresponding to each pixel in the image block. The semantic segmentation model generation module is used to train a preset deep learning model based on the image block training sample set and the remote sensing image classification label corresponding to each image block in the image block training sample set, and generate a semantic segmentation model.
[0062] Optional, also includes: The image patch test sample set acquisition module is used to acquire the image patch test sample set after the semantic segmentation model is generated. A segmentation performance evaluation module is used to determine the segmentation performance evaluation index of the semantic segmentation model based on the image patch test sample set, and to evaluate the segmentation performance of the semantic segmentation model based on the segmentation performance evaluation index; wherein, the segmentation performance evaluation index includes pixel accuracy and average intersection-over-union ratio (IoU); the pixel accuracy is the ratio of the number of pixels correctly classified by the semantic segmentation model in the image patch test sample set to the total number of pixels in the image patch test sample set; the average IoU is the average of the IoU of the number of pixels in all categories classified by the semantic segmentation model in the image patch test sample set.
[0063] The training sample determination device for intelligent classification of high-resolution remote sensing images provided in this application embodiment can execute the training sample determination method for intelligent classification of high-resolution remote sensing images provided in any embodiment of this application, and has the corresponding functional modules and beneficial effects of the execution method.
[0064] Figure 9 A schematic diagram of an electronic device 10, which can be used to implement embodiments of this application, is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices (such as helmets, glasses, watches, etc.), and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the application described and / or claimed herein.
[0065] like Figure 9As shown, the electronic device 10 includes at least one processor 11 and a memory, such as a read-only memory (ROM) 12 or a random access memory (RAM) 13, communicatively connected to the at least one processor 11. The memory stores computer programs executable by the at least one processor. The processor 11 can perform various appropriate actions and processes based on the computer program stored in the ROM 12 or loaded from storage unit 18 into the RAM 13. The RAM 13 can also store various programs and data required for the operation of the electronic device 10. The processor 11, ROM 12, and RAM 13 are interconnected via a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.
[0066] Multiple components in electronic device 10 are connected to I / O interface 15, including: input unit 16, such as keyboard, mouse, etc.; output unit 17, such as various types of monitors, speakers, etc.; storage unit 18, such as disk, optical disk, etc.; and communication unit 19, such as network card, modem, wireless transceiver, etc. Communication unit 19 allows electronic device 10 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.
[0067] Processor 11 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, digital signal processors (DSPs), and any suitable processor, controller, microcontroller, etc. Processor 11 performs the various methods and processes described above, such as training sample determination methods for intelligent classification of high-resolution remote sensing images.
[0068] In some embodiments, the training sample determination method for intelligent classification of high-resolution remote sensing images can be implemented as a computer program tangibly contained in a computer-readable storage medium, such as storage unit 18. In some embodiments, part or all of the computer program can be loaded and / or installed on electronic device 10 via ROM 12 and / or communication unit 19. When the computer program is loaded into RAM 13 and executed by processor 11, one or more steps of the training sample determination method for intelligent classification of high-resolution remote sensing images described above can be performed. Alternatively, in other embodiments, processor 11 can be configured to perform the training sample determination method for intelligent classification of high-resolution remote sensing images by any other suitable means (e.g., by means of firmware).
[0069] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.
[0070] Computer programs used to implement the methods of this application may be written in any combination of one or more programming languages. These computer programs may be provided to the processor of a general-purpose computer, a special-purpose computer, or other programmable training sample determination device for intelligent classification of high-resolution remote sensing images, such that when executed by the processor, the computer programs cause the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The computer programs may be executed entirely on a machine, partially on a machine, or as a standalone software package, partially on a machine and partially on a remote machine, or entirely on a remote machine or server.
[0071] In the context of this application, a computer-readable storage medium can be a tangible medium that may contain or store a computer program for use by or in conjunction with an instruction execution system, apparatus, or device. A computer-readable storage medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. Alternatively, a computer-readable storage medium can be a machine-readable signal medium. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0072] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the electronic device. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).
[0073] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as data servers), or middleware components (e.g., application servers), or frontend components (e.g., user computers with graphical user interfaces or web browsers through which users can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., communication networks). Examples of communication networks include local area networks (LANs), wide area networks (WANs), blockchain networks, and the Internet.
[0074] A computing system can include clients and servers. Clients and servers are generally located far apart and typically interact through a communication network. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or cloud host, which is a hosting product within the cloud computing service system to address the shortcomings of traditional physical hosts and VPS services, such as high management difficulty and weak business scalability.
[0075] This invention also provides a computer program product, including a computer program that, when executed by a processor, implements the training sample determination method for intelligent classification of high-resolution remote sensing images as provided in any embodiment of this application.
[0076] In implementing the computer program product, computer program code for performing the operations of this invention can be written in one or more programming languages or a combination thereof. Programming languages include object-oriented programming languages such as Java, Smalltalk, and C++, as well as conventional procedural programming languages such as C or similar languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or it can be connected to an external computer (e.g., via the Internet using an Internet service provider). It should be understood that the various forms of processes shown above can be used to rearrange, add, or delete steps. For example, the steps described in this application can be executed in parallel, sequentially, or in different orders, as long as the desired information of the technical solution of this application can be achieved, and this is not limited herein.
[0077] The specific embodiments described above do not constitute a limitation on the scope of protection of this application. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this application should be included within the scope of protection of this application.
Claims
1. A method for determining training samples for intelligent classification of high-resolution remote sensing images, characterized in that, The method includes: Acquire at least one remote sensing image, crop each remote sensing image into at least two image blocks, and construct an image block set based on each image block corresponding to each remote sensing image; The spectral feature complexity and spatial structure complexity of each image block in the image block set are determined respectively; Determine at least two spectral-spatial two-dimensional complexity levels and the stratification criteria corresponding to each spectral-spatial two-dimensional complexity level; For each of the spectral-spatial two-dimensional complexity levels, based on the hierarchical determination criteria, the spectral feature complexity, and the spatial structure complexity corresponding to the spectral-spatial two-dimensional complexity level, a subset of image blocks belonging to the spectral-spatial two-dimensional complexity level is determined from the image block set. A predetermined number of image block samples are selected from each of the image block subsets, and an image block training sample set for the semantic segmentation model is constructed based on the image block samples selected from each of the image block subsets; wherein, the semantic segmentation model is a deep learning model used to perform remote sensing image classification tasks.
2. The method according to claim 1, characterized in that, Determining the spectral feature complexity of each image block in the image block set includes: For each image block in the image block set, the image block is converted into a grayscale image, and a grayscale histogram corresponding to the image block is determined based on the grayscale image. The probability of occurrence of each gray value in the image block is determined based on the gray-level histogram; wherein, the probability of occurrence is the proportion of pixels with the gray value in the image block; The gray entropy of the image block is determined based on the probability of occurrence of each gray value, and the gray entropy is used as the spectral feature complexity of the image block.
3. The method according to claim 1, characterized in that, Determine the spatial structural complexity of each image block in the image block set, including: For each image block in the image block set, the image block is converted into a grayscale image, and edge detection is performed on the grayscale image to determine a binary edge image; The number of edge pixels in the image block is determined based on the binary edge image; The edge complexity of the image block is determined based on the number of edge pixels, and the edge complexity is used as the spatial structure complexity of the image block.
4. The method according to claim 1, characterized in that, Determine at least two spectral-spatial two-dimensional complexity levels and the corresponding hierarchical determination criteria for each spectral-spatial two-dimensional complexity level, including: Determine a pre-defined spectral complexity level and a spatial complexity level; wherein the number of spectral complexity levels is at least two, and / or the number of spatial complexity levels is at least two; Determine the spectral stratification threshold range corresponding to each spectral complexity level and the spatial stratification threshold range corresponding to each spatial complexity level; Based on the spectral complexity level and the spatial complexity level, at least two spectral-spatial two-dimensional complexity levels are determined, and based on the spectral stratification threshold range corresponding to each spectral complexity level and the spatial stratification threshold range corresponding to each spatial complexity level, a stratification determination criterion corresponding to each spectral-spatial two-dimensional complexity level is determined.
5. The method according to claim 4, characterized in that, Determining the spectral stratification threshold range corresponding to each of the spectral complexity levels and the spatial stratification threshold range corresponding to each of the spatial complexity levels includes: The spectral feature complexity and the spatial structure complexity are respectively used as target complexity, and the effective value range of the complexity index is determined according to the target complexity of each image block in the image block set. Based on the number of complexity index levels corresponding to the target complexity and the effective value range of the complexity index, the level step size of each complexity index level is determined, and the layer threshold range of the complexity index corresponding to the complexity index level is determined based on the level step size and the effective value range of the complexity index.
6. The method according to claim 1, characterized in that, After constructing the image patch training sample set for the semantic segmentation model based on the image patch samples selected from each of the image patch subsets, the method further includes: For each image block in the image block training sample set, a remote sensing image classification label corresponding to the image block is determined; wherein, the remote sensing image classification label includes the urban land cover type corresponding to each pixel in the image block; The preset deep learning model is trained based on the image patch training sample set and the remote sensing image classification label corresponding to each image patch in the image patch training sample set to generate a semantic segmentation model.
7. The method according to claim 6, characterized in that, After generating the semantic segmentation model, the following is also included: Obtain the image patch test sample set; The segmentation performance evaluation index of the semantic segmentation model is determined based on the image patch test sample set, and the segmentation performance of the semantic segmentation model is evaluated based on the segmentation performance evaluation index; wherein, the segmentation performance evaluation index includes pixel accuracy and average intersection-over-union ratio (IoU); the pixel accuracy is the ratio of the number of pixels correctly classified by the semantic segmentation model in the image patch test sample set to the total number of pixels in the image patch test sample set; the average IoU is the average of the IoU of the number of pixels in all categories classified by the semantic segmentation model in the image patch test sample set.
8. A training sample determination device for intelligent classification of high-resolution remote sensing images, characterized in that, include: The image block set determination module is used to acquire at least one remote sensing image, crop each remote sensing image into at least two image blocks, and form an image block set based on each image block corresponding to each remote sensing image. The complexity determination module is used to determine the spectral feature complexity and spatial structure complexity of each image block in the image block set, respectively. The stratification determination criterion module is used to determine at least two spectral-spatial two-dimensional complexity levels and the stratification determination criterion corresponding to each of the spectral-spatial two-dimensional complexity levels. The image block subset determination module is used to determine, for each of the spectral-spatial two-dimensional complexity levels, an image block subset consisting of image blocks belonging to the spectral-spatial two-dimensional complexity level from the image block set, based on the hierarchical determination criteria, the spectral feature complexity, and the spatial structure complexity corresponding to the spectral-spatial two-dimensional complexity level. The image patch training sample set determination module is used to select a preset number of image patch samples from each of the image patch subsets, and to construct an image patch training sample set for the semantic segmentation model based on the image patch samples selected from each of the image patch subsets; wherein, the semantic segmentation model is a deep learning model used to perform remote sensing image classification tasks.
9. An electronic device, characterized in that, The electronic device includes: At least one processor; and A memory communicatively connected to the at least one processor; wherein, The memory stores a computer program executable by the at least one processor, which enables the at least one processor to perform the training sample determination method for intelligent classification of high-resolution remote sensing images according to any one of claims 1-7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions that cause a processor to execute the training sample determination method for intelligent classification of high-resolution remote sensing images as described in any one of claims 1-7.
Citation Information
Patent Citations
Space-spectrum information combined spaceborne hyperspectral image segmentation and clustering method
CN113902759A
Hyperspectral technology-based eggplant core tobacco leaf grading method and system
CN117788920A