Construction method of UAV remote sensing image classification model

By using technologies such as hierarchical feature decomposition and graph convolutional networks in the drone remote sensing image classification model, the shortcomings of the existing technology in handling complex geographic scenarios and large-scale data are solved, and higher classification accuracy and robustness are achieved, and the dependence on labeled data is reduced.

CN119399655BActive Publication Date: 2025-05-13深圳市南湖勘测技术有限公司
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202411484395.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-10-23
Publication Date
2025-05-13
Estimated Expiration
2044-10-23

AI Technical Summary

Technical Problem

The existing drone remote sensing image classification model has shortcomings in processing complex geographic scenarios, processing large-scale data, solving the scarcity of labeling data, and capturing dynamic changes of geographic objects.

Method used

The structured representation method based on hierarchical feature decomposition is adopted to integrate dynamic feature capture of spatial topological relationships, complex spatial and temporal information is extracted through graph convolution networks and dynamic graph convolution networks, and self-supervised learning is performed on labelless data to reduce dependence on labeled data.

Benefits of technology

It improves classification accuracy and robustness, reduces dependence on labeled data, enhances the adaptability and generalization capabilities of the model, and can better handle complex geographic scenarios and large-scale data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119399655B_ABST
    Figure CN119399655B_ABST
Patent Text Reader

Abstract

The present invention relates to a method for constructing a classification model for remote sensing images of unmanned aerial vehicles. Structured representation of remote sensing images based on hierarchical feature decomposition; UAV images are divided into image blocks of different levels through a pyramid decomposition algorithm, and hierarchical combination is performed for the features of each level, and a multi-scale structured feature vector is formed through a feature weighting mechanism; dynamic feature capture integrating spatial topological relationships: after constructing a graph structure, a graph convolutional network (GCN) is used to aggregate features of each node, propagate spatial relationship information, so that the features of each node contain its own information, and integrate the influence of surrounding objects; for multi-temporal image data, a dynamic graph convolutional network (D-GCN) is introduced to model spatial relationships at different times, capture the changes of objects over time and their influence on classification; self-supervised clustering learning of object features: for objects that are difficult to distinguish, a soft clustering method is introduced to make the objects belong to multiple categories at the same time.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of UAV remote sensing image classification, and in particular to a method for constructing a UAV remote sensing image classification model. Background Art

[0002] In China Invention Publication No. CN117893927A, an intelligent pixel-level classification method suitable for UAV hyperspectral remote sensing images is proposed. This method processes UAV hyperspectral remote sensing images, constructs an attention network based on the reconstruction module, and trains the hyperspectral image training data set. Finally, the UAV images are classified through a lightweight neural network to generate a classification map. Although this method has the advantages of extracting high-level features, reducing network complexity, and improving classification performance, especially in meeting the real-time and accuracy requirements of UAVs, this method still has some shortcomings and drawbacks, which are mainly reflected in the following aspects.

[0003] First, although this method uses the attention network of the reconstruction module to classify hyperspectral images, it has limitations when dealing with complex ground object scenes. Although hyperspectral images provide rich spectral information, when dealing with the geometric characteristics and spatial distribution of various ground object types (such as buildings, vegetation, water bodies, etc.), the classification model that relies solely on spectral information may not be able to effectively capture the spatial relationship between these ground objects. UAV remote sensing images have strong spatial complexity, especially in high-resolution images, the shape, texture and other information of the ground objects are also crucial for the classification task. The existing technology mainly classifies based on pixel-level spectral information, ignoring the spatial topological structure and shape similarity between the ground objects, which makes it difficult for the model to accurately classify when dealing with complex ground object scenes. For example, for vegetation and water body edges that also have similar spectral characteristics, it may be difficult to distinguish based on pixel-based spectral information alone, while the shape and spatial distribution characteristics of the ground objects can provide important supplementary information for the classification task. Therefore, the model of the existing technology may show a decrease in classification accuracy in complex scenes.

[0004] Secondly, although this method uses lightweight processing to reduce network complexity, the processing speed and computing efficiency may still be limited when the data scale is large. UAV remote sensing images usually cover large areas of scenes, and the amount of image data is huge, especially hyperspectral images, each pixel of which contains spectral information of hundreds of bands, and the data dimension is extremely high. Although lightweight networks can partially alleviate the pressure on computing resources, in actual applications, when the amount of data obtained by drones increases significantly, the training and inference speed of the network may still be difficult to meet the requirements of real-time processing.

[0005] Thirdly, the classification methods used in the prior art rely heavily on labeled data and fail to effectively solve the problem of data labeling. In the classification task of UAV remote sensing images, especially hyperspectral image classification, it is often very difficult and expensive to obtain labeled data. Each pixel in the hyperspectral data represents different spectral reflectance information, and accurately labeling these pixels requires professional domain knowledge and a large amount of human resources. However, the invention relies on hyperspectral image training data sets for supervised learning, requiring a large amount of labeled data to train the model. This dependence limits its promotion and use in application scenarios where labeled data is insufficient, especially when there is no labeled data, the classification ability of the model will be greatly reduced. In current practical applications, labeled data is usually insufficient, which greatly limits the universality and practicality of the model.

[0006] In addition, existing models fail to fully consider the dynamic characteristics of objects changing over time. UAV remote sensing images usually cover multi-temporal images of the same area. These images not only reflect the changes of objects in space, but also include the evolution of objects over time. For example, crop growth in agriculture and changes in urban expansion are all dynamic features.

[0007] In summary, although the existing technology has improved the classification performance of hyperspectral remote sensing images through attention networks and lightweight processing technology, it still has shortcomings in processing complex scenes, dealing with large-scale data, solving the scarcity of labeled data, capturing dynamic changes of ground objects, and using more advanced technical means. Summary of the invention

[0008] The purpose of the present invention is to provide a method for constructing a UAV remote sensing image classification model, thereby solving some of the drawbacks and shortcomings pointed out in the background technology.

[0009] The present invention solves the above-mentioned technical problems by adopting the following technical solutions: A method for constructing a classification model of UAV remote sensing images, comprising: S1, a structured representation of remote sensing images based on hierarchical feature decomposition:

[0010] S1.1, UAV images are divided into image blocks of different levels through pyramid decomposition algorithm, each image block represents a single different spatial resolution;

[0011] S1.2. At each level, key features at different scales are extracted through feature clustering and edge detection methods;

[0012] S1.3, hierarchical combination of features at each level, forming a multi-scale structured feature vector through feature weighting mechanism;

[0013] S2. Dynamic feature capture integrating spatial topological relationships:

[0014] S2.1, using the hierarchical feature decomposition results in S1, each area block in the image is used as a node of the graph, and the edges between the nodes represent the topological relationships between different area blocks, including distance and shape similarity;

[0015] S2.2. After constructing the graph structure, use the graph convolutional network (GCN) to aggregate the features of each node and propagate spatial relationship information, so that the features of each node contain its own information and integrate the influence of surrounding objects;

[0016] S2.3. For multi-temporal image data, the dynamic graph convolutional network D-GCN is introduced to model the spatial relationship at different times, capturing the changes of objects over time and their impact on classification;

[0017] S3. Self-supervised clustering learning of ground features:

[0018] S3.1. Use a self-supervised learning framework to perform comparative learning on unlabeled data to learn the internal structure and similarity of various ground features in the image;

[0019] S3.2, after the self-supervised feature learning is completed, the feature distance-based clustering algorithm is used to cluster the features of the objects and classify similar objects into one category;

[0020] S3.3. For objects that are difficult to distinguish, a soft clustering method is introduced to make the objects belong to multiple categories at the same time.

[0021] Furthermore, the dynamic feature capture method of integrating spatial topological relationships includes:

[0022] The image is decomposed into hierarchical features; the decomposition process is based on pixel values, combined with local spatial continuity and shape characteristics of the object; the methods used include superpixel segmentation and region growing algorithm; after segmentation, each region block is regarded as a single node in the graph, and its feature vector reflects the color and texture characteristics of the region block, as well as shape and position information; the specific feature vector It is expressed as:

[0023]

[0024] in, is the feature vector of the node of the ith region block, which contains four main parts: f color is the color feature, which represents the spectral information of the area block; f texture It is a texture feature, which is used to describe the structural characteristics of the surface of the object; f shape represents the shape feature, which is used to capture the geometric outline of the area; and f location It is a location feature that reflects the specific geographical location of the area block in the image.

[0025] Furthermore, the dynamic feature capture method of integrating spatial topological relationships includes:

[0026] The Euclidean distance between geometric centers is used to calculate the spatial distance between regional blocks; the influence of neighboring nodes on the topological structure is reflected, and a nonlinear function is introduced to adjust the weight of the spatial distance; the definition of the nonlinear distance function is as follows:

[0027]

[0028] in, is the spatial distance weight between node i and node j, d ij Represents the Euclidean distance between the geometric centers of two nodes, and the calculation formula is:

[0029]

[0030] where x i ,y i is the position coordinate of node i, x j ,y j is the position coordinate of node j; constant k 1 is a distance adjustment coefficient that controls the distance scaling effect; function Ensure that the edge weights are larger between closer nodes, while the weights of distant nodes gradually decrease.

[0031] Furthermore, the dynamic feature capture method of integrating spatial topological relationships includes:

[0032] A shape similarity calculation method based on cosine similarity is used to capture the local geometric differences between regional blocks and ensure that regional blocks with similar shapes have higher connection weights in the graph structure. Shape similarity is calculated using the following formula:

[0033]

[0034] in, is the shape similarity weight between node i and node j, and Represent the shape feature vectors of node i and node j respectively; the symbol <·, ·> represents the inner product operation of the vector, which is used to measure the similarity of two vectors in direction; the modulus of the vector and Represents the size of the respective shape feature vector.

[0035] Furthermore, the dynamic feature capture method of integrating spatial topological relationships includes:

[0036] After calculating the spatial distance and shape similarity, the constructed graph structure will be fed into the graph neural network GCN as input; GCN recursively updates the features by propagating information between nodes; the features of each node are updated based on its own information, and the features are fused and propagated in combination with the information of neighboring nodes; the recursive update formula of the features of the graph neural network is:

[0037]

[0038] in, is the feature vector of node i in the lth layer, indicating the feature state of the node in the current layer; Represents the set of neighboring nodes of node i; the features of the neighboring nodes are integrated into the feature update of node i by weighted summation, and the weight is determined by shape similarity. and spatial distance The decision ensures that similar and adjacent nodes have a greater impact on the feature update of node i; the parameters α and β control the weight ratio of the self-node features and the neighbor node features, and σ(·) is a nonlinear activation function.

[0039] Furthermore, the self-supervised clustering learning method of the land feature includes:

[0040] By creating multiple versions of the object under different conditions, the feature invariance of the object under different viewing angles, scales and lighting conditions is learned; data enhancement includes multiple operations such as image rotation, scaling, lighting change, color perturbation, and cropping; the operation generates multiple versions of the same image as positive samples for self-supervised contrast learning and is used for feature learning; the mathematical expression of the data enhancement is:

[0041]

[0042] in, Represents the multi-view data enhancement of the input image x, Aug i (x) represents the i-th data including rotation, scaling, cropping, and illumination change enhancement operations, γ i It is the weight factor of each enhancement operation, which is used to balance the impact of different enhancement operations on feature extraction.

[0043] Furthermore, the self-supervised clustering learning method of the land feature includes:

[0044] By comparing positive samples of different transformed versions of the same image with negative samples from images of other objects, the model maximizes the similarity between positive samples and minimizes the similarity with negative samples; the model projects the extracted high-dimensional features into a low-dimensional embedding space; in the low-dimensional space, the model optimizes the expression of image features through feature contrast; the loss function of feature projection and contrast learning is defined as:

[0045]

[0046] in, is the contrastive learning loss function in the feature projection stage, which is used to optimize the distance relationship in the feature space; for each sample i, is the low-dimensional representation of sample i in the embedding space, It is the embedding vector of the same image block of the positive sample corresponding to the sample after different transformations; Represents the negative sample embedding vector, which comes from images of other objects;

[0047] function Used to calculate the similarity between two embedding vectors; the contrastive loss function is designed to maximize the similarity between positive samples and minimize the similarity between negative samples.

[0048] Furthermore, the self-supervised clustering learning method of the land feature includes:

[0049] After completing self-supervised learning, the model is fine-tuned with a small amount of labeled data to improve classification accuracy. The fine-tuning process guides the classifier to optimize the feature expression in the embedding space through a small amount of labeled data. During the fine-tuning process, the optimization objective function is:

[0050]

[0051] in, represents the optimization target in the fine-tuning phase, α and β are adjustment coefficients;

[0052] Item 1 Represents the model prediction value With the true label The square error between is used to optimize the classification accuracy of the model; the second is the regularization term, Used to limit overfitting in low-dimensional embedding spaces.

[0053] The method for constructing a classification model for unmanned aerial vehicle remote sensing images of the present invention has significant beneficial effects in improving classification accuracy, reducing dependence on labeled data, and enhancing the robustness and adaptability of the model:

[0054] 1. Improve classification accuracy and robustness: By introducing hierarchical feature decomposition technology, the model can capture the multi-scale features of images at different spatial resolutions. The pyramid decomposition method divides the image into different levels, allowing the model to focus on both global structure and local details.

[0055] 2. Integrate spatial topological relationships to enhance model understanding: Treat each area block in the image as a node of the graph, and use spatial distance and shape similarity to build edges between nodes, forming a graph structure that reflects the spatial relationship in the real world. Through the application of graph convolutional networks (GCN) and dynamic graph convolutional networks (D-GCN), the model can capture the spatial and temporal dynamic relationships between objects and improve the recognition ability of complex objects.

[0056] 3. Reduce reliance on labeled data and reduce costs: Using a self-supervised learning framework, comparative learning is performed on unlabeled data, and the model can automatically learn the internal structure and similarity of ground features. This method reduces the need for a large amount of manually labeled data and reduces the cost of data acquisition and processing.

[0057] 4. Improve the generalization and adaptability of the model: Through data enhancement technology, multiple versions of objects under different viewing angles, scales and lighting conditions are generated, and the model learns the robustness of the features of the objects. In the fine-tuning stage, a small amount of labeled data is used to optimize the model, further improving the adaptability and classification accuracy of the model in the new environment. BRIEF DESCRIPTION OF THE DRAWINGS

[0058] Figure 1 The present invention is a flowchart of the method for constructing the UAV remote sensing image classification model.

[0059] Figure 2 The figure is a flow chart of the dynamic feature capture method integrating spatial topological relationships of the present invention.

[0060] Figure 3 This is a flow chart of the self-supervised clustering learning method of land feature of the present invention. DETAILED DESCRIPTION

[0061] The specific implementation modes of the present invention will be described in detail below in conjunction with the accompanying drawings.

[0062] Combined with Figure 1 ,The method for constructing the classification model of UAV remote sensing images of the ,present invention is based on the structured representation of ,remote sensing images by hierarchical feature decomposition, ,aims to comprehensively capture the information in remote sensing images through different ,levels and spatial resolutions to improve the performance of the ,classification model.

[0063] S1.1, the pyramid decomposition algorithm is used to divide the drone image into image blocks of different levels. The pyramid decomposition algorithm downsamples and processes the image layer by layer to produce image representations of different resolutions at each level. Each image block represents the content of the image at that specific spatial resolution. The higher-resolution image blocks retain more detailed information, while the low-resolution image blocks summarize the overall structure of the image. Through this decomposition, the details and global features in the image can be effectively captured, ensuring that the model can process information at different scales.

[0064] S1.2, at each level, feature clustering and edge detection methods are used to extract key features. Feature clustering is an unsupervised learning method that can group pixels or area blocks with similar features, thereby helping the model understand the structure and content in the image. Edge detection technology is used to capture the geometric shape and boundary information in the image, especially the boundaries between objects are particularly important in remote sensing images. By combining these two technologies, image blocks at different scales can extract the most critical feature information, so that the representations of these levels not only contain texture and spectral features, but also reflect the shape and spatial structure characteristics of the image. This step improves the classification ability of the model by taking into account both detailed features and overall structure.

[0065] S1.3, hierarchically combine the features extracted from each level. This step is achieved through a feature weighting mechanism, which weights the features at each level according to their importance to generate a multi-scale structured feature vector. The core of the feature weighting mechanism is that image information of different resolutions contributes differently to the final classification. Generally, detail features are more suitable for identifying local objects, while low-resolution overall structural information helps to understand the global background. Therefore, through feature weighting, the model can adjust the importance of features at each level to ensure that in multi-scale information fusion, the overall classification result will not be affected by the excessive strength or weakness of features at a certain level. The multi-scale structured feature vector finally formed is a high-dimensional feature representation that integrates information from each level. It can not only reflect the local details of the image, but also take into account the global spatial relationship, providing a richer and more complete feature expression for subsequent classification models.

[0066] Step S2 mainly involves capturing dynamic features that integrate spatial topological relationships. This process is based on the hierarchical feature decomposition in S1, further constructing a graph structure and using graph convolutional networks (GCNs) and dynamic graph convolutional networks (D-GCNs) to extract complex spatial and temporal information.

[0067] First, in S2.1, using the results of hierarchical feature decomposition in S1, each regional block in the image is used as a node in the graph structure. In this process, the geometric information, spectral features, texture features, etc. of the regional blocks are used to construct the feature vector of the node. The edges between nodes are used to represent the topological relationship between different regional blocks. Specifically, the weight of the edge can be determined by calculating the geometric distance, shape similarity and other spatial properties between regional blocks. For example, regional blocks with closer distances have larger edge weights, and regional blocks with similar shapes or spectral features will also be given higher connection strengths. Through the construction of this graph structure, each node in the image not only contains its own features, but also reflects the spatial relationship with neighboring regional blocks, thereby providing structured input for subsequent classification tasks.

[0068] In S2.2, after the complete graph structure is constructed, the graph convolutional network (GCN) is used to aggregate the features of each node. GCN is a graph-based neural network that can effectively use the topological structure between nodes to propagate and fuse features. In this process, the features of each node are not only updated according to its own information, but also the features of the surrounding nodes are integrated into its own features through the information propagation mechanism of the adjacent nodes. This mechanism ensures that when the model extracts features, it can not only rely on the local information of each regional block, but also capture the contextual relationship in a larger spatial range through the information fusion of neighboring nodes. Specifically, the feature aggregation process of GCN gradually transfers the information of neighboring nodes to the central node through graph convolution operations, and the features of the node are updated after each convolution operation. Therefore, after multiple layers of convolution, the feature vector of each node not only contains its own attributes, but also integrates the ground object information and spatial relationships of its neighborhood. This propagation of spatial information greatly improves the model's ability to understand complex ground object classification.

[0069] In S2.3, a dynamic graph convolutional network (D-GCN) is further introduced to process multi-temporal image data. Multi-temporal images refer to remote sensing images of the same area collected at different time points. These data contain information about the changes of objects in the time dimension, such as seasonal changes, vegetation growth, and changes in buildings. D-GCN captures the dynamic characteristics of objects changing over time through graph convolution operations at multiple time points. Unlike static GCN, D-GCN aggregates the features of spatial relationships at each time point, and can also model node changes in the time dimension and capture the evolution of regional block features at different times. This means that D-GCN can not only understand the spatial characteristics of objects at a single time point, but also capture trends and change patterns in time series. This is especially important for objects that show significant changes in different periods (such as farmland, rivers, urban expansion, etc.). Under the action of D-GCN, the model can use spatial information and combine information in the time dimension to more comprehensively understand the dynamic characteristics of objects, and ultimately improve the accuracy and robustness of remote sensing image classification.

[0070] The S3 step mainly revolves around the self-supervised clustering learning of land feature features, with the focus on using the self-supervised learning framework to effectively learn and cluster land feature features in an environment without labeled data, thereby improving the performance of the classification model.

[0071] First, in S3.1, a self-supervised learning framework is used to perform contrastive learning on unlabeled data. The core of self-supervised learning is to allow the model to learn features without manual annotation by designing appropriate pre-tasks. Contrastive learning is a typical self-supervised learning strategy. By constructing positive and negative sample pairs, the model can learn the internal structure and similarity of various ground features in the image. In drone remote sensing images, different ground objects (such as vegetation, buildings, water bodies, etc.) have different shapes, spectra, and texture features. Through self-supervised learning, the model can extract these important feature representations from the original images without relying on manual annotation. In the specific implementation process, the positive samples are usually composed of versions of the same image under different transformations (such as rotation, scaling, etc.), while the negative samples come from other unrelated ground images. By maximizing the similarity between positive samples and minimizing the similarity between negative samples, the model can gradually capture the intrinsic patterns and similar features of the ground objects.

[0072] Then in S3.2, after the self-supervised feature learning is completed, the model has generated a high-dimensional feature representation of each feature. At this time, in order to further classify, these features need to be clustered. A clustering algorithm based on feature distance is used to classify similar features of features into one category. Common clustering algorithms include K-means, hierarchical clustering, etc. The core idea of ​​these algorithms is to classify similar features into the same category based on the similarity of the features of the features in the feature space (usually measured by Euclidean distance, cosine similarity, etc.). Since self-supervised learning has extracted the key features of the features, clustering can effectively distinguish different categories of features. For example, the same vegetation has similar spectral features, buildings have similar shape features, and water bodies are significantly different from other features in spectrum and morphology. Through clustering, the model can identify the similarities of these features and classify them into the same category.

[0073] In actual operation, for some difficult-to-distinguish objects, the traditional hard clustering method has limitations, that is, one object can only belong to one category. However, in remote sensing images, many objects have complex features and have the characteristics of multiple categories at the same time. For example, a wetland has both the characteristics of water and vegetation. At this time, the traditional hard clustering method cannot accurately classify such objects. Therefore, in S3.3, a soft clustering method is introduced to allow objects to belong to multiple categories at the same time. Soft clustering assigns a probability distribution to each object so that it has the weight of belonging to multiple categories, rather than simply forcing it to be classified into a single category. For example, for the aforementioned wetland, the soft clustering model will assign it 50% water features and 50% vegetation features. This method is more flexible and can better capture the complex features of objects in remote sensing images and improve the classification accuracy and adaptability of the model. Through soft clustering, the model can not only handle clear category boundaries, but also effectively deal with objects with fuzzy boundaries or multiple attributes, thereby improving the robustness and generalization ability of the entire classification system.

[0074] Embodiment 1:

[0075] See also Figure 2 In this embodiment, the dynamic feature capture method integrating spatial topological relationships is a key step. In an agricultural / land survey / mapping project, a drone is used to shoot remote sensing images. The goal is to distinguish different crops in the farmland, such as corn, rice, and wheat, through image classification technology. The drone regularly flies to shoot a farmland of about 10 square kilometers and generates high-resolution multispectral images. The image data will be processed by the hierarchical feature decomposition method proposed in this paper, and finally an effective classification model will be constructed.

[0076] First, the image processing starts with hierarchical feature decomposition. The decomposition process is based on the spectral value of each pixel, combined with the spatial continuity and shape characteristics of the objects in the image. The image is segmented using superpixel segmentation and region growing algorithms. The role of superpixel segmentation is to aggregate similar pixels in the image to form multiple superpixel area blocks. The region growing algorithm starts from a seed point and gradually merges the pixels adjacent to it until certain conditions (such as color similarity or edge strength) are met. If the resolution of the image taken by the drone is set to 0.1 meters, the entire farmland image contains about 1 million pixels. After superpixel segmentation and region growing algorithms, these pixels are divided into 5,000 different area blocks, each of which contains about 200 pixels.

[0077] After segmentation, each region block is regarded as a node in the graph structure. Each node is described by its internal features through a feature vector. It is used to represent the multi-dimensional features of each node (i.e., region block). The dimension of the feature vector is set to 4, including the color feature f color , texture feature f texture , shape feature f shape , and the position feature f location .

[0078] For the color feature f color , assuming that the drone uses a multispectral camera with 4 bands (red, green, blue and near infrared), then the color feature vector f of each area block is color It is composed of the average spectral values ​​of these four bands. For example, the color characteristics of a certain area block can be expressed as f color =[150,120,110,180], these values ​​reflect the reflection intensity of the area under the red, green, blue and near-infrared spectra.

[0079] Texture feature f texture Describes the structure of the surface of the object, usually extracted by texture analysis methods, such as using local binary patterns (LBP) or gray-level co-occurrence matrix (GLCM). Assuming that the GLCM method is used to describe the texture characteristics of the region block, the contrast and energy are calculated to be 50 and 0.8 respectively, then the texture feature vector of the region block is f texture =[50,0.8].

[0080] Shape feature shape It is mainly used to describe the geometric outline of the region block, such as aspect ratio, perimeter and area. For example, if the perimeter of a region block is 40 pixels and the area is 200 pixels, then its aspect ratio is approximately 2:1, and the shape feature vector is f shape= [2.0, 40, 200]. It can be used to capture the shape of the ground and help the model distinguish different types of crop areas.

[0081] Position feature f location It reflects the location of each block based on geographic coordinates. Assuming that the drone image data is calibrated by GPS, the center position of each block can be represented by longitude and latitude coordinates. For a block, its geographic location is 104.0123 degrees longitude and 31.5678 degrees latitude, then the position feature vector is f location =[104.0123,31.5678].

[0082] Therefore, the entire eigenvector Expressed as For example: In the 5,000 regional blocks, each node will have a similar feature vector representation, and all of these feature vectors will be used to construct the subsequent graph structure. In this graph structure, the feature vector of each node not only reflects its internal information, but also forms a complex spatial topological relationship through the feature propagation of neighboring nodes.

[0083] We further use the dynamic feature capture method that integrates spatial topological relationships to handle the classification task of UAV remote sensing images. We build a graph structure by calculating the spatial distance between area blocks to help the model better understand the spatial relationship of the objects, and use nonlinear functions to adjust the weight of the distance to improve the accuracy of classification.

[0084] First, among the 5,000 segmented blocks, each node represents a block and carries its feature vector, including color, texture, shape, and location information. Next, the spatial distance between every two nodes (blocks) is calculated to construct a graph structure based on geometric relationships.

[0085] The calculation of spatial distance depends on the Euclidean distance formula, which is:

[0086]

[0087] Among them, d ij represents the spatial distance between node i and node j, x i ,y i is the position coordinate of node i, and x j ,y j is the position coordinate of node j. Assuming the center coordinate of region block i is 104.0123, 31.5678 and the center coordinate of region block j is 104.0135, 31.5685, the Euclidean distance d between them is ij for:

[0088]

[0089] Next, a nonlinear function is used to adjust the weight of the spatial distance so that closer blocks contribute more to the graph structure, while farther blocks have less impact. The definition of the nonlinear distance function is as follows:

[0090]

[0091] in, is the spatial distance weight between nodes i and j, d ij is the Euclidean distance between the two, k 1 It is a constant that controls the distance scaling effect, usually between 0.01 and 10, and the specific value is adjusted according to the data scale. 1 =1, substitute the Euclidean distance d above ij ≈0.0014:

[0092]

[0093] It can be seen that due to the close distance between the region blocks, the calculated weight is close to 1, which indicates that the two region blocks will have a strong connection in the graph structure.

[0094] Similarly, set another farther node k with coordinates of 104.0500, 31.6000, and calculate the Euclidean distance between node i and node k:

[0095]

[0096] Now substitute the nonlinear distance function and get:

[0097]

[0098] Since node k is farther away from node i, its weight is significantly smaller than the weight between node i and node j. This weight setting ensures that in the graph structure, closer blocks have a greater impact on the classification model, while the influence of farther blocks gradually weakens.

[0099] In this example, the Euclidean distance calculation of the geometric center and the nonlinear distance weight adjustment are used to construct a graph structure that reflects the actual spatial relationship. The advantage of this graph structure is that it can effectively capture the spatial correlation between different landforms, and through weight adjustment, the model pays more attention to those areas that are closer and have stronger correlation.

[0100] In addition to capturing the geometric relationship between blocks through spatial distance, it is also necessary to capture the local geometric differences between blocks through shape similarity, so as to further enhance the expressive power of the graph structure and ensure that the model not only relies on the spatial relationship between blocks, but also optimizes the classification performance through shape features. In the aforementioned drone remote sensing image data, it is observed that the shapes of crop areas in some farmlands are very similar. For example, the boundary contours of the blocks of two corn fields are very close. In order to capture these blocks of similar shapes in the graph structure, a shape similarity calculation method based on cosine similarity is adopted.

[0101] Specifically, there are two region blocks i and j, and their shape feature vectors are and The shape feature vector can contain a variety of geometric attributes, such as the aspect ratio, perimeter, area, etc. of the region. Suppose the shape feature vector of region block i is It means that the aspect ratio of the region block is 2:1, the perimeter is 40 pixels, and the area is 200 pixels; the shape feature vector of region block j is Indicates that its aspect ratio is 2.1:1, the perimeter is 42 pixels, and the area is 210 pixels. In order to calculate the shape similarity between these two area blocks, the cosine similarity formula is used, as follows:

[0102]

[0103] in, represents the shape similarity weight between node i and node j, and are the shape feature vectors of node i and node j respectively, Represents the inner product of vectors, which is used to measure the similarity of two vectors in direction. and Represents the modulus of the shape eigenvector, that is, the length of the vector.

[0104] First, calculate the inner product of the two vectors:

[0105]

[0106] Next, calculate the magnitude of each vector:

[0107]

[0108]

[0109] Finally, substitute the cosine similarity formula to calculate the shape similarity:

[0110]

[0111] It can be seen that the shape similarity of the two area blocks is close to 1, indicating that they are very similar in shape. In the graph structure, nodes i and j with higher shape similarity will obtain higher edge weights, ensuring that the model can effectively distinguish crop areas with similar shapes while capturing the geometric morphology of the area blocks.

[0112] Then set the shape feature vector of another region block k to Calculate the shape similarity of region blocks i and k. First, calculate their inner product:

[0113]

[0114] Then, calculate

[0115]

[0116] Finally, the shape similarity of nodes i and k is calculated:

[0117]

[0118] Although the shape features are slightly different, the similarity is still high. This means that even if the shapes of the objects are slightly different, their edge weights are still high and can be strongly connected in the classification model. This weight distribution mechanism based on shape similarity ensures that the model can correctly capture local geometric changes in the graph structure through shape features and enhance the recognition ability of similar object areas through topological connections.

[0119] Furthermore, the constructed graph structure is processed by a graph neural network (GCN). By calculating the spatial distance and shape similarity between the regional blocks, the topological structure between the regional blocks has been obtained. Now this information needs to be input into the GCN to recursively update the feature vector of each node and perform feature propagation.

[0120] Graph Neural Network (GCN) propagates information and fuses features between nodes. The features of each node are updated not only based on its own attributes, but also based on the features of its neighboring nodes. In the classification task of UAV remote sensing images, GCN uses multiple layers of recursive calculations to make the features of each block contain the global information of its neighborhood. It is assumed that a graph structure has been built for 5,000 blocks, and the feature vector of each node is Indicates the characteristic state of the region block at layer l.

[0121] The feature recursive update formula of GCN is as follows:

[0122]

[0123] in, represents the feature vector of node i in the l+1th layer, α and β are weight parameters that control the influence of self-features and neighbor features, σ(·) is a nonlinear activation function (such as ReLU or Sigmoid), represents the set of neighbor nodes of node i, is the shape similarity between node i and node j, is the spatial distance weight between node i and node j.

[0124] The feature vector of region block i at the initial moment is set to represents color, texture and shape features. The feature vector of the adjacent region block j is In the previous step, the shape similarity between region blocks i and j has been calculated The spatial distance weight

[0125] Setting parameters α = 0.7 and β = 0.3 means that the node's own features contribute more to the update, while the features of neighboring nodes have a relatively small impact on the update. Next, substitute these values ​​into the recursive update formula of GCN:

[0126]

[0127] First, calculate the weighted features of neighbor nodes:

[0128]

[0129] Then, the weighted sum of the neighbor features and the node’s own features is calculated:

[0130] 0.7 · [150, 0.5, 0.7] = [105, 0.35, 0.49]

[0131] 0.3 · [140, 0.6, 0.8] = [42, 0.18, 0.24]

[0132] The sum is:

[0133] [105,0.35,0.49]+[42,0.18,0.24]=[147,0.53,0.73]

[0134] Finally, after a nonlinear activation function (set to ReLU activation function), the updated feature vector of node i at layer l+1 is obtained as:

[0135]

[0136] Through recursive updating, the feature vector of node i not only retains its own information, but also obtains a more comprehensive representation through the feature propagation of neighboring node j. This feature update mechanism ensures that in the graph structure, each node can not only express local features, but also integrate the influence information of surrounding objects, thereby improving the classification accuracy of the model for different object areas.

[0137] After recursive updates of multiple layers of GCN, the feature vector of each node will gradually integrate the information in the neighborhood, and finally form a comprehensive feature that integrates spatial distance, shape similarity and neighborhood topological relationship.

[0138] Embodiment 2:

[0139] Combination Figure 3 In this embodiment, the drone took remote sensing images of a corn field, a rice field, and a wheat field. The goal was to classify these different crops and improve the model's ability to recognize the features of the objects through self-supervised learning methods. In order to adapt the model to complex environmental conditions (such as different lighting, changes in viewing angles, shooting conditions of different scales, etc.), a self-supervised clustering learning method of object features was adopted. Multiple versions of the objects under different conditions were created through data enhancement technology to learn the invariance of the features of these objects under multiple environments.

[0140] First, in the data enhancement process, the original image x is processed using multiple operations to generate multiple versions of the same object. These versions have been changed in terms of lighting, rotation, scaling, and cropping, so as to ensure that the model can learn the characteristics of the object under different conditions. The image x is set to undergo four data enhancement operations: rotation, scaling, lighting change, and cropping. These operations generate multiple versions of image data, which are used as positive sample comparison learning for self-supervised learning.

[0141] The mathematical expression of data enhancement is:

[0142]

[0143] in, Represents the multi-view data enhancement of the input image x, Aug i (x) represents the i-th data augmentation operation (such as rotation, scaling, cropping, illumination change, etc.), γ i is the weight factor of each enhancement operation, which controls the influence of different enhancement operations on feature extraction. In this example, n=4 (i.e., four enhancement operations), including rotation (90 degrees, 180 degrees), scaling (1.2 times, 0.8 times), illumination change (increase brightness, reduce brightness), and cropping (random cropping 10% or 20%). The weight factors of these operations are i It can be set according to the impact of each operation on model learning, such as γrotate =0.3,γ scale =0.2,γ brightness =0.25,γ crop =0.25.

[0144] For example, assume that the image x of a certain area taken by a drone is a cornfield. On the original image, we rotate it 90 degrees to generate version 1, scale it 1.2 times to generate version 2, increase the brightness by 10% to generate version 3, and randomly crop it by 15% to generate version 4. After data enhancement, the four versions of images obtained are Aug 1 (x), Aug 2 (x), Aug 3 (x), Aug 4 (x). By using these different versions of images as positive samples for comparative learning, the model can learn the multi-angle and multi-environmental features of the cornfield without manual annotation.

[0145] Let's take the change of illumination as an example. The brightness range of the original image is 0,255. After increasing the brightness, the pixel value of the image is increased by 10%, that is, the brightness value of a certain pixel in the original image is set to 150. After enhancement, the brightness value of this pixel becomes 150*1.1=165. Similarly, for the scaling operation, if the width of the original image is 1000 pixels and the height is 800 pixels, after the scaling operation is proportionally enlarged by 1.2 times, the width of the image will become 1000*1.2=1200 pixels, and the height will become 800*1.2=960 pixels. Through these enhancement operations, the model will learn the characteristic performance of cornfields under different lighting, scales, and viewing angles, thereby improving the robustness of the model.

[0146] These enhanced versions of the images are input as positive samples into the self-supervised contrastive learning framework. The model learns the similarities between these positive samples by comparing them, ensuring that different versions of the same object are close to each other in the feature space, while using negative samples (images from rice fields or wheat fields) to increase the feature distance between different objects. The similarity of positive samples can be optimized by contrastive loss functions, such as the NT-Xent loss function, to maximize the feature similarity between similar samples and minimize the similarity between negative samples.

[0147] Through this self-supervised learning method, the model can efficiently learn the characteristics of different land objects (such as corn, rice fields, and wheat) without human labeled data, and can maintain a high classification accuracy when facing different environmental changes.

[0148] The self-supervised clustering learning method of ground feature features was further used to classify the remote sensing images of corn fields, rice fields, and wheat fields taken by drones. In this step, through the self-supervised learning method, the model can not only extract the key features of ground features from image data under different environmental conditions, but also maximize the similarity between positive samples and minimize the similarity with negative samples through the comparative learning framework.

[0149] It is assumed that the images taken by the drone have been augmented with multi-view data, including rotation, scaling, illumination change, and cropping operations. Through these enhancement operations, multiple versions of cornfield images in different environments are generated, for example, a version generated by rotating 90 degrees, a version generated by scaling 1.2 times, a version generated by increasing the brightness by 10%, and a version generated by randomly cropping 15%. These different versions of images are used as positive samples in the self-supervised contrastive learning framework of the model.

[0150] In the process of self-supervised contrastive learning, the model aims to maximize the similarity between positive samples (i.e., different transformed versions of the same image) while minimizing the similarity between negative samples (i.e., images from different objects, such as rice fields or wheat fields). Through this mechanism, the model gradually learns how to distinguish the performance characteristics of similar objects under different conditions. To achieve this, the high-dimensional features of the image are projected into a low-dimensional embedding space through dimensionality reduction. In this low-dimensional space, the model can perform feature comparison more effectively, thereby optimizing the expression of image features.

[0151] The loss function for feature projection and contrastive learning is defined as:

[0152]

[0153] in, is the contrastive learning loss function in the feature projection stage, is the low-dimensional representation of sample i in the embedding space, is the positive sample corresponding to sample i, representing different versions of the same object image, It represents the negative sample embedding vector from other objects.

[0154] Assume that cornfield image i and its positive sample i have been extracted + , and extracted the rice field images - The negative sample embedding vector of . Using the function To calculate the similarity between two embedded vectors. This function is used to measure the similarity between two vectors in a low-dimensional space. Usually, cosine similarity or the inverse of Euclidean distance can be used as a similarity measure.

[0155] Assume that after feature extraction, the low-dimensional embedding vector of cornfield image i is The positive samples i of different transformed versions of the same cornfield + The embedding vector of Calculate the similarity between them You can use cosine similarity, the formula is:

[0156]

[0157] where <·,·> is the inner product of vectors, and is the respective modulus. Compute the inner product:

[0158]

[0159] Next, calculate the magnitude of the vector:

[0160]

[0161]

[0162] Finally, the cosine similarity is:

[0163]

[0164] This shows that the similarity between positive samples is very high, close to 1, reflecting that the same type of objects have similar characteristic performances in different environments.

[0165] On the other hand, the negative sample embedding vector from the rice field is set to Calculate the similarity between cornfield and rice field images:

[0166]

[0167] Since corn fields and rice fields are different types of objects, their similarity is significantly lower than the similarity between positive samples.

[0168] Through the above calculations, we can see the effectiveness of the contrast loss function. By maximizing the similarity of positive samples (close to 1) and minimizing the similarity of negative samples (significantly less than 1), the model can gradually learn the characteristic expressions of different objects, thereby improving the classification performance. In the low-dimensional embedding space, the features of multiple transformed versions of corn fields are clustered together, while negative samples of rice fields or wheat fields are kept far away in the feature space.

[0169] Through this self-supervised clustering learning, the model can learn the characteristic differences between corn, rice fields and wheat fields without labeling, and has the robustness to cope with different environmental changes (such as lighting, perspective, scale changes, etc.). In practical applications, this method not only reduces the dependence on labeled data, but also significantly improves the accuracy and efficiency of remote sensing image classification.

[0170] Through self-supervised learning methods, the model has been able to effectively extract the features of the ground objects and distinguish different crop areas such as corn, rice fields and wheat fields. In order to further improve the classification accuracy of the model, fine-tuning with a small amount of labeled data is required.

[0171] It is assumed that there is a small amount of annotated data from actual scenes, for example, 50 samples are annotated from corn fields, rice fields, and wheat fields respectively. These annotated data will be used in the fine-tuning stage. The model has formed the characteristic distribution of objects in the low-dimensional embedding space, but these small amounts of annotated data are needed to adjust the decision boundary of the classifier so that the model has higher classification accuracy in the new environment.

[0172] During fine-tuning, the optimization goal is to minimize the gap between the model's predicted value and the true label while preventing overfitting. The optimization objective function of this process is defined as follows:

[0173]

[0174] in, is the optimization target in the fine-tuning phase, α and β are adjustment coefficients that control the sensitivity of the model to the classification error and the regularization term, respectively. Represents the model prediction value With the true label The square error between , which is used to improve the classification accuracy of the model; the second is the regularization term, It is used to limit the overfitting phenomenon in the low-dimensional embedding space and maintain the stability of the feature space.

[0175] Set the true label in the labeled sample represents cornfield, and the class probability predicted by the model This means that the model predicts that the probability of it being a corn field is 80%, the probability of it being a rice field is 15%, and the probability of it being a wheat field is 5%. The squared error part can be calculated as:

[0176]

[0177] This indicates that the model has a classification error of 0.065 on this sample.

[0178] Next, the regularization term is used to control the complexity of the model in the low-dimensional embedding space. The regularization term usually uses L2 regularization, which prevents feature overfitting by calculating the sum of the squares of the feature vectors. For example, if a feature vector in the embedding space is set to Its regularization term It can be expressed as:

[0179]

[0180] Set the sum of all regularization terms and set the regularization adjustment coefficient β = 0.01, then the value of the regularization part is:

[0181]

[0182] Therefore, the loss function calculation result of the entire fine-tuning process is:

[0183]

[0184] Among them, setting α = 0.9 is used to increase the impact of classification error on optimization, and β = 0.01 is used to limit the impact of regularization terms.

[0185] Through this fine-tuning method, the model can significantly reduce classification errors under the guidance of a small amount of labeled data. At the same time, through the role of the regularization term, the model will not overfit due to fine-tuning of a small amount of data. In this example, the fine-tuned model can more accurately classify corn fields, rice fields, and wheat fields, further improving the generalization ability and classification performance of the model.

[0186] The above shows and describes the basic principles, main features and advantages of the present invention. It should be understood by those skilled in the art that the present invention is not limited to the above embodiments, and the above embodiments and descriptions are only for explaining the principles of the present invention. Without departing from the spirit and scope of the present invention, the present invention may have various changes and improvements, which fall within the scope of the present invention to be protected. The scope of protection of the present invention is defined by the attached claims and their equivalents.

Claims

1. A method for constructing a classification model for UAV remote sensing images, characterized in that: S1. Structured representation of remote sensing images based on hierarchical feature decomposition: S1.1, UAV images are divided into image blocks of different levels through pyramid decomposition algorithm, each image block represents a single different spatial resolution; S1.

2. At each level, key features at different scales are extracted through feature clustering and edge detection methods; S1.3, hierarchical combination of features at each level, forming a multi-scale structured feature vector through feature weighting mechanism; S2. Dynamic feature capture integrating spatial topological relationships: S2.1, using the hierarchical feature decomposition results in S1, each area block in the image is used as a node of the graph, and the edges between the nodes represent the topological relationships between different area blocks, including distance and shape similarity; S2.

2. After constructing the graph structure, use the graph convolutional network (GCN) to aggregate the features of each node and propagate spatial relationship information, so that the features of each node contain its own information and integrate the influence of surrounding objects; S2.

3. For multi-temporal image data, the dynamic graph convolutional network D-GCN is introduced to model the spatial relationship at different times, capturing the changes of objects over time and their impact on classification; S3. Self-supervised clustering learning of ground features: S3.

1. Use a self-supervised learning framework to perform comparative learning on unlabeled data to learn the internal structure and similarity of various ground features in the image; S3.2, after the self-supervised feature learning is completed, the feature distance-based clustering algorithm is used to cluster the features of the objects and classify similar objects into one category; S3.

3. For objects that are difficult to distinguish, soft clustering methods are introduced to make the objects belong to multiple categories at the same time; The dynamic feature capture method of integrating spatial topological relationships includes: The Euclidean distance between geometric centers is used to calculate the spatial distance between regional blocks; the influence of neighboring nodes on the topological structure is reflected, and a nonlinear function is introduced to adjust the weight of the spatial distance; the definition of the nonlinear distance function is as follows: in, is the spatial distance weight between node i and node j, d ij Represents the Euclidean distance between the geometric centers of two nodes, and the calculation formula is: where x i ,y i is the position coordinate of node i, x j ,y j is the position coordinate of node j; the constant k1 is a distance adjustment coefficient that controls the distance scaling effect; the function Ensure that relatively large edge weights are formed between relatively close nodes, while the weights of distant nodes gradually decrease; A shape similarity calculation method based on cosine similarity is used to capture the local geometric differences between regional blocks and ensure that regional blocks with similar shapes have higher connection weights in the graph structure; After calculating the spatial distance and shape similarity, the constructed graph structure will be sent as input to the graph neural network GCN; GCN recursively updates the features by propagating information between nodes; the features of each node are updated based on its own information, and the information of neighboring nodes is combined for feature fusion and propagation.

2. The method for constructing a classification model for unmanned aerial vehicle remote sensing images according to claim 1, characterized in that The dynamic feature capture method of integrating spatial topological relationships includes: The image is decomposed into hierarchical features; the decomposition process is based on pixel values, combined with local spatial continuity and shape characteristics of the object; the methods used include superpixel segmentation and region growing algorithm; after segmentation, each region block is regarded as a single node in the graph, and its feature vector reflects the color and texture characteristics of the region block, as well as shape and position information; the specific feature vector It is expressed as: in, is the feature vector of the node of the ith region block, which consists of four parts: f color is the color feature, which represents the spectral information of the area block; f texture It is a texture feature, which is used to describe the structural characteristics of the surface of the object; f shape represents the shape feature, which is used to capture the geometric outline of the area; and f location It is a location feature that reflects the specific geographical location of the area block in the image.

3. The method for constructing a classification model for unmanned aerial vehicle remote sensing images according to claim 1, characterized in that The self-supervised clustering learning method of the ground feature comprises: By creating multiple versions of the object under different conditions, the feature invariance of the object under different viewing angles, scales and lighting conditions is learned; data enhancement includes multiple operations such as image rotation, scaling, lighting change, color perturbation, and cropping; the operation generates multiple versions of the same image as positive samples for self-supervised contrast learning and is used for feature learning; the mathematical expression of the data enhancement is: in, Represents the multi-view data enhancement of the input image x, Aug i (x) represents the i-th data including rotation, scaling, cropping, and illumination change enhancement operations, γ i It is the weight factor of each enhancement operation, which is used to balance the impact of different enhancement operations on feature extraction.

4. The method for constructing a classification model for UAV remote sensing images according to claim 3 is characterized in that The self-supervised clustering learning method of the ground feature comprises: By comparing positive samples of different transformed versions of the same image with negative samples from images of other objects, the model maximizes the similarity between positive samples and minimizes the similarity with negative samples; the model projects the extracted high-dimensional features into a low-dimensional embedding space; in the low-dimensional space, the model optimizes the expression of image features through feature comparison.

5. The method for constructing a classification model for unmanned aerial vehicle remote sensing images according to claim 4 is characterized in that The self-supervised clustering learning method of the ground feature includes: after completing the self-supervised learning, the model is fine-tuned by a small amount of labeled data to improve the classification accuracy; the fine-tuning process guides the classifier to optimize the feature expression in the embedding space by a small amount of labeled data, and the optimization objective function in the fine-tuning process is: in, represents the optimization target in the fine-tuning phase, α and β are adjustment coefficients; Item 1 Represents the model prediction value With the true label The square error between is used to optimize the classification accuracy of the model; the second is the regularization term, Used to limit overfitting in low-dimensional embedding spaces.

Citation Information

Patent Citations

  • Intelligent pixel-level classification method suitable for hyperspectral remote sensing image of unmanned aerial vehicle

    CN117893927A

  • High-resolution remote sensing image segmentation method based on inter-scale mapping

    CN104361589A

  • Remote sensing image building change detection method based on conditional adversarial network

    CN114708501A