Coarse-to-fine land utilization change intelligent detection method and system

By constructing an integrated deep learning model and combining multi-CNN collaboration, RCVA, and GLCM methods, the accuracy and efficiency problems of land use change detection in traditional technologies have been solved, enabling high-frequency, multi-scale land use change identification and improving the ability to perceive and manage dynamic changes in agricultural areas.

CN120931977APending Publication Date: 2025-11-11WUCHANG SHOUYI UNIV +4
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510835381.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-20
Publication Date
2025-11-11

AI Technical Summary

Technical Problem

Traditional land use change detection technologies struggle to meet both macro-level control and micro-level extraction needs, resulting in coarse granularity, large identification bias, and slow response speed in change identification. In particular, the monitoring accuracy decreases in hilly and mountainous areas, affecting the ability of intelligent agricultural management.

Method used

An integrated deep learning model for scene-level classification and pixel-level extraction is constructed. Spectral and texture change intensity information is extracted through multi-CNN collaboration, RCVA, and GLCM methods. Combined with a DBN model of multi-layer RBM and feedforward backpropagation network, high-frequency, multi-scale, and high-precision identification of land use change is achieved.

Benefits of technology

It improves the accuracy and efficiency of land use change detection, provides high-quality training samples, ensures the accuracy of the mathematical expression of change characteristics and the efficiency of parameter iteration, and supports the dynamic change perception and refined management of agricultural areas.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120931977A_ABST
    Figure CN120931977A_ABST
Patent Text Reader

Abstract

The invention discloses a coarse-to-fine land utilization change intelligent detection method and system. The method comprises the following steps: making a land utilization scene classification data set; according to the classification precision, carrying out adaptability distinguishing on the CNN model to complete screening and carrying out model fine tuning; determining a land utilization scene category through a multi-CNN collaborative scene category identification mechanism; spectral change intensity information is extracted based on an RCVA method; texture change intensity information is extracted based on a GLCM method; combining a scene classification result to complete the selection of a training sample; constructing a DBN model; defining joint probability distribution of an explicit layer and a hidden layer of the RBM based on an energy function; determining a neuron activation probability based on the structural characteristics of the RBM; fitting training data through a maximized log-likelihood function; and completing model training. And the processing efficiency, the monitoring precision and the comprehensive application value of change information are effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of remote sensing image processing and deep learning applications, and more specifically, relates to an intelligent detection method and system for land use change from coarse to fine. Background Technology

[0002] Land use change is a direct manifestation of the interplay of multiple factors, including agriculture, ecology, and urban-rural development. Its spatiotemporal dynamics have a significant impact on land resource management, agricultural production scheduling, and ecological environmental protection. In particular, under the dual pressures of intensifying global climate change and food security challenges, my country's agricultural development is facing prominent problems such as increasingly tight resource constraints, shrinking ecological red lines, and labor shortages.

[0003] The intelligent transformation of modern agricultural production has become a crucial support for promoting high-quality agricultural development and building a strong agricultural nation. However, this transformation still faces many key challenges: First, fragmented regional environmental monitoring makes it difficult to grasp land use dynamics in a timely manner; second, delayed identification of crop growth status affects the adjustment of arable land planting structure and agricultural management decisions; and third, the extensive use of agricultural resources restricts the improvement of agricultural production efficiency and the conservation of resources. Furthermore, agricultural remote sensing monitoring also faces technical bottlenecks such as large differences in data modalities, low registration accuracy, difficulty in integrating multi-source data, and difficulty in monitoring dynamic changes in arable land.

[0004] Currently, most traditional land use change detection technologies employ uniform granularity and static rules, making it difficult to simultaneously meet the dual needs of macro-level control and micro-level extraction. This results in problems such as coarse granularity, large identification bias, and slow response speed in change identification. Particularly in hilly and mountainous areas and in southern regions where small-scale farming is widespread, fragmented land use units and blurred boundaries significantly reduce the accuracy of change monitoring, severely restricting the capabilities of intelligent agricultural monitoring and management.

[0005] To address these challenges, intelligent land use change detection methods based on high-resolution remote sensing imagery, employing a coarse-to-fine approach, have become a research hotspot. By constructing a land use scene-level classification remote sensing dataset with high spatial resolution and multi-target information, and simultaneously using a pixel-level land use change extraction dataset, combined with artificial intelligence technologies such as deep learning scene classification and pixel-level change extraction, high-frequency, multi-scale, and high-precision identification of land use changes can be achieved. The coarse-to-fine change detection strategy first quickly identifies scene categories in the study area, and then achieves fine-grained change information extraction within a localized area, effectively improving processing efficiency, monitoring accuracy, and the comprehensive application value of change information.

[0006] In summary, a coarse-to-fine intelligent detection method for land use change can not only break through the current technical bottleneck of high-precision agricultural monitoring in large areas and multiple scenarios, but also significantly improve the ability to perceive dynamic changes in agricultural areas and the level of refined management. It provides solid technical support for farmland protection, the "storing grain in the land" strategy, and intelligent agricultural decision-making, and has important scientific research value and broad application prospects. Summary of the Invention

[0007] This invention aims to provide an intelligent detection method and system for land use change from coarse to fine. By constructing an integrated deep learning model that combines scene-level classification and pixel-level extraction, and combining a rich and large-scale scene classification dataset with a dataset of farmland changes under multiple scene and type change conditions, the invention achieves the identification and precise segmentation of farmland changes, providing technical support for the monitoring and management of farmland resources in modern agriculture, and improving the level of perception and refined management of dynamic changes in agricultural areas.

[0008] To address the aforementioned deficiencies or improvement needs of existing technologies, as a first aspect of this invention, the present invention provides an intelligent land use change detection method that progresses from coarse to fine, comprising:

[0009] S1. Creation of a land use scenario classification dataset;

[0010] S2. Train no fewer than 10 CNN networks and differentiate the models based on their classification accuracy to complete the CNN network selection; select the pre-trained CNN networks after selection and fine-tune the models based on the dataset obtained in S1; determine the land use scene category through a multi-CNN collaborative scene category identification mechanism;

[0011] S3. Extract spectral change intensity information based on the RCVA method; extract texture change intensity information based on the GLCM method; and then, in combination with the scene classification results, use multi-scale thresholding to divide the two intensity maps and take the union to complete the selection of training samples.

[0012] S4. A DBN model is constructed by combining a multi-layer Restricted Boltzmann Machine (RBM) with a feedforward backpropagation network. The joint probability distribution of the explicit and hidden layers of the RBM is defined based on the energy function. The activation probability of neurons is determined based on the structural characteristics of the RBM. The training data is fitted by maximizing the log-likelihood function. The model parameters are iteratively optimized by solving the contrastive divergence method to initialize the RBM parameters. Based on the error backpropagation method, the network weights are updated by calculating the learning error of each RBM layer, and the entire network is fine-tuned to complete the model training.

[0013] Furthermore, the scene category identification mechanism of multi-CNN collaboration in S2 is specifically as follows:

[0014] Let the scene assignment set be S = {S1, S2, ..., S...} k}, where S i Let k be the set of scenes assigned to the i-th CNN, and k be the number of CNNs; the set of prediction results is p = {p1, p2, ..., p...} k}, where p i Let be the predicted category of the i-th CNN for scene x;

[0015] The historical precision set is acc = {acc1, acc2, ..., acc k}, where acc i Let be the prediction accuracy of the i-th CNN in the corresponding scene;

[0016] Final Category C final The rules for determining it are as follows:

[0017] When there exists a class c such that all p i When =c, then C final=c ;

[0018] When there exists i such that p i ∈S i All p i When inconsistencies exist, then C final To satisfy p i ∈S i The CNN prediction category with the highest fitness is selected if multiple i satisfy the criteria, according to the preset priority.

[0019] When all When, then C final The CNN output class with the highest prediction accuracy, i.e., C final =p j , where j is the value that makes acc j The largest index.

[0020] Furthermore, the specific process of extracting spectral change intensity information based on the RCVA method in S3 is as follows:

[0021] The RCVA method eliminates the influence of registration error by considering the neighborhood information of the pixel and selecting the pixel pair with the smallest spectral difference for detection.

[0022] The method is based on the following assumption: Given a pixel x1(j,k) in image 1, find the pixel x2 with the smallest spectral difference from x1 within the range of image 2 (j±ω,k±ω). Then, consider x2 as a corresponding image point of x1, and calculate the difference x. d-a Similarly, based on image 2, find the corresponding image points in image 1 and calculate the difference x. d-b Finally, with x d-a and xd-b The smaller of the two is taken as the intensity of change, and the specific calculation process is as follows:

[0023]

[0024] In the formula, (j,k) represents the coordinate index of the target pixel in the image; (p,q) represents the coordinate index of the pixel traversed within the window range [j-ω,k+ω] centered at (j,k); ω represents the window radius; and n represents the number of bands involved in the calculation. These correspond to the pixel values ​​at coordinates (j,k) in the i-th band of the two types of data (or images); similarly... It is the pixel value at the neighboring point (p,q), which belongs to the "original data value";

[0025] By iterating through all pixels in both images, the intensity value M of the spectral change for all pixels can be obtained. S This leads to the spectral information intensity map that takes into account neighborhood information.

[0026] Furthermore, the specific process of extracting texture change intensity information based on the GLCM method in S3 is as follows:

[0027] GLCM extracts texture through the conditional probability density between image gray levels, using variance as a scalar to reflect the differences between different ground features. The calculation process is as follows:

[0028]

[0029] Where p(l,g,d,θ,i) is the frequency of occurrence of a pixel with gray value l in the i-th band of the image and a pixel with gray value g at a distance d from it; θ is the direction for generating the gray-level co-occurrence matrix; N is the gray level; The mean of the frequencies;

[0030] After obtaining the texture variance feature values ​​of image 1 and image 2, the texture variation intensity M can be calculated. T The calculation process is as follows:

[0031]

[0032] In the formula, These are variance-based feature values ​​calculated from the gray-level co-occurrence matrix. is the difference between the two variance features mentioned above; n is the total number of "feature sequences", that is, the total number of bands.

[0033] Furthermore, in S3, the specific process of selecting training samples by combining the scene classification results, using multi-scale thresholding to divide and combine the two intensity maps, and taking the union is as follows:

[0034] Based on the CNNs trained by S2, class detection is performed on remote sensing scene pairs obtained through multi-scale segmentation and overlapping cropping, and the changed scene is obtained according to the class. On the change intensity maps of RCVA and GLCM, the changed patches are divided into training samples by multi-scale thresholding, and the union of the samples divided by the two is taken as the final training sample.

[0035] Furthermore, the specific process of defining the joint probability distribution of the explicit and hidden layers of the RBM based on the energy function in S4 is as follows:

[0036] An RBM consists of two layers: a visible layer (v) and a hidden layer (h), forming an undirected graphical model. Each layer is defined as a vector, with the dimension of the vector being the number of neurons in that layer. Neurons in different layers are connected by a weight matrix W.

[0037] For each RBM, its energy as a system is:

[0038]

[0039] Where n and m are the number of neurons in the visible and hidden layers, respectively; v i and h j These represent the states of neurons in the visible and hidden layers, respectively; θ = {W ij ,a i ,b j} represents the model parameters, where W ij Let a be the weight between neurons i and j. i and b j These represent the biases of neurons in the visible and hidden layers, respectively.

[0040] Based on the energy function, the joint probability distribution of (v,h) can be obtained:

[0041]

[0042] Where Z(θ)=∑ v,h e -E(v,h|θ) This is the normalization factor.

[0043] Furthermore, the specific process for determining the neuron activation probability based on the structural characteristics of RBM in S4 is as follows:

[0044] Based on the unique structure of the RBM (Relative Building Block) – no intralayer connections and fully connected interlayer connections – when the state of the visible layer neurons is determined, the activation probabilistics of each hidden layer neuron are independent. Therefore, the activation probability of the j-th hidden layer neuron can be obtained as follows:

[0045]

[0046] in, For the logistic function;

[0047] The RBM structure is symmetrical. When the state of the hidden layer neurons is determined, the probabilistic activation states of each visible layer neuron are also conditionally independent. Therefore, the activation probability of the i-th visible layer neuron is:

[0048]

[0049] In the formula, σ is the logistic function.

[0050] Furthermore, the specific process of fitting the training data by maximizing the log-likelihood function in S4 is as follows:

[0051] To fit the training sample data, it is necessary to maximize the log-likelihood function θ of the RBM on the training dataset. * That is, to find The maximum value of θ, where the gradient of the log-likelihood function with respect to θ is:

[0052]

[0053] Where t is the circular index; v (t) θ represents the observation sample corresponding to step t; θ is the set of parameters for the model; <·> P This represents finding the expected value of distribution P; P(h|v) (t) ,θ) represents the explicit layer and the training samples v (t) The probability distribution of the hidden layer; P(v,h|θ) is the joint distribution of neurons in the visible layer and neurons in the hidden layer;

[0054] Log-likelihood function for W ij a i and b j The partial derivative is:

[0055]

[0056] To address the issues of inability to directly obtain the parameters and the low efficiency of Gibbs sampling, the sampling-contrast divergence method is used to solve for the model parameters. The parameter update criterion is as follows:

[0057]

[0058] Among them, <·> C ∈ represents the distribution defined by the model after one-step reconstruction using the contrastive divergence method; ∈ represents the learning rate.

[0059] As a second aspect of the present invention, the present invention provides an intelligent land use change detection system that ranges from coarse to fine, comprising:

[0060] A classification dataset creation unit, used for creating land use scenario classification datasets;

[0061] The category detection unit is used to train no fewer than 10 CNN networks and differentiate the models based on their classification accuracy to complete the CNN network selection; the pre-trained CNN networks selected after selection are fine-tuned based on the dataset obtained in S1; and the land use scene category is determined through a multi-CNN collaborative scene category identification mechanism.

[0062] The training sample extraction unit is used to extract spectral change intensity information based on the RCVA method and texture change intensity information based on the GLCM method. Then, combined with the scene classification results, the training samples are selected by dividing the two intensity maps using multi-scale thresholds and taking the union.

[0063] A pixel-level change information extraction model training unit is used to form a DBN model by combining a multi-layer Restricted Boltzmann Machine (RBM) with a feedforward backpropagation network. The joint probability distribution of the explicit and hidden layers of the RBM is defined based on an energy function. The activation probability of neurons is determined based on the structural characteristics of the RBM. Training data is fitted by maximizing the log-likelihood function. The model parameters are iteratively optimized using the contrastive divergence method to initialize the RBM parameters. Based on the error backpropagation method, the network weights are updated by calculating the learning error of each RBM layer, and the entire network is fine-tuned to complete model training.

[0064] As a third aspect of the invention, the invention provides a computer-readable storage medium having a computer program stored thereon, the computer program being executed by a processor of any step of the above-described intelligent land use change detection method from coarse to fine.

[0065] In summary, compared with the prior art, the above-described technical solutions conceived by this invention can achieve the following beneficial effects:

[0066] 1. The intelligent land use change detection method of the present invention, from coarse to fine, trains no fewer than 10 CNN networks, selects and fine-tunes them based on classification accuracy, and determines the land use scene category by relying on a multi-CNN collaborative scene category identification mechanism. The training phase allows each CNN network to learn the characteristic patterns of land use scenes. After selection and adaptation, the network model that is more effective in identifying land use scenes is retained. The fine-tuning process further adapts the pre-trained network to specific datasets, strengthening its classification ability for target scenes. The multi-CNN collaborative mechanism integrates the judgment results of different networks, compensating for the limitations of a single network in feature capture and category determination, accurately defining the land use scene category, providing a clear reference for subsequent change detection, improving detection accuracy from the scene classification dimension, and making land use type identification more reliable.

[0067] 2. The intelligent land use change detection method of this invention, from coarse to fine, extracts spectral change intensity information using the RCVA method and texture change intensity information using the GLCM method. Combining scene classification results, it employs multi-scale thresholding on two intensity maps and performs union analysis to select training samples, providing high-quality data for model training. RCVA captures the radiative characteristics of land use change from a spectral perspective, while GLCM characterizes changes in land cover structure from a texture perspective. Multi-scale thresholding avoids sample confusion caused by a single threshold, and union analysis ensures the comprehensiveness of both changed and unchanged samples, providing rich and accurate training materials for subsequent model learning and improving the samples' ability to represent real land use changes.

[0068] 3. The intelligent land use change detection method of this invention, from coarse to fine, constructs a DBN model composed of a multi-layer RBM and a feedforward backpropagation network. Training is completed by defining the joint probability distribution, determining the activation probability of neurons, maximizing the log-likelihood function to fit the data, iteratively optimizing the initial RBM parameters using the contrastive divergence method, and fine-tuning the network by updating the weights through error backpropagation. This achieves refined detection of land use change. The DBN model utilizes multi-layer RBM pre-training to extract deep features, combined with backpropagation fine-tuning to improve the model's learning ability for land use change patterns. The definition of the energy function and probability distribution ensures the accuracy of the mathematical expression of change features, while the optimization mechanism of contrastive divergence and backpropagation guarantees the efficiency of parameter iteration. Ultimately, the model accurately captures subtle features of land use change, outputting high-precision detection results and providing technical support for dynamic monitoring of land resources. Attached Figure Description

[0069] Figure 1 This is a flowchart of the intelligent land use change detection method of the present invention, which is based on coarse to fine parameters, according to an embodiment of the present invention.

[0070] Figure 2 This is an example diagram of the NS-55 dataset from an embodiment of the present invention;

[0071] Figure 3 This is a schematic diagram of the intelligent detection method according to an embodiment of the present invention;

[0072] Figure 4 This is a system unit diagram of an embodiment of the present invention. Detailed Implementation

[0073] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention. Furthermore, the technical features involved in the various embodiments of this invention described below can be combined with each other as long as they do not conflict with each other.

[0074] Example 1

[0075] Please refer to Figure 1 This embodiment 1 provides an intelligent detection method for land use change from coarse to fine, including:

[0076] S1. Creation of a land use scenario classification dataset;

[0077] S2. Train no fewer than 10 CNN networks and differentiate the models based on their classification accuracy to complete the CNN network selection; select the pre-trained CNN networks after selection and fine-tune the models based on the dataset obtained in S1; determine the land use scene category through a multi-CNN collaborative scene category identification mechanism;

[0078] S3. Extract spectral change intensity information based on the RCVA method; extract texture change intensity information based on the GLCM method; and then, in combination with the scene classification results, use multi-scale thresholding to divide the two intensity maps and take the union to complete the selection of training samples.

[0079] S4. A DBN model is constructed by combining a multi-layer Restricted Boltzmann Machine (RBM) with a feedforward backpropagation network. The joint probability distribution of the explicit and hidden layers of the RBM is defined based on the energy function. The activation probability of neurons is determined based on the structural characteristics of the RBM. The training data is fitted by maximizing the log-likelihood function. The model parameters are iteratively optimized by solving the contrastive divergence method to initialize the RBM parameters. Based on the error backpropagation method, the network weights are updated by calculating the learning error of each RBM layer, and the entire network is fine-tuned to complete the model training.

[0080] This embodiment 1 further elaborates on the above steps.

[0081] (1) Creation of classification datasets

[0082] In a preferred embodiment, a self-made land use scene classification dataset, NS-55, is used, which contains richer scenes and a larger amount of data. Mainstream open-source scene classification datasets, NWPU-RESISC45, UCMerced_LandUse, WHU-RS19, and AID, contain 45, 21, 19, and 30 scene categories respectively, with some overlapping categories among the datasets. To enrich the scene categories and improve the model's ability to detect land types, this embodiment 1, based on the above four datasets, uses image graphics operations such as attribute transformation, resolution adjustment, and image cropping to create the NS-55 scene classification dataset. The basic data consists of 44,605 ​​RGB images, with each scene containing between 100 and 1,266 images. Through data augmentation, the number of images for each scene category is increased to 3,000.

[0083] (2) Category Detection

[0084] The intelligent land use change detection method described in Embodiment 1, progressing from coarse to fine, includes scene-level classification of land use categories, i.e., "coarse" detection, and pixel-level extraction of land use changes, i.e., "fine" detection. This step is the "coarse" detection part, achieved by training a scene-level classification deep learning model. Specifically, it includes the following steps:

[0085] 2.1 Selection of Deep Networks for Image Scene Complexity

[0086] There is a certain adaptive relationship between the complexity of remote sensing images and the depth of CNNs. That is, when using shallow and deep CNNs to classify low-complexity and high-complexity images respectively, effectively utilizing the correspondence between CNNs and images can fully leverage the inherent characteristics of multiple CNNs and improve classification accuracy. Ten CNN networks—AlexNet, ResNet50, ResNet101, ResNet152, DenseNet121, DenseNet161, DenseNet169, Inception-v3, Inception-Res-v2, and Xception—were trained using the NWPU-RESISC45 dataset.

[0087] The models are differentiated based on their classification accuracy: AlexNet has a high recognition rate for macroscopic scenes such as snow-capped mountains, beaches, clouds, highways, forests, and bridges; ResNet50 has a high recognition rate for scenes with fixed shapes such as basketball courts, baseball fields, golf courses, athletic fields, and airplanes; DenseNet169 has a high recognition rate for abstract scenes with more semantic features such as parks, ports, jungles, sea ice, and islands; and ResNet152 has a high recognition rate for scenes such as lakes, terraced fields, thermal power plants, and mobile home parking lots.

[0088] 2.2 Model Fine-tuning Based on NS-55 Dataset

[0089] To accelerate model convergence and improve training efficiency, the concept of transfer learning (i.e., using existing knowledge to solve problems in different but related domains) was incorporated. AlexNet, ResNet50, ResNet152, and DenseNet169 were used as pre-trained CNNs, specifically networks pre-trained on the 1000-class ImageNet dataset. The NS-55 dataset was used by resetting the last layer of the network to match the number of classes in the new dataset, and then updating the weights of the CNNs during training on the new dataset.

[0090] 2.3 Confirmation of Land Use Scenarios

[0091] Based on the analysis in step 2.1, AlexNet demonstrates high validation accuracy for macroscopic scenes such as snow-capped mountains and beaches, ResNet50 for fixed-shape scenes such as basketball courts and baseball fields, DenseNet169 for scenes with multiple semantic features such as parks and ports, and ResNet152 for more complex scenes such as lakes and terraced fields. This verifies the adaptive relationship between the complexity of remote sensing images and the depth of CNNs. To fully leverage the recognition capabilities of different CNNs for remote sensing images and achieve higher scene recognition accuracy, this embodiment 1 designs a scene category labeling mechanism based on different CNNs.

[0092] In a preferred embodiment, the scene category identification mechanism for multi-CNN collaboration is specifically as follows:

[0093] Let the scene assignment set be S = {S1, S2, ..., S...} k}, where S i Let k be the set of scenes assigned to the i-th CNN, and k be the number of CNNs; the set of prediction results is p = {p1, p2, ..., p...} k}, where p i Let be the predicted category of the i-th CNN for scene x;

[0094] The historical precision set is acc = {acc1, acc2, ..., acc k}, where acc i Let be the prediction accuracy of the i-th CNN in the corresponding scene;

[0095] Final Category C final The rules for determining it are as follows:

[0096] When there exists a class c such that all p i When =c, then C final=c ;

[0097] When there exists i such that p i ∈S i All p i When inconsistencies exist, then C final To satisfy p i ∈S i The CNN prediction category with the highest fitness is selected if multiple i satisfy the criteria, according to the preset priority.

[0098] When all When, then C final The CNN output class with the highest prediction accuracy, i.e., C final =p j , where j is the value that makes acc j The largest index.

[0099] (3) Training sample extraction

[0100] By using both RCVA and GLCM methods, the differences in spectral and texture features of remote sensing images are taken into account. This approach addresses to some extent the impact of image configuration errors and feature extraction on change detection results. Furthermore, by using a custom threshold, training samples that are more likely to have changed and those that have not changed are selected.

[0101] 3.1 Extraction of Spectral Change Intensity Information Based on RCVA Method

[0102] The RCVA method considers the neighborhood information of pixels and selects the pixel pair with the smallest spectral difference for detection, which greatly eliminates the impact of registration error.

[0103] The method is based on the following assumption: Given a pixel x1(j,k) in image 1, find the pixel x2 with the smallest spectral difference from x1 within the range of image 2 (j±ω,k±ω). Then, consider x2 as a corresponding image point of x1, and calculate the difference x. d-a Similarly, based on image 2, find the corresponding image points in image 1 and calculate the difference x. d-b Finally, with x d-a and x d-b The smaller of the two is taken as the intensity of change, and the specific calculation process is as follows:

[0104]

[0105] In the formula, (j,k) represents the coordinate index of the target pixel in the image; (p,q) represents the coordinate index of the pixel traversed within the window range [j-ω,k+ω] centered at (j,k); ω represents the window radius; and n represents the number of bands involved in the calculation. These correspond to the pixel values ​​at coordinates (j,k) in the i-th band of the two types of data (or images); similarly... It is the pixel value at the neighboring point (p,q), which belongs to the "original data value";

[0106] By iterating through all pixels in both images, the intensity value M of the spectral change for all pixels can be obtained. S This leads to the spectral information intensity map that takes into account neighborhood information.

[0107] 3.2 Extraction of Texture Variation Intensity Information Based on GLCM Method

[0108] GLCM is a classic method for texture extraction and is currently the most widely used and effective texture feature analysis method. GLCM primarily extracts texture through the conditional probability density between image gray levels. It is a special matrix that describes the gray-level relationship between adjacent pixels and neighboring pixels at a certain distance. Using variance as a scalar best reflects the differences between different ground features. The calculation process is as follows:

[0109]

[0110] Where p(l,g,d,θ,i) is the frequency of occurrence of a pixel with gray value l in the i-th band of the image and a pixel with gray value g at a distance d from it; θ is the direction for generating the gray-level co-occurrence matrix; N is the gray level; The mean of the frequencies;

[0111] After obtaining the texture variance feature values ​​of image 1 and image 2, the texture variation intensity M can be calculated. T The calculation process is as follows:

[0112]

[0113] In the formula, These are variance-based feature values ​​calculated from the gray-level co-occurrence matrix. is the difference between the two variance features mentioned above; n is the total number of "feature sequences", that is, the total number of bands.

[0114] Using RCVA and GLCM change intensity maps, changed and unchanged regions are divided into samples using thresholds. However, using a single threshold often leads to high confusion between changed and unchanged samples, resulting in low-quality training samples. Therefore, this embodiment 1 employs a multi-scale threshold strategy for selecting training samples. Based on the CNNs trained in S2, category detection is performed on remote sensing scene pairs obtained through multi-scale segmentation and overlapping cropping, and changed scenes are identified based on the categories. On the RCVA and GLCM change intensity maps, changed patches are segmented using multi-scale thresholds to obtain training samples, and the union of the samples from both methods is taken as the final training sample.

[0115] (4) Training of the pixel-level change information extraction model

[0116] Using the training samples selected by S3, a DBN model is trained to achieve pixel-level detection of change regions. DBN, developed from biological neural networks and shallow neural networks, is a probabilistic generative model that infers the distribution of sample data through a joint probability distribution. It extracts features from a large amount of unlabeled sample data through layer-by-layer unsupervised training and optimizes the model using a small amount of labeled sample data. Finally, it obtains the optimal weights for the network, enabling it to generate training data with the highest probability, forming high-level abstract features and resulting in a high-performance classification model. DBN's input is a vector composed of continuously input real-valued variables, overcoming the limitation of CNNs that require images of a certain size as input. It can be easily applied to pixel-based change detection, enabling more detailed image detection.

[0117] DBN mainly consists of two parts: the first part is a multi-layer Restricted Boltzmann Machine (RBM) used for pre-training the network; the second part is a feedforward backpropagation network, which can refine the stacked RBM network. The basic component of the DBN model, the RBM, contains two layers (explicit layer v and hidden layer h), and is an undirected graphical model. Each layer can be defined as a vector, and the dimension of the vector is the number of neurons in that layer. Neurons in different layers are connected by a weight matrix W.

[0118] For each RBM, its energy as a system is:

[0119]

[0120] Where n and m are the number of neurons in the visible and hidden layers, respectively; v i and h j These represent the states of neurons in the visible and hidden layers, respectively; θ = {W ij ,a i ,b j} represents the model parameters, where W ij Let a be the weight between neurons i and j. i and b j These represent the biases of neurons in the visible and hidden layers, respectively.

[0121] Based on the energy function, the joint probability distribution of (v,h) can be obtained:

[0122]

[0123] Where Z(θ)=∑ v,h e -E(v,h|θ) This is the normalization factor.

[0124] At this point, the marginal distribution (likelihood function) of the joint probability distribution of v and h defined by RBM is:

[0125]

[0126] Based on the unique structure of the RBM (Relative Building Block) – no intralayer connections and fully connected interlayer connections – when the state of the visible layer neurons is determined, the activation probabilistics of each hidden layer neuron are independent. Therefore, the activation probability of the j-th hidden layer neuron can be obtained as follows:

[0127]

[0128] in, For the logistic function;

[0129] The RBM structure is symmetrical. When the state of the hidden layer neurons is determined, the probabilistic activation states of each visible layer neuron are also conditionally independent. Therefore, the activation probability of the i-th visible layer neuron is:

[0130]

[0131] In the formula, σ is the logistic function.

[0132] To fit the training sample data, it is necessary to maximize the log-likelihood function θ of the RBM on the training dataset. * That is, to find The maximum value of θ, where the gradient of the log-likelihood function with respect to θ is:

[0133]

[0134] Where t is the circular index; v (t) θ represents the observation sample corresponding to step t; θ is the set of parameters for the model; <·> P This represents finding the expected value of distribution P; P(h|v) (t) ,θ) represents the explicit layer and the training samples v (t) The probability distribution of the hidden layer; P(v,h|θ) is the joint distribution of neurons in the visible layer and neurons in the hidden layer;

[0135] Log-likelihood function for W ij a i and b j The partial derivative is:

[0136]

[0137]

[0138] To address the issues of inability to directly obtain the parameters and the low efficiency of Gibbs sampling, the sampling-contrast divergence method is used to solve for the model parameters. The parameter update criterion is as follows:

[0139]

[0140] Among them, <·> C ∈ represents the distribution defined by the model after one-step reconstruction using the contrastive divergence method; ∈ represents the learning rate.

[0141] The contrastive divergence method can quickly find an approximate solution, and the network parameters θ can be continuously optimized through iteration. * The goal is to initialize each RBM parameter so that the feature vectors in the RBM network can retain the most feature information when mapped to different feature spaces. Based on the error backpropagation method, the network weights are updated by calculating the learning error of each RBM layer using the labels corresponding to the sample data and the actual output values ​​of the model, and the entire network is fine-tuned.

[0142] Based on the changed and unchanged samples identified by S3, in a preferred embodiment, this paper uses a vector of normalized RGB brightness values ​​of pixels within a 2×2 area of ​​the image, arranged sequentially, as input, and establishes corresponding labels for each sample. During DBN training, the model is saved after the training accuracy reaches the ideal value; the saving process involves saving the network structure, weights, and bias values. Then, the samples to be detected are extracted from both images using a 2×2 window, and the saved model is called to judge the samples to be detected, dividing the changed and unchanged regions based on the judgment results.

[0143] The technical solution of Embodiment 1 will be described in detail below with reference to the accompanying drawings and more specific implementation methods.

[0144] Please refer to Figure 2 The NS-55 dataset constructed in this embodiment 1 contains 55 categories of remote sensing scenes: farmland, airplane, airport, wasteland, baseball field, basketball court, beach, bridge, building, center, jungle, church, circular farmland, cloud, commercial area, dense residential area, desert, forest, highway, golf course, track and field, port, industrial area, intersection, island, lake, grassland, medium-density residential area, mobile home parking lot, mountain, overpass, palace, park, parking lot, pond, railway, train station, rectangular farmland, resort area, river, roundabout, airport runway, school, sea ice, boat, snow mountain, sparse residential area, square, open-air sports center, oil storage tank, tennis court, terraced fields, thermal power plant, viaduct, and swamp.

[0145] Please refer to Figure 3This embodiment 1 proposes a coarse-to-fine intelligent land use change detection method, consisting of three steps. The first step involves training a change category detection neural network. AlexNet, ResNet50, ResNet152, and DenseNet169 neural networks trained on the NS-55 self-made dataset are used to classify scenes in the experimental image. The experimental scenes are acquired using multi-scale segmentation and overlapping cropping. The second step extracts training samples for change range detection. The introduced RCVA and GLCM methods are used to extract spectral and texture change features from the image, respectively. Based on the changed scenes identified in the first step, change training samples are divided using a multi-scale thresholding method in the spectral and texture change intensity maps. The third step detects the change range. A deep belief network is trained using the selected training samples and the change range is detected across the entire image. The specific steps are as follows:

[0146] Step 1: Training the change category detection neural network;

[0147] This step aims to achieve category detection of land use change, specifically the "coarse" stage in the "coarse-to-fine" detection process. First, a land use scene-level classification neural network is constructed. By integrating four existing mainstream scene classification datasets—NWPU-RESISC45, UCMerced_LandUse, WHU-RS19, and AID—and combining image attribute transformation, resolution unification, and cropping techniques, a custom NS-55 scene classification dataset is built, covering 55 typical scene categories and a total of 44,605 ​​RGB remote sensing images. Each category of images is expanded to 3,000 images through data augmentation to improve the model's robustness and generalization ability.

[0148] Subsequently, several deep convolutional neural networks (CNNs), including AlexNet, ResNet50, ResNet152, and DenseNet169, were selected. Combined with transfer learning strategies, the fully connected layers of the networks were fine-tuned based on the ImageNet pre-trained models to adapt to the output categories of the NS-55 dataset, completing model training. After training, an adaptive classification mechanism was designed based on the accuracy performance of different CNN models in various remote sensing scenarios: different scene types are identified by different CNN models. When multiple models predict the same result, the consistent category is used as the final classification; if they are inconsistent, the prediction result of the model corresponding to the task assignment or the model with the highest prediction accuracy is prioritized. This ultimately achieves the detection of major scene categories in changing areas, providing candidate regions for subsequent pixel-level change detection.

[0149] Step 2: Extract training samples for range of change detection;

[0150] After completing the change category detection, for the detected change areas, it is necessary to further extract high-quality training samples for pixel-level change detection. To this end, this embodiment 1 adopts a training sample extraction strategy based on spectral-texture joint features, combining the RCVA method and the GLCM method to extract representative samples from the remote sensing images before and after the change.

[0151] The RCVA method obtains a change intensity map by searching for spectral differences between neighboring pixels in two image periods, effectively overcoming the interference caused by registration errors. The GLCM method analyzes pixel texture features through gray-level co-occurrence matrix analysis, quantifies the changes between different land cover types using texture variance, and generates a texture change intensity map.

[0152] In this embodiment 1, a multi-scale threshold partitioning strategy is employed to extract potentially changed and unchanged samples from the two change intensity maps mentioned above, and their unions are fused to form the final training sample set. The sample extraction process, combined with the CNN scene classification results from step one, performs scene judgment and block cropping on the changed regions, improving the spatial representativeness and category diversity of the training samples, and ensuring that the subsequent model learning has sufficient discriminative power.

[0153] Step 3: Detect the range of change;

[0154] After obtaining high-quality training samples, this embodiment 1 uses a deep belief network (DBN) to construct a pixel-level change detection model, completing the "fine" stage of the "coarse-to-fine" detection process. As a multi-layer unsupervised pre-trained probabilistic generative model, DBN has the ability to extract deep features from high-dimensional complex data and can adapt to nonlinear changes in spatial, spectral, and textural features of remote sensing images.

[0155] Specifically, DBN consists of multiple Restricted Boltzmann Machines (RBMs). First, it pre-trains layer by layer on unlabeled training samples to capture deep features. Then, it performs supervised fine-tuning using a small number of labeled samples that show changes or no changes, optimizing the network parameters. The input data is a vector of normalized RGB channel brightness values ​​within a 2×2 neighborhood of each pixel, and the model output is a predicted label indicating whether the pixel has changed.

[0156] After training, the DBN model is used to perform sliding window processing on the remote sensing image to be detected, extract all pixel blocks to be classified, and input them one by one into the trained model for inference. Finally, a change probability map is output, and a binary change detection map is generated by threshold judgment to complete high-precision pixel-level change range detection.

[0157] Example 2

[0158] Please refer to Figure 4 This embodiment 2 provides an intelligent land use change detection system that ranges from coarse to fine, including:

[0159] A classification dataset creation unit, used for creating land use scenario classification datasets;

[0160] The category detection unit is used to train no fewer than 10 CNN networks and differentiate the models based on their classification accuracy to complete the CNN network selection; the pre-trained CNN networks selected after selection are fine-tuned based on the dataset obtained in S1; and the land use scene category is determined through a multi-CNN collaborative scene category identification mechanism.

[0161] The training sample extraction unit is used to extract spectral change intensity information based on the RCVA method and texture change intensity information based on the GLCM method. Then, combined with the scene classification results, the training samples are selected by dividing the two intensity maps using multi-scale thresholds and taking the union.

[0162] A pixel-level change information extraction model training unit is used to form a DBN model by combining a multi-layer Restricted Boltzmann Machine (RBM) with a feedforward backpropagation network. The joint probability distribution of the explicit and hidden layers of the RBM is defined based on an energy function. The activation probability of neurons is determined based on the structural characteristics of the RBM. Training data is fitted by maximizing the log-likelihood function. The model parameters are iteratively optimized using the contrastive divergence method to initialize the RBM parameters. Based on the error backpropagation method, the network weights are updated by calculating the learning error of each RBM layer, and the entire network is fine-tuned to complete model training.

[0163] Example 3

[0164] This embodiment 3 also provides a computer-readable storage medium storing a computer program, which, when executed by a processor, can implement any step of an intelligent land use change detection method from coarse to fine.

[0165] The computer-readable storage medium may include various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0166] For a description of the computer-readable storage medium provided in this application, please refer to the above method embodiments; further details will not be repeated here.

[0167] Those skilled in the art will readily understand that the above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.

Claims

1. An intelligent detection method for land use change from coarse to fine, characterized in that, include: S1. Creation of a land use scenario classification dataset; S2. Train no fewer than 10 CNN networks and differentiate the models based on their classification accuracy to complete the CNN network selection; The pre-trained CNN networks selected after screening were fine-tuned based on the dataset obtained in S1. Land use scene categories are determined through a multi-CNN collaborative scene category identification mechanism; S3. Extract spectral change intensity information based on the RCVA method; extract texture change intensity information based on the GLCM method; and then, in combination with the scene classification results, use multi-scale thresholding to divide the two intensity maps and take the union to complete the selection of training samples. S4. A DBN model is constructed by combining a multilayer restricted Boltzmann machine (RBM) with a feedforward backpropagation network; the joint probability distribution of the explicit and hidden layers of the RBM is defined based on the energy function; the activation probability of neurons is determined based on the structural characteristics of the RBM; and the training data is fitted by maximizing the log-likelihood function. The RBM parameters are initialized by iterative optimization using the comparative divergence method to solve for the model parameters. Based on the backpropagation method, the network weights are updated by calculating the learning error of each RBM layer, and the entire network is fine-tuned to complete the model training.

2. The intelligent detection method for land use change from coarse to fine as described in claim 1, characterized in that, The scene category identification mechanism in S2 involving multi-CNN collaboration is specifically as follows: Let the scene assignment set be S = {S1, S2, ..., S...} k }, where S i Let k be the set of scenes assigned to the i-th CNN, and k be the number of CNNs; the set of prediction results is p = {p1, p2, ..., p...} k }, where p i Let be the predicted category of the i-th CNN for scene x; The historical precision set is acc = {acc1, acc2, ..., acc k }, where acc i Let be the prediction accuracy of the i-th CNN in the corresponding scene; Final Category C final The rules for determining this are as follows: When there exists a class c such that all p i When =c, then C final=c ; When there exists i such that p i ∈S i All p i When inconsistencies exist, then C final To satisfy p i ∈S i The CNN prediction category with the highest fitness is selected if multiple i satisfy the criteria, according to the preset priority. When all When, then C final The CNN output class with the highest prediction accuracy, i.e., C final =p j , where j is the value that makes acc j The largest index.

3. The intelligent detection method for land use change from coarse to fine as described in claim 1, characterized in that, The specific process of extracting spectral change intensity information based on the RCVA method in S3 is as follows: The RCVA method eliminates the influence of registration error by considering the neighborhood information of the pixel and selecting the pixel pair with the smallest spectral difference for detection. The method is based on the following assumption: Given a pixel x1(j,k) in image 1, find the pixel x2 with the smallest spectral difference from x1 within the range of image 2 (j±ω,k±ω). Then, consider x2 as a corresponding image point of x1, and calculate the difference x. d-a Similarly, based on image 2, find the corresponding image points in image 1 and calculate the difference x. d-b Finally, with x d-a and x d-b The smaller of the two is taken as the intensity of change, and the specific calculation process is as follows: In the formula, (j,k) represents the coordinate index of the target pixel in the image; (p,q) represents the pixel coordinate index traversed within the window range [j-ω,k+ω] centered at ((j,k); ω represents the window radius; n represents the number of bands involved in the calculation; These correspond to the pixel values ​​at coordinates (j,k) in the i-th band of the two types of data (or images); similarly... It is the pixel value at the neighboring point ((p,q), which belongs to the "original data value"; By iterating through all pixels in both images, the intensity value M of the spectral change for all pixels can be obtained. S This leads to the spectral information intensity map that takes into account neighborhood information.

4. The intelligent detection method for land use change from coarse to fine as described in claim 1, characterized in that, The specific process of extracting texture change intensity information based on the GLCM method in S3 is as follows: GLCM extracts texture through the conditional probability density between image gray levels, using variance as a scalar to reflect the differences between different ground features. The calculation process is as follows: Where p(l,g,d,θ,i) is the frequency of occurrence of a pixel with gray value l in the i-th band of the image and a pixel with gray value g at a distance d from it; θ is the direction for generating the gray-level co-occurrence matrix; N is the gray level; The mean of the frequencies; After obtaining the texture variance feature values ​​of image 1 and image 2, the texture variation intensity M can be calculated. T The calculation process is as follows: In the formula, These are variance-based feature values ​​calculated from the gray-level co-occurrence matrix. It is the difference between the two variance features mentioned above; n is the total number of "feature sequences", that is, the total number of bands.

5. The intelligent detection method for land use change from coarse to fine as described in claim 1, characterized in that, In S3, the process of selecting training samples by combining scene classification results, using multi-scale thresholding to divide and combine the two intensity maps, and taking the union is as follows: Based on the CNNs trained by S2, class detection is performed on remote sensing scene pairs obtained through multi-scale segmentation and overlapping cropping, and the changed scene is obtained according to the class. On the change intensity maps of RCVA and GLCM, the changed patches are divided into training samples by multi-scale thresholding, and the union of the samples divided by the two is taken as the final training sample.

6. The intelligent detection method for land use change from coarse to fine as described in claim 1, characterized in that, The specific process of defining the joint probability distribution of the explicit and hidden layers of the RBM based on the energy function in S4 is as follows: An RBM consists of two layers: a visible layer (v) and a hidden layer (h), forming an undirected graphical model. Each layer is defined as a vector, with the dimension of the vector being the number of neurons in that layer. Neurons in different layers are connected by a weight matrix (W). For each RBM, its energy as a system is: Where n and m are the number of neurons in the visible and hidden layers, respectively; v i and h j These represent the states of neurons in the visible and hidden layers, respectively; θ = {W ij ,a i ,b j } represents the model parameters, where W ij Let a be the weight between neurons i and j. i and b j These represent the biases of neurons in the visible and hidden layers, respectively. Based on the energy function, the joint probability distribution of (v,h) can be obtained: Where Z(θ)=∑ v,h e -E(v,h|θ) This is the normalization factor.

7. The intelligent detection method for land use change from coarse to fine as described in claim 1, characterized in that, The specific process for determining the neuron activation probability based on the structural characteristics of RBM in S4 is as follows: Based on the unique structure of the RBM (Relative Building Block) – no intralayer connections and fully connected interlayer connections – when the state of the visible layer neurons is determined, the activation probabilistics of each hidden layer neuron are independent. Therefore, the activation probability of the j-th hidden layer neuron can be obtained as follows: in, For the logistic function; The RBM structure is symmetrical. When the state of the hidden layer neurons is determined, the probabilistic activation states of each visible layer neuron are also conditionally independent. Therefore, the activation probability of the i-th visible layer neuron is: In the formula, σ is the logistic function.

8. The intelligent detection method for land use change from coarse to fine as described in claim 1, characterized in that, The specific process of fitting the training data by maximizing the log-likelihood function in S4 is as follows: To fit the training sample data, it is necessary to maximize the log-likelihood function θ of the RBM on the training dataset. * That is, to find The maximum value of θ, where the gradient of the log-likelihood function with respect to θ is: Where t is the circular index; v (t) These are the observation samples corresponding to step t; θ is the set of parameters for the model; <·> P This represents finding the expected value of distribution P; P(h|v) (t) ,θ) represents the explicit layer and the training samples v (t) The probability distribution of the hidden layer; P(v,h|θ) is the joint distribution of neurons in the visible layer and neurons in the hidden layer; Log-likelihood function for W ij a i and b j The partial derivative is: To address the issues of inability to directly obtain the parameters and the low efficiency of Gibbs sampling, the sampling-contrast divergence method is used to solve for the model parameters. The parameter update criterion is as follows: Among them, <·> C ∈ represents the distribution defined by the model after one-step reconstruction using the contrastive divergence method; ∈ represents the learning rate.

9. An intelligent land use change monitoring system that ranges from coarse to fine, characterized in that, include: The classification dataset creation unit is used for creating land use scenario classification datasets. The category detection unit is used to train no fewer than 10 CNN networks and differentiate the models based on their classification accuracy to complete the CNN network selection. The pre-trained CNN networks selected after screening were fine-tuned based on the dataset obtained in S1. Land use scene categories are determined through a multi-CNN collaborative scene category identification mechanism; The training sample extraction unit is used to extract spectral change intensity information based on the RCVA method and texture change intensity information based on the GLCM method. Then, combined with the scene classification results, the training samples are selected by dividing the two intensity maps using multi-scale thresholds and taking the union. A pixel-level change information extraction model training unit is used to form a DBN model by combining a multi-layer Restricted Boltzmann Machine (RBM) with a feedforward backpropagation network; the joint probability distribution of the explicit and hidden layers of the RBM is defined based on the energy function; the activation probability of neurons is determined based on the structural characteristics of the RBM; and the training data is fitted by maximizing the log-likelihood function. The RBM parameters are initialized by iterative optimization using the comparative divergence method to solve for the model parameters. Based on the backpropagation method, the network weights are updated by calculating the learning error of each RBM layer, and the entire network is fine-tuned to complete the model training.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, The computer program is executed by a processor as described in any one of claims 1-8: a coarse-to-fine intelligent detection method for land use change.

Citation Information

Patent Citations

  • SAR image classification method based on texture features and DBN

    CN107506699A

  • Remote sensing image land utilization classification method based on residual network

    CN118537653A

  • Method for classifying hyperspectral images on basis of adaptive multi-scale feature extraction model

    US20230252761A1