Remote sensing image data screening method

By performing multidimensional spatial information measurement and principal component analysis on remote sensing images and classified raster images, an information transformation relationship model was established, which solved the problems of insufficient accuracy and adaptability in remote sensing image data screening, and achieved efficient image data screening and change detection.

CN121746921APending Publication Date: 2026-03-27CENT SOUTH UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-19
Publication Date
2026-03-27

AI Technical Summary

Technical Problem

Existing remote sensing image data screening methods suffer from poor accuracy and adaptability. In particular, when historical remote sensing image data is lacking, it is difficult to accurately screen out the true changes in the land surface, and the storage of massive amounts of data is also challenging.

Method used

By acquiring remote sensing image data and classified raster images of the current time phase, multidimensional spatial information measurement is performed. Principal component analysis is used to remove redundant information and reduce dimensionality, an information transformation relationship model is established, and image data is screened in combination with regional variation rate.

Benefits of technology

In the absence of historical remote sensing image data, it improves the accuracy and efficiency of remote sensing image data screening, effectively conveys multi-dimensional change information of ground features, reduces noise and redundancy, and improves the efficiency of change detection and regional positioning accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121746921A_ABST
    Figure CN121746921A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a remote sensing image data screening method, and belongs to the technical field of data processing, and the method specifically comprises the steps: obtaining remote sensing image data of a current time phase and a classification raster image of a corresponding region; performing multi-dimensional spatial information measurement on the remote sensing image data and the classified raster image, and extracting corresponding multi-dimensional spatial information components; respectively carrying out principal component analysis on the multi-dimensional spatial information components of the remote sensing image data and the classified grid image to obtain a remote sensing image information principal component and a classified grid information principal component; according to the remote sensing image information principal component of the current time phase, utilizing a preset information conversion relation model to predict and obtain a corresponding classification grid information principal component; and screening the remote sensing image data of the current time phase by using the difference between the principal component of the classification grid information and the principal component of the reference classification grid information in combination with the regional variation rate. Through the scheme disclosed by the invention, the screening precision and adaptability are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of data processing technology, and in particular to a method for filtering remote sensing image data. Background Technology

[0002] Currently, existing methods typically use the difference in spectral information between two sets of remote sensing image data, i.e., the change vector method, for remote sensing image data screening. This method has two shortcomings: 1) Due to the "different spectra for the same object, different objects for the same spectrum" characteristic of remote sensing image data, a large number of "spurious changes" occur. Furthermore, remote sensing images are also subject to interference from factors such as season, light intensity, and imaging angle, which easily affect the true spectral information of ground objects. Therefore, relying solely on the spectral information of two sets of remote sensing images is insufficient to reflect the true changes on the Earth's surface, resulting in poor accuracy in remote sensing image screening; 2) In most scenarios, earlier remote sensing images are often inaccessible to ordinary users or researchers due to data quality, offline storage, or access restrictions. In addition, the recent surge in both the "order of magnitude" and "data volume" of newly added images has drastically reduced the cost-effectiveness of storage systems, forcing massive amounts of data to be archived offline, making it increasingly difficult to pinpoint image data for specific time phases, regions, and imaging conditions. Therefore, due to the lack of earlier remote sensing image data, existing methods cannot be used for image data screening.

[0003] It is evident that there is an urgent need for a remote sensing image data screening method with high accuracy and adaptability. Summary of the Invention

[0004] In view of this, the present disclosure provides a remote sensing image data filtering method, which at least partially solves the problems of poor filtering accuracy and adaptability in the prior art.

[0005] This disclosure provides a method for filtering remote sensing image data, including:

[0006] Step 1: Acquire the remote sensing image data of the current time phase and the corresponding classified raster image of the region;

[0007] Step 2: Perform multidimensional spatial information measurement on the remote sensing image data and the classified raster image respectively, and extract the corresponding multidimensional spatial information components;

[0008] Step 3: Perform principal component analysis on the multidimensional spatial information components of remote sensing image data and classified raster images respectively, remove redundant information and reduce dimensionality to obtain the principal components of remote sensing image information and the principal components of classified raster information.

[0009] Step 4: Based on the principal components of the remote sensing image information at the current time, use the preset information transformation relationship model between the principal components of the remote sensing image information and the principal components of the classification raster information to predict the corresponding principal components of the classification raster information.

[0010] Step 5: Utilize the differences between the principal components of the classification raster information and the principal components of the baseline classification raster information, combined with the regional variation rate, to filter the remote sensing image data of the current time phase.

[0011] According to a specific implementation of this disclosure, the multidimensional spatial information components include texture information, topological information, and spatial distribution information corresponding to remote sensing image data, as well as geometric information, thematic information, topological information, and spatial distribution information corresponding to classified raster images.

[0012] According to a specific implementation of this disclosure, before step 4, the method further includes:

[0013] Acquire multiple sets of training samples, each set including remote sensing image data and classified raster images of the same area;

[0014] Multidimensional spatial information components of remote sensing image data and classified raster images in each training sample are extracted, and principal component analysis is performed to reduce dimensionality.

[0015] Using the principal components of the dimensionality-reduced remote sensing image information as input features and the principal components of the dimensionality-reduced classified raster information as output labels, a multiple linear regression model is used for training to obtain an information transformation relationship model.

[0016] According to a specific implementation of this disclosure, the expression of the multiple linear regression model is as follows:

[0017]

[0018] Here, Y is the output variable, which is the principal component of the categorical raster information. As the output variable, namely the principal components of information from remotely sensed images, For regression coefficients, This is the error term.

[0019] According to a specific implementation of this disclosure, the expression for the regional variation rate is:

[0020]

[0021] Where N is the total number of pixels within the region of interest A; and These are the category attribute values ​​of the i-th pixel in the previous and subsequent time-phase land cover data, respectively. Summation is performed by iterating through all pixels.

[0022] According to a specific implementation of an embodiment of this disclosure, step 5 specifically includes:

[0023] Based on the magnitude of the regional variation rate, the image regions are divided into three levels: slight variation, moderate variation, and significant variation.

[0024] Remote sensing image data are filtered based on different levels, with priority given to image areas showing significant and moderate changes.

[0025] According to a specific implementation of an embodiment of this disclosure, after step 1, the method further includes:

[0026] Preprocessing is performed on remote sensing image data and classified raster images, wherein the preprocessing includes one or more of geometric correction, registration, fusion, color balancing and cropping.

[0027] According to one specific implementation of this disclosure, the remote sensing image data is Landsat series satellite imagery, and the classified raster image is GlobeLand30 land cover data product.

[0028] The remote sensing image data filtering scheme in this embodiment includes: Step 1, acquiring remote sensing image data of the current time phase and the corresponding regional classification raster image; Step 2, performing multidimensional spatial information measurement on the remote sensing image data and the classification raster image respectively, and extracting the corresponding multidimensional spatial information components; Step 3, performing principal component analysis on the multidimensional spatial information components of the remote sensing image data and the classification raster image respectively, removing redundant information and reducing dimensionality to obtain the principal components of remote sensing image information and the principal components of classification raster information; Step 4, based on the principal components of remote sensing image information of the current time phase, using a preset information transformation relationship model between the principal components of remote sensing image information and the principal components of classification raster information, predicting the corresponding principal components of classification raster information; Step 5, using the difference between the principal components of classification raster information and the baseline principal components of classification raster information, combined with the regional variation rate, filtering the remote sensing image data of the current time phase.

[0029] The beneficial effects of this disclosure are as follows: The solution of this disclosure only requires a single-period remote sensing image, enabling remote sensing image data screening even when historical remote sensing image data is lacking and existing solutions cannot be used. As an auxiliary means for change detection operations, image data screening can identify image data rich in change information in advance, greatly improving the overall efficiency of change detection. Simultaneously, with the development of imaging technology and classification algorithms, it is expected that in the future, while ensuring production accuracy, the area of ​​the data screening cell grid will continue to shrink, and the regional positioning accuracy of change information will continue to improve, thereby further enhancing the application value of this solution. From the perspective of information transmission, the existing two-period image change vector method essentially only utilizes the single spectral information content of the remote sensing image, failing to reflect the multi-dimensional spatial information and changes in spatial relationships, geometric structures, and spatial distribution of the remote sensing image. Therefore, change information is difficult to effectively transmit through pixel grayscale difference maps. This solution, based on the multi-dimensional information content of the image, can not only effectively transmit various change information of ground features, but also uses PCA to reduce noise and redundancy in the original multi-dimensional information components, thereby significantly improving the image screening accuracy. Attached Figure Description

[0030] To more clearly illustrate the technical solutions of the embodiments of this disclosure, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this disclosure. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0031] Figure 1 A flowchart illustrating a remote sensing image data filtering method provided in this embodiment of the disclosure;

[0032] Figure 2 This is a schematic diagram illustrating the technical concept of a remote sensing image data filtering method provided in an embodiment of the present disclosure;

[0033] Figure 3 A flowchart illustrating the modeling of the information conversion relationship between optical remote sensing images and classified raster images provided in this disclosure embodiment;

[0034] Figure 4 A flowchart illustrating a remote sensing image data filtering technique based on a classification raster and an information transformation relationship model, provided for embodiments of this disclosure. Detailed Implementation

[0035] The embodiments of this disclosure will now be described in detail with reference to the accompanying drawings.

[0036] The following specific examples illustrate the implementation of this disclosure. Those skilled in the art can easily understand other advantages and effects of this disclosure from the content disclosed in this specification. Obviously, the described embodiments are only a part of the embodiments of this disclosure, and not all of them. This disclosure can also be implemented or applied through other different specific embodiments, and the details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of this disclosure. It should be noted that, in the absence of conflict, the following embodiments and features in the embodiments can be combined with each other. Based on the embodiments in this disclosure, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this disclosure.

[0037] It should be noted that various aspects of embodiments within the scope of the appended claims are described below. It will be apparent that the aspects described herein can be embodied in a wide variety of forms, and any particular structure and / or function described herein is merely illustrative. Based on this disclosure, those skilled in the art will understand that one aspect described herein can be implemented independently of any other aspect, and two or more of these aspects can be combined in various ways. For example, any number of aspects set forth herein can be used to implement the device and / or practice the method. Additionally, this device and / or method can be implemented using structures and / or functionalities other than one or more of the aspects set forth herein.

[0038] It should also be noted that the illustrations provided in the following embodiments are only schematic representations of the basic concept of this disclosure. The illustrations only show the components related to this disclosure and are not drawn according to the number, shape and size of the components in actual implementation. In actual implementation, the form, quantity and proportion of each component can be arbitrarily changed, and the layout of the components may also be more complex.

[0039] Furthermore, specific details are provided in the following description to facilitate a thorough understanding of the examples. However, those skilled in the art will understand that the described aspects can be practiced without these specific details.

[0040] This disclosure provides a method for filtering remote sensing image data, which can be applied to the process of filtering remote sensing data in remote sensing image scenes.

[0041] See Figure 1 This is a flowchart illustrating a remote sensing image data filtering method provided in an embodiment of this disclosure. Figure 1 As shown, the method mainly includes the following steps:

[0042] Step 1: Acquire the remote sensing image data of the current time phase and the corresponding classified raster image of the region;

[0043] Information content, as a reference indicator reflecting the spatial information content of an image, is widely used in remote sensing image data screening, such as band selection for hyperspectral images and data screening for multi-source images. In these image screening applications, the spatial region of the image is treated as a fixed quantity, and the information content of images from different sources, different bands, or different resolutions is used as a reference indicator for screening, thereby selecting image data with high information content within the same region. However, in data screening tasks aimed at change detection, the main approach is to use the difference in change information between two time periods to screen image data from different regions. However, existing screening methods cannot calculate the content of change information when historical remote sensing image data is lacking. Therefore, this scheme uses classified raster images of the same region as substitute data, and calculates the magnitude of change information by establishing a data information content conversion relationship between remote sensing images and classified raster images, thereby guiding remote sensing image data screening. The overall idea and technical process are as follows: Figure 2 As shown.

[0044] In practice, remote sensing image data of the current time phase and classified raster images of the corresponding area can be acquired first to facilitate subsequent operations.

[0045] Step 2: Perform multidimensional spatial information measurement on the remote sensing image data and the classified raster image respectively, and extract the corresponding multidimensional spatial information components;

[0046] In practice, the information content of remote sensing images and classified raster images can be measured using existing multidimensional spatial information measurement methods to obtain the texture information, topological information, and spatial distribution information of remote sensing image data, and the geometric information, thematic information, topological information, and spatial distribution information of classified raster images.

[0047] Step 3: Perform principal component analysis on the multidimensional spatial information components of remote sensing image data and classified raster images respectively, remove redundant information and reduce dimensionality to obtain the principal components of remote sensing image information and the principal components of classified raster information.

[0048] In practice, due to spatial correlation, there is a strong correlation between the spatial information components of the two types of raster images. This scheme proposes to use Principal Component Analysis (PCA) to remove redundant information from the multidimensional spatial information components and to reduce the dimensionality of the spatial information components. Finally, the principal components of the remote sensing image and the classified raster image are used as input data and labels to train a machine learning model.

[0049] Principal Component Analysis (PCA), also known as the Karhunen-Loeve transform (KL transform), is a widely used linear transformation method in statistics. This method utilizes dimensionality reduction to concentrate the information contained in the original random variables into a few independent components of the transformed result, thereby eliminating redundant data and focusing the effective information. Typically, the composite index generated by the transformation is called the principal component, where each principal component is a linear combination of the original variables, and each principal component is independent of the others. This gives the principal components some superior performance compared to the original variables.

[0050] The multidimensional information components of the two types of raster data are considered as input data matrix X, with a dimension of... Where M and N are the number of samples and the number of features, respectively. According to algebraic principles, given certain conditions, X can be decomposed as follows:

[0051] (1)

[0052] in, and They are all orthogonal and uniform. Matrix, i.e. , ( (representing the identity matrix); yes Quasi-diagonal matrix, in In the following circumstances:

[0053] (2)

[0054] Without loss of generality, it is usually assumed that each It is a singular value. Let it be... , Then the above equation (1) can be written as:

[0055] (3)

[0056] In the formula, and These are called the left singular vector and the right singular vector, respectively, and their dimensions are respectively and Thus, singular value decomposition was performed on the original data X. Taking the covariance matrix of X (assuming the mean of each row vector is 0), we have:

[0057] (4)

[0058] In the formula:

[0059] = (5)

[0060] say These are the eigenvalues. Each column is called a feature vector. Equation (5-6) is the expression for... Principal component decomposition. The components derived from the principal component decomposition are sorted according to their energy magnitude. Therefore, there is Furthermore, if the rank of the original data is less than M (i.e., the M N-dimensional vectors are not linearly independent), some singular values ​​and eigenvalues ​​will be equal to 0. This allows us to reduce the dimensionality of the data while preserving the core information of the original data by choosing to retain most of the variance.

[0061] Step 4: Based on the principal components of the remote sensing image information at the current time, use the preset information transformation relationship model between the principal components of the remote sensing image information and the principal components of the classification raster information to predict the corresponding principal components of the classification raster information.

[0062] In specific implementation, this embodiment proposes a modeling approach for information conversion relationships driven by raster image data. The overall process is as follows: Figure 3 As shown.

[0063] To model the information conversion relationship between two types of raster image data, a sufficiently large training sample library can be constructed using an image classification dataset to train the information conversion relationship model using a data-driven approach. The accuracy and generalization ability of this conversion relationship model depend on the diversity and representativeness of the sample data. Therefore, a dataset containing rich samples and various scenarios allows the model to be exposed to a wider range of scenes and variations during training, thereby fully utilizing the features, distribution, and patterns of the data to establish the information conversion relationship between the two types of raster image data.

[0064] Multiple linear regression (MLP) is a classic method in machine learning, belonging to the regression problem within supervised learning. Originating in statistics, MLP is a method used to establish and analyze the relationship between a dependent variable and multiple independent variables. As part of feature engineering, it helps machine learning models better understand data structures. In multiple regression, a multiple linear model can be built, containing multiple independent variables to explain variations in the dependent variable. Assume that the information components of the current remote sensing image data and the classification raster information components have already undergone data standardization and PCA dimensionality reduction. The principal components of the remote sensing image information serve as the input features. Let the principal components of the classified raster information be the output Y after image classification. Then, we can perform multivariate regression fitting on the principal components of Y. Suppose we have m samples and n input features, the mathematical expression of the multivariate linear regression model is:

[0065] (6)

[0066] Here, Y is the output variable, which is the principal component of the categorical raster information; As the output variable, that is, the principal components of information in remote sensing images; These are the regression coefficients; This is the error term.

[0067] The least squares method is used to fit the regression coefficients, i.e., the objective function is to minimize the sum of the observed value Y and the predicted value. The residual sum of squares (RSS) between the two values ​​is used to calculate the regression coefficients. Solve for it.

[0068] (7)

[0069] (8)

[0070] Where X is a of size The matrix Y contains all the input variables and an intercept term that is all 1s; Y is a... The column vector contains all observed values ​​Y. In constructing the information transformation relationship, the principal components of remote sensing image data are used as input variables (features X), and the principal components of classified raster image data are used as output variables (labels Y). The sample data are divided into training and test sets at a ratio of 4:1, and then the training data is input into a multiple regression model for fitting training. The model is evaluated with minimizing the error as the objective function, and the fitting coefficients are solved using the least squares method. Finally, the functional transformation relationship between the principal components of the two types of raster data is obtained.

[0071] Step 5: Utilize the differences between the principal components of the classification raster information and the principal components of the baseline classification raster information, combined with the regional variation rate, to filter the remote sensing image data of the current time phase.

[0072] In practical implementation, to quantify the overall degree of change in the region, this scheme uses the Regional Rates of Change (RRC) as a quantitative indicator. RRC is defined as the proportion of pixels at the same location in land cover data from different periods within the same region whose category changes between periods. Let A be the Region of Interest (ROI) within a certain area, and P1 and P2 be the pixel attribute values ​​at the same location in the two images, respectively. Then, the formula for calculating the Regional Rates of Change (RRC) is as follows:

[0073] (9)

[0074] Where N is the total number of pixels within the region of interest A; and These are the category attribute values ​​of the i-th pixel in the pre- and post-temporal land cover data, respectively. Summation is performed by iterating through all pixels. Clearly, the RRC value ranges from 0 to 1, where 0 indicates no overall change in the area, and 1 indicates changes in all land features. Combining the image information difference value and the regional variation rate as two core quantification parameters, a mapping relationship between the two basic variables is constructed through machine learning. The overall process is as follows: Figure 4 As shown.

[0075] The remote sensing image data filtering method provided in this embodiment can filter remote sensing image data even when existing methods are not feasible due to a lack of historical remote sensing image data, by utilizing only a single period of remote sensing imagery. As an auxiliary means for change detection operations, image data filtering can identify image data rich in change information in advance, which can greatly improve the overall efficiency of change detection operations. At the same time, with the development of imaging technology and classification algorithms, it is expected that in the future, while ensuring production accuracy, the area of ​​the cell grid for data filtering will continue to shrink, and the accuracy of locating the area where change information is generated will continue to improve, thereby further enhancing the application value of this solution. From the perspective of information transmission, the existing two-period image change vector method essentially only utilizes the single spectral information content of remote sensing images, and cannot reflect the multi-dimensional spatial information and its changes in spatial relationships, geometric structures, and spatial distribution of remote sensing images. Therefore, change information is difficult to effectively transmit through pixel grayscale difference maps. This solution, based on the multi-dimensional information content of images, can not only effectively transmit various change information of ground objects, but also use PCA to reduce noise and redundancy in the original multi-dimensional information components, thereby significantly improving the image filtering accuracy.

[0076] The method of this disclosure will be further described below with reference to a specific embodiment, using Landsat satellite imagery. Using the GlobeLand 30 land cover data as a case study, this experiment analyzed remote sensing imagery data of four grid sizes (10km, 20km, 40km, and 80km). The results showed that the image selection accuracy of this method gradually improved with increasing dataset size. The overall accuracy (OA) for the four datasets was 0.74, 0.82, 0.90, and 0.92, respectively, and the Kappa coefficient increased from 0.47 to 0.81, significantly outperforming existing methods. Therefore, applying this method to remote sensing imagery data with a grid size of approximately 40km will yield optimal accuracy. The specific experimental procedure is as follows:

[0077] Among the numerous publicly available remote sensing images and land cover data products online, Landsat remote sensing images with 30m spatial resolution provide rich texture details and spatial structure information, effectively reflecting or depicting most human land use activities and the resulting landscape patterns. This makes it the optimal scale for macroscopic and mesoscopic descriptions of global land cover. Maintaining the timeliness of 30m land cover products and achieving more accurate and efficient updates is a pressing issue for the industry. Therefore, this solution utilizes two phases of GlobeLand 30 (GLC-30) land cover products from 2010 and 2020, along with six Landsat satellite images, to create training sample data. The 2010 satellite imagery comes from Landsat 5 B1-B5 (5-band multispectral imagery), and the 2020 satellite imagery comes from Landsat 8 B2-B7 (6-band multispectral imagery). The stitched land cover data products from both phases and the six remote sensing images are all located in the central provinces of my country.

[0078] After downloading the two phases of image data and land cover products, necessary data preprocessing operations were performed on the raw image data, and standard datasets of different sizes were created. First, after performing preprocessing operations such as geometric correction, registration, fusion, color balancing, and cropping on the two phases of Landsat remote sensing images, according to Scheme 1, cell grids of different scales were directly deployed on the raw remote sensing images and GLC30 data. Then, the effective image data from the six images was obtained by segmenting according to the size of the cell grid. This scheme designed four scale segmentation levels, with the cell grid sizes at each level as follows: , , as well as At a spatial resolution of 30m, each Within the area, approximately Each pixel. It's important to note that during the creation of the sample database, if a segmented sub-block of remote sensing imagery contains background pixels, that sub-block and its corresponding land cover data will be discarded. Ultimately, two temporal remote sensing image-land cover sample datasets were obtained for four map sheet sizes. Furthermore, to verify the superiority of Scheme 2 during implementation, a multi-level segmentation using a tile quadtree structure was performed. A sub-block with an area of ​​[missing information] was randomly selected from the six remote sensing image datasets. The selected sub-image regions were then used as the maximum scale for the quadtree structure hierarchies, and layer-by-layer segmentation was performed to generate multi-level sub-block images and GLC30 data with 40km, 20km, and 10km square cell networks. Since the multi-level segmentation of this quadtree structure has the same sample data size as each level in Scheme 1, the sample datasets produced by the two methods can be merged, which not only expands the number of basic sample databases but also establishes links between some sample image data through the quadtree structure.

[0079] Screening accuracy analysis:

[0080] (1) Accuracy indicators

[0081] Unlike change detection or classification tasks that require incremental extraction and updating of specific land cover changes, this remote sensing image data screening method primarily uses image information difference values ​​as the basic variable. The image variation rate (RRC) of two land cover data periods is defined as a quantitative indicator of the overall degree of change. Then, the RRC is predicted using the image information difference value, serving as the core parameter for change cues. Therefore, the image data screening process guided by spatial information difference values ​​is essentially a mathematical numerical prediction process. The image screening accuracy guided by this method is determined by whether the degree of change in the screened images matches the actual results. Furthermore, considering practical application scenarios, this scheme divides the image's RRC into three gradient intervals based on its numerical value. Each interval represents the level of overall change in the region, namely, the significant change area (III), the moderate change area (II), and the slight change area (I) described in the text. Based on the distribution of RRC values ​​in different samples, this scheme uses 0.2 and 0.5 as thresholds for dividing the gradient intervals. If the RRC value of land cover under the region of interest is in the range of [0, 0.2), then the overall degree of change in the region represented by the image is assessed as slight (I); when the RRC value calculated under the image region is in [0.2, 0.5) and [0.5, 1.0], then the overall degree of change in the region will be assessed as moderate (II) and significant (III), respectively.

[0082] After dividing the overall degree of change into intervals using RRC quantification values, the task transformation of "information difference calculation—RRC numerical prediction—RRC category prediction—remote sensing image data screening" was achieved, which is beneficial for the output of image data screening results and accuracy evaluation. Therefore, the accuracy evaluation indicators for image screening can refer to the accuracy evaluation of classification tasks. In this type of task, common accuracy evaluation methods and indicators include confusion matrix, ROC curve, intersection-over-union ratio, etc. The confusion matrix is ​​a widely used accuracy evaluation method. Indicators such as overall classification accuracy, misclassification error, omission error, and Kappa coefficient can all be calculated using the confusion matrix. In this experiment, the following accuracy indicators were used: Intersection over Union (IoU), Overall Accuracy (OA), Precision (Pr), Recall (Rc), F1 score (F1Score), and Kappa coefficient (Kappa). Since recall and precision are mutually constrained, the F1 score is used to comprehensively evaluate the accuracy of RRC classification. The calculation formulas for the above accuracy indicators are as follows:

[0083] (10)

[0084] (11)

[0085] (12)

[0086] (13)

[0087] (14)

[0088] (15)

[0089] Where TP, TN, FP, and FN represent true positive, true negative, false positive, and false negative in the prediction task, respectively. It represents the sum of correctly classified samples in each class divided by the total number of samples (i.e., OA). The calculation formula is as follows:

[0090] (16)

[0091] in and These represent the number of samples that have changed and those that have not changed in the actual change label of the region, respectively; and These represent the number of samples that changed and the number that remained unchanged in the prediction results, respectively.

[0092] (2) Results Analysis

[0093] Using the established information transformation model, remote sensing image data at four cell sizes were filtered. The confusion matrix and classification accuracy results are shown in Tables 1 and 2. The actual degree of land surface change was graded and evaluated using the RRC calculated from the two periods of land cover data; the predicted degree of land surface change was graded and evaluated using the RRC estimated by the principal component difference (I-PCA1) of the two periods of image information. Finally, the accuracy between the prediction results and the true values ​​was analyzed, and confusion matrices for datasets of different sizes were constructed for accuracy evaluation.

[0094] Table 1

[0095]

[0096] Table 2

[0097]

[0098] As shown in Table 1, a large number of samples still exhibit classification errors when using image information content change values ​​to predict the overall degree of change in a region, especially in test samples from small areas. Overall, 1) image samples showing slight and moderate changes are more easily confused. For example, in image regions with small areas, a large number of image samples that actually show slight changes (I) are missed, with most being misclassified as moderate changes (II) and a small portion as significant changes (III). This results in the number of slightly changed image samples predicted by I-PCA1 being less than the number of samples evaluated by the actual RRC; 2) image samples with moderate changes (II) are more likely to be falsely detected, with many image samples with significant changes (III) being misclassified as moderate changes (II), resulting in a higher number of detected samples in this category than the actual number; 3) as the image size increases, the types of land cover contained in the region of interest gradually become richer, which is beneficial to improving the accuracy of the image information content conversion model. Therefore, the RRC predicted based on I-PCA1 gradually approaches the actual RRC value, significantly reducing missed detections and misclassifications, demonstrating that this scheme is effective in large-scale (… In the process of discovering clues of change in a region, better and more robust data filtering accuracy can be achieved.

[0099] Table 2 summarizes the accuracy evaluation results of existing methods (variable vector method) and our proposed method on four image datasets. It can be seen that our proposed method significantly outperforms existing methods. Overall, existing methods have an average accuracy of only 0.35 for some important accuracy metrics, such as OA, and Kappa coefficients are all less than 0. The image selection accuracy of our proposed method gradually increases with the increase of dataset size. The OA values ​​for the four datasets are 0.74, 0.82, 0.90, and 0.92, respectively, and the Kappa coefficient increases from 0.47 to 0.81. Therefore, applying our proposed method to remote sensing image data selection with a grid range of approximately 40 km will yield optimal accuracy.

[0100] To quickly identify the approximate range of changes occurring over a large area in real-world applications, the experiment further utilized... The dataset undergoes multi-level evaluation using a tile quadtree structure, and then cell sub-image data with significant changes are selected based on the degree of change in the image data. Furthermore, slightly invariant regions can be drawn to form overlay masks, providing a convenient product service for subsequent change detection operations and manual visual interpretation.

[0101] After adopting a multi-level data filtering scheme with a quadtree structure, the image filtering results of four sample data at multiple scales are presented intuitively. Simultaneously, the sub-regions generating change information are gradually located, ultimately reducing the working area by 79.7%, 64.1%, 85.9%, and 56.3%, respectively. Each row in the figure represents an original... The experimental samples are presented in columns: land cover at time T1 (baseline data), multi-level image change prediction results under a quadtree structure, and image filtering results after masking areas with minor changes. It can be seen that performing a step-by-step image filtering operation on the original image not only assigns level labels to the predicted change levels of each cell network region at different levels and filters out sub-image data generated by localized changes missed at large scales, but also allows for precise positioning of the approximate locations of change information at each level. Ultimately, this scheme effectively discovers and filters out sub-regions generating a large amount of change information, and by masking areas with minor changes, the filtered image data significantly reduces the area required for subsequent change detection.

[0102] It should be understood that the various parts of this disclosure can be implemented in hardware, software, firmware, or a combination thereof.

[0103] The above description is merely a specific embodiment of this disclosure, but the scope of protection of this disclosure is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this disclosure should be included within the scope of protection of this disclosure. Therefore, the scope of protection of this disclosure should be determined by the scope of the claims.

Claims

1. A method for filtering remote sensing image data, characterized in that, include: Step 1: Acquire the remote sensing image data of the current time phase and the corresponding classified raster image of the region; Step 2: Perform multidimensional spatial information measurement on the remote sensing image data and the classified raster image respectively, and extract the corresponding multidimensional spatial information components; Step 3: Perform principal component analysis on the multidimensional spatial information components of remote sensing image data and classified raster images respectively, remove redundant information and reduce dimensionality to obtain the principal components of remote sensing image information and the principal components of classified raster information. Step 4: Based on the principal components of the remote sensing image information at the current time, use the preset information transformation relationship model between the principal components of the remote sensing image information and the principal components of the classification raster information to predict the corresponding principal components of the classification raster information. Step 5: Utilize the differences between the principal components of the classification raster information and the principal components of the baseline classification raster information, combined with the regional variation rate, to filter the remote sensing image data of the current time phase.

2. The method according to claim 1, characterized in that, The multidimensional spatial information components include texture information, topological information, and spatial distribution information corresponding to remote sensing image data, as well as geometric information, thematic information, topological information, and spatial distribution information corresponding to classified raster images.

3. The method according to claim 1, characterized in that, Before step 4, the method further includes: Acquire multiple sets of training samples, each set including remote sensing image data and classified raster images of the same area; Multidimensional spatial information components of remote sensing image data and classified raster images in each training sample are extracted, and principal component analysis is performed to reduce dimensionality. Using the principal components of the dimensionality-reduced remote sensing image information as input features and the principal components of the dimensionality-reduced classified raster information as output labels, a multiple linear regression model is used for training to obtain an information transformation relationship model.

4. The method according to claim 3, characterized in that, The expression for the multiple linear regression model is: Here, Y is the output variable, which is the principal component of the categorical raster information. As the output variable, namely the principal components of information from remotely sensed images, For regression coefficients, This is the error term.

5. The method according to claim 1, characterized in that, The expression for the regional variation rate is: Where N is the total number of pixels within the region of interest A; and These are the category attribute values ​​of the i-th pixel in the previous and subsequent time-phase land cover data, respectively. Summation is performed by iterating through all pixels.

6. The method according to claim 5, characterized in that, Step 5 specifically includes: Based on the magnitude of the regional variation rate, the image regions are divided into three levels: slight variation, moderate variation, and significant variation. Remote sensing image data are filtered based on different levels, with priority given to image areas showing significant and moderate changes.

7. The method according to claim 1, characterized in that, Following step 1, the method further includes: Preprocessing is performed on remote sensing image data and classified raster images, wherein the preprocessing includes one or more of geometric correction, registration, fusion, color balancing and cropping.

8. The method according to claim 1, characterized in that, The remote sensing image data is Landsat series satellite imagery, and the classified raster image is GlobeLand30 land cover data product.