Hyperspectral lithology intelligent identification method based on fuzzy clustering

By applying a fuzzy clustering algorithm in hyperspectral remote sensing data processing and combining the principle of neighborhood similarity, unsupervised intelligent lithology recognition is achieved, solving the problem of large workload and relying on a large number of training samples in traditional methods, and improving the recognition accuracy and the ability to identify mineralization-related geological factors.

CN119915746APending Publication Date: 2025-05-02BEIJING RES INST OF URANIUM GEOLOGY
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202411946851.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-27
Publication Date
2025-05-02

AI Technical Summary

Technical Problem

The existing remote sensing lithologic recognition methods have problems such as high workload and high requirements for personnel and prior knowledge, and the accuracy of the supervised classification method based on machine learning depends on a large number of training samples.

Method used

The intelligent identification method of hyperspectral lithology based on fuzzy clustering is adopted, and the hyperspectral remote sensing data is combined with the principle of neighborhood similarity, and the intelligent identification of lithology is achieved through unsupervised classification methods to reduce the influence of human interference factors.

Benefits of technology

It improves lithology recognition accuracy and mineralization related geological factors recognition ability, reduces artificial interference, and reduces dependence on sample number.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119915746A_ABST
    Figure CN119915746A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of remote sensing image processing, and particularly relates to a hyperspectral lithology intelligent identification method based on fuzzy clustering, which comprises the following steps: acquiring a geological map of a research area, vectorizing the geological map of the research area, and acquiring initial lithology distribution vector data; preprocessing the hyperspectral data; performing data dimension reduction on the preprocessed hyperspectral data by using principal component analysis transformation to obtain dimension-reduced data; clustering the dimension-reduced data by adopting a spatial fuzzy C-means clustering algorithm, and segmenting a clustering result based on a neighborhood similarity criterion to obtain the clustering result; and according to the initial lithology distribution vector data, determining the lithology category with the maximum proportion in different clustered patches from the clustering result, and taking the category as the lithology category corresponding to the clustered patches to obtain an intelligent lithology identification result. According to the method, the influence of human interference factors can be effectively reduced, and the lithology identification precision and the mineralization-related geological element identification capability are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention belongs to the technical field of remote sensing image processing, and in particular relates to a hyperspectral lithology intelligent recognition method based on fuzzy clustering. Background Art

[0002] Lithology refers to the comprehensive expression of the material composition, structural characteristics, and physical and chemical properties of rocks. The difference in lithology directly reflects the changes in geological history and geological conditions. Through the study of lithology, we can better understand the earth and its complex natural processes and provide support for the sustainable development of mankind. Lithology identification is the process of determining the type and nature of underground rock formations by analyzing the physical and chemical characteristics of geological materials. The identification and classification of lithology can reveal the characteristics of underground structures, determine the paleogeographic environment, etc., and provide accurate data and scientific basis for exploration work. There are many ways to identify lithology, such as remote sensing technology, data analysis, rock sample analysis, etc.

[0003] In the related research on lithology identification, the combination of lithology identification methods and remote sensing technology is one of the important research contents. Remote sensing has the characteristics of large detection range, large amount of information, fast information acquisition speed and short update cycle. It plays an important role in the field of earth science, especially in large-area regional geological survey and research. Among them, the sensors used in hyperspectral remote sensing can capture multiple bands of data from visible light to near infrared and even short-wave infrared. This multi-band information acquisition capability enables hyperspectral remote sensing to provide rich spectral features, which is convenient for the classification and identification of rocks and minerals. Based on the unique spectral characteristics of different types of rocks and minerals, hyperspectral remote sensing can identify and distinguish different lithologies by analyzing the reflectance spectrum. With the advancement of data processing technology, combined with machine learning algorithms, the data analysis of hyperspectral remote sensing has been automated, which can effectively identify and classify a large amount of lithology information and improve the efficiency and accuracy of identification.

[0004] Based on the current research status at home and abroad, the methods currently used are mainly traditional remote sensing lithology identification and classification and supervised classification based on machine learning. However, the existing technologies have the following problems: the workload of traditional visual interpretation is large, and the requirements for staff and prior knowledge are high, while the accuracy of supervised classification based on machine learning depends on a sufficient number of training samples to build the model. Summary of the invention

[0005] The purpose of the present invention is to provide a hyperspectral lithology intelligent identification method based on fuzzy clustering. The method starts from the fuzzy clustering principle, is based on hyperspectral remote sensing data, combines the neighborhood similarity principle, and uses an unsupervised classification method to realize the intelligent identification of lithology. It can effectively reduce the influence of human interference factors and improve the accuracy of lithology identification and the ability to identify mineralization-related geological elements.

[0006] The technical solution to achieve the purpose of the present invention is:

[0007] A hyperspectral lithology intelligent identification method based on fuzzy clustering, the method comprising:

[0008] Step 1: Obtain the geological map of the study area, then vectorize the geological map of the study area to obtain the initial lithology distribution vector data;

[0009] Step 2: Preprocess the hyperspectral data;

[0010] Step 3: Use principal component analysis to reduce the dimension of the preprocessed hyperspectral data to obtain the reduced-dimensional data;

[0011] Step 4: Use the spatial fuzzy C-means clustering algorithm to cluster the reduced-dimensional data and segment the clustering results based on the neighborhood similarity criterion to obtain the clustering results;

[0012] Step 5: According to the initial lithology distribution vector data, determine the lithology category with the largest proportion in different cluster patches from the clustering results, use this category as the lithology category corresponding to the cluster patches, and obtain the lithology intelligent identification result.

[0013] The step 1 specifically includes: assigning corresponding attributes to each polygon according to the lithology information on the geological map, determining the initial lithology distribution of the study area, and obtaining initial lithology distribution vector data.

[0014] The preprocessing in step 2 includes atmospheric correction and geometric alignment.

[0015] The step 3 comprises:

[0016] Step 3.1: Center the preprocessed hyperspectral data and calculate the covariance matrix;

[0017] Step 3.2: Use the orthogonal decomposition method to perform eigenvalue decomposition on the covariance matrix C to obtain the eigenvalues ​​and corresponding eigenvectors;

[0018] Step 3.3: The obtained eigenvectors are combined into a projection matrix to reduce the dimension of the centralized data set to obtain the reduced-dimensional data;

[0019] Step 3.4: Save the reduced dimension data in ENVI standard format.

[0020] In step 3.1, the pre-processed hyperspectral data is centralized, and the calculation formula is:

[0021]

[0022] Among them, X is the preprocessed hyperspectral dataset, X iis the preprocessed hyperspectral data of each sample in each dimension, n is the total number of samples, and X1 is the centralized dataset.

[0023] The formula for calculating the covariance matrix in step 3.1 is:

[0024]

[0025] Among them, A is the covariance matrix, X1 T is the transpose of X1.

[0026] The calculation formula for reducing the dimension of the centralized data set in step 3.3 is:

[0027] Y=W*X1

[0028] W=[e1,e2,...,e k ]

[0029] Among them, Y is the data after dimensionality reduction, W is the projection matrix, X1 is the centralized data set, and e i is the i-th eigenvector, and k is the number of eigenvectors.

[0030] The step 4 comprises:

[0031] Step 4.1: Based on the dimension-reduced data obtained in step 3, set the number of clusters, fuzzy factor, maximum number of iterations, and band number. The number of initial cluster centers can be set based on the number of assigned attributes of the lithology distribution vector data in step 1. The value of the fuzzy factor is generally 1 to 3.

[0032] Step 4.2: Randomly select a sample data point as the initial cluster center, calculate the Euclidean distance from the remaining samples to the cluster center, and update the membership of the sample to the cluster center based on the calculated Euclidean distance;

[0033] Step 4.3: Based on the maximum number of iterations set in step 4.1, iterative update is performed to obtain the final membership matrix, which is used to determine the clustering affiliation of the samples and complete the preliminary segmentation of the samples, so that similar samples can be more accurately classified into one category;

[0034] Step 4.4: Define the neighborhood of each point by K nearest neighbor method;

[0035] Step 4.5: For each point in the cluster, use the Euclidean distance as an indicator to calculate the similarity measure between it and other points in the neighborhood;

[0036] Step 4.6: Based on the principle of proximity, the similarity between the pixels and the neighborhood in the cluster image obtained in step 4.5 is used for re-segmentation, and the extracted pixels are assigned to the nearest clusters to ensure that the correctly classified pixels are assigned to the clusters and the incorrectly assigned pixels are assigned to the most appropriate clusters. The final segmentation results are output to obtain the clustering results.

[0037] In step 4.3, the final membership matrix is ​​obtained to determine the clustering of the sample: for a set of data patterns x = {x1, x2, ..., x N}, the fuzzy C-means clustering algorithm calculates the cluster center c i And the membership matrix U, and minimize the objective function J about the cluster center and membership to divide the data space. The calculation formula of the objective function J is as follows:

[0038]

[0039] Among them, u ij is the sample x j For cluster center c i The membership degree, d(x j ,c i ) is the calculation element x j To the cluster center c i The distance metric, C is the number of cluster centers, N is the number of samples, and m is the fuzzy factor.

[0040] The step 5 comprises:

[0041] Step 5.1: Extract the spatial region of each cluster from the clustering results. The spatial region includes the position, area and corresponding pixels or sample points;

[0042] Step 5.2: Perform coordinate transformation on the initial lithology distribution vector data to ensure that the coordinate systems of the clustered spatial region data and the initial lithology distribution vector data are consistent, and perform spatial intersection operation on the spatial region of each cluster with the initial lithology distribution vector data to obtain the lithology category in the region;

[0043] Step 5.3: Count all the intersecting lithology categories in the spatial area of ​​each cluster, and based on the statistical results, select the lithology category with the highest frequency as the representative lithology of the cluster.

[0044] Step 5.4: Based on the comparison between the identification results and the existing lithology annotation data, the input parameters of the clustering algorithm are adjusted to optimize the clustering effect and the accuracy of lithology distribution.

[0045] The beneficial technical effects of the present invention are:

[0046] 1. Compared with the prior art, the hyperspectral lithology intelligent identification method based on fuzzy clustering provided by the present invention is an unsupervised classification method based on machine learning. Compared with the previous supervised classification method which is easily limited by the size of the study area and the number of samples, the unsupervised classification method has less human intervention and is not affected by the number of samples, thus realizing the intelligent identification of lithology in a true sense.

[0047] 2. In a hyperspectral lithology intelligent identification method based on fuzzy clustering provided by the present invention, the spatial fuzzy C-means clustering algorithm (SFCM) is a clustering method that combines fuzzy C-means (FCM) clustering with spatial information. This method not only focuses on the distance between the data point and the cluster center when dividing the lithology patches, but also considers the spatial relationship between the data points. The introduction of this spatial information makes it easier for similar samples to be classified into one category, while flexibly processing data points with fuzzy boundaries, and showing higher clustering accuracy when processing data with obvious spatial structures. At the same time, the setting of the fuzzy factor can significantly increase the flexibility of the clustering results, improve the accuracy of clustering, and help the model to maintain good clustering performance after dimensionality reduction. Finally, the initial lithology distribution vector data is matched with the cluster segmentation results to more accurately obtain the lithology category of each cluster patch, providing a reference for the intelligent identification of lithology. BRIEF DESCRIPTION OF THE DRAWINGS

[0048] Figure 1 This is a flow chart of a hyperspectral lithology intelligent identification method based on fuzzy clustering provided by the present invention. DETAILED DESCRIPTION

[0049] The present invention is further described in detail below with reference to the accompanying drawings and embodiments.

[0050] like Figure 1 As shown, the present invention provides a hyperspectral lithology intelligent identification method based on fuzzy clustering, which specifically includes the following steps:

[0051] Step 1: Obtain the geological map of the study area, then vectorize the geological map of the study area to obtain the initial lithology distribution vector data;

[0052] According to the lithology information on the geological map, each polygon is assigned corresponding attributes, the initial lithology distribution of the study area is determined, and the initial lithology distribution vector data is obtained.

[0053] Step 2: Preprocess the hyperspectral data;

[0054] Preprocessing includes atmospheric correction and geometric alignment.

[0055] Step 3: Use principal component analysis to reduce the dimension of the preprocessed hyperspectral data, aiming to extract features, reduce the amount of data, and obtain the reduced-dimensional data;

[0056] Step 3.1: Center the preprocessed hyperspectral data and calculate the covariance matrix;

[0057] First, the data is centralized, that is, the mean value of each dimension is subtracted from the data of that dimension to obtain the centralized data, as shown in formula (1);

[0058]

[0059] Among them, X is the preprocessed hyperspectral dataset, X i is the preprocessed hyperspectral data of each sample in each dimension, n is the total number of samples, and X1 is the centralized dataset.

[0060] Then calculate the covariance matrix as formula (2);

[0061]

[0062] Among them, A is the covariance matrix, X1 T is the transpose of X1.

[0063] Step 3.2: Use the orthogonal decomposition method to perform eigenvalue decomposition on the covariance matrix A to obtain the eigenvalues ​​and corresponding eigenvectors;

[0064] Step 3.3: The obtained eigenvectors are combined into a projection matrix to reduce the dimension of the centralized data set to obtain the reduced-dimensional data;

[0065] The selected k eigenvectors are combined into a projection matrix W = [e1, e2, ..., e k ], where e i Represents the i-th eigenvector. Perform dimensionality reduction on the centralized data set, Y=W*X1, where Y is the data after dimensionality reduction.

[0066] Step 3.4: Save the reduced dimension data in ENVI standard format.

[0067] Step 4: Use the Spatial Fuzzy C-Means (SFCM) algorithm to cluster the reduced-dimensional data and segment the clustering results based on the neighborhood similarity criterion to obtain the clustering results;

[0068] Step 4.1: Based on the dimension-reduced data obtained in step 3, set the number of clusters, fuzzy factor, maximum number of iterations, and band number. The number of initial cluster centers can be set based on the number of assigned attributes of the lithology distribution vector data in step 1. The value of the fuzzy factor is generally 1 to 3.

[0069] Step 4.2: Randomly select a sample data point as the initial cluster center, calculate the Euclidean distance from the remaining samples to the cluster center, and update the membership of the sample to the cluster center based on the calculated Euclidean distance;

[0070] Step 4.3: Based on the maximum number of iterations set in step 4.1, the final membership matrix is ​​iteratively updated to determine the clustering of the sample, and the sample is assigned to the cluster with the largest membership, completing the preliminary segmentation of the sample so that similar samples can be more accurately classified into one category;

[0071] The principle of obtaining the final membership matrix to determine the clustering of samples is as follows: for a set of data patterns x = {x1, x2, ..., x N}, the fuzzy C-means clustering algorithm calculates the cluster center (c i ) and the membership matrix (U), and minimize the objective function J about the cluster center and the membership to divide the data space, as shown in formula (3):

[0072]

[0073] Among them, u ij is the sample x j For cluster center c i The membership degree, d(x j ,c i ) is the calculation element x j To the cluster center c i The distance metric, C is the number of cluster centers, N is the number of samples, and m is the fuzzy factor.

[0074] Step 4.4: Define the neighborhood of each point by K nearest neighbor method;

[0075] Step 4.5: For each point in the cluster, use the Euclidean distance as an indicator to calculate the similarity measure between it and other points in the neighborhood;

[0076] Step 4.6: Based on the principle of proximity, the similarity between the pixels and the neighborhood in the cluster image obtained in step 4.5 is used for re-segmentation, and the extracted pixels are assigned to the nearest clusters to ensure that the correctly classified pixels are assigned to the clusters and the incorrectly assigned pixels are assigned to the most appropriate clusters. The final segmentation results are output to obtain the clustering results.

[0077] Step 5: According to the initial lithology distribution vector data, determine the lithology category with the largest proportion in different cluster patches from the clustering results, use this category as the lithology category corresponding to the cluster patches, and obtain the lithology intelligent identification result;

[0078] Step 5.1: Extract the spatial region of each cluster from the clustering results. The spatial region includes the position, area and corresponding pixels or sample points;

[0079] Step 5.2: Perform coordinate transformation on the initial lithology distribution vector data to ensure that the coordinate systems of the clustered spatial region data and the initial lithology distribution vector data are consistent, and perform spatial intersection operation on the spatial region of each cluster with the initial lithology distribution vector data to obtain the lithology category in the region;

[0080] Step 5.3: Count all the intersecting lithology categories in the spatial area of ​​each cluster, and based on the statistical results, select the lithology category with the highest frequency as the representative lithology of the cluster.

[0081] Step 5.4: According to the comparison between the recognition results and the existing lithology annotation data, the input parameters of the clustering algorithm are adjusted, such as increasing the number of iterations and adjusting the fuzzy factor to optimize the clustering effect and the accuracy of lithology distribution.

[0082] The present invention is described in detail above with reference to the accompanying drawings and embodiments, but the present invention is not limited to the above embodiments, and various changes can be made within the knowledge of ordinary technicians in the field without departing from the purpose of the present invention. The contents not described in detail in the present invention can adopt the existing technology.

Claims

1. A hyperspectral lithology intelligent identification method based on fuzzy clustering, characterized in that: The method comprises: Step 1: Obtain the geological map of the study area, then vectorize the geological map of the study area to obtain the initial lithology distribution vector data; Step 2: Preprocess the hyperspectral data; Step 3: Use principal component analysis to reduce the dimension of the preprocessed hyperspectral data to obtain the reduced-dimensional data; Step 4: Use the spatial fuzzy C-means clustering algorithm to cluster the reduced-dimensional data and segment the clustering results based on the neighborhood similarity criterion to obtain the clustering results; Step 5: According to the initial lithology distribution vector data, determine the lithology category with the largest proportion in different cluster patches from the clustering results, use this category as the lithology category corresponding to the cluster patches, and obtain the lithology intelligent identification result.

2. The hyperspectral lithology intelligent identification method based on fuzzy clustering according to claim 1 is characterized in that: The step 1 specifically includes: assigning corresponding attributes to each polygon according to the lithology information on the geological map, determining the initial lithology distribution of the study area, and obtaining initial lithology distribution vector data.

3. The hyperspectral lithology intelligent identification method based on fuzzy clustering according to claim 1 is characterized in that: The preprocessing in step 2 includes atmospheric correction and geometric alignment.

4. The hyperspectral lithology intelligent identification method based on fuzzy clustering according to claim 1 is characterized in that: The step 3 comprises: Step 3.1: Center the preprocessed hyperspectral data and calculate the covariance matrix; Step 3.2: Use the orthogonal decomposition method to perform eigenvalue decomposition on the covariance matrix C to obtain the eigenvalues ​​and corresponding eigenvectors; Step 3.3: The obtained eigenvectors are combined into a projection matrix to reduce the dimension of the centralized data set to obtain the reduced-dimensional data; Step 3.4: Save the reduced dimension data in ENVI standard format.

5. The hyperspectral lithology intelligent identification method based on fuzzy clustering according to claim 4 is characterized in that: In step 3.1, the pre-processed hyperspectral data is centralized, and the calculation formula is: Among them, X is the preprocessed hyperspectral dataset, X i is the preprocessed hyperspectral data of each sample in each dimension, n is the total number of samples, and X1 is the centralized dataset.

6. The hyperspectral lithology intelligent identification method based on fuzzy clustering according to claim 5 is characterized in that: The formula for calculating the covariance matrix in step 3.1 is: Among them, A is the covariance matrix, X1 T is the transpose of X1.

7. The hyperspectral lithology intelligent identification method based on fuzzy clustering according to claim 5 is characterized in that: The calculation formula for reducing the dimension of the centralized data set in step 3.3 is: Y=W*X1 W=[e1,e2,...,e k ] Among them, Y is the data after dimensionality reduction, W is the projection matrix, X1 is the centralized data set, and e i is the i-th eigenvector, and k is the number of eigenvectors.

8. The hyperspectral lithology intelligent identification method based on fuzzy clustering according to claim 1 is characterized in that: The step 4 comprises: Step 4.1: Based on the dimension-reduced data obtained in step 3, set the number of clusters, fuzzy factor, maximum number of iterations, and band number. The number of initial cluster centers can be set based on the number of assigned attributes of the lithology distribution vector data in step 1. The value of the fuzzy factor is generally 1 to 3. Step 4.2: Randomly select a sample data point as the initial cluster center, calculate the Euclidean distance from the remaining samples to the cluster center, and update the membership of the sample to the cluster center based on the calculated Euclidean distance; Step 4.3: Based on the maximum number of iterations set in step 4.1, iterative update is performed to obtain the final membership matrix, which is used to determine the clustering affiliation of the samples and complete the preliminary segmentation of the samples, so that similar samples can be more accurately classified into one category; Step 4.4: Define the neighborhood of each point by K nearest neighbor method; Step 4.5: For each point in the cluster, use the Euclidean distance as an indicator to calculate the similarity measure between it and other points in the neighborhood; Step 4.6: Based on the principle of proximity, the similarity between the pixels and the neighborhood in the cluster image obtained in step 4.5 is used for re-segmentation, and the extracted pixels are assigned to the nearest clusters to ensure that the correctly classified pixels are assigned to the clusters and the incorrectly assigned pixels are assigned to the most appropriate clusters. The final segmentation results are output to obtain the clustering results.

9. The hyperspectral lithology intelligent identification method based on fuzzy clustering according to claim 8 is characterized in that: In step 4.3, the final membership matrix is ​​obtained to determine the clustering of the sample: for a set of data patterns x = {x1, x2, ..., x N }, the fuzzy C-means clustering algorithm calculates the cluster center c i And the membership matrix U, and minimize the objective function J about the cluster center and membership to divide the data space. The calculation formula of the objective function J is as follows: Among them, u ij is the sample x j For cluster center c i The membership degree, d(x j ,c i ) is the calculation element x j To the cluster center c i The distance metric, C is the number of cluster centers, N is the number of samples, and m is the fuzzy factor.

10. The hyperspectral lithology intelligent identification method based on fuzzy clustering according to claim 1 is characterized in that: The step 5 comprises: Step 5.1: Extract the spatial region of each cluster from the clustering results. The spatial region includes the position, area and corresponding pixels or sample points; Step 5.2: Perform coordinate transformation on the initial lithology distribution vector data to ensure that the coordinate systems of the clustered spatial region data and the initial lithology distribution vector data are consistent, and perform spatial intersection operation on the spatial region of each cluster with the initial lithology distribution vector data to obtain the lithology category in the region; Step 5.3: Count all the intersecting lithology categories in the spatial area of ​​each cluster, and based on the statistical results, select the lithology category with the highest frequency as the representative lithology of the cluster. Step 5.4: Based on the comparison between the identification results and the existing lithology annotation data, the input parameters of the clustering algorithm are adjusted to optimize the clustering effect and the accuracy of lithology distribution.

Citation Information

Cited By

  • Mining area rock and ore classification method and system based on thermal infrared hyperspectral remote sensing technology

    CN120808140A

  • Chlorogenic acid component detection method fused with deep learning

    CN121476086A