Fast wave band screening and dimension reduction method and system for hyperspectral data cube

By performing correlation analysis and clustering on hyperspectral data cubes, the optimal representative bands are selected, solving the balance problem between efficiency and feature fidelity in existing band selection and dimensionality reduction methods, and achieving efficient and reliable data processing and mineral identification.

CN122087402APending Publication Date: 2026-05-26AERIAL PHOTOGRAMMETRY & REMOTE SENSING CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
AERIAL PHOTOGRAMMETRY & REMOTE SENSING CO LTD
Filing Date
2026-04-22
Publication Date
2026-05-26

AI Technical Summary

Technical Problem

Existing methods for band selection and dimensionality reduction of hyperspectral data cubes struggle to balance data processing efficiency with the fidelity of effective features, impacting the reliability and efficiency of subsequent analysis.

Method used

By performing correlation analysis on the spectral dimensions of the hyperspectral data cube, a correlation matrix between bands is generated. Based on the correlation matrix, clustering is performed to select the optimal representative bands of each band cluster. Duplicate features are eliminated through cross-cluster feature complementarity evaluation, and finally, a dimensionality-reduced hyperspectral data subset is formed.

Benefits of technology

While achieving computational lightweighting, it ensures high fidelity in geological features of the dimensionality-reduced dataset, improves data processing efficiency, and enhances the reliability of subsequent mineral identification and information extraction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122087402A_ABST
    Figure CN122087402A_ABST
Patent Text Reader

Abstract

The invention discloses a rapid wave band screening and dimension reduction method and system for a hyperspectral data cube, particularly relates to the technical field of data processing, and is used for solving the problem that in the prior art, processing efficiency and effective feature fidelity are difficult to consider in the hyperspectral data dimension reduction process. The method comprises the following steps: acquiring a hyperspectral data cube to be processed, performing spectral dimension correlation analysis on the hyperspectral data cube to generate an inter-band correlation matrix, and clustering all bands based on the matrix to obtain a plurality of band clusters; then, the response matching degree of wave bands in each wave band cluster is calculated according to the typical spectral response interval of the target geological object, the optimal representative wave band of each cluster is screened out, and then the cross-cluster characteristic complementarity of the optimal representative wave bands is evaluated to eliminate characteristic repeated wave bands; and finally, integrating the screened optimal representative wave band set to form a dimension-reduced hyperspectral data subset. The data processing efficiency can be ensured, and meanwhile, key spectral characteristics with indicating significance on a geological exploration target can be effectively reserved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of data processing technology, and more specifically, to a method and system for rapid band selection and dimensionality reduction of hyperspectral data cubes. Background Technology

[0002] Hyperspectral data cubes integrate massive amounts of continuous band information in both spatial and spectral dimensions, possessing significant application potential in remote sensing geological exploration, mineral identification, and prediction of prospective metallogenic areas. Band selection and dimensionality reduction are core steps in hyperspectral data preprocessing. Existing technologies typically employ traditional statistical analysis methods such as principal component analysis and band ratios, or combine them with machine learning algorithms such as support vector machines and random forests. By extracting, analyzing, selecting, or transforming the band features of hyperspectral data, data dimensionality compression is achieved, thereby reducing the computational overhead of subsequent model training and adapting to the needs of various data analysis tasks.

[0003] Existing methods for band selection and dimensionality reduction of hyperspectral data cubes struggle to balance data processing efficiency with the fidelity of effective features, thus affecting the reliability and efficiency of subsequent analyses based on the processed data. Summary of the Invention

[0004] In order to overcome the above-mentioned defects of the prior art, the present invention provides a rapid band screening and dimensionality reduction method and system for hyperspectral data cubes to solve the problems mentioned in the background art.

[0005] To achieve the above objectives, the present invention provides the following technical solution:

[0006] A fast band selection and dimensionality reduction method for hyperspectral data cubes includes the following steps:

[0007] S1. Obtain the hyperspectral data cube to be processed;

[0008] S2. Perform spectral correlation analysis on the hyperspectral data cube to generate a correlation matrix between bands;

[0009] S3. Based on the correlation matrix, clustering is performed on all bands to obtain several band clusters;

[0010] S4. By comparing the typical spectral response range of the target geological object, calculate the response matching degree of each band in the band cluster, and select the optimal representative band corresponding to each band cluster.

[0011] S5. Evaluate the cross-cluster feature complementarity of each optimal representative band, remove the optimal representative bands with duplicate features, and obtain the set of optimal representative bands after cross-cluster feature complementarity screening.

[0012] S6. Integrate the optimal representative band set after cross-cluster feature complementarity screening to form a dimensionality-reduced hyperspectral data subset.

[0013] Further, the hyperspectral data cube to be processed is obtained, including:

[0014] Acquire raw hyperspectral data containing spatial and spectral information;

[0015] Perform data validity verification on the raw hyperspectral data;

[0016] Invalid data that does not have continuous band characteristics are removed to obtain the hyperspectral data cube to be processed.

[0017] Furthermore, correlation analysis along the spectral dimension is performed on the hyperspectral data cube to generate a correlation matrix between bands, including:

[0018] Extract all band features of the spectral dimension of the hyperspectral data cube;

[0019] Calculate the correlation coefficient between any two band features in all band features;

[0020] A correlation matrix between bands is constructed based on all correlation coefficients.

[0021] Furthermore, clustering was performed on all bands based on the correlation matrix to obtain several band clusters, including:

[0022] The correlation coefficient in the correlation matrix is ​​used as the basis for determining band similarity;

[0023] Bands whose correlation coefficients meet similarity criteria are categorized.

[0024] After merging and classifying, the band groups are divided into several band clusters.

[0025] Furthermore, by comparing the typical spectral response range of the target geological object, the response matching degree of each band within the band cluster is calculated, and the optimal representative bands corresponding to each band cluster are selected, including:

[0026] Retrieve the typical spectral response range of the target geological object;

[0027] The degree of fit between each band within a band cluster and a typical spectral response range is calculated as the response matching degree.

[0028] Based on the response matching degree, the bands are sorted from high to low, and the band with the highest ranking is selected as the optimal representative band for the band cluster.

[0029] Furthermore, the target geological objects are mineralized alteration minerals in the field of remote sensing geological exploration, and the typical spectral response range of mineralized alteration minerals is the characteristic spectral wavelength range of pre-collected mineralized alteration minerals.

[0030] Furthermore, the cross-cluster feature complementarity of each optimal representative band is evaluated, and optimal representative bands with overlapping features are removed, resulting in a set of optimal representative bands after cross-cluster feature complementarity screening, including:

[0031] Calculate the characteristic overlap between any two optimal representative bands;

[0032] Determine whether the feature overlap meets the preset repetition condition, remove the best representative band that meets the preset repetition condition, and retain the best representative band that does not meet the preset repetition condition.

[0033] The best representative bands that are integrated and retained are used to obtain the set of best representative bands after cross-cluster feature complementarity screening.

[0034] Furthermore, the optimal representative band set after cross-cluster feature complementarity screening is integrated to form a dimensionality-reduced hyperspectral data subset, including:

[0035] Extract the spectral and spatial dimension information of all the best representative bands in the set of best representative bands after cross-cluster feature complementarity screening;

[0036] The extracted information is reconstructed according to the original data structure of the hyperspectral data cube, and the reconstructed dataset is output as a subset of the dimensionality-reduced hyperspectral data.

[0037] Furthermore, after forming the dimensionality-reduced hyperspectral data subset, a spectral dimension consistency check is performed on the dimensionality-reduced hyperspectral data subset. Band data with abnormal spectral responses are removed, and band data that conform to the spectral characteristics are retained as the final dimensionality-reduced dataset.

[0038] On the other hand, the present invention provides a rapid band selection and dimensionality reduction system for hyperspectral data cubes, comprising the following modules:

[0039] The data acquisition module is used to acquire the hyperspectral data cube to be processed;

[0040] The correlation analysis module is used to perform spectral-dimensional correlation analysis on hyperspectral data cubes and generate correlation matrices between bands.

[0041] The band clustering module is used to perform clustering processing on all bands based on the correlation matrix to obtain several band clusters;

[0042] The band selection module is used to compare the typical spectral response range of the target geological object, calculate the response matching degree of each band in the band cluster, and select the optimal representative band corresponding to each band cluster.

[0043] The complementarity evaluation module is used to evaluate the cross-cluster feature complementarity of each optimal representative band, remove the optimal representative bands with duplicate features, and obtain the set of optimal representative bands after cross-cluster feature complementarity screening.

[0044] The data integration module is used to integrate the best representative band set after cross-cluster feature complementarity screening to form a dimensionality-reduced hyperspectral data subset.

[0045] Compared with the prior art, the present invention has the following beneficial effects:

[0046] 1. By first automatically clustering massive bands based on spectral correlation, then selecting the most indicative representative bands from each group based on geological prior knowledge, and finally performing global feature redundancy removal across groups, a hierarchical and clearly guided band selection and dimensionality reduction process is formed. This effectively overcomes the contradictions of traditional methods, which either lose effective features when blindly compressing data or incur excessive computational costs when retaining features. By combining data-driven clustering analysis with target-driven band selection, the resulting dimensionality-reduced data subset is not only complementary in mathematical and statistical features but also directly correlated with exploration targets in a physical spectral sense. This significantly improves data processing efficiency while ensuring the reliability of subsequent mineral identification and information extraction.

[0047] 2. By rapidly identifying redundant bands and completing preliminary clustering through spectral dimensional correlation analysis, the number of units requiring subsequent fine processing is significantly reduced. Based on this, the typical spectral response range of the target geological object is introduced as a screening criterion, ensuring that the representative bands selected within each band cluster can retain spectral information sensitive to specific mineralization and alteration to the greatest extent. Further cross-cluster feature complementarity assessment eliminates potentially overlapping feature bands between different clusters from a global perspective, ensuring the compactness and diversity of the data set after dimensionality reduction. The final output hyperspectral data subset combines computational lightweighting with high-fidelity geological features, laying a high-quality data foundation for subsequent efficient and reliable quantitative remote sensing analysis. Attached Figure Description

[0048] Figure 1 This is a flowchart of the rapid band selection and dimensionality reduction method for hyperspectral data cubes according to the present invention;

[0049] Figure 2 This is a schematic diagram of the structure of the rapid band selection and dimensionality reduction system for hyperspectral data cubes according to the present invention. Detailed Implementation

[0050] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of the present invention.

[0051] Example 1: Figure 1 This invention presents a rapid band selection and dimensionality reduction method for hyperspectral data cubes, comprising the following steps:

[0052] S1. Obtain the hyperspectral data cube to be processed;

[0053] S2. Perform spectral correlation analysis on the hyperspectral data cube to generate a correlation matrix between bands;

[0054] S3. Based on the correlation matrix, clustering is performed on all bands to obtain several band clusters;

[0055] S4. By comparing the typical spectral response range of the target geological object, calculate the response matching degree of each band in the band cluster, and select the optimal representative band corresponding to each band cluster.

[0056] S5. Evaluate the cross-cluster feature complementarity of each optimal representative band, remove the optimal representative bands with duplicate features, and obtain the set of optimal representative bands after cross-cluster feature complementarity screening.

[0057] S6. Integrate the optimal representative band set after cross-cluster feature complementarity screening to form a dimensionality-reduced hyperspectral data subset.

[0058] According to step S1, the hyperspectral data cube to be processed is obtained, and the specific implementation process is as follows:

[0059] Obtain hyperspectral raw data containing spatial dimension information and spectral dimension information. The hyperspectral raw data is sourced from the original digital quantization values obtained by spaceborne or airborne hyperspectral imagers for earth observation. The hyperspectral raw data is a three-dimensional data block, containing two spatial dimensions and one spectral dimension. Read the hyperspectral raw data from the data storage medium supporting the hyperspectral imaging system or from the data stream transmitted through the data receiving station, to obtain an original data set containing the responses of each pixel in 100 to 300 consecutive narrow bands. Perform radiometric calibration processing on the obtained hyperspectral raw data. The radiometric calibration processing utilizes the pre-calibrated gain coefficient and bias coefficient of the imaging system. Convert the original digital value of each band into the apparent radiance value of the corresponding band. The conversion calculation is that the radiance of each pixel in a specific band is equal to the original digital value of the pixel multiplied by the calibration gain coefficient corresponding to the band plus the calibration bias coefficient corresponding to the band. Perform atmospheric correction processing on the data after radiometric calibration processing. The atmospheric correction processing adopts the FLAASH method based on the radiative transfer model or the empirical line method based on the image itself. The atmospheric correction processing eliminates the influence of atmospheric scattering and absorption on the surface signal, converts the radiance data into surface reflectance data, and forms an initial hyperspectral data cube. Each pixel point in the initial hyperspectral data cube contains a spectral curve obtained by continuously sampling from the visible light to the short-wave infrared region.

[0060] Data validity verification is performed on the raw hyperspectral data to ensure the integrity and basic rationality of the input data. Data validity verification includes file integrity verification. File integrity verification checks whether the data file is complete and whether the key dimension parameters specified in the file header information, such as the number of spatial rows, spatial columns, and the total number of bands, match the actual amount of data stored in the file. Data validity verification includes metadata verification. Metadata verification confirms that the metadata accompanying the data includes key parameters such as imaging time, geographical location, sensor type, solar altitude angle, and radiometric calibration coefficients. Metadata verification checks that the values ​​of key parameters such as imaging time, geographical location, sensor type, solar altitude angle, and radiometric calibration coefficients are within a reasonable physical range, and the solar altitude angle value is greater than 0 degrees and less than or equal to 90 degrees. Data validity verification includes data range verification. Data range verification checks the hyperspectral data cube after radiometric calibration and atmospheric correction. Data range verification checks that the data values ​​of all pixels in each band are within a theoretically reasonable range. For surface reflectance data, the theoretically reasonable range is defined by a lower theoretical reflectance threshold and a higher theoretical reflectance threshold. The theoretical reflectance lower and upper thresholds are set based on the physical reflectance characteristics of common ground features in the visible to shortwave infrared bands. Common ground features include vegetation, soil, water bodies, and rocks, most of which have reflectance between 0 and 1. For example, the theoretical reflectance lower threshold can be set to 0, and the theoretical reflectance upper threshold can be set to 1.0. If more than a certain proportion of pixel values ​​in a band exceed the range defined by the theoretical reflectance lower and upper thresholds, it is determined that the data in that band has a calibration or correction error. The certain proportion is set based on the allowable proportion of abnormal pixels, for example, 20%. Data validity verification includes outlier detection. Outlier detection calculates the mean and standard deviation of each band's data. Outlier detection identifies extreme pixel values ​​that differ from the band mean by more than a certain number of standard deviations. The number of standard deviations is set based on statistical outlier identification principles, for example, 5 standard deviations, marking extreme pixel values ​​as potential anomalous noise points. Data validity verification also includes preliminary verification of band continuity. The initial band continuity check calculates the average reflectance difference between adjacent bands. If the average reflectance difference between adjacent bands changes abruptly (meaning the difference exceeds a preset threshold), the threshold is set based on the characteristic that surface reflectance between adjacent bands in hyperspectral data typically does not change drastically (e.g., 0.1 reflectance units), and a warning message is issued. Through a verification process including file integrity check, metadata check, data range check, outlier detection, and initial band continuity check, missing, abnormal, or unreasonable elements in the data are identified and recorded, generating a data quality report.

[0061] Invalid data lacking continuous band characteristics are removed, resulting in a hyperspectral data cube for processing. This step applies to the hyperspectral data cube retained after validity verification. The main goal of removing invalid data lacking continuous band characteristics is to eliminate invalid bands whose spectral curves lose their continuous smoothness due to sensor malfunctions in specific bands or severe noise contamination. Continuous band characteristics are defined as the overall smoothness and continuity of the spectral curve corresponding to each spatial pixel, excluding specific absorption features. The sliding window method is used to detect the local continuity of the spectral curves. The sliding window method sets a window containing an odd number of bands that slides along the spectral dimension. The odd number of bands is determined by the window size that effectively assesses local continuity, such as 3, 5, or 7 bands. For each spectral curve in the hyperspectral data cube (the spectral curve being the data sequence of each pixel across all bands), the absolute value of the difference between the spectral value of the center band of the window and the spectral value of the preceding band is calculated. The absolute value of the difference between the spectral value of the center band of the window and the spectral value of the following band is also calculated. The absolute values ​​of the two calculated differences are compared with a preset continuity difference threshold. The continuity difference threshold is set based on the spectral resolution of the hyperspectral data and the local rate of change of the spectral curves of typical ground features. Higher spectral resolution results in smaller reflectance variations between adjacent bands, and the continuity difference threshold should be set accordingly. For example, for data with a spectral resolution of 10 nanometers, the reflectance variation of typical ground features within a 10-nanometer interval is typically less than 0.05, so the continuity difference threshold can be set to 0.05 reflectance units. If either of the absolute values ​​of the two differences exceeds the continuity difference threshold, the data at this pixel in that central band is marked as a potential discontinuity. The standard deviation of all band data within the sliding window is calculated. The standard deviation of all band data within the sliding window is compared with a preset noise level threshold. The noise level threshold is set based on the signal-to-noise ratio of the hyperspectral imaging system and the average noise level of the band data, and can be estimated from sensor performance parameters or by calculating the standard deviation of uniform regions in the image. For example, the noise level threshold can be set to 0.02 reflectance units. If the standard deviation of all band data within the sliding window exceeds the noise level threshold, the data at this pixel in the center band of the window is marked as a potential high-noise point. After traversing all pixels and all bands, the proportion of pixels marked as discontinuous or high-noise points in each band is calculated. If the proportion of marked pixels in a certain band exceeds a preset invalid band threshold, the entire band is determined to be an invalid band lacking continuous band characteristics and should be removed from the hyperspectral data cube. The invalid band threshold is set based on the application's tolerance for data continuity and the total number of hyperspectral bands. When the tolerance is low or the total number of bands is large, the invalid band threshold can be set more strictly. For example, the invalid band threshold can be set to 30%.For individual labeled pixel data within bands that were not entirely eliminated, linear interpolation of the corresponding pixel values ​​in adjacent valid bands is used for replacement and repair. Linear interpolation replacement and repair involves using the pixel's values ​​in the previous and next valid bands to estimate its value in the current band through linear calculation. Alternatively, the spatial location of the pixel can be marked, and these marked pixels will be noted in subsequent analyses. Finally, after eliminating invalid bands and identifying or repairing obviously discontinuous and abnormal pixels, a hyperspectral data cube with good continuity and physical plausibility in each band along the spectral dimension and completeness in the spatial dimension is obtained. This hyperspectral data cube serves as input for the spectral correlation analysis in subsequent step S2. The entire S1 step ensures the reliability and quality of the processed data, laying an accurate data foundation for subsequent band selection and dimensionality reduction.

[0062] According to step S2, correlation analysis of the spectral dimension is performed on the hyperspectral data cube to generate a correlation matrix between bands. The specific implementation process is as follows:

[0063] The input hyperspectral data cube originates from the hyperspectral data cube to be processed obtained in step S1. This data cube has undergone data validity verification and invalid data removal. Its spectral dimension contains multiple continuous bands, and its spatial dimension contains multiple pixels. Each pixel has a surface reflectance value or an atmospherically corrected radiance value in each band. All band features of the spectral dimension of the hyperspectral data cube are extracted. Band features refer to the distribution and statistical characteristics of the ground object radiation or reflectance information recorded in each band across the entire map space, represented in vector form. For a specific band, the band feature extraction method involves sequentially arranging all pixel values ​​in the corresponding two-dimensional spatial image to form a one-dimensional feature vector with a length equal to the total number of spatial pixels. Each element in this feature vector corresponds to the surface reflectance value or atmospherically corrected radiance value of a specific pixel in that specific band in the spatial dimension. The process involves traversing all bands of the hyperspectral data cube, performing the above operation on each band, thereby extracting multiple one-dimensional feature vectors of the same length as the total number of spatial pixels. This set of feature vectors forms the basis of the subsequent inter-band correlation analysis. Essentially, extracting features from all bands involves unfolding the three-dimensional hyperspectral data cube along the spectral dimension, flattening each band image into a one-dimensional array to facilitate the calculation of statistical relationships between bands.

[0064] Calculate the correlation coefficient between any two band features from all band features. The correlation coefficient is a statistic used to quantify the strength and direction of the linear relationship between two random variables. In this scenario, the correlation coefficient is calculated between the feature vectors corresponding to any two bands. The Pearson product-moment correlation coefficient is used as the method for calculating the correlation coefficient. The Pearson product-moment correlation coefficient measures the degree of linear correlation between two variables, and its value is between -1 and +1. For any two different bands, select the first band and the second band, and their corresponding feature vectors are the first feature vector and the second feature vector, respectively. The length of the first feature vector and the second feature vector is equal to the total number of spatial pixels. To calculate the Pearson product-moment correlation coefficient between the first band and the second band, first calculate the average of all elements in the first feature vector, and then calculate the average of all elements in the second feature vector. Then, calculate the deviation of each element in the first feature vector from its average value, and then calculate the deviation of each element in the second feature vector from its average value. Finally, calculate the sum of the products of all deviation values ​​of the first feature vector and all corresponding deviation values ​​of the second feature vector. Next, the sum of squares of all deviation values ​​for the first eigenvector and the sum of squares of all deviation values ​​for the second eigenvector are calculated separately. Finally, the Pearson product-moment correlation coefficient between the first and second bands is equal to the sum of the products of the deviations of the first and second eigenvectors, divided by the square root of the product of the sum of squares of the deviations of the first and second eigenvectors. This calculation process measures the linear similarity of the spatial distribution patterns of the two bands. If the spectral responses of the two bands increase or decrease synchronously in space, the correlation coefficient approaches positive 1, indicating a strong positive correlation; if one band increases while the other decreases, the correlation coefficient approaches negative 1, indicating a strong negative correlation; if there is no obvious linear relationship between the changes in the two bands, the correlation coefficient approaches 0. Traversing all bands, calculate the Pearson product-moment correlation coefficient between each pair of different band combinations. Since the correlation is symmetrical—that is, the correlation coefficient between the first and second bands is equal to the correlation coefficient between the second and first bands—the total number of correlation coefficients that actually needs to be calculated is the total number of bands multiplied by the total number of bands minus 1, then divided by 2. All the calculated correlation coefficients will be used to construct the matrix for the next stage.

[0065] A correlation matrix is ​​constructed based on all correlation coefficients. The correlation matrix is ​​a square matrix used to systematically store and display the pairwise correlation relationships between all bands. Let the total number of bands in the hyperspectral data cube be a specific value; then the constructed correlation matrix is ​​a square matrix with both rows and columns equal to the total number of bands. The element value in the first row and first column of the matrix is ​​the Pearson product-moment correlation coefficient between the band at the corresponding row number and the band at the corresponding column number. The specific process of constructing this matrix is ​​as follows: First, initialize a matrix with both rows and columns equal to the total number of bands and all values ​​being 0. Then, define the correlation coefficient between any band and itself as 1; that is, for all elements on the main diagonal of the matrix, where the row number and column number are equal, set the matrix element value to 1, indicating that any band is perfectly correlated with itself. Next, for all matrix element positions where the row number and column number are not equal (i.e., non-main diagonal elements), fill in the Pearson product-moment correlation coefficient between the corresponding two bands calculated in step S2. Since the correlation matrix is ​​a symmetric matrix (meaning the element value in the first row and first column equals the element value in the first row and first column), during filling, the upper triangular portion can be calculated and filled, then copied to the lower triangular portion; alternatively, all band pairs can be directly calculated and filled into their corresponding positions. The resulting correlation matrix is ​​a real symmetric matrix with diagonal elements equal to 1 and off-diagonal elements between -1 and +1. This matrix fully represents the linear correlation between any two spectral bands in the hyperspectral data cube in terms of spatial information distribution. The correlation matrix serves as the foundational input data for band clustering in subsequent step S3, using the similarity between bands represented by the elements in this matrix as the basis for the clustering algorithm. Step S2 condenses the complex relationships between bands in the hyperspectral data into a structured mathematical matrix, providing a quantitative analytical foundation for subsequent automated band selection and dimensionality reduction. Correlation analysis along the spectral dimension can identify band groups with high information redundancy, namely band pairs whose correlation coefficients are very close to positive 1 or negative 1. These band pairs may carry highly similar or repetitive spatial-spectral information, which points the way to compressing data dimensions while preserving effective information.

[0066] According to step S3, clustering is performed on all bands based on the correlation matrix to obtain several band clusters. The specific implementation process is as follows:

[0067] The input correlation matrix is ​​derived from the hyperspectral data cube band correlation matrix generated in step S2. This matrix is ​​a real symmetric matrix with both rows and columns equal to the total number of bands. Each element in the matrix is ​​the Pearson product-moment correlation coefficient between two corresponding bands, with values ​​ranging from -1 to +1. All elements on the main diagonal of the matrix are 1. The correlation coefficient in the correlation matrix is ​​used as the basis for determining band similarity. The similarity determination criterion refers to using the absolute value of the Pearson product-moment correlation coefficient between two bands as a quantitative indicator of the similarity in spatial information distribution between the two bands. The closer the absolute value of the Pearson product-moment correlation coefficient is to 1, the stronger the linear relationship between the spatial variation patterns of the corresponding feature vectors of the two bands, meaning the spatial information carried by the two bands is more similar or exhibits a mirror-image relationship, with higher information redundancy. Conversely, the closer the absolute value of the Pearson product-moment correlation coefficient is to 0, the weaker the linear relationship between the variation patterns of the two bands, meaning the stronger the independence of the spatial information carried by the two bands. In clustering, the focus is primarily on band combinations with high positive or negative correlations, as these typically represent bands with similar or complementary spectral response characteristics and high information overlap. Therefore, a similarity threshold is set as a quantitative criterion for determining whether two bands are sufficiently similar to be grouped into the same cluster. The similarity threshold is determined based on the tolerance for band information redundancy in the specific application scenario and the required degree of data dimensionality reduction, typically choosing a value close to 1. For example, a similarity threshold of 0.8 can be set. When the absolute value of the Pearson product-moment correlation coefficient between any two bands is greater than or equal to this similarity threshold, the two bands are considered to meet the similarity condition, possessing high spectral information similarity, and should be considered for grouping into the same class during the clustering process.

[0068] Bands whose correlation coefficients meet the similarity criteria are grouped. The grouping operation employs a threshold-based agglomerative hierarchical clustering algorithm. At the beginning of the algorithm, each band in the hyperspectral data cube is considered an independent initial band cluster; the number of band clusters equals the total number of bands. The similarity between all pairwise initial band clusters is calculated. Since each initial band cluster contains only one band, the similarity between two clusters is directly defined by the absolute value of the Pearson product-moment correlation coefficient corresponding to these two bands in the correlation matrix. Based on the calculated initial cluster similarity, the pair of band clusters with the highest similarity value is found among all cluster pairs. It is then determined whether this highest similarity value is greater than or equal to a pre-set similarity threshold. If the highest similarity value is greater than or equal to the similarity threshold, this pair of band clusters is merged into a new band cluster. The new band cluster contains all bands from the two merged original band clusters. If the highest similarity value is less than the similarity threshold, it means that no pair of band clusters currently meets the classification criteria, and the clustering process can terminate. After successfully merging a pair of band clusters, the total number of band clusters decreases by one. At this point, it's necessary to recalculate the similarity between this newly merged band cluster and all other existing band clusters. Calculating the similarity between the new cluster and other clusters requires defining a rule to measure inter-cluster similarity. The maximum similarity link rule is adopted, which defines the similarity between two band clusters as the maximum absolute value of the Pearson product-moment correlation coefficient between all possible band pairs in those two clusters. Specifically, for the newly merged band cluster A and another existing band cluster B, iterate through every band in band cluster A and every band in band cluster B, find the absolute values ​​of the correlation coefficients between all these band pairs from the correlation matrix, and then take the maximum value as the similarity between band cluster A and band cluster B. After updating the entire inter-cluster similarity relationship, repeat the previous steps, continuing to find the pair of clusters with the highest similarity among all existing band clusters, determining whether their similarity meets the similarity threshold, and if so, continuing the merging process and updating the similarity. This process is an iterative loop.

[0069] After merging and classifying the band groups, several band clusters are obtained. The termination condition of the iterative loop is that no pair of band clusters is found whose similarity is greater than or equal to a preset similarity threshold. When the termination condition is met, the agglomerative hierarchical clustering algorithm stops, and the resulting independent band groups that no longer meet the merging condition are the final band clusters. Each final band cluster contains one or more spectral bands, which are connected to each other through direct or indirect similarity links. That is, any two bands within a cluster either directly satisfy the similarity condition or satisfy a transitive similarity relationship through other bands within the cluster, thus ensuring that the bands within the same cluster have a high degree of homogeneity or redundancy in spectral spatial information. Any two bands belonging to different band clusters must have a similarity less than the similarity threshold, indicating that the spectral spatial information they carry has relative independence and complementarity. The number of band clusters obtained in the end depends on the setting of the similarity threshold and the inherent correlation structure between the bands in the original hyperspectral data. A higher similarity threshold, such as increasing it from 0.8 to 0.9, results in stricter merging conditions, meaning only extremely similar bands will be clustered together. This typically leads to a larger number of smaller band clusters. Conversely, a lower similarity threshold, such as decreasing it to 0.7, results in more lenient merging conditions, potentially leading to fewer but larger band clusters. In practical applications, the specific value of the similarity threshold can be pre-set using an empirical value, such as 0.85, by analyzing the correlation characteristics of typical ground cover spectral curves in adjacent bands and combining this with the dimensionality reduction objective. Alternatively, a desired range for the final number of band clusters can be set in the algorithm, and the number of clusters can be adjusted and repeated until it falls within the desired range. For example, if the goal is to reduce 200 bands to approximately 20 to 30 feature bands, the similarity threshold can be adjusted to ensure that the number of band clusters generated by clustering falls between 20 and 30. The final output of several band clusters forms the basis for representative band selection and feature complementarity analysis in subsequent steps S4 and S5. Each band cluster will participate as a whole in the subsequent selection of the optimal representative band. Step S3, through unsupervised clustering, automatically divides the massive hyperspectral bands into several band groups that are internally similar but externally dissimilar, based on objective correlation measures between bands. This establishes a structured grouping framework for systematic band selection and dimensionality reduction, effectively reducing the computational complexity of pairwise comparisons across the entire band range in subsequent steps. It is a key step in achieving rapid data compression while preserving effective features.

[0070] According to step S4, by comparing the typical spectral response range of the target geological object, the response matching degree of each band within the band cluster is calculated, and the optimal representative band corresponding to each band cluster is selected. The specific implementation process is as follows:

[0071] The input band clusters are derived from the final band cluster set obtained after the clustering process in step S3. Each band cluster is a group containing one or more spectral bands that are highly similar in spatial information. The target geological object is mineralized alteration minerals in the field of remote sensing geological exploration. Mineralized alteration minerals are a class of minerals with indicative mineralization potential, formed during hydrothermal or weathering processes. Common mineralized alteration minerals include sericite, chlorite, limonite, kaolinite, and alunite. These minerals have diagnostic absorption or reflection characteristics in the visible to shortwave infrared spectral range, which are manifested as absorption valleys or reflection peaks at specific wavelengths. The typical spectral response range of mineralized alteration minerals is the characteristic spectral wavelength range of pre-collected mineralized alteration minerals. The typical spectral response range refers to the continuous wavelength range corresponding to the diagnostic spectral characteristics of a specific mineralized alteration mineral. For example, sericite exhibits a distinct aluminum hydroxyl absorption feature near 2200 nm, and its typical spectral response range can be defined as the wavelength range from 2150 nm to 2250 nm around 2200 nm; chlorite exhibits an iron-magnesium hydroxyl absorption feature near 2250 nm, and its typical spectral response range can be defined as the wavelength range from 2220 nm to 2280 nm. These typical spectral response ranges are obtained by consulting authoritative spectral databases, such as the USGS spectral library or the ASTER thermal emission and reflection radiometer spectral library. Standard reflectance spectra of the target minerals are extracted, and then the wavelength width corresponding to the bottom half depth of the diagnostic absorption feature in the curve is identified manually or by algorithms as a typical spectral response range for that mineral. For minerals with multiple diagnostic features, multiple typical spectral response ranges can be extracted. The typical spectral response ranges of all target minerals to be detected are pre-organized and stored in a data list, where each record is associated with a mineral name and one or more response ranges defined by a start wavelength and an end wavelength.

[0072] Retrieve typical spectral response intervals for the target geological object. When selecting representative bands, based on the specific geological exploration objective, retrieve typical spectral response intervals of one or more relevant mineralization and alteration minerals from a pre-stored data list. For example, if the exploration objective is to find copper-gold mineralization associated with hydrothermal alteration, typical spectral response intervals of minerals such as sericite, chlorite, and kaolinite might be retrieved. These retrieved intervals constitute the spectral characteristic reference benchmark for this band selection. Calculate the degree of fit between each band within a band cluster and the typical spectral response interval as the response matching degree. For a given band cluster, which contains several spectral bands, each band corresponds to a specific center wavelength and a finite bandwidth of the hyperspectral imager. The center wavelength of a band refers to the wavelength value corresponding to the peak value of the band's passband response, and the bandwidth refers to the wavelength span corresponding to the full width at half maximum (FWHM) of the band's passband response. The response matching degree is an indicator that quantifies the degree of alignment between the center wavelength and bandwidth of a band and a typical spectral response interval in spectral position. One method for calculating the response matching degree is based on distance metrics. First, for a typical spectral response range being retrieved, the center wavelength of that range is calculated, which is the average of the starting and ending wavelengths of the range. Then, for a band within the band cluster to be evaluated, the absolute difference between the center wavelength of that band and the center wavelength of the typical spectral response range is calculated. To ensure the matching degree is a normalized value between 0 and 1, a normalization coefficient needs to be defined. The normalization coefficient can be set to half the width of the entire spectral coverage of the hyperspectral data; for example, for data covering 400 nm to 2500 nm, the normalization coefficient can be set to 1050 nm. Therefore, the initial distance matching degree of this band relative to this response range can be calculated as 1 minus the absolute difference between the band's center wavelength and the range's center wavelength divided by the normalization coefficient. However, considering only the center distance may ignore the overlap between the band width and the range width. Therefore, a more optimized response matching degree calculation needs to consider the overlap ratio between the band range and the response range. The overlap length between the band range and the typical spectral response interval is calculated. The band range begins with the center wavelength minus half the bandwidth and ends with the center wavelength plus half the bandwidth. The overlap length is equal to the difference between the maximum value of the starting wavelength and the minimum value of the ending wavelength of both the band range and the response interval. When the difference is less than 0, the overlap length is 0. The response matching degree can then be calculated as the overlap length divided by the band width. If the band falls entirely within the typical spectral response interval, the overlap length equals the band width, and the response matching degree is 1. If the band partially overlaps with the interval, the response matching degree is a decimal between 0 and 1. If the band does not overlap with the interval, the response matching degree is 0. For a band within a band cluster, if multiple typical spectral response intervals exist, the response matching degree between the band and each interval is calculated, and the maximum value is taken as the final response matching degree for that band.This means that a band is considered to have high geological significance as long as it corresponds well with the characteristic range of any target mineral in terms of spectral position.

[0073] Based on the response matching degree, bands are sorted from highest to lowest, and the band with the highest ranking is selected as the optimal representative band for the band cluster. For a specific band cluster, all spectral bands within the cluster are traversed, and the response matching degree of each band is calculated using the aforementioned method. The response matching degree values ​​of all bands are arranged in descending order to generate an ordered list. The band with the highest response matching degree, i.e., the band with the highest ranking, is selected as the optimal representative band for that band cluster. The optimal representative band means that among all bands with similar information within the cluster, the spectral position of this band best matches the diagnostic spectral characteristics of the target geological mineral. Therefore, selecting it as the representative can preserve the most valuable part of the cluster's spectral information for the interpretation of the target geological features to the greatest extent. In special cases, if multiple bands have the same response matching degree value and are all the highest value, a secondary selection rule is needed to determine the unique optimal representative band. Secondary selection rules can prioritize the band whose center wavelength is closer to the center of the typical spectral response interval, or prioritize the band with the smaller band number, or prioritize the band with the higher average reflectance signal-to-noise ratio in the entire hyperspectral data cube. Through this step, each band cluster will generate an optimal representative band. The optimal representative bands generated by all band clusters constitute an initial set of candidate representative bands. The number of bands in this set is equal to the number of band clusters generated in step S3, thus achieving the first round of screening and significant dimensionality reduction from hundreds or thousands of original bands based on spectral feature similarity and geological significance. Step S4 integrates geological prior knowledge into the automated band screening process in the form of typical spectral response intervals, ensuring that the dimensionality-reduced data subset can prominently reflect the target mineralization and alteration information, significantly improving the effectiveness and relevance of subsequent mineral identification or mineralization prediction analysis tasks. This target response-based screening mechanism, unlike dimensionality reduction methods that solely rely on statistical features, is one of the key designs for achieving a balance between data processing efficiency and the fidelity of effective features.

[0074] According to step S5, the cross-cluster feature complementarity of each optimal representative band is evaluated, and the optimal representative bands with overlapping features are removed to obtain the set of optimal representative bands after cross-cluster feature complementarity screening. The specific implementation process is as follows:

[0075] The input optimal representative bands are derived from the optimal representative bands corresponding to each band cluster obtained after screening in step S4. These optimal representative bands constitute an initial set of candidate representative bands. The total number of bands in this set is equal to the number of band clusters generated in step S3. The feature overlap between any two optimal representative bands is calculated. Feature overlap is a quantitative indicator used to measure the similarity or redundancy of the spectral spatial information carried by two different bands. In this embodiment, the feature overlap is calculated based on the similarity between the feature vectors of the original hyperspectral data corresponding to two optimal representative bands. Specifically, from the final hyperspectral data cube obtained in step S1, the full spatial pixel values ​​corresponding to each optimal representative band are extracted to form the same band feature vector as in step S2. For an optimal representative band A and another optimal representative band B, the feature vector of band A is denoted as vector A, and the feature vector of band B is denoted as vector B. The lengths of vector A and vector B are both equal to the total number of spatial pixels. The calculation of feature overlap can directly reuse the method for calculating the Pearson product-moment correlation coefficient in step S2, but here we focus on the degree of similarity represented by the absolute value. Calculate the Pearson product-moment correlation coefficient between feature vectors A and B of band A and band B, and then take the absolute value of the Pearson product-moment correlation coefficient as the preliminary value of feature overlap between these two optimal representative bands. The preliminary value of feature overlap ranges from 0 to 1. The closer the value is to 1, the closer the response change patterns of band A and band B are to linear correlation across all spatial pixels, meaning that the spatial distribution information of ground features reflected by the two bands is highly similar or completely redundant; the closer the value is to 0, the weaker the linear relationship between the change patterns of the two bands, meaning that the spatial information they carry is independent and their features are highly complementary. However, using only the correlation of spatial feature vectors may not be sufficient to comprehensively measure the feature overlap in the spectral dimension. Therefore, another method for calculating feature overlap also incorporates the spectral proximity of the bands. Calculate the absolute difference between the center wavelength of band A and the center wavelength of band B, and normalize this difference to the interval between 0 and 1. The normalization method divides the wavelength difference by the total span of the entire operating wavelength range of the hyperspectral sensor. For example, if the sensor's operating range is from 400 nm to 2500 nm, the total span is 2100 nm. Therefore, the spectral position overlap can be calculated as 1 minus the normalized wavelength difference. The final feature overlap can be defined as the weighted average of the absolute value of spatial feature vector correlation and spectral position overlap. Weighting coefficients are used to balance the importance of spatial information similarity and spectral position proximity in evaluating feature overlap. The weighting coefficients are set according to specific application requirements. If the uniqueness of spatial information is considered more critical, a higher weight is given to the absolute value of spatial correlation, for example, setting the spatial correlation weight greater than the spectral position weight. A simple setting is to set both weights to 0.5, indicating equal importance.Using the above method, all possible pairwise band combinations in the initial candidate representative band set are traversed, and the feature overlap between each pair of combinations is calculated.

[0076] The process involves determining whether the feature overlap meets a preset repetition condition. Optimal representative bands that meet the preset repetition condition are removed, while those that do not are retained. The preset repetition condition is a logical criterion used to identify optimal representative bands whose features are too similar, leading to redundant information contributions. This condition is typically defined by setting a feature overlap threshold. The threshold is based on the desired feature diversity of the dimensionality-reduced data subset and the upper limit of tolerance for information redundancy. Higher desired feature diversity necessitates a lower threshold to remove more similar bands. For example, a threshold of 0.7 can be used. The process is as follows: For the initial set of candidate representative bands, a band graph is constructed, where nodes represent each optimal representative band. All calculated feature overlap values ​​are iterated. When the feature overlap between band A and band B is found to be greater than or equal to the threshold, a connection edge is established between the nodes of band A and band B, indicating a high degree of feature overlap between the two bands. After completing all the determinations, the connected components of the band graph are analyzed. Within a connected component, any two nodes can be directly or indirectly connected by an edge. This means that all the best representative bands within that component have a high degree of feature overlap, forming a set of feature-redundant bands. The strategy for eliminating bands with pre-defined duplication conditions is to retain only one best representative band in each such connected component and eliminate all other best representative bands within that component. Choosing which band to retain requires establishing rules. One feasible rule is to retain the best representative band with the highest response matching degree in the connected component. The response matching degree is calculated from step S4, ensuring that the retained band is both unique and geologically significant. Another rule is to retain the band with the lowest average feature overlap with all other bands in the connected component; that is, to select the one least similar to other members in the group as the representative, maximizing feature complementarity. Alternatively, simply retaining the band with the smallest band number is also acceptable. By applying preset repetition conditions and corresponding elimination rules, the optimal representative bands that highly overlap with the features of other bands will be removed, thereby breaking the cross-cluster feature redundancy that may be caused by the clustering in step S3 and the screening in step S4, and ensuring that the remaining bands have sufficient differences and complementarity with each other.

[0077] The best representative bands retained are integrated to obtain the set of best representative bands after cross-cluster feature complementarity screening. After completing the above judgment and elimination operations, all the retained best representative bands are gathered together to form a new set, which is the set of best representative bands after cross-cluster feature complementarity screening. The number of bands in this set will be less than or equal to the number of the initial candidate representative band set, depending on the severity of feature overlap in the initial set and the strictness of the preset repetition conditions. After this screening step, each best representative band in the final set satisfies two conditions: first, it is the band in its band cluster that best matches the spectral response of the target geological object; second, its feature overlap with any other retained band in the entire set is lower than the preset feature overlap threshold, thus ensuring cross-cluster feature complementarity. This final set is representative in the spectral dimension, targeted in geological significance, and complementary in information content, avoiding information redundancy and laying the foundation for subsequently constructing a compact and information-rich hyperspectral data subset. Step S5 further optimizes and refines the results of step S4. By introducing a cross-cluster global feature complementarity evaluation, it addresses the issue that different band clusters may select representative bands with similar spectral positions or spatial information. This balances feature fidelity and data compression rate at a finer granularity, making it a key step for achieving high-efficiency and high-reliability dimensionality reduction. The final optimal set of representative bands will serve as input to step S6 to generate the final dimensionality-reduced data subset.

[0078] According to step S6, the optimal representative band set after cross-cluster feature complementarity screening is integrated to form a dimensionality-reduced hyperspectral data subset. The specific implementation process is as follows:

[0079] The input set of optimal representative bands after cross-cluster feature complementarity filtering is derived from the optimal representative bands retained after processing in step S5. This set contains multiple spectral bands, far fewer than the total number of bands in the original hyperspectral data cube. The spectral and spatial dimensions of all optimal representative bands in the set after cross-cluster feature complementarity filtering are extracted. Spectral dimension information refers to the center wavelength, bandwidth, and spectral channel number or index of each optimal representative band in the entire hyperspectral imaging system. Spatial dimension information refers to the two-dimensional spatial image data corresponding to each optimal representative band. This image data is a two-dimensional matrix, where each element represents the surface reflectance or atmospherically corrected radiance value of a pixel at a specific spatial location on the Earth's surface within that optimal representative band. The specific extraction method involves first reading the unique identifier of each optimal representative band from the set after cross-cluster feature complementarity filtering. This identifier is typically the band number in the original hyperspectral data cube. Then, based on this band number, the original hyperspectral data cube stored in memory or on disk is accessed. This data cube is the hyperspectral data cube to be processed obtained after preprocessing in step S1. Using the band number index, the entire two-dimensional spatial image data block corresponding to that band number is directly extracted from the spectral dimension of the original hyperspectral data cube. This data block contains the pixel values ​​of all spatial rows and columns, i.e., complete spatial dimension information. Simultaneously, the center wavelength and bandwidth parameters associated with that band number are retrieved from the metadata file attached to the original hyperspectral data cube. These parameters constitute the spectral dimension information of that band. For each optimal representative band in the optimal representative band set after cross-cluster feature complementarity filtering, the above indexing and extraction operations are performed sequentially, thereby obtaining multiple two-dimensional spatial image data blocks equal in number to the optimal representative bands, and the same number of multiple sets of spectral wavelength parameters.

[0080] The extracted information is reconstructed based on the original data structure of the hyperspectral data cube, and the reconstructed dataset is output as a subset of the dimensionality-reduced hyperspectral data. The original data structure of the hyperspectral data cube is a three-dimensional array, whose three dimensions are, in order, the number of spatial rows, the number of spatial columns, and the number of spectral bands. The goal of the reconstructing operation is to construct a new, smaller three-dimensional array. The specific reconstructing process is as follows: First, the dimensionality of the new three-dimensional array is determined. The number of spatial rows and columns of the new array are completely inherited from the original hyperspectral data cube and remain unchanged. The number of spectral bands in the new array is equal to the total number of bands in the optimal representative band set after cross-cluster feature complementarity screening. Then, a new three-dimensional array is created with an initial value of empty, and its shape is the number of spatial rows multiplied by the number of spatial columns multiplied by the total number of optimal representative bands. Next, multiple extracted two-dimensional spatial image data blocks are sequentially filled into the spectral dimension of the new three-dimensional array according to the specific order of their corresponding optimal representative bands in the final set. This specific order can be arranged according to the center wavelength values ​​of the bands from smallest to largest, or according to the original index order of the bands in the original data cube. Regardless of the order, the mapping relationship between this order and the spectral dimension information must be recorded to ensure the metadata of the new data subset is complete and correct. For example, an ascending order of center wavelength can be used, placing the 2D image data block corresponding to the optimal representative band with the smallest center wavelength in the first position of the new 3D array's spectral dimension, the second smallest center wavelength in the second position, and so on, until the largest center wavelength is placed in the last position. When filling each 2D image data block, it is necessary to ensure that its spatial row and column arrangement is consistent with the original data. That is, the pixel value of a certain row and column in the data block is accurately placed in the corresponding row position, corresponding column position, and position determined by the index of the currently filled spectral band in the new 3D array. After filling all the 2D data blocks, a new hyperspectral data cube is obtained, which is the dimensionality-reduced hyperspectral data subset. Simultaneously, corresponding metadata needs to be generated for this new data subset. This metadata must explicitly record the original band number, center wavelength, band width, and new band number for each spectral band in the new cube within the dimensionality-reduced subset. Finally, this new three-dimensional array data and its metadata are output in a standard hyperspectral data file format, such as an ENVI format binary data file and its corresponding header file, thus completing the generation and storage of the dimensionality-reduced hyperspectral data subset.

[0081] After forming the dimensionality-reduced hyperspectral data subset, a spectral dimension consistency check is performed on the subset. Band data with abnormal spectral responses are removed, and band data that conform to the spectral characteristics are retained as the final dimensionality-reduced dataset. The purpose of the spectral dimension consistency check is to ensure that, during the dimensionality reduction process, due to band selection and reorganization, no abnormal data points that would cause severe distortion of the spectral curve in local spatial regions have been introduced or omitted, thus guaranteeing the spectral quality of the final output dataset. The check is performed on the spectral curve of each spatial pixel in the dimensionality-reduced hyperspectral data subset. A spectral curve refers to the sequence of values ​​of a pixel at a specific spatial location in the dimensionality-reduced hyperspectral data subset across all the optimal representative bands. The check consists of two main parts. The first part is the spectral curve smoothness check. For a spectral curve, its first-order difference value between adjacent optimal representative bands is calculated, i.e., the value of the latter band minus the value of the former band. Since the reflectance spectrum of ground objects changes smoothly overall after excluding absorption features, the absolute values ​​of these difference values ​​should not show drastic jumps. Calculate the absolute values ​​of the differences between all adjacent bands and find the maximum value. Compare this maximum value with a preset spectral curve smoothness difference threshold. The spectral curve smoothness difference threshold is set based on the maximum reasonable rate of change of the spectral curve of a typical ground object within the corresponding wavelength interval. This rate of change can be obtained by analyzing data in a standard ground object spectral library. For example, for two bands approximately 20 nanometers apart in the shortwave infrared region, the reflectance change of a typical rock spectrum is usually less than 0.1, so the spectral curve smoothness difference threshold can be set to 0.15 reflectance units. If the maximum absolute value of the difference of a spectral curve exceeds the spectral curve smoothness difference threshold, the spectral curve marking that pixel location has a potential local anomalous jump. The second part is the global shape rationality check of the spectral curve. Calculate the mean and standard deviation of the values ​​of the spectral curve over all optimal representative bands. Check if any band's value deviates from the mean by more than a certain number of standard deviations. This number of deviations is usually set to 3 or 4 times the standard deviation, based on the assumption of a normal distribution of the data. If such a band exists, the value of that pixel in that specific band is marked as an isolated outlier. After traversing all spatial pixels in the dimensionality-reduced hyperspectral data subset, statistical analysis is performed. For a specific optimal representative band, the number of times it is marked as an isolated outlier in the spectral curves of all pixels is counted. If the number of times it is marked exceeds a preset threshold for the proportion of outlier pixels in a band, for example, 5% of the total number of pixels in that band, it is considered that the optimal representative band may have systematic data quality problems across the entire spatial range, such as residual stripe noise or incompletely corrected bad lines. Therefore, the entire band is removed from the dimensionality-reduced hyperspectral data subset. The threshold for the proportion of outlier pixels in a band is set based on image quality requirements and the total number of spatial pixels.For pixels marked due to smoothness testing, the values ​​of these pixels at their corresponding band positions are replaced and repaired using linear interpolation of the values ​​in their adjacent bands. After completing all verification, removal, and repair operations, the remaining band data constitutes the final dimensionality-reduced dataset. This final dimensionality-reduced dataset is a representative, complementary, and reliable subset of hyperspectral data in the spectral dimension and complete in the spatial dimension. Its data volume is much smaller than the original data, but it retains the spectral spatial information that is crucial for the identification and analysis of target geological objects to the greatest extent. It can be directly used for subsequent advanced applications such as mineral mapping, alteration information extraction, or mineralization prediction, effectively balancing data processing efficiency and feature fidelity. The entire step S6 completes the transformation from the band selection list to the final usable data product and is the output link of the entire methodology.

[0082] Example 2: Figure 2 A schematic diagram of the fast band selection and dimensionality reduction system for hyperspectral data cubes according to the present invention is given. The fast band selection and dimensionality reduction system for hyperspectral data cubes includes the following modules:

[0083] The data acquisition module is used to acquire the hyperspectral data cube to be processed;

[0084] The correlation analysis module is used to perform spectral-dimensional correlation analysis on hyperspectral data cubes and generate correlation matrices between bands.

[0085] The band clustering module is used to perform clustering processing on all bands based on the correlation matrix to obtain several band clusters;

[0086] The band selection module is used to compare the typical spectral response range of the target geological object, calculate the response matching degree of each band in the band cluster, and select the optimal representative band corresponding to each band cluster.

[0087] The complementarity evaluation module is used to evaluate the cross-cluster feature complementarity of each optimal representative band, remove the optimal representative bands with duplicate features, and obtain the set of optimal representative bands after cross-cluster feature complementarity screening.

[0088] The data integration module is used to integrate the best representative band set after cross-cluster feature complementarity screening to form a dimensionality-reduced hyperspectral data subset.

[0089] All calculations involved in the embodiments are dimensionless numerical calculations, and the preset parameters and thresholds in the calculations are set by those skilled in the art according to the actual situation.

[0090] The above embodiments can be implemented, in whole or in part, by software, hardware, firmware, or any other combination thereof. When implemented using software, the above embodiments can be implemented, in whole or in part, in the form of a computer program product.

[0091] Those skilled in the art will recognize that the modules and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and inventive constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0092] In addition, the functional modules in the various embodiments of this application can be integrated into one processing module, or each module can exist physically separately, or two or more modules can be integrated into one module.

[0093] In the several embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of modules is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple modules or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or modules may be electrical, mechanical, or other forms.

[0094] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

[0095] In conclusion, the above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.

Claims

1. A rapid band selection and dimensionality reduction method for hyperspectral data cubes, characterized in that, Includes the following steps: S1. Obtain the hyperspectral data cube to be processed; S2. Perform spectral correlation analysis on the hyperspectral data cube to generate a correlation matrix between bands; S3. Based on the correlation matrix, clustering is performed on all bands to obtain several band clusters; S4. By comparing the typical spectral response range of the target geological object, calculate the response matching degree of each band in the band cluster, and select the optimal representative band corresponding to each band cluster. S5. Evaluate the cross-cluster feature complementarity of each optimal representative band, remove the optimal representative bands with duplicate features, and obtain the set of optimal representative bands after cross-cluster feature complementarity screening. S6. Integrate the optimal representative band set after cross-cluster feature complementarity screening to form a dimensionality-reduced hyperspectral data subset.

2. The rapid band selection and dimensionality reduction method for hyperspectral data cubes according to claim 1, characterized in that, Obtain the hyperspectral data cube to be processed, including: Acquire raw hyperspectral data containing spatial and spectral information; Perform data validity verification on the raw hyperspectral data; Invalid data that does not have continuous band characteristics are removed to obtain the hyperspectral data cube to be processed.

3. The rapid band selection and dimensionality reduction method for hyperspectral data cubes according to claim 1, characterized in that, Correlation analysis along the spectral dimension is performed on the hyperspectral data cube to generate a correlation matrix between bands, including: Extract all band features of the spectral dimension of the hyperspectral data cube; Calculate the correlation coefficient between any two band features in all band features; A correlation matrix between bands is constructed based on all correlation coefficients.

4. The rapid band selection and dimensionality reduction method for hyperspectral data cubes according to claim 1, characterized in that, Clustering was performed on all bands based on the correlation matrix, resulting in several band clusters, including: The correlation coefficient in the correlation matrix is ​​used as the basis for determining band similarity; Bands whose correlation coefficients meet similarity criteria are categorized. After merging and classifying, the band groups are divided into several band clusters.

5. The rapid band selection and dimensionality reduction method for hyperspectral data cubes according to claim 1, characterized in that, By comparing the typical spectral response range of the target geological object, the response matching degree of each band within the band cluster is calculated, and the optimal representative bands corresponding to each band cluster are selected, including: Retrieve the typical spectral response range of the target geological object; The degree of fit between each band within a band cluster and a typical spectral response range is calculated as the response matching degree. Based on the response matching degree, the bands are sorted from high to low, and the band with the highest ranking is selected as the optimal representative band for the band cluster.

6. The rapid band selection and dimensionality reduction method for hyperspectral data cubes according to claim 5, characterized in that, The target geological objects are mineralized alteration minerals in the field of remote sensing geological exploration. The typical spectral response range of mineralized alteration minerals is the characteristic spectral wavelength range of pre-collected mineralized alteration minerals.

7. The rapid band selection and dimensionality reduction method for hyperspectral data cubes according to claim 1, characterized in that, The cross-cluster feature complementarity of each optimal representative band is evaluated, and optimal representative bands with overlapping features are removed to obtain the set of optimal representative bands after cross-cluster feature complementarity screening, including: Calculate the characteristic overlap between any two optimal representative bands; Determine whether the feature overlap meets the preset repetition condition, remove the best representative band that meets the preset repetition condition, and retain the best representative band that does not meet the preset repetition condition. The best representative bands that are integrated and retained are used to obtain the set of best representative bands after cross-cluster feature complementarity screening.

8. The rapid band selection and dimensionality reduction method for hyperspectral data cubes according to claim 1, characterized in that, The optimal representative band set, after cross-cluster feature complementarity screening, is integrated to form a dimensionality-reduced hyperspectral data subset, including: Extract the spectral and spatial dimension information of all the best representative bands in the set of best representative bands after cross-cluster feature complementarity screening; The extracted information is reconstructed according to the original data structure of the hyperspectral data cube, and the reconstructed dataset is output as a subset of the dimensionality-reduced hyperspectral data.

9. The rapid band selection and dimensionality reduction method for hyperspectral data cubes according to claim 8, characterized in that, After forming a dimensionality-reduced hyperspectral data subset, a spectral dimension consistency check is performed on the dimensionality-reduced hyperspectral data subset. Band data with abnormal spectral responses are removed, and band data that conform to the spectral characteristics are retained as the final dimensionality-reduced dataset.

10. A rapid band selection and dimensionality reduction system for hyperspectral data cubes, used to implement the rapid band selection and dimensionality reduction method for hyperspectral data cubes as described in any one of claims 1-9, characterized in that, Includes the following modules: The data acquisition module is used to acquire the hyperspectral data cube to be processed; The correlation analysis module is used to perform spectral-dimensional correlation analysis on hyperspectral data cubes and generate correlation matrices between bands. The band clustering module is used to perform clustering processing on all bands based on the correlation matrix to obtain several band clusters; The band selection module is used to compare the typical spectral response range of the target geological object, calculate the response matching degree of each band in the band cluster, and select the optimal representative band corresponding to each band cluster. The complementarity evaluation module is used to evaluate the cross-cluster feature complementarity of each optimal representative band, remove the optimal representative bands with duplicate features, and obtain the set of optimal representative bands after cross-cluster feature complementarity screening. The data integration module is used to integrate the best representative band set after cross-cluster feature complementarity screening to form a dimensionality-reduced hyperspectral data subset.

Citation Information

Patent Citations

  • Hyperspectral image waveband selection method, device and system

    CN113486876A

  • Intelligent waveband selection method for hyperspectral image target detection

    CN118212514A

  • InSAR hidden danger point automatic identification method based on hotspot analysis

    CN120446951A

  • Automobile part defect detection method based on multispectrum

    CN121904040A