Hyperspectral signal prediction method based on hyperspectral remote sensing image and pseudo label guidance

By employing a cascaded correction algorithm and a pseudo-label-guided bi-branch fusion model, the problems of noise interference and low quality in remote sensing spectral data were solved, achieving high-precision soil composition prediction and improving the model's generalization ability and robustness.

CN121834477AActive Publication Date: 2026-04-10KUNMING UNIV OF SCI & TECH

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-03-13
Publication Date
2026-04-10

AI Technical Summary

Technical Problem

Existing technologies for predicting soil composition using remote sensing spectral data suffer from complex noise interference and low quality, making it difficult to guarantee the accuracy of predictions. Furthermore, traditional correction methods have limited ability to handle complex nonlinear spectral variations, and machine learning methods rely on manual feature extraction, resulting in insufficient model generalization ability and robustness.

Method used

A cascaded correction algorithm based on hyperspectral remote sensing imagery is used to correct the remote sensing spectral data. The model is trained using a pseudo-label-guided two-branch fusion model, and imaging is performed using the GAF method. Supervised contrast loss is introduced to improve model performance.

Benefits of technology

It effectively improved the quality of remote sensing spectral data, enhanced the prediction accuracy and robustness of the model, strengthened the ability to represent the characteristics of soil components, and significantly improved the accuracy of soil property inversion and model performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121834477A_ABST
    Figure CN121834477A_ABST
Patent Text Reader

Abstract

The invention relates to a hyperspectral signal prediction method based on a hyperspectral remote sensing image and pseudo label guidance, and belongs to the technical field of hyperspectral signal prediction. The method comprises the following steps: collecting remote sensing spectral data and soil spectral data, and carrying out primary preprocessing on the remote sensing spectral data; based on a designed series correction algorithm, correcting the remote sensing spectral data after the primary preprocessing by using the soil spectral data, and performing secondary preprocessing; performing imaging on the remote sensing spectrum data after the secondary preprocessing; constructing a double-branch fusion model comprising a spectrum branch and an image branch; constructing a classification pseudo label, and introducing supervised contrast loss to train the double-branch fusion model; and inputting remote sensing spectral data to be predicted into the trained double-branch fusion model to obtain a hyperspectral signal prediction result so as to realize soil content prediction. The objective of the invention is to solve the technical problems of complex noise, low quality and difficulty in guaranteeing prediction accuracy when remote sensing spectral data is acquired in the prior art.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application relates to a hyperspectral signal prediction method based on hyperspectral remote sensing images and pseudo-label guidance, and belongs to the technical field of hyperspectral signal prediction. BACKGROUND

[0002] Red soil regions benefit from excellent natural climate conditions and have great land resource production potential. However, the fertility of red soil is usually low and easy to degrade. The highland acid red soil is formed under the combined influence of subtropical monsoon climate and complex topography, and has heavy texture and limited nutrients. The contents of soil organic matter (SOM) and nitrogen (N) can directly affect soil fertility and crop productivity by reducing soil bulk density, increasing total porosity and enhancing soil particle heterogeneity. Therefore, it is significant for precision agriculture, land resource management and environmental change research to quickly and widely accurately quantify soil organic carbon and total nitrogen and obtain spatial distribution information.

[0003] The main process of the traditional method for detecting soil content includes investigation, sampling and chemical measurement. Although high accuracy can be achieved, the high cost limits the application of the method. The traditional method mainly obtains distribution information based on spatial autocorrelation interpolation method. However, in the highland red soil region which is different in topography, land use and human activities, it is difficult to obtain accurate detection due to spatial heterogeneity. Studies have shown that visible and near-infrared spectroscopy (VIS-NIRS) has been proved to be a suitable method for SOM prediction in the laboratory, which can effectively establish a direct relationship between soil spectral data and soil component content within a specific wavelength (350-2500 nm), and is beneficial to the rapid and accurate estimation of soil organic matter and total nitrogen.

[0004] Satellite spectral data provides a valuable data source for soil property monitoring at a macro scale. However, satellite spectral data is disturbed by various factors in practical application, including atmospheric scattering and absorption, differences in sensor spectral response functions, mixed pixel effects, and observation geometry changes, resulting in significant deviations between the spectral response and the true surface reflectance characteristics. If such uncorrected satellite spectra are directly used to build inversion models, the generalization ability is poor, and the prediction accuracy is limited. In contrast, laboratory or near-ground hyperspectral measurements are carried out in a controlled environment, effectively avoiding atmospheric effects and pixel mixing problems, and can obtain pure, continuous and high signal-to-noise ratio spectral information of the target ground object, with high measurement accuracy and repeatability, and are often regarded as a reliable reference. Therefore, many studies use ground spectral data to correct satellite spectra to improve prediction accuracy. Although these studies have achieved good performance, traditional single correction methods, such as direct standardization (DS), can effectively correct global multiplicative bias, but have limited ability to handle complex nonlinear spectral changes. While piecewise direct standardization (PDS) improves local adaptability through piecewise processing, but has limitations in global spectral deformation correction.

[0005] Existing studies generally use machine learning (ML) and deep learning (DL) methods to build an implicit mapping function between soil properties and spectral data. Many studies use machine learning methods to build models. Although this method achieves high prediction accuracy to some extent, it relies on manual feature extraction, and the model generalization ability and robustness are limited. In order to improve the model generalization ability and robustness, and solve the problem of soil spatial heterogeneity and data redundancy, many studies use deep model building methods and have achieved better performance. These methods mostly ignore the inherent class differences and distribution structure between different soil characteristic samples in the modeling process, and usually assume that the samples are independent and identically distributed, thereby limiting the full learning of discriminative features by the model. SUMMARY

[0006] The purpose of the present application is to provide a hyperspectral signal prediction method based on hyperspectral remote sensing image and pseudo-label guidance, aiming to solve the technical problems of existing technology in obtaining remote sensing spectral data with complex noise and low quality, and difficulty in ensuring prediction accuracy.

[0007] To achieve the above purpose, the technical scheme of the present application is as follows: a hyperspectral signal prediction method based on hyperspectral remote sensing image and pseudo-label guidance, comprising the following steps: Step 1: Collect remote sensing spectral data and soil spectral data, and perform initial preprocessing on the remote sensing spectral data; wherein the initial preprocessing includes radiation correction, geometric correction and atmospheric correction; Step2: correcting the remotely sensed spectral data after the first preprocessing based on a designed serial correction algorithm using the soil spectral data, and performing secondary preprocessing on the corrected remotely sensed spectral data; wherein the serial correction algorithm is a serial use of a direct correction algorithm and a segmented direct correction algorithm, and the secondary preprocessing includes standard normal transformation, Savitzky-Golay smoothing, and first-order derivation; Step3: imaging the remotely sensed spectral data after the secondary preprocessing by a GAF method to obtain the imaged remotely sensed spectral data; Step4: constructing a double-branch fusion model including a spectral branch and an image branch; wherein the spectral branch is used to extract spectral features of the remotely sensed spectral data after the secondary preprocessing, the image branch is used to extract image features of the imaged remotely sensed spectral data, and the fusion features obtained by fusing the spectral features and the image features are taken as the output of the double-branch fusion model; Step5: constructing a classification pseudo-label, and introducing a supervised contrast loss based on the classification pseudo-label to train the double-branch fusion model; Step6: inputting the remotely sensed spectral data to be predicted into the trained double-branch fusion model to obtain a hyperspectral signal prediction result output by the double-branch fusion model, and realizing soil content prediction based on the hyperspectral signal prediction result.

[0008] Optionally, the correction of the remotely sensed spectral data after the first preprocessing based on the designed serial correction algorithm specifically includes: performing first correction on the remotely sensed spectral data after the first preprocessing by a direct correction algorithm using the soil spectral data, and the first correction specifically includes: training global linear transformation parameters between the soil spectral data and the remotely sensed spectral data after the first preprocessing by a partial least squares (PLS) method to establish a preliminary mapping between the ground and the remotely sensed spectral data, and the expression is:

[0009] wherein, is the remotely sensed spectral data after the correction by the direct correction algorithm, is a sample matrix composed of all the remotely sensed spectral data to be corrected, , are conversion matrix coefficients and residual matrices respectively trained by the partial least squares method; performing second correction on the remotely sensed spectral data after the first correction by a segmented direct correction algorithm using the soil spectral data, and the second correction specifically includes: establishing a local connection between the soil spectral data and the remotely sensed spectral data after the first correction to eliminate the intensity error between the ground and the remotely sensed spectral data, and the expression is:

[0010] wherein, is the reflectance value of the i-th band after correction by the piecewise direct correction algorithm, is the reflectance value of the i-th band after correction by the piecewise direct correction algorithm, represents the radius of the sliding window, represents the spectral reflectance value of the i-th band in the represents the local conversion weight from the i-th band to the j-th band obtained by the partial least squares method, , is the total number of bands and the local intercept term of the i-th band, respectively; Thus, the expression of the designed serial correction algorithm is:

[0011]

[0012] wherein, T is the band conversion matrix obtained by training the piecewise direct correction algorithm, is the remote sensing spectral data after correction by the designed serial correction algorithm, is the intercept vector of all bands.

[0013] Optionally, the imaging of the secondary preprocessed remote sensing spectral data by the GAF method is specifically: First, the secondary preprocessed remote sensing spectral data is aggregated by the GAF method to reach a preset imaging scale , and the original bands are replaced by average bands; Second, a normalization operation is performed on the values of the aggregated remote sensing spectral data to ensure the stability of the values; Then, the aggregated average bands are used as the radius , the inverse cosine of the aggregated reflectance value corresponding to each average band is used as the angle , and are used as the polar coordinates of one average band of the aggregated remote sensing spectral data; Finally, a matrix is constructed by using the corresponding to the and are the i-th and j-th average bands, respectively; and the value of a pixel in the image is .​​​​​​​​​ the angle corresponding to the first average wave band and the angle corresponding to the second average wave band. the angle corresponding to the first average wave band and the angle corresponding to the second average wave band.

[0014] Optionally, the Step 4 is specifically: using a one-dimensional convolution module as a feature extraction module of the spectrum branch, using a larger size of the convolution kernel in a shallow structure of the spectrum branch to extract the wave peak information of the remotely sensed spectrum data after the secondary preprocessing, and using a preset number of one-dimensional convolution modules to extract semantic information to obtain spectrum features, wherein the larger size is greater than a preset size; using a series of down-sampling two-dimensional convolution modules as a feature extraction module of the image branch, all two-dimensional convolution modules are designed as a smaller size of the convolution kernel to quickly expand the receptive field and extract the temporal prior between the wave bands, and the image branch is designed as a shallower network structure to avoid overfitting, thereby obtaining image features, wherein the smaller size is smaller than a preset size, and the shallower network structure is a network layer smaller than a preset layer number; channel splicing and fusing the spectrum features and the image features to obtain fusion features as an output of the double-branch fusion model.

[0015] Optionally, the constructing the classification pseudo label is specifically: determining that the distribution of the remotely sensed spectrum data satisfies a normal distribution based on a quantile-quantile diagram of the collected remotely sensed spectrum data, a skewness and a kurtosis of the distribution of the soil component label, and a histogram; selecting an equal-frequency binning discretization method according to the normal distribution to uniformly divide the remotely sensed spectrum data samples into four categories: first, sorting all remotely sensed spectrum data samples according to the soil component content, and then dividing the remotely sensed spectrum data samples into four categories according to the total amount of the remotely sensed spectrum data samples calculating the number of remotely sensed spectrum data samples contained in each bin , and finally adding a classification pseudo label to the sorted remotely sensed spectrum data samples.

[0016] Optionally, the training the double-branch fusion model based on the supervised contrastive loss of the classification pseudo label is specifically: selecting the feature of a certain sample as an anchor point to pull closer the features of the same label and push away the features of different categories in the feature space, and defining the classification pseudo label as , wherein, is a function of the content of the discrete label and represents the content of soil organic matter or soil total nitrogen, and the supervised contrastive loss of the coarse-grained classification pseudo label information is expressed as: , wherein, represents the supervised contrast loss calculated by taking the first sample as an anchor point, represents the index set of all samples, represents all other indexes with the same sample content level as the sample with index , wherein is the index set obtained after sampling, represent the classification pseudo-label of the first sample and the classification pseudo-label of the first sample, respectively, , represents all sampling indexes except the current index, represent the features corresponding to the first sample, the first sample and the first sample, respectively, is a temperature factor used to control the degree of nonlinear scaling of the similarity score; The obtained supervised contrast loss is introduced in the form of a weighted sum to obtain a total loss for training the dual-branch fusion model, and the total loss is The expression is: In the formula, is the mean absolute error loss used to supervise the difference between the predicted content and the real content, are weights used to adjust

[0017] The beneficial effects of the present application are: 1. The present application uses ASD spectral correction to correct remote sensing spectral data by designing a series correction algorithm, effectively improving the quality of remote sensing spectrum, and confirming the effectiveness of correcting remote sensing spectral data.

[0018] 2. The present application designs for one-dimensional spectral data with complete information and imaged spectral data with band-level time series information by constructing a dual-branch fusion model, improving the performance of the model.

[0019] 3. The present application assists the training of the dual-branch fusion model based on the pseudo-label guided method of supervised contrast learning, further improves the performance of the model by explicitly supervising the model to learn the invariant features of spectral signals and soil components. BRIEF DESCRIPTION OF DRAWINGS

[0020] Figure 1 is a step flowchart of the present application; ​​​​​​​​​​​​Figure 2 A schematic diagram of a sampling environment in a research area in an embodiment of the present application is shown in FIG. 1. Figure 3 A spectral curve diagram of remote sensing spectra after radiation correction, geometric correction and atmospheric correction in an embodiment of the present application is shown in FIG. 2. Figure 4 A spectral curve diagram of each stage of the serial correction algorithm in an embodiment of the present application is shown in FIG. 3. Figure 5 A spectral curve diagram of the corrected remote sensing data after standard normal transformation, SG smoothing and first-order derivative preprocessing in an embodiment of the present application is shown in FIG. 4. Figure 6 A data flow diagram of the dual-branch fusion model in an embodiment of the present application is shown in FIG. 5. DETAILED DESCRIPTION

[0021] The content of the present application will be further described below in combination with the drawings and specific embodiments, but this does not limit the scope of the present application.

[0022] Embodiment 1: As shown in FIG. 6, a hyperspectral signal prediction method based on hyperspectral remote sensing image and pseudo-label guidance is provided, which comprises the following steps: Figure 1 Firstly, a spectral two-dimensional image converted based on the Gram angle field (GAF) technology is used to establish a convolutional neural network (CNN) dual-branch model of the spectrum and the two-dimensional image to solve the problem that the in-situ image corresponding to the satellite-borne spectrum is difficult to obtain. Then, a sample pseudo-label assisted training mechanism based on contrast learning is introduced in the training stage. By constructing a pseudo-class constraint with different soil characteristic levels, the model is guided to explicitly perceive and learn the difference characteristics between samples during the training process, thereby enhancing the distinguishability and robustness of the feature representation and further improving the accuracy and model performance of soil property inversion. Step 1: Collect remote sensing spectral data and soil spectral data, and perform initial preprocessing on the remote sensing spectral data; wherein the initial preprocessing includes radiation correction, geometric correction and atmospheric correction; Optionally, based on a stratified sampling strategy, representative pixel positions are selected in the research area of the remote sensing spectral data collected by the ZY-102D satellite, soil samples corresponding to the point positions are collected in the field, and the organic matter and total nitrogen contents are determined. At the same time, the ASD ground object spectrometer is used to synchronously acquire the soil spectral data of each sampling point, wherein the remote sensing spectral data are used to establish a quantitative inversion model between soil properties and spectral features, and the data acquired by the ground object spectrometer are used for subsequent correction of the remote sensing image to improve the quality of the spectral data.

[0023] It can be understood that the embodiment Step1 eliminates various distortions attached to the radiation brightness in the image data by radiation correction, eliminates or corrects the geometric error of the remote sensing image by geometric correction, and eliminates the radiation error caused by atmospheric absorption and scattering by atmospheric correction, so that the complex interference received by the remote sensing spectral data is eliminated by the initial preprocessing.

[0024] Step2: correcting the remote sensing spectral data after the initial preprocessing by using the soil spectral data based on the designed serial correction algorithm, and performing secondary preprocessing on the corrected remote sensing spectral data; wherein the serial correction algorithm is a serial use of a direct correction algorithm (DS) and a segmented direct correction algorithm (PDS), and the secondary preprocessing includes standard normal transformation, Savitzky-Golay smoothing and first-order derivation. It should be understood that the core idea of the serial correction algorithm is to first establish a preliminary mapping relationship between the spectral systems by a global model, and then finely correct the prediction residual by using a localized model for each band, so as to realize effective conversion of satellite spectral data to ground spectral standard.

[0025] Optionally, the correction of the remote sensing spectral data after the initial preprocessing by using the soil spectral data based on the designed serial correction algorithm is specifically: The first correction of the remote sensing spectral data after the initial preprocessing by using the soil spectral data through the direct correction algorithm is specifically: The global linear transformation parameters between the soil spectral data and the remote sensing spectral data after the initial preprocessing are trained by the partial least squares method PLS, a preliminary mapping between the ground and the remote sensing spectral data is established, and the expression is:

[0026] wherein, is the remote sensing spectral data after the correction by the direct correction algorithm, is a sample matrix composed of all the to-be-corrected remote sensing spectral data, , are the conversion matrix coefficients and the residual matrix respectively trained by the partial least squares method; The second correction of the remote sensing spectral data after the first correction by using the soil spectral data through the segmented direct correction algorithm is specifically: The local relationship between the soil spectral data and the remote sensing spectral data after the first correction is established to finely correct the spectrum, so as to eliminate the intensity error between the ground and the remote sensing spectrum, and the expression is:

[0027] wherein, is the remote sensing spectral data after the correction by the segmented direct correction algorithm, Reflectance values ​​for each band, This represents the radius of the sliding window. express The Middle Spectral reflectance values ​​for each band, This indicates the band obtained using the partial least squares method. to band Local transformation weights, , These are the total number of bands and the number of... Local intercept term for each band; Therefore, the expression for the series correction algorithm of the design is:

[0028]

[0029] in, T This is the strip transformation matrix obtained by training using the piecewise direct correction algorithm. For remote sensing spectral data after the designed cascade correction algorithm, This is the intercept vector for all bands.

[0030] It is understandable that the cascaded correction algorithm designed in this embodiment performs correction by establishing a linear relationship between the corrected band and the entire ground band. By mining the covariance information between the entire band, preliminary spectral mapping is achieved. The PDS algorithm adopts a sliding window strategy, establishing a local correction model for each target band of the remote sensing spectrum, separately matching its corresponding band value with the ground ASD spectrum. The first band is used to extract remote sensing spectral data from its adjacent bands as independent variables, with the first band of the ground ASD spectrum as the independent variable. Using the band values ​​as dependent variables, PLS regression is performed to establish a linear relationship between the corrected bands and some terrestrial bands, allowing for more flexible and accurate spectral correction.

[0031] It is understood that this embodiment eliminates the systematic errors introduced by sensor response and environmental factors through the designed series correction algorithm, that is, it eliminates the influence of signal coupling, low spatial resolution and low spectral resolution that exist during remote sensing spectral data acquisition, improves the quality of remote sensing spectral data, and thus effectively improves the reliability and usability of remote sensing spectral data; then the corrected remote sensing spectral data is preprocessed a second time to reduce the noise in the remote sensing spectral data and enhance the signal-to-noise ratio.

[0032] Step 3: Image the preprocessed remote sensing spectral data using the GAF method to obtain the imaged remote sensing spectral data. Optionally, in order to introduce band-level temporal information into the subsequently constructed dual-branch fusion model and fully leverage the performance advantages of convolutional neural networks (CNNs) in extracting local features to improve the model's inversion accuracy, Step 3 of this embodiment uses the GAF method to image the remote sensing spectral data after secondary preprocessing, specifically: First, the preprocessed remote sensing spectral data are aggregated using the GAF method to achieve the preset imaging scale. And replace the original bands with the average bands; Secondly, a normalization operation is applied to the numerical values ​​of the aggregated remote sensing spectral data to ensure numerical stability. Then, the converged average band is used as the radius using the GAF method. The inverse cosine of the converged reflection value corresponding to each average band is used as the angle. and will ( , The polar coordinates of one of the average bands of the aggregated remote sensing spectral data are used as the reference. Finally, using The average band corresponds to indivual A matrix is ​​constructed, which is then used to generate an image; where the value of a pixel in the image is... for , and The first The angle corresponding to the first average band and the first The angle corresponding to each average band.

[0033] Step 4: Construct a dual-branch fusion model that includes a spectral branch and an image branch; wherein, the spectral branch is used to extract the spectral features of the remote sensing spectral data after secondary preprocessing, the image branch is used to extract the image features of the remote sensing spectral data after imaging, and the fused feature obtained by fusing the spectral features and the image features is used as the output of the dual-branch fusion model; Optionally, in Step 4 of this embodiment, a network structure with different branches is designed, and the features extracted from different branches are fused to enhance the representation of spectral features, specifically as follows: One-dimensional convolutional modules are used as feature extraction modules for spectral branches. Larger-sized convolutional kernels are used in the shallow structure of spectral branches to extract peak information of remote sensing spectral data after secondary preprocessing. At the same time, a preset number of one-dimensional convolutional modules are used to extract semantic information to obtain spectral features. The larger size is greater than the preset size. The image branch is composed of a series of down-sampling two-dimensional convolution modules as a feature extraction module, all the two-dimensional convolution modules are designed to have a small size of convolution kernel to quickly expand the receptive field and extract the temporal prior between bands, and the image branch is designed to have a shallow network structure to avoid overfitting, so as to obtain image features, wherein the small size is smaller than a preset size, and the shallow network structure has a number of network layers smaller than a preset number of network layers; The spectral features and the image features are channel spliced and fused to obtain fused features as the output of the double-branch fusion model.

[0034] It can be understood that, by the design of the double-branch structure in Step 4, the spectral branch can fully extract deep physical and chemical features in complete spectral information, and the image branch can assist in extracting temporal information between bands. Channel splicing and fusion of the two extracted features can effectively enhance the feature representation of the spectral signal, and thus enhance the performance of the model.

[0035] Step 5: constructing a classification pseudo-label, and training the double-branch fusion model based on a supervised contrastive loss introduced based on the classification pseudo-label; Optionally, the constructing a classification pseudo-label specifically includes: Based on the quantile-quantile diagram of the collected remote sensing spectral data, the skewness and kurtosis of the distribution of the soil component label, and the histogram, it is determined that the distribution of the remote sensing spectral data satisfies a normal distribution; According to the equal-frequency binning discretization method selected according to the normal distribution, the remote sensing spectral data samples are uniformly divided into four categories: first, all remote sensing spectral data samples are sorted according to the soil component content, and then the remote sensing spectral data samples are divided into four categories according to the total amount of the remote sensing spectral data samples The number of remote sensing spectral data samples contained in each bin is calculated , and finally, classification pseudo-labels are added to the sorted remote sensing spectral data samples.

[0036] Optionally, the double-branch fusion model is trained based on a supervised contrastive loss (SCL) introduced based on the classification pseudo-label, which explicitly supervises the model to learn the invariance of features between different categories of samples in the feature space, so as to improve the generalization and robustness of the model, and specifically includes: The features of a certain sample are selected as anchor points to pull the features of the same label closer and push the features of different categories farther apart in the feature space, and the classification pseudo-label is defined as , wherein, is a function of the content of the discrete label and represents the content of soil organic matter or soil total nitrogen, and the supervised contrastive loss of the coarse-grained classification pseudo-label information The expression is: in, Indicates the first The supervised contrast loss is calculated using 1 sample as the anchor point. Represents the set of indices for all samples. , indicating that the index is All other indices with the same sample content level, where, This is the set of indices obtained after sampling. and They represent the first The classification pseudo-label of the first sample and the first The pseudo-label of each sample. This indicates all sampled indexes except the current index. , and They respectively represent the corresponding number The sample, the first The first sample and the first Features of each sample This is a temperature factor used to control the degree of non-linear scaling of the similarity score; The supervised contrastive loss is introduced in a weighted sum manner to obtain the total loss for training the two-branch fusion model. The expression is: In the formula, The mean absolute error loss is used to monitor the difference between predicted and actual content. and For adjustment and The weight.

[0037] Understandably, the supervision signals commonly used in model training (such as mean absolute error (MAE) are only explicit supervision, which is not conducive to extracting noise-invariant features that are independent of soil content from spectral signals. Therefore, Step 5 of this embodiment improves the model inversion performance by constructing classification pseudo-labels and combining them with supervised contrastive learning to explicitly supervise the model in learning invariant features. This solves the following two problems: First, the spectral signals of samples with similar soil composition often exhibit diversity due to signal coupling, low spectral resolution of remote sensing satellite acquisition equipment, and low spatial resolution of remote sensing satellite acquisition equipment. Second, in natural scenes, samples with extreme content levels are often few, which constitutes a long-tailed distribution of labels, and the model may not be able to fully learn the knowledge of these samples.

[0038] Step 6: Input the remote sensing spectral data to be predicted into the trained dual-branch fusion model to obtain the hyperspectral signal prediction result output by the dual-branch fusion model, and realize soil content prediction based on the hyperspectral signal prediction result.

[0039] Based on the specific implementation details, the effectiveness of the technical solution of this invention will be demonstrated through experiments. The specific process is as follows: S1: Soil organic matter and total nitrogen content of some pixels in remote sensing spectral data were collected based on stratified sampling, and ground spectra were collected using a ground object spectral acquisition instrument; Specifically, this experiment extracted 266 pixels from the remote sensing spectral data through stratified sampling, such as... Figure 2 As shown, this displays information about the study area and sampling points.

[0040] S2: Radiometric correction, geometric correction, and atmospheric correction were performed on the remote sensing spectral data to further improve the spectral quality; Specifically, radiometric correction eliminates radiometric distortion in remote sensing spectral data caused by sensor characteristics, atmospheric transmission, and solar radiation. This aims to convert the relative measurements recorded by the sensor into calibrated reflectance or emissivity values. Geometric correction is used to eliminate geometric distortions, and atmospheric correction is employed to eliminate the influence of atmospheric substances such as water vapor, oxygen, carbon dioxide, methane, and ozone on ground object reflectance. The 1325nm-1474nm and 1779nm-1962nm bands were removed, resulting in 139 bands. The spectrum obtained after various corrections is shown below. Figure 3 As shown, the curves of different colors represent different samples.

[0041] S3: Based on ground spectral data, the remote sensing spectral data is corrected using a series correction algorithm, and the corrected remote sensing spectral data is preprocessed using standard normal transformation, Savitzky-Golay (SG) smoothing, and first-order derivative method. Specifically, such as Figure 4 All corrected data were displayed. Figure 5 The image shows the corrected satellite spectral data after secondary preprocessing, with different colored curves representing different samples.

[0042] S4: In order to make the distribution of samples in the training set as uniform as possible in the sample space so as to make full use of effective data to learn generalized knowledge, this experiment uses the Kennard-Stone algorithm (KS algorithm) to divide the spectral data after secondary preprocessing into training set and test set in a ratio of 7:3. Then, the class pseudo-label of the training set is obtained by using the equal frequency bin discretization method, and the Gram angle field (GAF) technology is used to image the spectral data and use it as a joint dataset with the spectral data. Specifically, the KS algorithm first selects the sample farthest from the sample mean as the starting point, and then iteratively selects the remaining sample with the largest minimum Euclidean distance from the selected sample set until a specified number of training samples are selected, thereby ensuring that the training set is evenly distributed in the feature space.

[0043] S5: A two-branch fusion model is constructed as the model for predicting soil content. Subsequently, a pseudo-label-guided method based on contrastive learning is used to eliminate distributional bias between samples and improve the predictive performance of the method. The structure and data flow of the two-branch fusion model are as follows: Figure 6 As shown. Finally, the coefficient of determination (COP) of the model trained on the training set is calculated on the test set. ), root mean square error ( ), relative percentage difference ( Three metrics are used to evaluate the model's performance.

[0044] Furthermore, this experiment constructed lightweight convolutional neural network models for one-dimensional spectral data and two-dimensional spectral imaging data, respectively. In the spectral branch, larger convolutional kernels and deeper layers were used to capture rich features, while in the image branch, the natural advantages of convolutional modules were used to construct simple branch networks to avoid overfitting. Finally, feature fusion was performed through connections. Table 1 shows the specific composition of the dual-branch fusion model proposed in this embodiment.

[0045] Table 1: Model Structure Table

[0046] Furthermore, this experiment introduced supervised contrastive loss based on the calibrated and preprocessed training set data, using pseudo-labels to guide the training process, and then obtained performance indicators on the test set. Tables 2 and 3 show the final results of different models, respectively.

[0047] Table 2: Performance of Soil Organic Matter Prediction Model

[0048] Table 3: Performance of Soil Total Nitrogen Prediction Model

[0049] Specifically, in Tables 2 and 3, S-Model represents the model that uses only the spectral branch of the two-branch fusion model for inversion, I-Model represents the model that uses only the image branch of the two-branch fusion model for inversion, F-Model represents the two-branch fusion model, and SCLF-Model represents the two-branch fusion model that uses the SCL method for auxiliary training.

[0050] As shown in Tables 2 and 3, the F-Model, a fusion model using two branches, further improves performance compared to single-branch networks (S-Model and I-Model). Specifically, compared to the best-performing single-branch network, I-Model, the quantitative model for SOM components shows improvements of 0.0166, 0.2364, and 0.0307 in R², RMSE, and RPD, respectively; while the quantitative model for N components shows improvements of 0.0108, 7.7544, and 0.0195. Furthermore, the SCLF-Model, a fusion model using the SCL method for assisted training, further improves performance compared to the F-Model. Analysis indicates that both the two-branch feature fusion and contrastive learning strategies are effective. To demonstrate the performance advantage of the proposed method, the proposed two-branch fusion model is compared with other commonly used deep learning models. The comparative experimental results are shown in Tables 4 and 5.

[0051] Table 4: Comparison of different models in soil organic matter retrieval task

[0052] Table 5: Comparison of different models in soil total nitrogen retrieval task

[0053] As shown in Tables 4 and 5, the SCLF-Model proposed in this invention achieves the best performance. In the soil organic matter inversion task, it achieves R², RMSE, and RPD scores of 0.6022, 11.6296, and 1.5854, respectively; in the soil total nitrogen inversion task, it achieves scores of 0.6284, 57, 0.5907, and 1.6404, respectively. Compared to the lightweight MobileNetV2 model, it only requires 0.12M more parameters and 0.14G more computing power in the soil organic matter inversion task, resulting in performance improvements of 35.5%, 15.4%, and 18.1% in the organic matter inversion task, and 19.56%, 11.49%, and 12.98% in the total nitrogen inversion task. These results fully demonstrate the effectiveness of the proposed model, outperforming one-dimensional models using single spectral data, two-dimensional models using image data transformed from single spectral data, and machine models.

[0054] Furthermore, comparative experiments on different models in soil organic matter inversion and soil total nitrogen inversion tasks fully demonstrated the advantages of the fusion model based on coarse-grained contrastive learning: not only are the number of parameters and required computing power sufficiently low, but the performance also reached the best. In both soil component inversion tasks, the dual-branch fusion model F-Model outperformed the single-branch model. This result proves that fusing spectral data and image data based on Gram angle field imaging can leverage the natural advantage of convolutional neural networks in extracting image features, further improving the performance of convolutional neural networks. Simultaneously, the results comparing deep models with machine models also fully illustrate the advantage of deep learning in automatically extracting features.

[0055] In summary, remote sensing spectral data is subject to various interferences in practical applications, leading to significant deviations between its spectral response and the actual surface reflectance characteristics. Furthermore, the distribution offset between samples is a major factor affecting the performance of remote sensing spectral prediction models. Therefore, to improve the performance of prediction models, this invention proposes a hyperspectral signal prediction method based on hyperspectral remote sensing imagery and pseudo-label guidance. This method first uses a concatenated correction algorithm to correct the remote sensing spectral data based on the ground spectral data of the sampling points to improve spectral quality. Second, to leverage the advantages of convolutional neural networks, this invention constructs a dual-branch fusion model, using the corresponding images of the sample spectra as input based on sequential imaging technology. Finally, pseudo-labels for the sample categories are generated through a discrete method, and supervised contrastive learning loss is introduced to improve the performance of the prediction method. Experimental results show that the method of this invention can effectively improve spectral quality, significantly enhance the prediction performance of the prediction method, and simultaneously maintain real-time performance and lightweight design.

[0056] The specific embodiments of the present invention have been described in detail above with reference to the accompanying drawings. However, the present invention is not limited to the above embodiments. Within the scope of knowledge possessed by those skilled in the art, various changes can be made without departing from the spirit of the present invention.

Claims

1. A method for predicting hyperspectral signals based on hyperspectral remote sensing imagery and pseudo-tag guidance, characterized in that, The method includes the following steps: Step 1: Collect remote sensing spectral data and soil spectral data, and perform initial preprocessing on the remote sensing spectral data; wherein, the initial preprocessing includes radiometric correction, geometric correction and atmospheric correction; Step 2: Based on the designed tandem correction algorithm, the soil spectral data is used to correct the remote sensing spectral data after the initial preprocessing, and the corrected remote sensing spectral data is then subjected to secondary preprocessing. The tandem correction algorithm is a combination of a direct correction algorithm and a piecewise direct correction algorithm. The secondary preprocessing includes standard normal transformation, Savitzky-Golay smoothing, and first-order differentiation. Step 3: Image the preprocessed remote sensing spectral data using the GAF method to obtain the imaged remote sensing spectral data. Step 4: Construct a dual-branch fusion model that includes a spectral branch and an image branch; wherein, the spectral branch is used to extract the spectral features of the remote sensing spectral data after secondary preprocessing, the image branch is used to extract the image features of the remote sensing spectral data after imaging, and the fused feature obtained by fusing the spectral features and the image features is used as the output of the dual-branch fusion model; Step 5: Construct classification pseudo-labels, and train the dual-branch fusion model based on the classification pseudo-labels by introducing supervised contrast loss; Step 6: Input the remote sensing spectral data to be predicted into the trained dual-branch fusion model to obtain the hyperspectral signal prediction result output by the dual-branch fusion model, and realize soil content prediction based on the hyperspectral signal prediction result.

2. The hyperspectral signal prediction method based on hyperspectral remote sensing imagery and pseudo-tag guidance according to claim 1, characterized in that, The design-based cascade correction algorithm uses the soil spectral data to correct the remote sensing spectral data after the initial preprocessing, specifically as follows: The initial correction of the preprocessed remote sensing spectral data is performed using soil spectral data and a direct correction algorithm. The initial correction specifically involves: A global linear transformation parameter between soil spectral data and pre-processed remote sensing spectral data was trained using partial least squares (PLS) to establish a preliminary mapping between the ground and remote sensing spectra. The expression is as follows: ; in, The remote sensing spectral data after correction using the direct correction algorithm. The sample matrix consists of all the remote sensing spectral data to be corrected. , These are the transformation matrix coefficients and residual matrix obtained through partial least squares training, respectively. A secondary correction is performed on the remote sensing spectral data after the initial correction using soil spectral data and a segmented direct correction algorithm. The secondary correction specifically involves: To establish a local correlation between soil spectral data and the remote sensing spectral data after initial correction, and to eliminate intensity errors between the ground and remote sensing spectra, the expression is as follows: ; in, The first part after correction by the piecewise direct correction algorithm Reflectance values ​​for each band, This represents the radius of the sliding window. express The Middle Spectral reflectance values ​​for each band, This indicates the band obtained using the partial least squares method. to band Local transformation weights, , These are the total number of bands and the number of... Local intercept term for each band; Therefore, the expression for the series correction algorithm of the design is: ; ; in, T This is the strip transformation matrix obtained by training using the piecewise direct correction algorithm. For remote sensing spectral data after the designed cascade correction algorithm, This is the intercept vector for all bands.

3. The hyperspectral signal prediction method based on hyperspectral remote sensing imagery and pseudo-tag guidance according to claim 1, characterized in that, The imaging of the preprocessed remote sensing spectral data using the GAF method specifically involves: First, the preprocessed remote sensing spectral data are aggregated using the GAF method to achieve the preset imaging scale. And replace the original bands with the average bands; Secondly, a normalization operation is applied to the numerical values ​​of the aggregated remote sensing spectral data to ensure numerical stability. Then, the converged average band is used as the radius using the GAF method. The inverse cosine of the converged reflection value corresponding to each average band is used as the angle. and will ( , The polar coordinates of one of the average bands of the aggregated remote sensing spectral data are used as the reference. Finally, using The average band corresponds to indivual A matrix is ​​constructed, which is then used to generate an image; where the value of a pixel in the image is... for , and The first The angle corresponding to the first average band and the first The angle corresponding to each average band.

4. The hyperspectral signal prediction method based on hyperspectral remote sensing imagery and pseudo-tag guidance according to claim 1, characterized in that, Step 4 specifically refers to: One-dimensional convolutional modules are used as feature extraction modules for spectral branches. Larger-sized convolutional kernels are used in the shallow structure of spectral branches to extract peak information of remote sensing spectral data after secondary preprocessing. At the same time, a preset number of one-dimensional convolutional modules are used to extract semantic information to obtain spectral features. The larger size is greater than the preset size. The feature extraction module for the image branch is composed of a series of downsampled two-dimensional convolutional modules. All two-dimensional convolutional modules are designed with small kernel size to quickly expand the receptive field and extract temporal priors between bands. At the same time, the image branch is designed with a shallow network structure to avoid overfitting, thereby obtaining image features. The smaller size is smaller than a preset size, and the shallower network structure is a network layer with fewer than a preset number of layers. The spectral features and the image features are channel-wise concatenated and fused to obtain the fused features as the output of the dual-branch fusion model.

5. The hyperspectral signal prediction method based on hyperspectral remote sensing imagery and pseudo-tag guidance according to claim 1, characterized in that, The specific steps for constructing the classification pseudo-labels are as follows: Based on the quantile-quantile plots of the collected remote sensing spectral data, the skewness and kurtosis of the distribution of soil component labels, and histograms, it was determined that the distribution of the remote sensing spectral data follows a normal distribution. Based on the normal distribution, an equal-frequency binning discretization method was selected to uniformly divide the remote sensing spectral data samples into four categories: First, all remote sensing spectral data samples were sorted according to the soil component content; then, based on the total number of remote sensing spectral data samples... Calculate the number of remote sensing spectral data samples contained in each bin. Finally, classification pseudo-labels are added to the sorted remote sensing spectral data samples.

6. The hyperspectral signal prediction method based on hyperspectral remote sensing imagery and pseudo-tag guidance according to claim 5, characterized in that, The specific steps for training the dual-branch fusion model using supervised contrastive loss based on the classification pseudo-labels are as follows: By selecting a feature of a sample as an anchor point in the feature space, features with the same label are brought closer together and features of different categories are pushed further apart. The classification pseudo-label is defined as follows: ,in, It is a function of the sum of discrete label contents. The supervised comparison loss, representing the content of soil organic matter or total nitrogen, utilizes coarse-grained classification pseudo-label information. The expression is: ; in, Indicates the first The supervised contrast loss is calculated using 1 sample as the anchor point. Represents the set of indices for all samples. , indicating that the index is All other indices with the same sample content level, where, This is the set of indices obtained after sampling. and They represent the first The classification pseudo-label of the first sample and the first The pseudo-label of each sample. This indicates all sampled indexes except the current index. , and They respectively represent the corresponding number The sample, the first The first sample and the first Features of each sample This is a temperature factor used to control the degree of non-linear scaling of the similarity score; The supervised contrastive loss is introduced in a weighted sum manner to obtain the total loss for training the two-branch fusion model. The expression is: ; In the formula, The mean absolute error loss is used to monitor the difference between predicted and actual content. and For adjustment and The weight.

Citation Information

Patent Citations

  • Hyperspectral image classification method and device based on spatial-spectral double-branch convolutional network

    CN115249332A

  • LIBS-NIR dual-spectrum integrated system for soil component detection and detection method

    CN119064310A

  • Semi-supervised multispectral remote sensing image scene classification method, device, equipment and medium

    CN120014370A

  • Soil heavy metal inversion method and system integrating satellite remote sensing and near-end sensing

    CN120847006A

  • Object-oriented method for identifying and classifying surface lithology in hyperspectral remote sensing image

    US20250209814A1

Cited By

  • Remote sensing image relative radiation calibration method and system based on quantile

    CN122042562A