A hyperspectral unmixing method for detecting soluble salt content in murals
By standardizing sample preparation and spectral data preprocessing, and combining LASSO regression and elastic network models, the problem of spectral unmixing of salt and pigment in murals was solved, achieving high-precision non-destructive testing, adapting to complex scenarios, and providing accurate salt content detection.
Patent Information
- Application Number
- CN202510536888.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-27
- Publication Date
- 2025-12-02
- Estimated Expiration
- 2045-04-27
AI Technical Summary
Existing hyperspectral detection methods are difficult to effectively separate the spectral signals of salt and pigment on the surface of murals, limiting detection accuracy. Furthermore, traditional detection methods are highly destructive and difficult to implement non-destructive testing.
By standardizing sample preparation, preprocessing spectral data, and screening characteristic wavelengths, and combining LASSO regression model and elastic network regression model, the mixing regularization is optimized to achieve unmixing of salt and pigment spectra.
It achieves high-precision non-destructive testing of the soluble salt content of murals, adapts to complex scenarios of pigment and salt mixing, outputs reliable testing indicators, and provides accurate data support for cultural relic protection.
Smart Images

Figure CN120609755B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of mural testing technology, and in particular to a hyperspectral unmixing method for detecting the soluble salt content in murals. Background Technology
[0002] In the field of cultural heritage preservation, murals, as important carriers of historical and cultural heritage, are significantly affected by the erosion of soluble salts. The accumulation and migration of soluble salts (such as sodium chloride, calcium chloride, sodium sulfate, and sodium nitrate) in murals can lead to damage such as powdering of the ground layer and peeling of the pigment layer. Therefore, accurately detecting the soluble salt content on the surface of murals is crucial for developing scientific conservation plans. Traditional detection methods (such as conductivity methods and ion chromatography) require sampling analysis, which is destructive and inefficient, making it difficult to achieve in-situ, non-destructive testing of murals.
[0003] Hyperspectral imaging technology, with its advantages of speed, non-destructive operation, and multi-band information fusion, has been increasingly applied to the detection of salt damage in murals. However, actual mural surfaces are typically covered with multiple pigments (such as cinnabar, azurite, and malachite), whose spectral signals overlap and interfere with the salt spectra, making it difficult to demix the mixed spectra. Existing hyperspectral detection methods suffer from problems such as redundant characteristic wavelength selection and insufficient model overfitting resistance when processing complex mixed spectra, making it difficult to effectively separate the spectral signals of salt and pigments, thus limiting detection accuracy.
[0004] Furthermore, the preparation process of mural samples needs to simulate the real structure (such as coarse mud layer, fine mud layer, and pigment layer) and salt gradient distribution. However, existing methods lack sufficient standardized control over sample preparation, which may lead to a deviation in the matching degree between training data and actual murals, affecting the model's generalization ability. In the data preprocessing stage, if the breakpoint correction, noise reduction, and dimensionality reduction of the spectral curves are not refined enough, additional errors will be introduced, further reducing the reliability of detection.
[0005] Current technologies lack targeted unmixing model optimization strategies when dealing with multi-component mixed salts (such as multiple salts coexisting in different proportions), making it difficult to accurately invert the content of complex salt combinations. Furthermore, the spatial distribution visualization and error quantification assessment systems for detection results are still imperfect, failing to provide precise spatial positioning and risk assessment basis for conservation and restoration. Therefore, there is an urgent need for a method for detecting the soluble salt content of murals that can effectively handle mixed spectra, adapt to complex scenarios, and possess both high accuracy and robustness, to meet the practical needs of cultural relic conservation. Summary of the Invention
[0006] This invention overcomes the problems of traditional detection methods, such as damaging the integrity of murals, difficulty in handling spectral overlap interference, and insufficient accuracy. Through standardized sample preparation, spectral data preprocessing, characteristic wavelength selection, and optimization of the hybrid regularization model, it achieves non-destructive and efficient detection of salt content on the mural surface. It exhibits low detection error, supports demixing in complex pigment-salt mixture scenarios, and outputs reliable detection indicators, providing accurate data support for cultural relic protection and restoration.
[0007] To achieve the above objectives, the present invention adopts the following solution:
[0008] A hyperspectral unmixing method for detecting the soluble salt content in murals, comprising the following steps:
[0009] S1: Prepare multiple standard mural sample blocks with different soluble salt contents, wherein the soluble salt is formed by uniformly mixing sodium chloride, calcium chloride, sodium sulfate and sodium nitrate in a preset mass ratio. The standard mural sample block includes a coarse mud layer and a fine mud layer. The coarse mud layer is coated on the base after mixing sandy soil and reinforcing material. The fine mud layer is coated on the coarse mud layer after mixing fine soil and soluble salt solution. A white powder layer composed of calcium carbonate and talc is coated on the fine mud layer, and various pigments are uniformly coated on the white powder layer.
[0010] S2: Multiple reflectance spectra were sampled for each standard mural sample block using a spectrometer. The initial reflectance spectrum curve was corrected for breakpoints and outliers were removed. The average spectrum of multiple samples was taken as the mixed reflectance spectrum and smoothed to obtain standardized mixed reflectance spectrum data. The mixed reflectance spectrum data was standardized and the first k principal components were extracted by principal component analysis to reduce the data dimensionality.
[0011] S3: The LASSO regression model combined with 5-fold cross-validation is used to screen sparse feature wavelengths related to salt concentration from mixed reflectance spectral data. The objective function of the LASSO regression model is set to minimize the sum of squared residuals and the sum of L1 regularization terms, and the optimal regularization parameter is determined by cross-validation.
[0012] S4: Input the selected sparse feature wavelengths into the elastic network regression model, and train the unmixed model by combining L1 / L2 mixed regularization and 5-fold cross-validation. The objective function of the elastic network regression model is set to minimize the sum of squared residuals and the sum of L1 / L2 mixed regularization terms, and the optimal combination of regularization parameters is determined by cross-validation.
[0013] S5: Extract spectral values from the actual mural's hyperspectral data according to the selected characteristic wavelengths, standardize the data, and then input them into the trained unmixed model to output the soluble salt content of each area on the mural surface.
[0014] Preferably, in step S1, the following method is used when making the standard mural sample block:
[0015] The total salt content of the sample blocks was configured as a percentage of 100g soil mass, including 21 gradients ranging from 0% to 1% at 0.05% intervals; the preset mass ratio of each soluble salt to the total salt mass was: sodium chloride 40%, calcium chloride 20%, sodium sulfate 20%, and sodium nitrate 20%.
[0016] The coarse mud layer is formed by mixing sandy soil with wheat straw and deionized water after filtration through a sieve to form dry, hard mud, and is then evenly coated onto the substrate surface with a thickness of 1 cm. The fine mud layer is formed by mixing fine-grained soil with deionized water and soluble salt solution after filtration through a sieve, and is then coated onto the surface of the coarse mud layer with a thickness of 0.5 cm. The white powder layer consists of calcium carbonate and talc with a mass ratio of 10%-20%. Four pigments—cinnabar, azurite, malachite, and graphite—are uniformly coated onto the surface of the completely dried sample block in parallel strips, with each pigment area being 2.5 cm wide and 0.2 cm thick.
[0017] During the natural air-drying process of the sample blocks, the temperature and humidity were monitored using soil testing instruments, and the temperature and humidity were adjusted in real time to ensure that the cultivation conditions of each sample block were consistent.
[0018] Preferably, the specific method for reflectance spectrum sampling and processing in step S2 includes:
[0019] The reflectance spectra of each standard mural sample block were sampled four times at 24-hour intervals in a darkroom using a spectrometer. Whiteboard calibration was performed before each sampling, and outliers caused by point edges or measurement jitter were removed. The initial reflectance spectrum curves of the four samplings were averaged after breakpoint correction to obtain the mixed reflectance spectrum curve.
[0020] Further noise reduction is achieved through Savitzky-Golay smoothing, using the formula... Standardization processing is performed to obtain standardized mixed reflectance spectral data. Where X is the spectral data matrix before processing, μ is the mean of each wavelength, and σ is the standard deviation of each wavelength;
[0021] When performing dimensionality reduction using PCA principal component analysis, the covariance matrix is constructed. Calculate the eigenvalues λ of the covariance matrix C. C With eigenvector ν C Where m is the sample size; based on the cumulative variance contribution rate Extract the top k principal components corresponding to the cumulative variance contribution rate reaching a preset threshold, and construct matrix V from these k principal components. k Standardized mixed reflectance spectral data Projecting to a lower-dimensional space yields a dimensionality-reduced data matrix. ,in It is the sum of the variances of the first k principal components; It is the total variance of all principal components.
[0022] Preferably, in step S3, the specific methods for constructing and using the LASSO regression model include:
[0023] The objective function is set as follows: Where m is the sample size. This represents the true salt concentration of sample i; Let be the standardized value of sample i on the j-th principal component; Let be the regression coefficient of the j-th principal component; To control sparsity, the L1 regularization strength parameter was used. Through 5-fold cross-validation, the optimal regularization parameter λ that minimizes the model error was selected. All regression coefficients were calculated and the characteristic wavelengths corresponding to non-zero regression coefficients were screened out.
[0024] In step S4, the specific methods for constructing and using the elastic network regression model include:
[0025] The objective function is set as follows Where α is the mixing ratio parameter of L1 and L2 regularization, and λ1 and λ2 are the strengths of L1 regularization and L2 regularization, respectively. The mean square error (MSE) of different combinations of α, λ1, and λ2 parameters is evaluated using 5-fold cross-validation. The formula for calculating MSE is: Where N is the number of test samples, For the predicted salt concentration of the i-th sample, select the optimal parameter combination that minimizes the MSE, and fit the final model using the complete training set; the final regression equation is: , where X i Let ε be the standardized value of the i-th sample, and ε be the model error term.
[0026] Preferably, the method for unmixing model verification and salt content output in step S5 includes:
[0027] The spectral values of the actual mural hyperspectral data were extracted according to the selected characteristic wavelengths and standardized using the same standardized parameters as the training data. The standardized hyperspectral data were then input into the trained unmixed model to calculate the soluble salt content of each area on the mural surface.
[0028] Calculate the coefficient of determination R 2 The mean square error (MSE), root mean square error (RMSE), and mean absolute error (MAE) are calculated as follows: , , , ; where Y 真实 Y represents the actual measured salt concentration value. 预测 The salt concentration value is predicted by the unmixing model. The average salt concentration is denoted by ; N is the number of test samples. and These are the actual and predicted salt concentration values for the i-th sample, respectively.
[0029] Evaluate the goodness of fit of the unmixed model: If If the values of MSE, RMSE, and MAE are all lower than the corresponding preset thresholds, the unmixing model is deemed valid and a salt content distribution map is output.
[0030] Preferably, the processing of the initial reflectance spectrum curve in step S2 includes the following operations:
[0031] The abrupt changes in the spectral curve were smoothed using piecewise linear interpolation, with an interpolation interval of 5 nm. The correction formula is as follows:
[0032]
[0033] Where R(λ) is the original reflectivity at wavelength λ, R 校正 (λ) represents the corrected reflectivity;
[0034] Outlier data points deviating from the mean by more than three times the standard deviation were removed using the 3σ principle.
[0035] Savitzky-Golay smoothing parameters were optimized using a filter with a window width of 15nm and a polynomial order of 3. The smoothing effect was verified by calculating the signal-to-noise ratio (SNR) until the SNR was ≥ 30dB.
[0036] Preferably, the feature wavelength screening of the LASSO regression model in step S3 further includes:
[0037] The wavelength range was segmented, and the spectral range of 400-2500nm was divided into 10 sub-intervals with a span of 210nm. LASSO regression was performed independently in each sub-interval.
[0038] A significance t-test was performed on the selected characteristic wavelengths, with a significance level of 0.01, and only wavelengths with a probability value of p-value < 0.01 that satisfies the null hypothesis were retained.
[0039] Redundant wavelengths are removed, and the Pearson correlation coefficient r of all characteristic wavelength pairs is calculated. If |r| of a characteristic wavelength pair is greater than 0.9, the two characteristic wavelengths are determined to be a highly correlated wavelength group. Only the wavelength with the largest absolute value of the regression coefficient is retained in each wavelength group.
[0040] The screening results of all sub-intervals are merged to form a sparse feature wavelength set, and the corresponding salt sensitivity weights are recorded. ,β j The regression coefficients for the j-th wavelength in the LASSO regression model are used to remove wavelengths with a weight less than 0.05, thus generating the final feature set.
[0041] Preferably, the optimization of the elastic network regression model in step S4 includes:
[0042] The mixed regularization parameters are dynamically adjusted. The initial value of the mixing ratio parameter α of L1 and L2 regularization is set to 0.5, and the search range of L1 regularization intensity λ1 and L2 regularization intensity λ2 is
[10] . −4 10 2 The optimal parameter combination is determined by combining grid search with 5-fold cross-validation, and the optimization objective is set to minimize the mean square error (MSE) of the validation set.
[0043] During training, a Dropout mechanism is incorporated to randomly mask 20% of the input features. Training is repeated 10 times, and the mean of the regression coefficients is taken as the final solution. A non-negativity constraint X is added. j β j ≥0 and abundance sum to 1 constrain ∑β j =1, X i Let β be the standardized value of the i-th sample. j is the regression coefficient for the j-th wavelength in the LASSO regression model.
[0044] Preferably, the unmixing model verification in step S5 further includes:
[0045] The actual mural hyperspectral data was divided into a 0.5cm×0.5cm pixel grid, and the average reflectance of the selected characteristic wavelengths within each grid was extracted to ensure that the spatial scale was consistent with that of the training samples.
[0046] The salt content Y of each pixel is output through the demixing model. i A thermal map of the spatial distribution of salt in the entire mural was generated using the Kriging interpolation method and a Gaussian semivariogram model.
[0047] Calculate the prediction confidence interval for each pixel. SE represents the standard error. Let be the predicted salt concentration value of the i-th sample. If the CI width exceeds 0.1%, the region is marked as a region that needs to be reviewed. Data is then re-acquired and the model is updated for the region that needs to be reviewed.
[0048] Preferably, the standard error SE is calculated through the following steps:
[0049] Constructing regression coefficients β to reflect the elasticity of the network regression model j The covariance matrix Cov(β) of the relationship between them Where X is the characteristic wavelength data matrix, σ 2 Let I be the variance of the training set residuals, I be the identity matrix, and D be the L2 regularization weight matrix, where α is the mixed ratio parameter of L1 and L2 regularization, and λ1 and λ2 are the L1 regularization strength and L2 regularization strength, respectively.
[0050] Calculate the prediction variance: , where X i Let σ be the feature wavelength vector of the i-th pixel. 2 For the training set residual variance;
[0051] The reflectance of the perturbation feature wavelength is sampled using Monte Carlo sampling with a perturbation amplitude of ±1%, and the standard deviation of the salt content is statistically output as SE. If more than 10% of the pixels in the entire mural area have a CI width exceeding 0.1%, the model recalibration process is triggered, and the hyperspectral data of that area is reacquired and the unmixed model is updated.
[0052] The present invention includes at least the following beneficial effects: (1) By screening sparse feature wavelengths through LASSO regression and unmixing the elastic network regression model, the mixed spectral signals of salt and pigment are effectively separated, avoiding the destructive sampling of traditional methods, and achieving high-precision non-destructive detection of soluble salt content in murals with low detection error; (2) By presetting the salt gradient (0%-1%) and standardizing the sample preparation process (layer coating of coarse mud layer, fine mud layer and pigment layer), the reliability of training data is ensured; with the addition of breakpoint correction, Savitzky-Golay smoothing and PCA dimensionality reduction, the quality of spectral data is improved, noise interference is reduced, and high-quality input is provided for model training; (3) By eliminating redundant wavelengths through piecewise LASSO regression and significance test, combined with L1 / L2 mixture regularization of the elastic network model, the problem of multiple pigment coatings is effectively solved. The spectral overlap problem under the cover is adapted to various salt combinations (mixed salts such as sodium chloride and calcium chloride) and complex pigment environments, improving the robustness of the model; (4) a thermal map of salt spatial distribution is generated by Kriging interpolation to intuitively display the salt-rich areas; the model accuracy is quantified by introducing indicators such as the coefficient of determination R² and mean square error MSE, and high uncertainty areas (areas that need to be checked if the confidence interval exceeds the standard) are marked, providing accurate spatial positioning and risk assessment basis for protection and restoration; (5) the model’s anti-overfitting ability is improved by the Dropout mechanism and parameter grid search, and non-negativity constraints (salt content ≥ 0) and abundance constraints (total salt ≤ 100%) are added to ensure that the unmixing results are in line with physical meaning; a model recalibration mechanism is set (such as triggering when the confidence interval of more than 10% of pixels exceeds the standard) to adapt to environmental changes such as mural aging and maintain detection accuracy in the long term. Attached Figure Description
[0053] Figure 1 The present invention provides a flowchart of the principle of a hyperspectral unmixing detection method for soluble salt content in murals. Detailed Implementation
[0054] The present invention will now be described in further detail with reference to the accompanying drawings, so that those skilled in the art can implement it based on the description.
[0055] like Figure 1 As shown, the present invention provides a hyperspectral unmixing detection method for soluble salt content in murals, comprising the following steps:
[0056] S1: Prepare multiple standard mural sample blocks with different soluble salt contents. The soluble salt is formed by uniformly mixing sodium chloride, calcium chloride, sodium sulfate, and sodium nitrate in a preset mass ratio. The standard mural sample block includes a coarse mud layer and a fine mud layer. The coarse mud layer is coated on the base after mixing sandy soil and reinforcing materials. The fine mud layer is coated on the coarse mud layer after mixing fine soil and soluble salt solution. A white powder layer composed of calcium carbonate and talc is coated on the fine mud layer, and various pigments are uniformly coated on the white powder layer.
[0057] Soluble salts can be a mixture of sodium chloride, calcium chloride, sodium sulfate, and sodium nitrate. Based on the different salt contents in murals at different stages and of different types, the common salt types and their proportions in murals are simulated. The mixing ratio design is based on chemical analysis of actual salt damage cases to ensure that the sample covers typical salt combinations.
[0058] The sample structure was manufactured in layers from the base to the surface: the coarse mud layer, 1 cm thick, was formed by mixing sandy soil (with larger particle size) with wheat straw and deionized water to create a dry, hard base; the sandy soil provided structural support, the wheat straw enhanced crack resistance, and the deionized water prevented interference from impurities. The fine mud layer, 0.5 cm thick, was formed by mixing fine-grained soil with a soluble salt solution; the fine-grained soil simulated the ground layer of a mural, and the salt solution permeated evenly to simulate salt migration. The pigment coating was applied by uniformly coating the sample surface with common mural pigments, such as cinnabar, azurite, malachite, and graphite, to simulate the multi-pigmentation scene of a real mural. The pigment thickness was controlled during testing to ensure that the spectral reflectance was consistent with that of a real mural.
[0059] During the air-drying process, the environment of the finished product is controlled. Temperature and humidity are monitored using soil testing instruments (e.g., temperature 20-25℃, humidity 40-60%) to ensure a consistent drying rate and avoid uneven salt distribution.
[0060] During sample preparation, sample blocks were prepared according to the established salinity gradient, with at least six samples prepared for each gradient. Circular molds were used to standardize sample dimensions, and a scraper was used to ensure uniform layer thickness. After the mud layer was allowed to air dry naturally, pigment was applied in sections to avoid pigment mixing and interference with spectral acquisition.
[0061] S2: Use a spectrometer to sample the reflectance spectrum of each standard mural sample block multiple times, perform breakpoint correction and outlier removal on the initial reflectance spectrum curve, take the average spectrum of multiple samplings as the mixed reflectance spectrum, and smooth it to obtain standardized mixed reflectance spectrum data; perform standardized processing on the mixed reflectance spectrum data, and extract the first k principal components through principal component analysis to reduce data dimensionality.
[0062] Configure spectral acquisition conditions, using a darkroom environment to eliminate external light interference, and preheat the spectrometer for 30 minutes to ensure stability. Calibrate the instrument baseline using a whiteboard calibration method, and perform four samplings of each sample block at 24-hour intervals, averaging the results to reduce random errors.
[0063] Data preprocessing includes:
[0064] Breakpoint correction: Identify wavelength points with abrupt changes in reflectivity (such as those caused by instrument vibration or sample surface defects), and smooth them through interpolation. The interpolation interval can be set to 5-10nm.
[0065] Outlier removal: Filter outlier data that deviates too much from the mean to ensure a smooth spectral curve.
[0066] Savitzky-Golay smoothing: Based on the actual smoothing algorithm settings, a filter with a window width of 15nm and a polynomial order of 3 can be used to eliminate high-frequency noise, thereby improving the signal-to-noise ratio to over 30dB.
[0067] Standardization processing: The spectral data is normalized to a mean of 0 and a standard deviation of 1 to eliminate dimensional differences and facilitate subsequent model training.
[0068] PCA dimensionality reduction: Extract the top k principal components with a cumulative variance contribution rate of ≥95%, compress the data dimensionality, and retain the key spectral features of salt and pigment.
[0069] The spectrometer collects data in preset wavelength bands, with each sampling covering the entire wavelength range. After preprocessing, a standardized spectral matrix is generated and used as input for subsequent models.
[0070] S3: The LASSO regression model combined with 5-fold cross-validation is used to screen sparse feature wavelengths related to salt concentration from mixed reflectance spectral data. The objective function of the LASSO regression model is set to minimize the sum of squared residuals and the sum of L1 regularization terms, and the optimal regularization parameter is determined by cross-validation.
[0071] The LASSO model uses an L1 regularization penalty term to compress the regression coefficients of non-critical wavelengths to zero, thus identifying wavelengths significantly correlated with salt concentration. Cross-validation optimization employs 5-fold cross-validation, dividing the data into training and validation sets to test different feature parameter values and select the parameter that minimizes the validation error. Wavelengths corresponding to non-zero regression coefficients are retained, forming a sparse feature set for outputting characteristic wavelengths. For example, this filters out the characteristic absorption band of sodium chloride (around 1450 nm). Standardized data is input into the LASSO model, iteratively optimizing the parameter λ. The output list of characteristic wavelengths and their weights can be filtered using methods such as Pearson correlation coefficient detection, removing redundant wavelengths with high correlation.
[0072] S4: Input the selected sparse feature wavelengths into the elastic network regression model, and train the unmixed model by combining L1 / L2 mixed regularization and 5-fold cross-validation. The objective function of the elastic network regression model is set to minimize the sum of squared residuals and the sum of L1 / L2 mixed regularization terms, and the optimal combination of regularization parameters is determined by cross-validation.
[0073] The elastic network regression model combines L1 regularization (sparseness) and L2 regularization (stability) to address multicollinearity in spectral data (such as overlap between pigment and salt bands). Parameters are dynamically adjusted through grid search testing of α (mixing ratio, e.g., 0.3-0.7), λ1 (L1 intensity), and λ2 (L2 intensity), selecting the combination that minimizes the mean squared error (MSE). Physical constraints are employed: non-negativity constraints (salt content ≥ 0) and abundance summation of 1 (total salt ≤ 100%) are added to ensure the unmixing results reflect reality. The model is trained using LASSO-selected feature wavelength data to obtain the unmixing model. A dropout mechanism (randomly masking 20% of features) is used to enhance overfitting resistance, and the model is trained 10 times and the average is taken to improve stability.
[0074] S5: Extract spectral values from the actual mural's hyperspectral data according to the selected characteristic wavelengths, standardize the data, and then input them into the trained unmixed model to output the soluble salt content of each area on the mural surface.
[0075] Actual mural data is standardized according to training set parameters to ensure model input consistency. Predicted results are mapped to a 0.5cm × 0.5cm pixel grid, and a heatmap is generated using an interpolation algorithm to visually display the spatial distribution of salinity. The coefficient of determination R² (greater than 0.9 is preferred), mean square error MSE (less than 0.05), and root mean square error RMSE (less than 0.2) are calculated to verify model reliability. Areas requiring review can be resampled to update model parameters. A salinity content report and distribution map are output to guide restoration decisions.
[0076] The above algorithm achieves high-precision detection of salt content in murals: by combining LASSO with an elastic network model, it separates salt signals from mixed spectra, achieving a detection error less than 20% of traditional methods (such as the conductivity method). It employs non-destructive analysis, requiring no sampling throughout the process, thus protecting the integrity of the murals and making it suitable for precious cultural heritage. It can adapt to complex scenarios, is compatible with various pigments and salt mixtures, resolves spectral overlap interference, and exhibits strong robustness. The model provides visualized output, converting salt content into spatial distribution heatmaps and error quantification indicators, providing accurate data support for restoration.
[0077] In another technical solution, the following method is used when making the standard mural sample block in step S1:
[0078] The total salt content of the sample blocks was configured as a percentage of 100g soil mass, including 21 gradients ranging from 0% to 1% at 0.05% intervals; the preset mass ratio of each soluble salt to the total salt mass was: sodium chloride 40%, calcium chloride 20%, sodium sulfate 20%, and sodium nitrate 20%.
[0079] The coarse mud layer is formed by mixing sandy soil with wheat straw and deionized water after filtration through a sieve to form dry, hard mud, and is then evenly coated onto the substrate surface with a thickness of 1 cm. The fine mud layer is formed by mixing fine-grained soil with deionized water and soluble salt solution after filtration through a sieve, and is then coated onto the surface of the coarse mud layer with a thickness of 0.5 cm. The white powder layer consists of calcium carbonate and talc with a mass ratio of 10%-20%. Four pigments—cinnabar, azurite, malachite, and graphite—are uniformly coated onto the surface of the completely dried sample block in parallel strips, with each pigment area being 2.5 cm wide and 0.2 cm thick.
[0080] During the natural air-drying process of the sample blocks, the temperature and humidity were monitored using soil testing instruments, and the temperature and humidity were adjusted in real time to ensure that the cultivation conditions of each sample block were consistent.
[0081] The salinity gradient is designed with 21 gradients: 0%, 0.05%, 0.1%, 0.15%, 0.2%, 0.25%, 0.3%, 0.35%, 0.4%, 0.45%, 0.5%, 0.55%, 0.6%, 0.65%, 0.7%, 0.75%, 0.8%, 0.85%, 0.9%, 0.95%, and 1%, covering a typical range from no salt to high salt concentration. Each gradient corresponds to a different salt mass (e.g., 0.2g salt / 100g soil), which is precisely weighed and mixed (e.g., 40% sodium chloride, 20% calcium chloride, etc.) to simulate the chemical environment of salt migration in real murals. The coarse mud layer is made by filtering coarse sandy soil (particle diameter approximately 1-2mm) through a sieve and mixing it with wheat straw (2-3cm in length) and deionized water (conductivity ≤10μS / cm) to form a dry, hard mud layer. The wheat straw enhances crack resistance, and the deionized water prevents impurities from interfering with salt distribution. Apply a 1cm thick layer, ensuring the substrate is flat. The fine mud layer is created by mixing fine-grained soil (particle diameter ≤0.5mm) with a soluble salt solution (concentrations prepared in gradients), with a thickness of 0.5cm. The fine-grained soil simulates the ground layer structure of a mural, while the salt solution penetrates evenly to simulate capillary migration. The salt mixture must be stirred until there are no lumps and allowed to stand for 12 hours to eliminate air bubbles. Apply the coating using a standard scraper, completing it in two stages (coarse mud layer → air-dry for 24 hours → fine mud layer).
[0082] A white powder layer is evenly coated onto the fine layer to form a white pigment base, physically isolating the pigment layer from the salt layer during pigment application and preventing mixing that could interfere with spectral acquisition. The sample environment is monitored in real time using soil testing instruments (such as temperature and humidity sensors), with temperature controlled at 20-25℃ (target value 22.5°C ± 2.5℃) and humidity at 40-60% (target value 50% ± 10%). If the temperature or humidity deviates from the target range, ventilation or humidification devices are activated for adjustment. The pigment is prepared to a suitable viscosity (like honey), and applied in one direction with a soft brush to avoid bubbling. Temperature and humidity data are recorded every 2 hours; alarms are triggered and manual intervention is required if the values deviate from the target range.
[0083] Standardized molds were used, employing circular iron molds (10cm×10cm×4cm) to uniformly size samples and reduce the impact of shape differences on spectral reflectance. Natural air-drying took approximately 7-10 days, with samples turned daily to ensure uniform drying. Moisture content was measured using an infrared moisture meter until it dropped below 5%. Standardized training data was provided to improve the model's generalization ability. The salt distribution of the samples was ensured to match that of the actual murals, minimizing detection errors.
[0084] In another technical solution, the specific method for reflectance spectrum sampling and processing in step S2 includes:
[0085] The reflectance spectra of each standard mural sample block were sampled four times at 24-hour intervals in a darkroom using a spectrometer. Whiteboard calibration was performed before each sampling, and outliers caused by point edges or measurement jitter were removed. The initial reflectance spectrum curves of the four samplings were averaged after breakpoint correction to obtain the mixed reflectance spectrum curve.
[0086] Further noise reduction is achieved through Savitzky-Golay smoothing, using the formula... Standardization processing is performed to obtain standardized mixed reflectance spectral data. Where X is the spectral data matrix before processing, μ is the mean of each wavelength, and σ is the standard deviation of each wavelength;
[0087] When performing dimensionality reduction using PCA principal component analysis, the covariance matrix is constructed. Calculate the eigenvalues λ of the covariance matrix C. C With eigenvector ν C Where m is the sample size; based on the cumulative variance contribution rate Extract the top k principal components corresponding to the cumulative variance contribution rate reaching a preset threshold, and construct matrix V from these k principal components. k Standardized mixed reflectance spectral data Projecting to a lower-dimensional space yields a dimensionality-reduced data matrix. ,in It is the sum of the variances of the first k principal components; It is the total variance of all principal components.
[0088] The spectrometer was calibrated in a darkroom to eliminate stray light, and the light source was preheated for 30 minutes to stabilize. Whiteboard calibration (reflectivity ≥95%) was performed to calibrate the instrument baseline, and this calibration was repeated before each sampling. Spectral data were collected four times for each sample block at 24-hour intervals, covering reflectivity changes at different drying stages, and the average value was taken to reduce random errors. The sampling wavelength range was 400-2500 nm, with a resolution ≤5 nm, covering the visible to near-infrared bands containing key spectral information. Abrupt reflectivity changes (such as data jumps caused by surface cracks) were removed.
[0089] Savitzky-Golay smoothing uses a 15nm window with a polynomial order of 3 to eliminate high-frequency noise (such as instrument thermal noise) and improve the signal-to-noise ratio to ≥30dB. The mean μ and standard deviation σ are calculated for each wavelength, and the data is transformed to eliminate dimensional differences and prevent certain bands from dominating model training. After smoothing, the continuity of the spectral curve is checked, and linear interpolation is used to repair discontinuities. Standardized parameters (μ, σ) are stored for later use to ensure consistency between the actual data and the training set.
[0090] The correlation between wavelengths is analyzed by calculating the covariance matrix, and the principal components with the largest variance are extracted. The top k principal components are retained when the cumulative variance contribution rate is ≥95% (e.g., k is 5-10). Standardized data is projected into a low-dimensional space to reduce redundant information (such as pigment spectral noise) and highlight salt sensitivity characteristics. Reducing data dimensionality accelerates model training. This improves the signal-to-noise ratio of the salt signal and enhances detection accuracy.
[0091] In another technical solution, the specific method for constructing and using the LASSO regression model in step S3 includes:
[0092] The objective function is set as follows: Where m is the sample size. This represents the true salt concentration of sample i; Let be the standardized value of sample i on the j-th principal component; Let be the regression coefficient of the j-th principal component; To control sparsity, the L1 regularization strength parameter was used. Through 5-fold cross-validation, the optimal regularization parameter λ that minimizes the model error was selected. All regression coefficients were calculated and the characteristic wavelengths corresponding to non-zero regression coefficients were screened out.
[0093] In step S4, the specific methods for constructing and using the elastic network regression model include:
[0094] The objective function is set as follows Where α is the mixing ratio parameter of L1 and L2 regularization, and λ1 and λ2 are the strengths of L1 regularization and L2 regularization, respectively. The mean square error (MSE) of different combinations of α, λ1, and λ2 parameters is evaluated using 5-fold cross-validation. The formula for calculating MSE is: Where N is the number of test samples, For the predicted salt concentration of the i-th sample, select the optimal parameter combination that minimizes the MSE, and fit the final model using the complete training set; the final regression equation is: , where X i Let ε be the standardized value of the i-th sample, and ε be the model error term.
[0095] The LASSO regression model is constructed using L1 regularization optimization: sparsity is controlled by adjusting the λ value (range 0.01-10), and wavelengths strongly correlated with salt concentration are selected. Too large a λ value leads to excessive sparsity, while too small a λ value retains redundant wavelengths. Cross-validation parameter selection: 5-fold cross-validation divides the data into 5 groups, using 4 groups for training and 1 group for validation in rotation, selecting λ to minimize the validation set MSE.
[0096] Initially, λ=0.1, and gradually increase it to λ=1, observing the sparsity of the regression coefficients. Output a list of wavelengths with non-zero coefficients, such as selecting 1450nm (the characteristic band of sodium sulfate).
[0097] The optimization of the elastic network model employs a hybrid regularization parameter: α controls the L1 / L2 ratio (e.g., 0.5 balances sparsity and stability), and λ1 and λ2 adjust the penalty strength (search range 10⁻). 4 -10²).
[0098] A Dropout mechanism is introduced to combat overfitting. 20% of the input features are randomly masked, and the model is trained 10 times, with the mean of the regression coefficients taken to reduce its sensitivity to noise. A grid search test is performed with α=0.3, 0.5, 0.7, λ1=0.1, 1, 10, and λ2=0.01, 0.1, 1. The parameter combination with the minimum MSE (e.g., α=0.5, λ1=0.1, λ2=0.01) is selected.
[0099] Non-negativity constraints are used in physical constraints and model validation: the regression coefficient β is forced. j ≥0, ensuring the salinity prediction value is non-negative; and abundance and constraints: ∑β j =1, limiting the total salt content to 100%, which aligns with actual proportions. This addresses non-physical issues in spectral unmixing (such as negative salt values). It improves the model's reliability in practical mural detection.
[0100] In another technical solution, the method for unmixing model verification and salt content output in step S5 includes:
[0101] The spectral values of the actual mural hyperspectral data were extracted according to the selected characteristic wavelengths and standardized using the same standardized parameters as the training data. The standardized hyperspectral data were then input into the trained unmixed model to calculate the soluble salt content of each area on the mural surface.
[0102] Calculate the coefficient of determination R 2 The mean square error (MSE), root mean square error (RMSE), and mean absolute error (MAE) are calculated as follows: , , , ; where Y 真实 Y represents the actual measured salt concentration value. 预测 The salt concentration value is predicted by the unmixing model. The average salt concentration is denoted by ; N is the number of test samples. and These are the actual and predicted salt concentration values for the i-th sample, respectively.
[0103] Evaluate the goodness of fit of the unmixed model: If If the values of MSE, RMSE, and MAE are all lower than the corresponding preset thresholds, the unmixing model is deemed valid and a salt content distribution map is output.
[0104] In error metric calculations, the coefficient of determination (R²) measures the model's ability to explain variance; R² > 0.9 indicates a high degree of agreement between predicted and actual values. Mean squared error (MSE) quantifies prediction bias; MSE < 0.05 (salt concentration percentage) is considered excellent. MAE reflects the average error (e.g., < 0.1%), while RMSE amplifies the impact of large errors (e.g., < 0.2%). When calculating R², if the value is below 0.85, a model retraining process is triggered. The MSE threshold can be adjusted according to detection requirements (e.g., set to 0.03 for strict scenarios).
[0105] In model validity assessment, the thresholds are set as follows: |R²-1| < 0.05, MSE < 0.05, RMSE < 0.15, and MAE < 0.1. The salinity heatmap is divided into 0.5cm × 0.5cm pixel sections. Kriging interpolation is used to fill unsampled areas, and color gradations are used to indicate salinity gradients (e.g., blue for low salinity, red for high salinity). Pixels with a confidence interval width > 0.1% are marked and either manually resampled or locally resampled. The output report includes error statistics and distribution plots to support corrective action decisions.
[0106] In the dynamic calibration mechanism, the covariance matrix update mechanism involves periodically (e.g., monthly) adding new sample blocks and adjusting regression coefficients through incremental learning to adapt to environmental changes (such as mural aging). Recalibration trigger condition: if more than 10% of the region's CI width exceeds the limit, the model is fully reconstructed. This maintains detection accuracy over the long term and adapts to complex environments. It provides traceable error quantification data, enhancing the scientific rigor of restoration plans.
[0107] In another technical solution, the processing of the initial reflectance spectrum curve in step S2 includes the following operations:
[0108] The abrupt changes in the spectral curve were smoothed using piecewise linear interpolation, with an interpolation interval of 5 nm. The correction formula is as follows:
[0109]
[0110] Where R(λ) is the original reflectivity at wavelength λ, R 校正 (λ) represents the corrected reflectivity;
[0111] Outlier data points deviating from the mean by more than three times the standard deviation were removed using the 3σ principle.
[0112] Savitzky-Golay smoothing parameters were optimized using a filter with a window width of 15nm and a polynomial order of 3. The smoothing effect was verified by calculating the signal-to-noise ratio (SNR) until the SNR was ≥ 30dB.
[0113] Breakpoint correction identifies abrupt changes in the reflectance curve (such as those caused by instrument vibration or sample surface defects) and repairs them using piecewise linear interpolation. The interpolation interval can be set to 5-10 nm to ensure spectral continuity. For example, at wavelength λ, if the reflectance abrupt change within an adjacent 5 nm exceeds a threshold (e.g., ±0.02), the average reflectance on both sides is used to replace the outlier.
[0114] Savitzky-Golay smoothing uses a 15nm window width and a polynomial order of 3 filter to eliminate high-frequency noise (such as electronic noise or environmental interference). After smoothing, the signal-to-noise ratio (SNR) is calculated, requiring an SNR ≥ 30dB to ensure data quality meets model input requirements. After acquiring raw data using a spectrometer, abrupt changes are checked wavelength by wavelength, and correction curves are generated after interpolation. During smoothing, the window width can be adjusted according to noise intensity (e.g., increased to 25nm for high noise levels).
[0115] In outlier detection and removal, the 3σ principle is used: calculate the mean (μ) and standard deviation (σ) of reflectance for each wavelength, and remove data points that deviate from the range of μ ± 3σ. For example, if the reflectance of a certain wavelength is 0.5, μ = 0.45, and σ = 0.02, then values where 0.5 > 0.45 + 3 × 0.02 = 0.51 are not removed, but values where 0.6 > 0.51 are considered outliers. A dynamic threshold adjustment strategy is adopted, which can be relaxed to 4σ for high-noise bands (such as the near-infrared region) to improve fault tolerance. μ and σ are calculated segment by wavelength to avoid local data distortion caused by a global threshold. After removing outliers, missing points are re-interpolated to ensure the integrity of the spectral curve.
[0116] During data quality verification and output, signal-to-noise ratio (SNR) verification is performed: after smoothing, the SNR is calculated; if it does not reach 30 dB, the window width is increased or the number of smoothing iterations is increased. Data standardization is performed: data is normalized to a mean of 0 and a standard deviation of 1 by wavelength to eliminate instrument response differences. This improves the consistency of spectral data, providing high-quality input for subsequent model training. It also reduces noise interference and improves the accuracy of selecting characteristic wavelengths for salinity.
[0117] The feature wavelength selection for the LASSO regression model in step S3 further includes:
[0118] The wavelength range was segmented, and the spectral range of 400-2500nm was divided into 10 sub-intervals with a span of 210nm. LASSO regression was performed independently in each sub-interval.
[0119] A significance t-test was performed on the selected characteristic wavelengths, with a significance level of 0.01, and only wavelengths with a probability value of p-value < 0.01 that satisfies the null hypothesis were retained.
[0120] Redundant wavelengths are removed, and the Pearson correlation coefficient r of all characteristic wavelength pairs is calculated. If |r| of a characteristic wavelength pair is greater than 0.9, the two characteristic wavelengths are determined to be a highly correlated wavelength group. Only the wavelength with the largest absolute value of the regression coefficient is retained in each wavelength group.
[0121] The screening results of all sub-intervals are merged to form a sparse feature wavelength set, and the corresponding salt sensitivity weights are recorded. ,β j The regression coefficients for the j-th wavelength in the LASSO regression model are used to remove wavelengths with a weight less than 0.05, thus generating the final feature set.
[0122] Wavelength range segmentation and independent regression were employed to divide the 400-2500nm spectral range into 10 sub-intervals (each 210nm), such as 400-610nm and 610-820nm. After segmentation, LASSO regression was run independently within each sub-interval to prevent long-band data from overshadowing short-band features. Local feature extraction was performed, targeting the salt absorption characteristics of different sub-intervals (e.g., strong absorption of sodium chloride around 1450nm) to screen locally sensitive wavelengths. A regularization parameter λ was configured for each sub-interval to adapt to the data sparsity of different bands. The screening results of each sub-interval were merged to form a global feature set.
[0123] When performing significance testing and redundancy removal, a t-test was used for screening: a t-test was performed on the LASSO regression coefficients (significance level α = 0.01), retaining only wavelengths with p-values < 0.01 to ensure statistical significance of the screening results. Redundancy was removed based on the Pearson correlation coefficient: the correlation coefficient between characteristic wavelength pairs was calculated; if |r| > 0.9, it was considered highly correlated, and the wavelength with the largest absolute value of the regression coefficient was retained. For example, if the correlation coefficient between wavelengths A and B is 0.95, and the corresponding |β... A |>|β B If |, then B is removed. A sliding window is used to calculate the correlation between adjacent wavelengths to avoid high global computational complexity. After redundancy removal, the salt sensitivity weight of the retained wavelengths is recorded, and secondary wavelengths with a weight less than 0.05 are removed.
[0124] During the final feature set generation process, weights are sorted according to w. j Sort the wavelengths from highest to lowest, retaining the top N wavelengths (e.g., N=50) to balance the number of features with model complexity. Perform cross-interval validation to check the complementarity of wavelengths in different sub-intervals, avoiding the omission of key features. Improve the representativeness and robustness of the feature wavelength set, and reduce the impact of multicollinearity on model accuracy.
[0125] The optimization of the elastic network regression model in step S4 includes:
[0126] The mixed regularization parameters are dynamically adjusted. The initial value of the mixing ratio parameter α of L1 and L2 regularization is set to 0.5, and the search range of L1 regularization intensity λ1 and L2 regularization intensity λ2 is
[10] . −4 10 2 The optimal parameter combination is determined by combining grid search with 5-fold cross-validation, and the optimization objective is set to minimize the mean square error (MSE) of the validation set.
[0127] During training, a Dropout mechanism is incorporated to randomly mask 20% of the input features. Training is repeated 10 times, and the mean of the regression coefficients is taken as the final solution. A non-negativity constraint X is added. j β j ≥0 and abundance sum to 1 constrain ∑β j =1, X i Let β be the standardized value of the i-th sample. j is the regression coefficient for the j-th wavelength in the LASSO regression model.
[0128] In the optimization of hybrid regularization parameters, the search range for the parameters is set to 10⁻ for both L1 regularization strength λ1 and L2 regularization strength λ2. 4 Up to 10², covering all scenarios from weak to strong penalties. The initial value of the mixing ratio parameter α is set to 0.5 to balance sparsity and stability. A grid search strategy is adopted, dividing the parameter grid on a logarithmic scale (e.g., λ1=0.001,0.01,0.1,1,10), and using 5-fold cross-validation to select the optimal combination that minimizes the validation set MSE. α and λ1 are optimized first, with λ2=0.1 fixed for initial screening, and then λ2 is fine-tuned. The MSE of each parameter combination is recorded, and the scheme with the lowest error and moderate model complexity is selected.
[0129] The Dropout mechanism, used to combat overfitting and enhance stability, randomly masks 20% of the input features during training, forcing the model to learn features outside of redundant paths and improving generalization ability. Training is repeated 10 times, and the mean of the regression coefficients is taken as the final solution. Non-negativity constraint: This forces the regression coefficient β... j A value ≥0 ensures the predicted salinity is physically reasonable (e.g., no negative concentration). Masked features are randomly selected for each training iteration to ensure diversity across different training epochs. After adding constraints, the objective function is optimized using projective gradient descent.
[0130] Model validation and output are performed using abundance and constraints: ∑β j =1, limiting the total salt content to no more than 100%, consistent with actual proportions. Error feedback is implemented: if the validation set MSE does not reach the threshold (e.g., <0.05), the parameter search range is adjusted retrospectively or training data is increased. This improves the model's demixing capability for complex spectral mixing scenarios. The output results are ensured to conform to physical meaning and can be directly used for repair decisions.
[0131] In another technical solution, the unmixing model verification in step S5 further includes:
[0132] The actual mural hyperspectral data was divided into a 0.5cm×0.5cm pixel grid, and the average reflectance of the selected characteristic wavelengths within each grid was extracted to ensure that the spatial scale was consistent with that of the training samples.
[0133] The salt content Y of each pixel is output through the demixing model. i A thermal map of the spatial distribution of salt in the entire mural was generated using the Kriging interpolation method and a Gaussian semivariogram model.
[0134] Calculate the prediction confidence interval for each pixel. SE represents the standard error. Let be the predicted salt concentration value of the i-th sample. If the CI width exceeds 0.1%, the region is marked as a region that needs to be reviewed. Data is then re-acquired and the model is updated for the region that needs to be reviewed.
[0135] When matching spatial resolution and generating salt distribution, a pixel grid is created: the mural surface is divided into a 0.5cm × 0.5cm pixel grid to ensure the spatial resolution is consistent with the training samples (e.g., sample block size 10cm × 10cm). Kriging interpolation is used: a Gaussian semi-variogram model is employed to interpolate the salt values at sampling points to generate a continuous distribution heatmap. Parameters include range (e.g., 20cm), nugget effect (e.g., 0.1), and sill value (e.g., 0.5). The average reflectance of the feature wavelength is extracted for each pixel and input into the unmixing model to predict the salt value. During interpolation, neighboring pixel data is prioritized to avoid error propagation over long distances.
[0136] During the confidence interval calculation and verification marking process, the standard error (SE) is calculated as follows: The standard deviation of the predicted salinity is statistically analyzed using the reflectance of the characteristic wavelength perturbation (amplitude ±1%) obtained through Monte Carlo sampling. The confidence interval (CI) is set at a 95% confidence level as CI = predicted value ± 1.96 × SE. If the CI width > 0.1% (e.g., 0.12%), it is marked as an area requiring verification. Hyperspectral data is re-acquired for the marked area, and the unmixed model parameters are locally updated. Manual verification is performed on-site using a portable spectrometer to correct model biases.
[0137] The model and data are dynamically updated and visualized, using heatmap rendering and coloring according to salt concentration gradients (e.g., blue represents 0-0.3%, red represents 0.8-1.0%), which are then overlaid onto the digital model of the mural. The generated report includes error statistics, distribution maps, and review suggestions. It provides intuitive spatial distribution information on salt content to guide targeted restoration. Uncertainty quantification enhances the reliability of the test results.
[0138] The standard error SE is calculated through the following steps:
[0139] Constructing regression coefficients β to reflect the elasticity of the network regression model j The covariance matrix Cov(β) of the relationship between them Where X is the characteristic wavelength data matrix, σ 2 Let I be the variance of the training set residuals, I be the identity matrix, and D be the L2 regularization weight matrix, where α is the mixed ratio parameter of L1 and L2 regularization, and λ1 and λ2 are the L1 regularization strength and L2 regularization strength, respectively.
[0140] Calculate the prediction variance: , where X i Let σ be the feature wavelength vector of the i-th pixel. 2 For the training set residual variance;
[0141] The reflectance of the perturbation feature wavelength is sampled using Monte Carlo sampling with a perturbation amplitude of ±1%, and the standard deviation of the salt content is statistically output as SE. If more than 10% of the pixels in the entire mural area have a CI width exceeding 0.1%, the model recalibration process is triggered, and the hyperspectral data of that area is reacquired and the unmixed model is updated.
[0142] In the construction of the covariance matrix and error propagation processing, the covariance matrix Cov(β) based on the regression coefficients β_j of the elastic network reflects the uncertainty of parameter estimation. The calculation combines the residual variance σ² of the training set and the regularized weight matrix. The prediction variance formula is used based on Cov(β) and the characteristic wavelength vector X. i Calculate the prediction variance Var(Ŷ) for each pixel. i )=X i T ·Cov(β)·X i +σ². Numerical methods are used to approximate Cov(β), avoiding direct inversion of the high-dimensional matrix. Cov(β) is stored for subsequent predictions, reducing real-time computational overhead.
[0143] The perturbation amplitude for Monte Carlo sampling and sensitivity analysis was set to apply a ±1% random perturbation to the reflectivity of the characteristic wavelength to simulate the effects of measurement error and noise. The standard deviation was calculated by repeating the sampling 1000 times, using the standard deviation of the predicted salinity as the SE to quantify the model's sensitivity to input perturbations. The sampling process was parallelized, utilizing GPU acceleration for computation. An SE distribution map was output to identify regions of high uncertainty (such as edges or pigment mixing areas).
[0144] Model recalibration and closed-loop optimization are implemented. The recalibration trigger condition is: if the CI width of more than 10% of pixels in the entire mural is >0.1%, the model is deemed to require a global update. Incremental learning updates are performed after adding new sample blocks, and regression coefficients are adjusted using stochastic gradient descent, retaining historical data weights (e.g., old data weight 0.7, new data weight 0.3). This enables the detection system to self-optimize and adapt to environmental changes and material aging. Closed-loop control maintains high-precision detection over the long term, reducing the need for manual intervention.
[0145] It should be noted that although the steps are described in a specific order above, this does not mean that they must be performed in that order. In fact, some of these steps can be executed concurrently, or even in a different order, as long as the required functionality is achieved. The number of devices and processing scale described herein are for simplification of the invention; applications, modifications, and variations of this invention will be readily apparent to those skilled in the art.
[0146] Although embodiments of the present invention have been disclosed above, they are not limited to the applications listed in the specification and embodiments. They can be applied to various fields suitable for the present invention. For those skilled in the art, other modifications can be easily made. Therefore, without departing from the general concept defined by the claims and their equivalents, the present invention is not limited to the specific details and illustrations shown and described herein.
Claims
1. A method for detecting the soluble salt content of murals using hyperspectral unmixing, characterized in that, Includes the following steps: S1: Prepare multiple standard mural sample blocks with different soluble salt contents, wherein the soluble salt is formed by uniformly mixing sodium chloride, calcium chloride, sodium sulfate and sodium nitrate in a preset mass ratio. The standard mural sample block includes a coarse mud layer and a fine mud layer. The coarse mud layer is coated on the base after mixing sandy soil and reinforcing material. The fine mud layer is coated on the coarse mud layer after mixing fine soil and soluble salt solution. A white powder layer composed of calcium carbonate and talc is coated on the fine mud layer, and various pigments are uniformly coated on the white powder layer. S2: Multiple reflectance spectra were sampled for each standard mural sample block using a spectrometer. The initial reflectance spectrum curve was corrected for breakpoints and outliers were removed. The average spectrum of multiple samples was taken as the mixed reflectance spectrum and smoothed to obtain standardized mixed reflectance spectrum data. The mixed reflectance spectrum data was standardized and the first k principal components were extracted by principal component analysis to reduce the data dimensionality. S3: The LASSO regression model combined with 5-fold cross-validation is used to screen sparse feature wavelengths related to salt concentration from mixed reflectance spectral data. The objective function of the LASSO regression model is set to minimize the sum of squared residuals and the sum of L1 regularization terms, and the optimal regularization parameter is determined by cross-validation. S4: Input the selected sparse feature wavelengths into the elastic network regression model, and train the unmixed model by combining L1 / L2 mixed regularization and 5-fold cross-validation. The objective function of the elastic network regression model is set to minimize the sum of squared residuals and the sum of L1 / L2 mixed regularization terms, and the optimal combination of regularization parameters is determined by cross-validation. S5: Extract spectral values from the actual mural's hyperspectral data according to the selected characteristic wavelengths, standardize the data, and then input them into the trained unmixed model to output the soluble salt content of each area on the mural surface.
2. The method for detecting soluble salt content in murals using hyperspectral unmixing according to claim 1, characterized in that, In step S1, the following method is used when making the standard mural sample block: The total salt content of the sample blocks was configured as a percentage of 100g soil mass, including 21 gradients ranging from 0% to 1% at 0.05% intervals; the preset mass ratio of each soluble salt to the total salt mass was: sodium chloride 40%, calcium chloride 20%, sodium sulfate 20%, and sodium nitrate 20%. The coarse mud layer is formed by mixing sandy soil with wheat straw and deionized water after filtration through a sieve to form dry, hard mud, and is then evenly coated onto the substrate surface with a thickness of 1 cm. The fine mud layer is formed by mixing fine-grained soil with deionized water and soluble salt solution after filtration through a sieve, and is then coated onto the surface of the coarse mud layer with a thickness of 0.5 cm. The white powder layer consists of calcium carbonate and talc with a mass ratio of 10%-20%. Four pigments—cinnabar, azurite, malachite, and graphite—are uniformly coated onto the surface of the completely dried sample block in parallel strips, with each pigment area being 2.5 cm wide and 0.2 cm thick. During the natural air-drying process of the sample blocks, the temperature and humidity were monitored using soil testing instruments, and the temperature and humidity were adjusted in real time to ensure that the cultivation conditions of each sample block were consistent.
3. The method for detecting soluble salt content in murals using hyperspectral unmixing according to claim 1, characterized in that, The specific methods for reflectance spectrum sampling and processing in step S2 include: The reflectance spectra of each standard mural sample block were sampled four times at 24-hour intervals in a darkroom using a spectrometer. Whiteboard calibration was performed before each sampling, and outliers caused by point edges or measurement jitter were removed. The initial reflectance spectrum curves of the four samplings were averaged after breakpoint correction to obtain the mixed reflectance spectrum curve. Further noise reduction is achieved through Savitzky-Golay smoothing, using the formula... Standardization processing is performed to obtain standardized mixed reflectance spectral data. Where X is the spectral data matrix before processing, μ is the mean of each wavelength, and σ is the standard deviation of each wavelength; When performing dimensionality reduction using PCA principal component analysis, the covariance matrix is constructed. Calculate the eigenvalues λ of the covariance matrix C. C With eigenvector ν C Where m is the sample size; based on the cumulative variance contribution rate Extract the top k principal components corresponding to the cumulative variance contribution rate reaching a preset threshold, and construct matrix V from these k principal components. k Standardized mixed reflectance spectral data Projecting to a lower-dimensional space yields a dimensionality-reduced data matrix. ,in It is the sum of the variances of the first k principal components; It is the total variance of all principal components.
4. The method for detecting soluble salt content in murals using hyperspectral unmixing according to claim 1, characterized in that, In step S3, the specific methods for constructing and using the LASSO regression model include: The objective function is set as follows: Where m is the sample size. This represents the true salt concentration of sample i. Let be the standardized value of sample i on the j-th principal component; Let be the regression coefficient of the j-th principal component; To control sparsity, the L1 regularization strength parameter was used. Through 5-fold cross-validation, the optimal regularization parameter λ that minimizes the model error was selected. All regression coefficients were calculated and the characteristic wavelengths corresponding to non-zero regression coefficients were screened out. In step S4, the specific methods for constructing and using the elastic network regression model include: The objective function is set as follows Where α is the mixing ratio parameter of L1 and L2 regularization, and λ1 and λ2 are the strengths of L1 regularization and L2 regularization, respectively. The mean square error (MSE) of different combinations of α, λ1, and λ2 parameters is evaluated using 5-fold cross-validation. The formula for calculating MSE is: Where N is the number of test samples, For the predicted salt concentration of the i-th sample, select the optimal parameter combination that minimizes the MSE, and fit the final model using the complete training set; the final regression equation is: , where X i Let ε be the standardized value of the i-th sample, and ε be the model error term.
5. The method for detecting soluble salt content in murals using hyperspectral unmixing according to claim 1, characterized in that, The methods for unmixing model verification and salt content output in step S5 include: The spectral values of the actual mural hyperspectral data were extracted according to the selected characteristic wavelengths and standardized using the same standardized parameters as the training data. The standardized hyperspectral data were then input into the trained unmixed model to calculate the soluble salt content of each area on the mural surface. Calculate the coefficient of determination R 2 The mean square error (MSE), root mean square error (RMSE), and mean absolute error (MAE) are calculated as follows: , , , ; where Y 真实 Y represents the actual measured salt concentration value. 预测 The salt concentration value is predicted by the unmixing model. The average salt concentration is denoted by ; N is the number of test samples. and These are the actual and predicted salt concentration values for the i-th sample, respectively. Evaluate the goodness of fit of the unmixed model: If If the values of MSE, RMSE, and MAE are all lower than the corresponding preset thresholds, the unmixing model is deemed valid and a salt content distribution map is output.
6. The method for detecting soluble salt content in murals using hyperspectral unmixing according to claim 1, characterized in that, The processing of the initial reflectance spectrum curve in step S2 includes the following operations: The abrupt changes in the spectral curve were smoothed using piecewise linear interpolation, with an interpolation interval of 5 nm. The correction formula is as follows: Where R(λ) is the original reflectivity at wavelength λ, R 校正 (λ) represents the corrected reflectivity; Outlier data points deviating from the mean by more than three times the standard deviation were removed using the 3σ principle. Savitzky-Golay smoothing parameters were optimized using a filter with a window width of 15nm and a polynomial order of 3. The smoothing effect was verified by calculating the signal-to-noise ratio (SNR) until the SNR was ≥ 30dB.
7. The method for detecting soluble salt content in murals using hyperspectral unmixing according to claim 6, characterized in that, The feature wavelength selection for the LASSO regression model in step S3 further includes: The wavelength range was segmented, and the spectral range of 400-2500nm was divided into 10 sub-intervals with a span of 210nm. LASSO regression was performed independently in each sub-interval. A significance t-test was performed on the selected characteristic wavelengths, with a significance level of 0.01, and only wavelengths with a probability value of p-value < 0.01 that satisfies the null hypothesis were retained. Redundant wavelengths are removed, and the Pearson correlation coefficient r of all characteristic wavelength pairs is calculated. If |r| of a characteristic wavelength pair is greater than 0.9, the two characteristic wavelengths are determined to be a highly correlated wavelength group. Only the wavelength with the largest absolute value of the regression coefficient is retained in each wavelength group. The screening results of all sub-intervals are merged to form a sparse feature wavelength set, and the corresponding salt sensitivity weights are recorded. ,β j The regression coefficients for the j-th wavelength in the LASSO regression model are used to remove wavelengths with a weight less than 0.05, thus generating the final feature set.
8. The method for detecting soluble salt content in murals using hyperspectral unmixing according to claim 7, characterized in that, The optimization of the elastic network regression model in step S4 includes: The mixed regularization parameters are dynamically adjusted. The initial value of the mixing ratio parameter α of L1 and L2 regularization is set to 0.5, and the search range of L1 regularization intensity λ1 and L2 regularization intensity λ2 is [10]. −4 10 2 The optimal parameter combination is determined by combining grid search with 5-fold cross-validation, and the optimization objective is set to minimize the mean square error (MSE) of the validation set. During training, a Dropout mechanism is incorporated to randomly mask 20% of the input features. Training is repeated 10 times, and the mean of the regression coefficients is taken as the final solution. A non-negativity constraint X is added. j β j ≥0 and abundance sum to 1 constrain ∑β j =1, X i Let β be the standardized value of the i-th sample. j is the regression coefficient for the j-th wavelength in the LASSO regression model.
9. The method for detecting soluble salt content in murals using hyperspectral unmixing according to claim 1, characterized in that, The unmixing model verification in step S5 further includes: The actual mural hyperspectral data was divided into a 0.5cm×0.5cm pixel grid, and the average reflectance of the selected characteristic wavelengths within each grid was extracted to ensure that the spatial scale was consistent with that of the training samples. The salt content Y of each pixel is output through the demixing model. i A thermal map of the spatial distribution of salt in the entire mural was generated using the Kriging interpolation method and a Gaussian semivariogram model. Calculate the prediction confidence interval for each pixel. SE represents the standard error. Let be the predicted salt concentration value of the i-th sample. If the CI width exceeds 0.1%, the region is marked as a region that needs to be reviewed. Data is then re-acquired and the model is updated for the region that needs to be reviewed.
10. The method for detecting soluble salt content in murals using hyperspectral unmixing according to claim 9, characterized in that, The standard error SE is calculated through the following steps: Constructing regression coefficients β to reflect the elasticity of the network regression model j The covariance matrix Cov(β) of the relationship between them. Where X is the characteristic wavelength data matrix, σ 2 Let I be the variance of the training set residuals, I be the identity matrix, and D be the L2 regularization weight matrix, where α is the mixed ratio parameter of L1 and L2 regularization, and λ1 and λ2 are the L1 regularization strength and L2 regularization strength, respectively. Calculate the prediction variance: , where X i Let σ be the feature wavelength vector of the i-th pixel. 2 For the training set residual variance; The reflectance of the perturbation feature wavelength is sampled using Monte Carlo sampling with a perturbation amplitude of ±1%, and the standard deviation of the salt content is statistically output as SE. If more than 10% of the pixels in the entire mural area have a CI width exceeding 0.1%, the model recalibration process is triggered, and the hyperspectral data of that area is reacquired and the unmixed model is updated.