Hyperspectral unmixing detection method for content of soluble salt in mural
Through standardized sample preparation and spectral data preprocessing, combined with LASSO regression and elastic network models, the difficult problem of spectral separation of salt and pigment in murals was solved, high-precision non-destructive testing was achieved, adapting to complex scenes, and providing accurate salt distribution maps and risk assessments.
Patent Information
- Application Number
- CN202510536888.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-27
- Publication Date
- 2025-09-09
- Estimated Expiration
- 2045-04-27
AI Technical Summary
Existing hyperspectral detection methods have difficulty in effectively separating the spectral signals of salt and pigment on the surface of murals, and the detection accuracy is limited. In addition, traditional detection methods are highly destructive and difficult to achieve non-destructive testing.
Through standardized sample preparation, spectral data preprocessing and characteristic wavelength screening, combined with LASSO regression model and elastic network regression model, the unmixing model is optimized to achieve the separation of salt and pigment spectra.
It achieves high-precision non-destructive testing of the soluble salt content of murals, reduces detection errors, adapts to complex pigment and salt mixing scenarios, and provides accurate spatial positioning and risk assessment basis.
Smart Images

Figure CN120609755A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of mural detection, and in particular to a hyperspectral unmixing detection method for soluble salt content in murals. Background Art
[0002] In the field of cultural heritage preservation, murals, as important historical and cultural relics, are significantly affected by soluble salt erosion. The accumulation and migration of soluble salts (such as sodium chloride, calcium chloride, sodium sulfate, and sodium nitrate) within murals can lead to defects such as crumbling of the base layer and loss of pigment. Therefore, accurate measurement of soluble salt content on mural surfaces is crucial for developing scientific conservation plans. Traditional testing methods (such as conductivity and ion chromatography) require destructive sampling and analysis, resulting in low efficiency and difficulty in achieving in-situ non-destructive testing of murals.
[0003] Hyperspectral imaging technology, due to its advantages of rapidity, non-destructiveness, and multi-band information fusion, has been increasingly applied to salt damage detection in murals. However, mural surfaces are often covered with multiple pigments (such as cinnabar, azurite, and malachite), whose spectral signals overlap and interfere with the salt spectrum, making it difficult to deconvolute the mixed spectra. Existing hyperspectral detection methods, when processing complex mixed spectra, suffer from redundant feature wavelength screening and insufficient model overfitting resistance. This makes it difficult to effectively separate the spectral signals of salt and pigments, limiting detection accuracy.
[0004] Furthermore, the preparation of mural samples requires simulating real-world structures (such as coarse and fine mud layers, and pigment layers) and salt gradients. However, existing methods lack standardized control over sample preparation, which can lead to mismatches between training data and the actual murals, affecting the model's generalization capabilities. In the data preprocessing phase, inadequate spectral curve breakpoint correction, noise removal, and dimensionality reduction can introduce additional errors, further reducing detection reliability.
[0005] Current technologies lack targeted unmixing model optimization strategies for multi-component salt mixtures (e.g., multiple salts coexisting in varying proportions), making it difficult to accurately invert the content of complex salt combinations. Furthermore, the spatial distribution visualization and error quantification assessment systems for detection results are still incomplete, making it difficult to provide precise spatial positioning and risk assessment for conservation and restoration. Therefore, a method for measuring the soluble salt content of murals that can effectively handle mixed spectra, adapt to complex scenarios, and achieve both high accuracy and robustness is urgently needed to meet the needs of practical cultural relic protection. Summary of the Invention
[0006] This method overcomes the problems of traditional detection methods, such as damage to mural integrity, difficulty dealing with spectral overlap, and insufficient precision. Through standardized sample preparation, spectral data preprocessing, characteristic wavelength screening, and hybrid regularization model optimization, it achieves non-destructive and efficient detection of salt content on mural surfaces. It has low detection error, supports demixing of complex pigment and salt mixtures, and outputs reliable detection indicators, providing accurate data support for cultural relic conservation and restoration.
[0007] In order to achieve the above object, the present invention adopts the following scheme: A method for detecting the soluble salt content of a mural through hyperspectral unmixing includes the following steps: S1: Producing multiple standard mural sample blocks with different soluble salt contents, wherein the soluble salts are formed by uniformly mixing sodium chloride, calcium chloride, sodium sulfate, and sodium nitrate in a preset mass ratio. The standard mural sample blocks include a coarse mud layer and a fine mud layer. The coarse mud layer is formed by mixing sandy soil and a reinforcing material and then applying it to a substrate. The fine mud layer is formed by mixing fine-grained soil and a soluble salt solution and then applying it to the coarse mud layer. A white powder layer composed of calcium carbonate and talc is applied on the fine mud layer, and a variety of pigments are evenly applied on the white powder layer. S2: Use a spectrometer to sample the reflection spectrum of each standard mural sample block multiple times, perform breakpoint correction and outlier removal on the initial reflection spectrum curve, take the average spectrum of multiple samples as the mixed reflection spectrum, and smooth it to obtain standardized mixed reflection spectrum data; standardize the mixed reflection spectrum data, and extract the first k principal components through principal component analysis to reduce the data dimension; S3: A LASSO regression model combined with 5-fold cross validation is used to screen sparse characteristic wavelengths related to salt concentration from the mixed reflectance spectrum data. The objective function of the LASSO regression model is set to minimize the sum of the residual sum of squares and the L1 regularization term, and the optimal regularization parameter is determined by cross validation; S4: Input the screened sparse characteristic wavelengths into the elastic network regression model, and train an unmixing model in combination with L1 / L2 mixed regularization and 5-fold cross validation. The objective function of the elastic network regression model is set to minimize the sum of the residual sum of squares and the L1 / L2 mixed regularization term, and the optimal regularization parameter combination is determined by cross validation; S5: After extracting the spectral values of the actual mural's hyperspectral data according to the screened characteristic wavelengths and performing normalization processing, it is input into the trained unmixing model to output the soluble salt content of each area on the mural surface.
[0008] Preferably, in step S1, when preparing the standard mural sample block, the following method is adopted: The total salt content of the sample block was configured as a percentage of 100g of soil mass, including 21 gradients set from 0% to 1% in 0.05% steps. The preset mass ratios of each soluble salt to the total salt mass were: sodium chloride 40%, calcium chloride 20%, sodium sulfate 20%, and sodium nitrate 20%. The coarse mud layer is formed by filtering sandy soil through a mesh, mixing it with wheat straw and ionized water to form dry hard soil, and then evenly coating the surface of the substrate with a thickness of 1 cm. The fine mud layer is formed by filtering fine-grained soil through a mesh, mixing it with ionized water and a soluble salt solution, and then coating the surface of the coarse mud layer with a thickness of 0.5 cm. The white powder layer is composed of calcium carbonate and 10%-20% by mass of talc. The four pigments of cinnabar, azurite, malachite, and graphite are evenly coated in parallel strips on the surface of the completely dried sample block. Each pigment area is 2.5 cm wide and 0.2 cm thick. During the natural drying process of the sample blocks, the temperature and humidity are observed and cultivated using soil testing instruments, and the temperature and humidity are adjusted in real time to ensure that the cultivation conditions of each sample block are consistent.
[0009] Preferably, in step S2, the specific method of sampling and processing the reflection spectrum includes: A spectrometer was used to sample the reflectance spectrum of each standard mural sample four times in a darkroom, 24 hours apart. Whiteboard calibration was performed before each sampling, and outliers caused by point edges or measurement jitter were eliminated. The initial reflectance spectrum curves of the four samples were corrected for breakpoints and averaged to obtain a mixed reflectance spectrum curve. Further noise is eliminated by Savitzky-Golay smoothing, using the formula Perform standardization processing to obtain standardized mixed reflectance spectrum data , where X is the spectral data matrix before processing, μ is the mean of each wavelength, and σ is the standard deviation of each wavelength; When performing PCA principal component analysis dimensionality reduction, construct the covariance matrix , calculate the eigenvalue λ of the covariance matrix C C With the eigenvector ν C , where m is the number of samples; according to the cumulative variance contribution rate , extract the first k principal components corresponding to the cumulative variance contribution rate reaching the preset threshold, and construct the matrix V with these k principal components k , normalize the mixed reflectance spectrum data Projecting to low-dimensional space to obtain the reduced-dimensional data matrix ,in is the sum of the variances of the first k principal components; is the total variance of all principal components.
[0010] Preferably, in step S3, the specific method of constructing and using the LASSO regression model includes: The optimization objective function is set as: , where m is the number of samples, is the true value of salt concentration of sample i; is the standardized value of sample i on the jth principal component; is the regression coefficient of the jth principal component; The L1 regularization strength parameter is used to control sparsity. The optimal regularization parameter λ that minimizes the model error is selected through 5-fold cross-validation testing. All regression coefficients are calculated and the characteristic wavelengths corresponding to non-zero regression coefficients are screened. In step S4, the specific method of constructing and using the elastic network regression model includes: The objective function is set as , where α is the mixing ratio parameter of L1 and L2 regularization, λ1 and λ2 are the L1 regularization strength and L2 regularization strength respectively. The mean square error (MSE) of different α, λ1, and λ2 parameter combinations is evaluated by 5-fold cross validation. The MSE calculation formula is: , where N is the number of test samples, For the predicted value of the salt concentration of the i-th sample, the optimal parameter combination that can minimize the MSE is selected, and the final model is fitted using the complete training set; the final regression equation is , where X i is the standardized value of the i-th sample, and ε is the model error term.
[0011] Preferably, the method for verifying the unmixing model and outputting the salt content in step S5 includes: The spectral values of the actual mural hyperspectral data were extracted according to the selected characteristic wavelengths and normalized according to the same normalization parameters as the training data. The normalized hyperspectral data were input into the trained unmixing model to calculate the soluble salt content of each area on the mural surface. Calculate the coefficient of determination R 2 , mean square error MSE, root mean square error RMSE and mean absolute error MAE, calculated as: , , , Among them, Y 真实 is the actual measured salt concentration value, Y 预测 is the salt concentration value predicted by the unmixing model, is the average value of salt concentration; N is the number of test samples, and are the actual and predicted values of the salt concentration of the i-th sample respectively; Evaluate the goodness of fit of the unmixing model: If If the values of MSE, RMSE, and MAE are all lower than the corresponding preset thresholds, the unmixing model is judged to be valid and the salt content distribution map is output.
[0012] Preferably, the processing of the initial reflection spectrum curve in step S2 includes the following operations: The mutation points of the spectral curve are smoothed by piecewise linear interpolation. The interpolation interval is set to 5nm. The correction formula is: Where R(λ) is the original reflectivity at wavelength λ, R 校正 (λ) is the reflectivity after correction; The 3σ principle is used to eliminate abnormal data points that deviate from the mean by more than 3 times the standard deviation; The Savitzky-Golay smoothing parameters were optimized, and a filter with a window width of 15 nm and a polynomial order of 3 was used for smoothing. The smoothing effect was verified by calculating the signal-to-noise ratio (SNR), and the SNR was processed to SNR ≥ 30 dB.
[0013] Preferably, the characteristic wavelength screening of the LASSO regression model in step S3 further includes: The wavelength range is segmented, and the spectral range of 400-2500 nm is divided into 10 sub-intervals with a span of 210 nm. LASSO regression is performed independently in each sub-interval. A t-test was performed on the selected characteristic wavelengths, with the significance level set to 0.01, and only wavelengths with a probability value p-value < 0.01 that satisfied the rejection of the null hypothesis were retained; Redundant wavelengths were eliminated and the Pearson correlation coefficient r of all characteristic wavelength pairs was calculated. If |r| of a characteristic wavelength pair was greater than 0.9, the two characteristic wavelengths were considered to be a highly correlated wavelength group. Only the wavelength with the largest absolute value of the regression coefficient was retained in each wavelength group. Combine the screening results of all sub-intervals to form a sparse characteristic wavelength set and record its corresponding salt sensitivity weight , β j is the regression coefficient of the jth wavelength in the LASSO regression model. The wavelengths with weights lower than 0.05 are eliminated to generate the final feature set.
[0014] Preferably, the optimization of the elastic network regression model in step S4 includes: The mixed regularization parameters are dynamically adjusted, and the initial value of the mixed ratio parameter α of L1 and L2 regularization is set to 0.5. The search range of L1 regularization strength λ1 and L2 regularization strength λ2 is [10 −4 ,10 2], the optimal parameter combination is determined by grid search combined with 5-fold cross validation, and the optimization goal is set to minimize the mean square error (MSE) of the validation set; During the training process, the Dropout mechanism is added to randomly block 20% of the feature inputs, and the training is repeated 10 times and the mean of the regression coefficient is taken as the final solution; the non-negativity constraint X is added. j β j ≥0 and the abundance sum is 1 constraint ∑β j =1,X i is the standardized value of the i-th sample, β j is the regression coefficient of the jth wavelength in the LASSO regression model.
[0015] Preferably, the unmixing model verification in step S5 further includes: The actual mural hyperspectral data was divided into 0.5cm×0.5cm pixel grids, and the average reflectance of the selected characteristic wavelengths within each grid was extracted to ensure consistency with the spatial scale of the training samples. Output the salt content Y of each pixel through the unmixing model i , and the Kriging interpolation method was used to generate the spatial distribution heat map of salt in the whole mural through the Gaussian semivariogram model; Compute prediction confidence intervals for each pixel , SE is the standard error, is the predicted value of the salt concentration of the i-th sample. If the CI width exceeds 0.1%, the area is marked as an area requiring review, and data is recollected and the unmixing model is updated for the area requiring review.
[0016] Preferably, the standard error SE is calculated by the following steps: Construct the regression coefficient β of the elastic network regression model j The covariance matrix Cov(β) of the relationship between , where X is the characteristic wavelength data matrix, σ 2 is the residual variance of the training set, I is the identity matrix, D is the L2 regularization weight matrix, where α is the mixing ratio parameter of L1 and L2 regularization, λ1 and λ2 are the L1 regularization strength and L2 regularization strength respectively; Calculate the prediction variance: , where X i is the characteristic wavelength vector of the i-th pixel, σ 2 is the residual variance of the training set; The reflectance of the characteristic wavelength is perturbed by Monte Carlo sampling with a perturbation amplitude of ±1%, and the standard deviation of the salt content is statistically output as SE. If more than 10% of the pixels in the entire mural area have a CI width exceeding 0.1%, the model recalibration process is triggered, and the hyperspectral data of the area is recollected and the unmixing model is updated.
[0017] The present invention has at least the following beneficial effects: (1) by screening sparse characteristic wavelengths through LASSO regression and combining with elastic network regression model unmixing, the mixed spectral signals of salt and pigment are effectively separated, thus avoiding the destructive sampling of traditional methods and realizing high-precision non-destructive detection of the soluble salt content of murals with low detection error; (2) by presetting the salt gradient (0%-1%) and standardizing the sample preparation process (layered coating of coarse mud layer, fine mud layer and pigment layer), the reliability of training data is ensured; combined with pre-processing such as breakpoint correction, Savitzky-Golay smoothing and PCA dimensionality reduction, the quality of spectral data is improved, noise interference is reduced, and high-quality input is provided for model training; (3) by removing redundant wavelengths through segmented LASSO regression and significance test, and combining with L1 / L2 mixed regularization of elastic network model, the problem of multi-pigment coverage is effectively solved. The spectral overlap problem under the cover is solved, and the model is adapted to various salt combinations (mixed salts such as sodium chloride and calcium chloride) and complex pigment environments, thereby improving the robustness of the model. (4) A heat map of the spatial distribution of salt is generated through Kriging interpolation to intuitively display the salt-enriched area. Indicators such as the determination coefficient R² and the mean square error MSE are introduced to quantify the model accuracy, and high uncertainty areas (areas that need to be reviewed if the confidence interval exceeds the standard) are marked to provide accurate spatial positioning and risk assessment basis for protection and restoration. (5) The model's ability to resist overfitting is improved through the Dropout mechanism and parameter grid search. Non-negative constraints (salt content ≥ 0) and abundance and constraints (total salt ≤ 100%) are added to ensure that the unmixing results are consistent with physical meaning. A model recalibration mechanism is set (such as triggering when the confidence interval exceeds the standard by more than 10%) to adapt to environmental changes such as mural aging and maintain detection accuracy over a long period of time. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] Figure 1 This is a principle flow chart of a hyperspectral unmixing detection method for soluble salt content in murals provided by the present invention. DETAILED DESCRIPTION
[0019] The present invention will be described in further detail below in conjunction with the accompanying drawings so that those skilled in the art can implement the invention with reference to the description.
[0020] like Figure 1 As shown, the present invention provides a method for detecting the soluble salt content of a mural through hyperspectral unmixing, comprising the following steps: S1: Prepare multiple standard mural sample blocks with different soluble salt contents, wherein the soluble salt is formed by uniformly mixing sodium chloride, calcium chloride, sodium sulfate, and sodium nitrate in a preset mass ratio. The standard mural sample blocks include a coarse mud layer and a fine mud layer. The coarse mud layer is formed by mixing sandy soil and reinforcing materials and then applied to a substrate. The fine mud layer is formed by mixing fine-grained soil and a soluble salt solution and then applied to the coarse mud layer. A white powder layer composed of calcium carbonate and talc is applied on the fine mud layer, and a variety of pigments are evenly applied on the white powder layer.
[0021] Soluble salts can be mixed salts of sodium chloride, calcium chloride, sodium sulfate, and sodium nitrate. According to the different salt content in murals at different stages and types, the common salt types and proportions in murals are simulated. The mixing ratio design is based on the chemical analysis of actual salt damage cases to ensure that the samples cover typical salt combinations.
[0022] The sample structure is layered from base to surface. The coarse mud layer is a dry, hard base layer, 1 cm thick, formed by mixing sandy soil (larger particles) with wheat straw and deionized water. The sandy soil provides structural support, the wheat straw enhances crack resistance, and the deionized water prevents impurities from interfering. The fine mud layer is a 0.5 cm thick mixture of fine-grained soil and a soluble salt solution. The fine-grained soil simulates the mural ground layer, and the salt solution penetrates evenly to simulate salt migration. The pigment coating layer is applied by applying common mural pigments, such as cinnabar, azurite, malachite, and graphite, evenly and parallel to the sample surface, simulating the multi-pigment coverage of a real mural. The pigment thickness is adjusted during testing to ensure spectral reflectance consistent with that of a real mural.
[0023] During the drying process of the samples, the environment of the finished product is controlled, and the temperature and humidity are monitored by soil testing instruments (e.g. temperature 20-25°C, humidity 40-60%) to ensure a consistent drying rate and avoid uneven salt distribution.
[0024] During sample preparation, sample blocks were prepared according to a defined salt gradient, with at least six samples per gradient. A circular mold was used to standardize sample size, and a spatula was applied to ensure uniform layer thickness. After the mud layer dried naturally in the shade, pigment was applied in separate areas to prevent pigment mixing from interfering with spectral acquisition.
[0025] S2: Use a spectrometer to sample the reflection spectrum of each standard mural sample block multiple times, perform breakpoint correction and outlier removal on the initial reflection spectrum curve, take the average spectrum of multiple samples as the mixed reflection spectrum, and smooth it to obtain standardized mixed reflection spectrum data; standardize the mixed reflection spectrum data, and extract the first k principal components through principal component analysis to reduce the data dimension.
[0026] Configure the spectrum acquisition conditions, using a darkroom environment to eliminate external light interference. Preheat the spectrometer for 30 minutes to ensure stability. Calibrate the instrument baseline with a white plate. Each sample block is sampled four times, 24 hours apart, and the average is taken to reduce random errors.
[0027] Data preprocessing includes: Breakpoint correction: Identify wavelength points where reflectivity changes suddenly (e.g. due to instrument jitter or sample surface defects) and smooth them out through interpolation. The interpolation interval can be set to 5-10nm.
[0028] Outlier removal: Filter outlier data that deviates too much from the mean to ensure a smooth spectral curve.
[0029] Savitzky-Golay smoothing: Set parameters based on the actual smoothing algorithm. A filter with a window width of 15 nm and a polynomial order of 3 can be used to eliminate high-frequency noise and improve the signal-to-noise ratio to above 30 dB.
[0030] Standardization processing: Normalize the spectral data to a mean of 0 and a standard deviation of 1 to eliminate dimensional differences and facilitate subsequent model training.
[0031] PCA dimensionality reduction: Extract the top k principal components with cumulative variance contribution ≥ 95%, compress the data dimension, and retain the key spectral characteristics of salt and pigment.
[0032] The spectrometer collects data in preset wavelength bands, covering the entire wavelength range with each sampling. After preprocessing, a standardized spectral matrix is generated as input for subsequent models.
[0033] S3: The LASSO regression model combined with 5-fold cross-validation was used to screen sparse characteristic wavelengths related to salt concentration from the mixed reflectance spectral data. The objective function of the LASSO regression model was set to minimize the sum of the residual sum of squares and the L1 regularization term, and the optimal regularization parameter was determined by cross-validation.
[0034] The LASSO model uses an L1 regularization penalty term to compress the regression coefficients of non-critical wavelengths to 0, thereby screening out wavelengths significantly correlated with salt concentration. Cross-validation optimization utilizes a 5-fold cross-validation algorithm, dividing the data into a training set and a validation set. Different characteristic parameter values are tested, and the parameters that minimize the validation error are selected. Wavelengths corresponding to non-zero regression coefficients are retained to form a sparse feature set for characteristic wavelength output. For example, the characteristic absorption band of sodium chloride (around 1450nm) is selected. Normalized data is input into the LASSO model, and the parameter λ is iteratively optimized. A list of characteristic wavelengths and their weights is output, which can be filtered using methods such as the Pearson correlation coefficient test, with redundant wavelengths with high correlations removed.
[0035] S4: Input the screened sparse feature wavelengths into the elastic network regression model, and train the unmixing model in combination with L1 / L2 mixed regularization and 5-fold cross validation. The objective function of the elastic network regression model is set to minimize the sum of the residual sum of squares and the L1 / L2 mixed regularization term, and the optimal regularization parameter combination is determined by cross-validation.
[0036] The elastic net regression model combines L1 regularization (for sparsity) and L2 regularization (for stability) to address multicollinearity in spectral data (e.g., overlap between pigment and salt bands). Dynamic parameter adjustment is performed through a grid search test of α (mixing ratio, e.g., 0.3-0.7), λ1 (L1 intensity), and λ2 (L2 intensity), selecting the combination that minimizes the mean squared error (MSE). Physical constraints are applied: non-negativity (salt content ≥ 0) and the sum of abundances to 1 (total salt content ≤ 100%) are added to ensure realistic unmixing results. The elastic net model is trained using characteristic wavelength data selected by LASSO to generate the unmixing model. A dropout mechanism (randomly masking 20% of features) is used to enhance overfitting resistance. The training is repeated 10 times and the average is taken to improve stability.
[0037] S5: After extracting the spectral values of the actual mural's hyperspectral data according to the screened characteristic wavelengths and performing normalization processing, it is input into the trained unmixing model to output the soluble salt content of each area on the mural surface.
[0038] Actual mural data is standardized according to the training set parameters to ensure model input consistency. Prediction results are mapped onto a 0.5 cm x 0.5 cm pixel grid, and an interpolation algorithm is used to generate a heat map, visualizing the spatial distribution of salt. Metrics such as the coefficient of determination (R²) (greater than 0.9 is considered excellent), mean square error (MSE) (less than 0.05), and root mean square error (RMSE) (less than 0.2) are calculated to verify model reliability. Areas requiring review can also be resampled to update model parameters. A salt content report and distribution map are output to guide restoration decisions.
[0039] The above algorithm achieves high-precision detection of salt content in murals. By combining the LASSO and elastic network models, it separates the salt signal from mixed spectra, achieving a detection error that is 20% less than traditional methods (such as conductivity). This non-destructive analysis eliminates the need for sampling, preserving the integrity of the murals and making it suitable for use in precious cultural heritage. It can adapt to complex scenarios, handle a variety of pigment-salt mixtures, and mitigate interference from spectral overlap, demonstrating strong robustness. The model provides visual output, converting salt content into a spatial distribution heat map and error quantification metrics, providing precise data support for restoration.
[0040] In another technical solution, in step S1, when preparing the standard mural sample block, the following method is used: The total salt content of the sample block was configured as a percentage of 100g of soil mass, including 21 gradients set from 0% to 1% in 0.05% steps. The preset mass ratios of each soluble salt to the total salt mass were: sodium chloride 40%, calcium chloride 20%, sodium sulfate 20%, and sodium nitrate 20%. The coarse mud layer is formed by filtering sandy soil through a mesh, mixing it with wheat straw and ionized water to form dry hard soil, and then evenly coating the surface of the substrate with a thickness of 1 cm. The fine mud layer is formed by filtering fine-grained soil through a mesh, mixing it with ionized water and a soluble salt solution, and then coating the surface of the coarse mud layer with a thickness of 0.5 cm. The white powder layer is composed of calcium carbonate and 10%-20% by mass of talc. The four pigments of cinnabar, azurite, malachite, and graphite are evenly coated in parallel strips on the surface of the completely dried sample block. Each pigment area is 2.5 cm wide and 0.2 cm thick. During the natural drying process of the sample blocks, the temperature and humidity are observed and cultivated using soil testing instruments, and the temperature and humidity are adjusted in real time to ensure that the cultivation conditions of each sample block are consistent.
[0041] A 21-step salt gradient was designed: 0%, 0.05%, 0.1%, 0.15%, 0.2%, 0.25%, 0.3%, 0.35%, 0.4%, 0.45%, 0.5%, 0.55%, 0.6%, 0.65%, 0.7%, 0.75%, 0.8%, 0.85%, 0.9%, 0.95%, and 1%, covering the typical range from no salt to high salt concentrations. Each gradient corresponds to a different salt mass (e.g., 0.2g salt / 100g soil). Precision weighing and mixing (e.g., 40% sodium chloride, 20% calcium chloride) simulates the chemical environment of salt migration in real murals. The coarse soil layer was created by filtering large-grained sandy soil (approximately 1-2mm in diameter) through a sieve and mixing it with wheat straw (2-3cm long) and deionized water (conductivity ≤10μS / cm) to form a dry, hardened soil layer. The wheat straw enhances crack resistance, while the deionized water prevents impurities from interfering with salt distribution. Apply a 1cm thick layer, ensuring a smooth surface. The fine mud layer is created by mixing fine-grained soil (particle diameter ≤ 0.5mm) with a soluble salt solution (prepared in a gradient concentration) and applying a 0.5cm thick layer. The fine-grained soil simulates the structure of the mural ground layer, allowing the salt solution to penetrate evenly, simulating capillary migration. The salt mixture should be stirred until free of lumps and allowed to stand for 12 hours to eliminate air bubbles. Apply using a standard spatula in two steps (coarse mud layer → air-dry for 24 hours → fine mud layer).
[0042] A white powder layer is evenly applied over the fine layer, forming a white pigment base. This physically isolates the pigment layer from the salt layer during pigment application, preventing mixing and interference with spectral acquisition. The sample environment is monitored in real time using soil testing equipment (such as temperature and humidity sensors), maintaining a temperature of 20-25°C (target value 22.5°C ± 2.5°C) and a humidity of 40-60% (target value 50% ± 10%). If the temperature and humidity deviate, ventilation or humidification is activated. The pigment is mixed to a suitable consistency (like honey) and applied in a single direction with a soft brush to avoid bubbling. Temperature and humidity data are recorded every two hours, triggering an alarm and requiring manual intervention if they deviate from the target range.
[0043] Standardize the molds, using round iron molds (10cm×10cm×4cm) to uniformly size the samples and minimize the impact of shape differences on spectral reflectance. Naturally dry the samples in the shade for approximately 7-10 days, turning them daily to ensure even drying. Use an infrared moisture meter to monitor moisture content until it drops below 5%. This provides standardized training data to improve model generalization. Ensure that the salt distribution of the samples is consistent with the actual murals, reducing detection errors.
[0044] In another technical solution, in step S2, the specific method of sampling and processing the reflection spectrum includes: A spectrometer was used to sample the reflectance spectrum of each standard mural sample four times in a darkroom, 24 hours apart. Whiteboard calibration was performed before each sampling, and outliers caused by point edges or measurement jitter were eliminated. The initial reflectance spectrum curves of the four samples were corrected for breakpoints and averaged to obtain a mixed reflectance spectrum curve. Further noise is eliminated by Savitzky-Golay smoothing, using the formula Perform standardization processing to obtain standardized mixed reflectance spectrum data , where X is the spectral data matrix before processing, μ is the mean of each wavelength, and σ is the standard deviation of each wavelength; When performing PCA principal component analysis dimensionality reduction, construct the covariance matrix , calculate the eigenvalue λ of the covariance matrix C C With the eigenvector ν C , where m is the number of samples; according to the cumulative variance contribution rate , extract the first k principal components corresponding to the cumulative variance contribution rate reaching the preset threshold, and construct the matrix V with these k principal components k , normalize the mixed reflectance spectrum data Projecting to low-dimensional space to obtain the reduced-dimensional data matrix ,in is the sum of the variances of the first k principal components; is the total variance of all principal components.
[0045] The spectrometer was calibrated in a darkroom to eliminate stray light and preheat for 30 minutes to stabilize the light source. The instrument baseline was calibrated using a white plate (reflectance ≥ 95%) and repeated before each sampling. Spectral data was collected four times with 24-hour intervals for each sample block to cover reflectance changes at different drying stages. The average was used to reduce random errors. The sampling wavelength range was 400-2500 nm with a resolution of ≤5 nm, covering the visible to near-infrared region, which contains critical spectral information. Reflectance jumps (e.g., data jumps caused by surface cracks) were eliminated.
[0046] The Savitzky-Golay smoothing window width is 15 nm, and the polynomial order is 3. This eliminates high-frequency noise (such as instrument thermal noise) and improves the signal-to-noise ratio to ≥30 dB. The mean μ and standard deviation σ are calculated by wavelength, and the data is transformed to eliminate dimensional differences and prevent certain wavelengths from dominating the model training. After smoothing, the spectral curve is checked for continuity, and discontinuities are repaired using linear interpolation. The normalization parameters (μ, σ) are stored for future use to ensure that the actual data are consistent with the training set.
[0047] The covariance matrix is calculated to analyze inter-wavelength correlations and extract the principal components with the largest variance. The top k principal components (e.g., k is 5-10) are retained when the cumulative variance contribution rate is ≥95%. Normalized data is projected into a low-dimensional space to reduce redundant information (such as pigment spectral noise) and highlight salt-sensitive features. This reduces data dimensionality and accelerates model training. This improves the signal-to-noise ratio of the salt signal and enhances detection accuracy.
[0048] In another technical solution, in step S3, the specific method of constructing and using the LASSO regression model includes: The optimization objective function is set as: , where m is the number of samples, is the true value of salt concentration of sample i; is the standardized value of sample i on the jth principal component; is the regression coefficient of the jth principal component; The L1 regularization strength parameter is used to control sparsity. The optimal regularization parameter λ that minimizes the model error is selected through 5-fold cross-validation testing. All regression coefficients are calculated and the characteristic wavelengths corresponding to non-zero regression coefficients are screened. In step S4, the specific method of constructing and using the elastic network regression model includes: The objective function is set as , where α is the mixing ratio parameter of L1 and L2 regularization, λ1 and λ2 are the L1 regularization strength and L2 regularization strength respectively. The mean square error (MSE) of different α, λ1, and λ2 parameter combinations is evaluated by 5-fold cross validation. The MSE calculation formula is: , where N is the number of test samples, For the predicted value of the salt concentration of the i-th sample, the optimal parameter combination that can minimize the MSE is selected, and the final model is fitted using the complete training set; the final regression equation is , where X i is the standardized value of the i-th sample, and ε is the model error term.
[0049] The LASSO regression model was constructed using L1 regularization optimization. Sparsity was controlled by adjusting the λ value (range 0.01-10) to select wavelengths strongly correlated with salinity. A λ value that was too large resulted in excessive sparsity, while a λ value that was too small retained redundant wavelengths. Cross-validation parameters were selected using a 5-fold cross-validation approach, where the data was divided into five groups, with four groups used for training and one for validation. λ was selected to minimize the mean square error (MSE) on the validation set.
[0050] Initially set λ to 0.1 and gradually increase it to λ = 1 to observe the sparsity of the regression coefficients. Output a list of wavelengths with non-zero coefficients, such as filtering out 1450nm (the characteristic wavelength of sodium sulfate).
[0051] Elastic net model optimization uses a hybrid regularization parameter: α controls the L1 / L2 ratio (such as 0.5 to balance sparsity and stability), λ1 and λ2 adjust the penalty intensity (search range 10⁻ 4 -10²).
[0052] The Dropout mechanism was introduced to combat overfitting. 20% of the feature inputs were randomly blocked, and the regression coefficients were averaged after 10 training iterations to reduce the model's sensitivity to noise. A grid search was performed with α = 0.3, 0.5, and 0.7, λ1 = 0.1, 1, and 10, and λ2 = 0.01, 0.1, and 1. The parameter combination with the lowest mean square error (MSE) was selected (e.g., α = 0.5, λ1 = 0.1, and λ2 = 0.01).
[0053] Physical constraints and non-negativity constraints in model validation: enforcing the regression coefficient β j ≥0, ensuring that the predicted salinity value is non-negative; and abundance and constraint: ∑β j = 1, limiting the total salt content to 100%, which is consistent with the actual proportion. This solves non-physical problems in spectral unmixing (such as negative salt values). This improves the reliability of the model in actual mural detection.
[0054] In another technical solution, the method for verifying the unmixing model and outputting the salt content in step S5 includes: The spectral values of the actual mural hyperspectral data were extracted according to the selected characteristic wavelengths and normalized according to the same normalization parameters as the training data. The normalized hyperspectral data were input into the trained unmixing model to calculate the soluble salt content of each area on the mural surface. Calculate the coefficient of determination R 2, mean square error MSE, root mean square error RMSE and mean absolute error MAE, calculated as: , , , Among them, Y 真实 is the actual measured salt concentration value, Y 预测 is the salt concentration value predicted by the unmixing model, is the average value of salt concentration; N is the number of test samples, and are the actual and predicted values of the salt concentration of the i-th sample respectively; Evaluate the goodness of fit of the unmixing model: If If the values of MSE, RMSE, and MAE are all lower than the corresponding preset thresholds, the unmixing model is judged to be valid and the salt content distribution map is output.
[0055] In error metric calculations, the coefficient of determination (R²) measures the model's ability to explain variation. An R² > 0.9 indicates a high degree of agreement between the predicted and true values. The mean squared error (MSE) quantifies prediction bias, with an MSE < 0.05 (in percent salt concentration) considered excellent. MAE reflects the average error (e.g., < 0.1%), while RMSE amplifies the impact of large errors (e.g., < 0.2%). When calculating R², if the value falls below 0.85, a model retraining process is triggered. The MSE threshold can be adjusted based on detection requirements (e.g., 0.03 for stringent scenarios).
[0056] Model validity is determined by setting thresholds for |R²-1| < 0.05, MSE < 0.05, RMSE < 0.15, and MAE < 0.1. The salinity heat map is divided into 0.5 cm x 0.5 cm pixels, and unsampled areas are filled using kriging interpolation. The salinity gradient is indicated by a color scale (e.g., blue for low salinity, red for high salinity). Pixels with a confidence interval width > 0.1% are flagged and manually retested or locally resampled. The output report includes error statistics and distribution maps to support remediation decisions.
[0057] The dynamic calibration mechanism includes a covariance matrix update mechanism: New sample blocks are added regularly (e.g., monthly), and regression coefficients are adjusted through incremental learning to adapt to environmental changes (e.g., aging of murals). Recalibration is triggered if the CI width exceeds the standard in more than 10% of the regions, and the model is fully rebuilt. This maintains detection accuracy over the long term and adapts to complex environments. It also provides traceable error quantification data, enhancing the scientific nature of restoration solutions.
[0058] In another technical solution, the processing of the initial reflection spectrum curve in step S2 includes the following operations: The mutation points of the spectral curve are smoothed by piecewise linear interpolation. The interpolation interval is set to 5nm. The correction formula is: Where R(λ) is the original reflectivity at wavelength λ, R 校正 (λ) is the reflectivity after correction; The 3σ principle is used to eliminate abnormal data points that deviate from the mean by more than 3 times the standard deviation; The Savitzky-Golay smoothing parameters were optimized, and a filter with a window width of 15 nm and a polynomial order of 3 was used for smoothing. The smoothing effect was verified by calculating the signal-to-noise ratio (SNR), and the SNR was processed to SNR ≥ 30 dB.
[0059] Breakpoint correction identifies sudden changes in the reflectance curve (e.g., caused by instrument jitter or sample surface defects) and repairs them using piecewise linear interpolation. Interpolation intervals can be set to 5-10 nm to ensure spectral continuity. For example, at wavelength λ, if the reflectance change within 5 nm exceeds a threshold (e.g., ±0.02), the outlier point is replaced by the average of the reflectance values on both sides.
[0060] Savitzky-Golay smoothing uses a 15nm window width and a polynomial order 3 filter to remove high-frequency noise (such as electronic noise or environmental interference). After smoothing, the signal-to-noise ratio (SNR) is calculated, with an SNR requirement of ≥30dB, to ensure data quality meets model input requirements. After the spectrometer acquires raw data, the spectrometer checks for discontinuities at each wavelength, then interpolates and corrects them to generate a calibration curve. During smoothing, the window width can be adjusted based on noise intensity (e.g., increasing it to 25nm for high noise).
[0061] Outlier detection and removal utilizes the 3σ principle: the mean (μ) and standard deviation (σ) of reflectance at each wavelength are calculated, and data points deviating from the μ ± 3σ range are removed. For example, if the reflectance at a certain wavelength is 0.5, μ = 0.45, and σ = 0.02, then data points with a value greater than 0.51 (0.45 + 3 × 0.02 = 0.51) are not removed, but data points with a value greater than 0.61 (0.51) are considered outliers. A dynamic threshold adjustment strategy is employed, and for high-noise bands (such as the near-infrared region), the threshold can be relaxed to 4σ to improve fault tolerance. μ and σ are calculated segmentally by wavelength to avoid local data distortion caused by global thresholding. After outlier removal, missing points are reinterpolated to ensure a complete spectral curve.
[0062] During data quality verification and output, perform signal-to-noise ratio verification: calculate the SNR after smoothing. If it is less than 30dB, increase the window width or smoothing times. Normalize the data by wavelength to a mean of 0 and a standard deviation of 1 to eliminate instrument response variations. This improves spectral data consistency and provides high-quality input for subsequent model training. It also reduces noise interference and improves the accuracy of screening for characteristic wavelengths of salt.
[0063] The characteristic wavelength screening of the LASSO regression model in step S3 further includes: The wavelength range is segmented, and the spectral range of 400-2500 nm is divided into 10 sub-intervals with a span of 210 nm. LASSO regression is performed independently in each sub-interval. A t-test was performed on the selected characteristic wavelengths, with the significance level set to 0.01, and only wavelengths with a probability value p-value < 0.01 that satisfied the rejection of the null hypothesis were retained; Redundant wavelengths were eliminated and the Pearson correlation coefficient r of all characteristic wavelength pairs was calculated. If |r| of a characteristic wavelength pair was greater than 0.9, the two characteristic wavelengths were considered to be a highly correlated wavelength group. Only the wavelength with the largest absolute value of the regression coefficient was retained in each wavelength group. Combine the screening results of all sub-intervals to form a sparse characteristic wavelength set and record its corresponding salt sensitivity weight , β j is the regression coefficient of the jth wavelength in the LASSO regression model. The wavelengths with weights lower than 0.05 are eliminated to generate the final feature set.
[0064] Wavelength interval segmentation and independent regression are performed to divide the 400-2500nm spectral range into 10 subintervals (each 210nm), such as 400-610nm and 610-820nm. After segmentation, LASSO regression is independently performed within each subinterval to prevent long-wavelength data from overwhelming short-wavelength features. Local feature extraction is performed to select locally sensitive wavelengths based on the salt absorption characteristics of each subinterval (for example, sodium chloride strongly absorbs around 1450nm). A regularization parameter λ is configured for each subinterval to accommodate data sparsity within the respective bands. The results of the subinterval screening are combined to form a global feature set.
[0065] When performing significance testing and redundancy elimination, a t-test was used for screening: a t-test (significance level α = 0.01) was performed on the LASSO regression coefficient, and only wavelengths with a p-value < 0.01 were retained to ensure statistical significance of the screening results. Redundancy was eliminated based on the Pearson correlation coefficient: the correlation coefficient of the characteristic wavelength pair was calculated. If |r| > 0.9, it was determined to be highly correlated, and the wavelength with the largest absolute value of the regression coefficient was retained. For example, the correlation coefficient between wavelengths A and B was 0.95, and if the corresponding |β A |>|β B |, then remove B. A sliding window is used to calculate the correlation between adjacent wavelengths to avoid high global computational complexity. After redundancy removal, the salt sensitivity weights of the retained wavelengths are recorded, and minor wavelengths with weights less than 0.05 are removed.
[0066] In the process of generating the final feature set, weight sorting is performed, and the weights are sorted by w jSort from high to low, retaining the top N wavelengths (e.g., N = 50), to balance the number of features with model complexity. Perform cross-interval validation to check the complementarity of wavelengths across different subintervals to avoid missing key features. This improves the representativeness and robustness of the feature wavelength set and reduces the impact of multicollinearity on model accuracy.
[0067] The optimization of the elastic network regression model in step S4 includes: The mixed regularization parameters are dynamically adjusted, and the initial value of the mixed ratio parameter α of L1 and L2 regularization is set to 0.5. The search range of L1 regularization strength λ1 and L2 regularization strength λ2 is [10 −4 ,10 2 ], the optimal parameter combination is determined by grid search combined with 5-fold cross validation, and the optimization goal is set to minimize the mean square error (MSE) of the validation set; During the training process, the Dropout mechanism is added to randomly block 20% of the feature inputs, and the training is repeated 10 times and the mean of the regression coefficient is taken as the final solution; the non-negativity constraint X is added. j β j ≥0 and the abundance sum is 1 constraint ∑β j =1,X i is the standardized value of the i-th sample, β j is the regression coefficient of the jth wavelength in the LASSO regression model.
[0068] In the hybrid regularization parameter optimization, the parameter search range is set to 10⁻ for the L1 regularization strength λ1 and the L2 regularization strength λ2. 4 To 10², covering the full range of scenarios from weak to strong penalties. The initial value of the mixing ratio parameter α is set to 0.5 to balance sparsity and stability. A grid search strategy is used, with the parameter grid divided on a logarithmic scale (e.g., λ1 = 0.001, 0.01, 0.1, 1, 10). Combined with 5-fold cross-validation, the optimal combination that minimizes the mean square error (MSE) on the validation set is selected. Prioritize optimizing α and λ1, fix λ2 = 0.1 for preliminary screening, and then fine-tune λ2. The MSE of each parameter combination is recorded, and the solution with the lowest error and moderate model complexity is selected.
[0069] The Dropout mechanism is used to combat overfitting and enhance stability: 20% of the feature inputs are randomly blocked during training, forcing the model to learn features outside of redundant paths and improve generalization. The training is repeated 10 times, and the mean of the regression coefficients is taken as the final solution. Non-negativity constraint: The regression coefficient β is forced to be j ≥ 0 to ensure that the salinity predictions are physically plausible (e.g., no negative concentrations). Masked features are randomly selected during each training run to ensure diversity across training rounds. After adding these constraints, the objective function is optimized using projected gradient descent.
[0070] Perform model validation and output using abundance and constraints: ∑β j = 1, limiting the total salt content to no more than 100%, consistent with actual proportions. Provide error feedback: If the validation set MSE falls below a threshold (e.g., < 0.05), retroactively adjust the parameter search range or increase training data. This improves the model's ability to demix complex spectral mixtures. Ensure that the output results are physically meaningful and can be directly used for repair decisions.
[0071] In another technical solution, the unmixing model verification in step S5 further includes: The actual mural hyperspectral data was divided into 0.5cm×0.5cm pixel grids, and the average reflectance of the selected characteristic wavelengths within each grid was extracted to ensure consistency with the spatial scale of the training samples. Output the salt content Y of each pixel through the unmixing model i , and the Kriging interpolation method was used to generate the spatial distribution heat map of salt in the whole mural through the Gaussian semivariogram model; Compute prediction confidence intervals for each pixel , SE is the standard error, is the predicted value of the salt concentration of the i-th sample. If the CI width exceeds 0.1%, the area is marked as an area requiring review, and data is recollected and the unmixing model is updated for the area requiring review.
[0072] To match the spatial resolution and generate the salinity distribution, a grid was performed: the mural surface was divided into a 0.5 cm × 0.5 cm pixel grid to ensure the spatial resolution was consistent with the training samples (e.g., a sample block size of 10 cm × 10 cm). Kriging interpolation was performed: a Gaussian semivariogram model was used to interpolate salinity values at sampling points to generate a continuous distribution heatmap. Parameters included range (e.g., 20 cm), nugget effect (e.g., 0.1), and sill value (e.g., 0.5). The average reflectance at the characteristic wavelength was extracted for each pixel and input into the unmixing model to predict the salinity value. Interpolation prioritized neighboring pixel data to prevent long-distance error propagation.
[0073] During the confidence interval calculation and review process, the standard error (SE) is calculated by perturbing the reflectance at the characteristic wavelength using Monte Carlo sampling (amplitude ±1%) and calculating the standard deviation of the predicted salinity as the SE. At a 95% confidence level, the confidence interval (CI) is set as the predicted value ±1.96 × SE. If the CI width is greater than 0.1% (e.g., 0.12%), the area is marked for review. Hyperspectral data is recollected for the marked area, and the unmixing model parameters are locally updated. Manual review is performed using a portable spectrometer for on-site verification to correct model deviations.
[0074] The model and data are dynamically updated and visualized using a heatmap rendering process, colored by salt concentration gradient (e.g., blue represents 0-0.3%, red represents 0.8-1.0%), and overlaid onto the digital mural model. The report generation process includes error statistics, distribution maps, and review recommendations. This provides intuitive spatial distribution of salt, guiding targeted remediation efforts. Uncertainty quantification enhances the credibility of test results.
[0075] The standard error SE is calculated by the following steps: Construct the regression coefficient β of the elastic network regression model j The covariance matrix Cov(β) of the relationship between , where X is the characteristic wavelength data matrix, σ 2 is the residual variance of the training set, I is the identity matrix, D is the L2 regularization weight matrix, where α is the mixing ratio parameter of L1 and L2 regularization, λ1 and λ2 are the L1 regularization strength and L2 regularization strength respectively; Calculate the prediction variance: , where X i is the characteristic wavelength vector of the i-th pixel, σ 2 is the residual variance of the training set; The reflectance of the characteristic wavelength is perturbed by Monte Carlo sampling with a perturbation amplitude of ±1%, and the standard deviation of the salt content is statistically output as SE. If more than 10% of the pixels in the entire mural area have a CI width exceeding 0.1%, the model recalibration process is triggered, and the hyperspectral data of the area is recollected and the unmixing model is updated.
[0076] In the covariance matrix construction and error propagation processing, the covariance matrix Cov(β) based on the elastic network regression coefficient β_j reflects the uncertainty of parameter estimation. The training set residual variance σ² and the regularization weight matrix are combined in the calculation. The prediction variance formula is used according to Cov(β) and the characteristic wavelength vector X i , calculate the prediction variance Var(Ŷ for each pixel i )=X i T Cov(β)X i +σ². Use numerical methods to approximate Cov(β), avoiding the direct inversion of high-dimensional matrices. Store Cov(β) for subsequent predictions, reducing real-time computational overhead.
[0077] The perturbation amplitude for Monte Carlo sampling and sensitivity analysis was set to ±1% random perturbation of the reflectance at the characteristic wavelength to simulate the effects of measurement error and noise. The standard deviation statistics were calculated by repeating the sampling 1000 times. The standard deviation of the predicted salinity was calculated as SE to quantify the model's sensitivity to input perturbations. The sampling process was parallelized and GPU-accelerated. SE distribution plots were output to identify areas of high uncertainty (such as edges or pigment mixing).
[0078] Model recalibration and closed-loop optimization were implemented. Recalibration was triggered when the CI width exceeded 0.1% for more than 10% of the pixels in the entire mural, indicating a global model update was required. Incremental learning updates were performed after adding new sample blocks, adjusting the regression coefficients using stochastic gradient descent while retaining the weighting of historical data (e.g., weighting old data 0.7, new data 0.3). This enabled the detection system to self-optimize and adapt to environmental changes and material aging. Closed-loop control maintained high-precision detection over the long term, reducing the need for manual intervention.
[0079] It should be noted that although the steps are described above in a specific order, this does not necessarily mean that the steps must be performed in this specific order. In fact, some of these steps can be performed concurrently or even in a different order, as long as the required functions can be achieved. The number of devices and processing scales described here are intended to simplify the description of the present invention. Applications, modifications, and variations of the present invention will be apparent to those skilled in the art.
[0080] Although the embodiments of the present invention have been disclosed above, they are not limited to the applications listed in the description and implementation methods. They can be fully applied to various fields suitable for the present invention. For those familiar with the art, additional modifications can be easily implemented. Therefore, without departing from the general concept defined by the claims and the scope of equivalents, the present invention is not limited to the specific details and illustrations shown and described herein.
Claims
1. A hyperspectral unmixing detection method for soluble salt content in murals, characterized in that: The following steps are involved: S1: Producing multiple standard mural sample blocks with different soluble salt contents, wherein the soluble salts are formed by uniformly mixing sodium chloride, calcium chloride, sodium sulfate, and sodium nitrate in a preset mass ratio. The standard mural sample blocks include a coarse mud layer and a fine mud layer. The coarse mud layer is formed by mixing sandy soil and a reinforcing material and then applying it to a substrate. The fine mud layer is formed by mixing fine-grained soil and a soluble salt solution and then applying it to the coarse mud layer. A white powder layer composed of calcium carbonate and talc is applied on the fine mud layer, and a variety of pigments are evenly applied on the white powder layer. S2: Use a spectrometer to sample the reflection spectrum of each standard mural sample block multiple times, perform breakpoint correction and outlier removal on the initial reflection spectrum curve, take the average spectrum of multiple samples as the mixed reflection spectrum, and smooth it to obtain standardized mixed reflection spectrum data; standardize the mixed reflection spectrum data, and extract the first k principal components through principal component analysis to reduce the data dimension; S3: A LASSO regression model combined with 5-fold cross validation is used to screen sparse characteristic wavelengths related to salt concentration from the mixed reflectance spectrum data. The objective function of the LASSO regression model is set to minimize the sum of the residual sum of squares and the L1 regularization term, and the optimal regularization parameter is determined by cross validation; S4: Input the screened sparse characteristic wavelengths into the elastic network regression model, and train an unmixing model in combination with L1 / L2 mixed regularization and 5-fold cross validation. The objective function of the elastic network regression model is set to minimize the sum of the residual sum of squares and the L1 / L2 mixed regularization term, and the optimal regularization parameter combination is determined by cross validation; S5: After extracting the spectral values of the actual mural's hyperspectral data according to the screened characteristic wavelengths and performing normalization processing, it is input into the trained unmixing model to output the soluble salt content of each area on the mural surface.
2. The hyperspectral unmixing detection method for soluble salt content in murals according to claim 1 is characterized in that: In step S1, when preparing the standard mural sample block, the following method is used: The total salt content of the sample block was configured as a percentage of 100g of soil mass, including 21 gradients set from 0% to 1% in 0.05% steps. The preset mass ratios of each soluble salt to the total salt mass were: sodium chloride 40%, calcium chloride 20%, sodium sulfate 20%, and sodium nitrate 20%. The coarse mud layer is formed by filtering sandy soil through a mesh, mixing it with wheat straw and ionized water to form dry hard soil, and then evenly coating the surface of the substrate with a thickness of 1 cm. The fine mud layer is formed by filtering fine-grained soil through a mesh, mixing it with ionized water and a soluble salt solution, and then coating the surface of the coarse mud layer with a thickness of 0.5 cm. The white powder layer is composed of calcium carbonate and 10%-20% by mass of talc. The four pigments of cinnabar, azurite, malachite, and graphite are evenly coated in parallel strips on the surface of the completely dried sample block. Each pigment area is 2.5 cm wide and 0.2 cm thick. During the natural drying process of the sample blocks, the temperature and humidity are observed and cultivated using soil testing instruments, and the temperature and humidity are adjusted in real time to ensure that the cultivation conditions of each sample block are consistent.
3. The hyperspectral unmixing detection method for soluble salt content in murals according to claim 1 is characterized in that: In step S2, the specific method of sampling and processing the reflection spectrum includes: A spectrometer was used to sample the reflectance spectrum of each standard mural sample four times in a darkroom, 24 hours apart. Whiteboard calibration was performed before each sampling, and outliers caused by point edges or measurement jitter were eliminated. The initial reflectance spectrum curves of the four samples were corrected for breakpoints and averaged to obtain a mixed reflectance spectrum curve. Further noise is eliminated by Savitzky-Golay smoothing, using the formula Perform standardization processing to obtain standardized mixed reflectance spectrum data , where X is the spectral data matrix before processing, μ is the mean of each wavelength, and σ is the standard deviation of each wavelength; When performing PCA principal component analysis dimensionality reduction, construct the covariance matrix , calculate the eigenvalue λ of the covariance matrix C C With the eigenvector ν C , where m is the number of samples; according to the cumulative variance contribution rate , extract the first k principal components corresponding to the cumulative variance contribution rate reaching the preset threshold, and construct the matrix V with these k principal components k , normalize the mixed reflectance spectrum data Projecting to low-dimensional space to obtain the reduced-dimensional data matrix ,in is the sum of the variances of the first k principal components; is the total variance of all principal components.
4. The hyperspectral unmixing detection method for soluble salt content in murals according to claim 1 is characterized in that: In step S3, the specific method of constructing and using the LASSO regression model includes: The optimization objective function is set as: , where m is the number of samples, is the true value of salt concentration of sample i; is the standardized value of sample i on the jth principal component; is the regression coefficient of the jth principal component; The L1 regularization strength parameter is used to control sparsity. The optimal regularization parameter λ that minimizes the model error is selected through 5-fold cross-validation testing. All regression coefficients are calculated and the characteristic wavelengths corresponding to non-zero regression coefficients are screened. In step S4, the specific method of constructing and using the elastic network regression model includes: The objective function is set as , where α is the mixing ratio parameter of L1 and L2 regularization, λ1 and λ2 are the L1 regularization strength and L2 regularization strength respectively. The mean square error (MSE) of different α, λ1, and λ2 parameter combinations is evaluated by 5-fold cross validation. The MSE calculation formula is: , where N is the number of test samples, For the predicted value of the salt concentration of the i-th sample, the optimal parameter combination that can minimize the MSE is selected, and the final model is fitted using the complete training set; the final regression equation is , where X i is the standardized value of the i-th sample, and ε is the model error term.
5. The hyperspectral unmixing detection method for soluble salt content in murals according to claim 1 is characterized in that: The method for verifying the unmixing model and outputting the salt content in step S5 includes: The spectral values of the actual mural hyperspectral data were extracted according to the selected characteristic wavelengths and normalized according to the same normalization parameters as the training data. The normalized hyperspectral data were input into the trained unmixing model to calculate the soluble salt content of each area on the mural surface. Calculate the coefficient of determination R 2 , mean square error MSE, root mean square error RMSE and mean absolute error MAE, calculated as: , , , Among them, Y 真实 is the actual measured salt concentration value, Y 预测 is the salt concentration value predicted by the unmixing model, is the average value of salt concentration; N is the number of test samples, and are the actual and predicted values of the salt concentration of the i-th sample respectively; Evaluate the goodness of fit of the unmixing model: If If the values of MSE, RMSE, and MAE are all lower than the corresponding preset thresholds, the unmixing model is judged to be valid and the salt content distribution map is output.
6. The hyperspectral unmixing detection method for soluble salt content in murals according to claim 1 is characterized in that: The processing of the initial reflection spectrum curve in step S2 includes the following operations: The mutation points of the spectral curve are smoothed by piecewise linear interpolation. The interpolation interval is set to 5nm. The correction formula is: Where R(λ) is the original reflectivity at wavelength λ, R 校正 (λ) is the reflectivity after correction; The 3σ principle is used to eliminate abnormal data points that deviate from the mean by more than 3 times the standard deviation; The Savitzky-Golay smoothing parameters were optimized, and a filter with a window width of 15 nm and a polynomial order of 3 was used for smoothing. The smoothing effect was verified by calculating the signal-to-noise ratio (SNR), and the SNR was processed to SNR ≥ 30 dB.
7. The method for detecting soluble salt content in murals by hyperspectral unmixing according to claim 6, characterized in that: The characteristic wavelength screening of the LASSO regression model in step S3 further includes: The wavelength range is segmented, and the spectral range of 400-2500 nm is divided into 10 sub-intervals with a span of 210 nm. LASSO regression is performed independently in each sub-interval. A t-test was performed on the selected characteristic wavelengths, with the significance level set to 0.01, and only wavelengths with a probability value p-value < 0.01 that satisfied the rejection of the null hypothesis were retained; Redundant wavelengths were eliminated and the Pearson correlation coefficient r of all characteristic wavelength pairs was calculated. If |r| of a characteristic wavelength pair was greater than 0.9, the two characteristic wavelengths were considered to be a highly correlated wavelength group. Only the wavelength with the largest absolute value of the regression coefficient was retained in each wavelength group. Combine the screening results of all sub-intervals to form a sparse characteristic wavelength set and record its corresponding salt sensitivity weight , β j is the regression coefficient of the jth wavelength in the LASSO regression model. The wavelengths with weights lower than 0.05 are eliminated to generate the final feature set.
8. The hyperspectral unmixing detection method for soluble salt content in murals according to claim 7 is characterized in that: The optimization of the elastic network regression model in step S4 includes: The mixed regularization parameters are dynamically adjusted, and the initial value of the mixed ratio parameter α of L1 and L2 regularization is set to 0.
5. The search range of L1 regularization strength λ1 and L2 regularization strength λ2 is [10 −4 ,10 2 ], the optimal parameter combination is determined by grid search combined with 5-fold cross validation, and the optimization goal is set to minimize the mean square error (MSE) of the validation set; During the training process, the Dropout mechanism is added to randomly block 20% of the feature inputs, and the training is repeated 10 times and the mean of the regression coefficient is taken as the final solution; the non-negativity constraint X is added. j β j ≥0 and the abundance sum is 1 constraint ∑β j =1,X i is the standardized value of the i-th sample, β j is the regression coefficient of the jth wavelength in the LASSO regression model.
9. The hyperspectral unmixing detection method for soluble salt content in murals according to claim 1, characterized in that: The unmixing model verification in step S5 further includes: The actual mural hyperspectral data was divided into 0.5cm×0.5cm pixel grids, and the average reflectance of the selected characteristic wavelengths within each grid was extracted to ensure consistency with the spatial scale of the training samples. Output the salt content Y of each pixel through the unmixing model i , and the Kriging interpolation method was used to generate the spatial distribution heat map of salt in the whole mural through the Gaussian semivariogram model; Compute prediction confidence intervals for each pixel , SE is the standard error, is the predicted value of the salt concentration of the i-th sample. If the CI width exceeds 0.1%, the area is marked as an area requiring review, and data is recollected and the unmixing model is updated for the area requiring review.
10. The hyperspectral unmixing detection method for soluble salt content in murals according to claim 9, characterized in that: The standard error SE is calculated by the following steps: Construct the regression coefficient β of the elastic network regression model j The covariance matrix Cov(β) of the relationship between , where X is the characteristic wavelength data matrix, σ 2 is the residual variance of the training set, I is the identity matrix, D is the L2 regularization weight matrix, where α is the mixing ratio parameter of L1 and L2 regularization, λ1 and λ2 are the L1 regularization strength and L2 regularization strength respectively; Calculate the prediction variance: , where X i is the characteristic wavelength vector of the i-th pixel, σ 2 is the residual variance of the training set; The reflectance of the characteristic wavelength is perturbed by Monte Carlo sampling with a perturbation amplitude of ±1%, and the standard deviation of the salt content is statistically output as SE. If more than 10% of the pixels in the entire mural area have a CI width exceeding 0.1%, the model recalibration process is triggered, and the hyperspectral data of the area is recollected and the unmixing model is updated.
Citation Information
Patent Citations
Monitoring method of salt ion content in saline soil based on hyperspectral technology
CN104897592A
Mural soluble salt content detection method based on reflection spectrum
CN114279976A
Earthen ruin cultural relic surface water content detection method based on hyperspectral imaging
CN114878508A
FOD-based hyperspectral prediction method for phosphate content of mural ground layer
CN117990616A
Method for detecting phosphate content of mural ground layer based on three-band spectrum
CN118937249A
Cited By
Method for hyperspectral inversion of sodium sulfate concentration in mural by considering temperature and humidity
CN121253517A
Method for retrieving concentration of sodium sulfate in fresco considering temperature and humidity by hyperspectral
CN121253517B