Water vapor infrared spectrum humidity correction method based on big data analysis
By constructing a water vapor infrared spectral humidity correction method based on big data analysis, the three-dimensional porosity and humidity gradient data of porous ceramic materials are obtained, and combined with distributed feature engineering, the spectral drift problem caused by micropore structure in high humidity environments is solved, and high-precision humidity correction and microcrack detection are achieved.
Patent Information
- Application Number
- CN202510761718.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-09
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2045-06-09
AI Technical Summary
The prior art cannot effectively compensate for local absorption differences caused by the microporous pore structure of porous ceramic materials in high humidity environments, and the traditional humidity correction method is separated from crack detection, making it difficult to meet the needs of high-precision and high-speed detection.
By acquiring the three-dimensional porosity distribution and surface humidity gradient data of porous ceramic materials, combined with big data analysis and distributed feature engineering, a regression and classification model of the combined output spectral humidity correction coefficient and microcrack classification is constructed, and the organic fusion of micropore structure parameters acquisition, surface humidity gradient acquisition, big data feature extraction, distributed regression modeling and microcrack detection is realized.
It significantly improves the robustness of spectral characteristics and the accuracy of crack detection, meets the online defect diagnosis needs of porous ceramics in high humidity environments, and realizes high-precision humidity correction and automatic microcrack detection.
Smart Images

Figure CN120253750A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of near-infrared spectroscopy humidity measurement, and specifically to a water vapor infrared spectroscopy humidity correction method based on big data analysis. Background Technique
[0002] In a high-humidity environment, the adsorption of water vapor by the internal pores of porous ceramic materials will cause a significant shift in the near-infrared absorption peak, resulting in an increase in the deviation of traditional infrared spectroscopy analysis results. In recent years, researchers have tried to perform macroscopic correction through the overall temperature and humidity sensor data, or use humidity compensation algorithms based on physical models to reduce the influence of water vapor on the spectral signal; at the same time, with the rise of big data and distributed computing platforms, pattern recognition and machine learning methods based on large-scale data samples have been applied to spectral preprocessing and feature extraction. However, these studies mostly focus on single-dimensional humidity compensation or crack detection, and have not yet formed a linkage correction and online diagnosis process that can take into account the differences in microscopic pore structures and macroscopic environmental humidity;
[0003] In the prior art, the publication number is CN115128034B, and the name is a method for co-correcting near-infrared spectral data with temperature and humidity. Its core idea: Use a near-infrared spectrometer to collect spectral data of samples under different temperature and humidity conditions, calculate the mean value of the spectral data of samples under different temperature and humidity conditions, then calculate the temperature and humidity spectral correction coefficient by combining the sample temperature and humidity values and the corresponding spectral mean value, and calculate the temperature and humidity spectral correction formula by combining this coefficient, and finally combine the temperature and humidity spectral correction formula to correct the spectral data of the sample to be corrected to complete the correction of the spectral data.
[0004] There are still many deficiencies in the prior art:
[0005] 1. First, relying only on the overall temperature and humidity values for correction cannot effectively compensate for the local absorption differences caused by the microscopic pore structure of the material, resulting in poor stability of the corrected spectral features; secondly, traditional humidity correction methods and crack detection methods belong to different modules, lacking a unified data processing and modeling process, making the system integration low and the real-time performance insufficient;
[0006] 2. Thirdly, one-way algorithms are difficult to take into account the heterogeneous features and online calculation requirements of large-scale samples, and it is difficult to meet the comprehensive requirements of high precision, high reliability and high-speed detection in industrial production.
[0007] The above information disclosed in the background art section is only used to strengthen the understanding of the background of the present disclosure, so it may include information that does not constitute the prior art known to those of ordinary skill in the art. Summary of the Invention
[0008] The purpose of the present invention is to provide a water vapor infrared spectroscopy humidity correction method based on big data analysis to solve the problems raised in the above background technique.
[0009] To achieve the above object, the present invention provides the following technical solutions:
[0010] A method for correcting water vapor infrared spectrum humidity based on big data analysis, the specific steps include:
[0011] Step S1: Obtain the three-dimensional porosity distribution data of P samples of materials to be measured;
[0012] Collect the local humidity gradient data of multiple preset sampling points on the surface of the samples of materials to be measured;
[0013] Perform preprocessing on the three-dimensional porosity distribution data and the local humidity gradient data within the same output range;
[0014] Step S2: Based on the preprocessing results, obtain the porosity and humidity gradient data of P samples of materials to be measured, and upload them to the big data analysis platform for distributed feature extraction and dimensionality reduction processing to obtain the dimensionality-reduced features and spectral offset data;
[0015] Step S3: Based on the dimensionality-reduced features and spectral offset data, construct a distributed regression model and obtain the regression coefficients to obtain the predicted values of spectral offset;
[0016] Step S4: Based on the predicted values of spectral offset, construct an infrared spectrum humidity correction coefficient and apply it to the original infrared spectrum data to obtain the corrected infrared spectrum data;
[0017] Step S5: Extract the peak shape features based on the corrected infrared spectrum data and classify the microcracks of the samples of materials to be measured.
[0018] Compared with the prior art, the beneficial effects of the present invention are as follows: By introducing the synchronous acquisition of the three-dimensional porosity distribution of porous materials and the surface humidity gradient, combining distributed feature engineering and dimensionality reduction processing, constructing a big data regression and classification model with joint outputs of "spectral humidity correction coefficient" and "microcrack classification result", the organic integration of six technical elements, namely, obtaining microscopic pore structure parameters, collecting surface humidity gradients, extracting big data features, constructing distributed regression models, constructing correction formulas, and detecting microcracks, is realized, significantly improving the robustness of the corrected spectral features and the accuracy of crack detection, and meeting the technical requirements for online defect diagnosis of porous ceramics in high humidity environments. Description of the Drawings
[0019] Figure 1 It is a schematic diagram of the overall method flow of the present invention;
[0020] Figure 2 It is a schematic diagram of the application of the porous ceramic material of the present invention. Detailed Embodiments
[0021] To make the above objects, features, and advantages of the present invention more apparent and understandable, the following provides a detailed description of the specific embodiments of the present invention in conjunction with the accompanying drawings of the specification.
[0022] In the following description, many specific details are set forth to facilitate a thorough understanding of the present invention. However, the present invention may be practiced in other ways different from those described herein. Those skilled in the art can make similar generalizations without departing from the spirit of the present invention. Therefore, the present invention is not limited by the specific embodiments disclosed below.
[0023] For porous ceramic materials in a high-humidity environment, moisture adsorbed in the internal pores causes a shift in the near-infrared spectral absorption peak; and relying solely on the overall temperature and humidity values for correction cannot compensate for the local absorption differences caused by the microscopic pore structure. This solution has six defined technical features: "acquisition of microscopic pore structure parameters", "collection of surface humidity gradients", "big data feature engineering", "distributed regression modeling", "construction of correction formulas", and "microcrack detection", to achieve the linkage of high-precision humidity correction and automatic detection of material microcracks; specifically including the following embodiments:
[0024] Embodiment 1:
[0025] Please refer to Figure 1 and Figure 2 , the present invention provides a technical solution:
[0026] A method for correcting water vapor infrared spectrum humidity based on big data analysis, applied to porous ceramic materials in a high-humidity environment, the specific steps include:
[0027] Step S1: Obtain three-dimensional porosity distribution data of P samples of the material to be measured;
[0028] Collect local humidity gradient data of multiple preset sampling points on the surface of the material sample to be measured;
[0029] Perform preprocessing on the three-dimensional porosity distribution data and the local humidity gradient data with the same output range;
[0030] Further explanation: Number the P samples of the material to be measured as {1, 2,..., i,..., P}; where i represents the index of the material sample to be measured;
[0031] Place the material sample to be measured into an X-ray micro-computed tomography scanner (micro-CT, hereinafter referred to as "micro-CT scan") to obtain three-dimensional pore structure data of the material sample to be measured;
[0032] Specifically, select a Bruker SkyScan 1275 type of micro-CT scan, and the specific parameter settings are as follows:
[0033] Tube voltage: 80 kV, current: 100 µA; Voxel resolution: 10 µm;
[0034] The scanning parameters for each sample of the material to be measured are 360° rotation, 0.2° step, and single-sample scanning for 15 minutes;
[0035] Based on the acquired three-dimensional pore structure data, the Feldkamp–Davis–Kress (FDK) algorithm implemented in NRecon software is used to perform three-dimensional reconstruction on the sample of the material to be measured, thereby generating a voxel matrix containing gray values ;
[0036] It should be noted that: The implementation steps of the Feldkamp–Davis–Kress (FDK) algorithm in this embodiment are as follows:
[0037] 1.1) Data acquisition: By rotating the cone-beam X-ray source and detector, two-dimensional projection images at different angles are acquired;
[0038] 1.2) Filtering process: Use the Ram-Lak filter or other high-pass filters to perform frequency-domain filtering on each projection image to compensate for the image blur caused by the cone-beam geometry.
[0039] 1.3) Projection data acquisition: Place the sample i of the material to be measured in the micro-CT scanner, and acquire two-dimensional projection images at different angles every 0.2° .
[0040] 1.4) Filtering process: Apply the Ram-Lak filter h6 to each projection image to obtain the filtered projection ;
[0041] 1.5) Weighted backprojection: According to the FDK algorithm, backproject the filtered projection images according to the weights to cumulatively generate a three-dimensional volume image .
[0042] 1.6) Porosity calculation: By setting the gray threshold T, the voxel matrix is binarized into solid phase and pores, and the average porosity and the standard deviation of porosity of the sample of the material to be measured are calculated. Specifically, it includes:
[0043] Set the gray threshold T for image segmentation to obtain ;
[0044] When ; it means that the sample i of the material to be measured is the solid phase;
[0045] When ; it means that there are pores in the sample i of the material to be measured;
[0046] The solid phase refers to the continuous and solid part that constitutes the main structure of the material sample. In the porosity distribution data, the solid phase represents the actual material part of the material, without any pores or holes.
[0047] Pores refer to the cavity areas such as holes, cracks or other forms existing in the material sample. These areas do not contain any actual material and are the "blank" parts inside the material.
[0048] Set the total number of voxels to ; Calculate the average porosity and the standard deviation of porosity ;
[0049]
[0050]
[0051] Among them, represents the pore identification of the v-th voxel on the material sample i to be measured, and ;
[0052] Further explanation: The local humidity gradient data of multiple preset sampling points are obtained by arranging a micro humidity sensor array on the surface of the material sample to be measured; specifically including:
[0053] Utilize the characteristic absorption peak of water vapor in the mid-infrared band, and obtain the humidity distribution by measuring the absorbance of multiple preset sampling points on the surface of the material sample i to be measured in real time through a spectrometer;
[0054] The spectral measurement instrument in this embodiment uses Thermo Nicolet iS50 (FTIR); wavelength range 2.5 µm - 25 µm; resolution 4 cm -1 ;
[0055] Set the distance of the material sample i to be measured as: the laser beam is perpendicular to the surface, and the optical path length l = 0.5 cm;
[0056] Apply the Beer–Lambert law to characterize the absorbance–humidity relationship: ; Among them, is the absorbance of the preset sampling point j on the material sample i to be measured; is the molar absorption coefficient of water vapor at the absorption peak ; is the water vapor concentration in the air at the preset sampling point j on the material sample i to be measured; is the optical path length;
[0057] The at the preset sampling point j on the material sample i to be measured is converted to relative humidity ;
[0058]
[0059] where is the saturated water vapor concentration at 25 °C; in this embodiment is 0.023 mol·L -1 ; is 0.12 L·cm -1 ; is 0.5 cm;
[0060] There are multiple preset sampling points in this embodiment points, and the single-point repeated measurement is K3 = 3 times, and the average value is taken;
[0061] Calculate the average relative humidity of the sample i of the material to be measured and the standard deviation of the humidity gradient ;
[0062]
[0063]
[0064] where, as increases, it indicates that increases, and then increases; M is the total number of preset sampling points, K3 is the total number of repeated measurements, and k3 is the index of the number of repeated measurements;
[0065] When the fluctuation range of increases,
[0066] increases, indicating uneven humidity distribution. , , the average porosity and the standard deviation of the porosity are associated and stored in the database.
[0067] Further explanation: The preprocessing of the same output range includes: respectively performing Min–Max normalization on , , and to obtain the normalized output results , , and .
[0068]
[0069] Output , to ensure the stability of the subsequent model input:
[0070] When As it approaches 0 more closely, the porosity is lower; As it approaches 1 more closely, the porosity is higher;
[0071] When As it approaches 0 more closely, the humidity is more uniform; As it approaches 1 more closely, the humidity distribution is more uneven;
[0072] It should be noted that , and The normalization method is the same as above and will not be elaborated.
[0073] Step S2: Based on the preprocessing results, obtain the porosity and humidity gradient data of P samples of the material to be tested, and upload them to the big data analysis platform for distributed feature extraction and dimensionality reduction processing to obtain the dimensionality-reduced features and spectral offset data;
[0074] Further explanation: In this embodiment, the number of samples P of the material to be tested is 500;
[0075] Obtain the porosity and humidity gradient features of P = 500 samples of the material to be tested after normalization, that is , , and ;
[0076] For the P samples of the material to be tested , , and , as well as the corresponding infrared spectral main peak offset, summarize and upload them to the big data analysis platform for distributed feature cleaning, standardization, and principal component analysis (PCA) dimensionality reduction, and at the same time store the dimensionality-reduced features and offsets together in the feature library; specifically including:
[0077] 2.1) Integrate the four-dimensional normalized feature vector of each sample i of the material to be tested with the corresponding infrared spectral offset into a row record, generate a CSV file containing P rows and 5 columns (4 features + 1 offset), and upload it to the HDFS path of the big data platform through the secure file transfer protocol (SFTP);
[0078] 2.2) Configure a distributed storage and computing environment on a 20-node Hadoop / Spark cluster, specifically:
[0079] HDFS storage: Import the CSV into the HDFS directory;
[0080] Spark Environment: Spark version 3.2.1, Java 8, configured with 4 cores × 16GB memory per node for Executor; when configured with 20 nodes, it can parallelize and accelerate data cleaning and matrix operations when processing 500×5-dimensional data, taking into account resource utilization and response timeliness.
[0081] 2.3) According to the distributed environment deployment, use PySpark to complete missing value filling and standardization processing, specifically as follows:
[0082] For each feature column Perform mean filling: Replace missing entries with the mean of the entire sample . r is a positive integer from 1 to 4, respectively representing , , and ;
[0083] Define the standardization formula as ; where is the r-th original feature of the material sample i to be measured; and are the mean and standard deviation of the r-th original feature in the P = 500 samples respectively; is the standardized feature;
[0084] When increases, if , then is positive and its absolute value increases with the increase of the difference;
[0085] Conversely, when , is negative and its absolute value increases with the increase of the difference;
[0086] The Z-score standardization in this embodiment can eliminate the dimensionality differences between different features, making the weights of each feature fair when the subsequent PCA selects the principal components according to the variance size;
[0087] 2.4) Perform principal component analysis PCA for dimensionality reduction, specifically including covariance matrix calculation, eigenvalue decomposition, principal component selection, and dimensionality reduction mapping;
[0088] According to the standardization processing, realize dimensionality reduction to the principal component space with a cumulative variance contribution rate ≥ 95% by calculating the covariance matrix and solving the eigenvalue problem:
[0089] The covariance matrix calculation is as follows:
[0090]
[0091] where ; T1 represents the matrix transpose operation, specifically transposing a column vector into a row vector; P represents the total number of samples of the material to be measured, and the eigenvalue decomposition is as follows:
[0092] Solve , and sort by ;
[0093] The principal components are selected as follows:
[0094] Calculate the cumulative variance contribution rate , and take K = 2;
[0095] Obtain the first two principal component vectors .
[0096] The dimensionality reduction mapping is as follows:
[0097]
[0098] Among them, is the score of the sample i of the material to be measured on the r-th principal component;
[0099] If the loading coefficient of the input feature in the principal component vector is relatively high, the positive change of this feature will significantly improve ;
[0100] The comprehensive change trend of the input features can be reflected by the principal component scores.
[0101] According to the PCA dimensionality reduction, the principal component scores are Min–Max normalized and stored in the feature library together with the offset for the model to quickly read: The normalization formula is:
[0102]
[0103] Among them, , are respectively the minimum and maximum scores of all samples of the material to be measured on the r-th principal component; ; is the normalized output value;
[0104] When gets closer to 0, the comprehensive feature of the sample i of the material to be measured in the direction of the r-th principal component is weaker;
[0105] When gets closer to 1, the comprehensive feature of the sample i of the material to be measured in the direction of the r-th principal component is stronger;
[0106] Normalizing the scores helps with the numerical stability in the subsequent multi-objective regression model.
[0107] Set the offset of the normalized infrared spectrum to ;
[0108] Characterize the normalized dimensionality-reduced features and the spectral offset data as .
[0109] The offset of the normalized infrared spectrum The increase or decrease directly reflects the degree of influence of humidity on the position of the main peak of the infrared spectrum;
[0110] The offset of the normalized infrared spectrum An increase indicates that the influence of humidity on the spectral main peak is enhanced and stronger humidity correction is required;
[0111] The offset of the normalized infrared spectrum A decrease indicates that the influence of humidity on the spectral main peak is weakened and the humidity correction requirement is low.
[0112] This embodiment forms a complete and implementable technical process from the original normalized features to the low-dimensional stable features, providing a solid data foundation for the efficient training of the subsequent spectral humidity correction model.
[0113] Step S3: Construct a distributed regression model based on the dimensionality-reduced features and the spectral offset data and obtain the regression coefficients to get the predicted value of the spectral offset;
[0114] Further explanation: According to the dimensionality-reduced features and the spectral offset data, use the normalized principal component scores and the spectral offset as input and output variables to construct a multiple linear regression model in a distributed computing environment, and extract the regression coefficients for online spectral offset prediction; specifically including:
[0115] 3.1) Load the normalized principal component scores and the spectral offset stored in the distributed file system as analysis data, and divide the data set according to the ratio of 80% training / 20% testing;
[0116] Read from the HDFS path the normalized principal component scores represented by the column fields "z1", "z2", "y" respectively , and ;
[0117] After randomly shuffling all the data, allocate the training set and the test set according to the ratio;
[0118] Select 80% as the model training data to ensure the statistical reliability of the model parameter estimation; the remaining 20% is used for model performance evaluation.
[0119] 3.2) The parameters of the distributed regression model are set as follows:
[0120] Based on the training set, a multiple linear regression model with L2 regularization is selected, and the regularization coefficient and the iteration termination condition are determined;
[0121] The multiple linear regression model uses Ridge regression to suppress the possible collinearity between different principal components;
[0122] Regularization coefficient : Selected in cross-validation to balance the bias and variance of the model;
[0123] Maximum number of iterations and convergence threshold : Ensure that the parameter solution converges sufficiently while taking into account the computational efficiency.
[0124] 3.3) According to the model configuration, perform least squares fitting on the training set and extract the intercept and the regression coefficients of each principal component to construct a prediction function.
[0125] Using the data in the training set , fit the Ridge regression model to solve the regression coefficients , and .
[0126] Intercept term represents the reference value of the predicted spectral offset;
[0127] Regression coefficient and : Correspond to the weights of the dimensionality-reduced features and respectively, reflecting their influence on the predicted spectral offset value;
[0128] Denote the output value of the distributed regression model of the material sample i to be measured as the predicted value ;
[0129] If , then when the corresponding dimensionality-reduced feature increases, the predicted value also increases;
[0130] If , then when the corresponding dimensionality-reduced feature increases, the predicted value decreases;
[0131] Based on the test set and the extracted regression coefficients, calculate the predicted value of each test sample and restore it to the actual offset through linear mapping.
[0132] Predicted value The calculation formula is:
[0133]
[0134] Denote the predicted spectral offset of the material sample \(i\) to be measured as ; and perform the following inverse normalization mapping:
[0135]
[0136] where and are the maximum and minimum values of the infrared spectral offset before normalization, respectively;
[0137] When gets closer to 0, gets closer to the minimum infrared spectral offset, indicating that the humidity effect on the material sample to be measured is weaker;
[0138] When gets closer to 1, gets closer to the maximum infrared spectral offset, indicating that the humidity effect on the material sample to be measured is stronger.
[0139] Furthermore, according to the prediction results and the true values of the test set, calculate the coefficient of determination and the root mean square error to evaluate the model performance, and persist the model parameters to the model library to support online prediction.
[0140] Set the evaluation indicators as:
[0141]
[0142]
[0143] where, is the actual offset of the material sample \(i\) in the test set; is the predicted spectral offset of the material sample \(i\) in the test set; is the average offset of all material samples to be measured in the test set; is the number of samples in the test set, and is less than the total number \(P\) of the material samples to be measured;
[0144] The coefficient of determination measures the explanatory ability of the model for the actual offset. The closer the value is to 1, the better the model fitting effect;
[0145] The root mean square error RMSE reflects the average deviation between the predicted value and the actual value of the model. The closer the value is to 0, the higher the model prediction accuracy;
[0146] Compare with and Deposit it into the model registry and load it in the distributed service for real-time prediction.
[0147] Step S4: Based on the predicted value of the spectral offset, construct an infrared spectral humidity correction coefficient and apply it to the original infrared spectral data to obtain the corrected infrared spectral data;
[0148] Further explanation: In industrial on-line quality inspection, the original infrared spectral data often shows peak position offset due to the influence of environmental water vapor; by quickly generating the correction coefficient through the predicted value of the spectral offset and making real-time correction, the detection accuracy and low latency can be maintained in the high-throughput detection scenario;
[0149] Obtain the predicted value of the sample i of the material to be tested from the output of the distributed regression model ;
[0150] For each sample i of the material to be tested, read the predicted value of the spectral offset calculated and stored for it ;
[0151] This embodiment ; directly use It can ensure the numerical stability and consistency of the correction coefficient calculation;
[0152] According to the read predicted value of the spectral offset , adopt linear mapping as the correction coefficient to ensure ;
[0153] where is the spectral correction coefficient of the sample i of the material to be tested, and its range is limited to the interval (0, 1);
[0154] When gets closer to 0, gets closer to 1, indicating that the required correction amplitude is smaller;
[0155] When gets closer to 1, gets closer to 0, indicating that the required correction amplitude is larger;
[0156] Based on the spectral correction coefficient , multiply the original infrared spectral data of the sample i of the material to be tested point by point by to generate the corrected infrared spectral data;
[0157] The corrected infrared spectral data includes the corrected wavelength value of the sample i of the material to be tested at the jth preset sampling point ;
[0158] The specific correction steps include:
[0159] The original infrared spectral data is represented as: the original spectral wavelength sequence of the sample i of the material to be measured , where j is the index of the preset sampling point;
[0160]
[0161] where is the original wavelength value of the sample i of the material to be measured at the j-th preset sampling point; is the corrected wavelength value of the sample i of the material to be measured at the j-th preset sampling point.
[0162] When increases, each tends to
[0163] When decreases, each
[0164] is point-by-point multiplication, which is simple and efficient, suitable for implementation on streaming computing or device-side FPGA / ASIC, and meets the real-time requirements.
[0165] Further explanation: The corrected infrared spectral data is used as the input of the defect detection module for subsequent analysis.
[0166] Push to the subsequent defect detection or feature extraction pipeline in the same data format as the original infrared spectral data.
[0167] Further explanation: The following embodiments are based on the content described in step S4, and the innovation and advantages of the present invention in infrared spectral humidity correction are demonstrated through specific data;
[0168] To verify the effectiveness of the proposed method, six porous material samples are selected for indoor environment simulation tests. The relative humidity of the test environment is maintained at 60% ± 2%, and the spectral measurement instrument uses a Fourier transform infrared spectrometer (FTIR) with a resolution of 0.5 nm, and the measurement band is 1000 cm -1 at the main absorption peak position offset. Each sample is implemented according to the following process:
[0169] First, obtain the reference absorption peak position of the six samples under standard dry conditions and record it as the reference peak position. Subsequently, under 60% humidity conditions, each sample is measured three times repeatedly, and the peak position offset is recorded, and the average value is taken as the original peak position offset. Substitute into the linear mapping formula of the present invention to calculate the correction coefficient. Then multiply the peak position offset of each sample of the material to be measured by the corresponding , the corrected peak position offset is obtained. Finally, the offset errors before and after correction are compared to evaluate the correction effect.
[0170] During the experiment, all data are recorded in a spreadsheet for subsequent statistical analysis; no interface modification is required for the downstream defect detection module throughout the process. The corrected wavelength data are directly input to verify the universality and real-time performance of this method on different samples. The data obtained in this embodiment are shown in Table 1:
[0171] Table 1 Research on the effectiveness of the correction coefficient:
[0172] As can be seen from Table 1, the average reduction rate of the peak position offset error of the method of the present invention for six samples reaches about 62.5%, which is significantly better than the traditional method without correction or with a fixed correction coefficient. The data in the embodiment fully prove that:
[0173] The linear mapping correction coefficient calculated dynamically using the normalized predicted offset value can adjust the correction intensity according to the degree of humidity influence;
[0174] This method not only ensures a significant reduction in the offset error after correction, but also has good universality for different materials;
[0175] When the predicted value approaches 0 (Sample 4), the correction coefficient approaches 1, and the error after correction is only 0.18 nm, maintaining the authenticity of the spectrum; when the predicted value approaches 1 (Sample 6), it approaches 0, and the error after correction drops to 0.18 nm, achieving the maximum correction amplitude.
[0176] Step S5: Extract the peak shape features based on the corrected infrared spectrum data and classify the microcracks of the sample to be measured.
[0177] Further explanation: The defect detection module includes classifying the microcracks of the sample to be measured, specifically including:
[0178] Based on the corrected wavelength value , extract the peak width and peak intensity in the peak shape of each sample i to be measured; the specific operation steps are as follows:
[0179] 5.1) Perform Gaussian fitting on the corrected wavelength value of the sample i to be measured near the target absorption peak. In this embodiment, the range near the target absorption peak is within ±5 nm of the absorption peak; the fitting function is:
[0180]
[0181] where, is the corrected wavelength value; is the peak center wavelength after fitting; is the Gaussian standard deviation; A is the fitting amplitude;
[0182] Peak width The calculation formula is as follows:
[0183]
[0184] Peak intensity is taken as the fitting amplitude ; when the crack causes the peak to become wider, that is, increases, the peak width increases; when the crack causes the absorption to weaken, the peak intensity decreases; the peak width and peak intensity are input into the support vector machine classifier for crack detection;
[0185] Use a pre-trained radial basis function (RBF) kernel support vector machine to classify the microcracks of the material sample i to be tested, so as to obtain a classification decision function for determining whether there are microcracks in the material sample i to be tested ; classification decision function is used to output the classification label of the material sample i to be tested; the specific logic includes setting the classification function of microcrack classification as:
[0186]
[0187] Among them, is the feature vector of the material sample i to be tested; are the support vector weights and true value labels in the training set respectively and the corresponding feature vectors; is the classification decision function for determining whether there are microcracks in the material sample i to be tested; L is the total number of support vectors, and k2 is the index of the support vector;
[0188] is the RBF kernel width hyperparameter, and the value in this embodiment is ; b1 is the classifier bias term; it controls the attenuation speed of the kernel function, and the larger the value, the more concentrated the kernel function;
[0189] is the true value label, +1 means "there is a crack", and -1 means "no crack";
[0190] When decreases, the kernel function value increases, and the influence of the support vector increases;
[0191] When increases, the response to data far from the support vector weakens, and the model pays more attention to the local;
[0192] Map the classification labels to the coordinates of the sensor array and perform 3D visualization; specifically including:
[0193] Map the coordinates of the sensor units of the test material sample i determined to be "crack present", and generate an annotation report in combination with the 3D pore structure data; the specific logic includes:
[0194] If , then the corresponding sensor unit of the test material sample i has a microcrack; is the spatial position of the test material sample i in the array;
[0195] The sensor array coordinates of this embodiment are obtained during equipment calibration and stored uniformly. And execute the following visualization process:
[0196] Load the 3D pore structure data; highlight the crack coordinate points in the model, and the color depth is proportional to the absolute value of the classification decision function value of ;
[0197] Output the annotated 3D view and generate a report in PDF or HTML format.
[0198] The absolute value of the classification decision function output reflects the crack confidence level.
[0199] When the absolute value of the classification decision function increases, the crack confidence level is higher, and the annotation color in the report is deeper;
[0200] The characteristic value range description is ;
[0201] When the peak width or the peak intensity significantly deviates from the crack-free reference value, decreases, the kernel function value increases, the absolute value of the decision function value increases, and finally , determine a crack;
[0202] When the peak shape feature approaches the normal reference range, near the decision boundary, the value range approaches zero, and "awaiting review" is marked on the production line.
[0203] The following embodiments, based on step S5, verify the innovation and beneficial effects of the present invention in the extraction of modified infrared spectral peak shape features and microcrack classification through specific experimental processes and data;
[0204] Six porous ceramic material samples were selected, numbered from "Sample 1" to "Sample 6". Among them, "Sample 1", "Sample 3", and "Sample 5" were pre-prepared with artificial microcracks with a width of about 0.05 mm on the surface, and the remaining three samples had no cracks. The size of all samples was 20 mm × 20 mm × 5 mm. First, all samples were placed in a standard drying chamber (relative humidity ≤ 5%) for 20 min to stabilize. A Fourier transform infrared spectrometer (FTIR) with a resolution of 0.2 nm was used to obtain the center wavelength of the reference absorption peak -1 at 1000 cm , and the baseline peak width and peak intensity were recorded. Subsequently, the samples were transferred to a simulated production line environmental chamber (relative humidity 60% ± 2%, temperature 25 °C), and the original infrared spectral data of each sample were continuously collected three times at 1-min intervals , with the preset number of sampling points M = 100. For each collected data, the predicted value of the spectral offset and the correction coefficient in step S4 were calculated to generate the corrected wavelength value .
[0205] After obtaining , it entered the peak shape feature extraction process of step S5: The data within ±5 nm of the target absorption peak were taken, and an improved fast Gaussian fitting algorithm (based on Levenberg–Marquardt iteration, with a maximum of 50 iterations and a residual threshold of 1×10 -6 ) was called, and the fitting function was used to output the center wavelength of the fitted peak, the standard deviation , and the fitting amplitude A. According to the formula , the peak width w (unit: nm) was calculated, and the peak intensity I was directly taken as A. To improve real-time performance, the algorithm was implemented on an FPGA acceleration module, and each fitting took 8 ms; after fitting, the feature vector x = [w, I] was normalized to the interval (0, 1) and sent to a radial basis kernel SVM classifier that had been trained on 100 groups of known crack samples in advance. The kernel width hyperparameter , and the number of support vectors L = 30;
[0206] where represented "crack present" and "no crack" respectively. The classifier was embedded in an embedded CPU, and the CPU and FPGA performed parallel computing, with a total time consumption of about 15 ms. The classification result was judged as "crack present", was judged as "no crack"; at the same time, the average decision function value was output for subsequent visualization confidence annotation. The average decision value of the three measurements of all samples was used as the final judgment basis to reduce accidental errors.
[0207] After the experiment, the average peak width, average peak intensity, average decision function value of each sample , SVM determination label, actual crack state, and whether the determination is correct and other indicators are summarized and the average determination time is statistically calculated, and shared to the data processing server for comparative analysis. In this embodiment, by comparing with the control group without humidity correction, the advantages of the method of the present invention in terms of classification accuracy, determination stability and real-time performance are quantified.
[0208] Table 2 Research on the extraction of infrared spectrum peak shape characteristics and microcrack classification after correction:
[0209] It can be seen from the above data:
[0210] 1. Classification accuracy: The average decision function values of all crack samples (Samples 1, 3, 5) are all greater than 0, and the determination label is +1, accurately reflecting the actual existence of cracks.
[0211] The average decision function values of the crack-free samples (Samples 2, 4, 6) are all less than 0, and the determination label is -1, accurately reflecting the actual absence of cracks.
[0212] The classification accuracy reaches 100%, significantly higher than the level of 80% - 85% without humidity correction, demonstrating a significant improvement in the classification accuracy of the method of the present invention.
[0213] 2. Feature stability: The peak width and peak intensity extracted from the corrected infrared spectrum data are highly correlated with the actual crack state, reducing the interference of humidity changes on spectral features and enhancing the robustness of feature extraction.
[0214] 3. Real-time performance: The average determination time is only 15 ms, meeting the real-time requirements of industrial on-line high-speed quality inspection, and is suitable for real-time microcrack detection in large-scale production environments.
[0215] The average decision function value indicates that the feature vector is on the "crack exists" side of the classification boundary, and vice versa.
[0216] Through the absolute value of the decision function value, the classification confidence can be quantified. The larger the value, the higher the confidence; the closer it is to 0, the higher the uncertainty, and the more manual review is required.
[0217] The advantage of this solution is the linkage control: the same set of data processing and modeling processes not only outputs the humidity correction coefficient, but also directly drives the automatic crack detection, realizing the seamless linkage of the two major technical effects.
[0218] It should be noted that all the calculation formulas in this application document adopt regression analysis including but not limited to machine learning algorithms to deeply analyze the relevant parameters collected, identify their natural trends and interrelationships. Using professional software such as the Scikit-learn library of Python or the R language, a mathematical model matching the data is automatically generated. Then, the performance of the model is objectively evaluated through methods such as cross-validation, and combined with continuous feedback and optimization to ensure that the created formula truly reflects the internal laws of the data, thereby ensuring its effectiveness and accuracy. In all the calculation formulas of this application, the parameters in each formula are processed by dimensionless normalization within a consistent range to ensure that different physical quantities can be compared on the same scale; the technical means of dimensionless normalization include but not limited to Min-Max-Normalization and Z-Score standardization;
[0219] Essentially, the technical solution of the present invention, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as a floppy disk, read-only memory (ROM), random access memory (RAM), flash memory, hard disk or optical disc of a computer, etc., and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods of various embodiments of the present invention.
[0220] The logic and / or steps represented in the flowchart or described in other ways herein, for example, can be considered as a definite sequence list of executable instructions for implementing logical functions, and can be specifically implemented in any computer-readable medium for use by an instruction execution system, apparatus or device (such as a computer-based system, a system including a processor, or other systems that can fetch and execute instructions from the instruction execution system, apparatus or device), or in combination with these instruction execution systems, apparatus or devices. For the purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate or transport a program for use by or in combination with an instruction execution system, apparatus or device.
[0221] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the technical solutions of the present invention, and they should all be covered within the scope of the claims of the present invention.
Claims
1. A water vapor infrared spectrum humidity correction method based on big data analysis, which is applied to a porous ceramic material sample in a high humidity environment, is characterized in that The specific steps include: Step S1: Obtain the three-dimensional porosity distribution data of P samples of materials to be tested; Collect the local humidity gradient data of multiple preset sampling points on the surface of the samples of materials to be tested; Perform preprocessing on the three-dimensional porosity distribution data and the local humidity gradient data within the same output range; Step S2: Based on the preprocessing results, obtain the porosity and humidity gradient data of P samples of materials to be tested, and upload them to the big data analysis platform for distributed feature extraction and dimensionality reduction processing to obtain the dimensionality-reduced features and spectral offset data; Step S3: Based on the dimensionality-reduced features and spectral offset data, construct a distributed regression model and calculate the regression coefficients to obtain the predicted values of spectral offset; Step S4: Based on the predicted values of spectral offset, construct an infrared spectrum humidity correction coefficient and apply it to the original infrared spectrum data to obtain the corrected infrared spectrum data; Step S5: Extract the peak shape features from the corrected infrared spectrum data and classify the microcracks of the samples of materials to be tested.
2. The water vapor infrared spectrum humidity correction method based on big data analysis according to claim 1, characterized in that: Number the P samples of materials to be tested as {1, 2, …, i, …, P}; where i represents the index of the sample of material to be tested; obtain the three-dimensional pore structure data of the sample of material to be tested; Based on the three-dimensional pore structure data obtained from the sample of the material to be tested, three-dimensional reconstruction is performed on the sample of the material to be tested, thereby generating a voxel matrix containing gray values ; Voxel matrix is binarized into solid phase and pores, and the average porosity of the sample of the material to be measured is calculated and the standard deviation of porosity .
3. The water vapor infrared spectrum humidity correction method based on big data analysis according to claim 2, characterized in that: The local humidity gradient data of multiple preset sampling points are obtained by arranging a micro humidity sensor array on the surface of the sample of material to be tested; specifically including: Utilize the characteristic absorption peak of water vapor in the mid-infrared band, and measure the absorbance of multiple preset sampling points on the surface of the sample i of material to be tested in real time by a spectrometer to obtain the humidity distribution; Apply the Beer–Lambert law to characterize the absorbance–humidity relationship: ; where is the absorbance at the preset sampling point j on the sample i of the material to be measured; is the molar extinction coefficient of water vapor at the absorption peak ; is the water vapor concentration in the air at the preset sampling point j on the sample i of the material to be measured; is the optical path length; Convert the at the preset sampling point j on the material sample i to be measured into relative humidity ; Calculate the average relative humidity of the sample i of the material to be measured and the standard deviation of the humidity gradient .
4. The water vapor infrared spectrum humidity correction method based on big data analysis according to claim 3, characterized in that: Preprocessing of the same output range, including: separately for , , and using Min–Max normalization to obtain the normalized output result , , and .
5. The water vapor infrared spectrum humidity correction method based on big data analysis according to claim 4, characterized in that: For P samples of materials to be measured , , and , together with the corresponding infrared spectrum main peak offset, are summarized and uploaded to the big data analysis platform for distributed feature cleaning, standardization, and principal component analysis for dimensionality reduction. At the same time, the features after dimensionality reduction and the spectral offset data are stored in the feature library together; Characterize the dimension-reduced features after normalization and the spectral offset data as .
6. The method for correcting water vapor infrared spectrum humidity based on big data analysis according to claim 5, characterized in that: According to the dimensionality-reduced features and spectral offset data, use the normalized principal component scores and spectral offset as input and output variables, construct a multiple linear regression model in a distributed computing environment, and extract the regression coefficients for online spectral offset prediction; Denote the output value of the distributed regression model of the material sample i to be measured as the predicted value ; Based on the predicted value , the predicted value of the spectral offset of the material sample i to be measured is denoted as ; When approaches 0 more and more closely, it approaches the minimum infrared spectral shift more and more closely, indicating that the humidity effect of the sample i of the material to be measured is weaker; When getting closer to 1, it gets closer to the maximum infrared spectrum offset, indicating a stronger humidity effect on the sample i of the material to be measured.
7. The method for correcting water vapor infrared spectrum humidity based on big data analysis according to claim 6, characterized in that: Obtain the predicted value of the material sample i to be measured from the output of the distributed regression model ; For each sample i of the material to be measured, read the predicted value of the calculated and stored spectral offset ; Predict the value according to the read spectral offset , and use linear mapping as the correction coefficient to ensure ; wherein is the spectral correction coefficient of the sample i of the material to be measured, and the range is limited to the interval (0, 1); When approaches 0 more and more closely, it approaches 1 more and more closely, indicating that the required correction amplitude is smaller; When getting closer to 1, getting closer to 0, indicating a greater need for correction amplitude; Based on the spectral correction coefficient , multiply each point of the original infrared spectral data of the material sample i to be measured by to generate the corrected infrared spectral data; The corrected infrared spectral data includes the corrected wavelength value of the sample i of the material to be measured at the j-th preset sampling point .
8. The water vapor infrared spectrum humidity correction method based on big data analysis according to claim 7, characterized in that: Use the corrected infrared spectrum data as the input of the defect detection module for subsequent analysis.
9. The method for correcting water vapor infrared spectrum humidity based on big data analysis according to claim 8, characterized in that: The defect detection module includes classifying the microcracks of the samples of materials to be tested, specifically including: Based on the corrected wavelength value , extract the peak width and peak intensity ; Use a pre-trained radial basis function kernel support vector machine to classify the microcracks of the material sample i to be measured, so as to obtain a classification decision function for determining whether there are microcracks in the material sample i to be measured ; The classification decision function is used to output the classification label of the material sample i to be measured; Map the classification labels to the sensor array coordinates and perform three-dimensional visualization; Map the coordinates of the sensor units of the sample i of material to be tested determined to "have cracks", and generate an annotation report in combination with the three-dimensional pore structure data.
Citation Information
Patent Citations
Method and system for analyzing substance concentration based on near infrared spectrum technology
CN114611582A
Near infrared spectrum rapid detection system and real-time analysis method
CN119510347A
Improved spectroscopic device and method for sample characterization
EP3385703A1
Method for classifying scientific materials such as silicate materials, polymer materials and / or nanomaterials
US20080177481A1
Transport of intensity diffraction tomography microscopic imaging method based on non-interferometric synthetic aperture
WO2023221741A1
Cited By
Low-porosity anode carbon block porosity real-time estimation system and low-porosity anode carbon block finished product thereof
CN121113828A