Humidity correction method of water vapor infrared spectrum based on big data analysis

By acquiring the three-dimensional porosity and local humidity gradient data of porous materials, and combining with the big data analysis platform for feature extraction and modeling, the absorption difference caused by pore structure in traditional infrared spectroscopy is solved, and high-precision humidity correction and microcrack detection are achieved, which improves the accuracy and robustness of the detection.

CN120253750BActive Publication Date: 2025-08-26GUANGZHOU KINDOU ELECTRONICS TECH
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510761718.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-06-09
Publication Date
2025-08-26
Estimated Expiration
2045-06-09

AI Technical Summary

Technical Problem

In high humidity environment, in traditional infrared spectroscopy, the adsorption of water vapor by porous ceramic materials causes absorption peak drift. The existing technology cannot effectively compensate for local absorption differences caused by microscopic pore structures, and lacks a unified data processing and modeling process, making it difficult to meet the requirements of high precision, high reliability and high speed detection.

Method used

By obtaining the three-dimensional porosity distribution and local humidity gradient data of porous materials, combining with the big data analysis platform for distributed feature extraction and dimensionality reduction processing, building a distributed regression model, obtaining the spectral offset prediction value, constructing infrared spectral humidity correction coefficients, and performing microcrack classification.

Benefits of technology

It realizes high-precision humidity correction and microcrack detection, improves the robustness and detection accuracy of spectral characteristics, and meets the online defect diagnosis needs in high humidity environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120253750B_ABST
    Figure CN120253750B_ABST
Patent Text Reader

Abstract

The present invention provides a water vapor infrared spectrum humidity correction method based on big data analysis, which relates to the technical field of near-infrared spectrum humidity measurement. The present invention introduces the synchronous acquisition of the three-dimensional porosity distribution and surface humidity gradient of porous materials, combines distributed feature engineering with dimensionality reduction processing, and constructs a big data regression and classification model that jointly outputs "spectral humidity correction coefficient" and "microcrack classification result". It realizes the organic integration of six major technical elements: microscopic pore structure parameter acquisition, surface humidity gradient acquisition, big data feature extraction, distributed regression modeling, correction formula construction and microcrack detection, significantly improves the robustness of the corrected spectral characteristics and the accuracy of crack detection, and meets the technical requirements of online defect diagnosis of porous ceramics in high humidity environments.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of near-infrared spectrum humidity measurement, and in particular to a water vapor infrared spectrum humidity correction method based on big data analysis. Background Art

[0002] In a high humidity environment, the adsorption of water vapor in the pores of porous ceramic materials can cause a significant drift in the near-infrared absorption peak, resulting in an increase in the deviation of traditional infrared spectroscopy analysis results. In recent years, researchers have attempted to reduce the impact of water vapor on spectral signals through macroscopic correction of overall temperature and humidity sensor data, or humidity compensation algorithms based on physical models; at the same time, with the rise of big data and distributed computing platforms, pattern recognition and machine learning methods based on large-scale data samples have been applied to spectral preprocessing and feature extraction. However, these studies have mostly focused on single-dimensional humidity compensation or crack detection, and have not yet formed a linkage correction and online diagnosis process that can take into account both microscopic pore structure differences and macroscopic environmental humidity.

[0003] In the prior art, the announcement number is CN115128034B, and its name is a method for collaborative correction of near-infrared spectral data with temperature and humidity. Its core idea is: use a near-infrared spectrometer to collect spectral data under different temperature and humidity samples, and calculate the mean of the spectral data of different temperature and humidity samples, and then calculate the temperature and humidity spectrum correction coefficient based on the sample temperature and humidity values ​​and the corresponding spectral mean, and calculate the temperature and humidity spectrum correction formula based on the coefficient, and finally correct the spectral data of the sample to be corrected based on the temperature and humidity spectrum correction formula to complete the correction of the spectral data.

[0004] The existing technology still has many shortcomings:

[0005] 1. First, relying solely on overall temperature and humidity values ​​for correction cannot effectively compensate for local absorption differences caused by the material's microscopic pore structure, resulting in poor stability of the corrected spectral characteristics. Second, traditional humidity correction methods and crack detection methods belong to different modules and lack a unified data processing and modeling process, resulting in low system integration and insufficient real-time performance.

[0006] 2. Again, one-way algorithms are difficult to take into account the heterogeneous characteristics of large-scale samples and the needs of online computing, and are difficult to meet the comprehensive requirements of high precision, high reliability and high-speed detection in industrial production.

[0007] The above information disclosed in this Background section is only for enhancement of understanding of the background of the present disclosure and therefore it may contain information that does not form the prior art that is already known to a person of ordinary skill in the art. Summary of the Invention

[0008] The purpose of the present invention is to provide a water vapor infrared spectrum humidity correction method based on big data analysis to solve the problems raised in the above background technology.

[0009] To achieve the above object, the present invention provides the following technical solutions:

[0010] The humidity correction method of water vapor infrared spectrum based on big data analysis includes the following steps:

[0011] Step S1: Obtain three-dimensional porosity distribution data of P samples of the material to be tested;

[0012] Collect local humidity gradient data at multiple preset sampling points on the surface of the material sample to be tested;

[0013] Preprocess the three-dimensional porosity distribution data and local humidity gradient data to the same output range;

[0014] Step S2: Based on the preprocessing results, the porosity and humidity gradient data of P samples of the material to be tested are obtained and uploaded to the big data analysis platform for distributed feature extraction and dimensionality reduction to obtain the reduced-dimensionality features and spectral offset data;

[0015] Step S3: constructing a distributed regression model based on the reduced-dimensional features and the spectral offset data and obtaining the regression coefficient to obtain the predicted value of the spectral offset;

[0016] Step S4: constructing an infrared spectrum humidity correction coefficient based on the predicted value of the spectrum offset and applying it to the original infrared spectrum data to obtain corrected infrared spectrum data;

[0017] Step S5: extracting peak features based on the corrected infrared spectrum data and classifying microcracks of the material sample to be tested.

[0018] Compared with the existing technology, the beneficial effects of the present invention are: by introducing the synchronous collection of the three-dimensional porosity distribution and surface humidity gradient of porous materials, combining distributed feature engineering and dimensionality reduction processing, a big data regression and classification model is constructed to jointly output "spectral humidity correction coefficient" and "microcrack classification result", which realizes the organic integration of six major technical elements: microscopic pore structure parameter acquisition, surface humidity gradient collection, big data feature extraction, distributed regression modeling, correction formula construction and microcrack detection, significantly improving the robustness of the corrected spectral characteristics and the accuracy of crack detection, and meeting the technical requirements of online defect diagnosis of porous ceramics in high humidity environments. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] Figure 1 Schematic diagram of the overall method flow of the present invention;

[0020] Figure 2 Schematic diagram of the application of porous ceramic materials in the present invention. DETAILED DESCRIPTION

[0021] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, the specific embodiments of the present invention are described in detail below with reference to the accompanying drawings.

[0022] In the following description, many specific details are set forth to facilitate a full understanding of the present invention. However, the present invention may also be implemented in other ways different from those described herein. Those skilled in the art may make similar generalizations without violating the connotation of the present invention. Therefore, the present invention is not limited to the specific embodiments disclosed below.

[0023] In porous ceramic materials, moisture adsorption in internal pores causes near-infrared spectral absorption peaks to drift in high humidity environments. Relying solely on overall temperature and humidity numerical corrections cannot compensate for local absorption differences caused by pore microstructures. This solution leverages six defining technical features: "micropore structure parameter acquisition," "surface humidity gradient acquisition," "big data feature engineering," "distributed regression modeling," "correction formula construction," and "microcrack detection" to achieve the linkage between high-precision humidity correction and automatic material microcrack detection. Specifically, this solution includes the following embodiments:

[0024] Example 1:

[0025] See also Figure 1 and Figure 2 , the present invention provides a technical solution:

[0026] The humidity correction method for water vapor infrared spectroscopy based on big data analysis is applied to porous ceramic materials in high humidity environments. The specific steps include:

[0027] Step S1: Obtain three-dimensional porosity distribution data of P samples of the material to be tested;

[0028] Collect local humidity gradient data at multiple preset sampling points on the surface of the material sample to be tested;

[0029] Preprocess the three-dimensional porosity distribution data and local humidity gradient data to the same output range;

[0030] Further explanation: P samples of the material to be tested are numbered as {1, 2, ..., i, ..., P}; where i represents the index of the sample of the material to be tested;

[0031] The material sample to be tested is placed in an X-ray micro-computed tomography scanner (micro-CT, hereinafter referred to as "micro-CT scanning") to obtain three-dimensional pore structure data of the material sample to be tested;

[0032] Specifically, a Bruker SkyScan 1275 microCT scanner was used, and the specific parameters were set as follows:

[0033] Tube voltage 80 kV, current 100 µA; voxel resolution 10 µm;

[0034] The scanning parameters for each material sample to be tested are 360° rotation, 0.2° step, and 15 min for a single sample;

[0035] Based on the acquired 3D pore structure data, the Feldkamp–Davis–Kress (FDK) algorithm implemented in the NRecon software is used to perform 3D reconstruction of the material sample to generate a voxel matrix containing grayscale values. ;

[0036] It should be noted that the implementation steps of the Feldkamp–Davis–Kress (FDK) algorithm in this embodiment are as follows:

[0037] 1.1) Data acquisition: By rotating the cone-beam X-ray source and detector, two-dimensional projection images at different angles are acquired;

[0038] 1.2) Filtering: Use a Ram-Lak filter or other high-pass filter to perform frequency-domain filtering on each projection image to compensate for the image blur caused by cone-beam geometry.

[0039] 1.3) Projection data acquisition: Place the material sample i in the microCT scanner and obtain two-dimensional projection images at different angles every 0.2°. .

[0040] 1.4) Filtering: Apply the Ram-Lak filter h6 to each projection image to obtain the filtered projection ;

[0041] 1.5) Weighted back projection: According to the FDK algorithm, the filtered projection image is weighted Perform back projection and accumulate to generate a three-dimensional volume image .

[0042] 1.6) Porosity calculation: By setting the grayscale threshold T, the voxel matrix Binarize into solid phase and pores, and calculate the average porosity of the material sample to be tested and porosity standard deviation . Specifically including:

[0043] Set the grayscale threshold T for image segmentation and get ;

[0044] when ; Indicates that the material sample i to be tested is a solid phase;

[0045] when ; Indicates that the material sample i to be tested has pores;

[0046] The solid phase refers to the continuous, solid portion of a material sample that makes up its bulk structure. In porosity distribution data, the solid phase represents the actual physical portion of the material, without any pores or voids.

[0047] Porosity refers to the presence of holes, cracks, or other forms of cavities within a material sample. These areas do not contain any actual substance and are the "empty" parts of the material.

[0048] Set the total number of voxels to ; Calculate the average porosity and porosity standard deviation ;

[0049]

[0050]

[0051] in, represents the pore identification of the vth voxel on the material sample i to be tested, and ;

[0052] Further explanation: Local humidity gradient data at multiple preset sampling points is obtained by arranging an array of micro humidity sensors on the surface of the material sample to be tested; specifically, including:

[0053] Utilizing the characteristic absorption peak of water vapor in the mid-infrared band, the spectrometer measures the absorbance of multiple preset sampling points on the surface of the material sample i in real time to obtain the humidity distribution;

[0054] The spectral measurement instrument used in this embodiment is Thermo Nicolet iS50 (FTIR); the wavelength range is 2.5µm–25µm; the resolution is 4cm -1 ;

[0055] The distance between the sample and the material to be tested is set as follows: the laser beam is perpendicular to the surface, and the optical path length is l = 0.5 cm;

[0056] The Beer–Lambert law is used to characterize the absorbance–humidity relationship: ;in, is the absorbance of the preset sampling point j on the material sample i to be tested; The water vapor absorption peak Molar absorptivity; The water vapor concentration in the air at the preset sampling point j on the material sample i to be tested; is the optical path length;

[0057] The preset sampling point j on the material sample i to be tested Convert to relative humidity ;

[0058]

[0059] in is the saturated water vapor concentration at 25°C; 0.023 mol·L -1 ; 0.12 L·cm -1 ; 0.5cm;

[0060] In this embodiment, multiple preset sampling points Repeat the measurement of a single point K3=3 times and take the average value;

[0061] Calculate the average relative humidity of the material sample i to be tested Standard deviation of humidity gradient ;

[0062]

[0063]

[0064] Among them, Increase, indicating Increase, and then Increase; M is the total number of preset sampling points, K3 is the total number of repeated measurements, and k3 is the index of the number of repeated measurements;

[0065] When the fluctuation amplitude of increases, Increased, indicating uneven moisture distribution.

[0066] Will 、 , average porosity and porosity standard deviation The association is stored in the database.

[0067] Further explanation: The preprocessing of the same output range includes: , , and Use Min–Max normalization to get the normalized output result , , and .

[0068]

[0069] Output , ensuring the stability of subsequent model input:

[0070] when The closer it is to 0, the lower the porosity; The closer it is to 1, the higher the porosity;

[0071] when The closer it is to 0, the more uniform the humidity is; The closer it is to 1, the more uneven the humidity distribution is;

[0072] It should be noted that , and The normalization method is the same as above and will not be described in detail.

[0073] Step S2: Based on the preprocessing results, the porosity and humidity gradient data of P samples of the material to be tested are obtained and uploaded to the big data analysis platform for distributed feature extraction and dimensionality reduction to obtain the reduced-dimensionality features and spectral offset data;

[0074] Further explanation: In this embodiment, the material sample P to be tested is 500;

[0075] Obtain the porosity and humidity gradient characteristics of P=500 material samples after normalization, which is , , and ;

[0076] P samples of the material to be tested , , and , and the corresponding infrared spectrum main peak offset, are summarized and uploaded to the big data analysis platform for distributed feature cleaning, standardization, principal component analysis (PCA) dimensionality reduction, and the reduced dimensionality features and offsets are stored in the feature library at the same time; specifically including:

[0077] 2.1) The four-dimensional normalized feature vector of each material sample i to be tested is and the corresponding infrared spectrum offset The data is integrated into a row of records, generating a CSV file containing P rows and 5 columns (4 features + 1 offset), and uploaded to the HDFS path of the big data platform via the secure file transfer protocol (SFTP);

[0078] 2.2) Configure a distributed storage and computing environment on a 20-node Hadoop / Spark cluster. Specifically:

[0079] HDFS storage: import CSV into HDFS directory;

[0080] Spark environment: Spark version 3.2.1, Java 8, configured with 4 Executor cores and 16 GB of memory per node; a 20-node configuration can accelerate data cleaning and matrix operations in parallel when processing 500 × 5-dimensional data, balancing resource utilization and response time.

[0081] 2.3) Based on the distributed environment deployment, use PySpark to complete missing value filling and standardization processing, specifically:

[0082] For each feature column Perform mean filling: replace missing entries with the mean of the entire sample . r=positive integers from 1 to 4, representing , , and ;

[0083] Define the normalization formula as ;in, is the rth original feature of the material sample i to be tested; and are the mean and standard deviation of the rth original feature in P=500 samples; is the standardized feature;

[0084] when When increasing, if ,but It is positive and its absolute value increases as the difference increases;

[0085] On the contrary, when hour, It is negative and its absolute value increases as the difference increases;

[0086] The Z-score standardization in this embodiment can eliminate the dimensional differences between different features, so that when the subsequent PCA selects the principal components according to the variance, the weights of each feature are fair;

[0087] 2.4) Perform principal component analysis (PCA) dimensionality reduction, including covariance matrix calculation, eigenvalue decomposition, principal component selection, and dimensionality reduction mapping;

[0088] According to the standardization process, by calculating the covariance matrix and solving the eigenvalue problem, the dimension is reduced to the principal component space with a cumulative variance contribution rate of ≥ 95%:

[0089] The covariance matrix is ​​calculated as follows:

[0090]

[0091] in ; T1 represents the matrix transposition operation, specifically transposing the column vector into a row vector; P represents the total number of material samples to be tested, and the eigenvalue decomposition is as follows:

[0092] Solution , and press Sorting;

[0093] The main components are selected as follows:

[0094] Calculate the cumulative variance contribution rate , take K=2;

[0095] Get the first two principal component vectors .

[0096] The dimensionality reduction mapping is as follows:

[0097]

[0098] in, is the score of the material sample i on the rth principal component;

[0099] If the input features are in the principal component vector The higher the loading coefficient in , the more positive the change of the feature will be. ;

[0100] The principal component score can reflect the comprehensive change trend of the input features.

[0101] According to PCA dimensionality reduction, the principal component scores are normalized by Min–Max and stored in the feature library together with the offset so that the model can read it quickly: The normalization formula is:

[0102]

[0103] in, 、 are the minimum and maximum scores of all tested material samples on the rth principal component; ; yes The normalized output value of ;

[0104] when The closer it is to 0, the weaker the comprehensive characteristics of the material sample i in the direction of the rth principal component;

[0105] when The closer it is to 1, the stronger the comprehensive characteristics of the material sample i in the direction of the rth principal component;

[0106] Normalizing the scores helps in numerical stability in subsequent multi-target regression models.

[0107] The normalized infrared spectrum offset is set to ;

[0108] The normalized dimension-reduced features and spectral offset data are represented as .

[0109] Normalized infrared spectrum offset The increase or decrease directly reflects the influence of humidity on the position of the main peak of infrared spectrum;

[0110] Normalized infrared spectrum offset Increases indicate that the influence of humidity on the main peak of the spectrum is increasing, and a stronger humidity correction is required;

[0111] Normalized infrared spectrum offset It decreases, indicating that the influence of humidity on the main peak of the spectrum is weakened and the need for humidity correction is low.

[0112] This embodiment forms a complete and implementable technical process from original normalized features to low-dimensional stable features, providing a solid data foundation for the efficient training of the subsequent spectral humidity correction model.

[0113] Step S3: constructing a distributed regression model based on the reduced-dimensional features and the spectral offset data and obtaining the regression coefficient to obtain the predicted value of the spectral offset;

[0114] Further explanation: Based on the dimensionality-reduced features and spectral offset data, the normalized principal component scores and spectral offsets are used as input and output variables. A multivariate linear regression model is constructed in a distributed computing environment, and the regression coefficients are extracted for online spectral offset prediction. Specifically, the following steps are performed:

[0115] 3.1) Load the normalized principal component scores and spectral offsets stored in the distributed file system as analysis data, and split the dataset into an 80% training / 20% test ratio.

[0116] Read from the HDFS path with the columns "z1", "z2", and "y" representing the normalized principal component scores. , and ;

[0117] After randomly shuffling all the data, the training set and test set are allocated in proportion;

[0118] 80% of the data was selected as model training data to ensure the statistical reliability of model parameter estimation; the remaining 20% ​​was used for model performance evaluation.

[0119] 3.2) The parameters of the distributed regression model are set as follows:

[0120] Based on the training set, a multivariate linear regression model with L2 regularization is selected, and the regularization coefficient and iteration termination condition are determined;

[0121] The multiple linear regression model uses Ridge regression to suppress the possible collinearity between different principal components;

[0122] Regularization coefficient : Select in cross-validation so that the model strikes a balance between bias and variance;

[0123] Maximum number of iterations And the convergence threshold : Ensure that the parameter solution fully converges while taking into account computational efficiency.

[0124] 3.3) Based on the model configuration, perform least squares fitting on the training set and extract the intercept and regression coefficients of each principal component to construct a prediction function.

[0125] Using data from the training set , Fit the Ridge regression model and solve the regression coefficient 、 and .

[0126] Intercept term represents the reference value of the predicted value of the spectral shift;

[0127] Regression coefficient and :Corresponding to dimensionality reduction features and The weights of ,reflect their influence on the predicted value of spectral offset;

[0128] The output value of the distributed regression model of the material sample i to be tested is recorded as the predicted value ;

[0129] like , then the corresponding dimensionality reduction feature When it increases, the predicted value It also increases;

[0130] like , then the corresponding dimensionality reduction feature When it increases, the predicted value reduce;

[0131] Based on the test set and the extracted regression coefficients, the predicted value of each test sample is calculated and restored to the actual offset through linear mapping.

[0132] Predicted value The calculation formula is:

[0133]

[0134] The predicted value of the spectrum shift of the material sample i to be tested is recorded as ; and perform the following anti-normalization mapping:

[0135]

[0136] in 、 are the maximum and minimum values ​​of infrared spectrum offset before normalization;

[0137] when The closer it is to 0, The closer it is to the minimum infrared spectrum offset, the weaker the humidity effect of the material sample to be tested;

[0138] when The closer it is to 1, The closer it is to the maximum infrared spectrum shift, the stronger the humidity influence of the material sample to be tested is.

[0139] Furthermore, based on the prediction results and the true values ​​of the test set, the coefficient of determination and root mean square error are calculated to evaluate the model performance, and the model parameters are persisted to the model library to support online prediction.

[0140] The evaluation indicators are set as:

[0141]

[0142]

[0143] in, is the actual offset of the material sample i in the test set; is the predicted value of the spectral shift of the material sample i in the test set; is the average offset of all material samples to be tested in the test set; is the number of samples in the test set, and Less than the total number P of material samples to be tested;

[0144] Coefficient of determination Measures the model's ability to explain the actual offset. The closer the value is to 1, the better the model fit is.

[0145] The root mean square error (RMSE) reflects the average deviation between the model prediction value and the actual value. The closer the value is to 0, the higher the model prediction accuracy.

[0146] Will and 、 Store it in the model registry and load it in the distributed service for real-time prediction.

[0147] Step S4: constructing an infrared spectrum humidity correction coefficient based on the predicted value of the spectrum offset and applying it to the original infrared spectrum data to obtain corrected infrared spectrum data;

[0148] Further explanation: In industrial online quality inspection, raw infrared spectral data often exhibit peak shifts due to the influence of ambient water vapor. By quickly generating correction coefficients based on the predicted spectral shift and performing real-time corrections, detection accuracy and low latency can be maintained in high-throughput detection scenarios.

[0149] Get the predicted value of the material sample i from the distributed regression model output ;

[0150] For each material sample i to be tested, read the calculated and stored spectral shift prediction value ;

[0151] This embodiment ; Direct use It can ensure the numerical stability and consistency of the correction coefficient calculation;

[0152] Predicted value based on the spectral shift read , using linear mapping as a correction factor to ensure ;

[0153] in is the spectral correction coefficient of the material sample i to be tested, and its range is limited to the interval (0,1);

[0154] when The closer it is to 0, The closer it is to 1, the smaller the correction needs;

[0155] when The closer it is to 1, The closer it is to 0, the greater the correction needs;

[0156] Based on spectral correction factor , multiply the original infrared spectrum data of the material sample i by To generate corrected infrared spectrum data;

[0157] The corrected infrared spectrum data includes the corrected wavelength value of the material sample i at the jth preset sampling point ;

[0158] The specific correction steps include:

[0159] The original infrared spectrum data is expressed as: the original spectrum wavelength sequence of the material sample i to be tested , j is the preset sampling point index;

[0160]

[0161] in, is the original wavelength value of the material sample i at the jth preset sampling point; is the corrected wavelength value of the material sample i at the jth preset sampling point.

[0162] When increasing, each Approach ;

[0163] When decreasing, each Decrease, reflecting a strong correction of the offset.

[0164] It is a point-by-point multiplication with simple and efficient operation, suitable for implementation on streaming computing or device-side FPGA / ASIC, meeting real-time requirements.

[0165] Further explanation: The corrected infrared spectrum data is used as input to the defect detection module for subsequent analysis.

[0166] Will The data is pushed to the subsequent defect detection or feature extraction pipeline in the same data format as the original infrared spectroscopy data.

[0167] Further explanation: The following examples are based on the content described in step S4 and demonstrate the innovation and advantages of the present invention in infrared spectrum humidity correction through specific data;

[0168] To verify the effectiveness of the proposed method, six porous material samples were selected for indoor environment simulation testing. The relative humidity of the test environment was maintained at 60% ± 2%. The spectral measurement instrument used was a Fourier transform infrared spectrometer (FTIR) with a resolution of 0.5 nm and a measurement band of 1000 cm -1 The main absorption peak position shifts. Each sample is implemented according to the following process:

[0169] First, the six samples were measured under standard drying conditions to obtain the reference absorption peak position, which was recorded as the reference peak position. Then, each sample was measured three times under 60% humidity conditions, and the peak position shift was recorded. The average value was taken as the original peak position shift. Substitute into the linear mapping formula of the present invention Calculate the correction factor. Then multiply the peak position offset of each material sample by the corresponding , and obtain the corrected peak offset. Finally, compare the offset errors before and after correction to evaluate the correction effect.

[0170] During the test, all data was recorded in a spreadsheet to facilitate subsequent statistical analysis. The entire process did not require any interface modifications to the downstream defect detection module, and the corrected wavelength data was directly input to verify the universality and real-time performance of the method on different samples. The data obtained in this example are shown in Table 1:

[0171] Table 1 Validity study of correction coefficients:

[0172]

[0173] As can be seen from Table 1, the average reduction rate of the peak position shift error of the six samples by the method of the present invention reaches about 62.5%, which is significantly better than the traditional non-correction or fixed correction coefficient method. The data in the examples fully demonstrate that:

[0174] The linear mapping correction coefficient dynamically calculated using the normalized predicted offset value can adjust the correction intensity according to the degree of humidity influence;

[0175] This method not only ensures that the offset error after correction is greatly reduced, but also has good universality for different materials;

[0176] When the predicted value approaches 0 (sample 4), the correction coefficient approaches 1, and the error after correction is only 0.18nm, maintaining the authenticity of the spectrum; when the predicted value approaches 1 (sample 6), It approaches 0, and the error after correction is reduced to 0.18nm, achieving the maximum correction.

[0177] Step S5: extracting peak features based on the corrected infrared spectrum data and classifying microcracks of the material sample to be tested.

[0178] Further explanation: The defect detection module includes micro-crack classification of the material sample to be tested, specifically including:

[0179] Based on the corrected wavelength value , extract the peak width of each material sample i in the peak shape Sum peak intensity ; The specific steps are as follows:

[0180] 5.1) Corrected wavelength value of the material sample i to be tested Gaussian fitting is performed near the target absorption peak. In this embodiment, the target absorption peak is within ±5 nm of the absorption peak. The fitting function is:

[0181]

[0182] in, is the corrected wavelength value; is the peak center wavelength after fitting; is the Gaussian standard deviation; A is the fitted amplitude;

[0183] Peak width The calculation formula is as follows:

[0184]

[0185] Peak intensity Take the fitting amplitude ; When the crack causes the peak to broaden, that is When the peak width increases When the cracks cause the absorption to weaken, the peak intensity Reduce; input the peak width and peak intensity into the support vector machine classifier for crack detection;

[0186] A pre-trained radial basis function (RBF) kernel support vector machine is used to classify the microcracks of the material sample i to obtain a classification decision function for determining whether the material sample i has microcracks. ; Classification decision function Used to output the classification label of the material sample i to be tested; the specific logic includes setting the classification function of microcrack classification as:

[0187]

[0188] in, is the characteristic vector of the material sample i to be tested; They are support vector weights and true value labels in the training set. and the corresponding eigenvectors; is a classification decision function used to determine whether the material sample i to be tested has microcracks; L is the total number of support vectors, and k2 is the index of the support vector;

[0189] is the RBF kernel width hyperparameter, and in this embodiment, the value ; b1 is the classifier bias term; it controls the decay speed of the kernel function. The larger the value, the more concentrated the kernel function.

[0190] is the true value label, +1 means “there is a crack”, –1 means “there is no crack”;

[0191] when When it decreases, the kernel function value increases and the support vector influence increases;

[0192] when When it increases, the response to data far away from the support vector weakens, and the model pays more attention to the local area;

[0193] Mapping classification labels to sensor array coordinates and performing 3D visualization; specifically:

[0194] Map the sensor unit coordinates of the material sample i that is judged to have cracks, and generate an annotation report based on the 3D pore structure data. The specific logic includes:

[0195] like , then the material sample i to be tested corresponds to the sensor unit Micro cracks occur; is the spatial position of the material sample i in the array;

[0196] The sensor array coordinates of this embodiment It is obtained during device calibration and stored uniformly. The following visualization process is executed:

[0197] Load 3D pore structure data; highlight crack coordinates in the model, color depth and The absolute value of the classification decision function is proportional to ;

[0198] Export annotated 3D views and generate reports in PDF or HTML format.

[0199] The absolute value output by the classification decision function reflects the crack confidence.

[0200] When the absolute value of the classification decision function increases, the crack confidence is higher and the color of the annotation in the report is darker;

[0201] The characteristic value range is described as ;

[0202] When the peak width or peak intensity When the crack-free reference value is greatly deviated, decreases, the kernel function value increases, the absolute value of the decision function value increases, and finally , determine the crack;

[0203] When the peak shape characteristics approach the normal reference range, near the decision boundary, the value range approaches zero, and it is marked as "pending review" on the production line.

[0204] The following examples, based on step S5, verify the innovation and beneficial effects of the present invention in extracting peak shape features of corrected infrared spectra and classifying microcracks through specific experimental processes and data.

[0205] Six porous ceramic samples were selected, numbered "Sample 1" to "Sample 6." Samples 1, 3, and 5 had artificial microcracks with a width of approximately 0.05 mm pre-prepared on their surfaces, while the remaining three samples were crack-free. All samples measured 20 mm × 20 mm × 5 mm. The experiment began by placing all samples in a standard drying chamber (relative humidity ≤ 5%) for 20 minutes. The samples were then analyzed using a Fourier transform infrared spectrometer (FTIR) at 1000 cm-1 with a resolution of 0.2 nm. -1 Get the reference absorption peak center wavelength The baseline peak width and peak intensity were recorded. The samples were then transferred to a simulated production line environment chamber (relative humidity 60% ± 2%, temperature 25 °C) and three consecutive raw infrared spectral data were collected for each sample at 1 min intervals. , the preset number of sampling points M = 100. Execute the spectrum offset prediction value and correction coefficient calculation in step S4 for each collected data to generate the corrected wavelength value .

[0206] In obtaining Then, the peak shape feature extraction process of step S5 is entered: the data within the range of ±5 nm of the target absorption peak are obtained, and the improved fast Gaussian fitting algorithm (based on Levenberg–Marquardt iteration, maximum iteration 50 times, residual threshold 1×10 -6 ), fitting function , output the fitted peak center wavelength , standard deviation And the fitting amplitude A. According to the formula The peak width w (unit: nm) is calculated, and the peak intensity I is directly taken as A. To improve real-time performance, the algorithm is implemented on the FPGA acceleration module, and each fitting takes 8 ms. After the fitting is completed, the feature vector x = [w, I] is normalized to the interval (0, 1) and sent to the radial basis kernel SVM classifier trained on 100 sets of known crack samples in advance. The kernel width hyperparameter , the number of support vectors L=30;

[0207] in The classifier is embedded in the embedded CPU, and the CPU and FPGA are used for parallel calculation, with a total time consumption of about 15 ms. Determine "there is a crack", Determine "no crack"; at the same time output the average decision function value , which is used for subsequent visualization confidence annotation. The average decision value of three measurements of all samples is taken as the final judgment basis to reduce accidental errors.

[0208] After the experiment, the average peak width, average peak intensity, and average decision function value of each sample were calculated. Indicators such as the SVM judgment label, actual crack state, and judgment accuracy are summarized and the average judgment time is calculated and shared with the data processing server for comparative analysis. This example quantifies the advantages of the method in classification accuracy, judgment stability, and real-time performance by comparing it with a control group without humidity correction.

[0209] Table 2 Corrected infrared spectrum peak shape feature extraction and microcrack classification research:

[0210]

[0211] From the above data we can see that:

[0212] 1. Classification accuracy: the average decision function value of all crack samples (samples 1, 3, and 5) are all greater than 0, and the judgment label is +1, which accurately reflects the actual existence of cracks.

[0213] Average decision function value of crack-free samples (samples 2, 4, and 6) are all less than 0, and the label is judged as −1, which accurately reflects that there is no crack.

[0214] The classification accuracy reached 100%, which is significantly higher than the level of 80%–85% when no humidity correction was made, demonstrating a significant improvement in the classification accuracy of the method of the present invention.

[0215] 2. Feature stability: The peak width and peak intensity extracted from the corrected infrared spectral data are highly correlated with the actual crack state, which reduces the interference of humidity changes on the spectral characteristics and enhances the robustness of feature extraction.

[0216] 3. Real-time performance: The average judgment time is only 15 ms, which meets the real-time requirements of industrial online high-speed quality inspection and is suitable for real-time micro-crack detection in large-scale production environments.

[0217] Average decision function value Indicates that the feature vector is on the "cracked" side of the classification boundary, and vice versa.

[0218] The absolute value of the decision function can be used to quantify the confidence of the classification. The larger the value, the higher the confidence. The closer it is to 0, the higher the uncertainty and the more manual review is needed.

[0219] The advantage of this solution is linkage control: the same set of data processing and modeling processes can output the humidity correction coefficient and directly drive automatic crack detection, achieving seamless linkage between the two major technical effects.

[0220] It should be noted that: All calculation formulas in this application document use regression analysis including but not limited to machine learning algorithms to deeply analyze the relevant parameters collected and identify their natural trends and relationships. Use professional software, such as Python's Scikit-learn library or R language, to automatically generate mathematical models that match the data. Then, objectively evaluate the performance of the model through methods such as cross-validation, and combine continuous feedback and optimization to ensure that the created formula truly reflects the inherent laws of the data, thereby ensuring its effectiveness and accuracy. In all calculation formulas in this application, the parameters in each formula are dimensionally non-dimensionalized within a consistent range to ensure that different physical quantities are compared on the same scale; dimensionless technical means include but are not limited to Min-Max-Normalization and Z-Score standardization;

[0221] The technical solution of the present invention, or the part that contributes to the prior art, can be embodied in the form of a software product. The computer software product can be stored in a computer-readable storage medium, such as a computer floppy disk, read-only memory (ROM), random access memory (RAM), flash memory (FLASH), hard disk or optical disk, etc., and includes a number of instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute the methods of various embodiments of the present invention.

[0222] The logic and / or steps represented in the flowcharts or otherwise described herein, for example, can be considered as a sequenced list of executable instructions for implementing the logical functions, and can be embodied in any computer-readable medium for use by, or in conjunction with, an instruction execution system, apparatus, or device (e.g., a computer-based system, a system including a processor, or other system that can fetch and execute instructions from an instruction execution system, apparatus, or device). For purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transport a program for use by, or in conjunction with, an instruction execution system, apparatus, or device.

[0223] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit the present invention. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solutions of the present invention may be modified or replaced by equivalents without departing from the spirit and scope of the technical solutions of the present invention, which should all be included in the scope of the claims of the present invention.

Claims

1. A moisture correction method for infrared spectroscopy of water vapor based on big data analysis, applied to porous ceramic material samples in high humidity environments, characterized by: The specific steps include: Step S1: Obtain three-dimensional porosity distribution data of P samples of the material to be tested; Numbering P material samples to be tested as {1, 2, ..., i, ..., P}; where i represents the index of the material sample to be tested; obtaining three-dimensional pore structure data of the material sample to be tested; Based on the three-dimensional pore structure data obtained from the material sample to be tested, the material sample to be tested is reconstructed in three dimensions to generate a voxel matrix containing grayscale values ; The voxel matrix Binarize into solid phase and pores, and calculate the average porosity of the material sample to be tested and porosity standard deviation ; Collect local humidity gradient data at multiple preset sampling points on the surface of the material sample to be tested; The local humidity gradient data of multiple preset sampling points is obtained by arranging a micro humidity sensor array on the surface of the material sample to be tested; specifically, it includes: Utilizing the characteristic absorption peak of water vapor in the mid-infrared band, the spectrometer measures the absorbance of multiple preset sampling points on the surface of the material sample i in real time to obtain the humidity distribution; The Beer–Lambert law is used to characterize the absorbance–humidity relationship: ;in, is the absorbance of the preset sampling point j on the material sample i to be tested; The water vapor absorption peak Molar absorptivity; The water vapor concentration in the air at the preset sampling point j on the material sample i to be tested; is the optical path length; The preset sampling point j on the material sample i to be tested Convert to relative humidity ; Calculate the average relative humidity of the material sample i to be tested Standard deviation of humidity gradient ; Preprocess the three-dimensional porosity distribution data and local humidity gradient data to the same output range; Preprocessing of the same output range includes: , , and Use Min–Max normalization to get the normalized output result , , and ; Step S2: Based on the preprocessing results, the porosity and humidity gradient data of P samples of the material to be tested are obtained and uploaded to the big data analysis platform for distributed feature extraction and dimensionality reduction to obtain the reduced-dimensionality features and spectral offset data; Step S3: Based on the dimensionality-reduced features and the spectral offset data, a distributed regression model is constructed and the regression coefficient is calculated to obtain the predicted value of the spectral offset; Step S4: constructing an infrared spectrum humidity correction coefficient based on the predicted value of the spectrum offset and applying it to the original infrared spectrum data to obtain corrected infrared spectrum data; Step S5: extracting peak features based on the corrected infrared spectrum data and classifying microcracks of the material sample to be tested.

2. The method for correcting humidity in water vapor infrared spectrum based on big data analysis according to claim 1, characterized in that: P samples of the material to be tested , , and , and the corresponding infrared spectrum main peak offset, are summarized and uploaded to the big data analysis platform for distributed feature cleaning, standardization, and principal component analysis dimensionality reduction. At the same time, the reduced dimensionality features and spectral offset data are stored in the feature library together; The normalized dimension-reduced features and spectral offset data are represented as .

3. The method for correcting humidity in water vapor infrared spectrum based on big data analysis according to claim 2, characterized in that: Based on the dimensionality-reduced features and spectral offset data, the normalized principal component scores and spectral offsets were used as input and output variables to construct a multivariate linear regression model in a distributed computing environment, and the regression coefficients were extracted for online spectral offset prediction. The output value of the distributed regression model of the material sample i to be tested is recorded as the predicted value ; Based on the predicted value , the predicted value of the spectrum shift of the material sample i to be tested is recorded as ; when The closer it is to 0, The closer it is to the minimum infrared spectrum offset, the weaker the humidity influence of the material sample i is. when The closer it is to 1, The closer it is to the maximum infrared spectrum offset, the stronger the humidity influence of the material sample i is.

4. The method for correcting humidity in water vapor infrared spectrum based on big data analysis according to claim 3, characterized in that: Get the predicted value of the material sample i from the distributed regression model output ; For each material sample i to be tested, read the calculated and stored spectral shift prediction value ; Predicted value based on the spectral shift read , using linear mapping as a correction factor to ensure ; in is the spectral correction coefficient of the material sample i to be tested, and its range is limited to the interval (0,1); when The closer it is to 0, The closer it is to 1, the smaller the correction needs; when The closer it is to 1, The closer it is to 0, the greater the correction needs; Based on spectral correction factor , multiply the original infrared spectrum data of the material sample i by To generate corrected infrared spectrum data; The corrected infrared spectrum data includes the corrected wavelength value of the material sample i at the jth preset sampling point .

5. The method for correcting humidity in water vapor infrared spectrum based on big data analysis according to claim 4, characterized in that: The corrected infrared spectrum data is used as input to the defect detection module for subsequent analysis.

6. The method for correcting humidity in water vapor infrared spectrum based on big data analysis according to claim 5, characterized in that: The defect detection module includes micro-crack classification of the material sample to be tested, specifically including: Based on the corrected wavelength value , extract the peak width of each material sample i in the peak shape Sum peak intensity ; The pre-trained radial basis function kernel support vector machine is used to classify the microcracks of the material sample i to obtain the classification decision function for determining whether the material sample i has microcracks. ; Classification decision function Used to output the classification label of the material sample i to be tested; Mapping classification labels to sensor array coordinates and performing 3D visualization; The sensor unit coordinates of the material sample i that is judged to have "cracks" are mapped, and an annotation report is generated in combination with the 3D pore structure data.

Citation Information

Patent Citations

  • A method for coordinating near-infrared spectral data with temperature and humidity

    CN115128034B

  • Method and system for analyzing substance concentration based on near infrared spectrum technology

    CN114611582A

  • Near infrared spectrum rapid detection system and real-time analysis method

    CN119510347A