Method for rapidly analyzing wax content in crude oil based on infrared spectrum fusion
Through the fusion technology of Fourier transform infrared FTIR and near-infrared NIR spectroscopy, combined with the XGBoost correction model, the complex operation, high cost and environmental pollution of wax content detection in crude oil in the prior art is solved, and a fast, accurate and lossless analysis effect is achieved.
Patent Information
- Application Number
- CN202510433217.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-08
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2045-04-08
AI Technical Summary
The existing wax content detection technology in crude oil has complex operation, high cost, high equipment requirements and potential pollution to the environment, and lacks targetedness and rapidity.
The fusion technology based on Fourier transform infrared FTIR and near infrared NIR spectroscopy is adopted to collect data from crude oil samples through Fourier transform infrared FTIR spectrometers, and pre-process and feature variable screening is used using the XGBoost correction model. Finally, the model hyperparameters are optimized by leaving a cross-validation and grid search to achieve rapid analysis of wax content in crude oil.
It improves the accuracy and reliability of the wax content analysis results in crude oil, realizes fast, non-destructive and accurate quantitative analysis, reduces operational complexity and cost, and reduces environmental pollution.
Smart Images

Figure CN119935950A_ABST
Abstract
Description
Technical Field
[0001] The invention belongs to the technical field of spectral analysis, and in particular relates to a rapid analysis method for wax content in crude oil based on infrared spectrum fusion. Background Art
[0002] In fact, more than half of crude oil production belongs to the category of waxy crude oil. The high wax content of waxy crude oil makes it easy for wax crystallization and precipitation to occur during transportation, which not only blocks pipelines and equipment, but also poses major technical difficulties in various links such as crude oil detection, mining, transportation, storage and refining. Therefore, controlling and managing petroleum wax content has become a key technical challenge facing the oil industry.
[0003] Existing technologies for detecting wax content in crude oil include differential scanning calorimetry (DSC), gas chromatography (GC) and thermal desorption, each of which has its own advantages and disadvantages: Differential scanning calorimetry (DSC) is based on the linear correlation between the thermal enthalpy of wax precipitation and the wax content. The wax content of the sample is inferred through the thermal spectrum. This method can provide detailed information about the thermal properties of the sample, but the equipment cost is high and the operation is complicated; Gas chromatography (GC) uses a gas chromatograph to analyze the sample to determine the wax content in the sample. This method can provide detailed information about the chemical composition of the sample, but the equipment cost is high and professional operating skills are required; Thermal desorption calculates the wax content by dissolving the crude oil sample in petroleum ether and separating the wax. This method has high sample processing technology requirements and is prone to human errors.
[0004] The invention patent with the publication (announcement) number CN107966420B discloses a method for predicting crude oil properties by near-infrared spectroscopy, which identifies a group of virtual library spectra that are closest to the spectrum of the crude oil to be tested, and then calculates the properties of the crude oil to be tested by weighted similarity calculation based on the similarity between the spectrum of the crude oil to be tested and each spectrum in the group of virtual library spectra. This method requires the establishment of a near-infrared spectral database of crude oil samples, and detects all conventional physical properties of crude oil samples, including density, acid value, residual carbon, sulfur content, nitrogen content, wax content, colloid content, asphaltene content, and actual boiling point distillation data, but does not detect and analyze the wax content in crude oil.
[0005] The above methods generally have problems such as complex operation, high cost, and lack of specificity. In addition, the use of large amounts of reagents may have adverse effects on the environment. Therefore, it is particularly important to find a simpler, faster and more environmentally friendly technology for detecting the wax content in crude oil. Summary of the invention
[0006] The purpose of the present invention is to provide a rapid analysis method for the wax content in crude oil based on infrared spectroscopy fusion. Based on the complementarity of Fourier transform infrared FTIR and near infrared NIR spectroscopy, the spectral data obtained by the two technologies are fused to obtain more comprehensive spectral information, thereby improving the accuracy and reliability of the analysis results of the wax content in crude oil.
[0007] The present invention provides a method for rapid analysis of wax content in crude oil based on infrared spectrum fusion, comprising the following steps: Step 1: Use Fourier transform infrared FTIR spectroscopy and near infrared NIR spectroscopy instruments to collect MIR and NIR spectral data of several crude oil samples; Step 2: Divide the number of crude oil sample sets in step 1 into calibration set samples and test set samples in a ratio of (3-4):1; Step 3: For the calibration set samples in step 2, construct an XGBoost calibration model, preprocess the MIR and NIR spectral data collected in step 1 based on a combination of different preprocessing methods, and determine the optimal preprocessing method based on leave-one-out cross-validation. The function definition of leave-one-out cross-validation is as follows: in, is the number of crude oil samples, is the loss function, is the true value of the wax content of the crude oil sample, Is not included in the The predicted value on the training set of crude oil samples; Step 4: extract and screen the characteristic variables based on the crude oil MIR and NIR spectral data processed by the leave-one-out cross-validation optimal preprocessing method in step 3; Step 5: The XGBoost correction model in step 4 is optimized by leave-one-out cross validation and grid search to obtain the optimal hyperparameters of the XGBoost correction model. The function definition of the grid search is as follows: in, is the hyperparameter space of the XGBoost calibration model; is the performance index of the prediction model for wax content in crude oil calculated using leave-one-out cross validation; The XGBoost correction model with the optimal preprocessing method and optimal hyperparameters established in steps three and five is used to predict the wax content in the test set sample crude oil in step two.
[0008] In the above technical solution, further, the several crude oil samples in step one are prepared from one crude oil sample through the following steps: collecting one crude oil sample, placing the crude oil sample in a constant temperature water bath at 60°C for 1.5 hours, then mixing it with an ultrasonic oscillator for 1 hour, and then using a vortex mixer to fully mix it for 0.5 hours, adding paraffin, and gradient ratio to 50%, to obtain several crude oil samples.
[0009] In the above technical solution, further, the number of the crude oil samples is not less than 15.
[0010] In the above technical scheme, further, the step three preprocesses the MIR and NIR spectral data based on a combination of different preprocessing methods, including: based on the maximum intensity normalization method of crude oil MIR and NIR spectra, the standard normal transformation method, the multivariate scattering correction method, the derivative method and the wavelet transform method, and preprocesses the MIR and NIR spectral data based on a combination of any two of the above methods.
[0011] In the above technical scheme, further, in the step 4, variable importance projection is used to extract and screen the characteristic variables of the pretreated crude oil MIR and NIR spectral data respectively. After screening, the wax characteristic peaks screened in the crude oil MIR spectral data are 3248nm, 3383nm, 3429nm, 3496nm, 6086nm, 6813nm, 7252nm, 8468nm, 10081nm, 10121nm, 10656nm, 11274nm and 13870nm; in the NIR spectral data, the screened CH combination bands are 2222-2500nm and 1250-1450nm, and the screened CH first and second harmonic frequency intervals are 1600-1850nm and 1100-1250nm, respectively.
[0012] In the above technical solution, further, when the XGBoost correction model in step 5 is optimized by leave-one-out cross-validation and grid search methods, two indicators, namely, coefficient of determination and root mean square error, are used as model performance evaluation parameters.
[0013] In the above technical solution, further, the determination coefficient function and the root mean square error in step 5 are defined as follows: Determination coefficient function: , root mean square error: in, It is the actual value of wax content in crude oil; is the predicted value of wax content in crude oil; is the mean of the actual values of wax content in crude oil; is the sample size of crude oil.
[0014] In the above technical solution, further, the spectral range of the Fourier transform infrared FTIR spectrometer used in the step 1 is 2500 to 20000 nm, and the spectral range of the near infrared NIR spectrometer is 800 to 2500 nm.
[0015] In the above technical solution, further, the optimal hyperparameters of the XGBoost correction model obtained by the leave-one-out cross-validation and grid search method in step 5 include three hyperparameters: n_estimators, subsample and colsample_bytree.
[0016] In the above technical solution, further, the parameter range of n_estimators is 50 to 300, the parameter range of subsample is 0.1 to 0.9, and the parameter range of colsample_bytree is 0.1 to 0.9.
[0017] Compared with the prior art, the present invention has the following advantages: based on the complementarity of Fourier transform infrared FTIR and near infrared NIR spectroscopy, the spectral data obtained by the two technologies are integrated to obtain more comprehensive spectral information, thereby improving the accuracy and reliability of the analysis results of the wax content in crude oil, and through the leave-one-out cross-validation method, the spectral data of the calibration set are optimized to obtain the optimal model hyperparameters to establish an XGBoost calibration model, and an XGBoost calibration model based on feature variable selection after optimal preprocessing is established to predict the wax content in the crude oil of the test set, thereby improving the accuracy of the XGBoost calibration model, thereby establishing a method for rapid, non-destructive and accurate quantitative analysis of the wax content in crude oil. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] Figure 1 is a spectrum of the MIR and NIR wavelengths of the crude oil sample in the embodiment of the present invention in the range of 2500-20000 nm; Figure 2 is a spectrum of the MIR and NIR wavelengths of the crude oil sample in the embodiment of the present invention in the range of 800-2500 nm; Figure 3 is a MIR and NIR feature selection spectrum of a crude oil sample in an embodiment of the present invention; Figure 4 A diagram showing the idea of constructing a crude oil quantitative analysis model in an embodiment of the present invention; Figure 5 It is a scatter plot of the prediction performance of the crude oil quantitative analysis model in the embodiment of the present invention. DETAILED DESCRIPTION
[0019] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0020] The crude oil described in this embodiment is an actual crude oil sample collected from an oil production plant.
[0021] like Figure 4 The present invention provides a method for rapid analysis of wax content in crude oil based on infrared spectrum fusion, comprising the following steps: Step 1: Use Fourier transform infrared FTIR spectroscopy and near infrared NIR spectroscopy instruments to collect MIR and NIR spectral data of several crude oil samples; Step 2: Divide the number of crude oil sample sets in step 1 into calibration set samples and test set samples in a ratio of (3-4):1; Step 3: For the calibration set samples in step 2, construct an XGBoost calibration model, preprocess the MIR and NIR spectral data collected in step 1 based on a combination of different preprocessing methods, and determine the optimal preprocessing method based on leave-one-out cross-validation. The function definition of leave-one-out cross-validation is as follows: in, is the number of crude oil samples, is the loss function, is the true value of the wax content of the crude oil sample, Is not included in the The predicted value on the training set of crude oil samples; Step 4: extract and screen the characteristic variables based on the crude oil MIR and NIR spectral data processed by the leave-one-out cross-validation optimal preprocessing method in step 3; Step 5: The XGBoost correction model in step 4 is optimized by leave-one-out cross validation and grid search to obtain the optimal hyperparameters of the XGBoost correction model. The function definition of the grid search is as follows: in, is the hyperparameter space of the XGBoost calibration model; is the performance index of the prediction model for wax content in crude oil calculated using leave-one-out cross validation; The XGBoost correction model with the optimal preprocessing method and optimal hyperparameters established in steps three and five is used to predict the wax content in the test set sample crude oil in step two.
[0022] This embodiment will introduce the details of each step in detail according to the step sequence of the above method.
[0023] The several crude oil samples in step 1 are prepared from one crude oil sample by the following steps: collecting one crude oil sample, heating the crude oil sample at 60° C. in a constant temperature water bath for 1.5 hours, mixing it with an ultrasonic oscillator for 1 hour, and then mixing it with a vortex mixer for 0.5 hours, adding paraffin, and gradient mixing to 50%, to obtain several crude oil samples.
[0024] like Figure 1-2 As shown, the calibration set samples are analyzed using a Fourier transform infrared FTIR spectrometer and a near infrared NIR spectrometer to obtain corresponding spectral data.
[0025] The infrared spectrometer used in this example is a FTIR spectrometer equipped with a DTGS detector (VERTEX 70, Bruker, Germany), and the spectrum acquisition area is 2500-20000nm. During the measurement process, the sample is directly placed on the germanium crystal ATR, and the background spectrum is collected with air as a blank before each sample measurement. Each sample is measured 16 times, and the average spectrum of the 16 measurements is used as the analytical spectrum.
[0026] The near-infrared spectrometer used in this example is Shimadzu UV-3600 Plus, and the spectrum acquisition area is 800-2500 nm. During the measurement process, the sample is evenly applied to the sample table with barium sulfate as the base. Before each sample is measured, the background spectrum is collected with barium sulfate as the blank, and then the spectrum is collected.
[0027] The MIR and NIR spectrometers were used to collect spectral data under indoor lighting conditions at a temperature of 22 to 26 °C.
[0028] In step 2, the wax content determination in crude oil is taken as an example. In this example, 17 crude oil samples are used, and the wax content reference values are shown in Table 1. When modeling, 13 samples are randomly selected from the 17 samples as the calibration set, and the remaining 4 samples are used as the test set; in this embodiment, the sample numbers of the prediction set are selected as #1, #2, #6, and #16 respectively.
[0029] Table 1 Reference values of crude oil wax content
[0030] In the step 3, an XGBoost correction model is constructed for the calibration set samples in step 2, wherein XGBoost is an existing machine learning algorithm that integrates multiple decision trees to construct a prediction model through a gradient boosting method, has strong prediction ability and efficient computing performance, and is widely used in regression and classification tasks; in the present invention, the NIR and MIR spectral intensities of the crude oil sample are corresponding to the corresponding wax content to construct the XGBoost correction model, and the optimal hyperparameters of the XGBoost correction model include three hyperparameters: n_estimators, subsample and colsample_bytree, and in the model optimization process, the determination coefficient and the root mean square error are used as evaluation indicators of the model performance; the calibration set data in step 2 is based on different preprocessing methods, including based on crude oil M IR and NIR spectra maximum intensity normalization method, standard normal transformation method, multivariate scattering correction method, derivative method and wavelet transform method, the above single preprocessing methods are all existing technologies, for example: the maximum intensity normalization method is a method for eliminating the difference between the dimensions of spectral data, the standard normal transformation method is a method for eliminating the spectral scattering effect caused by the size of the microscopic particles of the sample, the multivariate scattering correction method is a method for eliminating the spectral scattering effect caused by the difference in the uniformity of the microscopic particles of the sample, the derivative method includes the first-order derivative method and the second-order derivative method is used to correct the baseline drift phenomenon of the spectral data, the wavelet transform method is a method for processing the spectral data signal from the two dimensions of time domain and frequency domain at the same time, and the combination of any two of these methods is used to preprocess the spectral data collected from the crude oil sample, and the optimal preprocessing method is determined based on the leave-one-out cross validation. The results show that the better combination of spectral data preprocessing methods for crude oil samples is the combination of wavelet transform method and standard normal transformation method.
[0031] The preprocessed spectral data was optimized by grid search and leave-one-out cross-validation, and the coefficient of determination and root mean square error were used as evaluation parameters to obtain the optimal hyperparameters of the XGBoost correction model. The prediction accuracy of the model was improved by adjusting n_estimators, subsample, and colsample_bytree, where n_estimators ranged from 50 to 300, and this range could balance performance and complexity; subsample ranged from 0.1 to 0.9, which helped prevent overfitting and ensure that the model had enough data for training; colsample_bytree ranged from 0.1 to 0.9, and lower values (such as 0.1 and 0.3) helped reduce the correlation between features, thereby preventing overfitting, while larger values retained more information and avoided underfitting.
[0032] In the step 4, if Figure 3As shown, characteristic variables are extracted from the optimal preprocessing method, the input variables are further optimized, and the characteristic variables of the preprocessed spectral data are screened using variable importance projection. Among them, the wax characteristic peaks screened in the crude oil MIR spectral data are 3248nm, 3383nm, 3429nm, 3496nm, 6086nm, 6813nm, 7252nm, 8468nm, 10081nm, 10121nm, 10656nm, 11274nm and 13870nm; in the NIR spectral data, the combined bands of CH screened are 2222-2500nm and 1250-1450nm, and the first and second harmonic frequency intervals of CH screened are 1600-1850nm and 1100-1250nm, respectively.
[0033] In the step five, the XGBoost correction model established under the optimization preprocessing method, model hyperparameters and variable importance projection method predicts the wax content in the crude oil of the test set sample in step two. When the XGBoost correction model in step five is optimized by leave-one-out cross-validation and grid search methods, the two indicators of determination coefficient and root mean square error are used as model performance evaluation parameters.
[0034] The determination coefficient function and the root mean square error in step 5 are defined as follows: Coefficient of determination function: , root mean square error: in, It is the actual value of wax content in crude oil; is the predicted value of wax content in crude oil; is the mean of the actual values of wax content in crude oil; is the sample size of crude oil.
[0035] Model performance prediction scatter plot Figure 5 The wax content prediction model for crude oil was constructed. For wax content, the XGBoost model R 2 CV The value increased to 0.9844, RMSE CV Reduced to 1.6793%, R 2 P The above data show that the quantitative analysis model constructed based on infrared spectral fusion data has a very good prediction performance for the wax content in crude oil, and the determination coefficient function value R 2cv and root mean square error RMSEcv show that the model has very good stability. In summary, this method has the advantages of simple operation, small manual intervention, rapid analysis, low cost, accurate prediction and non-destructive analysis of samples, and is suitable for on-site rapid detection of wax content in complex crude oil matrices.
Claims
1. A rapid analysis method for wax content in crude oil based on infrared spectrum fusion, characterized in that: The following steps are involved: Step 1: Use Fourier transform infrared FTIR spectroscopy and near infrared NIR spectroscopy instruments to collect MIR and NIR spectral data of several crude oil samples; Step 2: Divide the number of crude oil sample sets in step 1 into calibration set samples and test set samples in a ratio of (3-4):1; Step 3: For the calibration set samples in step 2, construct an XGBoost calibration model, preprocess the MIR and NIR spectral data collected in step 1 based on a combination of different preprocessing methods, and determine the optimal preprocessing method based on leave-one-out cross-validation. The function definition of leave-one-out cross-validation is as follows: in, is the number of crude oil samples, is the loss function, is the true value of the wax content of the crude oil sample, Is not included in the The predicted value on the training set of crude oil samples; Step 4: extract and screen the characteristic variables based on the crude oil MIR and NIR spectral data processed by the leave-one-out cross-validation optimal preprocessing method in step 3; Step 5: The XGBoost correction model in step 4 is optimized by leave-one-out cross validation and grid search to obtain the optimal hyperparameters of the XGBoost correction model. The function definition of the grid search is as follows: in, is the hyperparameter space of the XGBoost calibration model; is the performance index of the prediction model for wax content in crude oil calculated using leave-one-out cross validation; The XGBoost correction model with the optimal preprocessing method and optimal hyperparameters established in steps three and five is used to predict the wax content in the test set sample crude oil in step two.
2. The method for rapid analysis of wax content in crude oil based on infrared spectrum fusion according to claim 1, characterized in that: The several crude oil samples in step 1 are prepared from one crude oil sample by the following steps: collecting one crude oil sample, heating the crude oil sample at 60° C. in a constant temperature water bath for 1.5 hours, mixing it with an ultrasonic oscillator for 1 hour, and then mixing it with a vortex mixer for 0.5 hours, adding paraffin, and gradient mixing to 50%, to obtain several crude oil samples.
3. The method for rapid analysis of wax content in crude oil based on infrared spectrum fusion according to claim 2, characterized in that: The number of crude oil samples shall not be less than 15.
4. The method for rapid analysis of wax content in crude oil based on infrared spectrum fusion according to claim 1, characterized in that: The step three preprocesses the MIR and NIR spectral data based on a combination of different preprocessing methods, including: preprocessing the MIR and NIR spectral data based on the maximum intensity normalization method of crude oil MIR and NIR spectra, the standard normal transformation method, the multivariate scattering correction method, the derivative method and the wavelet transformation method, and preprocessing the MIR and NIR spectral data based on a combination of any two of the above methods.
5. The method for rapid analysis of wax content in crude oil based on infrared spectrum fusion according to claim 1, characterized in that: In the step 4, variable importance projection is used to extract and screen the characteristic variables of the pretreated crude oil MIR and NIR spectral data. After screening, the wax characteristic peaks screened in the crude oil MIR spectral data are 3248nm, 3383nm, 3429nm, 3496nm, 6086nm, 6813nm, 7252nm, 8468nm, 10081nm, 10121nm, 10656nm, 11274nm and 13870nm; in the NIR spectral data, the combined bands of CH screened are 2222-2500nm and 1250-1450nm, and the first and second harmonic frequency intervals of CH screened are 1600-1850nm and 1100-1250nm, respectively.
6. The method for rapid analysis of wax content in crude oil based on infrared spectrum fusion according to claim 1, characterized in that: In the step 5, when the XGBoost correction model is optimized by leave-one-out cross validation and grid search, two indicators, the coefficient of determination and the root mean square error, are used as model performance evaluation parameters.
7. The method for rapid analysis of wax content in crude oil based on infrared spectrum fusion according to claim 6, characterized in that: The determination coefficient function and the root mean square error in step 5 are defined as follows: Coefficient of determination function: , root mean square error: in, It is the actual value of wax content in crude oil; is the predicted value of wax content in crude oil; is the mean of the actual values of wax content in crude oil; is the sample size of crude oil.
8. The method for rapid analysis of wax content in crude oil based on infrared spectrum fusion according to claim 1, characterized in that: The spectral range of the Fourier transform infrared FTIR spectrometer used in the step 1 is 2500-20000nm, and the spectral range of the near infrared NIR spectrometer is 800-2500nm.
9. The method for rapid analysis of wax content in crude oil based on infrared spectrum fusion according to claim 1, characterized in that: The optimal hyperparameters of the XGBoost correction model obtained by the leave-one-out cross-validation and grid search method in step 5 include three hyperparameters: n_estimators, subsample, and colsample_bytree.
10. The method for rapid analysis of wax content in crude oil based on infrared spectrum fusion according to claim 9, characterized in that: The parameter range of n_estimators is 50 to 300, the parameter range of subsample is 0.1 to 0.9, and the parameter range of colsample_bytree is 0.1 to 0.9.
Citation Information
Patent Citations
A method for predicting crude oil properties using near-infrared spectroscopy
CN107966420B
Method for rapidly detecting methanol content in methanol gasoline based on Raman-near-infrared spectroscopy fusion technology
CN110361373A
Frying oil quality rapid detection method based on spectrum technology
CN111103259A
Nondestructive testing method for caprolactam content in sauce food
CN112179871A
Rice category identification model, construction method thereof and method for identifying rice category
CN115452758A
Cited By
Liquid sample crystallization precipitation real-time detection method based on intelligent sensor
CN121027203A