Fast prediction system of crude oil SARA composition based on low-field nuclear magnetic resonance
By combining low-field nuclear magnetic resonance with a partial least squares regression model, the problems of long analysis time, high cost, and poor model applicability of crude oil SARA composition analysis have been solved, achieving rapid and accurate prediction of crude oil composition, which is suitable for field detection in oil fields.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- NORTHEAST GASOLINEEUM UNIV
- Filing Date
- 2026-03-23
- Publication Date
- 2026-06-02
AI Technical Summary
Existing technologies for crude oil SARA composition analysis suffer from problems such as long detection time, complex operation, high cost, poor model applicability, and environmental unfriendliness. Traditional methods cannot meet the needs of rapid on-site detection in oil fields.
By employing low-field nuclear magnetic resonance combined with a partial least squares regression model, the transverse relaxation time data of crude oil is obtained through a low-field nuclear magnetic resonance spectrometer. Data preprocessing and feature extraction are performed to establish a partial least squares regression model, enabling rapid prediction of the SARA composition of crude oil.
It enables rapid and accurate prediction of crude oil SARA composition, shortening the single analysis time to 5-10 minutes. It has high prediction accuracy, strong applicability, and is environmentally friendly, making it suitable for on-site oilfield testing and improving testing efficiency and safety.
Smart Images

Figure CN122135813A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of petrochemical analysis and energy resource development, specifically to a rapid prediction system for crude oil SARA composition based on low-field nuclear magnetic resonance. Background Technology
[0002] Crude oil is a complex natural mixture composed of thousands of hydrocarbons and non-hydrocarbon organic compounds. Its composition and structure have a decisive influence on reservoir fluid properties, extraction methods, and subsequent refining processes. Based on chemical polarity and solubility, the industry typically classifies crude oil into saturated hydrocarbons, aromatic hydrocarbons, gums, and asphaltenes, known as SARA components. SARA composition is not only an important indicator for assessing crude oil properties and processing potential, but also key data for crude oil classification, transportation stability analysis, refinery production optimization, and enhanced oil recovery design. For example, high asphaltenes and gums content lead to increased crude oil viscosity, decreased fluidity, and increased pipeline deposition; the proportion of aromatic hydrocarbons affects the efficiency of cracking and hydrogenation processes; and the proportion of saturated hydrocarbons is closely related to the yield of lubricating oil and fuel fractions. Therefore, rapidly and accurately obtaining SARA composition is a prerequisite for reservoir engineering, crude oil rheology research, and refining process control.
[0003] Currently, the mainstream international techniques for SARA determination include column chromatography (such as ASTM D2007 and D4124), high-performance liquid chromatography (HPLC), and thermogravimetric analysis (TGA). These traditional analytical methods have advantages such as high comparability and standardized measurement, but they have revealed the following shortcomings in practical applications:
[0004] Long experimental cycle: Taking column chromatography as an example, a single sample analysis usually takes 8–12 hours, and complex samples even take more than 24 hours, which cannot meet the rapid detection needs of oil fields or refining units.
[0005] The operation is complex and relies on experience: it involves many steps, frequent solvent switching, high skill requirements for experimental personnel, and poor reproducibility between different laboratories.
[0006] High reagent consumption and significant environmental and safety concerns: It requires large amounts of organic solvents (such as n-hexane, toluene, and chloroform), increasing costs and creating safety and environmental hazards.
[0007] Traditional experiments rely on fixed laboratories and large solvent systems, which are not suitable for deployment in oilfields, offshore platforms, or small analysis stations.
[0008] In recent years, nuclear magnetic resonance (NMR) has gained attention in petroleum testing due to its ability to directly detect the motion of hydrogen nuclei without destroying samples. Low-field NMR equipment, in particular, is significantly cheaper than high-field NMR and easier to deploy in the field. It can quickly obtain longitudinal and transverse relaxation time (T1, T2) spectra related to molecular motion, which can be used to infer crude oil properties. However, there are technical challenges in directly using LF-NMR signals for SARA quantitative analysis.
[0009] The signals are complex and highly overlapping: the relaxation time ranges of hydrogen nuclei of different molecular components overlap, and there is a lack of characteristic peaks that can be directly separated.
[0010] Sample differences lead to poor model transferability: The physical properties of crude oil vary significantly under different oil fields and geological conditions, making it difficult for a single empirical model to be applied across samples.
[0011] Existing multivariate methods are not robust enough: Although some studies have attempted to introduce multivariate statistical methods such as principal component regression and partial least squares regression, the sample size is limited, the data preprocessing is not standardized, the model generalization ability is insufficient, and the prediction accuracy and consistency cannot meet industrial requirements. Summary of the Invention
[0012] The purpose of this invention is to provide a rapid prediction system for crude oil SARA composition based on low-field nuclear magnetic resonance, so as to solve the problems mentioned in the background art.
[0013] To achieve the above objectives, the present invention provides the following technical solution: a rapid prediction system for crude oil SARA based on low-field nuclear magnetic resonance, the system comprising the following steps:
[0014] S1. Preparation and preliminary treatment of crude oil samples
[0015] Select crude oil samples to be tested, which can include light oil, medium oil, heavy oil and high asphaltene crude oil. After the samples are mixed at room temperature, they are loaded into a special NMR sample tube.
[0016] S2, Low-field NMR data acquisition
[0017] Transverse relaxation time of S1-treated samples was measured using a low-field nuclear magnetic resonance spectrometer. Measurement: Pulse excitation was performed using a CPMG sequence to obtain the echo sequence and invert the T2 relaxation time distribution curve to form the relaxation time spectrum; To ensure the applicability of the model, the temperature, echo interval, and echo number parameters of the acquired data were standardized.
[0018] S3. Data Preprocessing and Feature Extraction
[0019] The relaxation time spectrum obtained by S2 is normalized, mean centered and denoised; key relaxation information can be highlighted by stretching or compressing time scales, signal truncation and spectral region selection strategies; the signal can be partitioned and integrated or derived variables can be derived according to experimental requirements to enhance the robustness of subsequent modeling.
[0020] S4. Construction of Partial Least Squares Regression Model
[0021] Using the SARA components determined by column chromatography using traditional methods as reference values, and the preprocessed NMR spectrum data in S3 as input variables, a prediction model was established using the partial least squares regression algorithm.
[0022] S5. Model Optimization and Validation
[0023] The preprocessed data in S3 is divided into training and test sets. During the training process, the prediction model generated by S4 needs to respond to the T2 relaxation spectrum X and the SARA component Y of each sample. The PLSR model is built using the training set data, and the optimal number of latent variables is selected through cross-validation to minimize the root mean square error of cross-validation and maximize it. Then, the predictive ability of the PLSR model is verified using the test set data, and the R² value is calculated to prove the accuracy of the model.
[0024] S6. Rapid Prediction and Result Output
[0025] After the PLSR model in S5 is trained, the LF-NMR data of the unknown crude oil sample is input into the established PLSR model to quickly output the contents of saturated hydrocarbons, aromatic hydrocarbons, gums and asphaltenes.
[0026] Furthermore, the echo attenuation expression for CPMG in S1 is as follows:
[0027]
[0028] Among them, M( ) is the first Each echo signal strength; For the first The time corresponding to each echo; N is the time of each echo; Relaxation component fraction; For the first indivual The signal amplitude of the component; For the first Each relaxation component Relaxation time; This is the noise term, which includes system noise and baseline offset.
[0029] Furthermore, the T2 interval partitioning integration feature in S3 is as follows:
[0030]
[0031] in: The j-th Integral characteristics of an interval; , Distribution function; , the left endpoint of the j-th interval; , the right endpoint of the j-th interval;
[0032] Normalization process:
[0033]
[0034] Where: X is the original feature matrix; μ is the mean of each column (each feature); σ is the standard deviation of each column; This is the standardized matrix.
[0035] Furthermore, the partial least squares regression model in S4 establishes LF-NMR by using the partial least squares regression method. The relationship between the relaxed signal and the SARA composition; by projecting the independent variable LF-NMR signal and the dependent variable SARA composition into a new space, the covariance between the independent and dependent variables is maximized, thereby enabling the prediction of the SARA composition.
[0036] The general multivariate underlying model of partial least squares is:
[0037]
[0038]
[0039] Where: X is a The prediction matrix, Y is The response matrix; and yes The matrices are the projections of X and Y, respectively; P and Q are the projections of X and Y, respectively. and The orthogonal loading matrix, and matrices E and F are error terms, which are normally distributed random variables that are independent and identically distributed; the decomposition of X and Y is used to maximize the covariance between T and U.
[0040] Furthermore, the objective function for finding the maximum covariance in S5 is as follows:
[0041]
[0042] Where w is the weight vector in the X direction; c is the weight vector in the Y direction; and cov is the covariance.
[0043] PLS regression prediction equation:
[0044]
[0045] in, B is the model's predicted output; Q is the regression coefficient matrix; Root Mean Square Error (RMSE), Mean Absolute Error (MAE), and Coefficient of Determination (R²) are also included.
[0046]
[0047]
[0048]
[0049] Where n is the number of samples; : Actual value; Predicted value; : Average of the true values;
[0050] Furthermore, S5 also includes the following steps: using independent test samples to perform external validation of the model, evaluating the prediction accuracy and mean absolute error (MAE) index; and fine-tuning the data preprocessing based on the validation results until the model performs stably on samples with different API strengths and different viscosities.
[0051] The following is Molecular correlation time Relationship and Relationship with viscosity η:
[0052]
[0053]
[0054] Where T2 is the lateral relaxation time; These are molecular structure-related constants; Molecular correlation time; For the Larmor frequency; For fluid viscosity; The effective volume of the molecule; Boltzmann's constant; For temperature.
[0055] Furthermore, S5 also includes the following step: From spectral density The weighting and decision, with the relevant time Increase, spectral density The item rapidly increased and became dominant. ,lead to Significantly shortened;
[0056] Under low field conditions, Smaller size allows lightweight samples to approach an extremely narrow spectral region. ), while the heavy sample is still in slow motion ( ), Maintain high discrimination of SARA composition.
[0057] The beneficial effects of this invention: By organically combining low-field nuclear magnetic resonance with partial least squares regression model, this invention breaks through the bottlenecks of long time consumption, high cost and difficulty in interpreting LF-NMR data in traditional crude oil SARA analysis. It shows significant technical advantages in terms of analysis speed, prediction accuracy, model applicability and environmental friendliness.
[0058] Traditional SARA analysis usually relies on column chromatography or high performance liquid chromatography, and the detection time for a single sample often takes 8–12 hours or even longer. However, this invention rapidly acquires T2 relaxation data through LF-NMR and predicts it through PLSR modeling. A single complete analysis process can be completed in 5–10 minutes, improving efficiency by more than 90%, which can significantly reduce the workload of the laboratory and is suitable for rapid screening of large-scale samples and real-time process monitoring.
[0059] In tests on 22 Daqing crude oil samples with different API strengths (4.6°–42°), viscosities, and origins, the PLSR model prediction results of this invention showed high consistency with standard column chromatography: saturated hydrocarbons (S) had a prediction coefficient of determination R² = 0.973, with a mean absolute error (MAE) of less than 2 wt%; aromatic hydrocarbons (A) had a prediction R² = 0.961, with an MAE of less than 3 wt%; resins (R) had a prediction R² = 0.958; and asphaltenes (A) had a prediction R² = 0.952. The overall predictive performance achieved an average R² of over 0.97, remaining stable across samples of different viscosities and origins, outperforming conventional schemes combining LF-NMR with simple linear models.
[0060] This invention optimizes the number of latent variables through systematic data preprocessing and cross-validation, making the model applicable not only to light and conventional crude oils but also effectively adaptable to difficult-to-test samples such as high-asphaltite and heavy oils. Validation results show that even with significant fluctuations in high resin and asphaltenite content, the prediction error remains within ±5wt%, far superior to the generally ±10wt% fluctuation range of existing publicly available methods. This method is completely independent of large amounts of organic solvents, generating almost no chemical waste during the detection process, avoiding the potential hazards to the experimental environment and personnel health posed by traditional column chromatography; it also reduces the storage and handling of flammable organic solvents, significantly enhancing safety. It can be directly detected on-site in oilfields, providing real-time assessment of the properties of reservoir-produced crude oil and offering data support for adjusting recovery measures; it rapidly obtains the SARA distribution of feedstock oil, assisting in adjusting processing parameters and improving the operating efficiency of catalytic cracking and hydrocracking units; and it provides quantifiable compositional data, improving the objectivity and fairness of crude oil grading and trading pricing. Attached Figure Description
[0061] Figure 1 Nuclear magnetic resonance (NMR) experiment flowchart;
[0062] Figure 2 A schematic diagram of the experimental procedure for separating SARA components by column chromatography;
[0063] Figure 3 T2 relaxation distribution curve of crude oil sample;
[0064] Figure 4 Correlation between predicted values and actual SARA measurements;
[0065] Figure 5 Flowchart of the prediction system. Detailed Implementation
[0066] Please see Figure 1 , 2 5. The present invention provides a technical solution: a rapid prediction system for crude oil SARA based on low-field nuclear magnetic resonance, the system comprising the following steps:
[0067] S1. Preparation and preliminary treatment of crude oil samples
[0068] Select crude oil samples to be tested, which can include light oil, medium oil, heavy oil and high asphaltene crude oil. After the samples are mixed at room temperature, they are loaded into a special NMR sample tube. No complicated chemical treatment is required. Degassing is only performed when necessary to reduce the interference of free gas on the signal.
[0069] S2, Low-field NMR data acquisition
[0070] Transverse relaxation time of S1-treated samples was measured using a low-field nuclear magnetic resonance spectrometer. Measurement: Pulse excitation was performed using a CPMG sequence to obtain the echo sequence and invert the T2 relaxation time distribution curve to form the relaxation time spectrum; To ensure the applicability of the model, the temperature, echo interval, and echo number parameters of the acquired data were standardized.
[0071] The echo attenuation expression for CPMG in S1:
[0072]
[0073] Among them, M( ) is the first Each echo signal strength; For the first The time corresponding to each echo; N is the time of each echo; Relaxation component fraction; For the first indivual The signal amplitude of the component; For the first Each relaxation component Relaxation time; This is the noise term, which includes system noise and baseline offset.
[0074] S3. Data Preprocessing and Feature Extraction
[0075] The relaxation time spectrum obtained by S2 is normalized, mean centered and denoised; key relaxation information can be highlighted by stretching or compressing time scales, signal truncation and spectral region selection strategies; the signal can be partitioned and integrated or derived variables can be derived according to experimental requirements to enhance the robustness of subsequent modeling.
[0076] The integral characteristics of the T2 interval in S3 are as follows:
[0077]
[0078] in: The j-th Integral characteristics of an interval; , Distribution function; , the left endpoint of the j-th interval; , the right endpoint of the j-th interval;
[0079] Normalization process:
[0080]
[0081] Where: X is the original feature matrix; μ is the mean of each column; σ is the standard deviation of each column; This is the standardized matrix.
[0082] S4. Construction of Partial Least Squares Regression Model
[0083] Using the SARA components determined by column chromatography using traditional methods as reference values, and the preprocessed NMR spectrum data in S3 as input variables, a prediction model was established using the partial least squares regression algorithm.
[0084] Partial Least Squares Regression (PLSR) is a well-known multivariate forecasting method that establishes a regression model by projecting the target variable and observable variables into a new space. This bilinear factor model typically works by defining covariance pairwise. PLSR models establish LF-NMR using the partial least squares regression method. The relationship between the relaxed signal and the SARA composition; by projecting the independent variable LF-NMR signal and the dependent variable SARA composition into a new space, the covariance between the independent and dependent variables is maximized, thereby enabling the prediction of the SARA composition.
[0085] The general multivariate underlying model of partial least squares is:
[0086]
[0087]
[0088] Where: X is a The prediction matrix, Y is The response matrix; and yes The matrices are the projections of X and Y, respectively; P and Q are the projections of X and Y, respectively. and The orthogonal loading matrix, and matrices E and F are error terms, which are normally distributed random variables that are independent and identically distributed; the decomposition of X and Y is used to maximize the covariance between T and U.
[0089] S5. Model Optimization and Validation
[0090] The preprocessed data in S3 is divided into training and test sets. During training, the prediction model generated in S4 needs to respond to the T2 relaxation spectrum X and the SARA component Y of each sample. The PLSR model is built using the training set data, and the optimal number of latent variables is selected through cross-validation to minimize the root mean square error of cross-validation and maximize CV-Q². Then, the predictive ability of the PLSR model is verified using the test set data, and the R² value is calculated to prove the accuracy of the model.
[0091] Solving the objective function for maximum covariance in S5:
[0092]
[0093] Where w is the weight vector in the X direction; c is the weight vector in the Y direction; and cov is the covariance.
[0094] PLS regression prediction equation:
[0095]
[0096] in, B is the model's predicted output; Q is the regression coefficient matrix; Root Mean Square Error (RMSE), Mean Absolute Error (MAE), and Coefficient of Determination (R²) are also included.
[0097]
[0098]
[0099]
[0100] Where n is the number of samples; : Actual value; : Predicted value; : Average of the true values;
[0101] S5 also includes the following steps: using independent test samples to externally validate the model, evaluate the prediction accuracy and mean absolute error (MAE) index; and fine-tune the data preprocessing based on the validation results until the model performs stably on samples with different API strengths and viscosities.
[0102] The following is Molecular correlation time Relationship and Relationship with viscosity η:
[0103]
[0104]
[0105] Where T2 is the lateral relaxation time; These are molecular structure-related constants; Molecular correlation time; For the Lamour frequency; For fluid viscosity; The effective volume of the molecule; Boltzmann's constant; For temperature.
[0106] S5 also includes the following steps: From spectral density The weighting and decision, with the relevant time Increase, spectral density The item rapidly increased and became dominant. ,lead to Significantly shortened;
[0107] Under low field conditions, Smaller size allows lightweight samples to approach an extremely narrow spectral region. ), while the heavy sample is still in slow motion ( ), Maintain high discrimination of SARA composition.
[0108] S6. Rapid Prediction and Result Output
[0109] After the PLSR model in S5 is trained, the LF-NMR data of the unknown crude oil sample is input into the established PLSR model to quickly output the contents of saturated hydrocarbons, aromatic hydrocarbons, gums and asphaltenes.
[0110] Reference Figure 3 The T2 relaxation time spectrum of the crude oil sample was obtained by performing low-field nuclear magnetic resonance (LF-NMR) T2 relaxation measurement. Figure 3 The T2 relaxation curves for each unit mass of crude oil samples are shown, revealing significant differences in the T2 spectra among different samples. These differences in relaxation distribution reflect the kinematic characteristics and interactions of different molecular components within the crude oil samples.
[0111] The T2 relaxation distribution curves from NMR show that the longer T2 relaxation time indicates that these components have low viscosity and high fluidity, which is generally considered an advantageous characteristic in oilfield development because these crude oils are easy to recover and process.
[0112] Heavy components typically exhibit shorter T2 relaxation times, which is related to their complex molecular structure, high viscosity, and poor fluidity. These components are generally more difficult to handle during crude oil processing and often require special recovery techniques in oilfield development, such as steam flooding or chemical flooding, to improve recovery rates.
[0113] The broad T2 relaxation signal distribution indicates the presence of multiple different molecular motion modes. This suggests significant differences in molecular size and structure within these samples, encompassing a wide range of components from light to heavy. This complexity may affect the overall processing and handling difficulty of the crude oil, and also suggests the need for more diversified and flexible technological approaches in the development of these oil fields.
[0114] The presence of two significant peaks within a short timeframe indicates complex molecular motion characteristics, suggesting the possible inclusion of multiple components, including light, heavy, and possibly other intermediate components. This complex T2 relaxation time distribution suggests the multiphase and component diversity of these samples, potentially requiring the comprehensive use of various processing and treatment techniques to improve recovery rates.
[0115] Reference Figure 4 As shown, the correlation between actual SARA composition and model predictions is illustrated. It is evident that most data points are close to the diagonal, indicating a strong linear correlation between the model's predictions and actual measurements. This further demonstrates the effectiveness and potential of the PLSR model in crude oil composition prediction, showcasing its broad application prospects in crude oil exploration and development.
[0116] The predicted values were compared with the actual measured values, and the R² value for each SARA component and each sample was calculated separately. Table 5 shows the correlation (R² value) between the predicted and actual values for each SARA component. Table 6 shows the R² value for each validation data. The R² value reflects the proportion of the variance of the independent variable explained by the model; the closer the value is to 1, the better the model fits the data.
[0117] Overall, the R² value reached 0.8185 for saturated hydrocarbons, 0.7431 for aromatics, 0.7901 for non-hydrocarbons, and 0.7718 for asphaltenes, demonstrating the model's good performance in predicting various components. In particular, the prediction accuracy for saturated hydrocarbons and non-hydrocarbons was high, which is related to the relatively simple structures and small intermolecular interactions of these components.
[0118] Example 1: Based on the above prediction system;
[0119] Sample source and preparation: Crude oil samples with API values ranging from 4.6° to 42° were selected, covering light, medium, heavy, and high-asphalt crude oils. All samples were mixed at room temperature and loaded into standard NMR sample tubes (10 mm in diameter). No salt or metal removal or solvent separation was required; degassing was only performed on some high-gas-content samples.
[0120] LF-NMR data acquisition: Instrument model: MesoMR23-040V, Numatron Corporation; Magnet frequency: 22MHz; Probe diameter: 25mm; Pulse sequence: CPMG; Parameters: Waiting time TW=6s, echo interval τ=0.6ms, echo number 18,000, accumulation count 32; Test temperature: 25℃; Data export: Obtain the T2 relaxation time distribution curve after baseline subtraction.
[0121] Data preprocessing: Environment: Computer software Python 3.10, main libraries include numpy, pandas, scipy, scikit-learn; Workflow: Use scipy.signal to denoise and correct the baseline of the T2 signal; use StandardScaler for area normalization and mean centering; perform partitioned integration on the T2 time interval to extract 30–50 feature variables and construct the input matrix X; SARA measured values (column tomography, SY / T5119-2016) are used as supervision variables Y.
[0122] PLSR model prediction:
[0123] In Python, you can directly read the trained PLSRegression while maintaining the predetermined number of latent variables and hyperparameters.
[0124] Input the preprocessed features X of the current sample into the model, and output the predicted values of S, A, R, and As; the entire process for a single sample takes about 5 minutes.
[0125] If a small number of reference SARA values for this batch of samples are obtained on-site, incremental calibration / small sample refitting is used (only the model intercept is fine-tuned or domain adaptive weight updates are performed), without changing the core hyperparameters;
[0126] The evaluation metrics are mainly R², MAE, and RMSE; fine-tuning is triggered only when the performance is below a preset threshold to avoid overfitting and model drift.
[0127] Results output: Generates the content of four components and error assessment for each sample, supporting tabular / graphical export.
[0128] The method of this invention can rapidly and accurately predict the SARA composition of various API crude oil samples without the use of chemical solvents; it is operable and scalable in both laboratory and oilfield settings.
[0129] Example 2: Extended to heavy crude oil and high asphaltene samples;
[0130] Based on the relevant steps of Example 1 above, the difference is as follows: eight samples were selected, including extra-heavy crude oil with an API value of <10° and asphaltene content >20%; high-resolution integration of short relaxation time periods (<1 ms) was added to the T2 data to optimize the latent variables to eight; and the model was subjected to sample weighting to improve its sensitivity to asphaltene signals.
[0131] The prediction error for high asphalt content samples was controlled within ±5 wt%, which is better than the generally ±10 wt% fluctuation range of existing publicly available LF-NMR models.
[0132] Example 3: System integration and rapid on-site testing;
[0133] Based on the relevant steps of Embodiment 1 above, the difference lies in: integrating the LF-NMR measurement module, data processing unit, and wireless transmission module into a portable device, with a pre-trained PLSR model built in; directly installing the device for testing after collecting crude oil at the oilfield site, and displaying the SARA prediction results on a tablet terminal within 5 minutes; the device can complete the analysis without professional operators; and it supports real-time oil property screening, providing data basis for dynamically adjusting the harvesting process and crude oil grading and pricing.
Claims
1. A rapid prediction system for crude oil SARA based on low-field nuclear magnetic resonance, characterized in that: The system includes the following steps: S1. Preparation and preliminary treatment of crude oil samples Select crude oil samples to be tested, which can include light oil, medium oil, heavy oil and high asphaltene crude oil. After the samples are mixed at room temperature, they are loaded into a special NMR sample tube. S2, Low-field NMR data acquisition Transverse relaxation time of S1-treated samples was measured using a low-field nuclear magnetic resonance spectrometer. Measurement: Pulse excitation was performed using a CPMG sequence to obtain the echo sequence and invert the T2 relaxation time distribution curve to form the relaxation time spectrum; To ensure the applicability of the model, the temperature, echo interval, and echo number parameters of the acquired data were standardized. S3. Data Preprocessing and Feature Extraction The relaxation time spectrum obtained by S2 is normalized, mean centered and denoised; strategies such as stretching or compressing time scales, signal truncation and spectral region selection can be used to highlight key relaxation information. Based on experimental requirements, the signal can be partitioned and integrated or derived into variants to enhance the robustness of subsequent modeling. S4. Construction of Partial Least Squares Regression Model Using the SARA components determined by column chromatography using traditional methods as reference values, and the preprocessed NMR spectrum data in S3 as input variables, a prediction model was established using the partial least squares regression algorithm. S5. Model Optimization and Validation The preprocessed data in S3 is divided into training set and test set. During the training process, the prediction model generated by S4 needs to respond to the T2 relaxation spectrum X and the SARA component Y of each sample. The PLSR model is built using the training set data, and the optimal number of latent variables is selected through cross-validation to minimize the root mean square error of cross-validation and maximize it. Then, the predictive ability of the PLSR model is validated using the test set data, and the R² value is calculated to prove the accuracy of the model. S6. Rapid Prediction and Result Output After the PLSR model in S5 is trained, the LF-NMR data of the unknown crude oil sample is input into the established PLSR model to quickly output the contents of saturated hydrocarbons, aromatic hydrocarbons, gums and asphaltenes.
2. The rapid prediction system for crude oil SARA composition based on low-field nuclear magnetic resonance according to claim 1, characterized in that: The echo attenuation expression for CPMG in S1 is as follows: Among them, M( ) is the first Each echo signal strength; For the first The time corresponding to each echo; N is the time of each echo; Relaxation component fraction; For the first indivual The signal amplitude of the component; For the first Each relaxation component Relaxation time; This is the noise term, which includes system noise and baseline offset.
3. The rapid prediction system for crude oil SARA composition based on low-field nuclear magnetic resonance according to claim 1, characterized in that: The T2 interval partitioning integral characteristics in S3 are as follows: in: The j-th Integral characteristics of an interval; , Distribution function; , the left endpoint of the j-th interval; , the right endpoint of the j-th interval; Normalization process: Where: X is the original feature matrix; μ is the mean of each column (each feature); σ is the standard deviation of each column; This is the standardized matrix.
4. The rapid prediction system for crude oil SARA composition based on low-field nuclear magnetic resonance according to claim 1, characterized in that: The partial least squares regression model in S4 establishes LF-NMR by using the partial least squares regression method. The relationship between the relaxed signal and the SARA composition; by projecting the independent variable LF-NMR signal and the dependent variable SARA composition into a new space, the covariance between the independent and dependent variables is maximized, thereby enabling the prediction of the SARA composition. The general multivariate underlying model of partial least squares is: Where: X is a The prediction matrix, Y is The response matrix; and yes The matrices are the projections of X and Y, respectively; P and Q are the projections of X and Y, respectively. and The orthogonal loading matrix, and matrices E and F are error terms, which are normally distributed random variables that are independent and identically distributed; the decomposition of X and Y is used to maximize the covariance between T and U.
5. The rapid prediction system for crude oil SARA composition based on low-field nuclear magnetic resonance according to claim 1, characterized in that: The objective function for finding the maximum covariance in S5 is as follows: Where w is the weight vector in the X direction; c is the weight vector in the Y direction; and cov is the covariance. PLS regression prediction equation: in, B is the model's predicted output; Q is the regression coefficient matrix; Root Mean Square Error (RMSE), Mean Absolute Error (MAE), and Coefficient of Determination (R²) are also included. Where n is the number of samples; : Actual value; Predicted value; : Average of the true values.
6. The rapid prediction system for crude oil SARA composition based on low-field nuclear magnetic resonance according to claim 5, characterized in that: S5 also includes the following steps: using independent test samples to perform external validation of the model, evaluating the prediction accuracy and mean absolute error (MAE) index; and fine-tuning the data preprocessing based on the validation results until the model performs stably on samples with different API strengths and different viscosities. The following is Molecular correlation time Relationship and Relationship with viscosity η: Where T2 is the lateral relaxation time; These are molecular structure-related constants; Molecular correlation time; For the Larmor frequency; For fluid viscosity; The effective volume of the molecule; Boltzmann's constant; For temperature.
7. The rapid prediction system for crude oil SARA composition based on low-field nuclear magnetic resonance according to claim 6, characterized in that: The S5 further includes the following steps: From spectral density The weighting and decision, with the relevant time Increase, spectral density The item rapidly increased and became dominant. ,lead to Significantly shortened; Under low field conditions, Smaller size allows lightweight samples to approach an extremely narrow spectral region. ), while the heavy sample is still in slow motion ( ), Maintain high discrimination of SARA composition.