Petroleum hydrocarbon viscosity prediction and uncertainty quantification method based on machine learning

By acquiring multi-source molecular descriptors of petroleum hydrocarbon mixtures through machine learning methods, and using improved Boruta feature screening and Delta method to quantify uncertainty, the universality and cost issues of liquid viscosity prediction are solved, achieving high-precision viscosity prediction and risk control, applicable to the variable operating conditions of petroleum hydrocarbons.

CN121862235APending Publication Date: 2026-04-14TSINGHUA UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-18
Publication Date
2026-04-14

AI Technical Summary

Technical Problem

Existing technologies for liquid viscosity prediction suffer from poor universality, strong parameter dependence, high computational cost, and insufficient adaptability to new structures, making it difficult to achieve reliable viscosity prediction and risk control under the variable operating conditions of petroleum hydrocarbons.

Method used

A machine learning-based approach is adopted to acquire multi-source molecular descriptor data of petroleum hydrocarbon mixtures, use an improved Boruta feature screening algorithm to select core features, construct a virtual single-fluid model, and use the Delta method to quantify uncertainty, output viscosity prediction values ​​and their confidence intervals.

Benefits of technology

It improves the accuracy of viscosity prediction and the generalization ability of the model, enabling reliable viscosity assessment and risk control under varying operating conditions, reducing the cost of engineering data acquisition, and enhancing the fluidity and safety of the petroleum industry.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121862235A_ABST
    Figure CN121862235A_ABST
Patent Text Reader

Abstract

The invention discloses a petroleum hydrocarbon viscosity prediction and uncertainty quantification method based on machine learning, which can improve the precision and adaptability of petroleum hydrocarbon pure component and mixture viscosity prediction, and realizes uncertainty quantification of a prediction result through a Delta method and conformal prediction. And a reliable support is provided for viscosity evaluation and risk control under variable working conditions in the petroleum industry.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of petroleum hydrocarbon property prediction, and in particular to a method and system for predicting the viscosity of petroleum hydrocarbons and quantifying their uncertainty based on machine learning. Background Technology

[0002] Viscosity is a fundamental physical property parameter characterizing the flowability of a liquid. Essentially, it measures the rate of deformation caused by shear force and, under normal processing and transportation conditions, obeys Newton's law of viscosity. ,in, It is a constant that is independent of the shear rate, but depends on state parameters such as molecular structure, temperature and pressure. Shear stress represents the shear force exerted on a unit area of ​​fluid. Viscosity, or velocity gradient (or shear rate), represents the rate of change of velocity perpendicular to the flow direction. In process industries, viscosity is not only a fundamental parameter describing the flow behavior of matter, but also a key indicator in heat transfer, mass transfer, and reaction kinetics analysis. As a core parameter of the Navier-Stokes equations, viscosity directly affects the velocity distribution of the flow field, thereby determining the energy and mass transport efficiency within equipment. Especially under continuous and large-scale production conditions, the increased scale of equipment makes the viscosity-dominated effect more significant, exerting a considerable influence on system pressure drop, flow stability, and heat transfer efficiency; its essence is an energy dissipation mechanism.

[0003] Based on the global initiative of the Sustainable Development Goals (SDGs), viscosity control has become a key variable in promoting cleaner production, improving energy efficiency, and driving industrial innovation. In the oil and gas sector, viscosity is directly related to crude oil seepage behavior in reservoirs, surface gathering and transportation efficiency, and energy consumption in long-distance pipeline transportation. As one of the two key parameters for measuring the transportability of oil products, it is crucial for ensuring fluidity and operational safety. With the co-transportation of crude oil from multiple sources becoming the norm, dynamic fluctuations in physical properties make viscosity a core variable for scheduling optimization and risk management. Furthermore, liquid viscosity also plays a significant role in lubricants, coatings, personal care products, and pharmaceutical formulations, directly affecting the rheological properties, film-forming characteristics, and end-user experience of these materials.

[0004] Due to the complexity of intermolecular interactions in liquids, theoretical models for liquid viscosity are still immature and lack a unified, universally applicable modeling framework. Existing research has primarily employed three approaches to liquid viscosity modeling, but all have significant limitations.

[0005] The first type of method consists of semi-empirical models built upon physical assumptions. Experimental measurements of liquid viscosity are costly and inefficient, especially under wide temperature and pressure ranges, making them unsuitable for practical applications and often resulting in missing viscosity data for the target environment. Existing theoretical equations lack universality and primarily rely on empirically modified semi-empirical models. These models are based on physical assumptions including Eyring's absolute rate theory, free volume theory, friction theory, and the expansion fluid assumption, leading to classical viscosity-temperature relationships such as the Andrade equation and the Vogel-Fulcher-Tammann (VFT) model. Property correlation models based on the contrastive state principle (CSP) normalize viscosity through critical parameters, offering simplicity and versatility. However, they rely on macroscopic parameters rather than molecular mechanisms, limiting prediction accuracy to the accuracy of input parameters, exhibiting weak adaptability to novel substances, and often exceeding 30% error in strongly polar systems.

[0006] The second type of method is molecular dynamics simulation (MD), which predicts the viscosity of a system by explicitly simulating particle motion and interaction processes at the atomic or molecular scale. Based on microscopic mechanisms, it can provide theoretical support when experimental data is scarce. Although it has theoretical support, its prediction accuracy is limited by the accuracy of the force field model and the degree of simplification of the system. It can usually only reach the accuracy level of experimental values ​​by one order of magnitude, and it consumes huge computational resources, making it difficult to apply on a large scale in engineering practice.

[0007] The third category of methods focuses on establishing structure-property relationship models for viscosity based on molecular structure, mainly including group contribution method (GC) and quantitative structure-activity relationship (QSPR) modeling. The GC method decomposes the compound structure into predefined groups and estimates properties by summing the contribution values. It performs well for compounds with well-defined groups, but faces significant challenges in isomer identification and handling of new groups. The QSPR method utilizes molecular descriptors and statistical learning algorithms to construct a nonlinear mapping between molecular structure and viscosity, demonstrating excellent prediction accuracy and model flexibility. However, it relies on a manually defined descriptor set, facing the risk of overfitting due to the curse of dimensionality and multicollinearity, or the trade-off between insufficient low-dimensional spatial information weakening prediction accuracy, thus limiting its generalization ability in complex chemical spaces.

[0008] All three types of methods have significant drawbacks: semi-empirical models have poor universality and strong parameter dependence, molecular dynamics simulations are costly, and structure-property methods are not adaptable to new structures. These factors collectively limit the reliability and promotion value of viscosity prediction models in practical industrial applications. Summary of the Invention

[0009] To overcome the shortcomings of semi-empirical models, such as poor universality and strong parameter dependence, high cost of molecular dynamics simulation, and poor adaptability of structure-property methods to new structures, this invention proposes a machine learning-based method for predicting the viscosity of petroleum hydrocarbons and quantifying its uncertainty. This method outputs both the predicted viscosity value and the corresponding prediction range, thus better adapting to the variable operating conditions in petroleum transportation and processing.

[0010] The second objective of this invention is to propose a system for predicting the viscosity of petroleum hydrocarbons and quantifying its uncertainty based on machine learning.

[0011] The third objective of this invention is to provide a computer device.

[0012] A fourth objective of this invention is to provide a non-transitory computer-readable storage medium.

[0013] To achieve the above objectives, a first aspect of the present invention proposes a method for predicting the viscosity of petroleum hydrocarbons and quantifying its uncertainty based on machine learning, comprising: S1, acquire multi-source molecular descriptor data of petroleum hydrocarbon mixture, wherein the multi-source molecular descriptor includes two-dimensional molecular descriptor, three-dimensional molecular descriptor and molecular structure fingerprint features; S2, based on the improved Boruta feature selection algorithm, processes multi-source molecular descriptors, automatically selects features and iterates through random forest importance measurement and shadow feature test, terminates the iteration when there are no undetermined features or the preset number of iterations is reached, and finally determines the core feature set; S3. A virtual single-fluid model is constructed using the core feature set. Molecular descriptors of each component of the mixture are weighted and combined using a molar weighting or energy weighting strategy to generate a virtual pure substance descriptor of the mixture. S4. Based on the virtual pure substance descriptor and the preset viscosity equation parameter distribution, the confidence interval of the viscosity prediction value is calculated using the Delta method, and the viscosity prediction result and its corresponding uncertainty quantification interval are output.

[0014] In one embodiment of the present invention, S2 includes: Two-dimensional molecular descriptors were calculated using RDKit and the Mordred cheminformatics module. Molecular fingerprints are generated based on PubChem substructure definitions, and three-dimensional molecular descriptors are calculated after geometric optimization of the molecular conformation. Two-dimensional molecular descriptors, PubChem molecular structure fingerprints, and three-dimensional molecular descriptors are converted into NumPy array format and concatenated along the feature dimension to form a unified high-dimensional feature matrix.

[0015] In one embodiment of the present invention, S2 further includes: The Boruta algorithm was rewritten specifically to replace native Python loop operations with NumPy's broadcast mechanism and matrix operations. In each iteration, shadow features are generated in batches using NumPy's random module and vectorized concatenation is performed. The iteration marks the confirmed salient features and irrelevant features to be removed until all features are judged or the preset maximum number of iterations is reached, and the filtered salient feature subset is output.

[0016] In one embodiment of the present invention, using screened molecular features and experimental viscosity data, four regression models—multiple linear regression, Gaussian process regression, random forest, and extreme learning machine—are constructed and optimized, with the goal of fitting the following viscosity equation. and And optimize it:

[0017] To address the uncertainty of the model, the linear regression model employs error propagation analysis, the Gaussian process regression directly utilizes its Bayesian posterior distribution, the random forest combined with the Jackknife+ method evaluates the residual distribution, and the extreme learning machine designs a direct interval prediction scheme based on the characteristics of the network.

[0018] In one embodiment of the present invention, the method further includes constructing a mixture descriptor by weighted combination of the descriptors of each component, so as to transform the QSPR modeling of the mixture into the modeling problem of a quasi-pure substance: Among them, the viscosity mixing rules for petroleum hydrocarbons are as follows: Kendall–Monroe model: The Kendall–Monroe model is proposed based on the additivity of the cube root of viscosity, and its expression is:

[0019] Arrhenius model: The Arrhenius model assumes a linear relationship between the logarithm of viscosity and composition, in the form of:

[0020] Bingham model: The Bingham model uses fluid flow as an additive physical quantity, and its form is as follows:

[0021] Confidence interval prediction: For machine learning models, the distribution that the predicted parameters may follow is given as follows:

[0022] use replace ,use replace This yields a point estimate of the logarithmic viscosity:

[0023] exist Item, cannot be determined The distribution follows a normal distribution. Using the delta method, we can estimate its asymptotic normal approximation. Let:

[0024] According to the standard asymptotic theory:

[0025] in for and Given the covariance matrices of the variables, and considering that they are independent, we have:

[0026] definition:

[0027] Its variance is approximately:

[0028] in:

[0029] Therefore, we have given The standard deviation estimate is:

[0030] as well as The 95% confidence interval is estimated as follows:

[0031] To achieve high-precision, interpretable, and uncertainty-quantified prediction of petroleum hydrocarbon viscosity.

[0032] To achieve the above objectives, a second aspect of this application proposes a system for predicting the viscosity of petroleum hydrocarbons and quantifying its uncertainty based on machine learning, comprising: A multi-source molecular descriptor acquisition module is used to acquire multi-source molecular descriptor data of petroleum hydrocarbon mixtures. The multi-source molecular descriptors include two-dimensional molecular descriptors, three-dimensional molecular descriptors, and molecular structure fingerprint features. An improved Boruta feature selection module is used to process the multi-source molecule descriptors based on the improved Boruta feature selection algorithm. Features are automatically selected and iterated through random forest importance measurement and shadow feature test. The iteration is terminated when there are no undetermined features or the preset number of iterations is reached, and the core feature set is finally determined. The virtual single-fluid model construction module is used to construct a virtual single-fluid model using the core feature set. It uses a molar weighting or energy weighting strategy to weight and combine the molecular descriptors of each component of the mixture to generate a virtual pure substance descriptor of the mixture. The Delta method uncertainty quantification module is used to calculate the confidence interval of the viscosity prediction value based on the virtual pure substance descriptor and the preset viscosity equation parameter distribution, and output the viscosity prediction result and its corresponding uncertainty quantification interval.

[0033] To achieve the above objectives, a third aspect of this application provides a computer device, including a processor and a memory; wherein the processor runs a program corresponding to the executable program code by reading executable program code stored in the memory, for implementing the method described in the first aspect embodiment.

[0034] To achieve the above objectives, a fourth aspect of this application provides a non-transitory computer-readable storage medium having a computer program stored thereon that, when executed by a processor, implements the method described in the first aspect.

[0035] The methods, systems, devices, and storage media of this invention can improve the prediction accuracy and model generalization ability of the viscosity of pure components and mixtures of petroleum hydrocarbons, and realize the quantification of the uncertainty of the prediction results through the Delta method, providing reliable support for viscosity assessment and risk control in the petroleum industry under variable operating conditions. Attached Figure Description

[0036] The above and / or additional aspects and advantages of the present invention will become apparent and readily understood from the following description of the embodiments taken in conjunction with the accompanying drawings, wherein: Figure 1 A flowchart illustrating a method for predicting the viscosity of petroleum hydrocarbons and quantifying its uncertainty based on machine learning, provided as an embodiment of the present invention; Figure 2 A framework diagram of a method for predicting the viscosity of petroleum hydrocarbons and quantifying its uncertainty based on machine learning, provided in an embodiment of the present invention; Figure 3 Homologues of PONA and Comparison chart of prediction results; Figure 4 For the four models A comparison chart of prediction results; Figure 5 For the four models A comparison chart of prediction results; Figure 6 A comparison chart showing the prediction performance of GPR and ELM models for different blended oil products; Figure 7 A structural diagram of a system for predicting the viscosity of petroleum hydrocarbons and quantifying its uncertainty based on machine learning, provided in an embodiment of the present invention; Figure 8 The computer device provided in the embodiments of the present invention. Detailed Implementation

[0037] It should be noted that, unless otherwise specified, the embodiments and features described in the present invention can be combined with each other. The present invention will now be described in detail with reference to the accompanying drawings and embodiments.

[0038] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.

[0039] The following description, with reference to the accompanying drawings, describes a method, system, computer device, and storage medium for predicting the viscosity of petroleum hydrocarbons and quantifying its uncertainty based on machine learning, according to embodiments of the present invention.

[0040] Example 1 This embodiment provides a method for predicting the viscosity of petroleum hydrocarbons and quantifying its uncertainty based on machine learning. For example... Figure 1 As shown, the method includes the following steps: S1, acquire multi-source molecular descriptor data of petroleum hydrocarbon mixture, wherein the multi-source molecular descriptor includes two-dimensional molecular descriptor, three-dimensional molecular descriptor and molecular structure fingerprint features.

[0041] Specifically, in some implementations, acquiring multi-source molecular descriptor data of petroleum hydrocarbon mixtures is a key preprocessing step in the technical solution of this invention. This technical implementation is based on a joint processing flow of cheminformatics and computational chemistry. Specifically, this step integrates three types of features: two-dimensional molecular descriptors, three-dimensional molecular descriptors, and substructure fingerprints, to construct a high-dimensional, multi-scale molecular characterization system, providing the basic input for structure-property mapping in subsequent machine learning modeling.

[0042] Two-dimensional descriptors were primarily extracted using RDKit and the Mordred Chemical Information Toolkit, encompassing over 200 candidate features, including topological indices, molecular weight, number of hydrogen bond donors / acceptors, polar surface area (PSA), and lipophilicity parameters (logP). Three-dimensional descriptors were calculated using Gaussian-16 software after molecular conformation optimization at the B3LYP / 6-31G* theoretical level, including van der Waals volume, molecular shape index (MSI), dipole moment, and charge distribution, as shown in Table 1. Molecular structural fingerprints were generated using PubChem fingerprints, with a length of 881 bits, representing the presence of specific substructure patterns in the molecule, such as aromatic rings, branched structures, and functional groups, in binary vector form.

[0043] Table 1 Predefined quantum chemical descriptors

[0044] This step is applicable to viscosity prediction modeling of petroleum hydrocarbon mixtures, and has significant advantages, especially in multi-component, non-ideal mixture systems. By treating the mixture as a "virtual single fluid," the molecular descriptors of each component are fused using mole fraction or energy weighting to form a unified feature vector, thus avoiding the prediction bias of traditional mixing rules (such as Kendall-Monroe, Arrhenius, and Bingham) when there are large structural differences. This method can be used as an input feature preprocessing module in scenarios such as crude oil scheduling in refineries, pipeline transportation optimization, and lubricant formulation design, providing a feature set with high structural sensitivity and broad information dimensions for subsequent model training.

[0045] The fusion of multi-source molecular descriptors significantly enhances the model's sensitivity to and predictive ability regarding changes in molecular structure. By introducing three-dimensional descriptors, spatial interactions and conformational effects between molecules can be captured, compensating for the lack of molecular configuration information in two-dimensional descriptors. Molecular structural fingerprints enhance the ability to identify local chemical environments, helping to capture subtle differences between isomers. This step lays a solid data foundation for subsequent Boruta feature optimization and regression modeling, and is a key prerequisite for achieving high-precision, low-bias viscosity prediction.

[0046] S2 processes multi-source molecular descriptors based on the improved Boruta feature selection algorithm. It automatically selects features and iterates through random forest importance measurement and shadow feature test. The iteration terminates when there are no undetermined features or when the preset number of iterations is reached, and finally determines the core feature set.

[0047] Specifically, this step employs an improved Boruta feature selection algorithm to process multi-source molecular descriptors, aiming to automatically identify core feature sets highly correlated with viscosity response, thereby improving the prediction accuracy and generalization ability of subsequent machine learning models. The Boruta algorithm is a feature selection method based on random forests. Its core idea is to introduce "shadow features" to compare with the original features, evaluate their importance, and ultimately select features with significant statistical meaning.

[0048] In its implementation, this invention optimizes the Boruta algorithm for efficient operation in the NumPy environment. Input features include two-dimensional molecular descriptors extracted by RDKit and Mordred, PubChem molecular structure fingerprints, and three-dimensional molecular descriptors calculated after conformation optimization using Gaussian-16 (B3LYP / 6-31G*). These descriptors encompass multi-dimensional information such as molecular topology, electronic properties, and spatial configuration, with the total number of features typically ranging from hundreds to thousands. The improved Boruta algorithm calculates the importance of each feature or Mean Decrease Impurity (MDI) using a random forest model and compares it with randomly shuffled shadow features. If the importance of the original feature is significantly higher than its corresponding shadow feature, the feature is considered statistically significant and retained; otherwise, it is discarded.

[0049] The number of trees in the random forest model can be set to 100, using the default Gini impurity as the splitting criterion. Shadow features are generated using a random column permutation strategy, with one shadow feature generated for each original feature, for a total of 100 iterations. The algorithm terminates if no new "pending" features are added during iteration or if the preset iteration limit is reached. The output of this step is a filtered subset of features, whose dimensionality can typically be reduced to 4%-10% of the original features, significantly alleviating the "curse of dimensionality" problem caused by high-dimensional data.

[0050] This feature selection step plays a crucial role in the entire technical solution. On the one hand, it improves the training efficiency and prediction stability of the model by eliminating redundant and noisy features; on the other hand, the selected core feature set can more accurately reflect the physical relationship between molecular structure and viscosity, providing high-quality input for subsequent viscosity modeling and uncertainty quantification. This method is particularly suitable for complex molecular systems of petroleum hydrocarbon mixtures in practical applications, effectively addressing challenges such as isomers and multi-source data fusion, and exhibits good engineering adaptability and scalability.

[0051] S3. A virtual single-fluid model is constructed using the core feature set. Molecular descriptors of each component of the mixture are weighted and combined using a molar weighting or energy weighting strategy to generate a virtual pure substance descriptor for the mixture.

[0052] Specifically, in some implementations, this invention transforms the viscosity prediction problem of mixed petroleum hydrocarbon systems into a modeling process of a "virtual pure substance" by constructing a virtual single-fluid model. Specifically, this step first involves weighting the molecular descriptors of each component in the mixture based on a pre-selected core molecular feature set (such as highly relevant features extracted from RDKit, Mordred, and 3D molecular descriptors using an improved Boruta algorithm). The weighting strategy can optionally employ molar weighting or energy weighting to reflect the relative contributions of different components in the mixed system. In the molar weighting strategy, the descriptors of each component are weighted according to their mole fraction. Perform a linear combination, that is:

[0053] in, A vector of virtual pure substance descriptors representing a mixture. For the first Molecular descriptor vectors of each component, This refers to the number of components in the mixture. The energy-weighted strategy further introduces the energy percentage of each component in the mixture. To more accurately reflect the effect of intermolecular interactions on viscosity, it takes the form:

[0054] This step relies on the standardized processing of molecular descriptors and the efficient implementation of weighted algorithms at the implementation level. In practical applications, the component information of the mixture needs to be obtained through gas chromatography analysis or mass spectrometry data, and its mole fraction or energy ratio needs to meet certain requirements. The normalization conditions are obtained. This virtual single-fluid model simplifies the complex multi-component problem of mixtures into a viscosity prediction problem for a single virtual substance, thereby significantly improving the model's computational efficiency and prediction stability.

[0055] Furthermore, this step plays a crucial role in this invention, serving as a bridge between the preceding and subsequent steps. On one hand, it relies on the highly relevant molecular descriptors extracted by the preceding feature screening module; on the other hand, it provides a unified input format for subsequent viscosity prediction models (such as GPR, ELM, etc.), enabling the models to achieve robust viscosity prediction for complex systems without needing to be retrained for each mixture. This method demonstrates good adaptability and generalization ability in viscosity prediction of petroleum hydrocarbon mixtures, and is particularly suitable for industrial mixing systems with diverse components and high non-ideal properties.

[0056] S4. Based on the virtual pure substance descriptor and the preset viscosity equation parameter distribution, the confidence interval of the viscosity prediction value is calculated using the Delta method, and the viscosity prediction result and its corresponding uncertainty quantification interval are output.

[0057] Specifically, in this invention, calculating the confidence interval of the viscosity prediction value using the Delta method based on a virtual pure substance descriptor and a preset viscosity equation parameter distribution is one of the key steps in achieving uncertainty quantification (UQ). This step quantifies the uncertainty of the viscosity prediction model output through statistical derivation and error propagation theory, thereby providing a statistically significant prediction interval for engineering applications.

[0058] This step first transforms the mixed petroleum hydrocarbon system into a "virtual pure substance" based on the virtual single-fluid approximation method. Its descriptor is constructed through weighted combinations of components (e.g., molar weighting, energy weighting). Subsequently, a trained regression model (e.g., GPR, ELM) is used to analyze two key parameters in the viscosity equation. and Make a prediction, assuming it follows a normal distribution, that is:

[0059] Based on parameter estimation, a viscosity prediction function is defined. ,in This is the current operating temperature. Because... and There is a nonlinear relationship between them, and their joint distribution is not strictly normal. Therefore, the Delta method is used for an asymptotically normal approximation. The Delta method approximates the variance of the nonlinear function through a first-order Taylor expansion, and its gradient vector is:

[0060] Combining covariance matrix Calculate Standard deviation estimation:

[0061] Finally, based on this standard deviation, a 95% confidence interval can be constructed as follows:

[0062] In practical applications, this step can be embedded into a petroleum hydrocarbon viscosity prediction system to output predicted viscosity values ​​and their corresponding confidence intervals under different temperature, pressure, and composition conditions. For example, in oil pipeline design, this interval can be used to assess the impact of viscosity fluctuations at specific operating temperatures on pumping energy consumption and flow stability, thereby providing a basis for equipment redundancy design and operational optimization.

[0063] By introducing statistical uncertainty analysis, the reliability and engineering applicability of viscosity prediction are improved. Compared with traditional point prediction methods, the confidence interval output by this invention can more comprehensively reflect the credibility of the model prediction, especially in scenarios with sparse data or limited model generalization ability, which has significant engineering guidance value.

[0064] Example 2 This invention proposes a framework for predicting the viscosity of petroleum hydrocarbons, as follows: Figure 2 As shown, this integrates efficient feature selection, uncertainty quantification based on conformal prediction, and confidence interval estimation using the Delta method. Specifically: Optimize feature filtering: The Boruta feature selection algorithm was rewritten to adapt it to the current NumPy environment, enabling efficient processing of 2D molecular descriptors extracted by RDKit and Mordred, PubChem molecular structure fingerprints, and 3D molecular descriptors calculated after conformation optimization using Gaussian-16 (B3LYP / 6-31G*). The 3D molecular descriptors are shown in Table 1. This method automatically filters features using random forest importance metrics and "shadow feature" checks, terminating when no new "pending" features are added or when a preset number of iterations is reached, extracting a core feature set highly correlated with viscosity response. Specifically, two-dimensional molecular descriptors are calculated using RDKit and Mordred cheminformatics modules, and molecular fingerprints are generated based on PubChem substructure definitions. After geometric optimization of the molecular conformation, three-dimensional molecular descriptors are calculated. The two-dimensional molecular descriptors, PubChem molecular structure fingerprints, and three-dimensional molecular descriptors are converted into NumPy array format and concatenated along the feature dimension to form a unified high-dimensional feature matrix. The Boruta algorithm is rewritten to replace native Python loop operations with NumPy's broadcast mechanism and matrix operations. In each iteration, shadow features are generated in batches using NumPy's random module and vectorized concatenation is performed. The iterative marking of confirmed significant features and irrelevant features to be removed continues until all features are judged or the preset maximum number of iterations is reached, and the filtered significant feature subset is output.

[0065] Model building: Using the screened molecular features and experimental viscosity data, four regression models—Multiple Linear Regression (MLR), Gaussian Process Regression (GPR), Random Forest (RF), and Extreme Learning Machine (ELM)—were constructed and optimized to fit the following viscosity equation. and And optimize it.

[0066]

[0067] To quantify the uncertainty of the models, error propagation analysis was used for linear regression, Gaussian process regression directly utilized its Bayesian posterior distribution, random forest combined with the Jackknife+ method evaluated the residual distribution, and extreme learning machine designed a direct interval prediction scheme based on network characteristics. The results are shown below. Figure 3 , Figure 4 and Figure 5 As shown.

[0068] Petroleum hydrocarbon viscosity mixing rules: Kendall–Monroe model: This model is proposed based on the additivity of the cube root of viscosity, and its expression is:

[0069] It is applicable to ideal or near-ideal hydrocarbon mixtures, and its prediction effect is particularly good when the molecular structure and polarity of the components are similar.

[0070] Arrhenius model: This model assumes a linear relationship between the logarithm of viscosity and composition, and its general form is:

[0071] It is widely used in low- to medium-polarity mixed systems, but its applicability to systems with strong intermolecular interactions or microphase separation is limited.

[0072] Bingham model: This model uses fluid flowability (i.e., the reciprocal of viscosity) as the summation physical quantity, and its basic form is:

[0073] It is suitable for systems dominated by flowability, such as dispersions or microemulsions, but has a larger error in high-viscosity or heterogeneous mixtures.

[0074] This invention proposes a virtual single-fluid approximation: This method treats mixtures as "virtual pure substances," constructing mixture descriptors by weighting the descriptors of each component (e.g., molar weighting, energy weighting), thus transforming mixture QSPR modeling into a modeling problem for quasi-pure substances. Its advantages lie in its ability to embed machine learning and statistical methods, its applicability to a wide range of operating conditions and complex components, and its suitability, especially for multi-component, highly non-ideal industrial systems such as petroleum hydrocarbons.

[0075] Confidence interval prediction: For machine learning models, the distribution that the predicted parameters may follow can be given:

[0076] use replace ,use replace This yields a point estimate of the logarithmic viscosity:

[0077] Therefore, it exists. Item, we cannot be sure Since it follows a distribution, we can use the Delta method to estimate its asymptotic normal approximation, let:

[0078] According to the standard asymptotic theory:

[0079] in for and Given the covariance matrices of the variables, and considering that they are independent, we have:

[0080] definition:

[0081] Its variance is approximately:

[0082] in:

[0083] Therefore, the following is given The standard deviation estimate is:

[0084] as well as The 95% confidence interval is estimated as follows:

[0085] The three modules work together to achieve high-precision, interpretable, and uncertainty-quantifiable prediction of petroleum hydrocarbon viscosity.

[0086] Compared with existing technologies, the advantages of this invention lie in improved computational efficiency and accuracy. It can not only quickly and accurately predict the viscosity of mixed petroleum hydrocarbons at different temperatures, but also reduce the cost of engineering data acquisition; the specific effects are evident. Figure 6 As shown in Table 2, the uncertainty of the prediction results can also be quantified and presented in the form of confidence intervals. This quantification of uncertainty is particularly important in engineering practice. For example, in the design of oil pipelines, confidence intervals can be considered together with oil pipeline heating and power equipment redundancy, which facilitates the design optimization of oil pipeline equipment.

[0087] Table 2 Comparison of AARD values ​​for oil viscosity prediction from different studies

[0088] Furthermore, the accurate prediction of viscosity and its uncertainty is valuable throughout the entire petroleum hydrocarbon industry chain, encompassing extraction, processing, transportation, and application. In extraction, it provides a reliable basis for optimizing reservoir development strategies; in processing and transportation, viscosity data helps guide refining process adjustments and optimize pipeline heating strategies, thereby reducing energy consumption and enhancing system reliability; and at the end-product level, such as lubricants, asphalt, and some care products, accurate viscosity assessment is fundamental to ensuring stable product performance and enhancing market competitiveness.

[0089] Therefore, constructing a viscosity model that can simultaneously output predicted values ​​and uncertainty ranges is one of the key technical details for promoting the quality and efficiency improvement and risk control of the entire petrochemical industry chain.

[0090] This invention presents a machine learning-based method for predicting the viscosity of petroleum hydrocarbons and quantifying its uncertainty. Compared to existing QSPR methods, it is the first to integrate 2D / 3D multi-source molecular descriptors and achieve efficient feature selection through an improved Boruta algorithm, avoiding the curse of dimensionality. It overcomes the limitations of traditional point prediction by integrating machine learning and a conformal prediction framework to output confidence intervals. Unlike traditional mixed rules, it maps multiple components to a single pseudo-component to directly predict viscosity. The model's robustness against interference is improved.

[0091] Example 3 like Figure 7 As shown, this invention proposes a system for predicting the viscosity of petroleum hydrocarbons and quantifying its uncertainty based on machine learning, comprising: The multi-source molecular descriptor acquisition module 100 is used to acquire multi-source molecular descriptor data of petroleum hydrocarbon mixtures. The multi-source molecular descriptors include two-dimensional molecular descriptors, three-dimensional molecular descriptors, and molecular structure fingerprint features. An improved Boruta feature selection module 200 is used to process the multi-source molecular descriptors based on the improved Boruta feature selection algorithm. It automatically selects features and iterates through random forest importance measurement and shadow feature test. When there are no undetermined features or the preset number of iterations is reached, the iteration is terminated, and the core feature set is finally determined. The virtual single-fluid model construction module 300 is used to construct a virtual single-fluid model using the core feature set, and to generate a virtual pure substance descriptor of the mixture by weighting and combining the molecular descriptors of each component of the mixture through a molar weighting or energy weighting strategy. The Delta method uncertainty quantification module 400 is used to calculate the confidence interval of the viscosity prediction value based on the virtual pure substance descriptor and the preset viscosity equation parameter distribution, and output the viscosity prediction result and its corresponding uncertainty quantification interval.

[0092] Furthermore, the Boruta feature filtering module has been improved for: Two-dimensional molecular descriptors are calculated using RDKit and Mordred cheminformatics modules, and molecular fingerprints are generated based on PubChem substructure definitions. After geometric optimization of the molecular conformation, three-dimensional molecular descriptors are calculated. The two-dimensional molecular descriptors, PubChem molecular structure fingerprints, and three-dimensional molecular descriptors are converted into NumPy array format and concatenated along the feature dimension to form a unified high-dimensional feature matrix. The Boruta algorithm is rewritten to replace the native Python loop operation with NumPy's broadcast mechanism and matrix operations. In each iteration, shadow features are generated in batches through NumPy's random module and vectorized concatenation is performed. The confirmed significant features and irrelevant features to be removed are iteratively marked until all features are judged or the preset maximum number of iterations is reached, and the filtered significant feature subset is output.

[0093] Furthermore, the system is also used to construct and optimize four regression models—multivariate linear regression, Gaussian process regression, random forest, and extreme learning machine—using the screened molecular features and experimental viscosity data, with the goal of fitting the following viscosity equation. and And optimize it:

[0094] To address the uncertainty of the model, the linear regression model employs error propagation analysis, the Gaussian process regression directly utilizes its Bayesian posterior distribution, the random forest combined with the Jackknife+ method evaluates the residual distribution, and the extreme learning machine designs a direct interval prediction scheme based on the characteristics of the network.

[0095] Furthermore, the system is also used to construct a mixture descriptor by weighted combination of the descriptors of each component, thereby transforming mixture QSPR modeling into a modeling problem of quasi-pure substances: Among them, the viscosity mixing rules for petroleum hydrocarbons are as follows: Kendall–Monroe model: The Kendall–Monroe model is proposed based on the additivity of the cube root of viscosity, and its expression is:

[0096] Arrhenius model: The Arrhenius model assumes a linear relationship between the logarithm of viscosity and composition, in the form of:

[0097] Bingham model: The Bingham model uses fluid flow as an additive physical quantity, and its form is as follows:

[0098] Confidence interval prediction: For machine learning models, the distribution that the predicted parameters may follow is given as follows:

[0099] use replace ,use replace This yields a point estimate of the logarithmic viscosity:

[0100] exist Item, cannot be determined The distribution follows a normal distribution. Using the delta method, we can estimate its asymptotic normal approximation. Let:

[0101] According to the standard asymptotic theory:

[0102] in for and Given the covariance matrices of the variables, and considering that they are independent, we have:

[0103] definition:

[0104] Its variance is approximately:

[0105] in:

[0106] Therefore, we have given The standard deviation estimate is:

[0107] as well as The 95% confidence interval is estimated as follows:

[0108] To achieve high-precision, interpretable, and uncertainty-quantified prediction of petroleum hydrocarbon viscosity.

[0109] The system for predicting the viscosity of petroleum hydrocarbons and quantifying its uncertainty based on machine learning in this invention can improve the accuracy and computational efficiency of petroleum hydrocarbon viscosity prediction. At the same time, by providing prediction confidence intervals through uncertainty quantification, it enhances the applicability and engineering reliability of the model in complex working conditions and multi-component systems.

[0110] Example 4 To implement the methods of the above embodiments, the present invention also provides a computer device, such as... Figure 8 As shown, the computer device 600 includes a memory 601 and a processor 602; wherein, the processor 602 reads executable program code stored in the memory 601 to run a program corresponding to the executable program code, so as to implement the various steps of the method described above.

[0111] Example 5 To implement the above embodiments, this application also proposes a non-transitory computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the method described in the foregoing embodiments.

[0112] The above description is merely a preferred embodiment of the present invention and is not intended to limit the invention. Various modifications and variations can be made to the present invention by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.

[0113] In the description of this specification, the references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of the present invention. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples. Moreover, without contradiction, those skilled in the art can combine and integrate the different embodiments or examples described in this specification, as well as the features of different embodiments or examples.

[0114] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one of that feature. In the description of this invention, "a plurality of" means at least two, such as two, three, etc., unless otherwise explicitly specified.

Claims

1. A method for predicting the viscosity of petroleum hydrocarbons and quantifying its uncertainty based on machine learning, characterized in that, include: S1, acquire multi-source molecular descriptor data of petroleum hydrocarbon mixture, wherein the multi-source molecular descriptor includes two-dimensional molecular descriptor, three-dimensional molecular descriptor and molecular structure fingerprint features; S2, based on the improved Boruta feature selection algorithm, processes multi-source molecular descriptors, automatically selects features and iterates through random forest importance measurement and shadow feature test, terminates the iteration when there are no undetermined features or the preset number of iterations is reached, and finally determines the core feature set; S3. A virtual single-fluid model is constructed using the core feature set. Molecular descriptors of each component of the mixture are weighted and combined using a molar weighting or energy weighting strategy to generate a virtual pure substance descriptor of the mixture. S4. Based on the virtual pure substance descriptor and the preset viscosity equation parameter distribution, the confidence interval of the viscosity prediction value is calculated using the Delta method, and the viscosity prediction result and its corresponding uncertainty quantification interval are output.

2. The method as described in claim 1, characterized in that, S2 includes: Two-dimensional molecular descriptors were calculated using RDKit and the Mordred cheminformatics module. Molecular fingerprints are generated based on PubChem substructure definitions, and three-dimensional molecular descriptors are calculated after geometric optimization of the molecular conformation. Two-dimensional molecular descriptors, PubChem molecular structure fingerprints, and three-dimensional molecular descriptors are converted into NumPy array format and concatenated along the feature dimension to form a unified high-dimensional feature matrix.

3. The method as described in claim 2, characterized in that, S2 also includes: The Boruta algorithm was rewritten specifically to replace native Python loop operations with NumPy's broadcast mechanism and matrix operations. In each iteration, shadow features are generated in batches using NumPy's random module and vectorized concatenation is performed. The iteration marks the confirmed salient features and irrelevant features to be removed until all features are judged or the preset maximum number of iterations is reached, and the filtered salient feature subset is output.

4. The method as described in claim 3, characterized in that, Using the screened molecular features and experimental viscosity data, four regression models—multiple linear regression, Gaussian process regression, random forest, and extreme learning machine—were constructed and optimized, with the goal of fitting the following viscosity equation. and And optimize it: To address the uncertainty of the model, the linear regression model employs error propagation analysis, the Gaussian process regression directly utilizes its Bayesian posterior distribution, the random forest combined with the Jackknife+ method evaluates the residual distribution, and the extreme learning machine designs a direct interval prediction scheme based on the characteristics of the network.

5. The method as described in claim 4, characterized in that, The method further includes constructing a mixture descriptor by weighting and combining the descriptors of each component, thereby transforming mixture QSPR modeling into a modeling problem of quasi-pure substances: Among them, the viscosity mixing rules for petroleum hydrocarbons are as follows: Kendall–Monroe model: The Kendall–Monroe model is proposed based on the additivity of the cube root of viscosity, and its expression is: Arrhenius model: The Arrhenius model assumes a linear relationship between the logarithm of viscosity and composition, in the form of: Bingham model: The Bingham model uses fluid flow as an additive physical quantity, and its form is as follows: Confidence interval prediction: For machine learning models, the distribution that the predicted parameters may follow is given as follows: use replace ,use replace This yields a point estimate of the logarithmic viscosity: exist Item, cannot be determined The distribution follows a normal distribution. Using the Delta method, we can estimate its asymptotic normal approximation. Let: According to the standard asymptotic theory: in for and Given the covariance matrices of the variables, and considering that they are independent, we have: definition: Its variance is approximately: in: Therefore, we have given The standard deviation estimate is: as well as The 95% confidence interval is estimated as follows: To achieve high-precision, interpretable, and uncertainty-quantified prediction of petroleum hydrocarbon viscosity.

6. A system for predicting the viscosity of petroleum hydrocarbons and quantifying its uncertainty based on machine learning, characterized in that, include: A multi-source molecular descriptor acquisition module is used to acquire multi-source molecular descriptor data of petroleum hydrocarbon mixtures. The multi-source molecular descriptors include two-dimensional molecular descriptors, three-dimensional molecular descriptors, and molecular structure fingerprint features. An improved Boruta feature selection module is used to process the multi-source molecule descriptors based on the improved Boruta feature selection algorithm. Features are automatically selected and iterated through random forest importance measurement and shadow feature test. The iteration is terminated when there are no undetermined features or the preset number of iterations is reached, and the core feature set is finally determined. The virtual single-fluid model construction module is used to construct a virtual single-fluid model using the core feature set. It uses a molar weighting or energy weighting strategy to weight and combine the molecular descriptors of each component of the mixture to generate a virtual pure substance descriptor of the mixture. The Delta method uncertainty quantification module is used to calculate the confidence interval of the viscosity prediction value based on the virtual pure substance descriptor and the preset viscosity equation parameter distribution, and output the viscosity prediction result and its corresponding uncertainty quantification interval.

7. The system as described in claim 6, characterized in that, The system is also used to construct and optimize four regression models—multivariate linear regression, Gaussian process regression, random forest, and extreme learning machine—using the screened molecular features and experimental viscosity data, with the goal of fitting the following viscosity equation. and And optimize it: To address the uncertainty of the model, the linear regression model employs error propagation analysis, the Gaussian process regression directly utilizes its Bayesian posterior distribution, the random forest combined with the Jackknife+ method evaluates the residual distribution, and the extreme learning machine designs a direct interval prediction scheme based on the characteristics of the network.

8. The system as described in claim 7, characterized in that, The system is also used to construct a mixture descriptor by weighted combination of the descriptors of each component, thereby transforming mixture QSPR modeling into a modeling problem of quasi-pure substances: Among them, the viscosity mixing rules for petroleum hydrocarbons are as follows: Kendall–Monroe model: The Kendall–Monroe model is proposed based on the additivity of the cube root of viscosity, and its expression is: Arrhenius model: The Arrhenius model assumes a linear relationship between the logarithm of viscosity and composition, in the form of: Bingham model: The Bingham model uses fluid flow as an additive physical quantity, and its form is as follows: Confidence interval prediction: For machine learning models, the distribution that the predicted parameters may follow is given as follows: use replace ,use replace This yields a point estimate of the logarithmic viscosity: exist Item, cannot be determined The distribution follows a normal distribution. Using the delta method, we can estimate its asymptotic normal approximation. Let: According to the standard asymptotic theory: in for and Given the covariance matrices of the variables, and considering that they are independent, we have: definition: Its variance is approximately: in: Therefore, we have given The standard deviation estimate is: as well as The 95% confidence interval is estimated as follows: To achieve high-precision, interpretable, and uncertainty-quantified prediction of petroleum hydrocarbon viscosity.

9. A computer device, characterized in that, Including processor and memory; The processor reads executable program code stored in the memory to run a program corresponding to the executable program code, so as to implement the method of predicting the viscosity of petroleum hydrocarbons and quantifying its uncertainty based on machine learning as described in any one of claims 1-4.

10. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When executed by a processor, the program implements the method for predicting the viscosity of petroleum hydrocarbons and quantifying its uncertainty based on machine learning as described in any one of claims 1-4.