Prediction method and device for crude oil cutting distillate oil yield, medium and product
By acquiring near-infrared spectral data of crude oil, identifying key spectral variables and building a proprietary prediction model, the problem of inaccurate prediction of the yield of mixed crude oil cut fractions in traditional methods was solved, achieving fast and accurate prediction results, and improving the efficiency and economic benefits of the refining process.
Patent Information
- Application Number
- CN202510835295.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-20
- Publication Date
- 2025-10-03
AI Technical Summary
Traditional methods are difficult to accurately predict the yield of cut fractions from mixed crude oil or unconventional crude oil, resulting in a large deviation between the predicted results and the actual results, which affects the process control and economic benefits of the refining process.
By obtaining the near-infrared spectral data of the crude oil to be predicted, determining the key spectral variables, and inputting them into the target prediction model, a dedicated prediction model for different types of crude oil is constructed, and a machine learning algorithm is used to achieve fast and accurate prediction of the cut distillate oil yield.
It improves the prediction accuracy and efficiency of cut fraction oil yield, reduces the interference of heterogeneity between crude oil categories on prediction, adapts to the changes in the characteristics of different crude oils, and supports the rapid evaluation and optimization of the refining process.
Smart Images

Figure CN120741401A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of petrochemical technology, and in particular to a method, device, medium and product for predicting the yield of crude oil cut distillate oil. Background Art
[0002] As international crude oil prices fluctuate more and more, refineries are widely using multi-source blended crude oil to optimize costs. However, the components of blended crude oil are complex and their physical properties vary significantly, resulting in their distillation characteristics being highly nonlinear and difficult to predict.
[0003] Traditional fraction yield predictions rely on interpolation calculations or empirical formulas. However, when the measured data points are sparsely distributed or the crude oil properties change, interpolation algorithms often have difficulty accurately reflecting the actual distillation behavior of crude oil, resulting in large deviations between the simulation results and the actual cutting results. This is especially true when dealing with mixed crude oils or unconventional crude oils. Traditional methods are limited in their ability to express the complex physical properties of crude oil, and have obvious uncertainties and insufficient applicability, resulting in low prediction accuracy for unconventional crude oils.
[0004] Therefore, there is an urgent need for a method for predicting the yield of cut distillate oil to improve the prediction efficiency and accuracy of the yield of crude oil cut distillate oil. Summary of the Invention
[0005] The present invention provides a method, a device and a storage medium for predicting the yield of crude oil cut distillate oil, which are used to improve the prediction efficiency and accuracy of the crude oil cut distillate oil yield.
[0006] In a first aspect, the present application provides a method for predicting the yield of crude oil cut fractions, the method comprising:
[0007] Acquiring near-infrared spectral data of the crude oil to be predicted, and determining corresponding key spectral variables based on the near-infrared spectral data; the near-infrared spectral data characterizes the spectral absorption characteristics of the corresponding crude oil within a specific wavelength band; the key spectral variables characterize near-infrared spectral data whose correlation with the cut fraction yield exceeds a preset threshold;
[0008] The key spectral variables are input into a target prediction model to obtain the cut fraction oil yield output by the target prediction model; the target prediction model corresponds to the target crude oil category of the crude oil to be predicted, and the target crude oil category is determined based on the key spectral variables.
[0009] Optionally, the target prediction model is trained in the following manner:
[0010] Obtaining training data sets corresponding to a plurality of crude oil samples, each training data set including near-infrared spectral data and cut fraction oil yield of the corresponding crude oil sample;
[0011] Based on a preset stepwise linear regression strategy, feature selection is performed on the near-infrared spectral data to obtain corresponding key spectral variables;
[0012] Based on a preset clustering strategy and in combination with key spectral variables, the plurality of crude oil samples are classified into categories to determine a plurality of crude oil categories;
[0013] Based on the training data set corresponding to each crude oil category, the prediction model of the corresponding category is trained, and when the training is completed, the prediction model of each crude oil category is output.
[0014] Optionally, the feature selection of the near-infrared spectral data based on a preset stepwise linear regression strategy to obtain corresponding key spectral variables includes:
[0015] Perform wavelength band division based on the near-infrared spectrum data to determine a target set and a candidate set; the target set includes two initial wavelength bands, and the candidate set includes other wavelength bands in the near-infrared spectrum data except the two initial wavelength bands;
[0016] Based on a preset significance threshold, the wavelength bands of the target set and the candidate set are updated, and the updated target set is determined as the key spectral variable; the significance level of each wavelength band in the updated target set is not less than the significance threshold, and the significance level represents the degree of correlation between the near-infrared spectral data of a specific wavelength band and the distillate oil yield.
[0017] Optionally, updating the wavelength bands of the target set and the candidate set based on a preset significance threshold includes:
[0018] Based on the significance threshold, screening each wavelength band in the candidate set and adding a first wavelength band to the target set; the significance level of the first wavelength band is greater than the significance threshold;
[0019] Based on the significance threshold, the wavelength bands in the target set are screened, and a second wavelength band is removed from the target set; the significance level of the second wavelength band is less than the significance threshold.
[0020] Optionally, the method of classifying the plurality of crude oil samples based on a preset clustering strategy and combining key spectral variables to determine a plurality of crude oil categories includes:
[0021] Based on a preset number of crude oil categories, a plurality of central parameters are determined from the key spectral variables; each central parameter corresponds to a crude oil category;
[0022] For each key spectral variable, the category assignment process is iteratively performed until the crude oil category of each key spectral variable no longer changes. Each category assignment process includes:
[0023] Based on the central parameters of this processing, the relative distance between each key spectral variable and each central parameter is determined;
[0024] Based on the center parameter corresponding to the minimum relative distance, the crude oil category of the corresponding key spectral variable is determined;
[0025] Based on the parameter average value of each crude oil category, the central parameter of each crude oil category is updated; the parameter average value represents the weighted average value of each key spectral variable of the corresponding crude oil category.
[0026] Optionally, obtaining a training data set corresponding to each of the plurality of crude oil samples includes:
[0027] Performing near-infrared spectral scanning on the plurality of crude oil samples to obtain original near-infrared spectral data of each of the plurality of crude oil samples;
[0028] Based on a preset smoothing strategy, the raw near-infrared spectral data is smoothed;
[0029] The smoothed original infrared spectrum data is subjected to denoising and scattering correction processing to obtain preprocessed infrared spectrum data.
[0030] Optionally, the step of training the prediction model of the corresponding category based on the training data set corresponding to each crude oil category includes:
[0031] Using the key spectral variables as model input, and adjusting parameters of the prediction model based on a preset loss function and back-propagation strategy;
[0032] When the loss function meets the preset termination condition, a trained prediction model is obtained; the preset termination condition is that the loss value is not greater than the preset loss threshold, or the number of iterations is greater than the preset number threshold.
[0033] In a third aspect, the present application provides a computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, the method for predicting the yield of crude oil fractions described in any one of the first aspects above is implemented.
[0034] In a fourth aspect, the present application provides a computer storage medium, wherein the computer-readable storage medium stores computer program instructions, and the computer program instructions are executed by a processor to implement any one of the crude oil cutting distillate oil yield prediction methods described in the first aspect.
[0035] In a fifth aspect, an embodiment of the present application provides a computer program product, comprising computer program instructions, which, when executed by a processor, implement any one of the methods for predicting the yield of crude oil fractions described in the first aspect.
[0036] The beneficial effects of the present invention are as follows:
[0037] The present application provides a method for predicting the cutoff distillate yield of crude oil. This method obtains near-infrared spectral data of the crude oil to be predicted, determines corresponding key spectral variables, and inputs these key spectral variables into a target prediction model corresponding to the crude oil category to be predicted, thereby obtaining the cutoff distillate yield as output by the model. Thus, by using key spectral variables that are highly correlated with the cutoff distillate yield, the interference of other parameters on prediction accuracy is reduced. Furthermore, by constructing dedicated prediction models for different crude oil categories, the interference of heterogeneity between crude oil categories on prediction accuracy is reduced, and the prediction model's adaptability to changes in the characteristics of different crude oils is improved. This allows the prediction model to quickly and accurately predict the cutoff distillate yield of the crude oil, thereby improving both the accuracy and efficiency of cutoff distillate yield predictions. BRIEF DESCRIPTION OF THE DRAWINGS
[0038] In order to more clearly illustrate the technical solutions in the embodiments of the present application or related technologies, the following briefly introduces the drawings required for use in the embodiments or related technical descriptions. Obviously, the drawings described below are only embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without paying any creative work.
[0039] Figure 1 A schematic diagram of a training process of a distillate oil yield prediction model provided in an embodiment of the present application;
[0040] Figure 2 Schematic diagram of near-infrared spectroscopic images of various crude oils provided in the embodiments of this application;
[0041] Figure 3 A schematic diagram of the relationship between crude oil temperature and yield provided in an embodiment of the present application;
[0042] Figure 4 A simplified neural network diagram provided in an embodiment of the present application;
[0043] Figure 5 A schematic flow chart of a method for predicting the yield of crude oil cut fractions provided in an embodiment of the present application;
[0044] Figure 6 A schematic diagram of the structure of a computer device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0045] In order to make the purpose, technical solutions and advantages of the present application clearer, the technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of this application. Unless there is a conflict, the embodiments in the present application and the features in the embodiments can be combined with each other in any way. In addition, although a logical order is shown in the flowchart, in some cases, the steps shown or described can be performed in an order different from that here.
[0046] The terms "first" and "second" in the specification and claims of this application and the above-mentioned drawings are used to distinguish different objects, rather than to describe a specific order. In addition, the term "comprising" and any of its variations are intended to cover non-exclusive protection. For example, a process, method, system, product or device that includes a series of steps or units is not limited to the listed steps or units, but optionally also includes steps or units that are not listed, or optionally also includes other steps or units inherent to these processes, methods, products or devices. "Multiple" in this application can mean at least two, for example, two, three or more, and the embodiments of this application are not limited thereto.
[0047] The term "and / or" in the embodiments of this application is simply a description of the association relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent the following three situations: A exists alone, A and B exist at the same time, and B exists alone. In addition, the character " / " in this document generally indicates that the related objects are in an "or" relationship.
[0048] It is understood that in the following specific implementation methods of this application, crude oil data and other related data are involved. When the various embodiments of this application are applied to specific products or technologies, relevant licenses or consents need to be obtained, and the collection, use and processing of relevant data need to comply with the relevant laws, regulations and standards of relevant countries and regions. For example, it is possible to recruit relevant volunteers and sign relevant agreements for volunteer authorization data, and then use the data of these volunteers for implementation; or, by implementing within the scope of an authorized organization, data management is carried out by implementing the following implementation methods using the data of internal members of the organization; or, the relevant data used in the specific implementation are all simulated data, such as simulated data generated in a virtual scene.
[0049] The embodiments of the present application relate to artificial intelligence and machine learning (ML) technology, and are mainly designed based on machine learning in artificial intelligence.
[0050] Artificial intelligence (AI) refers to the theories, methods, techniques, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, to perceive the environment, acquire knowledge, and use that knowledge to achieve optimal results. In other words, AI is a comprehensive technology within computer science that seeks to understand the essence of intelligence and produce new intelligent machines that can respond in a manner similar to human intelligence. AI also involves studying the design principles and implementation methods of various intelligent machines, enabling them to possess the capabilities of perception, reasoning, and decision-making.
[0051] Artificial intelligence (AI) technology is a comprehensive discipline encompassing a wide range of fields, encompassing both hardware and software technologies. Foundational AI technologies generally include sensors, specialized AI chips, cloud computing, distributed storage, big data processing, operating / interaction systems, and mechatronics. AI software technologies primarily encompass computer vision, speech processing, natural language processing, and machine learning / deep learning.
[0052] Machine learning is a multidisciplinary field that encompasses probability theory, statistics, approximation theory, convex analysis, and algorithmic complexity theory. It specifically studies how computers can simulate or implement human learning behaviors to acquire new knowledge or skills and reorganize existing knowledge structures to continuously improve their performance. Machine learning is at the core of artificial intelligence and the fundamental way to make computers intelligent. Its applications span all areas of artificial intelligence. Machine learning and deep learning typically include techniques such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and self-learning.
[0053] Machine learning is at the core of artificial intelligence and the fundamental way to make computers intelligent. Its applications span all areas of AI. Machine learning and deep learning typically include techniques such as artificial neural networks, belief networks, reinforcement learning, transfer learning, and inductive learning. Artificial neural networks (ANNs) abstract the neural networks in the human brain from an information processing perspective, building a simple model that forms different networks based on different connection structures. A neural network is a computational model composed of a large number of interconnected nodes (or neurons). Each node represents a specific output function, called an activation function. Each connection between two nodes represents a weighted value for the signal passing through that connection, called a weight. This serves as the memory of the artificial neural network. The network's output varies depending on the network's connection structure, weights, and activation function. The network itself is often an approximation of a natural algorithm or function, or it may express a logical strategy.
[0054] The following is a brief introduction to the design concept of the embodiments of this application.
[0055] With the continued volatility of international crude oil prices, refineries are increasingly blending imported crude oil from multiple sources to reduce procurement costs. Due to the complex composition and significant differences in physical properties of blended crude oils, their distillation characteristics are highly uncertain. Traditional empirical methods struggle to accurately reflect their actual cutting behavior, making it difficult to accurately predict the distillate yields of crude oil cuts. This directly impacts process control of distillate yields, and consequently, the product mix, energy efficiency, and economic performance of the entire refinery.
[0056] Currently, rapid prediction methods for narrow fraction yields of mixed crude oils mostly rely on interpolation calculations or empirical formulas. However, when the measured data points are sparsely distributed or the crude oil properties change, such methods often have difficulty accurately reflecting the actual distillation behavior of the crude oil, resulting in large deviations between the simulation results and the actual cutting results. Especially when faced with mixed crude oils or unconventional crude oils, traditional methods are limited in their ability to express the complex physical properties of crude oils, and have obvious uncertainties and insufficient applicability.
[0057] In view of the above problems, the present invention provides a method for predicting the yield of crude oil cut fractions. The method obtains near-infrared spectral data of the crude oil to be predicted, determines the corresponding key spectral variables, and inputs the key spectral variables into a target prediction model corresponding to the crude oil category to be predicted, thereby obtaining the cut fraction yield output by the model. In this way, the method can quickly obtain crude oil information through spectroscopy and accurately predict the crude oil cut fraction yield by combining it with a machine learning algorithm, without the need for actual distillation experiments. The method has the advantages of fast prediction speed and high accuracy. In addition, by using key spectral variables with a high degree of correlation with the cut fraction yield, the interference of other parameters on the prediction accuracy is reduced. Dedicated prediction models are constructed for different crude oil categories, reducing the interference of heterogeneity between crude oil categories on prediction accuracy. The predictive model's adaptability to changes in different crude oil characteristics is further improved, enabling the predictive model to quickly and accurately predict the cut fraction yield of the crude oil, thereby improving the prediction accuracy and efficiency of the cut fraction yield. This method provides effective technical support for rapid crude oil evaluation and refining process optimization, and provides scientific decision support and production optimization basis for the refining process.
[0058] The following briefly introduces the application scenarios to which the technical solutions of the embodiments of the present application can be applied. It should be noted that the application scenarios described below are only used to illustrate the embodiments of the present application and are not limiting. In the specific implementation process, the technical solutions provided by the embodiments of the present application can be flexibly applied according to actual needs.
[0059] The solutions provided in the embodiments of this application can be applied to most oil refining production scenarios and used to accurately obtain the distillate yield of crude oil. For example, in the crude oil distillation process of an oil refinery, the distillate yield of different crude oils can be quickly predicted based on the prediction method of the cut distillate yield provided by this application. In particular, in the case of mixed processing of multiple crude oils, the distillate yield distribution of the mixed crude oil can be accurately predicted based on the prediction method of the cut distillate yield provided by this application, providing data support for the optimization of the batching scheme, thereby improving the processing efficiency and finished oil yield of the mixed crude oil by rationally allocating the ratios of various crude oils, such as light and heavy.
[0060] Of course, the method provided in the embodiment of the present application is not limited to the above application scenarios, and can also be used in other possible application scenarios, which are not limited by the embodiment of the present application. The functions that can be implemented by each device in the above application scenarios will be described in the subsequent method embodiments, and will not be described in detail here.
[0061] Below, in combination with the application scenarios described above, the method provided by the exemplary embodiment of the present application is described with reference to the accompanying drawings. It should be noted that the above application scenarios are only shown to facilitate understanding of the spirit and principles of the present application, and the implementation methods of the present application are not limited in this respect.
[0062] In the embodiments of the present application, a prediction model can be used to implement the above-mentioned prediction process of the cut distillate oil yield. Before the prediction model is put into use, it is necessary to pre-train the model so that the prediction model converges. Next, in order to facilitate the description of the model application process, the training process of the prediction model is first introduced.
[0063] Please refer to Figure 1 , is a schematic diagram of a training process of a prediction model provided in an embodiment of the present application. The specific implementation process of the training process is as follows:
[0064] Step 101: Obtain training data sets corresponding to multiple crude oil samples.
[0065] In the embodiment of the present application, the training data set of each crude oil sample includes the near-infrared spectral data of the crude oil sample and the cut distillate oil yield. Among them, the near-infrared spectral data represents the spectral absorption characteristics of the crude oil in a specific band, and the key spectral variables represent the near-infrared spectral data whose correlation with the cut distillate oil yield exceeds a preset threshold. The near-infrared spectral analysis technology of the present application, with its advantages of being fast, non-destructive, and requiring no pre-processing, can provide new ideas for the fine modeling of crude oil fraction distribution, enabling the present application to obtain full spectrum information of crude oil within seconds, reflecting the absorption characteristics of different molecular structures in crude oil, and identifying trace component differences through spectral modeling, and is suitable for crude oil samples of different types and sources. In this way, the prediction model constructed by the present application based on near-infrared spectroscopy not only greatly improves the analysis efficiency, but also can achieve high-precision estimation of crude oil cutting characteristics and distillate yield without destroying the sample. It provides effective technical support for rapid evaluation of crude oil and optimization of refining process, and provides scientific decision-making support and production optimization basis for the refining process.
[0066] Specifically, refer to Figure 2 Schematic diagram of near-infrared spectra of various crude oils provided in the examples of this application, wherein the near-infrared spectral data has wavelength points in the range of 4000-12000nm-1, reflecting the spectral absorption characteristics of crude oil in this band, and the cut distillate oil yield data includes the above Figure 2 The crude oil cut distillate yield data with a cut interval of 5°C is shown in the figure. Figure 2 The crude oil distillate yield curve is obtained by combining these 5°C yield data. Figure 3 This is a schematic diagram of the relationship between crude oil temperature and yield provided in the examples of this application. Figure 3The broken line in the figure reflects the core law of the temperature-yield cumulative curve during crude oil distillation. For example, in the 0-400°C range, the yield shows a slow upward trend, indicating that low-boiling-point components such as naphtha and kerosene are mainly precipitated in this stage. In the 400-500°C turning zone, the yield growth rate increases significantly, corresponding to the concentrated precipitation of C20-C40 hydrocarbon components in crude oil. When the temperature exceeds 500°C, the curve approaches the 100% platform, indicating that the remaining high-boiling-point components are difficult to further separate under conventional distillation conditions. This curve verifies the modeling ability of the multi-layer perceptron neural network in this application for nonlinear distillation characteristics, and can accurately predict the yield distribution of different fraction cut points in mixed crude oil, providing a visual decision-making basis for the optimization of temperature parameters in the refining process.
[0067] In one possible implementation, the embodiment of the present application can perform near-infrared spectral scanning on multiple crude oil samples to obtain original near-infrared spectral data of each of the multiple crude oil samples, and smooth the original near-infrared spectral data through a smoothing strategy, and perform denoising and scattering correction on the smoothed original infrared spectral data to achieve preprocessing of the spectral data.
[0068] Specifically, the present application can use the Savitzky-Golay (SG) smoothing filter method to smooth the original near-infrared spectral data to reduce random noise and smooth the spectral curve, thereby improving the signal-to-noise ratio. The Standard Normal Variate (SNV) method is used to remove and eliminate baseline drift and systematic differences between spectra to improve the comparability of the data. The Multiplicative Scatter Correction (MSC) method is used to correct the baseline drift and amplitude changes caused by optical path differences in the spectral data.
[0069] In one possible implementation, the present invention can use an SG smoothing method to smooth the raw spectral data. SG smoothing reduces random noise by applying polynomial fitting to the spectral data, thereby smoothing the spectral curve and improving the signal-to-noise ratio.
[0070] Specifically, the SG smoothing formula in the embodiment of the present application is as follows:
[0071]
[0072] in, is the smoothed data point, a j is the weight coefficient fitted by the least squares method, L is the half width of the smoothing window, y i+j is the data point with offset j in the spectral data.
[0073] Next, the present application can use the SNV method to perform denoising on the smoothed spectral data. The SNV method eliminates baseline drift and systematic differences between spectra by standardizing each spectral data point so that the mean of the spectral data is 0 and the standard deviation is 1.
[0074] Specifically, the standardized data in this application is as follows:
[0075]
[0076] in, is the standardized data, x i is the data of each sample point, u i is the mean of each sample, σ i is the standard deviation.
[0077] Finally, the present application can apply the MSC method to correct for baseline drift and amplitude variation caused by optical pathlength differences in spectral data. MSC corrects for scattering effects by comparing the sample spectrum with the standard spectrum, thereby reducing errors caused by scattering.
[0078] Specifically, the calculation formula of MSC in this application is as follows:
[0079]
[0080] in, is the corrected data, x i is the data of each sample point, is the mean of the standard spectrum, and k is the standard deviation of the standard spectrum.
[0081] Step 102: Based on a preset stepwise linear regression strategy, feature selection is performed on the near-infrared spectral data to obtain corresponding key spectral variables.
[0082] In the embodiment of the present application, after obtaining the spectral data of each crude oil sample, a stepwise linear regression strategy is used to perform feature selection on the spectral data to determine key spectral variables, thereby eliminating irrelevant wavelengths and enhancing the robustness of the constructed model.
[0083] In a possible embodiment, the present application embodiment can adopt the feature selection method of stepwise Bayesian linear regression (SBLR) to perform feature selection of spectral data. First, the present application can divide the wavelength band according to the wavelength of the near-infrared spectral data to determine the target set and the candidate set. Among them, the target set includes two randomly selected initial wavelength bands, and the other wavelength bands in the near-infrared spectral data except these two initial wavelength bands are placed in the candidate set. And according to the preset significance threshold, each wavelength band in the candidate set is screened to add the first wavelength band with a significance level greater than the threshold to the target set, then each wavelength band in the target set is screened and the second wavelength band with a significance level less than the threshold is removed from the target set, and finally the target set containing the key spectral variables is screened. In this way, selection is made in all the spectral data by stepwise linear regression to screen out the spectral variables that are highly correlated with the distillate yield and have a significant contribution to the prediction effect, and the spectral data of other parts are eliminated.
[0084] Specifically, the present application can randomly divide the near-infrared spectral data into several wavelength bands of equal length, which are divided into a "target set" and a "candidate set." The target set initially contains two wavelength bands, which are randomly selected, and the remaining wavelength bands are placed in the candidate set. Next, the threshold Fin, corresponding to the significance level α of the F test, is set to 0.95 to determine the significance of a wavelength band. If the F value of a candidate wavelength band is greater than the threshold Fin, the wavelength band is considered to have a significant impact on the model and is added to the target set; otherwise, it is removed. A stepwise Bayesian linear regression algorithm is then used to calculate the F values of all wavelength bands in the candidate set. The wavelength band with the largest F value is selected. If the F value of the wavelength band is greater than the threshold Fin, it is transferred from the candidate set to the target set, thereby introducing a new wavelength band that significantly contributes to the model. After the transfer, the F values of the remaining wavelength bands in the candidate set are recalculated, and new wavelength bands are selected based on the threshold Fin to be introduced into the target set. Next, the SBLR algorithm is used to recalculate the F values of the wavelength bands in the target set. All wavelength bands in the target set are examined, especially those with the smallest F values. If the F value of the wavelength band is less than the threshold of 0.95, it will be removed to ensure that the target set only contains the wavelength bands that are most useful for model prediction. In this way, the application can use the F value to represent whether the improvement in the model's fitting ability after the introduction of the new variable is significantly greater than the complexity introduced by it. Each time a wavelength band is added to the candidate set, F is recalculated and compared with the critical value Fin. If it is greater than the critical value, it is added to the wavelength band of the target set; otherwise, it is removed.
[0085] Specifically, the calculation of the F value in this application is as follows:
[0086]
[0087] Where RSS0 is the residual sum of squares in the original model without the candidate wavelength band, RSS1 is the residual sum of squares in the original model with the wavelength band added, p1 is the number of parameters in the new model including the wavelength band, p0 is the number of parameters in the old model, and n is the number of samples.
[0088] In this way, the above steps are repeated until there are no more wavelength bands that need to be removed in the target set.
[0089] Step 103: Based on a preset clustering strategy and in combination with key spectral variables, the multiple crude oil samples are classified into categories to determine multiple crude oil categories.
[0090] In the embodiment of the present application, after determining the key spectral variables of each crude oil sample, each crude oil sample can be divided into crude oil categories according to a preset clustering strategy, thereby determining multiple crude oil categories, laying the foundation for the subsequent establishment and training of prediction models for each crude oil category.
[0091] Specifically, after obtaining the key spectral variables after feature selection processing, the present application can use the K-means (K-Means, KMeans) clustering algorithm to establish a clustering model for the important properties of each crude oil sample, and cluster multiple crude oil samples to divide the crude oil samples into at least two crude oil categories, so that the crude oil data of different categories can be modeled separately during modeling to avoid errors caused by differences between oil products.
[0092] In one possible implementation, the embodiment of the present application can determine multiple center parameters from the key spectral variables based on a preset number of crude oil categories, and each center parameter corresponds one-to-one to a crude oil category. For each key spectral variable, the category assignment process is iteratively performed until the crude oil category of each key spectral variable no longer changes. Each category assignment process includes: determining the relative distance between each key spectral variable and each center parameter based on each center parameter processed this time; determining the crude oil category of the corresponding key spectral variable based on the center parameter corresponding to the minimum relative distance; and updating the center parameter of each crude oil category based on the average parameter value of each crude oil category, wherein the average parameter value represents the weighted average value of each key spectral variable of the corresponding crude oil category.
[0093] Specifically, the number of crude oil types is preset to K. For example, based on the different distillate yields of different types of crude oil in various temperature ranges, it is planned to divide crude oil into three categories, and K is set to 3. From the key property parameters, K parameters are randomly selected as the initial cluster centers of the crude oil types, i.e., the center parameters. The Euclidean distance from each key property parameter to each cluster center is calculated. The crude oil type with the smallest Euclidean distance is the crude oil type to which the key property parameter should belong. Each key property parameter is thus assigned to the crude oil type to which the nearest center point belongs, thereby updating the cluster center of each crude oil type. In this way, the latest center point of each crude oil type is repeatedly calculated, and the parameters are reallocated, gradually reducing the distance between each parameter and the center of the type to which it belongs, until the type assignment of each key property parameter no longer changes, thereby improving the modeling accuracy of the subsequent artificial neural network.
[0094] In one possible implementation, the embodiment of the present application can cluster crude oil into three categories based on the data differences in the distillate oil yields of different crude oils in the high distillate segment, the middle distillate segment, and the low distillate segment, and calculate the center point of each category and the Euclidean distance of each data point to each cluster center respectively, so as to assign each data point to the category to which the nearest center point belongs, and update the cluster center of each cluster until the category assignment no longer changes.
[0095] The formula for updating the cluster center is as follows:
[0096]
[0097] Among them, u j is the jth cluster center, C j is the jth cluster, |C j | is the number of data points in the jth cluster, x i is the i-th data point in the j-th cluster.
[0098] Every time you update, use u j It is equal to the weighted average of all parameters in crude oil type Cj, so that the updated uj can better represent the overall characteristics of crude oil type j.
[0099] Step 104: Based on the training data set corresponding to each crude oil category, the prediction model of the corresponding category is trained, and after the training is completed, the prediction model of each crude oil category is output.
[0100] In the present embodiment, after screening the key spectral variables of crude oil and classifying it into different crude oil categories, the present invention constructs a corresponding prediction model for each crude oil category and performs model training, thereby obtaining trained prediction models corresponding to multiple crude oil categories. By constructing and training prediction models based on crude oil categories, the predictive model's adaptability to different crude oil characteristics is enhanced, preventing the impact of differences between crude oil categories on model fitting.
[0101] In one possible implementation, the present application may use key spectral variables as model inputs and adjust the parameters of the prediction model for each crude oil type based on a preset loss function and backpropagation strategy, thereby obtaining a trained prediction model when the loss function meets a preset termination condition. The preset termination condition is that the loss value is no greater than a preset loss threshold, or the number of iterations exceeds a preset threshold.
[0102] Specifically, this application can use the clustered spectral data of various crude oil categories for modeling, use the key spectral data corresponding to these crude oil categories as model input, and use the mean squared error (MSE) as the loss function. The learning rate is adjusted through the backpropagation algorithm to optimize the weight and bias terms to minimize the loss function and improve the performance of the model. At the same time, multiple hyperparameter lists are set, including the number of neurons in the hidden layer, the structure of the hidden layer, the type of activation function, the weight optimizer, the regularization parameter and the learning rate. These hyperparameters are then tuned through the grid search method to obtain the optimal parameter combination. The final neural network model maps the spectral data to 102 output nodes, each of which represents the distillate oil yield in different temperature ranges.
[0103] Specifically, refer to Figure 4 The figure shows a simplified neural network diagram provided by an embodiment of the present application. Figure 4 The neuron structure shown includes an input layer, two hidden layers, and an output layer, where ∈R n Indicates that the input has n input features and the hidden layer ∈R n Indicates that the number of neurons in this part is n, and R in the output layer 102 The output is a 102-dimensional output vector. The input layer corresponds to the near-infrared spectral data after feature selection. The two hidden layers realize high-dimensional nonlinear transformation through the fully connected structure. The R 102 The dimension accurately corresponds to the predicted value of the distillate oil yield at 102 cutting temperature points, and its linear mapping relationship (the last 102 nodes) is optimized through the MSE loss function, ultimately achieving end-to-end modeling from spectral features to the yield distribution over the entire distillation range.
[0104] In one possible implementation, after clustering, the present application can establish a corresponding multilayer perceptron neural network prediction model for each classified crude oil sample based on its corresponding key spectral variables. The multilayer perceptron regression model established for each category takes as input the spectral data corresponding to the classified crude oil category and outputs the corresponding fraction yields. The model structure can consist of an input layer, several hidden layers, and an output layer. The neurons in each layer introduce nonlinearity through activation functions, enabling the network to learn complex mapping relationships.
[0105] Specifically, the structure and steps of establishing a multilayer perceptron (MLP) in this application are as follows:
[0106]
[0107] Among them, y k is the output of the kth neuron in the output layer, f k is the activation function, w ki is the weight connecting the input and output layers, x i is the i-th feature of the input layer, b k Is the bias term of the output layer. At this time, choose to set the weight w ki and the bias term b k Random sampling from Gaussian distribution is used so that each neuron in the model can learn different features.
[0108] Furthermore, the prediction model in this application can use mean square error as the loss function. The loss function measures the squared difference between the model output and the true value, and calculates the gradient through the backpropagation algorithm to update the optimized weights and bias terms to reduce the loss:
[0109]
[0110] Where n is the number of samples, Y i is the true value, is the predicted value of the model, w ki To optimize the weight, b k is the bias term, and η is the learning rate.
[0111] In one possible implementation, the embodiment of the present application can optimize model parameters by adjusting the hyperparameters of various parts of the model.
[0112] Specifically, the present application may set multiple hyperparameter lists, including the number of hidden layer neurons, the structure of the hidden layer, the activation function type, the weight optimizer, the regularization parameter, and the learning rate, with multiple candidate values for each hyperparameter. Different combinations of these hyperparameters are evaluated using a grid search method to select the optimal hyperparameter combination.
[0113] After obtaining the trained prediction models corresponding to the respective crude oil categories, the embodiment of the present application can use the prediction models to predict the cut distillate oil yield of the crude oil to be split, so as to obtain the cut distillate oil yield output by the prediction model.
[0114] refer to Figure 5 FIG. 1 is a flow chart of a method for predicting the yield of crude oil fractions provided in an embodiment of the present application. The specific implementation process of the method is as follows:
[0115] Step 501: Acquire near-infrared spectral data of the crude oil to be predicted, and determine corresponding key spectral variables based on the near-infrared spectral data.
[0116] In the embodiments of the present application, near-infrared spectral data characterizes the spectral absorption characteristics of the corresponding crude oil within a specific wavelength band, and key spectral variables characterize near-infrared spectral data whose correlation with the distillate oil yield exceeds a preset threshold. Through the aforementioned model training process, the present application meticulously screens and classifies the near-infrared spectral data of each type of crude oil, thereby obtaining multiple key spectral variables that are highly correlated with distillate oil yield. This allows the present embodiment of the application, when applying the model, to screen the near-infrared spectral data of the crude oil to be predicted and determine the key spectral variables of the crude oil to be predicted.
[0117] Step 502: Input the key spectral variables into the target prediction model to obtain the cut distillate oil yield output by the target prediction model.
[0118] In the embodiment of the present application, the target prediction model corresponds to the target crude oil category of the crude oil to be predicted, and the target crude oil category is determined based on the key spectral variables.
[0119] Specifically, the present embodiment acquires near-infrared spectral data of a new oil product, performs spectral preprocessing and feature selection on it, and then inputs it into a trained kmeans clustering model to determine its crude oil classification. After determining the new oil product's classification, the spectral data is then input into a corresponding trained multi-layer perceptron model to predict the yield of its cut distillate oil.
[0120] In one possible implementation, the present application example analyzed 444 types of crude oil used for testing in the database, and the mean square error of the final cutting yield was 0.263, and the average R-square of the regression model reached 0.89, meeting the production needs of the enterprise.
[0121] It is worth mentioning that in the embodiment of the present application, the prediction process of the distillate oil yield during the training process is the same as the prediction process during the actual application process. Therefore, the process can refer to the detailed introduction of the aforementioned training process and will not be repeated here.
[0122] See Figure 6 As shown, based on the same technical concept, an embodiment of the present application further provides a computer device 60 , which may include a memory 601 and a processor 602 .
[0123] The so-called memory 601 is used to store computer programs executed by the processor 602. The memory 601 may mainly include a program storage area and a data storage area, wherein the program storage area may store an operating system, an application required for at least one function, etc.; the data storage area may store data created according to the use of the computer device, etc. The processor 602 may be a central processing unit (CPU), or a digital processing unit, etc. The specific connection medium between the above-mentioned memory 601 and the processor 602 is not limited in the embodiment of the present application. The embodiment of the present application is Figure 6 In the embodiment, the memory 601 and the processor 602 are connected via a bus 603. The bus 603 is connected to the processor 602 via a bus 603. Figure 6 The connections between the other components are shown in bold lines, which are only for illustration and are not intended to be limiting. The so-called bus 603 can be divided into an address bus, a data bus, a control bus, etc. For ease of illustration, Figure 6 Only one thick line is used in the diagram, but this does not mean that there is only one bus or one type of bus.
[0124] Memory 601 may be a volatile memory, such as random-access memory (RAM); a non-volatile memory, such as read-only memory, flash memory, a hard disk drive (HDD), or a solid-state drive (SSD); or any other medium capable of carrying or storing desired program code in the form of instructions or data structures and accessible by a computer, but is not limited thereto. Memory 601 may be a combination of the above memories.
[0125] The processor 602 is configured to call the computer program stored in the so-called memory 601 to execute the crude oil cutting distillate oil yield prediction method executed by the device in each embodiment of the present application.
[0126] In some possible embodiments, various aspects of the method for predicting the yield of crude oil cutting distillates provided in the present application can also be implemented in the form of a program product, which includes program code. When the program product is run on a computer device, the program code is used to enable the computer device to execute the steps of the method for predicting the yield of crude oil cutting distillates according to various exemplary embodiments of the present application described above in this specification. For example, the computer device can execute the steps of each embodiment.
[0127] The program product may employ any combination of one or more readable media. The readable medium may be a readable signal medium or a readable storage medium. The readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or component, or any combination thereof. More specific examples of readable storage media (a non-exhaustive list) include: an electrical connection with one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof.
[0128] The program product of the embodiment of the present application may be a portable compact disc read-only memory (CD-ROM) and include program code, and can be run on a computing device. However, the program product of the present application is not limited thereto. In the present application, a readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with a command execution system, device, or apparatus.
[0129] A readable signal medium may include a data signal transmitted in baseband or as part of a carrier wave, which carries readable program code. Such a transmitted data signal may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A readable signal medium may also be any readable medium other than a readable storage medium that can transmit, propagate, or transfer a program for use by or in conjunction with a command execution system, apparatus, or device.
[0130] The program code embodied on the readable medium may be transmitted using any appropriate medium, including but not limited to wireless, wireline, optical fiber cable, RF, etc., or any suitable combination of the foregoing.
[0131] The program code for performing the operations of the present application can be written in any combination of one or more programming languages, including object-oriented programming languages such as Java, C++, etc., and conventional procedural programming languages such as "C" or similar programming languages. The program code can be executed entirely on the user computing device, partially on the user device, as a stand-alone software package, partially on the user computing device and partially on a remote computing device, or entirely on a remote computing device or server. In the case of a remote computing device, the remote computing device can be connected to the user computing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computing device (e.g., via the Internet using an Internet service provider).
[0132] It should be noted that although several units or subunits of the device are mentioned in the detailed description above, this division is merely exemplary and not mandatory. In fact, depending on the embodiment of the application, the features and functions of two or more units described above can be embodied in a single unit. Conversely, the features and functions of a single unit described above can be further divided and embodied by multiple units.
[0133] Furthermore, although the operations of the method of the present application are described in a particular order in the accompanying drawings, this does not require or imply that the operations must be performed in this particular order, or that all illustrated operations must be performed to achieve the desired results. Additionally or alternatively, some steps may be omitted, multiple steps may be combined into one step, and / or one step may be decomposed into multiple steps.
[0134] Those skilled in the art will appreciate that the embodiments of the present application can be provided as methods, systems, or computer program products. Therefore, the present application can adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment in combination with software and hardware. Moreover, the present application can adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.
[0135] Although the preferred embodiments of the present application have been described, those skilled in the art may make additional changes and modifications to these embodiments once they have learned the basic creative concept. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications that fall within the scope of the present application.
[0136] Obviously, those skilled in the art may make various changes and modifications to this application without departing from the spirit and scope of this application. Thus, if these modifications and variations of this application fall within the scope of the claims of this application and their equivalents, this application is intended to include these modifications and variations.
Claims
1. A method for predicting the yield of crude oil cut fractions, characterized in that: The method comprises: Acquiring near-infrared spectral data of the crude oil to be predicted, and determining corresponding key spectral variables based on the near-infrared spectral data; the near-infrared spectral data characterizes the spectral absorption characteristics of the corresponding crude oil within a specific wavelength band; the key spectral variables characterize near-infrared spectral data whose correlation with the cut fraction yield exceeds a preset threshold; The key spectral variables are input into a target prediction model to obtain the cut fraction oil yield output by the target prediction model; the target prediction model corresponds to the target crude oil category of the crude oil to be predicted, and the target crude oil category is determined based on the key spectral variables.
2. The method according to claim 1, wherein The target prediction model is trained in the following way: Obtaining training data sets corresponding to a plurality of crude oil samples, each training data set including near-infrared spectral data and cut fraction oil yield of the corresponding crude oil sample; Based on a preset stepwise linear regression strategy, feature selection is performed on the near-infrared spectral data to obtain corresponding key spectral variables; Based on a preset clustering strategy and in combination with key spectral variables, the plurality of crude oil samples are classified into categories to determine a plurality of crude oil categories; Based on the training data set corresponding to each crude oil category, the prediction model of the corresponding category is trained, and when the training is completed, the prediction model of each crude oil category is output.
3. The method according to claim 2, wherein The method of performing feature selection on the near-infrared spectral data based on a preset stepwise linear regression strategy to obtain corresponding key spectral variables includes: Perform wavelength band division based on the near-infrared spectrum data to determine a target set and a candidate set; the target set includes two initial wavelength bands, and the candidate set includes other wavelength bands in the near-infrared spectrum data except the two initial wavelength bands; Based on a preset significance threshold, the wavelength bands of the target set and the candidate set are updated, and the updated target set is determined as the key spectral variable; the significance level of each wavelength band in the updated target set is not less than the significance threshold, and the significance level represents the degree of correlation between the near-infrared spectral data of a specific wavelength band and the distillate oil yield.
4. The method according to claim 3, wherein The updating of the wavelength bands of the target set and the candidate set based on a preset significance threshold comprises: Based on the significance threshold, screening each wavelength band in the candidate set and adding a first wavelength band to the target set; the significance level of the first wavelength band is greater than the significance threshold; Based on the significance threshold, the wavelength bands in the target set are screened, and a second wavelength band is removed from the target set; the significance level of the second wavelength band is less than the significance threshold.
5. The method according to claim 2, wherein The method of classifying the plurality of crude oil samples based on a preset clustering strategy and combining key spectral variables to determine a plurality of crude oil categories includes: Based on a preset number of crude oil categories, a plurality of central parameters are determined from the key spectral variables; each central parameter corresponds to a crude oil category; For each key spectral variable, the category assignment process is iteratively performed until the crude oil category of each key spectral variable no longer changes. Each category assignment process includes: Based on the central parameters of this processing, the relative distance between each key spectral variable and each central parameter is determined; Based on the center parameter corresponding to the minimum relative distance, the crude oil category of the corresponding key spectral variable is determined; Based on the parameter average value of each crude oil category, the central parameter of each crude oil category is updated; the parameter average value represents the weighted average value of each key spectral variable of the corresponding crude oil category.
6. The method according to claim 2, wherein The step of obtaining a training data set corresponding to each of the plurality of crude oil samples comprises: Performing near-infrared spectral scanning on the plurality of crude oil samples to obtain original near-infrared spectral data of each of the plurality of crude oil samples; Based on a preset smoothing strategy, the raw near-infrared spectral data is smoothed; The smoothed original infrared spectrum data is subjected to denoising and scattering correction processing to obtain preprocessed infrared spectrum data.
7. The method according to claim 2, wherein The method of training the prediction model of the corresponding category based on the training data set corresponding to each crude oil category includes: Using the key spectral variables as model input, and adjusting parameters of the prediction model based on a preset loss function and back-propagation strategy; When the loss function meets the preset termination condition, a trained prediction model is obtained; the preset termination condition is that the loss value is not greater than the preset loss threshold, or the number of iterations is greater than the preset number threshold.
8. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 7 are implemented.
9. A computer storage medium having computer program instructions stored thereon, characterized in that: When the computer program instructions are executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.
10. A computer program product comprising computer program instructions, characterized in that When the computer program instructions are executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.
Citation Information
Cited By
Pipe network sludge component detection method, device, medium and equipment
CN121558669A
Crude oil property prediction method based on distillate oil depth characteristic modeling
CN121583376A
Oil substance quality detection method and system based on artificial intelligence and spectrum detection
CN122259506A