Fault risk prediction method for oil refining process device and related device
By constructing a random forest model based on real historical and virtual data, and using molecular refining simulation and the SOL method to generate sample datasets, the problems of low accuracy and slow response speed of traditional refining process risk detection methods are solved, and efficient, accurate prediction and rapid response to refining process unit failures are achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- RICHFIT INFORMATION TECH
- Filing Date
- 2024-10-29
- Publication Date
- 2026-05-01
Smart Images

Figure CN121961189A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of oil refining process risk detection technology, and in particular to a method and related apparatus for predicting the failure risk of oil refining process equipment. Background Technology
[0002] Oil refining is a complex and critical process involving numerous physicochemical reactions. During actual oil refining production, various factors can cause malfunctions or anomalies in refining units. This not only affects the quality of petroleum products and reduces production efficiency but can also cause serious damage to the refining equipment and even trigger major safety accidents. Traditional methods for detecting risks in oil refining processes mainly rely on experience-based judgment, periodic inspections, and post-event analysis. These methods suffer from drawbacks such as low predictive accuracy, slow response speed, and inability to cover all types of failures, failing to meet the demands of modern oil refining for safe production and efficient operation. Summary of the Invention
[0003] In view of the above problems, the purpose of this invention is to provide a method and related apparatus for predicting the failure risk of oil refining process equipment.
[0004] In a first aspect, embodiments of the present invention provide a method for predicting the failure risk of an oil refining process unit, comprising:
[0005] Based on real historical data of oil refining process units and fictitious virtual data, a sample dataset is constructed by simulating the molecular refining process; the samples in the sample dataset include historical samples and virtual samples.
[0006] Based on the historical samples and virtual samples, a random forest model is constructed and trained;
[0007] The trained random forest model is used to predict the fault type and fault severity of the process unit.
[0008] In one embodiment, the historical sample is obtained by the following method:
[0009] Obtain real historical data of the process unit during operation; real historical data includes the fault type and fault severity of the process unit, as well as the values of refining process parameters under different fault severity.
[0010] Based on each fault type, the fault severity under each fault type, and the corresponding refining process parameter values for each fault severity, corresponding historical samples are generated.
[0011] By using molecular refining simulation and structure-guided lumped SOL method, extended attribute petroleum molecular structure data of each historical sample were obtained;
[0012] For each historical sample, the fault type, refining process parameters, and petroleum molecular structure data of the historical sample are identified as characteristics of the historical sample; the fault type-fault severity of the historical sample is identified as the label of the historical sample.
[0013] In one embodiment, the virtual sample is obtained by the following method:
[0014] Based on the fault types in the obtained real historical data, several fault degrees different from the real historical data are fabricated, as are several refining process parameter values different from the real historical data.
[0015] Based on each fault type, the fictitious fault severity under that fault type, and the corresponding fictitious refining process parameter values for each fault severity, corresponding virtual samples are generated.
[0016] Through molecular refining simulation and the SOL method, extended attribute petroleum molecular structure data of each virtual sample are obtained.
[0017] For each virtual sample, the fault type, refining process parameters, and petroleum molecular structure data of the virtual sample are determined as the characteristics of the virtual sample; the fault type-fault degree of the virtual sample is determined as the label of the virtual sample.
[0018] In one embodiment, the sample feature fault type is converted into a numerical feature through one-hot encoding.
[0019] In one embodiment, the feature value structure of the sample feature petroleum molecular structure data is AB; where A is the crude oil molecular structure data before refining, and B is the product molecular structure data after refining.
[0020] In one embodiment, the label value structure of the sample label is fault type i - fault degree C%; where i is the i-th fault type and C ranges from 0 to 100.
[0021] In one embodiment, the random forest model includes multiple decision trees;
[0022] The root node and internal nodes of the decision number are the test results of different pairs of sample features;
[0023] The leaf nodes of the decision number represent the decision results for the sample labels.
[0024] In one embodiment, predicting the fault type and severity of the process unit includes:
[0025] The newly generated operating data of the oil refining unit is input into multiple decision trees in the random forest model, and multiple fault types and corresponding fault severity values are output; the newly generated operating data includes the oil refining process parameters and petroleum molecular structure data newly obtained by the unit during operation.
[0026] From the multiple fault types and their corresponding fault degrees, the final fault type and fault degree value of the process unit are determined.
[0027] In one embodiment, determining the final fault type and fault severity value of the process unit from the plurality of fault types and corresponding fault severity values includes:
[0028] For multiple fault types and corresponding fault degrees obtained from multiple decision trees, a voting method is used to obtain the fault type and fault degree with the highest probability; or, the probabilities of the multiple fault types are sorted from high to low to determine several most likely fault types and corresponding fault degree values.
[0029] In one embodiment, the method further includes: periodically evaluating and updating the random forest model, including:
[0030] Obtain the actual results of the fault types and fault severity during the operation of the process unit;
[0031] The actual results are compared with the fault type-fault severity with the highest probability obtained by the voting method to determine the evaluation factors precision, recall and F1 score, where the F1 score is the harmonic mean of the precision and the recall.
[0032] New samples generated during the operation of the process unit are added to the sample set, and the hyperparameters of the random forest model are adjusted to continuously improve the value of the evaluation factor.
[0033] Secondly, embodiments of the present invention provide an apparatus related to a method for predicting the failure risk of an oil refining process unit, comprising:
[0034] The sample dataset construction module is used to construct a sample dataset by simulating a molecular refining process based on real historical data of the refining process unit and fictitious virtual data; the samples in the sample dataset include historical samples and virtual samples.
[0035] The random forest model building module is used to build and train a random forest model based on the historical samples and virtual samples.
[0036] The prediction module is used to predict the fault type and fault severity of the process unit using the trained random forest model.
[0037] Thirdly, embodiments of the present invention provide a computing device, including: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the aforementioned method for predicting the failure risk of an oil refining process unit.
[0038] Fourthly, embodiments of the present invention provide a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the aforementioned method for predicting the failure risk of an oil refining process unit.
[0039] Fifthly, embodiments of the present invention provide a computer program product, including: a computer program, which, when executed by a processor, implements the aforementioned method for predicting the failure risk of an oil refining process unit.
[0040] The beneficial effects of the above-described technical solutions provided in the embodiments of the present invention include at least the following:
[0041] This invention provides a method and related apparatus for predicting the failure risk of oil refining process units. Based on real historical data of the oil refining process unit, this invention constructs samples by simulating the molecular refining process, enabling accurate identification of petroleum characteristics at the molecular level. These samples are then used to construct and train a random forest model, improving the accuracy of the random forest model. The trained random forest model can then perform prior predictions of the failure status of the oil refining process unit. Compared to past methods such as experience-based judgment and post-event inspection and analysis, its prediction process offers higher safety and efficiency. The model prediction can also determine the specific failure type and severity of the oil refining process unit, thus providing engineers with more practical failure status information and facilitating the rapid location and response to potential failure risks in the oil refining process unit.
[0042] Furthermore, in the process of using the random forest model to predict oil refining process units, the random forest model can be continuously updated and optimized, thereby continuously improving the accuracy of the prediction.
[0043] The technical solution of the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. Attached Figure Description
[0044] The accompanying drawings are provided to further illustrate the invention and form part of the specification. They are used in conjunction with embodiments of the invention to explain the invention and do not constitute a limitation thereof. In the drawings:
[0045] Figure 1 This is a flowchart of the oil refining process unit failure risk prediction method in an embodiment of the present invention;
[0046] Figure 2 This is a technical flowchart for predicting the failure risk of refining process equipment in an embodiment of the present invention;
[0047] Figure 3 These are the various groups represented in the SOL representation in the embodiments of this invention;
[0048] Figure 4 This is a schematic diagram illustrating the construction and prediction of the random forest model in an embodiment of the present invention;
[0049] Figure 5 This is a schematic diagram of the structure of the relevant device in the oil refining process unit failure risk prediction method in the embodiments of the present invention. Detailed Implementation
[0050] This invention provides a method and related apparatus for predicting the failure risk of an oil refining process unit. Although exemplary embodiments of this disclosure are shown in the accompanying drawings, it should be understood that this disclosure can be implemented in various forms and should not be limited to the embodiments set forth herein. Rather, these embodiments are provided to enable a more thorough understanding of this disclosure and to fully convey the scope of this disclosure to those skilled in the art.
[0051] This invention provides a method for predicting the failure risk of an oil refining process unit, referring to... Figure 1 As shown, the method includes the following steps:
[0052] S11. Based on real historical data of the oil refining process unit and fictitious virtual data, a sample dataset is constructed by simulating the molecular oil refining process; the samples in the sample dataset include historical samples and virtual samples.
[0053] S12. Based on the historical samples and virtual samples, construct and train a random forest model.
[0054] S13. Using the trained random forest model, predict the fault type and fault severity of the process unit.
[0055] This invention provides a method for predicting the failure risk of oil refining process units. It uses a constructed random forest model to predict the types and severity of potential failures that may occur in oil refining process units during operation. The prediction method includes three steps: First, establishing a sample set for constructing the random forest model; second, constructing and training the random forest model; and third, using the trained random forest model to make prior predictions about the types and severity of potential failures in the process unit over a future period. Compared to traditional methods such as experience-based judgment and post-event inspection and analysis, the prediction method provided by this invention not only has the advantage of prior prediction, improving the safety, energy efficiency, and effectiveness of the prediction process, but also can identify and predict specific failure types and their corresponding severity, enabling engineers to quickly formulate reasonable measures to handle and intervene in process unit failures.
[0056] Step S11 involves constructing a sample set, which can be derived from historical samples and virtual samples.
[0057] Historical samples were obtained through the following method: Common failure types of refining process units were collected and organized to form a failure type list. Taking catalytic cracking units as an example, failure types include: catalyst deactivation, temperature sensor malfunction, heat exchanger leakage, pipeline blockage, etc. Simultaneously, an associated attribute—failure severity—was assigned to each failure type to measure the degree of danger or harm of the refining process unit failure. The failure severity value was preset to a range of 0-100%, with a higher percentage indicating a more severe failure of the corresponding refining process unit.
[0058] Based on the actual operating conditions of refining process units in the refinery and combined with expert experience, different fault severity values are assigned to different operating conditions of the refining process units. A fault severity value of 0 is assigned to a refining process unit in a healthy state. In addition, the values of actual refining process parameters corresponding to different fault types and severity levels can be collected. Taking a catalytic cracking unit as an example, its refining process parameters include reaction temperature, regeneration temperature, operating pressure, regeneration pressure, catalyst type, catalyst activity, feed flow rate, hydrogen flow rate, and residence time, etc.
[0059] The above is the actual historical data obtained during the operation of the oil refining process unit, including: the type and severity of the oil refining process unit failures, and the values of the oil refining process parameters under different failure severity levels. (Refer to...) Figure 2As shown, the real historical data includes fault status data (fault severity value greater than 0) and health status data (fault severity value 0). Based on this real historical data, various historical samples are generated; each historical sample contains attributes such as fault type, fault severity, and corresponding refining process parameter values.
[0060] Molecular refining simulation is a controllable refining method at the petroleum molecule level. The structure-guided lumped-array (SOL) method represents petroleum molecules based on specific structural features or functional groups, referencing... Figure 3 As shown, these groups are organized into a vector, where the elements of the vector represent different petroleum molecule structures, and the number of elements indicates the number of specific structural groups in the molecule. The SOL method can provide detailed information on the petroleum molecule structure during the refining process.
[0061] Molecular refining simulation is used to initialize the process before simulation. Different simulation parameters are then set to correspond to different fault types and degrees in the refining unit, thereby simulating various fault conditions in the refining unit. (Refer to...) Figure 2 As shown, extended attribute petroleum molecular structure data for each historical sample were extracted using molecular refining simulation and the structure-guided lumped SOL method. This includes the molecular structure data of crude oil before refining and the molecular structure data of the product after refining. Both the crude oil and product molecular structure data are 24-dimensional vectors. For example, the molecular structure data of a certain crude oil is illustrated in Table 1.
[0062] Table 1
[0063] quality score A6 A4 ... KO Ni V 2.36E-07 0 0 ... 1 0 0 0.011892688 0 0 ... 0 0 0 0.002997709 1 0 ... 0 0 0 ... ... ... ... ... ... ...
[0064] For each historical sample, the original attribute fault type, refining process parameters, and simulated extended attribute petroleum molecular structure data are identified as the characteristics of the historical sample. The attribute fault type is combined with the fault severity, i.e., fault type-fault severity, to determine the label of the historical sample.
[0065] In one embodiment, historical samples are obtained through the following method:
[0066] Obtain real historical data of the process unit during operation; real historical data includes the type and severity of the process unit failure, as well as the values of refining process parameters under different failure severity levels;
[0067] Based on each fault type, the degree of each fault under each fault type, and the corresponding values of refining process parameters for each fault degree, corresponding historical samples are generated.
[0068] By using molecular refining simulation and structure-guided lumped SOL method, extended attribute petroleum molecular structure data of each historical sample were obtained;
[0069] For each historical sample, the failure type, refining process parameters, and petroleum molecular structure data of the historical sample are identified as characteristics of the historical sample; the failure type-failure severity of the historical sample is identified as the label of the historical sample.
[0070] To obtain a random forest model with better learning performance, it is necessary to provide a sufficient number of training samples. To compensate for the possibility of insufficient historical samples, the inventors constructed virtual samples as a supplement to the sample dataset based on the acquired historical data.
[0071] In one embodiment, the virtual sample is obtained by the following method:
[0072] Based on the fault types in the real historical data, several fault degrees and several refining process parameter values that are different from the real historical data are fabricated.
[0073] Based on each fault type, the fictitious fault severity under each fault type, and the corresponding fictitious refining process parameter values for each fault severity, corresponding virtual samples are generated.
[0074] Extended attribute petroleum molecular structure data of each virtual sample were obtained through molecular refining simulation and SOL method.
[0075] For each virtual sample, the fault type, refining process parameters, and petroleum molecular structure data of the virtual sample are identified as the characteristics of the virtual sample; the fault type-fault degree of the virtual sample is identified as the label of the virtual sample.
[0076] The sample dataset constructed using the above method includes historical samples and virtual samples. Both historical and virtual samples are constructed based on the obtained historical data, with virtual samples supplementing the dataset when historical samples are insufficient. Using molecular refining simulation and the SOL method, new extended attributes—petroleum molecular structure data—can be accurately extracted from historical and virtual samples, including crude oil molecular structure data before refining and product molecular structure data after refining. This method of identifying petroleum characteristics at the molecular level enables precise localization of potential failure risks in refining processes and significantly improves the accuracy of subsequent random forest model construction.
[0077] Step S12 involves constructing and training a random forest model.
[0078] Reference Figure 2 As shown, before building and training a random forest model, the samples need to be preprocessed, including data cleaning and unit unification, in order to improve the quality and reliability of the samples and increase the accuracy of modeling.
[0079] The sample feature, fault type, is an unordered feature. The inventors used one-hot encoding to convert it into a numerical feature suitable for training random forest models. For example, a new binary feature (i.e., a "one-hot feature") is created for each fault type, where only one feature is activated at any given time and marked as 1, while all other features are marked as 0. When there are n fault types, n one-hot features will be generated accordingly. For example, if there are only 3 fault types: catalyst deactivation, heat exchanger leakage, and pipe blockage, then n=3, and the results after one-hot encoding are shown in Table 2 below.
[0080] Table 2
[0081]
[0082] In one embodiment, the sample feature fault type of the input sample of the random forest model is converted into a numerical feature through one-hot encoding.
[0083] In one embodiment, the feature structure of the sample feature petroleum molecular structure data of the input sample of the random forest model is AB, where A is the crude oil molecular structure data before refining and B is the product molecular structure data after refining.
[0084] In one embodiment, the label value structure of the sample labels of the samples output by the random forest model is fault type i - fault degree C%; where i is the i-th fault type and C ranges from 0 to 100.
[0085] A random forest model consists of multiple decision trees. For a single decision tree, samples are first drawn from historical and virtual samples to form multiple first subsets. Then, a decision tree is constructed using each first subset. Each decision tree includes a root node, internal nodes, and leaf nodes. The root node is the starting node, and the leaf nodes are the ending nodes. Internal nodes are all nodes except the root and leaf nodes. During the construction of the decision tree, a subset of features is randomly selected for each node to split, until a leaf node is reached. When constructing a single decision tree, the root node includes the entire first subset, containing all samples and features within it, as shown below. Figure 4As shown, the first subset includes "Sample 1", "Sample 2", ..., "Sample n". The decision tree selects a feature for splitting. Each node, except for the leaf nodes, corresponds to a test result for a feature. The leaf nodes correspond to the decision results. The quality of different feature splits is evaluated using criteria such as the Gini index or information gain. The first subset is continuously split into multiple second subsets based on the selected feature. After splitting, the second subsets form the next layer of nodes in the decision tree, and the above process is repeated recursively. When the number of samples in a node reaches a preset minimum number (e.g., 10), or the purity of a node (e.g., the Gini index) reaches a preset standard, the splitting stops, and that node becomes a leaf node. Following this method, a preset number of decision trees are gradually formed, resulting in the random forest model. When constructing multiple decision trees, each decision tree randomly selects sample features for testing and splitting during the construction process, ensuring that each decision tree is different. This randomness gives each decision tree a unique structure, thereby improving the generalization ability of the random forest model. To ensure a relative balance in the number of samples, a random oversampling method is used to increase the number of a few scarce samples when training each decision tree, thereby improving the stability and accuracy of model training.
[0086] In one embodiment, the root node and internal nodes of the decision tree represent the test results of different sample features.
[0087] In one embodiment, the leaf nodes of the decision tree represent the decision results for the sample labels.
[0088] The random forest model constructed and trained in this embodiment of the invention fully utilizes the advantages of random forest models in handling high-dimensional data, resisting overfitting, and parallel computing. At the same time, the sample data is preprocessed before model training, such as one-hot encoding, to ensure the consistency of the sample data and the effectiveness of the constructed random forest model.
[0089] Step S13 involves using the trained random forest model to make prior predictions about the fault types and severity of the oil refining process unit.
[0090] Reference Figure 4As shown, newly acquired operating data of refining process units are input into a random forest model to predict the failure status of the process units. For example, for a catalytic cracking unit in a refinery, periodic predictions are made on a monthly basis. Using the new operating data generated by the unit in the previous month, at the beginning of the current month, the possible failure types and severity of the unit in the current month are predicted. Multiple decision trees in the random forest model will produce multiple different failure type-failure severity predictions. The predictions from all decision trees are then aggregated to form the final prediction result. For example, a voting method can be used to obtain the most likely failure type-failure severity from the prediction results of each decision tree, which is then determined as the final failure type-failure severity. Alternatively, the output failure type probabilities can be sorted from high to low to determine several high-probability failure types and their corresponding severity. For example, after sorting the failure types by probability from high to low, the top ten most likely failure types and their corresponding severity values can be determined. Figure 4 The top ten results are respectively called the best classification, the second best classification, and so on. (Refer to...) Figure 2 As shown, the refinery can promptly issue early warnings of fault risks based on the prediction results of the random forest model and formulate corresponding handling and intervention measures.
[0091] In one embodiment, predicting the type and severity of process unit failures includes:
[0092] The newly generated operational data of the refining process unit is input into multiple decision trees in the random forest model, and the output is multiple fault types and corresponding fault severity values; the newly generated operational data includes newly acquired refining process parameters and petroleum molecular structure data during the operation of the process unit.
[0093] The final fault type and fault severity value of the process unit are determined from multiple fault types and their corresponding fault severity values.
[0094] In one embodiment, determining the final fault type and fault severity value of the process unit from multiple fault types and corresponding fault severity values includes:
[0095] For multiple fault types and corresponding fault degrees obtained from multiple decision trees, a voting method is used to obtain the fault type and fault degree with the highest probability; or, the probabilities of multiple fault types are sorted from high to low to determine several most likely fault types and corresponding fault degree values.
[0096] In the prediction process using the random forest model, new operational data is generated to evaluate and update the model. For example, for a refining process unit in an oil refinery, the failure risk of the unit is predicted on a monthly basis. The newly generated operational data from the previous month is input into the random forest model to obtain the predicted results for this month (the most probable failure type and its severity). Next, the actual results of the refining process unit's operation during this month (i.e., the actual failure type and severity) are obtained, and the predicted results are compared with the actual results. For example, evaluation factors such as accuracy, recall, and F1 score are used to evaluate the model.
[0097] Accuracy refers to the proportion of samples correctly predicted by the model out of the total number of samples. Recall refers to the proportion of samples correctly predicted as positive out of the actual number of positive samples. The F1 score is the harmonic mean of accuracy and recall. Higher values for these three evaluation factors indicate better model performance. The expressions for the evaluation factors are shown below:
[0098]
[0099] in:
[0100] TP (TruePositives): The predicted value is 1, and the actual value is also 1;
[0101] TN (TrueNegatives): The predicted value is 0, and the actual value is also 0;
[0102] FP (FalsePositives): The predicted value is 1, and the true value is 0;
[0103] FN(FalseNegatives): The predicted value is 0, and the true value is 1.
[0104] After evaluating the random forest model, new operating data from the refining unit from the previous month and this month are processed to obtain new samples. These new samples are then added to the sample set, and the hyperparameters of the random forest model are adjusted accordingly, such as the number of decision trees, the maximum depth of a single decision tree, and the minimum number of splits. This approach allows the constructed random forest model to continuously improve the values of evaluation factors during prediction, enabling ongoing optimization of the random forest model and enhancing its accuracy in predicting refining unit failure risks.
[0105] In one embodiment, the random forest model is periodically evaluated and updated, including:
[0106] Obtain the actual results of the fault types and fault severity during the operation of the process unit;
[0107] The actual results are compared with the fault type-fault severity with the highest probability obtained by the voting method to determine the evaluation factors precision, recall and F1 score, where the F1 score is the harmonic mean of the precision and the recall.
[0108] New samples generated during the operation of the process unit are added to the sample set, and the hyperparameters of the random forest model are adjusted to continuously improve the value of the evaluation factor.
[0109] The oil refining process unit failure risk prediction method provided in this invention has several advantages: The training samples constructed using molecular refining simulation and the SOL method accurately obtain extended attribute petroleum molecular structure data, achieving precise localization of oil refining process unit failure risks at the molecular level. Furthermore, by expanding upon historical samples, new hypothetical samples are generated, whose sample characteristics and failure severity comprehensively cover the failure types and severity of all existing process units. Additionally, the constructed random forest model can be continuously updated and optimized during use.
[0110] Based on the same inventive concept, this invention also provides a related device for the method of predicting the failure risk of oil refining process equipment. Since the principle of the problem solved by this device is similar to that of the aforementioned method of predicting the failure risk of oil refining process equipment, the implementation of this device can refer to the implementation of the aforementioned method, and the repeated parts will not be described again.
[0111] This invention provides an apparatus related to a method for predicting the failure risk of an oil refining process unit, referring to... Figure 5 As shown, it includes:
[0112] The sample dataset construction module 51 is used to construct a sample dataset by simulating a molecular refining process based on real historical data of the refining process unit and fictitious virtual data; the samples in the sample dataset include historical samples and virtual samples.
[0113] Random forest model building module 52 is used to build and train a random forest model based on the historical samples and virtual samples.
[0114] The prediction module 53 is used to predict the fault type and fault severity of the process unit using the trained random forest model.
[0115] This invention provides a computing device, including: a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the program, it implements the aforementioned method for predicting the failure risk of an oil refining process unit.
[0116] This invention provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the aforementioned method for predicting the failure risk of an oil refining process unit.
[0117] This invention provides a computer program product, including: a computer program that, when executed by a processor, implements the aforementioned method for predicting the failure risk of an oil refining process unit.
[0118] Obviously, those skilled in the art can make various modifications to this invention without departing from its spirit and scope. Therefore, if these modifications fall within the scope of the claims and their equivalents, this invention is also intended to include these modifications.
Claims
1. A method for predicting failure risks of oil refining process units, characterized in that, include: Based on real historical data of oil refining process equipment and fictitious virtual data, a sample dataset is constructed by simulating the molecular refining process; the samples in the sample dataset include historical samples and virtual samples. Based on the historical samples and virtual samples, a random forest model is constructed and trained; The trained random forest model is used to predict the fault type and fault severity of the process unit.
2. The prediction method as described in claim 1, characterized in that, The historical samples were obtained through the following method: Obtain real historical data of the process unit during operation; real historical data includes the fault type and fault severity of the process unit, as well as the values of refining process parameters under different fault severity. Based on each fault type, the fault severity under each fault type, and the corresponding refining process parameter values for each fault severity, corresponding historical samples are generated. By using molecular refining simulation and structure-guided lumped SOL method, extended attribute petroleum molecular structure data of each historical sample were obtained; For each historical sample, the failure type, refining process parameters, and petroleum molecular structure data of the historical sample are identified as characteristics of the historical sample. The fault type and fault severity of the historical samples are determined as the labels of the historical samples.
3. The prediction method as described in claim 2, characterized in that, The virtual sample is obtained through the following method: Based on the fault types in the obtained real historical data, several fault degrees different from the real historical data are fabricated, as are several refining process parameter values different from the real historical data. Based on each fault type, the fictitious fault severity under the fault type, and the corresponding fictitious refining process parameter values for each fault severity, corresponding virtual samples are generated. Through molecular refining simulation and the SOL method, extended attribute petroleum molecular structure data of each virtual sample are obtained. For each virtual sample, the fault type, refining process parameters, and petroleum molecular structure data of the virtual sample are determined as the characteristics of the virtual sample; the fault type-fault degree of the virtual sample is determined as the label of the virtual sample.
4. The prediction method as described in claim 3, characterized in that, The sample feature fault type is converted into a numerical feature through one-hot encoding.
5. The prediction method as described in claim 4, characterized in that, The eigenvalue structure of the sample characteristic petroleum molecular structure data is AB; where A is the crude oil molecular structure data before refining and B is the product molecular structure data after refining.
6. The prediction method as described in claim 5, characterized in that, The label value structure of the sample label is fault type i - fault degree C%; where i is the i-th fault type and C ranges from 0 to 100.
7. The prediction method according to any one of claims 1-6, characterized in that, The random forest model includes multiple decision trees; The root node and internal nodes of the decision number are the test results of different pairs of sample features; The leaf nodes of the decision number represent the decision results for the sample labels.
8. The prediction method as described in claim 7, characterized in that, The prediction of the fault type and fault severity of the process unit includes: The newly generated operating data of the oil refining unit is input into multiple decision trees in the random forest model, and multiple fault types and corresponding fault severity values are output; the newly generated operating data includes the newly acquired oil refining process parameters and petroleum molecular structure data during the operation of the unit. From the multiple fault types and their corresponding fault degrees, the final fault type and fault degree value of the process unit are determined.
9. The prediction method as described in claim 8, characterized in that, Determining the final fault type and fault severity value of the process unit from the plurality of fault types and corresponding fault severity values includes: For multiple fault types and corresponding fault degrees obtained from multiple decision trees, a voting method is used to obtain the fault type and fault degree with the highest probability; or, the probabilities of the multiple fault types are sorted from high to low to determine several most likely fault types and corresponding fault degree values.
10. The prediction method as described in claim 9, characterized in that, Also includes: The random forest model is periodically evaluated and updated, including: Obtain the actual results of the fault types and fault severity during the operation of the process unit; The actual results are compared with the fault type-fault severity with the highest probability obtained by the voting method to determine the evaluation factor precision, recall and F1 score, where the F1 score is the harmonic mean of the precision and the recall. New samples generated during the operation of the process unit are added to the sample set, and the hyperparameters of the random forest model are adjusted to continuously improve the value of the evaluation factor.
11. A related apparatus for a method of predicting failure risks in an oil refining process unit, characterized in that, include: The sample dataset construction module is used to construct a sample dataset by simulating a molecular refining process based on real historical data of the refining process unit and fictitious virtual data; the samples in the sample dataset include historical samples and virtual samples. The random forest model building module is used to build and train a random forest model based on the historical samples and virtual samples. The prediction module is used to predict the fault type and fault severity of the process unit using the trained random forest model.
12. A computing device, characterized in that, include: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor, when executing the program, implements the method for predicting the failure risk of an oil refining process unit as described in any one of claims 1-10.
13. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, implements the method for predicting the failure risk of an oil refining process unit as described in any one of claims 1-10.
14. A computer program product, characterized in that, The computer program product includes a computer program that, when executed by a processor, implements the oil refining process unit failure risk prediction method according to any one of claims 1-10.