METHOD FOR PREDICTING THE PERFORMANCE OF A CATALYTIC CRACKING UNIT (FCC)
A predictive method using phenomenological modeling and machine learning optimizes FCC unit performance by addressing the complexity of catalyst interactions and deactivation, ensuring precise yield predictions and real-time adjustments.
Patent Information
- Authority / Receiving Office
- BR · BR
- Patent Type
- Applications
- Current Assignee / Owner
- PETROLEO BRASILEIRO SA PETROBRAS
- Filing Date
- 2024-12-27
- Publication Date
- 2026-07-07
AI Technical Summary
Existing methods struggle to accurately predict the performance of fluid catalytic cracking (FCC) units due to the complex interaction between operational variables and catalyst properties, which are subject to frequent changes, and the impact of catalyst deactivation and contaminants on yield efficiency and product quality.
A predictive method combining phenomenological modeling with machine learning, using a Multilayer Perceptron (MLP) neural network, pre-trained with simulated data and refined with real historical data, to optimize catalyst formulation and adjust operations in real-time based on market demands and operating conditions.
The method provides accurate predictions and enables real-time adjustments, enhancing the efficiency and adaptability of the FCC process to meet varying market demands and operational conditions.
Smart Images

Figure 00000000_0000_ABST
Description
1 / 27 METHOD FOR PREDICTING THE PERFORMANCE OF A CATALYTIC CRACKING UNIT (FCC) FIELD OF THE INVENTION
[01] The present invention falls within the technical field of chemical engineering and process technology, specifically in the monitoring and optimization of fluid catalytic cracking (FCC) units in oil refineries.
[02] The invention relates to a method for predicting the yield and performance of catalysts in FCC units, aiming to optimize the conversion of hydrocarbons into higher value products such as LPG (Liquefied Petroleum Gas), propene, cracked naphtha, and other derivatives.
[03] The invention is especially useful for refineries seeking greater efficiency and precision in catalytic cracking operations, meeting varying market demands, as well as aiming to maximize catalyst performance over time. FUNDAMENTALS OF THE INVENTION
[04] The objective of the invention is to solve the limitations associated with predicting performance in FCC units, where there is a complex interaction between operational variables and catalyst properties, which are subject to frequent changes. These changes make it challenging to predict yield and adjust operation in real time.
[05] Another problem is the need to consider catalyst deactivation and the cumulative effects of contaminants, which directly impact yield efficiency and the quality of the final products.
[06] In this sense, the present invention aims to offer a predictive model that can capture the complexity of the catalytic cracking process, considering both the operating conditions and the characteristics of the catalyst in use, and that allows dynamic adjustments according to market demand. Petition 870250058707, dated 10 / 07 / 2025, page 5 / 64 2 / 27 and the current state of the catalyst, combining physical modeling with machine learning. STATE OF THE ART
[07] Document US11676061B2 describes a method that uses a real-time predictive model to control the sulfur level in gasoline produced by a fluid catalytic cracking (FCC) unit. The proposal is to adjust the temperature of the feed pretreatment reactor of the FCC unit, based on the sulfur levels measured in the final product (gasoline). The system uses machine learning and real-time data analysis to predict the sulfur level of gasoline based on the characteristics of the feed before treatment. In contrast, the present invention deals with a method for modeling the physicochemical properties of catalysts with the aim of optimizing the yields of a catalytic cracking (FCC) unit, acting on the formulation of the catalyst to be loaded into the FCC reactor.Furthermore, the automation system described in document US11676061B2 is not implemented in a process simulator architecture that acts on various FCC process variables to meet the production demands of its products. Additionally, the automation system described in the document is detailed to use a generic model not specified in the patent.
[08] Document US10838412B2 describes systems and methods for improving the accuracy of olefin yield predictions in a cracking process (steam reformer - stream cracker) using a hybrid approach that combines machine learning techniques with phenomenological models. The phenomenological model, which is less computationally demanding, provides an initial prediction of olefin yield from a hydrocarbon stream, while machine learning corrects this prediction based on historical errors. An Experimental Design is used to generate training data. Petition 870250058707, dated 10 / 07 / 2025, p. 6 / 64 3 / 27 (DOE), in which both the phenomenological and detailed kinetic radical models are run under different conditions to identify discrepancies in the results. In contrast, the present invention relates to a method for modeling the physicochemical properties of catalysts with the aim of optimizing the yields of a catalytic cracking (FCC) unit, acting on the formulation of the catalyst to be loaded into the FCC reactor. The learning model used in document US10838412B2 proposes to use the results of two predictions from independent reaction models: a phenomenological model and a kinetic free radical model. The difference in yield prediction between the models is determined and used to train a machine learning model. The machine learning model is then applied in real time, allowing for the correction of the predicted yield.In other words, the result of the machine training model repositions the unit's yields, but does not modify the previous reaction models.
[09] Document US11987758B2 describes a method involving the development and implementation of a predictive model for the catalytic performance of waste FCC units, applicable to any type of catalyst, including virgin and flushing catalysts (Purchased Equilibrium Catalyst). This model has been integrated into process simulators such as PETRO-SIM™, FCC-SIM™, and HYSYS. The objective is to improve the adherence of process simulators to the real behavior of FCC units, specifically with regard to variations in the quality and quantity of flushing catalysts. By modifying existing simulators, the model ensures that these variations, which were previously not accurately predicted, can now be represented more faithfully, optimizing the operation of the units. In this sense, document US11987758B2 deals with the allocation of Petition 870250058707, dated 10 / 07 / 2025, page 7 / 64 4 / 27 virgin catalysts and flushing without considering their formulation or physicochemical properties, in order to optimize the allocation of flushing for a set of FCC and RFCC units simultaneously, as proposed by the present invention. The mixture of virgin catalyst with the flushing catalysts was carried out for any proportion between them, and the quality of the flushing used was varied for application in waste FCC (RFCC) units. The differences between the catalysts were determined by catalytic tests in a pilot plant and the formulation or physicochemical properties of the catalysts were not used to estimate their model, nor was any form of machine learning modeling used, which would require costly experimental resources that are not necessary in the present invention.
[010] Document US11669063B2 refers to a method for modeling a chemical production process using simulations and machine learning models. The described computational system receives simulation event logs, which contain inputs and outputs of a simulated chemical process. These logs represent the mass flow and are limited by the physical characteristics of the equipment in a real chemical production facility. The method consists of training a surrogate model, which can be a linear regression model, neural networks, among others, using simulation logs. The surrogate model is trained to predict optimal outputs based on hypothetical inputs, and the inputs are limited by constraints, such as the minimum and maximum values of the input variables observed in the simulation logs.
[011] Thus, document US11669063B2 describes a method for improving the optimization of the selection of oils for a refinery, implementing constraints for the optimization responses on surrogate models, Petition 870250058707, dated 10 / 07 / 2025, p. 8 / 64 5 / 27 that replace refining planning models, limiting their response to production plans with stream flow rates compatible with the capacities of the refinery units. For example, if the plant's production capacity is to produce 1000 gallons / hour of aviation fuel, and the replacement model finds a solution of 1200 gallons / hour, this means that the model needs to be retrained based on new results from the traditional simulation (refining planning model), since the simulated model has this clear restriction and the replacement model does not. However, document US11669063B2 does not apply to the modeling of a process unit, nor to the modeling of catalyst properties for the performance response of an FCC unit, thus differing from the present invention.
[012] The article Artificial Intelligence for Hybrid Modeling in Fluid Catalytic Cracking (FCC) discusses a modeling approach for predicting FCC yields based on a 3D CPFD (Computational Fluid Dynamics of Particles) fluid dynamics study and FCC process data to train and validate a machine learning (ML) model. CPFD is used to simulate the behavior of fluids and particles in FCC reactors, allowing the analysis of complex operating conditions on an industrial scale. The document does not describe an optimization method, but rather the direct prediction of plant results. The document studies the interaction of equipment geometries and the interaction between the fluid phase and catalyst particles, without specifying the physicochemical properties of the catalyst and its formulation. The result of the hybrid ML-CFD modeling serves to accelerate the results of CFD modeling.In contrast, the present invention uses the prediction of a large database simulated by an FCC phenomenological process simulator to pre-train a machine learning model and then retrains the model with real data. Petition 870250058707, dated 10 / 07 / 2025, page 9 / 64 6 / 27 considering the physicochemical properties of catalysts. The result is a unique model capable of predicting and optimizing different FCC catalyst formulations, resulting in an economic optimization of the catalyst formulation, a step also absent in this article.
[013] Thus, it is evident that none of the cited State of the Art documents are capable of anticipating a prediction method that contemplates a hybrid model trained with simulation data and real data to optimize the formulation of catalysts, being able to predict performance in an industrial plant in a more precise and economical way, as proposed by the present invention. SUMMARY OF THE INVENTION
[014] The present invention proposes a predictive method to optimize the performance of a catalytic cracking (FCC) unit, essential for the production of LPG (Liquefied Petroleum Gas), propene and naphtha, for example. The technical problem addressed is the difficulty in accurately predicting the unit's yields due to the physicochemical properties of the virgin catalytic system concomitant with operational variables and catalyst deactivation, which impact the process outcome.The method claimed in this invention includes: pre-training a model with simulated data (BD_FCCSIM - refers to the first database); refinement with real historical data (BD_Histórica - refers to the second database); adjusting the model to consider process variables and physicochemical properties of the catalysts; validation with error and correlation metrics; and implementation in a monitoring system that includes modifications to the hybrid model in the process simulator, allowing its real-time use in optimizing the catalytic system of the unit in conjunction with the operating conditions. This method, which includes the hybrid model adapted to the process simulator, combines phenomenology and... Petition 870250058707, dated 10 / 07 / 2025, page 10 / 64 7 / 27 neural networks, offering accurate predictions and enabling real-time adjustments, increasing the efficiency and adaptability of the FCC process according to market demand and operating conditions. BRIEF DESCRIPTION OF THE DRAWINGS
[015] The present invention will now be described with reference to its typical embodiments and also with reference to the accompanying drawings, in which:
[016] Figure 1 is a flowchart of the present invention illustrating the steps of the method.
[017] Figure 2 is a graphical representation of the adjustment of the hybrid model for fuel gas yield in mass percent, according to an exemplary embodiment of the present invention.
[018] Figure 3 is a graphical representation of the adjustment of the hybrid model for hydrogen yield in mass percent, according to an exemplary embodiment of the present invention.
[019] Figure 4 is a graphical representation of the adjustment of the hybrid model for LPG yield in mass percent, according to an exemplary embodiment of the present invention.
[020] Figure 5 is a graphical representation of the hybrid model fit for the yield of propene in mass percent, according to an exemplary embodiment of the present invention.
[021] Figure 6 is a graphical representation of the adjustment of the hybrid model for the yield of cracked naphtha in mass percent, according to an exemplary embodiment of the present invention.
[022] Figure 7 is a graphical representation of the hybrid model fit for LCO (light recycled oil) yield in mass percent, according to an exemplary embodiment of the present invention. Petition 870250058707, dated 10 / 07 / 2025, page 11 / 64 8 / 27
[023] Figure 8 is a graphical representation of the adjustment of the hybrid model for the yield of decanted oil in mass percent, according to an exemplary embodiment of the present invention.
[024] Figure 9 is a graphical representation of the adjustment of the hybrid model for coke yield in mass percent, according to an exemplary embodiment of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[025] The present invention relates to a predictive modeling method that combines phenomenological modeling and Multilayer Perceptron (MLP) neural networks to simulate and predict the performance of an FCC unit. The method comprises several sequential steps involving pretraining the model with a simulated database, refinement with a historical database of the FCC unit, tuning and configuring the model to predict unit yields from catalyst characteristics, model validation based on error and correlation metrics, and finally, implementation of the hybrid model in the process simulator that allows a monitoring and optimization system for the FCC unit, jointly adjusting the catalytic system and the unit conditions. Pre-Training of the Predictive Model
[026] The first step of the method consists of pre-training a predictive model using a simulated database (BD_FCCSIM). This database is generated by means of a specific simulator for FCC, which represents varied operational scenarios, integrating different process parameters and catalyst characteristics.
[027] The BD_FCCSIM database includes input variables that represent the operating conditions of the FCC unit, such as ZSM-5 zeolite content, reaction temperature (TRX), dense phase temperature (TFD), dilute phase temperature Petition 870250058707, dated 10 / 07 / 2025, page 12 / 64 9 / 27 (TFDil), combined charge temperature (TCC), charge flow rate (VCarga), vacuum heavy gas oil flow rate (VGop), coking heavy gas oil flow rate (VGopk), riser naphtha flow rate (VNR), 20 / 4 density of vacuum heavy gas oil (d20Carga), 20 / 4 density of coking gas oil (d20Gopk), basic nitrogen of the charge (NBCarga), and specific properties of fresh catalysts, such as specific surface area (VCatAE), pore volume (VCatVP), apparent density (VCatDA), and metal content such as sodium (VCatNa), iron (VCatFe), and rare earths (VCatRE). In addition, this basis considers equilibrium catalyst (Ecat) variables, including activity, coke content, accessibility, and contaminant content, such as vanadium and nickel.
[028] To generate the BD_FCCSIM, a factorial simulation design is applied, which allows the creation of multiple scenarios with variations in operating conditions and catalyst parameters. This design must be carried out in more than one distinct period, each with a different calibration. In both periods, a common base catalyst is used, with one period including the ZSM-5 zeolite additive, while the other preferably does not. This approach allows the predictive model to learn the influence of different compositions and process conditions on the performance of the FCC unit.
[029] During pretraining, the MLP model is configured with training and test datasets from the simulated database BD_FCCSIM. Overfitting prevention techniques, such as early stopping, are applied to ensure that the model is not overfitted to the pretraining data, maintaining greater generalization for the next step. Refinement with Historical Database (Historical DB)
[030] After pretraining, the predictive model is refined with a historical database (Historical_DB), composed of real data from the FCC unit. This database covers a long Petition 870250058707, dated 10 / 07 / 2025, page 13 / 64 The 10 / 27 operating period of the unit includes process variables, parameters of virgin and equilibrium catalysts (Ecat), and product yields. The Historical Database contains yield data for products such as fuel gas, LPG, cracked naphtha, LCO, decanted oil, and coke, as well as gas chromatography variables (hydrogen and C1 to C4), CO and CO2 content in the combustion gases, reaction temperature, feedstock density, and GOP flow rate. This data allows the model to incorporate the real conditions of the unit, reflecting the influences of contaminants on the catalyst and operational variables.
[031] To ensure data consistency and quality, the Historical_Database is filtered to remove records where the mass balance is outside the 98 to 102% range and to exclude data with missing values. Then, the data are divided into three sets: training, test, and validation, using data from the same two months of each year to compose the validation set. This refinement adjusts the MLP model to consider characteristics of catalysts in real-world use, adjusting it to improve the accuracy of predictions and the representativeness of optimizations. Model Adjustment and Configuration
[032] The next step involves tuning and configuring the MLP neural network to predict the yields of the FCC unit and the characteristics of the catalysts in use. The model is structured with an input layer, two hidden layers, and an output layer. The input layer receives the operational variables and characteristics of the catalysts, while the hidden layers are configured with a range of 50 to 200 neurons, depending on the adjustments that minimize the root mean square error (RMSE) and maximize the correlation coefficient (R2) for the test and validation data. The activation functions used can be of the Rectified Linear Unit (ReLU) and hyperbolic tangent type, with Petition 870250058707, dated 10 / 07 / 2025, page 14 / 64 11 / 27 normalized weight initializations to optimize the learning process.
[033] To ensure that the model maintains accuracy in predicting yields, the mass balance is standardized to 100% before and after the modeling process, and the input and output variables are normalized to the range of 0 to 1. Normalization and batch size adjustment to 10 samples per iteration optimize training, ensuring that the model learns effectively and requires less adjustment when applied to future data. Model Validation
[034] The validation stage is fundamental to ensuring the accuracy and applicability of the predictive model. The model is validated based on two main metrics: the Root Mean Square Error (RMSE) and the coefficient of determination (R2), applied to the test and validation datasets. The RMSE measures the accuracy of the predictions, while the R2 indicates the strength of the correlation between the predicted and actual values, providing insight into the model's effectiveness in explaining variations in the data.
[035] Furthermore, validation is conducted with multiple training and test iterations to calculate the mean and standard deviation of the RMSE and R2 values, ensuring consistency in the results. The comparison between the RMSE and R2 values obtained in the pre-training (simulated) and validation (real) datasets ensures that the model is applicable in both simulated and real conditions, selecting the configuration with the lowest RMSE and highest R2.
[036] Finally, the predictive model is implemented in an FCC unit monitoring system. This system integrates the model in real time with the unit, allowing for continuous collection of operational data and the use of the model's predictions for Petition 870250058707, dated 10 / 07 / 2025, page 15 / 64 12 / 27 Adjustment of the catalytic system formulation according to operating conditions. EXEMPLARY MODEL
[037] In an exemplary embodiment, only the MLP model with pre-training was selected as an example for predicting yields and Catalyst Factors (CFs) for CDB / FCC-SIMTM, based on the properties of virgin catalysts used commercially in the U-220 unit of the REPLAN refinery.
[038] The following models used two databases. The first database was generated from a factorial design that considered the catalyst parameters, the process parameters, and the yields obtained through the FCC-SIM™ process simulator. This design was carried out with data from two different dates; that is, two distinct calibrations in the process simulator. The chosen periods have the same base catalyst, but one period used ZSM-5 additive, and the other did not. The response variables were modeled in the FCC-SIM™ simulator according to the complete factorial design listed in Table I. In this design, 11,664 cases were simulated for each calibration: 1) calibration of June 15, 2022, with a catalyst containing 1.5% ZSM-5 in its formulation and a replenishment rate of 4.5 t / d; 2) calibration of April 26, 2017, considering a formulation without ZSM-5 and with a replenishment rate of 7.5 t / d. Therefore, the complete pre-training database for the model contains 23.328 cases. Variable Minimum Value Maximum Value Step Levels TCC 220 280 30 3 TRX 525 540 7.5 3 TFD 690 720 30 2 VNR 0 2000 1000 3 Vgopk 0 2000 1000 3 VCarga 7300 7800 500 2 Petition 870250058707, dated 10 / 07 / 2025, page 16 / 64 13 / 27 d20Carga 0, 92 0, 945 0,0125 3 NBCarga 800 1500 700 2 ECatV 750 1500 375 3 ECatNa 0,4 0, 601 0,2 2 Table I - Factorial design of the simulated cases in FCC- SIM™.
[039] The second database contains the actual data from the U-220 unit of REPLAN, and covers the period from January 2011 to June 2022, and contains information on the unit parameters, virgin catalyst and equilibrium catalyst (Ecat) and their respective yields. For the analyses, the yields of fuel gas, LPG, cracked naphtha, LCO, decanted oil, and coke were considered. In addition, gas chromatography was considered, covering hydrogen and the range from C1 to C4.
[040] Unit parameters: CO content in flue gases (FGCO), CO2 content in flue gases (FGCO2), 20 / 4 density of the feed (d20Feed), 20 / 4 density of coking gas oil (d20Gopk), basic nitrogen in the feed (NBFeed), reaction temperature of riser 1 (TRX), reaction temperature of riser 2 (TRX2), dense phase temperature (TFD), dilute phase temperature (TFDil), combined feed temperature (TCC), blower air flow rate (ArSop), feed flow rate (VFeed), GOP flow rate (VGop), GOPK flow rate (VGopk), naphtha riser flow rate (VNR), and catalyst replenishment (RepCat);
[041] Virgin catalyst parameters: catalyst specific surface area (VCatAE), catalyst deactivated specific surface area (788 °C / 100% vapor for 5 h) (VCatAED), catalyst pore volume (VCatVP), catalyst apparent density (VCatDA), catalyst unit cell size (VCatTCU), catalyst accessibility (VCatAAI), catalyst alumina content (VCatAl), sodium content (VCatNa), iron content (VCatFe), rare earth content (VCatRE), phosphorus content (VCatP), Petition 870250058707, dated 10 / 07 / 2025, page 17 / 64 14 / 27 deactivated catalyst activity (788 °C / 100% vapor for 5 h) (VCatMAT), and ZSM-5 zeolite content (ZSM); and
[042] Equilibrium catalyst (Ecat) parameters: Ecat activity (ECatMAT), Ecat accessibility (ECatAAI), Ecat specific surface area (ECatAE), Ecat micropore volume (ECatMiPV), Ecat mesopore area (ECatMSA), Ecat coke content (ECatK), Ecat alumina content (ECatAl), Ecat rare earth content (ECatRE), Ecat sodium content (ECatNa), Ecat iron content (ECatFe), Ecat vanadium content (ECatV), Ecat nickel content (ECatNi), Ecat phosphorus content (ECatP), and Ecat platinum content (ECatPt).
[043] The databases used in this study were designated as follows: • BD_FCCSIM: Refers to the first database, which consists of the yields obtained through simulations using the factorial design mentioned in Table I (cases: 23,328). • Historical Database: Refers to the second database, which covers the period from January 2011 to June 2022 and contains unit data, including virgin catalyst and Ecat parameters and process parameters with their respective yields throughout that period. The strategy adopted was to filter data where the unit's material balance was outside the range of 98 to 102 and for cases with missing data. In this way, the original database with 4,974 cases was reduced to 2,918 cases; however, the quality of the data information improved.
[044] These two databases were used to perform the analyses and develop the necessary models. The objective is to combine phenomenological modeling with machine learning. This approach seeks to take advantage of the strengths of both methods, using phenomenological modeling to capture prior knowledge and existing physicochemical relationships, while machine learning is Petition 870250058707, dated 10 / 07 / 2025, page 18 / 64 15 / 27 is used to deal with the complexity and uncertainty of the systems. Therefore, this approach allows considering the underlying physicochemical laws of the systems studied. In this way, the estimation of the neural network parameters follows the flowchart in Figure 1.
[045] The MLP (Multilayer Perceptron) Pre-Train model uses the BD_FCCSIM database from the factorial design for simulation with FCC-SIMTM (Table I). This approach incorporates partial physical and kinetic elements to define the network parameters, which will later be refined with real data from BD_Histórica. The data were divided into training (70%) and test (30%) sets.
[046] The properties of the virgin catalyst were represented only by the composition of two catalysts and with variation only in the ZSM-5 zeolite content. The parameters of contaminant metals in the equilibrium catalyst (Ecat) and the process variables were evaluated as independent variables, and the yields and C1-C4 chromatography as dependent variables. The mass balance of the yields was standardized to 100 before and after modeling. The input and output variables were normalized in the range of 0 to 1. To avoid overfitting, the early stopping technique was applied. The batch size was fixed at 10, without considering computational time limitations.
[047] The proposed MLP model has 26 input variables (process, virgin catalyst, and Ecat contamination) and 19 output variables (yields and chromatography). The input layer consists of the initialization layer, which receives the 26 input variables and the bias, using normal initialization models, a ReLU (Rectified Linear Unit) activation function (package kernel_initializer='normal', activation='relu'), and contains 28 neurons. Two more hidden layers were used with normal initialization and a ReLU activation function. Petition 870250058707, dated 10 / 07 / 2025, page 19 / 64 16 / 27 hyperbolic tangent (kernel_initializer='normal', activation='tanh') and with a varying number of neurons depending on the evaluated architecture. The output layer is a combination of the 19 modeled variables and is represented by normal initialization models (kernel_initializer='normal'). The description of the strategy used in each MPL performed is presented in Table II. The evaluation and validation criterion for the models was the RMSE (Root Mean Square Error) in model validation, analyzing different numbers of neurons (50, 100, 200) in the two varied hidden layers. Five distinct trainings were performed to obtain the average RMSE and the correlation coefficient (R2) for each choice of hyperparameters. Table II also shows the average RMSE of the main yields and the correlation coefficient (R2) of the trained models.The models obtained showed a higher R² value and lower RMSE of the individual yields using 50 neurons in the hidden layers, except for the cracked naphtha yield, in which the model with 100 neurons showed the lowest RMSE. Since the data that generated this pre-training MLP model were based on simulated data, the errors are small.
[048] The parameters in Table II are defined as follows: 26 input variables encompassing process and catalyst variables; 19 output variables main yields and gases C1 to C4.
[049] Considering an example with 50 neurons in the hidden layers, the parameters of each layer are defined as follows: • Initialization layer (input layer): (26 input variables + 1 bias) x (input layer (28)) = 756 (Eq. 1) • Dense layer 1 is given by: (input layer (28) + 1 bias) x Dense layer 1 (50) = 1450 (Eq. 2) Petition 870250058707, dated 10 / 07 / 2025, page 20 / 64 17 / 27 • Dense layer 2 is given by: (Densal layer (50) + 1 bias) x Densal layer2 (50) = 2550 (Eq. 3) • Output layer: (Dense layer 2 (50) + 1 bias) x 19 output variables = 969 (Eq. 4) Neuron Parameters Neuron Parameters Neuron Parameters Model 50n 100n 200n Input 28 756 28 756 28 756 Dense 1 50 1450 100 2900 200 5800 Dense 2 50 2550 100 10100 200 40200 Output 19 969 19 1919 19 3819 Total 5725 15675 50575 Average STD Average STD Average STD R2 train 0.9989 0.0001 0.9986 0.0002 0.9987 0.0005 R2 test 0.9989 0.0001 0.9986 0.0002 0.9987 0.0005 RMSE GC (%m) 0.0210 0.0029 0.0253 0.0039 0.0198 0.0039 RMSE LPG (%m) 0.1177 0.0106 0.1289 0.0131 0.1352 0.0246 RMSE NC (%m) 0.1706 0.0422 0.1893 0.0343 0.2117 0.0981 RMSE LCO (%m) 0.1125 0.0117 0.1308 0.0171 0.1252 0.0181 RMSE OD(%m) 0.1264 0.0133 0.1368 0.0158 0.1449 0.0388 RMSE Coke(%m) 0.0277 0.0031 0.0322 0.0037 0.0281 0.0034 Table II - RMSE and R2 of the test versus the number of neurons for the pre-training MLP model (BD_FCCSIM)
[050] Although the variation in RMSE between the number of neurons is reduced, the network structure with 50 neurons in the inner layers was chosen for the MLP models.
[051] In constructing the pre-trained MLP model, the historical database, BD_Histórica, was used, along with the structure of the pre-trained MLP model obtained in the previous section. In this step, the network parameters are refined with real data from the REPLAN U-220 unit using its database. Petition 870250058707, dated 10 / 07 / 2025, page 21 / 64 18 / 27 (BD_Histórica). The database was divided into three sets: training (58.3%), test (25.0%), and validation (16.7%, corresponding to the months of June and July of each year). The training and test data are randomized by the model's own strategy.
[052] In this step, the properties of the virgin catalyst were represented, as well as the replacement of virgin catalyst and the contamination of the Ecat by contaminating metals, and the process variables were evaluated as independent variables. Yields and chromatography of C1-C4 gases were used as dependent variables. The mass balance of yields was standardized to 100 before and after modeling. Input and output variables were normalized in the range of 0 to 1. To avoid overfitting, the early stopping technique was applied. The batch size was reduced to 2, since the database is smaller compared to that used for pre-training. The MLP model re-estimated its parameters from the best model obtained in the previous step (Pre-Training MLP Model based on FCCSIMTM). That is, the MLP model has the same parameters and initiation and activation functions.
[053] The criterion for evaluating and validating the models was the RMSE (Root Mean Square Error) in the model validation, for a number of neurons of 50 in the two inner layers. Ten iterations were performed to obtain the average RMSE and the correlation coefficient (R2). Table III also shows the RMSE of the iterations and the correlation coefficient (R2) of the model. The models converged between 11 and 39 epochs, even though 200 epochs were offered to the algorithm. Training Model Test Validation RMSE RMSE RMSE 0 0.668 0.670 0.784 1 0.643 0.651 0.755 2 0.642 0.646 0.729 Petition 870250058707, dated 10 / 07 / 2025, page 22 / 64 19 / 27 3 0.634 0.641 0.794 4 0.628 0.637 0.771 5 0.618 0.629 0.779 6 0.614 0.628 0.778 7 0.595 0.615 0.791 8 0.604 0.621 0.766 9 0.602 0.621 0.808 Table III - RMSE and R2 of the MLP model with pre-training and data from the U-220 unit (Historical Database) Model RMSE GC RMSE H2 RMSE LPG RMSE Propylene RMSE NC RMSE LCO RMSE OD RMSE Coke 0 0.417 0.0123 1.264 0.614 1.811 1.339 1.950 0.247 1 0.409 0.0124 1.203 0.575 1, 683 1.380 1.875 0.270 2 0.437 0.0125 1.350 0.560 1.673 1.338 1.558 0.253 3 0.429 0.0122 1.453 0.632 1.668 1.402 1,931 0.255 4 0.429 0.0118 1.250 0.550 1.845 1.343 1.843 0.241 5 0.441 0.0127 1.199 0.536 1.866 1.312 1.950 0.255 6 0.413 0.0129 1.298 0.589 1.876 1.351 1.801 0.261 7 0.440 0.0125 1.404 0.617 1.821 1.361 1.861 0.257 8 0.442 0.0125 1.172 0.550 1.932 1.333 1.741 0.255 9 0.419 0.0125 1.314 0.566 1.979 1.344 1.927 0.252 Table IV - RMSE of MLP model yields with pretraining and data from unit U-220 (Historical Database)
[054] Figures 2 to 9 illustrate the model fit with real data, in which: the black dots are from the real data of the U-220 unit; the gray square dots are the training prediction; the gray cross dots are the test prediction; and the gray cross dots are the validation prediction. More specifically, Figure 2 represents the model results for fuel gas yield, Figure 3 the model results for hydrogen yield, Figure 4 the model results for LPG yield, Figure 5 the model results for propene yield, Figure 6 the Petition 870250058707, dated 10 / 07 / 2025, page 23 / 64 Figure 20 / 27 shows the model results for cracked naphtha yield, Figure 7 shows the model results for LCO yield, Figure 8 shows the model results for decanted oil yield, and Figure 9 shows the model results for coke yield. It can be observed that, although the validation model assumes low R² values, the graphical visualization shows that the high dispersion of the data and the predicted model data are within this range for all yields. The high dispersion of the experimental data did not allow for a better R² value for model validation. Comparison of performance prediction between the Hybrid MLP model and industrial assessments of U-220 catalysts
[055] A comparison was made between existing industrial assessments of catalysts by changing catalytic systems at U220 and the prediction of the Hybrid MLP model developed in this work. The commercial assessments carried out followed the PETROBRAS standard for industrial assessment of catalysts (BASTIANI, R. et al. Industrial Assessment Standard for Catalysts. Rio de Janeiro: PETROBRAS. CENPES. PDAB. TFCC, 2011. 76 p. Technical Communication (CT TFCC n° (007 / 2011). The performance comparison of the catalysts is presented in terms of yield delta between the evaluated catalysts and, according to the standard, this evaluation can be performed by: 1. Group analysis, when there is a group with both catalysts under similar operating conditions and according to the PETROBRAS standard for evaluation; 2. Pure MLP-type Neural Networks with data from the period in which the catalysts are being evaluated; 3. Evaluation of Ecats of catalytic systems in laboratory units (ACE unit: 535 °C, 100% GOP) and / or pilot plant (DCR unit: 540 °C, 100% GOP). In industrial assessments of FCC catalytic systems, the goal is to rank the performance of the catalysts. To this end, Petition 870250058707, dated 10 / 07 / 2025, page 24 / 64 21 / 27 the techniques are applied in the evaluation; and whenever possible, the greatest agreement between them is sought.
[056] The Hybrid MLP model prediction was performed in three ways: 1. Catalytic performance prediction for the same catalytic systems analyzed in the industrial evaluation, considering the average performance for the operation with 4974 data points; that is, for the entire operating period on which the model estimate was based (January 2011 to June 2022, REPLAN unit U-220); 2. Catalytic performance prediction for the same catalytic systems analyzed in the industrial evaluation, considering only the period in which the catalysts are being evaluated; 3.Catalyst Factors (CFs) were calculated from the prediction of the Hybrid MLP model for the catalytic systems used in U-220, and these CFs were fed into the CDB / FCC-SIM™ model in order to alter the CDB factors, modifying the FCC-SIM™ riser / reactor expressions that predict the yields and properties of the products with the change of catalytic systems, thus being able to predict the performance with the change of catalytic system of any U-220 calibration in the FCC-SIM™ software. Table V shows the characterization of virgin catalysts, used for prediction and estimation of the catalyst CFs by the hybrid MLP model for catalytic systems used in the U-220 unit of REPLAN in the period from January 2011 to June 2022. SIRIUS 2727 SIRIUS 2147 SIRIUS ZOOM 2176 Catalyst Specific surface area of catalyst (m² / g) 316 313 312 Specific surface area of deactivated catalyst (m² / g) 180 176 175 Pore volume (ml / g) 0.412 0.409 0.412 Petition 870250058707, dated 10 / 07 / 2025, page 25 / 64 22 / 27 Apparent density (g / ml) 0.74 0.75 0.75 Unit cell size (nm) 2.460 2.460 2.460 Accessibility (%) 8.2 11.7 11.0 Alumina content (% m / m) 49.35 53.68 52.36 Sodium content (% m / m) 0.35 0.39 0.35 Iron content (% m / m) 0.35 0.40 0.42 Rare earth content (% m / m) 1.94 2.06 1.84 Phosphorus content (% m / m) 0.42 0.22 0.62 ZSM-5 content (% m / m) 0.00 0.00 1.50 Table V - Characterization of virgin catalysts used for the prediction and estimation of the catalyst coefficients (FCs) of the hybrid MLP model.
[057] The CFs that modify the FCC-SIM™ reaction model follow an experimental plan for the evaluation of catalysts in a laboratory unit or pilot plant (FCC-SIM Technology Manual Version 2002. KBC PROFIMATICS. 2004).
[058] To estimate the CFs factors for the CDB / FCCSIMTM model, the following strategy was adopted: First, a standard condition was adopted (Table VI) and the same database was repeated 9 times. Variable TRX (Riser 1 and 2) 533 Naphtha riser flow rate (m³ / d) 600 TFD 707 d20 / 4 feed 0.93 6 TFDil 710 Basic nitrogen in feed (ppm) 1100 TCC 255 Catalyst replenishment (t / d) 5 Blower air flow rate (t / h) 192 Sodium content in Ecat (ppm) 4300 Feed flow rate (m³ / d) 7400 Iron content in Ecat (ppm) 4200 GOP DD flow rate (m³ / d) 6000 Vanadium content in Ecat (ppm) 1150 GOPK flow rate (m³ / d) 800 Nickel content in Ecat (ppm) 700 Table VI - Arbitrary standard operating condition. Petition 870250058707, dated 10 / 07 / 2025, p. 26 / 64 23 / 27
[059] Arbitrary default condition: • Two sets of data; only the replacement of 4 and 6 t / d was varied; • One of the sets adopted TRXs of 545 °C; • One of the sets adopted: TCC of 200 °C, vanadium content of 500 ppm, nickel content of 400 ppm, and replacement of 6 t / d; • One of the sets adopted: TCC of 200 °C, vanadium content of 500 ppm, nickel content of 400 ppm, and replacement of 4 t / d; • One of the sets adopted: vanadium content of 500 ppm, nickel content of 400 ppm; • One of the sets adopted: vanadium content of 1800 ppm, nickel content of 1000 ppm; • One of the sets adopted: vanadium content of 1800 ppm, nickel content of 1000 ppm;
[060] For all the above conditions, yields and CFs were estimated for the catalysts whose properties are in Table V. The CFs factors were transformed into a file and used directly in the simulation in the CDB / FCC-SIMTM module.
[061] Table VII shows the yield deltas obtained both in the industrial evaluation of catalysts and in the prediction of the developed Hybrid MLP model, in which the catalytic systems evaluated are: • SIRIUS 2727 base catalyst (August 19, 2015 to May 4, 2016): 50% VEGA 2640 + 50% OPAL SC LRT; • Catalyst evaluated SIRIUS 2147 (11 / 01 / 2017 to 25 / 07 / 2018): 25% VEGA 2742 + 75% OPAL SC LRT. Evaluation Conventional evaluation of catalysts
[17] Hybrid MLP model Yields Group Neural networks MLP1 MLP H.2 model MLP H.1 model CDB / FCCSIM BD used 87 4974 87 Petition 870250058707, dated 10 / 07 / 2025, page 27 / 64 24 / 27 GC -0.41 +0.003 + 0.06 + 0.06 -0.04 H2 +0.012 +0.014 +0.008 0.00 +0.002 LPG + 1.32 + 1.08 + 0.32 + 0.19 + 0.69 Propene + 0.21 + 0.33 + 0.20 + 0.09 + 0.14 NC + 1.58 + 0.32 + 0.30 + 0.10 + 0.23 LCO -1.88 -0.59 -0.33 -0.22 -0.28 OD -1.09 -0.89 -0.55 -0.34 -0.63 Coke + 0.39 + 0.06 + 0.21 + 0.21 + 0.02 Delta coke + 4.6% -0.2% + 1.2% Table VII - Comparison of catalytic performance in industrial evaluation of the SIRIUS 2147 catalyst in relation to SIRIUS 2727, and equivalent prediction of the hybrid MLP model.
[062] The performance evaluation of the SIRIUS 2147 catalyst in relation to the SIRIUS 2727 catalyst was carried out by two different analyses: evaluation by similar groups and neural networks (Table VII). The industrial evaluation of the catalytic systems indicated that the performance of the SIRIUS 2147 catalyst is superior to that of the SIRIUS 2727 catalyst for LPG, propene and naphtha yields, with a reduction in LCO and decanted oil (DO) yields. The predictions of the Hybrid MLP model applied to the catalytic systems indicated agreement with the industrial evaluation for the increase in naphtha yield with a reduction in LCO and funds.
[063] Table VIII shows the yield deltas obtained both in the industrial evaluation of catalysts and in the prediction of the developed Hybrid MLP model, in which the catalytic systems evaluated are: • SIRIUS 2147 base catalyst (11 / 01 / 2017 to 25 / 07 / 2018): 25% VEGA 2742 + 75% OPAL SC LRT; • Catalyst evaluated SIRIUS ZOOM 2176 (02 / 01 / 2019 to 11 / 17 / 2021): 24.6% VEGA 2742 + 72.19% OPAL SC LRT + 3.75% ZOOM. PETROBRAS Standard Assessment of Catalysts - Hybrid MLP Model Petition 870250058707, dated 10 / 07 / 2025, page 28 / 64 25 / 27 Yields Neural Networks MLP1 Simplified GOP Model ZSM-5 1.0 MLP H.2 Model MLP H.1 CDB / FCCSIM BD used 140 4974 140 GC -0.26 -0.13 -0.22 -0.13 -0.48 H2 -0.002 -0.007 -0.001 GLP +4.08 +3.26 +4.02 +4.21 +3.82 Propene +1.72 +2.18 +1.69 +2.09 +1.67 NC -2.80 -2.89 -3.21 -3.28 -3.59 LCO +0.16 -0.14 +0.23 +0.46 +1.41 OD -1.41 -0. 11 -1.01 -1.49 -1.15 Coke +0.22 0.00 +0.19 +0.24 -0.01 Delta coke -0.4% -11.8% -3.31% Table VIII - Comparison of catalytic performance in evaluation industrial performance of the SIRIUS ZOOM 2176 catalyst compared to SIRIUS 2147, and equivalent prediction of the hybrid MLP model.
[064] The performance evaluation of the SIRIUS 2147 catalyst in relation to the SIRIUS ZOOM 2176 catalyst was performed using only neural networks. This evaluation was complemented by yield prediction using the Simplified ZSM-5 1.0 GOP Model (Table VIII). The Simplified ZSM-5 1.0 GOP Model was developed to predict the performance of ZSM-5 zeolite under various operating conditions, and is based on a wide range of tests performed in a pilot plant - DCR (PINHO, AR Maximization of Light Olefins with Vacuum Heavy Gas Oil. Rio de Janeiro: PETROBRAS. CENPES. PDAB. TFCC, 2012. 15 p. Technical Report (RT TFCC n° 001 / 2012)). Industrial evaluation of the catalytic systems indicated that the performance of the SIRIUS ZOOM 2176 catalyst is superior to that of the SIRIUS 2727 catalyst for LPG and propene yields, with a reduction in fuel gas and cracked naphtha yields and, unexpectedly, a reduction in decanted oil yield.The predictions of the Hybrid MLP model applied to the catalytic systems indicated agreement with the industrial evaluation and the Simplified GOP ZSM-5 1.0 Model, presenting an equivalent performance profile. Petition 870250058707, dated 10 / 07 / 2025, page 29 / 64 26 / 27
[065] The modeling of the yields of the U-220 catalytic cracking unit at REPLAN as a function of its operation and the catalytic systems used was carried out according to the method proposed by the present invention. This approach seeks to take advantage of the benefits of both methods, using physical modeling to capture prior knowledge and existing physical relationships, while machine learning is used to deal with the complexity and uncertainty of the systems, increasing the accuracy, interpretation and robustness of the models.
[066] The modeling was performed in the following steps: 1) Pre-training using simulated data in the FCC-SIM™ simulator, which takes into account physical fundamentals, such as material and energy balance in the effects of operational variables on unit performance; 2) Incorporation of operating data, catalytic system quality, and U-220 yields from January 2011 to June 2022. The trained models were evaluated by RMSE (Root Mean Square Error) and their correlation coefficient (R²), with small errors (low RMSE values) and high R² values observed for the pre-training model, which was the best model for a number of 50 neurons in the two inner layers. This strategy was used to re-estimate the MLP model parameters with historical data from the U-220 unit. In this step, 10 iterations were performed to obtain the RMSE and R² of the Hybrid MLP model of the present invention.The best model with the lowest individual yield error values and the highest R2 was used to predict and evaluate the method proposed in this work.
[067] The Hybrid MLP model was compared in four catalyst reformulations from the U-220 unit and compared with industrial assessments carried out following the PETROBRAS standard for industrial catalyst evaluation. The only catalyst information provided to the Hybrid MLP model was the properties of the virgin catalysts. The Hybrid MLP model was Petition 870250058707, dated 10 / 07 / 2025, page 30 / 64 27 / 27 was evaluated in three different ways: one predicting catalyst performance for the entire database, another predicting performance for the dataset during the catalyst's usage period in the unit, and another considering the adapted CDB / FCCSIMT model with the CFs predicted by the model. All comparisons correctly predicted the conversion trend of funds in both the industrial evaluation and the Hybrid MLP model of the present invention. It is possible to affirm that the Hybrid MLP model performed well in predicting the increase in TOPAZ technology (FCC catalyst manufacturing technology) and the use of ZSM-5 based additive, correctly predicting the performance direction of the main yields: LPG, propene, naphtha, LCO, and funds.
[068] Therefore, the results of the Hybrid MLP model allowed us to estimate the performance of the FCC unit by providing only the properties of the virgin catalyst and the operating conditions, with extremely reduced costs when compared to other current alternatives. The proposed method will be implemented in the FCC model of PETROBRAS' Digital Twin and will allow us to indicate the best formulation based on market demand for products. Petition 870250058707, dated 10 / 07 / 2025, page 31 / 64
Claims
1 / 4 CLAIMS 1. A method for predicting the performance of a catalytic cracking (FCC) unit, characterized in that it comprises: pre-training a predictive model using a simulated database (BD_FCCSIM) representing various operational scenarios of an FCC unit; refining the predictive model with a historical database (BD_Histórica) of real-world operation of an FCC unit, including virgin catalyst and Ecat parameters and process parameters with their respective yields over a ten-year period; adjusting and configuring the model to predict the yields of the FCC unit and catalyst characteristics, based on operating conditions and the properties of the catalysts used; validating the model based on error and correlation metrics to ensure its application in simulated and real-world scenarios;Implementation of the predictive model in a monitoring system to adjust the operations of the FCC unit, according to changes in operating conditions.
2. Method, according to claim 1, characterized in that the simulated database (BD_FCCSIM) and historical data (BD_Histórica) which includes input variables such as ZSM-5 zeolite content, reaction temperature (TRX), dense phase temperature (TFD), dilute phase temperature (TFDil), combined charge temperature (TCC), charge flow rate (VCarga), GOP flow rate (VGop), GOPK flow rate (VGopk), naphtha riser flow rate (VNR), 20 / 4 charge density (d20Carga), 20 / 4 density of Petition 870240110839, dated 12 / 27 / 2024, p.62 / 71 2 / 4 coking gas oil (d20Gopk), basic nitrogen of the feedstock (NBCarga), specific area of the catalyst (VCatAE), specific area of the deactivated catalyst (VCatAED), catalyst pore volume (VCatVP), apparent density of the catalyst (VCatDA), catalyst unit cell size (VCatTCU), catalyst accessibility (VCatAAI), catalyst alumina content (VCatAl), sodium content (VCatNa), iron content (VCatFe), rare earth content (VCatRE), phosphorus content (VCatP), vanadium content in the catalyst (ECatV), nickel content in the catalyst (ECatNi), phosphorus content in the catalyst (ECatP), coke content in the catalyst (ECatK), deactivated catalyst activity (VCatMAT), catalyst activity (ECatMAT), and blower air flow rate (ArSop).
3. Method, according to claim 1, characterized in that the simulated database (BD_FCCSIM) is created by means of a factorial simulation design, which generates multiple scenarios by varying the conditions of the catalyst and process parameters.
4. Method, according to claim 1, characterized in that the pre-training step uses a factorial design carried out in two different periods with two distinct calibrations, such that the chosen periods have the same base catalyst, one period using ZSM-5 additive, and the other not.
5. Method according to claim 1, characterized in that the historical database (BD_Histórica) is composed of actual operating data from the FCC unit over an extended period, including information on process conditions and parameters of virgin and equilibrium catalysts (Ecat). Petition 870240110839, dated 12 / 27 / 2024, pp. 63 / 71 3 / 4 6. Method, according to claim 1, characterized in that the historical database (BD_Histórica) contains product yield data including: fuel gas, LPG, cracked naphtha, LCO, decanted oil, coke, gas chromatography (hydrogen and C1 to C4).
7. Method, according to claim 1, characterized in that the adjustment of the predictive model includes the configuration of a Multilayer Perceptron (MLP) type neural network with one input layer, two hidden layers and one output layer, the input layer being composed of operational and catalyst variables.
8. Method, according to claim 7, characterized in that the hidden layers of the predictive model are configured to include between 50 and 200 neurons in each layer, with ReLU and tanh type activation functions, applying normalized weight initializations.
9. Method, according to claim 1, characterized in that the validation of the predictive model includes: the calculation of the root mean square error (RMSE) to evaluate the accuracy of the model's predictions in relation to the real data; and the calculation of the correlation coefficient (R2) between the values predicted by the model and the real values obtained in the test and validation scenarios.
10. Method, according to claim 1, characterized in that the model validation includes selecting a model configuration that minimizes the RMSE and maximizes the R2 on the training, test, and validation sets.
11. Method, according to claim 1, characterized in that the monitoring system uses the model predictions to adjust operational variables of the FCC unit, Petition 870240110839, dated 12 / 27 / 2024, page 64 / 71 4 / 4, such as reaction temperature, feed flow rate, and catalyst replacement rate, according to the operating conditions and characteristics of the catalyst in use.
12. Method, according to claim 1, characterized in that the implementation of the predictive model in the monitoring system allows the storage and processing of the predicted Catalyst Factors (CFs) for each catalyst, which are used to adjust the reaction factors allowing the prediction and optimization of the FCC unit yields.
13. Method, according to claim 1, characterized in that the monitoring system with the predictive model is connected to the digital twin of the FCC unit, allowing simulations of future scenarios and operational adjustments based on forecasts of catalyst yield and behavior. Petition 870240110839, dated 12 / 27 / 2024, pp. 65 / 71