Failure probability prediction method and system for safety instrument systems in petroleum refining units
Through the combination of KPCA and GAM, the nonlinear and high-dimensional data processing problems in the prediction of DU failure probability of SIS equipment in petroleum refining equipment are solved, and the identification and accurate prediction of key factors are achieved, which improves the safety and operation efficiency of the equipment.
Patent Information
- Application Number
- CN202411862549.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-17
- Publication Date
- 2025-08-26
- Estimated Expiration
- 2044-12-17
AI Technical Summary
When the prior art predicts the undetectable hazard failure (DU failure) probability of safety instrumentation system (SIS) equipment in petroleum refining equipment, it is difficult to capture the differences and complex nonlinear relationships between devices, and it is highly dependent on high-quality data and lacks explanatory nature, resulting in inaccurate predictions.
The kernel principal component analysis (KPCA) is used to reduce the dimensionality of high-dimensional data, extract key features, and combine generalized additive model (GAM) to describe the nonlinear relationship between principal component scores and failure probability, and establish a failure probability prediction model.
It realizes accurate prediction of the failure probability of SIS equipment DU, improves the accuracy and reliability of the prediction, strong adaptability, reduces equipment maintenance costs and safety accident risks, and improves the safety and operation efficiency of petroleum refining equipment.
Smart Images

Figure CN119670572B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of safety engineering in petrochemical production, and in particular relates to a failure probability prediction method and system for a safety instrument system of a petroleum refining device. Background Art
[0002] In petrochemical production, safety instrumented systems (SIS) are used to execute predetermined protective measures in extremely dangerous situations, such as closing valves, shutting down equipment, or initiating emergency shutdown procedures. Their purpose is to protect personnel, equipment, and the environment from potential hazards, and they serve as the last line of defense for ensuring petrochemical production and personnel safety. Failure of SIS equipment can trigger major safety incidents such as explosions, fires, or hazardous chemical leaks, resulting in significant economic losses for petrochemical production. Therefore, predicting the failure probability of SIS equipment in petroleum refining units and promptly replacing or reinforcing the equipment is a crucial means of ensuring safe operation.
[0003] The international functional safety standard IEC 61508 categorizes SIS equipment failures into four categories: dangerous detectable failures (DD), dangerous undetectable failures (DU), safe failures, and no-component / no-impact failures. DD and DU failures are critical to the safe operation of the equipment. The difference lies in the fact that DD failures are detectable and, upon occurrence, can be detected through automated diagnostic procedures, leading to a rapid switchover to a backup SIS device, resulting in a minimal impact. DU failures, on the other hand, are latent and typically only revealed during control commands, shutdown tests, or occasional inspections. Because DU failures cannot be detected immediately and must wait until the next scheduled test for restoration, they pose the greatest threat to the safe operation of SIS equipment. Therefore, accurately predicting the probability of DU failure in newly connected SIS equipment has become a key issue in the design and risk management of safety instrumentation systems in petroleum refining.
[0004] Among existing failure probability prediction methods, traditional statistical models have been widely used. For example, patent CN117708703A uses a BP neural network to predict the failure probability of electronic detonators. This method acquires key data through multi-stage testing, establishes a sample library, and then trains the neural network. However, this method relies on a complex network structure, resulting in a computationally intensive training process and a lack of interpretability, making it difficult to identify the role of specific influencing factors. Another method, patent CN114065674B, predicts the failure probability of CMOS devices based on a linear regression model, quantifying the importance of different influencing factors by assigning weights. However, this method assumes a linear relationship between variables, making it difficult to capture complex nonlinear characteristics and highly dependent on data quality. Furthermore, most statistical models rely on multiple sets of device data and assume that items within the same set have similar functions and the same failure rates, but differ in design (such as measurement principle), location, and environment. This assumption is flawed. Even in the same operating environment, similar devices can have varying failure rates. For example, the failure mode "fail to open" (FTO) of a valve is affected by the temperature of the medium flowing through the valve, and its failure probability often changes dynamically. Many factors in petrochemical production have a certain impact on the failure probability of SIS equipment. Among them, the factors with the greatest impact on failure probability are called key influencing factors. In addition, existing methods have difficulty in providing accurate predictions when dealing with high-dimensional factors with complex nonlinear relationships.
[0005] Through the above analysis, the problems and defects of the existing technology are as follows:
[0006] Most statistical models assume that devices in the same group have similar functions and the same failure rates. However, in reality, devices often differ in design (such as measurement principles), installation location, and operating environment. This assumption makes it difficult to accurately reflect the differences between devices. In addition, traditional models are not adaptable enough to dynamic changes in operating conditions. For example, the failure mode "fail to open (FTO)" of a valve can affect the failure probability due to fluctuations in the medium temperature, but existing methods have difficulty capturing real-time changes. More importantly, many models assume that the relationship between variables is linear and lack the ability to model complex nonlinear relationships. At the same time, they are highly dependent on high-quality data and lack clear interpretability, which is not conducive to optimizing equipment design and operating conditions. Summary of the Invention
[0007] In view of the problems existing in the prior art, the present invention provides a failure probability prediction method and system for a safety instrument system of a petroleum refining device.
[0008] The present invention is implemented as follows: a failure probability prediction method for a safety instrument system of a petroleum refining device includes:
[0009] Step 1: Data collection and preprocessing: Demarcate the safety instrument system boundaries and collect equipment information and historical failure data from the safety instrument system of the oil refining unit to determine failure modes and causes, and perform classification and preprocessing.
[0010] Step 2: Data dimensionality reduction and failure probability modeling. KPCA is used to reduce the dimensionality of high-dimensional data and extract key features. KPCA maps complex data into a low-dimensional principal component space, generating a principal component score matrix to capture the nonlinear correlation between DU failure and influencing factors. Based on this, GAM is used to describe the relationship between principal component scores and DU failure probability, and a failure probability prediction model is established.
[0011] Step 3, failure probability prediction, is used to import the probability distribution, correlation, and important influencing factors of DU failure nodes obtained by kernel principal component analysis after data dimensionality reduction and failure probability modeling. GAM is then used for modeling to predict the failure probability of newly connected equipment in the safety instrumented system, and the prediction results are verified and interpreted to evaluate the accuracy and reliability of the failure probability. The failure probability of SIS equipment in the newly connected facility is calculated according to the failure rate prediction formula.
[0012] Furthermore, the data collection and preprocessing includes:
[0013] Data Collection: Data sources for collection include equipment maintenance records, safety standards and specifications, process and instrumentation diagrams, control instruction configurations, safety analysis reports, manufacturer specifications, safety ratings, and standards. These data sources are crucial for analyzing failure component data and influencing component data. The selected equipment must be accompanied by sufficient data to achieve the required statistical confidence level, and the time scale must be greater than a complete equipment life cycle.
[0014] Data review and classification: Historical failure data is reviewed based on failure causes, failure modes, and detection methods to avoid data that invalidates the overall results. The reviewed historical failure data is divided into failure component data and impact component data. Failure component data includes failure causes, failure times, failure modes, and failure probabilities, while impact component data includes equipment attributes, environmental attributes, and maintenance activities. The impact component data type primarily includes equipment attributes, operating environment, and maintenance activities. Equipment attributes are used to describe equipment information related to manufacturer data and design characteristics.
[0015] Data preprocessing: To ensure that there is no invalid impact on the overall results, data preprocessing is required to eliminate some invalid data. For example, repeated failures caused by a specific problem can be deduplicated. In addition, the classification type of the equipment needs to be determined in advance based on expert advice to enable appropriate data grouping in the analysis. In some cases, there will be missing data, which can be filled by making reasonable assumptions. If information on the flow medium inside the valve is not provided, it can be assumed that valves installed in the same specific system share the same medium.
[0016] Furthermore, the data dimension reduction and failure probability modeling are carried out:
[0017] DU failure component data includes key factors influencing failure, such as failure causes and failure modes. Influencing component data primarily involves equipment attributes, environmental factors, and maintenance activities. Kernel principal component analysis (KPCA) is used to reduce the dimensionality of high-dimensional data and extract key features. This generates a low-dimensional principal component score matrix. This matrix not only reveals the correlation between DU failure component data and influencing component data but also provides simplified and accurate input features for subsequent failure probability modeling based on the generalized additive model (GAM).
[0018] Specifically, the influencing factors can be defined as a set of explanatory variables, expressed as X = [X1, X2, ... X n ] T Assume that there are m equipment samples used to describe the various influencing factors and the situation related to the DU failure state, and use '1' and '0' to represent the detection of DU failure, where the number '1' represents the detection of DU failure and '0' represents the non-detection. Then, the high-dimensional space matrix X is mapped to the low-dimensional principal component space to obtain the score matrix Z = [Z1, Z2…Z k ], the process is as follows: Use kernel principal component analysis (KPCA) to calculate the kernel matrix:
[0019]
[0020] Among them, the kernel matrix K(x i ,x j ) Calculate the sample x i and x j Similarity in high-dimensional space, δ is the kernel parameter, then the kernel matrix K is centered to remove the mean effect and obtain the centralized kernel matrix:
[0021]
[0022] Among them, 1 is a full The matrix is used to realize data centering in the kernel space, and the kernel matrix after centering Perform eigenvalue decomposition to obtain the eigenvalue λ i and the eigenvector υ i :
[0023]
[0024] Eigenvalue λ i Represents the weight of each principal component, the eigenvector υ i Then describe the direction of the principal component. In the high-dimensional kernel space, project each sample onto the kernel principal component to obtain the principal component score matrix Z:
[0025]
[0026] Here, Z i is the score of the i-th sample on the principal component, which is used to model the subsequent failure probability.
[0027] Combined with the principal component score matrix Z obtained in the above steps, the generalized additive model GAM is used to build a failure probability prediction model;
[0028] Specifically, the response variable Y=[Y1,Y2,…Y m ] T Represents the predicted target variable, which is the DU failure probability in this case. GAM is used to model the effect of each principal component on the failure probability Y. The basic form of the GAM model is as follows:
[0029]
[0030] Among them, α is the intercept term, f j (Z j ) is the score Z for the jth principal component j The smoothing function of GAM is used to capture the relationship between the principal component and the failure probability. ε is the error term. The smoothing function f of GAM is used to capture the relationship between the principal component and the failure probability. j (Z j ) can be further expressed as a spline basis expansion form:
[0031]
[0032] Among them, B m (Z j ) is the spline basis function, β j,m is the parameter to be estimated, the principal component Z j Contribution to the failure probability. Further, in order to optimize the GAM model and determine the optimal parameters of each smoothing function, the following loss function can be constructed and minimized:
[0033]
[0034] Among them, the first Represents the sum of squares of the prediction errors, which is used to minimize the model prediction deviation. The second term is a smoothing regularization term used to control the complexity of the smoothing function and prevent overfitting. λ is a regularization parameter that controls the smoothness. By optimizing the loss function, the optimal smoothing parameter β can be found. j,m and regularization parameter λ, thus obtaining a more stable and accurate GAM model.
[0035] Furthermore, the failure probability prediction:
[0036] The DU failure rate Y DU According to different failure modes, they are divided into i groups and the failure rate Y under different failure modes is calculated respectively. DU,i , which are grouped as follows:
[0037] Y DU =Y DU,1 +Y DU,2 …+Y DU,i
[0038] According to the kernel principal component analysis (KPCA) in step 2, the input high-dimensional data is reduced in dimension to obtain the score matrix Z of each sample in the kernel principal component space. Based on the principal component scores, the generalized additive model (GAM) is used to model the failure probability of each failure mode. For each failure mode i, the GAM modeling formula is:
[0039]
[0040] Finally, the failure probabilities of all failure modes are combined to obtain the overall failure probability of the new device; as shown below:
[0041]
[0042] Another object of the present invention is to provide a failure probability prediction system for a safety instrument system of a petroleum refining device, comprising:
[0043] The data acquisition and preprocessing module is used to demarcate the boundaries of the safety instrument system and collect equipment information and historical failure data of the safety instrument system of the oil refining unit to determine the failure mode and cause, and perform classification and preprocessing;
[0044] The data dimensionality reduction and failure probability modeling module is used to reduce the dimensionality of high-dimensional data and extract key features through KPCA. KPCA maps complex data into a low-dimensional principal component space, generating a principal component score matrix to capture the nonlinear correlation between DU failure and influencing factors. Based on this, GAM is combined to describe the relationship between principal component scores and DU failure probability, and a failure probability prediction model is established.
[0045] The failure probability prediction module is used to import the probability distribution, correlation and important influencing factors of DU failure nodes obtained by kernel principal component analysis after data dimensionality reduction and failure probability modeling. Then, GAM is used for modeling to predict the failure probability of newly connected equipment in the safety instrumented system, verify and interpret the prediction results, evaluate the accuracy and reliability of the failure probability, and calculate the failure probability of SIS equipment in the newly connected facility according to the failure rate prediction formula.
[0046] Another object of the present invention is to provide a computer device, which includes a memory and a processor, wherein the memory stores a computer program, and when the computer program is executed by the processor, the processor executes the steps of the failure probability prediction method for the safety instrument system of the petroleum refining device.
[0047] Another object of the present invention is to provide a computer-readable storage medium storing a computer program, which, when executed by a processor, causes the processor to execute the steps of the failure probability prediction method for a safety instrument system of a petroleum refining device.
[0048] Another object of the present invention is to provide an information data processing terminal, which is used to implement the failure probability prediction system for the safety instrument system of the petroleum refining device.
[0049] In combination with the above technical solutions and the technical problems solved, the advantages and positive effects of the technical solutions to be protected by the present invention are as follows:
[0050] First, the present invention provides a failure rate prediction method for the safety instrument system of a petroleum refining unit. In view of the shortcomings of the existing technology in processing nonlinear data and high-dimensional variables, the method combines the nonlinear dimensionality reduction capability of KPCA and the flexible modeling and interpretation capability of GAM to achieve accurate prediction of the DU probability of SIS equipment. The method uses KPCA to extract the nonlinear principal components of the failure data, fundamentally solving the limitations of traditional PCA in processing high-dimensional and nonlinear data, and can more comprehensively capture the complex characteristics of the DU failure mode. At the same time, by modeling the nonlinear relationship between the principal component score and the failure probability through GAM, an explanatory analysis of the influence of the principal component is provided while performing failure prediction, providing a scientific basis for the identification of important influencing factors and the optimization of the prediction model. The results show that the method has significant accuracy and reliability in the prediction of failure probability, solves the problem of insufficient processing capability of traditional statistical models for nonlinear data, and shows strong adaptability in the failure prediction of newly connected equipment. The specific description is as follows:
[0051] The present invention adopts a data-driven approach, namely, the identification of important influencing factors and the prediction of SIS equipment failure probability are completed through the kernel principal component analysis (KPCA) method and the generalized additive model (GAM). Among them, KPCA is a technology for simplifying data sets. It maps data to a high-dimensional space through a nonlinear kernel function, captures the nonlinear structure of the data in this space, and thus performs dimensionality reduction processing, while retaining the features that contribute the most to the variance and capturing important influencing factors. GAM aims to establish a nonlinear model that captures the nonlinear relationship between the principal component score and the target variable through a smoothing function. Compared with the traditional multiple regression method, GAM can adaptively adjust the influence of each principal component without explicit linear assumptions, thereby accurately predicting the relationship between the independent variable and the dependent variable. GAM uses the principal component scores obtained by KPCA for modeling, and seeks to maximize the influence of the principal component on the failure probability, and finally converts the original high-dimensional data set into a feature space that can capture the main nonlinear relationships. This invention combines a data-driven model based on empirical data with traditional statistical models. It can effectively identify the factors that have the greatest impact on the probability of SIS failure, predict the failure probability of newly connected SIS equipment, and preventively reinforce the potential failure behavior of old equipment, thereby ensuring the safety of the petroleum refining production process.
[0052] Second, the expected benefits and commercial value after the transformation of the technical solution of the present invention are as follows: The failure probability prediction method based on KPCA+GAM proposed in the present invention significantly improves the equipment safety and operating efficiency in industrial scenarios such as petroleum refining units by accurately predicting the DU failure probability of SIS equipment. This improvement directly reduces the equipment maintenance costs, unplanned downtime costs and economic losses caused by safety accidents, and improves overall operational efficiency. In addition, the core algorithm of this technology is universal and can be extended to fields such as electricity, aviation, and medical care to achieve accurate prediction of failures of other equipment in complex systems, and has broad commercial potential. In particular, with the increasing attention paid to the safety of high-reliability industrial equipment in China, this method has significant market promotion prospects and economic value.
[0053] The technical solution of the present invention fills the technical gap in the industry at home and abroad: traditional failure probability prediction methods are mostly based on linear statistical models (such as PLSR), which have obvious limitations when processing high-dimensional and nonlinear data, especially for the prediction of DU modes in complex industrial systems. There is a lack of reliable and accurate solutions. The present invention combines the nonlinear dimensionality reduction capability of KPCA and the flexible modeling capability of GAM to propose a new failure probability prediction framework, which fundamentally solves the problem that traditional methods are difficult to capture complex nonlinear modes, and fills the technical gap in the accurate prediction of nonlinear DU failure modes in complex industrial scenarios. It has important scientific value and practical significance.
[0054] Third, in the safety instrument systems (SIS) of petroleum refining plants, the complexity of the equipment and the diversity of operating conditions make failure prediction a significant challenge. Traditional failure prediction methods rely primarily on simple statistical models or expert experience, which cannot fully capture the nonlinear relationship between equipment failure modes and multidimensional influencing factors. Furthermore, existing methods have limited processing capabilities for high-dimensional data, making it difficult to extract key features, resulting in insufficient prediction accuracy and making it difficult to meet the demand for real-time early warning and accurate prediction of equipment failures in actual industrial scenarios. This deficiency directly impacts the safety and reliability of refining plants, increasing maintenance costs and safety risks.
[0055] To address the shortcomings of the existing technology, the present invention proposes a failure probability prediction method based on kernel principal component analysis (KPCA) and generalized additive models (GAM). KPCA is used to reduce the dimensionality of high-dimensional data, extract key features related to failure modes, and utilize the GAM model to describe the nonlinear relationship between principal component scores and failure probabilities. This effectively addresses the inability of traditional methods to handle complex nonlinear relationships. Furthermore, the present invention can dynamically analyze failure modes and predict the failure probability of newly connected equipment in real time, providing a reliable basis for monitoring equipment operating status and optimizing maintenance strategies.
[0056] This invention has achieved several technological advancements in industrial applications: (1) Data dimensionality reduction and nonlinear modeling have significantly improved the accuracy and robustness of failure prediction; (2) A dynamic failure probability prediction method that can adapt to different equipment characteristics has been established to meet the diverse industrial scenario requirements of petroleum refining units; (3) The introduction of intelligent modeling technology has enabled the automatic processing of high-dimensional complex data and the extraction of key features, reducing reliance on expert experience and improving prediction efficiency and real-time performance. These technological advancements have significantly enhanced the early warning capabilities of SIS systems and reduced the safety risks caused by equipment failures.
[0057] This invention provides an efficient and reliable failure prediction tool for the safety management of petroleum refining equipment. Its advanced modeling methods and data processing capabilities not only optimize equipment operation and maintenance processes, but also reduce the risk of downtime caused by equipment failures, improving the overall safety and economic benefits of the system. Furthermore, this invention provides a general technical framework for the intelligent monitoring of industrial equipment and can be widely applied to the safety management of other complex industrial systems. It has significant industrial value and widespread application prospects, providing important support for the digital transformation and intelligent upgrading of the petrochemical industry. BRIEF DESCRIPTION OF THE DRAWINGS
[0058] Figure 1 This is a flow chart of a failure probability prediction method for a safety instrument system of a petroleum refining device provided by an embodiment of the present invention.
[0059] Figure 2 This is a structural block diagram of a failure probability prediction system for a safety instrument system of a petroleum refining unit provided by an embodiment of the present invention.
[0060] Figure 3 This is a failure probability prediction method for a safety instrument system of a petroleum refining unit and an overall system flow chart provided by an embodiment of the present invention;
[0061] Figure 4 This is a schematic diagram of data collection and preprocessing provided by an embodiment of the present invention;
[0062] Figure 5 This is a schematic diagram of data dimensionality reduction and failure probability modeling provided by an embodiment of the present invention;
[0063] Figure 6 This is a schematic diagram of failure probability prediction provided by an embodiment of the present invention;
[0064] Figure 7 It is a brief schematic diagram of the process of evaluating the explanatory power of the model provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0065] In order to make the purpose, technical solutions and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with the embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.
[0066] like Figure 1 As shown, the embodiment of the present invention provides a failure probability prediction method for a safety instrument system of a petroleum refining device, comprising the following steps:
[0067] S101, data collection and preprocessing: demarcate the safety instrument system boundaries, collect equipment information and historical failure data of the safety instrument system of the oil refining unit to determine the failure mode and cause, and perform classification and preprocessing;
[0068] S102, Data Dimensionality Reduction and Failure Probability Modeling: KPCA is used to reduce the dimensionality of high-dimensional data and extract key features. KPCA maps complex data into a low-dimensional principal component space, generating a principal component score matrix to capture the nonlinear correlation between DU failure and influencing factors. Based on this, GAM is used to describe the relationship between principal component scores and DU failure probability, and a failure probability prediction model is established.
[0069] S103, failure probability prediction, after data dimensionality reduction and failure probability modeling, imports the probability distribution, correlation and important influencing factors of DU failure nodes obtained by kernel principal component analysis, and then uses GAM modeling to predict the failure probability of newly connected equipment in the safety instrumented system. The prediction results are verified and interpreted, the accuracy and reliability of the failure probability are evaluated, and the failure probability of SIS equipment in the newly connected facility is calculated according to the failure rate prediction formula.
[0070] This method first defines the boundaries of the safety instrument system (SIS) within a petroleum refinery unit and identifies the equipment and subsystems within it. By collecting operational information, historical failure records, and device attribute data from relevant equipment, common failure modes and their causes are identified. The collected raw data is then categorized and sorted, missing values are filled, outliers are eliminated, and multi-source data is uniformly formatted. The goal of data preprocessing is to generate high-quality, standardized input data, laying the foundation for subsequent dimensionality reduction and modeling.
[0071] Because failure data from SIS equipment typically has complex characteristics such as high dimensionality, multivariate, and nonlinearity, direct modeling makes it difficult to extract key features. This method uses kernel principal component analysis (KPCA) to reduce the dimensionality of high-dimensional data and map complex nonlinear relationships into a low-dimensional principal component space. The principal component score matrix generated by KPCA can effectively capture the nonlinear correlation between failure modes and influencing factors. On this basis, a generalized additive model (GAM) is used to model the relationship between principal component scores and DU failure probability, and a failure probability prediction model is constructed. GAM approximates complex relationships through flexible nonlinear functions, enabling the model to accurately describe the dynamic relationship between failure probability and influencing factors.
[0072] After completing data dimensionality reduction and modeling, the principal component score matrix generated by KPCA, the probability distribution of DU failure nodes, and correlation information were imported to predict the failure probability of newly connected devices using the GAM model. During the prediction process, the model calculates the DU failure probability of the device based on its historical data and current operating status and generates a prediction result. Simultaneously, the model output is verified and interpreted. By comparing it with known failure data, the accuracy and reliability of the prediction are evaluated, and the mechanisms of the main influencing factors are analyzed.
[0073] Based on the predicted failure probabilities and the failure rate prediction formula, the specific failure probabilities of key SIS devices in the newly connected equipment are calculated. The prediction results can be used to proactively identify devices with high failure risks, support the optimization of preventive maintenance strategies for the SIS system, and improve the overall operational safety and reliability of the system. Through feedback and iteration of the prediction process, the parameters of KPCA and GAM are further optimized, improving the model's generalization capabilities and adapting it to a wider range of industrial scenarios, providing reliable technical support for the safe operation of petroleum refining units.
[0074] The data collection and preprocessing provided by the embodiment of the present invention include:
[0075] Data Collection: Data sources for collection include equipment maintenance records, safety standards and specifications, process and instrumentation diagrams, control instruction configurations, safety analysis reports, manufacturer specifications, safety ratings, and standards. These data sources are crucial for analyzing failure component data and influencing component data. The selected equipment must be accompanied by sufficient data to achieve the required statistical confidence level, and the time scale must be greater than a complete equipment life cycle.
[0076] Data review and classification: Historical failure data is reviewed based on failure causes, failure modes, and detection methods to avoid data that invalidates the overall results. The reviewed historical failure data is divided into failure component data and impact component data. Failure component data includes failure causes, failure times, failure modes, and failure probabilities, while impact component data includes equipment attributes, environmental attributes, and maintenance activities. The impact component data type primarily includes equipment attributes, operating environment, and maintenance activities. Equipment attributes are used to describe equipment information related to manufacturer data and design characteristics.
[0077] Data preprocessing: To ensure that there is no invalid impact on the overall results, data preprocessing is required to eliminate some invalid data. For example, repeated failures caused by a specific problem can be deduplicated. In addition, the classification type of the equipment needs to be determined in advance based on expert advice to enable appropriate data grouping in the analysis. In some cases, there will be missing data, which can be filled by making reasonable assumptions. If information on the flow medium inside the valve is not provided, it can be assumed that valves installed in the same specific system share the same medium.
[0078] Data dimensionality reduction and failure probability modeling provided by the embodiment of the present invention:
[0079] DU failure component data includes key factors influencing failure, such as failure causes and failure modes. Influencing component data primarily involves equipment attributes, environmental factors, and maintenance activities. Kernel principal component analysis (KPCA) is used to reduce the dimensionality of high-dimensional data and extract key features. This generates a low-dimensional principal component score matrix. This matrix not only reveals the correlation between DU failure component data and influencing component data but also provides simplified and accurate input features for subsequent failure probability modeling based on the generalized additive model (GAM).
[0080] Specifically, the influencing factors can be defined as a set of explanatory variables, expressed as X = [X1, X2, ... X n ] T Assume that there are m equipment samples used to describe the various influencing factors and the situation related to the DU failure state, and use '1' and '0' to represent the detection of DU failure, where the number '1' represents the detection of DU failure and '0' represents the non-detection. Then, the high-dimensional space matrix X is mapped to the low-dimensional principal component space to obtain the score matrix Z = [Z1, Z2…Z k ], the process is as follows: Use kernel principal component analysis (KPCA) to calculate the kernel matrix:
[0081]
[0082] Among them, the kernel matrix K(x i ,x j ) Calculate the sample x i and x j Similarity in high-dimensional space, δ is the kernel parameter, then the kernel matrix K is centered to remove the mean effect and obtain the centralized kernel matrix:
[0083]
[0084] Among them, 1 is a full The matrix is used to realize data centering in the kernel space, and the kernel matrix after centering Perform eigenvalue decomposition to obtain the eigenvalue λ i and the eigenvector υ i :
[0085]
[0086] Eigenvalue λ i Represents the weight of each principal component, the eigenvector υ i Then describe the direction of the principal component. In the high-dimensional kernel space, project each sample onto the kernel principal component to obtain the principal component score matrix Z:
[0087]
[0088] Here, Zi is the score of the i-th sample on the principal component, which is used to model the subsequent failure probability.
[0089] Combined with the principal component score matrix Z obtained in the above steps, the generalized additive model GAM is used to build a failure probability prediction model;
[0090] Specifically, the response variable Y=[Y1,Y2,…Y m ] T Represents the predicted target variable, which is the DU failure probability in this case. GAM is used to model the effect of each principal component on the failure probability Y. The basic form of the GAM model is as follows:
[0091]
[0092] Among them, α is the intercept term, f j (Z j ) is the score Z for the jth principal component j The smoothing function of GAM is used to capture the relationship between the principal component and the failure probability. ε is the error term. The smoothing function f of GAM is used to capture the relationship between the principal component and the failure probability. j (Z j ) can be further expressed as a spline basis expansion form:
[0093]
[0094] Among them, B m (Z j ) is the spline basis function, β j,m is the parameter to be estimated, the principal component Z j Contribution to the failure probability. Further, in order to optimize the GAM model and determine the optimal parameters of each smoothing function, the following loss function can be constructed and minimized:
[0095]
[0096] Among them, the first Represents the sum of squares of the prediction errors, which is used to minimize the model prediction deviation. The second term is a smoothing regularization term used to control the complexity of the smoothing function and prevent overfitting. λ is a regularization parameter that controls the smoothness. By optimizing the loss function, the optimal smoothing parameter β can be found. j,m and regularization parameter λ, thus obtaining a more stable and accurate GAM model.
[0097] Failure probability prediction provided by the embodiment of the present invention:
[0098] The DU failure rate Y DU According to different failure modes, they are divided into i groups and the failure rate Y under different failure modes is calculated respectively. DU,i, which are grouped as follows:
[0099] Y DU =Y DU,1 +Y DU,2 …+Y DU,i
[0100] According to the kernel principal component analysis (KPCA) in step 2, the input high-dimensional data is reduced in dimension to obtain the score matrix Z of each sample in the kernel principal component space. Based on the principal component scores, the generalized additive model (GAM) is used to model the failure probability of each failure mode. For each failure mode i, the GAM modeling formula is:
[0101]
[0102] Finally, the failure probabilities of all failure modes are combined to obtain the overall failure probability of the new device; as shown below:
[0103]
[0104] like Figure 2 As shown, an embodiment of the present invention provides a failure probability prediction system for a safety instrument system of a petroleum refining device, including:
[0105] The data acquisition and preprocessing module is used to demarcate the boundaries of the safety instrument system and collect equipment information and historical failure data of the safety instrument system of the oil refining unit to determine the failure mode and cause, and perform classification and preprocessing;
[0106] The data dimensionality reduction and failure probability modeling module is used to reduce the dimensionality of high-dimensional data and extract key features through KPCA. KPCA maps complex data into a low-dimensional principal component space, generating a principal component score matrix to capture the nonlinear correlation between DU failure and influencing factors. Based on this, GAM is combined to describe the relationship between principal component scores and DU failure probability, and a failure probability prediction model is established.
[0107] The failure probability prediction module is used to import the probability distribution, correlation and important influencing factors of DU failure nodes obtained by kernel principal component analysis after data dimensionality reduction and failure probability modeling. Then, GAM is used for modeling to predict the failure probability of newly connected equipment in the safety instrumented system, verify and interpret the prediction results, evaluate the accuracy and reliability of the failure probability, and calculate the failure probability of SIS equipment in the newly connected facility according to the failure rate prediction formula.
[0108] Another object of the present invention is to provide a computer device, which includes a memory and a processor, wherein the memory stores a computer program, and when the computer program is executed by the processor, the processor executes the steps of the failure probability prediction method for the safety instrument system of the petroleum refining device.
[0109] Another object of the present invention is to provide a computer-readable storage medium storing a computer program, which, when executed by a processor, causes the processor to execute the steps of the failure probability prediction method for a safety instrument system of a petroleum refining device.
[0110] Another object of the present invention is to provide an information data processing terminal, which is used to implement the failure probability prediction system for the safety instrument system of the petroleum refining device.
[0111] The present invention is specifically implemented:
[0112] The data-driven KPCA and GAM models described in this paper can reduce the dimensionality of collected multidimensional data and analyze correlations between them. This allows for more accurate prediction of the failure rates of SIS equipment in new facilities, enabling developers to build systems that meet safety integrity level requirements when needed. Furthermore, because this model is a general one, it can be used not only for predicting the failure rates of oil refining safety instrumentation systems, but also for other equipment.
[0113] The following will be combined with the accompanying drawings, taking the shut-off valve in the safety instrument system (SIS) as a specific example to explain in detail the technical solution and effects of the present invention. Among them, the SIS control strategy is "three-choose-one", that is, three SIS loops jointly control a shut-off valve, and when the sensor acquisition data of any of the SIS loops exceeds the safety threshold, the shut-off valve will immediately operate. It should be noted that in order to prevent SIS from malfunctioning, some manufacturers adopt a "three-choose-two" strategy when configuring the system configuration, that is, when the sensor acquisition data of at least two SIS loops exceeds the safety threshold, the actuator will only operate. Although the failure probability of SIS equipment obtained under different control strategies is different, it can be processed through data normalization, so it does not affect the feasibility of the method proposed in this patent for predicting the failure probability of SIS. Figure 3 The overall failure rate prediction process is shown as follows:
[0114] Step 1: Collect and process data of existing comparable equipment for the specific equipment to be predicted. Figure 4 A schematic diagram of the data processing system is shown, which includes:
[0115] Obtain historical failure data for SIS equipment from fault notifications and maintenance records, including failure component data and influencing component data. Failure component data includes failure cause, failure time, failure mode, and failure probability. This data is sourced from safety standards and specifications, control instructions and configurations, safety manuals and safety analysis reports, and manufacturer specifications. Influencing component data includes equipment attributes, environmental attributes, and maintenance activities. This data is sourced from maintenance notifications, process and instrumentation diagrams, safety ratings, and relevant standards. Furthermore, when collecting historical failure data, discussions with technical advisors and process engineers are necessary to gain a more comprehensive understanding and document this data. The collected data needs to be preprocessed, including data cleaning, missing data filling, and failure probability normalization under different control strategies. First, data cleaning is performed to remove duplicate, invalid or inconsistent data and unify the data format. Then, missing data is filled through expert experience, interpolation and other methods. Next, failure probability normalization is performed according to different control strategies and operating conditions. For example, the failure probability of a certain SIS device is 0.08 under high load and 0.03 under low load. In order to make a unified comparison, it can be normalized to a standard operating condition (such as average load) to obtain a new failure probability for subsequent analysis. Common normalization methods include maximum-minimum normalization or Z-score normalization. Finally, the collected SIS device failure history data is organized and recorded to ensure that the failure mode, failure cause, failure probability and its influencing components of each device are archived in detail. Taking the shut-off valve as an example: the DU failure mode of the shut-off valve mainly includes three types: fail to close (FTC), leakage in closed position (LCP) and delayed operation (LCP). After detailed records of the data of these three types of failure modes are recorded, they will be helpful for subsequent safety assessment, maintenance optimization and system upgrade decisions. The failure record examples are shown in Table 1:
[0116] Table 1 Example of failure component data
[0117]
[0118]
[0119] (2) In order to avoid invalid effects on the overall results, repeated failure data caused by some specific problems should be deleted during data preprocessing; in addition, the equipment classification type needs to be determined in advance according to expert recommendations to ensure the accuracy and reliability of subsequent analysis; the equipment attributes, environmental attributes and maintenance activities of the shut-off valve are shown in Table 2, Table 3, and Table 4 below. Since equipment attributes have been shown to be very important in explaining the differences in the empirical reliability performance of SIS equipment, the main focus was on equipment attributes when collecting data;
[0120] Table 2 Example of equipment attributes for shut-off valves
[0121]
[0122]
[0123] Table 3 Examples of environmental properties of shut-off valves
[0124]
[0125] Table 4 Examples of maintenance activities for shut-off valves
[0126]
[0127] Step 2: Use the data-driven module to perform data dimension reduction and failure probability model building. Figure 5 The derivation process of data dimensionality reduction and failure probability modeling is demonstrated, including:
[0128] In the analysis, each of the above influencing factors is defined as an explanatory variable, and its symbol is represented by X = [X1, X2, ... X n ] T Here, X is a column vector containing the values of various influencing factors. The shut-off valve is considered as a device sample and is used to describe the relationship between various influencing factors and DU failure status. In this scenario, '1' represents a detected DU failure and '0' represents an undetected DU failure. These data are distributed in the variable space and used to build a model to further explore the relationship between the influencing factors and DU failure status.
[0129] KPCA is used to determine the overall correlation between DU failure and different influencing factors. Here, the Gaussian kernel function is used to calculate the kernel matrix K, and the kernel matrix is obtained as follows:
[0130]
[0131] Among them, δ is the kernel parameter (), and then the kernel matrix K is centered to obtain the centered kernel matrix:
[0132]
[0133] Then, the centered kernel matrix Perform eigenvalue decomposition to obtain the eigenvalue λ i and the eigenvector υ i , through the projection operation, the principal component score matrix Z can be obtained:
[0134]
[0135] Eigenvalue λ i The size of reflects the amount of information contained in the principal component, that is, it is related to the score of each influencing factor. The larger the eigenvalue, the higher the score and the more information it carries, that is, the greater the influence. By selecting the principal components with larger eigenvalues, the data dimension can be reduced, and the most influential information for failure probability modeling can be retained.
[0136] GAM modeling is used to further analyze the nonlinear relationship between the principal component Z and the DU failure probability Y. The basic form of the GAM model is:
[0137]
[0138] f j (Z j ) is a smooth function for the jth principal component, by fitting f j (Z j ) can explain the nonlinear contribution of principal components to failure probability and identify important influencing factors;
[0139] Finally, various evaluation metrics can be used to assess the performance and goodness of fit of the GAM model, such as the root mean square error (RMSE) and the coefficient of determination (R 2 ); the root mean square error (RMSE) is used to measure the deviation between the predicted value and the actual value, and its formula is as follows:
[0140]
[0141] Where n is the number of samples, y i is the actual value, The predicted value is the smaller the calculated RMSE is, the smaller the prediction error of the model is and the higher the prediction accuracy of the model is. When the RMSE is 0, it means that the model prediction is completely consistent with the actual value.
[0142] Coefficient of determination (R 2 ) is used to evaluate the explanatory power of the model, that is, how well the model fits the data. The formula is as follows:
[0143]
[0144] in, is the residual sum of squares of the model (SS_res), is the total sum of squares (SS_tot), is the mean of the actual values, R 2 The value range of R is [0,1]. The closer to 1, the better the model fit is, and the closer to 0, the worse the model fit is. 2 =1 means that the model can perfectly explain the variation of the data, while R 2 = 0 means that the model fails to explain any data changes. Since this step is mainly to evaluate whether the established model meets the requirements, after considering the advantages and disadvantages of the two, the present invention uses R 2 To evaluate the accuracy and explanatory power of model predictions, the brief flowchart is as follows Figure 7 shown.
[0145] Step 3: Combine the results obtained from the data-driven module and the safety integrity level (SIL) requirements to predict the case failure rate. The steps are as follows: Figure 6 As shown, its contents include:
[0146] DU failures are subdivided into the three failure modes discussed: DOP, FTC, and LCP. For the data of each failure mode, KPCA and GAM are used to analyze the failure rate of each mode. Make a prediction:
[0147]
[0148] Table 5 below lists examples of DU failures and corresponding failure rates for each failure mode of the shut-off valve, where OTH stands for other or unknown failure mode;
[0149] Table 5 Examples of failure distribution and corresponding failure rates
[0150]
[0151] As described in the previous steps, when collecting data, we focused on equipment attributes. The principal component score matrix extracted by KPCA and the GAM modeling results can be used to evaluate the contribution of different principal components to the failure rate. The GAM model directly generates the predicted value of each failure mode by capturing the nonlinear relationship between the principal components and the failure probability. The formula is as follows:
[0152]
[0153] The above formula takes into account the operating conditions under different DU failure modes. That is, data from different DU failure modes is extracted to predict the failure rates of different failure modes and then integrated. This helps to accurately identify the characteristics of each mode, avoid data mixing, and improve prediction accuracy; at the same time, it reduces the complexity of the model and the risk of data overfitting. In addition, this method enhances the adaptability and flexibility of the model. It can select appropriate variables and modeling strategies for different modes, facilitate the decomposition of the total failure rate, clarify the contribution of each mode, and provide a clearer basis for failure management and optimization decisions, thereby formulating more focused and effective improvement measures to meet the needs of complex industrial scenarios.
[0154] Example 1:
[0155] Failure prediction of safety instrument systems in petrochemical plants
[0156] At a large petrochemical plant, the Safety Instrumented System (SIS) controls and monitors the operating status of critical equipment to ensure safe and stable production. Due to the long-term exposure of these equipment to harsh, high-temperature, and high-pressure environments, some equipment frequently fails, potentially leading to serious safety incidents. To mitigate these failures, the plant implemented the aforementioned failure probability prediction method.
[0157] Step 1: Data collection and preprocessing
[0158] The factory collected a large amount of data including equipment maintenance records, historical failure data, and safety standards and specifications, and preprocessed this data to eliminate invalid and redundant data to ensure the integrity and accuracy of the model input data.
[0159] Step 2: Data dimensionality reduction and failure probability model building
[0160] Kernel principal component analysis (KPCA) is used to reduce the dimensionality of preprocessed high-dimensional data and extract key features (such as principal component scores) that influence failure probability. Subsequently, a generalized additive model (GAM) is used to construct a failure probability prediction model. This model captures nonlinear relationships in the data and describes the association between principal component features and failure probability, providing a basis for quantitative modeling of equipment failure risk.
[0161] Step 3: Failure probability prediction
[0162] Based on the modeling results, the plant predicted the failure probability of newly installed critical equipment (such as emergency shutoff valves and pressure controllers) under different operating conditions. By continuously verifying and adjusting the model, the plant predicted the failure risk of these devices under different operating conditions, and implemented preventive maintenance accordingly, significantly reducing the incidence of equipment failures.
[0163] Example 2:
[0164] Failure Probability Prediction in Natural Gas Pipeline Safety Monitoring System
[0165] In a natural gas pipeline system, a safety instrument system monitors pipeline pressure, flow, and temperature in real time to ensure the safety of natural gas transportation. To prevent major safety incidents caused by pipeline failures (such as ruptures and leaks), the gas transmission company deployed this failure probability prediction method to improve equipment reliability.
[0166] Step 1: Data collection and preprocessing
[0167] The pipeline company collected historical operating data and sensor data from gas transmission equipment, including pressure change records, temperature sensor data, flow data, maintenance records, and fault history, and classified and preprocessed this data to remove noise data and duplicate records.
[0168] Step 2: Data dimensionality reduction and failure probability model building
[0169] KPCA was used to reduce the dimensionality of the preprocessed high-dimensional data and extract the key features that influence pipeline failure (such as pipeline age, corrosiveness of the transported medium, and pressure fluctuation amplitude). Subsequently, a GAM was used to establish a failure probability prediction model, capturing the nonlinear relationship between characteristic data and failure probability, providing a key basis for accurate prediction of pipeline failure.
[0170] Step 3: Failure probability prediction
[0171] Based on the established predictive model, the gas transmission company accurately predicted the probability of pipeline failure under various pressure fluctuations and temperature conditions. The model accurately predicted the risk of pipeline failure under these conditions. By optimizing pipeline maintenance plans and proactively repairing high-risk sections, the risk of pipeline leaks and ruptures was significantly reduced.
[0172] These two examples demonstrate how to use this failure probability prediction method in the fields of petrochemicals and natural gas pipelines to prevent safety accidents and improve system reliability and safety.
[0173] 2. Relevant evidence of the technical effects obtained by the embodiments of the present invention.
[0174] KPCA's dimensionality reduction effect: KPCA can extract key features from high-dimensional nonlinear data, reducing complex input data to a low-dimensional principal component score matrix, while preserving key information and eliminating redundant data. This dimensionality reduction improves the computational efficiency of subsequent modeling and reduces the risk of overfitting.
[0175] GAM's nonlinear modeling capabilities: GAM can use smooth functions to capture the nonlinear relationship between principal components and failure probability, avoiding errors caused by linear assumptions.
[0176] Effect of mode-based prediction: DU failure is subdivided into multiple failure modes and modeled separately to avoid the impact of data congestion on prediction accuracy and reduce modeling complexity.
[0177] Nonlinear feature extraction: KPCA's nonlinear dimensionality reduction capability can more accurately capture DU failure characteristics.
[0178] Flexibility and adaptability: The modeling flexibility of GAM significantly improves the accuracy of failure probability prediction.
[0179] Highly targeted: Mode-specific modeling provides a clearer basis for failure management and optimization.
[0180] It should be noted that the embodiments of the present invention can be implemented by hardware, software, or a combination of software and hardware. The hardware portion can be implemented using dedicated logic; the software portion can be stored in a memory and executed by an appropriate instruction execution system, such as a microprocessor or dedicated design hardware. Those skilled in the art will appreciate that the above-mentioned devices and methods can be implemented using computer-executable instructions and / or contained in processor control code, for example, such as a carrier medium such as a disk, CD or DVD-ROM, a programmable memory such as a read-only memory (firmware), or a data carrier such as an optical or electronic signal carrier. The devices and modules of the present invention can be implemented by hardware circuits such as very large-scale integrated circuits or gate arrays, semiconductors such as logic chips, transistors, or programmable hardware devices such as field programmable gate arrays, programmable logic devices, etc., can also be implemented by software executed by various types of processors, or can be implemented by a combination of the above-mentioned hardware circuits and software, such as firmware.
[0181] The above description is only a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any modifications, equivalent substitutions and improvements made by any technician familiar with this technical field within the technical scope disclosed by the present invention and within the spirit and principles of the present invention should be covered by the scope of protection of the present invention.
Claims
1. A failure probability prediction method for a safety instrument system of a petroleum refining unit, characterized by: The method comprises the following steps: Step 1, Data Collection and Preprocessing: Delineate the safety instrumented system boundaries, collect equipment information and historical failure data, determine failure modes and causes, and classify and preprocess the data; Step 2: Data Dimensionality Reduction and Failure Probability Modeling: KPCA is used to reduce the dimensionality of high-dimensional data, extract key features, and generate a principal component score matrix. Based on this, GAM is used to describe the relationship between the principal component scores and DU failure probability, and a failure probability prediction model is established. Step 3, Failure Probability Prediction: Based on the KPCA and GAM analysis results, predict the failure probability of the newly connected device and calculate the failure probability according to the failure rate prediction formula; The data dimensionality reduction and failure probability modeling uses KPCA to identify the nonlinear correlation of failure component data and calculates the principal component score matrix through the kernel matrix, thereby screening out low-dimensional features that have a significant impact on the equipment. GAM analyzes the nonlinear relationship between the failure probability of safety instrumented system equipment and the principal component characteristics, and captures the quantitative contribution of each principal component to the failure probability through a smoothing function.
2. The failure probability prediction method for a safety instrument system of a petroleum refining device according to claim 1 is characterized in that: The data collection and preprocessing includes data from equipment maintenance records, safety standards and specifications, process and instrumentation diagrams, control instruction configurations, safety analysis reports and manufacturer specifications, and the time scale is greater than a complete equipment life cycle.
3. The failure probability prediction method for a safety instrument system of a petroleum refining device according to claim 1 is characterized in that: Based on the prediction results of different failure modes, the failure probability of each mode is integrated to calculate the overall failure probability of the equipment, and the accuracy and reliability of the prediction are improved through verification and adjustment.
4. A failure probability prediction system for a safety instrument system of a petroleum refinery device implementing the failure probability prediction method for a safety instrument system of a petroleum refinery device as claimed in any one of claims 1 to 3, characterized in that: The failure probability prediction system for the safety instrument system of the petroleum refining device includes: The data acquisition and preprocessing module is used to demarcate the boundaries of the safety instrument system and collect equipment information and historical failure data of the safety instrument system of the oil refining unit to determine the failure mode and cause, and perform classification and preprocessing; The data dimensionality reduction and failure probability modeling module uses KPCA to reduce the dimensionality of high-dimensional data and extract key features. KPCA maps complex data into a low-dimensional principal component space, generating a principal component score matrix to capture the nonlinear correlation between DU failure and influencing factors. Based on this, GAM is combined to describe the relationship between principal component scores and DU failure probability, and a failure probability prediction model is established. The failure probability prediction module is used to import the probability distribution, correlation and important influencing factors of DU failure nodes obtained by kernel principal component analysis after data dimensionality reduction and failure probability modeling. Then, GAM is used for modeling to predict the failure probability of newly connected equipment in the safety instrumented system, verify and interpret the prediction results, evaluate the accuracy and reliability of the failure probability, and calculate the failure probability of SIS equipment in the newly connected facility according to the failure rate prediction formula.
5. A computer device, characterized in that: The computer device includes a memory and a processor, the memory stores a computer program, and when the computer program is executed by the processor, the processor executes the steps of the failure probability prediction method for a safety instrument system of a petroleum refining device as described in any one of claims 1 to 3.
6. A computer-readable storage medium storing a computer program, wherein when the computer program is executed by a processor, the processor executes the steps of the failure probability prediction method for a safety instrument system of a petroleum refining device according to any one of claims 1 to 3.
7. An information data processing terminal, characterized in that: The information data processing terminal is used to implement the failure probability prediction system for the safety instrument system of the petroleum refining device as claimed in claim 4.
Citation Information
Patent Citations
Method and system for predicting failure of hydrogen-doped natural gas pipeline based on uncertainty perception
CN117172095A
Predictive Model Data Stream Prioritization
US20230123322A1