Method and device for determining fault source of hydrogen production equipment, equipment and medium
The principal component analysis model is used to utilize the operating data matrix of the alkaline electrolytic cell to determine the contribution rate of the characteristic variables to the fault, which solves the problems of low accuracy and efficiency in determining the fault source of the alkaline electrolytic cell in the existing technology and realizes efficient and accurate fault source positioning.
Patent Information
- Application Number
- CN202510898066.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-30
- Publication Date
- 2025-09-12
AI Technical Summary
Existing methods for determining the source of alkaline electrolyzer faults rely on high-precision fault operation data matrices, resulting in low accuracy and efficiency. Non-model-based methods are susceptible to noise interference or require additional hardware, increasing costs and introducing measurement errors.
The principal component analysis model is adopted to obtain the operating data matrix of the alkaline electrolyzer to be tested. The contribution rate of each characteristic variable to the fault is determined using the T2 statistic and the SPE statistic. The fault source is determined using only the existing sensor data.
It avoids hardware modification costs and measurement errors, improves the accuracy and efficiency of fault source determination, and circumvents the problem of low modeling accuracy caused by scarce fault data.
Smart Images

Figure CN120625112A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of water electrolysis hydrogen production equipment, and in particular to a method, device, equipment and medium for determining a fault source of hydrogen production equipment. Background Art
[0002] The volatility of renewable energy can easily cause operational failures in alkaline electrolyzers, even leading to safety hazards. Therefore, improving the efficiency of identifying the source of alkaline electrolyzer failures has become a core requirement for renewable energy-coupled water electrolysis hydrogen production technology.
[0003] Currently, methods for determining the source of alkaline electrolytic cell faults are primarily categorized into model-based and non-model-based methods. Model-based methods require inputting the alkaline electrolytic cell's fault operating data matrix and the fault source into a machine learning model. By minimizing the loss between the model's output and the actual fault source, the machine learning model is trained to obtain a fault diagnosis model. In practice, the test data matrix can then be directly input into the fault diagnosis model to determine the fault source. However, this method relies on a large amount of high-precision fault operating data matrices, which are often scarce and have limited accuracy in real-world applications, resulting in low accuracy and efficiency in fault source determination. Non-model-based methods for determining the source of faults include signal processing methods (such as fast Fourier transform) and electrochemical testing methods (such as electrochemical impedance spectroscopy). However, signal processing methods are highly susceptible to noise, similarly resulting in low accuracy in fault source determination. Electrochemical testing methods require additional testing equipment, which not only increases hardware costs but also may introduce measurement errors.
[0004] Therefore, how to obtain a method for determining the source of a fault that does not require additional hardware and is highly accurate and efficient has become a technical problem that needs to be solved urgently. Summary of the Invention
[0005] Based on the above problems, the present application provides a method, apparatus, device and medium for determining the fault source of hydrogen production equipment, which can perform fault source determination with high accuracy and efficiency without the need for additional hardware.
[0006] The embodiments of this application disclose the following technical solutions:
[0007] In a first aspect, the present application discloses a method for determining a fault source of a hydrogen production device, the method comprising:
[0008] Get the alkaline electrolytic cell's operating data matrix X to be tested n×m , where n is the number of sampling moments, m is the number of characteristic variables, and both m and n are integers greater than 1;
[0009] The running data matrix to be tested is input into the principal component analysis model to obtain T 2 Statistics and SPE statistics, wherein the principal component analysis model is obtained by training the normal operation data matrix of the alkaline electrolytic cell;
[0010] If the T 2 If the statistical value is greater than the first threshold and / or the SPE statistical value is greater than the second threshold, then determining the contribution rate of each characteristic variable to the fault;
[0011] The characteristic variable with the highest contribution rate is determined as the fault source.
[0012] Optionally, determining the contribution rate of each characteristic variable to the fault includes:
[0013] Determine a principal component score vector at the time of the fault, wherein the principal component score vector includes k principal component scores, where k is a positive integer less than m;
[0014] Determining the principal component direction of the fault according to the k principal component scores;
[0015] For each of the fault principal component directions, calculating the local contribution rate of the characteristic variable to the fault;
[0016] The local contribution rates of all the fault principal component directions are summed to obtain the contribution rate of each characteristic variable to the fault.
[0017] Optionally, the first threshold is determined by the following method:
[0018] Obtain the normal operation data matrix X of the alkaline electrolyzer n×m , where n is the number of sampling moments, m is the number of characteristic variables, and both m and n are integers greater than 1;
[0019] Performing principal component analysis and dimensionality reduction on the normal operating data matrix to obtain a normal operating data matrix after dimensionality reduction;
[0020] According to the normal operation data matrix after dimension reduction, the first threshold is determined by the following formula:
[0021] ;
[0022] Among them, T a is the first threshold, It follows the F distribution with the first degree of freedom k and the second degree of freedom nk, and 1-α is the confidence level.
[0023] Optionally, the second threshold is determined by the following method:
[0024] According to the normal operation data matrix after dimension reduction, the second threshold is determined by the following formula:
[0025] ;
[0026] Among them, Q a is the second threshold, θ1 is the sum of the characteristic variable values discarded by principal component analysis, and c a is the confidence limit of the standard normal distribution, θ2 is the sum of squares of the eigenvalues discarded by principal component analysis, and h0 is the distribution correction coefficient.
[0027] Optionally, performing principal component analysis dimensionality reduction on the normal operation data matrix to obtain a normal operation data matrix after dimensionality reduction includes:
[0028] Determining the mean and variance of n sample values of each characteristic variable vector of the normal operating data matrix;
[0029] Normalizing the normal operation data matrix according to the mean and variance of the n sampling values to obtain a processed normal operation data matrix;
[0030] The processed normal operation data matrix is subjected to principal component analysis and dimensionality reduction to obtain a normal operation data matrix after dimensionality reduction.
[0031] Optionally, performing principal component analysis dimensionality reduction on the normal operation data matrix to obtain a normal operation data matrix after dimensionality reduction includes:
[0032] determining a covariance matrix of the normal operating data matrix;
[0033] Performing eigenvariate decomposition on the covariance matrix to obtain m eigenvariate values arranged from small to large;
[0034] Selecting the first k characteristic variable values from the m characteristic variable values, wherein the variance contribution rate of the first k characteristic variable values is greater than or equal to a contribution rate threshold;
[0035] According to the characteristic variable vector matrix corresponding to the first k characteristic variable values, principal component analysis dimensionality reduction is performed on the normal operating data matrix to obtain the normal operating data matrix after dimensionality reduction.
[0036] In a second aspect, the present application discloses a device for determining a fault source of a hydrogen production device, the device comprising: a data acquisition module, a value determination module, a contribution determination module, and a source determination module;
[0037] The data acquisition module is used to obtain the operating data matrix X to be tested of the alkaline electrolytic cell. n×m , where n is the number of sampling moments, m is the number of characteristic variables, and both m and n are integers greater than 1;
[0038] The numerical determination module is used to input the running data matrix to be detected into the principal component analysis model to obtain T 2 Statistics and SPE statistics, wherein the principal component analysis model is obtained by training the normal operation data matrix of the alkaline electrolytic cell;
[0039] The contribution determination module is used to determine if the T 2 If the statistical value is greater than the first threshold and / or the SPE statistical value is greater than the second threshold, then determining the contribution rate of each characteristic variable to the fault;
[0040] The source determination module is used to determine the characteristic variable with the highest contribution rate as the fault source.
[0041] Optionally, the contribution determination module includes: a first determination module, a second determination module, a third determination module and a fourth determination module;
[0042] The first determining module is configured to determine a principal component score vector at the time of the fault, wherein the principal component score vector includes k principal component scores, where k is a positive integer less than m;
[0043] The second determining module is configured to determine a fault principal component direction based on the k principal component scores;
[0044] The third determination module is configured to calculate, for each of the fault principal component directions, a local contribution rate of the characteristic variable to the fault;
[0045] The fourth determination module is configured to sum the local contribution rates of all the fault principal component directions to obtain the contribution rate of each characteristic variable to the fault.
[0046] Optionally, the first threshold is determined by the following unit:
[0047] The first determining unit is used to obtain the normal operation data matrix X of the alkaline electrolytic cell. n×m , where n is the number of sampling moments, m is the number of characteristic variables, and both m and n are integers greater than 1;
[0048] A second determining unit is configured to perform principal component analysis dimensionality reduction on the normal operating data matrix to obtain a normal operating data matrix after dimensionality reduction;
[0049] The third determining unit is configured to determine a first threshold value according to the normal operating data matrix after dimension reduction by using the following formula:
[0050] ;
[0051] Among them, T a is the first threshold, It follows the F distribution with the first degree of freedom k and the second degree of freedom nk, and 1-α is the confidence level.
[0052] Optionally, the second threshold is determined by the following unit:
[0053] The fourth determining unit is configured to determine a second threshold value according to the normal operation data matrix after dimension reduction by using the following formula:
[0054] ;
[0055] Among them, Q a is the second threshold, θ1 is the sum of the characteristic variable values discarded by principal component analysis, and c a is the confidence limit of the standard normal distribution, θ2 is the sum of squares of the eigenvalues discarded by principal component analysis, and h0 is the distribution correction coefficient.
[0056] Optionally, the second determination unit is specifically used to: determine the mean and variance of n sampling values of each characteristic variable vector of the normal operation data matrix; normalize the normal operation data matrix according to the mean and variance of the n sampling values to obtain a processed normal operation data matrix; perform principal component analysis on the processed normal operation data matrix to reduce the dimension of the normal operation data matrix to obtain a reduced dimension normal operation data matrix.
[0057] Optionally, the second determination unit is specifically used to: determine the covariance matrix of the normal operating data matrix; perform characteristic variable decomposition on the covariance matrix to obtain m characteristic variable values arranged from small to large; select the first k characteristic variable values from the m characteristic variable values, wherein the variance contribution rate of the first k characteristic variable values is greater than or equal to the contribution rate threshold; perform principal component analysis dimensionality reduction on the normal operating data matrix according to the characteristic variable vector matrix corresponding to the first k characteristic variable values to obtain the normal operating data matrix after dimensionality reduction.
[0058] In a third aspect, the present application discloses a device for determining a fault source of a hydrogen production device, the device comprising: a memory and a processor;
[0059] The memory is used to store programs;
[0060] The processor is used to execute the program to implement the various steps of the method for determining the fault source of the hydrogen production equipment as described in the first aspect.
[0061] In a fourth aspect, the present application discloses a computer-readable medium having a computer program stored thereon. When the computer program is executed by a processor, the computer program implements the various steps of the method for determining the source of a fault of a hydrogen production device as described in the first aspect.
[0062] Compared with the existing technology, this application has the following beneficial effects:
[0063] The present application discloses a method, device, equipment and medium for determining the source of a fault in a hydrogen production device. The method comprises: obtaining a matrix X of operating data to be tested of an alkaline electrolyzer; n×m , where n is the number of sampling moments, m is the number of characteristic variables, and both m and n are integers greater than 1; the running data matrix to be tested is input into the principal component analysis model to obtain T 2 Statistics and SPE statistics, where the principal component analysis model is trained through the normal operation data matrix of the alkaline electrolyzer; if T 2 If the statistical value is greater than the first threshold value and / or the SPE statistical value is greater than the second threshold value, the contribution rate of each characteristic variable to the fault is determined; the characteristic variable with the highest contribution rate is determined as the source of the fault. Therefore, the method for determining the source of the fault of the hydrogen production equipment provided in the embodiment of the present application only uses the operating data matrix collected by the existing sensors to determine the source of the fault, which not only avoids the cost of hardware modification, but also eliminates the measurement error that may be introduced by the additional device. In addition, the method for determining the source of the fault of the hydrogen production equipment provided in the embodiment of the present application quantifies the contribution rate of each characteristic variable to the fault, thereby avoiding the problem of low modeling accuracy caused by the scarcity of fault data in actual applications, and improving the accuracy and efficiency of fault source determination. BRIEF DESCRIPTION OF THE DRAWINGS
[0064] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.
[0065] Figure 1A A flow chart of a method for determining a fault source of hydrogen production equipment provided in an embodiment of the present application;
[0066] Figure 1B A flowchart of a method for obtaining a principal component analysis model provided in an embodiment of the present application;
[0067] Figure 2A and Figure 2B A schematic diagram of the first statistical change and contribution rate provided in an embodiment of the present application;
[0068] Figure 3A and Figure 3B A schematic diagram of the second statistical change and contribution rate provided in an embodiment of the present application;
[0069] Figure 4A 、 Figure 4B and Figure 4C A schematic diagram of the third statistical change and contribution rate provided in an embodiment of the present application;
[0070] Figure 5A and Figure 5B A schematic diagram of the fourth statistical change and contribution rate provided in an embodiment of the present application;
[0071] Figure 6 A schematic diagram of a device for determining a fault source of hydrogen production equipment provided in an embodiment of the present application;
[0072] Figure 7 A schematic diagram of a computer-readable medium provided in an embodiment of the present application. DETAILED DESCRIPTION
[0073] As described above, current methods for determining the source of alkaline electrolytic cell faults are primarily categorized into model-based and non-model-based methods. Model-based methods require inputting the alkaline electrolytic cell's fault operating data matrix and the fault source into a machine learning model. By minimizing the loss between the model's output and the actual fault source, the machine learning model is trained to obtain a fault diagnosis model. In practice, the test data matrix can then be directly input into the fault diagnosis model to determine the fault source. However, this method relies on a large amount of high-precision fault operating data matrices, which are often scarce and have limited accuracy in real-world applications, resulting in low accuracy and efficiency in fault source determination. Non-model-based methods for determining the source of faults include signal processing methods (such as fast Fourier transform) and electrochemical testing methods (such as electrochemical impedance spectroscopy). However, signal processing methods are highly susceptible to noise, similarly resulting in low accuracy in fault source determination. Electrochemical testing methods require additional testing equipment, which not only increases hardware costs but also may introduce measurement errors.
[0074] After research, the inventors have proposed a method, device, equipment and medium for determining the source of a fault in a hydrogen production device. The method comprises: obtaining a matrix X of operating data to be tested of an alkaline electrolyzer; n×m , where n is the number of sampling moments, m is the number of characteristic variables, and both m and n are integers greater than 1; the running data matrix to be tested is input into the principal component analysis model to obtain T 2 Statistics and SPE statistics, where the principal component analysis model is trained through the normal operation data matrix of the alkaline electrolyzer; if T 2If the statistical value is greater than the first threshold value and / or the SPE statistical value is greater than the second threshold value, the contribution rate of each characteristic variable to the fault is determined; the characteristic variable with the highest contribution rate is determined as the source of the fault. Therefore, the method for determining the source of the fault of the hydrogen production equipment provided in the embodiment of the present application only uses the operating data matrix collected by the existing sensors to determine the source of the fault, which not only avoids the cost of hardware modification, but also eliminates the measurement error that may be introduced by the additional device. In addition, the method for determining the source of the fault of the hydrogen production equipment provided in the embodiment of the present application quantifies the contribution rate of each characteristic variable to the fault, thereby avoiding the problem of low modeling accuracy caused by the scarcity of fault data in actual applications, and improving the accuracy and efficiency of fault source determination.
[0075] In order to help those skilled in the art better understand the present invention, the following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative work are within the scope of protection of this application.
[0076] See also Figure 1A , which is a flow chart of a method for determining the source of a fault in a hydrogen production device provided in an embodiment of the present application. The method includes:
[0077] S101: Obtain the operating data matrix X of the alkaline electrolytic cell to be tested n×m , where n is the number of sampling moments, m is the number of characteristic variables, and both m and n are integers greater than 1.
[0078] The alkaline electrolyzer operating data matrix to be tested (i.e., the test data set) refers to the data matrix needed to determine the source of the alkaline electrolyzer fault. This data matrix includes m characteristic variables (m is an integer greater than 1). These characteristic variables comprehensively reflect the operating status of the alkaline electrolyzer, such as current, voltage, temperature, pressure, and gas purity. For each characteristic variable, data is collected at n different sampling times to ensure comprehensiveness and timeliness of the data.
[0079] S102: Input the running data matrix to be tested into the principal component analysis model to obtain T 2 Statistics and SPE statistics, where the principal component analysis model is trained using the normal operation data matrix of the alkaline electrolyzer.
[0080] First, the training process of the principal component analysis model is explained in detail. Figure 1B , which is a flow chart of a method for obtaining a principal component analysis model provided in an embodiment of the present application. The training steps of the principal component analysis model are as follows:
[0081] A1: Obtain the normal operation data matrix of the alkaline electrolyzer.
[0082] The normal operation data matrix (i.e., training dataset) of an alkaline electrolyzer refers to the data matrix of the alkaline electrolyzer's normal, fault-free operation. This data matrix includes m characteristic variables (m is an integer greater than 1). These characteristic variables comprehensively reflect the normal operation of the alkaline electrolyzer, such as current, voltage, temperature, pressure, and gas purity. Each characteristic variable is collected at n different sampling times, ensuring comprehensiveness and timeliness of the data.
[0083] See Table 1, which is a schematic diagram of a normal operation data matrix provided in an embodiment of the present application. In Table 1, the normal operation data matrix includes three characteristic variables: current, voltage, and temperature, and each characteristic variable includes five sample values. Therefore, the normal operation data matrix constitutes a 5×3 matrix.
[0084] Table 1
[0085]
[0086] A2: Normalize the normal operation data matrix to obtain a processed normal operation data matrix.
[0087] Since the normal operation data matrix of the alkaline electrolyzer includes multiple characteristic variables such as current, voltage, temperature, pressure, and gas purity, there are usually dimensional differences and numerical range differences between the sampling values of these characteristic variables. Such differences will have an adverse effect on the subsequent determination of the source of the fault. Therefore, it is necessary to normalize the normal operation data matrix so that all characteristic variables have a unified scale and distribution characteristics, so as to carry out subsequent analysis more accurately and effectively.
[0088] Specifically, the normal operation data matrix is represented as X n×m dimensional data, where n is the number of sampling moments and m is the number of characteristic variables. First, the mean and variance of the n sampling values of each characteristic variable are calculated. The formula for calculating the mean is shown in the following formula (1), and the formula for calculating the variance is shown in the following formula (2):
[0089] (1)
[0090] (2)
[0091] in, is the average value of all sample values of the i-th characteristic variable, is the jth sampling value of the i-th characteristic variable, is the variance of all sample values of the i-th feature variable.
[0092] Then, the normal operation data matrix is normalized according to the mean and variance of the n sample values of each characteristic variable. The normalization formula is shown in the following formula (3):
[0093] (3)
[0094] in, is the jth sample value of the i-th feature variable after processing.
[0095] It's understandable that after normalization, all feature variables will follow a distribution with a mean of 0 and a standard deviation of 1. This approach effectively eliminates the impact of differences in dimension and numerical range between feature variables. Furthermore, after normalization, further operations such as noise removal and outlier processing can be performed on the processed normal operation data matrix to improve data quality and ensure the accuracy and effectiveness of fault diagnosis results.
[0096] A3: Perform principal component analysis and dimensionality reduction on the processed normal operation data matrix to obtain a normal operation data matrix after dimensionality reduction.
[0097] Principal Components Analysis (PCA) dimensionality reduction involves reducing the m-dimensional processed normal operation data matrix to a k-dimensional matrix (k is a positive integer less than m). This simplifies the calculation process and removes redundant characteristic variables, thereby improving the accuracy and efficiency of subsequent fault source determination.
[0098] Specifically, first, the covariance matrix is calculated based on the processed normal operation data matrix X. The covariance matrix can reflect the degree of linear relationship between different feature variables. The calculation formula of the covariance matrix is shown in the following formula (4):
[0099] (4)
[0100] Where R is the covariance matrix of the processed normal operation data matrix, and X is the normal operation data matrix after processing.
[0101] Next, the covariance matrix is decomposed into eigenvalues to obtain m eigenvalues of the covariance matrix, and the m eigenvalues are arranged in descending order, i.e., λ1, λ2...λ m And, according to the order of eigenvalues, determine the vector matrix P mm =[p1, p2...p m ].
[0102] Then, the first k eigenvalues (i.e., the larger k eigenvariate values) are selected so that their variance contribution rate is greater than or equal to the contribution rate threshold. It should be noted that the contribution rate threshold is a value greater than 0 and less than 1. Taking the contribution rate threshold of 0.85 as an example, it can be expressed as the following formula (5):
[0103] (5)
[0104] Then, based on the first k eigenvalues (k is the dimension after dimensionality reduction), construct the diagonal matrix S kk , and construct the projection matrix P based on the vector matrix corresponding to the first k eigenvalues mk . Diagonal matrix S kk and the projection matrix P mk The formula can be shown as follows:
[0105] (6)
[0106] (7)
[0107] Finally, according to the projection matrix P mk The matrix X of the normal operation data matrix after processing n×m Perform principal component analysis dimensionality reduction to obtain a matrix after dimensionality reduction. The matrix after dimensionality reduction includes the normal operation data matrix after dimensionality reduction. The formula for principal component analysis dimensionality reduction is shown in the following formula (8):
[0108] (8)
[0109] in, is the matrix after dimensionality reduction.
[0110] It can be understood that the principal component analysis dimensionality reduction processing can effectively convert the high-dimensional operating data matrix into a low-dimensional operating data matrix, while retaining the main information, removing redundant and secondary information, greatly simplifying the complexity of subsequent data processing, and improving the efficiency and accuracy of fault source diagnosis.
[0111] A4: Determine the first threshold and the second threshold based on the normal operation data matrix after dimensionality reduction.
[0112] In the method for determining the source of hydrogen production equipment failures provided in the embodiments of this application, two key statistics, the T² statistic and the SPE statistic, are used to effectively detect anomalies in the data matrix of the alkaline electrolyzer's operating state. The T² statistic is primarily used to monitor anomalies within the principal component space, such as systemic failures; the SPE statistic focuses on monitoring anomalies within the residual space, such as random noise or sensor failures.
[0113] For the T 2 To determine whether there is a fault, the first threshold T is determined based on the normal operation data matrix after dimensionality reduction. a (ie T 2 The first threshold T a The determination formula is shown in the following formula (9):
[0114] (9)
[0115] Among them, T a is the first threshold, The probability follows an F distribution with k as the first degree of freedom and nk as the second degree of freedom, and 1-α is the confidence level. Typically, α=0.01.
[0116] To determine whether a fault occurs based on the SPE statistic, the second threshold Q needs to be determined based on the normal operation data matrix after dimensionality reduction. a (i.e., the control limit of the SPE statistic). The second threshold Q a The determination formula is shown in the following formula (10):
[0117] (10)
[0118] Among them, Q a is the second threshold, , which reflects the power sum of the eigenvalues in the residual space and is used to describe the statistical characteristics of the residual. , is the distribution correction coefficient, which is used to adjust the distribution characteristics of the SPE statistic, c a is the confidence limit of the standard normal distribution, θ1 is the sum of the characteristic variable values discarded by the principal component analysis, and θ2 is the sum of the squares of the characteristic variable values discarded by the principal component analysis.
[0119] A5: Construct a principal component analysis model based on the first threshold and the second threshold.
[0120] After the principal component analysis model is constructed, the alkaline electrolytic cell operation data matrix to be tested obtained in step S101 can be normalized to obtain a processed operation data matrix to be tested. Subsequently, the processed operation data matrix to be tested is input into the principal component analysis model to determine the T 2 Statistics and SPE statistics are used to determine whether the alkaline electrolyzer has a fault and locate the source of the fault.
[0121] Specifically, determine the T of the running data matrix to be tested 2 The formula for the statistical value can be shown as follows (11):
[0122] (11)
[0123] in, T 2 Statistics, used to monitor abnormal conditions in the principal component space, such as systematic failures, is the vector of the characteristic variables of the processed running data matrix to be tested (hereinafter referred to as the vector to be tested, which represents the representation of the data to be tested in the feature space), P mk is the projection matrix, S kk is a diagonal matrix.
[0124] The formula for determining the SPE statistic value of the running data matrix to be tested can be shown as follows (12):
[0125] (12)
[0126] in, is the SPE statistic, which is used to monitor anomalies in the residual space, such as random noise or sensor failure. is the vector to be detected, is the m-dimensional identity matrix, P mk is the projection matrix.
[0127] S103: If T 2 If the statistical value is greater than the first threshold and / or the SPE statistical value is greater than the second threshold, the contribution rate of each characteristic variable to the fault is determined.
[0128] If T 2 The statistical value is higher than the first threshold T a , it is determined that the principal component space is abnormal (such as overall performance degradation). If the SPE statistic is higher than the second threshold Q a , then it is determined to be an abnormality in the residual space (such as sensor drift or noise mutation). Therefore, combining the above two aspects, if T 2 The statistical value is greater than the first threshold T a , and / or, the SPE statistic is greater than the second threshold Q a , then it is determined that there is an abnormality, and it is necessary to determine the contribution rate of each characteristic variable to the fault.
[0129] Specifically, first, determine the principal component score vector t=[t1,t2,…,t k ] T , where k is the number of principal components retained by the PCA model. Secondly, the direction of the fault principal component is determined based on the scores of the k principal components. For example, if the ratio of the principal component score to the characteristic variable value is t i / λ i Less than T a / k, it is determined as the fault principal component direction.
[0130] Then, for each fault principal component direction, the local contribution rate of the calculated characteristic variable to the fault is determined by the following formula (13):
[0131] (13)
[0132] in, is the local contribution rate, which indicates the local contribution of the jth characteristic variable to the fault under the direction of the i-th fault principal component, si is the i-th of the u principal component scores, and j is the dimension of the vector to be detected.
[0133] Finally, the following formula (14) sums the local contribution rates of all fault principal component directions to obtain the contribution rate of each characteristic variable to the fault:
[0134] (14)
[0135] in, The contribution rate comprehensively reflects the total contribution of the characteristic variable to the fault under different principal component directions. By comparing the contribution rates of different characteristic variables, we can identify the key characteristic variables that have the greatest impact on the fault, thus providing a basis for locating the fault source.
[0136] S104: Determine the characteristic variable with the highest contribution rate as the fault source.
[0137] In one specific implementation, to more intuitively demonstrate the impact of each target feature variable on the fault, the contribution rate of all target feature variables can be displayed using a histogram. In the histogram, the target feature variable with the highest contribution rate is the source of the fault.
[0138] For example, if the contribution rate of the characteristic variable "alkaline solution temperature" is significantly higher than that of other characteristic variables, it can be inferred that there may be a temperature sensor failure or an abnormality in the cooling system. In this way, the key factors that may cause alkaline electrolyzer failure can be quickly and accurately located, providing clear guidance for subsequent troubleshooting and repair work.
[0139] See also Figure 2A and Figure 2B , Figure 2A and Figure 2BSchematic diagram of the first statistical change and contribution rate provided in the embodiment of the present application. The training data set was collected from an alkaline electrolytic cell that operated normally for 168 hours, with a sampling interval of once every 2 minutes. The electrolytic cell consists of 4 chambers, and the characteristic variables involved include total voltage, voltage of each chamber, alkali solution temperature, chamber plate temperature, total current, and alkali solution flow rate. Based on this, the training data set forms a 5040×12 matrix (5040 = 168×60÷2, that is, the total sampling time converted to minutes and divided by the sampling interval to obtain the number of sampling points). The test data set was also collected for an alkaline electrolytic cell that ran for 479 hours, and the sampling interval and characteristic variables were consistent with the training set. At about 227 hours, the thermocouple measuring the alkali solution temperature malfunctioned, resulting in abnormal alkali solution temperature. Figure 2A and Figure 2B The changes in statistics in this case and the contribution rate of each characteristic variable to the fault are shown.
[0140] See also Figure 3A and Figure 3B , Figure 3A and Figure 3B Schematic diagram of the second statistical change and contribution rate provided in the embodiment of the present application. The training data set comes from an alkaline electrolytic cell that has been operating normally for 168 hours, with sampling every 2 minutes. The electrolytic cell structure consists of 4 chambers, and the characteristic variables include total voltage, voltage of each chamber, alkali temperature, chamber plate temperature, total current and alkali flow rate. The training data set is a 5040×12 matrix. The test data set was collected from an alkaline electrolytic cell that has been running for 479 hours, and the sampling interval and data characteristics are the same as those of the training set. At about 236 hours, the thermocouple for measuring the plate temperature malfunctioned, causing the plate temperature to become abnormal. Figure 3A and Figure 3B The changes in statistics in this case and the contribution rate of characteristic variables are shown.
[0141] See also Figure 4A 、 Figure 4B and Figure 4C , Figure 4A 、 Figure 4B and Figure 4CSchematic diagram of the third statistical change and contribution rate provided in the embodiment of the present application. The training data set was taken from an alkaline electrolytic cell that operated normally for 168 hours, with a sampling interval of 2 minutes. The electrolytic cell consists of 4 chambers, and the characteristic variables are total voltage, voltage of each chamber, alkali solution temperature, chamber plate temperature, total current and alkali solution flow rate. The training data set is a 5040×12 matrix. The test data set was collected from an alkaline electrolytic cell that ran for 479 hours, and the sampling interval and data characteristics were consistent with the training set. At about 227 hours, the thermocouple measuring the alkali solution temperature malfunctioned, resulting in abnormal alkali solution temperature; at about 236 hours, the thermocouple measuring the plate temperature malfunctioned, resulting in abnormal plate temperature. Figure 4A 、 Figure 4B and Figure 4C The changes in statistics and the contribution rate of characteristic variables in this case are shown.
[0142] See also Figure 5A and Figure 5B , Figure 5A and Figure 5B Schematic diagram of the fourth statistical change and contribution rate provided in the embodiment of the present application. The training data set was collected from an alkaline electrolyzer that operated normally for 11 hours, and samples were taken every 1 minute. The characteristic variables include system pressure, liquid level difference, lye flow rate, oxygen content in hydrogen, hydrogen content in oxygen, electrolysis current, electrolysis voltage, hydrogen tank temperature, oxygen tank temperature and lye temperature. The training data set forms a 646×11 matrix (646 = 11×60÷1, that is, the total sampling time is converted to minutes and divided by the number of sampling points obtained by the sampling interval). The test data set was collected from an alkaline electrolyzer that ran for 15 hours, and the sampling interval and data characteristics were the same as the training set. At about 3 hours, an abnormal hydrogen content in oxygen occurred. Figure 5A and Figure 5B The changes in statistics in this case and the contribution rate of characteristic variables are shown.
[0143] Through the above statistical changes and contribution rate analysis under different situations, we can intuitively observe the fluctuation of statistics and the contribution of each characteristic variable to the fault when different faults occur, which helps to accurately determine the source of the fault.
[0144] In summary, this application discloses a method for determining the source of faults in hydrogen production equipment. The method provided in the embodiments of this application utilizes only the operating data matrix collected by existing sensors to determine the source of the fault, thereby avoiding both hardware modification costs and measurement errors that may be introduced by additional devices. Furthermore, the method provided in the embodiments of this application circumvents the problem of low modeling accuracy caused by a scarcity of fault data in practical applications by quantifying the contribution of each characteristic variable to the fault, thereby improving the accuracy and efficiency of fault source determination.
[0145] See also Figure 6 , which is a schematic diagram of a device for determining the source of a hydrogen production device provided in an embodiment of the present application. The device 600 for determining the source of a hydrogen production device comprises: a data acquisition module 601, a value determination module 602, a contribution determination module 603, and a source determination module 604;
[0146] The data acquisition module 601 is used to obtain the operating data matrix X to be tested of the alkaline electrolytic cell. n×m , where n is the number of sampling moments, m is the number of characteristic variables, and both m and n are integers greater than 1;
[0147] The numerical determination module 602 is used to input the running data matrix to be tested into the principal component analysis model to obtain T 2 Statistics and SPE statistics, where the principal component analysis model is trained using the normal operating data matrix of the alkaline electrolyzer;
[0148] Contribution determination module 603, for 2 If the statistical value is greater than the first threshold and / or the SPE statistical value is greater than the second threshold, the contribution rate of each characteristic variable to the fault is determined;
[0149] The source determination module 604 is configured to determine the characteristic variable with the highest contribution rate as the fault source.
[0150] In a specific implementation, the contribution determination module 603 includes: a first determination module, a second determination module, a third determination module, and a fourth determination module;
[0151] A first determination module is used to determine a principal component score vector at the time of the fault, where the principal component score vector includes k principal component scores, where k is a positive integer less than m;
[0152] The second determination module is used to determine the direction of the fault principal component according to the k principal component scores;
[0153] The third determination module is used to calculate the local contribution rate of the characteristic variable to the fault for each main component direction of the fault;
[0154] The fourth determination module is used to sum the local contribution rates of all fault principal component directions to obtain the contribution rate of each characteristic variable to the fault.
[0155] In a specific implementation, the first threshold is determined by the following units:
[0156] The first determining unit is used to obtain the normal operation data matrix X of the alkaline electrolytic cell n×m , where n is the number of sampling moments, m is the number of characteristic variables, and both m and n are integers greater than 1;
[0157] The second determining unit is configured to perform principal component analysis dimensionality reduction on the normal operating data matrix to obtain a normal operating data matrix after dimensionality reduction;
[0158] The third determining unit is configured to determine a first threshold value according to the normal operating data matrix after dimensionality reduction by using the following formula:
[0159] ;
[0160] Among them, T a is the first threshold, It follows the F distribution with the first degree of freedom k and the second degree of freedom nk, and 1-α is the confidence level.
[0161] In a specific implementation, the second threshold is determined by the following units:
[0162] The fourth determining unit is configured to determine a second threshold value according to the normal operation data matrix after dimensionality reduction by using the following formula:
[0163] ;
[0164] Among them, Q a is the second threshold, θ1 is the sum of the characteristic variable values discarded by principal component analysis, and c a is the confidence limit of the standard normal distribution, θ2 is the sum of squares of the eigenvalues discarded by principal component analysis, and h0 is the distribution correction coefficient.
[0165] In a specific implementation method, the second determination unit is specifically used to: determine the mean and variance of the n sampling values of each characteristic variable vector of the normal operation data matrix; normalize the normal operation data matrix according to the mean and variance of the n sampling values to obtain the processed normal operation data matrix; perform principal component analysis on the processed normal operation data matrix to reduce the dimension to obtain the normal operation data matrix after dimension reduction.
[0166] In a specific implementation method, the second determination unit is specifically used to: determine the covariance matrix of the normal operating data matrix; perform characteristic variable decomposition on the covariance matrix to obtain m characteristic variable values arranged from small to large; select the first k characteristic variable values from the m characteristic variable values, wherein the variance contribution rate of the first k characteristic variable values is greater than or equal to the contribution rate threshold; perform principal component analysis dimensionality reduction on the normal operating data matrix according to the characteristic variable vector matrix corresponding to the first k characteristic variable values to obtain the normal operating data matrix after dimensionality reduction.
[0167] In summary, this application discloses a device for determining the source of a hydrogen production device fault. The device, provided in an embodiment of this application, determines the source of a fault using only the operating data matrix collected by existing sensors. This avoids hardware modification costs and eliminates measurement errors that may be introduced by additional devices. Furthermore, by quantifying the contribution of each characteristic variable to the fault, the device, provided in an embodiment of this application, avoids the problem of low modeling accuracy caused by a scarcity of fault data in actual applications, thereby improving the accuracy and efficiency of fault source determination.
[0168] The embodiment of the present application also provides a corresponding device for determining the source of a fault of a hydrogen production device and a computer-readable medium for implementing the method for determining the source of a fault of a hydrogen production device provided in the embodiment of the present application.
[0169] Among them, the fault source determination device of the hydrogen production equipment includes a memory and a processor, the memory is used to store instructions or codes, and the processor is used to execute instructions or codes so that the device executes a fault source determination method of a hydrogen production equipment in any embodiment of the present application.
[0170] See also Figure 7 , which is a schematic diagram of a computer-readable medium provided in an embodiment of the present application. The computer-readable medium 400 stores a computer program 411, which, when executed by a processor, implements the steps of the method for determining the source of a fault in the hydrogen production equipment of FIG. 1 .
[0171] It should be noted that in the context of the present application, a machine-readable medium may be a tangible medium that may contain or store a program for use by an instruction execution system, device or equipment or used in conjunction with an instruction execution system, device or equipment. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium may include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0172] It should be noted that the machine-readable medium mentioned above in this application can be a computer-readable signal medium or a computer-readable storage medium, or any combination of the two. A computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or device, or any combination of the above. More specific examples of computer-readable storage media can include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In this application, a computer-readable storage medium can be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, device, or device. In this application, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, which carries computer-readable program code. This propagated data signal can take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium that can transmit, propagate, or transport a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium may be transmitted using any suitable medium, including but not limited to wires, optical cables, RF (radio frequency), etc., or any suitable combination thereof.
[0173] The computer-readable medium may be included in the electronic device, or may exist independently without being incorporated into the electronic device.
[0174] Although the subject matter has been described in language specific to structural features and / or methodological logical acts, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or acts described above. Rather, the specific features and acts described above are merely example forms of implementing the claims.
[0175] Although several specific implementation details are included in the above discussion, these should not be construed as limiting the scope of this application. Certain features described in the context of separate embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented in multiple embodiments individually or in any suitable sub-combination.
[0176] The above description is merely a preferred embodiment of the present application and an illustration of the technical principles employed. Those skilled in the art should understand that the scope of disclosure in this application is not limited to the technical solutions formed by a specific combination of the above-mentioned technical features, but also encompasses other technical solutions formed by any combination of the above-mentioned technical features or their equivalents without departing from the above-mentioned disclosed concepts. For example, a technical solution formed by replacing the above-mentioned features with (but not limited to) technical features with similar functions disclosed in this application.
Claims
1. A method for determining the source of a fault in a hydrogen production device, characterized in that: The method comprises: Get the alkaline electrolytic cell's operating data matrix X to be tested n×m , where n is the number of sampling moments, m is the number of characteristic variables, and both m and n are integers greater than 1; The running data matrix to be tested is input into the principal component analysis model to obtain T 2 Statistics and SPE statistics, wherein the principal component analysis model is obtained by training the normal operation data matrix of the alkaline electrolytic cell; If the T 2 If the statistical value is greater than the first threshold and / or the SPE statistical value is greater than the second threshold, then determining the contribution rate of each characteristic variable to the fault; The characteristic variable with the highest contribution rate is determined as the fault source.
2. The method according to claim 1, characterized in that Determining the contribution rate of each characteristic variable to the fault includes: Determine a principal component score vector at the time of the fault, wherein the principal component score vector includes k principal component scores, where k is a positive integer less than m; Determining the principal component direction of the fault according to the k principal component scores; For each of the fault principal component directions, calculating the local contribution rate of the characteristic variable to the fault; The local contribution rates of all the fault principal component directions are summed to obtain the contribution rate of each characteristic variable to the fault.
3. The method according to claim 1, characterized in that The first threshold is determined by the following method: Obtain the normal operation data matrix X of the alkaline electrolyzer n×m , where n is the number of sampling moments, m is the number of characteristic variables, and both m and n are integers greater than 1; Performing principal component analysis and dimensionality reduction on the normal operating data matrix to obtain a normal operating data matrix after dimensionality reduction; According to the normal operation data matrix after dimension reduction, the first threshold is determined by the following formula: ; Among them, T a is the first threshold, It follows the F distribution with the first degree of freedom k and the second degree of freedom nk, and 1-α is the confidence level.
4. The method according to claim 3, characterized in that The second threshold is determined by the following method: According to the normal operation data matrix after dimension reduction, the second threshold is determined by the following formula: ; Among them, Q a is the second threshold, θ1 is the sum of the characteristic variable values discarded by principal component analysis, and c a is the confidence limit of the standard normal distribution, θ2 is the sum of squares of the eigenvalues discarded by principal component analysis, and h0 is the distribution correction coefficient.
5. The method according to claim 3, characterized in that Performing principal component analysis and dimensionality reduction on the normal operating data matrix to obtain a normal operating data matrix after dimensionality reduction, including: Determining the mean and variance of n sample values of each characteristic variable vector of the normal operating data matrix; Normalizing the normal operation data matrix according to the mean and variance of the n sampling values to obtain a processed normal operation data matrix; The processed normal operation data matrix is subjected to principal component analysis and dimensionality reduction to obtain a normal operation data matrix after dimensionality reduction.
6. The method according to claim 3, characterized in that The performing principal component analysis dimensionality reduction on the normal operating data matrix to obtain the normal operating data matrix after dimensionality reduction includes: determining a covariance matrix of the normal operating data matrix; Performing eigenvariate decomposition on the covariance matrix to obtain m eigenvariate values arranged from small to large; Selecting the first k characteristic variable values from the m characteristic variable values, wherein the variance contribution rate of the first k characteristic variable values is greater than or equal to a contribution rate threshold; According to the characteristic variable vector matrix corresponding to the first k characteristic variable values, principal component analysis dimensionality reduction is performed on the normal operating data matrix to obtain the normal operating data matrix after dimensionality reduction.
7. A device for determining the source of a fault in a hydrogen production device, characterized in that: The device includes: a data acquisition module, a value determination module, a contribution determination module and a source determination module; The data acquisition module is used to obtain the operating data matrix X to be tested of the alkaline electrolytic cell. n×m , where n is the number of sampling moments, m is the number of characteristic variables, and both m and n are integers greater than 1; The numerical determination module is used to input the running data matrix to be detected into the principal component analysis model to obtain T 2 Statistics and SPE statistics, wherein the principal component analysis model is obtained by training the normal operation data matrix of the alkaline electrolytic cell; The contribution determination module is used to determine if the T 2 If the statistical value is greater than the first threshold and / or the SPE statistical value is greater than the second threshold, then determining the contribution rate of each characteristic variable to the fault; The source determination module is used to determine the characteristic variable with the highest contribution rate as the fault source.
8. The device according to claim 7, characterized in that The contribution determination module includes: a first determination module, a second determination module, a third determination module and a fourth determination module; The first determining module is configured to determine a principal component score vector at the time of the fault, wherein the principal component score vector includes k principal component scores, where k is a positive integer less than m; The second determining module is configured to determine a fault principal component direction based on the k principal component scores; The third determination module is configured to calculate, for each of the fault principal component directions, a local contribution rate of the characteristic variable to the fault; The fourth determination module is configured to sum the local contribution rates of all the fault principal component directions to obtain the contribution rate of each characteristic variable to the fault.
9. A device for determining the source of a hydrogen production device, characterized in that: The device includes: a memory and a processor; The memory is used to store programs; The processor is configured to execute the program to implement each step of the method for determining a fault source of hydrogen production equipment according to any one of claims 1 to 6.
10. A computer-readable medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, each step of the method for determining a fault source of a hydrogen production equipment according to any one of claims 1 to 6 is implemented.