Transformer oil chromatography fault diagnosis method based on improved whale algorithm optimized BP neural network
By improving the whale algorithm and optimizing the BP neural network, integrating multi-source heterogeneous data, constructing a self-attention module and a fault knowledge graph, the problem of identifying complex fault types in transformer fault diagnosis was solved, achieving high-precision and fast fault diagnosis results.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- HANSHAN COUNTY POWER SUPPLY CO OF STATE GRID ANHUI ELECTRIC POWER CO LTD
- Filing Date
- 2025-12-11
- Publication Date
- 2026-05-12
AI Technical Summary
Existing transformer fault diagnosis methods lack accuracy when faced with complex faults, cannot effectively identify compound fault types, and traditional BP neural networks have reduced diagnostic accuracy in noisy environments, making them difficult to adapt to the fusion of multi-source heterogeneous data and real-time diagnosis requirements.
An improved whale algorithm is used to optimize the BP neural network. By fusing numerical, time-series, and textual data, a multi-source heterogeneous data fusion model is constructed. Combined with a self-attention module and a fault knowledge graph, the parameters of the BP diagnostic model are optimized to enhance the robustness and generalization ability of the diagnostic model.
It improves the accuracy and response speed of transformer fault diagnosis, especially in the case of complex fault types, and can effectively identify faults. The diagnostic accuracy is improved by 12% and the convergence speed is accelerated by 40%. It is suitable for real-time monitoring and maintenance management of power transformers.
Smart Images

Figure CN122024918A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of power system fault diagnosis technology, specifically a transformer oil chromatography fault diagnosis method based on an improved whale algorithm and optimized BP neural network. Background Technology
[0002] Power transformers are crucial equipment in power systems, responsible for voltage conversion and power transmission. Due to their long-term operation and complex working environment, transformers are prone to various faults, with insulation faults and overheating faults being the most common. Transformer faults not only affect the stability of the power system but can also lead to equipment damage, shutdowns, and even more serious accidents. Therefore, early detection and accurate diagnosis of transformer faults are key to ensuring the reliability and stability of the power system. Traditional transformer fault diagnosis methods mainly rely on transformer oil chromatography (DGA) analysis, which determines the fault type by analyzing changes in the composition and concentration of dissolved gases in the oil. According to the IEC 60599 standard, DGA analysis methods identify transformer fault types based on the ratios of different gases, typically including gas ratio methods (such as the Rogers method and the Dornenburg method) and empirical judgment methods. While these methods are simple and easy to implement, their limitations make them less effective in handling complex faults. The following section details several main methods for current transformer fault diagnosis and their limitations.
[0003] The three-ratio method proposed in the IEC 60599 standard is commonly used to determine transformer fault types by analyzing the ratios of dissolved gases (such as H2, CH4, C2H6, etc.). The basic principle of the three-ratio method is to infer the possible fault type of the transformer based on changes in the ratios of different gas concentrations. Although this method is simple to implement and low in cost, its shortcomings are also obvious. First, the three-ratio method is only applicable to single fault types and has weak diagnostic capabilities for complex faults. For example, for simultaneous "low-temperature overheating" and "low-energy discharge" faults, the three-ratio method may not accurately identify the specific situation, leading to insufficient accuracy in fault diagnosis. Second, the three-ratio method cannot provide the specific location and severity of the fault, limiting its early warning capability. Furthermore, the three-ratio method is sensitive to measurement errors in gas concentration; changes in the external environment may also affect its diagnostic results, further reducing the reliability of the method. Besides the three-ratio method, traditional empirical and statistical methods are also commonly used fault diagnosis tools. These methods are usually based on historical fault data and statistical models, combined with the variation patterns of gas concentrations in transformer oil for fault prediction. However, the main problem with these methods is their inability to handle complex fault modes and the interaction of multiple variables. When faced with novel or complex faults (such as combined faults like "multi-point grounding of the core" and "winding overheating"), these empirical and statistical methods show poor identification accuracy. Furthermore, the applicability of traditional methods is limited by data quality and acquisition accuracy, making them ineffective in handling the diverse operating conditions in large-scale transformer monitoring systems.
[0004] With the development of artificial intelligence technology, intelligent fault diagnosis methods based on machine learning have gradually become a research hotspot, especially transformer fault diagnosis systems based on BP neural networks and other deep learning methods. These methods learn the characteristics of different faults through training datasets and can adaptively improve the accuracy of diagnosis. Compared with traditional methods, neural networks and machine learning methods can extract information from higher-dimensional features, thus theoretically improving the accuracy and robustness of fault diagnosis. However, in practical applications, neural network-based fault diagnosis still faces many challenges.
[0005] Although backpropagation (BP) neural networks theoretically possess powerful learning capabilities, in practical applications, their convergence speed is often slow due to the extensive weight adjustments and gradient calculations required. In high-dimensional feature spaces, BP neural networks are prone to getting trapped in local optima, thus failing to reach the global optimum and leading to decreased model accuracy. Experiments show that standard BP neural networks typically require over 800 convergence iterations during training with low convergence accuracy, impacting their effectiveness in real-time diagnostics.
[0006] Transformer oil chromatography data are often affected by external noise, especially during long-term monitoring, where data quality fluctuates significantly due to equipment aging, temperature changes, or environmental interference. Traditional BP neural networks often struggle to maintain high diagnostic accuracy in such noisy environments. Some studies have shown that the accuracy of standard BP neural networks decreases by more than 20% under noise interference, and their diagnostic capability for complex fault types is further weakened.
[0007] Currently, most fault diagnosis methods rely solely on static data (such as gas concentration) for fault assessment, neglecting the temporal characteristics of transformer faults and the influence of historical data. Time-series data, such as gas production rates and gas concentration variation curves, often provide more accurate fault diagnosis information. In recent years, deep learning-based time-series analysis methods (such as LSTM networks) have been increasingly applied to fault diagnosis. However, effectively integrating multi-source heterogeneous data and improving the robustness and generalization ability of models in complex environments remains a significant challenge for current technology. Summary of the Invention
[0008] The purpose of this invention is to address the problems existing in the prior art by providing a transformer oil chromatography fault diagnosis method based on an improved whale algorithm and optimized BP neural network.
[0009] The objective of this invention is achieved through the following technical solution:
[0010] A method for fault diagnosis of transformer oil chromatography based on an improved whale algorithm and optimized BP neural network is characterized by the following steps:
[0011] S1. Collect numerical data to represent fault types, time-series data to capture fault development trends, and text data to represent historical fault cases.
[0012] S2. Numerical data, time series data, and text data are preprocessed and fused to generate a 51-dimensional fused feature vector, and the sample set of the 51-dimensional fused feature vector is augmented to form a standardized training set.
[0013] S3. Input the standardized training set into the BP diagnostic model. The BP diagnostic model is constructed by combining a BP neural network with an embedded self-attention module and a fault knowledge graph. The objective function of the BP diagnostic model is defined by the difference between the predicted probability and the true probability when performing cross-entropy loss.
[0014] S4. The parameters of the BP diagnostic model are iteratively optimized using the improved whale algorithm to minimize the objective function and obtain the fault diagnosis results.
[0015] The numerical data in step S1 is acquired and detected using a gas chromatograph equipped with a flame ionization detector and a thermal conductivity detector. The five characteristic gases for numerical data acquisition are hydrogen, methane, ethane, ethylene, and acetylene. During acquisition, 50 mL of oil sample is drawn from the sampling valve at the bottom of the transformer, degassed in the headspace at 80°C for 30 minutes, and then separated by an HP-PLOTQ column. The detection is performed according to the procedure of holding at 40°C for 5 minutes, then increasing the temperature at 5°C / min to 150°C and holding for 10 minutes. Finally, the volume concentration of the five characteristic gases is calculated by the peak area normalization method. After calculating the concentrations of the five characteristic gases, a 5-dimensional concentration vector is obtained and output. : The unit is μL / L.
[0016] The time-series data in step S1 includes 24-hour gas production rate and concentration change curves. The 24-hour gas production rate reflects the rate of fault deterioration, and the calculation formula is as follows: ,in: For the first Such gas in Gas production rate at any given time, in μL / (L・h), For the first Such gas in Volume concentration at time, For the first Such gas in Volume concentration 24 hours ago; concentration change curves based on 24-hour gas production rate, recording the volume concentrations of five gases at 6 time points within 1 hour with a time granularity of 10 minutes, forming a 5×6 time series data matrix. .
[0017] The text data in step S1 comes from historical fault cases in the substation operation and maintenance system, including fault ID, equipment model, fault type, gas data, handling plan and operating conditions; the historical fault cases are formatted into JSON structure to form a sample set of text data.
[0018] In step S2, the numerical data is normalized to eliminate dimensional differences, resulting in a 5-dimensional normalized vector. PCA is then used to reduce the dimensionality of this 5-dimensional normalized vector to obtain a 3-dimensional normalized vector representing the principal components. The time-series data in step S2 requires weighted interpolation to repair outliers in the 24-hour gas production rate and smoothing abrupt changes in the concentration curve using moving average filtering. The repaired time-series data is then extracted using an LSTM network to obtain a 16-dimensional time-series feature vector. The textual data in step S2 is encoded one-to-one using the Word2Vec semantic embedding method, resulting in a one-to-one 32-dimensional semantic vector. The 3-dimensional normalized vector, 16-dimensional time-series feature vector, and 32-dimensional semantic vector are then fused using the standard whale algorithm to optimize weights, generating a 51-dimensional fused feature vector. CGAN is used to generate simulated data to expand the sample set of the 51-dimensional fused feature vector, resulting in a standardized training set for the expanded 51-dimensional fused feature vector.
[0019] Outliers in the time-series data refer to gas production rates exceeding three times the historical average over 24 hours; abrupt changes in the time-series data refer to time-series data with a volume concentration change rate > 5% within 10 minutes.
[0020] The BP neural network in step S3 includes an input layer, hidden layer 1, hidden layer 2, hidden layer 3, and an output layer. A self-attention module that can enhance the weights of key features is embedded between hidden layer 2 and hidden layer 3.
[0021] In step S3, the BP diagnostic model receives a standardized training set containing 51-dimensional fused feature vectors through a BP neural network with an embedded self-attention module and outputs the initial probabilities of 6 types of faults. The fault knowledge graph built into the BP diagnostic model in step S3 is encoded into 128-dimensional knowledge vectors using a graph attention network. The 128-dimensional knowledge vectors are mapped to six-dimensional propensity scores through a fully connected layer. Then, Softmax normalization is used to convert the six-dimensional propensity scores into six knowledge probabilities that correspond one-to-one with the initial probabilities of the 6 types of faults. The knowledge probabilities and the one-to-one corresponding initial probabilities are fused with a weight of 0.2:0.8 to obtain the predicted probabilities.
[0022] The objective function for cross-entropy loss in step S3 is: ,in: For the sample size, For the sample The True label for the type of fault (0 or 1). To predict probabilities, the objective function aims to minimize the cross-entropy loss. .
[0023] The specific steps of the improved whale algorithm in step S4 are as follows:
[0024] S41. Initialize the parameters of the improved whale algorithm;
[0025] S42. Randomly generate the initial population;
[0026] S43. Enter the iteration loop;
[0027] S44. Calculate the fitness of each individual, i.e., the cross-entropy loss;
[0028] S45. Calculate the variation threshold;
[0029] S46. Use inertia weights and improved contraction factors to perform a dynamic weight secondary decay mechanism to balance global and local searches;
[0030] S47. Using elite back-learning, calculate the fitness value of the current generation population to obtain the optimal individual. And generate the inverse solution To increase population diversity;
[0031] S48. Determine whether the population diversity is less than the variation threshold. If yes, proceed to step S49; otherwise, proceed to step S10.
[0032] S49. Trigger Gaussian mutation or Cauchy mutation, proceed to step S10;
[0033] S410, Quantum-inspired optimization expands the search space, with individuals represented as quantum superposition states. ), through the revolving door Adjusting probability amplitude This guides the population to converge toward the optimal solution, improving the search efficiency for high-dimensional parameters;
[0034] S411, Update the whale population location and record the current iteration's optimal solution;
[0035] S412. Determine whether the maximum number of iterations has been reached. If yes, proceed to step S414; otherwise, proceed to step S413.
[0036] S413, Return to step S43;
[0037] S414. Output the globally optimal combination of BP neural network parameters and assign it to the BP diagnostic model.
[0038] The present invention has the following advantages over the prior art:
[0039] This invention improves the Whale Algorithm to optimize the BP network, enhances the feature extraction capability for complex faults, and constructs a multi-source heterogeneous data fusion model to integrate information such as gas concentration, gas production rate, and historical fault cases, thereby enhancing diagnostic generalization. At the same time, it starts from model optimization and hardware perception design to adapt to real-time diagnostic scenarios of edge devices, systematically solving the bottlenecks of existing technologies and providing a more reliable and efficient technical solution for intelligent transformer diagnosis.
[0040] The transformer oil chromatography fault diagnosis method provided by this invention is applicable to the intelligent identification and early warning of insulation faults in power transformers. Especially when facing complex fault types, it can effectively improve the accuracy of diagnosis and response speed, and is suitable for real-time monitoring and maintenance management of power transformers.
[0041] The transformer oil chromatography fault diagnosis method provided by this invention optimizes the initial parameters of the BP network by constructing a dynamic weight adjustment and elite reverse learning mechanism, thereby achieving high-precision diagnosis of transformer faults and solving the problems of slow convergence and easy trapping in local optima in traditional BP networks. After verification with 100 sets of actual data, the diagnostic accuracy is improved by 12% and the convergence speed is accelerated by 40%, making it suitable for intelligent diagnosis of complex fault scenarios. Attached Figure Description
[0042] Appendix Figure 1 A flowchart of a transformer oil chromatography fault diagnosis method based on an improved whale algorithm and optimized BP neural network is provided for this invention.
[0043] Appendix Figure 2 This is a flowchart of the data acquisition process in the transformer oil chromatography fault diagnosis method provided by the present invention;
[0044] Appendix Figure 3 This is a data processing flowchart of the transformer oil chromatography fault diagnosis method provided by the present invention;
[0045] Appendix Figure 4 The flowchart of model construction in the transformer oil chromatography fault diagnosis method provided by the present invention;
[0046] Appendix Figure 5 A flowchart for parameter optimization in the transformer oil chromatography fault diagnosis method provided by the present invention;
[0047] Appendix Figure 6 The spherical function test iteration diagram provided by this invention;
[0048] Appendix Figure 7 The Schweifer 2.22 problem test iteration diagram provided for the experimental testing of this invention;
[0049] Appendix Figure 8 Schweifer 1.2 problem test iteration diagram provided for experimental testing of the present invention;
[0050] Appendix Figure 9 The Schweifer 2.21 problem test iteration diagram provided for experimental testing of this invention;
[0051] Appendix Figure 10 One of the test iteration graphs of the generalized Rosenbrock function provided for experimental testing of the present invention;
[0052] Appendix Figure 11 The second test iteration graph of the generalized Rosenbrock function provided for the experimental testing of this invention;
[0053] Appendix Figure 12 A test iteration graph of a quartic function (including noise) provided for experimental testing of the present invention;
[0054] Appendix Figure 13 The Ackley function test iteration graph provided for the experimental testing of this invention;
[0055] Appendix Figure 14 Test diagram of the WOA-BP neural network model based on the elite adaptation strategy provided for experimental testing of the present invention;
[0056] Appendix Figure 15 A schematic diagram comparing the convergence curves of the three algorithms provided for experimental testing of this invention;
[0057] Appendix Figure 16 A schematic diagram showing the comparison of the accuracy of six types of fault diagnosis provided for experimental testing of this invention. Detailed Implementation
[0058] Exemplary embodiments will now be described more fully with reference to the accompanying drawings. However, these exemplary embodiments can be implemented in many forms and should not be construed as limited to the embodiments set forth herein; rather, they are provided so that the invention will be thorough and complete, and the concept of the exemplary embodiments will be fully conveyed to those skilled in the art. The same reference numerals in the drawings denote the same or similar structures, and therefore their detailed description will be omitted.
[0059] The terms “a,” “one,” “the,” and “the” are used to indicate the existence of one or more elements / components / etc.; the terms “including” and “having” are used to indicate an open-ended meaning of inclusion and that other elements / components / etc. may exist in addition to the listed elements / components / etc.
[0060] The flowchart of a transformer oil chromatography fault diagnosis method based on an improved whale algorithm and optimized BP neural network provided by this invention is as follows: Figure 1As shown, the steps of the transformer oil chromatography fault diagnosis method are as follows: S1, collect numerical data to represent the fault type, time-series data to capture the fault development trend, and textual data to represent historical fault cases; S2, preprocess and fuse the numerical data, time-series data, and textual data to generate a 51-dimensional fusion feature vector, and perform data augmentation on the sample set of the 51-dimensional fusion feature vector to form a standardized training set; S3, input the standardized training set into the BP diagnostic model, which is constructed by combining a BP neural network with an embedded self-attention module and a fault knowledge graph. The BP diagnostic model defines the objective function for cross-entropy loss based on the difference between the predicted probability and the true probability; S4, use an improved whale algorithm to iteratively optimize the parameters of the BP diagnostic model to minimize the objective function and obtain the fault diagnosis result.
[0061] Numerical data were acquired and detected using a gas chromatograph equipped with a flame ionization detector and a thermal conductivity detector. The five characteristic gases collected were hydrogen, methane, ethane, ethylene, and acetylene. During acquisition, a 50 mL oil sample was drawn from the bottom sampling valve of the transformer, degassed in the headspace at 80°C for 30 minutes, and then separated using an HP-PLOTQ column. The sample was detected using a program of holding at 40°C for 5 minutes, then increasing the temperature at 5°C / min to 150°C and holding for 10 minutes. The volume concentrations of the five characteristic gases were calculated using the peak area normalization method. After calculating the concentrations of each of the five characteristic gases, a 5-dimensional concentration vector was obtained and output. : The unit is μL / L; numerical data are normalized by range normalization to eliminate dimensional differences and obtain a 5-dimensional normalized vector. The 5-dimensional normalized vector is then subjected to PCA dimensionality reduction to obtain a 3-dimensional normalized vector that reflects the principal components.
[0062] Time-series data includes 24-hour gas production rate and concentration variation curves. The 24-hour gas production rate reflects the rate of fault deterioration, and the calculation formula is as follows: ,in: For the first Such gas in Gas production rate at any given time, in μL / (L・h), For the first Such gas in Volume concentration at time, For the first Such gas in Volume concentration 24 hours ago; concentration change curves based on 24-hour gas production rate, recording the volume concentrations of five gases at 6 time points within 1 hour with a time granularity of 10 minutes, forming a 5×6 time series data matrix. Weighted interpolation was used to repair outliers where the 24-hour gas production rate exceeded three times the historical average, and moving average filtering was used to smooth abrupt changes in the concentration change curve where the volume concentration change rate within 10 minutes was >5%. The repaired time-series data was obtained, and an LSTM network was used to extract trend features from the time-series data to obtain a 16-dimensional time-series feature vector representing the time-series data.
[0063] The text-based data originates from historical fault cases in the substation operation and maintenance system, including fault ID, equipment model, fault type, gas data, handling plan, and operating conditions. These historical fault cases are formatted into a JSON structure to form a sample set of text-based data. The text-based data in the sample set is then encoded one-to-one using the Word2Vec semantic embedding method, resulting in a corresponding 32-dimensional semantic vector.
[0064] The 3-dimensional normalized vector representing numerical data, the 16-dimensional temporal feature vector representing time-series data, and the 32-dimensional semantic vector representing textual data are fused using the standard whale algorithm with optimized weights to generate a 51-dimensional fused feature vector. The CGAN is then used to generate simulated data to expand the samples in the 51-dimensional fused feature vector sample set, resulting in a standardized training set of the 51-dimensional fused feature vector with increased data volume.
[0065] The BP neural network provided by this invention includes an input layer, hidden layer 1, hidden layer 2, hidden layer 3, and an output layer. A self-attention module that enhances the weights of key features is embedded between hidden layer 2 and hidden layer 3. The BP diagnostic model receives a standardized training set containing a 51-dimensional fused feature vector through the BP neural network with the embedded self-attention module and outputs the initial probabilities of six types of faults. In step S3, the fault knowledge graph built into the BP diagnostic model is encoded using a graph attention network to obtain a 128-dimensional knowledge vector. This 128-dimensional knowledge vector is mapped to a six-dimensional propensity score through a fully connected layer. Then, Softmax normalization is used to convert the six-dimensional propensity score into six knowledge probabilities that correspond one-to-one with the initial probabilities of the six types of faults. The knowledge probabilities and the one-to-one corresponding initial probabilities are fused with a weight of 0.2:0.8 to obtain the predicted probabilities. The objective function for cross-entropy loss is... ,in: For the sample size, For the sample The True label for the type of fault (0 or 1). To predict probabilities, the objective function aims to minimize the cross-entropy loss. .
[0066] The specific steps of the improved whale algorithm provided by this invention are as follows: S41, initialize the parameters of the improved whale algorithm; S42, randomly generate the initial population; S43, enter the iteration loop; S44, calculate the fitness of each individual, i.e., the cross-entropy loss; S45, calculate the mutation threshold; S46, use inertia weight and improved shrinkage factor to perform a dynamic weight secondary decay mechanism to balance the global and local search; S47, use elite back learning to calculate the fitness value of the current generation population to obtain the optimal solution individual. And generate the inverse solution To increase population diversity; S48, determine if the population diversity is less than the mutation threshold, if yes, proceed to step S49, otherwise proceed to step S10; S49, trigger Gaussian mutation or Cauchy mutation, proceed to step S10; S410, quantum-inspired optimization expands the search space, individuals are represented as quantum superposition states. ), through the revolving door Adjusting probability amplitude The process involves guiding the population towards the optimal solution and improving the search efficiency for high-dimensional parameters. Steps are as follows: S411: Update the whale population position and record the current iteration's optimal solution; S412: Determine if the maximum number of iterations has been reached. If yes, proceed to step S414; otherwise, proceed to step S413; S413: Return to step S43; S414: Output the globally optimal BP neural network parameter combination and assign it to the BP diagnostic model.
[0067] The following sections will detail the present invention’s method for transformer oil chromatography fault diagnosis based on an improved whale algorithm and optimized BP neural network, comprising four parts: data acquisition, data processing, model building, and parameter optimization.
[0068] I. Data Collection
[0069] The specific process of data collection is as follows: Figure 2 As shown, data acquisition is the foundation of fault diagnosis. It is necessary to accurately capture multi-dimensional information that reflects the internal state of the transformer to provide raw input for subsequent processing and modeling. Transformer faults cause the decomposition of insulating oil to produce characteristic gases. Different fault types correspond to specific gas compositions and variation patterns. According to IEC60599 and DL / T722 standards, the core data to be collected includes three categories: numerical data, time-series data, and text data.
[0070] Numerical data focuses on the concentrations of five key gases: hydrogen (H2), methane (CH4), ethane (C2H6), ethylene (C2H4), and acetylene (C2H2). Changes in these gas concentrations are directly related to fault types; for example, H2 is a sensitive indicator of insulation breakdown, while C2H2 is a specific marker of high-energy discharge. The data acquisition equipment uses a gas chromatograph (GC-2014) equipped with a flame ionization detector (FID) and a thermal conductivity detector (TCD), with a detection limit of 0.01 μL / L, capable of capturing low-concentration fault signals. During acquisition, a 50 mL oil sample is drawn from the bottom sampling valve of the transformer, degassed in the headspace at 80℃ for 30 minutes, and then separated using an HP-PLOTQ column. The sample is analyzed according to the procedure of "holding at 40℃ for 5 minutes → increasing the temperature at 5℃ / min to 150℃ and holding for 10 minutes." Finally, the volume concentration of each numerical characteristic gas is calculated using the peak area normalization method, with the following formula: ,in: For the first time collected Volume concentration of a characteristic gas; The peak area of the sample; As a correction factor; The standard gas is a standard gas with a known volume concentration (such as a 20 μL / L C2H2 standard). The peak area is for the standard gas. For the standard gas, the correction factor is... For the first time collected The correction factors for the five gases are calculated separately, and a 5-dimensional concentration vector is obtained and output. :
[0071] (1).
[0072] Time-series data includes 24-hour gas production rate and concentration variation curves to capture fault development trends. The 24-hour gas production rate reflects the rate of fault deterioration, calculated using the following formula: ,in: For the first Such gas in Gas production rate at any given time (unit: μL / (L・h)). For the first Such gas in Volume concentration at time, For the first Such gas in The volume concentration 24 hours ago, for example, when the current volume concentration of H2 is 280 μL / L and it was 220 μL / L 24 hours ago, the gas production rate is 2.5 μL / (L・h).
[0073] The concentration change curves were based on the 24-hour gas production rate, recording the volume concentrations of five gases at six time points within one hour with a time granularity of 10 minutes, forming a 5×6 time series data matrix. It is used to capture short-term mutations (such as C2H2 rising from 30 μL / L to 50 μL / L within 10 minutes, indicating increased discharge) and obtain mutation data (volume concentration change rate > 5% within 10 minutes); short-term mutation data is used to avoid invalid noise interference and retain key characteristics of early / sudden faults.
[0074] The time-series data acquisition device is an online DGA sensor (SST-DGA-900). The integrated temperature compensation module of the online DGA sensor can ensure measurement accuracy in environments ranging from 30 to 70°C. The data is transmitted to the InfluxDB time-series database in real time via 4G.
[0075] Text-based data originates from historical fault cases in substation operation and maintenance systems (such as PMS2.0), including fault ID, equipment model, fault type (e.g., "high-energy discharge + winding overheating"), gas data, handling solutions, and operating conditions (load rate, oil temperature). During data collection, valid cases from 500kV transformers between 2013 and 2023 were selected, removing records with missing information or logical contradictions, ultimately retaining 482 records. These records were formatted into a JSON structure to form a sample set of text-based data (e.g., {"fault type":"high-energy discharge + winding overheating","gas data":{"H2":260,"C2H2":28},...}), used for subsequent semantic encoding to supplement expert experience.
[0076] II. Data Processing
[0077] The flowchart of data processing is as follows Figure 3 As shown, the goal of data processing is to transform the raw data collected in the first step into standardized model input features. Through three stages—preprocessing, fusion, and enhancement—noise is eliminated, information is integrated, and samples are expanded, ultimately generating a 51-dimensional fused feature vector. The preprocessing stage mainly addresses the issues of data noise and dimensional differences.
[0078] For numerical data based on gas concentrations, the concentration ranges of the five gases differ significantly (e.g., the volume concentration of H2 is typically 0-500 μL / L, and the volume concentration of C2H2 is typically 0-50 μL / L). Directly inputting these into the model would cause the weights to be biased towards high-concentration gases. Therefore, the volume concentrations of the characteristic gases are corrected to range-normalized data with an offset. :
[0079] (2)
[0080] In equation (2), For the first time collected The normalized concentration of the gas, For the first time collected The volume concentration of the gas, and The collected data are respectively the first The minimum and maximum volume concentrations of the gas. As an offset (to avoid denominators of 0 or values extremely close to 0), the normalized data is mapped to the interval [0.01, 1.01], forming a 5-dimensional normalized vector. .
[0081] For time-series data, outliers and abrupt noise need to be handled. An outlier is identified when the 24-hour gas production rate exceeds three times the historical average, and weighted interpolation is used for repair. The repaired 24-hour gas production rate is then calculated. :
[0082] (3)
[0083] In equation (3), and The volume concentrations are shown for one hour before and one hour after the abnormal event. weight 2 and The weight 3 is set based on "the later time point is closer to the fault trend" (e.g., when the value before the anomaly point is 280 and the value after the anomaly point is 320, the value after the repair is 304).
[0084] Meanwhile, for abrupt changes in volume concentration with a change rate >5% within 10 minutes in the time series data, a moving average filter is used to smooth the abrupt changes.
[0085] (4)
[0086] In equation (4), Indicates the first The gas in the first Filtered volume concentration at each time point Or 5 (depending on the mutation magnitude). For the first Original volume concentration at each time point (e.g., sequence) exist (After filtering, the value is 100), ensuring the continuity and smoothness of the time series data.
[0087] The integration phase requires the integration of three types of features to improve information utilization.
[0088] For 5-dimensional normalized vectors PCA dimensionality reduction is performed by calculating the covariance matrix:
[0089] (5)
[0090] In equation (5), The first element in the covariance matrix Line number Column elements, For the first In the nth sample The normalized mean concentration of the gases. For the first In the nth sample The normalized concentration of the gas, For the first In the nth sample The normalized mean concentration of the gases. For the first In the nth sample The normalized concentrations of the gases are decomposed using eigenvalue decomposition, and the top three principal components with a cumulative variance contribution rate > 95% are retained to form a 3D principal component vector Y.
[0091] For time-series gas production rate vectors, an LSTM network is used to extract trend features, and its core gating mechanism includes a forgetting gate. (Retain historical information), Input portal (Updated with new information) and output gate (Generate the current hidden state), where The state was hidden in the previous moment. For the current moment Gas production rate at that time Weight matrix, Let be the bias vectors for the forget gate, input gate, and output gate, where for The output of the forget gate at any moment, for Input gate output at any time, for The output gate outputs at each time step, ultimately producing a 16-dimensional temporal feature vector. .
[0092] For text-based data, the Word2Vec semantic embedding method is used for encoding: First, keywords are extracted from the case text, retaining core terms related to "fault type," "characteristic gas," and "treatment plan" (such as "high-energy discharge," "C2H2," and "replace tap changer"). Keywords are then vectorized using the Word2Vec model, with a model window size of 5 (covering the semantic association of the five characters before and after the keyword) and an output vector dimension of 32. The mean of all keyword vectors for a single case is then calculated to obtain the 32-dimensional semantic vector for that case. This enables the conversion of textual experience into numerical features, avoids the dependence of complex pre-trained models on hardware resources, and preserves the basic semantic association of "fault-feature-solution".
[0093] The three types of features are fused using the standard Whale Algorithm (WOA) with optimized weights. The fusion formula is as follows:
[0094] (6)
[0095] In equation (6), the weight constraint condition The sum is 1, that is ; It is a 51-dimensional fused feature vector.
[0096] Using the "fault diagnosis accuracy after inputting three types of features into the BP neural network" as the fitness function, the optimal weight combination is searched through the standard WOA (Whale Enclosure Analysis)—WOA simulates whale encirclement and predation behavior and bubble net attack behavior. During the iteration process, a shrinkage factor is used. ( (K=300 is the total number of iterations) Adjust the search range to ultimately obtain the weight combination that optimizes diagnostic accuracy. After fusion, a 3+16+32=51-dimensional feature vector is formed. .
[0097] To address the scarcity of composite fault samples (accounting for only 20%), CGAN is used to generate simulated data to augment the sample set. Generator Input a 100-dimensional noise z and a fault label y, output a 51-dimensional fused feature vector with increased data volume; discriminator Distinguishing between real and generated data, the objective function is:
[0098] (7)
[0099] In equation (7), For generator, For discriminator, It is 100-dimensional random noise. It's a fault label. It is a 51-dimensional fused feature vector , It is the actual data distribution, reflecting The feature probability, For the distribution of random noise, For expectation operator, Given fault label At that time, the discriminator determines The probability of real data. Given fault label At that time, the generator is based on random noise The generated simulated feature vectors are used to filter samples with a model confidence score > 0.9, increasing the original sample set from 482 to 675 samples as a standardized training set to improve the model's generalization ability.
[0100] III. Model Construction
[0101] The flowchart of model construction is as follows Figure 4 As shown, the model construction uses the 51-dimensional fused feature vector output from the second step as input to build a fusion architecture of "BP neural network + self-attention module + fault knowledge graph". High-precision fault diagnosis is achieved by optimizing parameters through an improved Integral Whale Algorithm (IWOA). The topology of the BP neural network is optimized through grid search and consists of 5 layers: the input layer contains 51 nodes and receives the 51-dimensional fused feature vector. Hidden layer 1 has 16 nodes and uses the ReLU activation function. This effectively preserves weak fault characteristics of low-concentration gases (such as the slight increase of C2H2 in the early stages of a fault); Hidden layer 2 has 12 nodes and uses the LeakyReLU activation function. To address the "dead neuron" problem of ReLU, negative features (such as the relative decrease in concentration during fault relief) are captured; hidden layer 3 has 8 nodes and uses the Sigmoid activation function. The features are mapped to a probability distribution in the 0-1 range; the output layer has 6 nodes, corresponding to 6 types of faults (4 single faults + 2 composite faults, including: low temperature overheating, high temperature overheating, low energy discharge, high energy discharge, high energy discharge + winding overheating, and low energy discharge + low temperature overheating), and the Softmax activation function is used. Output the initial probabilities of the 6 types of faults. ( For the first Initial probability of class-specific faults). For the output layer The input value of each node, For the output layer For each node's input value, Softmax maps the output layer input to a probability distribution, facilitating fault category determination.
[0102] To enhance the weights of key features (such as C2H2 features during high-energy discharge), a self-attention module is embedded between hidden layer 2 and hidden layer 3, as shown in the formula:
[0103] (8)
[0104] In equation (8), This is the 12-dimensional output of hidden layer 2. The dimension of the key vector. This is used to prevent gradient vanishing. The self-attention module can automatically increase the weight of key features by 2-3 times (e.g., the weight of C2H2 feature increases from 0.15 to 0.345) and suppress redundant information (e.g., the weight of CH4 decreases from 0.12 to 0.08), thereby improving the accuracy of complex fault diagnosis by 12%.
[0105] To supplement prior knowledge (such as the association rule "C2H2 concentration > 10 μL / L → high-energy discharge"), a fault knowledge graph containing 1023 triples (such as <high-energy discharge, characteristic gas, C2H2>) is embedded and converted into a 128-dimensional knowledge vector through a graph attention network (GAT).
[0106] (9)
[0107] In equation (9), For nodes Embedded vector, for The adjacent nodes (e.g., the adjacent nodes of "high-energy discharge" are "C2H2" and "processing scheme"), The LeakyRuLU activation function is used. For attention weights calculate, This is the weight matrix. For bias.
[0108] The 128-dimensional knowledge vector is mapped to a six-dimensional propensity score through a fully connected layer. ,in It is a weight matrix. This is a bias. Then, Softmax normalization is used to transform the six-dimensional propensity scores into six categories of knowledge probabilities. : ,in It is the normalized first Knowledge probability of class-specific faults; The purpose of exponential operations on propensity scores is to amplify the differences in scores and ensure that the result is non-negative; This is the exponential sum of the six fault tendency scores, used for normalization.
[0109] The knowledge probability and the initial probability of one-to-one correspondence are fused with a weight of 0.2:0.8 to obtain the predicted probability. The predicted probability retains the accuracy of data-driven approaches while improving the interpretability of the model.
[0110] To measure the difference between the predicted probability and the true label, the objective function of the cross-entropy loss model is defined as:
[0111] (10)
[0112] In equation (10), For the sample size, For the sample The True label for the type of fault (0 or 1). To predict probabilities, the objective function aims to minimize the cross-entropy loss. .
[0113] IV. Parameter Optimization
[0114] The flowchart for parameter optimization is as follows: Figure 5 As shown, the parameters (1152 weights and biases) of the BP neural network are iteratively optimized using the improved Whale Algorithm (IWOA) to minimize the objective function L, and finally output the fault diagnosis result. IWOA makes four improvements to address the shortcomings of the standard WOA (slow convergence and susceptibility to local optima), significantly improving optimization performance.
[0115] First, a dynamic weight decay mechanism balances global and local searches, using inertial weights. and improved contractility factor :
[0116] (11)
[0117] In equation (11), The current iteration number (1≤ ≤300), This represents the total number of iterations. , Early stage of iteration ( <150) Inertia Weight With contractile factor If the value is large, expand the search range to explore the global optimum; later ( ≥150) Secondary attenuation, focusing on local fine-grained search, avoiding excessive oscillation.
[0118] Second, elite reverse learning increases population diversity: the fitness value of the current generation population is calculated to obtain the optimal individual. And generate the inverse solution. :
[0119] (12)
[0120] In equation (12), The upper and lower bounds of the parameter search space are defined, the optimal solution individual and its reverse solution are calculated, the fitness value of the individual is retained and better solutions are retained, effectively avoiding the population from getting trapped in local optima.
[0121] Third, the dynamic mutation threshold mechanism triggers the mutation threshold when population diversity is low. :
[0122]
[0123] As the number of iterations decreases (initially 0.8, eventually approaching 0), population diversity... Calculate according to the following formula
[0124] (13)
[0125] In equation (13), For population size, For parameter dimensions, For the first Dimensional parameter mean, when At this time, Gaussian mutation or Cauchy mutation is triggered, enhancing the ability to escape local optima.
[0126] The Gaussian mutation is:
[0127] (14)
[0128] Cauchy mutation is:
[0129] (15)
[0130] In equations (14) and (15), For the first The first individual Dimension value, It follows a standard normal distribution. These are random numbers distributed according to the standard Cauchy distribution.
[0131] Fourth, quantum-inspired optimization expands the search space: individuals are represented as quantum superposition states. ), through the revolving door Adjusting probability amplitude This guides the population towards the optimal solution, improving the search efficiency for high-dimensional parameters.
[0132] The optimization process is as follows: initialize 50 individuals (each group corresponds to 1 set of BP parameters), calculate the fitness value of each individual, and 50 individuals correspond to 50 fitness values, among which the individual with the lowest fitness value represents the optimal solution.
[0133] (16)
[0134] In equation (16), For cross-entropy loss, The inference time is specifically the time from when the model receives the input features of one sample to be diagnosed to when it outputs the complete response time of the probability distribution of six types of faults. Model size refers to the total space occupied by the model file (including all parameters, topology, and auxiliary module configuration) that can be directly deployed after training. Weights reflect priority. The iteration is performed 300 times, with dynamic weight adjustment, elite back learning, mutation, and quantum rotation gate updates performed in each round to retain the best individual. After the iteration ends, the optimal parameters are output.
[0135] When iteratively optimizing the parameters of the BP neural network using the Improved Whale Algorithm (IWOA), after 300 iterations (each round executes formula (11) dynamic weight second decay, formula (12) elite back learning, formula (13)-(15) dynamic mutation and quantum rotation gate update), the optimal whale individual (containing 1149 optimized BP network weights and biases) is finally selected. Its solution space is assigned to the BP diagnostic model to form the final diagnostic model. The core of the BP diagnostic model depends on the 51-dimensional fusion feature input of formula (6), the self-attention module of formula (8) to enhance key features, the knowledge graph of formula (9) to supplement prior knowledge, and the cross-entropy loss of formula (10) to measure prediction bias. The final model is packaged in a lightweight format (such as .tflite) (meeting the requirement of model size S < 50MB in the fitness function of formula (16). The input is the 51-dimensional fusion feature that has been standardized by formula (2)-(6). Afterwards, the probability of 6 types of faults can be output in milliseconds (meeting the requirement of inference time T<50ms in formula (16)) to complete the fault diagnosis of transformer oil chromatography.
[0136] Experimental Test
[0137] To evaluate the performance of the improved whale algorithm (IWOA) provided by this invention, benchmark tests were conducted as shown in Table 1. This experiment compared the proposed improved whale algorithm (IWOA) with five metaheuristic algorithms (SWO, WOA, KOA, SCA, and GWO).
[0138] Table 1. Names of Benchmark Functions
[0139]
[0140] The final convergence results of each algorithm on the test function are shown in Table 2 and... Figure 6 , Figure 7 , Figure 8 , Figure 9 , Figure 10 , Figure 11 , Figure 12 and Figure 13 As shown.
[0141] Table 2 Numerical Test Results
[0142]
[0143] Benchmark function testing and analysis show that IWOA exhibits excellent convergence performance and speed. It achieves optimal convergence results on all benchmark functions, demonstrating outstanding convergence speed. Despite... Figure 10 The generalized Rosenbrock function test iteration graph shown did not reach the optimal result, but still exhibited excellent convergence speed and a small difference from the optimal solution.
[0144] To verify the performance of the improved whale algorithm-optimized BP neural network model (improved IWOA-optimized BP) proposed in this invention in transformer oil chromatography fault diagnosis, experiments were conducted. Partial datasets are shown in Table 3, and experimental results are as follows: Figure 14 , Figure 15 and Figure 16 As shown.
[0145] Table 3. Partial data from oil chromatography.
[0146]
[0147] Figure 14 The training / validation / test curves of the IWOA-BP neural network model show that, during training, the cross-entropy of the IWOA-BP neural network model provided by this invention converges to 0.0877 within 77 iterations; the number of fault diagnosis errors ("valfail") in the validation set stabilizes at 1-2 after 50 iterations; and the weight update gradient stabilizes below 1e-4. This confirms that the IWOA-BP neural network model has stable convergence and strong generalization ability. Figure 15 and Figure 16 (Performance comparison chart of the three algorithms) further demonstrates the advantages of the IWOA-BP neural network model: in terms of convergence speed ( Figure 15 The IWOA-BP neural network model converges in just 48 iterations (meeting the mean squared error MSE < 1e-5), which is 52% faster than the traditional BP neural network (800 iterations) and 40% faster than the standard WOA-BP model (80 iterations); in terms of fault diagnosis accuracy ( Figure 16 The IWOA-BP neural network model achieved an overall accuracy of 94%, a 12 percentage point improvement over the traditional BP neural network (82%) and a 6 percentage point improvement over the standard WOA-BP model (88%). The improvement was particularly significant in complex fault diagnosis (96% accuracy for high-temperature overheating faults and 88% accuracy for high-energy discharge faults). This achievement provides a new method for intelligent operation and maintenance of power equipment. In the future, the model's generalization ability under complex operating conditions can be further enhanced through convolutional neural networks and transfer learning.
[0148] In this embodiment of the invention, the term "multiple" refers to two or more, unless otherwise explicitly defined. The terms "install," "connect," and "fix" should be interpreted broadly. For example, "connect" can mean a fixed connection, a detachable connection, or an integral connection. Those skilled in the art can understand the specific meaning of the above terms in this embodiment of the invention based on the specific circumstances.
[0149] In the description of the embodiments of the present invention, it should be understood that the terms "upper" and "lower" indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings. They are only for the convenience of describing the embodiments of the present invention and simplifying the description, and do not indicate or imply that the device or unit referred to must have a specific orientation or be constructed and operated in a specific orientation. Therefore, they should not be construed as limitations on the embodiments of the present invention.
[0150] In the description of this specification, the terms "an embodiment," "a preferred embodiment," etc., refer to a specific feature, structure, material, or characteristic described in connection with that embodiment or example, which is included in at least one embodiment or example of the present invention. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples.
[0151] The above embodiments are merely illustrative of the technical concept of the present invention and should not be construed as limiting the scope of protection of the present invention. Any modifications made to the technical solutions based on the technical concept proposed in this invention shall fall within the scope of protection of this invention. Technologies not covered in this invention can be implemented using existing technologies.
Claims
1. A method for transformer oil chromatographic fault diagnosis based on an improved whale algorithm and optimized BP neural network, characterized in that: The steps of this fault diagnosis method are as follows: S1. Collect numerical data to represent fault types, time-series data to capture fault development trends, and text data to represent historical fault cases. S2. Numerical data, time series data, and text data are preprocessed and fused to generate a 51-dimensional fused feature vector, and the sample set of the 51-dimensional fused feature vector is augmented to form a standardized training set. S3. Input the standardized training set into the BP diagnostic model. The BP diagnostic model is constructed by combining a BP neural network with an embedded self-attention module and a fault knowledge graph. The objective function of the BP diagnostic model is defined by the difference between the predicted probability and the true probability when performing cross-entropy loss. S4. The parameters of the BP diagnostic model are iteratively optimized using the improved whale algorithm to minimize the objective function and obtain the fault diagnosis results.
2. The transformer oil chromatographic fault diagnosis method based on the improved whale algorithm and optimized BP neural network according to claim 1, characterized in that: The numerical data in step S1 is acquired and detected using a gas chromatograph equipped with a flame ionization detector and a thermal conductivity detector. The five characteristic gases for numerical data acquisition are hydrogen, methane, ethane, ethylene, and acetylene. During acquisition, 50 mL of oil sample is drawn from the sampling valve at the bottom of the transformer, degassed in the headspace at 80°C for 30 minutes, and then separated by an HP-PLOTQ column. The detection is performed according to the procedure of holding at 40°C for 5 minutes, then increasing the temperature at 5°C / min to 150°C and holding for 10 minutes. Finally, the volume concentration of the five characteristic gases is calculated by the peak area normalization method. After calculating the concentrations of the five characteristic gases, a 5-dimensional concentration vector is obtained and output. : The unit is μL / L.
3. The transformer oil chromatographic fault diagnosis method based on the improved whale algorithm and optimized BP neural network according to claim 1, characterized in that: The time-series data in step S1 includes 24-hour gas production rate and concentration change curves. The 24-hour gas production rate reflects the rate of fault deterioration, and the calculation formula is as follows: ,in: For the first Such gas in Gas production rate at any given time, in μL / (L・h), For the first Such gas in Volume concentration at time, For the first Such gas in Volume concentration 24 hours ago; concentration change curves based on 24-hour gas production rate, recording the volume concentrations of five gases at 6 time points within 1 hour with a time granularity of 10 minutes, forming a 5×6 time series data matrix. .
4. The transformer oil chromatographic fault diagnosis method based on the improved whale algorithm and optimized BP neural network according to claim 1, characterized in that: The text data in step S1 comes from historical fault cases in the substation operation and maintenance system, including fault ID, equipment model, fault type, gas data, handling plan and operating conditions; the historical fault cases are formatted into JSON structure to form a sample set of text data.
5. The transformer oil chromatographic fault diagnosis method based on an improved whale algorithm optimized BP neural network according to any one of claims 1-4, characterized in that: In step S2, the numerical data is normalized to eliminate dimensional differences, resulting in a 5-dimensional normalized vector. PCA is then used to reduce the dimensionality of this 5-dimensional normalized vector to obtain a 3-dimensional normalized vector representing the principal components. The time-series data in step S2 requires weighted interpolation to repair outliers in the 24-hour gas production rate and smoothing abrupt changes in the concentration curve using moving average filtering. The repaired time-series data is then extracted using an LSTM network to obtain a 16-dimensional time-series feature vector. The textual data in step S2 is encoded one-to-one using the Word2Vec semantic embedding method, resulting in a one-to-one 32-dimensional semantic vector. The 3-dimensional normalized vector, 16-dimensional time-series feature vector, and 32-dimensional semantic vector are then fused using the standard whale algorithm to optimize weights, generating a 51-dimensional fused feature vector. CGAN is used to generate simulated data to expand the sample set of the 51-dimensional fused feature vector, resulting in a standardized training set for the expanded 51-dimensional fused feature vector.
6. The transformer oil chromatographic fault diagnosis method based on the improved whale algorithm and optimized BP neural network according to claim 5, characterized in that: Outliers in the time-series data refer to gas production rates exceeding three times the historical average over 24 hours; abrupt changes in the time-series data refer to time-series data with a volume concentration change rate > 5% within 10 minutes.
7. The transformer oil chromatographic fault diagnosis method based on an improved whale algorithm optimized BP neural network according to any one of claims 1-4, characterized in that: The BP neural network in step S3 includes an input layer, hidden layer 1, hidden layer 2, hidden layer 3, and an output layer. A self-attention module that can enhance the weights of key features is embedded between hidden layer 2 and hidden layer 3.
8. The transformer oil chromatographic fault diagnosis method based on the improved whale algorithm and optimized BP neural network according to claim 7, characterized in that: In step S3, the BP diagnostic model receives a standardized training set containing 51-dimensional fusion feature vectors through a BP neural network with an embedded self-attention module and outputs the initial probabilities of 6 types of faults. The fault knowledge graph built into the BP diagnostic model in step S3 is encoded into a 128-dimensional knowledge vector using a graph attention network. The 128-dimensional knowledge vector is mapped to a six-dimensional propensity score through a fully connected layer. Then, Softmax normalization is used to convert the six-dimensional propensity score into six knowledge probabilities that correspond one-to-one with the initial probabilities of the 6 types of faults. The knowledge probability and the initial probability of one-to-one correspondence are fused with a weight of 0.2:0.8 to obtain the predicted probability.
9. The transformer oil chromatographic fault diagnosis method based on the improved whale algorithm and optimized BP neural network according to claim 1, characterized in that: The objective function for cross-entropy loss in step S3 is: ,in: For the sample size, For the sample The True label for the type of fault (0 or 1). To predict probabilities, the objective function aims to minimize the cross-entropy loss. .
10. The transformer oil chromatographic fault diagnosis method based on the improved whale algorithm and optimized BP neural network according to claim 1, characterized in that: The specific steps of the improved whale algorithm in step S4 are as follows: S41. Initialize the parameters of the improved whale algorithm; S42. Randomly generate the initial population; S43. Enter the iteration loop; S44. Calculate the fitness of each individual, i.e., the cross-entropy loss; S45. Calculate the variation threshold; S46. Use inertia weights and improved contraction factors to perform a dynamic weight secondary decay mechanism to balance global and local searches; S47. Using elite back-learning, calculate the fitness value of the current generation population to obtain the optimal individual. And generate the inverse solution To increase population diversity; S48. Determine whether the population diversity is less than the variation threshold. If yes, proceed to step S49; otherwise, proceed to step S10. S49. Trigger Gaussian mutation or Cauchy mutation, proceed to step S10; S410, Quantum-inspired optimization expands the search space, with individuals represented as quantum superposition states. ), through the revolving door Adjusting probability amplitude This guides the population to converge toward the optimal solution, improving the search efficiency for high-dimensional parameters; S411, Update the whale population location and record the current iteration's optimal solution; S412. Determine whether the maximum number of iterations has been reached. If yes, proceed to step S414; otherwise, proceed to step S413. S413, Return to step S43; S414. Output the globally optimal combination of BP neural network parameters and assign it to the BP diagnostic model.