A data pre-cleaning method for hard connection contact point AI training data hot analysis
By constructing a three-dimensional model and thermal performance surface of the high-voltage switch disconnector, and combining normalized coefficient labels and statistical discrimination methods, the training data of the hard connection contact points at the outgoing terminals of the high-voltage switch disconnector was cleaned, thus solving the data error problem and improving the data quality of AI training and the accuracy of thermal analysis.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-15
- Publication Date
- 2026-03-24
AI Technical Summary
Existing technologies lack effective methods for error processing of heating data at the hard-connected contact points of high-voltage switch disconnectors, leading to contamination of AI training data and affecting the accuracy of thermal analysis.
By establishing a three-dimensional model and performing finite element analysis, a thermal performance surface model is constructed. The training data is cleaned using normalized coefficient labels and statistical discrimination methods to remove data containing gross errors, thereby obtaining a high-quality training dataset.
It effectively removes gross errors caused by multiple factors, improves the quality of AI training data, shortens training time, and enhances the accuracy of thermal analysis.
Smart Images

Figure CN121301939B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of data processing technology, and specifically to a data pre-cleaning method for thermal analysis of AI training data at hard-connected contact points. Background Technology
[0002] Existing traditional methods for training neural network AI to perform thermal analysis on other electrical equipment and components involve designing a neural network method, acquiring a large amount of measured data in the field, starting with an arbitrary initial performance planar model, training the neural network model, obtaining a surface model of the electrical equipment's thermal performance system through the system identification of the trained neural network, and then predicting the node temperature based on the surface model.
[0003] Existing AI technologies require a large amount of accurate training data, and currently there is no method for heat analysis and data processing of the hard-connected contact points at the outgoing terminals of high-voltage switchgear. However, there are methods for diagnosing the heat parameters of the contact heads in the switchgear body, such as patent CN 114819194 A (a method for predicting heat defects in switchgear). This method involves acquiring historical data for each switchgear, preprocessing the historical data to obtain a training dataset; using a number as an index for the corresponding switchgear; establishing various base learners based on different algorithms, training each base learner using the training dataset to obtain trained base learners; acquiring the data to be predicted for each switchgear, preprocessing the data to be predicted to obtain a dataset to be predicted; using each base learner to predict the dataset to be predicted, obtaining classification discriminant values; performing probability transformations on the classification discriminant values obtained by each base learner to obtain probability values; calculating the average of each probability value to obtain a predicted probability value; the data to be predicted includes parameter A of the switchgear; and the predicted probability value represents the probability of a heat defect occurring in the switchgear.
[0004] This method also requires training with a large amount of actual field data, and currently there is no information available on error handling methods for hard connection contact points at the outgoing terminals of high-voltage switches. Summary of the Invention
[0005] The purpose of this invention is to provide a data pre-cleaning method for thermal analysis of AI training data at hard-connected contact points, in order to address the shortcomings in the prior art.
[0006] To achieve the above objectives, the present invention provides the following technical solution: a data pre-cleaning method for thermal analysis of AI training data at hard-connected contact points, the cleaning method comprising the following steps:
[0007] The high-voltage switch disconnectors on site are classified into frames, and then three-dimensional models are established for the classified high-voltage switch disconnectors. The thermal shape simulation model is simulated using finite element analysis to obtain the thermal performance surface corresponding to the thermal analysis structure of the three-dimensional model of the high-voltage switch disconnector under a certain contact resistance condition.
[0008] When analyzing the hard connection contact point of the high-voltage switch knife switch, a thermal shape simulation model is selected, and a simulation thermal parameter surface model is constructed using basic data. The simulation thermal parameter surface model includes a temperature rise surface model under contact resistance.
[0009] The training data collected on-site is used as the dataset to be cleaned. The temperature rise surface model is sliced to obtain curves in combination with the training data. The normalization coefficient labels are obtained by comparing the curves with the training data. Normalization coefficient labels are then assigned to each training data.
[0010] The normalized coefficient labels of the training data are used as statistical data X samples. They are listed in a table to create a sample set of coefficient labels to be cleaned. The statistical discriminant method is used to identify the sample set of coefficient labels to be cleaned, find the normalized coefficient labels whose variance exceeds the confidence interval, and remove the training data corresponding to the normalized coefficient labels to obtain the reconstructed and cleaned on-site training dataset.
[0011] Preferably, the training data collected on-site is used as the dataset to be cleaned. The temperature rise surface model is sliced using the training data to obtain curves. The curves are compared with the training data to obtain normalized coefficient labels. Each training data point is then labeled with a normalized coefficient label, including the following steps:
[0012] Determine the operating conditions corresponding to the data to be cleaned, and obtain the standard temperature rise response curve corresponding to the operating conditions by slicing the temperature rise surface model through the on-site wind speed, contact resistance or switching current of the hard-connected contact point.
[0013] The temperature response sequence curves in the training data are feature-aligned and compared with the standard temperature rise response curves obtained by slicing the temperature rise surface model.
[0014] The two curves are compared and analyzed point by point or as a whole. The normalization coefficient is calculated based on the comparison results to quantitatively characterize the degree of fit or deviation of a single training data point relative to the standard temperature rise response curve.
[0015] Each training data point is labeled with a normalized coefficient. The calculated normalized coefficient serves as the data label for that training data point and is bound to the original data record for storage.
[0016] Preferably, the comparative analysis of the two curves, either point-by-point or overall, includes the following steps:
[0017] Extract the temperature values of the two curves at time nodes and calculate the absolute or relative difference between the corresponding time nodes; or select the temperature series within a time window and calculate the mean, variance, and maximum deviation of the temperature difference within the time window; or compare the changing trends of the training data and the standard curve segment by segment using a sliding window method.
[0018] Preferably, a normalization coefficient is calculated based on the comparison results to quantitatively characterize the degree of fit or deviation of a single training data point relative to the standard temperature rise response curve. The calculation logic is as follows:
[0019] Select one or more indicators that can reflect the degree of bias; integrate the indicators into a single value through normalization mapping, which serves as the normalization coefficient for the training data.
[0020] If the training data is consistent with the standard temperature rise response curve in all features, then the normalization coefficient approaches the preset ideal reference value.
[0021] Conversely, if there is a deviation, the normalization coefficient will deviate from the ideal value.
[0022] Preferably, the training data collected on-site includes time-series temperature samples, switching current, ambient temperature, wind speed, contact resistance, and high-voltage switch type.
[0023] Preferably, the normalized coefficient labels of the training data are used as statistical data X samples, and after being listed in a table, a sample set of coefficient labels to be cleaned is created, including the following steps:
[0024] Construct a statistical data table of normalized coefficient labels. Arrange and record the normalized coefficient labels corresponding to each training data point in a structured table according to the data organization rules, forming a statistical sample table with data samples as rows and normalized coefficients as statistical observations.
[0025] Preferably, the statistical discriminant method is used to identify the sample set of coefficient labels to be cleaned, find the normalized coefficient labels whose variance exceeds the confidence interval, and remove the training data corresponding to the normalized coefficient labels to obtain the reconstructed and cleaned on-site training dataset, including the following steps:
[0026] Based on the completed sample set of coefficient labels to be cleaned, statistical discriminant logic is used to perform overall distribution characteristic analysis and outlier detection on all normalized coefficient labels, evaluate the overall distribution pattern of all normalized coefficient labels, and identify outliers that deviate from the pattern.
[0027] Calculate the global statistical characteristics of the normalized coefficient labels, and based on the global statistical characteristics, set a confidence interval to define whether the normalized coefficient labels are within the normal fluctuation range or deviate abnormally.
[0028] Each normalized coefficient label is evaluated one by one to determine whether its value falls within the preset confidence interval. If the value of a normalized coefficient label exceeds the upper or lower limit of the interval, it is determined to be an abnormal label.
[0029] Locate all original training data records associated with the normalized coefficient labels that are judged to be abnormal, and perform data cleanup operations;
[0030] After identifying and removing anomalous data, the remaining training data constitutes the reconstructed and cleaned on-site training dataset.
[0031] Preferably, the statistical sample table includes a unique identifier for the data sample, the category of the high-voltage switch disconnector, the collected operating parameters, the index of the original data file, and a normalization coefficient label field.
[0032] Preferably, the global statistical features include mean, variance, range, skewness, and kurtosis.
[0033] Preferably, the high-voltage switchgear on site is classified into frames, and then a three-dimensional model is established for the classified high-voltage switchgear. The thermal shape simulation model is generated using finite element analysis, and the thermal performance surface corresponding to the thermal analysis structure of the three-dimensional model of the high-voltage switchgear is obtained under a certain contact resistance condition. This includes the following steps:
[0034] The high-voltage switchgear involved in the field data collection is divided into several categories according to the framework;
[0035] After classifying the high-voltage switchgear frame, a three-dimensional model reflecting the physical structure and material distribution of each high-voltage switchgear in each category is constructed based on its structural design parameters.
[0036] Material properties and boundary conditions were set for the three-dimensional model, and multiple sets of simulation results were obtained by numerical simulation of the high-voltage switch disconnector using finite element analysis software.
[0037] Based on multiple sets of simulation results, thermal performance surfaces are extracted and constructed. By integrating and visualizing the temperature response data obtained under different contact resistance conditions, thermal performance surfaces reflecting the correspondence between contact resistance and thermal response are obtained.
[0038] The technical effects and advantages provided by the present invention in the above technical solution are as follows:
[0039] This invention addresses the heating data of hard-connected contact points at the outgoing terminals of high-voltage switchgear. It removes data with gross errors from data collected under various parameter conditions. Essentially, it uses a multi-dimensional performance surface model to clean the collected system identification data. This model is then compared with the normalized coefficients of the simulated surface using the collected data. This process removes out-of-quality data containing gross errors, preventing contamination during subsequent AI neural network training and accelerating training convergence.
[0040] This method can be used in the data collection of probability distribution of independent events under the influence of multiple factors. It can clean and process multi-influence factor data and multi-dimensional training data to remove training data containing gross errors. Attached Figure Description
[0041] To more clearly illustrate the technical solutions in this embodiment or the prior art, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in this invention. For those skilled in the art, other drawings can be obtained based on these drawings.
[0042] Figure 1 This is a flowchart of the cleaning method of the present invention.
[0043] Figure 2 A flowchart for cleaning data using existing technologies.
[0044] Figure 3 This is a current-temperature rise curve at 0 wind speed according to the present invention.
[0045] Figure 4 The figures show the current-temperature rise curves at 32 microohms and 0.5 microohms, respectively, according to the present invention.
[0046] Figure 5 This is a comparison chart of the predicted temperature rise and the actual temperature rise under different wind speeds according to the present invention. Detailed Implementation
[0047] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0048] The hard connection contact point on the outgoing side of the high-voltage switch disconnector is a key area for fault monitoring during power transmission and distribution operation and maintenance. Current theoretical research and technology do not provide heat analysis of this hard connection contact point.
[0049] The modeling technology for heat generation at connection contact points has evolved from simplified empirical formulas to high-precision multiphysics simulations. Early methods primarily relied on contact resistance models and steady-state thermal balance equations, neglecting microstructure and dynamic effects, resulting in significant estimation errors. After 2000, finite element analysis was introduced to achieve electro-thermal coupling simulation. To address the modeling and assessment of the health status of hard-connection contact points in actual field measurements, and to lay the groundwork for subsequent intelligent analysis, a standardized modeling scheme for high-voltage switchgear connections is needed to improve the speed and response of subsequent intelligent analysis of specific nodes.
[0050] Existing traditional methods for training neural network AI to perform thermal analysis on other electrical equipment and components involve designing a neural network method, acquiring a large amount of measured data in the field, starting with an arbitrary initial performance planar model, training the neural network model, obtaining a surface model of the electrical equipment's thermal performance system through the system identification of the trained neural network, and then predicting the node temperature based on the surface model.
[0051] This traditional training method requires training a surface model of the thermal performance system that conforms to practice from a large amount of field measurement data.
[0052] However, training data obtained from long-term actual measurements in the field may encounter the following problems due to issues with system sampling, communication, and monitoring:
[0053] 1. Due to instrument and equipment-related reasons, such as improper instrument calibration, the thermometer may develop deviations after prolonged use. Failure to perform regular calibration will lead to systematic measurement errors. When the battery power of an infrared thermometer is low, the laser indicator, display, and sensor may malfunction, resulting in unstable or erroneous readings. For example, a mismatch in the distance-to-spot (D:S) ratio, such as being too far from the target or aiming at a spot that is too large, may result in measuring the average temperature of the target object and its surrounding background instead of the precise temperature of the switch contact. Inaccurate emissivity settings, such as the low and variable emissivity of the metal oxide surface of the switch, can also lead to significantly lower readings if the default value of 0.95 is simply used. Furthermore, if the instrument's optical system experiences a sudden temperature change (such as moving from an air-conditioned room to direct sunlight), and measurements are started before reaching thermal stability, the readings will drift.
[0054] 2. In environments with strong interference and under certain measurement conditions, such as switch stations which are in strong electromagnetic environments, interference may occur to the electronic components of the infrared thermometer, resulting in jumps or erroneous readings. Additionally, background radiation interference, severe weather conditions, and the presence of dense fog, dust, drizzle, or smoke along the measurement path can absorb and scatter infrared radiation, leading to lower readings.
[0055] 3. Most importantly, the cause may be human error or the characteristics of the target. For example, aiming errors may occur, where the laser point is not precisely aligned with the critical area to be measured, but instead measures the support, spring, or background. The contact point being measured may be smaller than the instrument's minimum spot size, resulting in a measurement that is a mixture of the target and background temperatures. Alternatively, at the moment of measurement, the equipment's load current may be fluctuating, causing the temperature to change rapidly, making a single reading unrepresentative.
[0056] If data with large errors is not removed and cleaned when processing training data, it may cause AI training data pollution, resulting in inaccurate thermal analysis.
[0057] The conventional method involves collecting pyrography data multiple times under fixed conditions and using confidence interval error analysis to remove gross error data outside the intervals. However, when this method is applied to data with multiple influencing factors or multidimensional data, it is often impossible to obtain confidence intervals for samples under the same conditions due to inconsistent environmental conditions of the detection data. Therefore, it is impossible to clean and remove invalid detection data.
[0058] The heating data error of the hard connection contact point at the output terminal of a high-voltage switch generally includes three parts: systematic error, random error, and gross error.
[0059] 1. Random error is a measurement error that occurs when the same measurand is repeatedly measured under identical conditions, and varies in an unpredictable manner. Its causes include environmental fluctuations, noise from the measuring equipment itself, minute changes in contact resistance, and fluctuations in observer interpretation; slight differences may exist in each measurement. Its characteristic is unpredictability, but it possesses compensatory properties. Its distribution usually follows a normal distribution, meaning small errors have a high probability of occurrence, and large errors have a low probability of occurrence. It cannot be completely eliminated, but it can be reduced by methods such as performing multiple measurements.
[0060] 2. Systematic error is an error that remains constant or varies according to a certain definite pattern when the same measurand is repeatedly measured under the same measurement conditions. It is caused by instrument calibration errors, usually due to imperfections in the measurement method or theory, or incorrect emissivity settings. Its characteristics include constancy or regularity; the error value is constant under certain conditions, or varies according to a certain functional law (such as linearity or periodicity), usually causing the measurement results to systematically bias to one side. It possesses both irreparability and correctability: it cannot be detected and eliminated by increasing the number of measurements; the average of multiple measurements still contains systematic error. However, once the source and magnitude of the systematic error are discovered through calibration or theoretical analysis, the measurement results can be corrected using correction values or formulas, thereby essentially eliminating its influence.
[0061] 3. Gross errors are errors that significantly exceed the expected specifications. They are usually caused by mistakes, negligence, or sudden external interference during the measurement process, resulting in severe distortion of the measurement results. For AI neural networks, this can be considered contaminated data, and directly substituting it will affect the accuracy of the final surface model. The causes generally include operational errors, such as misreading data, incorrect aiming of the infrared thermometer, seriously incorrect parameter settings, improper use or malfunction of equipment, and sudden, strong interference. Its characteristic is anomaly; the absolute value of the error is very large, significantly abnormal compared to other normal measurement values, and it is not an inherent uncertainty of the measurement process. It is something that can and must be avoided. Measurement values containing gross errors are called "bad values" or "outliers," and once discovered, they should be discarded and cannot be used in the final data processing. In data processing, they can be identified using physical discrimination methods (analyzing measurement conditions) or statistical discrimination methods (such as the Laida criterion and Grubbs criterion).
[0062] After the heat generation data of the hard connection contact point at the outgoing terminal of the high-voltage switch is collected, the AI neural network can compensate for the influence of random errors through multiple training sessions. However, data containing gross errors should not be included in the AI training. Therefore, data cleaning should be performed to first filter out data containing gross errors.
[0063] Current field data cleaning methods utilize data stacking, involving multiple data collections at a single point under identical conditions and averaging the results. For large datasets, a normal distribution with confidence intervals can be used to filter out contaminated data caused by gross errors. The general process is as follows: Figure 2 As shown.
[0064] In actual high-voltage switchgear outgoing line hard connection contact point temperature measurement, the temperature rise data is not only affected by the current magnitude, but also by various environmental factors such as radiation, wind speed, wind direction, and humidity. Therefore, the measurement data obtained on site, even under the same current conditions, are mostly not temperature rise data obtained under the same conditions due to different wind speed and other conditions during measurement. They do not conform to the probability distribution of independent events under the same conditions. Therefore, it is not possible to directly use methods such as Wright's criterion to clean the data in the above process to eliminate bad values containing gross errors.
[0065] This application provides a data pre-cleaning method for thermal analysis of AI training data at hard-connected contact points. Please refer to [link to relevant documentation]. Figure 1 As shown, the cleaning method includes the following steps:
[0066] The high-voltage switchgear on site is classified into frames, and then a three-dimensional model is built for each classified high-voltage switchgear. A thermal shape simulation model is generated using finite element analysis to obtain the thermal performance surface corresponding to the thermal analysis structure of the three-dimensional model of the high-voltage switchgear under a certain contact resistance condition. When analyzing the hard connection contact points of the high-voltage switchgear, a thermal shape simulation model is selected, and a simulated thermal parameter surface model is constructed using basic data. This simulated thermal parameter surface model includes a temperature rise surface model under contact resistance. Training data collected on site is used as the dataset to be cleaned. The temperature rise surface model is sliced using the on-site training data to obtain curves. The curves are compared with the training data to obtain normalized coefficient labels. Each training data point is labeled with a normalized coefficient. These normalized coefficient labels are used as statistical data samples (X samples), and then tabulated to create a sample set of coefficient labels to be cleaned. Statistical discrimination methods (such as the Wright criterion) are used to identify the sample set of coefficient labels to be cleaned. Normalized coefficient labels with variances exceeding the confidence interval are identified, and the on-site training data corresponding to these normalized coefficient labels (i.e., bad values containing gross errors) are removed to obtain a reconstructed and cleaned on-site training dataset.
[0067] The following is a detailed description of the contents of this application:
[0068] Furthermore, the high-voltage switchgear on site is classified into frames, and then a three-dimensional model is established for the classified high-voltage switchgear. The thermal shape simulation model is generated by finite element analysis, and the thermal performance surface corresponding to the thermal analysis structure of the three-dimensional model of the high-voltage switchgear is obtained under a certain contact resistance condition.
[0069] A framework classification of high-voltage switchgear in the field is necessary. Because high-voltage switchgear in real-world applications may exhibit various structural designs, material selections, contact point configurations, heat dissipation conditions, and rated parameters (such as rated current, rated voltage, and mechanical life), these differences directly affect the equipment's thermal conductivity and thermal response behavior. Therefore, before subsequent modeling and simulation, it is necessary to formulate scientific and reasonable classification rules based on the equipment's structural type, functional purpose, key design parameters, and operating conditions.
[0070] The high-voltage switches involved in the field data collection are divided into several representative and homogeneous categories (e.g., single-column disconnect switches, double-column disconnect switches, three-column disconnect switches, etc.). This classification process is typically based on multi-source information such as equipment technical specifications, field installation parameters, historical operation and maintenance data, and structural drawings. Category division is achieved through expert experience rules, ensuring that switchgear within the same category has high similarity in thermophysical characteristics and operational behavior, while significant differences exist between different categories. The accuracy and rationality of the classification directly affect the relevance and effectiveness of subsequent 3D modeling and simulation.
[0071] A 3D geometric model is established for each type of switchgear. After classifying the equipment framework, a 3D geometric model that fully reflects the physical structure and material distribution of typical switchgear (or representative equipment models) within each category needs to be constructed based on its actual structural design parameters (such as contact point size, conductive path layout, housing shape, insulation material distribution, heat dissipation structure, etc.). This model should not only include the main conductive and contact components but also cover all key structures affecting heat conduction and convection heat dissipation, such as heat sinks, housing ventilation holes, insulation supports, and connecting conductors, ensuring that the model is highly consistent with the actual equipment in terms of geometry and material properties. During the modeling process, accurate physical property parameters, including thermal conductivity, specific heat capacity, density, and convective heat transfer coefficient, need to be assigned to different materials (such as copper contacts, steel housing, epoxy insulation components, etc.).
[0072] Thermal simulation calculations were performed using the finite element method (FEM). After obtaining a high-precision three-dimensional geometric model and setting reasonable material properties and boundary conditions (such as ambient temperature, initial contact resistance, current load, and heat dissipation conditions), FEM software was used to numerically simulate the heat conduction, convection, and radiation behavior of the switchgear under specific operating conditions. This simulation primarily focuses on how the Joule heating effect caused by contact resistance generates temperature distribution within the internal structure of the switchgear during energized operation, and how this temperature distribution, combined with heat conduction paths and external heat dissipation conditions, forms a steady-state or transient thermal response. During the simulation, different contact resistance values were set (covering typical ranges that may occur in actual operation, such as from normal low resistance to abnormally high resistance), and multiple independent thermal analysis calculations were performed to obtain key thermal performance parameters such as the internal temperature field distribution, hot spot locations, temperature rise amplitude, and thermal response time constant of the switchgear under different contact resistance conditions.
[0073] Based on the above simulation results, a thermal performance surface is extracted and constructed. This surface is a set of mapping relationships between a high-dimensional parameter space (with contact resistance as the main variable, but may also include other influencing factors such as current intensity and ambient temperature) and thermal response outputs (such as contact temperature, average temperature rise, maximum temperature difference, steady-state temperature rise curve, etc.). By integrating and visualizing the temperature response data obtained under different contact resistance conditions (for example, constructing a two-dimensional or three-dimensional surface with contact resistance as the horizontal axis and temperature rise amplitude or temperature value as the vertical axis), a thermal performance surface reflecting the correspondence between contact resistance and thermal response is obtained.
[0074] Finally, when analyzing the hard-connection contact points of specific high-voltage switch disconnectors, a suitable thermal shape simulation surface model is selected. Using the basic data, a simulation thermal parameter surface model incorporating the temperature rise surface model under contact resistance is constructed. Based on the aforementioned thermal performance surface, the key thermal parameters directly related to contact resistance and most significantly affecting the quality of AI training data—namely, temperature rise characteristics—are further screened and extracted to construct a temperature rise surface model. This model uses contact resistance as the input variable and outputs the temperature rise amplitude, temperature rise rate, or steady-state temperature value of key parts of the equipment (such as main contacts, connecting conductors, and casing), forming a high-dimensional mapping surface with clear physical meaning and engineering application value.
[0075] Furthermore, the training data collected on-site is used as the dataset to be cleaned. The temperature rise surface model is sliced to obtain curves by combining the on-site training data. The curves are compared with the training data to obtain normalized coefficient labels, and each training data is labeled with a normalized coefficient label.
[0076] Training data collected on-site typically contains multi-dimensional information, such as time-series temperature samples, switching current, ambient temperature, wind speed, contact resistance (or estimable resistance values), and switch type. This data reflects the transient or steady-state thermal response behavior of the switching equipment during actual operation. Before normalization, it is necessary to locate the most suitable model region from a pre-constructed temperature rise surface model based on the switching equipment type, operating parameters (such as load type and environmental conditions), and estimated or measured contact resistance for each data point. This means determining the subset of operating conditions corresponding to that data point. The temperature rise surface model is a comprehensive thermal response output set obtained through simulation using finite element analysis or other thermal simulation methods under multiple combinations of contact resistance and operating parameters. It typically uses contact resistance as the primary variable and may include other influencing dimensions. Its output is the corresponding temperature response curve or temperature field distribution.
[0077] The temperature rise surface model is sliced to extract the target temperature rise response curve. After determining the operating conditions corresponding to the data to be cleaned, the temperature rise surface model is sliced using the on-site wind speed, contact resistance, or switching current at the hard-connected contact points to obtain the standard temperature rise response curve corresponding to those conditions. This process is essentially a projection or truncation of a high-dimensional thermal performance surface onto a cross-section with specific parameters (such as contact resistance, current value, wind speed, etc.). The output is one or more standardized temperature rise response curves, reflecting the expected temperature rise behavior of the switchgear (usually critical contact points or hot spots) under given operating conditions as time or current changes.
[0078] This standard temperature rise response curve serves as a benchmark for subsequent comparisons, representing the thermal response characteristics that the equipment should exhibit under ideal or standard conditions. Its accuracy depends on the rationality of the preliminary simulation modeling and parameter settings. The precision and applicability of the slicing operation directly affect the reliability of the subsequent comparison results. Therefore, it is necessary to ensure that the parameter combinations selected for slicing are highly consistent with the actual operating conditions of the field data. If necessary, interpolation or fitting methods can be used to generate theoretical curves under intermediate conditions.
[0079] The temperature response sequence curves from the field training data are matched and compared with the standard temperature rise response curves obtained from the temperature rise surface model slices. Since the field training data are generally actual time-series temperature samples collected, and the standard curves output by the model slices are also time- or current-related temperature response sequences, it is necessary to first align the time axes or synchronize the operating conditions of the two to ensure that they are consistent in terms of time scale, sampling frequency, or key operating condition nodes (such as the current loading time, steady-state start point, etc.).
[0080] After alignment, a point-by-point or overall feature comparison analysis is performed on the two curves. For example, the temperature values of the two curves at key time points (such as the starting point of heating, peak temperature, steady-state temperature, etc.) are extracted, and the absolute or relative difference between the corresponding time points is calculated; or a temperature series within a certain time window is selected, and the statistical characteristics such as the mean, variance, and maximum deviation of the temperature difference within the window are calculated; or a sliding window method is used to compare the changing trends and slope differences between the field data and the standard curve over a short period of time. The goal of this step is to evaluate the similarity or deviation between the field data and the standard model in terms of thermal response behavior from multiple dimensions, providing multi-faceted input for the subsequent calculation of normalization coefficients.
[0081] Next, the normalization coefficient is calculated based on the comparison results. This normalization coefficient is a dimensionless scalar indicator used to quantitatively characterize the degree of fit or deviation of a single field training data point relative to the standard temperature rise response curve. Its calculation logic is as follows:
[0082] First, based on the aforementioned feature comparison results, select one or more key indicators that can effectively reflect the degree of deviation (such as mean temperature difference, peak deviation, steady-state error, and area under the curve difference). Then, integrate these indicators into a single value through normalization mapping, which serves as the normalization coefficient for this training data.
[0083] In this process, if the field training data is highly consistent with the standard curve in all key features, the normalization coefficient should approach the preset ideal reference value (e.g., 1.0 or 0.0, the specific definition can be adjusted according to actual needs); conversely, if there is a significant deviation (e.g., high temperature response, response lag, peak deviation, etc.), the normalization coefficient will deviate significantly from the ideal value.
[0084] Each piece of on-site training data is labeled with a normalized coefficient. The calculated normalized coefficient will serve as an independent attribute or metadata label for that training data piece, and will be stored in conjunction with the original data record. This label not only quantifies the degree to which the training data piece deviates from the standard model in terms of its thermal response characteristics, but also provides a direct quantitative basis for subsequent data quality assessment, anomaly detection, and screening. In practice, this label can be stored in an extended field of the dataset, or associated with the original data record as an independent label table, ensuring that it can be accurately retrieved and analyzed in the subsequent data cleaning process.
[0085] Furthermore, the normalized coefficient labels of the training data are used as statistical data X samples, and then listed in a table to create a sample set of coefficient labels to be cleaned. The statistical discriminant method is used to identify the sample set of coefficient labels to be cleaned, find the normalized coefficient labels whose variance exceeds the confidence interval, and remove the field training data corresponding to the normalized coefficient labels to obtain the reconstructed and cleaned field training dataset.
[0086] Construct a statistical data table of normalized coefficient labels (i.e., a sample set of coefficient labels to be cleaned). Arrange and record the normalized coefficient labels corresponding to each piece of field training data sequentially in a structured table according to predetermined data organization rules, forming a statistical sample table with data samples as rows and normalized coefficients as key statistical observations. This table typically contains multiple fields, such as, but not limited to: a unique identifier for the data sample, the category of the switching equipment, the collected operating parameters (such as current, ambient temperature, etc.), the original data file index, and the normalized coefficient label field.
[0087] The construction process of this statistical sample table requires that the data arrangement order strictly correspond to the sample order in the original training dataset, ensuring that each normalized coefficient label can be accurately traced back to its source data, providing a clear mapping relationship for subsequent anomaly localization and data cleanup operations. Furthermore, if there are multiple categories or operating condition groups, category labels or grouping fields can be added to the table to support group statistics and differential discrimination, further improving the accuracy of anomaly identification.
[0088] This step involves analyzing the distribution of normalized coefficient labels and identifying outliers using statistical discriminant analysis. Based on the constructed sample set of coefficient labels to be cleaned, statistical discriminant logic is employed to analyze the overall distribution characteristics and detect outliers of all normalized coefficient labels. The processing logic aims to evaluate the overall distribution pattern of all normalized coefficient labels and identify outliers that significantly deviate from this pattern.
[0089] Specifically, the global statistical characteristics of the normalized coefficient labels are calculated, including but not limited to descriptive statistics such as mean (or expected value), variance (or standard deviation), range, skewness, and kurtosis, to characterize the central tendency and dispersion of the group. Based on these statistical characteristics, a reasonable confidence interval (or reasonable range of fluctuation) is set. This confidence interval is usually based on statistical principles (such as the ±3σ principle under the assumption of normality, or based on the upper and lower bounds of percentiles, such as the 5% and 95% percentile intervals) to define which normalized coefficient labels are within the normal range of fluctuation and which may be abnormal deviations.
[0090] Identify outlier normalized coefficient labels whose variance exceeds the confidence interval. After completing the calculation of overall or grouped statistical characteristics and setting the confidence interval, each normalized coefficient label is evaluated one by one to determine whether its value falls within the preset confidence interval range. If the value of a normalized coefficient label exceeds the upper or lower limit of the interval, it is determined to be an outlier label, indicating that the corresponding field training data has a significant deviation from the group standard model in terms of thermal response characteristics, which may originate from various factors such as abnormal data acquisition, abnormal equipment status, model adaptation bias, or simulation condition mismatch. This discrimination process can be implemented by traversing each label, batch filtering, or automated labeling to ensure that each outlier label can be accurately identified and its corresponding data index or identifier recorded, providing a clear target for subsequent data cleaning operations.
[0091] Remove the field training data corresponding to abnormal normalization coefficient labels. Based on the aforementioned anomaly identification results, locate all original field training data records associated with normalization coefficient labels identified as abnormal, and perform data removal operations. This removal process can be physical deletion (removing the corresponding data sample from the dataset), logical marking (marking abnormal data as unavailable or pending verification), or isolated storage (transferring abnormal data to an independent backup area for subsequent review). The specific method is selected based on the actual application scenario and data management requirements. The core objective is to ensure that the normalization coefficients of all samples in the final retained training dataset are within a reasonable statistical fluctuation range, that is, their thermal response characteristics have high consistency with the standard simulation model, thereby significantly reducing the interference of data noise on the AI model training process.
[0092] Obtain the reconstructed and cleaned field training dataset. After identifying and removing outlier data, the remaining field training data (i.e., data samples where all normalized coefficient labels are within the confidence interval) will constitute a high-quality, highly consistent reconstructed and cleaned dataset. The reconstructed dataset retains key information fields from the original data (such as time-series temperature data, operating parameters, equipment categories, etc.), while implicitly containing high-quality signals conveyed through normalization coefficient filtering.
[0093] The following detailed implementation methods illustrate the content of this application: Example
[0094] For example, in handling gross errors in the heating data of the hard-connected contact points at the outgoing terminals of high-voltage switch disconnectors, the following steps are taken:
[0095] Step 1: Model the high-voltage switchgear object:
[0096] In hard-contact heating structures, the final temperature value is affected not only by environmental indices and node heating, but also by the heating generated by current flowing through the conductor and other high-temperature nodes. For double-column disconnect switches, since the disconnect switch contacts are relatively close to the outgoing hard-contact connection, their heating effect must be considered, and modeling and simulation should be performed according to the heat source.
[0097] Based on the preliminary analysis of disconnector drawings, disconnectors from different manufacturers generally need to meet the national standard current carrying capacity requirements. Therefore, for disconnectors of the same type and grade, the outgoing line installation requirements and the heat-conducting base are similar. The cross-sectional area of the conductive outgoing line material and the shape of the rigid connection to the base are also similar. Therefore, the most reasonable classification method is to establish a model based on the current rating of the disconnector frame. Thus, the classification according to the requirements of the project research object is shown in Table 1:
[0098] Table 1:
[0099] Major Categories name name name name Single-column simulation model Single column 2000A Single column 2500A Single-column 3150A Single column 4000A Two-column simulation model 2000A with two columns 2500A Twin-Column 3150A with two columns 4000A with two columns
[0100] Step 2: Establish a simulation structural model, set relevant thermally conductive materials for simulation, and perform finite element analysis of heat generation based on the established model:
[0101] Current input: Current I is used as the variable, with the unit being A, to facilitate simulation of multiple current values in the future;
[0102] Wind speed setting: Use wind speed V_wind as the variable, with the unit being m / s, to facilitate subsequent simulations of multiple wind speeds. The wind direction is the Z direction, i.e., the horizontal direction.
[0103] Contact resistance setting: The power P is used as the variable, and the unit is W. The calculation is performed to facilitate subsequent simulations of multiple contact resistances, where R is the contact resistance.
[0104] Set up contact surface temperature detection to detect the maximum temperature value;
[0105] Substitute the specific values into the three parameters to form a calculation matrix, calculate each combination, and obtain the heating temperature;
[0106] This allows us to obtain simulation data on temperature rise under different currents, contact resistances, and wind speeds. For example, the simulation data obtained at a wind speed of 0 is shown in Table 2.
[0107] Table 2:
[0108]
[0109] Its constituent thermal structure surfaces are as follows Figure 3 As shown, the simulation surface of thermal temperature rise under different wind speeds can be obtained by analogy. After fitting, the surface corresponding to the wind speed and current temperature rise performance under a certain resistance can be obtained (taking 0.5 microohms and 32 microohms as examples), such as... Figure 4 As shown.
[0110] Step 3: Collect training data on-site. For example, simplified initial on-site data is shown in Table 3:
[0111] Table 3:
[0112] Serial Number Switching current value (A) On-site wind speed (m / s) Temperature rise (degrees) 1 2000 4 31.5 2 2560 8 42.5 3 2805 6 33 4 2300 0.5 41 5 2850 1 56 6 2100 0.5 54 7 3400 0.5 78 8 3420 8 43 9 3380 2 52 10 2790 0.5 73
[0113] Step 4: In the simulated performance curve from Step 2, using a surface with a fixed contact resistance (e.g., 0.8 microohms), slice the surface by wind speed to obtain the performance curve. Calculate and find the simulated temperature rise value under the corresponding current, compare it with the training data, and divide the training data by the simulated temperature rise value to obtain the corresponding coefficient label, such as... Figure 5 As shown, Figure 5 The 10 data points are the data with the corresponding 10 serial numbers in Table 3. Then, the labels for the training data are recorded as shown in Table 4.
[0114] Table 4:
[0115] Serial Number Switching current value (A) On-site wind speed (m / s) Temperature rise (degrees) Simulated temperature value at this wind speed: 0.8 microohms Coefficient label value 1 2000 4 31.5 19.8 1.590909091 2 2560 8 42.5 28.5 1.49122807 3 2805 6 33 36 0.916666667 4 2300 0.5 41 34 1.205882353 5 2850 1 56 52 1.076923077 6 2100 0.5 54 28.5 1.894736842 7 3400 0.5 78 66 1.181818182 8 3420 8 43 50.1 0.858283433 9 3380 2 52 56.5 0.920353982 10 2790 0.5 73 51 1.431372549
[0116] Step 5: Use the data normalization coefficient labels as statistical data X samples, and use statistical discriminant analysis to identify data labels whose variance exceeds the confidence interval.
[0117] (1) Calculate the average value:
[0118] In the formula, n is the number of normalized coefficient labels. Let X be the label value of the i-th normalized coefficient in the statistical data sample X.
[0119] (2) Calculate the residuals The results are shown in Table 5:
[0120] Table 5:
[0121] Coefficient label value residual 1.590909091 0.371212963 1.49122807 0.271531942 0.916666667 -0.303029462 1.205882353 -0.013813775 1.076923077 -0.142773051 1.894736842 0.675040714 1.181818182 -0.037877947 0.858283433 -0.361412695 0.920353982 -0.299342146 1.431372549 0.211676421
[0122] (3) Calculate the standard deviation of the data .
[0123] (4) Using Grubbs' test , The maximum allowable value for filtering residuals is G, which is the cutoff coefficient set by the Grubbs test. It is generally set to 0 to 3. The larger the value, the wider the confidence interval. In this example, G=1 is used to set the confidence interval.
[0124] (5) It can be seen that the data with serial numbers 1, 6, and 8 exceed the confidence interval and should be filtered out because they contain large errors, as shown in Table 6:
[0125] Table 6:
[0126] Serial Number Switching current value (A) On-site wind speed (m / s) Temperature rise (degrees) Simulated temperature value at this wind speed: 0.8 microohms Coefficient label value residual Data cleaning and retention 1 2000 4 31.5 19.8 1.590909091 0.371212963 X Elimination 2 2560 8 42.5 28.5 1.49122807 0.271531942 reserve 3 2805 6 33 36 0.916666667 -0.303029462 reserve 4 2300 0.5 41 34 1.205882353 -0.013813775 reserve 5 2850 1 56 52 1.076923077 -0.142773051 reserve 6 2100 0.5 54 28.5 1.894736842 0.675040714 X Elimination 7 3400 0.5 78 66 1.181818182 -0.037877947 reserve 8 3420 8 43 50.1 0.858283433 -0.361412695 X Elimination 9 3380 2 52 56.5 0.920353982 -0.299342146 reserve 10 2790 0.5 73 51 1.431372549 0.211676421 reserve
[0127] By following the steps above, a new training dataset after data cleaning can be obtained, which will be used in subsequent AI neural network training, as shown in Table 7.
[0128] Table 7:
[0129] Serial Number Switching current value (A) On-site wind speed (m / s) Temperature rise (degrees) 2 2560 8 42.5 3 2805 6 33 4 2300 0.5 41 5 2850 1 56 7 3400 0.5 78 9 3380 2 52 10 2790 0.5 73
[0130] This application utilizes a finite element method (FEM) model for thermal conduction. First, the high-voltage switchgear in the field is classified into frames. Then, a three-dimensional model is established for each typical high-voltage switchgear. A thermal shape simulation model is generated using FEM, yielding a thermal performance surface under a specific contact resistance condition corresponding to the thermal analysis structure of the typical three-dimensional high-voltage switchgear model. When analyzing the hard-connection contact points of a specific high-voltage switchgear, a suitable typical thermal shape simulation model is selected, and a simulated thermal parameter surface model is constructed using basic data. This simulated thermal parameter surface model includes a temperature rise surface model under a specific contact resistance.
[0131] The process of slicing the performance surface using data: Combining on-site training data and based on the on-site data parameters, the simulated temperature rise surface model is sliced, and the obtained curves are compared with the training data to obtain normalized coefficient labels. Each training data point is then labeled with this normalized coefficient.
[0132] The normalized coefficient labels are used as statistical data samples X, and listed in a table. Statistical discriminant methods (such as Wright's criterion and Grubbs' criterion) are used to identify the normalized data labels whose variance exceeds the confidence interval.
[0133] The data to be cleaned is the field training data corresponding to the normalized data label (i.e., bad values containing gross errors). The training data is not used directly for cleaning calculations; the cleaned field training dataset is reconstructed.
[0134] In the description of this specification, references to terms such as "an embodiment," "example," "specific example," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of the invention. In this specification, illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples.
[0135] The preferred embodiments of the present invention disclosed above are merely illustrative of the invention. These preferred embodiments do not exhaustively describe all details, nor do they limit the invention to any specific implementation. Clearly, many modifications and variations can be made based on the content of this specification. This specification selects and specifically describes these embodiments to better explain the principles and practical applications of the invention, thereby enabling those skilled in the art to better understand and utilize the invention. The invention is limited only by the claims and their full scope and equivalents.
Claims
1. A data pre-cleaning method for thermal analysis of AI training data at hard-connected contact points, characterized in that: The cleaning method includes the following steps: The high-voltage switch disconnectors on site are classified into frames, and then three-dimensional models are established for the classified high-voltage switch disconnectors. The thermal shape simulation model is simulated using finite element analysis to obtain the thermal performance surface corresponding to the thermal analysis structure of the three-dimensional model of the high-voltage switch disconnector under a certain contact resistance condition. When analyzing the hard connection contact point of the high-voltage switch knife switch, a thermal shape simulation model is selected, and a simulation thermal parameter surface model is constructed using basic data. The simulation thermal parameter surface model includes a simulation temperature rise surface model under contact resistance. The training data collected on-site is used as the dataset to be cleaned. Curves are obtained by slicing the simulated temperature rise surface model using the training data. The curves are then compared with the training data to obtain normalized coefficient labels. Each training data point is labeled with a normalized coefficient label, including the following steps: Determine the operating conditions corresponding to the data to be cleaned, and obtain the standard temperature rise response curve corresponding to the operating conditions by slicing the temperature rise surface model through the on-site wind speed, contact resistance or switching current of the hard-connected contact point. The temperature response sequence curves in the training data are aligned and compared with the standard temperature response curves obtained by slicing the temperature rise surface model. The two curves are compared and analyzed point by point or as a whole, including the following steps: extracting the temperature values of the two curves at time nodes and calculating the absolute or relative difference between the corresponding time nodes; or selecting the temperature sequence within a time window and calculating the mean, variance, and maximum deviation of the temperature difference within the time window; or comparing the changing trends of the training data and the standard curve segment by segment using a sliding window method. The two curves are compared and analyzed point by point or as a whole. The normalization coefficient is calculated based on the comparison results to quantitatively characterize the degree of fit or deviation of a single training data point relative to the standard temperature rise response curve. Each training data point is labeled with a normalization coefficient. The calculated normalization coefficient serves as the data label for that training data point and is bound to the original data record for storage. Before normalization, based on the type of switching equipment, operating parameters, and estimated or measured contact resistance of each data point, the model region that best matches the data point is located from the pre-built simulation temperature rise surface model, thus determining the subset of operating conditions corresponding to the data. The simulation temperature rise surface model is a comprehensive thermal response output set obtained by simulating multiple combinations of contact resistance and operating parameters based on finite element analysis or other thermal simulation methods. With contact resistance as the main variable, the output is the corresponding temperature response curve or temperature field distribution. The normalized coefficient labels of the training data are used as statistical data X samples. They are listed in a table to create a sample set of coefficient labels to be cleaned. The statistical discriminant method is used to identify the sample set of coefficient labels to be cleaned, find the normalized coefficient labels whose variance exceeds the confidence interval, and remove the training data corresponding to the normalized coefficient labels to obtain the reconstructed and cleaned on-site training dataset.
2. The data pre-cleaning method for thermal analysis of AI training data at hard-connected contact points according to claim 1, characterized in that: The normalization coefficient is calculated based on the comparison results. It is used to quantitatively characterize the degree of fit or deviation of a single training data point relative to the standard temperature rise response curve. The calculation logic is as follows: Select one or more indicators that can reflect the degree of bias; integrate the indicators into a single value through normalization mapping, which serves as the normalization coefficient for the training data. If the training data is consistent with the standard temperature rise response curve in all features, then the normalization coefficient approaches the preset ideal reference value. Conversely, if there is a deviation, the normalization coefficient will deviate from the ideal value.
3. The data pre-cleaning method for thermal analysis of AI training data at hard-connected contact points according to claim 2, characterized in that: The training data collected on-site includes time-series temperature samples, switching current, ambient temperature, wind speed, contact resistance, and high-voltage switch type.
4. The data pre-cleaning method for thermal analysis of AI training data at hard-connected contact points according to claim 1, characterized in that: The normalized coefficient labels of the training data are used as statistical data X samples. After being listed in a table, a sample set of coefficient labels to be cleaned is created, including the following steps: Construct a statistical data table of normalized coefficient labels. Arrange and record the normalized coefficient labels corresponding to each training data point in a structured table according to the data organization rules, forming a statistical sample table with data samples as rows and normalized coefficients as statistical observations.
5. The data pre-cleaning method for thermal analysis of AI training data at hard-connected contact points according to claim 4, characterized in that: The statistical discriminant method is used to identify the sample set of coefficient labels to be cleaned, find the normalized coefficient labels whose variance exceeds the confidence interval, and remove the training data corresponding to the normalized coefficient labels to obtain the reconstructed and cleaned on-site training dataset. This includes the following steps: Based on the completed sample set of coefficient labels to be cleaned, statistical discriminant logic is used to perform overall distribution characteristic analysis and outlier detection on all normalized coefficient labels, evaluate the overall distribution pattern of all normalized coefficient labels, and identify outliers that deviate from the pattern. Calculate the global statistical characteristics of the normalized coefficient labels, and based on the global statistical characteristics, set a confidence interval to define whether the normalized coefficient labels are within the normal fluctuation range or deviate abnormally. Each normalized coefficient label is evaluated one by one to determine whether its value falls within the preset confidence interval. If the value of a normalized coefficient label exceeds the upper or lower limit of the interval, it is determined to be an abnormal label. Locate all original training data records associated with the normalized coefficient labels that are judged to be abnormal, and perform data cleanup operations; After identifying and removing anomalous data, the remaining training data constitutes the reconstructed and cleaned on-site training dataset.
6. The data pre-cleaning method for thermal analysis of AI training data at hard-connected contact points according to claim 5, characterized in that: The statistical sample table includes a unique identifier for each data sample, the category of the high-voltage switchgear, the collected operating parameters, the index of the original data file, and a normalization coefficient label field.
7. The data pre-cleaning method for thermal analysis of AI training data at hard-connected contact points according to claim 6, characterized in that: The global statistical features include mean, variance, range, skewness, and kurtosis.
8. The data pre-cleaning method for thermal analysis of AI training data at hard-connected contact points according to claim 1, characterized in that: The high-voltage switchgear on site is classified into frames, and then a three-dimensional model is built for the classified high-voltage switchgear. The thermal shape simulation model is generated using finite element analysis, and the thermal performance surface corresponding to the thermal analysis structure of the three-dimensional model of the high-voltage switchgear is obtained under a certain contact resistance condition. The process includes the following steps: The high-voltage switchgear involved in the field data collection is divided into several categories according to the framework; After classifying the high-voltage switchgear frame, a three-dimensional model reflecting the physical structure and material distribution of each high-voltage switchgear in each category is constructed based on its structural design parameters. Material properties and boundary conditions were set for the three-dimensional model, and multiple sets of simulation results were obtained by numerical simulation of the high-voltage switch disconnector using finite element analysis software. Based on multiple sets of simulation results, thermal performance surfaces are extracted and constructed. By integrating and visualizing the temperature response data obtained under different contact resistance conditions, thermal performance surfaces reflecting the correspondence between contact resistance and thermal response are obtained.
Citation Information
Patent Citations
Running state analysis method and system of dry-type transformer, server and medium
CN113553927A
Temperature field prediction model training method and temperature field reconstruction method
CN120493752A