A power distribution switchgear reliability evaluation method and system based on a large model
Patent Information
- Application Number
- CN202610554527.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-04-24
- Publication Date
- 2026-09-22
- Estimated Expiration
- 2046-04-24
AI Technical Summary
[0003]传统模型无法挖掘多维时序特征的内在关联,难以形成具备上下文感知能力的深度状态表征,特征利用效率较低
改进注意力机制能够对多维运行状态特征开展深层特征交互,依据各维度特征对设备状态的影响程度完成自适应重要性加权,强化特征间的关联耦合关系,弱化冗余特征与干扰特征的影响。综合数据治理后的标准化时序数据可保障特征输入的一致性与连续性,深度特征交互可挖掘多维特征间的内在关联,上下文感知可保留时序数据的连续变化特性,形成能够完整反映设备运行状态的深度状态表征,提升状态特征的表达精度与区分度,优化特征表征的有效性与针对性。
Smart Images

Figure CN122432559B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of intelligent evaluation technology for power equipment, specifically a reliability evaluation method and system for power distribution switchgear based on a large model. Background Technology
[0002] Current reliability assessments of power distribution switchgear employ traditional machine learning and shallow neural network models, which only perform simple processing on equipment monitoring data. They lack comprehensive data governance, including missing value imputation, dimension normalization, and time-series alignment, resulting in insufficient data standardization. Existing models utilize conventional attention mechanisms, enabling only independent feature extraction and shallow computation, failing to achieve deep interaction and precise importance weighting of multi-dimensional operational features. Current technologies separate potential failure mode identification and remaining service life prediction into independent models, employing a step-by-step inference approach to complete the assessment.
[0003] Traditional models fail to uncover the intrinsic relationships between multi-dimensional temporal features, making it difficult to form deep state representations with context-aware capabilities, resulting in low feature utilization efficiency. Independent modeling methods prevent the reuse of features between fault analysis and lifetime prediction, leading to fragmented reasoning logic for the two tasks and a lack of consistency in evaluation results. Furthermore, the lack of systematic temporal alignment and normalization of monitoring data results in insufficient continuity and comparability of state features, impacting model representation effectiveness.
[0004] Currently, there is a need to standardize and comprehensively manage monitoring data, achieving deep feature interaction and adaptive weighting through an optimized attention mechanism to form a context-aware deep state representation. Simultaneously, a unified network architecture is required to synchronously complete the analysis of potential failure mode probability distributions and the quantification of remaining lifetime, addressing the issues of fragmented assessment results and insufficient feature mining. Summary of the Invention
[0005] This invention aims to solve at least one of the technical problems existing in the prior art; Therefore, this invention proposes a reliability assessment method for power distribution switchgear based on a large model, including: Collect multi-dimensional monitoring data of the target power distribution switchgear during its historical operating cycle to form a raw monitoring data set; The original monitoring data set is subjected to comprehensive data governance, including missing value imputation, unit normalization, and time series alignment, to generate a standardized monitoring time series dataset; Multiple operational status features are extracted from the standardized monitoring time-series dataset to construct an operational status feature vector set; The set of running state feature vectors is input into a pre-trained reliability assessment large language model. The improved attention mechanism contained in the reliability assessment large language model is used to perform deep feature interaction and importance weighting to generate a context-aware deep state representation. Using the deep state representation, a reliability regression network is used to perform synchronous quantitative analysis of potential failure modes and remaining useful life, and outputs a reliability assessment report containing the probability distribution of potential failure modes and the quantified remaining useful life.
[0006] Furthermore, comprehensive data governance is performed on the original monitoring data set, including missing value imputation, dimension normalization, and time series alignment, including: The numerical gaps in different monitoring signal channels within the original monitoring data set are identified, and an interpolation algorithm based on the trend inference of adjacent time points is used to fill in the numerical gaps to generate a continuous monitoring sequence. For the different physical dimensions of the data in each channel of the continuous monitoring sequence, the monitoring data of each channel is subtracted from its own historical mean and divided by its own historical standard deviation, and all monitoring data are mapped to the dimensionless standard normal distribution space. All channels' data within the standard normal distribution space are resampled according to a unified time reference to ensure that each channel's data point has an aligned value at each sampling time, ultimately generating the standardized monitoring time series dataset.
[0007] Furthermore, various operational status features are extracted from the standardized monitoring time-series dataset to construct an operational status feature vector set, including: In the time domain dimension, the statistical characteristics of each signal channel data in the standardized monitoring time series dataset are calculated, and the statistical characteristics include mean, variance, peak value and waveform factor. In the frequency domain, a fast Fourier transform is performed on the standardized monitoring time series dataset to calculate the dominant frequency, centroid frequency, and spectral entropy of each signal channel data. Construct a joint time-frequency domain feature extraction window, and calculate the variation trends of the short-time energy and zero-crossing rate of the signal within the joint time-frequency domain feature extraction window; The time-domain statistical features, frequency-domain features, and time-frequency domain variation trend features calculated for each signal channel are concatenated into a high-dimensional feature vector in a preset order. The high-dimensional feature vectors of all signal channels together constitute the operating state feature vector set.
[0008] Furthermore, the improved attention mechanism works as follows: When the reliability assessment large language model processes the runtime feature vector set, the improved attention mechanism receives the initial query vector, initial key vector, and initial value vector transformed from the runtime feature vector set; The improved attention mechanism introduces a dynamic association weight matrix, which is dynamically generated based on the metadata of the device operation stage and the metadata of the feature type, and is used to modulate the similarity calculation between the initial query vector and the initial key vector. When calculating attention weights, the modulated similarity score is input into a nonlinear sharpening function, which strengthens the weights of highly relevant feature pairs while further suppressing the weights of low-relevance feature pairs, resulting in a sharpened attention distribution. The initial value vector is weighted and fused based on the sharpened attention distribution, and residual connections and layer normalization operations are introduced to finally generate the context-aware deep state representation with enhanced discriminative power.
[0009] Furthermore, utilizing the aforementioned deep state representation, a reliability regression network is used to perform synchronous quantitative analysis of potential failure modes and remaining useful life, including: The deep state representation is input into the multi-head analysis layer of the reliability regression network; In the multi-head analysis layer, one analysis head is dedicated to processing the feature subspace associated with the failure mode. The analysis head consists of multiple fully connected layers, the ends of which are connected to a classifier to output the probability distribution of the potential failure modes. In the multi-head analysis layer, another analysis head is dedicated to processing the feature subspace related to lifetime degradation. This analysis head consists of multiple fully connected layers, with its ends connected to a regressor, which outputs the quantified remaining lifetime. The potential failure mode probability distribution and the quantified remaining useful life are encapsulated together to form the reliability assessment report.
[0010] Furthermore, the process of specifically processing the analysis head end of the feature subspace related to the fault mode and connecting it to the classifier's output includes: The analysis head performs multi-level nonlinear transformations on the input feature subspace related to the fault mode to obtain high-level fault features; The advanced fault features are input into the Softmax layer of the classifier, which maps the advanced fault features to probability values corresponding to different predefined fault modes. The output probability values are sorted, and fault modes whose probability values exceed the preset activation threshold are selected. These, along with their corresponding probability values, are used as effective fault warnings, forming the core content of the potential fault mode probability distribution.
[0011] Furthermore, the process of connecting the output of the analysis head-end regressor, which specifically processes the feature subspace related to lifetime degradation, includes: The analysis head extracts features and compresses dimensions in the input feature subspace related to lifespan degradation to obtain the lifespan degradation index. The lifetime degradation index is input into the fully connected layer of the regressor, which linearly maps the lifetime degradation index to a preliminary estimate of the remaining useful life. The preliminary estimate is input into a monotonically decreasing constraint function, which ensures that the output value does not increase with the device's operating time and is mapped to the device's maximum rated lifespan to obtain the final quantified remaining lifespan.
[0012] Furthermore, the pre-training process of the reliability assessment large language model includes: Collect multi-source heterogeneous monitoring data and corresponding fault record documents of power distribution switchgear throughout its entire life cycle, and build a pre-training corpus. Comprehensive data governance and feature extraction are performed on the multi-source heterogeneous monitoring data in the pre-training corpus to construct a massive number of operational status feature vector samples; The fault record document is structured and parsed to extract fault mode labels and equipment final lifespan labels, and then aligned and labeled with the massive operating status feature vector samples at time points. Using the massive number of running state feature vector samples and their aligned labels, based on the joint training objective of masked language modeling and next state prediction, a large-scale pre-training of the Transformer architecture's large language model is performed to obtain a preliminary model. Based on the preliminary model, the preliminary model is fine-tuned in a supervised manner using finely labeled data for specific scenarios, and finally the pre-trained reliability evaluation large language model is obtained.
[0013] Furthermore, after the output includes a reliability assessment report with a potential failure mode probability distribution and quantified remaining useful life, the method further includes: The currently generated reliability assessment report and the historical reliability assessment report of the target power distribution switchgear are stored in the same assessment file to form an assessment history sequence; Perform trend analysis on the historical evaluation sequence in the evaluation file, and calculate the probability change rate of key modes in the probability distribution of potential failure modes and the decay rate of the quantified remaining useful life. When the probability change rate or the decay rate exceeds their respective dynamic thresholds, an early warning signal is triggered, and an early warning summary document containing key change indicators is automatically generated. The method for determining the dynamic threshold includes: Select the evaluation history subsequence of the most recent stable operation phase from the evaluation archive; Calculate the mean and standard deviation of the probability change rate of key patterns in the evaluated historical subsequence, and the mean and standard deviation of the decay rate of the quantified remaining lifetime; The mean of the probability change rate is added to a certain number of times its standard deviation, and the result is used as the dynamic threshold of the probability change rate. The mean of the attenuation rate is added to a certain number of times its standard deviation, and the result is used as the dynamic threshold of the attenuation rate.
[0014] Furthermore, the present invention also includes a reliability assessment system for power distribution switchgear based on a large model, the system including a memory, a processor, and a computer program stored in the memory and running on the processor, wherein the processor, when executing the computer program, implements the steps of the reliability assessment method for power distribution switchgear based on the large model described above.
[0015] Compared with the prior art, the beneficial effects of the present invention are: Improving the attention mechanism enables deep feature interaction on multi-dimensional operational status features. It adaptively weights the importance of each feature based on its impact on the device status, strengthening the correlation and coupling between features while mitigating the influence of redundant and interfering features. Standardized time-series data after comprehensive data governance ensures the consistency and continuity of feature inputs. Deep feature interaction uncovers the intrinsic relationships between multi-dimensional features, and context awareness preserves the continuous variation characteristics of time-series data, forming a deep state representation that fully reflects the device's operational status. This improves the accuracy and discriminative power of state feature representation, optimizing its effectiveness and relevance.
[0016] Leveraging deep state representation, the reliability regression network can simultaneously perform potential failure mode probability distribution analysis and remaining lifetime quantification, achieving parallel analysis of two tasks. Failure mode identification and remaining lifetime prediction employ unified feature input and associative reasoning logic, allowing feature resources to be shared and reused between the two tasks. This avoids feature loss and logical fragmentation caused by independent modeling and reduces error accumulation from step-by-step analysis. The potential failure mode probability distribution can intuitively reflect the likelihood of different failure types occurring, while the remaining lifetime can directly output quantified values. The two assessment results are generated synchronously and possess inherent consistency, improving the overall integrity and coherence of the assessment results, optimizing the compactness of the assessment process, and ensuring the accuracy and synergy of the reliability analysis results. Attached Figure Description
[0017] Figure 1 This is a state diagram of the reliability assessment method for power distribution switchgear based on a large model as described in this invention. Figure 2 A flowchart for generating standardized monitoring time-series datasets for comprehensive data governance; Figure 3A flowchart for improving the attention mechanism. Detailed Implementation
[0018] The technical solution of the present invention will be clearly and completely described below with reference to the embodiments. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0019] See Figure 1 This invention provides a reliability assessment method for power distribution switchgear based on a large model, the specific method including: For the target power distribution switchgear, the system collects various monitoring data from its historical operating cycles. This data includes, but is not limited to, mechanical characteristics, current, voltage, temperature, and partial discharge signals, forming a raw monitoring data set. This raw monitoring data set undergoes comprehensive data governance, including imputing any missing values, normalizing data with different physical dimensions, and aligning multi-channel data along the time axis, thereby generating a clean and comparable standardized monitoring time-series dataset. From this standardized monitoring time-series dataset, the system automatically extracts diverse features characterizing the equipment's operating status. These features cover the time domain, frequency domain, and time-frequency domain, and are sequentially concatenated to construct an operating status feature vector set that comprehensively describes the equipment's state. This operating status feature vector set is input into a pre-trained reliability assessment large language model. This model incorporates an improved attention mechanism that performs deep interactive analysis and importance reweighting of the input features, ultimately outputting a deep state representation that integrates global contextual information. This deep state representation is fed into a reliability regression network, which employs a multi-head architecture to simultaneously analyze the health status of the device. One analysis head outputs the probability distribution of various failure modes that the device may experience, while the other analysis head outputs the quantified remaining useful life of the device. Together, they constitute a complete reliability assessment report.
[0020] In one embodiment of the present invention, this embodiment relates to the specific process of comprehensively managing the original monitoring data set to generate a standardized monitoring time-series dataset, and extracting features from the dataset to construct a set of operational status feature vectors. See also... Figure 2The system identifies data gaps within each monitoring signal channel of the original monitoring dataset and fills these gaps using an interpolation algorithm based on trend inference from nearby time points, generating a continuous monitoring sequence without interruption. For channel data with different physical dimensions within the continuous monitoring sequence, the system calculates the historical mean and standard deviation for each channel. Each data point is subtracted from its channel mean and then divided by its standard deviation, mapping all channel data to a standard normal distribution space with a mean of 0 and a standard deviation of 1, thus achieving dimension normalization. Within the standard normal distribution space, the system resamples all channel data according to a unified time reference, ensuring that all channels have an aligned data value at each sampling time point, ultimately forming a standardized monitoring time-series dataset.
[0021] From the standardized monitoring time-series dataset, the system calculates the statistical characteristics of each signal channel in the time domain, including mean, variance, peak value, and waveform factor. In the frequency domain, it performs a Fast Fourier Transform on the standardized monitoring time-series dataset to calculate the dominant frequency, centroid frequency, and spectral entropy of each signal channel. Furthermore, the system constructs a joint time-frequency domain feature extraction window, within which it calculates the time-varying trends of the signal's short-time energy and zero-crossing rate. For each signal channel, its calculated time-domain statistical characteristics, frequency-domain characteristics, and time-frequency-domain trend characteristics are concatenated in a preset, fixed order to form a high-dimensional feature vector. The high-dimensional feature vectors of all signal channels are aggregated to form the operating state feature vector set used as input for subsequent models.
[0022] In practical implementation, the reliability assessment method for power distribution switchgear based on large models involves comprehensive data governance of the original monitoring data set to generate a standardized monitoring time series dataset. The system identifies the numerical gaps in different monitoring signal channels in the original monitoring data set, and uses an interpolation algorithm based on the trend inference of adjacent time points to fill in the numerical gaps to generate a continuous monitoring sequence. For the different physical dimensions of the data in each channel in the continuous monitoring sequence, the monitoring data of each channel is subtracted from its own historical mean and divided by its own historical standard deviation. All monitoring data are mapped to a dimensionless standard normal distribution space. The data of all channels in the standard normal distribution space are resampled according to a unified time benchmark to ensure that each channel data point has an aligned value at each sampling time, and finally a standardized monitoring time series dataset is generated. In practical implementation, various operational status features are extracted from the standardized monitoring time-series dataset to construct an operational status feature vector set. In the time domain dimension, statistical features of each signal channel data in the standardized monitoring time-series dataset are calculated, including mean, variance, peak value, and waveform factor. In the frequency domain dimension, a fast Fourier transform is performed on the standardized monitoring time-series dataset to calculate the dominant frequency, centroid frequency, and spectral entropy of each signal channel data. A joint time-frequency domain feature extraction window is constructed, and the changing trends of short-time energy and zero-crossing rate of the signal are calculated within the joint time-frequency domain feature extraction window. The time-domain statistical features, frequency-domain features, and time-frequency domain changing trend features calculated for each signal channel are concatenated into a high-dimensional feature vector in a preset order. The high-dimensional feature vectors of all signal channels together constitute the operational status feature vector set.
[0023] In some embodiments, the interpolation algorithm based on the trend inference of adjacent time points uses a linear interpolation method to fill the numerical gaps in the original monitoring data set, as expressed by the formula: in: This is the value after filling in the blanks. It is the value of the previous valid data point. It is the value of the next valid data point. It is the timestamp of the missing point. It is the timestamp of the previous valid data point. This is the timestamp of the next valid data point. In some embodiments, the calculation process of time-domain statistical features involves the mean reflecting the average level of the signal, the variance reflecting the degree of signal fluctuation, the peak reflecting the extreme values of the signal, the waveform factor reflecting the shape characteristics of the signal, and the frequency domain feature calculation being implemented through the Fast Fourier Transform. The dominant frequency refers to the frequency at which the signal energy is concentrated, the centroid frequency refers to the weighted average frequency of the spectrum, the spectral entropy refers to the complexity measure of the spectrum, and the time-frequency domain change trend features capture the dynamic characteristics of the signal by the changes in short-time energy and zero-crossing rate within a time window.
[0024] Optionally, the dimension normalization process converts the monitoring data of each channel into a standard normal distribution, eliminating scale differences caused by different physical dimensions. Mapping to the standard normal distribution space allows subsequent feature extraction and model processing to be performed in a consistent data space. Optionally, the length of the joint time-frequency domain feature extraction window is dynamically adjusted based on the signal sampling rate and the equipment operating cycle. The construction of the operating state feature vector set ensures structural consistency and comparability of the feature vectors by concatenating features in a preset order. In essence, the comprehensive data governance steps ensure data continuity, eliminate the influence of dimensions, and achieve data synchronization, providing a clean and comparable standardized monitoring time-series dataset for subsequent feature extraction. In essence, the feature extraction step captures multi-dimensional information about the equipment status in the time domain, frequency domain, and time-frequency domain, and the operating state feature vector set directly supports the model evaluation process as input to the subsequent large language model.
[0025] In one embodiment of the present invention, after the set of running state feature vectors is input into the model and converted into initial query vectors, key vectors, and value vectors, the improved attention mechanism begins to operate. This mechanism introduces a dynamic association weight matrix, which is not a fixed parameter but is dynamically generated based on the metadata of the current running stage of the device and the type metadata of the features themselves. This dynamic association weight matrix is used to modulate the interaction process between the initial query vector and the initial key vector during similarity calculation. When calculating attention weights, the similarity score obtained after modulation by the dynamic association weight matrix is input into a non-linear sharpening function for processing. This non-linear sharpening function strengthens the attention weights corresponding to highly relevant feature pairs while further suppressing the weights of low-relevance feature pairs, thereby producing a sharpened attention distribution with more significant discriminative power. Based on the sharpened attention distribution, the initial value vectors are weighted and fused. Before output, the mechanism also introduces residual connections and layer normalization operations to stabilize the training process and improve information flow efficiency, ultimately generating a context-aware deep state representation with stronger representational capabilities.
[0026] In practical implementation, the improved attention mechanism included in the large language model for reliability assessment works through a series of computational and information modulation steps, as described in [reference needed]. Figure 3In reliability assessment, when the large language model processes the operational state feature vector set, an improved attention mechanism receives the initial query vector, initial key vector, and initial value vector transformed from the operational state feature vector set. This improved attention mechanism introduces a dynamic association weight matrix, which is dynamically generated based on the device's operational stage metadata and feature type metadata. This dynamic association weight matrix is used to modulate the similarity calculation between the initial query vector and the initial key vector. When calculating the attention weights, the modulated similarity score is input into a nonlinear sharpening function. This function strengthens the weights of highly relevant feature pairs while further suppressing the weights of low-relevance feature pairs, generating a sharpened attention distribution. The initial value vector is then weighted and fused based on this sharpened attention distribution, and residual connections and layer normalization operations are introduced to ultimately generate a context-aware deep state representation with enhanced discriminative power.
[0027] In some embodiments, the generation of the dynamic association weight matrix depends on device operation phase metadata and feature type metadata. Device operation phase metadata includes the cumulative device operating time and the time interval since the last maintenance. Feature type metadata includes the category label indicating whether the feature belongs to the time domain, frequency domain, or time-frequency domain. The modulation of the initial query vector and initial key vector by the dynamic association weight matrix is reflected in element-level weighting operations. In some embodiments, the nonlinear sharpening function uses an exponential scaling function to process the modulated similarity score, expressed as: in: It is the sharpened attention distribution vector. It is a similarity score vector modulated by a dynamic association weight matrix. It is a sharpening factor greater than 1. It is an exponential function. This represents summing all elements of a vector.
[0028] Optionally, the residual connection operation adds the feature vector before the improved attention mechanism to the feature vector after weighted fusion, and the layer normalization operation standardizes the added feature vector. The residual connection and layer normalization operation work together to stabilize the training process of the model and promote gradient flow. Optionally, the device operation stage metadata and feature type metadata are converted into continuous vector representations through independent embedding layers. These continuous vector representations are then fused by a feedforward network to generate parameters for a dynamically correlated weight matrix. It can be understood that the improved attention mechanism introduces prior knowledge of device operation stage and feature type through the dynamically correlated weight matrix, making similarity calculation more task-specific. It can also be understood that the nonlinear sharpening function amplifies the weights of highly correlated features and suppresses the weights of low-correlation features, making the generated attention distribution more focused on key information, thereby improving the discriminative power of the deep state representation.
[0029] In one embodiment of the present invention, the reliability regression network receives deep state representation as input and feeds it into the network's multi-head analysis layer. In this multi-head analysis layer, an analysis head specifically processes the feature subspace associated with fault modes. This analysis head consists of multiple stacked fully connected layers, with its end connected to a classifier. The analysis head performs multi-layer nonlinear transformations on the input feature subspace to obtain high-level fault features, which are then input into the classifier's Softmax layer. The Softmax layer maps the high-level fault features to probability values corresponding to a series of predefined fault modes. These probability values are then sorted, and fault modes whose probability values exceed a preset activation threshold are selected. These fault modes, along with their corresponding probability values, are output as effective fault warning information, constituting the core content of the potential fault mode probability distribution.
[0030] Within the same multi-head analysis layer, another analysis head specifically processes the feature subspace related to lifetime degradation. This head, also composed of multiple fully connected layers, terminates at a regressor. This head extracts features and compresses the dimensionality of the input feature subspace, yielding a lifetime degradation index characterizing device health degradation. This lifetime degradation index is then input into the fully connected layer of the regressor, which linearly maps it to a preliminary estimate of the remaining useful life. To conform to the physical laws of device lifetime decay, this preliminary estimate is further input into a monotonically decreasing constraint function. This function ensures that the output lifetime value does not increase with device operating time and maps it within the device's maximum rated lifetime range, thus obtaining the final quantified remaining useful life value. Finally, the network encapsulates the parsed potential failure mode probability distribution and the quantified remaining useful life together to form a structured reliability assessment report.
[0031] In practical implementation, deep state representation is used to simultaneously quantify and analyze potential failure modes and remaining useful life through a reliability regression network. The deep state representation is input into the multi-head analysis layer of the reliability regression network. Within this multi-head analysis layer, one analysis head specifically processes the feature subspace related to failure modes. This analysis head consists of multiple fully connected layers, and its terminal is connected to a classifier, which outputs the probability distribution of potential failure modes. Another analysis head in the multi-head analysis layer specifically processes the feature subspace related to lifetime degradation. This analysis head also consists of multiple fully connected layers, and its terminal is connected to a regressor, which outputs the quantified remaining useful life. The probability distribution of potential failure modes and the quantified remaining useful life are jointly encapsulated to form a reliability assessment report.
[0032] In practical implementation, a dedicated analysis head for the feature subspace related to fault modes is connected to the output of the classifier. This analysis head performs multi-layer nonlinear transformations on the input feature subspace related to fault modes to obtain high-level fault features. These high-level fault features are then input into the classifier's Softmax layer. The Softmax layer maps these high-level fault features to probability values corresponding to different predefined fault modes. The output probability values are sorted, and fault modes with probability values exceeding a preset activation threshold are selected, along with their corresponding probability values, as effective fault warnings, constituting the core content of the potential fault mode probability distribution. In practical implementation, a dedicated analysis head for the feature subspace related to lifespan degradation is connected to the output of the regressor. This analysis head extracts features and compresses the dimensions of the input feature subspace related to lifespan degradation to obtain a lifespan degradation index. This lifespan degradation index is input into the fully connected layer of the regressor. The fully connected layer linearly maps the lifespan degradation index to a preliminary estimate of the remaining lifespan. This preliminary estimate is then input into a monotonically decreasing constraint function. The monotonically decreasing constraint function ensures that the output value does not increase with equipment operating time and is mapped to the maximum rated lifespan of the equipment, resulting in the final quantified remaining lifespan.
[0033] In some embodiments, the classifier's Softmax layer maps high-level fault features to probability values. A preset activation threshold is used to filter low-probability events to focus on high-risk potential faults. Fault patterns exceeding the preset activation threshold and their probabilities constitute fault warning information. In some embodiments, a monotonically decreasing constraint function processes the initial estimate output by the regressor to ensure that the quantized remaining useful life is monotonically non-increasing over time. The function has the following form: in: It is the final quantified remaining useful life. This is the preliminary estimate output by the regressor. This is the maximum rated lifespan of the equipment. It is an attenuation coefficient between 0 and 1. It is a bias term. and These are functions that take the minimum and maximum values, respectively.
[0034] Optionally, the preset activation threshold value can be set based on statistical analysis of historical fault data. Different activation thresholds can be set for different fault modes to suit their different occurrence frequencies and severity levels. Optionally, the output results of the fault mode probability distribution can be presented in tabular form, clearly showing the probability of each fault mode and whether an alert is triggered.
[0035] It is understandable that the design of the multi-head analysis layer enables the simultaneous analysis of information on two different attributes—fault and lifetime—from a single deep state representation. It is also understandable that the Softmax layer at the end of the classifier and the monotonically decreasing constraint function at the end of the regressor ensure the rationality and usability of the output results from the perspectives of probability normalization and physical law constraint, respectively (see Table 1).
[0036] Table 1: Probability Distribution of Potential Failure Modes In one embodiment of the present invention, pre-training first requires collecting a large amount of multi-source heterogeneous monitoring data generated by power distribution switchgear throughout its entire lifecycle, along with corresponding detailed fault record documents. These data and documents are then used to construct a large-scale pre-training corpus. The multi-source heterogeneous monitoring data in the corpus is processed using a comprehensive data governance and feature extraction process consistent with the aforementioned embodiments, constructing a massive number of operating state feature vector samples. Simultaneously, the fault record documents are structurally parsed to extract specific fault mode labels and the final lifespan labels of the equipment. These labels are then precisely aligned and labeled with the massive operating state feature vector samples based on their generation time. Using these massive operating state feature vector samples with alignment labels, a large-scale pre-training of a large language model based on a Transformer architecture is performed based on a joint training objective of masked language modeling and next-state prediction, resulting in a preliminary model with general equipment state understanding capabilities. Based on this preliminary model, a batch of finely labeled data from specific application scenarios is used to perform supervised fine-tuning of the model, making it more focused on the reliability assessment task of power distribution switchgear, ultimately obtaining a pre-trained large language model for reliability assessment.
[0037] In practical implementation, the pre-training process of the large language model for reliability assessment includes multiple steps. First, multi-source heterogeneous monitoring data and corresponding fault record documents of the power distribution switchgear throughout its entire lifecycle are collected to construct a pre-training corpus. Second, comprehensive data governance and feature extraction are performed on the multi-source heterogeneous monitoring data in the pre-training corpus to construct massive operational state feature vector samples. Third, fault record documents are structured and parsed to extract fault mode labels and equipment end-life labels. Fourth, fault mode labels, equipment end-life labels, and massive operational state feature vector samples are aligned and labeled at time points. Fifth, the massive operational state feature vector samples and their aligned labels are used to perform large-scale pre-training of the large language model based on the Transformer architecture, using a joint training objective of masked language modeling and next-state prediction, to obtain a preliminary model. Finally, based on the preliminary model, supervised fine-tuning is performed using finely labeled data from specific scenarios to finally obtain the pre-trained large language model for reliability assessment.
[0038] In practical implementation, the integrated data governance and feature extraction operations follow the same processing flow as the evaluation phase, ensuring that the data processing logic isomorphic between pre-training, fine-tuning, and application phases. Structured parsing of fault record documents involves natural language processing techniques to identify the fault types and occurrence times described in the documents. The alignment and labeling process associates operating state feature vector samples with fault mode labels and equipment end-life labels based on timestamps. For a single equipment sample, all operating state feature vector samples prior to the fault occurrence time are typically labeled "normal" or associated with specific early fault features. The equipment end-life label is a numerical value representing the total time from equipment commissioning to failure. The masking language modeling training objective randomly masks some feature values in the operating state feature vector samples, allowing the model to predict the masked original values. The next state prediction training objective requires the model to predict the operating state feature vector for the next time step based on the historical operating state feature vector sequence.
[0039] In some embodiments, the loss function for the joint training objective combines the masked language modeling loss and the next-state prediction loss, as expressed in the formula: in: It is the total loss. This is a loss for masked language modeling tasks. It is the loss of the next state prediction task. and These are hyperparameters used to balance the weights of the two tasks. In some embodiments, the multi-source heterogeneous monitoring data collected in the pre-training corpus covers various physical quantities such as mechanical, electrical, and thermal data, and their specific composition can be summarized in tabular form.
[0040] Optionally, the finely labeled data for specific scenarios comes from a small number of power distribution switchgear devices equipped with high-precision monitoring and clear fault records in the target application scenario. This type of data is much smaller than the pre-training corpus but has higher labeling accuracy. Optionally, the supervised fine-tuning process uses the cross-entropy loss function for the fault mode classification task and the mean squared error loss function for the remaining useful life regression task to further adjust the parameters of the initial model.
[0041] It is understandable that pre-training based on large-scale historical data and joint training objectives enables the model to learn general patterns and correlations between features in the evolution of device states. It is also understandable that fine-tuning using meticulously labeled data for specific scenarios allows the model to adapt to the specific failure modes and lifespan characteristics of the target application scenario, improving the accuracy of the assessment (see Table 2).
[0042] Table 2: Data Types for Multi-Source Heterogeneous Monitoring of Pre-trained Corpora In one embodiment of the present invention, the system stores the reliability assessment report currently generated for the target power distribution switchgear, along with all historical reliability assessment reports generated for that equipment, in the same assessment archive, arranged chronologically to form an assessment history sequence. The system periodically performs trend analysis on the assessment history sequence in the assessment archive, calculating the rate of change of the probability value of key failure modes in the potential failure mode probability distribution over time, and simultaneously calculating the rate of decay of the remaining service life. When the system detects that the rate of change of the probability of a certain key failure mode, or the rate of decay of the remaining service life of the equipment, exceeds its respective set dynamic threshold, it automatically triggers an early warning signal and generates an early warning summary document containing these key change indicators.
[0043] The dynamic threshold is determined as follows: The system selects the historical evaluation subsequence corresponding to the most recent stable operation phase of the equipment from the evaluation archive. The mean and standard deviation of the probability change rate of critical failure modes, and the mean and standard deviation of the quantified remaining useful life decay rate, are calculated within this subsequence. The mean of the probability change rate is added to a multiple of its standard deviation, and the result is used as the dynamic threshold for the probability change rate of that failure mode. Similarly, the mean of the decay rate is added to a multiple of its standard deviation, and the result is used as the dynamic threshold for the remaining useful life decay rate.
[0044] In practice, after outputting a reliability assessment report containing the probability distribution of potential failure modes and the quantified remaining useful life, subsequent early warning and analysis processes are executed. The currently generated reliability assessment report and the historical reliability assessment reports of the target power distribution switchgear are stored in the same assessment archive, forming an assessment history sequence. Trend analysis is performed on the assessment history sequence in the assessment archive to calculate the probability change rate of key modes in the probability distribution of potential failure modes and the decay rate of the quantified remaining useful life. When the probability change rate or decay rate exceeds its corresponding dynamic threshold, an early warning signal is triggered, and an early warning summary document containing key change indicators is automatically generated.
[0045] In practice, the determination of dynamic thresholds follows specific steps: First, a historical evaluation subsequence representing the most recent stable operating phase is selected from the evaluation archive. Then, the mean and standard deviation of the probability change rate of key patterns within the historical evaluation subsequence, as well as the mean and standard deviation of the decay rate of quantified remaining lifetime, are calculated. The mean of the probability change rate plus a certain multiple of its standard deviation is used as the dynamic threshold for the probability change rate, and the mean of the decay rate plus a certain multiple of its standard deviation is used as the dynamic threshold for the decay rate.
[0046] In some embodiments, the probability change rate is obtained by calculating the difference between the probability values of the same critical failure mode in two adjacent assessment reports or the average change slope within a sliding window, and the decay rate is obtained by calculating the ratio of the reduction in quantified remaining useful life to the time interval in two adjacent assessment reports. In some embodiments, the "multiple times" used in the dynamic threshold calculation is a configurable parameter coefficient, expressed by the formula: in: This represents the calculated dynamic threshold. This represents the mean of the probability change rate or decay rate calculated over the selected evaluation history subsequence. This represents the corresponding standard deviation. It is a configurable threshold coefficient.
[0047] Optionally, the evaluation history subsequence for the stable operation phase is determined by identifying consecutive time periods within the evaluation history sequence where the quantified remaining useful life declines gradually and there are no fault warnings. Optionally, the warning summary document includes the name of the fault mode that triggered the warning, its probability change rate, the current quantified remaining useful life, its decline rate, and the specific value exceeding the threshold. It can be understood that by maintaining the evaluation history sequence and performing trend analysis, continuous tracking of the equipment reliability degradation process can be achieved. It can be understood that a dynamic threshold calculated based on the equipment's own historical stable operation data provides an adaptive and individualized warning triggering mechanism, which better reflects the actual operating characteristics of a specific piece of equipment compared to a fixed threshold.
[0048] The above embodiments are only used to illustrate the technical methods of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical methods of the present invention without departing from the spirit and scope of the technical methods of the present invention.
Claims
1. A reliability assessment method for power distribution switchgear based on a large model, characterized in that, The method includes: Collect multi-dimensional monitoring data of the target power distribution switchgear during its historical operating cycle to form a raw monitoring data set; The original monitoring data set is subjected to comprehensive data governance, including missing value imputation, unit normalization, and time series alignment, to generate a standardized monitoring time series dataset; Multiple operational status features are extracted from the standardized monitoring time-series dataset to construct an operational status feature vector set; The set of running state feature vectors is input into a pre-trained reliability assessment large language model. The improved attention mechanism contained in the reliability assessment large language model is used to perform deep feature interaction and importance weighting to generate a context-aware deep state representation. Using the deep state representation, a reliability regression network is used to perform synchronous quantitative analysis of potential failure modes and remaining useful life, and outputs a reliability assessment report containing the probability distribution of potential failure modes and the quantified remaining useful life. The working principle of the improved attention mechanism is as follows: When the reliability assessment large language model processes the runtime feature vector set, the improved attention mechanism receives the initial query vector, initial key vector, and initial value vector transformed from the runtime feature vector set; The improved attention mechanism introduces a dynamic association weight matrix, which is dynamically generated based on the metadata of the device operation stage and the metadata of the feature type, and is used to modulate the similarity calculation between the initial query vector and the initial key vector. When calculating attention weights, the modulated similarity score is input into a nonlinear sharpening function, which strengthens the weights of highly relevant feature pairs while further suppressing the weights of low-relevance feature pairs, resulting in a sharpened attention distribution. The initial value vector is weighted and fused based on the sharpened attention distribution, and residual connections and layer normalization operations are introduced to finally generate the context-aware deep state representation with enhanced discriminativeness. Using the aforementioned deep state representation, a reliability regression network is used to perform synchronous quantitative analysis of potential failure modes and remaining useful life, including: The deep state representation is input into the multi-head analysis layer of the reliability regression network; In the multi-head analysis layer, one analysis head is dedicated to processing the feature subspace associated with the failure mode. The analysis head consists of multiple fully connected layers, the ends of which are connected to a classifier to output the probability distribution of the potential failure modes. In the multi-head analysis layer, another analysis head is dedicated to processing the feature subspace related to lifetime degradation. This analysis head consists of multiple fully connected layers, with its ends connected to a regressor, which outputs the quantified remaining lifetime. The potential failure mode probability distribution and the quantified remaining useful life are encapsulated together to form the reliability assessment report.
2. The reliability assessment method for power distribution switchgear based on a large model according to claim 1, characterized in that, Comprehensive data governance, including missing value imputation, dimension normalization, and time series alignment, is performed on the original monitoring data set, including: The numerical gaps in different monitoring signal channels within the original monitoring data set are identified, and an interpolation algorithm based on the trend inference of adjacent time points is used to fill in the numerical gaps to generate a continuous monitoring sequence. For the different physical dimensions of the data in each channel of the continuous monitoring sequence, the monitoring data of each channel is subtracted from its own historical mean and divided by its own historical standard deviation, and all monitoring data are mapped to the dimensionless standard normal distribution space. All channels' data within the standard normal distribution space are resampled according to a unified time reference to ensure that each channel's data point has an aligned value at each sampling time, ultimately generating the standardized monitoring time series dataset.
3. The reliability assessment method for power distribution switchgear based on a large model according to claim 1, characterized in that, Multiple operational status features are extracted from the standardized monitoring time-series dataset to construct an operational status feature vector set, including: In the time domain dimension, the statistical characteristics of each signal channel data in the standardized monitoring time series dataset are calculated, and the statistical characteristics include mean, variance, peak value and waveform factor. In the frequency domain, a fast Fourier transform is performed on the standardized monitoring time series dataset to calculate the dominant frequency, centroid frequency, and spectral entropy of each signal channel data. Construct a joint time-frequency domain feature extraction window, and calculate the variation trends of the short-time energy and zero-crossing rate of the signal within the joint time-frequency domain feature extraction window; The time-domain statistical features, frequency-domain features, and time-frequency domain variation trend features calculated for each signal channel are concatenated into a high-dimensional feature vector in a preset order. The high-dimensional feature vectors of all signal channels together constitute the operating state feature vector set.
4. The reliability assessment method for power distribution switchgear based on a large model according to claim 1, characterized in that, The process of connecting the output of the classifier to the analysis head end, which specifically processes the feature subspace related to fault modes, includes: The analysis head performs multi-level nonlinear transformations on the input feature subspace related to the fault mode to obtain high-level fault features; The advanced fault features are input into the Softmax layer of the classifier, which maps the advanced fault features to probability values corresponding to different predefined fault modes. The output probability values are sorted, and fault modes whose probability values exceed the preset activation threshold are selected. These, along with their corresponding probability values, are used as effective fault warnings, forming the core content of the potential fault mode probability distribution.
5. The reliability assessment method for power distribution switchgear based on a large model according to claim 1, characterized in that, The process of connecting the output of the analysis head end regressor, which specifically processes the feature subspace related to lifetime degradation, includes: The analysis head extracts features and compresses dimensions in the input feature subspace related to lifespan degradation to obtain the lifespan degradation index. The lifetime degradation index is input into the fully connected layer of the regressor, which linearly maps the lifetime degradation index to a preliminary estimate of the remaining useful life. The preliminary estimate is input into a monotonically decreasing constraint function, which ensures that the output value does not increase with the device's operating time and is mapped to the device's maximum rated lifespan to obtain the final quantified remaining lifespan.
6. The reliability assessment method for power distribution switchgear based on a large model according to claim 1, characterized in that, The pre-training process of the large language model for reliability assessment includes: Collect multi-source heterogeneous monitoring data and corresponding fault record documents of power distribution switchgear throughout its entire life cycle, and build a pre-training corpus. Comprehensive data governance and feature extraction are performed on the multi-source heterogeneous monitoring data in the pre-training corpus to construct a massive number of operational status feature vector samples; The fault record document is structured and parsed to extract fault mode labels and equipment final lifespan labels, and then aligned and labeled with the massive operating status feature vector samples at time points. Using the massive number of running state feature vector samples and their aligned labels, based on the joint training objective of masked language modeling and next state prediction, a large-scale pre-training of the Transformer architecture's large language model is performed to obtain a preliminary model. Based on the preliminary model, the preliminary model is fine-tuned in a supervised manner using finely labeled data for specific scenarios, and finally the pre-trained reliability evaluation large language model is obtained.
7. The reliability assessment method for power distribution switchgear based on a large model according to claim 1, characterized in that, After the output includes a reliability assessment report with potential failure mode probability distributions and quantified remaining useful life, the method further includes: The currently generated reliability assessment report and the historical reliability assessment report of the target power distribution switchgear are stored in the same assessment file to form an assessment history sequence; Perform trend analysis on the historical evaluation sequence in the evaluation file, and calculate the probability change rate of key modes in the probability distribution of potential failure modes and the decay rate of the quantified remaining useful life. When the probability change rate or the decay rate exceeds their respective dynamic thresholds, an early warning signal is triggered, and an early warning summary document containing key change indicators is automatically generated. The method for determining the dynamic threshold includes: Select the evaluation history subsequence of the most recent stable operation phase from the evaluation archive; Calculate the mean and standard deviation of the probability change rate of key patterns in the evaluated historical subsequence, and the mean and standard deviation of the decay rate of the quantified remaining lifetime; The mean of the probability change rate is added to a certain number of times its standard deviation, and the result is used as the dynamic threshold of the probability change rate. The mean of the attenuation rate is added to a certain number of times its standard deviation, and the result is used as the dynamic threshold of the attenuation rate.
8. A reliability assessment system for power distribution switchgear based on a large model, comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the reliability assessment method for power distribution switchgear based on any one of claims 1 to 7.
Citation Information
Patent Citations
Optical cable intelligent label full life cycle management method and system
CN121052276A
Power equipment fault intelligent diagnosis method and system based on deep learning
CN121278535A