Intelligent Monitoring Method for Full Insulation State of Mine Explosion-Proof High-Voltage Distribution Devices
Through the deep probability neural network combined with adaptive normalization and dual-path timing feature extraction methods, the problems of hysteresis and high false alarm rate of insulation status monitoring of explosion-proof high-voltage distribution devices for mining are solved, and accurate diagnosis and early warning are achieved under complex operating conditions.
Patent Information
- Application Number
- CN202510668104.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-23
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2045-05-23
AI Technical Summary
The insulation state monitoring of mining explosion-proof high-voltage distribution devices has problems of hysteresis and high false alarm rates. It is difficult for existing methods to achieve accurate identification and early warning in complex working conditions, especially in early deterioration stages or intermittent abnormalities. Traditional methods cannot adapt to multi-factor combination deterioration and timing non-stationarity.
Using a fully insulated state intelligent monitoring method based on deep probability neural network, data is collected through multi-source sensors, combined with adaptive normalization, dual-path timing feature extraction and probability hidden layer modeling, the capture and uncertainty quantization of long and short-term timing dependence are achieved, local timing fluctuations are dynamically perceived, and the self-attention mechanism and gating mechanism are fused to generate three types of state probability and uncertainty quantization values.
It significantly improves the accuracy and anti-interference ability of insulation state diagnosis, and can achieve early warning and false alarm rate reduction of equipment deterioration under complex working conditions, adapt to the dynamic changes of sensing signals such as voltage and current, and reduce the risk of false alarm and missed response in severe deterioration.
Smart Images

Figure CN120180283B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical fields of artificial intelligence and data processing, and particularly relates to an intelligent monitoring method for the full insulation state of a mine explosion-proof high-voltage distribution device. Background Art
[0002] During the long-term operation of a mine explosion-proof high-voltage distribution device, its insulation state is easily affected by various complex factors, such as load fluctuations, changes in ambient temperature and humidity, arc discharges, equipment aging, etc., resulting in the gradual accumulation of potential local insulation deterioration or breakdown, and it is extremely easy to cause serious electrical faults and safety accidents. However, at present, the insulation monitoring of mine high-voltage equipment still mainly relies on periodic manual inspections or fixed-threshold alarms, and it is difficult to identify hidden deterioration states in a timely and accurate manner. Especially in the early deterioration stage or during intermittent abnormalities, the existing methods show obvious lag and high false alarm rates. At the same time, the monitoring signals in the mine power distribution system have significant characteristics of temporal non-stationarity and multi-scale feature coupling. There are not only situations such as amplitude mutations, frequency drifts, and local disturbances, but also strong correlations and noise pollution between signals. Traditional state recognition methods based on deterministic deep learning methods or expert rules are difficult to model such complex multi-source, multi-temporal domain, and uncertain signal features, resulting in weak generalization ability of the diagnostic model, poor adaptability to complex working conditions, and inability to achieve stable and accurate insulation state recognition in the harsh underground environment.
[0003] Problems objectively existing in the prior art: Conventional deep learning classification methods such as CNN, RNN, and LSTM only output deterministic predictions, cannot express prediction uncertainty, and have poor adaptability to mutations and redundant features in time series; traditional signal processing or rule threshold judgment methods cannot learn complex feature representations, cannot adapt to working condition drift and multi-factor combined deterioration problems, and cannot be updated online with poor generalization ability; models that only use attention or gating mechanisms cannot perform probability representation and posterior inference, cannot generate confidence intervals or uncertainty estimates, and are difficult to be used for reliability judgment of boundary samples; conventional optimization methods such as SGD and Adam adopt fixed or round-by-round adjusted learning rates, do not consider the dynamic adaptability between the model state and the working condition complexity, and are difficult to adapt to high-variation industrial environments; the global normalization of conventional feature normalization methods cannot adapt to the temporal non-stationary characteristics and may mask key micro-variation signals.
[0004] Therefore, the present invention proposes an intelligent monitoring method for the full insulation state of a mine explosion-proof high-voltage distribution device to solve the above problems. Summary of the Invention
[0005] In view of the deficiencies of the prior art, the present invention develops an intelligent monitoring method for the full insulation state of mine explosion-proof high-voltage distribution devices. The present invention effectively solves the problems such as dynamic signal adaptation, long-term and short-term feature capture, and data noise interference in the monitoring of mine high-voltage equipment, and significantly improves the accuracy and anti-interference ability of insulation state diagnosis.
[0006] The technical solution for the present invention to solve the technical problem is an intelligent monitoring method for the full insulation state of mine explosion-proof high-voltage distribution devices, including the following steps:
[0007] S1. Collect training data: Collect the full-condition operation data of the mine explosion-proof high-voltage distribution device through a multi-source sensor system and construct a data set. Segment and store the collected data by the sliding window method to obtain a structured data set;
[0008] S2. Label the training data: Label the data in the structured data set by a labeling method combining expert experience and off-line detection to form a labeled data set containing the mapping relationship between multi-dimensional features and state labels;
[0009] S3. Preprocess the training data: Regularize the data in the labeled data set by using the industrial time series data standardization process. The regularization operations include missing value processing, deleting abnormal samples, and sample category balancing operations to obtain a preprocessed data set;
[0010] S4. Model construction and training: Build an intelligent monitoring model for the full insulation state of mine explosion-proof high-voltage distribution devices based on a deep probabilistic neural network. The model includes an adaptive normalization layer, a dual-channel time series feature module, a probabilistic hidden layer, and an output layer. Among them, the adaptive normalization layer adopts the sliding window dynamic normalization method. The inputs of the two branches of the dual-channel time series feature module are the long-term and short-term time series dependencies captured by the self-attention mechanism and the gated unit respectively. The probabilistic hidden layer is a multi-layer fully connected neuron hidden layer; Input the preprocessed data set into the model for training, and repeat the iterative training process until the preset stop iteration condition is met to obtain a trained model;
[0011] S5. Intelligent monitoring classification of the full insulation state: Input the preprocessed real-time data into the trained intelligent monitoring model for the full insulation state of mine explosion-proof high-voltage distribution devices. The output layer generates the probability distribution of three categories: normal operation, slight deterioration, and severe deterioration, and the corresponding uncertainty quantification values, and obtain the final diagnosis result according to the preset discrimination conditions;
[0012] S6. Model incremental update: Establish a dynamic incremental learning mechanism, deploy an online monitoring module, set a trigger value. When the number of continuously detected unseen data samples is greater than the trigger value, data collection is automatically triggered. After manual review, the data is stored in the incremental dataset, and the incremental dataset is regularly used for incremental training of the intelligent monitoring model for the fully insulated state of mine explosion-proof high-voltage distribution devices.
[0013] S1 is as follows:
[0014] Deploy a multi-source sensor system at the key nodes of the mine explosion-proof high-voltage distribution device for data collection. The multi-source sensor system includes a voltage transformer, a current transformer, a temperature sensor, and a partial discharge detector. The key monitoring nodes include the main circuit insulating sleeve, the breaker contact, and the cable joint.
[0015] The multi-source sensor system synchronously collects time-series signals at a set sampling rate. The time-series signals include voltage fluctuations, leakage current, temperature rise gradient, and partial discharge pulse sequences.
[0016] Simulate different working condition scenarios, continuously collect data under different working condition scenarios. The collected data covers three operating states: normal operation, slight deterioration, and severe deterioration. The working condition scenarios include load switching, overvoltage impact, environmental change, temperature change, and humidity change.
[0017] Adopt the sliding window method to segment and store the collected time-series signal data, set the window length and the overlapping rate of adjacent windows, and obtain a structured dataset with time-series alignment and multi-dimensional physical quantities.
[0018] S2 is as follows:
[0019] First, according to the IEC insulation state evaluation standard, electrical engineers combine the insulation resistance test records, partial discharge intensity thresholds, and historical fault cases in the equipment maintenance log to label the operating state of each data in the structured dataset. The operating states include normal operation, slight deterioration, and severe deterioration.
[0020] Then, for the boundary samples with disputes, an offline dielectric loss tangent value tester is used for supplementary measurement, and the accuracy of the label is verified through frequency-domain dielectric characteristic analysis.
[0021] Finally, a labeled dataset containing the mapping relationship between multi-dimensional features and state labels is formed.
[0022] S3 is as follows:
[0023] Missing value processing: For the zero-value or constant-value segments caused by the instantaneous failure of the sensor, the adjacent window linear interpolation method is used to complete them; for the abnormal values exceeding the physical range, threshold truncation is performed according to the rated parameters of the mine explosion-proof high-voltage distribution device.
[0024] Delete abnormal samples: Extract corresponding metrics from the data in the labeled dataset. The metric types are time-domain features, frequency-domain features, and time-frequency domain features. Based on the extracted metrics, perform clustering analysis on the data in the labeled dataset, set a screening strategy, and eliminate outliers according to the screening strategy;
[0025] Sample class balance operation: Balance the class distribution through the SMOTE oversampling technique, amplify the lacking data, so that the ratio of the three types of data is: normal operation: slight deterioration: severe deterioration = 3:2:1.
[0026] The training process of the intelligent monitoring model for the fully insulated state of the mine explosion-proof high-voltage distribution device is as follows:
[0027] Input the preprocessed dataset into the model for training. After the input data is dynamically normalized by a sliding window, capture long-term and short-term time series dependencies through the self-attention mechanism and the gating mechanism, and then input it into the probabilistic hidden layer. Model the mean and covariance of the latent variables in the form of a Gaussian distribution, and perform uncertainty quantification through the reparameterization technique. During this period, uncertainty is transmitted between probabilistic hidden layers through the probability propagation rule, combined with the dynamic estimation of process noise, to form a hierarchical uncertainty accumulation mechanism. Finally, fuse the probabilistic features and the gating output, and generate classification probabilities and uncertainty quantification values through a mixture density network.
[0028] The operation of normalizing the preprocessed dataset in the adaptive normalization layer is specifically as follows:
[0029] Adopt an adaptive normalization method based on a sliding window, and dynamically calculate the statistics within the window in combination with time decay weights; First, calculate the weighted mean and weighted standard deviation within the sliding window, where the weights are determined by the time decay weights, and the data closer to the current time has a greater weight. Then, normalize the data at the current time using the weighted mean and standard deviation, and adjust the normalized result through learnable scaling and offset parameters. Finally, obtain the normalized data.
[0030] The operation of extracting time series correlation features from the normalized data in the dual-channel time series feature extraction module is specifically as follows:
[0031] First, linearly transform the normalized data into query vectors, key vectors, and value vectors respectively. Then, use the self-attention mechanism to calculate the similarity between the query vector and the key vector, and reconstruct the features through the value vector to obtain the self-attention weighted feature representation. Then, concatenate the normalized data with the self-attention features, and screen and adjust the features through the gating mechanism to finally obtain the time series correlation features;
[0032] The calculation formula in the dual-channel time series feature extraction module is as follows:
[0033] ,
[0034] ,
[0035] wherein, is the query vector, is the key vector, is the value vector; is the normalized data; is the dimensional scaling factor, used to scale the dot product result to stabilize the gradient; is the vector concatenation; is the feature after self-attention weighting; is the gated feature; is the gated weight matrix; is the hyperbolic tangent function; is the projection weight matrix.
[0036] The operations in the probabilistic hidden layer are as follows:
[0037] (1) Input the temporal correlation feature into the probabilistic hidden layer. The temporal correlation feature is the output of the gated feature , and use probabilistic distribution modeling. Through the reparameterization trick, perform gradient backpropagation so that the model can simultaneously learn feature representation and uncertainty quantification;
[0038] First, input the temporal correlation feature into the probabilistic hidden layer. Perform a linear transformation on the feature through the weight matrix and bias term to obtain the mean matrix of the Gaussian distribution. At the same time, perform a non-linear transformation on the feature using the covariance generation matrix to obtain the diagonal covariance matrix of the Gaussian distribution, and represent the feature as a probabilistic form of the Gaussian distribution;
[0039] (2) Generate a prior distribution based on the statistical characteristics of the normalized data to enable the model to quickly converge to a reasonable probability space;
[0040] First, calculate the mean and covariance matrix of all sample gated features. Then, perform Cholesky decomposition on the diagonal covariance matrix of the Gaussian distribution to obtain the initial value of the covariance generation matrix, and use the statistical characteristics of the data to provide a reasonable initial value for the covariance generation matrix;
[0041] (3) Define the probability propagation chain rule between hidden layers to accumulate and transfer the uncertainty between hidden layers;
[0042] Define the probability propagation chain rule between hidden layers. For the hidden variables of each layer, calculate their posterior distribution through integration, where the integration range is all possible values of the hidden variables of the previous layer. During the integration process, update the mean and covariance of the hidden variables using the properties of the Gaussian distribution.
[0043] The operation of class prediction in the output layer is as follows:
[0044] Concatenate the output of the probability hidden layer with the gated features, and generate class probabilities through a mixture density network;
[0045] First, concatenate the last hidden variable of the deep probability neural network with the gated features. Then, perform a non-linear transformation on the gated features through a feature mapping function, and perform a weighted sum on the transformed features to obtain a linear combination of class probabilities. Then, use the Softmax function to normalize the linear combination to obtain the predicted probability for each class.
[0046] Optimize the intelligent monitoring model for the fully insulated state of the mine explosion-proof high-voltage distribution device:
[0047] Optimize the intelligent monitoring model for the fully insulated state of the mine explosion-proof high-voltage distribution device by calculating the loss function, updating parameters, and adjusting the learning rate;
[0048] (1) Calculate the loss function:
[0049] Define the total loss function of the deep probability neural network. The total loss function includes classification loss, calibration loss, and covariance regularization term. Measure the difference between the model's predicted class probabilities and the true labels through the classification loss, measure the difference between the model's predicted uncertainty and the true uncertainty through the calibration loss, and constrain the model's uncertainty estimation through the covariance regularization term. Then, combine the three parts of the loss in a weighted manner to obtain the total loss function. The calculation formula is as follows:
[0050] ,
[0051] ,
[0052] ,
[0053] Among them, is the total loss function of the deep probability neural network; is the classification loss weight coefficient, such as; is the cross-entropy loss; is the true class label; is the calibration loss; represents the expectation of distribution; represents the predicted variance; represents the matrix trace operation; is the covariance regularization coefficient; represents the covariance regularization term, is the diagonal covariance matrix of the Gaussian distribution, Denotes the number of layers of the deep probabilistic neural network;
[0054] (2) Parameter update:
[0055] Define the set of trainable parameters of the deep probabilistic neural network as , and use the probabilistic gradient update rule to update the training parameters of the deep probabilistic neural network;
[0056] During the update process, first calculate the gradient of the total loss function with respect to the parameters, and use the variational posterior distribution to weight-average the gradient. Then, consider the gradient of the trace of the covariance matrix for parameter update, and adjust the gradient through the momentum decay rate. The calculation formula is as follows:
[0057] ,
[0058] where, is the parameter of the deep probabilistic neural network at the -th iteration; is the parameter of the deep probabilistic neural network at the -th iteration; is the learning rate at the e-th iteration; Denotes the expectation of the variational posterior distribution; is the variational posterior distribution; Denotes the gradient of the total loss function of the deep probabilistic neural network with respect to the parameter θ; Denotes the gradient operation on the parameter θ; is the momentum decay rate; is the sign function; Denotes the sum of the traces of all hidden layer covariance matrices;
[0059] (3) Learning rate adjustment:
[0060] Adopt an adaptive strategy based on uncertainty entropy to adjust the learning rate;
[0061] First calculate the uncertainty entropy of the model at each iteration step, and then dynamically adjust the learning rate according to the average value of the uncertainty entropy. When the overall uncertainty of the model is high, reduce the learning rate, and when the uncertainty of the model is low, increase the learning rate:
[0062] The calculation formula of the learning rate is as follows:
[0063] ,
[0064] where, is the learning rate at the e-th iteration; is the base learning rate; is the learning rate decay coefficient; is the maximum number of iterations and is a positive integer; is the input data for the -th iteration; represents the predicted distribution of the input data model for the -th iteration represents the predicted distribution of the Shannon entropy.
[0065] The effects provided in the invention content are only the effects of the embodiments, rather than all the effects of the invention. The above technical solutions have the following advantages or beneficial effects:
[0066] The present invention discloses a full-insulation state intelligent monitoring method for mine explosion-proof high-voltage distribution devices. Through the sliding window weighted normalization method based on the time decay factor, it can dynamically perceive local time series fluctuations, break through the limitation that traditional Min-Max normalization cannot adapt to local non-stationary problems, and adapt to the dynamic change characteristics of sensing signals such as voltage, current, and partial discharge over time; through the dual-channel structure that combines the self-attention mechanism and the gating mechanism, it can take into account the long-term dependence modeling ability and local anomaly suppression ability, and capture early latent anomalies and short-term mutation characteristics;
[0067] By adopting learnable Gaussian distribution latent variables in the deep probability neural network, constructing a probabilistic hidden layer, and using the reparameterization technique to achieve end-to-end trainable uncertainty modeling, it can cope with problems such as fuzzy labels, sample disputes, and data perturbations in the collected data; through the "hidden layer conditional Gaussian propagation + process noise modeling" mechanism, it accumulates uncertainty layer by layer, breaks the fixed paradigm that traditional neural network hidden layers only propagate activation values, and more realistically simulates the industrial noise transmission effect in the multi-layer structure;
[0068] By proposing to use the hybrid density network + gating splicing method to output the three-category state probabilities, combined with the entropy threshold decision and the sliding window trend re-judgment, it realizes stable and reliable diagnosis, reduces the false alarm and missed alarm risks in the "severely deteriorated" state; adopts the dynamic adjustment learning rate method driven by the prediction entropy, which can automatically coordinate the "robustness in the initial training stage" and the "efficiency in the convergence stage", and cope with training instability phenomena such as heavy noise and data drift; the covariance matrix initialization adopts Cholesky decomposition and prior distribution, and joint regularization to avoid non-positive definite and overfitting problems in the covariance training process, and improve the stability and convergence speed of latent variable modeling.
[0069] In summary, the present invention effectively solves the problems of dynamic signal adaptation, long-term and short-term feature capture, and data noise interference in the monitoring of mine high-voltage equipment, significantly improves the accuracy and anti-interference ability of insulation state diagnosis, and can achieve a double breakthrough in early warning of equipment deterioration and reduction of false alarm rate under complex working conditions. BRIEF DESCRIPTION OF THE DRAWINGS
[0070] The accompanying drawings are used to provide a further understanding of the present invention and form a part of the specification. Together with the embodiments of the present invention, they are used to explain the present invention and do not constitute a limitation to the present invention.
[0071] Figure 1 It is a schematic diagram of the method flow of the present invention.
[0072] Figure 2 It is a schematic diagram of the original signal and the true trend collected by the current sensor.
[0073] Figure 3 It is a comparison chart of the effects of different normalization methods.
[0074] Figure 4 It is a comparison chart of different normalization methods regarding classification accuracy and uncertainty entropy.
[0075] Figure 5 It is a schematic diagram of the time-frequency spectrum of the original signal.
[0076] Figure 6 It is a schematic diagram of the time-frequency spectrum for feature extraction using the traditional LSTM method.
[0077] Figure 7 It is a schematic diagram of the time-frequency spectrum for feature extraction using the dual-channel method.
[0078] Figure 8 It is a comparison chart of the hidden layer uncertainty propagation paths. Specific embodiments
[0079] In order to clearly illustrate the technical features of the present solution, the present invention will be elaborated in detail below through specific embodiments and in conjunction with its accompanying drawings. The following disclosure provides many different embodiments or examples for implementing different structures of the present invention. To simplify the disclosure of the present invention, the components and settings of specific examples are described below.
[0080] Embodiment 1
[0081] An all-insulation state intelligent monitoring method for a mine explosion-proof high-voltage distribution device, comprising the following steps:
[0082] S1. Collect training data: Collect the full-condition operation data of the mine explosion-proof high-voltage distribution device through a multi-source sensor system and construct a data set. Segment and store the collected data by the sliding window method to obtain a structured data set;
[0083] S2. Label the training data: Label the data in the structured data set by using a labeling method combining expert experience and offline detection to form a labeled data set containing the mapping relationship between multi-dimensional features and state labels;
[0084] S3. Training data preprocessing: Regularize the data in the labeled dataset using the industrial time-series data standardization process. The regularization operations include missing value processing, deleting abnormal samples, and sample class balancing operations to obtain the preprocessed dataset.
[0085] S4. Model construction and training: Build an intelligent monitoring model for the full insulation state of mine explosion-proof high-voltage distribution devices based on a deep probabilistic neural network. The model includes an adaptive normalization layer, a dual-channel time-series feature module, a probabilistic hidden layer, and an output layer. Among them, the adaptive normalization layer uses a sliding window dynamic normalization method. The inputs of the two branches of the dual-channel time-series feature module are the long-term and short-term time-series dependencies captured by the self-attention mechanism and the gated unit respectively. The probabilistic hidden layer is a multi-layer fully connected neuron hidden layer. Input the preprocessed dataset into the model for training, and repeat the iterative training process until the preset stop iteration condition is met to obtain the trained model.
[0086] S5. Intelligent monitoring classification of full insulation state: Input the preprocessed real-time data into the trained intelligent monitoring model for the full insulation state of mine explosion-proof high-voltage distribution devices. The output layer generates the probability distributions of three categories: normal operation, slight deterioration, and severe deterioration, as well as the corresponding uncertainty quantification values, and obtain the final diagnosis result according to the preset discrimination conditions.
[0087] S6. Model incremental update: Establish a dynamic incremental learning mechanism, deploy an online monitoring module, set a trigger value. When the number of unseen data samples continuously detected is greater than the trigger value, data collection is automatically triggered. After manual review, it is stored in the incremental dataset, and the incremental dataset is regularly used for incremental training of the intelligent monitoring model for the full insulation state of mine explosion-proof high-voltage distribution devices.
[0088] In the specific implementation manner, S1 is specifically as follows:
[0089] Deploy a multi-source sensor system at the key nodes of the mine explosion-proof high-voltage distribution device for data collection. The multi-source sensor system includes a voltage transformer, a current transformer, a temperature sensor, and a partial discharge detector. The key monitoring nodes include the main circuit insulating sleeve, the breaker contact, and the cable joint.
[0090] The multi-source sensor system synchronously collects time-series signals at a set sampling rate. The time-series signals include voltage fluctuations, leakage current, temperature rise gradient, and partial discharge pulse sequences. The sampling rate is set to 1 kHz.
[0091] Simulate different working condition scenarios and continuously collect data under different working condition scenarios. The collected data covers three operating states: normal operation, slight deterioration, and severe deterioration. The working condition scenarios include load switching, overvoltage impact, environmental change, temperature change, and humidity change. The total duration of continuously collected data is not less than 400 hours.
[0092] The collected time-series signal data is segmented and stored using the sliding window method. The window length is set to 5 seconds and the overlapping rate of adjacent windows is 30%, resulting in a structured data set with time-series alignment and multi-dimensional physical quantities.
[0093] It should be noted that the above only introduces one data acquisition and storage method;
[0094] It should also be noted that the above introduces the acquisition of time-series signals such as voltage fluctuations, leakage current, temperature rise gradient, and partial discharge pulse sequences. When training a machine learning model, any one of these data formats and sources can be used, or multiple data formats and sources can be selected. If multi-source data formats and sources are selected, data connection can be considered in the form of signal data concatenation, or feature fusion can be considered using methods such as weighted average and neural network mapping.
[0095] In the specific implementation manner, S2 is as follows:
[0096] First, electrical engineers, according to the IEC insulation status assessment standard, combined with the insulation resistance test records, partial discharge intensity thresholds, and historical fault cases in the equipment maintenance log, label the operating status of each data in the structured data set. The operating status includes normal operation, slight deterioration, and severe deterioration;
[0097] Then, for the boundary samples with disputes, such as intermittent discharge signals, an off-line dielectric loss tangent value tester is used for supplementary measurement, and the accuracy of the label is verified through frequency-domain dielectric characteristic analysis;
[0098] Finally, a labeled data set containing the mapping relationship between multi-dimensional features and status labels is formed.
[0099] It should be noted that the above only introduces one data labeling method, and those skilled in the art can perform data labeling of different categories according to actual needs.
[0100] In the specific implementation manner, S3 is as follows:
[0101] Missing value processing: For the zero-value or constant-value segments caused by the instantaneous failure of the sensor, linear interpolation of adjacent windows is used to complete them; for the outliers beyond the physical range, threshold truncation is performed according to the rated parameters of the mine explosion-proof high-voltage distribution device;
[0102] Deleting abnormal samples: Extract the corresponding indicators from the data in the labeled data set. The indicator types are time-domain features, frequency-domain features, and time-frequency domain features. According to the extracted indicators, cluster analysis is performed on the data in the labeled data set, and a screening strategy is set to eliminate outliers according to the screening strategy;
[0103] Time-domain features include mean, kurtosis, waveform factor, etc.; frequency-domain features include fundamental frequency amplitude of fast Fourier transform, proportion of third harmonic, etc.; time-frequency domain features include wavelet packet energy entropy, etc.
[0104] Sample class balance operation: Balance the class distribution through the SMOTE oversampling technique, amplify the lacking data, so that the ratio of the three types of data is: normal operation: slight deterioration: severe deterioration = 3:2:1.
[0105] It should be noted that the SMOTE oversampling technique can also be replaced by methods / models such as bilinear interpolation, generative adversarial network, autoencoder, etc. to achieve data augmentation and then achieve sample class balance.
[0106] In the specific implementation manner, the training process of the intelligent monitoring model for the fully insulated state of the mine explosion-proof high-voltage distribution device is as follows:
[0107] Input the preprocessed data set into the model for training. After the input data is dynamically normalized by the sliding window, capture long-term and short-term time series dependencies through the self-attention mechanism and the gating mechanism, and then input it into the probability hidden layer. Model the mean and covariance of the latent variables in the form of a Gaussian distribution, and perform uncertainty quantification through the reparameterization technique. During this period, the uncertainty is propagated between the probability hidden layers through the probability propagation rule, combined with the dynamic estimation of the process noise, to form a hierarchical uncertainty accumulation mechanism. Finally, fuse the probability features and the gating output, and generate the classification probability and the uncertainty quantification value through the mixture density network. When the highest class probability exceeds 0.85 and the uncertainty entropy is lower than 0.2, directly output the diagnostic result; for the fuzzy samples with uncertainty between 0.2 and 0.5, trigger the secondary verification mechanism and perform sliding window weighted decision by synthesizing the feature trend in the recent 10 minutes.
[0108] In the specific implementation manner, the operation of normalizing the preprocessed data set in the adaptive normalization layer is specifically as follows:
[0109] Adopt the adaptive normalization method based on the sliding window, and dynamically calculate the statistics within the window in combination with the time decay weight; first, calculate the weighted mean and weighted standard deviation within the sliding window, where the weight is determined by the time decay weight, and the data closer to the current moment has a greater weight. Then, normalize the data at the current moment using the weighted mean and standard deviation, and adjust the normalized result through the learnable scaling and offset parameters. Finally, obtain the normalized data.
[0110] The calculation formula for normalization is as follows:
[0111] ,
[0112] where, is the Data after moment normalization; is the data at the th moment in the preprocessed dataset; is the th moment weighted moving window mean; is the th moment weighted standard deviation;
[0113] The voltage and current amplitudes of the mine high-voltage distribution device change significantly over time. For the non-stationary time series characteristics, the local statistics of the moving window are more adaptable to local fluctuations. The th moment weighted moving window mean is calculated as follows:
[0114] ,
[0115] ,
[0116] where, is the length of the moving window; is another moment point after the current moment ; is the time decay weight, which can suppress historical noise interference such as sensor instantaneous anomalies; is the current moment and the th moment within the window; is the decay rate coefficient, which can control the decline speed of the weight with the increase of the time distance, and set ; is the exponential function with the natural constant as the base; is the data at the th moment in the preprocessed dataset;
[0117] The weighted standard deviation is obtained by calculating the weighted standard deviation of the data within the calculation window, which can reflect the local fluctuation amplitude. The th moment weighted standard deviation is calculated as follows:
[0118] ,
[0119] ,
[0120] ,
[0121] where, is a learnable scaling parameter, dynamically generated through a fully connected layer; is the Sigmoid activation function; is the weight matrix of the fully connected layer of the scaling parameter; is a learnable offset parameter, dynamically generated through a fully connected layer; is the weight matrix of the fully connected layer for the offset parameter.
[0122] In the specific implementation, the operation of extracting the temporal correlation features from the normalized data in the dual-channel temporal feature extraction module is as follows:
[0123] First, the normalized data is respectively mapped to a query vector, a key vector, and a value vector through a linear transformation. Then, the self-attention mechanism is used to calculate the similarity between the query vector and the key vector, and the features are reconstructed through the value vector to obtain the feature representation after self-attention weighting. Then, the normalized data is concatenated with the self-attention features, and the features are screened and adjusted through a gating mechanism to finally obtain the temporal correlation features;
[0124] The calculation formula in the dual-channel temporal feature extraction module is as follows:
[0125] ,
[0126] ,
[0127] Among them, is the query vector, is the key vector, is the value vector; is the normalized data; is the dimensionality scaling factor, used to scale the dot product result to stabilize the gradient; is the vector concatenation; is the feature after self-attention weighting; is the gated feature; is the gated weight matrix; is the hyperbolic tangent function; is the projection weight matrix.
[0128] Through the query weight matrix perform a linear transformation on to capture the temporal feature requirements at the current moment, that is, the query vector ; Through the key weight matrix perform a linear transformation on to represent the similarity relationship between data, that is, the key vector ; Through the value weight matrix perform a linear transformation on to obtain the feature reconstruction information, that is, the value vector ; The calculation formula of the linear transformation is as follows:
[0129] ,
[0130] ,
[0131] ,
[0132] Among them, is the query weight matrix, which is used to map the input normalized data to the query space; is the key weight matrix, which is used to map the input to the key vector space representing the similarity between data; is the value weight matrix, which is used to map the input to a value vector to provide feature reconstruction information;
[0133] There is a large amount of steady-state noise in the mine sensor data. The gating mechanism can filter out irrelevant features through nonlinear activation. By adjusting the importance of features in different paths through the gating mechanism and controlling the flow of the final features, redundant information can be effectively filtered out and the response to key timing features can be enhanced;
[0134] It should be noted that the failure of mine high-voltage equipment may be caused by early cumulative effects. Self-attention can capture potential associations across time steps. In particular, using the key vector and the value vector to represent the similarity relationship and feature reconstruction information of historical moments.
[0135] In the specific implementation manner, the operations in the probability hidden layer are as follows:
[0136] (1) Input the time-series correlation features into the probability hidden layer. The time-series correlation features are the output of the gating features , and use probability distribution modeling. Through the reparameterization trick, gradient backpropagation is performed so that the model can simultaneously learn feature representation and uncertainty quantification;
[0137] First, input the time-series correlation features into the probability hidden layer. The features are linearly transformed through the weight matrix and bias term to obtain the mean matrix of the Gaussian distribution. At the same time, the features are nonlinearly transformed using the covariance generation matrix to obtain the diagonal covariance matrix of the Gaussian distribution, and the features are represented in the probability form of the Gaussian distribution;
[0138] The probability form of the Gaussian distribution is as follows:
[0139] ,
[0140] ,
[0141] ,
[0142] Among them, is the probability distribution of the th layer of hidden variables; To be subject to a distribution; is the mean matrix of the Gaussian distribution; is the diagonal covariance matrix of the Gaussian distribution; represents a mean of , and a covariance of Gaussian distribution; represents converting a vector into a diagonal matrix; is the covariance generation matrix;
[0143] It should be noted that when the covariance generation matrix is randomly initialized during initialization, it is likely to cause the probability hidden layer parameters to deviate from the true data distribution;
[0144] (2) Generate a prior distribution based on the statistical characteristics of the normalized data to enable the model to quickly converge to a reasonable probability space;
[0145] First, calculate the mean and covariance matrix of all sample gated features. Then, perform a Cholesky decomposition on the diagonal covariance matrix of the Gaussian distribution to obtain the initial value of the covariance generation matrix, and use the statistical characteristics of the data to provide a reasonable initial value for the covariance generation matrix;
[0146] The calculation formula is as follows:
[0147] ,
[0148] where is the initial value of the covariance generation matrix; represents performing a Cholesky decomposition operation on the covariance matrix, which can ensure positive definiteness. Combining with the prior statistical characteristics of the data, enables the model to quickly enter a reasonable probability space; is the total number of sample data for model training in the preprocessed input dataset; is the th gated feature of the input data; is the mean of all sample gated features; represents transpose;
[0149] (3) Define the probability propagation chain rule between hidden layers to accumulate and transfer the uncertainty between hidden layers;
[0150] Define the probability propagation chain rule between hidden layers. For the hidden variables of each layer, calculate their posterior distribution through integration, where the integration range is all possible values of the hidden variables in the previous layer. During the integration process, update the mean and covariance of the hidden variables using the properties of the Gaussian distribution;
[0151] The calculation process is as follows:
[0152] ,
[0153] Among them, is the -th layer of hidden variables; is the -th layer of hidden variables; represents the conditional probability distribution of the -th layer of hidden variables given the -th layer of hidden variables; represents the integral over ; represents the prior distribution of the -th layer; is the weight matrix from the -th layer to the -th layer of the hidden layer of the probabilistic neural network; is the transpose of; is the bias term from the -th layer to the -th layer of the hidden layer of the probabilistic neural network; is the process noise covariance; Let be , be , be , then represents that follows a Gaussian distribution with mean and covariance ;
[0154] Mine equipment is affected by environmental interference, and the inter-layer noise transmission needs to be dynamically modeled. The process noise covariance is used to simulate the real noise propagation, and the calculation formula is as follows:
[0155] ,
[0156] ,
[0157] Among them, is the noise coefficient vector, which is dynamically generated by a multi-layer perceptron; represents the diagonal matrix formed by squaring the elements of the noise coefficient vector; is the multi-layer perceptron network.
[0158] In the specific implementation manner, the operation of class prediction in the output layer is specifically as follows:
[0159] Concatenate the output of the probabilistic hidden layer with the gated features, and generate class probabilities through a mixture density network;
[0160] First, concatenate the last-layer latent variables of the deep probabilistic neural network with the gated features. Then, perform a non-linear transformation on the gated features through a feature mapping function, and perform a weighted sum on the transformed features to obtain a linear combination of class probabilities. Then, use the Softmax function to normalize the linear combination to obtain the predicted probability for each class.
[0161] The calculation formula is as follows:
[0162] ,
[0163] where, is the predicted probability that the sample belongs to different classes when the last-layer latent variables of the given deep probabilistic neural network are provided; is the Softmax function, which is used to convert the input vector into a probability distribution is the weight of the output layer; is the bias term of the output layer; is the last-layer latent variable of the deep probabilistic neural network; denotes element-wise addition; denotes performing an average operation on the features of samples; is the gated feature of the k-th input sample data; is the feature mapping function, which is used to map the input features to a high-dimensional space.
[0164] In the specific implementation manner, optimize the intelligent monitoring model for the fully insulated state of the mine explosion-proof high-voltage distribution device:
[0165] Optimize the intelligent monitoring model for the fully insulated state of the mine explosion-proof high-voltage distribution device by calculating the loss function, updating the parameters, and adjusting the learning rate;
[0166] (1) Calculate the loss function:
[0167] Define the total loss function of the deep probabilistic neural network. The total loss function includes the classification loss, calibration loss, and covariance regularization term. Measure the difference between the model's predicted class probabilities and the true labels through the classification loss, measure the difference between the model's predicted uncertainty and the true uncertainty through the calibration loss, and constrain the model's uncertainty estimation through the covariance regularization term. Then, combine the three parts of the loss in a weighted manner to obtain the total loss function. The calculation formula is as follows:
[0168] ,
[0169] ,
[0170] ,
[0171] Among them, is the total loss function of the deep probabilistic neural network; is the classification loss weight coefficient. For example, set ; is the cross-entropy loss; is the true class label; is the calibration loss; represents the expectation of distribution; represents the predicted variance; represents the matrix trace operation; is the covariance regularization coefficient, set ; represents the covariance regularization term, is the diagonal covariance matrix of the Gaussian distribution, represents the number of layers of the deep probabilistic neural network;
[0172] (2) Parameter update:
[0173] Define the set of trainable parameters of the deep probabilistic neural network as , and use the probabilistic gradient update rule to update the training parameters of the deep probabilistic neural network;
[0174] The set of trainable parameters includes , , , , , , , , , , , , , and the weights and biases inside the perceptron;
[0175] During the update process, first calculate the gradient of the total loss function with respect to the parameters, and use the variational posterior distribution to perform weighted averaging on the gradient. Then, consider the gradient of the trace of the covariance matrix for parameter update, and adjust the gradient through the momentum decay rate. The calculation formula is as follows:
[0176] ,
[0177] Among them, is the parameter of the deep probabilistic neural network at the th iteration; is the parameter of the deep probabilistic neural network at the th iteration; is the learning rate for the e-th iteration; represents the expectation of the variational posterior distribution; is the variational posterior distribution; represents the gradient of the total loss function of the deep probabilistic neural network with respect to the parameter θ; represents the gradient operation on the parameter θ; is the momentum decay rate, set ; is the sign function; represents the sum of the traces of all hidden layer covariance matrices;
[0178] (3)Learning rate adjustment:
[0179] An adaptive strategy based on uncertainty entropy is adopted to adjust the learning rate;
[0180] First, calculate the uncertainty entropy of the model at each iteration step, and then, dynamically adjust the learning rate according to the average value of the uncertainty entropy. When the overall uncertainty of the model is high, reduce the learning rate, and when the uncertainty of the model is low, increase the learning rate:
[0181] The calculation formula of the learning rate is as follows:
[0182] ,
[0183] where, is the learning rate for the e-th iteration; is the base learning rate; is the learning rate decay coefficient, set ; is the maximum number of iterations is a positive integer; is the input data for the e-th iteration; represents the predicted distribution of the model for the input data at the e-th iteration represents the predicted distribution of the Shannon entropy.
[0184] Example 2
[0185] As Figure 2 and Figure 3As shown in the comparison graph of the time-domain signal and the normalization effect, the trend capture capabilities of different methods are intuitively compared. The sub-graph of the original signal shows the typical characteristics of non-stationary time series, with a slowly changing trend term superimposed on high-frequency noise and burst pulses. The sub-graph of the normalization effect shows that the global method allows the influence of pulse interference to continue to spread. However, through the calculation of the weighted statistic (sliding Z-score) of the sliding window in the present invention, local anomalies are effectively suppressed while maintaining the integrity of the trend. At the pulse positions around 4.2 seconds and 7.8 seconds, only instantaneous perturbations appear in the adaptive normalization curve, verifying the inhibitory effect of the time decay weight on historical noise.
[0186] As Figure 4 shown, it is the adaptability of the adaptive normalization method adopted by the present invention to non-stationary time series data. By comparing the traditional global normalization, sliding window normalization and the adaptive normalization method proposed by the present invention, Figure 4 the two core indicators of classification accuracy and uncertainty entropy are presented in a double-ordinate bar chart. The left column shows the classification accuracy, and the right column represents the degree of uncertainty of the model prediction. The experimental results show that due to ignoring local fluctuations, the traditional global normalization has the lowest accuracy and the highest uncertainty. Although the sliding window normalization is partially improved, there is still a problem of statistical lag. The adaptive method of the present invention's technology significantly outperforms the previous two in terms of accuracy through dynamic weighted statistics and learnable parameters, and has the lowest uncertainty entropy. Figure 4 The difference in the height of the columns in it intuitively reflects the ability of dynamic normalization to capture local time series features, indicating that the synergistic effect of time decay weight and parameter adaptability can effectively suppress sensor noise interference and improve the robustness of feature expression.
[0187] Example 3
[0188] As Figures 5 to 7 shown, time-frequency analysis is used to compare the feature extraction effects of different methods. The time-frequency spectrum of the original signal shows a mixed state of low-frequency trend (0-2Hz) and high-frequency noise (>10Hz). The traditional method shows energy diffusion in the low-frequency region, while the time-frequency spectrum processed by the technology of the present invention shows clear energy aggregation at the fundamental frequency of 0.8Hz, and effectively suppresses the 15Hz high-frequency noise. Especially at the transient pulse near 4 seconds, the dual-channel processing retains the time-frequency features of the pulse without introducing false frequency components, proving that through the dual-channel structure that combines the self-attention mechanism and the gating mechanism, the long-term dependence modeling ability and the local anomaly suppression ability can be taken into account, and early latent anomalies and short-term mutation features can be captured.
[0189] Example 4
[0190] As Figure 8As shown, by visualizing the propagation path of hidden layer uncertainty, the differences in probability modeling of different methods are analyzed. The covariance ellipse of the technology of the present invention gradually expands with the network depth, showing a reasonable cumulative process of uncertainty. The ellipse of the deterministic network is almost invisible, indicating its inability to model uncertainty. The ellipse of the Bayesian network shows irregular changes, especially in the output layer. The major axis direction of the ellipse of the technology of the present invention is consistent with the feature abstraction dimension, verifying the conformity of the probability propagation chain rule to the noise transmission rule of the physical system.
[0191] Although the specific implementation manners of the invention are described above in conjunction with the accompanying drawings, they are not limitations on the protection scope of the present invention. Based on the technical solutions of the present invention, various modifications or deformations that can be made by those skilled in the art without creative efforts are still within the protection scope of the present invention.
Claims
1. An intelligent monitoring method for the full insulation state of a mine explosion-proof high-voltage power distribution device, characterized in that It includes the following steps: S1. Collect training data: Collect the full-condition operation data of the explosion-proof high-voltage distribution device for mines through a multi-source sensor system and construct a data set. Segment and store the collected data by the sliding window method to obtain a structured data set; S2. Label the training data: Label the data in the structured data set by using a labeling method combining expert experience and off-line detection to form a labeled data set containing the mapping relationship between multi-dimensional features and state labels; S3. Preprocess the training data: Regularize the data in the labeled data set by using the industrial time-series data standardization process. The regularization operations include missing value processing, deleting abnormal samples, and sample category balancing operations to obtain a preprocessed data set; S4. Construct and train the model: Based on the deep probability neural network, construct an intelligent monitoring model for the full-insulation state of the explosion-proof high-voltage distribution device for mines. This model includes an adaptive normalization layer, a dual-channel time-series feature module, a probability hidden layer, and an output layer. Among them, the adaptive normalization layer adopts the sliding window dynamic normalization method. The inputs of the two branches of the dual-channel time-series feature module are the long-term and short-term time-series dependencies captured by the self-attention mechanism and the gated unit respectively. The probability hidden layer is a multi-layer fully-connected neuron hidden layer; Input the preprocessed data set into the model for training, and repeat the iterative training process until the preset stop iteration condition is met to obtain a trained model; S5. Intelligent monitoring and classification of the full-insulation state: Input the preprocessed real-time data into the trained intelligent monitoring model for the full-insulation state of the explosion-proof high-voltage distribution device for mines. The output layer generates the probability distributions of three categories, namely normal operation, slight deterioration, and severe deterioration, and the corresponding uncertainty quantification values, and obtain the final diagnosis result according to the preset discrimination conditions; S6. Incremental update of the model: Establish a dynamic incremental learning mechanism, deploy an on-line monitoring module, set a trigger value. When the number of unseen data samples continuously detected is greater than the trigger value, automatically trigger data collection. After manual review, store it in the incremental data set, and regularly use the incremental data set for incremental training of the intelligent monitoring model for the full-insulation state of the explosion-proof high-voltage distribution device for mines.
2. The intelligent monitoring method for the fully insulated state of the explosion-proof high-voltage power distribution device for mine use according to claim 1, characterized in that, S1 is specifically as follows: Deploy a multi-source sensor system at the key nodes of the explosion-proof high-voltage distribution device for mines for data collection. The multi-source sensor system includes a voltage transformer, a current transformer, a temperature sensor, and a partial discharge detector. The key monitoring nodes include the main circuit insulating sleeve, the breaker contact, and the cable joint; The multi-source sensor system synchronously collects time-series signals at a set sampling rate. The time-series signals include voltage fluctuations, leakage current, temperature rise gradient, and partial discharge pulse sequences; Simulate different working condition scenarios, continuously collect data under different working condition scenarios. The collected data covers three operating states: normal operation, slight deterioration, and severe deterioration. The working condition scenarios include load switching, overvoltage impact, environmental change, temperature change, and humidity change; Use the sliding window method to segment and store the collected time-series signal data, set the window length and the overlapping rate of adjacent windows to obtain a structured data set with time-series alignment and multi-dimensional physical quantities.
3. The intelligent monitoring method for the full-insulation state of the explosion-proof high-voltage power distribution device for mine use according to claim 2, wherein S2 is specifically as follows: First, electrical engineers label the operating status of each data in the structured dataset according to the IEC insulation status assessment standard, in combination with the insulation resistance test records, partial discharge intensity thresholds, and historical fault cases in the equipment maintenance log. The operating status includes normal operation, slight deterioration, and severe deterioration; Then, for the boundary samples with disputes, an offline dielectric loss tangent value tester is used for supplementary measurement, and the annotation accuracy is verified through frequency-domain dielectric characteristic analysis; Finally, an annotated dataset containing the mapping relationship between multi-dimensional features and status labels is formed.
4. The intelligent monitoring method for the fully insulated state of a mine explosion-proof high-voltage power distribution device according to claim 3, characterized in that S3 Specifically as follows: Missing value processing: For the zero-value or constant-value segments caused by the instantaneous failure of the sensor, the adjacent window linear interpolation method is used to complete them; for the abnormal values exceeding the physical range, threshold truncation is performed according to the rated parameters of the mine explosion-proof high-voltage distribution device; Deleting abnormal samples: Extract the corresponding indicators from the data in the annotated dataset. The indicator types are time-domain features, frequency-domain features, and time-frequency domain features. According to the extracted indicators, cluster analysis is performed on the data in the annotated dataset, and a screening strategy is set to eliminate outliers according to the screening strategy; Sample category balance operation: The SMOTE oversampling technique is used to balance the category distribution, and the lacking data is amplified so that the ratio of the three types of data is: normal operation: slight deterioration: severe deterioration = 3:2:
1.
5. The intelligent monitoring method for the fully insulated state of the mine explosion-proof high-voltage power distribution device according to claim 4, characterized in that The training process of the intelligent monitoring model for the full insulation status of the mine explosion-proof high-voltage distribution device is as follows: The preprocessed dataset is input into the model for training. After the input data is dynamically normalized by a sliding window, the long-term and short-term time series dependencies are captured through the self-attention mechanism and the gating mechanism, and then input into the probabilistic hidden layer. The mean and covariance of the latent variables are modeled in the form of a Gaussian distribution, and uncertainty quantification is performed through the reparameterization technique. During this period, uncertainty is transmitted between the probabilistic hidden layers through the probability propagation rule, and combined with the dynamic estimation of the process noise, a hierarchical uncertainty accumulation mechanism is formed. Finally, the probabilistic features are fused with the gating output, and the classification probability and uncertainty quantification value are generated through the mixture density network.
6. The intelligent monitoring method for the fully insulated state of a mine explosion-proof high-voltage power distribution device according to claim 5, characterized in that it is adaptive The operation of normalizing the preprocessed dataset in the normalization layer is specifically as follows: An adaptive normalization method based on a sliding window is adopted, and the statistics within the window are dynamically calculated in combination with the time decay weight; first, the weighted mean and weighted standard deviation within the sliding window are calculated, where the weight is determined by the time decay weight, and the data closer to the current moment has a greater weight. Then, the data at the current moment is normalized using the weighted mean and standard deviation, and the normalized result is adjusted through learnable scaling and offset parameters. Finally, the normalized data is obtained.
7. The intelligent monitoring method for the fully insulated state of the explosion-proof high-voltage power distribution device for mine use according to claim 6, characterized in that, The operation of extracting the time series correlation features from the normalized data in the dual-channel time series feature extraction module is specifically as follows: First, the normalized data are respectively mapped into query vectors, key vectors, and value vectors through linear transformation. Then, the self-attention mechanism is used to calculate the similarity between the query vector and the key vector, and the features are reconstructed through the value vector to obtain the feature representation after self-attention weighting. Then, the normalized data are concatenated with the self-attention features, and the features are screened and adjusted through the gating mechanism to finally obtain the time series correlation features; The calculation formula in the dual-channel time series feature extraction module is as follows: , , Among them, is the query vector, is the key vector, is the value vector; is the normalized data; is the dimensional scaling factor, used to scale the dot product result to stabilize the gradient; is the vector concatenation; is the feature after self-attention weighting; is the gated feature; is the gated weight matrix; is the hyperbolic tangent function; is the projection weight matrix.
8. The intelligent monitoring method for the fully insulated state of a mine explosion-proof high-voltage power distribution device according to claim 7, characterized in that the probability The operations in the hidden layer are specifically as follows: (1) Input the temporal correlation features into the probabilistic hidden layer, where the temporal correlation features are the gated feature outputs , adopt probabilistic distribution modeling, and perform gradient backpropagation through the reparameterization trick, enabling the model to learn both feature representation and uncertainty quantification simultaneously; First, the time series correlation features are input into the probabilistic hidden layer, and the features are linearly transformed through the weight matrix and the bias term to obtain the mean matrix of the Gaussian distribution. At the same time, the features are nonlinearly transformed through the covariance generation matrix to obtain the diagonal covariance matrix of the Gaussian distribution, and the features are represented in the probabilistic form of the Gaussian distribution; (2)Generate a prior distribution based on the statistical characteristics of the normalized data to enable the model to quickly converge to a reasonable probability space; First, calculate the mean and covariance matrix of all sample gating features. Then, perform Cholesky decomposition on the diagonal covariance matrix of the Gaussian distribution to obtain the initial value of the covariance generation matrix, and use the statistical characteristics of the data to provide a reasonable initial value for the covariance generation matrix; (3)Define the probability propagation chain rule between hidden layers to accumulate and transfer the uncertainty between hidden layers; Define the probability propagation chain rule between hidden layers. For the hidden variables of each layer, calculate their posterior distribution through integration, where the integration range is all possible values of the hidden variables of the previous layer. During the integration process, use the properties of the Gaussian distribution to update the mean and covariance of the hidden variables.
9. The intelligent monitoring method for the fully insulated state of the explosion-proof high-voltage power distribution device for mine use according to claim 8, characterized in that, The operation of class prediction in the output layer is specifically as follows: Concatenate the output of the probabilistic hidden layer with the gating features, and generate class probabilities through a mixture density network; First, concatenate the last-layer hidden variables of the deep probabilistic neural network with the gating features. Then, perform a nonlinear transformation on the gating features through a feature mapping function, and perform a weighted sum on the transformed features to obtain a linear combination of class probabilities. Then, use the Softmax function to normalize the linear combination to obtain the prediction probability of each class.
10. The intelligent monitoring method for the fully insulated state of the mine explosion-proof high-voltage power distribution device according to claim 9, characterized in that, Optimize the intelligent monitoring model for the fully insulated state of the mine explosion-proof high-voltage distribution device: Optimize the intelligent monitoring model for the fully insulated state of the mine explosion-proof high-voltage distribution device by calculating the loss function, parameter updating, and learning rate adjustment; (1)Calculate the loss function: Define the total loss function of the deep probabilistic neural network. The total loss function includes classification loss, calibration loss, and covariance regularization term. Measure the difference between the model's predicted class probabilities and the true labels through the classification loss, measure the difference between the model's predicted uncertainty and the true uncertainty through the calibration loss, and constrain the model's uncertainty estimation through the covariance regularization term. Then, combine the three parts of the loss through weighting to obtain the total loss function. The calculation formula is as follows: , , , Among them, is the total loss function of the deep probabilistic neural network; is the classification loss weight coefficient, such as; is the cross-entropy loss; is the true class label; is the calibration loss; denotes the expectation of distribution; represents the predicted variance; represents the matrix trace operation; is the covariance regularization coefficient; represents the covariance regularization term, is the diagonal covariance matrix of the Gaussian distribution, represents the number of layers of the deep probabilistic neural network; (2)Parameter updating: Define the set of trainable parameters of the deep probabilistic neural network as , and update the training parameters of the deep probabilistic neural network using the probabilistic gradient update rule; During the update process, first calculate the total loss function the gradient with respect to the parameters, and use the variational posterior distribution to perform weighted averaging on the gradient. Then, consider the gradient of the trace of the covariance matrix for parameter update, and adjust the gradient through the momentum decay rate. The calculation formula is as follows: , Among them, is the parameter of the deep probabilistic neural network for the th iteration; is the parameter of the deep probabilistic neural network for the th iteration; is the learning rate for the e-th iteration; represents the expectation of the variational posterior distribution; is the variational posterior distribution; represents the gradient of the total loss function of the deep probabilistic neural network with respect to the parameter θ; represents the gradient operation on the parameter θ; is the momentum decay rate; is the sign function; represents the sum of the traces of all hidden layer covariance matrices; (3)Learning rate adjustment: Adopt an adaptive strategy based on uncertainty entropy to adjust the learning rate; First, calculate the uncertainty entropy of the model at each iteration step. Then, dynamically adjust the learning rate according to the average value of the uncertainty entropy. When the overall uncertainty of the model is high, reduce the learning rate. When the uncertainty of the model is low, increase the learning rate: The calculation formula of the learning rate is as follows: , Among them, is the learning rate for the e-th iteration; is the base learning rate; is the learning rate decay coefficient; is the maximum number of iterations is a positive integer; is the input data for the -th iteration; represents the predicted distribution of the input data model for the -th iteration represents the Shannon entropy of the predicted distribution.
Citation Information
Patent Citations
Deep fusion network production line fault prediction method based on deep learning
CN119357769A
Dynamic health adaptive monitoring method and system using artificial intelligence
CN119480112A