Deep belief network-based power communication safety production multi-source heterogeneous data adaptive fusion method

By combining principal component analysis, TF-IDF, CNN, and DBN, the problems of feature redundancy and insufficient cross-modal correlation of multi-source heterogeneous data in power communication systems are solved, enabling more comprehensive data analysis and fault detection, and improving the security and reliability of power communication systems.

CN121117934BActive Publication Date: 2026-03-24CHANGCHUN UNIV OF SCI & TECH
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-09-01
Publication Date
2026-03-24

AI Technical Summary

Technical Problem

Existing power communication systems suffer from data silos, feature redundancy, and insufficient cross-modal correlation mining when processing multi-source heterogeneous data, making it difficult to adapt to the ever-changing power communication environment and limiting the accuracy and reliability of data analysis results.

Method used

Principal component analysis is used to reduce the dimensionality of structured data. Unstructured data features are extracted using TF-IDF and a pre-trained CNN network. Feature selection is performed using recursive feature elimination. Cross-modal feature interaction and fusion are achieved using a multi-layer stacked deep belief network (DBN). The training process is optimized using the Nadam algorithm to generate a comprehensive feature vector.

Benefits of technology

It has improved the data processing capabilities and security management level of the power communication system, enhanced the accuracy of fault detection and real-time monitoring capabilities, reduced the false alarm rate, optimized resource allocation, improved operational efficiency and targeted equipment maintenance, and promoted the intelligent and automated development of power communication.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121117934B_ABST
    Figure CN121117934B_ABST
Patent Text Reader

Abstract

The application relates to the field of power communication safety production management and solves the problems that existing data processing methods adopt single-dimensional analysis and are mostly static, data islands exist, feature redundancy exists, cross-modal correlation mining is insufficient, the methods are difficult to adapt to the changing power communication environment, and the accuracy and reliability of data analysis results are limited. The application obtains structured and unstructured data related to power communication system safety generation, performs preprocessing, then applies PCA dimension reduction to the structured data; for the unstructured data, jointly uses TF-IDF and a pre-trained CNN to extract a feature vector; through RFE, key features are adaptively selected until the best feature set is found; and the DBN network and the Nadam optimization algorithm are used to fuse the multi-source features, improve the comprehensiveness and accuracy of analysis, and optimize power communication safety management.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of power communication safety production management, specifically to an adaptive fusion method for multi-source heterogeneous data in power communication safety production based on deep belief networks. Background Technology

[0002] With the rapid development of the power industry, power communication systems are playing an increasingly important role in ensuring power security and improving management efficiency. However, power communication systems face numerous challenges, including equipment failures, network attacks, and data loss. These problems not only affect the reliable supply of electricity but may also have serious social and economic impacts. Therefore, improving the security and risk management capabilities of power communication systems has become an urgent need for the industry.

[0003] Traditional data processing methods often focus on the analysis of structured data, neglecting the importance of unstructured data. With the diversification of data sources, effectively integrating and analyzing multi-source heterogeneous data has become crucial for improving decision support capabilities. Furthermore, existing feature selection methods are mostly static and struggle to adapt to the ever-changing power communication environment, limiting the accuracy and reliability of data analysis results.

[0004] The rapid development of deep learning technology has brought new opportunities to the field of power communications. Deep learning models excel at processing complex data and extracting high-dimensional features, effectively improving the comprehensiveness of data analysis. However, how to effectively select and fuse features in deep learning models to enhance model performance remains a problem that urgently needs to be solved.

[0005] Therefore, this invention develops a multi-source heterogeneous data processing method that combines an adaptive feature selection mechanism with deep learning technology. This method can not only improve the data processing capabilities of power communication systems but also enhance their security management level, providing a solid technical foundation for the intelligent transformation of the power industry. Summary of the Invention

[0006] This invention addresses the problems of existing data processing methods that employ single-dimensional analysis and are mostly static, resulting in data silos, feature redundancy, and insufficient cross-modal correlation mining. These problems make it difficult to adapt to the ever-changing power communication environment, thus limiting the accuracy and reliability of data analysis results. The invention provides a method for fusing multi-source heterogeneous data in power communication safety production based on deep belief networks.

[0007] A method for fusing multi-source heterogeneous data for safe power communication production based on deep belief networks, which is implemented by the following steps:

[0008] Step S1: collect data related to the safe operation of the power communication system at the same time, classify the data into structured data and unstructured data, and preprocess the structured data and unstructured data;

[0009] Step S2: using principal component analysis to reduce the dimensionality of the structured data, obtaining principal component features; for text data in unstructured data, extracting features by TF-IDF, and for image data, extracting features by pre-training CNN network model;

[0010] Step S3: using RFE method to select the optimal feature set from all features obtained in step S2;

[0011] Step S4: input the optimal feature set obtained in step S3 into the deep belief network DBN constructed by multi-layer stacking RBM, compress the features layer by layer and realize cross-modal interaction, and finally output the fusion features; specifically:

[0012] Using DBN model to fuse the selected features, including setting the network structure of DBN model, determining the number and node number of hidden layers to adapt to the complexity of features;

[0013] Then, input the extracted feature vector into the DBN model for pre-training, use unsupervised learning method to make the DBN model adjust and identify the relationship between features, and use supervised fine-tuning to adapt the fusion features to downstream tasks;

[0014] Finally, the Nadam algorithm is used to optimize the training process to generate a comprehensive feature vector; the optimization algorithm formula is as follows:

[0015] Momentum direction correction:

[0016]

[0017] In the formula, m t is the first moment estimate, is the corrected first moment, θ t is the model parameter of the t-th iteration, is the gradient of the loss function to the parameter, and β1=0.9 is the momentum decay coefficient;

[0018] Adaptive learning rate base:

[0019]

[0020]

[0021] In the formula, β2=0.999 is the adaptive rate decay coefficient, v t is the second moment estimate at the t-th iteration, is the corrected second moment;

[0022] Nesterov lookahead gradient calculation:

[0023]

[0024] where, is the lookahead gradient;

[0025] Nadam final update:

[0026]

[0027] where, η = 0.001 is the initial learning rate, ∈ = 10 -8 is a constant to prevent the denominator from being zero.

[0028] Advantages of the present application:

[0029] 1. The multi-source heterogeneous data adaptive fusion method described in the present application integrates the structural data and non-structural data related to power communication safety production together, providing a more comprehensive information perspective. This all-round data vision enables the system to capture more potential problems and abnormal situations, thereby improving the accuracy of fault detection.

[0030] 2. In real-time monitoring and risk warning, it can help the system to quickly identify abnormalities and issue timely warnings, significantly reducing the risk of system failure. Data fusion also optimizes resource allocation. By integrating data from different sources, the system can more effectively allocate resources, improve operational efficiency, reduce unnecessary expenses and time waste, and the operation and maintenance team can also conduct targeted equipment maintenance and repair based on the results of comprehensive analysis, reducing maintenance costs. At the same time, data fusion reduces the false alarm rate, and the integration of multi-dimensional data enables the system to reduce the misleading of isolated data, improve the accuracy of fault identification, and thus reduce the need for manual intervention, improving work efficiency and reducing the burden on operation and maintenance personnel.

[0031] 3. Data fusion enhances the efficiency of overall data analysis, enabling the system to operate efficiently even when faced with large-scale data, ensuring rapid response in complex environments. The implementation of the method of the present application will help to achieve more efficient risk warning and fault diagnosis, promoting the intelligent and automated development of power communication safety production management. In summary, data fusion significantly improves the safety and reliability of the power communication system, providing solid technical support for the intelligent development of the industry.

[0032] 4. The multi-source heterogeneous data adaptive fusion method according to the present application, through the cooperative processing and deep feature fusion of structured and unstructured data, the method organically integrates principal component analysis (PCA), recursive feature elimination (RFE) and deep belief network (DBN), structured data is dimensionally reduced by PCA to form a multi-modal feature set with text TF-IDF features and image CNN features, redundant information is removed through a dynamic feature selection mechanism, and then the deep nonlinear fusion of cross-domain features is realized by using the unsupervised hierarchical pre-training of DBN. This method significantly improves the comprehensiveness of power equipment state evaluation. Through the three mechanisms of data enhancement, dynamic feature optimization and deep learning fusion, not only the data utilization efficiency is improved, but also strong support is provided for practical applications such as intelligent monitoring, fault prediction and risk assessment. It provides an interpretable and adaptable cross-modal data analysis framework for power system safety monitoring, and fills the technical gap of multi-source heterogeneous data collaborative governance and intelligent decision-making in the power field. BRIEF DESCRIPTION OF DRAWINGS

[0033] Figure 1 The flowchart of the power communication safety production multi-source heterogeneous data fusion method based on deep belief network according to the present application.

[0034] Figure 2 The principle block diagram of the specific implementation of the power communication safety production multi-source heterogeneous data fusion method based on deep belief network according to the present application.

[0035] Figure 3 The ResNet-50 model provided by the present application extracts image features.

[0036] Figure 4 The data fusion process chart of the DBN network provided by the present application. DETAILED DESCRIPTION

[0037] Specific implementation one, combination Figures 1 to 4 To illustrate this embodiment, the power communication safety production multi-source heterogeneous data fusion method based on deep belief network is realized by the following steps:

[0038] Step S1: Collecting multi-source heterogeneous data related to the safe operation of the power communication system at the same time, dividing the multi-source heterogeneous data into structured data and unstructured data according to data types, and then pre-processing the structured data and unstructured data;

[0039] As shown in Figure 2 In this embodiment, in the power communication system safety monitoring scenario, first, multi-source heterogeneous data is collected at the same timestamp.

[0040] Structured data includes device information, network performance parameters, power operation parameters, security logs, environmental data, event records, and user information, etc. These data are usually stored in time series or discrete event form in time series databases or relational databases. Unstructured data includes device log files (text format) and video monitoring content (image format), which need to be synchronized in time across devices through a distributed message queue (such as Kafka) to ensure that the timestamp error of all data sources is controlled within milliseconds (for example, using NTP protocol to synchronize the clock, satisfying Δt≤10ms). After collection, structured data and unstructured data need to be associated through device ID and global timestamp to form a unified data index.

[0041] In this embodiment, power and environmental monitoring (dynamic environmental data) data as a key component of structured data, high-precision acquisition is realized through multi-level protocol fusion: in the physical layer, the UPS device in the machine room uploads the input and output voltage and the battery health status in real time through the Modbus TCP protocol, and the power distribution cabinet adopts the IEC 61850 standard to collect the three-phase current harmonic distortion rate, and the abnormal threshold is dynamically set to ≤5%; in terms of environmental data, the distributed temperature and humidity sensor (accuracy ±0.5℃ / ±3%RH) and the smoke detector (sensitivity 0.1dB / m) transmit data through the RS485 bus at a frequency of 1Hz, and the edge node is packaged as IEC60870-5-104 protocol message, and is time-aligned with the video monitoring stream (ONVIF protocol) and the firewall log at the millisecond level (NTP synchronization error ≤10ms). Dynamic environmental data in the transmission layer eliminates voltage transient noise through sliding window mean filtering, and adds CRC-32 check code to ensure integrity, and finally generates JSON time series data with device topology label, which is uniformly stored in the time series database with structured data such as network traffic and device alarm, and is associated with unstructured data (such as the path of infrared thermal imaging graph at abnormal time), forming a complete multi-modal data entity, providing accurate power communication machine room holographic image for subsequent feature fusion. The sliding window mean filtering formula is as follows:

[0042] Suppose the original dynamic environmental data sequence is x1, x2,...,x n , the window size is N (odd number is more common), and the output sequence y t after filtering is:

[0043]

[0044] In the formula, t represents the current time point (t≥N). N is the length of the sliding window (the larger the window, the stronger the smoothing effect, but the larger the delay). x t-v is the vth historical data point in the window.

[0045] In this embodiment, the specific process of preprocessing the collected structured data is as follows:

[0046] Preprocessing of structured data begins with handling missing values. If the missing value rate of a field exceeds 20% (e.g., a sensor has been malfunctioning for a long time), the field is directly deleted. For data with a low missing value rate, numerical fields (e.g., current values) are imputed using the mean or linear interpolation, while categorical fields (e.g., protocol type) are imputed using the mode. Subsequently, outliers are detected using the Z-score method; if the current value exceeds the equipment's rated range (e.g., 500A), it is truncated to the maximum value. Finally, the data is standardized using Min-Max normalization to unify the units of measurement. The mean formula is as follows:

[0047]

[0048] Among them, y i It is the fill value, x i Here, n represents the number of non-missing values. For time series data, linear interpolation is used, as shown in the following formula:

[0049]

[0050] Where x(t) is the interpolation result of the variable at the current time point t (the missing value to be filled), x(t) i ) and x(t i-1 ) is the previous discrete time t i-1 and the next discrete time t i The known observations. t is the current time at which interpolation is needed. i-1 and t i These are two known time points adjacent to t.

[0051] The formula for outlier detection using the Z-score method is as follows:

[0052]

[0053] Where X is the data point, μ is the mean, and σ is the standard deviation.

[0054] The Min-Max normalized unified dimensionless formula is as follows:

[0055]

[0056] Where, x min It is the minimum value of all samples in the dataset, x max It is the maximum value of all samples in the dataset, where x is the normalized original data in the dataset. norm It is the result of x after normalization, and its range is usually [0,1].

[0057] As Figure 2 shown in the following table, in this embodiment, the process of preprocessing non-structured data is as follows:

[0058] The unstructured data preprocessing needs to design cleaning and enhancement processes for text and image modalities respectively. The text data is cleaned to remove special characters, punctuation marks and stop words, and is segmented using a segmentation tool;

[0059] For text data (such as device logs), first, remove noise characters irrelevant to the business by regular expression matching, for example, remove special symbols (such as "#ERROR:", "$") in the log, redundant spaces and non-Chinese characters, and keep the core description information (such as "Transformer B phase temperature abnormally rises to 120℃"). Then load the customized stop word list in the power field, filter general function words (such as "of", "is") but keep professional terms (such as "insulator", "overload"), and perform fine-grained segmentation on the text through a segmentation tool (such as the zh_core_web_sm model of spaCy), for example, the original log "circuit breaker A phase current over limit" will be parsed into ['circuit breaker', 'A phase', 'current', 'over limit'], ensuring that the subsequent feature extraction can capture key entities and states.

[0060] For image data (such as video monitoring pictures), the preprocessing includes data enhancement and pixel normalization. Data enhancement expands sample diversity by random rotation (angle range ±15°), scaling (scale 0.8-1.2 times) and center cropping (fixed size 224x224 pixels) to simulate the observation conditions of the device under different angles and distances. The enhanced image needs to be normalized in pixel value, which can also be normalized by the Min-Max method described above. Wherein, x min =0 and x max =255 correspond to the 8-bit pixel value range (0-255) of the original image, and the normalized pixel value is compressed to the interval [0,1]. For example, if the original value of a pixel point in a monitoring picture is x=128, then the normalized result is x norm =128 / 255≈0.502. This operation can eliminate the pixel distribution deviation caused by light difference and provide standardized input for subsequent convolutional neural network (CNN) feature extraction.

[0061] The segmented text preprocessing result is stored in the Elasticsearch index in the form of word vector, supporting keyword search and context association; the image data and its metadata (such as device ID, timestamp) are stored in the MinIO object storage system, and the spatiotemporal alignment with structured data is realized through JSON files. Finally, the preprocessing effect needs to be verified: the text segmentation accuracy needs to be ≥90% (manual sampling inspection), and the enhanced image needs to completely retain the key components of the device (such as lightning arrestor, terminal) and be free of distortion.

[0062] Step S2: Principal component analysis (PCA) is used to reduce the dimensionality of the preprocessed structured data, removing redundant features and retaining the main information. For unstructured data, text data is quantified using term frequency-inverse document frequency (TF-IDF) to extract features, while image data is extracted using a pre-trained Inception CNN (initial network). The specific process is as follows:

[0063] Step S21: After the structured data is collected and preprocessed, it needs to be reduced in dimensionality and redundancy removed using Principal Component Analysis (PCA). First, the preprocessed structured data is standardized to eliminate dimensional differences. The standardized data must have a mean of 0 and a standard deviation of 1 to ensure that the influence weights of features of different magnitudes on the PCA direction are consistent. The data standardization formula is as follows:

[0064]

[0065] Where, x norm σ is the result after normalizing x, u is the mean of the feature column, and σ is the standard deviation of the feature column.

[0066] Calculate the covariance matrix of standardized data to reflect the linear correlation between features.

[0067]

[0068] Where Z is the standardized data matrix, with dimensions n×m (n is the number of samples, m is the number of features). C is the covariance matrix, with dimensions m×m, and is a symmetric positive semi-definite matrix.

[0069] The covariance matrix is ​​decomposed into eigenvalues ​​to extract the principal directions (eigenvectors) and their corresponding variances (eigenvalues). The eigenvalue decomposition formula is as follows:

[0070] Cv a =λ a v a

[0071] Where, λ a is the a-th eigenvalue, representing the variance of the data along the a-th principal direction. a These are the corresponding eigenvectors, which are unit vectors and orthogonal to each other.

[0072] After sorting by eigenvalue from largest to smallest (λ1≥λ2≥…≥λ) m The first k eigenvectors are selected to form the projection matrix. This process requires verifying the cumulative variance contribution rate of the principal components. Typically, dimensions with a cumulative contribution rate ≥ 85% are retained to balance information preservation and dimensionality reduction. The formula is as follows:

[0073]

[0074] Where k is the number of principal components selected, and m is the total number of original features.

[0075] By projecting the data into a low-dimensional space using a projection matrix, the resulting principal component features will be used for subsequent model training, reducing noise interference and improving computational efficiency. The formula is as follows:

[0076] X PCA =ZV k

[0077] Among them, V k X is the projection matrix composed of the first k eigenvectors, with dimensions m×k. PCA This is the data matrix after dimensionality reduction, with dimensions n×k.

[0078] Step S22: Extract features from unstructured data. Text and images need to be extracted using TF-IDF and pre-trained CNN network models, respectively.

[0079] Step S221: For text data (such as power equipment logs), firstly, based on the preprocessed word segmentation results (such as ['circuit breaker', 'phase A', 'current', 'over-limit']), calculate the frequency of occurrence (TF) of each word in a single log entry, using the following formula:

[0080]

[0081] Among them, f c,d N is the number of times word c appears in document d. d This is the total number of words in document d. The inverse document frequency (IDF) of this word is calculated using the global corpus, as shown in the following formula:

[0082]

[0083] Where H is the total number of documents, and n t This is the number of documents containing the word. Finally, the product of these two values ​​is used as the TF-IDF weight to generate a sparse vector representation, as shown in the following formula:

[0084] TF-IDF(c,d) = TF(c,d) × IDF(t)

[0085] TF-IDF values ​​highlight keywords with high discriminative power and suppress interference from common words. Through TF-IDF, textual data is transformed into numerical features, enabling machine learning models to effectively identify key fault modes in power logs.

[0086] Step S222: For image data, a pre-trained ResNet-50 model is used to extract deep features. ResNet-50 is a deep residual network containing 50 convolutional layers (including residual blocks). Its core design solves the gradient vanishing problem in deep networks through skip connections. The overall architecture is divided into 7 stages, such as... Figure 3 As shown.

[0087] The Bottleneck structure consists of three convolutional layers for each residual block, and gradients are directly passed through skip connections.

[0088] 1) 1×1 convolution: dimensionality reduction (reducing the number of channels).

[0089] 2) 3×3 convolution: extracting spatial features.

[0090] 3) 1×1 convolution: dimensionality increase (recovery of channel count).

[0091] 4) Skip connections: If the input and output dimensions are inconsistent, adjust the number of channels using a 1×1 convolution. The formula is shown below:

[0092]

[0093] Where y is the output feature tensor, i.e., the output of the residual block, x is the input feature tensor, i.e., the input of the residual block, F is the residual function (3-layer convolution operation), and W is the residual function. s For linear projections of skip connections (W when dimension matching) s (for identity mapping), W i This is the set of weight parameters (such as convolution kernels and fully connected weights) within the residual function.

[0094] In this implementation, a ResNet-50 model pre-trained on ImageNet is loaded using PyTorch. The normalized image is input into the network, the final classification layer (fully connected layer) is removed, and the output of the penultimate layer's global average pooling is extracted as a 2048-dimensional feature vector. This process utilizes transfer learning to retain the general visual features (such as edges and textures) trained on the ImageNet dataset and adapt them to specific scenarios of power equipment. For example, after processing a monitoring image with ResNet-50, high activation dimensions in the output vector may correspond to visual patterns such as insulator cracks or terminal overheating.

[0095] Step S3: Adaptive feature selection is performed on the structured and unstructured data features obtained in step S2 using the recursive feature elimination (RFE) method. Unimportant features are iteratively removed. After each iteration, the SVM model is retrained and its performance is evaluated until the optimal feature set is found.

[0096] In this implementation, after extracting features from both structured and unstructured data, all features are grouped together and adaptive feature selection is performed using the RFE (Resource-Free Evaluator) method. The RFE method iteratively trains the model, evaluates feature importance, and removes redundant features, ultimately selecting the subset of features that contribute most to the prediction of the target variable. This process begins with the full feature set. In each iteration, a base model of a Support Vector Machine (SVM) is trained. Feature importance is ranked based on the absolute value of the SVM's weight coefficients, and the feature with the lowest current importance is removed. The goal of the SVM is to find a hyperplane that maximizes the margin between different categories of data. For linearly separable data, the hyperplane equation is:

[0097] w·x+b=0

[0098] Here, w is the weight vector, which determines the orientation of the hyperplane. b is the bias term, which determines the position of the hyperplane.

[0099] SVM solves for w and b by minimizing the following loss function:

[0100]

[0101] Among them, ||w|| 2 ξ is the regularization term, controlling model complexity. D is the regularization parameter, balancing margin maximization with classification error. i As a slack variable, it allows for a small number of misclassifications.

[0102] In this embodiment, the weight vector w of the SVM can also be represented by a linear combination of the training data:

[0103]

[0104] Where, α i For Grand multipliers, only support vectors (α) i >0) contributes to w, y i x represents the sample class label (±1). i This is the feature vector of the sample.

[0105] After each feature removal, the SVM is retrained, and its performance metrics, accuracy and F1 score, are evaluated using cross-validation. Model performance is recorded for different numbers of features. The accuracy formula is as follows:

[0106]

[0107] Where TP (True Positive) represents the number of correctly predicted positive samples. TN (True Negative) represents the number of correctly predicted negative samples. FP (False Positive) represents the number of negative samples that were misclassified as positive. FN (False Negative) represents the number of positive samples that were misclassified as negative. The F1 score is as follows:

[0108]

[0109] in:

[0110]

[0111] Precision (balanced precision) is the proportion of true positive samples among those predicted as positive. Recall is the proportion of true positive samples that were correctly predicted.

[0112] Furthermore, the iteration continues until a preset termination condition is reached (such as the number of features decreasing to a threshold or a significant drop in model performance), ultimately selecting the feature set that optimizes the evaluation metrics. During this process, a balance must be struck between feature simplification and model performance to avoid information loss due to excessive removal. The formula for the termination condition is as follows:

[0113]

[0114] Among them, Metnic m δ is the evaluation metric value (e.g., F1) for the m-th iteration. δ is a preset tolerance threshold (e.g., 0.05, indicating that a 5% performance degradation is allowed).

[0115] Step S4: The optimal feature set obtained in Step S3 is processed through a multi-layer stacked Restricted Boltzmann Machine (RBM) Deep Belief Network (DBN) to compress features layer by layer and achieve cross-modal interaction. Unsupervised pre-training is performed using a contrastive divergence algorithm to minimize the energy function and learn the data distribution. Supervised fine-tuning is introduced after pre-training, utilizing the Nadam optimizer to automatically adjust the learning rate. A gradient update mechanism based on predicted future positions is also introduced to reduce directional oscillations during the update process, efficiently adapting to downstream tasks and generating high-order fusion features to improve decision-making accuracy.

[0116] like Figure 4As shown, in this embodiment, after completing the RFE screening, the feature set (joint feature vector) obtained through screening is input into the DBN for deep fusion. The DBN achieves cross-modal feature abstraction and interaction through the stacking of multiple Restricted Boltzmann Machines (RBMs). First, the number of nodes in the hidden layer is determined based on the input feature dimension. Typically, the number of hidden nodes in the first layer can be set to 0.5 to 2 times the input dimension. For example, when the input is 500-dimensional, the number of nodes in the first hidden layer can be set to 512 (approximately 1 times). At this time, the DBN is composed of 3 stacked RBMs, with the number of hidden layer nodes being 512, 256, and 128 respectively, compressing layer by layer to extract high-order interactive features. The visible layer and hidden layer of each RBM are connected by a weight matrix w, and the bias vectors are a (visible layer) and b (hidden layer). The specific process is as follows:

[0117] Step S41: In the unsupervised pre-training stage, a layer-by-layer greedy training strategy is adopted. Each RBM is trained using the contrastive divergence (CD-k) algorithm, with the goal of minimizing the energy function to learn the data distribution. Taking RBM1 as an example, the joint energy function of layer v (input features) and hidden layer h is:

[0118]

[0119] Among them, a i For visible layer bias, v i For visible layer features, b j For hidden layer bias, h j For hidden nodes, w i,j This is the weight matrix. The energy function measures the "cost" of jointly configuring the visible and hidden layers; a lower value indicates that the configuration is more likely. The training objective is to adjust the parameters to minimize the energy of the real data while increasing the energy of random noise.

[0120] Furthermore, the activation probabilities of the hidden and visible layers are determined by the Sigmoid function:

[0121]

[0122]

[0123] in, For visible layer features, the hidden node h j Weighted input. For the hidden layer to the visible layer features v i The feedback effect. Activation probability represents the likelihood of a node being activated given an input, and it conveys cross-modal feature correlation information through the weight matrix w.

[0124] Update weights and biases using contrastive divergence (CD-1):

[0125] ΔW i,j=∈( <v i h j > data - <v i h j > recon )

[0126] Where ∈ represents the learning rate (e.g., 0.01), and controls the step size of the parameter update. <v i h j > data This represents the expected co-occurrence of visible and hidden layer nodes under the input data distribution. <v i h j > recon The co-occurrence expectation is given after reconstructing the data using a single-step Gibbs sampling. By reducing the distribution difference between the real data and the reconstructed data, the RBM learns to generate a distribution similar to the input data. After completing the training of the first layer, its hidden layer output h(1) is used as the input to the visible layer of RBM2. The above process is repeated until all RBM pre-training is completed.

[0127] Step S42: After pre-training, supervised fine-tuning is performed using a loss function and backpropagation, based on the further requirements of the fused data, to adapt the fused features to downstream tasks. For example, in the equipment fault classification problem, the cross-entropy loss function is used, as shown in the following formula:

[0128]

[0129] Where N is the number of samples and C is the total number of categories; y l,c ∈{0,1}: the true label (one-hot encoded) of sample l, which is 1 if and only if the sample belongs to class c. p l,c Predict the probability that sample l belongs to class c for the model (Softmax output).

[0130] In this implementation, backpropagation gradient calculation is performed, and the weight gradient of each layer is calculated using the chain rule:

[0131]

[0132] in, Let be the error gradient of the j-th node in the l-th layer. It is the activation value of the i-th node in the (l-1)-th layer.

[0133] In this implementation, the Nadam optimization algorithm is used to improve convergence speed and resistance to local optima. The update rules are as follows:

[0134] Momentum direction correction (first moment):

[0135]

[0136] Adaptive learning rate basis (second moment):

[0137]

[0138] Nesterov look-ahead gradient calculation:

[0139]

[0140] Nadam final update:

[0141]

[0142] Where, θ t Let be the model parameters for the t-th iteration. β1 = 0.9 is the gradient of the loss function with respect to the parameters. β2 = 0.999 is the momentum decay coefficient, controlling the degree to which historical gradient directions are preserved. η = 0.001 is the adaptive learning rate decay coefficient, controlling the decay rate of the accumulated squared gradients. t This is a first-order moment estimate (momentum direction), which is an exponential moving average of the gradient. t The second-moment estimate (adaptive basis) is an exponential moving average of the squared gradient. The corrected first moment eliminates the zero bias in the early stages of the iteration. To obtain the corrected second moment, the zero bias from the initial iteration is eliminated. ∈=10 -8 To prevent constants with a denominator of zero. Forward gradient: at the predicted location The gradient is calculated. Nadam adjusts the update direction in advance through prospective gradient estimation (Nesterov momentum), approximating the optimal solution more accurately than Adam, and is particularly suitable for high-dimensional non-convex optimization problems of DBN. Finally, the fused features are applied to existing fault detection models in this field and compared with the original features under the same model and parameters. The quality of the fusion result is judged by the performance of the model.

[0143] In this implementation, DBN models the joint distribution of cross-modal features using an energy function, updates parameters using contrastive divergence to capture the intrinsic structure of the data, and then uses supervised fine-tuning to adapt the fused features to downstream tasks. The Nadam algorithm, with its fast convergence and resistance to local optima, is an ideal choice for optimizing DBN. The weight matrix w and bias terms a and b in the formula jointly encode the nonlinear relationships between features, while the loss function and optimization rules ensure that the model learns efficiently on complex data.

[0144] This invention integrates structured equipment data, unstructured logs, and video data from the field of power communication safety production. It utilizes Principal Component Analysis (PCA) for dimensionality reduction and Recursive Feature Elimination (RFE) to dynamically screen key features, and combines Deep Belief Networks (DBN) to achieve cross-modal feature fusion, thus constructing an intelligent analysis system for multi-source heterogeneous data. Its core function is to assist subsequent risk assessment research on power communication safety production, improving the comprehensiveness and effectiveness of power system safety monitoring: by extracting equipment fault signals, abnormal flow patterns, and visual features, it reduces fault response time from 45 minutes to 8 minutes, increases fault prediction accuracy to 89%, supports a data processing capacity of 5TB / day, and optimizes system expansion efficiency, providing end-to-end safety assurance for power communication networks from data acquisition and feature optimization to intelligent decision-making.

[0145] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on its differences from other embodiments. Similar or identical parts between embodiments can be referred to interchangeably. For the apparatus disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple; relevant parts can be referred to the method section.

[0146] The above description of the disclosed embodiments enables those skilled in the art to make or use the invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the invention. Therefore, the invention is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. An adaptive fusion method for multi-source heterogeneous data in power communication safety production based on deep belief networks, characterized by: This method is implemented by the following steps: Step S1: Collect relevant data on the safe operation of the power communication system at the same time, classify them into structured and unstructured data, and preprocess the structured and unstructured data; Step S2: Principal component analysis is used to reduce the dimensionality of the structured data to obtain principal component features; for unstructured text data, features are extracted using TF-IDF, and for image data, features are extracted using a pre-trained CNN network model. Step S3: Use the RFE method to perform adaptive feature selection on all features obtained in step S2, and select the optimal feature set; Step S4: Input the optimal feature set obtained in step S3 into a deep belief network (DBN) constructed by stacking multiple layers of RBMs, compress the features layer by layer and realize cross-modal interaction, and finally output the fused features; Specifically: The DBN model is used to perform data fusion on the selected features, including setting the network structure of the DBN model and determining the number of hidden layers and nodes to adapt to the complexity of the features. Then, the extracted feature vectors are input into the DBN model for pre-training. An unsupervised learning method is used to enable the DBN model to adjust itself and identify the relationships between features. Supervised fine-tuning is then used to adapt the fused features to downstream tasks. Finally, the training process is optimized using the Nadam algorithm to generate a comprehensive feature vector; the optimization algorithm formula is as follows: Momentum direction correction: ; ; In the formula, For first-order moment estimation, The corrected first moment, Let be the model parameters for the t-th iteration. The gradient of the loss function with respect to the parameters. The momentum decay coefficient; Adaptive learning rate basis: ; ; In the formula, This is the adaptive rate attenuation coefficient. This is the second-order moment estimate for the t-th iteration. The corrected second moment; Nesterov look-ahead gradient calculation: ; In the formula, Forward-looking gradient; Nadam Last Update: ; In the formula, The initial learning rate, To prevent constants with a denominator of zero.

2. The adaptive fusion method for multi-source heterogeneous data in power communication safety production based on deep belief networks according to claim 1, characterized in that: The structured data includes device information, network performance parameters, power operation parameters, security logs, environmental data, event records, and user information; the unstructured data includes log files and video surveillance content.

3. The adaptive fusion method for multi-source heterogeneous data in power communication safety production based on deep belief networks according to claim 1, characterized in that: Preprocessing of structured data includes handling missing values, identifying and handling outliers, and removing redundant data to ensure data format consistency.

4. The adaptive fusion method for multi-source heterogeneous data in power communication safety production based on deep belief networks according to claim 1, characterized in that: In step S2, principal component analysis is used to reduce the structured data to a low dimension, including standardized data, to calculate the covariance matrix, and then extract eigenvalues ​​and eigenvectors.

5. The adaptive fusion method for multi-source heterogeneous data in power communication safety production based on deep belief networks according to claim 1, characterized in that: Preprocessing of unstructured data includes cleaning text data to remove special characters, punctuation marks, and stop words, and segmenting it using word segmentation tools; image data undergoes data augmentation and pixel value normalization.

6. The adaptive fusion method for multi-source heterogeneous data in power communication safety production based on deep belief networks according to claim 1, characterized in that: To extract features from unstructured text data, TF-IDF is used to quantify the words in the text and extract their features. The formula for calculating TF is as follows: ; in, For words In the document The number of times it appears in For document The total word count; combined with the global corpus, the IDF of the word is calculated using the following formula: ; in, Total number of documents The number of documents containing the word is used; finally, the TF-IDF weights are calculated to generate a sparse vector representation, as shown in the following formula: 。 7. The adaptive fusion method for multi-source heterogeneous data in power communication safety production based on deep belief networks according to claim 1, characterized in that: Feature extraction is performed on image data in unstructured data. A pre-trained ResNet-50 model is used to extract deep features. The normalized image is then input into the pre-trained ResNet-50 model. The last classification layer is removed, and the feature vector of the penultimate global average pooling layer is extracted.

8. The adaptive fusion method for multi-source heterogeneous data in power communication safety production based on deep belief networks according to claim 1, characterized in that: In step S3, the SVM model is trained iteratively using the RFE method. The importance of features is ranked based on the absolute value of the weight coefficients of the SVM model. The features with the lowest current importance are removed until the optimal feature set is found.

Citation Information

Patent Citations

  • Data fusion processing method, system and equipment for power system

    CN114239683A

  • Geological disaster early warning method and accurate early warning system based on multi-source data fusion

    CN120452170A