Data fusion mining method and system based on multi-modal power cross-domain
Through the multimodal power cross-domain data fusion mining method, the fusion problem of multi-source heterogeneous data in the power system is solved by using technologies such as tensor decomposition, domain adaptation feature mapping and graph convolution networks, and the fusion problem of multi-source heterogeneous data in the power system is achieved, and data utilization efficiency is improved and model optimization is achieved.
Patent Information
- Application Number
- CN202510217727.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-26
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2045-02-26
AI Technical Summary
In modern power systems, the heterogeneity and diversification of multimodal data are high, resulting in increased difficulty in deep mining and fusion of data, especially in the problems of inconsistent data formats, dispersed data sources and low data utilization efficiency.
The data fusion mining method based on multimodal power cross-domain is adopted, and the core features of the power are reconstructed through tensor decomposition and feature selection, domain adaptive feature mapping and domain difference elimination are carried out, graph convolutional networks and distributed model training are constructed, and combined with federal average computing and teacher-student network construction, efficient data fusion and global model optimization are achieved.
It significantly improves the utilization efficiency of multi-source data in the power system, reduces data processing costs, solves the problems of inconsistent data formats and dispersed data sources, and improves the robustness and inference efficiency of the model.
Smart Images

Figure CN120123773A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of power data fusion, and in particular, to a multi-modal power cross-domain data fusion and mining method and system. Background Art
[0002] In modern power systems, the sources of data are becoming increasingly rich, covering all aspects from power generation, transmission, transformation to distribution and consumption. These data not only include traditional power parameters (such as voltage, current, power, etc.), but also cover equipment operation status, environmental monitoring data, user behavior data, and various signals collected by intelligent sensors. The heterogeneity and diversity of this multi-modal data are extremely high, and the data types range from structured data (such as tabular data) to semi-structured data (such as log files) and unstructured data (such as images, voices, and texts). Such a complex data environment greatly increases the difficulty of in-depth data mining and fusion. For example, consider the scenario of an intelligent substation. In this substation, there are not only traditional power equipment operation data, such as the measurement data of current transformers and voltage transformers, but also the appearance images of equipment collected by intelligent cameras, the temperature and humidity data recorded by environmental sensors, and the user power consumption behavior data collected by intelligent meters. These data come from different devices and sensors, have different formats, and are distributed in different systems. To achieve efficient fusion of these data, problems such as inconsistent data formats, scattered data sources, and low data utilization efficiency need to be solved. Summary of the Invention
[0003] Based on this, it is necessary for the present invention to provide a multi-modal power cross-domain data fusion and mining method and system to solve at least one of the above technical problems.
[0004] To achieve the above object, a multi-modal power cross-domain data fusion and mining method includes the following steps:
[0005] Step S1: Obtain a multi-source heterogeneous power data set; perform tensor decomposition on the multi-source heterogeneous power data set to obtain a power core feature tensor set; perform feature selection and reconstruction on the power core feature tensor set to obtain a reconstructed power feature data set;
[0006] Step S2: Perform domain adaptation feature mapping on the reconstructed power feature data set to obtain a cross-domain power feature mapping data set; perform domain difference elimination on the cross-domain power feature mapping data set to obtain a domain-aligned power feature data set; perform attention weight calculation on the domain-aligned power feature data set to obtain a cross-domain fusion power feature data set;
[0007] Step S3: Obtain the distributed node topology data; construct a graph convolutional network for the distributed node topology data to obtain the power node graph network structure data; perform distributed model training on the cross-domain fusion power feature dataset based on the power node graph network structure data to obtain the local power model parameter set; perform federated averaging calculation on the local power model parameter set to obtain the globally optimized model parameter set;
[0008] Step S4: Construct a teacher network for the globally optimized model parameter set to obtain the power multi-modal teacher model; perform student network training according to the power multi-modal teacher model to obtain the power lightweight student model; perform model fusion on the power lightweight student model and the power multi-modal teacher model to obtain the power feature extraction fusion model;
[0009] Step S5: Obtain the power multi-modal quality assessment data; construct a reinforcement learning strategy for the power multi-modal quality assessment data to obtain the power data dynamic fusion strategy; adjust the weights of the power feature extraction fusion model according to the power data dynamic fusion strategy to obtain the optimal power fusion decision model; use the optimal power fusion decision model to perform semantic association mining on the multi-source heterogeneous power dataset to obtain the power multi-modal semantic association dataset.
[0010] Through tensor decomposition and feature selection reconstruction, the present invention can extract core features from complex multi-source heterogeneous power datasets, effectively solving the problems of inconsistent data formats and scattered data sources. Through domain adaptation feature mapping and domain difference elimination, it can effectively solve the differences between multi-modal data in different domains. By constructing a graph convolutional network and distributed model training, combined with federated averaging calculation, it can make full use of the topological information of distributed nodes, improving the efficiency and robustness of model training. Through distributed training and federated learning, it avoids the computational pressure and data transmission bottlenecks brought by centralized processing, while ensuring the global optimization of the model. Through the construction of the teacher-student network and knowledge distillation strategy, the knowledge of the complex multi-modal teacher model can be transferred to the lightweight student model. This not only retains the semantic information and feature extraction ability of the teacher model, but also significantly reduces the computational complexity and storage requirements of the model, improving the inference efficiency of the model and making it more suitable for real-time applications and edge computing scenarios in actual power systems. By constructing a reinforcement learning strategy for the power data dynamic fusion strategy and adjusting the weights of the model according to the dynamic fusion strategy, it can achieve dynamic optimization of the multi-modal data fusion process. In addition, using the optimal power fusion decision model for semantic association mining can further explore the deep semantic information in multi-source heterogeneous data. In summary, the present invention can significantly improve the utilization efficiency of multi-source data in power systems.
[0011] Preferably, the present invention further provides a multi-modal power cross-domain data fusion and mining system for performing the above-mentioned multi-modal power cross-domain data fusion and mining method. The multi-modal power cross-domain data fusion and mining system includes:
[0012] A data preprocessing module, configured to obtain a multi-source heterogeneous power data set; perform tensor decomposition on the multi-source heterogeneous power data set to obtain a power core feature tensor set; perform feature selection and reconstruction on the power core feature tensor set to obtain a reconstructed power feature data set;
[0013] A feature mapping module, configured to perform domain adaptation feature mapping on the reconstructed power feature data set to obtain a cross-domain power feature mapping data set; perform domain difference elimination on the cross-domain power feature mapping data set to obtain a domain-aligned power feature data set; perform attention weight calculation on the domain-aligned power feature data set to obtain a cross-domain fusion power feature data set;
[0014] A model training module, configured to obtain distributed node topology data; construct a graph convolutional network for the distributed node topology data to obtain power node graph network structure data; perform distributed model training on the cross-domain fusion power feature data set based on the power node graph network structure data to obtain a local power model parameter set; perform federated average calculation on the local power model parameter set to obtain a globally optimized model parameter set;
[0015] A model fusion module, configured to construct a teacher network for the globally optimized model parameter set to obtain a multi-modal power teacher model; perform student network training according to the multi-modal power teacher model to obtain a lightweight power student model; perform model fusion on the lightweight power student model and the multi-modal power teacher model to obtain a power feature extraction and fusion model;
[0016] A data mining module, configured to obtain power multi-modal quality assessment data; construct a reinforcement learning strategy for the power multi-modal quality assessment data to obtain a power data dynamic fusion strategy; adjust the weights of the power feature extraction and fusion model according to the power data dynamic fusion strategy to obtain an optimal power fusion decision model; use the optimal power fusion decision model to perform semantic association mining on the multi-source heterogeneous power data set to obtain a power multi-modal semantic association data set.
[0017] In the present invention, the entire system significantly improves the utilization efficiency of multi-source data in the power system and reduces the data processing cost through the efficient fusion, cross-domain alignment, distributed training, and dynamic optimization of multi-modal data. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] Other features, objects, and advantages of the present invention will become more apparent by reading the following detailed description with reference to the accompanying drawings:
[0019] Figure 1 The figure shows a schematic flow chart of the steps of a data fusion mining method based on multi-modal power cross-domain in an embodiment.
[0020] Figure 2 The figure shows a detailed schematic flow chart of step S4 in an embodiment. Specific embodiments
[0021] The technical method of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all of them. All other embodiments obtained by those skilled in the art within the scope of the present invention without creative work based on the embodiments of the present invention belong to the scope of protection of the present invention.
[0022] In addition, the accompanying drawings are only schematic diagrams of the present invention and are not necessarily drawn to scale. The same reference numerals in the figures represent the same or similar parts, and thus their repeated description will be omitted. Some of the block diagrams shown in the drawings are functional entities and do not necessarily correspond to physically or logically independent entities. The functional entities can be implemented in software form, or implemented in one or more hardware modules or integrated circuits, or implemented in different networks and / or processor methods and / or microcontroller methods.
[0023] It should be understood that although the terms "first", "second", etc. may be used here to describe each unit, these units should not be limited by these terms. These terms are only used to distinguish one unit from another. For example, without departing from the scope of the exemplary embodiments, the first unit can be called the second unit, and similarly the second unit can be called the first unit. The term "and / or" used here includes any and all combinations of one or more of the listed related items.
[0024] To achieve the above object, please refer to Figures 1 to 2 , the present invention provides a data fusion mining method based on multi-modal power cross-domain, including the following steps:
[0025] Step S1: Obtain a multi-source heterogeneous power data set; perform tensor decomposition on the multi-source heterogeneous power data set to obtain a power core feature tensor set; perform feature selection and reconstruction on the power core feature tensor set to obtain a reconstructed power feature data set;
[0026] Step S2: Perform domain adaptation feature mapping on the reconstructed power feature data set to obtain a cross-domain power feature mapping data set; perform domain difference elimination on the cross-domain power feature mapping data set to obtain a domain-aligned power feature data set; perform attention weight calculation on the domain-aligned power feature data set to obtain a cross-domain fusion power feature data set;
[0027] Step S3: Obtain the distributed node topology data; construct a graph convolutional network for the distributed node topology data to obtain the power node graph network structure data; perform distributed model training on the cross-domain fusion power feature dataset based on the power node graph network structure data to obtain the local power model parameter set; perform federated average calculation on the local power model parameter set to obtain the global optimized model parameter set;
[0028] Step S4: Construct a teacher network for the global optimized model parameter set to obtain the power multi-modal teacher model; perform student network training according to the power multi-modal teacher model to obtain the power lightweight student model; perform model fusion on the power lightweight student model and the power multi-modal teacher model to obtain the power feature extraction fusion model;
[0029] Step S5: Obtain the power multi-modal quality assessment data; construct a reinforcement learning strategy for the power multi-modal quality assessment data to obtain the power data dynamic fusion strategy; adjust the weights of the power feature extraction fusion model according to the power data dynamic fusion strategy to obtain the optimal power fusion decision model; use the optimal power fusion decision model to mine semantic associations in the multi-source heterogeneous power dataset to obtain the power multi-modal semantic association dataset.
[0030] In this embodiment, multi-source heterogeneous data is first collected from the power system, including voltage, current, power, equipment status, ambient temperature, humidity, and equipment images, etc. These data come from multiple substations and power generation stations. The Tensorly library of Python is used to perform tensor decomposition on the data to extract the power core feature tensor set, and the SelectKBest method of scikit-learn is used to select important features. Then, an autoencoder is used for feature reconstruction to form a reconstructed power feature data set. Next, a domain adaptation network based on TensorFlow is designed. The reconstructed power feature data set is input, and a cross-domain power feature mapping data set is output. The inter-domain differences are eliminated through an adversarial training mechanism and a discriminator network to obtain a domain-aligned power feature data set. The self-attention mechanism (such as the Transformer architecture) is used to calculate the attention weights of each feature to generate a cross-domain fusion power feature data set. Then, the topological relationship data of distributed power nodes is extracted from the SCADA system, and a graph convolutional network (GCN) is constructed using PyTorch Geometric. Based on the power node graph network structure data, the cross-domain fusion power feature data set is distributedly trained to obtain a local power model parameter set, and the federated average calculation is performed through the torch.distributed module of PyTorch to obtain a globally optimized model parameter set. Subsequently, a power multi-modal teacher model is constructed using the globally optimized model parameter set, a lightweight student network is designed, and the key semantic features of the teacher model are transferred to the student model through knowledge distillation. Finally, the two are fused to obtain a power feature extraction and fusion model. Finally, the output of the power feature extraction and fusion model is reliability evaluated to generate power multi-modal quality assessment data, and the Stable Baselines3 framework is used for reinforcement learning to design a power data dynamic fusion strategy. According to this strategy, the weights of the model are adjusted to obtain an optimal power fusion decision model, and this model is used to mine semantic associations in the multi-source heterogeneous power data set to extract a power multi-modal semantic association data set.
[0031] Preferably, step S1 includes the following steps:
[0032] Step S11: Collect multi-modal power data for the target power grid to obtain an original multi-modal power data set, and perform timestamp alignment on the original multi-modal power data set to obtain a time-series power data set;
[0033] Specifically, an intelligent electricity meter can be used to collect users' electricity consumption behavior data (such as electricity quantity, power, etc.), current transformers and voltage transformers can be used to collect the operating parameters of power equipment (such as current, voltage), environmental sensors can be used to record environmental data such as temperature and humidity in the substation, and image data of the equipment appearance can be collected through high-definition cameras. These data are stored in different local databases respectively, forming an original multi-modal power dataset. A distributed time synchronization system (such as a GPS clock synchronization module) is used to ensure that the time bases of all sensors and acquisition devices are consistent. The GPS clock synchronization module is connected to each sensor device, and the internal clocks of the sensors are calibrated through the Network Time Protocol. A data preprocessing software (such as the Pandas library in Python) is used to read the original data files of each data source and perform alignment operations according to the timestamp fields. For example, for the intelligent electricity meter data and current transformer data, sampling alignment is performed at intervals of every 15 minutes; for the image data, marking is performed at intervals of every hour, and finally a time-series power dataset is obtained.
[0034] Step S12: Perform anomaly detection on the time-series power dataset based on preset data quality evaluation metrics to obtain a power data anomaly marking set, and perform data cleaning on the time-series power dataset according to the power data anomaly marking set to obtain a multi-source heterogeneous power dataset;
[0035] Specifically, data quality evaluation metrics can be selected, such as data integrity, consistency, and accuracy. Taking current data as an example, the integrity metric requires that there are no missing values in the current data in the time series; the consistency metric requires that the change trend of the current data matches the change trend of the voltage data; the accuracy metric determines whether the current data is within a reasonable range by comparing it with historical data. The NumPy and SciPy libraries in Python are used to check the integrity of the current data, and interpolation methods are used to fill in the missing values. For consistency detection, the Pearson correlation coefficient is used to calculate the correlation between the current data and the voltage data. When the correlation coefficient is lower than 0.8, it is marked as abnormal data. For accuracy detection, an anomaly detection model based on a long short-term memory network is constructed. The historical current data sequence is input, and a judgment on whether the current data is abnormal is output. By training the LSTM model, it can learn the normal change pattern of the current data. When data outside the normal range is detected, it is marked as abnormal data and recorded in the power data anomaly marking set. For the marked abnormal data points, if the data is missing, it is filled in by linear interpolation or the average value based on the surrounding data; if the data fluctuates abnormally, it is smoothed by median filtering or the moving window average method. Finally, a multi-source heterogeneous power dataset is obtained.
[0036] Step S13: Identify the data types of the multi-source heterogeneous power data set to obtain the power data type feature set, and standardize the multi-source heterogeneous power data set according to the power data type feature set to obtain the standard heterogeneous power data set;
[0037] Specifically, the multi-source heterogeneous power data set includes structured data (such as numerical data like current and voltage), semi-structured data (such as device log files), and unstructured data (such as device appearance images). For structured data, use Pandas to read the data file and identify numerical data and timestamp type data by checking the data types of the data columns. For semi-structured data, use regular expressions to match the key fields (such as timestamp, device ID, error code) in the log file and extract them as structured data. For unstructured data (such as image data), use the OpenCV library to load the image file, identify it as image type data, and extract feature information such as the size and color channels of the image to form the power data type feature set. For numerical data, adopt the Z-score standardization method. By calculating the mean and standard deviation of the data, convert the data into standardized data with a mean of 0 and a standard deviation of 1. For timestamp data, uniformly convert it to the UTC time format. For log data, numerically encode the extracted key fields. For example, map the error code "ERROR_001" to the numerical value 1. For image data, scale the pixel values to between 0 and 1 through normalization processing. Through the above operations, finally obtain the standard heterogeneous power data set.
[0038] Step S14: Construct a tensor for the standard heterogeneous power data set to obtain a multi-dimensional power tensor data set;
[0039] Specifically, the standard heterogeneous power dataset contains various data such as current, voltage, temperature, humidity, and device appearance images. Select the tensor operation tool in the TensorFlow framework to construct tensors. Arrange the numerical data such as current, voltage, temperature, and humidity in a time series to form a four-dimensional tensor, where the dimensions are time, device number, data type (current, voltage, temperature, humidity), and data value respectively. For example, assume the dataset contains current, voltage, temperature, and humidity data of 10 devices at 100 time points. Then the shape of the constructed tensor is [100, 10, 4, 1], where the last dimension represents the data value. For the device appearance image data, since it is a two-dimensional image, it needs to be converted into a tensor form compatible with the numerical data. Use the OpenCV library to convert the image data into a grayscale image and flatten it into a one-dimensional vector. Assume the size of each image is 224×224 pixels, then the length of the flattened vector is 50,176. Arrange these vectors in a time series to form a three-dimensional tensor with a shape of [100, 10, 50176], where 100 represents the time point, 10 represents the device number, and 50176 represents the length of the one-dimensional vector of the image data. Finally, merge the numerical data tensor and the image data tensor into a multi-dimensional power tensor dataset with a shape of [100, 10, 50180], where the last dimension contains the features of the numerical data and the image data.
[0040] Step S15: Perform rank evaluation on the multi-dimensional power tensor dataset to obtain a power tensor rank evaluation dataset;
[0041] Specifically, it can be assumed that the shape of the multi-dimensional power tensor dataset is [100, 10, 50180], where 100 represents the time point, 10 represents the device number, and 50180 represents the feature dimension (including numerical and image data features). Use the Tensorly rank evaluation library, load the multi-dimensional power tensor dataset into Tensorly, and convert it into a tensor format supported by Tensorly. Use the rank evaluation tool provided by Tensorly, such as the tensorly.decomposition.parafac function, to perform a preliminary rank estimation on the tensor. This function evaluates the rank by calculating the CP decomposition (Canberra-Polyadic Decomposition) of the tensor. In actual operation, set the initial estimation range of the rank to [10, 50], and gradually adjust the rank value to evaluate the reconstruction error at different rank values. By calculating the reconstruction error (such as the mean square error MSE), select a suitable rank value to minimize the reconstruction error. For example, when the rank value is 20, the reconstruction error is 0.05, and when the rank value increases to 30, the reconstruction error only slightly decreases to 0.045. Therefore, select the rank value of 20 as the optimal rank value and record this rank value in the power tensor rank evaluation dataset.
[0042] Step S16: Perform tensor decomposition on the multi-dimensional power tensor dataset according to the power tensor rank evaluation dataset to obtain a power core feature tensor set;
[0043] Specifically, the CP decomposition method can be selected for tensor decomposition. Input the multi-dimensional power tensor dataset into the tensorly.decomposition.parafac function of Tensorly, and pass the rank value (20) evaluated in step S15 as a parameter. The function decomposes the multi-dimensional power tensor into a set of factor matrices and a core tensor. The shapes of the factor matrices are [100, 20] (time dimension), [10, 20] (device dimension), and [50180, 20] (feature dimension) respectively, and the shape of the core tensor is [20, 20, 20]. Through the above operations, these core feature tensors are combined into a power core feature tensor set.
[0044] Step S17: Perform feature selection and reconstruction on the power core feature tensor set to obtain a reconstructed power feature dataset.
[0045] Specifically, for the detailed implementation process of this embodiment, please refer to the sub-steps of step S17.
[0046] Through timestamp alignment, the present invention can ensure the consistency and integrity of data from different sensors and devices in the time dimension. Through anomaly detection and data cleaning, noise, outliers, and error records in the data can be effectively identified and removed. Through data type identification and standardization processing, different types of data (such as structured, semi-structured, and unstructured data) can be unified into a standardized framework. Through tensor construction and rank evaluation, the multi-dimensional power dataset can be transformed into a more compact tensor form and its inherent rank structure can be evaluated. Through the feature selection and reconstruction phase, the key features most valuable for power system analysis and decision-making can be screened out.
[0047] Preferably, step S17 includes the following steps:
[0048] Step S171: Evaluate the importance of features for the power core feature tensor set to obtain a power feature importance score set;
[0049] Specifically, it can be assumed that the shape of the power core feature tensor set is [20, 20, 20], where each element represents the weight or intensity of a feature. Flatten the power core feature tensor set into a two-dimensional matrix, with each row representing a sample (such as a data point in time or of a device), and each column representing a feature. Then, use RandomForestClassifier or RandomForestRegressor in scikit-learn (select the classification or regression model according to the type of data) to evaluate the importance of features. In this embodiment, it is assumed that the goal is to predict the fault state of a power device, and RandomForestClassifier is selected. The parameters of the model are set as follows: the number of trees is 100, the maximum number of features is automatic, and the maximum depth is 10. Input the flattened feature matrix and the corresponding labels (such as whether the device is faulty) into the model for training. After training is completed, call the feature_importances_ attribute of the model to obtain the importance scores of each feature. These score values range from 0 to 1, and the higher the value, the more important the feature. For example, assume that in the feature score results, the score of the first feature is 0.25, the score of the second feature is 0.15, and the score of the 20th feature is 0.01. Store the score values as the power feature importance score set.
[0050] Step S172: Perform dynamic feature threshold iteration based on the power feature importance score set to obtain the power feature selection threshold set;
[0051] Specifically, it can be assumed that the power feature importance score set contains the score values of 20 features, ranging from 0.01 to 0.25. Select an initial feature threshold, such as 0.1. Mark all features with scores lower than 0.1 as "unimportant" and remove them from the dataset. Retrain a simple machine learning model (such as a logistic regression model) using the remaining features. In this embodiment, use the LogisticRegression model in scikit-learn with the parameters set to the default values. Calculate the average accuracy and recall rate of the model through cross-validation. Gradually adjust the feature threshold, for example, increase it from 0.1 to 0.15, and repeat the above feature selection and model training process. Record the performance metrics of the model after each adjustment. For example, when the threshold is 0.1, the accuracy of the model is 0.85 and the recall rate is 0.80; when the threshold is increased to 0.15, the accuracy increases to 0.88 and the recall rate slightly drops to 0.78. Continue to adjust the threshold until a balance point is found such that the model performance metrics reach the optimal. Assume that the finally determined threshold is 0.18, at which time the accuracy of the model is 0.90 and the recall rate is 0.85. Record this threshold as the power feature selection threshold set.
[0052] Step S173: Select a threshold set according to the power characteristics to perform feature screening on the power core feature tensor set, and obtain a key power feature data set;
[0053] Specifically, it can be assumed that the shape of the power core feature tensor set is [20, 20, 20], and at the same time, the feature selection threshold is assumed to be 0.18 (the threshold determined in step S172). Flatten the power core feature tensor set into a two-dimensional matrix, where each row represents a sample (such as a data at a time point or a device), and each column represents a feature. According to the power feature importance score set (the result of step S171) and the feature selection threshold, select the features with scores higher than the threshold. For example, assume that the feature importance score set is an array of length 20, and the indices of the features with scores higher than 0.18 are [0, 1, 3, 5, 9]. Through boolean indexing, select these key feature columns from the flattened feature matrix to form a key power feature data set. Finally, the shape of the key power feature data set is [N, 5], where N is the number of samples and 5 is the number of selected key features.
[0054] Step S174: Design an autoencoder network structure for the key power feature data set to obtain an encoder structure parameter set;
[0055] Specifically, it can be assumed that the shape of the key power feature data set is [N, 5], where N is the number of samples and 5 is the number of key features. Use the TensorFlow framework and its API Keras to design the autoencoder network structure. In this embodiment, the structure of the encoder is designed as follows: the input layer contains 5 neurons (consistent with the number of key features), the hidden layer contains 3 neurons, and the activation function is selected as ReLU; the latent layer contains 2 neurons, and the activation function is selected as tanh. The structure of the decoder is designed as follows: the input layer contains 2 neurons (consistent with the latent layer), the hidden layer contains 3 neurons, and the activation function is selected as ReLU; the output layer contains 5 neurons (consistent with the input layer), and the activation function is selected as sigmoid. Define the input layer of the encoder, whose dimension is consistent with the number of features of the key power feature data set. Compress the input features into a more compact representation through a hidden layer. The number of neurons in the hidden layer is empirically selected as 3, and the activation function is selected as ReLU. Further compress the features into a two-dimensional space through the latent layer, and the activation function is selected as tanh because the tanh function can limit the output value within the range of [-1, 1]. The optimizer of the autoencoder is selected as Adam. The loss function is selected as mean squared error. By training the autoencoder model, the model will automatically learn how to compress the input features into a low-dimensional latent space and reconstruct features as close as possible to the original input from the low-dimensional space. Finally, the autoencoder model parameters obtained through training (including the weights and biases of the encoder and decoder) constitute the encoder structure parameter set.
[0056] Step S175: Perform encoding training on the key power feature dataset according to the encoder structure parameter set to obtain a latent power feature encoding dataset;
[0057] Specifically, it can be assumed that the shape of the key power feature dataset is [N, 5], where N is the number of samples and 5 is the number of key features. The encoder structure parameter set has been defined in step S174. Extract the encoder part from the autoencoder model to construct an independent encoder model. The input of the encoder model is the key power feature dataset, and the output is the feature encoding of the latent layer. During the training process, the Adam optimizer is selected, the learning rate is set to 0.001, and the mean squared error (MSE) is chosen as the loss function. The training process includes the following steps: Divide the key power feature dataset into a training set and a validation set in a ratio of 8:2. Use the training set to train the encoder, with a batch size of 32 and 100 training epochs set during the training process. In each training epoch, the model calculates the MSE loss between the input features and the reconstructed features and updates the weights of the encoder through backpropagation. At the same time, use the validation set to evaluate the performance of the model. After training is completed, the encoder model can encode the input key power feature dataset into low-dimensional latent features. Finally, store the encoded latent features as the latent power feature encoding dataset, whose shape is [N, 2], where 2 is the feature dimension of the latent layer.
[0058] Step S176: Perform decoder reconstruction on the latent power feature encoding dataset to obtain a reconstructed power feature dataset.
[0059] Specifically, it can be assumed that the shape of the latent power feature encoding dataset is [N, 2]. The decoder structure parameter set has been defined in step S174. Use the TensorFlow framework and its API Keras for decoder reconstruction. Extract the decoder part from the autoencoder model to construct an independent decoder model. The input of the decoder model is the latent power feature encoding dataset, and the output is the reconstructed power feature data. During the reconstruction process, the decoder gradually decodes the low-dimensional latent features back to the original feature space according to the trained weight and bias parameters. The specific operation is as follows: Input the latent power feature encoding dataset into the decoder model. The input layer of the decoder receives 2D latent features, performs a non-linear transformation through the ReLU activation function of the hidden layer, and finally reconstructs 5D power feature data through the sigmoid activation function of the output layer. The reconstructed power feature dataset after decoder reconstruction has the same shape [N, 5] as the original key power feature dataset.
[0060] The present invention can accurately screen out the most valuable key features for power system analysis through feature importance evaluation and dynamic feature threshold iteration. Through encoding and decoding processing, the deep semantic information of the features can be further extracted. The autoencoder maps the features to a low-dimensional latent space through nonlinear transformation and reconstructs the features during the decoding process, which not only enhances the robustness and representation ability of the features, but also captures the implicit structure and pattern in the power data. The feature selection threshold set determined by feature importance evaluation and dynamic threshold iteration can dynamically adjust the standard of feature screening to adapt it to different data distribution and analysis requirements. Through the encoding training and decoding reconstruction process, complex feature data can be converted into a more compact and representative latent feature encoding data set. The autoencoder can automatically remove the influence of noise and abnormal features by learning the intrinsic structure of the data during the training process. The reconstructed power feature data set after feature screening and autoencoder processing not only retains the core information of the original data, but also improves the quality of the features through dimensionality reduction and optimization.
[0061] Preferably, step S2 comprises the following steps:
[0062] Step S21: quantifying the domain feature distribution of the reconstructed power feature data set to obtain a multi-domain power feature statistical data set;
[0063] Specifically, the reconstructed power feature dataset contains features from multiple domains, each domain corresponding to different power equipment operating states or environmental conditions. The reconstructed power feature dataset is divided into multiple domains. For example, the dataset is divided into "power generation equipment domain", "power transmission equipment domain" and "power distribution equipment domain" according to the equipment type. For each domain, the statistical distribution of the features is calculated, including mean, variance, skewness and kurtosis statistics. For example, the mean of the feature is calculated using numpy.mean(), the variance is calculated using numpy.var(), the skewness is calculated using scipy.stats.skew(), and the kurtosis is calculated using scipy.stats.kurtosis(). In the specific operation, it is assumed that the shape of the reconstructed power feature dataset is [N,D], where N is the number of samples and D is the feature dimension. For each domain, the corresponding sub-dataset is extracted, and the statistics of each feature are calculated. For example, for the power generation equipment domain, the mean of feature 1 is calculated to be 0.5, the variance is 0.1, the skewness is 0.2, and the kurtosis is 3.0; for the power transmission equipment domain, the mean of feature 1 is calculated to be 0.6, the variance is 0.15, the skewness is 0.1, and the kurtosis is 2.8. The statistics are summarized into a multi-domain power feature statistical data set.
[0064] Step S22: performing domain adaptive network mapping on the reconstructed power feature data set according to the multi-domain power feature statistical data set to obtain a power domain mapping network parameter set;
[0065] Specifically, the structure of the domain adaptation network can be designed based on the multi-domain power feature statistical dataset. Assuming that the feature dimension of the reconstructed power feature dataset is D, the input layer dimension of the domain adaptation network is also D. The network structure includes an input layer, several hidden layers (such as two fully connected layers), and an output layer. The number of neurons in the hidden layers can be set to D / 2 and D / 4, and the activation function is selected as ReLU. The dimension of the output layer is the same as that of the input layer (i.e., D), and the linear activation function is selected as the activation function to output the mapped features. Next, a discriminator network is constructed, and its goal is to distinguish which domain the input features come from. The input of the discriminator network is the output of the domain adaptation network, and the output is the probability distribution of the domain labels. Assuming there are two domains (Domain A and Domain B) in the multi-domain power feature statistical dataset, the output layer of the discriminator network will have two neurons, corresponding to the probability distributions of the two domains respectively. During the training process, an adversarial training mechanism is adopted. Specifically, by minimizing the loss function of the discriminator (such as cross-entropy loss), while maximizing the confusion loss of the domain adaptation network (that is, making the discriminator unable to distinguish the domains). During training, the learning rate is set to 0.001, the optimizer is selected as Adam, and the number of training epochs is 100. In each training epoch, first update the parameters of the discriminator network; then update the parameters of the domain adaptation network so that the features it outputs are more difficult to be distinguished by the discriminator. For example, during the training process, assuming the initial discriminator loss is 0.6, after 100 training epochs, the discriminator loss is reduced to 0.3, indicating that the domain adaptation network has successfully mapped the features of different domains into a similar distribution space. Finally, the parameter set of the domain adaptation network obtained through training includes the weights and biases of the network.
[0066] Step S23: Perform intra-domain sample clustering on the power domain mapping network parameter set to obtain a hierarchical intra-domain clustering feature dataset;
[0067] Specifically, the hierarchical clustering method can be used for intra-domain sample clustering, and the specific tool is the scikit-learn library in Python. Extract the power domain mapping network parameter set from the domain adaptation network. Take each domain as a unit and cluster the mapped features. For example, assume the dataset contains three domains: the power generation equipment domain, the power transmission equipment domain, and the power distribution equipment domain. For the power generation equipment domain, use the AgglomerativeClustering class in scikit-learn for hierarchical clustering, set the number of clusters to K = 6, the distance metric method to Euclidean distance, and select the Ward criterion as the connection method for clustering. Through hierarchical clustering, the samples in the power generation equipment domain are divided into 6 sub-clusters, and each sub-cluster represents a set of similar features within the domain. Repeat the above process to perform hierarchical clustering on the power transmission equipment domain and the power distribution equipment domain respectively to obtain the intra-domain clustering results for each domain. Finally, a hierarchical intra-domain clustering feature dataset is formed. Each sample in the dataset not only contains its original features but also the corresponding clustering labels.
[0068] Step S24: Calculate the inter-domain distance based on the intra-hierarchy domain clustering feature dataset to obtain an inter-domain metric distance dataset;
[0069] Specifically, the scipy library and scikit-learn library in Python can be used to calculate the inter-domain distance. Extract the clustering centers of each domain in the intra-hierarchy domain clustering feature dataset. For each domain, use the KMeans algorithm in scikit-learn to recalculate the clustering center, and use the clustering center as the representative feature of the domain. For example, for the 6 sub-clusters of the power generation equipment domain, calculate the clustering center of each sub-cluster to obtain 6 clustering center vectors. Use the scipy.spatial.distance.cdist function to calculate the Euclidean distance between the clustering centers of different domains. Assume that there are 6 clustering centers in the power generation equipment domain, 5 clustering centers in the power transmission equipment domain, and 4 clustering centers in the power distribution equipment domain. Through calculation, an inter-domain distance matrix is obtained, with shapes of [5,4] (distance matrix between the power generation equipment domain and the power transmission equipment domain), [5,3] (distance matrix between the power generation equipment domain and the power distribution equipment domain), and [4,3] (distance matrix between the power transmission equipment domain and the power distribution equipment domain). Finally, these distance matrices are summarized into an inter-domain metric distance dataset.
[0070] Step S25: Iteratively adjust the gradient of the power domain mapping network parameter set according to the inter-domain metric distance dataset to obtain an adaptive power domain mapping parameter set;
[0071] Specifically, the deep learning framework TensorFlow and its optimizer tool can be used for iterative gradient adjustment. Define a loss function to measure the difference in inter-domain distance. For example, select the MMD (Maximum Mean Discrepancy) loss function, which can effectively measure the difference in feature distributions between different domains. During the iteration process, use the Adam optimizer with a learning rate set to 0.0005. In the specific operation, use the inter-domain metric distance dataset as the input, calculate the loss value of the inter-domain distance (such as the MMD loss), and use backpropagation to update the weights and biases of the domain mapping network. In each iteration, the optimizer adjusts the network parameters according to the gradient value of the loss function to reduce the inter-domain difference. For example, assume that the initial inter-domain distance loss is 0.5, and after 100 iterations, the loss value is reduced to 0.1, indicating that the parameters of the domain mapping network can better align the feature distributions of different domains. Finally, through multiple iterations of optimization, an adaptive power domain mapping parameter set is obtained.
[0072] Step S26: Perform feature projection on the reconstructed power feature dataset according to the adaptive power domain mapping parameter set to obtain an inter-domain power feature mapping dataset;
[0073] Specifically, the reconstructed power feature dataset can be input into the optimized domain mapping network. The structure of the domain mapping network is the same as that defined in step S22, but the parameters have been updated to the adaptive power domain mapping parameter set. Through the forward propagation of the network, the input features are mapped to a new feature space. The feature vector of each sample is input into the domain mapping network, and after the non-linear transformation of the hidden layer and the linear transformation of the output layer, the mapped feature vector is obtained. For example, assume that the dimension of the input feature is D = 10. After passing through the domain mapping network, the output feature dimension is compressed to D 1 = 5. Through the above operations, the features of all samples are projected, and finally the cross-domain power feature mapping dataset is obtained. The shape of this dataset is [N, D 1 , where D 1 is the mapped feature dimension.
[0074] Step S27: Eliminate the domain differences in the cross-domain power feature mapping dataset to obtain a domain-aligned power feature dataset, and calculate the attention weights for the domain-aligned power feature dataset to obtain a cross-domain fusion power feature dataset.
[0075] Specifically, for the detailed implementation process of this embodiment, please refer to the sub-steps of step S27.
[0076] Through domain feature distribution quantization and domain adaptation network mapping, the present invention can align power feature data from different domains to the same feature space, significantly reducing the differences between domains. Through hierarchical clustering and inter-domain distance calculation, the internal structure and distribution characteristics of the data can be deeply explored. Through iterative gradient adjustment, the parameters of the domain mapping network can be dynamically adjusted to adapt to the differences between different domains. Through feature projection and domain difference elimination, the inconsistencies in cross-domain data can be further eliminated. Through attention weight calculation, the key features that are most valuable for power system analysis can be automatically identified and highlighted.
[0077] Preferably, step S27 includes the following steps:
[0078] Step S271: Construct a discriminator network for the cross-domain power feature mapping dataset to obtain an adversarial discriminator network;
[0079] Specifically, the deep learning framework TensorFlow and its API Keras can be used to construct the discriminator network. The structure design of the discriminator network is as follows: First, define the input layer of the discriminator network, whose dimension is the same as the feature dimension. The input features are non-linearly transformed through two hidden layers, and finally, the probability that the sample belongs to a certain domain is output through the Sigmoid activation function. For example, assume that the input feature dimension is D 1If = 10, the structure of the discriminator network can be expressed as follows: Input layer: 10 neurons; Hidden layer 1: 64 neurons, with the activation function ReLU; Hidden layer 2: 32 neurons, with the activation function ReLU; Output layer: 1 neuron, with the activation function Sigmoid; Through the above steps, a network model capable of performing domain discrimination on input features is constructed.
[0080] Step S272: Based on the adversarial discriminator network, perform adversarial training on the cross-domain power feature mapping dataset to obtain a domain adversarial loss power dataset;
[0081] Specifically, the cross-domain power feature mapping dataset can be input into the adversarial discriminator network to calculate the output probability of the discriminator. At the same time, an adversarial loss function is defined. For example, the binary cross-entropy loss function is selected, which can measure the difference between the probability output by the discriminator and the true domain label. During the adversarial training process, the parameters of both the discriminator network and the domain mapping network are updated simultaneously. The specific operations are as follows: Update the discriminator network: By minimizing the loss function of the discriminator (i.e., maximizing the discrimination ability of the discriminator), use the Adam optimizer to update the parameters of the discriminator network. For example, set the learning rate to 0.001 and the number of training epochs to 50. Update the domain mapping network: By maximizing the loss function of the discriminator (i.e., minimizing the discrimination ability of the discriminator), use the same optimizer to update the parameters of the domain mapping network. For example, during the training process, assume that the initial discriminator loss is 0.6, and after adversarial training, the discriminator loss gradually decreases to 0.3, indicating that the domain mapping network has successfully generated indistinguishable features. Finally, the domain adversarial loss power dataset obtained through adversarial training records the loss value of each training.
[0082] Step S273: According to the domain adversarial loss power dataset, perform feature adjustment on the cross-domain power feature mapping dataset to obtain a domain-aligned power feature dataset;
[0083] Specifically, the domain adversarial loss power dataset can be analyzed to extract the loss value at each iteration. According to the change trend of the loss value, the feature dimension to be adjusted is determined. For example, if the loss value of a certain feature dimension is high, it indicates that there are significant differences in this feature between different domains and needs to be adjusted. In the specific operation, a feature adjustment module is defined, with the input being the cross-domain power feature mapping dataset and the domain adversarial loss dataset. The goal of the feature adjustment module is to minimize the domain adversarial loss while maintaining the representational ability of the features. The gradient descent method is used to adjust the features. For example, for the feature dimension with a high loss value, by calculating the gradient of the loss function, the feature value is gradually adjusted to make it more consistent between different domains. Suppose the initial domain adversarial loss is 0.3, and after feature adjustment, the loss is reduced to 0.1, indicating that the feature adjustment is effective. Finally, the adjusted feature dataset is called the domain-aligned power feature dataset, and its shape is still [N, D 1 .
[0084] Step S274: Calculate the attention weights for the domain-aligned power feature dataset to obtain the multi-modal power feature weight dataset;
[0085] Specifically, the deep learning framework TensorFlow and its API Keras can be used to calculate the attention weights. Design an attention mechanism module with the input being the domain-aligned power feature dataset. The goal of the attention mechanism is to assign a weight to each feature. In the specific operation, an attention layer is defined, and its structure includes a fully connected layer and a Softmax activation function. The input dimension of the fully connected layer is D 1 , and the output dimension is also D 1 , and the ReLU activation function is selected. The Softmax activation function normalizes the output value to a weight within the range of [0, 1]. For example, suppose the input feature dimension is D 1 = 10, the structure of the attention layer can be expressed as: fully connected layer: input dimension is 10, output dimension is 10, activation function is ReLU; Softmax activation function: normalizes the output value to a weight; Through the above steps, a weight value is calculated for each feature. Finally, the calculated weight values are combined with the domain-aligned power feature dataset to obtain the multi-modal power feature weight dataset.
[0086] Step S275: Construct the weight optimization objective based on the multi-modal power feature weight dataset to obtain the power feature weight optimization objective dataset, and iteratively update the parameters of the multi-modal power feature weight dataset according to the power feature weight optimization objective dataset to obtain the optimized power feature weight dataset;
[0087] Specifically, the deep learning framework TensorFlow can be used to construct the weight optimization objective. Define a weight optimization objective function, the goal of which is to minimize the variance of the feature weights while maximizing the correlation between the feature weights and the target variable (such as the power system fault prediction label). The weight optimization objective function can be expressed as: Loss = α·Var(w) - β·Corr(w, y); where w is the feature weight, y is the target variable, and α and β are hyperparameters, which are set to 0.1 and 0.9 respectively. Use the Adam optimizer to iteratively update the feature weights. Input the multi-modal power feature weight dataset into the optimization objective function, calculate the loss value, and update the weight parameters through backpropagation. For example, the initial distribution of the feature weights is relatively dispersed, with a variance of 0.2. After 100 iterations, the variance is reduced to 0.05, and at the same time, the correlation between the weights and the target variable is increased from 0.6 to 0.8. Finally, the power feature weight optimization objective dataset is obtained, and the updated weight values are stored as the optimized power feature weight dataset.
[0088] Step S276: Perform weighted fusion on the domain-aligned power feature dataset according to the optimized power feature weight dataset to obtain the initial fused power feature dataset, and perform feature consistency verification on the initial fused power feature dataset to obtain the inter-domain feature consistency evaluation dataset;
[0089] Specifically, the domain-aligned power feature dataset can be multiplied element-wise by the optimized power feature weight dataset to obtain the weighted feature matrix. For example, assume that the shape of the feature dataset is [1000, 10], and the shape of the weight dataset is
[10] . By element-wise multiplication, the initial fused power feature dataset is obtained, and its shape is still [1000, 10]. Cosine similarity is used as the consistency evaluation metric for performing feature consistency verification on the initial fused power feature dataset. In the specific operation, calculate the cosine similarity between samples from different domains to generate the inter-domain feature consistency evaluation dataset. For example, assume that the dataset contains samples from two domains. By calculating the cosine similarity, the shape of the consistency evaluation dataset is [500, 500], where each element represents the similarity between a sample and other samples. Finally, through weighted fusion and consistency verification, the initial fused power feature dataset and the inter-domain feature consistency evaluation dataset are obtained.
[0090] Step S277: Perform consistency calibration on the initial fused power feature dataset according to the inter-domain feature consistency evaluation dataset to obtain the cross-domain fused power feature dataset.
[0091] Specifically, the inter-domain feature consistency evaluation dataset can be analyzed to calculate the consistency scores on each feature dimension. The specific method is to evaluate the similarity of each feature dimension between different domains by calculating the mean and standard deviation of the cosine similarity on each feature dimension. If the consistency score of a certain feature dimension is lower than a preset threshold (e.g., 0.8), it is considered that there is an inconsistency in this feature dimension and calibration is required. For the inconsistent feature dimensions, the following methods are used for calibration: Normalization processing: Normalize the inconsistent feature dimensions to limit their value range between [0, 1]. Weighted average calibration: According to the inter-domain consistency scores, perform weighted average calibration on the inconsistent feature dimensions. Specifically, for the values of a certain feature dimension between different domains, the weights are dynamically adjusted according to its consistency score. For example, if the consistency score of a certain feature dimension between domain A and domain B is 0.7, which is lower than the threshold of 0.8, then weighted average processing is performed on this feature dimension. Through the above steps, the finally obtained feature dataset after consistency calibration is called the cross-domain fusion power feature dataset.
[0092] By constructing an adversarial discriminator network and conducting adversarial training, the present invention can effectively identify and eliminate the inter-domain differences in cross-domain feature data. Through weight calculation and optimization, it can automatically identify and highlight the key features that are most valuable for power system analysis. By constructing a weight optimization target and iteratively updating parameters, it can dynamically adjust the feature weights to ensure that the fusion process always develops in the optimal direction. By performing feature consistency verification and calibration on the initial fusion dataset, potential inconsistencies in the fusion process can be further eliminated.
[0093] Preferably, step S3 includes the following steps:
[0094] Step S31: Collect the topological relationship of distributed power nodes to obtain a node connection relationship dataset;
[0095] Specifically, the target power grid includes multiple distributed nodes, such as substations, power generation stations, and customer-side devices, which are interconnected by transmission lines. Real-time operation data of the power grid, including voltage, current, and power information of the nodes, is obtained through the automated monitoring system of the power system (such as the SCADA system). The connection relationships between the nodes are modeled using network topology analysis tools (such as the power system analysis toolbox of MATLAB or the NetworkX library of Python). The identifiers of the nodes (such as node IDs) and connection information (such as the start and end points of the lines) are extracted from the SCADA system. For example, if there is a transmission line between node A and node B, it is recorded as a connection relationship (A, B), and the attributes of the line, such as the line length of 10.5 kilometers and the impedance of 0.25 ohms, are recorded. Through the above steps, a node connection relationship dataset is constructed. The dataset is stored in the form of a matrix or a table, and each row represents a connection relationship, including the start node ID, the end node ID, and the attributes of the connecting line.
[0096] Step S32: Based on the node connection relationship dataset, perform link state evaluation on the distributed power nodes to obtain a node link quality dataset;
[0097] Specifically, power system analysis software (such as SimPowerSystems of MATLAB) and data processing tools of Python can be used for link state evaluation. The attributes of each line are extracted from the node connection relationship dataset. Combining the operation data of the power system (such as current, voltage, and power), calculate the loss, voltage drop, and transmission efficiency of each line. For each line, calculate its transmission loss and voltage drop. For example, assume that the impedance of line (A, B) is 0.25 ohms and the current passing through this line is 100 amperes, then the line loss is I 2 ×R = (100) 2 ×0.25 = 2500 watts. At the same time, calculate the voltage drop of the line as I×R = 100×0.25 = 25 volts. According to the calculation results, evaluate the link quality of each line. For example, define the link quality index as the transmission efficiency divided by the sum of the loss and the voltage drop. Assume that the input power of line (A, B) is 100 kW and the output power is 95 kW, then the transmission efficiency is 95%. The link quality can be calculated through the above index. Finally, store the link quality index of each line as the node link quality dataset. The dataset contains the start node ID, end node ID, loss, voltage drop, transmission efficiency, and link quality information of each line.
[0098] Step S33: Calculate the connection weights of the distributed power nodes according to the node link quality dataset to obtain a node connection weight dataset, and construct a dynamic adjacency matrix according to the node connection weight dataset to obtain the distributed node topology data;
[0099] Specifically, the link quality metrics for each line can be extracted from the node link quality dataset. The higher the link quality, the higher the transmission efficiency and the lower the loss of the line, so a higher weight is assigned. For example, the connection weight is defined as the normalized value of the link quality, ranging from [0, 1]. Suppose the link quality of line (A, B) is 0.85, the link quality of line (B, C) is 0.88, and the link quality of line (C, D) is 0.82. Then the normalized weights are 0.85, 0.88, and 0.82 respectively. A dynamic adjacency matrix is constructed based on the node connection weight dataset. The adjacency matrix is a two-dimensional matrix, whose rows and columns represent nodes respectively, and the elements in the matrix represent the connection weights between nodes. The construction process of the dynamic adjacency matrix is as follows: For node A, it is connected to node B with a weight of 0.85. Therefore, in the adjacency matrix, the position from A to B is recorded as 0.85; similarly, the position from B to A is also recorded as 0.85. Node B is connected to node C with a weight of 0.88. Therefore, the position from B to C is recorded as 0.88, and the position from C to B is recorded as 0.88. Node C is connected to node D with a weight of 0.82. Therefore, the position from C to D is recorded as 0.82, and the position from D to C is recorded as 0.82. For nodes that are not directly connected, the corresponding positions in the adjacency matrix are recorded as 0. Finally, through the calculation of the connection weights and the construction of the dynamic adjacency matrix, the distributed node topology data is obtained.
[0100] Step S34: Perform feature extraction layer planning on the distributed power nodes according to the distributed node topology data to obtain the power node graph convolution layer parameter set;
[0101] Specifically, the deep learning framework TensorFlow and its API Keras can be used, combined with the graph neural network (GNN) toolkit for feature extraction layer planning. Define the structural parameters of the graph convolution layer, including the input feature dimension, the hidden layer dimension, and the output feature dimension. For example, suppose the input feature dimension is D in = 10 (such as information about the voltage and current of the node), the hidden layer dimension is D hiddden = 16, and the output feature dimension is D out = 8. Construct the graph convolution layer. The core operation of the graph convolution layer is to aggregate the feature information of adjacent nodes and perform a linear transformation through the weight matrix. In TensorFlow, tf.keras.layers.Dense is used to implement the transformation of the weight matrix, and feature aggregation is combined with the adjacency matrix. For example, for the graph convolution operation of the hidden layer, it can be expressed as: H (L+1) = σ(AH (L) W (L) ); where H (L) is the node feature matrix of the L-th layer, and W (L)is the weight matrix, A is the adjacency matrix, and σ is the activation function (such as ReLU). Through the above steps, the parameter set of the power node graph convolutional layer is defined, including the input feature dimension, the hidden layer dimension, the output feature dimension, and the activation function.
[0102] Step S35: Construct a heterogeneous message passing mechanism according to the power node graph convolutional layer parameter set to obtain a power node message passing rule set;
[0103] Specifically, the deep learning framework PyTorch and its graph neural network library PyTorch Geometric can be used to construct the heterogeneous message passing mechanism. According to the power node graph convolutional layer parameter set, the input and output feature dimensions of the message passing are defined. For example, the input feature dimension is D in = 10, the hidden layer dimension is D hiddden = 16, and the output feature dimension is D out = 8. In PyTorch Geometric, the MessagePassing module is used to implement a custom message passing mechanism. For example, define a message passing rule where each node aggregates the features of its adjacent nodes and performs a linear transformation through a learnable weight matrix. The specific rules are as follows: Message function: For each edge (i, j), calculate the message m ij from node i to node j: m ij = W msg · h i ; where h i is the feature of node i, and W msg is the message weight matrix. Aggregation function: Aggregate all the messages of node j's adjacent nodes: m j = AGGREGATE({m ij | i ∈ N(j)}); where N(j) is the set of adjacent nodes of node j, and the aggregation method can be selected as summation, average, or maximum. Update function: Update the feature of node j according to the aggregated message: h j ` = W update (h j + m j ); where W update is the update weight matrix. By defining these message passing rules, a power node message passing rule set is constructed.
[0104] Step S36: Construct a graph neural architecture according to the power node message passing rule set to obtain power node graph network structure data;
[0105] Specifically, the structure of the graph convolutional layer can be defined according to the power node message passing rule set. In PyTorchGeometric, the GCNConv module is used to implement graph convolution operations, or a graph convolutional layer based on MessagePassing can be customized. For example, a two-layer graph convolutional network is defined, where the input feature dimension of the first layer is D in = 10, the hidden layer dimension is D hiddden = 16, and the output feature dimension of the second layer is D out = 8. The specific operations are as follows: The first-layer graph convolutional layer: The input feature dimension is 10, and the output feature dimension is 16. The ReLU activation function is used to perform a non-linear transformation on the output features. The second-layer graph convolutional layer: The input feature dimension is 16, and the output feature dimension is 8. The output features of this layer will be used as the final node feature representation of the graph network. By stacking these two layers of graph convolutional layers, the power node graph network structure is constructed.
[0106] Step S37: Based on the power node graph network structure data, perform distributed model training on the cross-domain fusion power feature dataset to obtain a local power model parameter set, and perform federated averaging calculation on the local power model parameter set to obtain a globally optimized model parameter set.
[0107] Specifically, for the detailed implementation process of this embodiment, please refer to the sub-steps of step S37.
[0108] Through topology relationship acquisition and link state evaluation, the present invention can characterize the node connection relationship and link quality of the power network. Through connection weight calculation and dynamically constructing an adjacency matrix, the connection strength between nodes can be adaptively adjusted according to the actual quality of the link. Through feature extraction layer planning and the construction of a heterogeneous message passing mechanism, node features can be efficiently extracted and information can be transmitted between nodes. By adopting distributed model training combined with federated averaging calculation, model training can be performed in parallel on distributed power nodes, and global optimization can be achieved through the federated learning mechanism. Through the construction of a graph neural network architecture and distributed training, the model can better adapt to the complex topology structure and dynamic changes of the power network.
[0109] Preferably, step S37 includes the following steps:
[0110] Step S371: According to the power node graph network structure data, perform distributed training batch division on the cross-domain fusion power feature dataset to obtain a distributed training batch parameter set;
[0111] Specifically, based on the power node diagram network structure data, determine the data volume and computing resource allocation for each node. For example, assume that the entire dataset contains 100,000 samples, distributed across 10 distributed nodes, and the data volume of each node is approximately 10,000 samples. Further divide the data of each node into multiple batches. Assume that the size of each batch is 1,000 samples, then each node will be divided into 10 batches. Define a distributed training batch parameter set, including the size of each batch, the number of batches, and the node allocation. For example, for node 1, its training batch parameter set can be expressed as: node ID is 1, batch size is 1,000, and the total number of batches is 10. Similarly, generate corresponding batch parameter sets for other nodes.
[0112] Step S372: Perform asynchronous parallel scheduling on the cross-domain fusion power feature dataset according to the distributed training batch parameter set to obtain a parallel training task allocation dataset;
[0113] Specifically, according to the distributed training batch parameter set, training tasks can be allocated to each node. For example, assume that each node has 10 batches, and each batch contains 1,000 samples. The training tasks will be allocated to each node in the order of batches. Start asynchronous parallel scheduling. In the asynchronous mode, each node independently executes its allocated training task without waiting for other nodes to complete. For example, node 1 starts processing the first batch, and node 2 simultaneously starts processing the first batch, and so on. After each node completes the training of a batch, it immediately sends the result to the master node for aggregation, and then continues to process the next batch. Finally, a parallel training task allocation dataset is generated, recording the task allocation and completion status of each node.
[0114] Step S373: Perform distributed optimizer configuration according to the parallel training task allocation dataset to obtain a distributed optimizer configuration parameter set;
[0115] Specifically, an appropriate optimizer can be selected according to the training task volume and computing resources of each node. For example, for the power feature training task, the Adam optimizer is selected. Configure the parameters of the optimizer for each node, including the learning rate and weight decay. Assume the initial learning rate is 0.001 and the weight decay is 0.0001. In a distributed environment, the synchronization mechanism of the optimizer also needs to be configured. For example, use the DistributedDataParallel module of PyTorch, set gradient_as_bucket_view = True to optimize the gradient synchronization efficiency, and ensure that the gradients of all nodes are synchronously updated after each iteration through the all_reduce operation of the torch.distributed module. Finally, the generated set of distributed optimizer configuration parameters will include the optimizer type, learning rate, weight decay, and synchronization mechanism parameters of each node.
[0116] Step S374: Design a learning rate scheduling strategy for the set of distributed optimizer configuration parameters to obtain a power feature scheduling learning rate parameter set;
[0117] Specifically, a learning rate scheduler can be selected according to the characteristics of the training task and the characteristics of the distributed environment. For example, select the torch.optim.lr_scheduler.CosineAnnealingLR scheduler. Assume that the total number of iterations of the training task is 1000 times, the period of the learning rate scheduler is set to 500 times, and the minimum learning rate is set to 1 e-6 . Configure the parameters of the learning rate scheduler. For example, the initial learning rate is 0.001, the period is 500 times, and the minimum learning rate is 1 e-6 . During the training process, the learning rate will be adjusted according to the cosine annealing strategy. When the number of iterations reaches 500 times, the learning rate will gradually decrease to the minimum value of 1 e -6 , and then start a new cycle. Through the above steps, finally, the generated power feature scheduling learning rate parameter set will include the type of learning rate scheduler, initial learning rate, cycle length, and minimum learning rate parameters.
[0118] Step S375: Perform model distributed gradient calculation on the distributed power nodes according to the power feature scheduling learning rate parameter set to obtain a power node gradient update data set;
[0119] Specifically, the deep learning framework PyTorch and its distributed training module torch.distributed can be used for distributed gradient calculation of the model. According to the power feature scheduling learning rate parameter set, the learning rate for the current iteration is configured for each node. For example, in the 100th iteration of training, the learning rate scheduler adjusts the learning rate to 0.0005 according to the preset cosine annealing strategy. Each node independently performs forward propagation and backward propagation calculations on its assigned training batches. During forward propagation, the model generates predicted outputs based on the input power feature data; during backward propagation, the loss function (such as mean squared error or cross-entropy loss) between the predicted output and the true label is calculated, and the gradients of the model parameters are calculated according to the chain rule. For example, assume that the current batch of data of a certain node contains 1000 samples, and the loss of the model on this batch is 0.35. The gradients calculated through backward propagation will be used to update the model parameters. After each node completes gradient calculation, the calculated gradients are stored as the power node gradient update dataset, which contains the gradient information of each node on the current batch.
[0120] Step S376: Asynchronously aggregate the gradients in the power node gradient update dataset to obtain a global gradient aggregation dataset, and update the model parameters according to the global gradient aggregation dataset to obtain a local model update parameter set;
[0121] Specifically, a global gradient aggregation strategy can be defined using the torch.distributed module of PyTorch. In asynchronous mode, after each node completes local gradient calculation, it immediately sends the gradients to the master node (or parameter server). The master node is responsible for collecting the gradients from each node and performing the aggregation operation. For example, using the average aggregation strategy, the gradients of all nodes are averaged element-wise to obtain the global gradient. This global gradient aggregation dataset will be used to update the model parameters. Update the model parameters according to the global gradient aggregation dataset. The specific operation is to apply the global gradient to the current parameters of the model, and the parameter update is completed through an optimizer (such as Adam). For example, assume that the current model parameters are params, and the optimizer updates the model parameters according to the global gradient global_grad and the current learning rate (such as 0.0005) to obtain the new model parameters new_params. Finally, the updated model parameters new_params are broadcast back to each node as the local model update parameter set.
[0122] Step S377: Perform parameter weighted averaging on the local power model parameter set to obtain a federated average weight dataset, and perform federated aggregation on the local power model parameter set based on the federated average weight dataset to obtain a globally optimized model parameter set.
[0123] Specifically, the weight of each node can be determined. The weights can be assigned according to the data volume of the nodes or the model performance. For example, assume there are 4 nodes in a distributed system, and the data volumes of each node are 1000, 1200, 1100, and 900 samples respectively. The weight of each node can be set as the proportion of its data volume to the total data volume. The total data volume is 4200 samples. Therefore, the weight of node 1 is 1000 / 4200, the weight of node 2 is 1200 / 4200, and so on. The model parameters of each node are weighted averaged. The specific operation is to multiply the model parameters of each node by its corresponding weight, and then add up the weighted parameters of all nodes. Through the above steps, a federated average weight data set is obtained, which reflects the contribution of each node to the global model. Finally, the weighted average parameters are used as the new parameters of the global model, and these parameters are broadcast back to each node. Each node updates its local model using the global optimized model parameter set to prepare for the next round of training.
[0124] Through distributed training batch division and asynchronous parallel scheduling, the present invention can make full use of the computing resources of distributed power nodes to achieve efficient parallel training of cross-domain fusion power feature data. Through distributed optimizer configuration, the most suitable optimization strategy can be assigned to each distributed node. Through distributed gradient calculation, it can be ensured that the models of each node are updated at the optimal learning rate during the training process. Through asynchronous gradient aggregation and federated average mechanism, the local model parameters on distributed nodes can be effectively integrated.
[0125] Preferably, step S4 includes the following steps:
[0126] Step S41: Construct a neural network for the global optimized model parameter set to obtain a multi-modal teacher network, and perform power cross-modal feature enhancement on the multi-modal teacher network to obtain a power multi-modal teacher model;
[0127] Specifically, the deep learning framework TensorFlow and its API Keras can be used to build a multi-modal teacher network. First, define the input layer of the network to adapt to different modalities of power data, such as voltage, current, power, and image data. For structured data (such as voltage and current), the input layer can be a fully connected layer; for image data, the input layer can be a convolutional layer. Next, build a multi-modal fusion layer to fuse the features of different modalities. For example, use a fully connected layer with 128 neurons and select the ReLU activation function to map the features of different modalities to a shared feature space. Introduce an attention mechanism module to achieve cross-modal feature enhancement of power. For example, use a self-attention layer to assign weights by calculating the similarity between features and highlight important feature information. Through the above steps, after neural network construction and cross-modal feature enhancement, a power multi-modal teacher model is obtained.
[0128] Step S42: Perform feature-level decomposition on the power multi-modal teacher model to obtain a set of teacher semantic knowledge features;
[0129] Specifically, the structure of the power multi-modal teacher model can be analyzed to determine the feature extraction layers at different levels. For example, assume that the teacher model contains multiple convolutional layers and fully connected layers. The convolutional layers extract low-level features (such as edges and textures), and the fully connected layers extract high-level semantic features (such as device status and fault modes). Select a specific feature extraction layer as the benchmark for decomposition. For example, select the output of the penultimate fully connected layer as the representative of high-level semantic features. By inserting a hook function into the model, extract the feature output of this layer. Assume that the output dimension of this layer is 64, and the extracted features can be represented as a 64-dimensional vector, with each dimension corresponding to a semantic feature. Through the above steps, a set of teacher semantic knowledge features is obtained.
[0130] Step S43: Perform feature selection on the set of teacher semantic knowledge features to obtain a set of key power representation knowledge features, and perform neural architecture search and construction on the set of key power representation knowledge features to obtain a lightweight power student model;
[0131] Specifically, a method based on information gain can be used to evaluate the importance of each feature in the teacher semantic knowledge feature set. For example, calculate the mutual information between each feature and the target variable (such as the state of power equipment), and select the top 20% of the features with the highest information gain as the key power representation knowledge feature set. Suppose the dimension of the teacher semantic knowledge feature set is 100. After evaluation by information gain, 20 features with the highest information gain are selected as the key features. Use a neural architecture search (NAS) method based on reinforcement learning, such as ENAS (Efficient Neural Architecture Search). During the search process, define the search space, including different types of neural network layers (such as convolutional layers, fully connected layers) and the connection methods between layers. For example, the search space includes hyperparameters such as the number of filters in the convolutional layer (such as 32, 64, 128), the size of the convolutional kernel (such as 3×3, 5×5), and whether to use batch normalization. Through the ENAS algorithm, automatically search for the optimal network structure. For example, after multiple rounds of search, a network structure is found that includes two convolutional layers (each with 64 filters and a convolutional kernel size of 3×3) and two fully connected layers (each with 128 neurons). Finally, through feature selection and neural architecture search, a lightweight power student model is constructed.
[0132] Step S44: Align and extract the features of the lightweight power student model and the power multimodal teacher model to obtain a cross-modal feature alignment loss data set;
[0133] Specifically, the goal of feature alignment can be defined. Suppose the output dimension of the feature extraction layer of the teacher model is 64, and the output dimension of the corresponding layer of the student model is 32. Use a fully connected layer to map the features of the student model to the same dimension as the teacher model. For example, add a fully connected layer to map the 32-dimensional features to 64 dimensions. Calculate the feature alignment loss. Select the mean squared error as the loss function to measure the difference between the features of the student model and the teacher model. For example, suppose the features of the teacher model are teacher_features and the aligned features of the student model are student_features, then the feature alignment loss can be expressed as: alignment_loss = MSE(teacher_features, student_features); By calculating the feature alignment loss of each sample, a cross-modal feature alignment loss data set is generated, which records the errors of the student model in the process of learning the features of the teacher model.
[0134] Step S45: Design a knowledge distillation strategy based on the cross-modal feature alignment loss dataset to obtain a multi-level power knowledge distillation strategy parameter set, and optimize the power lightweight student model according to the multi-level power knowledge distillation strategy parameter set to obtain a model distillation optimization parameter set;
[0135] Specifically, the deep learning framework TensorFlow and its API Keras can be used to design the knowledge distillation strategy. Define the distillation loss function, combining the feature alignment loss and the soft target (softmax output) of the teacher model. For example, the distillation loss can be expressed as a weighted sum of the feature alignment loss and the soft target loss, with weights of 0.5 and 0.5 respectively. Assume that the soft target output of the teacher model is teacher_softmax and the output of the student model is student_output, then the soft target loss can be measured using the Kullback-Leibler Divergence:
[0136] distillation_loss = θ·MSE(teacher_features, student_features)+(1 - θ)·KL(teacher_softmax, student_output);
[0137] where θ is the weight of the feature alignment loss, set to 0.5. Adjust the parameters in the distillation strategy according to the cross-modal feature alignment loss dataset. For example, if the feature alignment loss of certain modalities is high, the feature alignment weight of that modality can be increased. Through multiple rounds of experiments, determine the optimal distillation strategy parameter set, including the feature alignment weight, soft target weight, and learning rate. Finally, optimize the power lightweight student model according to the multi-level power knowledge distillation strategy parameter set to obtain a model distillation optimization parameter set.
[0138] Step S46: Compress the model distillation optimization parameter set to obtain a lightweight compression model parameter set, and construct an inference model for the power multi-modal teacher model according to the lightweight compression model parameter set to obtain a power feature extraction and fusion model.
[0139] Specifically, the deep learning framework TensorFlow and its model compression tools can be used for model compression, such as the TensorFlow Model Optimization Toolkit. Pruning is performed on the distilled and optimized student model to remove redundant neurons and connections. For example, setting the pruning ratio to 50% means removing 50% of the weights in the model while keeping the model structure unchanged. Through pruning, the number of model parameters is significantly reduced, and the computational complexity is decreased. Quantization is performed on the pruned model to convert the model weights from floating-point numbers to low-bitwidth integers (such as 8-bit integers). For example, using the dynamic quantization tool in TensorFlow, the model weights and activation functions are quantized to 8-bit integers. Finally, a power feature extraction and fusion model is constructed based on the lightweight compressed model parameter set.
[0140] By constructing a multimodal teacher network and performing cross-modal feature enhancement, the present invention can effectively integrate power data features from different modalities and improve the model's comprehensive representation ability for multi-source heterogeneous data. By using feature hierarchical decomposition and neural architecture search to construct a lightweight power student model, the key semantic knowledge in the teacher model can be efficiently transferred to the student model. Through feature alignment, it can be ensured that the student model can accurately inherit the feature representation ability of the teacher model during the learning process. By designing a multi-level power knowledge distillation strategy based on a cross-modal feature alignment loss dataset, the parameters in the knowledge transfer process can be dynamically adjusted to ensure that the student model can efficiently learn the knowledge of the teacher model at different levels. Through model compression, the storage and computational requirements of the model can be further reduced while retaining the key feature extraction ability.
[0141] Preferably, step S5 includes the following steps:
[0142] Step S51: Perform a reliability assessment on the multi-source heterogeneous power dataset to obtain power multimodal quality assessment data;
[0143] Specifically, the IsolationForest algorithm in scikit-learn can be used to detect abnormal samples in the multi-source heterogeneous power dataset. For example, setting the parameter contamination of IsolationForest to 0.05 means that the proportion of outliers in the dataset is approximately 5%. Through this algorithm, a reliability score is given to each sample, and the reliability score ranges from 0 to 1. The closer the value is to 1, the more reliable the data is. Finally, according to the reliability score, high-quality data is selected as the power multimodal quality assessment data. For example, data with a reliability score greater than 0.8 is selected as high-quality data.
[0144] Step S52: Perform state space modeling on the power multimodal quality assessment data to obtain a power feature state space dataset;
[0145] Specifically, the statsmodels library and the PyEMD library in Python can be used for state space modeling. The SARIMAX model in statsmodels is used to model time series data (such as voltage, current, and power). For example, assuming that the voltage data has seasonal variations, the parameters of the SARIMAX model are selected as (p = 1, D = 1, q = 1) × (P = 1, D = 1, Q = 1, s = 24), indicating that the model takes into account daily periodic variations. For non-linear features (such as equipment status, environmental temperature, and humidity), the empirical mode decomposition (EMD) method in the PyEMD library is used to decompose them into multiple intrinsic mode functions (IMFs). For example, the EMD decomposition is performed on the equipment status data to extract the IMF components that reflect the changes in the equipment operating status. Each IMF component can be regarded as a state variable to describe the dynamic behavior of the system. Finally, the output of the time series model and the IMF components obtained from the EMD decomposition are integrated into the power feature state space dataset.
[0146] Step S53: Perform reinforcement learning on the power feature state space dataset to obtain a dynamic power data fusion strategy;
[0147] Specifically, a reinforcement learning environment can be defined, where the state space is the power feature state space dataset, and the action space is the parameter adjustment of the data fusion strategy. For example, the actions can include selecting the weights of different features or adjusting the hyperparameters of the fusion model. Select a reinforcement learning algorithm, such as the Soft Actor-Critic (SAC) algorithm. During the training process, the agent learns the optimal strategy through interactions with the environment. For example, the reward function is set as a weighted sum of the stability index of the system state (such as the negative value of voltage fluctuation) and the prediction accuracy. The goal of the agent is to maximize the cumulative reward, that is, to improve the accuracy of data fusion while ensuring the stability of the system. Through multiple rounds of training, the agent gradually learns the optimal dynamic power data fusion strategy. For example, after 1000 iterations of training, the agent can dynamically adjust the feature weights according to the current state to adapt to different operating conditions. Finally, the obtained dynamic power data fusion strategy can automatically adjust the data fusion process according to the real-time state, so as to maintain the optimal fusion effect in different scenarios.
[0148] Step S54: Conduct a sensitivity assessment on the power feature extraction and fusion model to obtain the power modal feature sensitivity set;
[0149] Specifically, the input of the power feature extraction and fusion model can be defined as multi-modal features, and the output as the predicted value of the system state (such as equipment failure probability or voltage stability). The Integrated Gradients (IG) method is used to calculate the contribution of each modal feature to the model output. For example, for the voltage feature, calculate its integrated gradient value with respect to the model output. Suppose the sensitivity of the voltage feature is 0.6, indicating that the voltage feature has a greater impact on the model output; while the sensitivity of the ambient temperature feature is 0.2, indicating a smaller impact. Through the above steps, a sensitivity assessment is performed on all modal features, and the results are summarized into a power modal feature sensitivity set.
[0150] Step S55: According to the power modal feature sensitivity set, weight assignment is performed on the power data dynamic fusion strategy to obtain a power feature weight assignment data set;
[0151] Specifically, each modal feature can be normalized according to the power modal feature sensitivity set so that the sum of their weight values is 1. For example, the normalized weight of the voltage feature is 0.36, the current feature is 0.30, the power feature is 0.24, the temperature feature is 0.12, and the humidity feature is 0.08. These weight values are combined with the power data dynamic fusion strategy. The specific operation is to assign the weight of each modal feature to the corresponding feature processing module in the dynamic fusion strategy. For example, in the fusion model, the processing module for the voltage feature will be assigned a higher weight (0.36), while the processing module for the humidity feature has a lower weight (0.08). Through the above steps, the generated power feature weight assignment data set will contain each modal feature and its corresponding weight value.
[0152] Step S56: According to the power feature weight assignment data set, parameter adjustment is performed on the power feature extraction and fusion model to obtain an optimal power fusion decision model.
[0153] Specifically, the parameters of each feature processing module in the model can be adjusted according to the weight assignment data set. For example, for the voltage feature processing module, increase its weight in the model (such as by increasing the number of neurons in the relevant layer or adjusting the learning rate). The specific operation can be to adjust the output weight of the voltage feature processing module to 0.36, the output weight of the current feature module to 0.30, and so on. Next, retrain the adjusted model. During the training process, use the weighted feature outputs to optimize the loss function of the model. For example, define a weighted loss function where the loss contribution of each modal feature is proportional to its weight. Through the above steps, after parameter adjustment and retraining, the obtained model is the optimal power fusion decision model.
[0154] Step S57: Use the optimal power fusion decision model to conduct semantic association mining on the multi-source heterogeneous power data set to obtain a power multi-modal semantic association data set.
[0155] Specifically, the multi-source heterogeneous power data set can be input into the optimal power fusion decision model. The model can identify the relationship between voltage fluctuations and equipment status, or the impact of ambient temperature changes on power. Define the goal of semantic association mining. For example, assume the goal is to mine the association between equipment failures and multi-modal data. Through the output layer of the model, extract the feature representations related to equipment failures. These feature representations can be the activation values of the middle layer of the model, or the weighted features obtained through the attention mechanism. For example, use the activation values of the second-to-last layer of the model as the feature representation, with a dimension of 128. Use cosine similarity or Pearson correlation coefficient to measure the association strength between different modal features and equipment failures. For example, calculate the cosine similarity between the voltage feature and the equipment failure feature, and obtain an association strength of 0.75; calculate the cosine similarity between the ambient temperature feature and the equipment failure feature, and obtain an association strength of 0.45. Through the above steps, it is possible to determine which modal features have a strong semantic association with equipment failures. Finally, summarize the mined semantic association information into a power multi-modal semantic association data set. This data set not only contains the original data, but also attaches the association strength and semantic relationship between different modal features. For example, the data set can record that the association strength between the voltage feature and equipment failures is 0.75, indicating a strong semantic association between voltage fluctuations and equipment failures.
[0156] Through reliability assessment, the present invention can screen out high-quality and highly reliable data. By using state space modeling and reinforcement learning, it can dynamically adjust the data fusion strategy according to the real-time state of the power system. Through sensitivity assessment, it can identify the features that have the greatest impact on the model performance. Based on these sensitive features, weight assignment is performed to further optimize the data fusion strategy, enabling the model to capture key information more accurately and enhancing the model's adaptability to complex data environments. Through parameter adjustment, the optimal power fusion decision model can be obtained. Through semantic association mining, it can efficiently extract the deep semantic information in the data.
[0157] Preferably, the present invention also provides a multi-modal power cross-domain data fusion mining system for performing the multi-modal power cross-domain data fusion mining method as described above. The multi-modal power cross-domain data fusion mining system includes:
[0158] A data preprocessing module, configured to obtain a multi-source heterogeneous power data set; perform tensor decomposition on the multi-source heterogeneous power data set to obtain a power core feature tensor set; perform feature selection and reconstruction on the power core feature tensor set to obtain a reconstructed power feature data set;
[0159] A feature mapping module, which is used to perform domain adaptation feature mapping on the reconstructed power feature dataset to obtain a cross-domain power feature mapping dataset; eliminate domain differences from the cross-domain power feature mapping dataset to obtain a domain-aligned power feature dataset; calculate attention weights for the domain-aligned power feature dataset to obtain a cross-domain fusion power feature dataset;
[0160] A model training module, which is used to obtain distributed node topology data; construct a graph convolutional network for the distributed node topology data to obtain power node graph network structure data; perform distributed model training on the cross-domain fusion power feature dataset based on the power node graph network structure data to obtain a local power model parameter set; perform federated average calculation on the local power model parameter set to obtain a globally optimized model parameter set;
[0161] A model fusion module, which is used to construct a teacher network for the globally optimized model parameter set to obtain a power multi-modal teacher model; train a student network according to the power multi-modal teacher model to obtain a lightweight power student model; fuse the lightweight power student model and the power multi-modal teacher model to obtain a power feature extraction fusion model;
[0162] A data mining module, which is used to obtain power multi-modal quality assessment data; construct a reinforcement learning strategy for the power multi-modal quality assessment data to obtain a power data dynamic fusion strategy; adjust the weights of the power feature extraction fusion model according to the power data dynamic fusion strategy to obtain an optimal power fusion decision model; use the optimal power fusion decision model to mine semantic associations in the multi-source heterogeneous power dataset to obtain a power multi-modal semantic association dataset.
Claims
1. A data fusion mining method based on multi-modal power cross-domain, characterized in that: The following steps are involved: Step S1: Acquire a multi-source heterogeneous power data set; perform tensor decomposition on the multi-source heterogeneous power data set to obtain a power core feature tensor set; Perform feature selection and reconstruction on the power core feature tensor set to obtain a reconstructed power feature data set; Step S2: performing domain adaptation feature mapping on the reconstructed power feature data set to obtain a cross-domain power feature mapping data set; Eliminate domain differences on the cross-domain power feature mapping dataset to obtain a domain-aligned power feature dataset; Attention weights are calculated for the domain-aligned power feature dataset to obtain a cross-domain fused power feature dataset; Step S3: Obtain distributed node topology data; Construct a graph convolutional network for distributed node topology data to obtain the power node graph network structure data; Based on the network structure data of the power node graph, a distributed model training is performed on the cross-domain fusion power feature data set to obtain a local power model parameter set; a federal average calculation is performed on the local power model parameter set to obtain a global optimization model parameter set; Step S4: constructing a teacher network for the global optimization model parameter set to obtain a power multimodal teacher model; training a student network based on the power multimodal teacher model to obtain a power lightweight student model; The power lightweight student model and the power multimodal teacher model are fused to obtain the power feature extraction fusion model; Step S5: Acquire power multimodal quality assessment data; construct a reinforcement learning strategy for the power multimodal quality assessment data to obtain a power data dynamic fusion strategy; adjust the weights of the power feature extraction fusion model according to the power data dynamic fusion strategy to obtain an optimal power fusion decision model; use the optimal power fusion decision model to perform semantic association mining on multi-source heterogeneous power data sets to obtain a power multimodal semantic association data set.
2. The data fusion mining method based on multi-modal power cross-domain according to claim 1 is characterized in that: Step S1 includes the following steps: Step S11: performing multimodal power data collection on the target power grid to obtain an original multimodal power data set, and performing time stamp alignment on the original multimodal power data set to obtain a time series power data set; Step S12: performing anomaly detection on the time series power data set based on preset data quality assessment indicators to obtain a power data anomaly tag set, and performing data cleaning on the time series power data set according to the power data anomaly tag set to obtain a multi-source heterogeneous power data set; Step S13: identifying data types of the multi-source heterogeneous power data sets to obtain a power data type feature set, and standardizing the multi-source heterogeneous power data sets according to the power data type feature set to obtain a standard heterogeneous power data set; Step S14: constructing a tensor for the standard heterogeneous power data set to obtain a multi-dimensional power tensor data set; Step S15: performing rank evaluation on the multidimensional power tensor data set to obtain a power tensor rank evaluation data set; Step S16: performing tensor decomposition on the multidimensional power tensor data set according to the power tensor rank evaluation data set to obtain a power core feature tensor set; Step S17: performing feature selection and reconstruction on the power core feature tensor set to obtain a reconstructed power feature data set.
3. The data fusion mining method based on multi-modal power cross-domain according to claim 2 is characterized in that: Step S17 includes the following steps: Step S171: performing feature importance evaluation on the power core feature tensor set to obtain a power feature importance score set; Step S172: performing dynamic feature threshold iteration according to the power feature importance score set to obtain a power feature selection threshold set; Step S173: performing feature screening on the power core feature tensor set according to the power feature selection threshold set to obtain a key power feature data set; Step S174: designing an autoencoder network structure for the key power feature data set to obtain an encoder structure parameter set; Step S175: performing encoding training on the key power feature data set according to the encoder structure parameter set to obtain a latent power feature encoding data set; Step S176: Perform decoder reconstruction on the latent power feature encoding data set to obtain a reconstructed power feature data set.
4. The data fusion mining method based on multi-modal power cross-domain according to claim 1 is characterized in that: Step S2 includes the following steps: Step S21: quantifying the domain feature distribution of the reconstructed power feature data set to obtain a multi-domain power feature statistical data set; Step S22: performing domain adaptive network mapping on the reconstructed power feature data set according to the multi-domain power feature statistical data set to obtain a power domain mapping network parameter set; Step S23: clustering samples within the domain for the power domain mapping network parameter set to obtain a hierarchical domain clustering feature data set; Step S24: performing inter-domain distance calculation based on the hierarchical domain clustering feature data set to obtain a cross-domain metric distance data set; Step S25: performing iterative gradient adjustment on the power domain mapping network parameter set according to the cross-domain metric distance data set to obtain an adaptive power domain mapping parameter set; Step S26: performing feature projection on the reconstructed power feature data set according to the adaptive power domain mapping parameter set to obtain a cross-domain power feature mapping data set; Step S27: performing domain difference elimination on the cross-domain power feature mapping dataset to obtain a domain-aligned power feature dataset, and performing attention weight calculation on the domain-aligned power feature dataset to obtain a cross-domain fused power feature dataset.
5. The data fusion mining method based on multi-modal power cross-domain according to claim 4 is characterized in that: Step S27 includes the following steps: Step S271: constructing a discriminator network for the cross-domain power feature mapping data set to obtain an adversarial discriminator network; Step S272: performing adversarial training on the cross-domain power feature mapping dataset based on the adversarial discriminator network to obtain a domain adversarial loss power dataset; Step S273: adjusting the cross-domain power feature mapping data set according to the domain-resistance loss power data set to obtain a domain-aligned power feature data set; Step S274: performing attention weight calculation on the domain-aligned power feature dataset to obtain a multimodal power feature weight dataset; Step S275: constructing a weight optimization target according to the multimodal power feature weight data set to obtain a power feature weight optimization target data set, and iteratively updating the parameters of the multimodal power feature weight data set according to the power feature weight optimization target data set to obtain an optimized power feature weight data set; Step S276: performing weighted fusion on the domain-aligned power feature data set according to the optimized power feature weight data set to obtain an initial fused power feature data set, and performing feature consistency verification on the initial fused power feature data set to obtain an inter-domain feature consistency evaluation data set; Step S277: performing consistency calibration on the initial fused power feature dataset according to the inter-domain feature consistency evaluation dataset to obtain a cross-domain fused power feature dataset.
6. The data fusion mining method based on multi-modal power cross-domain according to claim 1 is characterized in that: Step S3 includes the following steps: Step S31: collecting topological relationships of distributed power nodes to obtain a node connection relationship data set; Step S32: performing link status evaluation on distributed power nodes based on the node connection relationship data set to obtain a node link quality data set; Step S33: Calculating the connection weights of the distributed power nodes according to the node link quality data set to obtain a node connection weight data set, and constructing a dynamic adjacency matrix according to the node connection weight data set to obtain distributed node topology data; Step S34: performing feature extraction layer planning for distributed power nodes according to the distributed node topology data to obtain a power node graph convolution layer parameter set; Step S35: constructing a heterogeneous message transmission mechanism according to the power node graph convolution layer parameter set to obtain a power node message transmission rule set; Step S36: constructing a graph neural architecture according to the power node message transmission rule set to obtain power node graph network structure data; Step S37: Perform distributed model training on the cross-domain fused power feature data set based on the power node graph network structure data to obtain a local power model parameter set, and perform federal average calculation on the local power model parameter set to obtain a global optimization model parameter set.
7. The data fusion mining method based on multi-modal power cross-domain according to claim 6 is characterized in that: Step S37 includes the following steps: Step S371: dividing the cross-domain fusion power feature data set into distributed training batches according to the power node graph network structure data to obtain a distributed training batch parameter set; Step S372: asynchronously parallel scheduling the cross-domain fusion power feature data set according to the distributed training batch parameter set to obtain a parallel training task allocation data set; Step S373: performing distributed optimizer configuration according to the parallel training task allocation data set to obtain a distributed optimizer configuration parameter set; Step S374: Design a learning rate scheduling strategy for the distributed optimizer configuration parameter set to obtain a power feature scheduling learning rate parameter set; Step S375: performing model distributed gradient calculation on distributed power nodes according to the power feature scheduling learning rate parameter set to obtain a power node gradient update data set; Step S376: performing asynchronous gradient aggregation on the power node gradient update data set to obtain a global gradient aggregation data set, and performing model parameter update according to the global gradient aggregation data set to obtain a local model update parameter set; Step S377: perform parameter weighted averaging on the local power model parameter set to obtain a federal average weight data set, and perform federal aggregation on the local power model parameter set based on the federal average weight data set to obtain a global optimization model parameter set.
8. The data fusion mining method based on multi-modal power cross-domain according to claim 1 is characterized in that: Step S4 includes the following steps: Step S41: constructing a neural network for the global optimization model parameter set to obtain a multimodal teacher network, and performing power cross-modal feature enhancement on the multimodal teacher network to obtain a power multimodal teacher model; Step S42: performing feature hierarchical decomposition on the power multimodal teacher model to obtain a teacher semantic knowledge feature set; Step S43: performing feature selection on the teacher semantic knowledge feature set to obtain a key power representation knowledge feature set, and performing neural architecture search and construction on the key power representation knowledge feature set to obtain a power lightweight student model; Step S44: aligning and extracting features of the power lightweight student model and the power multimodal teacher model to obtain a cross-modal feature alignment loss data set; Step S45: designing a knowledge distillation strategy according to the cross-modal feature alignment loss data set to obtain a multi-level power knowledge distillation strategy parameter set, and performing distillation optimization on the power lightweight student model according to the multi-level power knowledge distillation strategy parameter set to obtain a model distillation optimization parameter set; Step S46: compress the model distillation optimization parameter set to obtain a lightweight compressed model parameter set, and construct an inference model for the power multimodal teacher model based on the lightweight compressed model parameter set to obtain a power feature extraction fusion model.
9. The data fusion mining method based on multi-modal power cross-domain according to claim 1 is characterized in that: Step S5 includes the following steps: Step S51: performing reliability assessment on multi-source heterogeneous power data sets to obtain power multimodal quality assessment data; Step S52: performing state space modeling on the power multimodal quality assessment data to obtain a power feature state space data set; Step S53: performing reinforcement learning on the power feature state space data set to obtain a power data dynamic fusion strategy; Step S54: performing sensitivity evaluation on the power feature extraction fusion model to obtain a power modal feature sensitivity set; Step S55: weighting the power data dynamic fusion strategy according to the power modal feature sensitivity set to obtain a power feature weighted distribution data set; Step S56: adjusting the parameters of the power feature extraction and fusion model according to the power feature weight allocation data set to obtain an optimal power fusion decision model. Step S57: using the optimal power fusion decision model to perform semantic association mining on the multi-source heterogeneous power data set to obtain a power multimodal semantic association data set.
10. A data fusion mining system based on multi-modal power cross-domain, characterized in that: Used to execute the data fusion mining method based on multi-modal power cross-domain as claimed in claim 1, the data fusion mining system based on multi-modal power cross-domain includes: The data preprocessing module is used to obtain a multi-source heterogeneous power data set; perform tensor decomposition on the multi-source heterogeneous power data set to obtain a power core feature tensor set; perform feature selection and reconstruction on the power core feature tensor set to obtain a reconstructed power feature data set; The feature mapping module is used to perform domain adaptation feature mapping on the reconstructed power feature data set to obtain a cross-domain power feature mapping data set; perform domain difference elimination on the cross-domain power feature mapping data set to obtain a domain-aligned power feature data set; perform attention weight calculation on the domain-aligned power feature data set to obtain a cross-domain fusion power feature data set; The model training module is used to obtain distributed node topology data; construct a graph convolution network on the distributed node topology data to obtain the power node graph network structure data; perform distributed model training on the cross-domain fusion power feature data set based on the power node graph network structure data to obtain the local power model parameter set; perform federal average calculation on the local power model parameter set to obtain the global optimization model parameter set; The model fusion module is used to construct a teacher network for the global optimization model parameter set to obtain a power multimodal teacher model; train a student network based on the power multimodal teacher model to obtain a power lightweight student model; and fuse the power lightweight student model with the power multimodal teacher model to obtain a power feature extraction fusion model. The data mining module is used to obtain power multimodal quality assessment data; to construct a reinforcement learning strategy for the power multimodal quality assessment data to obtain a power data dynamic fusion strategy; to adjust the weights of the power feature extraction and fusion model according to the power data dynamic fusion strategy to obtain an optimal power fusion decision model; to use the optimal power fusion decision model to perform semantic association mining on multi-source heterogeneous power data sets to obtain a power multimodal semantic association data set.
Citation Information
Patent Citations
Rolling bearing fault diagnosis method and system under different working conditions based on federal feature transfer learning
CN115560983A
Optical fiber signal fusion and reconstruction method for multi-source heterogeneous data
CN117786593A
Fusion networking method and system based on satellite communication and short-wave communication
CN118233936A
Power big data processing method based on data clustering algorithm
CN118503730A
Graph neural network-based power grid dispatching decision-making method and large model
CN119294872A
Cited By
Multi-modal database dynamic fusion optimization method and system based on large model
CN120670414A
Power system scheduling method and device based on multi-modal data fusion
CN120879789A
Small sample image classification method and system based on multi-modal data fusion
CN120974296A
Power equipment fault interaction method based on double-brain driving and related system
CN121188698A
Large model parameter fine tuning method and device for electric power multi-modal data fusion and medium
CN121350955A