Power data analysis method based on data fusion

By performing time-frequency analysis and feature fusion of power data, the power data characteristics are extracted using Graph SAGE network and neural network, the problem of low accuracy of multi-source heterogeneous data analysis is solved, and a higher precision of power data analysis is achieved.

CN120493087APending Publication Date: 2025-08-15HONGLAN (NINGXIA) ENERGY DEV CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510579989.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-07
Publication Date
2025-08-15

AI Technical Summary

Technical Problem

In the prior art, the analysis accuracy of multi-source heterogeneous power data is low, and traditional methods cannot fully utilize multi-source data information, resulting in limited accuracy and reliability of the analysis results.

Method used

By obtaining power data and dividing it into multiple data fragments, time-frequency analysis is performed, directed graphs are constructed, and features are extracted using Graph SAGE network and pyramid convolutional neural network, feature fusion is combined with bidirectional recurrent neural network, and gradient enhancement tree model is constructed for power data analysis.

Benefits of technology

It significantly improves the accuracy of power data analysis, fully explores the feature information of a single data segment, captures the topological relationship between data, uses the local and sequential characteristics of time-frequency domain data, establishes the correlation between data of different physical quantities, and improves the feature representation ability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120493087A_ABST
    Figure CN120493087A_ABST
Patent Text Reader

Abstract

The invention discloses an electric power data analysis method based on data fusion, and relates to the field of data processing, and the method comprises the steps: obtaining electric power data, and dividing the electric power data into a plurality of data segments; extracting time domain features and frequency domain features; constructing a directed graph; a Graph SAGE network is adopted for coding, and graph embedding features are obtained; acquiring time domain modal features by adopting a pyramid convolutional neural network; a bidirectional recurrent neural network is adopted to obtain frequency domain modal features; obtaining a multi-modal fusion feature according to the graph embedding feature, the time domain feature, the frequency domain feature, the time domain modal feature and the frequency domain modal feature; obtaining a cross-modal fusion feature according to the time domain modal feature and the frequency domain modal feature; according to the multi-modal fusion features and the cross-modal fusion features, splicing feature vectors of the corresponding data segments are obtained through feature interaction; performing nonlinear transformation on the spliced feature vector to obtain final feature representation; aiming at the problem of low power data analysis precision caused by multi-source heterogeneous data in the prior art, the data analysis precision is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of data processing, and in particular to a power data analysis method based on data fusion. Background Art

[0002] The continuous development and increasing intelligence of power systems, coupled with the widespread adoption of sensor technology, data acquisition equipment, and monitoring systems, has led to the generation of massive amounts of data. This data comes from diverse sources, including but not limited to various monitoring data from power generation, transmission, and distribution processes, sensor data, and SCADA system data. This data is generated in diverse forms, structures, and frequencies, forming multi-source, heterogeneous datasets.

[0003] However, this multi-source, heterogeneous dataset also presents a series of challenges, one of the most significant of which is how to fully utilize this data for accurate and efficient power data analysis. Traditional data analysis methods are often limited to analyzing a single data source or a specific data type, failing to fully utilize the information from multiple sources, resulting in limited accuracy and reliability of the analysis results.

[0004] In related technologies, for example, Chinese patent document CN117149846B provides a data fusion-based power data analysis method and system, which relates to the field of data processing technology. The method includes: utilizing multiple power data analysis networks to perform feature mining operations on a power data file to be processed to output multiple corresponding initial data feature representations, wherein each of the multiple power data analysis networks is used to output corresponding power anomaly characterization data based on the loaded data; performing a feature representation fusion operation on the multiple initial data feature representations to form a corresponding aggregated data feature representation; and analyzing target power anomaly characterization data corresponding to the power data file to be processed based on the aggregated data feature representation. The target power anomaly characterization data is used to reflect the abnormal state of the power system corresponding to the power data file to be processed. However, the target power anomaly characterization data analyzed based on the aggregated data feature representation is used to reflect the abnormal state of the power system corresponding to the power data file to be processed. However, the identification of the abnormal state may be affected by incomplete data representation or error propagation, resulting in limitations in the accurate identification of the abnormal state. Therefore, the accuracy of power data analysis of this scheme needs to be further improved. Summary of the Invention

[0005] In response to the problem of low accuracy of power data analysis caused by multi-source heterogeneous data in the existing technology, the present application provides a power data analysis method based on data fusion, which fully integrates the characteristics of multi-source data from the time-frequency domain, graph topology and modal complementarity, thereby improving the accuracy of data analysis.

[0006] The purpose of this application is achieved through the following technical solutions.

[0007] This specification provides a power data analysis method based on data fusion, including: acquiring power data, dividing the acquired power data into multiple data segments; performing time-frequency analysis on the multiple data segments, extracting the time domain features and frequency domain features of each data segment; constructing a directed graph based on the multiple data segments; the nodes in the directed graph represent the data segments, and the edges in the directed graph represent the correlation between the data segments; taking the constructed directed graph as input, using Graph The SAGE network is used for encoding, and the features of the central node are updated by aggregating the features of the neighboring nodes to obtain the aggregate vector of the node corresponding to each data segment, which is used as the graph embedding feature of the corresponding data segment; the pyramid convolutional neural network is used to extract the time domain features of each data segment to obtain the time domain modal features; the bidirectional recurrent neural network is used to extract the frequency domain features of each data segment to obtain the frequency domain modal features; the graph embedding features, time domain features, frequency domain features, time domain modal features and frequency domain modal features of each data segment are spliced to obtain multimodal fusion features; the time domain modal features and frequency domain modal features of each data segment are weighted fused to obtain the cross-modal fusion features of the corresponding data segment; for each data segment, the corresponding multimodal fusion features and cross-modal fusion features are multiplied, and the spliced feature vector of the corresponding data segment is obtained through feature interaction; the spliced feature vector of each data segment is input into the multi-layer perceptron network for nonlinear transformation to obtain the final feature representation of the corresponding data segment; based on the final feature representation of each data segment, a gradient boosting tree model is constructed as a prediction model for power data analysis.

[0008] Furthermore, the time domain features and frequency domain features of each data segment are obtained, including: performing a short-time Fourier transform on each data segment, converting the data segment into a time-frequency matrix by setting the window length and overlap rate parameters of the sliding window; the rows of the time-frequency matrix represent the time dimension, and the columns represent the frequency dimension; extracting the statistical features of the time-frequency matrix of each data segment as the time domain features; the statistical features include mean, variance and kurtosis; performing a wavelet transform and a Hilbert-Huang transform on the time-frequency matrix of each data segment to obtain the wavelet frequency domain features and the Hilbert frequency domain features of the corresponding data segment, and obtaining the frequency domain features based on the wavelet frequency domain features and the Hilbert frequency domain features of the corresponding data segment.

[0009] Preferably, a polynomial window or Gaussian window is used as the window function in the short-time Fourier transform. This allows the window function parameters to be adaptively adjusted based on the local characteristics of the data, improving the adaptability of time-frequency analysis. The optimal overlap ratio parameter is found through methods such as cross-validation or heuristic search to achieve a balance between time resolution and frequency resolution.

[0010] Furthermore, wavelet frequency domain features are obtained, including: performing continuous wavelet transform on the time-frequency matrix, converting the time-frequency matrix into a wavelet coefficient matrix through a wavelet function; the rows of the wavelet coefficient matrix represent the time dimension, and the columns represent the scale dimension; the wavelet function is one of Morlet wavelet, Mexican Hat wavelet, Daubechies wavelet and Symlet wavelet; performing energy spectrum analysis on the wavelet coefficient matrix, and obtaining the wavelet energy spectrum of the corresponding data segment by calculating the square sum of each column of the wavelet coefficient matrix; extracting the maximum energy scale, energy center of gravity scale and energy dispersion of the wavelet energy spectrum as the wavelet frequency domain features of the corresponding data segment.

[0011] Furthermore, frequency domain features are obtained, including: performing marginal spectrum analysis on the time-frequency matrix, obtaining the Fourier frequency domain marginal spectrum of the corresponding data segment by summing each column of the time-frequency matrix; performing Hilbert-Huang transform on the Fourier frequency domain marginal spectrum, solving the analytical signal by the frequency domain marginal spectrum, and extracting the instantaneous frequency and instantaneous amplitude; extracting the statistical features of the instantaneous frequency and instantaneous amplitude to obtain the Hilbert frequency domain features of the corresponding data segment; splicing the wavelet frequency domain features and the Hilbert frequency domain features of the data segment as the frequency domain features of the corresponding data segment. Preferably, a smoothing kernel function in the frequency domain, such as a Gaussian kernel or a triangular kernel, is used to perform convolution smoothing on the frequency domain marginal spectrum to obtain a smoother and more continuous spectrum line.

[0012] Furthermore, the graph embedding features of the corresponding data fragment are obtained, including: initializing the Graph SAGE network according to the constructed directed graph, and taking the features of each node in the directed graph as the initial feature vector; the features of each node include the time-frequency statistical features, wavelet frequency domain features and Hilbert frequency domain features of the corresponding data fragment; setting the aggregation function and update function of the Graph SAGE network; the aggregation function adopts the attention pooling method, and assigns different weights to the feature vectors of the neighboring nodes through the attention mechanism to obtain a weighted aggregate feature vector; the update function adopts the gated recursive method to perform nonlinear transformation on the weighted aggregate feature vector and update the features of the central node; through the forward propagation algorithm, the features of the nodes in the Graph SAGE network are recursively updated for K rounds; after K rounds of updates, the final features of each node are output as the graph embedding features of the corresponding data fragment.

[0013] Furthermore, the forward propagation algorithm is used to recursively update the features of the nodes in the Graph SAGE network for K rounds, including: in each round of update, all nodes in the directed graph are traversed, and the currently traversed node is used as the center node; for each center node, the first-order neighboring nodes and second-order neighboring nodes of the center node are sampled according to the adjacency matrix of the directed graph to obtain the first-order neighboring node set and the second-order neighboring node set; according to the first-order neighboring node set and the second-order neighboring node set, the nodes in the Graph SAGE network are updated. From the node feature matrix of the last round of update of the SAGE network, the current feature vectors corresponding to the first-order neighbor nodes and the second-order neighbor nodes are extracted; the current feature vectors of the first-order neighbor nodes and the current feature vectors of the second-order neighbor nodes are respectively input into the self-attention pooling function, and the correlation between the neighbor nodes is calculated through the multi-head self-attention mechanism to generate the self-attention weight matrices of the first-order neighbor nodes and the second-order neighbor nodes; according to the self-attention weight matrices of the first-order neighbor nodes and the second-order neighbor nodes, the current feature vectors of the first-order neighbor nodes and the current feature vectors of the second-order neighbor nodes are weighted summed to obtain the self-attention aggregate feature vectors of the first-order neighbor nodes and the second-order neighbor nodes; the self-attention aggregate feature vectors of the first-order neighbor nodes and the self-attention aggregate feature vectors of the second-order neighbor nodes are spliced, and the spliced aggregate feature vectors are residually connected with the feature vector of the center node in the last round of update to obtain the initial updated feature vector of the center node.

[0014] Furthermore, the features of the nodes in the Graph SAGE network are recursively updated for K rounds through a forward propagation algorithm, which also includes: inputting the initial updated feature vector of the central node into a multi-layer gated recursive unit, and obtaining the final updated feature vector of the central node after undergoing multi-layer nonlinear transformation; adding the final updated feature vector of the central node to the feature vector of the central node in the previous round of update through a residual connection to obtain the current round updated feature vector of the central node; and storing the current round updated feature vector of the central node in the corresponding position of the node feature matrix of the Graph SAGE network in the current round of update.

[0015] Among them, the multi-layer gated recurrent unit is a recurrent neural network structure for processing sequence data. It is a multi-layer extension of the gated recurrent unit (GRU). The gated recurrent unit (GRU) is a simplified version of the long short-term memory (LSTM) network, which is used to solve the gradient vanishing and gradient exploding problems faced by traditional recurrent neural networks (RNNs) when processing long sequence data. By introducing a gating mechanism, the GRU controls the flow and update of information, selectively retaining and forgetting historical information, and thus can better capture long-distance dependencies in sequence data. The multi-layer gated recurrent unit (Multi-layer GRU) stacks multiple GRU layers together to form a deeper recurrent neural network structure. The output of each GRU layer serves as the input to the next layer, and through multiple layers of nonlinear transformations, it gradually extracts and abstracts high-level features in the sequence data.

[0016] Furthermore, multimodal fusion features are obtained, including: taking the graph embedding features, time domain features, frequency domain features, time domain modal features and frequency domain modal features of each data segment as input features; performing linear transformation on the input features to obtain latent variables; generating query vectors from the latent variables through an attention mechanism; generating graph embedding feature key-value pairs, time domain feature key-value pairs, frequency domain feature key-value pairs, time domain modal feature key-value pairs and frequency domain modal feature key-value pairs through a multi-layer perceptron network based on the input features; using a scaled dot product attention mechanism, performing dot product operations on the query vector and the graph embedding feature key-value pairs, time domain feature key-value pairs, frequency domain feature key-value pairs, time domain modal feature key-value pairs and frequency domain modal feature key-value pairs respectively to obtain corresponding attention scores; according to the attention scores, obtaining corresponding attention weights through a softmax function; and according to the attention weights, performing weighted summation on the input features to obtain multimodal fusion features.

[0017] Furthermore, for each data segment, the corresponding multimodal fusion feature and the cross-modal fusion feature are multiplied, and the spliced feature vector of the corresponding data segment is obtained through feature interaction, including: for each data segment, the corresponding multimodal fusion feature and the cross-modal fusion feature are outer-product operated to generate a two-dimensional interaction feature matrix; wherein the number of rows of the two-dimensional interaction feature matrix is equal to the dimension of the multimodal fusion feature, and the number of columns is equal to the dimension of the cross-modal fusion feature; attention pooling operations are performed on the rows and columns of the two-dimensional interaction feature matrix respectively to obtain a row attention feature matrix and a column attention feature matrix; the row attention feature matrix and the column attention feature matrix are weightedly fused to obtain an attention pooling feature matrix; a gating operation is performed on the attention pooling feature matrix to obtain an enhanced attention pooling feature matrix; the enhanced attention pooling feature matrix is flattened into a one-dimensional vector as the spliced feature vector of the corresponding data segment; the spliced feature vector is fused with the multimodal fusion feature and the cross-modal fusion feature through a residual connection to obtain the final spliced feature vector of the data segment.

[0018] Furthermore, the concatenated feature vector of each data segment is input into a multi-layer perceptron network for nonlinear transformation to obtain a final feature representation of the corresponding data segment, including: sequentially inputting the concatenated feature vector of each data segment into multiple fully connected layers of the multi-layer perceptron network, performing forward propagation through a nonlinear activation function, and obtaining a final feature representation of the corresponding data segment; wherein the weight parameters of the multi-layer perceptron network are obtained through training.

[0019] Compared with the existing technology, the advantages of this application are:

[0020] By dividing the power data and extracting time-frequency features, the characteristic information of a single data segment can be fully mined, laying the foundation for subsequent data fusion.

[0021] Using directed graphs to represent the relationships between data segments, the Graph SAGE network learns graph embedding features to capture the topological relationships between data. This approach employs an attention mechanism to aggregate features of neighboring nodes and gates recursively update features of central nodes, enhancing information transfer between nodes and enabling graph embedding features to more accurately capture data associations.

[0022] Pyramid convolutional neural network and bidirectional recurrent neural network are used to extract modal features in the time domain and frequency domain respectively, making full use of the local and sequential characteristics of time and frequency domain data to learn richer and more abstract modal representations.

[0023] By concatenating and fusing graph embedding features, original time-frequency domain features, and modal features, we generate multimodal fusion features that comprehensively represent the characteristic information of a single data segment. We also perform weighted fusion on the time-frequency modal features to generate cross-modal fusion features, establishing associations between data of different physical quantities. This comprehensive utilization of multi-source heterogeneous data significantly enhances feature representation capabilities.

[0024] Through the outer product operation of multimodal and cross-modal fusion features, feature interactions within and between segments are explicitly modeled. Attention pooling and gating operations are used to enhance interactive features, highlighting key interactive information. The original features are then fused through residual connections to obtain a final comprehensive representation of the data segments, fully capturing the inherent connections within the data.

[0025] The data fragment representation is input into a multi-layer perceptron for nonlinear transformation and high-level abstraction, extracting deep semantic features. A gradient boosting tree model is constructed for analysis and prediction, fusing the prediction results of multiple decision trees. This model has strong nonlinear modeling capabilities and generalization performance, significantly improving the accuracy of power data analysis. BRIEF DESCRIPTION OF THE DRAWINGS

[0026] Figure 1 is an exemplary flow chart of a power data analysis method based on data fusion according to some embodiments of this specification;

[0027] Figure 2 is an exemplary flow chart of obtaining time domain features and frequency domain features according to some embodiments of this specification;

[0028] Figure 3 is an exemplary flow chart of obtaining graph embedding features according to some embodiments of this specification;

[0029] Figure 4 is an exemplary flow chart of obtaining an updated feature vector according to some embodiments of this specification;

[0030] Figure 5 This is an exemplary flowchart of obtaining multimodal fusion features according to some embodiments of this specification. DETAILED DESCRIPTION

[0031] The methods and systems provided in the embodiments of this specification are described in detail below with reference to the accompanying drawings.

[0032] Figure 1This is an exemplary flow chart of a power data analysis method based on data fusion according to some embodiments of this specification, a power data analysis method based on data fusion, including: acquiring power data, dividing the acquired power data into multiple data segments; performing time-frequency analysis on the multiple divided data segments, and extracting the time domain features and frequency domain features of each data segment; constructing a directed graph based on the multiple divided data segments; the nodes in the directed graph represent the data segments, and the edges in the directed graph represent the correlation between the data segments; using the constructed directed graph as input, encoding it using the GraphSAGE network, updating the features of the central node by aggregating the features of the neighboring nodes, and obtaining the aggregated vector of the node corresponding to each data segment as the graph embedding feature of the corresponding data segment; using the pyramid convolutional neural network to extract the time domain features of each data segment, and obtaining to time domain modal features; a bidirectional recurrent neural network is used to extract the frequency domain features of each data segment to obtain frequency domain modal features; the graph embedding features, time domain features, frequency domain features, time domain modal features and frequency domain modal features of each data segment are spliced to obtain multimodal fusion features; the time domain modal features and frequency domain modal features of each data segment are weightedly fused to obtain cross-modal fusion features of the corresponding data segment; for each data segment, the corresponding multimodal fusion features and cross-modal fusion features are multiplied, and the spliced feature vector of the corresponding data segment is obtained through feature interaction; the spliced feature vector of each data segment is input into the multi-layer perceptron network for nonlinear transformation to obtain the final feature representation of the corresponding data segment; according to the final feature representation of each data segment, a gradient boosting tree model is constructed as a prediction model to perform power data analysis.

[0033] First, power data is acquired and divided into multiple data segments. A short-time Fourier transform is performed on each data segment to extract the statistical characteristics of the time-frequency matrix as time-frequency features. The time-frequency matrix is then subjected to a wavelet transform and a Hilbert-Huang transform to extract the wavelet energy spectrum features and the instantaneous frequency and amplitude features, thereby obtaining frequency domain features.

[0034] Figure 2This is an exemplary flow chart for obtaining time domain features and frequency domain features according to some embodiments of this specification, wherein Short-time Fourier Transform (STFT): Short-time Fourier Transform is a time-frequency analysis method used to analyze the time-frequency characteristics of non-stationary signals. It performs a local Fourier transform on the signal through a sliding window to obtain the spectrum information of the signal in different time periods. STFT can generate a time-frequency matrix that reflects the energy distribution of the signal in the time and frequency dimensions. Hilbert-Huang Transform (HHT): Hilbert-Huang Transform is a method suitable for nonlinear and non-stationary signal analysis. It includes two steps: empirical mode decomposition (EMD) and Hilbert spectrum analysis. EMD decomposes the signal into several intrinsic mode functions (IMFs), each IMF representing an intrinsic mode of the signal. Then, each IMF is Hilbert transformed to obtain the instantaneous frequency and instantaneous amplitude, and a Hilbert spectrum is generated to reflect the time-frequency energy distribution of the signal. Frequency domain features refer to the characteristic representation of the signal in the frequency domain. By performing frequency domain analysis methods such as Fourier transform and wavelet transform on the time domain signal, the frequency domain features of the signal, such as spectral characteristics and wavelet coefficients, can be extracted. Frequency domain features reflect the energy distribution and variation law of the signal at different frequency components, and can characterize the frequency characteristics of the signal. Time domain features refer to the characteristic representation of the signal in the time domain. Time domain features are extracted directly from the time series of the original signal. Common time domain features include statistical features (such as mean, variance, kurtosis, etc.), morphological features (such as peak value, rise time, etc.) and energy features (such as root mean square energy). Time domain features reflect the waveform characteristics and change patterns of the signal in the time dimension. In this application, the data segment is converted into a time-frequency matrix by short-time Fourier transform, and the statistical features of the time-frequency matrix are extracted as time-frequency features. In this way, the characteristics of the data segment can be characterized in the time-frequency domain. At the same time, the wavelet frequency domain features and Hilbert frequency domain features of the data segment are extracted by wavelet transform and Hilbert-Huang transform to obtain a more comprehensive frequency domain feature representation. The combination of time domain features and frequency domain features can describe data segments from two dimensions: time and frequency, capture the diversity of data information, and provide richer feature representations for subsequent feature fusion and data analysis.

[0035] Among them, obtaining wavelet frequency domain features includes: performing continuous wavelet transform on the time-frequency matrix, converting the time-frequency matrix into a wavelet coefficient matrix through a wavelet function; the rows of the wavelet coefficient matrix represent the time dimension, and the columns represent the scale dimension; the wavelet function is one of Morlet wavelet, Mexican Hat wavelet, Daubechies wavelet and Symlet wavelet; performing energy spectrum analysis on the wavelet coefficient matrix, and obtaining the wavelet energy spectrum of the corresponding data segment by calculating the square sum of each column of the wavelet coefficient matrix; extracting the maximum energy scale, energy center of gravity scale and energy dispersion of the wavelet energy spectrum as the wavelet frequency domain features of the corresponding data segment.

[0036] Among them, Morlet wavelet: also known as Gabor wavelet, is a Gaussian modulated cosine function with good time-frequency localization characteristics, suitable for analyzing stationary signals and extracting instantaneous frequency features. Mexican Hat wavelet: also known as Ricker wavelet, is a negative derivative of a second-order Gaussian function with a symmetrical hat shape, suitable for detecting singular points and mutation points of signals. Daubechies wavelet: is a type of orthogonal wavelet function with compact support and vanishing moment characteristics. It can achieve a balance of time-frequency resolution by adjusting the order, and is suitable for multi-scale analysis and noise reduction of signals. Symlet wavelet: is a variant of the Daubechies wavelet with higher symmetry. While maintaining compact support and orthogonality, it can provide a smoother reconstructed signal.

[0037] Energy spectrum analysis is a time-frequency analysis method based on the wavelet transform. By calculating the energy of the wavelet coefficient matrix, the energy distribution of the signal at different scales (frequencies) can be obtained, namely the wavelet energy spectrum. The wavelet energy spectrum reflects the energy intensity of the signal at different frequency bands and can be used for feature extraction and signal analysis. The maximum energy scale refers to the scale (frequency) position with the highest energy in the wavelet energy spectrum. It indicates that the signal has the strongest energy component at this scale and reflects the main frequency components of the signal. The maximum energy scale can be used as an important frequency domain feature of the signal to characterize its frequency characteristics. The energy center of gravity scale refers to the scale value corresponding to the center of gravity position of the wavelet energy spectrum. It considers the distribution of energy at each scale and reflects the concentration of the signal energy. The energy center of gravity scale can characterize the energy distribution characteristics of the signal in the frequency domain and provide overall information about the signal's frequency components. Energy dispersion is a measure of the dispersion of the wavelet energy spectrum. It measures the degree of energy dispersion by calculating the difference between the energy spectrum and a uniform distribution. The larger the energy discreteness, the more uneven the energy distribution at different scales, that is, the more concentrated the energy distribution of the signal in the frequency domain; the smaller the energy discreteness, the more even the energy distribution at different scales, that is, the more dispersed the energy distribution of the signal in the frequency domain. In this application, a wavelet coefficient matrix is obtained by performing a continuous wavelet transform on the time-frequency matrix. Then, an energy spectrum analysis is performed on the wavelet coefficient matrix, and the energy at each scale is calculated to obtain a wavelet energy spectrum. The maximum energy scale, energy center of gravity scale, and energy discreteness are extracted from the wavelet energy spectrum as wavelet frequency domain features, which can characterize the energy distribution characteristics of the data segment in the frequency domain from different angles. The maximum energy scale reflects the main frequency component, the energy center of gravity scale reflects the overall frequency distribution, and the energy discreteness reflects the degree of concentration of energy.

[0038] Among them, obtaining frequency domain features includes: performing marginal spectrum analysis on the time-frequency matrix, obtaining the Fourier frequency domain marginal spectrum of the corresponding data segment by summing each column of the time-frequency matrix; performing Hilbert-Huang transform on the Fourier frequency domain marginal spectrum, solving the analytical signal by the frequency domain marginal spectrum, and extracting the instantaneous frequency and instantaneous amplitude; extracting the statistical features of the instantaneous frequency and instantaneous amplitude to obtain the Hilbert frequency domain features of the corresponding data segment; and splicing the wavelet frequency domain features and the Hilbert frequency domain features of the data segment as the frequency domain features of the corresponding data segment.

[0039] Figure 3This is an exemplary flow chart for obtaining graph embedding features according to some embodiments of this specification, wherein the Fourier frequency domain marginal spectrum is obtained by performing marginal spectrum analysis on the time-frequency matrix. Marginal spectrum analysis accumulates the time-frequency matrix in the time dimension to obtain the energy distribution in the frequency dimension, namely the Fourier frequency domain marginal spectrum. The Fourier frequency domain marginal spectrum reflects the energy level of the signal at each frequency component and provides the overall characteristics of the signal across the entire frequency domain. Solving the frequency domain marginal spectrum for the analytical signal is a step in the Hilbert-Huang transform. The analytical signal is obtained by performing the Hilbert transform on the frequency domain marginal spectrum. The analytical signal is a complex-valued signal whose real part is the original signal and whose imaginary part is the Hilbert transform of the original signal. The analytical signal can be used to extract the instantaneous characteristics of the signal, such as instantaneous frequency and instantaneous amplitude. The instantaneous frequency is a physical quantity that describes the frequency variation of a signal in the time-frequency plane. It represents the main frequency component of the signal at a given moment. The instantaneous frequency can be obtained by calculating the phase of the analytical signal and taking the derivative of the phase. The instantaneous frequency reflects the temporal variation of the signal's frequency components and provides the signal's dynamic characteristics in the time-frequency domain. The instantaneous amplitude represents the amplitude of the signal at a specific moment. It describes how the signal's energy changes over time. The instantaneous amplitude can be obtained by calculating the modulus of the analytical signal. The instantaneous amplitude reflects the temporal distribution of the signal's energy and provides the signal's dynamic characteristics in the time domain. In this application, a Fourier frequency domain marginal spectrum is obtained by performing marginal spectrum analysis on the time-frequency matrix, reflecting the signal's energy distribution across the entire frequency domain. The Hilbert-Huang transform is then performed on the Fourier frequency domain marginal spectrum to solve the analytical signal and extract the instantaneous frequency and instantaneous amplitude. The instantaneous frequency and instantaneous amplitude respectively characterize the dynamic variation characteristics of the signal's frequency and energy. By extracting the statistical characteristics of the instantaneous frequency and instantaneous amplitude, the Hilbert frequency domain features are obtained, reflecting the signal's dynamic characteristics in the time-frequency domain. Finally, the wavelet frequency domain features and the Hilbert frequency domain features are concatenated to obtain the frequency domain features, which comprehensively characterize the static and dynamic characteristics of the data segment in the frequency domain.

[0040] Next, a directed graph is constructed based on the data segments, and the time-frequency statistical features, wavelet frequency domain features, and Hilbert frequency domain features of each data segment are used as the initial feature vector of the node and input into the Graph SAGE network. Graph SAGE samples first-order and second-order neighboring nodes, aggregates neighborhood features using attention pooling, and recursively updates the center node features using gated updating to ultimately obtain the graph embedding features of each node. Graph SAGE (Graph S Ample and aggre GatE) is a neural network model for graph representation learning. It generates node embedding representations by aggregating node neighborhood information and can generate embeddings for unseen nodes. Graph SAGE can process large-scale graph data and supports unsupervised, semi-supervised, and supervised learning tasks.

[0041] Among them, the aggregation function is a function used to aggregate the features of neighborhood nodes in the Graph SAGE network. It summarizes the features of the neighborhood nodes of the central node to obtain a representation of the neighborhood information. Common aggregation functions include averaging, summing, pooling, etc. In this application, the aggregation function adopts the attention pooling method, and assigns different weights to the features of the neighborhood nodes through the attention mechanism to obtain a weighted aggregate feature vector, highlighting the contribution of important neighborhood nodes. The update function is a function used to update the features of the central node in the Graph SAGE network. It combines the aggregated neighborhood information with the current features of the central node to obtain the updated central node features. Common update functions include splicing, nonlinear transformation, etc. In this application, the update function adopts a gated recursive method, which controls the fusion of neighborhood information and current features through a gating mechanism, and performs nonlinear transformation to obtain the updated central node features. The gated recursive method is a way to implement the update function. It draws on the idea of the gated recurrent unit (GRU) and introduces update gates and reset gates to control the information flow. The update gate controls the influence of neighborhood information on the current feature, while the reset gate controls the influence of the current feature on the updated feature. This gating mechanism adaptively balances the importance of neighborhood information and the current feature, achieving selective feature updates. Attention pooling is an implementation of an aggregation function. It uses an attention mechanism to assign different weights to the features of neighboring nodes, highlighting the contributions of important neighboring nodes. Specifically, by calculating the attention weights between the central node and its neighbors, a weighted sum of the neighborhood node features is obtained, which serves as the aggregated neighborhood information. Attention weights can be calculated using methods such as dot product and cosine similarity, reflecting the relevance of neighboring nodes to the central node. Graph embedding features are node feature representations learned using the Graph SAGE network. They fuse graph structure and node attribute information into a low-dimensional vector space, ensuring that similar nodes in the embedded space have similar feature representations. Graph embedding features capture both local and global node structure and the mutual influence between nodes, providing an effective feature representation for subsequent graph analysis tasks. In this application, a directed graph is encoded using a Graph SAGE network. Node features are recursively updated using an attention-pooling aggregation function and a gated recursive update function. After K rounds of updates, a graph embedding is obtained for each node, reflecting the node's role and influence within the graph structure. This graph embedding combines the time-frequency and frequency-domain features of each data segment, as well as the correlation information between the data segments, providing a comprehensive representation of the data segments.

[0042] Figure 4This is an exemplary flowchart for obtaining and updating feature vectors according to some embodiments of this specification, wherein obtaining graph embedding features of corresponding data segments includes: initializing a Graph SAGE network based on a constructed directed graph, and using the features of each node in the directed graph as an initial feature vector; the features of each node include the time-frequency statistical features, wavelet frequency domain features, and Hilbert frequency domain features of the corresponding data segment; setting an aggregation function and an update function of the Graph SAGE network; the aggregation function adopts an attention pooling method, and assigns different weights to the feature vectors of neighboring nodes through the attention mechanism to obtain a weighted aggregate feature vector; the update function adopts a gated recursive method to perform a nonlinear transformation on the weighted aggregate feature vector and update the features of the central node.

[0043] The adjacency matrix is a commonly used matrix representation for graph structure. For a graph with N nodes, the adjacency matrix is an N×N square matrix. If there is an edge connecting node i and node j, then the element A[i, j] in the adjacency matrix equals 1; otherwise, A[i, j] equals 0. The adjacency matrix describes the direct connections between nodes in the graph and is a compact representation of the graph's structural information. In Graph SAGE networks, the adjacency matrix guides the sampling and aggregation of neighboring nodes. First-order neighbor nodes are nodes directly connected to the central node. In a graph, first-order neighbor nodes are directly connected to the central node by edges. For a central node v, its first-order neighbor node set can be represented as N(v) = {u|(u, v)∈E}, where E represents the set of edges in the graph. First-order neighbor nodes provide local structural information about the central node and reflect the nodes directly connected to the central node. Second-order neighbor nodes are nodes directly connected to the central node's first-order neighbor nodes, excluding the central node itself. They are indirectly connected to the central node by two edges. For a central node v, its second-order neighboring node set can be expressed as N(N(v)) = {w|(w,u)∈E,u∈N(v),w≠v}. The second-order neighboring nodes provide extended structural information of the central node and reflect the indirect related nodes of the central node.

[0044] Specifically, find the row vector corresponding to the central node in the adjacency matrix. Traverse each element in the row vector. If the element value is 1, the corresponding column index represents a first-order neighbor node. Collect the indexes of all first-order neighbor nodes into the first-order neighbor node set. For each node in the first-order neighbor node set, find the row vector corresponding to the node in the adjacency matrix. Traverse each element in the row vector. If the element value is 1, the corresponding column index represents a second-order neighbor node. Collect the indexes of all second-order neighbor nodes into the second-order neighbor node set and remove duplicate node indexes. Return the first-order neighbor node set and the second-order neighbor node set of the central node.

[0045] Preferably, the idea of random walk is introduced. When sampling neighborhood nodes, not only directly connected nodes but also nodes reachable through random walk are considered. The specific steps are as follows: For a given central node, its first-order neighborhood node set is determined based on the adjacency matrix. A node is randomly selected from the first-order neighborhood node set as the starting node and a random walk is performed. During the random walk, a node is chosen with a certain probability to proceed to the next neighborhood node or return to the central node. The random walk process is repeated multiple times, and the visited nodes are collected as the second-order neighborhood node set.

[0046] In this application, the Graph SAGE network is updated for K rounds using a forward propagation algorithm. In each round of updates, all nodes in the graph are traversed, and the currently traversed node is used as the central node. According to the adjacency matrix, the first-order neighbor nodes and second-order neighbor nodes of the central node are sampled to obtain their feature vectors in the previous round of updates. Then, the feature vectors of the first-order neighbor nodes and the second-order neighbor nodes are self-attention pooled respectively, and the correlation between the neighbor nodes is calculated through the multi-head self-attention mechanism to obtain the self-attention weight matrix. According to the self-attention weight matrix, the feature vectors of the first-order neighborhood and the second-order neighborhood are weighted and summed to obtain the aggregated neighborhood feature vector. Finally, the aggregated feature vectors of the first-order neighborhood and the second-order neighborhood are residually connected with the feature vector of the previous round of the central node to obtain the initial updated feature vector of the central node. In this way, by aggregating and propagating neighborhood information, the feature representation of the node can be updated to capture the role and influence of the node in the graph structure.

[0047] The forward propagation algorithm is used to recursively update the features of the nodes in the Graph SAGE network for K rounds. The forward propagation algorithm is used to recursively update the features of the nodes in the Graph SAGE network for K rounds, including: in each round of update, all nodes in the directed graph are traversed, and the currently traversed node is used as the center node; for each center node, according to the adjacency matrix of the directed graph, the first-order neighboring nodes and the second-order neighboring nodes of the center node are sampled to obtain the first-order neighboring node set and the second-order neighboring node set; according to the first-order neighboring node set and the second-order neighboring node set, the nodes in the Graph SAGE network are updated. From the node feature matrix of the last round of update of the SAGE network, the current feature vectors corresponding to the first-order neighbor nodes and the second-order neighbor nodes are extracted; the current feature vectors of the first-order neighbor nodes and the current feature vectors of the second-order neighbor nodes are respectively input into the self-attention pooling function, and the correlation between the neighbor nodes is calculated through the multi-head self-attention mechanism to generate the self-attention weight matrices of the first-order neighbor nodes and the second-order neighbor nodes; according to the self-attention weight matrices of the first-order neighbor nodes and the second-order neighbor nodes, the current feature vectors of the first-order neighbor nodes and the current feature vectors of the second-order neighbor nodes are weighted and summed to obtain the self-attention aggregation feature vectors of the first-order neighbor nodes and the second-order neighbor nodes; the self-attention aggregation feature vector of the first-order neighbor nodes and the self-attention aggregation feature vector of the second-order neighbor nodes are spliced, and the spliced aggregation feature vector is residually connected with the feature vector of the center node in the last round of update to obtain the initial updated feature vector of the center node; among them, the center node, in the graph neural network, refers to the target node that is currently undergoing feature aggregation and update. Graph SAGE updates the feature representation of a central node by aggregating information from its neighboring nodes. The aggregate vector is the result of aggregating node features in the Graph SAGE network. The features of the central node's neighboring nodes are aggregated using aggregation functions (such as averaging, pooling, and LSTM), resulting in a feature vector called the aggregate vector. The aggregate vector represents the local structural information of the node. After K rounds of updates, the final features of each node are output as the graph embedding features for the corresponding data segment.

[0048] Then, the time domain features of each data segment are input into a pyramid convolutional neural network, and the time domain modal features are extracted through multi-scale convolution and pooling operations. The frequency domain features are input into a bidirectional recurrent neural network, and the frequency domain modal features are extracted through sequence modeling. The pyramid convolutional neural network is a multi-scale convolutional neural network structure. It performs a convolution operation on the input through multiple convolution kernels with different receptive field sizes to extract features of different scales. The feature maps of different scales are then fused according to the pyramid structure to obtain a feature representation with multi-scale information. Time domain modal features refer to data features extracted from the time domain perspective. In this application, the time domain features of the data segments are learned through the pyramid convolutional neural network to extract higher-level and more abstract time domain modal features to characterize the patterns and regularities of the data in the time domain. Frequency domain modal features refer to data features extracted from the frequency domain perspective. In this application, the frequency domain features of the data segments are learned through a bidirectional recurrent neural network to extract sequence dependencies and global information in the frequency domain to obtain frequency domain modal features to characterize the patterns and regularities of the data in the frequency domain.

[0049] Figure 5This is an exemplary flow chart for obtaining multimodal fusion features according to some embodiments of this specification. Next, the graph embedding features, time domain features, frequency domain features, time domain modal features, and frequency domain modal features of each data segment are concatenated and fused through an attention mechanism to obtain multimodal fusion features. Simultaneously, a weighted fusion of the time domain modal features and the frequency domain modal features is performed to obtain cross-modal fusion features. Among them, for each data segment, the graph embedding modal features, time domain modal features, frequency domain modal features, time domain modal sub-features and frequency domain modal sub-features are input into the attention fusion module, and the importance weights of different modal features are adaptively learned through the attention mechanism, and the different modal features are weightedly fused according to the weights to obtain the multimodal fusion feature vector of the data segment; the attention fusion module includes a query, a key-value pair generator and an attention aggregator; the query receives context information and generates a query vector; the key-value pair generator receives features of different modalities and generates key-value pairs; the attention aggregator calculates the attention weight according to the query vector and the key-value pair, and performs weighted summation on the values corresponding to the keys according to the attention weights to obtain the multimodal fusion feature vector; the key-value pair generator adopts a multi-layer perceptron network to embed the graph modal features, time domain modal features, and frequency domain modal sub-features. The time domain and frequency domain modal features are input into different multi-layer perceptrons respectively to generate graph embedding modal key-value pairs, time domain modal key-value pairs and frequency domain modal key-value pairs; the time domain modal sub-features and frequency domain modal sub-features are input into two other multi-layer perceptrons respectively to generate time domain modal sub-key-value pairs and frequency domain modal sub-key-value pairs; the attention aggregator adopts the scaled dot product attention mechanism to perform dot product operations on the query vector with the graph embedding modal key, time domain modal key, frequency domain modal key, time domain modal sub-key and frequency domain modal sub-key respectively to obtain five attention scores; the five attention scores are divided by the key vector dimension under the square root, and then normalized by the Softmax function to obtain five attention weights; finally, the graph embedding modal value, time domain modal value, frequency domain modal value, time domain modal sub-value and frequency domain modal sub-value are weighted and summed according to the attention weights to obtain the multimodal fusion feature vector.The multimodal fusion feature vector is input into the gated fusion network, and the gating mechanism is used to adaptively control the flow and integration of multimodal information. After multiple layers of nonlinear transformation, an enhanced multimodal fusion feature vector is obtained. The gated fusion network adopts multiple layers of gated recursive units, each of which includes an update gate, a reset gate and a candidate state. The update gate and the reset gate receive the hidden state at the previous moment and the multimodal fusion feature vector at the current moment, respectively, and calculate the gating signal through the Sigmoid function. The candidate state receives the output of the reset gate and the multimodal fusion feature vector at the current moment, and calculates the candidate state through the hyperbolic tangent function. Finally, according to the output of the update gate, the hidden state at the previous moment and the candidate state are weightedly combined to obtain the hidden state at the current moment, which serves as the input of the next layer. The output layer of the gated fusion network receives the hidden state of the last gated recursive unit, and obtains the enhanced multimodal fusion feature vector through linear transformation and hyperbolic tangent activation function.

[0050] The attention fusion module employs a structure consisting of a query engine, a key-value generator, and an attention aggregator. It uses an attention mechanism to adaptively calculate the importance weights of features from different modalities. The query engine generates a query vector based on contextual information, while the key-value generator converts features from different modalities into key-value pairs. The attention aggregator uses the query vector and key-value pairs to calculate attention weights, and then performs a weighted summation based on the weights to obtain a multimodal fusion feature vector. Compared to simple concatenation, attention fusion can adaptively adjust the importance of modalities based on different samples and contexts, extracting more targeted fusion features. The gated fusion network employs multiple layers of gated recurrent units, using update and reset gates to control the transfer and integration of multimodal information at different times and levels. The update gate determines how much information from the previous layer is retained, while the reset gate determines how much information from the current layer is updated. Candidate states undergo nonlinear transformations based on the reset information. By stacking these multiple layers of gated recurrent units, the network can adaptively extract and enhance multimodal fusion features, uncovering deep inter-modal correlations.

[0051] Among them, latent variables refer to variables that are not directly observed in the machine learning model. They represent the potential characteristics or structure of the data by transforming or combining the observed variables. In this application, latent variables are obtained by linearly transforming the input features. The latent variables can be regarded as a compressed representation of the input features, capturing the potential patterns and relationships of the input features. The introduction of latent variables can enhance the expressiveness of the model, enabling it to learn more abstract and advanced feature representations. The scaled dot product attention mechanism is a commonly used attention mechanism used to calculate the correlation between the query vector and the key-value pair. It performs a dot product operation on the query vector and the key vector to obtain the similarity score between them, and then divides the similarity score by a scaling factor (usually the square root of the key vector dimension) to obtain a scaled attention score. The purpose of the scaling operation is to prevent the dot product result from being too large, which will lead to the gradient disappearance problem of the softmax function. Finally, the attention score is converted into an attention weight through the softmax function, which is used to perform weighted summation on the value vector to obtain the result of attention aggregation. In this application, the graph embedding features, time domain features, frequency domain features, time domain modal features and frequency domain modal features of each data segment are used as input features, latent variables are obtained through linear transformation, and the latent variables are used to generate query vectors through the attention mechanism. Then, according to the input features, the corresponding key-value pairs are generated through the multi-layer perceptron network. The scaled dot product attention mechanism is used to perform dot product operations on the query vector and the key-value pairs of each feature to obtain the corresponding attention score. The attention score is converted into attention weight by the softmax function, and the input features are weighted and summed according to the attention weight to obtain multimodal fusion features. In this way, through the attention mechanism, the features of different modalities can be adaptively fused to highlight the features that are more important and relevant to the current data segment, and obtain a comprehensive and targeted feature representation.

[0052] Among them, the time domain modal features and frequency domain modal features of each data segment are weightedly fused to obtain the cross-modal fusion features of the corresponding data segment, including the following steps: taking the time domain modal features and frequency domain modal features of each data segment as input, performing feature transformation through the fully connected layer, and obtaining the transformed time domain modal features and the transformed frequency domain modal features; wherein the weight parameters of the fully connected layer are obtained through training; the transformed time domain modal features and the transformed frequency domain modal features are spliced into feature vectors, and the cross-modal attention weights are calculated through the attention mechanism; the attention mechanism includes: converting the feature vector into a query vector, a key vector and a value vector through a linear transformation; performing a dot product operation on the query vector and the key vector to obtain the attention vector. force score, and then scale the attention score by dividing it by the square root of the key vector dimension; normalize the scaled attention score through the Softmax function to obtain the cross-modal attention weight; according to the cross-modal attention weight, perform weighted summation on the transformed time domain modal features and the transformed frequency domain modal features to obtain the cross-modal fusion feature vector; perform nonlinear transformation on the cross-modal fusion feature vector through a feedforward neural network to obtain the transformed cross-modal fusion feature vector; wherein the feedforward neural network includes at least one fully connected layer and a nonlinear activation function; the transformed cross-modal fusion feature vector is element-wise added to the cross-modal fusion feature vector through a residual connection to obtain the final cross-modal fusion feature.

[0053] For each data segment, perform an outer product operation on its multimodal fusion features and cross-modal fusion features, enhance the interactive features through row-column attention pooling and gating operations, flatten and concatenate to obtain the final feature representation vector of the data segment. Among them, for each data segment, perform a product operation on the corresponding multimodal fusion features and cross-modal fusion features, and obtain the concatenated feature vector of the corresponding data segment through feature interaction, including: for each data segment, perform an outer product operation on the corresponding multimodal fusion features and cross-modal fusion features to generate a two-dimensional interactive feature matrix; wherein the number of rows of the two-dimensional interactive feature matrix is equal to the dimension of the multimodal fusion features, and the number of columns is equal to the dimension of the cross-modal fusion features;

[0054] Perform linear transformation on the rows and columns of the two-dimensional interaction feature matrix respectively to obtain the row query matrix, row key matrix, row value matrix and column query matrix, column key matrix, column value matrix; perform matrix multiplication on the row query matrix and the transpose of the row key matrix to obtain the row attention score matrix; perform matrix multiplication on the column query matrix and the transpose of the column key matrix to obtain the column attention score matrix; perform softmax normalization on the row attention score matrix and the column attention score matrix respectively to obtain the row attention weight matrix and the column attention weight matrix; perform matrix multiplication on the row attention weight matrix and the row value matrix to obtain the row attention feature matrix; perform matrix multiplication on the column attention weight matrix and the column value matrix to obtain the column attention feature matrix.

[0055] Among them, the two-dimensional interaction feature matrix is a matrix obtained by performing an outer product operation on the multimodal fusion features and the cross-modal fusion features. The outer product operation multiplies each element of the two feature vectors by two by two to generate a two-dimensional matrix. The number of rows of the two-dimensional interaction feature matrix is equal to the dimension of the multimodal fusion features, and the number of columns is equal to the dimension of the cross-modal fusion features. It captures the interaction information between the multimodal fusion features and the cross-modal fusion features, and each element represents the interaction intensity between the two feature dimensions. Through the two-dimensional interaction feature matrix, the high-order interaction relationship between the features can be explicitly modeled to provide a richer feature representation. Gating operation is a commonly used feature enhancement technique used to control the flow of information and selectively update features. It draws on the idea of gating units (such as LSTM or GRU) and introduces gating signals to regulate the transmission and update of features. In this application, the attention pooling feature matrix is gated to obtain an enhanced attention pooling feature matrix. Specifically, the gating signal is calculated and multiplied element-by-element with the attention pooling feature matrix to obtain the enhanced feature matrix. The gating signal can be compressed to between 0 and 1 by the sigmoid function, indicating the importance of each feature dimension. Through the gating operation, the influence of the feature can be adaptively adjusted, important features can be highlighted, and unimportant features can be suppressed, thereby obtaining a more refined and effective feature representation. In this application, for each data segment, the corresponding multimodal fusion features and cross-modal fusion features are subjected to an outer product operation to generate a two-dimensional interaction feature matrix. Then, the rows and columns of the two-dimensional interaction feature matrix are subjected to attention pooling operations respectively to obtain a row attention feature matrix and a column attention feature matrix to capture the interaction pattern between the features. The row attention feature matrix and the column attention feature matrix are weightedly fused to obtain an attention pooling feature matrix. Next, the attention pooling feature matrix is gated, and the importance of the feature is adjusted by the gating signal to obtain an enhanced attention pooling feature matrix. Finally, the enhanced attention pooling feature matrix is flattened into a one-dimensional vector as the splicing feature vector of the corresponding data segment, and is fused with the multimodal fusion features and cross-modal fusion features through residual connection to obtain the final splicing feature vector. In this way, through feature interaction and gating operations, we can capture high-order relationships between multimodal features and adaptively adjust the importance of features to obtain a more comprehensive and effective feature representation.

[0056] The row pooling feature matrix and the column pooling feature matrix are summed row by row and column by column respectively to obtain the row pooling feature vector and the column pooling feature vector; the enhanced attention pooling feature matrix is adaptively selected to obtain the compressed attention pooling feature matrix; the adaptive feature selection includes: generating a feature selection vector through global average pooling and a fully connected layer; performing an outer product operation on the feature selection vector and the enhanced attention pooling feature matrix to obtain a feature selection mask matrix; performing element-by-element multiplication of the feature selection mask matrix and the enhanced attention pooling feature matrix to obtain a compressed attention pooling feature matrix; flattening the compressed attention pooling feature matrix into a one-dimensional vector as the splicing feature vector of the corresponding data segment;

[0057] Preferably, a multi-head self-attention mechanism is adopted to repeatedly perform the self-attention pooling operation K times to obtain K row pooling feature vectors and K column pooling feature vectors; wherein, the linear transformation matrices of the query matrix, key matrix and value matrix calculated each time are different; the K row pooling feature vectors and the K column pooling feature vectors are spliced respectively to obtain the spliced row pooling feature vector and the spliced column pooling feature vector; the spliced row pooling feature vector and the spliced column pooling feature vector are subjected to an outer product operation, and a pooling interaction feature matrix is obtained through feature interaction; the pooling interaction feature matrix is flattened into a one-dimensional vector as the spliced feature vector of the corresponding data segment. Specifically, after performing the outer product operation to obtain the interaction feature matrix, the self-attention pooling operation is performed on the rows and columns of the interaction feature matrix respectively. Self-attention pooling obtains the query matrix, key matrix and value matrix through linear transformation, and then calculates the attention score matrix and the attention weight matrix, and finally obtains the row pooling feature matrix and the column pooling feature matrix. This can better capture the correlation information between rows and columns and extract more discriminative features. Next, a multi-head self-attention mechanism is used to repeatedly perform K self-attention pooling operations to obtain K groups of row pooling feature vectors and column pooling feature vectors. Multi-head self-attention enhances the expressiveness of the model by computing multiple different attention maps in parallel, and is able to capture rich feature information from different subspaces. Finally, the K groups of row pooling feature vectors and column pooling feature vectors are concatenated, and then an outer product operation is performed to obtain the pooled interaction feature matrix, which is flattened into a one-dimensional vector as the final concatenated feature vector. By introducing self-attention pooling and multi-head self-attention mechanisms, the improved method can better mine the correlation information between multimodal fusion features and cross-modal fusion features, generate more discriminative and informative concatenated feature vectors, and thus improve the performance of power data analysis.

[0058] Finally, the final feature representation vectors of each data segment are input into a multilayer perceptron network. Deep features are obtained through multiple layers of nonlinear transformations, and a gradient boosting tree model is constructed for training and prediction, outputting analysis results for the power data. The weight parameters of the multilayer perceptron network are obtained through training, and the dataset is divided into training, validation, and test sets. The training set is used to train the weight parameters of the multilayer perceptron network, the validation set is used to adjust hyperparameters and evaluate model performance, and the test set is used for final model evaluation. The weight parameters of the multilayer perceptron network are randomly initialized, using methods such as Xavier initialization and He initialization, to ensure an appropriate initial weight distribution and avoid vanishing or exploding gradients. The concatenated feature vectors of the training data are input into the multilayer perceptron network. Forward propagation is performed through multiple fully connected layers and nonlinear activation functions to obtain the predicted feature representation of the corresponding data segment. The predicted feature representation is compared with the true feature representation, and a loss function is calculated. Common loss functions include mean squared error loss and cross entropy loss, and the appropriate loss function should be selected based on the characteristics of the task. The gradient of the loss function with respect to the weight parameters of the multilayer perceptron network is calculated using the backpropagation algorithm. Using the chain rule, the gradient of the loss function is propagated layer by layer to each layer of the network, obtaining the gradient value of each weight parameter. An optimization algorithm, such as stochastic gradient descent (SGD) or Adam, is used to update the weight parameters of the multilayer perceptron network based on the calculated gradient values. The optimization algorithm controls the step size and direction of the weight update by adjusting hyperparameters such as the learning rate and momentum. The training data is iterated over multiple epochs until the model converges or the predetermined number of iterations is reached. After each epoch, the model performance is evaluated using a validation set. Based on the validation results, hyperparameters are adjusted or early stopping is performed to prevent overfitting.

[0059] Preferably, regularization techniques can be introduced to prevent overfitting of the multilayer perceptron network. L1 and L2 regularization encourage the network to learn sparse or small weights by adding the L1 norm or L2 norm of the weights to the loss function, reducing the risk of overfitting. Dropout technology improves the generalization ability of the network by randomly dropping a portion of neurons during training. Batch Normalization can also be considered to normalize the features of each batch, accelerate network convergence, and improve training stability. The choice of activation function is crucial to the performance of the multilayer perceptron network. Consider using adaptive activation functions such as PReLU or Swish to allow the network to autonomously learn the parameters of the activation function and improve the adaptability of feature transformations. Introduce an attention module in the middle layer of the network to generate attention weights by calculating the correlation between features and perform a weighted combination of features. Attention mechanisms include additive attention, dot product attention, and self-attention.

[0060] The final feature representation vectors of each data segment are combined into a feature matrix as the input of the gradient boosting tree model. Each data segment corresponds to a sample, and its final feature representation vector corresponds to the characteristics of the sample. The output of the gradient boosting tree model is determined according to the task type of power data analysis (such as fault diagnosis, load forecasting, etc.). For classification tasks, the output is the fault category; for regression tasks, the output is the predicted value. The base learner of the gradient boosting tree model is constructed using the CART regression tree. The number of base learners is set to N, the maximum depth of each regression tree is set to D, and the minimum number of samples for leaf nodes is set to M. These hyperparameters can be tuned through methods such as cross-validation. Initialize the first regression tree, and obtain the tree structure and the predicted value of the leaf node by minimizing the mean square error loss function based on the characteristics and labels of the training samples. Calculate the residual of each training sample, that is, the difference between the true value of the sample and the predicted value of the current model. Use the residual as the new target value to fit the next regression tree. Repeat the iteration to generate N regression trees. In each iteration, a new regression tree is fitted based on the residuals from the previous round, and the predictions of the new tree are added to the predictions of the current model, continuously reducing the residuals and gradually approaching the true value. When generating each regression tree, feature importance assessment and feature subsampling are used to reduce the risk of overfitting. Feature importance is assessed by calculating the gain of each feature during the tree split. At each split, a subset of features is randomly sampled as candidate features to introduce randomness. For the test sample, its final feature representation vector is input into the trained gradient boosted tree model. Prediction is performed tree by tree, and the predictions of each tree are weighted averaged to obtain the final predicted output. The performance of the gradient boosted tree model on the test set is evaluated using metrics such as precision, recall, and mean squared error to verify the model's generalization ability. The trained gradient boosted tree model is then used to analyze and predict new power data. The feature representation of the new data is input into the model to obtain predictions, enabling intelligent monitoring and decision support for the power system. By constructing a gradient boosted tree model, the deep feature representation of the data segments is fully utilized to learn the complex nonlinear relationships between features and outputs. Gradient boosting trees continuously improve prediction accuracy through iterative optimization and integration of multiple decision trees, demonstrating strong nonlinear modeling capabilities and generalization performance. Furthermore, feature importance assessment and subsampling strategies help reduce overfitting and enhance model robustness. This technical solution can efficiently mine the key features and inherent patterns of power data, providing reliable decision support for the safe and stable operation of power systems.

Claims

1. A power data analysis method based on data fusion, comprising: Acquire power data, and divide the acquired power data into multiple data segments; Perform time-frequency analysis on the divided data segments to extract the time domain features and frequency domain features of each data segment; Construct a directed graph based on the divided data segments; the nodes in the directed graph represent the data segments, and the edges in the directed graph represent the correlation between the data segments; The constructed directed graph is used as input and encoded using the Graph SAGE network. The features of the central node are updated by aggregating the features of the neighboring nodes. The aggregated vector of the node corresponding to each data segment is obtained as the graph embedding feature of the corresponding data segment. A pyramid convolutional neural network is used to extract the time domain features of each data segment to obtain the time domain modal features; A bidirectional recurrent neural network is used to extract the frequency domain features of each data segment to obtain the frequency domain modal features; The graph embedding features, time domain features, frequency domain features, time domain modal features and frequency domain modal features of each data segment are spliced to obtain multimodal fusion features; The time domain modal features and frequency domain modal features of each data segment are weightedly fused to obtain the cross-modal fusion features of the corresponding data segment; For each data segment, the corresponding multimodal fusion feature and cross-modal fusion feature are multiplied, and the concatenated feature vector of the corresponding data segment is obtained through feature interaction; The concatenated feature vector of each data segment is input into the multi-layer perceptron network for nonlinear transformation to obtain the final feature representation of the corresponding data segment; According to the final feature representation of each data segment, a gradient boosting tree model is constructed as a prediction model for power data analysis.

2. The power data analysis method based on data fusion according to claim 1, characterized in that: Obtain the time domain and frequency domain features of each data segment, including: Perform a short-time Fourier transform on each data segment and convert the data segment into a time-frequency matrix by setting the window length and overlap rate parameters of the sliding window; the rows of the time-frequency matrix represent the time dimension and the columns represent the frequency dimension; Extract the statistical features of the time-frequency matrix of each data segment as time domain features; the statistical features include mean, variance and kurtosis; The time-frequency matrix of each data segment is subjected to wavelet transform and Hilbert-Huang transform to obtain the wavelet frequency domain features and Hilbert frequency domain features of the corresponding data segment, and the frequency domain features are obtained based on the wavelet frequency domain features and Hilbert frequency domain features of the corresponding data segment.

3. The power data analysis method based on data fusion according to claim 2, characterized in that: Get wavelet frequency domain features, including: Performing continuous wavelet transform on the time-frequency matrix, and converting the time-frequency matrix into a wavelet coefficient matrix through the wavelet function; the rows of the wavelet coefficient matrix represent the time dimension, and the columns represent the scale dimension; the wavelet function is one of the Morlet wavelet, Mexican Hat wavelet, Daubechies wavelet and Symlet wavelet; Perform energy spectrum analysis on the wavelet coefficient matrix and obtain the wavelet energy spectrum of the corresponding data segment by calculating the square sum of each column of the wavelet coefficient matrix; The maximum energy scale, energy center scale and energy dispersion of the wavelet energy spectrum are extracted as the wavelet frequency domain features of the corresponding data segment.

4. The power data analysis method based on data fusion according to claim 3 is characterized in that: Obtain frequency domain features, including: Perform marginal spectrum analysis on the time-frequency matrix and obtain the Fourier frequency domain marginal spectrum of the corresponding data segment by summing each column of the time-frequency matrix; Perform Hilbert-Huang transform on the Fourier frequency domain marginal spectrum, and extract the instantaneous frequency and instantaneous amplitude by solving the analytical signal on the frequency domain marginal spectrum; Extract the statistical features of instantaneous frequency and instantaneous amplitude to obtain the Hilbert frequency domain features of the corresponding data segment; The wavelet frequency domain features and Hilbert frequency domain features of the data segment are spliced as the frequency domain features of the corresponding data segment.

5. The power data analysis method based on data fusion according to claim 1, characterized in that: Get the graph embedding features of the corresponding data segment, including: Initialize the Graph SAGE network based on the constructed directed graph, and use the features of each node in the directed graph as the initial feature vector; the features of each node include the time-frequency statistical features, wavelet frequency domain features, and Hilbert frequency domain features of the corresponding data segment; Set the aggregation function and update function of the Graph SAGE network; the aggregation function uses attention pooling to assign different weights to the feature vectors of neighboring nodes through the attention mechanism to obtain a weighted aggregate feature vector; the update function uses a gated recursive method to perform a nonlinear transformation on the weighted aggregate feature vector and update the features of the central node; Through the forward propagation algorithm, the features of the nodes in the Graph SAGE network are recursively updated for K rounds; After K rounds of updates, the final features of each node are output as the graph embedding features of the corresponding data segment.

6. The power data analysis method based on data fusion according to claim 5, characterized in that: Through the forward propagation algorithm, the features of the nodes in the Graph SAGE network are recursively updated for K rounds, including: In each round of update, all nodes in the directed graph are traversed, and the currently traversed node is used as the central node; For each central node, according to the adjacency matrix of the directed graph, the first-order neighboring nodes and the second-order neighboring nodes of the central node are sampled to obtain the first-order neighboring node set and the second-order neighboring node set; According to the first-order neighbor node set and the second-order neighbor node set, the current feature vectors corresponding to the first-order neighbor node and the second-order neighbor node are extracted from the node feature matrix updated in the last round of the Graph SAGE network; The current feature vector of the first-order neighbor node and the current feature vector of the second-order neighbor node are input into the self-attention pooling function respectively. The correlation between the neighbor nodes is calculated through the multi-head self-attention mechanism to generate the self-attention weight matrix of the first-order neighbor node and the second-order neighbor node. According to the self-attention weight matrix of the first-order neighborhood and the second-order neighborhood, the current feature vector of the first-order neighborhood node and the current feature vector of the second-order neighborhood node are weighted summed to obtain the self-attention aggregate feature vector of the first-order neighborhood and the second-order neighborhood; The self-attention aggregated feature vector of the first-order neighborhood and the self-attention aggregated feature vector of the second-order neighborhood are spliced together, and the spliced aggregated feature vector is residually connected with the feature vector of the previous round of update of the central node to obtain the initial updated feature vector of the central node.

7. The power data analysis method based on data fusion according to claim 6, characterized in that: Through the forward propagation algorithm, the features of the nodes in the Graph SAGE network are recursively updated for K rounds, including: The initial updated feature vector of the central node is input into the multi-layer gated recursive unit, and after multiple layers of nonlinear transformation, the final updated feature vector of the central node is obtained; The final updated feature vector of the central node is added to the feature vector of the previous round of the central node through the residual connection to obtain the updated feature vector of the central node in this round; The updated feature vector of the central node in this round is stored in the corresponding position of the node feature matrix of the Graph SAGE network in this round of update.

8. The power data analysis method based on data fusion according to any one of claims 1 to 7, characterized in that: Obtain multimodal fusion features, including: The graph embedding features, time domain features, frequency domain features, time domain modal features and frequency domain modal features of each data segment are used as input features; Perform linear transformation on the input features to obtain latent variables; Generate query vectors from latent variables through attention mechanism; Based on the input features, a multi-layer perceptron network is used to generate graph embedding feature key-value pairs, time domain feature key-value pairs, frequency domain feature key-value pairs, time domain modal feature key-value pairs, and frequency domain modal feature key-value pairs. Using the scaled dot product attention mechanism, the query vector and the graph embedding feature key-value pairs, time domain feature key-value pairs, frequency domain feature key-value pairs, time domain modal feature key-value pairs, and frequency domain modal feature key-value pairs are respectively subjected to dot product operations to obtain the corresponding attention scores; According to the attention score, the corresponding attention weight is obtained through the softmax function; According to the attention weights, the input features are weighted and summed to obtain multimodal fusion features.

9. The power data analysis method based on data fusion according to claim 1, characterized in that: For each data segment, the corresponding multimodal fusion feature and cross-modal fusion feature are multiplied, and the concatenated feature vector of the corresponding data segment is obtained through feature interaction, including: For each data segment, the outer product operation is performed on the corresponding multimodal fusion features and cross-modal fusion features to generate a two-dimensional interaction feature matrix; the number of rows of the two-dimensional interaction feature matrix is equal to the dimension of the multimodal fusion features, and the number of columns is equal to the dimension of the cross-modal fusion features; Perform attention pooling operations on the rows and columns of the two-dimensional interaction feature matrix to obtain row attention feature matrix and column attention feature matrix; Perform weighted fusion of the row attention feature matrix and the column attention feature matrix to obtain the attention pooling feature matrix; Perform a gating operation on the attention pooling feature matrix to obtain the enhanced attention pooling feature matrix; Flatten the enhanced attention pooling feature matrix into a one-dimensional vector as the concatenated feature vector of the corresponding data segment; The spliced feature vector is fused with the multimodal fusion feature and the cross-modal fusion feature through residual connection to obtain the final spliced feature vector of the data segment.

10. The power data analysis method based on data fusion according to claim 9, characterized in that: The concatenated feature vector of each data segment is input into the multi-layer perceptron network for nonlinear transformation to obtain the final feature representation of the corresponding data segment, including: The concatenated feature vectors of each data segment are sequentially input into multiple fully connected layers of a multilayer perceptron network, and forward propagated through a nonlinear activation function to obtain the final feature representation of the corresponding data segment; wherein, the weight parameters of the multilayer perceptron network are obtained through training.

Citation Information

Patent Citations

  • A method and system for analyzing power data based on data fusion

    CN117149846B