Power construction deviation degree diagnosis method based on multi-modal time sequence data fusion

Through multi-channel synchronous sampling and long short-term memory network analysis, combined with the isolation forest algorithm and multi-layer perceptron model, the problem of integrating and analyzing multimodal time series data of the power system was solved, accurate anomaly identification and classification of the power system was achieved, and the reliability and safety of the system were improved.

CN120724345APending Publication Date: 2025-09-30GUANGDONG YUNFENG POWER INSTALLATION CO LTD

Patent Information

Application Number
CN202511055775.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-30
Publication Date
2025-09-30

AI Technical Summary

Technical Problem

Existing technologies have difficulty in effectively integrating and analyzing multimodal time series data of power systems, resulting in insufficient sensitivity in abnormal pattern recognition, affecting the rapid response capability of the power system and data encoding efficiency.

Method used

Through multi-channel synchronous sampling, multi-dimensional time series data of voltage, current and frequency are obtained, a unified feature space is constructed, and the time dependence of electrical variables is analyzed using long and short-term memory networks to capture dynamic change trends. The isolation forest algorithm and multi-layer perceptron model are used for anomaly detection and classification.

Benefits of technology

Accurate identification and classification of multimodal data are achieved, improving the operational reliability and safety of the power system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120724345A_ABST
    Figure CN120724345A_ABST
Patent Text Reader

Abstract

The invention relates to a power construction deviation degree diagnosis method based on multi-modal time series data fusion, and the method comprises the steps: collecting voltage, current and frequency data through a multi-channel synchronous sampling technology, filling missing data through cubic spline interpolation, and constructing an initial data matrix; based on a hierarchical feature extraction technology, mapping the voltage frequency domain features and the current time domain statistical features to a unified feature space through canonical correlation analysis, and generating a multi-modal feature vector with a time sequence tag in combination with a sliding window; analyzing the dynamic trend of the electrical variable under multiple time scales by adopting a long short-term memory network, and capturing a key time point through an attention mechanism; a sudden change point and a stationary section are defined, an isolated forest algorithm is combined to detect an abnormal point location deviating from a trajectory, anomaly is classified as transient disturbance or continuous deviation through a multi-layer perceptron, the evolution trend of regional continuous deviation is predicted, and the key problem of the power deviation degree diagnosis capability is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of electric variable technology, and in particular to a method for diagnosing deviation of electric power construction by fusion of multimodal time series data. Background Art

[0002] Research on multimodal time series data is of vital importance in the operation and management of power systems. This field is directly related to the stability and security of power systems. Through real-time measurement and analysis of electrical variables, potential anomalies and deviations can be detected in a timely manner to ensure the reliability of power supply. However, current research and application methods still have obvious shortcomings in processing complex time series data, especially in terms of efficient data representation and accurate identification of abnormal patterns. There is a general problem of finding a balance between accuracy and efficiency. In-depth analysis reveals that existing methods often fail to fully exploit the inherent patterns of power system time series data, resulting in insensitive identification of abnormal patterns. This is particularly true when dealing with multimodal data, which lacks effective integration and analysis methods. A deeper problem lies in the inadequacy of current methods in capturing the temporal characteristics of changes in electrical variables, which directly impacts the ability to rapidly respond to transient power system behavior. The intertwined effects of inefficient data encoding and the lack of temporal information make it difficult to accurately diagnose deviations in the face of complex dynamic changes. The solution proposed by the present invention to the shortcomings of the above-mentioned prior art is: to obtain multi-dimensional time series data of voltage, current and frequency through multi-channel synchronous sampling, complete time alignment, map different modal data into a unified feature space, and use long short-term memory networks to analyze the time dependence of electrical variables and capture dynamic change trends. Summary of the Invention

[0003] In order to solve the problems existing in the above-mentioned prior art, the purpose of this application is to provide a power construction deviation diagnosis method based on multimodal time series data fusion.

[0004] The present application discloses a method for diagnosing deviation of electric power construction using multimodal time series data fusion, comprising: S101 uses multi-channel synchronous sampling technology to acquire multi-dimensional time series data streams such as voltage, current, and frequency from sensors of different modalities, constructs an original multimodal data set, and completes timestamp alignment to generate a time-uniform initial data matrix for subsequent feature mapping processing; S102: Based on the heterogeneous characteristics of multimodal data, a hierarchical feature extraction method is applied to the initial data matrix to map the data streams of different modalities into a unified feature space, and a multimodal feature vector set is constructed for subsequent temporal dependency analysis, while retaining the time variation information of each modality; S103, for the multimodal feature vector set, a long short-term memory network is used to perform time series modeling, analyze the time dependence of electrical variables at different time scales, capture dynamic change trends, and generate a time series feature sequence containing time context for subsequent change pattern extraction; S104: Using statistical analysis methods to extract differential features between mutation points and plateaus, the time series feature sequences are analyzed. Mutation points are defined as points where the rate of change of values ​​exceeds a preset threshold, and plateaus are defined as intervals where the rate of change is stable. This generates a data set of multi-time-scale variation patterns of electrical variables for subsequent abnormal signal screening. S105: Applying a preliminary screening method based on a preset threshold to the multi-time-scale change pattern data set, identifying potential abnormal points, analyzing their initial trigger conditions and fluctuation characteristics, and generating a dynamic change data set containing potential abnormal signals for subsequent deep anomaly detection; S106: Using the Isolation Forest algorithm to perform anomaly detection on the dynamically changing data set, the algorithm identifies abnormal points that deviate from the normal operating trajectory based on the data distribution in the multidimensional feature space, and generates an abnormality tag sequence for subsequent anomaly type classification. S107: For the abnormal marking sequence, if the deviation of the abnormal point exceeds the preset threshold, it is classified through the multi-layer perceptron model. Based on the fluctuation characteristics of the mutation point and the time context information, it is determined whether the abnormality is a transient disturbance or a continuous deviation, and the abnormal type label is generated.

[0005] Preferably, in step S101, multi-channel synchronous sampling is used to obtain multimodal time series data such as voltage, current, and frequency. After timestamp alignment and interpolation to fill in the gaps, an initial matrix is ​​constructed, normalized and features are extracted, and PCA dimensionality reduction is performed and classified mapping is performed. Finally, the data is integrated into a multimodal fusion feature matrix through feature splicing technology.

[0006] Preferably, in the step S102, the multimodal data stream is separated by layered processing, and features are extracted layer by layer and mapped to a unified space to generate a feature vector set containing time series information; the time series features are corrected, and the LSTM mines the time series patterns; the features and time series information are integrated to construct the final multimodal time series feature description, and a structured feature data stream is generated for storage and backup.

[0007] Preferably, in the step S103, preprocessing, standardization and noise reduction are performed to generate a feature vector set; LSTM time series modeling is used to analyze the time dependence of electrical variables; fluctuation patterns are extracted to determine the change trend sequence; a time series feature set is constructed in combination with the time context; abnormal fluctuation points are screened and corrected; and a final time series feature representation is generated for subsequent analysis.

[0008] Preferably, in step S104, the time series features are statistically preliminarily decomposed to obtain the change rate to identify the mutation point and the stable interval; the sliding window is segmented to quantify the difference features; the multi-time scale analysis clusters the change pattern; the logistic regression is performed for secondary verification, and finally the abnormal signal is classified.

[0009] Preferably, in the step S105, hierarchical extraction is performed to construct a multi-dimensional feature matrix; threshold screening is performed to mark suspected abnormal points; trigger conditions and fluctuation characteristics are analyzed to obtain a fluctuation feature description set; SVM classification is performed to obtain an abnormal signal candidate set; trend analysis is performed to determine the final abnormal signal data set; feature mapping relationships are constructed to quantify correlations to obtain a dynamically changing feature matrix; cluster analysis is performed to group and determine priorities to obtain an abnormal signal priority list.

[0010] Preferably, in step S106, the dynamic data set is identified as anomaly and a label sequence is generated by the isolation forest algorithm; multidimensional features are extracted to construct an initial matrix; spatial distribution is analyzed to determine the anomaly distribution characteristics; preliminary clustering is performed to determine the aggregation trend; if there is significant aggregation, the grouping is refined to determine the potential category; associated business logic is used to generate a final anomaly classification sequence; and structured data is generated for subsequent calls.

[0011] Preferably, in step S107, the threshold is compared to screen abnormal points; the mutation fluctuation characteristics and time context are extracted to form a vector group; the multi-layer perceptron classifies transient disturbances and persistent deviations; a data set of abnormal type labels is generated; persistent deviation data is extracted to construct a feature matrix; a regression model is used to analyze the evolution trend; and structured prediction data is generated for subsequent analysis.

[0012] The method for diagnosing deviations of power construction using multimodal time series data fusion described in the present application has the advantages of acquiring multi-dimensional time series data such as voltage, current and frequency through multi-channel synchronous sampling, constructing the original multimodal data set and merging it to complete time alignment. Subsequently, hierarchical feature extraction is used to map different modal data into a unified feature space, and a long-short-term memory network is used to analyze the time dependence of electrical variables and capture dynamic change trends. The present invention further extracts the difference characteristics between mutation points and stable segments, identifies potential abnormal points, and detects anomalies that deviate from the normal operating trajectory through an isolation forest algorithm. Finally, based on the fluctuation characteristics and time context information, the anomaly type is judged and the evolution trend is predicted. This method can effectively fuse multimodal data, realize accurate identification and classification of power system anomalies, and improve the reliability and safety of system operation. BRIEF DESCRIPTION OF THE DRAWINGS

[0013] Figure 1 This is the process of a multi-modal time series data fusion power construction deviation diagnosis method described in this application Figure 1 ; Figure 2This is the process of a multi-modal time series data fusion power construction deviation diagnosis method described in this application Figure 2 . DETAILED DESCRIPTION

[0014] like Figure 1-Figure 2 As shown, the electric power construction deviation diagnosis method based on multimodal time series data fusion described in this application includes the following steps: like Figure 1-Figure 2 As shown, in step S101, multi-dimensional time series data streams such as voltage, current and frequency are obtained from sensors of different modes through multi-channel synchronous sampling technology, the original multimodal data set is constructed, and the timestamp alignment is completed to generate a time-uniform initial data matrix for subsequent feature mapping processing.

[0015] Furthermore, in step S101, multi-dimensional time series data streams such as voltage data, current data, and frequency data are acquired from sensor groups of different modalities through multi-channel synchronous sampling technology to construct an original multimodal data set and obtain preliminary acquisition results; Based on the preliminary collection results, perform timestamp alignment on the time series data, use the preset time synchronization algorithm to generate standardized data streams under a unified time axis, and determine the aligned time series; If there are missing data or abnormal points in the aligned time series, the missing parts are filled by interpolation method to obtain the complete time series data stream and determine whether the data integrity meets the preset standards; According to the complete time series data stream, the initial matrix is ​​constructed, and the data of different modes are normalized by data standardization to obtain the standardized data matrix; Through the standardized data matrix, multi-dimensional features are extracted, and the principal component analysis method is applied to reduce the dimension of the data matrix to obtain the feature set after dimensionality reduction; Based on the feature set after dimensionality reduction, classification mapping is performed on the data features of different modalities to generate classified feature subsets and determine the feature mapping results; Through the classified feature subsets, a multimodal fusion data structure is constructed, and the feature splicing technology is used to integrate the feature information of different modalities to obtain the final multimodal fusion feature matrix.

[0016] Specifically, in step S101, data is collected synchronously from the voltage sensor, current sensor, and frequency sensor at a sampling rate of 1 kHz using multi-channel synchronous sampling technology to construct an original multimodal data set including voltage (0-10 V), current (0-5 A), and frequency (50-60 Hz); A time synchronization algorithm based on the PTP protocol is used to align the timestamps of each sensor data and generate a standardized data stream with a time error of less than 1ms. If missing points or outliers are detected (current suddenly drops to 0A), cubic spline interpolation is used to fill the missing segments to ensure data continuity; Perform Z-score standardization on the complete time series data to make the mean of voltage, current and frequency data 0 and the variance 1, and construct an initial data matrix of 1000×3; The PCA algorithm is used to reduce the matrix dimension, retaining 95% of the variance contribution rate and reducing the feature dimension from 3 to 2; The features after dimensionality reduction were classified by K-means clustering (k=3) and divided into three feature subsets: voltage stable area, current fluctuation area and frequency abnormal area; The three feature subsets are integrated into a 1000×6 fusion matrix using horizontal splicing, and local features are extracted layer by layer using CNN and mapped into a 128-dimensional unified feature space. Finally, the temporal dependency is modeled through the LSTM network, and the time evolution law of the feature vector is analyzed with a sliding window (length 50), and the feature state sequence with timestamp is output.

[0017] In one embodiment, in the multi-channel synchronous sampling technology of step S101, the sensor configuration uses a voltage sensor (range 0-10V, accuracy ±0.1%), a current sensor (range 0-5A, accuracy ±0.2%), and a frequency sensor (range 50-60Hz, resolution 0.01Hz), and synchronously collects data at a sampling rate of 1kHz; In the time synchronization mechanism, based on the IEEE 1588 Precision Time Protocol (PTP), the master clock and each sensor's sub-clock are synchronized via Ethernet, with the timestamp alignment error controlled within ±1ms. In the specific implementation, hardware timestamp correction and software compensation algorithms are used to ensure the time consistency of cross-modal data. In data integrity processing, if missing data is detected (for example, the current suddenly drops to 0A), cubic spline interpolation is used to fill the missing segment. The interpolation formula is: , Among them, the coefficients ai, bi, ci, and di are solved by the boundary conditions of adjacent data points to ensure that the curve is smooth and continuous; Character description: ai, bi, ci, di: coefficients of the interpolation polynomial, solved by the boundary conditions of adjacent data points (such as continuity, first-order derivative continuity, second-order derivative continuity) to ensure the smoothness of the interpolation curve, xi, xi+1: the horizontal coordinates of the adjacent data points, defining the interpolation interval, x: the horizontal coordinate of the data point to be filled; In data standardization and dimensionality reduction, Z-score standardization performs normalization on voltage V, current I, and frequency f data. The formula is: , Among them, μ is the mean, σ is the standard deviation, and the mean of the processed data is 0 and the variance is 1; Character description: V, I, f: original voltage (unit: volt, V), current (unit: ampere, A), frequency (unit: Hertz, Hz), μV, μI, μf: mean of voltage, current, frequency, σV, σI, σf: standard deviation of voltage, current, frequency, V norm ,I norm ,f norm : dimensionless data after standardization; In principal component analysis (PCA), the dimensionality of the standardized 1000×3 data matrix was reduced to retain 95% of the variance contribution, and the feature dimension was reduced from 3 to 2. Specifically, the eigenvectors of the covariance matrix were calculated through singular value decomposition (SVD), and the first two principal components were selected to construct the reduced dimension matrix. In the clustering and feature splicing of feature mapping and fusion, the K-means algorithm (k=3, maximum number of iterations 100) is used to divide the dimensionality-reduced features into three categories: voltage stability area, current fluctuation area, and frequency anomaly area. Feature splicing uses horizontal splicing technology to integrate the three feature subsets into a 1000×6 fusion matrix. In deep learning feature extraction, a convolutional neural network (CNN, structure is Conv1D(64)-MaxPooling(2)-Flatten) is used to extract local features, map them to a 128-dimensional unified space, and model the temporal dependency through an LSTM network (64 hidden units, 50 time steps).

[0018] like Figure 1-Figure 2 As shown, in step S102, for the initial data matrix, based on the heterogeneous characteristics of multimodal data, a hierarchical feature extraction method is applied to map the data streams of different modes to a unified feature space, and a multimodal feature vector set is constructed for subsequent time series dependency analysis, while retaining the time change information of each mode.

[0019] Furthermore, in step S102, for the initial data, a multimodal data source is obtained, and the data matrix is ​​preliminarily decomposed by a layered processing method to separate data streams of different modes, thereby obtaining a preliminarily separated multimodal data stream set; Based on the preliminarily separated multimodal data stream set, hierarchical feature extraction technology is used to perform layer-by-layer feature decomposition for heterogeneous characteristics, map different modal data streams to a unified space, and determine a unified feature representation form; By unifying the feature representation in the space, we construct a multimodal feature vector and combine the time variation information of each modality to generate a feature vector set containing time series information. For the feature vector set, the correlation between time variation and temporal dependency is analyzed. If the time variation exceeds a preset threshold, the feature vector is time-corrected to obtain a corrected temporal feature set. Based on the corrected temporal feature set, the long short-term memory network model is used to deeply mine temporal dependencies, extract hidden temporal patterns, and determine potential temporal correlation features. By extracting the temporal correlation features, we integrate the multimodal feature vector set and the temporal pattern information to construct a comprehensive feature representation and determine the final multimodal temporal feature description. Based on the final multimodal time series feature description, a structured feature data stream is generated, and feature storage and call are performed for subsequent analysis tasks to obtain a callable feature library.

[0020] Specifically, in step S102, a hierarchical feature extraction method based on modal type is used for the initial data matrix, such as using wavelet transform to extract frequency domain features (such as 5-layer decomposition of Daubechies-4 wavelet basis) for the voltage data stream, and applying time domain statistical features (mean, variance, peak-to-peak value) to the current data stream to separate the multi-modal data stream set; Through hierarchical feature extraction technology, the voltage frequency domain features are reduced to 10 dimensions using PCA, and the current time domain features are mapped to 3 dimensions using t-SNE to generate intermediate feature representations; The canonical correlation analysis (CCA) algorithm is used to align the voltage PCA features and the current t-SNE features in a unified space to construct a 128-dimensional unified feature representation. Combining the 1ms timestamp information of each mode, a sliding window (window length 500ms, step size 100ms) is used to splice the voltage-current features to generate a multimodal feature vector set with time series labels; Calculate the Euclidean distance of feature vectors in adjacent time windows. If the distance exceeds the threshold of 0.7, linear interpolation is used to perform time series correction on the feature vectors. Use an LSTM network (64 hidden units, 10 time steps) to process the corrected features, extract the weight distribution of key time points (such as the weight of 0.85 at time t-3) through the attention mechanism, and identify temporal correlation features; The 128-dimensional temporal features output by LSTM are weighted and fused with the original multimodal features (weight coefficient α = 0.6) to generate a 256-dimensional comprehensive feature representation; Use Apache Parquet columnar storage format to store structured feature data, partitioned by time (one file per hour) and stored in HDFS; Build a feature index based on Elasticsearch, set the modality type (voltage / current) and time range (2023-01-01T00:00:00Z to 2023-01-02T00:00:00Z) as the joint primary key, and implement millisecond-level feature retrieval.

[0021] In one embodiment, in the multimodal hierarchical processing of step S102, the voltage frequency domain features: use Daubechies-4 wavelet basis to perform 5-layer wavelet decomposition on the voltage data, extract the detail coefficients (D1-D5) and approximation coefficient (A5) of each layer, calculate the energy entropy as the frequency domain features, the current time domain features: calculate the mean of the current data , forming a 3D time domain feature vector; In feature alignment and fusion: the voltage PCA features (10 dimensions) and the current t-SNE features (3 dimensions) are aligned in a 128-dimensional space through canonical correlation analysis (CCA). The formula is: , Among them, X v and X i are the voltage and current characteristic matrices respectively, w v and w i is the projection weight; Character Description: X v ,X i : characteristic matrices of voltage and current (dimensions are 1000×10 and 1000×3 respectively), w v ,w i : projection weight vector, used to map voltage and current features to a unified space, Corr: Pearson correlation coefficient, measuring the correlation between features after projection; In temporal dependency modeling, the LSTM network configuration is as follows: 64 hidden units, 10 time steps, a dropout rate of 0.2, and an Adam optimizer (learning rate of 0.001). The time point weights are calculated using the attention mechanism, using the formula: , Among them, h t is the LSTM hidden state, W and b are trainable parameters; Character description: h t : The hidden state of LSTM at time step t (dimension 64), W,b: trainable parameter matrix and bias term, α t : The attention weight at time step t, used to weight the key time points, T: the length of the time window (e.g. 10); In feature storage and retrieval, Apache Parquet columnar storage is used, and data is stored in HDFS by time partition (1 file per hour). A joint index (modality type + time range) is established through Elasticsearch to achieve millisecond-level queries.

[0022] like Figure 1-Figure 2 As shown, in step S103, a long short-term memory network is used to perform time series modeling on the multimodal feature vector set, analyze the time dependence of electrical variables at different time scales, capture dynamic change trends, and generate a time series feature sequence containing time context for subsequent change pattern extraction.

[0023] Furthermore, in step S103, a data preprocessing method is used to standardize and reduce noise on the original data based on the multimodal features and vector set, thereby generating a unified feature vector set. Through the generated feature vector set, the long short-term memory network is used to perform time series modeling on the data, analyze the performance of the electrical variable values ​​at different time scales, and obtain the time-dependent characteristics; According to the time-dependent characteristics, dynamic changes and trends are analyzed, the fluctuation patterns of data at different time points are extracted, and the trend sequence is determined; For the change trend sequence, combined with the time context information, a feature representation containing the temporal relationship is constructed to obtain the temporal feature set; If there are abnormal fluctuations in the time series feature set, the location and impact range of the abnormal feature points are determined by comparing them with the preset threshold range; After obtaining the location information of the abnormal feature points, the feature reconstruction method is used to correct the time series features to obtain the optimized feature sequence; The optimized feature sequence is combined with the time context and change trend to generate the final time series feature representation and determine the data basis for subsequent analysis.

[0024] Specifically, in step S103, an LSTM network (hidden layer dimension 128, time step 10) is used to perform time series modeling on the multimodal feature vector set. The time dependency of the electrical variables at the scales of 1 minute, 1 hour, and 1 day is calculated, and the dynamic change trend is captured through a gating mechanism. The 256-dimensional time series feature sequence containing the time context is output. Based on this sequence, a sliding window (window size 5, step size 1) is used to extract the fluctuation pattern of the electric variable, and the mean, variance and range are calculated based on the timestamp information to generate a change trend sequence; Through trend sequences, the Pearson correlation coefficient is used to analyze the association between dynamic change patterns and time dependencies, and a 128-dimensional feature representation that integrates temporal relationships is constructed to form a temporal feature set. If the fluctuation amplitude of a feature point in the set exceeds the preset threshold ±2σ, it is marked as an outlier and its position and influence radius are recorded; Based on the location of the outlier point, linear interpolation or weighted average based on neighboring points is used to reconstruct the features and output the corrected smooth sequence; The optimized sequence is concatenated with multimodal features (temperature, vibration signals) in the feature space and fused through a fully connected layer (dimension 64) to generate the final time series feature representation; Using hierarchical feature extraction (CNN+Attention), we mapped the different modal data into a unified space (dimension 32), calculated the cosine similarity between the modalities, and determined the comprehensive feature representation. By integrating features and using the self-attention mechanism to analyze the association between temporal dependencies and multimodal features, it extracts hidden patterns (such as electricity consumption patterns with a 24-hour cycle) and outputs a structured data stream. Finally, the feature library is stored in Apache Parquet format and indexed by time partition for subsequent analysis and calls.

[0025] In one embodiment, in the multi-time scale analysis of step S103, the time window is divided into: at the scales of 1 minute, 1 hour, and 1 day, a sliding window (length 5, step size 1) is used to extract the mean μ, variance σ2, and range R, respectively, to construct a 128-dimensional time series feature; In the outlier marking, if the fluctuation range of a feature point exceeds ±2σ (σ is the standard deviation of historical data), it is marked as an outlier, and its location and impact radius (such as the data 10 minutes before and after the impact) are recorded; In the anomaly correction of feature reconstruction and fusion, linear interpolation or weighted averaging of neighboring points (weight is inversely proportional to distance) is used to reconstruct data for the anomaly points. The formula is: , Among them, N(i) is the set of neighboring points, d(i,j) is the Euclidean distance; Character Description: : The corrected outlier value, N(i): The set of neighboring points of outlier i (e.g., data points 10 minutes before and after), d(i,j): The Euclidean distance between point i and neighboring point j, w j : The weight of neighboring point j, inversely proportional to the distance; In modal fusion, the corrected time series features are spliced ​​with the temperature and vibration signals, fused through a fully connected layer (dimension 64, activation function ReLU), and output as 32-dimensional comprehensive features.

[0026] like Figure 1-Figure 2As shown, in step S104, for the time series feature sequence, a statistical analysis method is used to extract the difference characteristics of the mutation point and the stable segment, the mutation point is defined as the point where the value change rate exceeds the preset threshold, and the stable segment is defined as the interval where the change rate is stable, and a multi-time scale change pattern data set of the electrical variable is generated for subsequent abnormal signal screening.

[0027] Furthermore, in step S104, by processing the time series feature sequence, a statistical analysis method is used to preliminarily decompose the data, obtain the distribution of the value change, and obtain a preliminary calculation result of the change rate value; Based on the calculation results of the change rate value, for the identification of the mutation point, if the change rate value exceeds the preset threshold, it is determined to be a mutation point, the specific location information of the mutation point is obtained, and its distribution characteristics in the time series feature are determined; Based on the distribution characteristics of the mutation points and combined with the identification of the stable interval, if the change rate value remains stable in the continuous interval, it is determined to be a stable interval, and the start and end point information of the stable interval is obtained; According to the start and end point information of the stable interval, combined with the extraction of difference features, the sliding window method is used to segment the time series feature sequence, obtain the change range between the mutation point and the stable interval, and determine the specific quantitative indicators of the difference features; By using the quantitative indicators of difference characteristics, a multi-time scale analysis framework is constructed, and the time window segmentation technology is used to perform multi-level decomposition of the temporal feature sequence to obtain the manifestation of the change pattern at different time scales; According to the manifestation of the change pattern at different time scales, the data set is integrated and the change pattern is grouped using cluster analysis method to obtain the potential distribution characteristics of abnormal signals and determine the preliminary screening range of abnormal signals; Through the preliminary screening range of abnormal signals, combined with the comprehensive analysis of the change rate value and difference characteristics, the logistic regression model is used to conduct a secondary verification of the abnormal signals to obtain the final abnormal signal classification results and determine the specific category of the abnormal signals.

[0028] Specifically, in step S104, a statistical analysis method is used to process the time series feature sequence to extract the difference characteristics between the mutation point and the stable segment. For example, by calculating the rate of change of the values ​​at adjacent time points, the specific distribution data of the rate of change is obtained. According to the distribution data of the value change rate, the preset threshold range is set to ±10%, and the points where the value change rate exceeds ±10% are marked as mutation points. For example, if the change rate at a certain time point is 15%, it is marked as a mutation point, and the location set of the mutation points is determined; Analyze the stability of the rate of change by using interval data outside the set of mutation point locations. For example, calculate the variance of the rate of change within the interval. If the variance is less than 0.5, mark the interval where the rate of change fluctuates within the stable range as a stable segment, and obtain the interval distribution of the stable segment. Using the mutation point location set and the stationary interval distribution, a multi-time-scale variation pattern data set of electrical variables is constructed. For example, 1-minute, 1-hour, and 1-day time-scale data are integrated to generate structured data for subsequent processing. Extract the change characteristics at different time scales based on the multi-time scale change pattern data set. For example, calculate the average change rate at a 1-hour time scale through a sliding window and combine it with the time context information to obtain the fluctuation pattern data of the time series characteristics. By analyzing the fluctuation pattern data of time series characteristics, the performance of mutation points at different time scales is analyzed. If the change rate of a mutation point continuously exceeds the threshold range of ±10% at both the 1-minute and 1-hour time scales, it is marked as a potential anomaly point, and a potential anomaly point set is obtained. For the set of potential outliers, combined with the interval distribution data of the stable segment, a comparative analysis method is used. For example, the difference in the average rate of change of the stable segment before and after the outlier is calculated. If the difference is greater than 5%, it is determined that the outlier affects the stability of the surrounding stable segments, and the preliminary screening results of the abnormal signal are determined; Based on the preliminary screening results of abnormal signals, the specific location and impact range of the abnormal point are extracted, such as the data 10 minutes before and after the abnormal point, to generate the feature vector group of the abnormal signal and obtain the input data for subsequent classification; The feature vector group of the abnormal signal is used in combination with the time context information to classify the abnormal signal through a classification model. For example, a support vector machine model is used. If the feature vector group shows short-term changes and the time context has no persistence, it is judged as a transient disturbance, and the classification label data of the abnormal signal is obtained.

[0029] In one embodiment, the isolation forest anomaly score formula in step S104 is: , Character description: h(x): path length of sample x in the isolation tree (number of edges from the root node to the leaf node), c(n): normalization factor, where H(n) is the harmonic number, s(x): anomaly score, the closer the value is to 1, the more likely it is an anomaly point.

[0030] like Figure 1-Figure 2 As shown, in step S105, a preliminary screening method based on a preset threshold is applied to the multi-time scale change pattern data set to identify potential abnormal points, analyze their initial trigger conditions and fluctuation characteristics, and generate a dynamic change data set containing potential abnormal signals for subsequent deep anomaly detection.

[0031] Furthermore, in step S105, for the data set at multiple time scales, a hierarchical extraction method is adopted to obtain a set of change patterns from different time windows, construct an initial multi-dimensional feature matrix, and obtain a preliminarily sorted feature data set; Based on the preliminarily sorted feature data set, a screening method with a preset threshold is applied to preliminarily identify potential abnormal points. If the feature value exceeds the preset threshold range, it is marked as a suspected abnormal point, and a set of suspected abnormal points is determined; For the set of suspected outliers, we analyze their triggering conditions and fluctuation characteristics. By detecting the continuity of the time series, we obtain the contextual fluctuation amplitude of each suspected outlier and obtain a set of fluctuation feature descriptions. Based on the fluctuation feature description set, the support vector machine algorithm is used to classify the fluctuation characteristics of potential anomalies, determine whether they conform to the pattern of abnormal signals, and obtain the classified abnormal signal candidate set; For the classified abnormal signal candidate set, through dynamic trend analysis, the evolution trajectory of the abnormal signal at multiple time scales is obtained to determine the final abnormal signal data set; Based on the final abnormal signal data set, a dynamic feature mapping relationship is constructed, and the correlation between the abnormal signal and the trigger condition is quantified to obtain the dynamic feature matrix of the abnormal signal; Aiming at the dynamic change feature matrix of abnormal signals, a cluster analysis method is used to group abnormal signals with similar fluctuation characteristics, determine their priority in deep detection, and obtain a priority list of grouped abnormal signals.

[0032] Specifically, in step S105, for the multi-time scale variation pattern data set, a preliminary screening method based on a preset threshold is adopted, for example, the characteristic value threshold is set to ±2σ, and points outside the range are identified to obtain a set of suspected abnormal points; Based on the set of suspected outliers, we analyze the fluctuation amplitude of each point within the previous and next 10 time windows through the continuity test of the time series to determine the fluctuation feature description set; For the fluctuation feature description set, the support vector machine algorithm is used, and the kernel function is set as the radial basis function. The fluctuation characteristics of potential anomalies are classified and processed to determine whether they meet the abnormal signal pattern, and the classified abnormal signal candidate set is obtained; Based on the classified abnormal signal candidate set, dynamic trend analysis is performed, such as calculating the rate of change of each signal on three time scales: hour, day, and week. This allows the evolution trajectory of the abnormal signals to be obtained and the final abnormal signal data set to be determined. For the final abnormal signal data set, a dynamically changing feature mapping relationship is constructed. For example, principal component analysis (PCA) is used for dimensionality reduction to analyze the correlation between abnormal signals and trigger conditions to obtain a dynamically changing feature matrix. Based on the dynamic change feature matrix, a cluster analysis method, such as K-means clustering, is used. The number of clusters is set to 5. Abnormal signals with similar fluctuation characteristics are grouped and processed. Their priority in deep detection is determined to obtain a priority list of grouped abnormal signals. Based on the grouped abnormal signal priority list, combined with business-related operating parameters such as temperature, pressure, and flow, the multi-dimensional features of high-priority abnormal signals are extracted to determine the high-priority abnormal signal feature set; If the feature distribution in the high-priority abnormal signal feature set shows a significant clustering trend, for example, if a densely populated area is detected by the DBSCAN algorithm, the feature classification is further refined through spatial distribution analysis to obtain a refined abnormal signal category label; Based on the refined abnormal signal category labels, the type classification rules in the business logic are associated, such as associating temperature anomalies with equipment overheating, generating a structured abnormality classification sequence, and completing the type classification of abnormal points.

[0033] like Figure 1-Figure 2 As shown, in step S106, the isolation forest algorithm is used to perform anomaly detection on the dynamically changing data set, and based on the data distribution in the multidimensional feature space, anomaly points that deviate from the normal operation trajectory are identified, and an anomaly label sequence is generated for subsequent anomaly type classification.

[0034] Furthermore, in step S106, the dynamically changing data set is processed, an isolation forest algorithm is used to identify abnormal points, a tag sequence is generated, and subsequent type classification is completed by combining multidimensional features and spatial distribution; Obtain dynamically changing data sets, extract multi-dimensional features from them, and form an initial feature matrix for subsequent anomaly detection processing; For the initial feature matrix, the isolation forest algorithm is used to analyze the spatial distribution, identify abnormal points that deviate from the normal trajectory, and obtain an abnormal marker set; Based on the abnormal mark set and combined with the characteristic analysis of operation deviation, the distribution pattern of abnormal points is extracted and the abnormal distribution characteristics are determined; Based on the abnormal distribution characteristics, the abnormal points are preliminarily clustered, and a simple distance calculation method is used to determine the aggregation trend of the abnormal points and obtain the clustering grouping results; If there is a significant clustering trend in the clustering grouping results, the grouping is further refined by combining multidimensional features to determine the potential category labels of the abnormal points; Based on the potential category labels and the business logic of the associated type classification, the final anomaly classification sequence is generated to complete the classification of the anomaly points. Through the final exception classification sequence, structured data of the exception marking sequence is generated for subsequent business process calls.

[0035] Specifically, in step S106, a dynamically changing data set is obtained, and an initial feature matrix is ​​constructed by extracting multidimensional features. For example, 10 features such as mean, variance, and peak are extracted from the time series data to form an initial feature matrix of 1000×10, thereby obtaining structured data for subsequent anomaly detection. For the initial feature matrix, the isolation forest algorithm is used to analyze the data distribution in the multidimensional feature space. The number of trees is set to 100 and the maximum number of samples is set to 256. The abnormal points that deviate from the normal operation trajectory are identified, and an abnormal label set containing 200 abnormal points is obtained. Based on the anomaly marker set and combined with feature analysis of operational deviations, the distribution pattern of anomaly points is extracted. For example, the spatial density of anomaly points is calculated through kernel density estimation to determine the anomaly distribution characteristics. Based on the abnormal distribution characteristics, the abnormal points are preliminarily clustered, and the clustering trend is analyzed using the Euclidean distance calculation method. The number of clusters is set to 3, and the clustering grouping results are obtained; If there is a significant clustering trend in the clustering results, for example, if a certain category accounts for more than 60%, the grouping is further refined by combining multidimensional features. The DBSCAN algorithm is used with a neighborhood radius of 0.5 and a minimum sample size of 5 to determine the potential category labels of the outliers. Based on the potential category labels, the business logic rules for the associated type classification are calculated. For example, high-density clustered abnormal points are labeled as "equipment failures." This generates the final abnormal classification sequence and completes the classification of abnormal points. The final anomaly classification sequence is used to construct structured data for the anomaly tag sequence. For example, JSON format data containing timestamp, anomaly type, and deviation is generated to obtain a data format that can be called by subsequent processes. For structured data in anomaly marker sequences, if the deviation of an anomaly point exceeds a preset threshold, for example, if the deviation is greater than 2.5, a multi-layer perceptron model is used for classification. The model sets the hidden layer to two layers, with 64 neurons per layer. The model combines the fluctuation characteristics of the mutation point with the temporal context information to determine the anomaly type. Based on the classification results of the multi-layer perceptron model, anomaly type labels are generated. For example, anomalies with significant fluctuation characteristics and short duration are marked as "transient disturbances", and anomalies with smooth fluctuation characteristics and long duration are marked as "persistent deviations" for subsequent processing.

[0036] like Figure 1-Figure 2As shown, in step S107, for the abnormal mark sequence, if the deviation of the abnormal point exceeds the preset threshold, it is classified through the multi-layer perceptron model, and based on the fluctuation characteristics of the mutation point and the time context information, it is judged whether the abnormality belongs to a transient disturbance or a continuous deviation, and an abnormal type label is generated.

[0037] Furthermore, in step S107, for the abnormal mark sequence, the deviation degree data therein is obtained, and by comparing it with the preset threshold, it is determined whether there are abnormal points exceeding the threshold, and a preliminary screened abnormal point set is obtained; From the initially screened set of abnormal points, extract the mutation fluctuation characteristics of each point, combine them with the corresponding time context information, form a feature vector group, and determine the input data for subsequent classification; A multi-layer perceptron model is used to classify the feature vector group. If the sudden fluctuation characteristics show short-term changes and the time context has no persistence, it is determined to be a transient disturbance. If the mutation fluctuation characteristics show long-term deviation and the temporal context is continuous, it is determined to be a persistent deviation and the classification result of the abnormal type is obtained; Based on the classification results, corresponding type labels are generated, and transient disturbances and persistent deviations are marked separately to form a structured anomaly type label dataset; Through the anomaly type label dataset, we extract point data related to persistent deviations and combine it with time context information to construct an input feature matrix for trend prediction. The input feature matrix of trend prediction is processed using a pre-established regression model to analyze the changing pattern of persistent deviations over time and determine future evolution trends. Based on the results of the evolution trend and combined with the anomaly type label, structured prediction data output is generated for subsequent system analysis.

[0038] Specifically, in step S107, for the abnormal mark sequence, the deviation degree data is extracted. For example, if the current value collected by a certain sensor deviates from the normal range by ±10%, the abnormal point set with a deviation of 15% is screened out. Extract mutation fluctuation features from this set. For example, calculate the mutation amplitude when the absolute value of the difference between adjacent time points exceeds 5%, and combine it with the mean of the time window before and after 30 minutes to form a feature vector group that includes the fluctuation intensity and time correlation. A multi-layer perceptron model (3 hidden layers, ReLU activation function) is used to classify feature vectors. If the fluctuation amplitude of a certain point falls back to the threshold within 3 time units and there is no continuous abnormality, it is marked as a transient disturbance. If the fluctuation lasts for more than 10 time units and is accompanied by cumulative deviation, it is marked as a persistent deviation; Generate a labeled dataset based on the classification results, such as "transient disturbance - voltage sag" or "sustained deviation - temperature drift"; Extract continuous deviation data, construct a feature matrix with timestamp, deviation, and moving average as columns, and input it into the LSTM regression model (with a time step of 5) to predict the trend slope for the next three periods; The prediction results are combined with spatial distribution characteristics, and the isolation forest algorithm (100 trees, contamination = 0.1) is used to detect the spatial clustering of abnormal points. If more than 60% of the points in a certain area show persistent deviation and the Euclidean distance is less than 0.5, the "regional persistent anomaly" label is output, and finally a structured dataset containing classification, trend and spatial distribution is generated.

[0039] In one embodiment, in the multi-layer perceptron classification rule of step S107, the input features include 128-dimensional features such as fluctuation intensity (difference absolute value) and time window mean, and the output labels are: the fluctuation amplitude falls back to the threshold within 3 time units and has no continuity, and the fluctuation lasts for more than 10 time units and is accompanied by cumulative offset; In the LSTM regression model, the input features are: timestamp, deviation (current deviates from the normal value by ±10%), and moving average, and the output is: the trend slope of the next three cycles (unit: deviation / time unit).

[0040] Those skilled in the art can make various other corresponding changes and deformations based on the technical solutions and concepts described above, and all of these changes and deformations should fall within the scope of protection of the claims of this application.

Claims

1. A method for diagnosing deviation of electric power construction based on multimodal time series data fusion, characterized in that: include: Through multi-channel synchronous sampling technology, voltage, current and frequency data streams are synchronously collected, and the initial multimodal data set is constructed based on timestamp alignment. The initial data matrix is ​​generated through interpolation to fill missing data and standardization processing; Performing hierarchical feature extraction on the initial data matrix, mapping the frequency domain features of voltage, the time domain statistical features of current, and frequency data into a unified feature space, and generating a multimodal time series feature vector set in combination with time context information; Modeling the multimodal time series characteristics, analyzing the dynamic change trends of electrical variables at different time scales, and generating a time series feature sequence including time dependence; Extracting the difference features between the mutation points and the stable segments in the time series feature sequence, constructing a change pattern data set through multi-time scale analysis, and screening potential abnormal signals; Identify abnormal points that deviate from the normal operating trajectory based on the isolation forest algorithm and generate an abnormal marker sequence; The multi-layer perceptron model combines the fluctuation characteristics of the mutation point and the time context information to classify the anomaly as a transient disturbance or a persistent deviation, and generates an anomaly type label. Combined with the abnormal type label, the evolution trend of the continuous deviation is predicted, and structured prediction data is output for risk warning.

2. The electric power construction deviation diagnosis method based on multimodal time series data fusion according to claim 1 is characterized in that: The multi-channel synchronous sampling technology uses the IEEE 1588 protocol for timestamp alignment, with a time error of ≤1ms. Missing data is filled using a cubic spline interpolation algorithm, and data normalization uses the Z-score method to ensure that the voltage, current, and frequency data have a mean of 0 and a variance of 1.

3. The electric power construction deviation diagnosis method based on multimodal time series data fusion according to claim 1 is characterized in that: The hierarchical feature extraction includes: extracting frequency domain energy entropy from voltage data using wavelet transform, calculating mean, variance, and peak-to-peak values ​​from current data as time domain features, aligning voltage and current features in a unified space through canonical correlation analysis, and generating a multimodal feature vector with time series labels using sliding window technology.

4. The electric power construction deviation diagnosis method based on multimodal time series data fusion according to claim 1 is characterized in that: The LSTM network is configured with 64 hidden layer units. The key time points are weighted by the attention mechanism to extract dynamic change trends. Abnormal fluctuation points are corrected using the weighted average method of neighboring points, and the weight is inversely proportional to the Euclidean distance.

5. The electric power construction deviation diagnosis method based on multimodal time series data fusion according to claim 1 is characterized in that: The mutation point is defined as the point where the value change rate exceeds ±10%, and the stable segment is a continuous interval with a variance ≤ 0.

5. The difference characteristics are quantified by sliding window segmentation, and the clustering algorithm is used to prioritize the abnormal signals.

6. The electric power construction deviation diagnosis method based on multimodal time series data fusion according to claim 1 is characterized in that: The isolation forest algorithm sets the number of trees to 100, calculates the anomaly score by path length, marks points with a deviation ≥ 2.5 as anomalies, refines the classification of anomaly points through spatial density analysis and DBSCAN clustering, and generates anomaly type labels by associating equipment operating parameters.

7. The electric power construction deviation diagnosis method based on multimodal time series data fusion according to claim 1 is characterized in that: The multi-layer perceptron model is configured with three hidden layers, the activation function is ReLU, and the classification rule is: if the fluctuation amplitude falls back within three time units and has no continuity, it is a transient disturbance; if the fluctuation lasts for more than 10 time units, it is a continuous deviation.

8. The electric power construction deviation diagnosis method based on multimodal time series data fusion according to claim 1 is characterized in that: The trend prediction adopts the LSTM regression model, with input features including timestamp, deviation and moving average, and outputs the trend slope of the next three periods. It combines the spatial distribution characteristics to generate regional anomaly warning labels.

Citation Information

Patent Citations

  • Real-time fault monitoring Internet of Things system for chemical production equipment cluster

    CN119232773A

  • Method and system for predicting health of battery pack

    CN119959781A

  • Photovoltaic module fault monitoring system and method

    CN120263107A

  • Power grid abnormal flow detection method based on multi-modal data fusion

    CN120354230A

Cited By

  • Intelligent pathology review method and system based on case history comparison

    CN120913888A

  • Safety monitoring system for electric vehicle charging

    CN121278607A

  • Method and equipment for monitoring liquid sulfur blockage of sulfur recovery device

    CN121300316A

  • Electric meter box fault detection method and system and electronic equipment

    CN121679256A