Power data anomaly detection method and system, computer equipment and storage medium

By constructing a multi-source data association matrix and a sliding window algorithm, combined with transfer learning and federated learning, and dynamically adjusting the baseline threshold, the problems of threshold interference, adaptability, and privacy security in power data anomaly detection are solved, achieving efficient and accurate anomaly detection and fault diagnosis.

CN122020582APending Publication Date: 2026-05-12HUIZHOU HUANENG ELECTRIC INSTALLATION CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
HUIZHOU HUANENG ELECTRIC INSTALLATION CO LTD
Filing Date
2025-12-18
Publication Date
2026-05-12

AI Technical Summary

Technical Problem

Existing methods for detecting anomalies in power data are susceptible to interference from sudden abnormal data during the threshold adjustment process. The window size cannot adapt to multi-scale fluctuations, ignores temporal correlations and the impact of non-power scenarios, and has insufficient cross-domain adaptability. Furthermore, they suffer from privacy and security issues as well as low detection accuracy.

Method used

By constructing a multi-source data association matrix, combining the sliding window algorithm and transfer learning, dynamically adjusting the baseline threshold, introducing domain-adaptive fusion coefficients and federated learning framework, cross-regional detection is achieved. Anomalies are screened using a combination of coarse screening and fine verification, and the source of anomalies is traced.

Benefits of technology

It improves the accuracy and cross-domain adaptability of power data anomaly detection, ensures data privacy and security, reduces misjudgments and network latency, and provides direction for fault diagnosis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122020582A_ABST
    Figure CN122020582A_ABST
Patent Text Reader

Abstract

The invention provides a power data anomaly detection method and system, computer equipment and a storage medium. The method comprises the following steps: dynamically associating power data with non-power scene data to construct a multi-source data association matrix; according to the time sequence evolution features of the power data and various scene influence variables, feature information corresponding to a specific area or equipment which is subjected to exception labeling is migrated to similar scenes, cooperative training of a cross-regional exception detection model is realized on the premise that original power data of each area is not shared, and a cross-regional exception detection model is obtained by adopting a coarse screening and fine testing mode. In combination with a topological connection structure of a power system, a propagation path of abnormal data is tracked and traced, and an initial node position of abnormal triggering is determined through reverse derivation. According to the method, real anomalies can be accurately distinguished, a direction is provided for troubleshooting, the detection adaptation capability of unlabeled data is improved, the cloud transmission quantity and network delay are reduced, frequent manual intervention is not needed, and the labor intensity is enhanced while the efficiency is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of power data anomaly detection technology, and in particular to a power data anomaly detection method, system, computer equipment, and storage medium. Background Technology

[0002] During the operation of the power system, accurate monitoring and anomaly identification of power data are crucial to ensuring the stable and efficient operation of the system. Power data includes core indicators such as load change curves, equipment operating status parameters, and voltage and current time-series monitoring data. Its data quality directly affects key tasks such as power dispatching, fault diagnosis, and equipment maintenance.

[0003] The sliding window algorithm used in existing technologies has several drawbacks in the threshold adjustment stage of power data anomaly detection. Firstly, if sudden abnormal data is mixed into the window, it directly interferes with the threshold calculation results, leading to baseline shift and misjudging normal data. Secondly, the fixed window size cannot adapt to the multi-scale fluctuation characteristics of power data, and it is poorly suited to complex scenarios with short-term sudden changes or long-term gradual changes. Furthermore, traditional methods rely only on the local features of the data within the window, ignoring the inherent temporal correlation of power data and the influence of non-power scenario data such as meteorological environment and traffic flow, resulting in a lag in threshold adjustment and an inability to accurately capture abnormal data features. Additionally, the distribution of power data varies across different regions and devices, and traditional detection methods lack effective cross-domain adaptation mechanisms, leading to low accuracy in screening unlabeled data. Moreover, sharing original data during cross-regional detection can easily raise privacy and security issues, and failures at some nodes can affect the training effect of the overall detection model. Finally, traditional detection methods do not perform sufficient verification of the authenticity of abnormal data, easily misjudging redundant data and false anomalies caused by transmission delays as real anomalies, and making it difficult to accurately trace the initial node of the anomaly, thus complicating troubleshooting.

[0004] Therefore, those skilled in the art are dedicated to providing a method, system, computer equipment, and storage medium for detecting power data anomalies that can effectively solve the above-mentioned technical problems. Summary of the Invention

[0005] In view of the above-mentioned deficiencies of the prior art, the technical problem to be solved by the present invention is to provide a method, system, computer equipment and storage medium for detecting power data anomalies, so as to solve the problems existing in the prior art.

[0006] To achieve the above objectives, the present invention provides a method for detecting power data anomalies, the method comprising:

[0007] S1: A multi-source data association matrix is ​​constructed by dynamically associating power data with non-power scenario data; wherein, the power data includes load change curves, equipment operating status parameters, and voltage and current time-series monitoring data, and the non-power scenario data includes meteorological environmental data and traffic flow statistics; the dynamic association process includes unified alignment processing of time granularity and dynamic adaptation and allocation of scenario weights.

[0008] S2: Based on the time-series evolution characteristics of power data and various scenario-related variables, the baseline threshold range is dynamically adjusted in real time using a sliding window algorithm. The baseline threshold range can exhibit non-linear scaling and adaptation characteristics as scenario variables fluctuate.

[0009] S3: Through transfer learning algorithms, the feature information corresponding to specific regions or devices that have completed anomaly labeling is transferred to similar scenarios. Combined with a domain adaptive model, the distribution differences between different data sources are eliminated. A domain adaptive fusion coefficient is introduced to dynamically balance the adaptability of data features between the source and target domains, achieving preliminary screening of anomalies in unlabeled data. The domain adaptive fusion coefficient introduced in S3 is... The expression is:

[0010] in For domain-adaptive fusion coefficients;

[0011] These are the feature values ​​of the source domain data;

[0012] For the feature values ​​of the target domain data,

[0013] The number of feature dimensions;

[0014] is the Pearson correlation coefficient between the source domain and the target domain;

[0015] Set migration trust factor The value range is [0.6, 0.9]), dynamically adjusting the weights for migrating features from the source domain to the target domain. The value is determined by the scene similarity and data distribution consistency between the source and target domains, and is expressed as:

[0016]

[0017] in For scene similarity;

[0018] To ensure data distribution consistency;

[0019] Scene similarity The expression is obtained by calculating the weighted sum of the meteorological type matching degree, traffic flow level matching degree, and topological structure similarity between the source and target domains:

[0020]

[0021] in For meteorological type matching degree,

[0022] Traffic flow level matching degree (levels are divided according to flow range, with 1 for consistent levels, 0.6 for adjacent levels, 0.3 for interval levels, and 0 for the rest).

[0023] The topological similarity is calculated based on node connection methods and distances, with a value range of [0,1].

[0024] Data distribution consistency pass The divergence is calculated by comparing the data distributions in the source and target domains, and then normalized to the [0,1] interval.

[0025]

[0026] in For source domain With the target domain of divergence,

[0027] The largest in history Divergence value;

[0028] Through the federated learning framework, collaborative training of cross-regional anomaly detection models can be achieved without sharing the original power data of each region. Multi-regional data features are aggregated through encrypted gradient transmission. The federated learning framework adopts an encrypted aggregation algorithm. Specifically, it ensures data privacy and security during gradient transmission through a secret shared gradient aggregation mechanism. At the same time, it sets up a fault tolerance mechanism for aggregation nodes. When some regional nodes fail, the model training is completed through the gradient data of the remaining nodes.

[0029] S4: A combination of coarse screening and fine verification is adopted. First, a coarse-grained algorithm is used to define the candidate set of anomalies; then, a fine-grained algorithm is used to verify the authenticity of the anomalies in the candidate set.

[0030] S5: Combining the topology of the power system, the initial node location of the anomaly is determined by tracing the propagation path of the abnormal data and deducing in reverse.

[0031] Furthermore, the dynamic correlation between power data and non-power scenario data in S1, and the subsequent construction of a multi-source data correlation matrix, specifically includes:

[0032] S11: Perform outlier removal and normalization on the power data, mapping it uniformly to the [0,1] interval; the expression is:

[0033] in Represents the raw values ​​of power data;

[0034] This represents the reasonable minimum value for power data;

[0035] This indicates the reasonable maximum value of the power data;

[0036] This represents the result after normalization;

[0037] Indicates the scene influence coefficient;

[0038] Indicates the scene adaptation factor;

[0039] Scene Influence System The value range is [0.1, 0.8] for meteorological data and [0.05, 0.5] for traffic flow data; scene adaptation factor. The value is dynamically adjusted by the system based on historical data and the associated accuracy, and the range is [0.8, 1.2].

[0040] The scene adaptation factor The dynamic adjustment mechanism is as follows: the system calculates the historical data correlation accuracy based on a preset time; if the accuracy is higher than the preset value, then... Increase the value based on the current value; if the accuracy is lower than the preset value, then... Reduce from the current value; if the accuracy is within the preset range, then Remain unchanged;

[0041] The data from non-electrical scenarios are standardized in format, meteorological data is converted into impact factor indices, and traffic flow data is converted into regional activity intensity values. The time granularity of the two types of data is uniformly adjusted through a timestamp alignment algorithm.

[0042] S12: By using a sliding time window, calculate the Pearson correlation coefficient between power data and non-power data within the window range, capture the correlation changes between the two in the time series dimension, construct a scenario impact factor library, configure initial correlation weights for different weather types and traffic conditions, and then dynamically adjust the weights based on real-time data feedback information.

[0043] By increasing the weight of abrupt change nodes in the time series of power data through a temporal attention mechanism, and constructing a multi-source data association matrix, the data association strength during abnormally sensitive periods is enhanced through an attention coefficient. The expression for the attention coefficient is:

[0044]

[0045] in for Attention coefficient at any moment;

[0046] It is a sensitive regulatory factor;

[0047] for Real-time power data;

[0048] The length of the time series;

[0049] S13: Using power data types as the row dimension of the matrix and non-power data types as the column dimension, the initial values ​​of the matrix elements are set as the real-time correlation coefficients of the corresponding dimensions. A scene weight layer is set on top of this base matrix. The scene adaptation weights are multiplied by the correlation coefficients to obtain the weighted correlation value. A power equipment topology correlation column is set and filled with the corresponding coefficients of the influence of non-power data on the power data of different topology nodes, forming a three-dimensional correlation matrix including data type, scene type and topology node.

[0050] S14: Trigger matrix update operation according to preset time intervals, recalculate correlation coefficients and scene weights using newly collected data, replace outdated elements in the matrix, and start the update process when a sudden change in data distribution is detected to complete dynamic matrix adjustment and establish an association effectiveness evaluation index system. If the association accuracy of the matrix is ​​lower than a preset threshold, the system will adjust the correlation coefficient calculation model and weight allocation rules. The association effectiveness evaluation index system includes association accuracy, time synchronization error rate, and weight fit. The association accuracy is the proportion of effective associated data in the matrix, the time synchronization error rate is the statistical value of data timestamp alignment deviation, and the weight fit is the degree of matching between scene weights and the actual data association strength.

[0051] Furthermore, in step S2, the real-time dynamic adjustment of the baseline threshold interval using the sliding window algorithm specifically includes:

[0052] S21: Set a first initial window, a second initial window, and a third initial window based on the power scenario. The first initial window has 15-20 data points, the second initial window has 30-50 data points, and the third initial window has 60-90 data points; standardize the format of data from non-power scenarios.

[0053] S22: Calculate the variance and trend slope of the data within the first initial window, the second initial window, and the third initial window;

[0054]

[0055] in Indicates the dynamic fusion weight of the window;

[0056] , ,and These represent the standard deviations of the data in the first initial window, the second initial window, and the third initial window, respectively.

[0057] , and These represent the window's basic weights;

[0058] A window adaptation evaluation model is established, and dynamic weights are assigned to the first, second, and third initial windows based on data fluctuations and changes. If the data fluctuates sharply in the short term, the weight of the first window is increased; if the data shows stable changes in the medium term, the weight of the second window is increased; if the data shows a stable long-term trend, the weight of the third window is increased. The data features of the first, second, and third initial windows are fused based on the weights to generate the calculation base for the initial baseline threshold interval.

[0059] S23: Real-time acquisition of physical constraint parameters of power equipment operation, including the rated voltage range, rated current range and maximum load threshold of the equipment, and conversion of them into corresponding threshold boundary values; fusion of the threshold boundary values ​​corresponding to the physical constraint parameters with the dynamically adjusted baseline threshold range, and taking the intersection of the two as the final baseline threshold range. If the two have no intersection, the threshold of the physical constraint parameters shall prevail, and the system shall trigger an alarm to indicate threshold conflict.

[0060] Furthermore, after S23, the method also includes introducing a physical constraint threshold for power equipment, trimming and calibrating the calculated threshold range to ensure that the threshold does not exceed the safe operation boundary of the equipment, monitoring the distribution density of data within the threshold range in real time, and triggering iterative adjustment of window parameters if the proportion of normal data distribution is low. The method calculates the difference between the time of occurrence of abnormal data and the response time of threshold adjustment. If the difference exceeds the preset threshold, the window weight allocation rule is optimized.

[0061] By deploying a lightweight anomaly screening algorithm on the power terminal equipment side through an edge computing node deployment mechanism, abnormal data can be identified and initially processed locally. This reduces the amount of data transmitted to the cloud, lowers network latency, and only uploads suspected complex anomaly data to the cloud for detailed verification.

[0062] Furthermore, the iterative adjustment of the window parameters is as follows: First, based on the scene weights and topological correlation coefficients in the three-dimensional correlation matrix, a dynamic proportion threshold model is constructed. This model is used to calculate the normal data distribution proportion threshold in real time. The model takes the current scene influence coefficient and the topological node correlation strength as inputs, and combines the benchmark value of the normal distribution proportion under the same scene in historical data. The dynamic threshold is output through an adaptive weighting algorithm. If the actual normal data proportion is lower than the dynamic threshold, a hierarchical adjustment mechanism is initiated: The number of data points in the first initial window is adaptively reduced by 2-4 according to the scene fluctuation intensity. Specifically, the upper limit of reduction is taken when the short-term fluctuation is severe, and the lower limit of reduction is taken when the fluctuation is gentle. The number of data points in the second initial window is dynamically increased or decreased by 1-3 according to the slope of the time-series mutation of the power data. Specifically, the number of data points is increased when the mutation slope exceeds the threshold and decreased when it is lower than the threshold. The number of data points in the third initial window is increased by 2-5 in combination with the long-term trend stability. After adjustment, a time-series attention coefficient is introduced to weight and correct the window data, strengthen the data weight of abnormally sensitive periods, and then the baseline threshold range is recalculated. At the same time, the threshold deviation value and scene adaptability change before and after adjustment are recorded.

[0063] Furthermore, the method of rapidly defining the anomaly candidate set using a coarse-grained algorithm specifically includes: setting a dynamic clustering radius benchmark value based on the spatiotemporal correlation characteristics of power data; the method for setting the dynamic clustering radius benchmark value is as follows: setting an initial benchmark value based on the average radius of normal data clusters in historical data, combined with the power data type, and then dynamically fine-tuning it according to the real-time scenario weights; using an improved density clustering algorithm, with the time-series nodes of power data in the three-dimensional correlation matrix as the core, dividing data that are adjacent in the time dimension and belong to the same topological correlation region in the spatial dimension into basic data clusters; calculating the core feature statistics of each data cluster, including the data mean, variance of fluctuation, and deviation from the baseline threshold interval, and setting the deviation threshold as... The The expression is

[0064]

[0065] in Indicates the overall deviation of the data;

[0066] The mean of the baseline threshold interval;

[0067] The slope of change in the time series segment containing the abnormal data;

[0068] This represents the maximum permissible slope for normal equipment operation.

[0069] The slope influence factor;

[0070] Data clusters with deviations exceeding a set threshold are selected, and all data within each cluster are marked as abnormal candidate data. These are then integrated to form an abnormal candidate set. Simultaneously, the spatiotemporal coordinates and associated scene information of each candidate data are recorded. The candidate set is then subjected to preliminary deduplication to remove redundant abnormal records caused by repeated data collection or transmission delays, while retaining uniquely identified abnormal candidate data entries.

[0071] Furthermore, the fine-grained algorithm verifies the authenticity of anomalies in the candidate set by extracting the physical operation constraint parameters of the power equipment corresponding to each abnormal data in the candidate set and constructing an equipment constraint verification library; introducing a time-series consistency verification model, comparing the change trends of adjacent time-series data before and after the abnormal data, calculating the abrupt change slope between the abnormal data and the preceding normal data, and if the slope exceeds the maximum allowable change rate under normal equipment operation, proceeding to the next verification step; otherwise, it is marked as a false anomaly.

[0072] By combining the scene weights and topological correlation coefficients in the three-dimensional correlation matrix, we analyze whether the abnormal data is related to the current scene conditions. If the abnormal behavior contradicts the scene's influence pattern, it is judged as a suspicious abnormality. We call the equipment operation mechanism model to simulate whether the operating state corresponding to the abnormal data conforms to the physical characteristics of the power equipment. If the deviation exceeds a preset threshold, it is judged as a real abnormality. We comprehensively score the candidate data after the above steps. If the score meets the standard, it is confirmed as a real abnormality. If it does not meet the standard, it is marked as pending verification or a false abnormality, thus completing the authenticity verification.

[0073] A power data anomaly detection system, the system comprising:

[0074] S100: Multi-source data association module, used to remove outliers, normalize and standardize the format of power data and non-power scenario data. It unifies the time granularity through timestamp alignment algorithm, calculates Pearson correlation coefficient and dynamically adjusts weights in combination with time series attention mechanism, constructs a three-dimensional association matrix containing data type, scenario type and topology node, and completes dynamic matrix update and effectiveness evaluation according to preset rules.

[0075] S200: Baseline threshold dynamic adjustment module, used to set the initial window with different numbers of data points, calculate the data variance and trend slope within the window and assign dynamic weights, integrate the physical constraint parameters of power equipment to generate the final baseline threshold range, deploy a lightweight coarse screening algorithm through edge computing nodes, and iteratively adjust the window parameters according to the data distribution.

[0076] S300: Cross-domain collaborative detection module, used to transfer labeled abnormal features to similar scenarios through transfer learning algorithms, combine domain adaptive models to eliminate differences in data source distribution, and realize cross-regional model collaborative training based on the encrypted aggregation mechanism of federated learning framework, ensuring data privacy and security and node fault tolerance;

[0077] S400: Anomaly Precision Screening Module, which uses a coarse-grained screening algorithm of improved density clustering to define anomaly candidate sets and remove duplicates, and then uses fine-grained algorithms of time series consistency verification, scene association analysis and equipment operation mechanism simulation to verify the authenticity of candidate set data and give a comprehensive score.

[0078] S500: Anomaly tracing and location module, used to combine the power system topology connection structure to track the propagation path of abnormal data, reverse deduce and determine the initial node location of the anomaly trigger, and synchronously record the spatiotemporal coordinates and related scene information of the abnormal data.

[0079] A computer device includes a memory, a processor, and a computer program stored in the memory, wherein the processor executes the computer program to implement the steps of the power data anomaly detection method according to any one of claims 1 to 7.

[0080] A computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the power data anomaly detection method according to any one of claims 1 to 7.

[0081] As described above, the present invention provides a method, system, computer equipment, and storage medium for detecting power data anomalies, which has the following beneficial effects:

[0082] 1. This invention integrates power data with non-power data such as meteorological and transportation data, and captures multi-dimensional correlations through a three-dimensional correlation matrix to reduce misjudgments based on a single data dimension. It combines coarse screening with fine verification, first quickly identifying the candidate set of anomalies through improved density clustering, and then accurately distinguishing between real anomalies, those awaiting verification, and false anomalies through fine-grained verification such as time series consistency verification and equipment operation mechanism simulation. At the same time, the anomaly source is clearly located, and the propagation path is traced by combining the power system topology to reverse deduce the initial anomaly node, providing guidance for fault diagnosis.

[0083] 2. Dynamic adjustment of baseline threshold: Dynamic weights are assigned through multi-scale initial windows, and physical constraint parameters of the device are integrated to make the threshold range expand and contract non-linearly with scene fluctuations, adapting to different data characteristics such as short-term sudden changes and long-term gradual changes; Iterative optimization of window parameters: The number of data points in the window is adjusted in real time according to the data distribution density and fluctuation intensity, strengthening the weight of abnormal sensitive periods and avoiding the threshold lag problem caused by fixed windows; Collaboration of domain adaptation and transfer learning: Eliminates the distribution differences of different data sources, transfers labeled abnormal features to similar scenes, and improves the detection adaptation capability of unlabeled data;

[0084] 3. This invention aggregates features from multiple regions through encrypted gradient transmission, without sharing the original power data, effectively improving detection performance and data privacy and security; it has a node fault tolerance mechanism, so that when some regional nodes fail, the model can be trained through the gradient data of the remaining nodes, enabling the system to operate effectively and continuously; in addition, edge computing and cloud collaboration, with lightweight coarse screening algorithms deployed on the terminal side, reduce cloud transmission volume and network latency, while triggering threshold conflict alarms to prevent thresholds from exceeding the device's safety boundary;

[0085] 4. By removing outliers, normalizing, and aligning timestamps, data quality and subsequent calculation efficiency are improved. Weights and correlation coefficients are adjusted in conjunction with real-time data feedback. The accuracy of associations is ensured through an effectiveness evaluation index system. This eliminates the need for frequent manual intervention, improving efficiency while reducing labor intensity. Attached Figure Description

[0086] Figure 1 This is a schematic diagram illustrating an exemplary system architecture for applying the technical solutions in one or more embodiments of the present invention;

[0087] Figure 2 This is a schematic flowchart of the power data anomaly detection method in this invention;

[0088] Figure 3 This is a schematic block diagram of the power data anomaly detection system module architecture in this invention;

[0089] Figure 4 This is a schematic diagram of the hardware structure suitable for implementing the computer device of the present invention. Detailed Implementation

[0090] The present invention will be further described below with reference to the accompanying drawings and embodiments:

[0091] The following specific examples illustrate the implementation of the present invention. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. The present invention can also be implemented or applied through other different specific embodiments, and various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present invention. It is understood that, without conflict, the following embodiments and features in the embodiments can be combined with each other. Furthermore, it is understood that the illustrations provided in the following embodiments are only schematically illustrating the basic concept of the present invention. Therefore, the drawings only show components related to the present invention and are not drawn according to the actual number, shape, and size of components in the actual implementation. In the actual implementation, the type, quantity, and proportion of each component can be arbitrarily changed, and the component layout may also be more complex.

[0092] Figure 1 A schematic diagram of an exemplary system architecture that can apply the technical solutions of one or more embodiments of the present invention is shown. Figure 1 As shown, the system architecture 100 may include terminal device 110, network 120, and server 130. Terminal device 110 may include various electronic devices such as smartphones, tablets, laptops, and desktop computers. Server 130 may be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing cloud computing services. Network 120 may be a communication medium of various connection types capable of providing a communication link between terminal device 110 and server 130, such as a wired communication link or a wireless communication link.

[0093] Depending on the implementation requirements, the system architecture in this embodiment of the invention can have any number of terminal devices, networks, and servers. For example, server 130 can be a server group composed of multiple server devices. Furthermore, the technical solutions provided in this embodiment of the invention can be applied to terminal device 110, or to server 130, or can be implemented jointly by terminal device 110 and server 130; this invention does not impose any special limitations on these applications.

[0094] like Figures 1 to 4As shown, in one embodiment of the present invention, the terminal device 110 or server 130 can construct a multi-source data association matrix by dynamically associating power data and non-power scenario data; wherein, the power data includes load change curves, equipment operating status parameters, and voltage and current time-series monitoring data, and the non-power scenario data includes meteorological environmental data and traffic flow statistics; the dynamic association process includes unified alignment processing of time granularity and dynamic adaptation and allocation of scenario weights; based on the time-series evolution characteristics of the power data itself and various scenario influencing variables, the baseline threshold interval is dynamically adjusted in real time using a sliding window algorithm, and the baseline threshold interval can exhibit non-linear scaling and adaptation characteristics as scenario variables fluctuate; through a transfer learning algorithm, the features that have been anomaly labeled are... Feature information corresponding to a specific region or device is transferred to similar scenarios. A domain-adaptive model is used to eliminate distribution differences between different data sources. A domain-adaptive fusion coefficient is introduced to dynamically balance the adaptability of data features between the source and target domains, enabling preliminary screening of anomalies in unlabeled data. Through a federated learning framework, collaborative training of cross-regional anomaly detection models is achieved without sharing original power data from different regions. Multi-regional data features are aggregated through encrypted gradient transmission. A coarse-grained and fine-grained approach is adopted. First, a coarse-grained algorithm is used to define anomaly candidate sets. Then, a fine-grained algorithm is used to verify the authenticity of anomalies in the candidate sets. Finally, by tracing the propagation path of anomaly data and back-deriving the initial node location that triggered the anomaly, the topological connection structure of the power system is considered.

[0095] The above section introduced an exemplary system architecture for applying the technical solution of this invention. Next, we will continue to introduce the power data anomaly detection method of this invention.

[0096] Figure 2 A flowchart illustrating a method for detecting anomalies in power data is shown. Specifically, in an exemplary embodiment, as follows... Figure 2 As shown, this embodiment provides a method for detecting power data anomalies, including the following steps:

[0097] S1: A multi-source data association matrix is ​​constructed by dynamically associating power data with non-power scenario data; wherein, the power data includes load change curves, equipment operating status parameters, and voltage and current time-series monitoring data, and the non-power scenario data includes meteorological environmental data and traffic flow statistics; the dynamic association process includes unified alignment processing of time granularity and dynamic adaptation and allocation of scenario weights.

[0098] S2: Based on the time-series evolution characteristics of power data and various scenario-related variables, the baseline threshold range is dynamically adjusted in real time using a sliding window algorithm. The baseline threshold range can exhibit non-linear scaling and adaptation characteristics as scenario variables fluctuate.

[0099] S3: Through transfer learning algorithms, feature information corresponding to specific regions or devices that have completed anomaly labeling is transferred to similar scenarios. Combined with a domain adaptive model, distribution differences between different data sources are eliminated. A domain adaptive fusion coefficient is introduced to dynamically balance the adaptability of data features between the source and target domains, achieving preliminary screening of anomalies in unlabeled data. The domain adaptive fusion coefficient introduced in S3... The expression is:

[0100] in For domain-adaptive fusion coefficients;

[0101] These are the feature values ​​of the source domain data;

[0102] For the feature values ​​of the target domain data,

[0103] The number of feature dimensions;

[0104] is the Pearson correlation coefficient between the source domain and the target domain;

[0105] A federated transfer learning fusion mechanism is introduced, which involves the federated transfer of data features across regions and the setting of a transfer trust coefficient. The value range is [0.6, 0.9], and the weights for migrating features from the source domain to the target domain are dynamically adjusted. The value is determined by the scene similarity and data distribution consistency between the source and target domains, and is expressed as:

[0106]

[0107] in For scene similarity;

[0108] To ensure data distribution consistency;

[0109] Scene similarity The expression is obtained by calculating the weighted sum of the meteorological type matching degree, traffic flow level matching degree, and topological structure similarity between the source and target domains:

[0110]

[0111] in For meteorological type matching degree,

[0112] Traffic flow level matching degree (levels are divided according to flow range, with 1 for consistent levels, 0.6 for adjacent levels, 0.3 for interval levels, and 0 for the rest).

[0113] The topological similarity is calculated based on node connection methods and distances, with a value range of [0,1].

[0114] Data distribution consistency pass The divergence is calculated by comparing the data distributions in the source and target domains, and then normalized to the [0,1] interval.

[0115]

[0116] in For source domain With the target domain of divergence,

[0117] The largest in history Divergence value;

[0118] Through the federated learning framework, collaborative training of cross-regional anomaly detection models can be achieved without sharing the original power data of each region. Multi-regional data features are aggregated through encrypted gradient transmission. The federated learning framework adopts an encrypted aggregation algorithm. Specifically, it ensures data privacy and security during gradient transmission through a secret shared gradient aggregation mechanism. At the same time, it sets up a fault tolerance mechanism for aggregation nodes. When some regional nodes fail, the model training is completed through the gradient data of the remaining nodes.

[0119] S4: A combination of coarse screening and fine verification is adopted. First, a coarse-grained algorithm is used to define the candidate set of anomalies; then, a fine-grained algorithm is used to verify the authenticity of the anomalies in the candidate set.

[0120] S5: Combining the topology of the power system, the initial node location of the anomaly is determined by tracing the propagation path of the abnormal data and deducing in reverse.

[0121] The dynamic correlation between power data and non-power scenario data in S1, and the subsequent construction of a multi-source data correlation matrix, specifically includes:

[0122] S11: Perform outlier removal and normalization on the power data, mapping it uniformly to the [0,1] interval; the expression is:

[0123] in Represents the raw values ​​of power data;

[0124] This represents the reasonable minimum value for power data;

[0125] This indicates the reasonable maximum value of the power data;

[0126] This represents the result after normalization;

[0127] Indicates the scene influence coefficient;

[0128] Indicates the scene adaptation factor;

[0129] Scene Influence System The value range is [0.1, 0.8] for meteorological data and [0.05, 0.5] for traffic flow data; scene adaptation factor. The value is dynamically adjusted by the system based on historical data and the associated accuracy, and the range is [0.8, 1.2].

[0130] The scene adaptation factor The dynamic adjustment mechanism is as follows: the system calculates the historical data correlation accuracy based on a preset time; if the accuracy is higher than the preset value, then... Increase the value based on the current value; if the accuracy is lower than the preset value, then... Reduce from the current value; if the accuracy is within the preset range, then Remain unchanged;

[0131] The data from non-electrical scenarios are standardized in format, meteorological data is converted into impact factor indices, and traffic flow data is converted into regional activity intensity values. The time granularity of the two types of data is uniformly adjusted through a timestamp alignment algorithm.

[0132] S12: By using a sliding time window, calculate the Pearson correlation coefficient between power data and non-power data within the window range, capture the correlation changes between the two in the time series dimension, construct a scenario impact factor library, configure initial correlation weights for different weather types and traffic conditions, and then dynamically adjust the weights based on real-time data feedback information.

[0133] By increasing the weight of abrupt change nodes in the time series of power data through a temporal attention mechanism, and constructing a multi-source data association matrix, the data association strength during abnormally sensitive periods is enhanced through an attention coefficient. The expression for the attention coefficient is:

[0134]

[0135] in for Attention coefficient at any moment;

[0136] It is a sensitive regulatory factor;

[0137] for Real-time power data;

[0138] The length of the time series;

[0139] S13: Using power data types as the row dimension of the matrix and non-power data types as the column dimension, the initial values ​​of the matrix elements are set as the real-time correlation coefficients of the corresponding dimensions. A scene weight layer is set on top of this base matrix. The scene adaptation weights are multiplied by the correlation coefficients to obtain the weighted correlation value. A power equipment topology correlation column is set and filled with the corresponding coefficients of the influence of non-power data on the power data of different topology nodes, forming a three-dimensional correlation matrix including data type, scene type and topology node.

[0140] S14: Trigger matrix update operation according to preset time intervals, recalculate correlation coefficients and scene weights using newly collected data, replace outdated elements in the matrix, and start the update process when a sudden change in data distribution is detected to complete dynamic matrix adjustment and establish an association effectiveness evaluation index system. If the association accuracy of the matrix is ​​lower than a preset threshold, the system will adjust the correlation coefficient calculation model and weight allocation rules. The association effectiveness evaluation index system includes association accuracy, time synchronization error rate, and weight fit. The association accuracy is the proportion of effective associated data in the matrix, the time synchronization error rate is the statistical value of data timestamp alignment deviation, and the weight fit is the degree of matching between scene weights and the actual data association strength.

[0141] In step S2, the real-time dynamic adjustment of the baseline threshold interval using the sliding window algorithm specifically includes:

[0142] S21: Set a first initial window, a second initial window, and a third initial window based on the power scenario. The first initial window has 15-20 data points, the second initial window has 30-50 data points, and the third initial window has 60-90 data points; standardize the format of data from non-power scenarios.

[0143] S22: Calculate the variance and trend slope of the data within the first initial window, the second initial window, and the third initial window;

[0144]

[0145] in Indicates the dynamic fusion weight of the window;

[0146] , ,and These represent the standard deviations of the data in the first initial window, the second initial window, and the third initial window, respectively.

[0147] , and These represent the window's basic weights;

[0148] A window adaptation evaluation model is established, and dynamic weights are assigned to the first, second, and third initial windows based on data fluctuations and changes. If the data fluctuates sharply in the short term, the weight of the first window is increased; if the data shows stable changes in the medium term, the weight of the second window is increased; if the data shows a stable long-term trend, the weight of the third window is increased. The data features of the first, second, and third initial windows are fused based on the weights to generate the calculation base for the initial baseline threshold interval.

[0149] S23: Real-time acquisition of physical constraint parameters of power equipment operation, including the rated voltage range, rated current range and maximum load threshold of the equipment, and conversion of them into corresponding threshold boundary values; fusion of the threshold boundary values ​​corresponding to the physical constraint parameters with the dynamically adjusted baseline threshold range, and taking the intersection of the two as the final baseline threshold range. If the two have no intersection, the threshold of the physical constraint parameters shall prevail, and the system shall trigger an alarm to indicate threshold conflict.

[0150] The S23 section further includes introducing a physical constraint threshold for power equipment, trimming and calibrating the calculated threshold range to ensure that the threshold does not exceed the safe operation boundary of the equipment, monitoring the distribution density of data within the threshold range in real time, triggering iterative adjustment of window parameters if the proportion of normal data distribution is low, calculating the difference between the time of abnormal data occurrence and the threshold adjustment response time, and optimizing the window weight allocation rule if the difference exceeds the preset threshold.

[0151] By deploying a lightweight anomaly screening algorithm on the power terminal equipment side through an edge computing node deployment mechanism, abnormal data can be identified and initially processed locally. This reduces the amount of data transmitted to the cloud, lowers network latency, and only uploads suspected complex anomaly data to the cloud for detailed verification.

[0152] The iterative adjustment of the window parameters is as follows: First, based on the scene weights and topological correlation coefficients in the three-dimensional correlation matrix, a dynamic proportion threshold model is constructed. This model calculates the normal data distribution proportion threshold in real time. The model takes the current scene influence coefficient and topological node correlation strength as inputs, and combines the benchmark value of the normal distribution proportion under the same scene in historical data. It outputs the dynamic threshold through an adaptive weighting algorithm. If the actual normal data proportion is lower than the dynamic threshold, a hierarchical adjustment mechanism is initiated: The number of data points in the first initial window is adaptively reduced by 2-4 according to the scene fluctuation intensity. Specifically, the upper limit of reduction is taken when the short-term fluctuation is severe, and the lower limit of reduction is taken when the fluctuation is gentle. The number of data points in the second initial window is dynamically increased or decreased by 1-3 according to the slope of the time-series mutation of the power data. Specifically, the number is increased when the mutation slope exceeds the threshold and decreased when it is lower than the threshold. The number of data points in the third initial window is increased by 2-5 based on the long-term trend stability. After adjustment, a time-series attention coefficient is introduced to weight and correct the window data, strengthen the data weight of abnormally sensitive periods, and then the baseline threshold range is recalculated. At the same time, the threshold deviation value and scene adaptability change before and after adjustment are recorded.

[0153] The method for rapidly defining the anomaly candidate set using a coarse-grained algorithm specifically includes: setting a dynamic clustering radius benchmark value based on the spatiotemporal correlation characteristics of power data; the method for setting the dynamic clustering radius benchmark value is as follows: setting an initial benchmark value based on the average radius of normal data clusters in historical data, combined with the power data type, and then dynamically fine-tuning it according to the real-time scenario weights; using an improved density clustering algorithm, with the time-series nodes of power data in the three-dimensional correlation matrix as the core, dividing data that are adjacent in the time dimension and belong to the same topological correlation region in the spatial dimension into basic data clusters; calculating the core feature statistics of each data cluster, including the data mean, variance of fluctuation, and deviation from the baseline threshold interval, and setting the deviation threshold as... The The expression is

[0154]

[0155] in Indicates the overall deviation of the data;

[0156] The mean of the baseline threshold interval;

[0157] The slope of change in the time series segment containing the abnormal data;

[0158] This represents the maximum permissible slope for normal equipment operation.

[0159] The slope influence factor;

[0160] Data clusters with deviations exceeding a set threshold are selected, and all data within each cluster are marked as abnormal candidate data. These are then integrated to form an abnormal candidate set. Simultaneously, the spatiotemporal coordinates and associated scene information of each candidate data are recorded. The candidate set is then subjected to preliminary deduplication to remove redundant abnormal records caused by repeated data collection or transmission delays, while retaining uniquely identified abnormal candidate data entries.

[0161] The fine-grained algorithm verifies the authenticity of anomalies in the candidate set by extracting the physical operation constraint parameters of the power equipment corresponding to each abnormal data in the candidate set and constructing an equipment constraint verification library; introducing a time-series consistency verification model, comparing the change trends of adjacent time-series data before and after the abnormal data, calculating the abrupt change slope between the abnormal data and the preceding normal data, and if the slope exceeds the maximum allowable change rate under normal equipment operation, proceeding to the next verification step; otherwise, it is marked as a false anomaly.

[0162] By combining the scene weights and topological correlation coefficients in the three-dimensional correlation matrix, we analyze whether the abnormal data is related to the current scene conditions. If the abnormal behavior contradicts the scene's influence pattern, it is judged as a suspicious abnormality. We call the equipment operation mechanism model to simulate whether the operating state corresponding to the abnormal data conforms to the physical characteristics of the power equipment. If the deviation exceeds a preset threshold, it is judged as a real abnormality. We comprehensively score the candidate data after the above steps. If the score meets the standard, it is confirmed as a real abnormality. If it does not meet the standard, it is marked as pending verification or a false abnormality, thus completing the authenticity verification.

[0163] A power data anomaly detection system, the system comprising:

[0164] S100: Multi-source data association module, used to remove outliers, normalize and standardize the format of power data and non-power scenario data. It unifies the time granularity through timestamp alignment algorithm, calculates Pearson correlation coefficient and dynamically adjusts weights in combination with time series attention mechanism, constructs a three-dimensional association matrix containing data type, scenario type and topology node, and completes dynamic matrix update and effectiveness evaluation according to preset rules.

[0165] S200: Baseline threshold dynamic adjustment module, used to set the initial window with different numbers of data points, calculate the data variance and trend slope within the window and assign dynamic weights, integrate the physical constraint parameters of power equipment to generate the final baseline threshold range, deploy a lightweight coarse screening algorithm through edge computing nodes, and iteratively adjust the window parameters according to the data distribution.

[0166] S300: Cross-domain collaborative detection module, used to transfer labeled abnormal features to similar scenarios through transfer learning algorithms, combine domain adaptive models to eliminate differences in data source distribution, and realize cross-regional model collaborative training based on the encrypted aggregation mechanism of federated learning framework, ensuring data privacy and security and node fault tolerance;

[0167] S400: Anomaly Precision Screening Module, which uses a coarse-grained screening algorithm of improved density clustering to define anomaly candidate sets and remove duplicates, and then uses fine-grained algorithms of time series consistency verification, scene association analysis and equipment operation mechanism simulation to verify the authenticity of candidate set data and give a comprehensive score.

[0168] S500: Anomaly tracing and location module, used to combine the power system topology connection structure to track the propagation path of abnormal data, reverse deduce and determine the initial node location of the anomaly trigger, and synchronously record the spatiotemporal coordinates and related scene information of the abnormal data.

[0169] It is understood that the power data anomaly detection system and the power data anomaly detection method provided in the above embodiments belong to the same concept. The specific execution method of the power data anomaly detection method has been described in detail in the above method embodiments and will not be repeated here. In practical applications, the power data anomaly detection system provided in the above embodiments can, as needed, allocate the above functions to different functional modules. That is, the internal structure of the power data anomaly detection system can be divided into different functional modules, and then all or part of the functions of the corresponding functional modules can be implemented through the power data anomaly detection method described in the above embodiments. No specific limitations are imposed here.

[0170] In another exemplary embodiment of the present invention, a computer device is also provided, which may include a memory, a processor, and a computer program stored in the memory, wherein the processor executes the computer program to cause the computer device to perform... Figure 2 The steps of the power data anomaly detection method are described above. Figure 4 A schematic diagram of the structure of a computer device 1000 is shown. (See attached diagram.) Figure 4 As shown, the computer device 1000 includes: a processor 1010, a memory 1020, a power supply 1030, a display unit 1040, and an input unit 1060.

[0171] The processor 1010 is the control center of the computer device 1000. It connects various components via various interfaces and lines, and executes various functions of the computer device 1000 by running or executing computer programs / instructions stored in the memory 1020, thereby performing overall monitoring of the computer device 1000. In this embodiment of the invention, when the processor 1010 calls the computer program stored in the memory 1020, it executes, for example... Figure 2The steps of the power data anomaly detection method are described above. Optionally, the processor 1010 may include one or more processing units; preferably, the processor 1010 may integrate an application processor and a modem processor, wherein the application processor mainly handles the operating system, user interface, and applications, and the modem processor mainly handles wireless communication. In some embodiments, the processor and memory can be implemented on a single chip; in some embodiments, they can also be implemented separately on independent chips.

[0172] The memory 1020 may primarily include a program storage area and a data storage area. The program storage area may store the operating system, various applications, etc.; the data storage area may store instruction data created based on the use of the computer device 1000, etc. In addition, the memory 1020 may include high-speed random access memory, and may also include non-volatile memory, such as at least one disk storage device, flash memory device, or other non-volatile solid-state storage device.

[0173] The computer device 1000 also includes a power supply 1030 (such as a battery) that supplies power to various components. The power supply can be logically connected to the processor 1010 through a power management system, thereby enabling the management of charging, discharging, and power consumption.

[0174] The display unit 1040 can be used to display information input by the user or information provided to the user, as well as various menus of the computer device 1000. In this embodiment of the invention, it is mainly used to display the display interfaces of various applications in the computer device 1000, as well as text, images, and other objects displayed on the display interfaces. The display unit 1040 may include a display panel 1050. The display panel 1050 may be configured in the form of a liquid crystal display (LCD), an organic light-emitting diode (OLED), or the like.

[0175] The input unit 1060 can be used to receive information such as numbers or characters input by the user. The input unit 1060 may include a touch panel 1070 and other input devices 1080. The touch panel 1070, also known as a touch screen, can collect touch operations on or near the user (such as operations performed by the user using a finger, stylus, or any suitable object or accessory on or near the touch panel 1070).

[0176] Specifically, the touch panel 1070 can detect user touch operations and the signals generated by these operations, convert them into touch point coordinates, send them to the processor 1010, and receive and execute commands from the processor 1010. Furthermore, the touch panel 1070 can be implemented using various types of sensors, including resistive, capacitive, infrared, and surface acoustic wave sensors. Other input devices 1080 can include, but are not limited to, one or more of the following: physical keyboard, function keys (such as volume control buttons, power buttons, etc.), trackball, mouse, joystick, etc.

[0177] Of course, the touch panel 1070 can cover the display panel 1050. When the touch panel 1070 detects a touch operation on or near it, it transmits the information to the processor 1010 to determine the type of touch event. Subsequently, the processor 1010 provides corresponding visual output on the display panel 1050 based on the type of touch event. Although in Figure 4 In this embodiment, the touch panel 1070 and the display panel 1050 are two separate components to realize the input and output functions of the computer device 1000. However, in some embodiments, the touch panel 1070 and the display panel 1050 can be integrated to realize the input and output functions of the computer device 1000.

[0178] The computer device 1000 may also include one or more sensors, such as pressure sensors, gravity acceleration sensors, proximity sensors, etc. Of course, depending on the specific application requirements, the computer device 1000 may also include other components such as cameras.

[0179] This invention also provides a computer-readable storage medium storing a computer program / instructions, which, when executed by a processor, enable the aforementioned device to perform the functions described in this invention. Figure 2 The steps of the power data anomaly detection method are described above.

[0180] It will be understood by those skilled in the art that Figure 4 This is merely an example of a computer device and does not constitute a limitation on the device. The device may include more or fewer components than illustrated, or a combination of certain components, or different components. For ease of description, the above parts are divided into modules (or units) according to their functions and described separately. Of course, in implementing this invention, the functions of each module (or unit) can be implemented in one or more software or hardware components.

[0181] Those skilled in the art will understand that the present invention can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code. The present invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention, and it should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be applied to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing device to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing device, produce implementations of the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 The computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The functions specified in one or more boxes. These computer program instructions may also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable apparatus for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0182] It is understood that although terms such as first, second, third, etc., may be used in the embodiments of the present invention to describe the preset range, these terms are only used to distinguish the preset ranges from one another. For example, without departing from the scope of the embodiments of the present invention, the first preset range may also be referred to as the second preset range, and similarly, the second preset range may also be referred to as the first preset range.

[0183] The above embodiments are merely illustrative of the principles and effects of the present invention and are not intended to limit the invention. Any person skilled in the art can modify or alter the above embodiments without departing from the spirit and scope of the present invention. Therefore, all equivalent modifications or alterations made by those skilled in the art without departing from the spirit and technical concept disclosed in the present invention should still be covered by the claims of the present invention.

[0184] The preferred embodiments of the present invention have been described in detail above. It should be understood that those skilled in the art can make numerous modifications and variations based on the concept of the present invention without creative effort. Therefore, all technical solutions that can be obtained by those skilled in the art based on the concept of the present invention through logical analysis, reasoning, or limited experimentation on the basis of existing technology should be within the scope of protection defined by the claims.

Claims

1. A method for detecting anomalies in power data, characterized in that, The method includes: S1: A multi-source data association matrix is ​​constructed by dynamically associating power data with non-power scenario data; wherein, the power data includes load change curves, equipment operating status parameters, and voltage and current time-series monitoring data, and the non-power scenario data includes meteorological environmental data and traffic flow statistics; the dynamic association process includes unified alignment processing of time granularity and dynamic adaptation and allocation of scenario weights. S2: Based on the time-series evolution characteristics of power data and various scenario-related variables, the baseline threshold range is dynamically adjusted in real time using a sliding window algorithm. The baseline threshold range can exhibit non-linear scaling and adaptation characteristics as scenario variables fluctuate. S3: Through transfer learning algorithms, feature information corresponding to specific regions or devices that have completed anomaly labeling is transferred to similar scenarios. Combined with domain adaptive models, the distribution differences between different data sources are eliminated. Domain adaptive fusion coefficients are introduced to dynamically balance the adaptability of data features between the source domain and the target domain, thereby achieving preliminary screening of anomalies in unlabeled data. S4: A combination of coarse screening and fine verification is adopted. First, a coarse-grained algorithm is used to define the candidate set of anomalies; then, a fine-grained algorithm is used to verify the authenticity of the anomalies in the candidate set. S5: Combining the topology of the power system, the initial node location of the anomaly is determined by tracing the propagation path of the abnormal data and deducing in reverse.

2. The power data anomaly detection method according to claim 1, characterized in that, The domain-adaptive fusion coefficient introduced in S3 The expression is: in For domain-adaptive fusion coefficients; These are the feature values ​​of the source domain data; For the feature values ​​of the target domain data, The number of feature dimensions; is the Pearson correlation coefficient between the source domain and the target domain; By federalizing the migration of cross-regional data characteristics, a migration trust coefficient is set. The value range is [0.6, 0.9], which dynamically adjusts the weights for migrating features from the source domain to the target domain. The value is determined by the scene similarity and data distribution consistency between the source and target domains, and is expressed as: in For scene similarity; To ensure data distribution consistency; Scene similarity The expression is obtained by calculating the weighted sum of the meteorological type matching degree, traffic flow level matching degree, and topological structure similarity between the source and target domains: in For meteorological type matching degree, For traffic flow level matching degree; For topological similarity; Data distribution consistency pass The divergence is calculated by comparing the data distributions in the source and target domains, and then normalized to the [0,1] interval. in For source domain With the target domain of divergence, The largest in history Divergence value; By using a federated learning framework, collaborative training of cross-regional anomaly detection models can be achieved without sharing the original power data of each region, and multi-regional data features can be aggregated through encrypted gradient transmission.

3. The power data anomaly detection method according to claim 2, characterized in that, The dynamic correlation between power data and non-power scenario data in S1, and the subsequent construction of a multi-source data correlation matrix, specifically includes: S11: Perform outlier removal and normalization on the power data, mapping it uniformly to the [0,1] interval; the expression is: in Represents the raw values ​​of power data; This represents the reasonable minimum value for power data; This indicates the reasonable maximum value of the power data; This represents the result after normalization; Indicates the scene influence coefficient; Indicates the scene adaptation factor; Scene Influence System The value range is [0.1, 0.8] for meteorological data and [0.05, 0.5] for traffic flow data; scene adaptation factor. The value is dynamically adjusted by the system based on historical data and the associated accuracy, and the range is [0.8, 1.2]. The scene adaptation factor The dynamic adjustment mechanism is as follows: the system calculates the historical data correlation accuracy based on a preset time; if the accuracy is higher than the preset value, then... Increase the value based on the current value; if the accuracy is lower than the preset value, then... Reduce from the current value; if the accuracy is within the preset range, then Remain unchanged; The data from non-electrical scenarios are standardized in format, meteorological data is converted into impact factor indices, and traffic flow data is converted into regional activity intensity values. The time granularity of the two types of data is uniformly adjusted through a timestamp alignment algorithm. S12: By using a sliding time window, calculate the Pearson correlation coefficient between power data and non-power data within the window range, capture the correlation changes between the two in the time series dimension, construct a scenario impact factor library, configure initial correlation weights for different weather types and traffic conditions, and then dynamically adjust the weights based on real-time data feedback information. By increasing the weight of abrupt change nodes in the time series of power data through a temporal attention mechanism, and constructing a multi-source data correlation matrix, the data correlation strength during abnormally sensitive periods is enhanced through an attention coefficient. The expression for the attention coefficient is: in for Attention coefficient at any moment; It is a sensitive regulatory factor; for Real-time power data; The length of the time series; S13: Using power data types as the row dimension of the matrix and non-power data types as the column dimension, the initial values ​​of the matrix elements are set as the real-time correlation coefficients of the corresponding dimensions. A scene weight layer is set on top of this base matrix. The scene adaptation weights are multiplied by the correlation coefficients to obtain the weighted correlation value. A power equipment topology correlation column is set and filled with the corresponding coefficients of the influence of non-power data on the power data of different topology nodes, forming a three-dimensional correlation matrix including data type, scene type and topology node. S14: Trigger matrix update operation according to preset time intervals, recalculate correlation coefficients and scene weights using newly collected data, replace outdated elements in the matrix, and start the update process when a sudden change in data distribution is detected to complete dynamic matrix adjustment and establish an association effectiveness evaluation index system. If the association accuracy of the matrix is ​​lower than a preset threshold, the system will adjust the correlation coefficient calculation model and weight allocation rules. The association effectiveness evaluation index system includes association accuracy, time synchronization error rate, and weight fit. The association accuracy is the proportion of effective associated data in the matrix, the time synchronization error rate is the statistical value of data timestamp alignment deviation, and the weight fit is the degree of matching between scene weights and the actual data association strength.

4. The power data anomaly detection method according to claim 3, characterized in that, In step S2, the real-time dynamic adjustment of the baseline threshold interval using the sliding window algorithm specifically includes: S21: Set a first initial window, a second initial window, and a third initial window based on the power scenario. The first initial window has 15-20 data points, the second initial window has 30-50 data points, and the third initial window has 60-90 data points; standardize the format of data from non-power scenarios. S22: Calculate the variance and trend slope of the data within the first initial window, the second initial window, and the third initial window; in Indicates the dynamic fusion weight of the window; , ,and These represent the standard deviations of the data in the first initial window, the second initial window, and the third initial window, respectively. , and These represent the window's base weights; A window adaptation evaluation model is established, and dynamic weights are assigned to the first, second, and third initial windows based on data fluctuations and changes. If the data fluctuates sharply in the short term, the weight of the first window is increased; if the data shows stable changes in the medium term, the weight of the second window is increased; if the data shows a stable long-term trend, the weight of the third window is increased. The data features of the first, second, and third initial windows are fused based on the weights to generate the calculation base for the initial baseline threshold interval. S23: Real-time acquisition of physical constraint parameters of power equipment operation, including the rated voltage range, rated current range and maximum load threshold of the equipment, and conversion of them into corresponding threshold boundary values; fusion of the threshold boundary values ​​corresponding to the physical constraint parameters with the dynamically adjusted baseline threshold range, and taking the intersection of the two as the final baseline threshold range. If the two have no intersection, the threshold of the physical constraint parameters shall prevail, and the system shall trigger an alarm to indicate threshold conflict.

5. The power data anomaly detection method according to claim 4, characterized in that, The S23 section further includes introducing a physical constraint threshold for power equipment, trimming and calibrating the calculated threshold range to ensure that the threshold does not exceed the safe operation boundary of the equipment, monitoring the distribution density of data within the threshold range in real time, triggering iterative adjustment of window parameters if the proportion of normal data distribution is low, calculating the difference between the time of abnormal data occurrence and the threshold adjustment response time, and optimizing the window weight allocation rule if the difference exceeds the preset threshold. By deploying edge computing nodes, a lightweight anomaly screening algorithm is deployed on the power terminal equipment side to achieve local identification and preliminary processing of abnormal data.

6. The power data anomaly detection method according to claim 5, characterized in that, The specific window parameter iterative adjustment is as follows: First, based on the scene weights and topological correlation coefficients in the three-dimensional correlation matrix, a dynamic proportion threshold model is constructed. The normal data distribution proportion threshold is calculated in real time through this model. The model takes the current scene influence coefficient and the topological node correlation strength as inputs, and combines the benchmark value of the normal distribution proportion under the same scene in historical data to output the dynamic threshold through an adaptive weighting algorithm. If the actual proportion of normal data is lower than this dynamic threshold, a tiered adjustment mechanism is activated: The number of data points in the first initial window is adaptively reduced by 2-4 based on the intensity of fluctuations in the scenario. Specifically, the upper limit of the reduction is used when short-term fluctuations are severe, and the lower limit of the reduction is used when fluctuations are mild. The number of data points in the second initial window is dynamically increased or decreased by 1-3 based on the slope of the time-series abrupt changes in power data. Specifically, the number of data points is increased when the slope of the abrupt change exceeds the threshold and decreased when it is below the threshold. The number of data points in the third initial window is increased by 2-5 based on the stability of the long-term trend. After adjustment, a temporal attention coefficient is introduced to weight and correct the window data, strengthen the data weight of abnormally sensitive periods, and then the baseline threshold range is recalculated. At the same time, the threshold deviation value and scene adaptability change before and after the adjustment are recorded.

7. The power data anomaly detection method according to claim 6, characterized in that, The method for rapidly defining the anomaly candidate set using a coarse-grained algorithm specifically includes: setting a dynamic clustering radius benchmark value based on the spatiotemporal correlation characteristics of power data; the method for setting the dynamic clustering radius benchmark value is as follows: setting an initial benchmark value based on the average radius of normal data clusters in historical data, combined with the power data type, and then dynamically fine-tuning it according to the real-time scenario weights; using an improved density clustering algorithm, with the time-series nodes of power data in the three-dimensional correlation matrix as the core, dividing data that are adjacent in the time dimension and belong to the same topological correlation region in the spatial dimension into basic data clusters; calculating the core feature statistics of each data cluster, including the data mean, variance of fluctuation, and deviation from the baseline threshold interval, and setting the deviation threshold as... The The expression is in Indicates the overall deviation of the data; The mean of the baseline threshold interval; The slope of change in the time series segment containing the abnormal data; This represents the maximum permissible slope for normal equipment operation. The slope influence factor; Data clusters with deviations exceeding a set threshold are selected, and all data within a cluster are marked as abnormal candidate data. These are then integrated to form an abnormal candidate set. Simultaneously, the spatiotemporal coordinates and associated scene information of each candidate data are recorded. The candidate set is then subjected to preliminary deduplication to remove redundant abnormal records caused by repeated data collection or transmission delays, while retaining uniquely identified abnormal candidate data entries. The fine-grained algorithm verifies the authenticity of anomalies in the candidate set by extracting the physical operation constraint parameters of the power equipment corresponding to each abnormal data in the candidate set and constructing an equipment constraint verification library; introducing a time-series consistency verification model, comparing the change trends of adjacent time-series data before and after the abnormal data, calculating the abrupt change slope between the abnormal data and the preceding normal data, and if the slope exceeds the maximum allowable change rate under normal equipment operation, proceeding to the next verification step; otherwise, it is marked as a false anomaly. By combining the scene weights and topological correlation coefficients in the three-dimensional correlation matrix, we analyze whether the abnormal data is related to the current scene conditions. If the abnormal behavior contradicts the scene's influence pattern, it is judged as a suspicious abnormality. We call the equipment operation mechanism model to simulate whether the operating state corresponding to the abnormal data conforms to the physical characteristics of the power equipment. If the deviation exceeds a preset threshold, it is judged as a real abnormality. We comprehensively score the candidate data after the above steps. If the score meets the standard, it is confirmed as a real abnormality. If it does not meet the standard, it is marked as pending verification or a false abnormality, thus completing the authenticity verification.

8. A power data anomaly detection system, characterized in that, The system includes: S100: Multi-source data association module, used to remove outliers, normalize and standardize the format of power data and non-power scenario data. It unifies the time granularity through timestamp alignment algorithm, calculates Pearson correlation coefficient and dynamically adjusts weights in combination with time series attention mechanism, constructs a three-dimensional association matrix containing data type, scenario type and topology node, and completes dynamic matrix update and effectiveness evaluation according to preset rules. S200: Baseline threshold dynamic adjustment module, used to set the initial window with different numbers of data points, calculate the data variance and trend slope within the window and assign dynamic weights, integrate the physical constraint parameters of power equipment to generate the final baseline threshold range, deploy a lightweight coarse screening algorithm through edge computing nodes, and iteratively adjust the window parameters according to the data distribution. S300: Cross-domain collaborative detection module, used to transfer labeled abnormal features to similar scenarios through transfer learning algorithms, combine domain adaptive models to eliminate differences in data source distribution, and realize cross-regional model collaborative training based on the encrypted aggregation mechanism of federated learning framework, ensuring data privacy and security and node fault tolerance; S400: Anomaly Precision Screening Module, which uses a coarse-grained screening algorithm of improved density clustering to define anomaly candidate sets and remove duplicates, and then uses fine-grained algorithms of time series consistency verification, scene association analysis and equipment operation mechanism simulation to verify the authenticity of candidate set data and give a comprehensive score. S500: Anomaly tracing and location module, used to combine the power system topology connection structure to track the propagation path of abnormal data, reverse deduce and determine the initial node location of the anomaly trigger, and synchronously record the spatiotemporal coordinates and related scene information of the abnormal data.

9. A computer device, characterized in that, The device includes a memory, a processor, and a computer program stored in the memory, wherein the processor executes the computer program to implement the steps of the power data anomaly detection method according to any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, It stores a computer program that, when executed by a processor, implements the steps of the power data anomaly detection method according to any one of claims 1 to 7.