A data anomaly tracing and locating method, system, device and storage medium

Through the knowledge graph and graph neural network, the data association relationship network is built, and the abnormal propagation path of power marketing data is identified, which solves the problems of abnormal detection accuracy and inefficient traceability, and realizes efficient and flexible abnormal traceability positioning, which improves the adaptability and stability of the system.

CN119416131BActive Publication Date: 2025-08-12STATE GRID ZHEJIANG ELECTRIC POWER CO MARKETING SERVICE CENT +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510022643.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-07
Publication Date
2025-08-12
Estimated Expiration
2045-01-07

AI Technical Summary

Technical Problem

There are problems in power marketing data such as insufficient accuracy of abnormal detection, low traceability positioning efficiency and incomplete knowledge correlation, which leads to difficulty in quickly positioning the abnormal root cause, affecting business response speed and system stability.

Method used

Using a knowledge-driven method, features are extracted through knowledge graphs, combined with cluster analysis and graph neural networks, a data association relationship network is built, an abnormal propagation path is identified, and the model is optimized through hyperparameter adjustment and feedback mechanism to achieve efficient and flexible abnormal traceability positioning.

Benefits of technology

It improves the accuracy and traceability efficiency of abnormal detection of power marketing data, enhances the correlation analysis capabilities of multi-source data, reduces manual intervention, and improves the adaptability and stability of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119416131B_ABST
    Figure CN119416131B_ABST
Patent Text Reader

Abstract

The present invention discloses a data anomaly tracing and locating method, system, device and storage medium. It relates to the technical field of power system data governance. At present, there are problems such as insufficient accuracy of data anomaly detection, low efficiency of tracing and locating, and incomplete knowledge association in the analysis and application of power marketing data. The present invention includes: preprocessing and feature extraction of power marketing data, constructing a data association relationship network through cluster analysis and statistical methods, analyzing the differences in data features, using path reasoning technology to analyze the propagation path of the abnormal points identified in the data stream, constructing a data anomaly tracing and locating model, using semi-supervised learning methods to improve the accuracy of the model, performing result optimization and hyperparameter adjustment, and introducing multiple conditional factors that affect business data information as hyperparameters into the model. This technical solution greatly improves the anomaly detection and tracing and locating capabilities of power marketing data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of power system data governance, and in particular to a data anomaly tracing and locating method, system, device and storage medium. Background Art

[0002] In recent years, knowledge graphs and graph neural network technologies have been widely applied to semantic data association and knowledge mining, enabling efficient analysis of complex data relationships. Against this backdrop, the analysis and application of marketing data are becoming increasingly important for business decision-making and service optimization. On the one hand, data anomaly traceability and location technology has rapidly developed in the field of big data monitoring. By constructing data link association models and causal networks, it facilitates the accurate identification and analysis of data anomalies. On the other hand, power marketing data is massive and diverse, and with policy changes and market fluctuations, data anomalies are frequent. Traditional data monitoring methods struggle to promptly identify the causes of anomalies, making the traceability process time-consuming and labor-intensive. Based on this, a knowledge-driven traceability and location method is employed. This method utilizes knowledge graphs to construct an association network for multi-source heterogeneous data and applies graph neural networks to infer the propagation paths of anomalies. This method can effectively improve the ability to identify and quickly locate anomalies in power marketing data.

[0003] However, due to the differences in multi-source data, the lack of data link anomaly diagnosis tools, and the high cost of data governance, the link anomaly diagnosis method for power marketing data faces many challenges in practical applications: (1) Insufficient accuracy of data anomaly detection. Traditional rule detection methods rely on fixed threshold settings and cannot adapt to the complex changes in data characteristics in a timely manner. Especially in a multi-source heterogeneous data environment, the fixed threshold detection method is prone to false positives or false negatives, which makes the anomaly detection process lack flexibility and difficult to cope with dynamically changing marketing data flows. (2) Low efficiency of traceability and positioning. In existing traceability analysis, due to the multiple sources and long links of power marketing data, traceability and positioning after anomalies often require a lot of manual intervention and time investment. A single data link monitoring method fails to cover the complete business scenario and cannot build an efficient traceability path, resulting in difficulty in quickly locating the root cause of the anomaly. The low traceability efficiency not only affects the business response speed, but also makes it impossible to timely control potential system risks, greatly affecting the stability of power marketing data and business continuity. (3) Incomplete knowledge association and lack of a global perspective. Existing power marketing data analysis is often limited to a single data dimension, lacking the global perspective supported by knowledge graphs. This makes it impossible to capture implicit relationships and associated knowledge across multiple data sources, limiting the comprehensiveness of anomaly diagnosis and preventing effective linkage analysis between anomalies and their underlying causes. Furthermore, data correlation analysis lacks the ability to learn from historical anomaly data, making it difficult for the system to recognize complex anomaly patterns, impacting the adaptability and flexibility of data processing. Summary of the Invention

[0004] The technical problem to be solved and the technical task to be addressed by this invention are to improve and enhance existing technical solutions by providing a data anomaly traceability and location method, system, device, and storage medium, with the goal of achieving knowledge-driven data anomaly traceability and location, thereby improving the accuracy and efficiency of anomaly detection. To this end, this invention adopts the following technical solution.

[0005] The first technical solution of the present invention is: a knowledge-driven data anomaly tracing and locating method, which includes the following steps:

[0006] 1) Preprocessing and feature extraction of power marketing data, including data cleaning, deduplication, and standardization, as well as extracting features of electricity consumption data using knowledge graphs and screening out key features;

[0007] 2) Use cluster analysis and statistical methods to construct a data association network, analyze the differences in data characteristics, distinguish the similarities and differences between different data results or application scenarios, and compare them;

[0008] 3) Using path reasoning technology and complex network analysis methods, we can identify and locate the root causes of data anomalies by analyzing the propagation paths of anomalies in the data flow;

[0009] 4) Build a data anomaly traceability and location model, use graph neural network methods for feature learning, and use knowledge graphs to identify the entire chain of abnormal data generation, realize the judgment and labeling of raw data anomalies, and perform semi-supervised anomaly location assessment;

[0010] 5) Optimize results and adjust hyperparameters. Introduce multiple conditional factors that affect business data information into the model as hyperparameters, adjust the model's hyperparameters, and use the feedback mechanism to dynamically adjust the model's output results.

[0011] This technical solution provides an efficient, flexible, and global-perspective knowledge-driven data anomaly traceability and location method, which solves the problems of low anomaly detection accuracy, low traceability efficiency, and incomplete knowledge association in power marketing data, and greatly improves the anomaly detection and traceability capabilities of power marketing data.

[0012] By leveraging knowledge graphs to extract key features of power marketing data and combining them with data cleaning, deduplication, and standardization, this invention can effectively identify potential anomalies in the data. Traditional rule-based detection methods rely on fixed thresholds and often fail to adapt to complex changes in data characteristics. However, this invention, by incorporating knowledge graphs for feature extraction and screening, enables the model to dynamically adapt to changes in the data stream, avoiding the false positives and false negatives associated with fixed thresholds, thereby improving the accuracy and flexibility of data anomaly detection.

[0013] In traditional anomaly tracing, tracing and locating the source after an anomaly occurs often relies on extensive manual intervention and time investment, and it is difficult to quickly locate the root cause of the anomaly. This invention, by constructing a data association network and combining cluster analysis and statistical methods, can analyze the differences in data characteristics from a global perspective and construct an efficient data tracing path. By using path reasoning technology and complex network analysis methods, the root cause of data anomalies can be quickly located by identifying the anomaly propagation path, significantly improving the efficiency of tracing and locating the source.

[0014] Traditional power marketing data analysis methods are typically limited to a single data dimension and lack the ability to deeply explore and comprehensively analyze the implicit relationships between heterogeneous data from multiple sources. This invention achieves the association of multi-source data through a knowledge graph, capturing the correlations and implicit relationships between data from a global perspective, thereby improving the comprehensiveness and accuracy of anomaly diagnosis. Furthermore, feature learning and semi-supervised anomaly location assessment based on graph neural networks can continuously optimize the ability to identify and judge anomaly patterns, avoiding the incomplete knowledge association problem of traditional methods.

[0015] By introducing hyperparameter adjustment and feedback mechanisms, this invention enables the model to dynamically adjust and optimize in complex data environments with multiple conditions. In particular, when multiple factors influence power marketing data operations, the introduction of external condition characteristics and business hyperparameter adjustments ensures the model's continued adaptability and flexibility, optimizes its predictive effectiveness, reduces systemic risk, and improves the stability and business continuity of power marketing data.

[0016] As a preferred technical means: In step 1), the statistical, computational, and analytical relationship system in the knowledge graph is used to strengthen the identification of data features, and key features are screened out through statistical analysis and dimensionality reduction methods to ensure the efficiency and effectiveness of model processing.

[0017] Knowledge graphs, through their rich semantic relationships and hierarchical structure, provide deep connections between data. This enables models to identify underlying patterns and key features within the data at a higher level, avoiding the limitations of traditional methods that rely solely on a single data dimension. Guided by knowledge graphs, more implicit connections between data can be captured, improving the accuracy and comprehensiveness of data feature recognition.

[0018] The knowledge graph not only provides rich feature data, but also provides semantic explanations for the selection of each feature, so that the model not only has efficient prediction capabilities, but also has better interpretability, which is easier to understand and debug. Users can trace the source of specific features and the logical relationship behind them through the knowledge graph, thereby improving the transparency of the model.

[0019] Power marketing data often contains a large number of redundant and irrelevant features. By combining statistical analysis with dimensionality reduction methods, we can remove redundant features without losing key information, reduce the data dimension, and improve the efficiency of subsequent model processing. This efficient feature selection and dimensionality reduction method effectively avoids the curse of dimensionality, reduces the size of training datasets, and improves model training speed and efficiency.

[0020] By selecting the most representative and influential key features and reducing unnecessary feature interference, the model can be trained and predicted more efficiently. Dimensionality reduction methods can reduce the amount of computation and improve computational efficiency, making the model more efficient when processing large amounts of data, delivering results in a shorter time, and improving business response speed.

[0021] Power marketing data is highly volatile and can be affected by a variety of factors, including policy changes and market fluctuations. By leveraging the statistical, computational, and analytical relationships within the knowledge graph, the model can more flexibly adjust feature recognition strategies as data characteristics change. Dimensionality reduction and feature screening methods help the model maintain adaptability and robustness in complex and changing environments, enabling it to cope with the dynamic flow of power marketing data.

[0022] By screening out key features through statistical analysis and dimensionality reduction, and combining it with the rich semantic information provided by the knowledge graph, the model can better capture potential abnormal patterns, making anomaly detection more accurate, avoiding over-reliance on detection methods based on a single data source, and improving the system's stability and anomaly detection capabilities.

[0023] As a preferred technical means: in step 2), the specific steps of constructing the data association network include:

[0024] 2.1) Use the K-means clustering algorithm to divide the data into K clusters and construct a data association network G = (V, E) based on the clustering results, where nodes V represent clusters and edges E represent the relationships between clusters.

[0025] 2.2) Perform feature vector comparison and similarity calculation, use cosine similarity to compare the feature differences between different clusters, and identify clusters with obvious feature differences based on the similarity matrix S, and analyze their possible abnormal behavior.

[0026] The K-means algorithm automatically identifies natural distribution patterns in data and clusters similar data points. Through clustering, complex and high-dimensional power marketing data can be transformed into a more easily processed structured form, helping the system discover underlying patterns within massive amounts of data. This avoids the difficulty of manually filtering features and adaptively reveals the inherent structure of the data. The association network constructed from the K-means clustering results transforms the relationships between data points from a single data dimension into a high-order association network. This networked structure is more expressive and helps the system more comprehensively understand and process complex relationships between data.

[0027] By mapping data nodes into clusters and establishing edges between these clusters to represent their relationships, complex datasets can be simplified into a hierarchical network structure. This structured representation not only makes clustering results intuitive to understand, but also enables better analysis and location of possible sources of data anomalies. The K-means algorithm iteratively optimizes cluster centers, resulting in stable and generalizable clustering results. In a multi-source, heterogeneous data environment, K-means can effectively cope with variations in different data types, improving data processing stability.

[0028] Cosine similarity accurately measures the differences in features between clusters and is particularly suitable for processing high-dimensional and sparse data. By comparing feature vectors based on cosine similarity, subtle differences between clusters can be more flexibly and precisely detected, effectively identifying clusters with significant feature variations and quickly targeting potential anomaly areas. The use of cosine similarity enables anomaly detection to dynamically adapt within the feature space. For clusters with significant feature differences, the model can detect anomalous trends earlier, allowing for timely labeling and location.

[0029] By constructing the similarity matrix S, we can globally compare the relationships between different clusters and quickly identify abnormal clusters that have significant differences in characteristics from most clusters. Without manual intervention, we can automatically detect and distinguish normal and abnormal behavior patterns, thereby improving the efficiency and accuracy of anomaly detection.

[0030] The combination of the K-means algorithm and cosine similarity offers excellent computational efficiency in big data environments. The K-means algorithm has a time complexity of O(nk), where n is the number of data points and k is the number of clusters. Cosine similarity calculations can be accelerated through matrix operations, making the entire analysis process more efficient when processing large-scale power marketing data, enabling data processing and anomaly detection to be completed in a shorter time. Clustered data has been aggregated and categorized, reducing data complexity, computational costs, and improving the efficiency of subsequent analysis.

[0031] Cluster analysis helps automate grouping and, by comparing the characteristics of different clusters, reveals the characteristics of abnormal patterns, which is particularly important for identifying complex and subtle abnormal behaviors. This method can capture subtle data deviations caused by factors such as the dynamic changes in the power marketing environment and market fluctuations, thereby more accurately identifying potential abnormal behaviors.

[0032] The combination of the K-means algorithm and similarity calculation offers excellent scalability and can be flexibly applied to various power marketing data scenarios. The number of clusters and similarity calculation method can be adjusted as needed. Furthermore, as the amount of data increases, the algorithm can be optimized by adjusting parameters (such as the K value), thereby improving clustering accuracy and anomaly detection reliability.

[0033] As a preferred technical means: in step 3), the specific steps of identifying the abnormal propagation path include:

[0034] 3.1) Construct a weighted directed graph G = (V, E), where nodes V represent data points and edges E represent the influence relationships between data points. Use a graph traversal algorithm to find the anomaly propagation path.

[0035] 3.2) Based on complex network analysis methods, network indicators are calculated, including node degree, centrality, and influence. A propagation model is established to simulate the process of anomalies spreading from one node to another.

[0036] Weighted directed graphs clearly represent the influence relationships between data points. Nodes V represent data points, and edges E represent the influence or relationship between data points. By weighting, the magnitude of influence between different data points can be more accurately characterized. In anomaly detection, using weighted directed graphs can reveal the propagation paths of anomalies in data streams, helping to accurately identify the root cause of anomalies and the scope of their impact.

[0037] Graph traversal algorithms (such as depth-first search (DFS) or breadth-first search (BFS)) can efficiently find the propagation path from an anomaly point in the graph. Graph traversal algorithms can help the system dynamically identify the anomaly propagation process, quickly locate the anomaly's propagation trajectory and key nodes, and reduce the time required to manually troubleshoot anomaly paths using traditional methods.

[0038] By constructing a weighted directed graph, the model can accurately identify the originating node of an anomaly and trace its propagation from one data point to other data points. The weight of each node and edge represents the varying degrees of influence, precisely determining which nodes play a key role in anomaly propagation, thereby helping to pinpoint the root cause of the data anomaly. Compared to traditional rule-based detection methods, this graph structure can handle complex data propagation paths, providing more refined and accurate anomaly tracing and location.

[0039] The propagation model simulates the diffusion of anomalies within a network, further enhancing our understanding of the anomaly propagation paths. By establishing a propagation model, we can better capture the dynamics of anomaly propagation, helping decision makers understand the underlying causes of data anomalies and implement targeted interventions.

[0040] Network analysis of data relationship graphs allows for deeper exploration and understanding of anomaly propagation patterns. Network analysis methods (such as node degree, centrality, and influence) not only help identify anomaly propagation paths but also help identify key nodes or data points within the network, which may be the epicenter or source of anomaly propagation. Network analysis can quantify the network structure to reveal which nodes play a significant role in the data flow, thereby helping to identify the root cause of data anomalies. Network metrics provide strong support for identifying anomaly propagation paths.

[0041] Complex network-based propagation models, such as the SIR model, can simulate the propagation of anomalies within a network. This simulation allows the system to predict how an anomaly spreads from one node to other nodes and assess its speed and scope. With this information, proactive measures can be taken to prevent the further spread of the anomaly, thereby preventing widespread system failures or data errors. By combining propagation models with actual anomaly data, the system can more precisely adjust anomaly detection mechanisms, effectively predict future anomaly propagation trends, and enhance anomaly early warning capabilities.

[0042] Weighted directed graph and complex network analysis methods can dynamically adapt to different power marketing data streams without requiring fixed rules or thresholds. This highly adaptable approach automatically adjusts anomaly detection strategies based on the varying characteristics of data streams (e.g., market fluctuations, policy changes, etc.), responding to anomalies in real time and avoiding both false positives and false negatives. Through graph traversal and complex network analysis, anomalies can be detected and traced in real time within data streams, ensuring the system's ability to respond and locate anomalies promptly, reducing manual intervention and improving exception handling efficiency.

[0043] Graph-based analysis methods provide a global perspective, enabling visibility into the relationships between different data points. This helps decision makers fully understand the propagation of anomalies at the system level. This not only helps quickly locate the root cause of anomalies, but also provides more comprehensive information support for subsequent decision-making.

[0044] The construction of weighted directed graphs and complex network analysis methods are well-suited for processing large-scale power marketing data. Even with extremely large data volumes, this method can efficiently construct relationships between data points and quickly identify anomalous propagation paths through graph traversal and network analysis, adapting to the demands of massive data processing. The structured representation of graphs and the calculation of network metrics can be implemented using efficient algorithms, offering excellent computational performance and scalability, enabling high real-time processing capabilities in complex big data environments.

[0045] As a preferred technical means: In step 3.1), first construct a weighted directed graph , where the node V represents the data point and the edge E represents the influence relationship between the data; secondly, the graph traversal algorithm is used to find the abnormal propagation path. Given an abnormal point , calculated from To other nodes The shortest path is:

[0046]

[0047] in, is the weight of each edge in the path;

[0048] By traversing all nodes, we can identify the abnormal points Most relevant nodes , thereby determining the propagation path of the anomaly.

[0049] In step 3.2),

[0050] First, network index calculations are performed to calculate indicators including node degree, centrality, and influence to help identify key nodes. Defined as a node The number of connected edges, the centrality calculation formula is: ,in, is a node and the distance between them;

[0051] Secondly, based on the propagation path, a propagation model is established. Assuming that the way data anomalies propagate between nodes is the SIR model, the formula is: ,in, is the number of susceptible nodes, is the number of infected nodes, is the number of recovery nodes, is the transmission rate, is the recovery rate, which simulates the process of anomaly propagation from one node to another with the help of the propagation model.

[0052] By constructing a weighted directed graph, the weights of nodes and edges can more accurately represent the influence relationships between data points. The weight of each edge reflects the strength of the association or influence between data points, providing a reliable basis for tracking the propagation path of anomalies. By traversing all nodes, the system can accurately identify the nodes most relevant to the anomaly, thereby determining the anomaly's propagation path and impact range. This graph-based structured analysis is more flexible and accurate than traditional rule-based or fixed threshold methods.

[0053] The shortest path algorithm helps find the shortest path along which an anomaly propagates, starting from an anomaly point. This helps identify the core nodes of the anomaly's propagation. By calculating this path, the system can effectively identify the key links in the anomaly's propagation chain, reducing false positives and missed positives, and improving anomaly tracing efficiency.

[0054] Using the SIR propagation model to simulate the propagation of anomalies in data streams can help better understand how anomalies spread from one node to other nodes. By modeling the relationships between susceptible nodes (S), infected nodes (I), and recovered nodes (R), the SIR model simulates the dynamics of anomaly propagation between nodes. By adjusting the propagation rate (β) and recovery rate (γ), the speed and scope of anomaly propagation can be predicted, providing decision support for anomaly detection. The SIR model provides dynamic modeling capabilities for anomaly propagation, capturing the diffusion trend of anomalies in the system and helping to predict the potential impact of anomalies in the future. This dynamic simulation capability supports the timely detection of potential anomalies and system risks, facilitating early warning and action to avoid system failures or business interruptions.

[0055] By calculating node degree, centrality, and influence, we can effectively identify key nodes in the network. These network metrics help us identify which nodes play a key role in the propagation of anomalies, thus supporting anomaly tracing and impact assessment.

[0056] Degree: The degree (k_i) represents the number of connections a node has with other nodes. Nodes with higher degrees usually have greater influence and may play an important role in anomaly propagation.

[0057] Centrality: Centrality calculation can help identify the central nodes of abnormal transmission by evaluating the position of nodes in the network. These nodes may be the source of transmission or the key link of transmission.

[0058] Influence: Nodes with high influence may play a greater role in promoting anomaly propagation. Calculating influence can help prioritize the identification of these nodes, thereby speeding up anomaly location and processing.

[0059] By constructing a weighted directed graph and calculating network metrics, we can provide a global perspective on anomaly propagation. Rather than focusing solely on local data points, we can identify anomaly paths from the perspective of the overall network structure. This global perspective can reveal potential, hidden anomaly propagation chains within data streams, effectively identifying anomalies that are difficult to detect due to the interplay of multiple nodes. By combining graph traversal, network metric calculation, and propagation models, the anomaly tracing process is no longer limited to a single data stream, but rather integrates analysis from multiple dimensions. This allows for more comprehensive and accurate localization of anomaly sources and propagation paths, significantly improving the accuracy of anomaly identification.

[0060] The SIR model and network metric calculation dynamically adjust the identification and evaluation strategies for anomaly propagation paths as data flows change. This allows the system to adapt to different types of data flows and anomaly patterns, avoiding the limitations of traditional methods that rely on fixed rules and thresholds. In the complex environment of power marketing data, anomaly patterns can be complex and difficult to predict. By combining graph structures and propagation models, it is possible to identify potentially complex anomalies and respond accordingly, enhancing the system's adaptability to different anomaly types.

[0061] Through graph traversal algorithms and propagation models, the system can quickly identify anomalies and key nodes, significantly improving the efficiency of anomaly tracing. Traditional methods often require extensive manual intervention, especially in power marketing environments with large data volumes and complex structures. This solution, through automated graph analysis and propagation models, can effectively reduce labor costs and improve tracing efficiency. Traditional rule-based detection methods can lead to false positives or false negatives due to their over-reliance on fixed thresholds. This solution, through weighted directed graphs and complex network analysis, can more accurately identify the propagation paths of anomalies, reduce false detections, and improve the accuracy of anomaly location.

[0062] As a preferred technical means: the specific steps of step 4) include:

[0063] 4.1) Graph Neural Network Feature Learning: Using graph neural networks to extract features from data nodes and capture the complex relationships between nodes:

[0064] First, aggregate neighbor information:

[0065] ,

[0066] in is a node In the The layer representation, Representation node Neighbors, is a trainable weight matrix, is the activation function, is an aggregate function;

[0067] Secondly, through multi-layer aggregation, the feature representation of each node is obtained , where L represents the number of layers, Represents a knowledge graph;

[0068] Finally, the extracted feature representation is used as the input for subsequent anomaly detection. In the last layer, the node features is passed to a classifier for anomaly detection:

[0069] ,

[0070] in represents the predicted anomaly label;

[0071] 4.2) Abnormal Feature Labeling and Semi-Supervised Localization Evaluation:

[0072] First, the features are judged using the set threshold and abnormal situations are marked;

[0073] Secondly, semi-supervised learning technology is used to train the model on labeled and unlabeled data, and data with known anomalies are used to guide model learning to improve the ability to detect unknown anomalies.

[0074] Graph neural networks are highly capable of capturing complex relationships and dependencies between nodes. Through multi-layer aggregation (from aggregating neighbor node information to learning global information), GNNs are able to learn high-dimensional features of nodes from a local to global hierarchy, accurately representing the attributes of each data point and its role in the network. For data such as power marketing data, which has complex structures and multi-level relationships, graph neural networks are effective in identifying potential anomalous patterns.

[0075] By aggregating node neighbors, GNN can combine local information of each node with global information to form a feature representation with context-awareness. This enables the model to capture the complex relationships between nodes, thereby better identifying abnormal data and abnormal propagation paths.

[0076] Through the multi-layered graph neural network structure, the node features ultimately extracted can deeply explore the potential characteristics of the data and provide them for subsequent anomaly detection tasks. This high-quality feature representation can improve the accuracy and stability of the anomaly detection model.

[0077] Graph neural networks not only extract features from static data but also adapt to dynamically changing data through multi-layer learning. In real-world applications, power marketing data is influenced by a variety of external factors, and data characteristics frequently change. The multi-layer aggregation nature of GNNs enables the system to adapt to these changes, adjusting its understanding of the data in real time, thereby enhancing its ability to identify new and unseen anomalies.

[0078] Traditional anomaly detection methods often rely on fixed rules or thresholds, while graph neural networks can better adapt to complex anomaly patterns by learning complex relationships and propagation patterns between nodes, reduce false positives and missed negatives, and significantly improve the accuracy and robustness of anomaly detection.

[0079] By setting thresholds to judge node features, the model can automatically flag potentially abnormal data. This effectively reduces manual intervention and improves the system's automation level. Furthermore, feature extraction based on graph neural networks can provide higher accuracy for anomaly tagging, avoiding the issues of fixed thresholds and manual adjustments that traditional rules rely on.

[0080] Traditional anomaly detection typically requires a large amount of labeled data, which can be difficult to obtain in real-world applications. Semi-supervised learning technology can improve anomaly detection capabilities by combining a small amount of labeled data with a large amount of unlabeled data. Data with known anomalies is used to guide model learning, thereby improving the ability to detect unknown anomalies and adapting to complex and incompletely labeled data environments. By incorporating unlabeled data, semi-supervised learning significantly reduces the cost of manual labeling, enabling anomaly detection models to be trained and deployed more efficiently in real-world applications. Between labeled and unlabeled data, semi-supervised learning can gradually discover new anomaly patterns through self-learning. This can effectively improve the ability to identify new anomalies, especially in scenarios where data is scarce or difficult to label.

[0081] Graph neural networks' feature learning isn't limited to local node attributes; it captures relationships, interdependencies, and global structures between nodes. This allows the system to identify anomalies in a broader context, identifying not only local anomalies but also those with potentially catastrophic effects within the global structure, thus providing more accurate judgment for anomaly location. By combining the deep features extracted by graph neural networks with anomaly labeling and evaluation using semi-supervised learning, the system can precisely pinpoint the source of anomalies, reducing the ambiguity inherent in traditional methods. This is crucial for anomaly monitoring in large-scale power marketing data, enabling rapid response and action.

[0082] Power marketing data is subject to numerous factors, resulting in large volumes and constant changes. The flexibility and adaptability of graph neural networks enable the system to adapt promptly to data changes, maintaining efficient anomaly detection capabilities even in complex, dynamic data streams. This enables the system to operate over the long term without frequent adjustments.

[0083] Through automated feature learning and anomaly detection, the system can efficiently identify abnormal nodes and propagation paths, thereby helping decision makers focus resources on processing the most critical anomalies and avoiding unnecessary waste of resources.

[0084] As a preferred technical means: the specific steps of step 5) include:

[0085] 5.1) Define multiple hyperparameters that affect the business data, denoted as , where each hyperparameter Represents a specific business condition and combines these hyperparameters with the input features H of the graph neural network:

[0086]

[0087] in, The eigenvector representing the external conditions, Represents the join operation of features;

[0088] 5.2) Use cross-validation to optimize and select hyperparameters: Divide the data into k subsets, use each subset as the validation set in turn, and use the remaining k-1 subsets as the training set. In each round, adjust the hyperparameter H and train the model. Record the performance indicators of the validation set and select the performance indicator P to evaluate the model effect:

[0089]

[0090] in, It's a true positive. It is a false positive; grid search and Bayesian optimization are used to find the best hyperparameter combination : ;

[0091] 5.3) Feedback the model’s output back to the model through a feedback mechanism to adjust and optimize the model’s parameters and hyperparameters. The feedback model is:

[0092]

[0093] in is the predicted output of the model, is the actual label, is the feedback function used to calculate the performance difference of the model;

[0094] Further parameter adjustment is performed to update the model parameters W and hyperparameters H based on the feedback information:

[0095]

[0096]

[0097] in, is the learning rate, is the gradient of the loss function, is the feedback influence coefficient.

[0098] In the task of tracing and locating anomalies in power marketing data, the diversity and complexity of the data require that the model can be flexibly adjusted under different business conditions. By defining multiple data hyperparameters that affect the business (such as business environment, market fluctuations, policy changes, etc.), this technical solution can optimize the model for different actual situations, so that the effect of anomaly detection is more in line with actual needs. By combining hyperparameters with external conditions, the model can simultaneously consider changes in data and the external environment. For example, when the power market fluctuates or policies change, the system can adjust the input features in real time and optimize the model to adapt to the changing business needs and external environment, thereby improving the accuracy and flexibility of anomaly detection and tracing and positioning.

[0099] Cross-validation effectively prevents model overfitting and underfitting, ensuring model generalization. By dividing the data into multiple subsets and training the model on different training and validation sets, the model's adaptability to different datasets is ensured, making the final selected hyperparameter combination more stable. Using the true positive rate (TP) / false positive rate (TP+FP) as the performance metric for model evaluation, the accuracy of the model in anomaly detection tasks can be effectively monitored and optimized, ensuring that the true accuracy of anomaly detection is maximized during the model optimization process. Grid search and Bayesian optimization methods can efficiently search for the optimal hyperparameter combination, especially when the parameter space is large. This can avoid the tedious manual steps of hyperparameter adjustment, thereby improving model optimization efficiency. Bayesian optimization can automatically select optimal hyperparameters through statistical models when dealing with complex high-dimensional problems, reducing computational costs and improving model performance.

[0100] By comparing the model's predictions with the actual labels, a feedback mechanism is used to adjust the model in real time. This feedback mechanism significantly improves the model's adaptability and avoids erroneous decisions caused by large deviations between model predictions and actual results. This mechanism enables the model to dynamically adjust parameters based on changes in the actual situation, making the system more intelligent and adaptive. During the model optimization process, model parameters and hyperparameters are adjusted by calculating the gradient of the loss function and feedback information. This step enables the model to adjust itself more precisely after each iteration, avoiding the error accumulation problem caused by static parameter settings. Furthermore, the introduction of the feedback coefficient λ further enhances the flexibility and accuracy of parameter adjustments during the feedback process. Because the system automatically adjusts parameters based on actual feedback, frequent manual intervention is unnecessary. Users only need to focus on the model's performance indicators (such as TP and FP), while the model itself automatically optimizes in practice, improving efficiency and reducing labor costs.

[0101] The business environment, market fluctuations, policy changes, and other dynamic and complex factors in power marketing data make traditional models ineffective in responding to these rapid changes. By incorporating external conditional features and adjusting hyperparameters, the model can quickly respond to environmental changes and optimize its decisions. This flexible adaptability is particularly important in a complex and ever-changing marketing environment. The model dynamically adjusts hyperparameters based on business needs and actual data during each iteration, ensuring that anomaly detection maintains high accuracy under varying conditions. This dynamic adjustment allows the system to adapt to various uncertainties and continuously provide high-quality detection results despite changes in data and business requirements.

[0102] By continuously optimizing hyperparameters and incorporating feedback mechanisms, the model's accuracy in identifying anomalies can be continuously improved, reducing false positives and missed negatives, thereby enhancing the efficiency and accuracy of anomaly detection. This is particularly important for extremely large and complex datasets like power marketing data, enabling rapid identification of potential risks and effective countermeasures. Through accurate anomaly detection and an efficient feedback mechanism, the model can provide power marketing decision-makers with more precise information support, helping them to respond quickly to sudden market changes and mitigate potential losses.

[0103] Through precise hyperparameter tuning and feedback mechanisms, the model can identify anomalies that truly require attention, avoiding wasting resources on irrelevant or unimportant anomalies. In power marketing scenarios, this refined management helps improve overall operational efficiency.

[0104] The second technical solution of the present invention is to provide a data anomaly tracing and locating system, which applies the aforementioned knowledge-driven data anomaly tracing and locating method, and the system includes:

[0105] Data preprocessing and feature extraction module: This module preprocesses power marketing data, including data cleaning, deduplication, and standardization. It leverages the statistical, computational, and analytical relationship systems within the knowledge graph to enhance data feature identification. Key features are screened through statistical analysis and dimensionality reduction methods to ensure efficient and effective model processing.

[0106] Similarity and Difference Analysis Module: Uses the K-means clustering algorithm to divide data into K clusters, and constructs a data association network based on the clustering results. Through feature vector comparison and similarity calculation, cosine similarity is used to compare the feature differences between different clusters. Clusters with obvious feature differences are identified based on the similarity matrix, and their possible abnormal behavior is analyzed. A data association network is constructed through cluster analysis and statistical methods to identify data clusters of the same location or type. Differences in data features are analyzed, and based on the location of data influence relationships, the similarities and differences of different data results or application scenarios are distinguished and compared. The Similarity and Difference Analysis Module constructs a weighted directed graph, uses a graph traversal algorithm to find the anomaly propagation path, and calculates network indicators based on complex network analysis methods. Network indicators include node degree, centrality, and influence. A propagation model is established to simulate the process of anomalies propagating from one node to another.

[0107] Anomaly Propagation Path Inference Module: This module uses path inference technology and complex network analysis methods to identify and locate the root causes of data anomalies by analyzing the propagation paths of anomalies in the data stream. It uses graph neural networks to extract features from data nodes, capturing the complex relationships between nodes. It then uses set thresholds to judge features and flag anomalies. It also uses semi-supervised learning technology for model training to improve its ability to detect unknown anomalies.

[0108] Anomaly tracing and location model construction module: This module builds a data anomaly tracing and location model, uses graph neural network methods for feature learning, and uses knowledge graphs to identify the entire chain of abnormal data generation. This module then identifies and labels anomalies in the original data and performs semi-supervised anomaly location assessment.

[0109] Result optimization and hyperparameter adjustment module: introduces multiple conditional factors that affect business data information into the model as hyperparameters, adjusts the model's hyperparameters, and uses a feedback mechanism to dynamically adjust the model's output results; the result optimization and hyperparameter adjustment module defines multiple hyperparameters of data that affect the business through a construction module, and combines them with the input features of the graph neural network. It uses a cross-validation method to optimize and select the hyperparameters, and feeds the model's output results back to the model through a feedback mechanism to adjust and optimize the model's parameters and hyperparameters.

[0110] The data preprocessing and feature extraction module effectively addresses noise, duplication, and scale discrepancies in the raw data by cleaning, deduplicating, and standardizing power marketing data. By leveraging the statistical, computational, and analytical relationship systems within the knowledge graph, it can deeply explore potential relationships within the data.

[0111] Statistical analysis and dimensionality reduction methods are used to screen out key features, ensuring that the model can focus on the most valuable information during processing, improving computational efficiency and anomaly detection accuracy. This ensures efficient data processing while capturing potential abnormal patterns.

[0112] By using the K-means clustering algorithm to partition the data, the power marketing data is divided into K clusters. This division helps identify potential relationships between data while reducing the interference of noisy data.

[0113] By using cosine similarity to analyze the differences in features between clusters, similar data can be grouped together, while different data can be identified immediately. The similarity matrix helps identify clusters with significant differences and further analyze possible abnormal behavior.

[0114] By constructing a weighted directed graph to represent the influence relationships between data and using a graph traversal algorithm to find the anomaly propagation path, this method can accurately simulate the propagation of anomalies in data streams, identify the root node of the anomaly, and track the spread of anomalies in the network.

[0115] By calculating network indicators (such as node degree, centrality, influence, etc.), this system can identify key nodes and key paths of data flows, help analyze the trend and scope of anomaly propagation, and improve the accuracy of anomaly source location.

[0116] By learning the features of data nodes through graph neural networks (GNNs), we can capture the complex relationships between nodes. Especially in the multi-dimensional and multi-source environment of power marketing data, GNNs can effectively handle non-Euclidean relationships between data and improve the accuracy of anomaly detection.

[0117] During the anomaly feature labeling and semi-supervised location evaluation phase, semi-supervised learning is used to improve the model's ability to detect unknown anomalies by combining labeled and unlabeled data. This not only labels known anomalies but also improves the model's ability to identify new or uncommon anomalies.

[0118] By leveraging knowledge graphs and incorporating relevant information from various links and data nodes into the analysis framework, the system can trace and locate anomalies throughout the entire process. This not only improves the accuracy of identifying anomaly sources, but also enables the capture of more complex anomaly patterns and interrelated abnormal behaviors.

[0119] A semi-supervised anomaly localization and evaluation method is adopted, which combines labeled and unlabeled data for training, enhancing the flexibility of anomaly detection. It is especially suitable for complex systems such as power marketing data that change dynamically.

[0120] By introducing multiple hyperparameters that affect business data information into the model and combining them with the input features of the graph neural network, the model's processing mode can be adjusted in a dynamic environment to adapt to different business needs and changes.

[0121] Hyperparameters are optimized and selected through cross-validation, and the model output is compared with actual results through a feedback mechanism to dynamically adjust model parameters and hyperparameters. This can continuously improve the adaptability and accuracy of the model, reduce human intervention, and optimize system performance.

[0122] The feedback mechanism compares model outputs with actual labels, and uses feedback functions to adjust and optimize the model in real time. Combined with hyperparameter adjustment, this continuously optimizes model accuracy, ensuring the system can respond to rapidly changing business needs in the power marketing data environment.

[0123] The third technical solution of the present invention is: a computer device, which includes one or more processors and one or more memories, and the one or more memories store at least one program code. When the program code is executed by the one or more processors, it implements the aforementioned knowledge-driven data anomaly tracing and positioning method.

[0124] The fourth technical solution of the present invention is: a storage medium, in which at least one program code is stored, characterized in that when the program code is executed by a processor, the steps of the aforementioned knowledge-driven data anomaly tracing and locating method are implemented.

[0125] Beneficial effects:

[0126] 1. Improved anomaly detection accuracy: This invention utilizes a knowledge-driven data anomaly tracing and location method. By constructing a knowledge graph and graph neural network for multi-source heterogeneous data, it overcomes the limitations of traditional rule-based detection, which relies on fixed thresholds. By leveraging dynamic feature analysis and deep learning models, it significantly improves anomaly detection accuracy, reduces false positives and missed negatives, and enables the system to more flexibly adapt to the complex and ever-changing power marketing data streams, ensuring timely response to potential risks and improving the effectiveness of overall data monitoring.

[0127] 2. Optimizing traceability and location efficiency: By establishing a data anomaly traceability and location model centered on a graph neural network, this invention significantly improves the efficiency of traceability and location. Combining path reasoning and complex network analysis techniques, the traceability process after an anomaly occurs no longer relies on extensive manual intervention, enabling rapid identification and location of the root cause of the anomaly. This efficient traceability analysis method not only reduces time and labor costs, but also improves the stability of power marketing data and business continuity, thereby optimizing the response speed of business decisions.

[0128] 3. Enhanced Knowledge Association Analysis: By constructing a knowledge graph, this invention enables deep association analysis across multi-source data, overcoming the limitations of traditional analysis methods. Through a global knowledge-driven mechanism, the system captures implicit relationships and underlying knowledge across data. This enhanced knowledge association analysis capability not only enables effective linkage analysis of anomalies and their underlying causes, but also improves the system's ability to recognize complex anomaly patterns, thereby optimizing the adaptability and flexibility of data processing. BRIEF DESCRIPTION OF THE DRAWINGS

[0129] Figure 1 It is a flow chart of the present invention. DETAILED DESCRIPTION

[0130] The technical solution of the present invention is further described in detail below with reference to the accompanying drawings.

[0131] Step 1: Data Preprocessing and Feature Extraction. To address the multi-source heterogeneity of power marketing data, data cleaning, deduplication, and standardization are performed to ensure data consistency and integrity. Furthermore, the existing power marketing knowledge graph, which integrates multi-source heterogeneous data, is used to extract the characteristics of electricity consumption data. The statistical, computational, and analytical relationships within the knowledge graph are leveraged to enhance data feature identification. Key features are then selected through statistical analysis and dimensionality reduction methods to ensure efficient and effective model processing.

[0132] Assume that the electricity consumption data matrix is , then the feature extraction process generates features represented by the neural network model as H, .in, is the original electricity consumption data, is the normalized electricity consumption data matrix, W is the weight matrix, and f(·) is the activation function.

[0133] Step 2: Similarities and differences analysis. Combining power consumption feature analysis, statistics, and calculation results, we construct a data association network and cluster data of similar locations or types. By analyzing data impact relationships, we distinguish and compare the similarities and differences between different data results or application scenarios, identify differences in data features, and then analyze similarities and differences. The specific contents are as follows:

[0134] Step 2.1: Clustering to build a data association network. When performing similarity and difference analysis, the extracted feature data must first be clustered to identify data clusters of the same location or type. The K-means clustering algorithm is used to divide the data into K clusters, maximizing the similarity of data within the clusters and maximizing the difference between clusters. First, randomly select K initial cluster centers. ; Secondly, for each data point , calculate its distance to each cluster center , assign it to the cluster with the nearest distance; then, update the cluster center and calculate the new cluster center as ,in for all data points in cluster j; finally, repeat the above steps until the cluster center no longer changes or the change is less than the set threshold.

[0135] Through the above steps, we finally get K clusters Next, we build a data association network based on the clustering results. , where nodes V represent clusters and edges E represent the relationships between clusters. The weight of the edges can be calculated based on the distance between cluster centers. To define, specifically: ,in To prevent the denominator from being a constant of 0.

[0136] Step 2.2: Perform data similarity and difference analysis. After clustering is completed, it is necessary to compare the similarities and differences of different data results or application scenarios based on the data influence relationship position. By constructing the data association network G, the relationship and feature differences between clusters can be analyzed. The specific method is as follows: First, perform feature vector comparison. For each cluster , calculate its eigenvector The mean and variance of , ; Secondly, similarity calculation is performed, and cosine similarity is used to compare the feature differences between different clusters, which is defined as ,in and Cluster and Finally, we conduct data similarity and difference analysis. Based on the similarity matrix S, we can identify clusters with obvious feature differences. For example, we can set a threshold ,if , it is considered a cluster and clustering There are significant differences in features. At this point, we can further analyze the differences between the two clusters in application scenarios and identify possible abnormal behaviors.

[0137] Step 3: Anomaly Propagation Path Reasoning. The goal of anomaly propagation path reasoning is to identify and locate the root cause of data anomalies. To this end, we will use path reasoning technology and complex network analysis methods to analyze the propagation paths of anomalies in the data flow and reveal their impact relationships:

[0138] Step 3.1: Study the path reasoning technology for the propagation of data anomalies. First, build a weighted directed graph. , where the node V represents the data point, the edge E represents the influence relationship between the data, and the edge weight Calculated from the result of the previous step 2. Secondly, use the graph traversal algorithm to find the abnormal propagation path. Given an abnormal point , we can calculate from To other nodes The shortest path is as follows:

[0139]

[0140] in, is the weight of each edge in the path.

[0141] By traversing all nodes, we can identify the outliers Most relevant nodes , thereby determining the propagation path of the anomaly.

[0142] Step 3.2: Complex network analysis. After completing the path reasoning, further analyze the propagation characteristics of the outliers, and use the complex network analysis method to understand the interaction between the nodes in the data flow and the mechanism of anomaly propagation. First, calculate the network indicators, such as the degree, centrality and influence of the nodes, to help identify key nodes. Defined as a node The number of connected edges, the centrality is calculated as follows: ,in, is a node and The distance between nodes; secondly, based on the propagation path, a propagation model is established. It is assumed that the way data anomalies propagate between nodes is the SIR (susceptible-infected-recovered) model. The formula is as follows: ,in, is the number of susceptible nodes, is the number of infected nodes, is the number of recovery nodes, is the transmission rate, is the recovery rate. With the help of this model, we can simulate the process of anomaly propagation from one node to another.

[0143] Step 4: Constructing anomaly tracing and location model. In the process of data anomaly tracing and location, we will comprehensively utilize graph neural network methods to build a preliminary data anomaly tracing and location model. We will also use knowledge graphs to identify the entire link of abnormal data generation, realize the judgment and labeling of raw data anomalies, and thus conduct semi-supervised anomaly location assessment. The specific contents are as follows:

[0144] Step 4.1: Graph Neural Network Feature Learning. Use graph neural network to extract features from data nodes and capture the complex relationships between nodes. First, aggregate neighbor information. ,in is a node In the The layer representation, Representation node Neighbors, is a trainable weight matrix, is the activation function, is the aggregation function; secondly, through multi-layer aggregation, the feature representation of each node is obtained , where L represents the number of layers, Represents the knowledge graph; Finally, the extracted feature representation is used as the input for subsequent anomaly detection. In the last layer, the node features is passed to a classifier for anomaly detection: ,in Represents the predicted anomaly label.

[0145] Step 4.2: Abnormal Feature Labeling and Semi-supervised Localization Evaluation. When anomalies are detected in the raw data, we identify the data issues and label the abnormal features of the raw data. First, we use a set threshold to judge the features and mark the anomalies. Second, we use semi-supervised learning techniques to train the model on both labeled and unlabeled data. This uses data with known anomalies to guide model learning, thereby improving the ability to detect unknown anomalies.

[0146] Step 5: Optimize results and adjust hyperparameters. Introduce multiple factors that affect business data (such as policy changes, abnormal events, and weather changes) into the model as hyperparameters to enhance its adaptability and accuracy.

[0147] Step 5.1: Define multiple hyperparameters that affect the business data, denoted as , where each hyperparameter Representing a specific business condition (such as policy changes, weather, etc.), these hyperparameters are combined with the input features H of the graph neural network:

[0148]

[0149] in, The eigenvector representing the external conditions, Represents a join operation on a feature.

[0150] Step 5.2: To achieve better model performance, further optimize and select hyperparameters using methods such as cross-validation: Divide the data into k subsets, use each subset as the validation set in turn, and use the remaining k-1 subsets as the training set. In each round, adjust the hyperparameter H and train the model, recording the performance indicators of the validation set. For example, select the performance indicator P to evaluate the model effect:

[0151]

[0152] in, It's a true positive. It is a false positive. Based on this, grid search, Bayesian optimization and other methods are used to find the best hyperparameter combination. : .

[0153] Step 5.3: Feed the model’s output back to the model through the feedback mechanism to adjust and optimize the model’s parameters and hyperparameters. The feedback model is as follows:

[0154]

[0155] in is the predicted output of the model, is the actual label, is the feedback function used to calculate the difference in model performance.

[0156] Further parameter adjustment is performed to update the model parameters W and hyperparameters H based on the feedback information:

[0157]

[0158]

[0159] in, is the learning rate, is the gradient of the loss function, is the feedback influence coefficient.

[0160] Through the above steps, the data anomaly tracing and positioning model can be effectively optimized to better adapt to the changing power marketing environment and complex data characteristics, thereby improving overall business efficiency and decision-making support capabilities.

[0161] To validate the technical solution of this invention, experimental simulation tests of the data anomaly tracing and location method were conducted. The simulation dataset included power marketing data, business scenario simulation data, and historical anomaly data. The simulation used over 500,000 data records. To ensure representative simulation results, the data was divided into a training set and a test set, with the training set comprising 80% of the data and the test set comprising 20%.

[0162] Table 1 Application effect of knowledge-driven data anomaly tracing and location method

[0163]

[0164] Through multiple rounds of simulation comparison, the superiority of the present invention in anomaly detection and source tracing and positioning was verified. Test results in different scenarios show that: in dynamically changing power marketing data, the false alarm rate and missed alarm rate of the present invention method are significantly lower than those of traditional methods, and it can accurately capture complex anomaly patterns; the source tracing and positioning time is significantly shortened. The present invention method greatly improves the speed and accuracy of anomaly source positioning through optimized data association path reasoning and graph neural network feature learning; processing accuracy and cost-effectiveness are optimized. Compared with traditional manual intervention and rule setting methods, the present invention not only improves the accuracy of anomaly processing, but also reduces the workload of business personnel through automated anomaly tracing, reducing operating costs.

[0165] Example 2

[0166] Accordingly, this embodiment provides a knowledge-driven data anomaly tracing and locating system, including a data preprocessing and feature extraction module, a similarity and difference analysis module, an anomaly propagation path reasoning module, an anomaly tracing and locating model construction module, and a result optimization and hyperparameter adjustment module.

[0167] The data preprocessing and feature extraction module primarily applies data cleaning, deduplication, and standardization techniques to ensure the consistency and integrity of power marketing data. It also uses statistical analysis methods to calculate the mean and variance of key features, and employs dimensionality reduction techniques such as principal component analysis to reduce data dimensionality. This ultimately generates a high-quality feature matrix for electricity consumption data, providing a reliable foundation for subsequent analysis. This module implements the functionality of Step 1 in Example 1 and will not be further elaborated here.

[0168] The Similarity and Difference Analysis module uses a clustering algorithm to identify clusters of data in the same location or of the same type. By clustering the extracted feature data, a data association network is formed. Cosine similarity is used to calculate feature differences between clusters, helping to analyze data influence relationships. This module ultimately outputs the clustering results and the data association network, clarifying the similarities and differences between different data results or application scenarios. This module is used to implement the functionality of Step 2 in Example 1 and will not be further described here.

[0169] The anomaly propagation path inference module utilizes path inference technology and complex network analysis methods to construct a weighted directed graph to represent the influence relationships between data points. By traversing the nodes in the graph, it identifies the propagation paths of anomalies and calculates the shortest paths between nodes. Furthermore, it applies the SIR model to analyze the propagation characteristics of anomalies, thereby revealing the anomaly propagation mechanism in the data stream. This module implements the functionality of step 3 in Example 1 and will not be further described here.

[0170] The anomaly tracing and location model construction module integrates graph neural networks and semi-supervised learning techniques to establish a preliminary anomaly tracing and location model by extracting features from data nodes. This module utilizes a knowledge graph to identify the entire chain of abnormal data generation, enabling the identification and labeling of anomalies in raw data. This improves the ability to detect unknown anomalies and enhances the model's practicality. This module is used to implement the functionality of Step 4 in Example 1 and will not be further described here.

[0171] The Result Optimization and Hyperparameter Adjustment module adjusts the model's hyperparameters through cross-validation and optimization methods. Taking into account multiple factors influencing the business, the module introduces these hyperparameters into the graph neural network, thereby improving the model's accuracy and adaptability. Furthermore, a feedback mechanism is used to dynamically adjust the model's output results, ensuring that the model's parameters and hyperparameters are continuously optimized to adapt to the complex power marketing environment. This module implements the functionality of Step 5 in Example 1 and will not be further described here.

[0172] Example 3

[0173] This embodiment provides an electronic device, comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the knowledge-driven data anomaly tracing and locating method as described in any embodiment of the present invention is implemented.

[0174] Example 4

[0175] This embodiment provides a computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the knowledge-driven data anomaly tracing and locating method as described in any embodiment of the present invention is implemented.

[0176] Those skilled in the art will appreciate that the various units and algorithm steps described in the embodiments disclosed in the present invention can be implemented by a combination of electronic hardware, computer software, and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0177] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.

[0178] In the several embodiments provided in this application, if any function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, or the part that contributes to the prior art, or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the method described in each embodiment of this application. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (Read-Only Memory; hereinafter referred to as: ROM), random access memory (Random Access Memory; hereinafter referred to as: RAM), magnetic disk or optical disk, and other media that can store program code.

[0179] The data anomaly tracing and locating method, system, device and storage medium shown above are specific embodiments of the present invention, which have reflected the essential characteristics and progress of the present invention. According to actual usage needs and under the guidance of the present invention, equivalent modifications in shape, structure, etc. can be made to them, which are all within the scope of protection of this solution.

Claims

1. A knowledge-driven data anomaly tracing and location method, characterized by: The following steps are involved: 1) Preprocessing and feature extraction of power marketing data, including data cleaning, deduplication, and standardization, as well as extracting features of electricity consumption data using knowledge graphs and screening out key features; 2) Use cluster analysis and statistical methods to construct a data association network, analyze the differences in data characteristics, distinguish the similarities and differences between different data results or application scenarios, and compare them; 3) Using path reasoning technology and complex network analysis methods, we can identify and locate the root causes of data anomalies by analyzing the propagation paths of anomalies in the data flow; 4) Build a data anomaly traceability and location model, use graph neural network methods for feature learning, and use knowledge graphs to identify the entire chain of abnormal data generation, realize the judgment and labeling of raw data anomalies, and perform semi-supervised anomaly location assessment; 5) Optimize results and adjust hyperparameters. Introduce multiple factors that affect business data information into the model as hyperparameters, adjust the model's hyperparameters, and use feedback mechanisms to dynamically adjust the model's output results. In step 2), the specific steps of constructing the data association network include: 2.1) Use the K-means clustering algorithm to divide the data into K clusters and construct a data association network G = (V, E) based on the clustering results, where nodes V represent clusters and edges E represent the relationships between clusters. 2.2) Compare feature vectors and calculate similarity. Use cosine similarity to compare feature differences between clusters. Identify clusters with significant feature differences based on the similarity matrix S and analyze their possible abnormal behavior. In step 3), the specific steps of identifying the abnormal propagation path include: 3.1) Construct a weighted directed graph G = (V, E), where nodes V represent data points and edges E represent the influence relationships between data points. Use a graph traversal algorithm to find the anomaly propagation path. 3.2) Based on complex network analysis methods, network indicators are calculated, including node degree, centrality, and influence. A propagation model is established to simulate the process of anomalies spreading from one node to another.

2. The knowledge-driven data anomaly tracing and location method according to claim 1 is characterized by: In step 1), the statistical, computational, and analytical relationship systems in the knowledge graph are used to enhance the identification of data features, and key features are screened out through statistical analysis and dimensionality reduction methods to ensure the efficiency and effectiveness of model processing.

3. The knowledge-driven data anomaly tracing and location method according to claim 1 is characterized by: In step 3.1), first construct a weighted directed graph , where the node V represents the data point and the edge E represents the influence relationship between the data; secondly, the graph traversal algorithm is used to find the abnormal propagation path. Given an abnormal point , calculated from To other nodes The shortest path is: in, is the weight of each edge in the path; By traversing all nodes, we can identify the abnormal points Most relevant nodes , thereby determining the propagation path of the anomaly; In step 3.2), First, network index calculations are performed to calculate indicators including node degree, centrality, and influence to help identify key nodes. Defined as a node The number of connected edges, the centrality calculation formula is: ,in, is a node and the distance between them; Secondly, based on the propagation path, a propagation model is established. Assuming that the way data anomalies propagate between nodes is the SIR model, the formula is: ,in, is the number of susceptible nodes, is the number of infected nodes, is the number of recovery nodes, is the transmission rate, is the recovery rate, which simulates the process of anomaly propagation from one node to another with the help of the propagation model.

4. The knowledge-driven data anomaly tracing and location method according to claim 3 is characterized by: The specific steps of step 4) include: 4.1) Graph Neural Network Feature Learning: Using graph neural networks to extract features from data nodes and capture the complex relationships between nodes: First, aggregate neighbor information: , in is a node In the The layer representation, Representation node Neighbors, is a trainable weight matrix, is the activation function, is an aggregate function; Secondly, through multi-layer aggregation, the feature representation of each node is obtained , where L represents the number of layers, Represents a knowledge graph; Finally, the extracted feature representation is used as the input for subsequent anomaly detection. In the last layer, the node features is passed to a classifier for anomaly detection: , in represents the predicted anomaly label; 4.2) Abnormal Feature Labeling and Semi-Supervised Localization Evaluation: First, the features are judged using the set threshold and abnormal situations are marked; Secondly, semi-supervised learning technology is used to train the model on labeled and unlabeled data, and data with known anomalies are used to guide model learning to improve the ability to detect unknown anomalies.

5. The knowledge-driven data anomaly tracing and location method according to claim 4 is characterized by: The specific steps of step 5) include: 5.1) Define multiple hyperparameters that affect the business data, denoted as , where each hyperparameter Represents a specific business condition and combines these hyperparameters with the input features H of the graph neural network: in, The eigenvector representing the external conditions, Represents the join operation of features; 5.2) Use cross-validation to optimize and select hyperparameters: Divide the data into m subsets, use each subset as the validation set in turn, and use the remaining m-1 subsets as the training set. In each round, adjust the hyperparameter H and train the model. Record the performance indicators of the validation set and select the performance indicator P to evaluate the model effect: in, It's a true positive. It is a false positive; grid search and Bayesian optimization are used to find the best hyperparameter combination : ; 5.3) Feedback the model’s output back to the model through a feedback mechanism to adjust and optimize the model’s parameters and hyperparameters. The feedback model is: in is the predicted output of the model, is the actual label, is the feedback function used to calculate the performance difference of the model; Further parameter adjustment is performed to update the model parameters W and hyperparameters H based on the feedback information: in, is the learning rate, is the gradient of the loss function, is the feedback influence coefficient.

6. A knowledge-driven data anomaly tracing and positioning system, characterized by: A knowledge-driven data anomaly tracing and locating method according to any one of claims 1 to 5 is applied, wherein the system comprises: Data preprocessing and feature extraction module: This module preprocesses power marketing data, including data cleaning, deduplication, and standardization. It leverages the statistical, computational, and analytical relationship systems within the knowledge graph to enhance data feature identification. Key features are screened through statistical analysis and dimensionality reduction methods to ensure efficient and effective model processing. Similarity and Difference Analysis Module: Uses the K-means clustering algorithm to divide data into K clusters, and constructs a data association network based on the clustering results. Through feature vector comparison and similarity calculation, cosine similarity is used to compare the feature differences between different clusters. Clusters with obvious feature differences are identified based on the similarity matrix, and their possible abnormal behavior is analyzed. A data association network is constructed through cluster analysis and statistical methods to identify data clusters of the same location or type. Differences in data features are analyzed, and based on the location of data influence relationships, the similarities and differences of different data results or application scenarios are distinguished and compared. The Similarity and Difference Analysis Module constructs a weighted directed graph, uses a graph traversal algorithm to find the anomaly propagation path, and calculates network indicators based on complex network analysis methods. Network indicators include node degree, centrality, and influence. A propagation model is established to simulate the process of anomalies propagating from one node to another. Anomaly Propagation Path Inference Module: This module uses path inference technology and complex network analysis methods to identify and locate the root causes of data anomalies by analyzing the propagation paths of anomalies in the data stream. It uses graph neural networks to extract features from data nodes, capturing the complex relationships between nodes. It then uses set thresholds to judge features and flag anomalies. It also uses semi-supervised learning technology for model training to improve its ability to detect unknown anomalies. Anomaly tracing and location model construction module: This module builds a data anomaly tracing and location model, uses graph neural network methods for feature learning, and uses knowledge graphs to identify the entire chain of abnormal data generation. This module then identifies and labels anomalies in the original data and performs semi-supervised anomaly location assessment. Result optimization and hyperparameter adjustment module: introduces multiple conditional factors that affect business data information into the model as hyperparameters, adjusts the model's hyperparameters, and uses a feedback mechanism to dynamically adjust the model's output results; the result optimization and hyperparameter adjustment module defines multiple hyperparameters of data that affect the business through a construction module, and combines them with the input features of the graph neural network. It uses a cross-validation method to optimize and select the hyperparameters, and feeds the model's output results back to the model through a feedback mechanism to adjust and optimize the model's parameters and hyperparameters.

7. A computer device, characterized in that: The device includes one or more processors and one or more memories, and at least one program code is stored in the one or more memories. When the program code is executed by the one or more processors, a knowledge-driven data anomaly tracing and locating method as described in any one of claims 1 to 5 is implemented.

8. A storage medium storing at least one program code, characterized in that: When the program code is executed by a processor, the steps of a knowledge-driven data anomaly tracing and locating method as described in any one of claims 1 to 6 are implemented.

Citation Information

Patent Citations

  • Abnormal data diagnosis method and device based on knowledge graph, equipment and medium

    CN117668733A

  • Communication data transmission method and device, equipment and storage medium

    CN118316790A