Data anomaly detection method, device and equipment

By collecting and folding multi-source heterogeneous data, the dynamic detection threshold is generated using the graph federal learning model, which solves the data island problem in the power generation industry and achieves high-precision anomaly detection and data security protection.

CN120387112APending Publication Date: 2025-07-29济南作为科技有限公司
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510305714.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-14
Publication Date
2025-07-29

AI Technical Summary

Technical Problem

The data island problem in the existing power generation industry leads to inaccurate abnormal detection, insufficient training data, lack of real-time characteristics, and one-sided feature expression, making it difficult to support the demand for high-precision abnormal detection.

Method used

Multi-source heterogeneous data is collected and folded, and dynamic detection thresholds are generated using the graph federated learning model, and abnormal detection is performed in combination with the graph neural network algorithm and the federated learning framework.

Benefits of technology

Improve the accuracy of anomaly detection, solve the data silo problem, protect data security, and conduct model training in a distributed environment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120387112A_ABST
    Figure CN120387112A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of machine learning, and discloses a data anomaly detection method, device and equipment, and the method comprises the steps: collecting multi-source heterogeneous data, carrying out the folding operation of the multi-source heterogeneous data, obtaining the folded data, and generating a dynamic detection threshold through a graph federal learning model, the graph federated learning model is a model constructed based on a graph neural network algorithm and a federated learning framework, and performing anomaly detection on the folded data according to a dynamic detection threshold to obtain an anomaly detection result. According to the method, the collection and folding operation of the multi-source heterogeneous data is combined, the actual condition of the data is comprehensively reflected, the redundancy and complexity of the data are reduced, the introduction of the graph neural network algorithm and the federated learning framework improves the accuracy of anomaly detection, the use of the federated learning framework allows model training in a distributed environment, and the accuracy of anomaly detection is improved. The centralized storage of data is avoided, the problem of data islands is solved, and the data security is protected.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of machine learning technology, and particularly to a method, device, and equipment for data anomaly detection. Background Art

[0002] In the process of digital transformation of the power generation industry, the problem of data islands is one of the core pain points leading to inaccurate anomaly detection. The hardware (servers, real-time databases) and software (operating systems, databases) that traditional power generation control systems have long relied on will cause data to be scattered in heterogeneous systems, forming multi-source heterogeneous data islands. The communication protocols and data formats between these systems are significantly different, making it difficult to uniformly access and analyze. Due to the existence of data islands, existing anomaly detection methods generally face problems such as insufficient training data, lack of real-time performance, and one-sided feature expression while ignoring multi-modal data, making it difficult to support the requirements of high-precision anomaly detection and further amplifying the detection deviation problem caused by data islands. Summary of the Invention

[0003] The main purpose of this application is to provide a method, device, and equipment for data anomaly detection, aiming to solve the technical problem of inaccurate data anomaly detection in the existing technology.

[0004] To achieve the above object, this application proposes a method for data anomaly detection, the method includes:

[0005] Collect multi-source heterogeneous data, and perform a folding operation on the multi-source heterogeneous data to obtain folded data;

[0006] Generate a dynamic detection threshold through a graph federated learning model, where the graph federated learning model is a model constructed based on the graph neural network algorithm and the federated learning framework;

[0007] Perform anomaly detection on the folded data according to the dynamic detection threshold to obtain an anomaly detection result.

[0008] In one embodiment, before the step of generating a dynamic detection threshold through the graph federated learning model, it further includes:

[0009] Based on the distributed nodes of the graph neural network algorithm and the federated learning framework, combine the differential privacy protocol and the homomorphic encryption protocol to deploy an initial model based on the spatio-temporal graph convolutional network;

[0010] Obtain a graph federated learning model according to the initial model and a preset operator library.

[0011] In one embodiment, the step of obtaining a graph federated learning model according to the initial model and a preset operator library includes:

[0012] In the initial model, an association graph is constructed based on the cosine similarity of the folded data, and the node features of the association graph are embedded into the initial graph structure to output a low-dimensional embedding vector;

[0013] The low-dimensional embedding vector is input into each distributed node of the initial model for training to generate a parameter set;

[0014] The federated learning weights of the initial model are assigned according to the node data quality;

[0015] Based on the federated learning weights, a weighted average is performed on the parameter set to generate global model parameters, and the initial model is updated through the global model parameters to obtain an updated model;

[0016] The updated model is accelerated in inference through a preset operator library to obtain a graph federated learning model.

[0017] In one embodiment, the step of generating a dynamic detection threshold through the graph federated learning model includes:

[0018] Based on the graph federated learning model, feature weights are assigned to the folded data through a gated cross-attention mechanism to generate a weighted fusion feature;

[0019] Based on the association graph of the graph federated learning model, a convolution operation is performed on the weighted fusion feature to obtain an anomaly score result;

[0020] The global benchmark statistic is determined through a historical data set, and based on the benchmark statistic and the anomaly score result, a dynamic detection threshold is generated according to preset constraint conditions.

[0021] In one embodiment, the step of determining the global benchmark statistic through the historical data set and generating a dynamic detection threshold based on the benchmark statistic and the anomaly score result according to preset constraint conditions includes:

[0022] Determine the statistic of the historical data set, and summarize the statistic through a secure multi-party calculation algorithm to obtain the global benchmark statistic;

[0023] Optimize the global benchmark statistic through a loss function to obtain an optimized global benchmark statistic;

[0024] Perform a normalization process on the anomaly score result to obtain a normalized anomaly score;

[0025] Based on the optimized global benchmark statistic and the normalized anomaly score, a dynamic detection threshold is generated according to preset constraint conditions and the Bayesian optimization algorithm.

[0026] In one embodiment, the steps of collecting multi-source heterogeneous data and performing a folding operation on the multi-source heterogeneous data to obtain the folded data include:

[0027] Collect multi-source heterogeneous data at a corresponding collection frequency according to the adaptive collection mechanism and the data change rate;

[0028] Store the multi-source heterogeneous data in the storage nodes of the target storage system, where the target storage system is a storage system constructed based on a hyper-converged storage architecture;

[0029] Based on the embedded computing unit of the storage node, convert the unstructured data of the multi-source heterogeneous data into a feature map, and perform temporal folding on the structured data of the multi-source heterogeneous data through a statistical feature compression algorithm to obtain a target statistical block;

[0030] Merge the feature map and the target statistical block through a materialized view algorithm to generate the folded data.

[0031] In one embodiment, the steps of, based on the embedded computing unit of the storage node, converting the unstructured data of the multi-source heterogeneous data into a feature map, and performing temporal folding on the structured data of the multi-source heterogeneous data through a statistical feature compression algorithm to obtain a target statistical block include:

[0032] Based on the embedded computing unit of the storage node, convert the unstructured data of the multi-source heterogeneous data into a feature map;

[0033] Perform temporal folding on the structured data of the multi-source heterogeneous data to obtain a statistical block, and perform manifold learning dimensionality reduction processing on the statistical block through an approximate nearest neighbor compression algorithm to obtain a processed statistical block;

[0034] Align the processed statistical block according to the dynamic time warping algorithm to obtain a target statistical block.

[0035] In one embodiment, after the step of performing anomaly detection on the folded data according to the dynamic detection threshold to obtain an anomaly detection result, the following steps are further included:

[0036] Post-process the anomaly detection result to obtain a processed detection result;

[0037] Judge the processed detection result through a multi-modal graph neural network to obtain a judgment result;

[0038] When the judgment result is a high-value anomaly sample, obtain an updated model by triggering model incremental training;

[0039] Use the high-value anomaly sample as labeled data, and return the labeled data to each node through the federated learning framework to obtain a retrained model;

[0040] Record the update process of the graph federated learning model through the blockchain network to obtain a traceability log.

[0041] In addition, to achieve the above object, the present application also proposes a data anomaly detection device, which includes:

[0042] A data processing module, configured to collect multi-source heterogeneous data and perform a folding operation on the multi-source heterogeneous data to obtain folded data;

[0043] A threshold generation module, configured to generate a dynamic detection threshold through a graph federated learning model, wherein the graph federated learning model is a model constructed based on a graph neural network algorithm and a federated learning framework;

[0044] An anomaly detection module, configured to perform anomaly detection on the folded data according to the dynamic detection threshold to obtain an anomaly detection result.

[0045] In addition, to achieve the above object, the present application also proposes a data anomaly detection device, which includes: a memory, a processor, and a computer program stored on the memory and executable on the processor, and the computer program is configured to implement the steps of the data anomaly detection method as described above.

[0046] In addition, to achieve the above object, the present application also proposes a storage medium, which is a computer-readable storage medium, and a computer program is stored on the storage medium, and when the computer program is executed by a processor, it implements the steps of the data anomaly detection method as described above.

[0047] In addition, to achieve the above object, the present application also provides a computer program product, which includes a computer program, and when the computer program is executed by a processor, it implements the steps of the data anomaly detection method as described above.

[0048] The technical solution proposed by the present application collects multi-source heterogeneous data, performs a folding operation on the multi-source heterogeneous data to obtain folded data, generates a dynamic detection threshold through a graph federated learning model, wherein the graph federated learning model is a model constructed based on a graph neural network algorithm and a federated learning framework, and performs anomaly detection on the folded data according to the dynamic detection threshold to obtain an anomaly detection result. By combining the collection and folding operation of multi-source heterogeneous data, it comprehensively reflects the actual situation of the data, reduces the redundancy and complexity of the data. The introduction of the graph neural network algorithm and the federated learning framework improves the accuracy of anomaly detection, and the use of the federated learning framework allows model training in a distributed environment, avoiding centralized data storage, solving the data island problem, and helping to protect data security. Description of the Drawings

[0049] The accompanying drawings herein are incorporated into and constitute a part of this specification, showing embodiments consistent with the present application and, together with the specification, are used to explain the principles of the present application.

[0050] To more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the accompanying drawings required for use in the description of the embodiments or the prior art. Obviously, for those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0051] Figure 1 It is a schematic flowchart provided for the first embodiment of the data anomaly detection method of the present application;

[0052] Figure 2 It is a schematic flowchart provided for the second embodiment of the data anomaly detection method of the present application;

[0053] Figure 3 It is a schematic flowchart provided for the third embodiment of the data anomaly detection method of the present application;

[0054] Figure 4 It is a schematic diagram of the module structure of the data anomaly detection device in the embodiment of the present application;

[0055] Figure 5 It is a schematic diagram of the device structure of the hardware operating environment involved in the data anomaly detection method in the embodiment of the present application.

[0056] The realization of the objectives, functional features and advantages of the present application will be further described with reference to the embodiments and the accompanying drawings. Detailed Embodiments

[0057] It should be understood that the specific embodiments described herein are only used to explain the technical solutions of the present application and are not used to limit the present application.

[0058] To better understand the technical solutions of the present application, the following will be described in detail in combination with the accompanying drawings of the specification and the specific embodiments.

[0059] Due to the existence of data islands, existing anomaly detection methods generally face problems such as insufficient training data, lack of real-time performance, one-sided feature expression, and neglect of multi-modal data, making it difficult to support the requirements of high-precision anomaly detection and further amplifying the detection deviation problem caused by data islands.

[0060] Therefore, to overcome the above deficiencies, the present application provides a solution. By combining the collection and folding operations of multi-source heterogeneous data, it comprehensively reflects the actual situation of the data, reduces data redundancy and complexity. The introduction of the graph neural network algorithm and the federated learning framework improves the accuracy of anomaly detection, and the use of the federated learning framework allows model training in a distributed environment, avoiding centralized data storage, solving the data silo problem, and helping to protect data security.

[0061] It should be noted that the execution subject of each embodiment of the present application can be a computing service system with data processing, network communication, and program running functions, such as an electronic system, a data anomaly detection system, etc. that can implement the above functions. Hereinafter, taking the data anomaly detection system as an example (hereinafter referred to as "system"), the following embodiments will be described.

[0062] Based on this, an embodiment of the present application provides a data anomaly detection method, referring to Figure 1 , Figure 1 which is a schematic flowchart of the first embodiment of the data anomaly detection method of the present application.

[0063] In this embodiment, the data anomaly detection method includes steps S10 to S30:

[0064] Step S10, collect multi-source heterogeneous data, and perform a folding operation on the multi-source heterogeneous data to obtain folded data.

[0065] It should be noted that the data anomaly detection method of the present application has high modularity and scalability, and can be widely applied to multi-industry scenarios such as energy, finance, healthcare, and smart cities. In the power generation field, the present application integrates multi-source heterogeneous system data through a unified data platform, solves the data silo problem, and combines federated learning and graph convolutional networks to achieve device data anomaly detection; in other industrial fields (such as intelligent manufacturing, chemical industry), a real-time data anomaly detection system can be quickly deployed by replacing the data collection protocol and adapting to the industry database to ensure production safety and efficiency; in the financial and healthcare industries, it supports the correlation analysis of multi-modal data (transaction records, images, electronic medical records), and can be used in scenarios such as anti-fraud detection and disease prediction.

[0066] In addition, it should be noted that multi-source heterogeneous data refers to structured (such as numerical tables), semi-structured (such as JSON logs), and unstructured data (such as images, videos) from different sources. Multi-source heterogeneous data (such as device sensor signals, production logs, video surveillance, etc.) usually has the characteristics of large format differences, discrete spatio-temporal distributions, and high dimensions. Direct fusion will lead to high computational complexity and difficult model training. The folding operation maps multi-source heterogeneous data into low-dimensional and compact statistical blocks and feature maps by unifying the data representation form and extracting key features, solves the data heterogeneity problem, and at the same time retains the spatio-temporal correlation of the data, providing high-quality input for subsequent anomaly detection.

[0067] Step S20: Generate a dynamic detection threshold through the graph federated learning model, where the graph federated learning model is a model constructed based on the graph neural network algorithm and the federated learning framework.

[0068] After obtaining the folded data, the system will then use the graph federated learning model to generate a dynamic detection threshold. The Federated Graph Learning (FGL) model is an advanced model that combines the graph neural network algorithm and the federated learning framework. The graph neural network algorithm is good at processing graph-structured data and can capture the complex relationships between data; while the federated learning framework allows model training in a distributed environment without transmitting data to a central node, thus protecting data privacy and security. Through the graph federated learning model, the system can learn the internal laws and characteristics of the data, and then generate a dynamic detection threshold based on these laws and characteristics.

[0069] Step S30: Perform anomaly detection on the folded data according to the dynamic detection threshold to obtain an anomaly detection result.

[0070] In this step, the system will compare the size relationship between each data point and the threshold one by one. If a certain data point exceeds the corresponding threshold, it is considered an anomaly point. Through this process, the system can identify anomalies or abnormal situations in the data, thus providing valuable insights and warnings. The anomaly detection result can be used to trigger an alarm, trigger further analysis or investigation, and for the continuous optimization and improvement of the model.

[0071] As an implementation manner, after the above step S30 in this embodiment, it further includes: post-processing the anomaly detection result to obtain a processed detection result; judging the processed detection result through a multi-modal graph neural network to obtain a judgment result; when the judgment result is a high-value anomaly sample, triggering model incremental training to obtain an updated model; using the high-value anomaly sample as labeled data, and flowing back the labeled data to each node through the federated learning framework to obtain a retrained model; recording the update process of the graph federated learning model through a blockchain network to obtain a traceability log.

[0072] It can be understood that after completing the anomaly detection of the folded data according to the dynamic detection threshold and obtaining the anomaly detection result, post-processing the anomaly detection result to obtain a processed detection result, so as to further screen and refine the initially obtained anomaly detection result. Among them, the post-processing can include operations such as removing false alarms, merging duplicate anomalies, classifying or scoring anomalies, etc., to improve the accuracy and practicality of the detection result.

[0073] After processing the anomaly detection result, the system uses a multi-modal graph neural network to further judge the processed data. The multi-modal graph neural network can integrate various types of data and information, including text, images, numerical values, etc., so as to more comprehensively understand the internal laws and characteristics of the data. Through this step, a more in-depth evaluation of the processed detection result is carried out, and a judgment result is generated to distinguish high-value anomaly samples from low-value anomaly samples. When the system determines that an anomaly sample is a high-value anomaly sample, it means that this sample has a higher degree of anomaly or importance. At this time, the system will trigger the model incremental training process, adding the high-value anomaly sample as new training data to the model to update and optimize the performance of the model. Incremental training can quickly adapt to new data without retraining the entire model, thereby improving the accuracy and generalization ability of the model. In addition to triggering model incremental training, the system will also use the high-value anomaly sample as labeled data and flow it back to each node through the federated learning framework. This step aims to utilize the data resources in the distributed environment to more comprehensively train and optimize the model. Through the federated learning framework, each node can share and utilize the labeled data for model training without exposing the original data, thereby generating a more robust and accurate retrained model.

[0074] Finally, the system records the update process of the graph federated learning model through a blockchain network to generate a traceability log. Blockchain technology has characteristics such as decentralization, immutability, and transparency, which can ensure the reliability and security of the model update process. By recording the traceability log, the system can conveniently track the update history of the model, verify the legality and effectiveness of the model, and provide a basis for subsequent model auditing and supervision.

[0075] For ease of understanding, an example scenario of power plant equipment anomaly detection and model iteration is used for illustration:

[0076] For example, a coal-fired power plant has deployed a multi-sensor network (monitoring boiler temperature, turbine vibration, transformer oil temperature, etc.) and a video surveillance system (covering key equipment areas). The system needs to detect equipment anomalies in real time (such as turbine bearing wear, transformer faults), and dynamically optimize the model through federated learning while ensuring data privacy and operation traceability. The specific steps include:

[0077] 1. Post-processing of anomaly detection results

[0078] Original detection results: Sensor A detected that the vibration amplitude of the turbine bearing exceeded the threshold (initially judged as an anomaly); the video surveillance system captured a blurred image of a local area of the turbine (possibly due to equipment displacement caused by vibration).

[0079] Post-processing operations: Convert the vibration signal into a time-series statistical block (such as the spectral features after short-time Fourier transform), and retain the periodic fluctuation pattern (structured data processing); extract the displacement feature map of the turbine blade edge in the video frame through a lightweight CNN (unstructured data processing); use the materialized view algorithm to merge the statistical block and the feature map into a multi-modal data packet, and label the equipment ID (turbine #3) and timestamp (2025-03-10 14:30:00) (data fusion).

[0080] 2. Judgment by multi-modal graph neural network

[0081] Input: Multi-modal data packet (vibration spectrum + blade displacement feature map).

[0082] Processing: Construct a spatio-temporal graph based on the equipment topology relationship (there is a process dependency between turbine #3 and adjacent equipment #4 (generator) and #5 (cooling tower)) (association graph construction); apply a graph convolutional layer to the spectral features to capture the collaborative anomaly patterns between the turbine and other equipment (such as vibration being transmitted to the generator through the coupling) (multi-modal reasoning); apply a spatio-temporal attention mechanism to the displacement feature map to focus on the matching relationship between the blade deformation area and the vibration amplitude.

[0083] Output: Comprehensive score of 0.92 (judged as a "high-value anomaly sample" with a confidence of 92%, triggering incremental training).

[0084] 3. Trigger incremental training and model update

[0085] Incremental training trigger: The system automatically freezes the current global model parameters, and only fine-tunes the parameters related to turbine #3 to reduce the training overhead; use the labeled data (vibration spectrum + displacement feature map) to train the newly added local model (such as adding 1-2 convolutional layers).

[0086] Annotated data feedback: After the artificial engineer confirms that the anomaly is the wear of the steam turbine bearing, the annotated multi-modal data packet is distributed to each node through the federated learning framework (the local models of Power Plants #1 - #5 are all updated).

[0087] 4. Blockchain Traceability Log Record

[0088] Log content: Time, triggering condition (high-value anomaly sample), participating nodes (Power Plants #1 - #5), model version number (training event); source device of the annotated data (Steam Turbine #3), processing operations (spectrum conversion, feature extraction) (data traceability); engineer review record and signature (operator information).

[0089] Blockchain storage: The log is stored in the private chain of the power plant and the nodes of the consortium chain in the form of an encrypted hash chain to ensure immutability and audit traceability.

[0090] In this embodiment, by combining the collection and folding operations of multi-source heterogeneous data, the actual situation of the data is comprehensively reflected, reducing data redundancy and complexity. The introduction of the graph neural network algorithm and the federated learning framework improves the accuracy of anomaly detection, and the use of the federated learning framework allows model training in a distributed environment, avoiding centralized data storage, solving the data island problem, and helping to protect data security.

[0091] Based on the first embodiment of this application, in the second embodiment of this application, the same or similar content as in the above-mentioned Embodiment 1 can be referred to the above introduction and will not be repeated hereinafter. On this basis, please refer to Figure 2 , before the step S10, the steps S01 - S02 may further be included:

[0092] Step S01, based on the distributed nodes of the graph neural network algorithm and the federated learning framework, combining the differential privacy protocol and the homomorphic encryption protocol, deploy an initial model based on the spatio-temporal graph convolutional network.

[0093] It should be understood that an initial model based on the spatio-temporal graph convolutional network is constructed based on the distributed nodes of the graph neural network algorithm and the federated learning framework. The spatio-temporal graph convolutional network processes spatio-temporal related data (such as time series and spatial location information) and realizes distributed collaboration through the federated learning framework. Each node retains the original data and only shares model parameters or encrypted intermediate results. The differential privacy protocol introduces noise during parameter update to prevent reverse inference of individual data, while homomorphic encryption encrypts the calculation process to protect data privacy during transmission and aggregation.

[0094] Step S02, obtain the graph federated learning model according to the initial model and the preset operator library.

[0095] It is understandable that the structure of the initial model is dynamically adjusted through a preset operator library (including modules such as graph convolution and attention mechanism) to adapt to specific task requirements. After local training of each node, the parameters processed by differential privacy and homomorphic encryption are uploaded to the federated learning center. The center aggregates the parameters of each node (such as weighted average or gradient optimization) to generate a globally unified graph federated learning model that takes privacy protection into account. This process ensures that the original data always remains local, and the model achieves security and robustness through encryption and noise processing.

[0096] As an implementation manner, step S02 in this embodiment may include: in the initial model, constructing an association graph according to the cosine similarity of the folded data, and embedding the node features of the association graph into the initial graph structure to output a low-dimensional embedding vector; inputting the low-dimensional embedding vector into each distributed node of the initial model for training to generate a parameter set; allocating the federated learning weights of the initial model according to the node data quality; performing weighted average on the parameter set based on the federated learning weights to generate global model parameters, and updating the initial model through the global model parameters to obtain an updated model; performing inference acceleration on the updated model through a preset operator library to obtain a graph federated learning model.

[0097] Specifically, an association graph between data is constructed based on the cosine similarity of the folded data (such as nodes representing data samples and edges representing similarity relationships), and the node features of the association graph (such as original data attributes or transformed low-dimensional representations) are embedded into the initial spatio-temporal graph convolution network structure to output a low-dimensional embedding vector as the model input. The cosine similarity is used to quantify the correlation between data, enhancing the model's learning ability for the association relationships of multi-source heterogeneous data. The low-dimensional embedding vector is distributed to each distributed node (such as a terminal device or an edge server) to perform the training task of the initial model locally, generating a parameter set containing gradients or weights. During the training process, the federated learning framework ensures that only parameters are shared between nodes instead of the original data, and the security and privacy of parameter transmission are guaranteed through differential privacy and homomorphic encryption.

[0098] The federated learning weights are dynamically allocated according to the node data quality (such as data volume, annotation accuracy, noise level, etc.), and the parameter sets generated by the nodes are weighted averaged to generate global model parameters. The global parameters are sent to each node through the federated learning framework to complete the update of the initial model, forming an updated model that integrates the knowledge of distributed nodes. Then, a preset operator library (such as lightweight convolution kernels, pruning strategies, quantization techniques, etc.) is used to perform inference acceleration optimization on the updated model, reducing the model's computational complexity and memory occupancy. Through the general algorithm modules provided by the operator library, the model structure is flexibly adjusted to adapt to real-time detection requirements, and finally an optimized graph federated learning model is output.

[0099] In this embodiment, by combining the graph neural network algorithm and the federated learning framework, and using the differential privacy protocol and the homomorphic encryption protocol, efficient collaborative training of distributed nodes is achieved, ensuring data privacy and security. Moreover, efficient processing of spatio-temporal data in a distributed environment is realized. The cosine similarity is used to construct an association graph, providing a more accurate input for the initial model, enhancing the model's ability to capture complex data relationships, and dynamically allocating federated learning weights according to the quality of node data, improving the efficiency and accuracy of model training, especially in a distributed environment with uneven data quality.

[0100] As an implementation manner, the step S20 may include the steps of:

[0101] Based on the graph federated learning model, allocate feature weights to the folded data through a gated cross-attention mechanism to generate weighted fusion features;

[0102] Perform a convolution operation on the weighted fusion features based on the association graph of the graph federated learning model to obtain an anomaly score result;

[0103] Determine the global benchmark statistic through the historical data set, and generate a dynamic detection threshold based on the benchmark statistic and the anomaly score result according to the preset constraint conditions.

[0104] It can be understood that, based on the graph federated learning model, the weight distribution of each feature in the folded data is dynamically calculated through a gated cross-attention mechanism. This mechanism adaptively adjusts the weights according to the relevance of the data to the current task (such as the sensitivity of the anomaly pattern), and performs weighted fusion of the local features of multi-source heterogeneous data with the global context to generate a comprehensive feature representation (i.e., weighted fusion features). The weight distribution process depends on the structure of the association graph, strengthening the semantic association between cross-node data (such as spatio-temporal proximity or business relevance).

[0105] Input the weighted fusion features into the association graph structure (constructed based on the association relationship of the folded data), and capture local feature interactions and global pattern propagation through graph convolution operations. The graph convolution operation automatically extracts the high-order neighborhood information of the data (such as the aggregation effect of the anomaly pattern) and outputs an anomaly score result, characterizing the degree to which the data deviates from the normal distribution. In this stage, the topological structure of the association graph optimizes the feature propagation path, improving the accuracy of anomaly detection. Then, establish a global benchmark reference value based on the historical data set statistics (such as mean, quantile, variance, etc.), and generate a dynamic detection threshold in combination with the current anomaly score result and the preset constraint conditions.

[0106] Specific constraint conditions may include, for example, confidence intervals: defining the threshold adjustment range to control the false alarm rate (such as the threshold fluctuation boundary at a 95% confidence level); threshold adjustment coefficients: dynamically scaling the threshold sensitivity according to data distribution changes or business requirements (such as widening the threshold when the coefficient increases to adapt to the noise environment); calculation delay boundary values: ensuring that the threshold update frequency does not exceed the real-time requirements (such as the threshold stability within the maximum allowable delay time). By balancing the detection accuracy, system response speed, and resource consumption through the constraint conditions, an anomaly detection threshold adapted to dynamic scenarios is finally generated.

[0107] Further, the step of determining the global benchmark statistic based on the historical data set and generating a dynamic detection threshold according to the preset constraint conditions based on the benchmark statistic and the anomaly score result may further include: determining the statistic of the historical data set, and aggregating the statistic through a secure multi-party calculation algorithm to obtain the global benchmark statistic; optimizing the global benchmark statistic through a loss function to obtain the optimized global benchmark statistic; performing normalization processing on the anomaly score result to obtain the normalized anomaly score; and generating a dynamic detection threshold according to the preset constraint conditions and the Bayesian optimization algorithm based on the optimized global benchmark statistic and the normalized anomaly score.

[0108] It can be understood that calculating statistics (such as mean, variance, quantile, etc.) based on the historical data set characterizes the long-term distribution characteristics of the data. Encrypting and aggregating the local statistics of distributed nodes through a secure multi-party calculation algorithm to generate the global benchmark statistic ensures that the original data of each node does not flow out, and only the aggregated result of the statistics is shared, meeting the privacy protection requirements.

[0109] Then, a loss function (such as mean square error, maximum likelihood estimation) is used to optimize the global benchmark statistic to make it more suitable for the dynamic changes of the current data distribution (such as the occurrence frequency of anomaly patterns or data drift). The optimization process is completed through a federated learning framework, and each node participates in parameter adjustment without exposing the details of local data. Further, the original anomaly score is mapped to the standard normal distribution (such as Z-score) to eliminate the influence of data scale differences on threshold generation. The normalization process is based on the global benchmark statistic (such as mean and variance) to ensure the comparability of anomaly scores under different data sources and scenarios. Finally, a dynamic detection threshold is generated according to the preset constraint conditions and the Bayesian optimization algorithm.

[0110] For ease of understanding, an example scenario of power grid equipment anomaly detection and dynamic threshold generation is used for illustration.

[0111] For example, a regional power grid deploys smart meters, current transformers, and meteorological sensors to monitor voltage fluctuations, line loads, and the impact of extreme weather on the grid in real time. The system needs to dynamically generate anomaly detection thresholds using a federated learning model to ensure detection accuracy and adapt to grid load fluctuations, while also protecting the privacy of data at each node (such as electricity usage data from different substations).

[0112] 1. Dynamic detection threshold generation process

[0113] Step 1: Gating feature weights across attention mechanisms

[0114] Input data: The folded data contains multi-source heterogeneous information, structured data: historical power consumption of smart meters, three-phase current waveform of current transformers (time series data); unstructured data: wind speed / rainfall images of meteorological sensors (generated through feature map conversion).

[0115] Gated cross-attention mechanism: Builds an association graph based on grid topology (such as the power supply path between substations); automatically assigns weights to different data types (for example, the weight of meteorological images on rainy days is increased by 30% to highlight their impact on line short circuits).

[0116] Output: weighted fusion features (e.g., the weight of the “high load + heavy rainfall” combination feature is 0.7).

[0117] Step 2: Convolution operation to obtain anomaly score

[0118] Correlation graph convolution: Performs a convolution operation on the weighted fusion features on the power grid topology to capture the propagation of abnormal patterns in local areas (such as a transmission line) (such as the tendency of voltage sags to spread along the line); outputs anomaly scores (for example, a substation's score = 0.85 indicates that its load anomaly risk is above the threshold).

[0119] Step 3: Dynamic Threshold Generation

[0120] Global benchmark statistics: Calculate the mean, standard deviation, and range of electricity consumption at each substation from historical data sets; aggregate local statistics at each node through secure multi-party computing (for example, substation A only submits the encrypted mean to prevent data leakage).

[0121] Optimization and standardization: Use a loss function (such as Huber loss) to optimize the global benchmark statistics to adapt them to recent changes in grid load (such as the increase in the proportion of renewable energy generation); standardize the anomaly score to the Z-score (Z = (anomaly score - optimized mean) / optimized standard deviation).

[0122] Bayesian optimization and constraints generate dynamic thresholds based on Z-score and preset constraints:

[0123] Confidence interval: Set the threshold as Z = 3 (99% confidence level, controlling the false alarm rate);

[0124] Adjustment coefficient: Due to typhoon warnings, temporarily relax the threshold to Z = 2.5 (allowing short-term overload);

[0125] Delay boundary: Update the threshold every 5 minutes to avoid frequent adjustments affecting real-time performance.

[0126] 2. Overall process example

[0127] Data collection and folding: Substation A collects local three-phase current data (structured), and generates statistical blocks through time series folding; Weather station B uploads rainfall images (unstructured), which are merged with the statistical blocks into folded data after feature map conversion.

[0128] Federated learning model inference: The gated cross-attention mechanism assigns weights to current and meteorological features (rainy day weight 0.7); The associated graph convolution operation finds that the current anomaly in Substation A is related to the load fluctuation in adjacent Substation B, and the anomaly score = 0.85.

[0129] Dynamic threshold generation: Secure multi-party computation aggregates the historical electricity consumption statistics of each substation (after encryption); The optimized mean = 1000 kW (original mean 950 kW, due to increased new energy generation); Z-score = (0.85 - 1.0) / optimized standard deviation = -0.15 (normalized score); Bayesian optimization combines with constraints, and the final threshold is Z = 2.5 (adjusted due to typhoon warnings).

[0130] Anomaly determination and model iteration: If the normalized score of Substation A exceeds the threshold (such as Z = 2.6), trigger incremental training; Distribute the labeled anomaly samples (high load + heavy rainfall) to each node through the federated learning framework to update the model parameters; The blockchain records the logic of this threshold adjustment (such as "During typhoons, the threshold is relaxed to Z = 2.5, valid for 2 hours").

[0131] In this embodiment, by combining graph neural networks and federated learning, using the gated cross-attention mechanism to weight and fuse features, obtaining anomaly scores through graph convolution operations, and generating dynamic detection thresholds based on historical data sets, efficient anomaly detection in a distributed environment is achieved. Moreover, the secure multi-party computation technology is used to ensure data privacy, and the global benchmark statistics are optimized through loss functions, further improving the accuracy and applicability of the statistics.

[0132] Based on the first embodiment of this application, in the third embodiment of this application, the content that is the same as or similar to the above-mentioned Embodiment 1 can be referred to the above introduction and will not be repeated hereinafter. On this basis, please refer to Figure 3 , the step S10 may include steps S101 to S104:

[0133] Step S101, collect multi-source heterogeneous data at a corresponding collection frequency according to the adaptive collection mechanism and the data change rate.

[0134] Step S102, store the multi-source heterogeneous data in the storage nodes of the target storage system, where the target storage system is a storage system constructed based on a hyper-converged storage architecture.

[0135] It should be noted that hyper-converged storage is a technology that highly integrates storage, computing, network, and management functions in a unified architecture, and realizes the coordination of data storage, resource scheduling, and task processing through software-defined means. Its core feature is to break the traditional architecture mode of "storage-computing" separation, tightly combine computing resources (such as CPUs, GPUs) with storage nodes, and allow data to complete preprocessing, analysis, or lightweight computing tasks locally, thereby reducing the overhead of data cross-node transmission and improving system efficiency and real-time performance.

[0136] In the traditional "storage-computing" separation architecture, data needs to be read from storage nodes and transmitted to independent computing nodes for processing, which has obvious performance bottlenecks - cross-layer transmission not only increases latency but also occupies a large amount of bandwidth resources. The improved hyper-converged storage system (i.e., the target storage system of this application) has added a "storage-computing integrated" unit and a fusion unit on the basis of retaining the traditional architecture (such as independent storage units, computing nodes, and network modules), achieving the following breakthroughs:

[0137] Computing sinking and local processing: The storage node is built with an embedded computing unit, which can directly complete preprocessing tasks such as feature extraction and time series folding at the data storage location without transmitting the original data to an external computing node. For example, when processing multi-source heterogeneous data, unstructured data (such as images) can be converted into feature maps through a lightweight neural network in the storage node, while structured data (such as sensor time series data) can directly generate target statistical blocks through statistical compression and dynamic time warping algorithms, avoiding data movement.

[0138] The synergistic effect of the fusion unit: The newly added fusion unit integrates storage, computing, and network functions, supporting dynamic resource allocation and task scheduling. For example, the fusion unit can automatically allocate storage space, computing tasks, and network bandwidth according to real-time load requirements, and at the same time optimize the data processing process through algorithm innovation (such as approximate nearest neighbor compression, federated learning parameter aggregation).

[0139] Efficiency and scalability improvement: Through the "in-memory computing" design, the system reduces the latency and bandwidth consumption of cross-layer communication, especially suitable for scenarios with high real-time requirements (such as anomaly detection). In addition, the fusion unit supports flexible expansion (such as adding storage capacity or computing acceleration modules as needed) to adapt to the growth of data scale and the requirements of complex applications.

[0140] Step S103: Based on the embedded computing unit of the storage node, convert the unstructured data of multi-source heterogeneous data into a feature map, and perform temporal folding on the structured data of multi-source heterogeneous data through a statistical feature compression algorithm to obtain a target statistical block.

[0141] It can be understood that in the embedded computing unit of the storage node, differential processing is performed on multi-source heterogeneous data. For unstructured data, it is transformed into a compact feature map representation through a feature map conversion algorithm (such as the feature extraction layer of a convolutional neural network), retaining key semantic information; for structured data, temporal folding is performed through a statistical feature compression algorithm (such as sliding window statistics, temporal difference aggregation) to compress the data dimension and retain temporal correlation (such as periodic fluctuation patterns). The processing results of the two types respectively generate a feature map and a target statistical block, which are used as intermediate products for subsequent merging.

[0142] As an implementation manner, step S103 in this embodiment may include: Based on the embedded computing unit of the storage node, convert the unstructured data of multi-source heterogeneous data into a feature map; perform temporal folding on the structured data of multi-source heterogeneous data to obtain a statistical block, and perform manifold learning dimensionality reduction processing on the statistical block through an approximate nearest neighbor compression algorithm to obtain a processed statistical block; align the processed statistical block according to the dynamic time warping algorithm to obtain a target statistical block.

[0143] In the unstructured data processing stage, the embedded computing unit of the storage node performs feature map conversion on the unstructured data (such as images, videos) in multi-source heterogeneous data, extracts its semantic features through a lightweight neural network (such as MobileNet, SqueezeNet) and compresses them into a low-dimensional feature map, retaining the core pattern information of the data to reduce the subsequent computational complexity.

[0144] In the structured data processing stage, time series folding is performed on the structured data to capture time correlations. Specifically, it may include, for example, time series folding: mapping time series data into compact statistical blocks through statistical feature compression algorithms (such as moving window mean, frequency domain transformation), and retaining periodic or trend features; manifold learning dimensionality reduction: using approximate nearest neighbor compression algorithms to reduce the dimensionality of statistical blocks in high-dimensional space, eliminating redundant information and retaining local similarities in data distribution; dynamic time warping: aligning the time axis of the dimensionality-reduced statistical blocks based on the dynamic time warping algorithm to solve the sampling rate differences or time offsets of different data sources, and ensuring the consistency of time series features globally.

[0145] Step S104, merge the feature map and the target statistical block through the materialized view algorithm to generate the folded data.

[0146] Logically merge the feature map and the target statistical block through the materialized view algorithm to eliminate semantic fragmentation caused by data heterogeneity. The materialized view pre-defines the association rules between the feature map and the statistical block (such as spatio-temporal alignment, business logic mapping), and generates a unified folded data format through an efficient query optimization engine, making it directly input into the subsequent anomaly detection module.

[0147] In this embodiment, the acquisition strategy is automatically adjusted according to data characteristics to balance real-time performance and storage cost. Differentiated conversion strategies are adopted for unstructured and structured data to maximize the information retention efficiency. The embedded capabilities of the hyper-converged storage node are used to implement proximal data processing, reducing data transmission overhead. Moreover, unstructured data extracts semantics through feature map conversion, and structured data retains time series patterns through time series folding and dimensionality reduction, optimizing data dimensions and time consistency, ensuring the effectiveness of multi-source data fusion. All processing is completed in the embedded unit of the storage node, reducing data transmission overhead and improving real-time performance.

[0148] It should be noted that the above examples are only for understanding this application and do not constitute a limitation on the data anomaly detection method of this application. Based on this technical concept, more forms of simple transformations are within the protection scope of this application.

[0149] This application also provides a data anomaly detection device. Please refer to Figure 4 , the data anomaly detection device includes:

[0150] A data processing module 10, configured to collect multi-source heterogeneous data and perform a folding operation on the multi-source heterogeneous data to obtain folded data;

[0151] A threshold generation module 20, configured to generate a dynamic detection threshold through a graph federated learning model, where the graph federated learning model is a model constructed based on a graph neural network algorithm and a federated learning framework;

[0152] Anomaly detection module 30 is configured to perform anomaly detection on the folded data according to the dynamic detection threshold to obtain an anomaly detection result.

[0153] The data anomaly detection device provided by the present application adopts the data anomaly detection method in the above embodiment, and can solve the technical problem of inaccurate data anomaly detection in the prior art. Compared with the prior art, the beneficial effects of the data anomaly detection device provided by the present application are the same as those of the data anomaly detection method provided by the above embodiment, and other technical features in the data anomaly detection device are the same as the features disclosed in the method of the above embodiment, which will not be elaborated here.

[0154] The present application provides a data anomaly detection device, which includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the data anomaly detection method in the first embodiment above.

[0155] Next, refer to Figure 5 , which shows a schematic structural diagram of a data anomaly detection device suitable for implementing the embodiments of the present application. The data anomaly detection device in the embodiments of the present application may include, but is not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Portable Application Desctions), PMPs (Portable Media Players), in-vehicle terminals (such as in-vehicle navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc. Figure 5 The data anomaly detection device shown is only an example and should not impose any limitation on the functions and usage scope of the embodiments of the present application.

[0156] As Figure 5As shown, the data anomaly detection device may include a processing device 1001 (such as a central processing unit, a graphics processing unit, etc.), which can perform various appropriate actions and processes according to a program stored in the read-only memory 1002 or a program loaded from the storage device 1003 into the random access memory 1004. In the random access memory 1004, various programs and data required for the operation of the data anomaly detection device are also stored. The processing device 1001, the read-only memory 1002, and the random access memory 1004 are connected to each other through a bus 1005. The input / output interface 1006 is also connected to the bus. Generally, the following systems can be connected to the input / output interface 1006: an input device 1007 including, for example, a touch screen, a touch pad, a keyboard, a mouse, an image sensor, a microphone, an accelerometer, a gyroscope, etc.; an output device 1008 including, for example, a liquid crystal display (LCD: Liquid Crystal Display), a speaker, a vibrator, etc.; a storage device 1003 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 1009. The communication device 1009 can allow the data anomaly detection device to communicate with other devices wirelessly or wiredly to exchange data. Although the figure shows a data anomaly detection device having various systems, it should be understood that it is not required to implement or have all the systems shown. Instead, more or fewer systems can be implemented or had.

[0157] In particular, according to the embodiments disclosed in the present application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, the embodiments disclosed in the present application include a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program contains program codes for executing the methods shown in the flowcharts. In such an embodiment, the computer program can be downloaded and installed from a network through the communication device, or installed from the storage device 1003, or installed from the read-only memory 1002. When the computer program is executed by the processing device 1001, the above functions defined in the methods of the embodiments disclosed in the present application are executed.

[0158] The data anomaly detection device provided by the present application adopts the data anomaly detection method in the above embodiments, and can solve the technical problem of inaccurate data anomaly detection in the prior art. Compared with the prior art, the beneficial effects of the data anomaly detection device provided by the present application are the same as those of the data anomaly detection method provided by the above embodiments, and other technical features in the data anomaly detection device are the same as those disclosed in the method of the previous embodiment, and will not be elaborated here.

[0159] It should be understood that each part disclosed in this application can be implemented by hardware, software, firmware, or a combination thereof. In the description of the above embodiments, specific features, structures, materials, or characteristics can be combined in a suitable manner in any one or more embodiments or examples.

[0160] As described above, this is only the specific implementation manner of this application, but the protection scope of this application is not limited thereto. Any person skilled in the art can easily think of changes or substitutions within the technical scope disclosed in this application, and all of them should be covered by the protection scope of this application. Therefore, the protection scope of this application should be subject to the protection scope of the claims.

[0161] This application provides a computer-readable storage medium having computer-readable program instructions (i.e., computer programs) stored thereon, and the computer-readable program instructions are used to execute the data anomaly detection method in the above embodiments.

[0162] The computer-readable storage medium provided by this application can be, for example, a USB flash drive, but is not limited to electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems or devices, or any combination of the above. More specific examples of computer-readable storage media can include, but are not limited to: electrical connections with one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM) or flash memory, optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the above. In this embodiment, the computer-readable storage medium can be any tangible medium that contains or stores a program, and this program can be used by or combined with an instruction execution system or device. The program code contained on the computer-readable storage medium can be transmitted by any appropriate medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination of the above.

[0163] The above computer-readable storage medium can be included in the data anomaly detection device; or it can exist alone without being assembled into the data anomaly detection device.

[0164] The above computer-readable storage medium stores one or more programs which, when executed by the data anomaly detection device, cause the data anomaly detection device to: collect multi-source heterogeneous data, perform a folding operation on the multi-source heterogeneous data to obtain folded data, generate a dynamic detection threshold through a graph federated learning model, where the graph federated learning model is a model constructed based on a graph neural network algorithm and a federated learning framework, and perform anomaly detection on the folded data according to the dynamic detection threshold to obtain an anomaly detection result.

[0165] Computer program code for performing the operations of the present application may be written in one or more programming languages or combinations thereof. The programming languages include object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, executed as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., by connecting through an Internet service provider via the Internet).

[0166] The readable storage medium provided by the present application is a computer-readable storage medium, and the computer-readable storage medium stores computer-readable program instructions (i.e., computer programs) for executing the above data anomaly detection method, which can solve the technical problem of inaccurate data anomaly detection in the prior art. Compared with the prior art, the beneficial effects of the computer-readable storage medium provided by the present application are the same as those of the data anomaly detection method provided in the above embodiment, and will not be elaborated here.

[0167] The present application also provides a computer program product, including a computer program which, when executed by a processor, implements the steps of the data anomaly detection method as described above.

[0168] The computer program product provided by the present application can solve the technical problem of inaccurate data anomaly detection in the prior art. Compared with the prior art, the beneficial effects of the computer program product provided by the present application are the same as those of the data anomaly detection method provided in the above embodiment, and will not be elaborated here.

[0169] The above are only some embodiments of the present application, and thus do not limit the patent scope of the present application. Any equivalent structural transformation made by using the content of the specification and drawings of the present application under the technical concept of the present application, or any direct / indirect application in other related technical fields, is included in the patent protection scope of the present application.

Claims

1. A method for detecting data anomalies, characterized in that, The method includes the following steps: Collect multi-source heterogeneous data, and perform a folding operation on the multi-source heterogeneous data to obtain folded data; Generate a dynamic detection threshold through a graph federated learning model, where the graph federated learning model is a model constructed based on a graph neural network algorithm and a federated learning framework; Perform anomaly detection on the folded data according to the dynamic detection threshold to obtain an anomaly detection result.

2. The data anomaly detection method according to claim 1, wherein Before the step of generating a dynamic detection threshold through the graph federated learning model, it further includes: Based on the distributed nodes of the graph neural network algorithm and the federated learning framework, combined with the differential privacy protocol and the homomorphic encryption protocol, deploy an initial model based on a spatio-temporal graph convolutional network; Obtain a graph federated learning model according to the initial model and a preset operator library.

3. The data anomaly detection method according to claim 2, wherein The step of obtaining a graph federated learning model according to the initial model and a preset operator library includes: In the initial model, construct an association graph according to the cosine similarity of the folded data, and embed the node features of the association graph into the initial graph structure to output a low-dimensional embedding vector; Input the low-dimensional embedding vector into each distributed node of the initial model for training to generate a parameter set; Allocate the federated learning weights of the initial model according to the node data quality; Perform weighted averaging on the parameter set based on the federated learning weights to generate global model parameters, and update the initial model through the global model parameters to obtain an updated model; Perform inference acceleration on the updated model through a preset operator library to obtain a graph federated learning model.

4. The data anomaly detection method according to claim 1, wherein The step of generating a dynamic detection threshold through the graph federated learning model includes: Based on the graph federated learning model, allocate feature weights to the folded data through a gated cross-attention mechanism to generate a weighted fusion feature; Perform a convolution operation on the weighted fusion feature based on the association graph of the graph federated learning model to obtain an anomaly score result; Determine the global benchmark statistic through a historical data set, and generate a dynamic detection threshold based on the benchmark statistic and the anomaly score result according to preset constraint conditions.

5. The data anomaly detection method according to claim 4, wherein, The step of determining the global benchmark statistic through a historical data set, and generating a dynamic detection threshold based on the benchmark statistic and the anomaly score result according to preset constraint conditions includes: Determine the statistic of the historical data set, and summarize the statistic through a secure multi-party calculation algorithm to obtain the global benchmark statistic; Optimize the global benchmark statistic through a loss function to obtain an optimized global benchmark statistic; Perform normalization processing on the anomaly score result to obtain a normalized anomaly score; Generate a dynamic detection threshold based on the optimized global benchmark statistic and the normalized anomaly score according to preset constraint conditions and a Bayesian optimization algorithm.

6. The data anomaly detection method according to any one of claims 1 to 5, characterized in that The step of collecting multi-source heterogeneous data, and performing a folding operation on the multi-source heterogeneous data to obtain folded data includes: Collect multi-source heterogeneous data at a corresponding collection frequency according to an adaptive collection mechanism and a data change rate; Store the multi-source heterogeneous data in the storage nodes of a target storage system, where the target storage system is a storage system constructed based on a hyper-converged storage architecture; An embedded computing unit based on a storage node converts the unstructured data of multi-source heterogeneous data into a feature map, and performs temporal folding on the structured data of the multi-source heterogeneous data through a statistical feature compression algorithm to obtain a target statistical block; The feature map and the target statistical block are merged through a materialized view algorithm to generate folded data.

7. The data anomaly detection method according to any one of claims 6, characterized in that The step of the embedded computing unit based on the storage node converting the unstructured data of the multi-source heterogeneous data into a feature map and performing temporal folding on the structured data of the multi-source heterogeneous data through a statistical feature compression algorithm to obtain a target statistical block includes: The embedded computing unit based on the storage node converts the unstructured data of the multi-source heterogeneous data into a feature map; Perform temporal folding on the structured data of the multi-source heterogeneous data to obtain a statistical block, and perform manifold learning dimensionality reduction processing on the statistical block through an approximate nearest neighbor compression algorithm to obtain a processed statistical block; Align the processed statistical block according to the dynamic time warping algorithm to obtain a target statistical block.

8. The data anomaly detection method according to any one of claims 1 to 5, characterized in that After the step of obtaining an anomaly detection result by performing anomaly detection on the folded data according to the dynamic detection threshold, it further includes: Perform post-processing on the anomaly detection result to obtain a processed detection result; Judge the processed detection result through a multi-modal graph neural network to obtain a judgment result; When the judgment result is a high-value anomaly sample, obtain an updated model by triggering model incremental training; Use the high-value anomaly sample as labeled data, and return the labeled data to each node through the federated learning framework to obtain a retrained model; Record the update process of the graph federated learning model through a blockchain network to obtain a traceability log.

9. A data anomaly detection device, characterized in that, The data anomaly detection device includes: A data processing module for collecting multi-source heterogeneous data and performing a folding operation on the multi-source heterogeneous data to obtain folded data; A threshold generation module for generating a dynamic detection threshold through a graph federated learning model, where the graph federated learning model is a model constructed based on a graph neural network algorithm and a federated learning framework; An anomaly detection module for performing anomaly detection on the folded data according to the dynamic detection threshold to obtain an anomaly detection result.

10. A data anomaly detection device, characterized in that, The data anomaly detection device includes: a memory, a processor, and a data anomaly detection program stored on the memory and executable on the processor. When the data anomaly detection program is executed by the processor, it implements the data anomaly detection method according to any one of claims 1 to 8.