Data anomaly detection method and device and electronic equipment
By constructing and sharing local Gaussian density functions in a privacy-preserving computing environment to generate a global Gaussian density function, and combining improved contrastive learning and knowledge distillation techniques, the accuracy problem of anomaly detection in privacy-preserving computing data is solved, improving the accuracy and robustness of detection.
Patent Information
- Application Number
- CN202511407662.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-29
- Publication Date
- 2026-01-06
AI Technical Summary
In existing technologies, the accuracy of local anomaly detection for privacy-preserving computation data is low, and it cannot effectively solve the problem of non-independent identical distribution, resulting in insufficient accuracy of data anomaly detection.
By constructing local Gaussian density functions and sharing them on the blockchain, the central node aggregates the local Gaussian density functions of multiple nodes to generate a global Gaussian density function. Anomaly detection is then performed using the global Gaussian density function, and the detection accuracy is improved by combining an improved contrastive learning method and knowledge distillation technology.
It improves the accuracy and robustness of data anomaly detection in a privacy-preserving computing environment, solves the detection bias under non-independent and identically distributed data, and enhances feature recognition capabilities and model transparency.
Smart Images

Figure CN121283718A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of data processing technology, and in particular to a data anomaly detection method, apparatus and electronic device. Background Technology
[0002] With the rapid development of the information age, data sharing has become a key factor driving innovation and economic growth, and the value and importance of data are constantly being demonstrated. However, increased awareness of personal privacy protection has made the collection of personal data and cross-organizational data sharing more difficult. In addition, technologies such as data statistics, deep learning algorithms, and large models are constantly increasing the demand for large-scale datasets. Therefore, privacy-preserving computation technology is used to solve the problem of data sharing between organizations. The "usable but not visible" nature of privacy-preserving computation data means that each party can only see its own data. Generally, anomaly detection is required for privacy-preserving computation data. During the anomaly detection process, to ensure security, nodes perform anomaly detection on their local data, for example, by detecting anomalies based on local data differences. However, local data processing cannot avoid the problem of non-independent and identically distributed data. Locally anomalous data may be normal data globally, which can easily lead to low accuracy in data anomaly detection. Summary of the Invention
[0003] This application provides a data anomaly detection method, apparatus, and electronic device to address the problem of low accuracy in existing data anomaly detection methods.
[0004] To solve the above-mentioned technical problems, this application is implemented as follows:
[0005] In a first aspect, embodiments of this application provide a data anomaly detection method, applied to a first node among multiple nodes, the method comprising:
[0006] Based on the first dataset local to the first node, a local Gaussian density function is constructed for the first node, where the first node is any node among the plurality of nodes;
[0007] The local Gaussian density function of the first node is sent to the blockchain, which is used to store the local Gaussian density functions of the plurality of nodes;
[0008] Send the contract address in the blockchain of the local Gaussian density function of the first node to the central node, and receive the global Gaussian density function sent by the central node. The mean parameter of the global Gaussian density function is the average of the mean parameters of the local Gaussian density functions of the multiple nodes obtained from the blockchain, and the covariance parameter of the global Gaussian density function is the average of the covariance parameters of the local Gaussian density functions of the multiple nodes.
[0009] Anomaly detection is performed on the first dataset based on the global Gaussian density function to obtain the detection results for the first dataset.
[0010] Secondly, embodiments of this application provide a data anomaly detection method applied to a central node, the method comprising:
[0011] The local Gaussian density function of multiple nodes is obtained from the blockchain. The local Gaussian density function of the first node is constructed by the first dataset local to the first node. The first node is any node among the multiple nodes.
[0012] Based on the local Gaussian density functions of the multiple nodes, the global Gaussian density function is determined;
[0013] The global Gaussian density function is sent to the plurality of nodes. The mean parameter of the global Gaussian density function is the average of the mean parameters of the local Gaussian density functions of the plurality of nodes, and the covariance parameter of the global Gaussian density function is the average of the covariance parameters of the local Gaussian density functions of the plurality of nodes. The global Gaussian density function is used for data anomaly detection.
[0014] Thirdly, embodiments of this application provide an anomaly detection device applied to a first node among multiple nodes, the device comprising:
[0015] A construction module is used to construct a local Gaussian density function for the first node based on a first dataset local to the first node, wherein the first node is any node among the plurality of nodes;
[0016] The first sending module is used to send the local Gaussian density function of the first node to the blockchain, wherein the blockchain is used to store the local Gaussian density functions of the plurality of nodes;
[0017] The second sending module is used to send the contract address in the blockchain of the first node's local Gaussian density function to the central node;
[0018] The first receiving module is used to receive the global Gaussian density function sent by the central node. The mean parameter of the global Gaussian density function is the average of the mean parameters of the local Gaussian density functions of the multiple nodes obtained from the blockchain, and the covariance parameter of the global Gaussian density function is the average of the covariance parameters of the local Gaussian density functions of the multiple nodes.
[0019] The detection module is used to perform anomaly detection on the first dataset based on the global Gaussian density function, and obtain the detection results of the first dataset.
[0020] Fourthly, embodiments of this application provide an anomaly detection device applied to a central node, the device comprising:
[0021] The first acquisition module is used to acquire the local Gaussian density function of multiple nodes from the blockchain. The local Gaussian density function of the first node is constructed through the first dataset local to the first node, and the first node is any node among the multiple nodes.
[0022] The first determining module is used to determine the global Gaussian density function based on the local Gaussian density functions of the multiple nodes;
[0023] The third sending module is used to send the global Gaussian density function to the plurality of nodes. The mean parameter of the global Gaussian density function is the average of the mean parameters of the local Gaussian density functions of the plurality of nodes, and the covariance parameter of the global Gaussian density function is the average of the covariance parameters of the local Gaussian density functions of the plurality of nodes. The global Gaussian density function is used for data anomaly detection.
[0024] Fifthly, embodiments of this application provide an electronic device, which is a first node, including a transceiver and a processor.
[0025] The processor is used for:
[0026] Based on the first dataset local to the first node, a local Gaussian density function is constructed for the first node, where the first node is any node among the plurality of nodes;
[0027] The local Gaussian density function of the first node is sent to the blockchain, which is used to store the local Gaussian density functions of the plurality of nodes;
[0028] Send the contract address in the blockchain of the local Gaussian density function of the first node to the central node, and receive the global Gaussian density function sent by the central node. The mean parameter of the global Gaussian density function is the average of the mean parameters of the local Gaussian density functions of the multiple nodes obtained from the blockchain, and the covariance parameter of the global Gaussian density function is the average of the covariance parameters of the local Gaussian density functions of the multiple nodes.
[0029] Anomaly detection is performed on the first dataset based on the global Gaussian density function to obtain the detection results for the first dataset.
[0030] Sixthly, embodiments of this application provide an electronic device, which is a central node and includes a transceiver and a processor.
[0031] The processor is used for:
[0032] The local Gaussian density function of multiple nodes is obtained from the blockchain. The local Gaussian density function of the first node is constructed by the first dataset local to the first node. The first node is any node among the multiple nodes.
[0033] Based on the local Gaussian density functions of the multiple nodes, the global Gaussian density function is determined;
[0034] The global Gaussian density function is sent to the plurality of nodes. The mean parameter of the global Gaussian density function is the average of the mean parameters of the local Gaussian density functions of the plurality of nodes, and the covariance parameter of the global Gaussian density function is the average of the covariance parameters of the local Gaussian density functions of the plurality of nodes. The global Gaussian density function is used for data anomaly detection.
[0035] In a seventh aspect, embodiments of this application provide an electronic device, including: a processor, a memory, and a program stored in the memory and executable on the processor. When the program is executed by the processor, it implements the steps of the data anomaly detection method described in the first aspect, or implements the steps of the data anomaly detection method described in the second aspect.
[0036] Eighthly, embodiments of this application provide a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the steps of the data anomaly detection method described in the first aspect, or implements the steps of the data anomaly detection method described in the second aspect.
[0037] Ninthly, embodiments of this application provide a computer program product including computer instructions that, when executed by a processor, implement the steps of the method as described in the first or second aspect above.
[0038] In this embodiment, the first node can construct a local Gaussian density function based on its local first dataset and send it to the blockchain. It also sends the contract address of the local Gaussian density function in the blockchain to the central node and receives a global Gaussian density function from the central node. The global Gaussian density function aggregates the local Gaussian density functions of multiple nodes. The mean parameter of the global Gaussian density function is the average of the mean parameters of the local Gaussian density functions of the multiple nodes obtained from the blockchain, and the covariance parameter of the global Gaussian density function is the average of the covariance parameters of the local Gaussian density functions of the multiple nodes. The first node uses the global Gaussian density function provided by the central node to perform anomaly detection on its local first dataset, improving the accuracy of data anomaly detection. Attached Figure Description
[0039] To more clearly illustrate the technical solutions of the embodiments of this application, the drawings used in the description of the embodiments of this application will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0040] Figure 1 This is one of the flowcharts of a data anomaly detection method provided in the embodiments of this application;
[0041] Figure 2 This is a second flowchart of a data anomaly detection method provided in an embodiment of this application;
[0042] Figure 3 This is a structural diagram of a system for implementing a data anomaly detection method provided in an embodiment of this application;
[0043] Figure 4 This is the third flowchart of a data anomaly detection method provided in the embodiments of this application;
[0044] Figure 5 This is a schematic diagram of the structure of an anomaly detection device provided in an embodiment of this application;
[0045] Figure 6 This is a schematic diagram of the structure of an anomaly detection device provided in an embodiment of this application;
[0046] Figure 7 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application;
[0047] Figure 8 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation
[0048] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0049] See Figure 1 , Figure 1 This is a flowchart of a data anomaly detection method provided in an embodiment of this application, applied to the first node among multiple nodes, where the first node can be any node among the multiple nodes. For example... Figure 1 As shown, the data anomaly detection method provided in this embodiment includes the following steps:
[0050] Step 101: Construct the local Gaussian density function of the first node based on the first dataset local to the first node.
[0051] It should be understood that each Gaussian density function includes a mean parameter and a covariance parameter, and the first node can also be called a member node.
[0052] Step 102: Send the local Gaussian density function of the first node to the blockchain. The blockchain is used to store the local Gaussian density functions of multiple nodes.
[0053] After constructing the local Gaussian density function of the first node, the first node can upload its local Gaussian density function to the blockchain so that subsequent central nodes can pull the local Gaussian density function from the blockchain, thereby improving security.
[0054] Step 103: Send the contract address of the local Gaussian density function of the first node in the blockchain to the central node, and receive the global Gaussian density function sent by the central node. The mean parameter of the global Gaussian density function is the average of the mean parameters of the local Gaussian density functions of multiple nodes obtained from the blockchain, and the covariance parameter of the global Gaussian density function is the average of the covariance parameters of the local Gaussian density functions of multiple nodes.
[0055] The central node, also known as the central server, receives the contract address of the local Gaussian density function returned by the nodes in the blockchain after the first node uploads its local Gaussian density function to the blockchain. This is the address where the local Gaussian density function is stored in the blockchain. The first node can send this contract address to the central node, so that the central node can obtain the local Gaussian density function of the first node from the blockchain based on the contract address sent by the first node. Each of the multiple nodes can upload its corresponding local Gaussian density function and send its contract address on the blockchain to the central node. The central node can then obtain the local Gaussian density functions from these multiple nodes based on their contract addresses. It can then aggregate the mean and covariance parameters of these local Gaussian density functions to determine the global Gaussian density function. The mean parameter of the global Gaussian density function is the average of the mean parameters of the local Gaussian density functions from all the nodes, and the covariance parameter of the global Gaussian density function is the average of the covariance parameters of the local Gaussian density functions from all the nodes. This global Gaussian density function is then sent to the multiple nodes. The first node can receive the global Gaussian density function sent by the central node.
[0056] Step 104: Perform anomaly detection on the first dataset based on the global Gaussian density function to obtain the detection results for the first dataset.
[0057] The first node receives the global Gaussian density function sent by the central node and uses the global Gaussian density function to perform anomaly detection on the first dataset on the first node's local machine.
[0058] In this embodiment, the first node can construct a local Gaussian density function based on its local first dataset and send it to the blockchain. It also sends the contract address of the local Gaussian density function in the blockchain to the central node and receives a global Gaussian density function from the central node. The global Gaussian density function aggregates the local Gaussian density functions of multiple nodes. The mean parameter of the global Gaussian density function is the average of the mean parameters of the local Gaussian density functions of the multiple nodes obtained from the blockchain, and the covariance parameter of the global Gaussian density function is the average of the covariance parameters of the local Gaussian density functions of the multiple nodes. The first node uses the global Gaussian density function provided by the central node to perform anomaly detection on its local first dataset, improving the accuracy of data anomaly detection.
[0059] In some embodiments, the number of local Gaussian density functions of the first node is k; constructing the local Gaussian density function of the first node based on the local first dataset of the first node includes:
[0060] Based on the global feature extraction model, features are extracted from the data in the first dataset to obtain the first feature information of the data in the first dataset. Among them, multiple nodes share the global feature extraction model.
[0061] Determine the extended dataset of the first dataset, and extract features from the extended dataset based on the global feature extraction model to obtain the feature information of the extended dataset;
[0062] Add the expanded dataset to the first dataset to update the first dataset;
[0063] Cluster the feature information of the data in the updated first dataset to obtain k clustering results, where k is a positive integer;
[0064] For k clustering results, construct k local Gaussian density functions for the first node.
[0065] It should be understood that the expanded dataset is determined based on the first dataset. Expansion can also be understood as data augmentation, and there are various methods, without specific limitations. For example, the expanded dataset might be noisy data samples from the first dataset. If the first dataset has more data in the first category (e.g., positive samples) than in the second category (e.g., negative samples), then the expanded dataset would be additional data in the second category (data different from the second category data in the first dataset). Adding the expanded dataset to the first dataset updates it. Then, the feature information of the updated data in the first dataset can be clustered to obtain k clustering results. A local Gaussian density function can be constructed for each clustering result, resulting in k local Gaussian density functions. It should be noted that after adding the expanded dataset to the first dataset to update it, subsequent uses of the first dataset will always use the first dataset updated with the expanded dataset.
[0066] In this embodiment, an extended dataset can be added to the first dataset. Since different datasets have their own characteristics, the feature information of the updated first dataset can be clustered to aggregate data with similar characteristics into a single clustering result. Thus, based on k clustering results, k local Gaussian density functions can be constructed to improve the accuracy of local Gaussian density function construction. The first node reports the k local Gaussian density functions to the blockchain so that the central node can determine the global Gaussian density function, thereby improving the accuracy of the global Gaussian density function. Subsequently, the global Gaussian density function is used for data anomaly detection, thereby improving the accuracy of data anomaly detection.
[0067] In some embodiments, the updated first dataset includes a first subset with normal detection results and a second subset with abnormal detection results; after performing anomaly detection on the first dataset based on the global Gaussian density function to obtain the detection results of the first dataset, it further includes:
[0068] Determine the first extended subset of the first subset, and determine the second extended subset of the second subset;
[0069] Anomaly detection is performed on the first and second extended subsets based on the global Gaussian density function, and the detection results of the first and second extended subsets are obtained.
[0070] Based on the second dataset and its detection results, the local feature extraction model of the first node is trained iteratively in multiple rounds. After each round of iteration, the similarity set between the second feature information and the first feature information of the second dataset is calculated, and the first loss value is determined based on the similarity set. The second feature information of the second dataset is the feature information of the second dataset output by the local feature extraction model, and the first feature information of the second dataset is the feature information of the second dataset output by the global feature extraction model. The second dataset includes the first dataset, the first extended subset, and the second extended subset.
[0071] The global feature extraction model is updated based on the minimum value among the first loss values under multiple iterations.
[0072] There are multiple ways to determine the first extended subset of the first subset and the second extended subset of the second subset. For example, the extended subset can be obtained by rotating, cropping and enlarging, or changing the grayscale of the data in the first dataset. This application does not make any specific limitation.
[0073] Each node maintains its own local feature extraction model. However, since multiple nodes share a single global feature extraction model, in this embodiment, the first node can update the global feature extraction model based on its local model. This allows the global model to learn from the local model, and subsequently, the updated global model is used for feature extraction. Multiple nodes share this updated global model. As an example, there are various ways to calculate the first loss value, as long as it represents the difference between the second feature information of the second dataset and the first feature information of the second dataset. This embodiment does not impose specific limitations; for example, the first loss value could be the cross-entropy value.
[0074] In this embodiment, the local feature extraction model of the first node is trained through multiple iterations. After each iteration, the similarity set between the second feature information of the second dataset and the first feature information of the second dataset under this iteration can be calculated, and the first loss value is determined based on the similarity set. The minimum value of the first loss value under multiple iterations is used to update the global feature extraction model to improve the accuracy of the global feature extraction model and the feature extraction accuracy, thereby improving the accuracy of data anomaly detection.
[0075] In some embodiments, before sending the local Gaussian density function of the first node to the blockchain, the method further includes:
[0076] Based on a preset shrinkage coefficient, the covariance parameter of the local Gaussian density function of the first node is subjected to covariance shrinkage in order to update the local Gaussian density function of the first node.
[0077] Due to limited node resources, a small sample size can easily lead to unstable covariance estimation. To improve the accuracy of covariance estimation, a shrinking covariance estimation method is adopted. This involves pre-setting a shrinkage coefficient to shrink the covariance parameter of the local Gaussian density function of the first node, thereby updating the local Gaussian density function of the first node and improving its accuracy.
[0078] In some embodiments, after updating the global feature extraction model, the method further includes:
[0079] Send the updated parameters of the global feature extraction model to the blockchain.
[0080] After the first node updates the global feature extraction model, it can upload the parameters of the updated global feature extraction model to the blockchain. This allows the subsequent central nodes to pull the updated global feature extraction model parameters from the blockchain. The central nodes then use these parameters to update the global feature extraction model they maintain and distribute the updated global feature extraction model parameters to multiple nodes, enabling them to synchronize the global feature extraction model parameters.
[0081] In some embodiments, anomaly detection is performed on the first dataset based on the global Gaussian density function to obtain the detection results for the first dataset, including:
[0082] The detection result of the first data in the first dataset is determined to be normal, wherein the distance between the first feature information of the first data and the global Gaussian density function is less than or equal to a preset threshold.
[0083] The detection result of the second data in the first dataset is determined to be abnormal; wherein the distance between the first feature information of the second data and the global Gaussian density function is greater than a preset threshold.
[0084] This distance ensures the degree of difference between the first feature information of the first data and the global Gaussian density function. If the distance between the first feature information of the data and the global Gaussian density function is less than or equal to a preset threshold, the anomaly detection result of the data can be determined as normal; otherwise, it is considered abnormal. Anomaly detection is achieved by comparing the distance with the preset threshold to determine whether the data is abnormal. There are various types of distances mentioned above, and this embodiment does not limit them. For example, as an example, the distance can be Mahalanobis distance, etc.
[0085] See Figure 2 , Figure 2 This is a flowchart of a data anomaly detection method provided in an embodiment of this application, applied to a central node. For example... Figure 2 As shown, the data anomaly detection method provided in this embodiment includes the following steps:
[0086] Step 201: Obtain the local Gaussian density function of multiple nodes from the blockchain. The local Gaussian density function of the first node is constructed by the first dataset local to the first node. The first node can be any node among the multiple nodes.
[0087] Step 202: Determine the global Gaussian density function based on the local Gaussian density functions of multiple nodes;
[0088] Step 203: Send the global Gaussian density function to multiple nodes. The mean parameter of the global Gaussian density function is the average of the mean parameters of the local Gaussian density functions of multiple nodes, and the covariance parameter of the global Gaussian density function is the average of the covariance parameters of the local Gaussian density functions of multiple nodes. The global Gaussian density function is used for data anomaly detection.
[0089] In some embodiments, before obtaining the local Gaussian density functions of multiple nodes from the blockchain, the method further includes:
[0090] The parameters of the global feature extraction model are sent to multiple nodes. The global feature extraction model is used by multiple nodes to extract features from their local datasets.
[0091] In some embodiments, after performing anomaly detection on the first dataset based on the global Gaussian density function and obtaining the detection results for the first dataset, the method further includes:
[0092] Obtain the updated parameters of the global feature extraction model from the blockchain;
[0093] Based on the updated parameters of the global feature extraction model, the global feature extraction model in the central node is updated, and the updated parameters of the global feature extraction model in the central node are sent to multiple nodes.
[0094] The following specific embodiments illustrate the process of the above method.
[0095] First, an introduction to the relevant technologies:
[0096] With the rapid development of the information age, data sharing has become a key factor driving innovation and economic growth, and the value and importance of data are constantly being demonstrated. On the one hand, increased awareness of personal privacy protection has made the collection of personal data and cross-organizational data sharing more difficult; on the other hand, technologies such as data statistics, deep learning algorithms, and large language models are constantly increasing the demand for large-scale datasets. Therefore, privacy computing technology is used to solve the problem of data sharing between organizations. The "usable but not visible" nature of privacy computing data means that each party can only see its own data, and anomaly detection is performed based on local data differences. However, local data processing cannot avoid the problem of non-independent and identically distributed data; local anomaly data may be normal data globally. Since nodes collect datasets based on their own usage patterns and local environments, anomaly data is relative to local data. The goal of anomaly detection in privacy computing nodes is to obtain local optima, but this may cause them to deviate from the global optimum. Under non-independent and identically distributed data, the average local model has higher divergence than the global model, and the divergence increases with the number of iterations. In addition, current methods for handling anomaly data in privacy computing datasets are not flexible enough and rely on manual intervention. Furthermore, anomaly detection tasks in related technologies often face the problem of data imbalance. Related methods may fail to learn sufficient discriminative features to distinguish between normal and abnormal samples, resulting in a situation where the number of normal samples far exceeds the number of abnormal samples. In this case, the model may be biased towards predicting the majority class (normal samples) while neglecting the identification of the minority class (abnormal samples).
[0097] The anomaly detection scheme proposed in this application can aggregate the local noise density estimation functions (local Gaussian density functions) of multiple nodes to construct a global noise density evaluation function (global Gaussian density function) for anomaly data. This solves the problem that anomaly detection member nodes can only detect data silos within their own datasets, enabling all participants in the privacy computing task to jointly maintain the global Gaussian density function and detect global anomalies. This avoids the impact of removing anomalies from member nodes due to non-independent and identically distributed data on the overall privacy computing task. Furthermore, the anomaly detection scheme proposed in this application introduces a learnable important feature matrix module in contrastive learning. Through backpropagation during the learning process, it integrates the contribution of corresponding features to the final detection task. This approach allows the model to automatically identify and emphasize the impact of more important features on the performance of the comparative learning algorithm, explicitly representing the importance of each feature. This not only improves the transparency and interpretability of the model but also enhances the quality of the learned feature representations. Furthermore, in the anomaly detection scheme proposed in this application, the global noise estimation function, model parameters, and aggregation method all involve parameter passing. The local feature extraction model update of each node contains specific features of its local data. During the aggregation process, these updates are combined into the global feature extraction model, which may indirectly reveal the unique data distribution of the node. Attackers may use this information to infer the node's sensitive data. Privacy issues are addressed by using noise to obscure the data distribution and blockchain technology to encrypt parameter passing.
[0098] like Figure 3 The diagram shown illustrates an architecture of a privacy computing platform provided in this application embodiment. This privacy computing platform comprises a central service node (central node) and member nodes (multiple nodes). The central service node primarily undertakes the transaction management module and the privacy computing network module for privacy computing. The transaction management module mainly matches privacy computing orders to generate privacy computing work orders and distributes these work orders to the supply and demand side member nodes for subsequent privacy computing operations. The privacy computing network module mainly manages and configures the network of the member nodes involved in the work orders, enabling network interconnection among the member nodes. Member nodes are divided into two categories (providers and demanders), with roles determined based on the requirements in the order. Demanders typically act as the initiators of privacy computing tasks, utilizing their own privacy computing engine and local datasets to execute privacy computing tasks. Providers can provide datasets for the privacy computing tasks of demanders through the privacy computing management module. The central service node can provide information on participation in the privacy computing process.
[0099] The anomaly detection method for privacy-preserving computation datasets (e.g., node-local datasets) based on federated learning provided in this application mainly involves constructing a global noise density estimation function to estimate anomalous data before the privacy-preserving computation task begins. By introducing a modular contrastive learning method that can learn important feature matrices, it provides secure and reliable global anomaly detection for the datasets of each participating member node. Finally, it improves detection accuracy through distillation and aggregation, thereby enhancing the accuracy and robustness of privacy-preserving computation results while protecting privacy.
[0100] This application proposes an anomaly detection method for privacy-preserving computation datasets based on federated learning and global noise density estimation. The method uses an improved contrastive learning approach for training and aggregates capabilities through ensemble distillation, extracting knowledge learned from different distributions into a global feature extraction model. Combined with blockchain-related knowledge, it ensures privacy and data security among privacy-preserving computation nodes. The overall flowchart is shown below. Figure 4 As shown, this invention includes constructing a global noise probability density function evaluation, introducing a contrastive learning training method that incorporates learnable important feature matrices, and knowledge distillation combined with ensemble learning. The technical solution of this invention will be described in detail below. The specific process is as follows:
[0101] First, a global noise probability density evaluation function is constructed to more accurately assess and process the data knowledge distribution and anomaly definitions in each node's data within the federated learning framework. Here, the noise corresponds to anomalous data in the privacy-preserving computation dataset. This evaluation function constructs a global noise density estimation function (global Gaussian density function) by analyzing the noise density estimation function (local Gaussian density function) of each cluster (clustering result) sent by nodes to the central server. Blockchain technology ensures the security of noise density function transmission between nodes and the server. The specific modeling process for obtaining the global noise density estimation function is as follows:
[0102] For each node in the federated model, a low-dimensional representation z(x) of the training dataset is obtained through a global feature extraction model f(·). The low-dimensional representation is computed through a projection layer attached after feature extraction.
[0103] Introducing samples of other classes into the local data of each node creates sample noise for the cluster. Then, an early-stopping clustering strategy is applied to the low-dimensional representation data of each node, stopping the clustering process before it fully converges—this is weak clustering. For each node, k clusters are generated through weak clustering.
[0104] For each cluster, model it using a Gaussian density function and estimate the mean μ of its Gaussian density function. k ∑ covariance kBecause clustered data introduces noise, each cluster contains some noisy samples belonging to other clusters, causing a difference between the cluster density function and the true density function, and its μ k In reality, it is the sum of the mean of the main class samples and the mean of the noise samples, ∑ k It is also the mixed covariance of the main category and the noise samples;
[0105] Shrinking covariance estimation. Due to limited node resources, the number of samples is usually small, leading to unstable covariance estimation. To improve the accuracy of covariance estimation, a shrinking covariance estimation method is adopted, which is defined as a linear combination of the empirical covariance and the identity matrix:
[0106]
[0107] D represents the covariance matrix Dimension size, I D Let ρ be an identity matrix of dimension D, and ρ be a preset shrinkage coefficient. To The updated covariance matrix obtained after shrinkage. express The trace of the covariance matrix. The matrix formed by the covariance parameters of the k local Gaussian density functions constructed for the nodes.
[0108] Each node sends its estimated density function to the server. To improve security, the node's local noise density function is uploaded to the blockchain via a smart contract. The hash value, i.e., the contract address, is obtained from the returned data. The server then uses this contract address value to retrieve the noise density function needed for subsequent processing.
[0109] The central server (central node) aggregates the local Gaussian density functions of all nodes, constructs a global density function, and shares it among the nodes. The central server obtains the local noise density functions of nodes through contract addresses, then aggregates the local noise density functions of all nodes to obtain a global noise density function (global Gaussian density function). The mean and covariance sets of the local Gaussian density functions of each node are used to evaluate normal and abnormal samples. The mean of the global Gaussian density function is the mean of the set of means of the local Gaussian density functions of each node, and the covariance of the global Gaussian density function is the mean of the set of covariances of the local Gaussian density functions of each node.
[0110] The global noise density function can be obtained through the above process. This function is used to determine whether a sample is normal or abnormal. This application determines whether a sample is normal by judging the Mahalanobis distance between the low-dimensional representation of the node's data (first feature information) and the global noise density function; this is the detection scoring function. Samples whose distance to the global noise density function is within a certain threshold are considered normal, while those that do not are considered abnormal. The detection score s(x) is determined by the following formula:
[0111]
[0112] Where z(x) represents the first feature information of data x, The mean of the global noise density function. Let be the covariance of the global noise density function.
[0113] Secondly, to improve the accuracy and robustness of anomaly detection, this invention designs a specific contrastive learning optimization method. The core idea is to construct similar and dissimilar sample pairs to narrow the distance between normal samples in the feature space and widen the distance between anomaly samples. This proposal suggests a contrastive learning method with a learnable important feature matrix for anomaly detection, which can eliminate linear transformations in feature representation, accurately identify normal sample features, and thus improve anomaly detection performance. The specific steps are as follows:
[0114] Nodes perform global selection based on the global noise density function to obtain predicted normal and abnormal samples. Predicted normal samples are categorized into positive and negative samples according to the data augmentation type, while predicted abnormal samples are augmented using only negative sample data. Data augmentation based on shift transformations (such as rotation) generates negative samples, while data augmentation based on normal transformations (such as cropping and grayscale transformation) generates positive samples.
[0115] Set the detection threshold (preset) to δ, calculate the detection score s(x) of the sample, and the sample with a score less than the detection threshold will be regarded as a positive sample and the sample with a score less than the detection threshold will be a negative sample.
[0116] To calculate the contrastive loss function for positive and negative samples, a learnable diagonal matrix S is introduced. Each element on the diagonal of S represents the importance weight of the corresponding feature dimension. The similarity between feature vectors is not directly calculated through the dot product. Instead of calculating, it is through Let's calculate it. S is learnable, meaning that the local feature extraction model automatically adjusts the importance weights of each feature dimension during training. An example formula is shown below:
[0117]
[0118] Where sim(z(x),z(x′))=z(x) T z(x′) / ‖z(x)‖‖z(x′)‖ represents the L2 norm of z(x) and z(x′), τ represents an adjustment parameter, |{x +}| represents a positive sample x + The set of matrices, |{x -}|Positive Sample x - The set of matrices, This represents the loss function value used during the training of the local feature extraction model;
[0119] Each node updates (iterates) its local feature extraction model with its local privacy data (local dataset), which is then used to update the global feature extraction model.
[0120] Then, an ensemble distillation method is used when updating the global model from the local feature extraction model. The local feature extraction model acts as the teacher model (distillation source), and the global feature extraction model acts as the student model (distillation target). Each node maintains an instance queue (the feature representation output by the global feature extraction model, i.e., the first feature information). The similarity score between the feature representation output by the local feature extraction model (i.e., the second feature information) and the feature representation stored in the instance queue is calculated. The cross-entropy (i.e., the cross-entropy loss function value, corresponding to the first loss value) between the similarity score distributions output by the local and global feature extraction models is minimized. These two distributions are aligned, and the parameters of the global model are updated through the cross-entropy loss function value, thereby effectively transferring the knowledge of each local feature extraction model to the global feature extraction model.
[0121] In summary, the overall steps of the anomaly detection method based on federated learning provided in this application embodiment are as follows:
[0122] In the initial stage, the model is trained using a standard contrastive loss function to learn the basic representation of the data;
[0123] Weak clustering is performed based on the representation extracted from the training model, and the noise density function of each cluster is estimated.
[0124] Share the local noise density function (local Gaussian density function), construct the global density function (global Gaussian density function), and align the data distribution of different nodes;
[0125] Anomalies are predicted based on a shared global Gaussian density function, and the data is augmented into positive and negative samples. Then, a contrastive learning local feature extraction model with a learnable important feature matrix is trained.
[0126] By leveraging knowledge distillation through ensemble learning, the capabilities of local feature extraction models are aggregated into a global feature extraction model;
[0127] Anomaly detection is achieved using a global feature extraction model and a global Gaussian density function.
[0128] This proposal presents an anomaly detection method for privacy-preserving computation datasets based on federated learning. The protection points are as follows:
[0129] 1. This proposal puts forward an anomaly detection method for privacy computing datasets based on federated learning. First, it uses a global noise density function to estimate the abnormal data noise in the datasets of all privacy computing participants. Then, it performs anomaly detection by improving the contrastive learning method. Finally, it refines the global model by integrating knowledge distillation, providing an automated model to solve the problem of global anomaly data detection that existing privacy computing schemes cannot solve.
[0130] 2. This proposal proposes a method for constructing a shared noise density function in a global model to align the data distribution of different nodes, improve the detection success rate of privacy computing non-independent synchronous datasets, and reduce the false negative and false positive rates of anomaly detection in privacy computing datasets;
[0131] 3. This proposal puts forward an improved contrastive learning training method, which introduces a learnable important feature matrix to improve the identifiability and interpretability of data anomaly detection sample features and improve the accuracy of data anomaly detection.
[0132] 4. Based on consortium blockchain technology, ensure the secure transmission and storage of data and models, and improve the security and privacy protection capabilities of anomaly detection in privacy computing datasets.
[0133] Compared with related technologies, the method of this application embodiment has the following advantages:
[0134] This application proposes a federated learning-based anomaly detection method for privacy-preserving computation datasets. Compared to related techniques (such as statistical methods and deep learning methods), this method effectively solves the data silo problem by sharing a noise density function. It performs global anomaly detection on the datasets of all participating nodes in the privacy-preserving computation task, mitigating the problems of non-independent and identically distributed datasets and limited data. It exhibits higher robustness when facing training sets containing noisy or anomaly data. Compared to other related methods, this application's scheme combines federated learning technology and enhances the feature representation learning capability through an enhanced contrastive learning method. This allows the model to better distinguish between normal and anomaly samples, and the introduction of a learnable important feature matrix improves the identifiability and interpretability of sample features, thereby enhancing detection performance. In summary, the method of this invention outperforms existing anomaly detection methods in terms of detection performance, mitigation of false positives and false negatives, robustness, representation capability, and overall performance.
[0135] This proposal presents an anomaly detection method for privacy-preserving computation datasets based on federated learning. Primarily used for handling anomalous data in privacy-preserving computation datasets, it can be integrated with Data Networking Spectrum Arrays (DSSN) platforms to improve the accuracy and performance of privacy-preserving computations, demonstrating significant appeal and competitiveness in practical commercial applications. First, this method leverages federated learning to share the model across different privacy-preserving computation nodes, enabling the model to have global detection capabilities and improving the overall performance of the detection model. Currently, there are no platforms on the market that automatically process anomalous data in privacy-preserving computation datasets, which is beneficial for the DSSN platform to capture market share. Second, by sharing the noise density function, the privacy data of the privacy-preserving computation nodes is effectively protected, which is particularly important in data-sensitive fields (such as government, finance, and healthcare), helping to gain the trust of customers and users. Third, this method trains the model locally and only shares model parameters, significantly reducing data transmission volume and communication costs, which is especially crucial for resource-constrained edge devices and network environments. Finally, this method has broad applicability and can be applied to different privacy-preserving computation datasets and application scenarios, providing customized anomaly detection solutions for various industries. In summary, the proposed method not only improves the performance and efficiency of anomaly detection in privacy-preserving computation datasets but also increases the appeal and competitiveness of DSSN in practical commercial applications through privacy protection and cost reduction.
[0136] like Figure 5 As shown, Figure 5 This is a schematic diagram of the structure of an anomaly detection device 500 provided in an embodiment of this application, as shown below. Figure 5 As shown, the device 500, applied to the first node of a plurality of nodes, includes:
[0137] Module 501 is used to construct the local Gaussian density function of the first node based on the first dataset local to the first node, where the first node can be any node among multiple nodes;
[0138] The first sending module 502 is used to send the local Gaussian density function of the first node to the blockchain, and the blockchain is used to store the local Gaussian density functions of multiple nodes.
[0139] The second sending module 503 is used to send the contract address of the local Gaussian density function of the first node in the blockchain to the central node;
[0140] The first receiving module 504 is used to receive the global Gaussian density function sent by the central node. The mean parameter of the global Gaussian density function is the average of the mean parameters of the local Gaussian density functions of multiple nodes obtained from the blockchain, and the covariance parameter of the global Gaussian density function is the average of the covariance parameters of the local Gaussian density functions of multiple nodes.
[0141] The detection module 505 is used to perform anomaly detection on the first dataset based on the global Gaussian density function, and obtain the detection results of the first dataset.
[0142] The anomaly detection device 500 provided in this embodiment can realize the various processes of the above-described data anomaly detection method applied to the first node. The technical features are one-to-one and can achieve the same technical effect. To avoid repetition, it will not be described again here.
[0143] See Figure 6 , Figure 6 This is a schematic diagram of the structure of an anomaly detection device 600 provided in an embodiment of this application, as shown below. Figure 6 As shown, the anomaly detection device 600, applied to the central node, includes:
[0144] The first acquisition module 601 is used to acquire the local Gaussian density function of multiple nodes from the blockchain. The local Gaussian density function of the first node is constructed through the first dataset local to the first node. The first node is any node among the multiple nodes.
[0145] The first determining module 602 is used to determine the global Gaussian density function based on the local Gaussian density functions of multiple nodes;
[0146] The third sending module 603 is used to send a global Gaussian density function to multiple nodes. The mean parameter of the global Gaussian density function is the average of the mean parameters of the local Gaussian density functions of multiple nodes, and the covariance parameter of the global Gaussian density function is the average of the covariance parameters of the local Gaussian density functions of multiple nodes. The global Gaussian density function is used for data anomaly detection.
[0147] The anomaly detection device 600 provided in this embodiment can realize the various processes of the above-described data anomaly detection method applied to the central node. The technical features are one-to-one and can achieve the same technical effect. To avoid repetition, it will not be described again here.
[0148] This application also provides an electronic device, including: a processor, a memory, and a program stored in the memory and executable on the processor. When the program is executed by the processor, it implements the various processes of the above-described data anomaly detection method embodiment applied to the first node and achieves the same technical effect. To avoid repetition, it will not be described again here.
[0149] For details, see Figure 7 This application also provides an electronic device, which is a first node, including a bus 701, a transceiver 702, an antenna 703, a bus interface 704, a processor 705, and a memory 706.
[0150] Specifically, based on the first dataset local to the first node, a local Gaussian density function is constructed for the first node, where the first node can be any node among multiple nodes;
[0151] Send the local Gaussian density function of the first node to the blockchain, which is used to store the local Gaussian density functions of multiple nodes.
[0152] Send the contract address of the local Gaussian density function of the first node in the blockchain to the central node, and receive the global Gaussian density function sent by the central node. The mean parameter of the global Gaussian density function is the average of the mean parameters of the local Gaussian density functions of multiple nodes obtained from the blockchain, and the covariance parameter of the global Gaussian density function is the average of the covariance parameters of the local Gaussian density functions of multiple nodes.
[0153] Anomaly detection is performed on the first dataset based on the global Gaussian density function, and the detection results for the first dataset are obtained.
[0154] exist Figure 7 In this document, a bus architecture (represented by bus 701) is used. Bus 701 can include any number of interconnected buses and bridges, linking various circuits including one or more processors represented by processor 705 and memory represented by memory 706. Bus 701 can also link various other circuits such as peripheral devices, voltage regulators, and power management circuits, which are well known in the art and therefore will not be described further herein. Bus interface 704 provides an interface between bus 701 and transceiver 702. Transceiver 702 can be a single element or multiple elements, such as multiple receivers and transmitters, providing a unit for communicating with various other devices over a transmission medium. Data processed by processor 705 is transmitted over a wireless medium via antenna 703, which further receives data and transmits it to processor 705.
[0155] Processor 705 manages bus 701 and general processing, and also provides various functions, including timing, peripheral interface, voltage regulation, power management, and other control functions. Memory 706 can be used to store data used by processor 705 during operation.
[0156] Optionally, the processor 705 can be a CPU, ASIC, FPGA, or CPLD.
[0157] This application also provides a computer-readable storage medium storing a computer program. When executed by a processor, this computer program implements the various processes described in the above-described data anomaly detection method embodiment applied to the first node, and achieves the same technical effect. To avoid repetition, it will not be described again here. The computer-readable storage medium may be a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.
[0158] This application also provides an electronic device, including: a processor, a memory, and a program stored in the memory and executable on the processor. When the program is executed by the processor, it implements the various processes of the above-described data anomaly detection method embodiment applied to the central node and achieves the same technical effect. To avoid repetition, it will not be described again here.
[0159] For details, see Figure 8 As shown in the figure, this application embodiment also provides an electronic device, which is a central node and includes a bus 801, a transceiver 802, an antenna 803, a bus interface 804, a processor 805, and a memory 806.
[0160] Among them, the local Gaussian density function of multiple nodes is obtained from the blockchain. The local Gaussian density function of the first node is constructed by the first dataset of the first node. The first node can be any node among the multiple nodes.
[0161] The global Gaussian density function is determined based on the local Gaussian density functions of multiple nodes.
[0162] A global Gaussian density function is sent to multiple nodes. The mean parameter of the global Gaussian density function is the average of the mean parameters of the local Gaussian density functions of the multiple nodes, and the covariance parameter of the global Gaussian density function is the average of the covariance parameters of the local Gaussian density functions of the multiple nodes. The global Gaussian density function is used for data anomaly detection.
[0163] exist Figure 8In this document, a bus architecture (represented by bus 801) is used. Bus 801 can include any number of interconnected buses and bridges, linking various circuits including one or more processors represented by processor 805 and memory represented by memory 806. Bus 801 can also link various other circuits such as peripheral devices, voltage regulators, and power management circuits, which are well known in the art and therefore will not be described further herein. Bus interface 804 provides an interface between bus 801 and transceiver 802. Transceiver 802 can be a single element or multiple elements, such as multiple receivers and transmitters, providing a unit for communicating with various other devices over a transmission medium. Data processed by processor 805 is transmitted over a wireless medium via antenna 803, which further receives data and transmits data to processor 805.
[0164] The processor 805 manages the bus 801 and handles general processing, and also provides various functions, including timing, peripheral interfaces, voltage regulation, power management, and other control functions. The memory 806 can be used to store data used by the processor 805 during operation.
[0165] Optionally, the processor 805 can be a CPU, ASIC, FPGA, or CPLD.
[0166] This application also provides a computer-readable storage medium storing a computer program. When executed by a processor, this computer program implements the various processes described in the above-described embodiments of the data anomaly detection method applied to a central node, and achieves the same technical effect. To avoid repetition, it will not be described again here. The computer-readable storage medium may be, for example, ROM, RAM, magnetic disk, or optical disk.
[0167] This application provides a computer program product, including computer instructions. When the computer instructions are executed by a processor, they implement various processes as described in the embodiments. The technical features are one-to-one and can achieve the same technical effect. To avoid repetition, they will not be described again here.
[0168] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.
[0169] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal device (which may be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods of the various embodiments of this application.
[0170] The embodiments of this application have been described above with reference to the accompanying drawings. However, this application is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of this application without departing from the spirit and scope of the claims, and all of these forms are within the protection scope of this application.
Claims
1. A data anomaly detection method, characterized by, The method is applied to a first node in a plurality of nodes, and the method comprises: constructing a local Gaussian density function of the first node based on a first data set local to the first node, the first node being any node in the plurality of nodes; sending the local Gaussian density function of the first node to a blockchain, the blockchain being used to store local Gaussian density functions of the plurality of nodes; sending a contract address of the local Gaussian density function of the first node in the blockchain to a central node, and receiving a global Gaussian density function sent by the central node, a mean parameter of the global Gaussian density function being an average of mean parameters of the local Gaussian density functions of the plurality of nodes obtained from the blockchain, a covariance parameter of the global Gaussian density function being an average of covariance parameters of the local Gaussian density functions of the plurality of nodes; performing anomaly detection on the first data set based on the global Gaussian density function to obtain a detection result of the first data set.
2. The method of claim 1, wherein, The number of the local Gaussian density functions of the first node is k; and the constructing of the local Gaussian density function of the first node based on the first data set local to the first node comprises: performing feature extraction on data in the first data set based on a global feature extraction model to obtain first feature information of the data in the first data set, wherein the plurality of nodes share the global feature extraction model; determining an extended data set of the first data set, and performing feature extraction on the extended data set based on the global feature extraction model to obtain feature information of the extended data set; adding the extended data set to the first data set to update the first data set; clustering the feature information of the data in the updated first data set to obtain k clustering results, k being a positive integer; constructing the k local Gaussian density functions of the first node based on the k clustering results.
3. The method of claim 2, wherein, The updated first data set comprises a first subset with a normal detection result and a second subset with an abnormal detection result; and after the performing of the anomaly detection on the first data set based on the global Gaussian density function to obtain the detection result of the first data set, the method further comprises: determining a first extended subset of the first subset and a second extended subset of the second subset; performing anomaly detection on the first extended subset and the second extended subset based on the global Gaussian density function to obtain a detection result of the first extended subset and a detection result of the second extended subset; and performing multi-round iterative training on the local feature extraction model local to the first node based on the second data set and a detection result of the second data set, and after each round of iteration is completed, calculating a similarity set between second feature information of the second data set and first feature information of the second data set in the round of iteration, and determining a first loss value according to the similarity set; wherein the second feature information of the second data set is feature information of the second data set output by the local feature extraction model, the first feature information of the second data set is feature information of the second data set output by the global feature extraction model, and the second data set includes the first data set, the first extended subset, and the second extended subset; updating the global feature extraction model according to a minimum value in the first loss values in multiple rounds of iteration.
4. The method of claim 2, wherein, Before the sending of the local Gaussian density function of the first node to the blockchain, further comprising: performing covariance shrinkage on the covariance parameter of the local Gaussian density function of the first node according to a preset shrinkage coefficient, to update the local Gaussian density function of the first node.
5. The method of claim 3, wherein, After the updating of the global feature extraction model, further comprising: sending the parameters of the updated global feature extraction model to the blockchain.
6. The method according to any one of claims 1-5, characterized in that, The abnormal detection of the first data set based on the global Gaussian density function to obtain a detection result of the first data set, comprising: determining that the detection result of first data in the first data set is normal, wherein the distance between the first feature information of the first data and the global Gaussian density function is less than or equal to a preset threshold value; determining that the detection result of second data in the first data set is abnormal; wherein the distance between the first feature information of the second data and the global Gaussian density function is greater than the preset threshold value.
7. A data anomaly detection method characterized by, The method applied to a center node, comprising: obtaining local Gaussian density functions of multiple nodes from a blockchain, wherein the local Gaussian density function of a first node is constructed by a first data set local to the first node, and the first node is any node in the multiple nodes; determining a global Gaussian density function based on the local Gaussian density functions of the multiple nodes; sending the global Gaussian density function to the multiple nodes, wherein a mean parameter of the global Gaussian density function is an average of mean parameters of the local Gaussian density functions of the multiple nodes, and a covariance parameter of the global Gaussian density function is an average of covariance parameters of the local Gaussian density functions of the multiple nodes, and the global Gaussian density function is used for data anomaly detection.
8. The method of claim 7, wherein, Before the obtaining of the local Gaussian density functions of the multiple nodes from the blockchain, further comprising: sending parameters of a global feature extraction model to the multiple nodes, wherein the global feature extraction model is used for feature extraction of data sets local to the multiple nodes.
9. The method of claim 8, wherein, After the abnormal detection of the first data set based on the global Gaussian density function to obtain a detection result of the first data set, further comprising: obtaining the parameters of the updated global feature extraction model from the blockchain; According to the updated parameters of the global feature extraction model, the global feature extraction model in the center node is updated, and the parameters of the updated global feature extraction model of the center node are sent to the plurality of nodes.
10. A data anomaly detection apparatus characterized by comprising: The device is applied to a first node in a plurality of nodes, and the device comprises: A construction module is configured to construct a local Gaussian density function of the first node based on a first data set locally stored in the first node, the first node being any node in the plurality of nodes; A first sending module is configured to send the local Gaussian density function of the first node to a blockchain, the blockchain being configured to store local Gaussian density functions of the plurality of nodes; A second sending module is configured to send a contract address of the local Gaussian density function of the first node in the blockchain to a center node; A first receiving module is configured to receive a global Gaussian density function sent by the center node, a mean parameter of the global Gaussian density function being an average of mean parameters of the local Gaussian density functions of the plurality of nodes obtained from the blockchain, and a covariance parameter of the global Gaussian density function being an average of covariance parameters of the local Gaussian density functions of the plurality of nodes; A detection module is configured to perform anomaly detection on the first data set based on the global Gaussian density function, and obtain a detection result of the first data set.
11. A data anomaly detection apparatus characterized by comprising: The device is applied to a center node, and the device comprises: A first obtaining module is configured to obtain local Gaussian density functions of a plurality of nodes from a blockchain, a local Gaussian density function of a first node being constructed based on a first data set locally stored in the first node, the first node being any node in the plurality of nodes; A first determining module is configured to determine a global Gaussian density function based on the local Gaussian density functions of the plurality of nodes; A third sending module is configured to send the global Gaussian density function to the plurality of nodes, a mean parameter of the global Gaussian density function being an average of mean parameters of the local Gaussian density functions of the plurality of nodes, and a covariance parameter of the global Gaussian density function being an average of covariance parameters of the local Gaussian density functions of the plurality of nodes, the global Gaussian density function being used for data anomaly detection.
12. An electronic device, comprising: The electronic device is a first node, comprising a transceiver and a processor, The processor is configured to: construct a local Gaussian density function of the first node based on a first data set locally stored in the first node, the first node being any node in the plurality of nodes; send the local Gaussian density function of the first node to a blockchain, the blockchain being configured to store local Gaussian density functions of the plurality of nodes; send a contract address of the local Gaussian density function of the first node in the blockchain to a center node, and receive a global Gaussian density function sent by the center node, a mean parameter of the global Gaussian density function being an average of mean parameters of the local Gaussian density functions of the plurality of nodes obtained from the blockchain, and a covariance parameter of the global Gaussian density function being an average of covariance parameters of the local Gaussian density functions of the plurality of nodes. Performing anomaly detection on the first data set based on the global Gaussian density function, to obtain a detection result of the first data set.
13. An electronic device, comprising: The electronic device is a center node, comprising a transceiver and a processor, The processor is configured to: obtain local Gaussian density functions of a plurality of nodes from a blockchain, a local Gaussian density function of a first node being constructed by a first data set local to the first node, the first node being any node in the plurality of nodes; determine a global Gaussian density function based on the local Gaussian density functions of the plurality of nodes; send the global Gaussian density function to the plurality of nodes, a mean parameter of the global Gaussian density function being an average of mean parameters of the local Gaussian density functions of the plurality of nodes, a covariance parameter of the global Gaussian density function being an average of covariance parameters of the local Gaussian density functions of the plurality of nodes, the global Gaussian density function being used for data anomaly detection.
14. An electronic device, comprising: comprise: a processor, a memory, and a program stored on the memory and executable on the processor, the program, when executed by the processor, implementing the steps of the method of any one of claims 1 to 6, or implementing the steps of the method of any one of claims 7 to 9.
15. A computer-readable storage medium, characterized in that, The computer program is stored on the computer readable storage medium, and when executed by the processor, implements the steps of the method of any one of claims 1 to 6, or implements the steps of the method of any one of claims 7 to 9.
16. A computer program product, characterised in that, comprise computer instructions, which, when executed by the processor, implement the steps of the method of any one of claims 1 to 6, or implement the steps of the method of any one of claims 7 to 9.
Citation Information
Patent Citations
Data acquisition method and device, storage medium and system
CN111078488A
Intelligent factory block chain anomaly detection method
CN116866017A
Dynamic neural distribution function machine learning architecture
WO2024145079A1