In-memory Computing and Training Integrated Method for Anomaly Detection in Cloud Computing Environment Based on Variational Autoencoder

By adopting distributed variational autoencoder model and virtualization technology in the cloud computing environment, the problem of abnormal detection of large-scale and distributed data is solved, and efficient and accurate abnormal detection is achieved to adapt to the complexity of the cloud computing environment.

CN119109825BActive Publication Date: 2025-06-13HUNAN UNIV OF SCI & ENG
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411133514.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-08-19
Publication Date
2025-06-13
Estimated Expiration
2044-08-19

AI Technical Summary

Technical Problem

When the prior art processes large-scale and distributed data in a cloud computing environment, there are problems such as difficulty in maintaining data consistency, insufficient processing capabilities of high-dimensional data, poor real-time detection performance, and difficulty in adapting to distributed environments.

Method used

The distributed variational autoencoder model based on TensorFlow is adopted to realize distributed storage and calculation of data through distributed computing architecture and virtualization technology, and the distributed gradient descent algorithm and parameter server architecture are used to ensure the consistency and synchronization of model parameters.

Benefits of technology

It realizes abnormal detection of large-scale and distributed data in the cloud computing environment, improves the system's response speed and computing efficiency, enhances the accuracy and robustness of the detection, and is suitable for diversified data and dynamic changes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119109825B_ABST
    Figure CN119109825B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for anomaly detection and integrated storage, computing, and training in a cloud computing environment based on a variational autoencoder. S1: Design and initialize a distributed variational autoencoder model based on TensorFlow to form reconstructed data; S2: In the cloud computing environment, construct a distributed data storage and computing architecture; S3: Deploy and configure the distributed computing framework of TensorFlow on multiple computing nodes, and configure parameter servers and worker computing nodes; S4: Each computing node uses the distributed strategy of TensorFlow to parallelly obtain the preprocessed data set stored locally; S5: After training is completed, generate reconstructed data through the decoder; S6: After each computing node compares the detected reconstruction error with a preset threshold, mark the data exceeding the threshold as abnormal. Through an integrated design, the present invention successfully realizes anomaly detection of large-scale and distributed data in a cloud computing environment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of cloud computing, and in particular, to a method for integrated storage, computing, and training of anomaly detection in a cloud computing environment based on a variational autoencoder. Background Art

[0002] In a cloud computing environment, with the explosive growth of data scale and the increasing complexity of computing requirements, anomaly detection, as a key security guarantee mechanism, has become particularly important. Existing anomaly detection methods mainly rely on traditional machine learning and statistical models. These methods perform well when dealing with small-scale and static data sets, but they often have significant limitations when facing large-scale, dynamic, and distributed data in a cloud computing environment.

[0003] First of all, traditional anomaly detection methods usually assume that data is stored and processed in a centralized manner, which is difficult to achieve in a distributed cloud computing environment. Data in a cloud computing environment is usually distributed and stored on multiple computing nodes and transmitted and processed through a network. This distributed architecture increases the complexity of data processing, making it difficult to effectively deploy traditional centralized anomaly detection models in practical applications. In addition, the data consistency problem brought about by distributed storage and computing also makes it difficult for traditional methods to maintain high accuracy and low false alarm rates in a multi-node environment.

[0004] Secondly, existing anomaly detection methods show poor adaptability when dealing with high-dimensional data. With the increase in data complexity, traditional methods based on rules or simple statistical models are difficult to capture complex patterns and potential features in the data. This results in insufficient accuracy and robustness of anomaly detection in practical applications, and false alarms or missed detections are likely to occur. Especially when facing a dynamically changing cloud computing environment, this problem is even more prominent.

[0005] In addition, the real-time performance and efficiency of traditional anomaly detection methods are also a major problem. In a cloud computing environment, the demand for real-time processing of data streams is increasing, and traditional batch processing methods can no longer meet this demand. In the prior art, although some machine learning-based anomaly detection methods can process large-scale data, their training process usually takes a long time and requires a large amount of computing resources. Even after the model training is completed, the real-time detection process still relies on a large amount of computing resource allocation and data transmission, which further increases the response time of the detection system and is difficult to meet the real-time detection requirements of high concurrency and high throughput.

[0006] In view of the deficiencies of the above-mentioned prior art, researchers have proposed some improvement schemes, such as using the autoencoder model in deep learning technology for anomaly detection. However, the traditional autoencoder model still has limited effectiveness when dealing with large-scale distributed data. Although the autoencoder model can capture the features of data by learning the low-dimensional representation of the data, it still shows certain deficiencies when processing high-dimensional and dynamic data in a distributed environment. In addition, when detecting anomalies, the traditional autoencoder often relies on a single error metric standard, lacking flexibility and adaptability.

[0007] In summary, when dealing with large-scale and distributed data in a cloud computing environment, the prior art has problems such as difficulty in maintaining data consistency, insufficient high-dimensional data processing ability, poor real-time detection performance, and the existing VAE model being difficult to adapt to a distributed environment. These technical deficiencies severely restrict the effectiveness and practicality of anomaly detection in a cloud computing environment. Summary of the Invention

[0008] An object of the present invention is to propose a method for integrated storage, computing, and training of anomaly detection in a cloud computing environment based on a variational autoencoder. Through an integrated design, the present invention successfully realizes anomaly detection of large-scale and distributed data in a cloud computing environment.

[0009] A method for integrated storage, computing, and training of anomaly detection in a cloud computing environment based on a variational autoencoder according to an embodiment of the present invention includes the following steps:

[0010] S1. Design and initialize a distributed variational autoencoder model based on TensorFlow. The encoder is responsible for encoding the input data into a probability distribution in the latent space, and the decoder is responsible for decoding the variables in the latent space back to the original data space to form reconstructed data;

[0011] S2. In a cloud computing environment, construct a distributed data storage and computing architecture, and multiple computing nodes perform resource management and task scheduling through virtualization technology;

[0012] S3. Deploy and configure the distributed computing framework of TensorFlow on multiple computing nodes, and use its distributed strategy to distribute the training task and anomaly detection task of the variational autoencoder model on multiple computing nodes for execution. Configure a parameter server and worker computing nodes to enable each computing node to work collaboratively in a distributed environment;

[0013] S4. Each computing node uses the distributed strategy of TensorFlow to parallelly obtain the preprocessed data set stored locally, execute the forward propagation and backward propagation processes of the distributed variational autoencoder model, and synchronously update the model parameters among the computing nodes using the distributed gradient descent algorithm. Through the parameter server architecture of TensorFlow, the consistency and synchronization of the model parameters of each computing node are achieved;

[0014] S5. After the training is completed, each computing node uses the trained distributed variational autoencoder model to perform anomaly detection on the real-time input data. Each computing node maps the input data to the latent space through the encoder and generates reconstructed data through the decoder;

[0015] S6. Through the distributed integration strategy of TensorFlow, the anomaly detection results of multiple computing nodes are fused. After each computing node compares the detected reconstruction error with the preset threshold, the data exceeding the threshold is marked as abnormal, and the detection results of multiple computing nodes are summarized through integration strategies such as weighted average or voting to obtain the global anomaly detection conclusion.

[0016] Optionally, the specific steps of S1 include:

[0017] S11. Using the TensorFlow framework, define the structure of the variational autoencoder model. The encoder consists of several layers of fully connected neural networks and is responsible for mapping the input data X to the probability distribution in the latent space:

[0018] Z~N(μ(X),σ 2 (X));

[0019] where μ(X) and σ(X) are the mean function and standard deviation function respectively, representing the distribution characteristics of the input data in the latent space;

[0020] S12. Use a distributed parameter initialization method to initialize the network parameters of each layer in the encoder. The parameters include the weight matrix W and the bias vector b. By distributing the global parameters W and b to multiple computing nodes, parallel initialization is achieved to meet the high concurrency requirements in the cloud computing environment;

[0021] S13. In the decoder part, use a distributed backpropagation neural network to decode the variable Z in the latent space back to the original data space The structure of the decoder is symmetric to that of the encoder, and finally the reconstructed data is output:

[0022]

[0023] where D(·) represents the mapping function of the decoder;

[0024] S14. Define an improved reconstruction error loss function Combine the input data X and the reconstructed data Introduce the synchronization consistency metric δ between the mean square errors between the computing nodes in the cloud computing environment. The synchronization consistency metric is used to ensure that in a distributed computing environment, the reconstruction results between the computing nodes are consistent:

[0025]

[0026] where N is the number of data samples, ∥·∥ represents the norm, λ 1 is the weight factor, K is the number of computing nodes, and δ k represents the deviation between the reconstruction result of the k-th computing node and the global synchronization result;

[0027] S15. Define an improved KL divergence loss function When measuring the difference between the latent distribution Z output by the encoder and the standard normal distribution, introduce the latent space distribution consistency constraint between the computing nodes in the distributed training environment to make the latent distributions of the computing nodes consistent globally:

[0028]

[0029] where d is the dimension of the latent space, η 1 is the weight factor, and Δ k represents the deviation between the latent distribution of the k-th computing node and the global latent distribution;

[0030] S16. Use the distributed Adam optimization algorithm to train the parameters of the variational autoencoder model, and the optimization goal is to minimize the total loss function:

[0031]

[0032] where β 1 is the hyperparameter that adjusts the weight between the reconstruction error and the KL divergence.

[0033] Optionally, the S2 specifically includes:

[0034] S21. In the cloud computing environment, use virtualization technology to build a distributed data storage and computing architecture. The distributed data storage and computing architecture includes multiple virtual computing nodes C k , k = 1, 2, …, K, where K is the total number of virtual computing nodes. The virtual computing nodes are connected through a network to achieve data transmission and distributed processing of computing tasks;

[0035] S22. In the virtual computing node C kOn it, a distributed storage system is deployed, and each virtual computing node uses the HDFS distributed file system to implement block storage and management of data. Each data block D i , i = 1, 2, …, M, is redundantly stored among multiple virtual computing nodes, where M is the total number of data blocks;

[0036] S23. In the distributed data storage and computing architecture, the computing resources are allocated and scheduled through a virtualization resource manager. The virtualization resource manager dynamically allocates CPU, memory, and storage according to the computing requirements of tasks, maximizing the utilization rate of computing resources of each virtual computing node, and adopting resource isolation technology to prevent resource competition and interference between different tasks;

[0037] S24. Use a distributed task scheduler to allocate the training task and anomaly detection task of the variational autoencoder model to each computing node. The distributed task scheduler dynamically adjusts the execution order and distribution of tasks based on task priority, node computing power, and network latency factors, optimizing the overall computing efficiency and parallelism of task execution;

[0038] S25. Coordinate task execution through the synchronization mechanism between virtual computing nodes. Each virtual computing node regularly exchanges task execution status and intermediate calculation results, and uses the Paxos distributed consistency protocol for global consistency control.

[0039] Optionally, the specific content of S3 includes:

[0040] S31. Deploy and configure the distributed computing framework of TensorFlow on multiple virtual computing nodes C k . The distributed computing framework includes a parameter server and worker computing nodes. The parameter server is responsible for centrally storing and updating global model parameters, and the worker computing nodes are responsible for distributed data processing and model calculation;

[0041] S32. Initialize the global parameters of the variational autoencoder model on the parameter server, including the weight matrix W, the bias vector b, and the latent space distribution parameters μ and σ:

[0042] W (0) , b (0) , μ (0) (X), σ (0) (X);

[0043] Among them, μ (0) (X) and σ (0) (X) are the initial mean and standard deviation functions of the input data X, and the parameters distribute the initial global parameters to each worker virtual computing node C through a distributed strategy k ;

[0044] S33. On each worker virtual computing node C k configure the distributed strategy of TensorFlow. Each worker node receives the global model parameters from the parameter server and uses the data blocks D i stored locally to perform forward propagation and backpropagation calculations for the local model:

[0045]

[0046] where represents the latent variable on the worker virtual computing node C k , ∈ (k) is a noise variable sampled from the standard normal distribution, and ⊙ represents element-wise multiplication. Backpropagation updates the local model parameters by calculating the reconstruction error and KL divergence;

[0047] S34. After completing the local model training on each worker virtual computing node C k calculate the gradients of the local model parameters and pass the gradient information g k back to the parameter server:

[0048]

[0049] where is the reconstruction error, is the KL divergence, represents the gradient calculation for the parameter weight matrix W, bias vector b, and latent space distribution parameters μ and σ. The parameter server updates the global model parameters according to the gradient information g k returned by all worker nodes;

[0050] S35. The process of the parameter server updating the global model parameters is expressed as:

[0051]

[0052] where α 2 is the learning rate, t represents the number of iterations. The parameter server distributes the updated global model parameters to all worker computing nodes, and each worker node uses the latest model parameters to continue the next round of local model training. Repeat steps S31 - S35 until the model training converges.

[0053] Optionally, the specific steps of S5 include:

[0054] S51. After training is completed, each virtual computing node C k uses the trained distributed variational autoencoder model to process the real-time input data stream X i ;

[0055] S52. Calculate the input data X i and the reconstructed data to obtain the reconstruction error value

[0056]

[0057] wherein, X i,m and respectively represent the m-th feature dimension of the input data X i and the reconstructed data , M is the total number of feature dimensions, α m is the weight factor of the m-th dimension, used to adjust the relative importance of each feature dimension in the reconstruction error calculation; γ m is the non-linear adjustment factor of the m-th dimension, λ m is the wavelength parameter related to the m-th dimension, used to introduce a non-linear oscillation component into the reconstruction error to capture potential complex patterns.

[0058] Optionally, the S6 specifically includes:

[0059] S61. After comparing the detected reconstruction error value by each virtual computing node with the preset anomaly detection threshold τ k , mark the data exceeding the threshold as abnormal data, and the marking result r k (X i ) is expressed as:

[0060]

[0061] wherein, r k (X i ) represents the detection result of the virtual computing node for the input data, 1 indicates abnormal, and 0 indicates normal;

[0062] S62. Each virtual computing node sends the detection result r k (X i ) to the parameter server, and the parameter server performs weighted average processing according to the detection results r k (X i ) received from each node to obtain the global detection result R(X i ) as follows:

[0063]

[0064] wherein, w k represents the weight of the virtual computing node, and the weight can be set according to the computing power of the node, the data processing volume, or the credibility of the detection result;

[0065] S63. According to the global detection result R(X i ) and the global anomaly detection threshold τ global for comparison, if R(X i ) > τ global , then mark the data X i as global anomaly data. If R(X i ) ≤ τ global , then mark it as normal data. The global anomaly detection threshold τ global is dynamically adjusted by the following formula:

[0066] τ global = μ R + σ R · λ 3 ;

[0067] where μ R and σ R are respectively the mean and standard deviation of the detection results R(X i ) of all nodes, and λ 3 is an adjustment factor used to control the sensitivity of global anomaly detection;

[0068] S64. The parameter server stores the global anomaly detection result R(X i ) in the global storage system and generates an anomaly detection report, which includes the detailed information of the anomaly data, the detection time, and the detection results of each node.

[0069] The beneficial effects of the present invention are:

[0070] (1) By closely integrating data storage, computing, and model training, the present invention realizes the integrated design of storage, computing, and training. In the cloud computing environment, through virtualization technology and distributed computing architecture, each computing node can work collaboratively to execute model training and inference tasks under the distributed TensorFlow framework, significantly reducing the overhead of data transmission and task scheduling, avoiding the performance bottleneck caused by the separation of storage and computing in traditional methods, thereby greatly improving the response speed and computing efficiency of the system. By using the distributed parameter server architecture and distributed gradient descent algorithm, the consistency of global model parameters is ensured, further enhancing the stability and efficiency of the system.

[0071] (2) The present invention introduces the VAE model into the anomaly detection task in a distributed cloud computing environment. By modeling the latent space of data through VAE, it can better capture the latent patterns and complex structures of the data. The VAE model training and inference processes executed in parallel on multiple virtual computing nodes can effectively handle the anomaly detection task of high-dimensional and complex data. At the same time, the latent space distribution of the VAE model is optimized through KL divergence regularization, making the model capture the normal data distribution more accurately, thereby having higher accuracy and robustness in detecting abnormal data. Combining the global and local reconstruction error metrics further reduces the probability of false positives and false negatives.

[0072] (3) The present invention proposes an improved method for calculating the reconstruction error. By introducing a non-linear adjustment factor and a dynamic threshold adjustment mechanism, anomaly detection can adapt to the complexity of different feature spaces in the cloud computing environment. In the calculation of the reconstruction error, a non-linear adjustment factor based on the feature dimension is added, enabling the system to capture complex patterns and abnormal features in the data. In addition, by combining the threshold dynamic adjustment mechanism with the mean and variance of the latent space distribution, the sensitivity of anomaly detection can be dynamically adjusted according to the distribution characteristics of the data, thus maintaining high detection performance in different data scenarios. The present invention is applicable to diverse data and dynamically changing features in the cloud computing environment, improving the adaptability and detection accuracy of the system.

[0073] (4) The present invention adopts a distributed integration strategy. By performing weighted averaging and global consistency processing on the detection results of multiple virtual computing nodes, the reliability and consistency of the final anomaly detection conclusion are ensured. In the cloud computing environment, there may be differences in the computing power and data processing volume of each node. By setting node weights and dynamically adjusting the global threshold, the detection results of each node can be effectively fused, avoiding the impact of false positives or false negatives of a single node on the overall detection result. It not only improves the accuracy of anomaly detection but also enhances the fault tolerance and reliability of the system in a distributed environment. BRIEF DESCRIPTION OF THE DRAWINGS

[0074] The drawings are used to provide a further understanding of the present invention and constitute a part of the specification. They are used together with the embodiments of the present invention to explain the present invention and do not constitute a limitation to the present invention. In the drawings:

[0075] Figure 1 is a flowchart of an in-memory computing, training, and anomaly detection method for cloud computing environments based on variational autoencoders proposed by the present invention;

[0076] Figure 2 is a schematic diagram of the distributed storage and computing architecture of data in a cloud computing environment in an in-memory computing, training, and anomaly detection method for cloud computing environments based on variational autoencoders proposed by the present invention. Detailed Implementation Manner

[0077] Now, the present invention will be further described in detail with reference to the accompanying drawings. These drawings are all simplified schematic diagrams, only illustrating the basic structure of the present invention in a schematic manner, so they only show the components related to the present invention.

[0078] Reference Figure 1-2 , an integrated memory, computing, and training method for anomaly detection in a cloud computing environment based on a variational autoencoder, includes the following steps:

[0079] S1. Design and initialize a distributed variational autoencoder model based on TensorFlow. The encoder is responsible for encoding the input data into a probability distribution in the latent space, and the decoder is responsible for decoding the variables in the latent space back to the original data space to form reconstructed data;

[0080] S2. In the cloud computing environment, build a distributed data storage and computing architecture, and multiple computing nodes perform resource management and task scheduling through virtualization technology;

[0081] S3. Deploy and configure the distributed computing framework of TensorFlow on multiple computing nodes. Utilize its distributed strategy to distribute the training task and anomaly detection task of the variational autoencoder model on multiple computing nodes for execution. Configure the parameter server and worker computing nodes to enable each computing node to work collaboratively in the distributed environment;

[0082] S4. Each computing node uses the distributed strategy of TensorFlow to parallelly obtain the preprocessed data set stored locally, execute the forward propagation and backward propagation processes of the distributed variational autoencoder model, and use the distributed gradient descent algorithm to synchronously update the model parameters among the computing nodes. Through the parameter server architecture of TensorFlow, achieve the consistency and synchronization of the model parameters of each computing node;

[0083] S5. After the training is completed, each computing node uses the trained distributed variational autoencoder model to perform anomaly detection on the real-time input data. Each computing node maps the input data to the latent space through the encoder and generates reconstructed data through the decoder;

[0084] S6. Through the distributed integration strategy of TensorFlow, fuse the anomaly detection results of multiple computing nodes. After each computing node compares the detected reconstruction error with a preset threshold, mark the data exceeding the threshold as abnormal, and summarize the detection results of multiple computing nodes through an integration strategy such as weighted average or voting to obtain a global anomaly detection conclusion.

[0085] In this embodiment, S1 specifically includes:

[0086] S11. Use the TensorFlow framework to define the structure of the variational autoencoder model. The encoder consists of several layers of fully connected neural networks and is responsible for mapping the input data X to a probability distribution in the latent space:

[0087] Z ∼ N(μ(X), σ 2 (X));

[0088] where μ(X) and σ(X) are the mean function and the standard deviation function respectively, representing the distribution characteristics of the input data in the latent space;

[0089] S12. Use a distributed parameter initialization method to initialize the network parameters of each layer in the encoder. The parameters include the weight matrix W and the bias vector b. By distributing the global parameters W and b to multiple computing nodes, parallel initialization is achieved to meet the high concurrency requirements in the cloud computing environment;

[0090] S13. In the decoder part, use a distributed backpropagation neural network to decode the variable Z in the latent space back to the original data space The decoder structure is symmetric to the encoder, and the final output is the reconstructed data:

[0091]

[0092] where D(·) represents the mapping function of the decoder;

[0093] S14. Define an improved reconstruction error loss function Combine the mean square error between the input data X and the reconstructed data Introduce the synchronization consistency metric δ between computing nodes in the cloud computing environment. The synchronization consistency metric is used to ensure that in a distributed computing environment, the reconstruction results between computing nodes are consistent:

[0094]

[0095] where N is the number of data samples, ∥·∥ represents the norm, λ 1 is the weight factor, K is the number of computing nodes, and δ k represents the deviation between the reconstruction result of the k-th computing node and the global synchronization result;

[0096] S15. Define an improved KL divergence loss function When measuring the difference between the latent distribution Z output by the encoder and the standard normal distribution, introduce the latent space distribution consistency constraint between computing nodes in the distributed training environment to make the latent distributions of each computing node consistent globally:

[0097]

[0098] where d is the dimension of the latent space, and η 1 is the weight factor, and Δ k represents the deviation between the latent distribution of the k-th computing node and the global latent distribution;

[0099] S16. Use the distributed Adam optimization algorithm to train the parameters of the variational autoencoder model, and the optimization objective is to minimize the total loss function:

[0100]

[0101] where β 1 is a hyperparameter that adjusts the weight between the reconstruction error and the KL divergence.

[0102] In this embodiment, S2 specifically includes:

[0103] S21. In the cloud computing environment, use virtualization technology to build a distributed data storage and computing architecture. The distributed data storage and computing architecture includes multiple virtual computing nodes C k , k = 1, 2, …, K, where K is the total number of virtual computing nodes. The virtual computing nodes are connected through a network to achieve data transmission and distributed processing of computing tasks;

[0104] S22. On the virtual computing node C k , deploy a distributed storage system. Each virtual computing node uses the HDFS distributed file system to achieve block storage and management of data. Each data block D i , i = 1, 2, …, M, is redundantly stored among multiple virtual computing nodes, where M is the total number of data blocks;

[0105] S23. In the distributed data storage and computing architecture, allocate and schedule computing resources through a virtualization resource manager. The virtualization resource manager dynamically allocates CPU, memory, and storage according to the computing requirements of tasks, maximizes the utilization rate of computing resources of each virtual computing node, and adopts resource isolation technology to prevent resource competition and interference between different tasks;

[0106] S24. Use a distributed task scheduler to allocate the training task and anomaly detection task of the variational autoencoder model to each computing node. The distributed task scheduler dynamically adjusts the execution order and distribution of tasks based on task priority, node computing power, and network latency factors, and optimizes the overall computing efficiency and parallelism of task execution;

[0107] S25. Coordinate task execution through the synchronization mechanism between virtual computing nodes. Each virtual computing node regularly exchanges task execution status and intermediate calculation results, and uses the Paxos distributed consistency protocol for global consistency control.

[0108] In this embodiment, S3 specifically includes:

[0109] S31. Deploy and configure the distributed computing framework of TensorFlow on multiple virtual computing nodes C k . The distributed computing framework includes a parameter server and worker computing nodes. The parameter server is responsible for centrally storing and updating global model parameters, and the worker computing nodes are responsible for distributed data processing and model calculation;

[0110] S32. Initialize the global parameters of the variational autoencoder model on the parameter server, including the weight matrix W, the bias vector b, and the latent space distribution parameters μ and σ:

[0111] W (0) , b (0) , μ (0) (X), σ (0) (X);

[0112] Among them, μ (0) (X) and σ (0) (X) are the initial mean and standard deviation functions with respect to the input data X. The parameters distribute the initial global parameters to each worker virtual computing node C through a distributed strategy k ;

[0113] S33. On each worker virtual computing node C k , configure the distributed strategy of TensorFlow. Each worker node receives the global model parameters from the parameter server and uses the local data block D i to perform forward and backward propagation calculations of the local model:

[0114]

[0115] Among them, represents the latent variable on the worker virtual computing node C k , ∈ (k) is a noise variable sampled from the standard normal distribution, and ⊙ represents element-wise multiplication. The backward propagation updates the local model parameters by calculating the reconstruction error and the KL divergence;

[0116] S34. After each worker virtual computing node C k completes the local model training, calculate the gradient of the local model parameters and transfer the gradient information g k back to the parameter server:

[0117]

[0118] Among them, is the reconstruction error, is the KL divergence, indicating the gradient calculation for the parameter weight matrix W, bias vector b, and the potential space distribution parameters μ and σ. The parameter server updates the global model parameters according to the gradient information g returned by all worker nodes k to update the global model parameters;

[0119] S35. The process of the parameter server updating the global model parameters is expressed as:

[0120]

[0121]

[0122] where α 2 is the learning rate, t represents the number of iterations. The parameter server distributes the updated global model parameters to all worker computing nodes, and each worker node uses the latest model parameters to continue the next round of local model training, repeating steps S31 - S35 until the model training converges.

[0123] In this embodiment, S5 specifically includes:

[0124] S51. After the training is completed, each virtual computing node C k uses the trained distributed variational autoencoder model to process the real-time input data stream X i for processing;

[0125] S52. Calculate the reconstruction error value i between the input data X and the reconstructed data

[0126]

[0127] where X i,m and respectively represent the m-th feature dimension of the input data X i and the reconstructed data , M is the total number of feature dimensions, α m is the weight factor for the m-th dimension, used to adjust the relative importance of each feature dimension in the reconstruction error calculation; γ m is the non-linear adjustment factor for the m-th dimension, and λ m is the wavelength parameter related to the m-th dimension, used to introduce a non-linear oscillation component in the reconstruction error to capture potential complex patterns.

[0128] In this embodiment, S6 specifically includes:

[0129] S61. Each virtual computing node will detect the reconstruction error value Compared with the preset anomaly detection threshold τ k After comparison, the data exceeding the threshold is marked as abnormal data, and the marking result r k (X i ) is expressed as:

[0130]

[0131] where r k (X i ) represents the detection result of the input data by the virtual computing node. 1 indicates abnormal, and 0 indicates normal;

[0132] S62. Each virtual computing node sends the detection result r k 9X i ) to the parameter server. The parameter server performs weighted average processing based on the detection results r k 9X i ) received from each node to obtain the global detection result R(X i ):

[0133]

[0134] where w k represents the weight of the virtual computing node, and the weight can be set according to the computing power of the node, the data processing volume, or the credibility of the detection result;

[0135] S63. Compare the global detection result R(X i ) with the global anomaly detection threshold τ global . If R(X i ) > τ global , then mark the data X i as global abnormal data. If R(X i ) ≤ τ global , then mark it as normal data. The global anomaly detection threshold τ global is dynamically adjusted by the following formula:

[0136] τ global = μ R + σ R · λ 3 ;

[0137] where μ R and σ R are the mean and standard deviation of the detection results R(X i ) of all nodes respectively, and λ 3 is the adjustment factor used to control the sensitivity of global anomaly detection;

[0138] S64. The parameter server sends the global anomaly detection result R(Xi ) It is stored in the global storage system and an anomaly detection report is generated, which includes detailed information of the abnormal data, the detection time, and the detection results of each node.

[0139] Example 1:

[0140] A large financial institution relies on a cloud computing platform to process a large amount of transaction data and user behavior data in its global business. The financial institution processes millions of transactions every day and generates more than thousands of log data per second. These data need to be distributed and stored and processed in multiple data centers around the world. To ensure the security of the system and the integrity of the data, the financial institution introduces the method of anomaly detection, storage, computing and training integration in the cloud computing environment based on variational autoencoder of the present invention to monitor the running state of the system in real time, detect and respond to potential cyberattacks and abnormal behaviors.

[0141] In Example 1, the data centers of the financial institution are simultaneously processing millions of financial transactions and user behavior data globally. The peak period of each day is from 12:00 noon to 2:00 pm, with huge data traffic and the highest pressure of anomaly detection faced by the system. One day at noon, the system monitoring module suddenly found an abnormal phenomenon: a host with the source IP address of "192.168.50.23" initiated a large number of database query requests within a short period of time, and the packet length, query frequency, and query content characteristics of its data packets all showed abnormalities.

[0142] According to the method of the present invention, the virtual computing nodes distributed in different data centers immediately process these log data. First, the virtual computing nodes input these log data into the pre-trained variational autoencoder model, and the encoder part of the model maps these data into the latent space. The system notices that the distribution of the latent variables is significantly different from the normal query behavior, and the reconstruction error between the reconstructed data generated by the decoder and the original data far exceeds the normal threshold of the system.

[0143] Specifically, the host with the source IP address of "192.168.50.23" initiated more than 10,000 query requests within just 30 seconds. Under normal circumstances, the average number of query requests within the same time period is 500. The packet lengths of these abnormal requests are all greater than 1024 bytes, and the query content involves sensitive information such as user passwords and account balances. Through the analysis of the latent space mapping of the variational autoencoder, the anomaly score of the host reaches 0.92, while the threshold set by the system is 0.75.

[0144] The system generated a detailed attack report, which included the detailed information of the IP address, the feature analysis of abnormal requests, the reconstruction error value, and the deviation degree of the potential space distribution. The report showed that the abnormal behavior of the IP address was highly similar to the known SQL injection attack pattern, and it was very likely that hackers were trying to attack the database through the IP address. The system immediately triggered the warning mechanism and sent the information to the network security team in real time.

[0145] After receiving the warning signal, the security team immediately took countermeasures. They first blocked the connection to the host "192.168.50.23" to prevent further malicious behavior. Subsequently, by analyzing the attack report generated by the system in detail, it was confirmed that this was a highly concealed SQL injection attack attempt. Based on the abnormal features in the report, the team further strengthened the system's protection measures, repaired possible security vulnerabilities, and reported the detailed information of the attack to a higher-level security monitoring system.

[0146] In this incident, the system completed the whole process from data collection, anomaly detection to warning sending within 20 seconds, greatly shortening the response time and ensuring the data security of financial institutions. Compared with the traditional rule-based anomaly detection method, the method of the present invention not only improves the detection accuracy, but also significantly reduces the false alarm rate. In the past, in similar situations, traditional systems might miss such attacks due to unclear abnormal features, resulting in serious consequences.

[0147] In addition, in another data center, the system detected abnormal request traffic initiated by another group of IP addresses "172.16.34.100". The host accessed multiple user accounts within a short period of time and tried to perform multiple password reset operations. The variational autoencoder model of the system identified the abnormal potential space distribution of these requests, and its reconstruction error value reached 0.88, while the normal reconstruction error value is usually below 0.6. After analysis, the system determined that these requests were consistent with the typical brute-force cracking attack pattern.

[0148] Similarly, the system quickly generated an anomaly report, recording the detailed behavior characteristics, request content, and timeline of the IP address. The security team immediately added the IP address to the blacklist and conducted a detailed trace of it, and finally confirmed that this was a distributed brute-force cracking attempt. The team further improved the system's detection ability for similar attacks by adjusting the threshold of the model and the reconstruction error calculation method.

[0149] In practical applications, financial institutions compared the performance of the method of the present invention with its traditional rule-based anomaly detection system when dealing with the above events. The comparison data is shown in Table 1 below:

[0150] Table 1 Performance data of the method of the present invention and its tradition when dealing with events

[0151]

[0152]

[0153] Through the comparison of the above cases and data, it can be seen that the application of the method of the present invention in the actual cloud computing environment not only effectively solves the deficiencies of traditional anomaly detection methods, especially when dealing with high concurrency and complex data, but also significantly improves the detection efficiency and accuracy, providing a strong guarantee for the network security of financial institutions.

[0154] By closely integrating data storage, computing, and model training, the present invention realizes the integrated design of storage, computing, and training. In the cloud computing environment, through virtualization technology and distributed computing architecture, each computing node can work collaboratively to execute model training and inference tasks under the distributed TensorFlow framework, significantly reducing the overhead of data transmission and task scheduling, avoiding the performance bottleneck caused by the separation of storage and computing in traditional methods, thereby greatly improving the response speed and computing efficiency of the system. By using the distributed parameter server architecture and distributed gradient descent algorithm, the consistency of global model parameters is ensured, further enhancing the stability and efficiency of the system.

[0155] The present invention introduces the VAE model into the anomaly detection task in the distributed cloud computing environment. By modeling the latent space of data through VAE, it can better capture the latent patterns and complex structures of data. The VAE model training and inference processes executed in parallel on multiple virtual computing nodes can effectively handle the anomaly detection tasks of high-dimensional and complex data. At the same time, the latent space distribution of the VAE model is optimized through KL divergence regularization, making the model capture the normal data distribution more accurately, thus having higher accuracy and robustness in detecting abnormal data. Combining global and local reconstruction error metrics further reduces the probability of false positives and false negatives.

[0156] The present invention proposes an improved reconstruction error calculation method. By introducing a non-linear adjustment factor and a dynamic threshold adjustment mechanism, anomaly detection can adapt to the complexity of different feature spaces in the cloud computing environment. In the reconstruction error calculation, a non-linear adjustment factor based on the feature dimension is added, enabling the system to capture complex patterns and abnormal features in the data. In addition, by combining the dynamic threshold adjustment mechanism with the mean and variance of the latent space distribution, the sensitivity of anomaly detection can be dynamically adjusted according to the distribution characteristics of the data, thus maintaining high detection performance in different data scenarios. The present invention is applicable to diverse data and dynamically changing features in the cloud computing environment, improving the adaptability and detection accuracy of the system.

[0157] The present invention adopts a distributed integration strategy. By performing weighted averaging and global consistency processing on the detection results of multiple virtual computing nodes, the reliability and consistency of the final anomaly detection conclusion are ensured. In a cloud computing environment, there may be differences in the computing power and data processing volume of each node. By setting node weights and dynamically adjusting the global threshold, the detection results of each node can be effectively integrated, avoiding the impact of false positives or false negatives of a single node on the overall detection result. This not only improves the accuracy of anomaly detection but also enhances the fault tolerance and reliability of the system in a distributed environment.

[0158] The above is only a preferred specific embodiment of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention, according to the technical solution and inventive concept of the present invention, making equivalent substitutions or changes should be covered within the protection scope of the present invention.

Claims

1. A storage-computation-training method for abnormality detection in cloud computing environment based on variational autoencoder, characterized in that: The steps include: S1. Design and initialize a distributed variational autoencoder model based on TensorFlow. The encoder is responsible for encoding the input data into a probability distribution in the latent space, and the decoder is responsible for decoding the variables in the latent space back to the original data space to form reconstructed data. S2. In a cloud computing environment, a distributed data storage and computing architecture is constructed, and multiple computing nodes use virtualization technology to manage resources and schedule tasks; S3. Deploy and configure the distributed computing framework of TensorFlow on multiple computing nodes, use its distributed strategy to distribute the training tasks and anomaly detection tasks of the variational autoencoder model on multiple computing nodes, configure the parameter server and worker computing nodes, and each computing node uses the distributed strategy of TensorFlow to obtain the pre-processed data set stored locally in parallel, execute the forward propagation and back-propagation process of the distributed variational autoencoder model, and use the distributed gradient descent algorithm to synchronously update the model parameters among each computing node; S4. After the training is completed, each computing node uses the trained distributed variational autoencoder model to perform anomaly detection on the real-time input data. Each computing node maps the input data to the latent space through the encoder and generates reconstructed data through the decoder. S5. Through the distributed integration strategy of TensorFlow, the anomaly detection results of multiple computing nodes are integrated and processed. After each computing node compares the detected reconstruction error with the preset threshold, the data exceeding the threshold is marked as abnormal. The detection results of multiple computing nodes are summarized through weighted average or voting integration strategy to obtain a global anomaly detection conclusion. The S3 specifically includes: S31, on multiple virtual computing nodes C k Deploy and configure the distributed computing framework of TensorFlow on GitHub. The distributed computing framework includes parameter servers and worker computing nodes. The parameter servers are responsible for centralized storage and updating of global model parameters, while the worker computing nodes are responsible for distributed data processing and model calculation. S32, initializing the global parameters of the variational autoencoder model on the parameter server, including the weight matrix W, the bias vector b, and the latent space distribution parameters μ and σ: W (0) ,b (0) ,m (0) (X),s (0) (X); Among them, μ (0) (X) and σ (0) (X) is the initial mean and standard deviation function of the input data X. The parameters are distributed to each worker virtual computing node C through a distributed strategy. k ; S33. In each worker virtual computing node C k On the , configure the distributed strategy of TensorFlow, each worker node receives the global model parameters from the parameter server and uses the locally stored data block D i Perform forward propagation and back propagation calculations for the local model: in, Indicates that in the worker virtual computing node C k The latent variables on (k) is a noise variable sampled from a standard normal distribution, ⊙ represents element-by-element multiplication. Back propagation updates local model parameters by calculating reconstruction error and KL divergence; S34. In each worker virtual computing node C k After completing the local model training, calculate the gradient of the local model parameters and pass the gradient information g k Passed back to the parameter server: in, is the reconstruction error, is the KL divergence, Represents the gradient calculation of the parameter weight matrix W, the bias vector b, and the potential space distribution parameters μ and σ. The parameter server calculates the gradient information g based on the gradient information returned by all worker nodes. k Update global model parameters; S35, the process of the parameter server updating the global model parameters is expressed as: Where α2 is the learning rate, t represents the number of iterations, the parameter server distributes the updated global model parameters to all worker computing nodes, and each worker node uses the latest model parameters to continue the next round of local model training, repeating steps S31-S35 until the model training converges; The S4 specifically includes: S41. After the training is completed, each virtual computing node C k Use the trained distributed variational autoencoder model to train the real-time input data stream X i to process; S42. Calculate input data X i and reconstructing data The reconstruction error between Among them, X i,m and Represents the input data X i and reconstructing data The mth feature dimension of , M is the total number of feature dimensions, α m is the weight factor of the mth dimension, which is used to adjust the relative importance of each feature dimension in the reconstruction error calculation; γ m is the nonlinear adjustment factor of the mth dimension, λ m is the wavelength parameter associated with the mth dimension, which is used to introduce nonlinear oscillation components in the reconstruction error and capture potential complex patterns.

2. According to claim 1, a storage-computation-training integrated method for abnormality detection in a cloud computing environment based on a variational autoencoder is characterized in that: The S1 specifically includes: S11. Using the TensorFlow framework, define the structure of the variational autoencoder model, where the encoder consists of several layers of fully connected neural networks, responsible for mapping the input data X into a probability distribution in the latent space: Z~N(μ(X),σ 2 (X)); Among them, μ(X) and σ(X) are the mean function and standard deviation function, respectively, which represent the distribution characteristics of the input data in the latent space; S12, using a distributed parameter initialization method to initialize the network parameters of each layer in the encoder, the parameters include a weight matrix W and a bias vector b, and by distributing the global parameters W and b to multiple computing nodes, parallel initialization is achieved to meet the high concurrency requirements in the cloud computing environment; S13. In the decoder part, a distributed back-propagation neural network is used to decode the variable Z in the latent space back to the original data space. The decoder structure is symmetrical to the encoder, and finally outputs the reconstructed data: Where D(·) represents the mapping function of the decoder; S14. Define the improved reconstruction error loss function Combine the input data X with the reconstructed data The mean square error between them is denoted as δ, and the synchronization consistency metric δ between the computing nodes in the cloud computing environment is introduced. The synchronization consistency metric is used to ensure that the reconstruction results between the computing nodes in the distributed computing environment are consistent: Among them, N is the number of data samples, ||·|| represents the norm, λ1 is the weight factor, K is the number of computing nodes, δ k represents the deviation between the reconstruction result of the kth computing node and the global synchronization result; S15. Define the improved KL divergence loss function When measuring the difference between the potential distribution Z of the encoder output and the standard normal distribution, a potential space distribution consistency constraint between computing nodes in a distributed training environment is introduced to make the potential distribution of each computing node consistent globally: Where d is the dimension of the latent space, η1 is the weight factor, Δ k represents the deviation between the potential distribution of the kth computational node and the global potential distribution; S16. Use the distributed Adam optimization algorithm to train the parameters of the variational autoencoder model. The optimization goal is to minimize the total loss function: Among them, β1 is a hyperparameter that adjusts the weight between the reconstruction error and the KL divergence.

3. According to claim 1, a storage-computation-training integrated method for abnormality detection in a cloud computing environment based on a variational autoencoder is characterized in that: The S2 specifically includes: S21. In a cloud computing environment, virtualization technology is used to build a distributed data storage and computing architecture, which includes multiple virtual computing nodes C k , k = 1, 2, ..., K, where K is the total number of virtual computing nodes, and the virtual computing nodes are connected through the network to achieve distributed processing of data transmission and computing tasks; S22, in the virtual computing node C k On the network, a distributed storage system is deployed. Each virtual computing node uses the HDFS distributed file system to implement block storage and management of data. Each data block D i , i = 1, 2, ..., M, redundant storage is performed among multiple virtual computing nodes, where M is the total number of data blocks; S23. In the distributed data storage and computing architecture, computing resources are allocated and scheduled through a virtualized resource manager, which dynamically allocates CPU, memory, and storage according to the computing requirements of the task to maximize the computing resource utilization of each virtual computing node, and uses resource isolation technology to prevent resource competition and interference between different tasks; S24. Using a distributed task scheduler to distribute the training tasks and anomaly detection tasks of the variational autoencoder model to each computing node, the distributed task scheduler dynamically adjusts the execution order and distribution of tasks based on task priority, node computing power, and network delay factors to optimize the overall computing efficiency and parallelism of task execution; S25. Coordinate task execution through the synchronization mechanism between virtual computing nodes. Each virtual computing node regularly exchanges task execution status and intermediate calculation results, and uses the Paxos distributed consistency protocol to perform global consistency control.

4. According to claim 1, a storage-computation-training integrated method for abnormality detection in a cloud computing environment based on a variational autoencoder is characterized in that: S5 specifically includes: S51, each virtual computing node detects the reconstruction error value and the preset anomaly detection threshold τ k After comparison, the data exceeding the threshold value will be Mark as abnormal data, marking result r k (X i ) is expressed as: Among them, r k (X i ) represents the detection result of the virtual computing node on the input data, 1 represents abnormality and 0 represents normality; S52, each virtual computing node will detect the result r k (X i ) is sent to the parameter server, and the parameter server receives the detection results of each node r k (X i ) is weighted averaged to obtain the global detection result R(X i ): Among them, w k Indicates the weight of the virtual computing node, which is set according to the node's computing power, data processing volume, or the credibility of the detection result; S53, according to the global detection result R(X i ) and the global anomaly detection threshold τ global For comparison, if R(X i )>τ global , then the data X i Marked as global abnormal data, if R(X i )≤τ global , then it is marked as normal data, and the global anomaly detection threshold τ global Dynamically adjusted by the following formula: t global =μ R +s R ·λ3; Among them, μ R and σ R are the detection results R(X i ), λ3 is the adjustment factor used to control the sensitivity of global anomaly detection; S54, the parameter server sends the global anomaly detection result R(X i ) is stored in the global storage system and an anomaly detection report is generated, which includes detailed information of the anomaly data, detection time and detection results of each node.

Citation Information

Patent Citations

  • Risk value prediction method and device based on neural network

    CN115222536A

  • Machine learning for high quality image processing

    CN115812206A