An Edge Cloud Anomaly Detection Method Based on Trustworthy Long Short-Term Memory Network
By adding new layers to the long and short-term memory network, providing additional information and comparing similar user behaviors, the problem of abnormal detection accuracy caused by insufficient new user data is solved, and the abnormal access detection capability of edge cloud services is improved.
Patent Information
- Application Number
- CN202110500177.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-05-08
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2041-05-08
AI Technical Summary
In an edge cloud environment, due to the lack of behavioral information when new users join, there is insufficient analytical data, which affects the accuracy of long and short-term memory networks in anomaly detection and attack prediction.
Add a new layer to the long and short-term memory network, providing additional information about newly joined users, comparing the behavior of similar users in the same domain with the behavior of newly joined users, and training through supervised learning to identify abnormalities and improve prediction accuracy.
Improve the accuracy of abnormal access detection of edge cloud services, reduce the degree of error caused by insufficient new user data, and is not limited to the user's configuration file status.
Smart Images

Figure CN113138904B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to an edge cloud anomaly detection method based on a reliable long short-term memory network. More specifically, it particularly relates to a method of adding a new layer on top of the long short-term memory network method to provide additional information about newly added users, comparing the interaction data between users and edge cloud servers with the interaction data of other similar users, and analyzing the historical record of the interaction data of this user to detect abnormal access to the edge cloud server in the case of data deficiency. Background Art
[0002] In 5G and more advanced network environments, low-latency services are one of the important requirements for emerging applications. Edge cloud computing processes data closer to local servers and edge data centers rather than the central location of the entire cloud server, thus greatly reducing latency. However, distributing the data of a large number of users across a large network of edge data centers poses huge security vulnerabilities, which may lead to anomalies in the infrastructure of the edge cloud, such as slow server operation or even crashing of the entire cloud computing system. Infrastructure anomalies occurring at a single node in the edge cloud can quickly spread to other edges in the cloud computing system, so it is often difficult to trace the root cause when an anomaly occurs. Usually, the threat of external attacks can be mitigated by passing through a firewall and adopting appropriate authentication to control access. When an attacker impersonates a trusted user or steals the identity of a valid user, measures need to be taken to defend against internal attacks, such as using machine learning (ML) methods to reduce risks. Currently, long-term memory or short-term memory is generally used to classify users as normal users or malicious users. The problem with the former is that it stores irrelevant data, while the problem with the latter is that it may not have enough data for proper analysis. To identify abnormal access to edge cloud services, time series data needs to be collected from the monitored edge cloud system. Currently, long short-term memory (LSTM) is an effective method for identifying abnormal behavior in cloud networks. In the literature "LSTM for anomaly-based network intrusion detection," (S.A. Althubiti, E.M. Jones, K. Roy, in 2018 28th International Telecommunication Networks and Applications Conference (ITNAC), pp: 1–3, 2018), the CIDDS001 dataset was used for experiments, and several ML techniques were adopted to evaluate the system performance. These results show that the accuracy of the LSTM method is higher than that of the support vector machine (SVM) method and the naive Bayes method, and the LSTM method is significantly better than SVM and the naive Bayes method. In recent years, deep learning and recurrent neural network (RNN) methods have been used to prevent threats of internal attacks.In the literature "Deep learning for unsupervised insider threat detection in structured cybersecurity data streams," (A. Tuor, S. Kaplan, B. Hutchinson, in arXiv Prepr, 2017.), it is proposed to use deep neural networks and RNNs to detect insider attacks, train an internal neural network to discover the behavior of functions performed by internal users, so as to classify internal users as normal users or abnormal users in real-time applications. Anomaly can be avoided when there is sufficient user behavior information, but when new users join the network, there is not enough user behavior information, so the accuracy of LSTM prediction will be very low.
[0003] In view of the situation that the lack of user behavior information leads to insufficient data for analysis, the present invention proposes an edge cloud anomaly detection method based on a Trustworthiness Long Short-Term Memory (TLSTM) network, which actively detects abnormal access behaviors and improves the accuracy of anomaly detection and attack prediction. In the method of the present invention, the LSTM method is improved by adding a new layer, and in view of the insufficient data of new users, the new layer provides additional information about the newly added users, compares the behaviors of similar users belonging to the same domain with the behaviors of the newly added users, and trains to obtain available data. The edge cloud anomaly detection method based on a trustworthiness long short-term memory network proposed by the present invention is not limited by the profile status of users, reduces the error degree caused by insufficient data of new users, and improves the accuracy of anomaly access detection of edge cloud services. Summary of the Invention
[0004] In order to overcome the shortcomings of the above conventional long short-term memory network method, the present invention proposes an edge cloud anomaly detection method based on a trustworthiness long short-term memory network.
[0005] An edge cloud anomaly detection method based on a trustworthiness long short-term memory network of the present invention is characterized in that it is realized through the following steps: Step 1, in the edge cloud, add a new layer B t to the state and perform self-training; Step 2, through the sigmoid gate of the LSTM, retain, add and delete units according to the information requirements; Step 3, update the old unit state C t-1 to the new unit state C t , and then determine the output value of the network.
[0006] An edge cloud anomaly detection method based on a trustworthiness long short-term memory network of the present invention, the Step 1 is realized through the following sub-steps:
[0007] a) Add the new layer B t to the state, where this layer provides additional information about the comparison of unit states, and this additional information is used to predict anomalies; The value of B t can be expressed as:
[0008] B t =σ(W B ·[h t-1 ,x t +b B ) (1)
[0009] In formula (1), σ is a neural network layer that outputs a number between 0 and 1, describing how much information can pass through each component. 0 means no information passes through, and 1 means all information passes through. W B is the weight of the added layer, h t-1 is the previous output of the t-th layer, that is, the output of the (t - 1)-th layer, x t is the input of the t-th layer, and b B is the bias of the added layer;
[0010] b) Train each user's data in a supervised learning manner, where the model is trained with prior data, and in the prior data, the model can be dynamically trained according to the behavior of users in the edge network; In this framework, the model uses Personal User Behavior Data (PUBD) and Comparable User Behavior Data (CUBD) to train itself to identify anomalies and improve prediction accuracy in the absence of sufficient data from newly added users.
[0011] An edge cloud anomaly detection method based on a reliable long short-term memory network according to the present invention, the step two is implemented by the following sub-steps:
[0012] c) Determine which information to delete from the unit state; In the LSTM, the first sigmoid gate is the forget gate, and the forget gate layer f t can be expressed as:
[0013] f t =σ(W f ·[h t-1, x t +b f ) (2)
[0014] In formula (2), σ is a neural network layer, W f is the weight of the forget gate, h t-1 is the previous output of the t-th layer, x tis the input of the t-th layer, b f is the bias of the forget gate; the forget gate f t outputs a value between 0 and 1 in the cell state C t-1 (the previous state of the t-th layer), and the closer the value of f t is to 1, the more the cell state is needed; when the value is 1, it means the cell state is most needed and should be maintained; when the value is 0, it means the cell state is completely deleted;
[0015] d) Determine which new information should be retained in the cell state; set the input gate layer i t and the tanh layer. The input gate layer i t determines which values will be updated, and then tanh creates new candidate values that can be added to the cell state vector, and then combines these two components to create an update to the cell state; the input gate layer i t can be expressed as:
[0016] i t = σ(W i · [h t-1 , x t + b i ) (3)
[0017] The new candidate values can be expressed as:
[0018]
[0019] In formulas (3) and (4), W i and W c are the weights of the input gate and the tanh gate respectively, h t-1 is the previous output of the t-th layer, x t is the input of the t-th layer, b i and b c are the biases of the input gate and the tanh gate respectively.
[0020] A method for edge cloud anomaly detection based on a trustworthy long short-term memory network according to the present invention, and the third step is implemented by the following sub-steps:
[0021] e) Update the cell state; in order to update the old cell state C t-1 to the new cell state C t , multiply the old state by f t to represent the part expected to be forgotten, and then add it to to obtain the new cell state; the new cell state C t can be expressed as:
[0022]
[0023] f) Determine the output value; the output will be based on the cell state but will be a filtered version. First, run an S layer to decide which parts of the cell state should be output; the output part of the cell state can be expressed as:
[0024] o t = σ(W o · [h t-1 , x t + b o ) (6)
[0025] Then, the cell state passes through a tanh gate and is multiplied by the output of the sigmoid gate. At this time, the output part updated through the entire system is output, and the output part h t can be expressed as:
[0026] h t = o t * tanh(C i ) (7)
[0027] The beneficial effects of the present invention are as follows: An edge cloud anomaly detection method based on a reliable long short-term memory network proposed in the present invention adds a new layer on top of the long short-term memory network method, providing additional information about newly added users, thereby enhancing the ability to detect anomalies in edge cloud server access in the case of less user behavior data. Using the edge cloud anomaly detection method based on a reliable long short-term memory network of the present invention, the interaction data of a user is compared with the data of other similar users, and the historical record of the user's interaction data is analyzed to actively detect high-precision attacks. This method is not limited by the profile status of the user and reduces the error degree caused by insufficient data of new users, improving the accuracy of edge cloud service anomaly access detection. BRIEF DESCRIPTION OF THE DRAWINGS
[0028] Figure 1 is a schematic diagram of the LSTM framework;
[0029] Figure 2 is a schematic diagram of the TLSTM framework method of the present invention;
[0030] Figure 3 is a graph of the accuracy of anomaly detection and prediction of the method of the present invention on the master node;
[0031] Figure 4 is a graph of the accuracy of anomaly detection and prediction of the LSTM method on the master node;
[0032] Figure 5 is a graph of the accuracy of anomaly detection and prediction of the method of the present invention on the worker node;
[0033] Figure 6 Accuracy graph of the LSTM method for anomaly detection and prediction on the working node;
[0034] Figure 7 Accuracy graph of the method of the present invention for anomaly detection and prediction on the virtual machine CPU;
[0035] Figure 8 Accuracy graph of the LSTM method for anomaly detection and prediction on the virtual machine CPU. Detailed implementation mode
[0036] The present invention will be further described below in conjunction with the accompanying drawings and embodiments.
[0037] Long short-term memory network is a special type of recurrent neural network that can learn sequential dependencies in sequence prediction. The schematic diagram of the LSTM framework is as Figure 1 shown. The LSTM has gate layers that can retain, add, and delete units according to information requirements. They consist of sigmoid neural network layers and pointwise multiplication operations. The sigmoid neural network layer outputs a value between 0 and 1, which describes how much each component is allowed to pass through. A value of 0 means deleting all information, while a value of 1 means retaining all information.
[0038] The first sigmoid gate is the forget gate, which is used to determine which information to delete from the cell state. The forget gate f in formula (1) t in the cell state C t-1 (the previous state of the t layer) outputs a value between 0 and 1. The closer the value of f t is to 1, the more the cell state is needed; when the value is 1, it means the cell state is most needed and should be maintained; when the value is 0, it means the cell state is completely deleted.
[0039] f t =σ(W f ·[h t-1 , x t +b f ) (1)
[0040] The next layer is used to determine which new information should be retained in the cell state. Set the input gate layer i t and the tanh layer. The input gate layer i t determines which values will be updated. Next, tanh creates a new candidate value vector that can be added to the cell state, and then combines these two components to create an update to the cell state; the input gate layer i t can be expressed as:
[0041] i t =σ(Wi · [h t-1 , x t + b i ) (2)
[0042] The new candidate value can be expressed as:
[0043]
[0044] To update the old cell state C t-1 to the new cell state C t , multiply the old state by f t to represent the part expected to be forgotten, and then add it to to obtain the new cell state; the new cell state C t can be expressed as:
[0045]
[0046] Finally, the output value needs to be determined. The output is based on the cell state but will be a filtered version. First, run an S layer to decide which parts of the cell state should be output; the output part of the cell state can be expressed as:
[0047] o t = σ(W o · [h t-1 , x t + b o ) (5)
[0048] Then, the cell state passes through the tanh gate and is multiplied by the output of the sigmoid gate. At this time, the output part updated through the entire system is output, and the output part h t can be expressed as:
[0049] h t = o t * tanh(C i ) (6)
[0050] In the above formulas, W f , W i , W c and W o are the weights of the forget gate, input gate, tanh gate, and output gate respectively. b f , b i , b c and b o are the biases of the forget gate, input gate, tanh gate, and output gate respectively.
[0051] When there is sufficient data, the current LSTM method can predict user behavior patterns relatively accurately. However, the prediction is accurate only when there is a large amount of interaction data. The reliable long short-term memory framework is used to actively detect abnormal behaviors and improve the accuracy of anomaly detection. The proposed method detects anomalies based on the behavior patterns of users. For this purpose, a new layer B t is added to the state, which provides additional information about the comparison of the cell state, and this additional information is used to predict anomalies. The value of B t can be expressed as:
[0052] B t = σ(W B · [h t-1 , x t + b B ) (7)
[0053] This method trains the data of each user in a supervised learning manner, where the model is trained with prior data, and in the prior data, the model can be dynamically trained according to the behaviors of users in the edge network; in this framework, the model uses personal user behavior data (PUBD) and comparable user behavior data (CUBD) to train itself to identify anomalies and improve the prediction accuracy in the case of insufficient data of newly added users. In dynamic behavior monitoring, the learning process is automatically executed. When the user behavior data is not sufficient to attract newly added users, false alarms will be triggered, assuming that malicious activities have occurred. Therefore, the comparison data of user behaviors is used to detect errors in the behaviors of new users.
[0054] To distinguish newly added users from malicious users, the new layer of this method contains comparable user behavior data. This layer compares the behaviors of similar users belonging to the same domain with the behaviors and interaction information of newly added users until the newly added users generate necessary personal behavior data. Figure 2 Figure for the TLSTM framework
[0055] This data test collects data from Raspberry Pi and host / virtual machine. The Raspberry Pi includes seven HP servers and a pair that respectively simulate the edge computing environment and edge devices; the host / virtual machine uses the K8s settings on the host and virtual machine respectively. Prometheus is used for monitoring and data collection, and the data collection intervals for the host / virtual machine and Raspberry Pi are 10 seconds and 30 seconds respectively. The collected data contains information on the CPU utilization of the network and packet loss. In the measured data, the data value is replaced by the difference between two consecutive data samples. In the rated data, the data value is replaced by the difference between the current value and the value of the previous time interval divided by the time interval.
[0056] First, loop anomalies and cumulative anomalies are injected into the network. The anomalies are injected for 60 seconds, and then the system is cooled for 180 seconds. These steps are repeated again until the experiment time is exceeded. For cumulative anomalies, Stress-ng is used to generate cumulative memory pressure, and TCP flooding and Ping flooding are used for network flooding. The cumulative anomalies on the host / virtual machine run in four steps, while those on the Raspberry Pi run in two steps, each step lasting 60 seconds. Next, the collected data is marked according to the anomaly injection timestamps. A data value of 0 for recurrent anomalies indicates no anomaly, and a value of 1 indicates an anomaly. In the cumulative anomaly data, 0 is non-anomaly, and 1 - 4 are four anomaly levels. For the cumulative anomalies of the edge devices (i.e., Raspberry Pi), 0 is non-anomaly, and 1 - 2 are two levels of anomalies. To evaluate the performance of this method, relevant data is collected and compared with the performance of the LSTM network. Figures 3 to 8 Shows the accuracy of anomaly detection for the virtual machine CPU, worker nodes, and master nodes in the LSTM method and this method. In LSTM, the accuracy of anomaly detection is high. However, the accuracy of the prediction results of this method is not very ideal, while the accuracy of both detection and prediction in this method is higher than that of the LSTM method.
[0057] Table 1 Comparison of the accuracy of the LSTM method and this method for the training dataset
[0058]
[0059] Table 2 Comparison of the accuracy of the LSTM method and this method for the test dataset
[0060]
[0061] Table 3 Packet loss situations in the LSTM method and this method for the training dataset
[0062]
[0063] Table 4 Packet loss situations in the LSTM method and this method for the test dataset
[0064]
[0065] Table 1 shows the numerical values of the accuracy of the LSTM for the training dataset and the present method. Table 2 provides the numerical values of the accuracy of the LSTM and the present method for the test dataset. Tables 3 and 4 provide the loss values of the LSTM method and the present method for the training and test datasets. The experimental data show that the accuracy of the present method is better than that of the LSTM method.
[0066] In summary, an edge cloud anomaly detection method based on a trustworthy long short-term memory network of the present invention proposes adding a new layer on top of the long short-term memory network method, providing additional information about newly added users, thereby improving the anomaly detection ability in the case of less user behavior data. Using this method, the transaction data of a user is compared with the data of other similar users, and the transaction history of the user is analyzed to actively detect high-precision attacks. This method is not limited to the profile status of the user, and reduces the error degree caused by insufficient data of new users, improving the accuracy.
[0067] The above technical solution is only one implementation manner of the present invention. For those skilled in the art, based on the disclosed application methods and principles of the present invention, it is very easy to make various types of improvements or deformations, not limited to the methods described in the above specific implementation manners of the present invention. Therefore, the above-described manner is only preferred and does not have a restrictive meaning.
Claims
1. An edge cloud anomaly detection method based on a reliable long short-term memory network, characterized in that, it is implemented through three sub-steps, which are elaborated as follows; Step one is implemented through the following sub-steps: a) Place the new layer B t Added to the state, this layer provides additional information about the contrast of the unit state, which is used to predict anomalies; B t The value of is expressed as: B t = σ(W B · [h t-1 , x t + b B ) (1) In formula (1), σ is a neural network layer that outputs a number between 0 and 1, describing how much information of each component passes through. 0 means no information passes through, and 1 means all information passes through. W B is the weight of the addition layer, h t-1 is the previous output of the t-th layer, that is, the output of the (t - 1)-th layer, x t is the input of the t-th layer, b B is the bias of the addition layer; b) Each user's data is trained in a supervised learning manner, where the model is trained with prior data. In the prior data, the model is dynamically trained according to the behavior of users in the edge network; in this framework, the model uses personal user behavior data and comparable user behavior data for its own training to identify anomalies and improve prediction accuracy in the absence of sufficient data from newly added users; Step two is implemented through the following sub-steps: c) Determine which information to delete from the cell state; in the LSTM, the first sigmoid gate is the forget gate, and the forget gate layer f t is expressed as: f t = σ(W f · [h t-1 , x t + b f ) (2) In formula (2), σ is a neural network layer, and W f is the weight of the forget gate, h t-1 is the previous output of the t-th layer, x t is the input of the t-th layer, and b f is the bias of the forget gate; the forget gate f t outputs a value between 0 and 1 in the cell state C t-1 , where the C t-1 is the previous state of the t-th layer. The closer the value of f t is to 1, the more the cell state is needed; when the value is 1, it indicates that the cell state is most needed and should be maintained; when the value is 0, it indicates that the cell state is completely deleted; d) Determine which new information should be retained in the cell state; set the input gate layer i t and the tanh layer, the input gate layer i t determines which values will be updated, and then tanh creates new candidate values to be added to the cell state of the vector, and then combines these two components to create an update to the cell state; the input gate layer i t is expressed as: i t = σ(W i ·[h t-1 , x t + b i ) (3) New candidate value Expressed as: In formulas (3) and (4), W i and W c are the weights of the input gate and the tanh gate respectively, h t-1 is the previous output of the t-th layer, x t is the input of the t-th layer, b i and b c are the biases of the input gate and the tanh gate respectively; Step three is implemented through the following sub-steps: e) Update the unit state; To update the old unit state C t-1 to the new unit state C t , multiply the old state by f t , to represent the part expected to be forgotten, and then add it to to obtain the new unit state; The new unit state C t is expressed as: f) Determine the output value; The output will be based on the cell state, but will be a filtered version. First, an S layer is run to decide which parts of the cell state should be output; the output part of the cell state is expressed as: o t = σ(W o · [h t-1 , x t + b o ) (6) Then, the unit state passes through the tanh gate and is multiplied by the output of the sigmoid gate, at which point the output part updated through the entire system is output, and the output part h t is denoted as h t = o t *tanh(C i ).
Citation Information
Patent Citations
Edge network load distribution algorithm based on application perception prediction
CN111049903A
Edge cloud fault detection method
CN112286749A