Abnormality detection device and abnormality detection method
The anomaly detection device uses a mixed Bernoulli distribution and adversarial learning to generate pseudo-anomaly information, addressing the challenge of detecting abnormal traffic without extensive historical data, thereby enhancing the detection of temporary traffic increases.
Patent Information
- Application Number
- JP2024094904
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-06-12
- Publication Date
- 2025-12-24
- Estimated Expiration
- 2044-06-12
AI Technical Summary
Conventional methods struggle to detect abnormal traffic patterns without a large amount of past abnormal traffic data, making it difficult to estimate or identify such events effectively.
An anomaly detection device utilizing a mixed probability model, specifically a mixed Bernoulli distribution, for traffic volume analysis, combined with adversarial learning of a generative model to generate pseudo-anomaly information and a classifier to distinguish between true and pseudo-anomaly information, enabling detection of temporary traffic increases without extensive historical data.
Enables effective detection of anomalous traffic patterns, such as burst traffic, by training a generator to produce pseudo-anomaly information that mimics true anomalies, allowing for accurate identification even with limited data availability.
Smart Images

Figure 2025186674000001_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to an anomaly detection device and an anomaly detection method, and more particularly to a technique for detecting anomalies in communication traffic. [Background technology]
[0002] Conventionally, there are known techniques for estimating abnormal traffic such as burst traffic and predicting the amount of traffic flowing through a communication network. For example, Patent Document 1 discloses a system that predicts the maximum value of traffic volume for a predetermined link using a statistical estimation method such as maximum likelihood estimation based on correlation data of multiple past traffic data for different links.
[0003] When estimating abnormal traffic such as burst traffic using statistical estimation such as maximum likelihood estimation, a large amount of data on past abnormal traffic is required. Also, when estimating abnormal traffic using machine learning based on a discriminative model, a large amount of learning data on abnormal traffic is required. Therefore, when it is not possible to collect a large amount of data on past abnormal traffic, it can be difficult to estimate or detect abnormal traffic. [Prior art documents] [Patent documents]
[0004] [Patent Document 1] Japanese Patent Application Laid-Open No. 2015-216585 Summary of the Invention [Problem to be solved by the invention]
[0005] As described above, with conventional techniques, it may be difficult to detect abnormal traffic without using a large amount of past abnormal traffic data.
[0006] The present invention has been made to solve the above-mentioned problems, and has an object to detect abnormal traffic without using a large amount of past abnormal traffic data. [Means for solving the problem]
[0007] In order to solve the above-described problems, the anomaly detection device of the present invention includes: a first setting unit configured to set a distribution of observation data, which is a set of observation values, for a mixed probability model, where the observed value is the presence or absence of a temporary increase in traffic volume in time-series data of traffic volume of a communication network; a second setting unit configured to set parameters of the mixed probability model based on the set distribution of observation data; a learning unit configured to perform adversarial learning of a generative model using the mixed probability model having the set parameters as true anomaly information representing an occurrence pattern of anomalous traffic, which is the temporary increase in traffic volume, and the generative model including: a generator that generates pseudo-anomaly information similar to the true anomaly information; and a classifier that distinguishes between the pseudo-anomaly information generated by the generator and the true anomaly information; a generation unit configured to generate the pseudo-anomaly information using the trained generator constructed by the learning unit; and a detection unit configured to detect the anomalous traffic occurring in the communication network based on the pseudo-anomaly information generated by the generation unit.
[0008] Furthermore, the anomaly detection device according to the present invention may further include a collection unit configured to collect time series data on traffic volume on the communication network, and the detection unit may detect the abnormal traffic when the time series data on traffic volume on the communication network, including the temporary increase in traffic volume indicated by the pseudo-anomaly information, matches the time series data on traffic volume on the communication network collected by the collection unit.
[0009] In addition, in the anomaly detection device according to the present invention, the mixed probability model may be a model that follows a mixed Bernoulli distribution, and the parameters of the mixed probability model may include a mixture ratio that represents the probability that each cluster will generate the observed value, and a probability that the observed value of each cluster indicates that there is a temporary increase in traffic volume.
[0010] In order to solve the above-mentioned problems, the anomaly detection method of the present invention includes: a first setting step of setting a distribution of observation data, which is a set of observation values, for a mixed probability model, whose observation value is the presence or absence of a temporary increase in traffic volume in time-series data of traffic volume of a communication network; a second setting step of setting parameters of the mixed probability model based on the set distribution of observation data; a learning step of performing adversarial learning of a generative model using the mixed probability model having the set parameters as true anomaly information representing an occurrence pattern of anomalous traffic, which is the temporary increase in traffic volume, and the generative model including: a generator that generates pseudo-anomaly information similar to the true anomaly information; and a classifier that distinguishes between the pseudo-anomaly information generated by the generator and the true anomaly information; a generation step of generating the pseudo-anomaly information using the trained generator constructed in the learning step; and a detection step of detecting the anomalous traffic occurring in the communication network based on the pseudo-anomaly information generated in the generation step.
[0011] Furthermore, the anomaly detection method according to the present invention may further include a collection step of collecting time series data of traffic volume on the communication network, and the detection step may detect the abnormal traffic when the time series data of traffic volume on the communication network, including the temporary increase in traffic volume indicated by the pseudo-anomaly information, matches the time series data of traffic volume on the communication network collected in the collection step. [Effects of the Invention]
[0012] According to the present invention, adversarial learning of a generative model is performed using a mixture probability model with set parameters as true anomaly information representing an occurrence pattern of anomalous traffic, which is a temporary increase in traffic volume, as well as a generator that generates pseudo-anomaly information similar to true anomaly information and a classifier that distinguishes between the pseudo-anomaly information generated by the generator and true anomaly information. This makes it possible to detect anomalous traffic without using a large amount of past anomalous traffic data. [Brief explanation of the drawings]
[0013] [Figure 1] FIG. 1 is a block diagram showing the configuration of an abnormality detection system including an abnormality detection device according to an embodiment of the present invention. [Figure 2] FIG. 2 is a diagram for explaining an overview of the anomaly detection system according to this embodiment. [Figure 3] FIG. 3 is a diagram for explaining an overview of the anomaly detection system according to this embodiment. [Figure 4] FIG. 4 is a diagram for explaining the learning unit included in the abnormality detection device according to this embodiment. [Figure 5] FIG. 5 is a diagram for explaining the learning unit included in the abnormality detection device according to this embodiment. [Figure 6] FIG. 6 is a diagram for explaining the learning unit included in the abnormality detection device according to this embodiment. [Figure 7] FIG. 7 is a block diagram showing the hardware configuration of the abnormality detection device according to this embodiment. [Figure 8] FIG. 8 is a flowchart showing the operation of the abnormality detection device according to this embodiment. [Figure 9] FIG. 9 is a flowchart showing the operation of the abnormality detection device according to this embodiment. DETAILED DESCRIPTION OF THE INVENTION
[0014] DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS Preferred embodiments of the present invention will now be described in detail with reference to FIGS.
[0015] [Configuration of anomaly detection system] First, an overview of an abnormality detection system including an abnormality detection device 1 according to an embodiment of the present invention will be described with reference to FIG.
[0016] The anomaly detection system is installed in a mobile communication system such as 4G / LTE, 5G, or 6G, and includes an anomaly detection device 1, a communication terminal 2, a base station 3, and a core network 4. The anomaly detection device 1 and the core network 4 are connected via a network NW such as a WAN or the Internet. The anomaly detection system detects abnormal traffic that occurs when a large number of communication terminals 2 use mobile communication services at the same time, causing a temporary increase in traffic volume.
[0017] The communication terminal 2 is equipped with a SIM and is realized as a mobile communication terminal such as a smartphone, a tablet computer, a laptop computer, etc. In this embodiment, there are multiple communication terminals 2. The communication terminal 2 also includes an IoT terminal having an IP address.
[0018] The base station 3 is composed of a wireless base station compatible with a predetermined communication standard, and relays communication between the communication terminals 2 present in the communication area and the core network 4. The base station 3 is connected to the core network 4 via a network such as a backhaul link.
[0019] The core network 4 manages and controls communication data and control signals, and connects the communication terminal 2 and the base station 3 to external networks such as the Internet. The core network 4 supports, for example, the 5G communication standard, and includes function nodes such as an Access and Mobility Management Function (AMF), a Unified Data Management (UDM), and a Unified Data Repository (UDR) in the control plane (C-plane).
[0020] The AMF provided in the core network 4 is a control plane function node that manages user terminal registration and wireless connections. The UDM is a function node that manages subscriber information in the core network and manages the mobility of user terminals. The UDR stores a subscriber profile that holds the IMSI and location information of the communication terminal 2.
[0021] The core network 4 includes functional nodes such as a User Plane Function (UPF) and a Packet Data Network Gateway (PGW) in the user plane (U-plane). The UPF manages traffic such as packet forwarding and routing in data communications. The PGW routes data traffic from the communication terminal 2 to external networks.
[0022] The core network 4 records traffic data on the communication network. For example, the AMF and UPF included in the core network 4 can record time-series data on the amount of traffic flowing through the communication network at predetermined time intervals. Alternatively, the core network 4 can include a network monitor device (not shown), which can record time-series data on the amount of traffic obtained by captured packets or flow-based analysis.
[0023] A communication terminal 2 transmits request packets such as a location registration request, data communication request, and session establishment request to the core network 4 via a base station 3, and receives response packets to these requests from the core network 4. In this embodiment, a case will be described in which abnormal traffic, which is a temporary increase in traffic volume, is detected from request packets transmitted by the communication terminal 2 to the core network 4. Furthermore, abnormal traffic is burst traffic in which the traffic volume increases sharply in a short period of time.
[0024] The anomaly detection system according to this embodiment, having the above-described configuration, employs a mixed probability model in which burst traffic occurrence patterns are modeled using a mixed Bernoulli distribution. The anomaly detection system uses the mixed probability model, whose parameters are set to reflect patterns of temporary increases in traffic volume that occur at set times and frequencies in time-series data of traffic volume, as true data for adversarial learning, to train a generator 121 that generates pseudo data similar to the true data. Furthermore, the anomaly detection system maintains the pseudo data generated using the trained generator 121' as a burst traffic database, and detects burst traffic by comparing it with collected actual traffic data.
[0025] [Function block of the anomaly detection device] Next, functional blocks of the anomaly detection device 1 according to this embodiment will be described with reference to the block diagram of Fig. 1. As shown in Fig. 1, the anomaly detection device 1 includes a first setting unit 10, a second setting unit 11, a learning unit 12, a first storage unit 13, a generation unit 14, a second storage unit 15, a collection unit 16, a detection unit 17, and a presentation unit 18.
[0026] The first setting unit 10 sets a distribution of observed data, which is a set of observed values, for a mixed probability model in which the observed values indicate whether or not there is a temporary increase in traffic volume in time-series data of traffic volume in the communication network. In this embodiment, a model following a mixed Bernoulli distribution is adopted as the mixed probability model. The observed value is a value that represents, as binary data, whether or not the traffic volume due to a request packet transmitted from the communication terminal 2 exceeds the set traffic volume. The set traffic volume can be a reference value such as the maximum traffic volume that the communication network can tolerate.
[0027] The mixed Bernoulli distribution employed in this embodiment is a model for clustering a set of N pieces of observed data formed by vectors of length D whose elements are 0 and 1. In this embodiment, the time periods on the time axis when traffic volume in a communication network exceeds a set value and the occurrence patterns of burst traffic represented by the number and frequency thereof are considered as multiple clusters, and the observed data is a set of observed values that express with 0 or 1 whether the traffic volume exceeds the set value in each of the multiple time periods on the time axis. The observed values are generated from multiple different clusters having a binary distribution.
[0028] Here, we have N binary vector observations x1,x2,…,x N The set of observed data X is given by the parameters μ1,μ2,…,μ of K clusters. K The set of observed data X follows a mixed Bernoulli distribution with a parameter data set M. That is, it is assumed that the observed data X is generated from multiple clusters, and each cluster is associated with a parameter μ that represents the probability that the attempt will be successful, i.e., the probability that the set traffic volume will be exceeded.
[0029] The observation data X and the parameter data set M are expressed as the following equations (1) and (2), respectively (T is a transpose symbol).
number
[0030] nth binary vector observation x n and the kth parameter vector μ k Each of these is further expressed as a vector having D elements, as shown in the following equations (3) and (4).
number
[0031] Observation x n element x of nAs described above, each of (i) (i=1, 2, ..., D) is a binary variable that takes the value of 0 or 1. In this embodiment, the element x n (i) is defined as follows:
number
[0032] The probability that the traffic volume in each time slot will exceed the set value and the probability that the traffic volume in each time slot will not exceed the set value are expressed by the following equations (6) and (7), respectively.
number
[0033] From the above equations (6) and (7), the observed value x of the observed data X n When cluster k is assigned to observation x n is the parameter μ of cluster k k It follows the Bernoulli distribution of the following equation (8):
number
[0034] Here, the parameter data set π of the occurrence frequency of each cluster (K clusters) is defined as follows: The parameter π is the observed value x n is a parameter that indicates how much of the data is distributed to each cluster. k is the observed value x of the observed data X n indicates the mixture ratio occurring in cluster k.
number
[0035] Also, the observed value x n When cluster k is assigned to observation x n is μ k It follows D Bernoulli distributions with parameters. Furthermore, each observation x nare independently generated by the mixed Bernoulli distribution of the following equation (10).
number
[0036] The parameter μ, which indicates the probability of the Bernoulli distribution in each cluster, is defined by the definition of the mixed Bernoulli distribution in equation (10). k and the cluster mixing ratio π k and from all clusters, observation x n The above equation (10) is a mixed probability model in this embodiment, and represents the probability distribution that the traffic volume of the communication network exceeds a set value in each time period.
[0037] 2(a) and 2(b) and 3(a) and 3(b) show the distribution of the observed data X set by the first setting unit 10 and the second setting unit 11, and the parameter μ k , π k 2(b) and 3(b) show the time periods when the traffic volume exceeds the set value and the frequency of that time period, i.e., the overall occurrence pattern of burst traffic. There are D clusters (K=D), which corresponds to the number of time slots, which are the D time periods. The time slot IDs in the vertical direction between clusters correspond to each other. The observed values of "1 to D" that make up each cluster are the observed values x that indicate whether the traffic volume of the communication network exceeds the set traffic volume in each of the 1 to D time periods. n In the example of Fig. 2(b), for each data item 1 to D in each cluster, a time period in which the traffic volume exceeds a set value (black dots in the figure) and a time period in which the traffic volume does not exceed the set value (data other than the black dots in the figure) are set.
[0038] In the example of FIG. 2(b), the first setting unit 10 sets one observation value of "exceeds the set traffic volume: 1" for each cluster, and the other observation values are "does not exceed the set traffic volume: 0." Therefore, the observation value of time slot "1" in cluster 1, the observation value of time slot "2" in cluster 2, and the observation value of time slot "D" in cluster D are "exceeds the set traffic volume: 1." Furthermore, the first setting unit 10 can set a probability (e.g., 90%) of "exceeding the set traffic volume: 1." In this way, the first setting unit 10 can set in each cluster which time slot a burst traffic exceeding the set traffic volume will occur.
[0039] The second setting unit 11 sets parameters of the mixture probability model based on the distribution of the observation data X set by the first setting unit 10. The second setting unit 11 sets parameters μ of the mixture probability model based on the time periods in which "the traffic volume exceeds the set value: 1" occurs in each cluster set by the first setting unit 10 and the number of such periods. k , π k You can set the value of
[0040] The second setting unit 11 also sets the observed value x of each cluster by the first setting unit 10. n If you set the probability of occurrence of x, the observed value of each cluster will be n The mixture ratio is set taking into account the probability of occurrence of k , π k The value of can be adjusted and determined.
[0041] For example, as shown in FIG. 2(b), the second setting unit 11 sets the observed value x of each cluster by the first setting unit 10. n Based on the setting of the above, the proportions of clusters 1 to D in the observation data X can be set to be equal. In order to equalize the probability that each of clusters 1 to D is selected, the second setting unit 11 sets the parameters π1 to π D The value of can be set to 1 / D. Furthermore, at this time, the second setting unit 11 sets a parameter π kand the observed value x n The parameter μ that represents the probability of occurrence of k These can be mutually coordinated and determined.
[0042] In this way, the second setting unit 11 sets each parameter μ based on the occurrence probability of an observed value in which the traffic volume exceeds a set value in each time period, which is set for each cluster by the first setting unit 10. k , π k By adjusting the value of x, we can extract the x from a specific cluster in the mixture probability model of equation (10). n The probability of this occurring can be further adjusted.
[0043] Figure 2(a) shows time-series data of traffic volume that represents the burst traffic generation pattern of clusters 1 to D in Figure 2(b). The vertical axis represents traffic volume [bps], and the horizontal axis represents time [s]. The traffic volume exceeds the set traffic volume TH throughout all time slots from "1" to "D." In other words, the traffic volume time-series data in Figure 2(a) represents a burst traffic generation pattern that corresponds to the traffic volume of one time slot in each cluster being set to exceed the set value TH, and furthermore, each cluster being set to be selected equally.
[0044] 3(b), the first setting unit 10 sets observed values that exceed the set traffic volume TH in time slot "2" of cluster 2, time slot "4" of cluster 4, and time slot "10" of cluster 10 among clusters 1 to D. In addition, the second setting unit 11 sets π2=π4=π to equalize the probability that clusters 2, 4, and 10 are selected among the parameters of clusters 1 to D. 10 = 1 / 3. The parameter π for clusters other than clusters 2, 4, and 10 k is set to 0.
[0045] By setting up such clusters, as shown in Figure 3(a), a burst traffic generation pattern is obtained in which the traffic volume exceeds the set TH during time slots "2," "4," and "10" indicated by dotted lines, but does not exceed the set TH during other time slots.
[0046] The first setting unit 10 and the second setting unit 11 select the observed value x from each cluster. n The mixed probability model in which the probability of occurrence and the probability of selection of each cluster are set in advance is used as true data during adversarial learning by the learning unit 12 described below.
[0047] The learning unit 12 learns the parameter μ set by the second setting unit 11. k , π k The adversarial learning of a generative model is performed using a mixture probability model having the following: a generator 121 that generates pseudo-anomaly information similar to true anomaly information, and a classifier 122 that distinguishes between the pseudo-anomaly information generated by the generator 121 and true anomaly information, with the mixture probability model having the following as true anomaly information representing abnormal traffic, which is a temporary increase in traffic volume; and a classifier 122 that distinguishes between the pseudo-anomaly information generated by the generator 121 and true anomaly information. The true anomaly information is information that reflects a predetermined occurrence pattern of burst traffic. The learning unit 12 can construct a trained generator 121' (trained generators 121'_1, . . . , 121'_K) for each of 1 to K clusters of the mixture probability model. In this embodiment, D clusters are provided, and therefore trained generators 121'_1, . . . , 121'_D are constructed.
[0048] As shown in FIG. 4, the learning unit 12 adversarially trains a GAN (Generative Adversarial Network) having a generator 121 and a classifier 122. In this embodiment, it is possible to provide as many pairs of generators 121 and classifiers 122 as the number of clusters (K pairs). Alternatively, as shown in FIG. 4, it is possible to provide one classifier 122 for D generators 121_1, . . . , 121_D. Through learning by the learning unit 12, D trained generators 121′ (trained generators 121′_1, . . . , 121′_D) are constructed.
[0049] 5 and 6 are diagrams schematically illustrating the neural network configuration of the generator 121 and the classifier 122 of the GAN used by the learning unit 12. As shown in FIG. 5, the generator 121 is configured as a neural network having an input layer, a hidden layer, and an output layer. Note that the K generators 121_1, . . . , 121_D each have the same network structure, and will be collectively referred to as the generator 121 below. The generator 121 is a model that generates pseudo-anomaly information from random noise. For example, m Gaussian noise vectors are randomly sampled and input to the input node of the generator 121 (z1 to z m ).
[0050] The generator 121 performs a product-sum operation on the input and weight parameters and performs threshold processing using an activation function to output an output G(z). The output G(z) from the generator 121 is calculated based on the parameter μ set by the second setting unit 11. k , π k The data is similar to the observed data X obtained by a mixture probability model having the set parameter μ k , π k The observed values x of each cluster obtained by a mixture probability model with n Each output G(z1) to G(z D ) is the observation x of each cluster n As the neural network that configures the generator 121, a CNN or a ResNet can be used.
[0051] 6 is configured as a neural network having an input layer, a hidden layer, and an output layer. In the example of FIG. 6, the parameter μ k , π k The observed values x of each cluster are obtained by a mixture probability model with n is given.
[0052] The classifier 122 performs a product-sum operation on the input and weight parameters and threshold processing using an activation function, and outputs a binary output of 1 or 0. When the classifier 122 correctly identifies the input training data related to true abnormal information as true abnormal information, it outputs an output y=1. On the other hand, when the classifier 122 correctly identifies the input training data related to pseudo abnormal information as pseudo abnormal information, it outputs an output y=0. In this way, the classifier 122 is a model that distinguishes the model distribution generated by the generator 121 from the data distribution of the training data, which is the true distribution. A CNN can be used as the neural network that constitutes the classifier 122.
[0053] FIG. 4 is a block diagram for explaining the adversarial learning of GAN by the learning unit 12. The generator 121 of the GAN adopted by the learning unit 12 is represented as a function G, and the classifier 122 is represented as a function D. Furthermore, true abnormality information is represented as x, the predicted value output by the classifier 122 is represented as y, and the correct label is represented as t. The correct label t is set to 1 for true abnormality information and 0 for pseudo abnormality information generated by the generator 121. In this case, the classifier 122 calculates the cross entropy E CE It can be expressed as:
[0054]
number
[0055] The first term in the brace of the above equation (11) represents t n lny n In this case, the predicted value y n is the correct label of the true anomaly information, t n = 1. On the other hand, the second term in the braces represents (1-t n )ln(1-y n ), the predicted value y n is the value of the correct label (1-t n ) = 0. In this way, the cross entropy E CE is the maximum value when the predicted value matches the correct label value.
[0056] Here, the generator 121 that constitutes the GAN has parameters w G ,θ G and the function G(w G ,θ G ) The classifier 122 also uses the parameter w D ,θ D and function D(w D ,θ D ) The cross entropy E in the above equation (11) CE The objective function E of the GAN including the generator 121 and the discriminator 122 based on the above can be expressed by the following equation (12).
number
[0057] The first term of the above equation (12) represents E D(x)=1 lnD(w D ,θ D ) is the expected value at which the classifier 122 classifies true abnormal information as true abnormal information. D(x)=0 ln(1-D(G(w G ,θ G ),w D ,θ D )) is the expected value at which the classifier 122 classifies the pseudo-anomaly information generated by the generator 121 as pseudo-anomaly information. In GAN learning, the generator 121 and the classifier 122 are trained adversarially through min-max optimization of the objective function E. Therefore, the generator 121 is trained to be able to generate pseudo-anomaly information that deceives the classifier 122, and the classifier 122 is trained to classify the pseudo-anomaly information generated by the generator 121 as pseudo-anomaly information.
[0058] In learning of the classifier 122, when true abnormality information is given, the classifier 122 outputs an output close to y=1, thereby maximizing the first term of the objective function E in the above equation (12). On the other hand, when pseudo abnormality information is given, the classifier 122 learns to output an output close to y=0, thereby maximizing the second term of the objective function E.
[0059] In the learning of the generator 121, D(G(w G ,θ G ),w D ,θ D ) (D(G(z)) in Figure 4) is close to 1. G ,θ G ) (G(z) in FIG. 4 ) to minimize the objective function E. The learning unit 12 uses a learning procedure that alternately updates the parameters of the generator 121 and the classifier 122. Details of the learning procedure of the generator 121 and the classifier 122 by the learning unit 12 will be described later.
[0060] The first storage unit 13 stores a trained generator 121′ in which the objective function E of the GAN has been optimized by the learning unit 12. More specifically, the first storage unit 13 stores D trained generators 121′ (trained generators 121′_1, . . . , 121′_D) corresponding to clusters 1 to D. The first storage unit 13 also stores a parameter μ k , π k The mixed probability model of the above equation (10) is stored.
[0061] The generation unit 14 generates pseudo-anomaly information using the trained generator 121′ constructed by the learning unit 12. More specifically, the generation unit 14 performs calculations on each of D trained generators 121′_1, . . . , 121′_D corresponding to clusters 1 to K to generate pseudo-anomaly information.
[0062] The second storage unit 15 stores the pseudo-anomaly information generated by the generation unit 14. The pseudo-anomaly information stored in the second storage unit 15 is a database of occurrence patterns of burst traffic.
[0063] The collection unit 16 collects time-series data of traffic volume on the communication network. For example, the collection unit 16 can collect the history of traffic volume recorded in the core network 4 via the network NW.
[0064] The detection unit 17 detects burst traffic occurring in the communication network based on the pseudo-anomaly information generated by the generation unit 14. The detection unit 17 can detect burst traffic when the time series data of the traffic volume of the communication network, including a temporary increase in traffic volume, indicated by the pseudo-anomaly information generated by the generation unit 14 matches the time series data of the traffic volume of the communication network collected by the collection unit 16.
[0065] More specifically, the detection unit 17 converts the time-series data of the traffic volume collected by the collection unit 16 into a time-series of binary data corresponding to the above formula (5). Specifically, the detection unit 17 converts the collected time-series data of the traffic volume into a binary representation of "exceeds the set traffic volume: 1" and "does not exceed the set traffic volume: 0" based on the set traffic volume value.
[0066] The detection unit 17 can detect burst traffic when the time series of binary data according to the above formula (5), which is based on the probability that the traffic volume in each time period indicated by the pseudo-anomaly information will exceed a set value, matches the time series of binary data of the collected traffic volume time-series data. Furthermore, the detection unit 17 can sequentially compare the time series of binary data of the collected traffic volume time-series data with the time series of binary data of the traffic volume indicated by the pseudo-anomaly information, and check for a match starting from the earliest time period. For example, as shown in FIG. 3(a), the occurrence of burst traffic can be predicted in advance when the time series of binary data of the traffic volume time-series data matches in a time period before the peak value indicating the occurrence of burst traffic.
[0067] The presentation unit 18 presents the detection result of abnormal traffic by the detection unit 17. For example, the presentation unit 18 can notify an external traffic management server via the network NW of a time period during which burst traffic is predicted to occur.
[0068] [Hardware configuration of the anomaly detection device] Next, an example of a hardware configuration for realizing the abnormality detection device 1 having the above-described functions will be described with reference to FIG.
[0069] 7, the abnormality detection device 1 can be realized by, for example, a computer including a processor 102, a main memory device 103, a communication interface 104, an auxiliary memory device 105, and an input / output (I / O) 106 connected via a bus 101, and a program that controls these hardware resources. Furthermore, the abnormality detection device 1 includes a display device 107.
[0070] The processor 102 is realized by a CPU, a GPU, an FPGA, an ASIC, or the like.
[0071] The main memory device 103 pre-stores programs for the processor 102 to perform various controls and calculations. The processor 102 and the main memory device 103 implement the functions of the anomaly detection device 1, such as the first setting unit 10, the second setting unit 11, the learning unit 12, the generation unit 14, the collection unit 16, the detection unit 17, and the presentation unit 18 shown in FIG.
[0072] The communication interface 104 is an interface circuit for connecting the abnormality detection device 1 to various external electronic devices via a network.
[0073] The auxiliary storage device 105 is composed of a readable / writable storage medium and a drive for reading and writing various information such as programs and data from and to the storage medium. The auxiliary storage device 105 can use a semiconductor memory such as a hard disk or flash memory as the storage medium.
[0074] The auxiliary storage device 105 has a program storage area for storing an anomaly detection program for detecting anomalous traffic. The auxiliary storage device 105 also has a program storage area for storing a mixed probability model calculation program for calculating a mixed probability model related to a mixed Bernoulli distribution. The auxiliary storage device 105 also has a program storage area for storing a learning program for performing GAN adversarial learning executed by the anomaly detection device 1. The auxiliary storage device 105 realizes the first storage unit 13 and the second storage unit 15 described in FIG. 1. Furthermore, for example, the auxiliary storage device 105 may have a backup area for backing up the above-mentioned data, programs, etc.
[0075] The input / output I / O 106 is an input / output device that inputs signals from external devices and outputs signals to external devices.
[0076] The display device 107 is configured by an organic EL display, a liquid crystal display, etc. The display device 107 can realize the presentation unit 18.
[0077] [Operation of the anomaly detection device] Next, the operation of the abnormality detection device 1 having the above-described configuration will be described with reference to the flowcharts of FIGS.
[0078] 8, first, the first setting unit 10 sets the distribution of the observed data X in the mixed probability model (step S1). Specifically, the first setting unit 10 sets the distribution of the observed value x for each cluster so that the set occurrence pattern of burst traffic is reflected. n As explained in Fig. 2 and Fig. 3, the number of clusters is the number of observations, that is, the number of time slots (K=D).
[0079] Next, the second setting unit 11 adjusts the parameter μ of the mixture probability model so as to obtain the distribution of the observation data X set in step S1. k , π kSpecifically, the second setting unit 11 sets the observed value x set for each cluster in the mixed probability model of the above formula (10) (step S2). n Given the observed data X with k , π k For example, the second setting unit 11 adjusts and determines the observed value x for each cluster set in step S1. n Considering the probability of occurrence of k and set the parameter μ k , π k The values of the parameters μ can be adjusted and determined. k , π k The mixed probability model in which the above is set is stored in the first storage unit 13.
[0080] Next, the learning unit 12 performs a learning process (step S3). Specifically, the learning unit 12 performs a learning process on the parameter μ k , π k As true anomaly information indicating a burst traffic occurrence pattern, a mixed probability model having the above formula is used, and adversarial learning of a GAN having a generator 121 that generates pseudo-anomaly information similar to the true anomaly information and a classifier 122 that distinguishes between the pseudo-anomaly information generated by the generator 121 and the true anomaly information is performed. Through the learning process, D trained generators 121'_1, . . . , 121'_D corresponding to the clusters 1 to D respectively are constructed and stored in the first storage unit 13. Details of the learning process in step S3 will be described later.
[0081] Next, the generation unit 14 generates pseudo-anomaly information using the trained generator 121' constructed by the learning unit 12 in step S3 (step S4). More specifically, in step S4, the generation unit 14 performs calculations on each of the D trained generators 121'_1, . . . , 121'_D corresponding to clusters 1 to D to generate pseudo-anomaly information. Thereafter, the second storage unit 15 stores the pseudo-anomaly information generated in step S4 (step S5).
[0082] Next, the collection unit 16 collects time-series data of traffic volume on the communication network (step S6). For example, the collection unit 16 can collect the history of traffic volume recorded in the core network 4 via the network NW. Subsequently, the detection unit 17 converts the time-series data of traffic volume collected in step S6 into time-series binary data corresponding to the above equation (5) (step S7). In step S7, based on the value of the traffic volume set as the reference value, the time-series data of traffic volume is converted into binary data of "exceeds the set traffic volume: 1" and "does not exceed the set traffic volume: 0".
[0083] Next, the detection unit 17 detects abnormal traffic occurring in the communication network based on the pseudo-anomaly information stored in the second storage unit 15 in step S5 (step S8). The detection unit 17 can detect abnormal traffic when the time series of binary data corresponding to the time series data of traffic volume indicating the occurrence pattern of burst traffic, indicated by the pseudo-anomaly information generated by the generation unit 14, matches the time series of binary data of traffic volume collected by the collection unit 16 in step S6 and further converted in step S7. In step S8, the detection unit 17 can predict future burst traffic when the time series of binary data corresponding to the time series data of traffic volume matches in a time period before the occurrence of a burst of traffic volume indicated by the pseudo-anomaly information.
[0084] Thereafter, in step S7, the presenting unit 18 presents the result of the abnormal traffic detection by the detecting unit 17 (step S9). For example, in step S8, the presenting unit 18 can notify an external traffic management server or the like of a time period during which burst traffic is predicted to occur via the network NW.
[0085] Next, the learning process (step S3) by the learning unit 12 of the anomaly detection device 1 described in FIG. 8 will be described with reference to FIG. 9. In the example of the learning process shown in FIG. 9, the learning unit 12 repeatedly learns true abnormality information for each of 1 to D clusters. First, the learning unit 12 inputs the true abnormality information to the classifier 122 as training data 124, and adjusts the parameter w of the classifier 122 so that the true abnormality information is distinguished from the true abnormality information (y=1). D ,θ D is learned and updated (step S20).
[0086] As shown in the block diagram of the learning unit 12 in FIG. 4, the parameter μ set in step S2 in FIG. k , π k The mixture probability model having the following is true anomaly information that indicates the occurrence pattern of burst traffic, and is used as training data 124 when training the classifier 122. In step S20, first, among clusters 1 to D, the observed values x n is used as the training data 124.
[0087] In step S20, the learning unit 12 can cause the classifier 122 to learn true abnormality information using, for example, backpropagation. By step S20, the classifier 122 that can distinguish true abnormality information in cluster 1 from true abnormality information is constructed in advance.
[0088] Next, the learning unit 12 generates Gaussian noise and provides a random vector of the generated Gaussian noise as an input to the generator 121 (step S21). Subsequently, the generator 121 generates a random vector of the input z and the weight parameter w based on the provided Gaussian noise. G ,θ G and threshold processing using an activation function to generate pseudo abnormality information G(z) (step S22). In step S22, as shown in Fig. 4, the learning unit 12 combines 1 to D outputs generated by the D generators 121_1, ..., 121_D (point a shown in Fig. 4) to generate one pseudo abnormality information G(z).
[0089] Next, the learning unit 12 performs learning of the classifier 122. The learning of the classifier 122 is performed by using the parameter w D ,θ D First, the learning unit 12 fixes the parameter μ set in step S2 of FIG. k , π k The true anomaly information is input to the classifier 122 as training data 124. Then, the learning unit 12 adjusts the parameter w by backpropagation or the like so that the objective function E in the above equation (12) is maximized. D ,θ D (Step S23). The label of the training data 124 is set to 1 (true abnormal information). In Step S23, first, the learning unit 12 updates the observed values x n The classifier 122 is trained using the true abnormality information.
[0090] Next, the learning unit 12 provides the pseudo-abnormal information generated by the generator 121 in step S22 to the discriminator 122 as an input, and calculates the parameter w by backpropagation or the like so that the objective function E in the above equation (12) is maximized. D ,θ D (Step S24). That is, in steps S23 and S24, in order to maximize the objective function E in the above equation (12), the first term is updated as D(w D ,θ D )=1 is output, and the second term is D(G(w G ,θ G ),w D ,θ D )=0. Note that the label 0 (pseudo abnormal information) is set in the training data 124. In step S24, the classifier 122 is trained using one pseudo abnormal information G(z) generated by each of the D generators 121_1, . . . , 121_D and combined.
[0091] The learning of the classifier 122 in steps S23 and S24 corresponds to the dashed arrows in the block diagram of the learning unit 12 shown in FIG. 4 , which indicate that a classifier error is calculated in block 125 of the objective function E based on output 123 from the classifier 122, and then the error is back-propagated to the classifier 122.
[0092] Next, the learning unit 12 learns the generators 121 (generators 121_1, . . . , 121_D). The learning of the generators 121 is performed with the parameters of the classifier 122 fixed. The learning unit 12 learns the generator 121 so that pseudo-anomaly information is generated when random Gaussian noise is given to the generator 121. Specifically, the learning unit 12 learns the parameters w G ,θ G is updated (step S25).
[0093] The learning in step S25 corresponds to the flow indicated by the dashed arrow indicating backpropagation of error to the generator 121 in the block diagram of the learning unit 12 in Fig. 4. That is, step S25 corresponds to the flow indicated by the dashed arrow in Fig. 4 in which pseudo anomaly information generated by the generator 121 is input to the discriminator 122, a generator error is calculated from the output 123 thereof in the block 125 of the objective function E, and the error is further backpropagated to the generator 121.
[0094] Thereafter, learning of the discriminator 122 and the generator 121 (generators 121_1 to 121_D) from step S22 to step S25 is repeated until the value of the objective function E reaches a Nash equilibrium and converges (step S26: NO). On the other hand, if the value of the objective function E converges (step S26: YES), the processing from step S20 to step S26 is repeated using the remaining true abnormality information from cluster 2 to cluster D, out of all the true abnormality information from cluster 1 to cluster D, until learning of D generators 121_1, . . . , 121_D and the discriminator 122 is performed (step S27: NO).
[0095] Thereafter, when the generator 121 (generator 121_1 to generator 121_D) and the discriminator 122 are trained using the true anomaly information of the remaining D-1 clusters from cluster 2 to cluster D (step S27: YES), the training unit 12 stores 1 to D trained generators 121'_1, ..., 121'_D in the first storage unit 13 (step S28). The trained generator 121' is constructed by the above-described processes from step S20 to step S28. After that, the process proceeds to step S4 in FIG. 8.
[0096] As described above, according to the anomaly detection device 1 of this embodiment, the occurrence of burst traffic is modeled using a mixed Bernoulli distribution, and the parameter μ is adjusted so as to obtain a distribution of the observed data X that represents the occurrence pattern of pseudo-generated burst traffic. k , π k The mixed probability model with the above setting is used as the true data in the adversarial learning of the GAN. Furthermore, the pseudo-anomaly information generated by the trained generator 121' constructed by the adversarial learning is compiled into a database as burst traffic occurrence patterns, and burst traffic is detected when the collected time-series data of actual traffic volume matches the burst traffic occurrence pattern in the database. Therefore, burst traffic can be detected without using a large amount of past burst traffic data.
[0097] Furthermore, according to the anomaly detection device 1 of this embodiment, the parameter μ k , π k The mixed probability model in which the above is set is used as true abnormal information, which is true data in the adversarial learning of the GAN, and a generator 121 is trained to generate pseudo abnormal information similar to the true abnormal information, i.e., pseudo abnormal information that does not necessarily match the true abnormal information. Therefore, by performing the adversarial learning of the GAN, it is possible to learn the latent variables of the mixed probability model that follows the mixed Bernoulli distribution.
[0098] Furthermore, according to the anomaly detection device 1 of this embodiment, whether or not the traffic volume exceeds a set value is used as the observed value for each cluster of the mixed Bernoulli distribution, so it is possible to specify for each cluster the frequency and pattern of a sudden increase in traffic volume in a short period of time, which is seen in the occurrence of burst traffic. This makes it possible to set the occurrence pattern of burst traffic in detail and to detect burst traffic more effectively.
[0099] Furthermore, according to the anomaly detection device 1 of this embodiment, the time series data of the traffic volume indicated by the pseudo-anomaly information generated by the learned generator 121′ is compared with the history of the actual traffic volume collected, thereby making it possible to predict burst traffic that will occur in the future.
[0100] In the embodiment described above, the first setting unit 10 sets a time period in which 1 to D observed values of each cluster are "exceeding the set traffic volume: 1" or "not exceeding the set traffic volume: 0", and further sets the occurrence probability of each observed value. Also, the second setting unit 11 sets the mixing ratio parameter π k Furthermore, the parameter μ k and the mixing ratio parameter π k and adjust the parameter μ k , π k However, the first setting unit 10 sets specific binary values for 1 to D observation values of each cluster in the observation data X, and the second setting unit 11 sets the occurrence probability of each observation value, and the parameter μ k , π k may be configured to adjust each other.
[0101] The above describes embodiments of the anomaly detection device and anomaly detection method of the present invention, but the present invention is not limited to the described embodiments, and various modifications that a person skilled in the art can conceive are possible within the scope of the invention described in the claims. [Explanation of symbols]
[0102] 1...anomaly detection device, 2...communication terminal, 3...base station, 4...core network, 10...first setting unit, 11...second setting unit, 12...learning unit, 13...first memory unit, 14...generation unit, 15...second memory unit, 16...collection unit, 17...detection unit, 18...presentation unit, 101...bus, 102...processor, 103...main memory unit, 104...communication interface, 105...auxiliary memory unit, 106...input / output I / O, 107...display device, 121...generator, 121'...trained generator, 122...discriminator, 123...output, 124...training data, 125...block of objective function E, NW...network.
Claims
1. a first setting unit configured to set a distribution of observed data, which is a set of observed values, for a mixed probability model in which the presence or absence of a temporary increase in traffic volume in time-series data of traffic volume of a communication network is used as an observed value; a second setting unit configured to set parameters of the mixed probability model based on the set distribution of the observation data; a learning unit configured to perform adversarial learning of a generative model including: a generator that generates pseudo-anomaly information similar to true anomaly information, using the mixture probability model having the set parameters as true anomaly information representing an occurrence pattern of abnormal traffic, which is the temporary increase in traffic volume; and a classifier that distinguishes between the pseudo-anomaly information generated by the generator and the true anomaly information; a generation unit configured to generate the pseudo anomaly information using a trained generator constructed by the learning unit; a detection unit configured to detect the abnormal traffic occurring in the communication network based on the pseudo-anomaly information generated by the generation unit; and An abnormality detection device comprising:
2. 2. The abnormality detection device according to claim 1, further comprising a collection unit configured to collect time series data of traffic volume on the communication network, The detection unit detects the abnormal traffic when time series data of the traffic volume of the communication network including the temporary increase in traffic volume indicated by the pseudo-abnormal information matches time series data of the traffic volume of the communication network collected by the collection unit. An abnormality detection device characterized by:
3. 2. The abnormality detection device according to claim 1, the mixed probability model is a model that follows a mixed Bernoulli distribution, The parameters of the mixture probability model include a mixture ratio representing the probability that each cluster generates the observed value, and a probability that the observed value of each cluster indicates that there is a temporary increase in traffic volume. An abnormality detection device characterized by:
4. a first setting step of setting a distribution of observed data, which is a set of observed values, for a mixed probability model in which the presence or absence of a temporary increase in traffic volume in time-series data of traffic volume of a communication network is used as an observed value; a second setting step of setting parameters of the mixed probability model based on the set distribution of the observation data; a learning step of performing adversarial learning of a generative model including a generator that generates pseudo-anomaly information similar to true anomaly information, using the mixture probability model having the set parameters as true anomaly information representing an occurrence pattern of abnormal traffic, which is the temporary increase in traffic volume, and a classifier that distinguishes between the pseudo-anomaly information generated by the generator and the true anomaly information; a generating step of generating the pseudo anomaly information using a trained generator constructed in the learning step; a detection step of detecting the abnormal traffic occurring in the communication network based on the pseudo-anomaly information generated in the generation step; An anomaly detection method comprising:
5. The abnormality detection method according to claim 4, further comprising a collection step of collecting time series data of traffic volume on the communication network, The detecting step detects the abnormal traffic when time series data of the traffic volume of the communication network including the temporary increase in traffic volume indicated by the pseudo-abnormal information matches the time series data of the traffic volume of the communication network collected in the collecting step.
1. An anomaly detection method comprising:
Citation Information
Patent Citations
Traffic volume upper limit value prediction device, method and program
JP2015216585A