Abnormality management device and abnormality management method

The abnormality management device and method address the challenge of managing program operation abnormalities by quantizing and modeling resource usage to identify and respond to operational anomalies effectively.

JP7752278B1Active Publication Date: 2025-10-09INTERNET INITIATIVE JAPAN INC

Patent Information

Application Number
JP2025114956
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2025-07-08
Publication Date
2025-10-09
Estimated Expiration
2045-07-08

AI Technical Summary

Technical Problem

Conventional methods struggle to easily manage program operation abnormalities due to the complexity of identifying error causes in high-performance systems, often requiring analysis of source code.

Method used

An abnormality management device and method that quantizes processing resource usage, uses probabilistic models to distinguish between normal and pseudo-normal data, and stores pseudo-normal data as anomaly information for easy identification of operational abnormalities.

Benefits of technology

Facilitates easy management of program operation anomalies by accurately identifying abnormal processing resource usage, enabling prompt response to resolve operational issues.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007752278000001_ABST
    Figure 0007752278000001_ABST
Patent Text Reader

Abstract

The purpose is to easily manage abnormal program operation. [Solution] The abnormality management device 1 includes a second learning unit 13 that updates only the classifier parameters of a classifier 132 that distinguishes between true normal data and pseudo-normal data, based on a probability parameter of normal processing resource usage estimated by a first learning unit 12, in a direction that maximizes the classification accuracy, while keeping fixed the generator parameters of a generator 131 that generates pseudo-normal data that is sufficiently deviated from the distribution of true normal data, taking each observed value of the normal processing resource usage series as true normal data, and an abnormal process database 14 that stores the pseudo-normal data output by the generator 131 after the second learning unit 13 has updated the classifier parameters of the classifier 132 as abnormal information of a process corresponding to a process ID in which abnormal processing resource usage that deviates from the range of normal processing resource usage is detected.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to an abnormality management device and an abnormality management method. [Background technology]

[0002] In recent years, the demands on software have become more sophisticated and complex in order to provide high-performance systems and services. The software implemented in high-performance systems consists of programs with huge amounts of source code, and the execution of these programs is becoming increasingly complex.

[0003] When a process failure or an abnormal operation occurs in an operating system, identifying the cause of the program failure or abnormal operation is complicated, time-consuming, and not easy. For example, Patent Document 1 discloses a method for using a core file to retroactively identify the location of the error and the variable values ​​at the time.

[0004] However, the technology disclosed in Patent Document 1 requires analysis of the source code because it is not possible to identify the cause or location of an error from log information. [Prior art documents] [Patent documents]

[0005] [Patent Document 1] Japanese Patent Application Laid-Open No. 2005-301570 Summary of the Invention [Problem to be solved by the invention]

[0006] As described above, with conventional techniques, it has not been possible to easily manage abnormal program operation.

[0007] The present invention has been made to solve the above-mentioned problems, and has as its object to easily manage abnormalities in program operation. [Means for solving the problem]

[0008] In order to solve the above-mentioned problems, the abnormality management device according to the present invention includes a quantization unit configured to quantize normal processing resource usage amounts used in the execution of a process corresponding to each process ID and convert the amount into a normal processing resource usage sequence of integer values; a first learning unit configured to treat each observed value of the normal processing resource usage sequence as a mutually independent discrete value, and to estimate probability parameters of a probabilistic model representing the normal processing resource usage amounts for each process ID after quantization based on the frequency of occurrence of each observed value; and a second learning unit configured to treat each observed value of the normal processing resource usage sequence as true normal data, and to estimate probability parameters of a probabilistic model representing the normal processing resource usage amounts for each process ID after quantization based on the frequency of occurrence of each observed value. a second learning unit configured to update only a classifier parameter of a classifier that distinguishes between the true normal data and the pseudo-normal data, in a direction that maximizes classification accuracy, based on the probability parameter of the normal processing resource usage estimated by the first learning unit, while keeping fixed a generator parameter of a generator that generates the pseudo-normal data; and a storage unit configured to store the pseudo-normal data output by the generator after the second learning unit updates the classifier parameter of the classifier as anomaly information of a process corresponding to a process ID in which an abnormal processing resource usage that deviates from a normal processing resource usage range is detected.

[0009] In addition, the abnormality management device of the present invention may further include a collection unit configured to collect processing resource usage for each process ID corresponding to the managed process, the quantization unit quantizing the collected processing resource usage for each process ID corresponding to the managed process and converting it into a processing resource usage series of integer values, and a determination unit configured to determine that an operational abnormality has occurred in the managed process when the processing resource usage series matches the pseudo-normal data stored in the memory unit.

[0010] Furthermore, in the abnormality management device of the present invention, when the judgment unit judges that an operational abnormality has occurred in the process to be managed, an instruction may be sent to execute a predetermined response to resolve the operational abnormality in the process.

[0011] In order to solve the above-mentioned problems, the anomaly management method of the present invention includes a quantization step of quantizing normal processing resource usage amounts used in the execution of a process corresponding to each process ID and converting the amount into a normal processing resource usage sequence of integer values; a first learning step of treating each observed value of the normal processing resource usage sequence as a mutually independent discrete value and estimating probability parameters of a probabilistic model representing the normal processing resource usage amounts for each process ID after quantization based on the frequency of occurrence of each observed value; and a pseudo-normal learning step of treating each observed value of the normal processing resource usage sequence as true normal data and estimating probability parameters of a probabilistic model representing the normal processing resource usage amounts for each process ID after quantization based on the frequency of occurrence of each observed value. a second learning step of updating only the classifier parameters of a classifier that distinguishes between the true normal data and the pseudo-normal data in a direction that maximizes classification accuracy, based on the probability parameters of the normal processing resource usage estimated in the first learning step, while keeping fixed generator parameters of a generator that generates data; and a storage step of storing, in a storage unit, the pseudo-normal data output by the generator after updating the classifier parameters of the classifier in the second learning step, as abnormality information of a process corresponding to a process ID in which abnormal processing resource usage that deviates from a range of normal processing resource usage is detected.

[0012] In addition, the abnormality management method of the present invention may further include a collection step of collecting processing resource usage for each process ID corresponding to the managed process, wherein the quantization step quantizes the collected processing resource usage for each process ID corresponding to the managed process and converts it into a processing resource usage series of integer values, and further includes a determination step of determining that an operational abnormality has occurred in the managed process if the processing resource usage series matches the pseudo-normal data stored in the memory unit.

[0013] Furthermore, the abnormality management method according to the present invention may further include an instruction step of sending an instruction to execute a predetermined response to resolve the operational abnormality of the process when it is determined in the determination step that an operational abnormality has occurred in the process to be managed. [Effects of the Invention]

[0014] According to the present invention, after the second learning unit updates the classifier parameters of the classifier, the pseudo-normal data output by the generator is stored as anomaly information for a process corresponding to a process ID in which an abnormal processing resource usage that deviates from the normal processing resource usage range is detected, thereby making it possible to easily manage program operation anomalies. [Brief explanation of the drawings]

[0015] [Figure 1] FIG. 1 is a block diagram showing the configuration of an abnormality management system including an abnormality management device according to an embodiment of the present invention. [Figure 2] FIG. 2 is a diagram for explaining the amount of processing resource usage for each process ID collected by the abnormality management device according to this embodiment. [Figure 3] FIG. 3 is a block diagram showing the configuration of the second learning unit included in the abnormality management device according to this embodiment. [Figure 4] FIG. 4 is a diagram for explaining the second learning unit included in the abnormality management device according to this embodiment. [Figure 5] FIG. 5 is a diagram for explaining the second learning unit included in the abnormality management device according to this embodiment. [Figure 6] FIG. 6 is a block diagram showing the hardware configuration of the abnormality management device according to this embodiment. [Figure 7] FIG. 7 is a flowchart showing the operation of the abnormality management device according to this embodiment. [Figure 8] FIG. 8 is a flowchart showing the operation of the abnormality management device according to this embodiment. [Figure 9]FIG. 9 is a flowchart showing the operation of the abnormality management device according to this embodiment. DETAILED DESCRIPTION OF THE INVENTION

[0016] DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS Preferred embodiments of the present invention will now be described in detail with reference to FIGS.

[0017] [Configuration of anomaly management system] First, with reference to FIG. 1, an overview of an abnormality management system including an abnormality management device 1 according to an embodiment of the present invention will be described.

[0018] The fault management system includes a fault management device 1 and an information processing device 2. The fault management device 1 and the information processing device 2 are connected via a network NW. In a situation where there is very little fault data indicating abnormal processing resource usage when a process operation fault occurs, while a large amount of normal data indicating normal processing resource usage in a normal operating state is available, the fault management system according to this embodiment generates fault data through a learning process based on the normal data and builds a database of fault data for detecting process faults.

[0019] The network NW includes, for example, wired networks such as LAN, WAN, the Internet, and ISDN, as well as wireless networks such as wireless LAN and mobile communication networks using LTE / 4G, 5G, and 6G wireless communication systems, but the scope of the present invention is not limited to these.

[0020] The information processing device 2 can be realized as a server, a gateway, a desktop computer, an embedded device, a mobile communication terminal such as a smartphone, a tablet computer, a laptop computer, etc. In this embodiment, the information processing device 2 is not limited to one device, but includes the case of multiple devices. In the case of multiple devices, each information processing device 2 executes the same process generated by the same application or the same executable file.

[0021] The information processing device 2 is uniquely identified by network identification information such as an IP address or a MAC address, or a device ID assigned by an anomaly management system. The information processing device 2 can be realized by a computer equipped with a processor, a main memory device, a communication interface, an auxiliary memory device, and an input / output (I / O), and a program that controls these hardware resources.

[0022] The information processing device 2 executes one or more applications (processes) on the OS, and each process is assigned a unique process ID (PID) by the OS. A process ID is assigned to each execution unit of a process, and even the same application may have a different process ID depending on the execution environment and timing. The information processing device 2 calculates the CPU utilization rate as the amount of processing resources used for each process ID corresponding to the running process, and records the calculated CPU utilization rate in memory. If the configurations of the CPUs of the multiple information processing devices 2, such as the number of cores and clock speed, are not the same, the CPU utilization rate recorded by each information processing device 2 is normalized or corrected using each CPU benchmark, etc. The normalization or correction process may be performed by the abnormality management device 1.

[0023] In this embodiment, CPU usage rate for each process ID is used as data on processing resource usage, but processing resource usage rate may also be memory usage rate for each process ID, I / O wait time for each process ID, etc. Figure 2(a) is a diagram showing normal data a1, which is a normal CPU usage rate, and abnormal data b1, which includes abnormal CPU usage rates. The horizontal axis represents process IDs, which are arranged in the order in which the processes were executed. The vertical axis represents CPU usage rate [%]. Within the range of process IDs in area c, the CPU usage rate of abnormal data b1 is an abnormal value. The abnormal value of CPU usage rate is caused by an abnormal operation of a process executed by information processing device 2.

[0024] A state in which an operational abnormality occurs in a process executed by the information processing device 2 refers to a state in which the process deviates from the normal load range, such as excessive consumption of processing resources. For example, this state applies when the CPU usage rate for each process ID exceeds a preset threshold, or when a normal CPU usage rate range is defined according to the attributes of the process and the CPU usage rate outside that range deviates from the normal load range. Even if the CPU usage rate does not exceed the threshold, a predefined CPU usage rate pattern can be used as a load state pattern when an operational abnormality occurs in the process, and can be used as an abnormal CPU usage rate that deviates from the normal CPU usage rate range.

[0025] In this way, when a process is behaving abnormally, it is assumed that abnormal processing such as an infinite loop or excessive recursive calls is occurring in the execution of the process associated with that process ID.On the other hand, normal CPU usage means that no abnormal processing is occurring in the execution of the process associated with each process ID, and CPU resource consumption is within a normal range.

[0026] [Function block of the abnormality management device] Next, functional blocks of the abnormality management device 1 according to this embodiment will be described with reference to the block diagram of Fig. 1. As shown in Fig. 1, the abnormality management device 1 includes a collection unit 10, a quantization unit 11, a first learning unit 12, a second learning unit 13, an abnormal process database (storage unit) 14, a determination unit 15, an instruction unit 16, and a storage unit 17. The abnormality management device 1 builds a database of abnormal data through a learning process using normal data.

[0027] The collection unit 10 collects, via the network NW, CPU utilization rates, which are normal processing resource utilization rates used in executing processes corresponding to each process ID recorded in the information processing device 2. The collection unit 10 can collect CPU utilization rates, which are processing resource utilization data, from multiple information processing devices 2. The processing resource utilization data is associated with a process ID. The CPU utilization rates collected by the collection unit 10 have M (e.g., 10,000) process IDs.

[0028] The collection unit 10 collects normal CPU utilization rates for each process ID used for learning by the first learning unit 12 and the second learning unit 13. The collection unit 10 also collects CPU utilization rates for each process ID corresponding to a managed process that is to be subjected to abnormality determination by the determination unit 15.

[0029] The quantizer 11 quantizes the normal processing resource usage and converts it into a normal processing resource usage sequence of integer values. The quantizer 11 converts the normal CPU usage rate included in the CPU usage rate for each process ID collected by the collector 10 into a predetermined discrete integer value, for example, a positive integer level value 1 to 7, as shown on the vertical axis of FIG. 2(b). The quantizer 11 rounds the value up or down by rounding. The quantizer 11 further quantizes the CPU usage rate for each process ID corresponding to the managed process collected by the collector 10 and converts it into a processing resource usage sequence of integer values.

[0030] The first learning unit 12 treats each observed value in the normal processing resource usage series as a mutually independent discrete value and estimates probability parameters of a probability model representing normal processing resource usage after quantization based on the occurrence frequency of each observed value. The first learning unit 12 sets a probability model based on a multinomial distribution for normal CPU usage, estimates the probability parameters, and calculates the log-likelihood of the series of each observed value based on the probability model. In other words, when there is very little data on abnormal CPU usage, the first learning unit 12 uses a large amount of available normal CPU usage data to determine a probability model of normal CPU usage, which is a statistical standard, and a corresponding normality index.

[0031] The first learning unit 12 focuses on the fact that the naive Bayes method treats each observation value as a conditionally independent discrete variable, and sets a probability model by extending the naive Bayes method to a multinomial distribution model so that the quantized CPU usage rate for each process ID can be regarded as a vector of occurrence counts by category. Here, as shown in the following formula (1), the CPU usage rate observed for each process ID is quantized, and the vector expressed by the occurrence counts of M categories is taken as the observation value X. X=x=(x1,x2,...,x M ) ···(1)

[0032] Furthermore, the class label to which the observed value X belongs, that is, the event Y, is defined as a binary variable by the following equation (2).

number

[0033] Here, Bayes' theorem is defined by the following equation (3).

number

[0034] In the above equation (3), P(X) is the marginal probability (outcome) of observing observed value X, P(Y) is the prior probability that event Y occurs, P(X|Y) is the conditional probability that observed value X occurs given that event Y occurs, and P(Y|X) is the conditional probability (posterior probability) that event Y occurs given that observed value X occurs. It is known that maximum likelihood estimation determines the parameters of the prior probability P(Y) so that the posterior probability P(Y|X) is maximized, using observed value X as the teacher signal, which is the correct value. In contrast, Bayesian estimation is a procedure for determining the parameters (probabilistic values) of the prior probability P(Y) that best explain observed value X, represented by a quantized processing resource usage sequence.

[0035] In other words, by arbitrarily setting the parameters of the prior probability P(Y) in advance (e.g., normal distribution), it is possible to estimate the parameters of the prior probability P(Y) that maximizes P(X|Y), i.e., that best matches the marginal probability P(X) (outcome). In this way, the advantage of Bayesian estimation is that the prior probability P(Y) can be arbitrarily specified in advance. When each of the P(X|Y) is independent, the event Y is predicted by the observed variable x under given conditions. i are conditionally independent of each other, and the likelihood ρ(x|y) is given by the naive Bayes estimation method of the following equation (4).

[0036]

number

[0037] In this embodiment, we deal with a series of integer values ​​obtained by quantizing the CPU usage rate for each process ID, so we use a vector of occurrence counts for each category x = (x1, x2, . . . , x M ) as one sample, and while maintaining the naive assumption that each process ID is conditionally independent, we extend this to a multinomial distribution probability model. The occurrence probability θ1, θ M are independent and the sum of their occurrence probabilities is 1. Under this constraint, the probability model of the multinomial distribution is expressed by the following equation (5).

[0038]

number

[0039] As shown in the above formula (5), the observed value x, which is the count value of the CPU usage rate for each process ID, i The sum of (x1+x2++x M ) is determined, the distribution is the product of the following equation (6), and the observed value x observed in the i-th process ID in the series is i Independently, probability θ i ^x i can be obtained.

number

[0040] Therefore, it can be seen that the relationship is similar to that of the naive Bayes method in equation (4). Here, the prior probability P(Y) is a binary problem of abnormal values ​​(Y=0) and normal values ​​(Y=1), and each unknown parameter is θ 0 ,θ 1 Also, D is defined as the marginal probability (outcome), and the observed value of the abnormal value (Y=0) is D 0 , the normal value (Y=1) observation is D 1 The above equation (6) can be expressed as the following equation (7) by converting the product form into a sum under the Naive Bayes independence assumption and decomposing the log likelihood for the observation set D, which is the data set, by class.

[0041]

number

[0042] Here, the constraint is expressed by the following equation (8).

number

[0043] Furthermore, by applying the Lagrange multiplier method, the outlier parameter θ 0 The maximum value of the logarithmic likelihood for the i-th component of is given by the following equation (9).

number

[0044] Normal value parameter θ 1 Similarly, if we calculate the maximum value of the log likelihood for 1 , θ 0 is expressed by the following equation (10).

number

[0045] In addition, x i Since takes integer values, x i To prevent the problem of the multinomial distribution diverging when ^(n) is 0, smoothing can be performed by specifying +α (e.g., α=1) smoothing. In this way, the first learning unit 12 regards the normal CPU usage rate series for each quantized process ID as a multinomial distribution, estimates probability parameters from the count of occurrence frequency, and measures the normality by likelihood. Note that the first learning unit 12 also calculates the abnormal value parameter θ 0 No estimates are made regarding

[0046] The second learning unit 13 regards each observed value of the normal processing resource usage series as true normal data and updates only the classifier parameters of the classifier 132 that distinguishes between true normal data and pseudo-normal data in a direction that maximizes the classification accuracy, while keeping the generator parameters of the generator 131 that generates pseudo-normal data that is sufficiently deviated from the distribution of true normal data fixed, based on the probability parameters of the normal CPU usage estimated by the first learning unit 12.

[0047] As shown in Fig. 3, the second learning unit 13 executes a Max learning phase of a GAN (Generative Adversarial Network) having a generator 131 and a classifier 132, in which only the parameters of the classifier 132 are updated while the generator 131 is fixed. When a normal processing resource usage sequence is regarded as true normal data, the second learning unit 13 aims to generate pseudo-normal data that deviates sufficiently from the distribution of normal data, that is, a processing resource usage sequence that can be treated as an abnormal processing resource usage sequence. For this reason, the Min learning phase in the normal GAN ​​adversarial learning procedure, in which the generator 131 is updated in the minimization direction, is not performed.

[0048] As shown in FIG. 3, a generative model according to this embodiment, which includes a generator 131 and a classifier 132, is provided for each observed value (process ID) of a processing resource usage sequence. Therefore, in this embodiment, M generative models are trained. Here, pseudo-normal data that deviates sufficiently from the distribution of normal data refers to a generated sequence that, as a statistical property, is located in a region that significantly deviates from the normal region of normal data, relative to normal data that represents a normal processing resource usage sequence. When the index of distance or deviation from normal data is log-likelihood, a sequence with a smaller likelihood corresponds to the pseudo-normal data. When the index is cross-entropy, a sequence with a larger entropy value corresponds to a larger deviation. Furthermore, when the index is KL distance, a sequence with a large deviation in the overall distribution is treated as a sufficiently deviated sequence.

[0049] 4 and 5 are diagrams schematically illustrating the neural network configuration of the generator 131 and the classifier 132 of the generative model used by the second learning unit 13. As shown in FIG. 4, the generator 131 is configured as a neural network having an input layer, a hidden layer, and an output layer. The generator 131 is a model that generates pseudo-normal data from random noise. For example, m randomly sampled Gaussian noise vectors (z1 to z m ).

[0050] The generator 131 outputs the output G(z) after performing a product-sum operation on the input and weight parameters and threshold processing using an activation function. The output G(z) from the generator 131 is pseudo-normal data that deviates from the distribution of true normal data. CNN or ResNet can be used as the neural network that constitutes the generator 131.

[0051] The classifier 132 shown in Fig. 5 is configured with a neural network having an input layer, a hidden layer, and an output layer. In the example of Fig. 5, the normal processing resource usage sequence collected by the collection unit 10 and calculated by the quantization unit 11 is given as true normal data as training data input.

[0052] The classifier 132 outputs a binary value of 1 or 0 after performing a product-sum operation on the input and weight parameters and threshold processing using an activation function. The classifier 132 outputs an output y=1 when it correctly identifies the training data related to the input true normal data as true normal data. On the other hand, it outputs an output y=0 when it correctly identifies the training data related to the input pseudo normal data as pseudo normal data. In this way, the classifier 132 is a model that distinguishes the model distribution generated by the generator 131 from the data distribution of the training data, which is the true distribution. A CNN can be used as the neural network that constitutes the classifier 132.

[0053] FIG. 3 is a block diagram for explaining the configuration of the generative model by the second learning unit 13. The generator 131 of the generative model adopted by the second learning unit 13 is represented as a function G, and the classifier 132 is represented as a function D. Furthermore, true normal data is represented as x, the predicted value output by the classifier 132 is represented as y, and the correct label is represented as t. The correct label t is set to 1 for true normal data and 0 for pseudo-normal data generated by the generator 131. In this case, the classifier 132 calculates the cross entropy E CE It can be expressed as:

[0054]

number

[0055] The first term in the brace of the above equation (11) represents t n lny n In this case, the predicted value y n is the correct label of the true normal data, t n = 1. On the other hand, the second term in the braces represents (1-t n )ln(1-y n ), the predicted value y n is the correct label value (1-t n ) = 0. In this way, the cross entropy E CEis the maximum value when the predicted value matches the correct label value.

[0056] Here, the generator 131 that configures the generative model uses parameters w G ,θ G and the function G(w G ,θ G ) The classifier 132 uses the parameter w D ,θ D and function D(w D ,θ D ) The cross entropy E in the above equation (11) CE The objective function E of the generative model including the generator 131 and the discriminator 132 based on the above can be expressed by the following equation (12).

number

[0057] The first term of the above equation (12) represents E D(x)=1 lnD(w D ,θ D ) is the expected value at which the classifier 132 classifies true normal data as true normal data. D(x)=0 ln(1-D(G(w G ,θ G ),w D ,θ D )) is the expected value at which the classifier 132 classifies the pseudo-normal data generated by the generator 131 as pseudo-normal data. Here, the objective function E of the generative model in the above formula (12) is set to the parameter θ using the prior probability of the classes of abnormal values ​​(Y=0) and normal values ​​(Y=1) in the multinomial distribution. 0 , θ 1 When these are multiplied and incorporated, the objective function E is expressed by the following equation (13).

[0058]

number

[0059] In the above equation (13), the parameter θ of the multinomial distribution 0 , θ1 The objective function E of the i-th component is expressed by the following equation (14).

number

[0060] By weighting the objective function (equation (12)) of a normal GAN, an adversarial loss that proportionally allocates the contributions of the normal value class and the abnormal value class can be obtained, as shown in the above equations (13) and (14). Here, the objective function E when the generator 131 is fixed is expressed by the following equation (15).

number

[0061] The convergence value (maximum value) of the discriminator 132 is expressed by the following equation (16).

number

[0062] Normal value parameter θ i ^1 is calculated using the above formula (10). The parameter θ of the abnormal value i For ^0, since there is a small amount of abnormal data related to abnormal CPU usage, we use empirical rules or a preset value. i ^1,θ i ^0 is the value of each observation x i The expected value ratio, E, varies depending on the ρ(x|y=1) :E ρ(x|y=0) For example, since the number of normal data is overwhelmingly large, the ratio is set in advance to 0.99:0.01.

[0063] In the learning of the generative model of this embodiment, only Max optimization of the objective function E is performed, and the parameters of the discriminator 132 with the generator 131 fixed are learned. Therefore, it is possible to prevent the output of the generator 131 from converging to the distribution of normal data, and conversely, to maintain a sequence that is sufficiently deviated from the distribution of normal data. When the update of the discriminator 132 has converged, the pseudo-normal data sequence output by the fixed generator 131 is a model of normal CPU usage rates for each process ID (parameters θ of the multinomial distribution 1 ) has a low likelihood. In this case, the generator 131 can generate only pseudo-normal data, which is an abnormal sequence expressed by the following equation (17).

number

[0064] The abnormal process database 14 stores the pseudo-normal data generated by the generator 131 after the second learning unit 13 updates the parameters of the classifier 132 as abnormal information of the process corresponding to the process ID in which an abnormal CPU usage rate outside the normal CPU usage rate range is detected. Specifically, when the generator 131 provided in the generative model corresponding to each process ID learns with, for example, 10,000 pieces of training data, it generates 10,000 pieces of pseudo-normal data, which are registered in the abnormal process database 14. The abnormal process database 14 registers the abnormal CPU usage rate for each process ID as a reference pattern for abnormality determination.

[0065] The determination unit 15 determines that an abnormal operation of the managed process has occurred when the processing resource usage series for each process ID corresponding to the managed process matches the pseudo-normal data stored in the abnormal process database 14. More specifically, when the processing resource usage series converted by the quantization unit 11 into the CPU usage for the process ID of the managed process matches the pseudo-normal data stored in the abnormal process database 14, the determination unit 15 determines that an abnormal operation of the managed process has occurred.

[0066] In addition to the case where the data matches the pseudo-normal data stored in the abnormal process database 14, the judgment unit 15 can judge that an abnormal operation of the process has occurred if the sum of the squares of the differences between the level values ​​of the processing resource usage amount series of CPU usage rates for each process ID corresponding to the managed process and the level values ​​of the pseudo-normal data stored in the abnormal process database 14 is within a set threshold value according to the following equation (18).

number

[0067] When the determination unit 15 determines that an abnormality in the process operation has occurred, the instruction unit 16 sends an instruction to execute a predetermined response to resolve the abnormality in the process operation. For example, the instruction unit 16 can identify the process ID of the abnormal CPU usage rate and send an instruction to terminate and restart the corresponding process to the information processing device 2 via the network NW.

[0068] The storage unit 17 stores parameters of a probabilistic model that indicates a normal CPU utilization rate for each process ID, estimated by learning by the first learning unit 12. The storage unit 17 also stores a generator 131 included in the trained generative model constructed by learning by the second learning unit 13.

[0069] [Hardware configuration of the fault management device] Next, an example of a hardware configuration for realizing the abnormality management device 1 having the above-described functions will be described with reference to FIG.

[0070] 6, the fault management device 1 can be realized by, for example, a computer including a processor 102, a main memory device 103, a communication interface 104, an auxiliary memory device 105, and an input / output (I / O) 106 connected via a bus 101, and a program for controlling these hardware resources. Furthermore, the fault management device 1 includes a display device 107.

[0071] The processor 102 is realized by a CPU, a GPU, an FPGA, an ASIC, or the like.

[0072] The main memory device 103 pre-stores programs for the processor 102 to perform various controls and calculations. The processor 102 and the main memory device 103 implement the functions of the abnormality management device 1, such as the collection unit 10, the quantization unit 11, the first learning unit 12, the second learning unit 13, the determination unit 15, and the instruction unit 16 shown in FIG.

[0073] The communication interface 104 is an interface circuit for connecting the abnormality management device 1 to various external electronic devices via a network.

[0074] The auxiliary storage device 105 is composed of a readable / writable storage medium and a drive for reading and writing various information such as programs and data from and to the storage medium. The auxiliary storage device 105 can use a semiconductor memory such as a hard disk or flash memory as the storage medium.

[0075] The auxiliary storage device 105 has a program storage area for storing an abnormality management program. The auxiliary storage device 105 also has a program storage area for storing a first learning program for estimating parameters related to normal CPU utilization for each process ID using a multinomial distribution probability model executed by the abnormality management device 1. The auxiliary storage device 105 also has a program storage area for storing a second learning program for training the classifier 132 of the generative model executed by the abnormality management device 1. The auxiliary storage device 105 realizes the abnormal process database 14 and the storage unit 17 described in FIG. 1. Furthermore, the auxiliary storage device 105 may have, for example, a backup area for backing up the above-mentioned data, programs, etc.

[0076] The input / output I / O 106 is an input / output device that inputs signals from external devices and outputs signals to external devices.

[0077] The display device 107 is configured by an organic EL display, a liquid crystal display, etc. The display device 107 can display the CPU utilization rate data collected from the information processing device 2 on the screen.

[0078] [Operation of the abnormality management device] Next, the operation of the abnormality management device 1 having the above-described configuration will be described with reference to the flowcharts of FIGS.

[0079] 7, first, the collection unit 10 collects, via the network NW, data on CPU utilization rates for each process ID, which includes a certain number of normal CPU utilization rates or more (step S1). For example, the collection unit 10 collects multiple pieces of normal CPU utilization rate data for each process ID, which includes a normal CPU utilization rate of 99% or more.

[0080] Next, the quantization unit 11 quantizes the normal CPU utilization rate for each process ID collected in step S2 and converts it into a normal processing resource utilization sequence of integer values ​​(step S2). Next, the first learning unit 12 performs a first learning process (step S3). FIG. 8 is a flowchart illustrating the first learning process of step S3 in more detail. As shown in step S30 of FIG. 8, the first learning unit 12 models the normal processing resource utilization sequence obtained by conversion in step S2 using a multinomial distribution (step S30). The first learning model sets the probability model of the above equation (5).

[0081] Next, the first learning unit 12 calculates a parameter θ for binary classification of normal CPU utilization rate and abnormal CPU utilization rate. 1 , θ 0 (Step S31). Next, the first learning unit 12 defines the log-likelihood in accordance with the above formula (7) (Step S32). In Step S32, the first learning unit 12 converts the objective function for optimization from a product form (formula (5)) to a sum form. Next, the first learning unit 12 determines a parameter θ related to a normal CPU utilization rate that maximizes the log-likelihood (formula (7)) under the constraint (formula (8)) by the Lagrange multiplier method of the above formula (9). 1 (Equation (10)) is estimated (step S33). 1 is stored in the storage unit 17, and the process proceeds to step S4 in FIG.

[0082] Next, the second learning unit 13 determines each observed value of the normal processing resource usage sequence as true normal data, and while keeping the generator parameters of the generator 131 that generates pseudo-normal data that is sufficiently deviated from the distribution of true normal data fixed, the second learning unit 13 determines the normal value parameters θ estimated by the first learning unit 12 in step S3. 1 Based on this, only the classifier parameters of the classifier 132 that distinguishes between true normal data and pseudo normal data are updated in a direction that maximizes the classification accuracy (second learning process) (step S4).

[0083] 9 is a flowchart for explaining the second learning process in step S4. First, the second learning unit 13 calculates the normal and abnormal value parameters θ of the probabilistic model of the multinomial distribution estimated in the first learning process in step S3. 1 ,θ 0 is set as the objective function (equation (15)) of the generative model (step S50). More specifically, the second learning unit 13 sets the parameter θ i ^1 is substituted into the objective function when the generator 131 in the above equation (15) is fixed, and the parameter θ i For ^0, we use an empirical rule or a preset value and substitute it into the above equation (15). Furthermore, the ratio of the expected values, E ρ(x|y=1) :E ρ(x|y=0) Regarding , since the number of normal data is overwhelmingly large, a preset value of, for example, 0.99:0.01 is adopted in the above equation (15).

[0084] Next, the second learning unit 13 acquires the normal processing resource usage sequence collected in step S1 and quantized and converted in step S2 as true normal data (step S51). Next, the second learning unit 13 inputs the true normal data as training data 134 to the classifier 132, and adjusts the parameter w D ,θ D (Step S52). In Step S52, the second learning unit 13 can cause the classifier 132 to learn true normal data using, for example, an error backpropagation method. In Step S51, the classifier 132 that can distinguish true normal data from true normal data is pre-trained.

[0085] Next, the second learning unit 13 generates Gaussian noise and provides a random vector of the generated Gaussian noise as an input to the generator 131 (step S53). Subsequently, the generator 131 calculates a random vector of the input z and the weight parameter w based on the provided Gaussian noise. G ,θ GThen, a product-sum operation and a threshold process using an activation function are performed to generate pseudo-normal data G(z) (step S54).

[0086] Next, the second learning unit 13 learns the classifier 132. The learning of the classifier 132 is performed by using the parameter w D ,θ D First, the second learning unit 13 provides true normal data as training data 134 as input to the classifier 132. Then, the second learning unit 13 adjusts the parameter w by backpropagation or the like so that the objective function E in the above equation (15) is maximized. D ,θ D (Step S55). The label of the training data 134 is set to 1 (true normal data).

[0087] Next, the second learning unit 13 provides the pseudo-normal data generated by the generator 131 in step S54 to the classifier 132, and calculates the parameter w by backpropagation or the like so that the objective function E in the above equation (15) is maximized. D ,θ D is updated (step S56).

[0088] The learning of the classifier 132 in steps S55 and S56 corresponds to the dashed arrows in the block diagram of the second learning unit 13 shown in FIG. 3 , which indicate that a classifier error is calculated in block 135 of the objective function E based on the output 133 from the classifier 132, and then the error is backpropagated to the classifier 132.

[0089] Thereafter, the learning of the classifier 132 from step S54 to step S56 is repeated until the value of the objective function E converges (step S57: NO). On the other hand, if the value of the objective function E converges to the optimal solution of the above equation (16) (step S57: YES), the processing from step S52 to step S57 is repeated using the remaining true normal data in order until the generator 131 and the classifier 132 are learned (step S58: NO).

[0090] Thereafter, when the classifier 132 has been trained using all true normal data (step S58: YES), the second learning unit 13 stores the generator 131 in the storage unit 17 (step S59). The second learning unit 13 also performs the processes from step S50 to step S59 for each generative model corresponding to M observed values ​​(process IDs) to train the classifier 132. Note that in steps S52, S55, and S56, batch processing may be performed to update the parameters. Similarly, in steps S53 and S54, noise may be generated in batch units to generate pseudo-normal data. Thereafter, the process proceeds to step S5 in FIG. 7.

[0091] Next, the abnormal process database 14 stores the pseudo-normal data generated by the generator 131 after the classifier 132 has been updated through learning by the second learning unit 13 (step S5). In step S5, the pseudo-normal data generated by the generator 131 corresponding to each process ID is registered in the abnormal process database 14. For example, assume that 1,000 process IDs (M=1,000) are generated. In this case, if learning is performed using, for example, 10,000 pieces of training data in each of 1,000 generative models corresponding to the 1,000 process IDs in the second learning process of step S4, the generator 131 of each generative model generates 10,000 pieces of pseudo-normal data. Therefore, 10,000 pieces of pseudo-normal data are registered in the abnormal process database 14 for each of the 1,000 process IDs.

[0092] Next, the collection unit 10 collects the CPU utilization rate for each process ID corresponding to the process to be managed (step S6). Thereafter, the quantization unit 11 quantizes the CPU utilization rate for each process ID corresponding to the process to be managed collected in step S6 and converts it into a processing resource utilization amount series (step S7).

[0093] Next, the determination unit 15 determines that an operational abnormality has occurred in the managed process if the processing resource usage series of the CPU usage rate for each process ID corresponding to the process obtained in step S7 matches the pseudo-normal data stored in the abnormal process database 14 (step S8). In step S8, it can be determined that an operational abnormality has occurred in the managed process if the series of CPU usage rate for each process ID corresponding to the managed process matches the pseudo-normal data completely or within a certain tolerance range.

[0094] Specifically, the judgment unit 15 can judge that the signal to be managed is an abnormal signal if the sum of the squares of the differences between the level values ​​of the CPU usage series for each process ID corresponding to the process to be managed and the level values ​​of the pseudo-normal data is within a threshold value according to the above equation (18).

[0095] In addition, even if it does not match the processing resource usage series of all process IDs, the pseudo-normal data for each process ID can be checked sequentially, and when the processing resource usage series of CPU usage of the process ID corresponding to the managed process partially matches the pseudo-normal data series, it can be determined that an operational abnormality has occurred in the process ID corresponding to the managed process.

[0096] Next, the instruction unit 16 identifies a process linked to the process ID determined to have experienced an operational abnormality, and sends an instruction to execute a predetermined response to resolve the operational abnormality (step S9). The instruction unit 16 can send an instruction to terminate and restart the process linked to the identified process ID in which the operational abnormality has occurred, to the information processing device 2 via the network NW.

[0097] As described above, the anomaly management device 1 according to this embodiment employs a Naive Bayes algorithm with a multinomial distribution as a probability model for a sequence of integer values ​​obtained by quantizing normal CPU utilization rates for each process ID, and estimates probability parameters for a normal processing resource usage sequence. Furthermore, the probability parameters for the estimated normal processing resource usage sequence are set as the objective function of the generative model, and learning is performed by updating only the classifier 132 while keeping the generator 131 fixed, so that pseudo-normal data that is sufficiently different from the distribution of true normal data indicating a normal processing resource usage sequence is output from the generator 131. The collected pseudo-normal data is registered as a database for anomaly detection, making it easy to manage program operation anomalies.

[0098] Furthermore, according to the abnormality management device 1 of this embodiment, the first learning process and the second learning process are performed based on normal data of CPU usage rate for each process ID, and a database of abnormal data is constructed, so that the process ID in which an operational abnormality is occurring can be easily identified.

[0099] In the embodiment described above, the second learning unit 13 has been described as having a generative model with a GAN configuration. However, the generative model can be configured not only based on a GAN but also based on a VAE (Variational Autoencoder), Energy-Based Models (EBMs), or the like.

[0100] The above describes embodiments of the abnormality management device and abnormality management method of the present invention, but the present invention is not limited to the described embodiments, and various modifications that a person skilled in the art can conceive are possible within the scope of the invention described in the claims. [Explanation of symbols]

[0101] 1...abnormality management device, 2...information processing device, 10...collection unit, 11...quantization unit, 12...first learning unit, 13...second learning unit, 14...abnormal process database, 15...judgment unit, 16...instruction unit, 17...memory unit, 101...bus, 102...processor, 103...main memory device, 104...communication interface, 105...auxiliary memory device, 106...input / output I / O, 107...display device, 131...generator, 132...identifier, NW...network.

Claims

1. a quantization unit configured to quantize normal processing resource usage amounts used in the execution of a process corresponding to each process ID and convert the amounts into a normal processing resource usage sequence of integer values; a first learning unit configured to treat each observed value of the normal processing resource usage sequence as a mutually independent discrete value, and to estimate probability parameters of a probability model representing the normal processing resource usage for each of the process IDs after quantization based on the occurrence frequency of each observed value; a second learning unit configured to update only a classifier parameter of a classifier that distinguishes between the true normal data and the pseudo-normal data, based on the probability parameter of the normal processing resource usage estimated by the first learning unit, in a direction that maximizes classification accuracy, while keeping fixed a generator parameter of a generator that generates pseudo-normal data that is sufficiently deviated from the distribution of the true normal data, with each observed value of the normal processing resource usage sequence considered as true normal data; a storage unit configured to store the pseudo-normal data output by the generator after the second learning unit updates the classifier parameters of the classifier as abnormal information of a process corresponding to a process ID in which an abnormal processing resource usage amount that deviates from a normal processing resource usage amount range is detected; An abnormality management device comprising:

2. 2. The abnormality management device according to claim 1, The system further comprises a collection unit configured to collect the amount of processing resource usage for each process ID corresponding to a process to be managed, the quantization unit quantizes the collected processing resource usage for each process ID corresponding to the managed process and converts it into a processing resource usage series of integer values; The system further includes a determination unit configured to determine that an abnormal operation has occurred in the managed process when the processing resource usage series matches the pseudo-normal data stored in the storage unit. An abnormality management device characterized by:

3. 3. The abnormality management device according to claim 2, Further, an instruction unit configured to, when the determination unit determines that an operational abnormality has occurred in the process to be managed, issue an instruction to execute a predetermined measure to resolve the operational abnormality in the process. An abnormality management device characterized by:

4. A computer-implemented anomaly management method, comprising: a quantization step of quantizing the normal processing resource usage amount used in the execution of a process corresponding to each process ID and converting the amount into a normal processing resource usage amount series of integer values; a first learning step of treating each observed value of the normal processing resource usage sequence as a mutually independent discrete value, and estimating probability parameters of a probabilistic model representing the normal processing resource usage for each process ID after quantization based on the occurrence frequency of each observed value; a second learning step of updating only a classifier parameter of a classifier that distinguishes between the true normal data and the pseudo-normal data, based on the probability parameter of the normal processing resource usage estimated in the first learning step, in a direction that maximizes classification accuracy, while keeping fixed a generator parameter of a generator that generates pseudo-normal data that is sufficiently deviated from the distribution of the true normal data, with each observed value of the normal processing resource usage sequence considered as true normal data; a storage step of storing, in a storage unit, the pseudo-normal data output by the generator after updating the classifier parameters of the classifier in the second learning step, as abnormality information of a process corresponding to a process ID in which an abnormal processing resource usage amount that deviates from a normal processing resource usage amount range is detected; An abnormality management method comprising:

5. 5. The abnormality management method according to claim 4, The system further comprises a collection step of collecting the amount of processing resource usage for each process ID corresponding to the process to be managed, the quantization step quantizes the collected processing resource usage for each process ID corresponding to the process to be managed and converts it into a processing resource usage series of integer values; Further, the method includes a determining step of determining that an abnormal operation has occurred in the managed process when the processing resource usage series matches the pseudo-normal data stored in the storage unit. An abnormality management method characterized by:

6. 6. The abnormality management method according to claim 5, Further, the method includes an instruction step of issuing an instruction to execute a predetermined measure to resolve the operational abnormality of the process when it is determined in the determination step that an operational abnormality has occurred in the process to be managed. An abnormality management method characterized by:

Citation Information

Patent Citations

  • Redundant resource management device, program, and redundant resource management method

    JP2007122434A

  • Information processor with process monitoring function, method and program for monitoring process

    JP2010165036A

  • Process restart device, process restart method and process restart program

    JP2012168816A

  • Switch device and information processing system

    JP2021040215A

  • Process dumping method, device and program

    JP2005301570A

Cited By

  • Abnormality management device and abnormality management method

    JP7813946B1