Abnormal detection system, abnormal detection method, and program
The proposed system addresses the challenges of log analysis in large-scale event networks by calculating time-series data of log file size increases without accessing log content, enabling efficient and confidential anomaly detection.
Patent Information
- Application Number
- JP2021173917
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2021-10-25
- Publication Date
- 2025-06-11
- Estimated Expiration
- 2041-10-25
AI Technical Summary
In large-scale event networks with multi-vendor network devices, existing log analysis methods face challenges due to varied log formats, high computational resource requirements, and difficulties in maintaining confidentiality and adapting to rapid trend changes.
A system comprising a log storage unit, an information acquisition unit, and an analysis unit that calculates time-series data of the increase rate of log file sizes without accessing log content, thereby performing anomaly detection with minimal computational resources while ensuring confidentiality.
The system enables efficient anomaly detection in large-scale event networks with reduced computational load and maintained confidentiality, effectively addressing the limitations of existing methods.
Smart Images

Figure 0007691056000001 
Figure 0007691056000002 
Figure 0007691056000003
Abstract
Description
Technical Field
[0001] The present invention relates to a technique for performing anomaly detection using logs of network devices.
Background Art
[0002] In general, messages output from network devices such as routers and servers are saved as logs, and by analyzing the logs, the operating status of the network and network devices is grasped.
[0003] As a log analysis method, there is a method of performing semantic analysis of logs to perform anomaly detection and the like. However, in a large-scale event network constructed using a large number of multi-vendor network devices, log analysis using semantic analysis is difficult due to the following factors.
[0004] Since there are various log formats due to multi-vendor devices, it is difficult to make a computer interpret the meaning of the logs. Also, unlike normal network operations, the operation time is short, and a large number of unknown logs are generated for the operator, so it is difficult to perform labeling (initial setting) for semantic analysis. In addition, semantic analysis requires a large amount of computing resources, but in large-scale events, there is little surplus computing resources in terms of security enhancement measures and budget. Furthermore, since the trend changes in a short period of time in an event, it becomes difficult to adapt to machine learning and the like.
Prior Art Documents
Non-Patent Documents
[0005]
Non-Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0006] In a large-scale event network where it is difficult to semantically analyze logs as described above, there is a conventional technique (Non-Patent Document 1) for detecting network anomalies by counting the number of log lines and detecting anomalies in the change of the number of lines.
[0007] However, in the conventional technique disclosed in Non-Patent Document 1, in order to obtain the number of log lines, it is necessary to expand the log itself in memory and count the line feed characters, which requires a large amount of computing resources. In addition, since it is necessary to refer to the content of the log, confidentiality cannot be maintained if the log contains confidential information.
[0008] That is, the conventional technique has problems of a large amount of calculation and difficulty in maintaining confidentiality. This is a problem that may occur as the network scale and usage expand, not limited to event networks.
[0009] The present invention has been made in view of the above points, and an object thereof is to provide a technique for performing anomaly detection using logs of messages output from network devices with a small amount of calculation while maintaining confidentiality.
Means for Solving the Problems
[0010] According to the disclosed technique, a log storage unit that stores a log file recording logs output from network devices, an information acquisition unit that periodically acquires file information including the file size and file name of the log file from the log storage unit, an analysis unit that calculates time-series data of the increase rate of the file size of the log file from the file information and performs anomaly detection from the time-series data, An anomaly detection system including the above is provided.
Effects of the Invention
[0011] According to the disclosed technology, anomaly detection using the logs of messages output from network devices can be performed with a small amount of computation while maintaining confidentiality.
Brief Description of the Drawings
[0012]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Embodiments for Carrying Out the Invention
[0013] Hereinafter, embodiments of the present invention (these embodiments) will be described with reference to the drawings. The embodiments described below are merely examples, and the embodiments to which the present invention is applied are not limited to the following embodiments.
[0014] (System Configuration Example, Operation Example) FIG. 1 shows an overall configuration example of the system in this embodiment. As shown in FIG. 1, in this system, a log management device 100 and an anomaly detection device 200 are provided, and these are connected by a network. Note that the log management device 100 and the anomaly detection device 200 may be realized by one device.
[0015] The log management device 100 can communicate with a plurality of network devices (network device group) via a network. The network devices constituting the network device group are, for example, routers, servers, switches, firewalls, and the like.
[0016] Each network device and the log management device 100 in the present embodiment have the function of syslog, and the log management device 100 may be referred to as a syslog server. However, using syslog is just an example, and the technology according to the present invention is also applicable to log acquisition technologies other than syslog. The anomaly detection device 200 performs anomaly detection based on the time-series data of the increase amount of the file size of the log file.
[0017] The outline of the operation of this system will be described according to the procedure of the flowchart in FIG. 2. In S101, the log management device 100 receives messages (syslog messages) transmitted from each of a plurality of network devices (network device group), and stores the messages as logs in a file for each network device. The file in which the logs are stored will be referred to as a "log file".
[0018] In S102, the log management device 100 periodically (at a predetermined time interval), for example, acquires the file information of each log file from the storage unit, and transfers the acquired file information to the anomaly detection device 200. The file information to be acquired includes the file size of the log file and does not include the information of the content of the log file. Since there is no access to the content of the file, the access time is a fixed short time, and confidentiality can be ensured.
[0019] The anomaly detection device 200 analyzes the file information received from the log management device 100 based on the file size (S103), and performs anomaly detection based on the analysis result (S104). In the present embodiment, since it is not necessary to count the number of lines of the log, anomaly detection can be performed with a small amount of calculation.
[0020] (Specific Examples of Analysis and Abnormality Detection) Referring to FIG. 3, the analysis process and the like will be described more specifically. In the log management apparatus 100 shown in FIG. 3, it is assumed that the logs of each network device are continuously stored in the log files of each network device.
[0021] For each log file, the log management apparatus 100 acquires, as file information, the file name, the file size, and the i-node number. The log management apparatus 100 transmits the "file name, file size, and i-node number" of each log file to the abnormality detection apparatus 200. Note that the i-node number is an example of an identifier that uniquely identifies a file. Here, an identifier that uniquely identifies a file such as the i-node number is referred to as a "file identifier". In the present embodiment, a file identifier other than the i-node number may be used.
[0022] The abnormality detection apparatus 200 performs abnormality detection of the network based on the temporal change of the increase amount per unit time of the file size. The method of abnormality detection is not limited to a specific method. For example, the method disclosed in Non-Patent Document 1, which uses the Bollinger Bands used for analyzing price movements in the financial field for the time-series analysis of the total amount of syslog, can be applied to the time-series analysis of the file size. Also, Spectral Residual (SR), which is one of the methods of time-series abnormality detection, can be used. Hereinafter, the increase amount per unit time is referred to as an increase rate.
[0023] (Regarding the Method of Obtaining Time-Series Data of the Increase Rate of File Size) The abnormality detection apparatus 200 may analyze each log file for a plurality of network devices, or may analyze the total file size of all log files for a plurality of network devices, or may group a plurality of network devices and analyze the total file size of the log files for each group.
[0024] Here, for the sake of convenience, a method for obtaining time-series data of the rate of increase in the file size of one log file will be described. When multiple log files are targeted, the values of the time-series data described below may be summed for the number of multiple log files.
[0025] In the present embodiment, it is assumed that the log management device 100 performs log rotation so that the size of the log file does not become excessive.
[0026] Specifically, for example, the log management device 100 periodically or when the size reaches a threshold value, changes the name of the target log file (for example, "ABCD.log") to another name (for example, "ABCD1.log"), and newly creates a log file with the same name ("ABCD.log"). Logs begin to be stored in the newly created "ABCD.log".
[0027] When log rotation occurs, the inode number corresponding to "ABCD.log" changes. That is, when it is detected that the inode number for a certain file name has changed, it can be determined that log rotation has occurred.
[0028] The anomaly detection device 200 calculates the amount of increase in the file size per predetermined time, that is, the rate of increase, based on the file information periodically received from the log management device 100. In addition, the anomaly detection device 200 performs complementation of the file size considering log rotation.
[0029] FIG. 4 shows an example of the calculation result of the amount of increase in the file size per predetermined time. The middle section of FIG. 4 is an example when complementation by log rotation is not performed, and the lower section of FIG. 4 shows an example when complementation is performed by log rotation.
[0030] The example in Fig. 4 shows the increase in file size every 5 minutes. For example, the value of 1.0 in the column of 00:00 indicates that the increase in file size from 00:00 to 00:05 is 1.0. This value can be obtained, for example, by subtracting the file size at 00:00 from the file size at 00:05.
[0031] When log rotation complementation is not performed, it is shown that the increase amount at 00:15 is 0. This means that, for example, since log rotation occurred between 00:15 and 00:20 and the file size became 0, it was determined that there was no increase in file size from 00:15 to 00:20.
[0032] Based on the file information (file name, file size, i-node number) that the anomaly detection device 200 in this embodiment periodically receives from the log management device 100, when it detects that the i-node number corresponding to the file name has changed from the previous i-node number at a certain point in time, it determines that there has been log rotation and calculates the increase in file size by means of a complementation process.
[0033] For example, assuming that the anomaly detection device 200 receives file information every 5 minutes, it is assumed that a change in the i-node number is detected at the last moment (00:20) in the time interval of 00:15 in Fig. 4 (that is, the time interval from 00:15 to 00:20).
[0034] In this case, it can be estimated that log rotation occurred at some time in the time interval from 00:15 to 00:20. Therefore, the anomaly detection device 200 assumes, for example, that log rotation occurred (the file size became 0) at the midpoint of the time interval from 00:15 to 00:20, and uses the file size at 00:20 to estimate the increase in size from 00:15 to 00:20. For example, assuming that the file size increases at a constant rate from 00:15 to 00:20, twice the file size at 00:20 is taken as the increase in size from 00:15 to 00:20.
[0035] Also, when the period for acquiring (receiving) file information is shorter than 5 minutes, more accurate estimation can be performed. For example, assuming that the abnormality detection device 200 acquires file information every minute, and it detects that the file size at 00:15 is 10, the file size at 00:17 is 12, the i-node number change is detected at 00:18, and the file size at 00:20 is 1, then the size increase amount from 00:15 to 00:20 can be estimated as, for example, (12 - 10) + 1 = 3.
[0036] In the above manner, the abnormality detection device 200 can obtain the time-series change of the size increase rate of the log file.
[0037] (Example of Abnormality Detection) The abnormality detection device 200 can perform abnormality detection using Bollinger Bands in the same manner as the method disclosed in Non-Patent Document 1, for example. For example, the abnormality detection device 200 sets +3σ in the Bollinger Bands as the upper limit value and -3σ as the lower limit value in the time progression of the size increase rate of the log file, and determines that an abnormality has occurred when it detects that the size increase rate exceeds the upper limit value or the size increase amount is below the lower limit value. Note that using 3σ is just an example.
[0038] Referring to FIG. 5, an example of abnormality detection using Bollinger Bands will be described. In the graph of FIG. 5, the line indicating Log rate shows the time-series change of the size increase rate of the file sizes of five network devices (the sum of the sizes of five log files). The upper line above the line indicating Log rate shows the upper limit value, and the lower line below the line indicating Log rate shows the lower limit value.
[0039] In the example of FIG. 5, at the time point indicated by the downward arrow, it is detected that the file size increase rate exceeds the upper limit value by a threshold or more, so an abnormality is detected at this time point.
[0040] After that, for example, the anomaly detection device 200 (or the operator) examines each log file. In this case, as shown in FIG. 6, since the file size increase rate of the A.A.A.A.log was particularly high among the five log files, it can be estimated that there was an attack (e.g., Dos attack, IP scan) on the network device corresponding to A.A.A.A.log, etc.
[0041] After that, for example, a system administrator or the like who is permitted to view the contents of the log file can view the contents of A.A.A.A.log and identify specific attack details, etc.
[0042] Also, when it is detected that the file size increase rate has fallen below the lower limit value, or when it is detected that the lower limit value has fallen below the threshold value by a certain amount, in the same manner as above, it is possible to estimate the cause of equipment failure, communication failure, etc. in a specific section and take corresponding measures.
[0043] It is also possible to perform anomaly detection using the Spectral Residual (SR method). The SR method is one of the anomaly detection methods for time series data. In anomaly detection using the SR method, a saliency map is calculated from the time series data, a score is calculated from the saliency map, and when the score exceeds the threshold value, it is determined as an anomaly.
[0044] An example is shown in FIG. 7. In the example of FIG. 7, points determined as anomalies are shown on a graph indicating the file size increase rate (Log rate). For example, by examining the log file at the time when an anomaly is determined, the specific content of the anomaly can be grasped.
[0045] (Example of grouping) As described above, in the present embodiment, in the log management device 100, a log file is created for each network device in the network device group, and a message from a certain network device is stored in the log file corresponding to the network device. In the present embodiment, as the name of the log file, a name that can identify the corresponding network device (for example, IP address, host name, device number, etc.) is used. By using a name that can identify the corresponding network device as the name of the log file, in the anomaly detection device 200 that receives the file name as file information, it is possible to perform processing based on grouping as shown in the following Examples 1 to 3.
[0046] (Example 1) The anomaly detection device 200 groups the file information for each model such as routers, servers, switches, etc., and for each group, executes processing for anomaly detection based on the above-described file size. For example, as a result of performing anomaly detection processing based on the file size for each group where Group 1 = router, Group 2 = server, and Group 3 = switch, if an anomaly is detected only in Group 1, it can be estimated that there is an anomaly related to the router.
[0047] (Example 2) Here, it is assumed that the file name is the IP address of the network device. The anomaly detection device 200 groups the file information for each range of the IP address that is the file name, and for each group, executes processing for anomaly detection based on the above-described file size. For example, as a result of performing anomaly detection processing based on the file size for each group where Group 1 = range A, Group 2 = range B, and Group 3 = range C, if an anomaly is detected only in Group 2, it can be estimated that there is an anomaly in the network device group with the IP address of range B.
[0048] (Example 3) Here, it is assumed that the vendor of the network device can be identified from the file name. The anomaly detection device 200, for example, groups file information for each vendor and executes processing for anomaly detection based on the above-described file size for each group. For example, as a result of performing anomaly detection processing based on the file size for each group where Group 1 = Vendor A, Group 2 = Vendor B, and Group 3 = Vendor C, if an anomaly is detected only in Group 3, it can be estimated that there is some anomaly in the devices of Vendor C.
[0049] (Example of device configuration) FIG. 8 shows a functional configuration example of the log management device 100. As shown in FIG. 8, the log management device 100 includes a log collection unit 110, a log storage unit 120, and an information acquisition unit 130. The log collection unit 110 receives messages transmitted from each network device and stores them in the log file in the log storage unit 120.
[0050] The information acquisition unit 130 periodically performs processing of acquiring file information of each log file from the log storage unit 120 and transmitting the acquired file information to the anomaly detection device 200.
[0051] FIG. 9 shows a functional configuration example of the anomaly detection device 200. As shown in FIG. 9, the anomaly detection device 200 includes an information acquisition unit 210, a data storage unit 220, an analysis unit 230, and an output unit 240.
[0052] The information acquisition unit 210 receives the file information transmitted from the log management device 100 and stores it in the data storage unit 220. The analysis unit 230 calculates time-series data of the file size increase rate from the file information stored in the data storage unit 220 and performs anomaly detection using the time-series data. The output unit 240 outputs the analysis result (anomaly detection result) by the analysis unit 230.
[0053] A system including a log management device 100 and an anomaly detection device 200 may be referred to as an anomaly detection system. Even in the case where the log management device 100 and the anomaly detection device 200 are configured as one device, the said device may also be referred to as an anomaly detection system.
[0054] In the anomaly detection system, there may be provided a log storage unit that stores a log file recording logs output from network devices, an information acquisition unit that periodically acquires file information including the file size and file name of the log file from the log storage unit, and an analysis unit that calculates time-series data of the increase rate of the file size of the log file from the file information and performs anomaly detection from the time-series data.
[0055] When the log management device 100 and the anomaly detection device 200 are configured as one device, the transmission of file information from the log management device 100 to the anomaly detection device 200 is a process of passing file information within the device.
[0056] (Hardware configuration example of the device) The log management device 100, the anomaly detection device 200, and the anomaly detection system can all be realized, for example, by causing one or more computers to execute a program. The computer may be a physical machine or a virtual machine on the cloud. Hereinafter, the log management device 100, the anomaly detection device 200, and the anomaly detection system are collectively referred to as "devices".
[0057] FIG. 10 is a diagram showing a hardware configuration example of the above computer in the present embodiment. Note that when the above computer is a virtual machine, the hardware configuration is a virtual hardware configuration. The computers in FIG. 10 each have a drive device 1000, an auxiliary storage device 1002, a memory device 1003, a CPU 1004, an interface device 1005, a display device 1006, an input device 1007, an output device 1008, etc., which are mutually connected by a bus B.
[0058] A program for realizing processing on the computer is provided, for example, by a recording medium 1001 such as a CD-ROM or a memory card. When the recording medium 1001 storing the program is set in the drive device 1000, the program is installed from the recording medium 1001 via the drive device 1000 into the auxiliary storage device 1002. However, the program does not necessarily have to be installed from the recording medium 1001, and it may be downloaded from another computer via a network. The auxiliary storage device 1002 stores the installed program and also stores necessary files, data, etc.
[0059] When an instruction to start the program is given, the memory device 1003 reads out and stores the program from the auxiliary storage device 1002. The CPU 1004 realizes the functions related to the device according to the program stored in the memory device 1003. The interface device 1005 is used as an interface for connecting to a network. The display device 1006 displays a GUI (Graphical User Interface) etc. by the program. The input device 1007 is composed of a keyboard, a mouse, buttons, or a touch panel, etc., and is used to input various operation instructions. The output device 1008 outputs the calculation result.
[0060] (Effects of the Embodiment) According to the technology related to the present embodiment described above, it is possible to perform anomaly detection using the log of messages output from network devices with a small amount of calculation while maintaining confidentiality.
[0061] <Supplementary Note> This specification discloses at least an anomaly detection system, an anomaly detection method, and a program according to the following respective items. (Item 1) A log storage unit that stores a log file recording logs output from a network device, An information acquisition unit that periodically acquires file information including the file size and file name of the log file from the log storage unit, An analysis unit that calculates time-series data of the increase rate of the file size of the log file from the file information and performs anomaly detection from the time-series data; An anomaly detection system comprising the above. (Item 2) The file information further includes a file identifier, and the analysis unit detects the occurrence of log rotation based on a change in the file identifier and executes complementation of the increase rate. The anomaly detection system according to Item 1. (Item 3) The information acquisition unit periodically acquires file information including the file size and file name from each of a plurality of log files corresponding to a plurality of network devices. The analysis unit calculates time-series data of the increase rate of the file size in the entire plurality of log files and performs anomaly detection from the time-series data. The anomaly detection system according to Item 1 or Item 2. (Item 4) The analysis unit groups the plurality of log files based on the file name of each log file in the plurality of log files and performs anomaly detection for each group. The anomaly detection system according to Item 3. (Item 5) An anomaly detection method executed by an anomaly detection system, comprising: An information acquisition step of periodically acquiring file information including the file size and file name of the log file from a log storage unit that stores a log file recording a log output from a network device; An analysis step of calculating time-series data of the increase rate of the file size of the log file from the file information and performing anomaly detection from the time-series data. An anomaly detection method comprising the above. (Item 6) A program for causing a computer to function as each unit in the anomaly detection system according to any one of Items 1 to 4.
[0062] As described above, the present embodiment has been explained. However, the present invention is not limited to such a specific embodiment, and various modifications and changes are possible within the scope of the gist of the present invention described in the claims.
Explanation of Signs
[0063] 100 Log management device 110 Log collection unit 120 Log storage unit 130 Information acquisition unit 200 Abnormality detection device 210 Information acquisition unit 220 Data storage unit 230 Analysis unit 240 Output unit 1000 Drive device 1001 Recording medium 1002 Auxiliary storage device 1003 Memory device 1004 CPU 1005 Interface device 1006 Display device 1007 Input device 1008 Output device
Claims
1. A log storage unit that stores a log file recording logs output from a network device, an information acquisition unit that periodically acquires file information including the file size and file name of the log file from the log storage unit, and an analysis unit that calculates time-series data of the increase rate of the file size of the log file from the file information and performs anomaly detection from the time-series data. An anomaly detection system comprising the above.
2. The file information further includes a file identifier, and the analysis unit detects the occurrence of log rotation based on the change in the file identifier and executes complementation of the increase rate. The anomaly detection system according to Claim 1.
3. The information acquisition unit periodically acquires file information including the file size and file name from each of a plurality of log files corresponding to a plurality of network devices, and the analysis unit calculates time-series data of the increase rate of the file size for the entire plurality of log files and performs anomaly detection from the time-series data. The anomaly detection system according to Claim 1 or 2.
4. The analysis unit groups the plurality of log files based on the file name of each log file in the plurality of log files and performs anomaly detection for each group. The anomaly detection system according to Claim 3.
5. An anomaly detection method executed by an anomaly detection system, an information acquisition step of periodically acquiring file information including the file size and file name of the log file from a log storage unit that stores a log file recording logs output from a network device, and an analysis step of calculating time-series data of the increase rate of the file size of the log file from the file information and performing anomaly detection from the time-series data. An anomaly detection method comprising the above.
6. A program for causing a computer to function as each unit in the anomaly detection system according to any one of Claims 1 to 4.
Citation Information
Patent Citations
Data management method and device
JP2005031715A
Log management system and log display system
JP2010039878A
Failure prediction management method and computer system applied with the same
JP2010102548A
Automatic rotation setting system and method for log file
JP2012159926A
Database testing method
JP2015172865A