Anomaly detection device and anomaly detection method
Patent Information
- Application Number
- PCT/JP2025/012485
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-03-27
- Publication Date
- 2026-10-01
Smart Images

Figure JP2025012485_01102026_PF_FP_ABST
Abstract
Description
Anomaly detection device and anomaly detection method
[0001] The present invention relates to an anomaly detection device and an anomaly detection method.
[0002] There are systems that analyze logs output by systems, software, or devices to detect anomalies. For example, there are known techniques that use template portions of unstructured data extracted by log parsing as input data, learn log occurrence patterns and trends using machine learning models, and determine whether the system is normal or abnormal (see, for example, Non-Patent Documents 1-3).
[0003] Conventional models used for log anomaly detection include, for example, DeepLog (Non-Patent Document 1) or LogAnomaly (Non-Patent Document 2), which use LSTM (Long Short-Term Memory) to predict the type of log template, or NeuralLog (Non-Patent Document 3), which uses a Transformer to predict whether a log is normal or abnormal.
[0004] M. Du, F. Li, G. Zheng and V. Srikumar, "DeepLog: Anomaly Detection and Diagnosis from System Logs through Deep Learning", Proceedings of the 2017 ACM SIGSAC Conference on Computer and Communications Security, pp. 1285-1298, Oct. 2017.R. Zhou, P. Sun, S. Tao, R. Zhang, W. Meng, Y. Liu, et IEEE, pp.492-504, 2021.
[0005] Conventional log anomaly detection methods capture log appearance patterns and trends, but do not consider time. Therefore, for example, if there are three types of log templates A, B, and C, even if the same appearance pattern is A → B → C, there is a problem in that anomaly detection cannot be performed while taking into account differences in appearance intervals or appearance times for each log template.
[0006] The embodiments of this disclosure have been made in view of the above-mentioned problems, and provide an anomaly detection system that uses a machine learning model to detect anomalies in logs, and enable anomalies to be detected by taking into account the time when the logs appear.
[0007] To solve the above problems, an anomaly detection device according to the embodiment of this disclosure is an anomaly detection device that detects anomalies in logs using a machine learning model, and comprises: an input calculation unit that calculates input data to be input to the machine learning model based on template information of the unstructured portion of the log data after log parsing and time information of the log data; and a loss calculation unit that calculates a loss based on output data output by the machine learning model that has been input the input data, and learns the machine learning model so that the loss becomes smaller.
[0008] According to embodiments of this disclosure, an anomaly detection system that detects log anomalies using a machine learning model can detect anomalies by taking into account the time at which the log appears.
[0009] This figure shows an example configuration of the anomaly detection system according to this embodiment. This flowchart shows an example of processing by the anomaly detection device according to Embodiment 1. This flowchart shows an example of processing by the anomaly detection device according to Embodiment 2. This figure shows an example of the computer hardware configuration. This figure shows the flow of automatic log analysis. This figure shows an example of log parsing.
[0010] Hereinafter, embodiments of the present invention (this embodiment) will be described with reference to the drawings. The embodiments described below are merely examples, and the embodiments to which the present invention is applied are not limited to the embodiments described below.
[0011] <Overview> This embodiment relates to an anomaly detection device, an anomaly detection system, an anomaly detection method, and a program that use a machine learning model to detect anomalies in logs.
[0012] (Conventional Technology) Logs, which are semi-structured text output from systems, software, or devices, are important data that records information about their execution and allows for confirmation of their state or operation during execution. Therefore, by analyzing logs, system failures and other problems can be detected early. However, logs obtained from large and complex systems, for example, are often voluminous and cumbersome, and there is a need to streamline log analysis, so research into automated log analysis is progressing (see, for example, Non-Patent Documents 1-3).
[0013] Figure 5 shows the flow of automatic log analysis. Automatic log analysis, for example as shown in Figure 5, acquires logs output by a system, software, or device 1, and outputs detection results indicating normality or abnormality through processes such as log parsing 2 and log anomaly detection 3.
[0014] Log Parsing 2 decomposes the semi-structured text of each log line into structured parts, which are known from, for example, system specifications and can be decomposed using regular expressions, and unstructured parts, which are written as sentences and are not structured. Furthermore, the unstructured parts are parsed to decompose them into template parts, which are fixed text, and parameter parts, which are to which variables, etc.
[0015] Figure 6 shows an example of log parsing. Figure 6 shows an example of log data 20 obtained by log parsing log 10. In the example in Figure 6, the log data 20 obtained by log parsing is decomposed into a structured part 21 and an unstructured part 22, and the unstructured part 22 is decomposed into a template 23 and parameters 24. The unstructured part 22 of the log data obtained by log parsing is the part that the system developer can write in free form in natural language, and the template 23 contains particularly important information from the log data obtained by log parsing 20.
[0016] In log anomaly detection 3, the template 23 of the unstructured portion 22 extracted by log parsing 2 is used as input data, and a machine learning model is used to learn the occurrence patterns and / or trends of the logs, and for example, to determine whether the system is normal or abnormal. Existing machine learning models used for log anomaly detection include DeepLog (see, for example, Non-Patent Document 1), which uses LSTM (Long Short-Term Memory) to predict the type of log template, or LogAnomaly (see, for example, Non-Patent Document 2). Existing machine learning models used for log anomaly detection include NeuralLog (see, for example, Non-Patent Document 3), which uses Transformer to predict whether the log is normal or abnormal.
[0017] (Problem) Conventional log anomaly detection captures the occurrence patterns, trends, and the like of logs, but does not take time into consideration. For example, assume that there are three types of log templates A, B, and C, and logs may occur in the order of A→B→C. Furthermore, there are cases where it is desired to perform learning such that when both the time intervals of A→B and B→C are 1 minute, and when A→B has an interval of 1 minute and B→C has an interval of 5 minutes, the latter is determined to be anomalous because it takes a long time before the log template C is output. However, since conventional log anomaly detection does not use time information, both cases are recognized as the same log occurrence pattern.
[0018] Furthermore, for the occurrence pattern in the order of A→B→C, it is assumed that it is normal for A to occur in the 10-minute range of each hour (0:10, 1:10, etc.), B in the 15-minute range, and C in the 25-minute range, and no logs occur in the order of A→B→C outside these times. In this case, for example, when A occurs at 1:35, B at 1:40, and C at 1:50, it is desired to detect this as an anomaly, but conventional log anomaly detection determines it as normal. As described above, conventional techniques have a problem in that they cannot detect log anomalies by taking into consideration the time at which logs occur.
[0019] (Outline of Processing) Therefore, the anomaly detection system according to the present embodiment changes input data to be input to a machine learning model to input data including time information, by using the time information of log data in addition to template information of an unstructured portion of log data that has been subjected to log parsing. This enables anomaly detection that takes into consideration the time at which logs occur, and improves the accuracy of anomaly detection.
[0020] <Configuration Example> FIG. 1 is a diagram showing a configuration example of the anomaly detection system according to the present embodiment. In the example of FIG. 1, an anomaly detection system 100 includes an anomaly detection apparatus 110. The anomaly detection apparatus 110 is an information processing apparatus having a computer configuration, or a system including a plurality of computers.
[0021] The abnormality detection device 110 implements each functional configuration as shown in, for example, FIG. 1 by executing a predetermined program on a computer included in the abnormality detection device 110. In the example of FIG. 1, the abnormality detection device 110 includes a database 111, an input calculation unit 112, a machine learning model 113, a loss calculation unit 114, a data processing unit 115, an abnormality detection unit 116, and the like. However, the configuration of the abnormality detection system 100 shown in FIG. 1 is an example. For example, in FIG. 1, each functional configuration included in the abnormality detection device 110 may be distributed and provided in a plurality of information processing devices.
[0022] The database 111 stores parsed log data obtained by performing log parsing in advance. The parsed log data includes, for example, N lines (N is an integer of 1 or more) of parsed log data for learning (hereinafter referred to as learning data) and M lines (M is an integer of 1 or more) of parsed log data for testing (hereinafter referred to as test data). The database 111 also stores the calculation result of the input calculation unit 112, the calculation result of the loss calculation unit 114, model parameters of the machine learning model 113, and the like. As described above, the database 111 according to the present embodiment functions as a storage unit that stores various information, data, or the like.
[0023] The input calculation unit 112 calculates input data to be input to the machine learning model 113 based on information of the template 23 of the unstructured portion 22 of the log data 20 subjected to log parsing and time information of the log data 20, and executes input calculation processing. In the present embodiment, the input data is changed from that in conventional log abnormality detection, and the input data is obtained by adding a vector of time information to the input vector of conventional log abnormality detection.
[0024] Specifically, in the conventional technology, when creating data from a log in a certain n-th line (n is an integer of 1 or more), as shown in formula (1), a numerical vector W obtained by converting information of the template 23 of the unstructured portion 22 of the log data 20 subjected to log parsing n was created (D is the number of dimensions).
[0025] On the other hand, in this embodiment, as shown in equation (2), the numerical vector W n Next, a numerical vector T is obtained by transforming the time column of the structured portion 21 of the log data 20 that has undergone log parsing. n Adding W n +T n Create.
[0026] Thus, the input data calculated by the input calculation unit 112 according to this embodiment is a first numerical vector W obtained by converting information from multiple templates. n And a second numerical vector T obtained by transforming information from multiple time points. n This includes [the above]. A specific example of calculations using the input data will be discussed later.
[0027] The machine learning model 113 is a neural network that, for example, takes input data calculated by the input calculation unit 112 as input and outputs detection results indicating whether the log is normal or abnormal. For example, the machine learning model 113 can be modified to use LSTM (see, for example, Non-Patent Documents 1 and 2) or Transformer (see, for example, Non-Patent Document 3), which are conventionally used for log anomaly detection.
[0028] The loss calculation unit 114 calculates the loss based on the output data output by the machine learning model 113 that has been input with the input data, and performs a loss calculation process to train the machine learning model 113 so that the loss becomes smaller. For example, the loss calculation unit 114 updates the model parameters of the machine learning model 113 so that the calculated loss becomes smaller. For example, the loss calculation method used by the loss calculation unit 114 can be a known cross-entropy or mean squared error.
[0029] The data processing unit 115 performs data processing, such as log parsing. For example, as explained in Figure 6, the data processing unit 115 creates log data 20 by parsing the log 10. As a specific example, the data processing unit 115 decomposes the log 10 into a structured part 21 and an unstructured part 22, and decomposes the unstructured part 22 into a template 23 and parameters 24 to create log-parsed log data 20.
[0030] The abnormality detection unit 116 uses the trained machine learning model 113 to execute abnormality detection processing that outputs a detection result indicating whether the log 10 subject to abnormality detection is normal or abnormal. For example, the abnormality detection unit 116 parses the target log 10 using the data processing unit 115, and obtains W serving as input data from the parsed log data 20 using the input calculation unit 112 n +T n is calculated. Further, the abnormality detection unit 116 calculates W based on the target log 10 n +T n is input to the trained machine learning model 113, whereby for example, a detection result indicating whether the target log 10 output by the machine learning model 113 is normal or abnormal is acquired and output.
[0031] <Processing Flow> Next, the flow of processing of the abnormality detection method according to the present embodiment will be described.
[0032] [Example 1] FIG. 2 is a flowchart showing an example of processing of the abnormality detection apparatus according to Example 1. This processing shows an example of learning processing executed by the abnormality detection apparatus 110 having each functional configuration described with reference to FIG. 1, for example.
[0033] In step S201, the input calculation unit 112 acquires, from the database 111, parsed log data 20 as illustrated in FIG. 6, for example.
[0034] In step S202, the input calculation unit 112 acquires, from the acquired parsed log data 20, a first numerical vector W obtained by converting information of a plurality of templates n and a second numerical vector T obtained by converting information of a plurality of times n input data including (W n +T n ) is calculated.
[0035] (Calculation Example) Numerical vector W obtained by converting information of template 23 nThis is calculated in the same way as the conversion method in conventional models. For example, the input calculation unit 112 uses one-hot encoding, word2vec, or embedding using a language model such as BERT to calculate the numerical vector W using equation (1). n The input calculation unit 112 calculates W for all lines of the learning log in advance. n Calculate the average W of the elements of those vectors. mean and standard deviation W std This is calculated using the following equations (3) and (4).
[0036]
[0037] Furthermore, the input calculation unit 112 converts the time sequence (information) of the structured portion 21 into a numerical vector T. n This is calculated using the following equations (5) and (6).
[0038]
[0039] Here, i = 0, 1, ..., (D / 2) - 1, where τ is a numerical value representing the period and t is a numerical value representing the time. t is calculated from the time column; for example, in the example of log data 20 after log parsing in Figure 6, the "time" column of the structured part 21 is used.
[0040] For example, if the period is a daily period (24-hour period) and measured in seconds, τ is calculated using the following formula.
[0041] τ = 24 × 60 × 60 = 86400
[0042] Furthermore, if t is calculated using the following formula, where hours, minutes, and seconds are represented as hours, minutes, and seconds respectively in "time".
[0043] t = 3600 × hour + 60 × min + sec
[0044] As another example, if the period is a time period (60-minute period) in seconds, τ is calculated using the following formula.
[0045] τ = 60 × 60 = 3600
[0046] Furthermore, if t represents minutes and seconds in "time" as min and sec respectively, it is calculated using the following formula.
[0047] t = 60 × min + sec
[0048] The same calculation is performed even if the unit of period or time is different. Using equations (5) and (6), the periodic time is given by W. n It is expressed on the same level.
[0049] Note that in equations (5) and (6), T n The value of W n It may be smaller than the value of W n +T n The value of W n The value may not change much, so T n By setting the elements of the following equations (7) and (8), T n The mean and standard deviation are W n You may also align them to the mean and standard deviation.
[0050]
[0051] Furthermore, W in equations (7) and (8) mean , W std Change to any number, T n The mean and standard deviation of can be changed. This allows us to do the same as in equations (7) and (8), T n The size (weighting) can be changed.
[0052] Next, the loss calculation unit 114 trains the machine learning model 113 using the input data calculated by the input calculation unit 112. For example, the loss calculation unit 114 executes the processes in steps S203 to S205.
[0053] In step S203, the loss calculation unit 114 inputs the input data to the machine learning model 113 and obtains the output data output by the machine learning model 113.
[0054] In step S204, the loss calculation unit 114 calculates the loss based on the acquired output data and trains the machine learning model 113 to reduce the calculated loss (updates the model parameters of the machine learning model 113).
[0055] For example, in DeepLog (see Non-Patent Document 1, for example), one of the conventional log anomaly detection methods, a numerical vector W is obtained by transforming the template portion of the log from a certain k line to k+j lines. k , ..., W k+j The data is input to the LSTM, and the output is a vector whose dimensions are the number of types of templates in the training data (let's call it C), where each element takes a value between 0 and 1, and the sum of the elements is 1. The machine learning model 113 is trained to minimize the loss calculated by taking a one-hot vector of dimension C representing the type of template in the log on the k+j+1th line and the mean squared error of the output, so that this output predicts the type of template in the log on the k+j+1th line.
[0056] In this embodiment, the numerical vector W k , ..., W k+j Instead, W k +T k , ..., W k+j +T k+j This will be used as input data.
[0057] In step S205, the loss calculation unit 114 determines whether a predetermined number of learning iterations has been reached. If the predetermined number of learning iterations has not been reached, the loss calculation unit 114 returns to step 203. On the other hand, if the predetermined number of learning iterations has been reached, the loss calculation unit 114 terminates the process shown in Figure 2. Note that the process in step S205 is just one example. The loss calculation unit 114 may also terminate the process shown in Figure 2, for example, when the loss reaches a target value or becomes the minimum value.
[0058] As shown in Figure 2, the anomaly detection device 110 uses not only the template portion but also the time column of the structured portion as input for log anomaly detection to train the machine learning model 113. This enables the anomaly detection system 100 to perform anomaly detection that takes into account the time when the log appears, thereby improving the accuracy of anomaly detection.
[0059] [Example 2] Figure 3 is a flowchart showing an example of the processing of the anomaly detection device according to Example 2. This processing shows another example of the learning process performed by the anomaly detection device 110 having the functional configurations described in Figure 1. The basic processing content is the same as that of Example 1 described in Figure 2, so a detailed explanation of the processing similar to that of Example 1 is omitted here.
[0060] In step S301, the input calculation unit 112 obtains parsed log data 20 from the database 111, for example, as shown in Figure 6.
[0061] In step S302, the input calculation unit 112 converts the information of multiple templates into a first numerical vector W from the acquired parsed log data 20. n And a second numerical vector T obtained by transforming information from multiple time points. n Input data W, which includes n +T n The following is calculated. Note that the processing in steps S301 and S302 may be the same as in Example 1.
[0062] In step S302, the loss calculation unit 114 inputs the input data into the machine learning model 113 and obtains output data with increased dimensionality from the output output of the machine learning model 113.
[0063] In Example 2, the dimensionality of the output of the machine learning model 113 is increased by L from Example 1. For example, by setting L = D, in addition to the output of the conventional machine learning model 113 (e.g., the type of log template in DeepLog or LogAnomaly, or normal / abnormal in NeuralLog), T n Output in a way that allows for reconstruction. For example, in the DeepLog example, the machine learning model 113 outputs a C-dimensional vector followed by T n It outputs a C+D dimension vector, which is formed by adding a D-dimensional vector that reconstructs the original.
[0064] In step S304, the loss calculation unit 114 calculates the loss based on the acquired output data and trains the machine learning model 113 to reduce the calculated loss (updates the model parameters of the machine learning model 113).
[0065] At this time, the loss calculation unit 114 assigns T to the label corresponding to the added dimension in the output of the machine learning model 113. n The loss is calculated using the following. For example, the machine learning model 113 may output t with L=1, or it may output log time, minutes, and seconds with L=3. In this case, the loss calculation unit 114 uses t, or log time, minutes, and seconds, as labels corresponding to the added dimension in the output of the machine learning model 113.
[0066] In step S305, the loss calculation unit 114 determines whether a predetermined number of learning iterations has been reached. If the predetermined number of learning iterations has not been reached, the loss calculation unit 114 returns to step 303. On the other hand, if the predetermined number of learning iterations has been reached, the loss calculation unit 114 terminates the process shown in Figure 3. Note that the process in step S305 is just one example. The loss calculation unit 114 may also terminate the process shown in Figure 2, for example, when the loss reaches a target value or becomes the minimum value.
[0067] Thus, in Example 2, the machine learning model 113 is modified to increase the output dimension and output numerical values related to time, thereby enabling the calculation of loss for time as well, and allowing the machine learning model 113 to better reflect time information. Note that the machine learning model 113 may also be configured with L=0, in which case it will be the same as in Example 1.
[0068] <Hardware Configuration> The anomaly detection device 110 according to this embodiment can be realized, for example, by having a computer execute a program. This computer may be a physical computer (physical machine) or a virtual machine on the cloud, etc.
[0069] In other words, the anomaly detection device 110 can be realized by using the CPU (Central Processing Unit) and other hardware resources such as memory built into the computer to execute a program corresponding to the processing performed by the device. The above program can be recorded on a computer-readable recording medium (such as portable memory) and saved or distributed. It is also possible to provide the above program via a network such as the Internet or email.
[0070] Figure 4 shows an example of the hardware configuration of the computer described above. In the example in Figure 4, the computer 400 has a drive device 1000, an auxiliary storage device 1002, a memory device 1003, a CPU 1004, an interface device 1005, a display device 1006, an input device 1007, and an output device 1008, all of which are interconnected by bus B. The computer 400 may also be equipped with other processors such as a GPU (Graphics Processing Unit).
[0071] The program that enables processing on the computer 400 is provided on a recording medium 1001, such as a CD-ROM or memory card. When the recording medium 1001 containing the program is set in the drive device 1000, the program is installed from the recording medium 1001 to the auxiliary storage device 1002 via the drive device 1000. However, the program does not necessarily have to be installed from the recording medium 1001; it may also be downloaded from another computer via a network. The auxiliary storage device 1002 stores the installed program as well as necessary files and data.
[0072] The memory device 1003 reads and stores a program from the auxiliary storage device 1002 when a program startup command is received. The CPU 1004 implements the functions related to the abnormality detection device 110 described in this embodiment according to the program stored in the memory device 1003. The interface device 1005 is used as an interface for connecting to a network, etc. The display device 1006 displays a GUI (Graphical User Interface) etc. based on a program. The input device 1007 consists of a keyboard, mouse, buttons, and / or touch panel etc., and is used to input various operation commands. The output device 1008 outputs the calculation results.
[0073] <Supplement> The functions of the elements disclosed herein may be implemented using circuits or processing circuitry that include general-purpose processors, special-purpose processors, integrated circuits, ASICs (Application Specific Integrated Circuits), FPGAs (Field Programmable Gate Arrays), conventional circuits, and / or combinations thereof that are programmed using one or more programs stored in one or more memories, or otherwise configured to perform the disclosed functions. A processor is considered processing circuitry or circuitry because it includes transistors and other circuits. A processor may be a programmed processor that executes programs stored in memory. In this disclosure, a circuit, unit, or means is hardware that performs the enumerated functions, or hardware programmed to perform the enumerated functions. Hardware may be any hardware disclosed herein that is programmed or configured to perform the enumerated functions.
[0074] There is a memory for storing a computer program that includes computer instructions. These computer instructions provide logic and routines that enable hardware (e.g., processing circuitry or circuitry) to perform the methods disclosed herein. This computer program can be implemented in commonly known forms, such as computer-readable storage media, computer program products, memory devices, recording media such as CD-ROMs and DVDs, and / or memory in FPGAs and ASICs.
[0075] <Effects of the Embodiment> According to the anomaly detection device 110 of this embodiment, in an anomaly detection system 100 that detects log anomalies using a machine learning model 113, it becomes possible to detect anomalies by taking into account the time when the log appears, thereby improving the accuracy of anomaly detection.
[0076] The following additional information is disclosed regarding the embodiments described above.
[0077] <Notes> (Note 1) An anomaly detection device that detects anomalies in logs using a machine learning model, comprising: an input calculation unit that calculates input data to be input to the machine learning model based on template information of the unstructured portion of log data after log parsing and time information of the log data; and a loss calculation unit that calculates a loss based on output data output by the machine learning model that has been input the input data, and learns the machine learning model so that the loss becomes smaller. (Note 2) The anomaly detection device according to Note 1, wherein the input data includes a first numerical vector obtained by converting the information of a plurality of templates and a second numerical vector obtained by converting the information of a plurality of times. (Note 3) The anomaly detection device according to Note 1 or 2, wherein the machine learning model further outputs the time information by increasing the dimension of the output data. (Note 4) An anomaly detection method comprising: an anomaly detection device that uses a machine learning model to detect anomalies in logs, which performs an input calculation process that calculates input data to be input to the machine learning model based on template information of the unstructured portion of the log data after log parsing and time information of the log data; and a loss calculation process that calculates a loss based on output data output by the machine learning model that has been input with the input data and ground truth data, and learns the machine learning model so that the loss becomes smaller. (Note 5) An anomaly detection system that uses a machine learning model to detect anomalies in logs, comprising: an input calculation unit that calculates input data to be input to the machine learning model based on template information of the unstructured portion of the log data after log parsing and time information of the log data; and a loss calculation unit that calculates a loss based on output data output by the machine learning model that has been input with the input data, and learns the machine learning model so that the loss becomes smaller.(Appendix 6) A program, or a storage medium storing a program, that causes an anomaly detection device that detects anomalies in logs using a machine learning model to execute: an input calculation process that calculates input data to be input to the machine learning model based on template information of the unstructured portion of the log data after log parsing and time information of the log data; and a loss calculation process that calculates a loss based on output data output by the machine learning model that has been input with the input data and the correct data, and trains the machine learning model so that the loss becomes smaller.
[0078] Although this embodiment has been described above, the present invention is not limited to this specific embodiment, and various modifications and changes are possible within the scope of the gist of the invention as described in the claims.
[0079] 20 Log data 22 Unstructured portion 23 Template 100 Anomaly detection system 110 Anomaly detection device 111 Database 112 Input calculation unit 113 Machine learning model 114 Loss calculation unit 115 Data processing unit 116 Anomaly detection unit
Claims
1. An anomaly detection device that detects anomalies in logs using a machine learning model, comprising: an input calculation unit that calculates input data to be input to the machine learning model based on template information of the unstructured portion of log data after log parsing and time information of the log data; and a loss calculation unit that calculates a loss based on output data output by the machine learning model that has been input the input data, and trains the machine learning model so that the loss becomes smaller.
2. The anomaly detection device according to claim 1, wherein the input data includes a first numerical vector obtained by converting information from a plurality of templates and a second numerical vector obtained by converting information from a plurality of time points.
3. The anomaly detection device according to claim 1 or 2, wherein the machine learning model further outputs time information by increasing the dimensionality of the output data.
4. An anomaly detection method comprising: an anomaly detection device that detects anomalies in logs using a machine learning model, which performs an input calculation process to calculate input data to be input to the machine learning model based on template information of the unstructured portion of the log data after log parsing and time information of the log data; and a loss calculation process that calculates a loss based on output data output by the machine learning model that has been input with the input data and the correct data, and trains the machine learning model so that the loss becomes smaller.