Detection device and detection method
Patent Information
- Application Number
- JP2025557438
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Filing Date
- 2023-11-16
- Publication Date
- 2025-05-22
AI Technical Summary
Existing anomaly detection methods, such as those using Bidirectional Encoder Representations from Transformers (BERT) for packet analysis, struggle to accurately detect anomalies in packet contents.
A detection device and method that convert packets into feature vectors, use a detection model for primary evaluation, extract similar normal packets using a natural language processing model, and perform a secondary evaluation based on anomaly scores calculated by comparing these packets.
This approach enables high-accuracy detection of anomalies in packet contents by narrowing down anomaly candidates through primary and secondary evaluations, reducing false positives and improving processing speed.
Abstract
Description
Detection device and detection method
[0001] The present invention relates to a detection device and a detection method.
[0002] The total number and methods of cyber attacks are increasing year by year, and in order to respond to the new attacks that are constantly being created, attention is being paid to communication anomaly detection technology using unsupervised deep learning, which automatically learns the characteristics of normal communication from the network, detects data with deviating characteristics as anomalies, and outputs an alert.
[0003] For example, one anomaly detection method is to apply natural language processing techniques such as BERT (Bidirectional Encoder Representations from Transformers) to packet analysis to extract information from the payload of any protocol and perform anomaly detection (Patent Document 1).
[0004] International Publication No. 2022 / 059209
[0005] However, although the method described in Patent Document 1 operates at high speed, it has a problem in that it is not possible to accurately detect abnormalities occurring in the contents of packets.
[0006] The present invention has been made in view of the above, and has an object to provide a detection device and a detection method that can detect anomalies occurring in packet content with high accuracy.
[0007] In order to solve the above-mentioned problems and achieve the object, the detection device is characterized by having a conversion unit that converts a packet to be evaluated into a feature vector, a first detection unit that detects whether the packet to be evaluated is anomalous based on the result of inputting the feature vector into a detection model, an extraction unit that acquires an abnormal candidate packet that indicates the packet to be evaluated for which an abnormality has been detected by the first detection unit, and extracts a predetermined number of similar normal packets that have a relatively high similarity to the abnormal candidate packet from among a plurality of normal packets based on a natural language processing model, and a second detection unit that detects whether the abnormal candidate packet is anomalous based on an anomaly score calculated by comparing the predetermined number of similar normal packets extracted by the extraction unit with the abnormal candidate packet.
[0008] According to the present invention, anomalies occurring in packet contents can be detected with high accuracy.
[0009] FIG. 1 is a diagram illustrating an example of the configuration of a detection system according to an embodiment. FIG. 2 is a diagram illustrating an example of the configuration of a learning device shown in FIG. 1. FIG. 3 is a flowchart illustrating the procedure of learning processing performed by the learning device shown in FIG. 2. FIG. 4 is a diagram illustrating a primary evaluation performed by the detection device. FIG. 5 is a diagram illustrating a secondary evaluation performed by the detection device (1). FIG. 6 is a diagram illustrating a secondary evaluation performed by the detection device (2). FIG. 7 is a diagram illustrating an example of the configuration of the detection device shown in FIG. 1. FIG. 8 is a diagram illustrating a first calculation process. FIG. 9 is a flowchart illustrating the procedure of detection processing according to an embodiment. FIG. 10 is a flowchart illustrating the procedure of primary evaluation. FIG. 11 is a flowchart illustrating the procedure of secondary evaluation. FIG. 12 is a flowchart illustrating the procedure of first calculation processing. FIG. 13 is a flowchart illustrating the procedure of second calculation processing. FIG. 14 is a diagram illustrating the results of an evaluation experiment. FIG. 15 is a diagram illustrating an example of a computer that implements a learning device and a detection device by executing a program.
[0010] DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS A detailed description of embodiments of a detection device and a detection method according to the present invention will be given below with reference to the accompanying drawings. However, the present invention is not limited to the embodiments described below.
[0011] A detection device according to an embodiment performs a primary evaluation, which is a lightweight process, on evaluation target packets to narrow down the anomaly candidates. The detection device then performs a secondary evaluation on the evaluation target packets that have been determined to be anomaly candidates as a result of the primary evaluation to determine whether they are anomalies.
[0012] (Detection System) Next, a description will be given of a detection system according to Embodiment 1. Fig. 1 is a diagram showing an example of the configuration of a detection system according to an embodiment.
[0013] 1, the detection system 1 includes a learning device 10 and a detection device 20. The learning device 10 is a device that executes learning of a model for the detection device 20 to perform primary and secondary evaluations. The detection device 20 executes primary and secondary evaluations using the model that has been trained by the learning device 10, and detects anomalies in communication data (payload).
[0014] (Configuration of the learning device) An example of the configuration of the learning device 10 will be described. Fig. 2 is a diagram showing an example of the configuration of the learning device 10 shown in Fig. 1. As shown in Fig. 2, the learning device 10 has an input unit 11, an output unit 12, a communication unit 13, a storage unit 14, and a control unit 15.
[0015] Input unit 11 is an input interface that accepts various operations from the operator of study device 10. For example, input unit 11 is configured with input devices such as a touch panel, a voice input device, a keyboard, or a mouse.
[0016] The output unit 12 is realized by, for example, a display device such as a liquid crystal display, a printing device such as a printer, an information communication device, or the like.
[0017] The communication unit 13 is a communication interface that transmits and receives various information to and from other devices connected via a network or the like. The communication unit 13 is implemented by a network interface card (NIC) or the like, and performs communication between other devices and the control unit 15 (described later) via telecommunication lines such as a local area network (LAN) or the Internet. For example, the communication unit 13 receives normal packets, which are learning target data, via the network and outputs them to the control unit 15. The communication unit 13 also outputs a trained detection model, parameters of a conversion model (encoder) that has learned rules for converting data into fixed-length feature vectors, and a comparison normal packet group 141 to the detection device 20 via the network.
[0018] The storage unit 14 is a storage device such as a hard disk drive (HDD), a solid state drive (SSD), or an optical disk. The storage unit 14 may also be a rewritable semiconductor memory such as a random access memory (RAM), a flash memory, or a non-volatile static random access memory (NVSRAM). The storage unit 14 stores an operating system (OS) and various programs executed by the learning device 10. Furthermore, the storage unit 14 stores various information used in the execution of the programs. For example, the storage unit 14 has a comparison normal packet group 141.
[0019] The comparison normal packet group 141 includes data generated by a comparison normal packet generator 155. The comparison normal packet generator 155 will be described later.
[0020] The control unit 15 controls the entire learning device 10. The control unit 15 is, for example, an electronic circuit such as a CPU (Central Processing Unit) or an MPU (Micro Processing Unit), or an integrated circuit such as an ASIC (Application Specific Integrated Circuit) or an FPGA (Field Programmable Gate Array). The control unit 15 has an internal memory for storing programs that define various processing procedures and control data, and executes each process using the internal memory. The control unit 15 also functions as various processing units by running various programs. The control unit 15 has a collection unit 151, a conversion unit 152, a detection unit 153, a learning control unit 154, and a comparison normal packet generation unit 155.
[0021] The collection unit 151 acquires a plurality of packets (variable length) in a normal state as learning packets. In the following description, learning packets collected by the collection unit 151 and in a normal state are referred to as "learning packets."
[0022] The learning control unit 154 uses the learning packets to make the conversion model 152m learn the packet conversion rules.
[0023] The conversion unit 152 converts the packet into a feature vector using the conversion model 152m. For example, the feature vector is data in which each byte of the packet is associated with a vector that represents the feature of the value of each byte.
[0024] The conversion model 152m is, for example, a natural language processing model using BERT (Bidirectional Encoder Representations from Transformers). The conversion model 152m converts the value of each byte of each communication packet, which is a training packet, into a 784-dimensional vector (feature vector). The conversion model 152m learns a conversion rule for a vector that associates the position of each byte of a communication packet with the 784-dimensional vector converted from that byte.
[0025] The learning control unit 154 performs unsupervised learning on the detection model 153m using the feature vectors of packets in a normal state obtained from the conversion unit 152. In this way, the detection model 153m learns patterns of feature vectors of packets in a normal state. For example, the detection model 153m is a variational autoencoder (VAE), an autoencoder (AE), a local outlier factor (LoF), or the like.
[0026] The detection unit 153 uses the detection model 153m to calculate the degree of anomaly (anomaly score) of a packet based on the feature vector converted by the conversion unit 152. For example, in the feature space, the further the position of the feature vector of the target packet is from the position of the feature vector of a learned packet in a normal state, the greater the degree of anomaly.
[0027] The comparison normal packet generator 155 generates the comparison normal packet group 141 using the plurality of learning packets collected by the collector 151 .
[0028] For example, the comparison normal packet generator 155 selects a group of packets sampled from a plurality of learning packets as the comparison normal packet group 141. The comparison normal packet generator 155 performs sampling using a sampling method such as kernel herding. In this case, random sampling is not performed to prevent performance degradation. Although sampling is desirable from the perspective of processing speed, all learning packets may be selected as the comparison normal packet group 141. Furthermore, the greater the number of comparison normal packets in the comparison normal packet group 141, the higher the detection accuracy of the detection device 20, which will be described later.
[0029] The comparison normal packet generator 155 stores the comparison normal packet group 141 in the storage unit 14 .
[0030] The learning device 10 notifies the detection device 20 of the parameters of the trained conversion model 152m, the parameters of the trained detection model 153m, and the data of the comparison normal packet group 141 via a network.
[0031] (Processing Procedure of Learning Device) A processing procedure of the learning process executed by the learning device 10 shown in Fig. 2 will be described below. Fig. 3 is a flowchart showing the processing procedure of the learning process executed by the learning device shown in Fig. 2.
[0032] 3, the collection unit 151 of the learning device 10 collects learning packets (step S51). The learning control unit 154 of the learning device 10 uses the learning packets to cause the conversion model 152m to learn packet conversion rules (step S52).
[0033] The learning control unit 154 performs unsupervised learning on the detection model 153m using the feature vectors of the normal packets output from the conversion model 152m (step S53). The comparison normal packet generation unit 155 of the learning device 10 generates the comparison normal packet group 141 based on the learning packets (step S54).
[0034] The learning device 10 notifies the detection device 20 of the parameters of the trained conversion model 152m, the parameters of the trained detection model 153m, and the data of the comparison normal packet group 141 (step S55).
[0035] (Detection Device) Next, a description will be given of the detection device 20 shown in Fig. 1. The detection device 20 narrows down the anomaly candidates by performing a primary evaluation on the evaluation target packets, and performs a secondary evaluation on the evaluation target packets that have been determined to be anomaly candidates as a result of the primary evaluation, to detect whether the evaluation target packets are anomalies.
[0036] (Primary Evaluation) The primary evaluation performed by the detection device 20 will now be described. Fig. 4 is a diagram for explaining the primary evaluation performed by the detection device. As shown in Fig. 4, when performing the primary evaluation, the detection device 20 uses a conversion model 252m and a detection model 253m.
[0037] The conversion model 252m is a natural language processing model using BERT. The conversion model 252m is applied with parameters of the conversion model 152m, which has learned vector conversion rules in the learning device 10.
[0038] The detection model 253m is VAE, AE, LoF, etc. The parameters of the detection model 153m that was the subject of unsupervised learning in the learning device 10 are applied to the detection model 253m.
[0039] The detection device 20 inputs the evaluation target packet 5 into the conversion model 252m to calculate a feature vector 5a of the evaluation target packet 5. The detection device 20 inputs the feature vector 5a into the detection model 253m to calculate an anomaly score (abnormal value) of the evaluation target packet 5.
[0040] If the anomaly score is greater than a preset threshold θ, the detection device 20 detects the evaluation target packet 5 as an abnormality candidate. The detection device 20 repeats the above process for other evaluation target packets. The detection device 20 performs a secondary evaluation, which will be described later, on the evaluation target packet that is an abnormality candidate.
[0041] For evaluation target packets that are not abnormal candidates, the detection device 20 sets the lower limit of the anomaly score to the evaluation target packets that are not abnormal candidates. For example, when the range of possible values for the anomaly score is "0 to 1," the detection device 20 sets the anomaly score of evaluation target packets that are not abnormal candidates to "0."
[0042] (Secondary Evaluation) The secondary evaluation performed by the detection device 20 will now be described. FIG. 5 is a diagram (1) for explaining the secondary evaluation performed by the detection device. As shown in FIG. 5, the detection device 20 uses a comparison normal packet group 141 acquired from the learning device 10. In addition, a packet to be evaluated that has been determined to be an abnormality candidate by the primary evaluation is designated as a packet to be evaluated 5.
[0043] The detection device 20 calculates the feature vector of each normal packet by inputting each packet included in the comparison normal packet group 141 (hereinafter, normal packets) into the conversion model 252m. In the description of FIG. 5, the feature vectors of each normal packet are collectively referred to as a feature vector group 6.
[0044] The detection device 20 calculates the cosine similarity between the feature vector 5a of the evaluation target packet 5 calculated in the description of Fig. 4 and each feature vector included in the feature vector group 6. The detection device 20 sorts the feature vectors included in the feature vector group 6 in descending order of cosine similarity, and extracts normal packets corresponding to the feature vectors included in the Mth (e.g., 100th) from the top as similar normal packets. Fig. 5 shows an example in which 100 similar normal packets are extracted.
[0045] 6 is a diagram (2) for explaining the secondary evaluation performed by the detection device. As shown in FIG. 6, the detection device 20 performs the following process based on the evaluation target packet 5 and the similar normal packets (100 packets).
[0046] The detection device 20 compares the length of M (e.g., 100) similar normal packets with the length of the evaluation target packet 5. The detection device 20 determines whether the number of similar normal packets having the same length as the evaluation target packet among the M similar normal packets is equal to or greater than a predetermined threshold (step S10).
[0047] If the number of similar normal packets having the same length as the packet to be evaluated among the M similar normal packets is equal to or greater than a predetermined threshold (Yes in step S10), the detection device 20 executes a first calculation process (step S11).
[0048] As a first calculation process, the detection device 20 performs one-dimensional detection based on a comparison of the numerical values for each byte position for the evaluation target packet and the same normal packet, and calculates the anomaly score for each byte position. The detection device 20 compares each byte (int) (0 to 255) one by one, detects any violations, and calculates the anomaly score. The detection device 20 outputs the total value of the anomaly scores for all byte positions as the first anomaly score. In this case, the detection device 20 sets the first anomaly score as the anomaly score for the evaluation target packet 5.
[0049] On the other hand, if the number of similar normal packets of the same length as the packet to be evaluated among the M similar normal packets is not greater than a predetermined threshold (step S10, No), the detection device 20 executes a second calculation process (step S12).
[0050] As a second calculation process, the detection device 20 calculates the edit distance between the evaluation target packet and each similar normal packet. The detection device 20 outputs the smallest edit distance among the calculated edit distances as the second anomaly score. In this case, the detection device 20 sets the second anomaly score as the anomaly score of the evaluation target packet 5. Furthermore, the detection device 20 estimates and detects inserted byte locations suspected of insertion or deleted byte locations suspected of deletion based on the calculated edit distances.
[0051] As described above, the detection device 20 performs primary and secondary evaluations on each evaluation target packet to set an anomaly score for each evaluation target packet. If the anomaly score set for any evaluation target packet exceeds a predetermined threshold, the detection device 20 detects that the corresponding evaluation target packet is abnormal and outputs an alert.
[0052] (Configuration of the detection device) An example configuration of the detection device 20 that executes the processes described in Figures 4 to 6 will be described. Figure 7 is a diagram showing an example configuration of the detection device 20 shown in Figure 1. As shown in Figure 7, the detection device 20 has an input unit 21, an output unit 22, a communication unit 23, a storage unit 24, and a control unit 25.
[0053] The input unit 21 is an input interface that accepts various operations from an operator of the detection device 20. For example, the input unit 21 is configured with input devices such as a touch panel, a voice input device, a keyboard, and a mouse.
[0054] The output unit 22 is realized by, for example, a display device such as a liquid crystal display, a printing device such as a printer, an information communication device, etc. The output unit 22 outputs an alert or the like when the control unit 25 detects an abnormality in the packet to be evaluated.
[0055] The communication unit 23 is a communication interface that transmits and receives various information to and from other devices connected via a network or the like. The communication unit 23 is realized by a NIC or the like, and communicates between other devices and a control unit 25 (described later) via telecommunications lines such as a LAN or the Internet. For example, the communication unit 23 receives evaluation target packets via the network and outputs them to the control unit 25. The communication unit 23 also acquires, from the learning device 10, parameters of a trained detection model, a conversion model (encoder) that has learned rules for converting packets into fixed-length feature vectors, and a comparison normal packet group 141.
[0056] The storage unit 24 is a storage device such as an HDD, SSD, or optical disk. The storage unit 24 may also be a data-rewritable semiconductor memory such as a RAM, flash memory, or NVSRAM. The storage unit 24 stores the OS and various programs executed by the detection device 20. The storage unit 24 also stores various information used in the execution of the programs. For example, the storage unit 24 has a comparison normal packet group 141.
[0057] The comparison normal packet group 141 is acquired from the learning device 10. The description of the comparison normal packet group 141 is the same as that given above.
[0058] The control unit 25 controls the entire detection device 20. The control unit 25 is, for example, an electronic circuit such as a CPU or MPU, or an integrated circuit such as an ASIC or FPGA. The control unit 25 has an internal memory for storing programs that define various processing procedures and control data, and executes each process using the internal memory. The control unit 25 functions as various processing units when various programs are run. The control unit 25 has an evaluation target packet collection unit 251, a conversion unit 252, a first detection unit 253, an extraction unit 254, and a second detection unit 255.
[0059] The evaluation target packet collection unit 251 acquires one evaluation target packet and outputs the evaluation target packet to the conversion unit 252.
[0060] The conversion unit 252 and the first detection unit 253 execute processing corresponding to the primary evaluation described in FIG. 4. The conversion unit 252 inputs the evaluation target packet to the conversion model 252m and calculates a feature vector. The conversion unit 252 outputs a pair of the calculated feature vector and the evaluation target packet to the first detection unit 253.
[0061] The first detection unit 253 calculates an anomaly score (abnormal value) of the evaluation target packet by inputting the feature vector into the detection model 235m. If the anomaly score is greater than the threshold θ, the first detection unit 253 detects the evaluation target packet for the feature vector as an evaluation target packet of an abnormal candidate. The first detection unit 253 outputs the evaluation target packet of the abnormal candidate and the feature vector of the evaluation target packet to the extraction unit 254.
[0062] On the other hand, if the anomaly score is less than the threshold value θ, the first detection unit 253 excludes the evaluation target packet from the abnormality candidates. Furthermore, the first detection unit 253 sets the anomaly score for the evaluation target packet excluded from the abnormality candidates to the lower limit value of the anomaly score.
[0063] The conversion unit 252 and the extraction unit 254 execute the processing for the secondary evaluation described in Fig. 5. The conversion unit 252 inputs each normal packet included in the comparison normal packet group 141 into a conversion model 252m, thereby calculating a feature vector of each normal packet (for example, feature vector group 6). The conversion unit 252 outputs to the extraction unit 254 a pair of the feature vector group and a normal packet corresponding to each feature vector included in the feature vector group.
[0064] The extraction unit 254 calculates the cosine similarity between the feature vector of the evaluation target packet that has been determined to be an abnormal candidate by the primary evaluation and each feature vector included in the feature vector group. The extraction unit 254 sorts the feature vectors included in the feature vector group in descending order of cosine similarity, and extracts normal packets corresponding to the feature vectors included in the Mth (e.g., 100th) top-ranked feature vectors as similar normal packets.
[0065] The extraction unit 254 outputs M (for example, 100) similar normal packets and the evaluation target packet that has been determined to be an abnormal candidate to the second detection unit 255.
[0066] The second detection unit 255 executes the process for the secondary evaluation described in Fig. 6 based on the evaluation target packet that has been determined to be an abnormal candidate and M (100) similar normal packets. The second detection unit 255 has a length comparison unit 255a, a first calculation unit 255b, and a second calculation unit 255c.
[0067] The length comparison unit 255a compares the lengths of the M similar normal packets with the evaluation target packet, and determines whether the number of similar normal packets having the same length as the evaluation target packet among the M similar normal packets is equal to or greater than a predetermined threshold.
[0068] When it is determined that the number of similar normal packets having the same length as the evaluation target packet among the M similar normal packets is equal to or greater than a predetermined threshold, the first calculation unit 255b executes a first calculation process to calculate a first anomaly score. The first calculation unit 255b sets the first anomaly score for the evaluation target packet that has been determined to be an abnormal candidate.
[0069] If it is determined that the number of similar normal packets having the same length as the evaluation target packet among the M similar normal packets is not equal to or greater than a predetermined threshold, the second calculation unit 255c executes a second calculation process to calculate a second anomaly score. The second calculation unit 255c sets the second anomaly score for the evaluation target packet that has been determined to be an abnormal candidate.
[0070] When the anomaly score set for any of the evaluation target packets exceeds a predetermined threshold, the control unit 25 detects that the corresponding evaluation target packet is abnormal and outputs an alert. Note that when the first anomaly score or the second anomaly score set for the evaluation target packet exceeds a predetermined threshold, the second detection unit 255 may detect that the corresponding evaluation target packet is abnormal and output an alert.
[0071] (First Calculation Process) The first calculation process executed by the first calculation unit 255b will be described below with reference to Fig. 8.
[0072] The first calculation unit 255b directly treats the value of the first byte of the similar normal packet extracted for comparison as a number between "0" and "255." The first calculation unit 255b calculates the anomaly score by applying one-dimensional anomaly detection using, for example, KDE (Kernel Density Estimator) (Reference 1). The first calculation unit 255b similarly calculates the anomaly scores for the second byte and subsequent bytes. The first calculation unit 255b performs one-dimensional detection based on a comparison of the numerical values for each byte position to calculate the anomaly score for each byte position. Reference 1: Duda, R. and Hart, P. (1973), "PATTERN CLASSIFICATION AND SCENE ANALYSIS," John Wiley & Sons. ISBN 0-471-22361-1.
[0073] The first calculation unit 255b outputs the anomaly score in the range of a minimum of "0" to a maximum of "1." For example, a case where the detection target is byte position L1 will be described ((1) in FIG. 8).
[0074] The first calculation unit 255b calculates, for each similar normal packet, a probability density distribution of the value of the byte appearing at byte position L1 of the similar normal packet. In other words, the first calculation unit 255b calculates a distribution centered on the number of bytes appearing in the similar normal packet and then calculates the sum of these distributions ((2) in FIG. 8). For example, among M similar normal packets, the first calculation unit 255b calculates a probability density distribution D1 for the first similar normal packet, a probability density distribution D2 for the second similar normal packet, and a probability density distribution D3 for the third similar normal packet as the probability density distribution of the value of the byte appearing at byte position L1.
[0075] The first calculation unit 255b calculates a probability density distribution Dt by summing the probability density distributions (e.g., probability density distributions D1 to D3) of the similar normal packets. The first calculation unit 255b observes the probability density of the byte values of the evaluation target packet on the probability density distribution Dt, and performs a normal / abnormal determination of byte position L1 of the evaluation target packet ((3) in FIG. 8).
[0076] Specifically, the first calculation unit 255b compares a predetermined threshold (for example, 1 / 1024) with the probability density of the byte value of the packet to be evaluated, and calculates the anomaly score.
[0077] If the probability density of the byte value of the packet to be evaluated is greater than or equal to a threshold value (for example, byte values "17" or "21"), the first calculation unit 255b determines that byte position L1 is normal and assigns an anomaly score of 0.
[0078] On the other hand, if the probability density of the byte value of the packet to be evaluated is less than the threshold, i.e., approximately "0" (for example, the byte value is "ff"), the first calculation unit 255b determines that this byte position L1 is abnormal and assigns an anomaly score of 1.
[0079] Furthermore, when the probability density of the evaluation target packet is equal to or greater than the threshold but is close to the threshold (for example, the byte value "30"), the first calculation unit 255b may interpolate the anomaly score between 0 and 1 according to a predetermined rule. For example, the first calculation unit 255b assigns an anomaly score of 0.3 to the byte value "30".
[0080] In this way, the first calculation unit 255b determines whether there is no violation (normal) or a violation (abnormal) for each byte position based on the probability density distribution Dt ((3) in FIG. 8). Therefore, even if the byte value changes randomly at a byte position, this does not lead to overdetection.
[0081] The first calculation unit 255b then calculates an anomaly score for each byte position and outputs the sum of the calculated anomaly scores as the final anomaly score (first anomaly score) for the packet to be evaluated.
[0082] The first calculation unit 255b may use any anomaly detection method that can handle one-dimensional data, such as a non-parametric method or a naive Bayes method.
[0083] (Second Calculation Process) Next, the second calculation process will be described. In the second calculation process, the second calculation unit 255c calculates the edit distance between the evaluation target packet and each similar normal packet using dynamic programming.
[0084] The second calculation unit 255c directly outputs the calculated edit distance as the anomaly score. Specifically, the second calculation unit 255c sets the smallest edit distance among the calculated edit distances as the second anomaly score.
[0085] Furthermore, by calculating the edit distance, the second calculation unit 255c can identify the location of any anomalous byte insertion and / or deletion, and directly identify how many packets have been inserted. Furthermore, if there are no similar normal packets, the edit distance is likely to be large.
[0086] The second calculation unit 255c acquires the evaluation target packet and M similar normal packets from the length comparison unit 255a, and compares the acquired similar normal packets with the evaluation target packet to identify inserted or deleted byte locations in the evaluation target packet.
[0087] Specifically, the second calculation unit 255c calculates the edit distance between the evaluation target packet and each similar normal packet using dynamic programming. By calculating the edit distance, the second calculation unit 255c can identify insertion byte locations where insertion is suspected or deletion byte locations where deletion is suspected.
[0088] If there are similar normal packets whose edit distance is less than the certain distance, the second calculation unit 255c selects the similar normal packet with the shortest edit distance from among the similar normal packets whose edit distance is less than the certain distance as the final similar normal packet.The second calculation unit 255c then estimates the insertion / deletion byte locations using the edit distance of the selected final similar normal packet.
[0089] On the other hand, if the edit distances of all the acquired similar normal packets are equal to or greater than the certain distance, the second calculation unit 255c processes the packet to be evaluated as not having a final similar normal packet corresponding to the packet. Here, the certain distance is a parameter that can specify, for example, approximately 1 / 3 to 1 / 2 of the packet length.
[0090] (Procedure of Detection Processing) A description will be given of an example of a procedure of detection processing executed by the detection device 20. Fig. 9 is a flowchart showing the procedure of detection processing according to the embodiment.
[0091] 9, the evaluation target packet collection unit 251 of the detection device 20 acquires one evaluation target packet (step S101), and then the detection device 20 executes a primary evaluation (step S102).
[0092] If the anomaly score does not exceed the threshold value θ (No at step S103), the detection device 20 sets a lower limit value for the anomaly score of the packet to be evaluated (step S104), and proceeds to step S106.
[0093] On the other hand, if the anomaly score exceeds the threshold value θ (Yes at step S103), the detection device 20 performs a secondary evaluation (step S105).
[0094] If the anomaly score set for the packet to be evaluated exceeds a threshold (a preset threshold) (Yes at step S106), the detection device 20 generates an alert (step S107).
[0095] On the other hand, if the anomaly score set for the packet to be evaluated does not exceed the threshold (No at step S106), the detection device 20 ends the process.
[0096] Next, a description will be given of the processing procedure of the primary evaluation shown in step S102 of Fig. 9. Fig. 10 is a flowchart showing the processing procedure of the primary evaluation.
[0097] As shown in FIG. 10, the conversion unit 152 of the detection device 20 inputs the packet to be evaluated into the conversion model 252m, thereby calculating a feature vector (step S201).
[0098] The first detection unit 253 of the detection device 20 inputs the feature vector into the detection model 253m and calculates an anomaly score (step S202). The detection device 20 sets an anomaly score for the packet to be evaluated (step S203).
[0099] Next, a description will be given of the procedure for the secondary evaluation shown in step S105 of Fig. 9. Fig. 11 is a flowchart showing the procedure for the secondary evaluation.
[0100] The conversion unit 252 of the detection device 20 reads the normal packet comparison group 141 (step S301), and calculates the feature vector of each normal packet using the conversion model 252m (step S302).
[0101] The extraction unit 254 of the detection device 20 identifies the feature vectors of M normal packets that have a relatively high similarity to the feature vector of the packet to be evaluated, and extracts the normal packets corresponding to the identified feature vectors as similar normal packets (step S303).
[0102] The length comparison unit 255a of the detection device 20 compares the lengths of the M similar normal packets with the evaluation target packet, and determines whether the number of similar normal packets having the same length as the evaluation target packet among the M similar normal packets is equal to or greater than a predetermined threshold (step S304).
[0103] If the number of similar normal packets of the same length as the packet to be evaluated is greater than or equal to a predetermined threshold (step S304, Yes), the first calculation unit 255b of the detection device 20 performs a first calculation process (step S305) and outputs a first anomaly score.
[0104] On the other hand, if the number of similar normal packets of the same length as the packet to be evaluated is less than a predetermined threshold (step S304, No), the second calculation unit 255c of the detection device 20 performs a second calculation process (step S306) and outputs a second anomaly score.
[0105] Next, the processing procedure of the first calculation process shown in step S305 of Fig. 11 will be described. Fig. 12 is a flowchart showing the processing procedure of the first calculation process. The first calculation unit 255b sets a parameter n, which indicates the position of the byte to be compared, to n = 1 (step S401).
[0106] The first calculation unit 255b calculates a probability density distribution Dt by summing the values at byte position n of the same normal packet (step S402). The first calculation unit 255b compares the probability density distribution Dt with the value at byte position n of the evaluation target packet (step S403). The first calculation unit 255b then compares the value at byte position n of the evaluation target packet with a predetermined threshold (e.g., 1 / 1024) to determine whether the value at byte position n of the evaluation target packet is equal to or greater than this threshold (step S404).
[0107] If the numerical value of byte position n of the evaluation target packet is equal to or greater than the threshold value (step S404, Yes), the first calculation unit 255b assigns 0 as the anomaly score of this byte position n (step S405). The first calculation unit 255b may interpolate the anomaly score between 0 and 1 according to a predetermined rule.
[0108] If the numerical value of the byte position n of the evaluation target packet is less than the threshold value (No at Step S404), the first calculation unit 255b assigns 1 as the anomaly score of this byte position n (Step S406).
[0109] If byte position n is not the last position (step S407, No), the first calculation unit 255b sets n = n + 1 (step S408), performs the processing from step S402 onwards for the numerical value of the next byte position, and calculates the anomaly score.
[0110] If byte position n is the last position (step S407, Yes), the first calculation unit 255b adds up the anomaly scores of all byte positions (step S409) and outputs the sum of the anomaly scores as the first anomaly score for the packet to be evaluated.
[0111] Next, a description will be given of the processing procedure of the second calculation process shown in step S306 in Fig. 11. Fig. 13 is a flowchart showing the processing procedure of the second calculation process.
[0112] The second calculation unit 255c selects one unselected similar normal packet from the predetermined number of similar normal packets (Step S501).
[0113] Next, the second calculation unit 255c calculates the edit distance between the evaluation target packet and the selected similar normal packet using dynamic programming (step S502).
[0114] Next, the second calculation unit 255c determines whether or not the selection of all of the predetermined number of similar normal packets has been completed (Step S503). If an unselected normal packet remains among the predetermined number of similar normal packets (No at Step S503), the second calculation unit 255c returns to Step S501.
[0115] On the other hand, if selection of all of the predetermined number of similar normal packets has been completed (Yes at step S503), the second calculation unit 255c determines whether or not there is a similar normal packet whose edit distance is less than a certain distance (step S504).
[0116] If there is a similar normal packet whose edit distance is less than the certain distance (Yes at step S504), the second calculation unit 255c executes the following process. In this case, the second calculation unit 255c selects the similar normal packet with the shortest edit distance from among the similar normal packets whose edit distance is less than the certain distance as the final similar normal packet. Then, the second calculation unit 255c estimates the insertion / deletion byte location 19 using the edit distance between the selected final similar normal packet and the packet to be evaluated (step S505).
[0117] On the other hand, if there is no similar normal packet whose edit distance is less than the certain distance (No at Step S504), the second calculation unit 255c determines that there is no final similar normal packet (Step S506).
[0118] The second calculation unit 255c then outputs the smallest edit distance among the calculated edit distances as the second anomaly score (step S507). The second calculation unit 255c also outputs the results of estimation and detection of insertion byte locations where insertion or deletion byte locations where deletion is suspected.
[0119] (Reference Technology) In the above embodiment, the detection device 20 detects anomalies in evaluation target packets by performing a secondary evaluation on evaluation target packets that have been identified as anomaly candidates in a primary evaluation. In contrast, a detection device may be considered that, when collecting evaluation target packets, skips the primary evaluation and performs only a secondary evaluation on all evaluation target packets. In the following description, a detection device that performs only a secondary evaluation without narrowing down the packets based on a primary evaluation will be referred to as a "reference device."
[0120] While the reference device is capable of detecting anomalies with high accuracy, performing a secondary evaluation on all evaluation target packets would require an extremely long calculation time for analysis. On the other hand, the detection device 20 of the embodiment performs a secondary evaluation on evaluation target packets that have been identified as anomaly candidates in the primary evaluation, thereby reducing calculation time while maintaining the accuracy of anomaly detection.
[0121] (Evaluation Experiment) Next, the results of an evaluation experiment on the detection device 20 and the reference device are shown. FIG. 14 is a diagram showing the results of the evaluation experiment. In table T1 of FIG. 14, the number of data indicates the number of processed evaluation target packets. For example, the number of data is set to "42,957." AUC (Area Under the) is one of the evaluation indices for machine learning. The normal data processing time is the time required to process normal evaluation target packets. For example, the number of normal evaluation target packets used in the evaluation is set to "42,957." The abnormal data processing time is the time required to process abnormal evaluation target packets. For example, the number of abnormal evaluation target packets used in the evaluation is set to "42,957."
[0122] Referring to Table T1 in Figure 14, we can see that the detection device 20 maintains detection accuracy while processing normal data 2.45 times faster. On the other hand, because the detection device 20 performs detection in two stages, it takes a long time to process abnormal data. However, since abnormalities are thought to be far less common than normal data in a typical network, if we assume that normal data accounts for 99% and abnormal data accounts for 1%, the overall processing time is approximately 2.41 times faster. This demonstrates that the detection device 20 can speed up processing while maintaining detection accuracy.
[0123] (Effects of the Embodiment) As shown in this evaluation experiment, according to the embodiment, it is possible to perform the process for detecting evaluation target packets at high speed while maintaining detection accuracy.
[0124] (Application Example) This embodiment can be applied to an anomaly detection system for IoT devices, for example.
[0125] Specifically, a network sensor is placed on the IoT network and packets are captured. The captured packets are then used to train the encoding unit (BERT). After training is complete, normal packets for comparison are set in the detection device 20. Packets acquired by the network sensor are then input into the detection device 20, and an anomaly score is calculated for each packet. This anomaly score is used to determine whether the packet to be evaluated is similar to the packet acquired during training (i.e., whether it is a normal packet).
[0126] (System Configuration of the Embodiment) The components of the learning device 10 and the detection device 20 are conceptual functional components and do not necessarily need to be physically configured as shown in the figure. In other words, the specific form of distribution and integration of the functions of the learning device 10 and the detection device 20 is not limited to that shown in the figure, and all or part of them can be functionally or physically distributed or integrated in any unit depending on various loads, usage conditions, etc.
[0127] Furthermore, all or any part of the processes performed by the learning device 10 and the detection device 20 may be realized by a CPU and a program analyzed and executed by the CPU. Furthermore, each process performed by the learning device 10 and the detection device 20 may be realized as hardware using wired logic.
[0128] Furthermore, among the processes described in the embodiments, all or part of the processes described as being performed automatically can be performed manually. Alternatively, all or part of the processes described as being performed manually can be performed automatically using a known method. In addition, the processing procedures, control procedures, specific names, and information including various data and parameters described above and illustrated can be changed as appropriate unless otherwise specified.
[0129] 15 is a diagram showing an example of a computer in which the learning device 10 and the detection device 20 are realized by executing a program. The computer 1000 has, for example, a memory 1010 and a CPU 1020. The computer 1000 also has a hard disk drive interface 1030, a disk drive interface 1040, a serial port interface 1050, a video adapter 1060, and a network interface 1070. These components are connected by a bus 1080.
[0130] The memory 1010 includes a ROM 1011 and a RAM 1012. The ROM 1011 stores a boot program such as a BIOS (Basic Input Output System). The hard disk drive interface 1030 is connected to a hard disk drive 1031. The disk drive interface 1040 is connected to a disk drive 1041. A removable storage medium such as a magnetic disk or optical disk is inserted into the disk drive 1041. The serial port interface 1050 is connected to a mouse 1051 and a keyboard 1052, for example. The video adapter 1060 is connected to a display 1061, for example.
[0131] The hard disk drive 1031 stores, for example, an OS 1091, an application program 1092, a program module 1093, and program data 1094. That is, the programs that define the processes of the learning device 10 and the detection device 20 are implemented as program modules 1093 in which code executable by the computer 1000 is written. The program modules 1093 are stored, for example, in the hard disk drive 1031. For example, the program modules 1093 for executing processes similar to the functional configurations of the learning device 10 and the detection device 20 are stored in the hard disk drive 1031. Note that the hard disk drive 1031 may be replaced by an SSD (Solid State Drive).
[0132] Furthermore, setting data used in the processing of the above-described embodiment is stored as program data 1094, for example, in memory 1010 or hard disk drive 1031. Then, CPU 1020 reads out program module 1093 or program data 1094 stored in memory 1010 or hard disk drive 1031 into RAM 1012 as necessary and executes them.
[0133] The program module 1093 and program data 1094 may not necessarily be stored in the hard disk drive 1031, but may also be stored in a removable storage medium, and read by the CPU 1020 via the disk drive 1041 or the like. Alternatively, the program module 1093 and program data 1094 may be stored in another computer connected via a network (such as a local area network (LAN) or a wide area network (WAN)). The program module 1093 and program data 1094 may then be read by the CPU 1020 from the other computer via the network interface 1070.
[0134] Although the present invention has been described above as an embodiment, the present invention is not limited to the description and drawings that form part of the disclosure of the present invention. In other words, other embodiments, examples, and operational techniques that can be made by those skilled in the art based on the present invention are all included in the scope of the present invention.
[0135] REFERENCE SIGNS LIST 10 Learning device 11, 21 Input unit 12, 22 Output unit 13, 23 Communication unit 14, 24 Storage unit 15, 25 Control unit 20 Detection device 141 Comparison normal packet group 151 Collection unit 152, 252 Conversion unit 152m, 252m Conversion model 153 Detection unit 153m, 253m Detection model 154 Learning control unit 155 Comparison normal packet generation unit 251 Evaluation target packet collection unit 253 First detection unit 254 Extraction unit 255 Second detection unit 255a Length comparison unit 255b First calculation unit 255c Second calculation unit
Claims
1. A detection device comprising: a conversion unit which converts a packet to be evaluated into a feature vector; a first detection unit which detects whether the packet to be evaluated is anomalous based on the result of inputting the feature vector into a detection model; an extraction unit which acquires anomaly candidate packets indicating the packet to be evaluated in which an abnormality has been detected by the first detection unit, and extracts a predetermined number of similar normal packets from a plurality of normal packets, the similar normal packets having a relatively high similarity to the abnormal candidate packets, based on a natural language processing model; and a second detection unit which detects whether the abnormal candidate packets are anomalous based on an anomaly score calculated by comparing the predetermined number of similar normal packets extracted by the extraction unit with the abnormal candidate packets.
2. The detection device described in claim 1, characterized in that the first detection unit obtains the degree of anomaly of the feature vector converted by the conversion unit, and detects an anomaly in the packet to be evaluated if the degree of anomaly is greater than or equal to a predetermined threshold.
3. The detection device according to claim 1 or 2, characterized in that the second detection unit has: a first calculation unit that extracts from the similar normal packets identical in packet length to the packet to be evaluated, and compares the packet to be evaluated with the identical-length packet on a byte-by-byte basis to calculate a first anomaly score; and a second calculation unit that, when the similar normal packet has a different packet length from the packet to be evaluated, calculates a second anomaly score based on an edit distance between the packet to be evaluated and the similar normal packet.
4. A detection method executed by a detection device, comprising: a step of converting a packet to be evaluated into a feature vector; a step of detecting whether or not the packet to be evaluated is anomalous based on the result of inputting the feature vector into a detection model; a step of acquiring an abnormal candidate packet indicating the packet to be evaluated in which an abnormality has been detected, and extracting a predetermined number of similar normal packets having a relatively high similarity to the abnormal candidate packet from among a plurality of normal packets based on a natural language processing model; and a step of detecting whether or not the abnormal candidate packet is anomalous based on an anomaly score calculated by comparing the predetermined number of extracted similar normal packets with the abnormal candidate packet.