Complexity calculation device, complexity calculation method, and complexity calculation program
The complexity calculation device addresses the challenge of evaluating multi-dimensional communication data continuity by calculating entropy across feature sets, enhancing analysis efficiency and device characterization.
Patent Information
- Application Number
- JP2022051846
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2022-03-28
- Publication Date
- 2025-05-21
- Estimated Expiration
- 2042-03-28
AI Technical Summary
Conventional complexity indices for time series data are designed for one-dimensional data and struggle to evaluate the complexity of communication data with multiple dimensions, failing to account for data continuity.
A complexity calculation device and method that includes a feature selection unit, feature set extraction unit, set creation unit, entropy calculation unit, and index output unit to evaluate the continuity of communication data by calculating entropy across multiple feature sets and normalizing the sum to a predetermined value range.
Enables the evaluation of communication data complexity considering continuity, improving analysis efficiency and enabling accurate anomaly detection and device characterization.
Smart Images

Figure 0007680980000005 
Figure 0007680980000006 
Figure 0007680980000007
Abstract
Description
[Technical field]
[0001] The present invention relates to an apparatus, a method, and a program for calculating the complexity of communication data. [Background technology]
[0002] Conventionally, the complexity of communication data has been evaluated and utilized to select a method for analyzing communication data, such as for anomaly detection, or to infer or group characteristics or models of communication devices.
[0003] For example, Patent Document 1 discloses that in device identification, a device detection process is performed based on the correlation between a plurality of packets, including a packet sequence, a packet size, and an interval between packets, or on entropy. The entropy calculated here is calculated for each record. [Prior art documents] [Patent documents]
[0004] [Patent Document 1] Special Publication No. 2020-518208 Summary of the Invention [Problem to be solved by the invention]
[0005] Conventional complexity indices for time series data are designed to measure one-dimensional data such as sensor data, making it difficult to directly apply one-dimensional methods to evaluate the complexity of communication data, which has multiple dimensions. Furthermore, although there are methods for calculating entropy as an index of the complexity of communication data, these calculations focus on specific fields within a single record. As a result, while it is possible to evaluate the complexity of individual records, it is not possible to evaluate the complexity that takes into account the continuity of communication.
[0006] An object of the present invention is to provide a complexity calculation device, a complexity calculation method, and a complexity calculation program capable of evaluating complexity taking into account the continuity of data. [Means for solving the problem]
[0007] The complexity calculation device of the present invention includes a feature selection unit that selects one or more fields as features for continuous data having multiple fields; a feature set extraction unit that sequentially extracts multiple feature sets each consisting of a predetermined number of consecutive data features; a set creation unit that creates a set of unique feature sets from the multiple extracted feature sets; an entropy calculation unit that calculates the entropy of each feature set in the set of unique feature sets and calculates the sum of the entropies; and an index output unit that normalizes the sum of the entropies to a predetermined value range and outputs it as a complexity index.
[0008] The continuous data may be time-series data of communication information.
[0009] The time series data may be flow data.
[0010] The feature amount selection section may perform a specific process among a plurality of fields and select a newly created field as a part of the feature amount.
[0011] The feature selection unit may select features in a plurality of patterns, and the index output unit may obtain statistics from a plurality of the complexity indices calculated from each of the plurality of patterns, and output the statistics as an output value.
[0012] The feature set extraction unit may set a plurality of the numbers of data, and the index output unit may obtain statistics from a plurality of the complexity indices calculated from each of the plurality of the numbers of data, and set the statistics as output values.
[0013] The complexity calculation method of the present invention includes a feature selection step of selecting one or more fields as features for continuous data having multiple fields; a feature set extraction step of sequentially extracting multiple feature sets each consisting of a predetermined number of consecutive data features; a set creation step of creating a set of unique feature sets from the multiple extracted feature sets; an entropy calculation step of calculating the entropy of each feature set in the set of unique feature sets and calculating the sum of the entropies; and an index output step of normalizing the entropy sum to a predetermined value range and outputting it as a complexity index.
[0014] A complexity calculation program according to the present invention is for causing a computer to function as the complexity calculation device. Effect of the Invention
[0015] According to the present invention, it is possible to evaluate the complexity taking into account the continuity of data. [Brief description of the drawings]
[0016] [Figure 1] FIG. 2 is a diagram illustrating a functional configuration of a complexity calculation device according to an embodiment. [Diagram 2] 1 is a flowchart showing the flow of a complexity calculation method according to an embodiment. [Diagram 3] 10 is a diagram illustrating an example of an IPFIX record acquired as continuous data in an embodiment. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0017] An example of an embodiment of the present invention will now be described. In the complexity calculation method of this embodiment, an index indicating the complexity taking into account the continuity of communication in communication sent and received by a device is output using aggregate information such as IPFIX or header information of communication data.
[0018] FIG. 1 is a diagram showing the functional configuration of a complexity calculation device 1 in this embodiment. The complexity calculation device 1 is an information processing device including a control unit 10, a storage unit 20, as well as input / output devices and communication devices for various data.
[0019] The control unit 10 is a part that controls the entire complexity calculation device 1, and realizes each function in this embodiment by appropriately reading and executing various programs stored in the storage unit 20. The control unit 10 may be a CPU. Specifically, the control unit 10 includes a feature amount selection unit 11 , a feature amount set extraction unit 12 , a set creation unit 13 , an entropy calculation unit 14 , and an index output unit 15 .
[0020] The storage unit 20 is a storage area for various programs and various data for making the hardware group function as the complexity calculation device 1, and may be a ROM, a RAM, a flash memory, a hard disk drive (HDD), etc. Specifically, the storage unit 20 stores a program (complexity calculation program) for making the control unit 10 execute each function of this embodiment, and further stores continuous data received as a processing target, calculation results, etc.
[0021] The feature selection unit 11 selects one or more fields as feature quantities for continuous data having a plurality of fields, in accordance with a selection input from a user or the like. Furthermore, the feature quantity selection unit 11 may perform specific processing, such as arithmetic operations, between a plurality of fields and select a newly created field as a part of the feature quantity. Furthermore, the feature quantity selecting unit 11 may select feature quantities in a plurality of patterns.
[0022] Here, the continuous data may be, for example, output data from an IoT device, etc., or may be time-series data of bidirectional communication information. The communication information may be, for example, packet header information, or a sequence of flow data in which statistical information is stored. Furthermore, the continuous data may be sampled at a fixed period, for example.
[0023] The feature set extraction unit 12 sequentially extracts a plurality of feature sets each including a predetermined number of consecutive data features. At this time, the feature set extraction unit 12 may set a plurality of data numbers. In this case, the feature set extraction unit 12 extracts a feature set for each data number.
[0024] The set creation unit 13 creates a set of unique feature sets from among the multiple extracted feature sets.
[0025] The entropy calculation unit 14 calculates the entropy of each feature set in the collection of unique feature sets, and calculates the sum of these entropies.
[0026] The index output unit 15 normalizes the calculated sum of entropy to a predetermined range and outputs it as a complexity index. At this time, the index output unit 15 may calculate a statistical quantity from multiple complexity indices calculated from the features selected in multiple patterns and the multiple set data numbers, and output the calculated statistical quantity as a final complexity index. The statistical amount may be, for example, an average value, a median value, a maximum value, a minimum value, etc. Furthermore, the index output unit 15 may calculate a standard deviation from a plurality of complexity indices and output the complexity index with a range.
[0027] FIG. 2 is a flowchart showing the flow of the complexity calculation method in this embodiment. Here, a specific procedure for calculating the complexity index when IPFIX is used as the communication data will be illustrated.
[0028] IPFIX has the following fields (IPFIX elements): flowEndMilliseconds: The time the flow ended. flowStartMilliseconds: The time the flow started. flowDurationMilliseconds: The duration of the flow. reverseFlowDeltaMilliseconds: The time difference between the first outgoing packet and the returning packet of the flow. protocolIdentifier: The protocol of the flow (ICMP, TCP, UDP, etc.). sourceIPv4Address: Source IPv4 address. sourceTransportPort: The source port number. packetTotalCount: The number of packets in the flow going in that direction. octetTotalCount: The number of octets contained in the flow in the outbound direction. sourceMacAddress: Source MAC address. destinationIPv4Address: The IPv4 address of the destination. destinationTransportPort: The destination port number. reversePacketTotalCount: The number of return packets contained in the flow. reverseOctetTotalCount: The number of return octets contained in the flow.
[0029] In step S1, the feature selection unit 11 selects c fields from these fields, and sets a combination of the c fields as a feature F. For example, when three fields are selected, the feature of the i-th record is F i =(packetTotalCount, octetTotalCount, reversePacketTotalCount) It will look like this. As described above, the feature quantity selecting unit 11 may use any of these fields to include, for example, a calculation result such as a transmission / reception ratio (number of transmitted packets / number of received packets) in the feature quantity.
[0030] FIG. 3 is a diagram illustrating an example of an IPFIX record acquired as continuous data in this embodiment. Here, only the three selected fields are shown, and the total number of records obtained for the period in which complexity was measured was eight (record numbers 1 to 8).
[0031] In step S2, the feature set extraction unit 12 treats the d consecutive data as one set, and extracts a feature set from each set. For example, in the IPFIX record shown in Figure 3, if d=3, the combinations of record numbers that can become feature sets are as follows: {(1,2,3), (2,3,4), (3,4,5) (4,5,6), (5,6,7), (6,7,8)} If the total number of records used to calculate the complexity index is z (8 in the example of FIG. 3), the number of feature sets is z-d+1 (=8-3+1=6).
[0032] Here, the feature set to be extracted is, for example, when the record number is a combination of (1, 2, 3), F (1,2,3) ={F 1 ,F 2 ,F 3}={(8, 400, 300), (8, 400, 300), (4, 160, 80)} It is.
[0033] In step S3, the set creation unit 13 creates a set of unique feature sets from the generated feature sets. In the example of Figure 3, F (1,2,3) =F (5,6,7) Therefore, the set of unique feature sets is {(1,2,3), (2,3,4), (3,4,5), (4,5,6), (6,7,8)} It becomes.
[0034] In step S4, the entropy calculation unit 14 calculates the frequency of occurrence (number of occurrences) of each feature set in the communication data included in the period for which complexity is to be measured. For example, in the example of FIG. 3, the occurrence frequency of each feature set is calculated as follows: (1,2,3)=2, (2,3,4)=1, (3,4,5)=1, (4,5,6)=1, (6,7,8)=1
[0035] In step S5, the entropy calculation unit 14 calculates the entropy (average information amount) for each element in the collection of unique feature sets, and further calculates a total value by adding up the entropies of each feature set.
[0036] Here, the number of elements in the collection of unique feature sets is N, and the occurrence frequency of the i-th element is V i Then, the sum of the entropy of each feature set is
number
number
[0037] In step S6, the index output unit 15 divides the total entropy value E by the maximum entropy according to the number of elements N, and normalizes the value to be in the range of 0 to 1 to obtain a complexity index E p Let us assume that.
number
number
[0038] As described above, when multiple patterns of fields are selected, or when multiple numbers of consecutive data d to be included in the feature set are set, steps S1 to S6 are executed for each pattern and number of data, and multiple complexity indices are calculated.
[0039] In step S7, the index output unit 15 outputs the calculated complexity index E p Outputs multiple complexity indices E p is calculated, the index output unit 15 obtains these statistics and outputs them as the final complexity index.
[0040] According to this embodiment, the complexity calculation device 1 sequentially extracts multiple feature sets, each consisting of a predetermined number of consecutive data features, from consecutive data having multiple fields, and calculates a complexity index by summing up the entropy of each unique feature set. Therefore, the complexity calculation device 1 can evaluate the complexity of continuous data acquired from sensor output, communication time-series data, etc., taking into account the continuity of the data, by performing simple calculations with a small processing load.
[0041] As a result, for example, when performing anomaly detection on communication data, by understanding the complexity of the communication data in advance, it is possible to narrow the range of analysis methods to be selected, such as using whitelists when the complexity is low, and decision trees or deep learning when the complexity is high. In addition, the characteristics or model of the device that is the source of communication can be inferred and classified based on the complexity of the communication data.
[0042] With regard to communication data, the complexity calculation device 1 targets flow data, and can obtain features from records in which multiple packets are aggregated, thereby making it possible to reduce the number of continuous data d and improve calculation efficiency.
[0043] The complexity calculation device 1 may perform specific processing between multiple fields and select the newly created field as part of the feature, thereby making it possible to appropriately evaluate the complexity according to the characteristics of the data using a variety of feature values.
[0044] The complexity calculation device 1 may select feature amounts in a plurality of patterns, or may set a plurality of consecutive data numbers d, thereby calculating a plurality of complexity indices and outputting statistics. As a result, even if the optimal feature amount and the number of data are unknown, the complexity calculation device 1 can improve the reliability of the evaluation result and can appropriately output a requested index such as a maximum value or a minimum value. The number of consecutive data d is preferably set to about 2 or more and 7 or less, which makes it possible to appropriately extract the characteristics of the consecutive data and evaluate the complexity while suppressing the amount of calculation.
[0045] The above-described embodiment, for example, can improve the efficiency of analysis of communication data, which can contribute to Goal 9 of the United Nations-led Sustainable Development Goals (SDGs), which is to "build resilient infrastructure, promote sustainable industrialization and foster innovation."
[0046] Although the embodiments of the present invention have been described above, the present invention is not limited to the above-described embodiments. Furthermore, the effects described in the above-described embodiments are merely a list of the most preferable effects resulting from the present invention, and the effects of the present invention are not limited to those described in the embodiments.
[0047] The complexity calculation method by the complexity calculation device 1 is realized by software. When realized by software, a program constituting this software is installed in an information processing device (computer). These programs may be recorded on a removable medium such as a CD-ROM and distributed to users, or may be distributed by being downloaded to the user's computer via a network. Furthermore, these programs may be provided to the user's computer as a Web service via a network without being downloaded. [Explanation of symbols]
[0048] 1 Complexity calculator 10 Control section 11 Feature selection section 12 Feature Set Extraction Unit 13 Set Creation Department 14 Entropy calculation section 15 Index output section 20 Memory section
Claims
1. a feature selection unit that selects one or more fields as features for continuous data having a plurality of fields; a feature set extraction unit that sequentially extracts a plurality of feature sets each including a predetermined number of consecutive data features; a set creation unit that creates a set of unique feature sets from among the plurality of extracted feature sets; an entropy calculation unit that calculates the entropy of each feature set in the collection of unique feature sets and calculates a sum of the entropies; an index output unit that normalizes the sum of the entropies to a predetermined range and outputs the normalized sum as a complexity index; The feature selection unit selects feature amounts in a plurality of patterns, The index output unit calculates statistics from the plurality of complexity indices calculated from the plurality of patterns, respectively, and sets the statistics as an output value.
2. a feature selection unit that selects one or more fields as features for continuous data having a plurality of fields; a feature set extraction unit that sequentially extracts a plurality of feature sets each including a predetermined number of consecutive data features; a set creation unit that creates a set of unique feature sets from among the plurality of extracted feature sets; an entropy calculation unit that calculates the entropy of each feature set in the collection of unique feature sets and calculates a sum of the entropies; an index output unit that normalizes the sum of the entropies to a predetermined range and outputs the normalized sum as a complexity index; The feature set extraction unit sets a plurality of the data numbers, The index output unit calculates statistics from a plurality of complexity indices calculated from the respective plurality of data counts, and sets the calculated statistics as an output value.
3. The complexity calculation device according to claim 1 or 2, wherein the continuous data is time-series data of communication information.
4. The complexity calculation device according to claim 3 , wherein the time series data is flow data.
5. 5. The complexity calculation device according to claim 1, wherein the feature quantity selection unit performs a specific process among a plurality of fields and selects a newly created field as part of the feature quantities.
6. A feature selection step of selecting one or more fields as features for continuous data having a plurality of fields; a feature set extraction step of sequentially extracting a plurality of feature sets each including a predetermined number of consecutive data features; a set creation step of creating a set of unique feature sets from the plurality of extracted feature sets; an entropy calculation step of calculating the entropy of each feature set in the collection of unique feature sets and calculating a sum of the entropies; an index output step of normalizing the sum of the entropies to a predetermined range and outputting the normalized sum as a complexity index; In the feature selection step, feature values are selected in a plurality of patterns; A complexity calculation method, wherein in the index output step, a statistic is calculated from the plurality of complexity indices calculated from each of the plurality of patterns, and the calculated statistic is used as an output value.
7. A feature selection step of selecting one or more fields as features for continuous data having a plurality of fields; a feature set extraction step of sequentially extracting a plurality of feature sets each including a predetermined number of consecutive data features; a set creation step of creating a set of unique feature sets from the plurality of extracted feature sets; an entropy calculation step of calculating the entropy of each feature set in the collection of unique feature sets and calculating a sum of the entropies; an index output step of normalizing the sum of the entropies to a predetermined range and outputting the normalized sum as a complexity index; In the feature set extraction step, the number of pieces of data is set to a plurality of pieces; A complexity calculation method in which, in the index output step, a statistic is calculated from a plurality of complexity indices calculated from each of a plurality of the data counts, and the calculated statistic is used as an output value.
8. A complexity calculation program for causing a computer to function as the complexity calculation device according to any one of claims 1 to 5.
Citation Information
Patent Citations
Traffic analysis system and traffic analysis method
JP2011244098A
Method and system for detecting event of vehicle cyber-attack
JP2019145081A
Device Identification
JP2020518208A
Wavelet decomposition of software entropy to identify malware
US20160292418A1
Traffic analysis device, method, and program
WO2019176997A1