Communication identification device, method, and program

The communication identification device addresses the challenge of dynamic source connection information by classifying and integrating flow data to identify steady-state communication, ensuring accurate identification in networks with address translation functions.

JP2026100483APending Publication Date: 2026-06-19KDDI CORP +2
View PDF 4 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
KDDI CORP
Filing Date
2024-12-09
Publication Date
2026-06-19

AI Technical Summary

Technical Problem

Existing methods for identifying steady-state communication fail in network environments where source connection information changes dynamically due to address translation functions like NAPT and NAT, as they rely on static IP addresses.

Method used

A communication identification device that collects flow data, classifies it based on source and destination connection information, calculates communication fluctuations, identifies sets of steady communication, determines address translation candidates, and integrates sets based on probability calculations to accurately identify steady-state communication despite dynamic address changes.

Benefits of technology

Enables accurate identification of steady-state communication even in networks with dynamic source connection information, enhancing communication security and stability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026100483000001_ABST
    Figure 2026100483000001_ABST
Patent Text Reader

Abstract

Identify steady-state communications in a network environment where the address of the source connection information changes dynamically. [Solution] The data collection unit 10 collects flow data from communications on the network. The data classification unit 20 classifies the collected flow data into a first set unique to the combination of source connection information and destination connection information. The communication variation calculation unit 30 calculates communication variation for each set based on its flow data. The steady-state communication identification unit 40 identifies sets of steady-state communications based on the communication variation. The address translation candidate pair discrimination unit 50 discriminates sets of address translation candidates from sets of steady-state communications. The address translation probability calculation unit 60 calculates the probability that a set of address translation candidates is a set of address translations that have been classified into different sets as a result of address translation. The integration unit 70 integrates set of address translations whose probabilities satisfy predetermined conditions.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to an apparatus, method, and program for analyzing communications on a network and identifying steady communications, and more particularly to a communication identification apparatus, method, and program for identifying steady communications with a destination in a network environment where the connection information of the source of the communication changes dynamically. [Background technology]

[0002] Patent documents 1-4 disclose a method for distinguishing between foreground and background communications in a communication terminal, as well as an improved version thereof.

[0003] Patent Document 1 discloses a technology for distinguishing between foreground communication, which is triggered by user operations, and background communication, which is conducted by an application independently of user operations, for mobile terminals and the like. Here, the steady-state communication targeted for identification in the present invention corresponds to a part of the background communication.

[0004] Patent Document 1 uniquely identifies communications by a combination of source IP address, destination IP address, and destination port number (SD group), and classifies each communication as a foreground communication or background communication based on autocorrelation and crosscorrelation regarding the timing of communication occurrence.

[0005] Patent Document 1 determines that communications with a high autocorrelation coefficient are likely to be background communications automatically and mechanically executed by the OS or application, regardless of user operation. Similarly, communications with a high cross-correlation coefficient with communications already classified as background communications are also considered likely to be background communications. On the other hand, even if cross-correlation is observed, if the timing of the communication is below a predetermined threshold, it may be a user-initiated communication and is therefore classified as a foreground communication.

[0006] Patent Document 2 discloses a method for reducing computational complexity based on the method described in Patent Document 1. The time axis is discretized into bins of width Δt, and the timing of communication occurrences is vectorized by setting a 1 in the nth digit of the bit sequence T when communication occurs in the nth bin. By representing each digit with a 1 as the distance from the starting bin, the original sparse vector can be compressed. By calculating the confidence level from the compressed vector, foreground and background communication can be distinguished with less computation.

[0007] Patent Document 3 discloses an invention that distinguishes between foreground and background communication even when the periodicity of communication occurrence times is disrupted due to network conditions or other reasons, in addition to the method described in Patent Document 2. Improvements have been made to the invention by using statistical quantities such as the amount of communication per Δt, the number of packets, and throughput characteristics as communication characteristics, in addition to the time at which communication occurs, so that even when the periodicity is considered to be disrupted, if the time-series fluctuation of the communication characteristics is small, it can be determined to be background communication.

[0008] Patent Document 4 discloses an invention for identifying background communication in cases where, in addition to the method described in Patent Document 3, content servers are distributed in cloud services, etc., and even if the source is the same and the session is related to the same service or application, the destination may differ. Specifically, by aggregating SD groups with the same source information but different destination information based on the correlation of the communication characteristics of each session, communications with different destinations are identified as background communication. [Prior art documents] [Patent Documents]

[0009] [Patent Document 1] Japanese Patent Publication No. 2015-162810 [Patent Document 2] Japanese Patent Publication No. 2016-082518 [Patent Document 3] Japanese Patent Publication No. 2016-116026 [Patent Document 4] Japanese Patent Publication No. 2016-127361 [Overview of the project] [Problems that the invention aims to solve]

[0010] In the field of communication security, the increasing speed and volume of communications have made packet-level processing difficult, leading to the use of lighter flow data such as IPFIX.

[0011] Communication data can be broadly classified into steady-state communication and transient communication. Steady-state communication, exemplified by Keep Alive, is communication where the unit amount and frequency of communication are constant, and it often appears in various types of communication data.

[0012] On the other hand, in tasks such as abnormal communication detection and application identification, communication characteristics often appear in non-stationary communications, and it is expected that classifying and handling stationary and non-stationary communications will improve the accuracy of the task.

[0013] Patent documents 1-4 disclose methods for distinguishing between foreground and background communication based on information about the source and destination of communication, as well as communication characteristics such as communication volume and packet size. These methods are based on autocorrelation regarding the timing of communication occurrence; in other words, if communication occurs steadily from the start of the session, it is determined to be background communication. Furthermore, patent document 3 takes into account fluctuations in communication characteristics such as communication volume and packet size in addition to the timing of communication occurrence.

[0014] However, large-scale networks that span the internet, such as vehicle networks and IoT platforms, require a large number of global IP addresses, so address translation functions such as NAPT and NAT are sometimes used.

[0015] While NAPT and NAT (sometimes referred to simply as NAPT) allow for the efficient use of a small number of global IP addresses, NAPT translates connection information such as the source IP address and port number. As a result, from the application's perspective, the source connection information changes dynamically. In cases where the source connection information changes dynamically, the methods described in Patent Documents 1-4 cannot be used to identify steady-state communication.

[0016] The objective of the present invention is to solve the above technical problems and to enable the identification of steady communication by quantifying the steadyness of unit communication volume and communication frequency as index values ​​in a network environment in which the address of the source connection information is translated, and by combining both index values. [Means for solving the problem]

[0017] To achieve the above objective, the present invention is characterized in that, in a communication identification device that identifies communication on a network where the address of the source connection information x is translated, the following configuration is provided.

[0018] (1) The system comprises means for collecting flow data from communication data on a network, means for classifying the collected flow data into sets specific to the combination of source connection information x and destination connection information y, means for calculating communication fluctuations based on the flow data for each set, means for identifying sets of steady communication based on the communication fluctuations, means for determining from the sets of steady communication sets sets of address translation candidates that may have been classified into different sets as a result of address translation, means for calculating the probability that a set of address translation candidates is a set of address translations classified into different sets as a result of address translation, and means for integrating set of address translations whose probabilities satisfy predetermined conditions.

[0019] (2) The integrated set is added to the set of steady-state communications, and the identification of the set pairs of address translation candidates, the calculation of probabilities, and integration are repeated until there are no more sets that can be integrated. [Effects of the Invention]

[0020] According to the present invention, even in a network environment where the connection information of the communication source changes dynamically due to address translation functions such as NAT and NAPT, it becomes possible to accurately identify steady-state communication between the communication source and the communication destination. [Brief explanation of the drawing]

[0021] [Figure 1] This figure shows the configuration of a network to which the communication identification device of the present invention is applied. [Figure 2] This is a functional block diagram showing the configuration of a communication identification device to which the present invention is applied. [Figure 3] This diagram schematically illustrates a method for classifying flow data. [Figure 4] This flowchart illustrates the method for identifying regular communications using a communication identification device. [Figure 5] This diagram illustrates the method for calculating communication frequency. [Figure 6] This diagram schematically illustrates the method (part 1) for identifying candidate pairs for address translation. [Figure 7] This diagram schematically illustrates the method (part 2) for identifying candidate pairs for address translation. [Figure 8] This diagram schematically illustrates an example of how to calculate the address translation probability. [Modes for carrying out the invention]

[0022] Embodiments of the present invention will be described in detail below with reference to the drawings. Figure 1 is a schematic diagram showing the configuration of a network to which the communication identification device of the present invention is applied, in which a client as a communication source terminal and an application (app) server as a communication destination terminal are interconnected via a router equipped with an address translation function (NAPT, NAT). By accessing the router with an address translation function, the communication source terminal can dynamically translate its own IP address and port number and continuously receive desired services from the application server.

[0023] Figure 2 is a functional block diagram showing the configuration of the main parts of the communication identification device 1 to which the present invention is applied. It is implemented in or connected to the router with address translation function and identifies steady communication based on flow data obtained by analyzing communication data. The communication identification device 1 mainly consists of a data collection unit 10, a data classification unit 20, a communication variation calculation unit 30, a steady communication identification unit 40, an address translation candidate pair determination unit 50, an address translation probability calculation unit 60, and an integration unit 70.

[0024] The data collection unit 10 captures all communication data relayed by the Dynamic DNS reverse proxy and collects connection information such as the IP addresses and port numbers of the source and destination terminals, as well as statistics such as communication volume, number of packets, and communication start time, as flow data using network traffic monitoring / analysis methods such as IPFIX.

[0025] As schematically shown in Figure 3, the data classification unit 20 takes the entire set of collected flow data as D, and records each flow data in set D that has the same combination of connection information unique to the source terminal (source connection information) x and connection information unique to the destination terminal (destination connection information) y, and assigns the corresponding first set D x,y They are classified into these categories.

[0026] The communication variation calculation unit 30 is the first set D x,y For each, the communication fluctuation is calculated based on the flow data. In this embodiment, the first set D x,y For each communication fluctuation, the unit communication volume calculation unit 301 calculates the coefficient of variation of the unit communication volume, and the communication frequency calculation unit 302 calculates the coefficient of variation of the communication frequency. The steady-state communication identification unit 40 is the first set D x,y Based on the coefficient of variation of the unit communication volume and the coefficient of variation of the communication frequency, a set of flow data that is a steady-state communication is identified for each.

[0027] The address translation candidate pair discrimination unit 50 determines that the flow data is a set S of steady communication. x,y From two sets S x1,y1 ,Sx2,y2 Select this set pair S x1,y1 , S x2,y2 Based on the similarity of the coefficient of variation of the unit traffic volume and the coefficient of variation of the communication frequency for this set pair, determine whether the set pair is a set pair of address conversion candidates classified into different sets as a result of the source information being rewritten by the address conversion function.

[0028] In this embodiment, as will be described in detail later, a set pair in which the difference between the coefficients of variation of the unit traffic volume and the communication frequency is smaller than a predetermined threshold (Condition 1) and the continuity as steady communication is recognized in the time series of the communication source information (Condition 2) is determined to be a set pair of address conversion candidates.

[0029] The address conversion probability calculation unit 60 calculates the address conversion probability P by an appropriate calculation method such that the probability increases as the continuous time (elapsed time), that is, the time for which the flow data classified into one of the sets in the time series continues to communicate with the same communication source address, in other words, the elapsed time from the last NAT table rewrite time, is longer, for the set pair S x1,y1 , S x2,y2 of the address conversion candidates, as will be described in detail later.

[0030] The integration unit 70 integrates the flow records of each set by aggregating the flow data of the set pair in which the address conversion probability P exceeds a predetermined threshold. The procedure for integrating the set pair S x1,y1 , S x2,y2 is repeated until all sets cannot be integrated.

[0031] Such a communication identification device 1 can be configured by implementing an application (program) that realizes each function described in detail below on a general-purpose computer or server equipped with a CPU, ROM, RAM, bus, interface, etc., or a portable smartphone or tablet terminal. Alternatively, it can also be configured as a dedicated machine or single-function machine in which part of the application is hardwareized or softwareized.

[0032] Figure 4 is a flowchart showing the procedure for communication identification by the communication identification device 1. In step S1, the flow data collection unit 10 captures all communication data to be identified and collects connection information such as the IP addresses and port numbers of the source and destination terminals, as well as statistics such as communication volume, number of packets, and communication start time, as flow data.

[0033] In step S2, the data classification unit 20 combines the corresponding source connection information x and destination connection information y in the collected flow data set D to create a record, and assigns each record with the same combination to a first set D unique to that combination. x,y They are classified into these categories.

[0034] Each combination of connection information x and y can use the IP addresses of the communication source and destination individually, or it may use a combination of an IP address and a port number. Alternatively, destination information can be categorized by function based on the functional classification of the communication destination using methods such as clustering. In this embodiment, a combination of the IP address of the communication source terminal and the IP address and port number of the communication destination terminal will be used.

[0035] In this embodiment, since the source connection information x changes dynamically, even if the source terminal and destination terminal are the same communication at this point, the flow data (records) will be different for multiple first sets D. x,y It is sometimes classified as such.

[0036] In step S3, the unit communication amount calculation unit 301 calculates the first set D x,y Coefficient of variation CV for each unit of data transfer x,,y volume This is calculated. In this embodiment, the amount of communication is calculated in packets, and the flow record k∈D x,y Regarding the number of packets, k , the amount of data transmitted is q k The amount of data transfer per packet in this case is v k This can be found using the following equation (1).

[0037]

number

[0038] Next, the amount of data transmitted per packet v k (k∈D x,y ) coefficient of variation CV x,y volume This is calculated using the following equation (2), and is used as one of the indicators for evaluating stability.

[0039]

number

[0040] In step S4, the communication frequency calculation unit 302 calculates the first set D x,y Coefficient of variation CV for each communication frequency x,y frequency This is calculated. In this embodiment, the flow record k∈D x,y The start time of each communication is t k The start time of the entire period for which you want to determine stationarity is s t , end time e t Let i ∈ {1, 2, ..., |D} be the time point. x,,y A set of timestamps T that combine |} x,y We can find this using the following equation (3).

[0041]

number

[0042] In this embodiment, by adding the start time st and end time et for the entire period over which we want to determine stationarity, we can exclude sets of flow records that communicate periodically only during a portion of the period, as shown in Figure 5. This makes it possible to identify only true stationary communications that exhibit stationarity over the entire period.

[0043] In this embodiment, T x,y Communication frequency f i The coefficient of variation CV can be expressed as shown in equation (4) below and obtained by equation (5) below. x,yfrequency This is used as another indicator for evaluating stability.

[0044]

number

[0045]

number

[0046] In step S5, the steady-state communication identification unit 40 is the first set D x,y For each unit, the coefficient of variation CV of the unit communication amount and communication frequency. x,y volume CV x,y frequency Based on this, regular communications are identified.

[0047] For the first set, CV x,y volume The smaller the value, the more stable the communication is considered to be, so the threshold ε volume When set, the first set D satisfies the following condition (6). x,y This is judged to be constant with respect to the unit amount of data transmitted.

[0048]

number

[0049] When a flow record is bidirectional, the number of packets and the amount of data transmitted have statistics for inbound (destination → source) and outbound (source → destination), respectively. In this case, the sum of the inbound and outbound values ​​can be used. At this time, the number of packets p and the amount of data transmitted q are calculated as shown in equations (7) and (8) below.

[0050]

number

[0051]

number

[0052] For the first set, CV x,y frequency The smaller the value, the more we can determine that the communication frequency is stationary, so the threshold ε frequency When this is set, the first set satisfying equation (9) is judged to be stationary with respect to communication frequency.

[0053]

number

[0054] And the CV, a stationary index related to unit communication volume. x,y volume and the CV (Continuous Computation Index) related to communication frequency x,y frequency Using this, the first set in which both the unit communication amount and communication frequency are determined to be constant, as shown in equation (10), is identified as constant communication.

[0055]

number

[0056] Following the above procedure, all sets of flow data that were determined to be steady communication are the second set S. x,y (S x,y and D x,y The records contained in match; d∈S x,y →d∈D x, y).

[0057] In the above embodiment, the coefficient of variation CV x , y volume CV x,y frequency The present invention was described as setting a fixed threshold and identifying each set for steady-state communication by comparison with the fixed threshold. However, the present invention is not limited to this, and the threshold ε volume and ε frequencyIt's also acceptable to allow this to be set dynamically.

[0058] In other words, by calculating the stationarity index of the unit communication amount and the stationarity index of the communication frequency for all sets, we can obtain a series like the following equation (11). Here, X represents the set of all communication sources, and Y represents the set of all communication destinations.

[0059]

number

[0060] Coefficient of variation (CV) volume CV frequency Since the smaller the values ​​of each coefficient of variation, the more stationary the communication is considered to be, if we plot the above series on a two-dimensional plane with each coefficient of variation on the vertical and horizontal axes, we can expect that the sets representing stationary communication will cluster near the origin.

[0061] Here, we classify the above sequence into two clusters using a clustering method that has the number of clusters as a hyperparameter, such as k-means. At this time, we give the initial values ​​of the center of each cluster as {(0, 0), (a, b)} (0≪a,0≪b). It is expected that the cluster that starts learning from (0, 0) will learn to include stationary communication, and the cluster that starts learning from (a, b) will learn to include transient communication.

[0062] After learning the clusters, a threshold value is set to separate the clusters with steady communication from those with transient communication. For example, if the cluster with steady communication is c0, the threshold can be determined by obtaining the maximum value as shown in equations (12) and (13).

[0063]

number

[0064]

number

[0065] Also, threshold ε volume and ε frequency Instead of defining each independently, a function F that separates the stationary communication cluster from the transient communication cluster may be used as the threshold. Examples of such a function F include methods such as SVM.

[0066] Returning to Figure 4, in step S6, the second set S x, y From two sets S x1,y1 ,S x2,y2 ∈S is selected. In step S7, the selected set pair S x1,y1 ,S x2,y2 If ∈S satisfies the first conditions of equations (14) and (15) and the second condition of equation (16) or (17), it is determined to be an address translation pair that may be undergoing address translation by NAT or NAPT.

[0067] The first condition is that there are two sets S x1,y1 ,S x2,y2 Coefficient of variation CV volume CV frequency The condition is that ε is sufficiently small, and in this embodiment, the first condition is met if equations (14) and (15) are satisfied. However, ε volume ,ε frequency This is a pre-set threshold.

[0068]

number

[0069] The second condition is two sets S x1,y1 ,S x2,y2 The flow data contained in the set has continuity. In this embodiment, as shown in Figure 6, the communication interval Δt1 of packets in one set that precedes the other in the time series (here, Sx1, y1), the communication interval Δt3 of packets in the other set that follows in the time series (here, Sx2, y2), and the set S x1,y1 From the end time of communication S x2,y2 Refer to the interval Δt2 until the start time of communication.

[0070] As a result, as shown in Fig. 7(a), if the intervals Δt1, Δt2, and Δt3 are sufficiently equal (Δt1≒Δt2≒Δt3), it is determined that the second condition is satisfied. On the other hand, as shown in Fig. 7(b), if the interval Δt1≒Δt3≠Δt2 or as shown in Fig. 7(c), if the interval Δt1≒Δt2≠Δt3, it is determined that the second condition is not satisfied.

[0071] In step S8, for the address conversion candidate pair S x1,y1 , S x2,y2 , the probability P(S x1,y1 , S x2,y2 ) that the pair is address-converted is calculated. As the probability P, any probability that receives S x1,y1 , S x2,y2 as input can be used. It may be learned using a model such as a neural network, or parameters of a well-known probability distribution may be estimated and used.

[0072] As an example of the latter, for example, an exponential distribution can be used. The exponential distribution can be regarded as a probability distribution that formulates the probability from the occurrence of an event to the occurrence of the next event. Although it is impossible to know in advance the timing when the IP address and port number change due to the rewriting of the NAPT table, it is considered that the possibility of the NAPT table being rewritten at the next time point increases as time passes from the previous table rewriting time point. Therefore, it can be assumed that the time from the previous table change to the next table change follows an exponential distribution. Thus, in the present embodiment, by assuming a specific exponential distribution in advance, as shown in Fig. 8, P(S x1,y1 , S x2,y2 ) is specifically calculated.

[0073] In step S9, based on the probability P(S x1,y1 , S x2,y2 ), it is determined whether the address conversion pair S x1,y1 , S x2,y2 is address-converted. In the present embodiment, if the following equation (16) is satisfied, the set pair S x1,y1 , Sx2,y2 It is determined that this is an address translation pair and the process proceeds to step S10. Here, ε NAPT This is a predetermined threshold.

[0074]

number

[0075] In step S10, the set pair S is determined to be address-translated. x1,y1 ,S x2,y2 D corresponding to ∈S x1,y1 ,D x2,y2 The two sets are merged by aggregating ∈D into one of the classifications. In step S11, it is determined whether there are any remaining combinations of sets that can be aggregated. The process returns to step S6 and repeats each of the above steps while switching the selected pair of sets until there are no more combinations of sets that can be aggregated.

[0076] Furthermore, according to each of the above embodiments, even when the connection information of the communication source changes dynamically, it becomes possible to accurately identify steady-state communication between the communication source and the communication destination. Therefore, it becomes possible to contribute to Goal 9, "Build resilient infrastructure and promote inclusive and sustainable industrialization," and Goal 11, "Make cities inclusive, safe, resilient and sustainable," which are led by the United Nations. [Explanation of symbols]

[0077] 1...Communication identification device, 10...Data acquisition unit, 20...Data classification unit, 30...Communication fluctuation calculation unit, 40...Steady communication identification unit, 50...Address translation candidate pair discrimination unit, 60...Address translation probability calculation unit, 70...Integration unit, 301...Unit communication volume calculation unit, 302...Communication frequency calculation unit

Claims

1. In a communication identification device that identifies communication on a network where the address of the source connection information x is translated, A means of collecting flow data from communication data on a network, The collected flow data is classified into sets unique to each combination of the source connection information x and destination connection information y, For each of the aforementioned sets, means for calculating communication fluctuations based on its flow data, A means for identifying a set of steady-state communications based on the aforementioned communication fluctuations, A means for determining from the set of steady-state communications a set of address translation candidates that may have been classified into different sets as a result of address translation, A means for calculating the probability that the pair of address translation candidates is a pair of address translations that have been classified into different sets as a result of address translation, A communication identification device characterized by comprising means for integrating a set of address translation pairs whose probabilities satisfy predetermined conditions.

2. The communication identification device according to claim 2, characterized in that the integrated set is added to the set of steady-state communications, and the determination of the set pairs of address translation candidates, the calculation of probabilities and integration are repeated until there are no more sets that can be integrated.

3. The means for calculating the communication fluctuations calculates the coefficients of variation for the unit communication amount and communication frequency based on the flow data for each set, The communication identification device according to claim 1 or 2, characterized in that the identification means identifies a communication as steady-state communication when both the coefficient of variation of the unit communication amount and the communication frequency fall below a predetermined threshold.

4. The aforementioned identification means is, A means for dynamically setting the boundary between two statistical classifications of the coefficient of variation of the unit communication volume of the aforementioned set as a threshold for the unit communication volume, The system comprises means for dynamically setting a threshold for communication frequency, which is the boundary when the coefficient of variation of the communication frequency of the aforementioned set is statistically classified into two categories. The communication identification device according to claim 3, characterized in that it identifies a communication as a steady-state communication when both the unit communication volume and the communication frequency fall below the corresponding thresholds set dynamically.

5. The communication identification device according to claim 3 or 4, characterized in that the communication frequency is the communication frequency during the period from the start time to the end time of communication.

6. The means for extracting the set pair of address translation candidates is characterized by repeatedly selecting two sets from a set of steady-state communications, and determining the set pair of address translation candidates based on the first condition that the difference in the coefficients of variation of the unit communication amount and communication frequency of the two sets falls below a predetermined threshold after each selection.

7. The means for extracting the set pair of address translation candidates is characterized in that it determines the set pair of address translation candidates, with the second condition being that the flow data of two sets that satisfy the first condition exhibits continuity as a steady-state communication.

8. The means for calculating the probability is a calculation method in which the probability increases as the elapsed time since the last address translation increases in the set of address translation candidates that precedes the previous set in the time series. This is the communication identification device according to claim 7.

9. In a communication identification method in which a computer identifies a communication on a network where the address of the source connection information x is translated, By collecting flow data from network communication data, The collected flow data is classified into sets unique to each combination of the source connection information x and destination connection information y. For each of the aforementioned sets, calculate the communication fluctuations based on the flow data. Based on the aforementioned communication fluctuations, the flow data identifies a set of steady-state communications. From the set of steady-state communications, we identify pairs of address translation candidates that may have been classified into different sets as a result of address translation. The probability that the pair of address translation candidates is a pair of address translations that have been classified into different sets as a result of the address translation is calculated. The pairs of address translations whose probabilities satisfy the predetermined conditions are combined into a single set. A communication identification method characterized by adding the integrated set to the set of steady-state communications, and repeating the identification of the set pairs of address translation candidates, the calculation of probabilities, and integration until there are no more sets that can be integrated.

10. In a communication identification program that identifies communication on a network where the address of the source connection information x is translated, Procedures for collecting flow data from network communication data, The procedure for classifying the collected flow data into sets unique to each combination of source connection information x and destination connection information y, A procedure for calculating communication fluctuations based on the flow data for each of the aforementioned sets, A procedure for identifying a set of steady communications based on the aforementioned communication fluctuations, A procedure for identifying pairs of address translation candidates from the set of steady-state communications that may have been classified into different sets as a result of address translation, A procedure for calculating the probability that the pair of address translation candidates is a pair of address translations that have been classified into different sets as a result of address translation, The procedure includes integrating a set of address translation pairs whose probabilities satisfy predetermined conditions, A communication identification program characterized by causing a computer to repeatedly perform the following actions: add the integrated set to the set of steady-state communications, and determine the set pairs of address translation candidates, calculate the probability, and integrate them until there are no more sets that can be integrated.