A network data processing method and device
Patent Information
- Application Number
- CN202110443471.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-04-23
- Publication Date
- 2026-09-25
- Estimated Expiration
- 2041-04-23
AI Technical Summary
处理高速链路需要耗费网络设备大量的设备资源,导致网络设备基本的转发性能容易受到影响
[0092]本申请实施例中,第三设备能够确定多个样本,之后根据聚类标签对多个样本进行聚类,根据得到的多个簇中各簇内的样本生成指示信息。该指示信息指示第一匹配项对应于第一类别,第二匹配项对应于第二类别,能够指示匹配第一匹配项的报文所属网络流的统计值小于匹配第二匹配项的报文所属网络流的统计值,因此,该指示信息有利于指示网络设备从接收的网络流量中识别所属网络流的统计值较大或较小的报文,从而有利于网络设备对统计值较大或较小的网络流采用适合的处理策略进行处理,从而有利于在保证对高速链路的处理性能的情况下节约设备资源。
Smart Images

Figure CN115242683B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of communication technology, and in particular to a method and apparatus for processing network data. Background Technology
[0002] With the rapid development of the internet and the continuous increase in network link speeds, supporting high-speed link processing has become a primary requirement for network devices. Processing high-speed links consumes significant network device resources, easily impacting the basic forwarding performance of these devices. Currently, there is an urgent need for a network data processing method that can conserve device resources. Summary of the Invention
[0003] This application provides a network data processing method and apparatus that can generate indication information for identifying the category of the network flow to which a packet belongs. This indication information helps to instruct network devices to adopt differentiated processing strategies for packets in different categories of network flows, thereby helping to save network device resources.
[0004] Firstly, embodiments of this application provide a method that can be applied to a first device. By executing the method provided in this application, the first device can perform different measurement strategies on different received messages according to control information. Since a message is a data unit exchanged and transmitted in a network, and is a type of network data, embodiments of this application refer to this method as a network data processing method.
[0005] The first device can be a physical entity in the network, such as a router, switch, firewall, or other network device or a chip within a network device; or it can be a computing device with computing capabilities (such as a mainframe computer or server) or a chip within a computing device. The first device can also be a virtual device (or functional module) running within a physical entity, such as a virtual router or virtual switch. This application does not limit the specific form of the first device, as long as it can execute the method provided in this application.
[0006] The control information includes at least two matching items, which include a first matching item and a second matching item. The control information instructs that a first measurement strategy be executed on a first message, and also instructs that a second measurement strategy be executed on a second message.
[0007] After receiving a message, the first device, based on the message being a first message matching the first matching item, can execute a first measurement strategy on the message according to the control information. Based on the message being a second message matching the second matching item, the first device can execute a second measurement strategy on the message according to the control information. Therefore, by implementing the method provided in this application embodiment, the first device can advantageously implement different measurement strategies for different messages.
[0008] In this embodiment, the first matching term and the second matching term respectively indicate the flow identification rules of the packets. Therefore, all packets in the network flow to which the first packet belongs match the first matching term, and all packets in the network flow to which the second packet belongs match the second matching term. Furthermore, in this embodiment, the first measurement strategy or the second measurement strategy is used to perform network flow-based measurements. Therefore, by executing the method provided in this embodiment, the first device can advantageously implement differentiated measurement strategies for multiple received network flows, thereby facilitating the execution of measurement strategies suitable for the network flows, thus balancing the measurement performance and forwarding performance of the first device.
[0009] In this embodiment, if a first message matches a first matching item, then the flow identifier in the first message satisfies the flow identification rule of the message indicated by the first matching item. For example, the first matching item includes a specific flow identifier, such as: 10.171.111.124 (source IP address), 445 (source port), 10.136.82.178 (destination IP address), 6 (transport layer protocol code), 9610 (destination port). The flow identification rule of the message indicated by the first matching item is: the 5-tuple of the message is the same as the flow identifier in the first matching item. It should be noted that this embodiment does not limit the first message to a guaranteed match if the flow identifier in the first message satisfies the flow identification rule of the message indicated by the first matching item. For example, in one possible implementation, the control information also indicates the matching order of the first and second matching items with the message. Assuming the first matching item precedes the second matching item, then even if the first message simultaneously satisfies the flow identification rules indicated by both the first and second matching items, because the first matching item precedes the second matching item, the first device determines that the first message matches the first matching item based on the control information, and will not determine that the first message matches the second matching item. In one possible implementation, any message can match at most one matching item in the control information.
[0010] In one possible implementation, the first matching term and the second matching term are used to distinguish network flows with different statistical values in the network traffic. Optionally, the statistical value of the network flow to which the first packet belongs is less than the statistical value of the network flow to which the second packet belongs. Alternatively, the first matching term is used to match packets in the network flow with the smaller statistical value in the received network traffic, and the second matching term is used to match packets in the network flow with the larger statistical value in the received network traffic. In one possible implementation, the statistical value is positively correlated with the flow size, or positively correlated with the flow rate, or positively correlated with both the flow size and the flow rate. Network flows with different statistical values generally require different suitable measurement strategies. After the first device identifies packets in the network flow with different statistical values in the network traffic based on the matching term, it executes the measurement strategy corresponding to the matching term. This facilitates the execution of a suitable measurement strategy for the packet, thereby helping to balance the measurement performance and forwarding performance of the first device.
[0011] In one possible implementation, the first device may generate the control information locally, or in another possible implementation, the first device may receive the control information sent by the third device.
[0012] The measurement strategy is used for network flow-based measurements. In conjunction with the first aspect, in a first possible implementation of the first aspect, the first and second measurement strategies utilize different device resources. When the first device measures a packet, employing device resources suitable for the network flow to which the packet belongs helps to conserve device resources while ensuring measurement performance.
[0013] The process of network devices measuring packets requires the use of one or more device resources, such as storage resources, computing resources, and network resources (e.g., network bandwidth). In a second possible implementation of the first aspect, in conjunction with the first possible implementation of the first aspect, the first measurement strategy and the second measurement strategy utilize different device resources, which can be reflected in the different device resources used in the sampling process.
[0014] In one possible implementation, a first measurement strategy is used to perform a first sampling process on the message, and a second measurement strategy is used to perform a second sampling process on the message, with the first and second sampling processes using different sampling parameters.
[0015] In one possible implementation, the sampling parameter is used to indicate a sampling counter or a storage address for recording the number of sampled data. When the first device performs the sampling process of existing network measurements, recording the number of samples for all received packets (e.g., including the first and second packets) at the same storage address or with the same sampling counter can easily lead to missed sampling of the first or second packet, resulting in the inability to measure the network flow to which the first or second packet belongs, thus reducing the accuracy of network measurements. Under the premise of using the same sampling rate, compared with performing the existing sampling process, after receiving the first and second packets, the first device, by increasing the storage resources or sampling counter used for recording the number of samples, can improve the sampling probability of the first and second packets, thereby facilitating the measurement of the network flow to which the first and second packets belong, and enabling the measurement of more received network flows, thus improving the accuracy of network measurements.
[0016] In one possible implementation, a first sampling process is used to sample packets according to a first sampling rate, and a second sampling process is used to sample packets according to a second sampling rate, wherein the first sampling rate is less than the second sampling rate. Based on the aforementioned possible implementation, where the statistical value of the network flow to which the first packet belongs is less than the statistical value of the network flow to which the second packet belongs, the first device uses a smaller sampling rate for packets in network flows with smaller statistical values. This is beneficial for improving the measurement accuracy of network flows while increasing the number of samples obtained.
[0017] In one possible implementation, the first measurement strategy and the second measurement strategy utilize different equipment resources, which can be reflected in the different equipment resources used in the recording process.
[0018] In one possible implementation, the first measurement strategy is used to instruct the first device to notify the second device to record statistical information of the network flow to which the first packet belongs. Based on the first device sampling the first packet after receiving it, the first measurement strategy is also used to instruct the first device to notify the second device to record the sampled statistical information of the network flow to which the first packet belongs. Based on the aforementioned possible implementation, where the statistical value of the network flow to which the first packet belongs is less than the statistical value of the network flow to which the second packet belongs, the first device instructing the second device to record the statistical information of the network flow to which the first packet belongs can save computational resources required for recording statistical information while sacrificing a small amount of network resources.
[0019] In one possible implementation, the first measurement strategy is also used to send message digest information to the second device. The message digest information indicates the network flow to which the first message belongs and the time when the first message was received. By sending this message digest information to the second device, the first device can instruct the second device to accurately record the statistical information of the network flow to which the first message belongs.
[0020] In one possible implementation, the first device and the second device are respectively a forwarding plane device and a control plane device. This helps to ensure the forwarding performance of the forwarding plane device and make full use of the device resources of the control plane device.
[0021] In one possible implementation, the first device includes a first storage component and a second storage component, wherein the memory access performance of the second storage component is superior to that of the first storage component. A first measurement strategy utilizes the first storage component to record statistical information of the network flow to which the first packet belongs, and a second measurement strategy utilizes the second storage component to record statistical information of the network flow to which the second packet belongs. Furthermore, based on the first device sampling the first packet after receiving it, the first measurement strategy also utilizes the first storage component to record the statistical information of the network flow to which the sampled first packet belongs. Similarly, based on the first device sampling the second packet after receiving it, the second measurement strategy also utilizes the second storage component to record the statistical information of the network flow to which the sampled second packet belongs. Based on one possible implementation described above, where the statistical value of the network flow to which the first packet belongs is less than the statistical value of the network flow to which the second packet belongs, the first device uses storage components with better memory access performance to measure packets in network flows with larger statistical values, and uses storage components with poorer memory access performance to measure packets in network flows with smaller statistical values. This helps to save storage resources with better memory access performance while ensuring the efficiency of network flow measurement, thereby helping to ensure the forwarding performance of the first device.
[0022] In one possible implementation, the first measurement strategy and the second measurement strategy utilize different equipment resources, which can be reflected in the fact that the equipment resources used in the sampling process and the recording process are different. In one possible implementation, the understanding of the different equipment resources used in the sampling process can refer to the third or fourth possible implementation in the first aspect above, and the understanding of the different equipment resources used in the recording process can refer to any one of the sixth to ninth possible implementations in the first aspect above; these will not be elaborated upon here.
[0023] In one possible implementation, the control information can be determined based on statistical values of multiple network flows. Optionally, these multiple network flows are those received by the first device before receiving the packet. Network traffic flowing through different nodes generally has different characteristics. Determining the control information based on the statistical values of the network flows flowing through the first device helps guide the first device to more accurately identify the magnitude of the statistical values of the network flows to which the received packet belongs, thereby facilitating the execution of appropriate measurement strategies for the received packet.
[0024] In one possible implementation, the control information is generated based on indication information used to classify packets in network traffic according to the statistical value of their respective network flows. In this embodiment, the indication information may be the same as or different from the control information. In this embodiment, the indication information may be generated by a first device, or it may be generated by another device (referred to as a third device). The method for a third device to generate indication information is described below through a second aspect. Possible methods for a first device to generate indication information can be obtained by replacing the third device in the second aspect with the first device, and will not be elaborated here. It should be noted that this embodiment does not limit the indication information determined by the method of the second aspect to be used to generate the control information involved in the first aspect.
[0025] Secondly, embodiments of this application provide a method that can be applied to a third device. By executing the method provided in these embodiments, the third device can generate indication information based on multiple network flows in the network traffic. Since network traffic consists of all packets arriving at a certain observation point within a certain time period, and a network flow is one or more packets with the same flow identifier within the network traffic, and a packet is a data unit exchanged and transmitted in the network, thus constituting network data, embodiments of this application refer to this method as a network data processing method.
[0026] The third device can be a physical entity in the network, such as a router, switch, firewall, or a chip within a network device; or a computing device with computing capabilities (such as a mainframe computer or server) or a chip within a computing device. The first device can also be a virtual device (or functional module) running within a physical entity, such as a virtual router or virtual switch. This application does not limit the specific form of the third device, as long as it can execute the methods provided in this application.
[0027] The third device can identify multiple samples, each of which corresponds to a network flow. Each sample includes the flow identifier and statistics of the corresponding network flow.
[0028] In one possible implementation, a third device can acquire network traffic information and then determine the multiple samples based on this information. To easily distinguish it from other network traffic, the acquired network traffic is referred to as the first network traffic. The information of the first network traffic can be used to determine the flow identifier and statistical values of each network flow within it. The flow identifier is, for example, a five-tuple in a packet. The flow identifier corresponds to multiple information units; taking a five-tuple as an example, the flow identifier includes five information units, describing the source IP address, source port, destination IP address, transport layer protocol, and destination port, respectively.
[0029] The third device determines multiple samples based on the information of the first network traffic. Each sample corresponds to a network flow within the first network traffic, and each sample includes the flow identifier and statistical value of the corresponding network flow. After determining multiple samples, the third device selects a subset of information units from the multiple information units corresponding to the flow identifier as clustering labels for clustering the multiple samples. Optionally, the number of information units corresponding to the clustering label is less than the number of information units corresponding to the flow identifier. Then, the third device can cluster the multiple samples according to the clustering labels to obtain multiple clusters. The clustering labels of any two samples within each cluster are identical. Through the clustering process, the third device can obtain clusters corresponding to at least two network flows. Then, based on the flow identifier and statistical value of each sample within each cluster, the third device can discover multiple matching items, including a first matching item and a second matching item, wherein the statistical value of the network flow to which the packet matching the first matching item belongs is greater than the statistical value of the network flow to which the packet matching the second matching item belongs. The third device generates indication information based on at least two discovered matches, the indication information indicating the category corresponding to each of the plurality of matches in a plurality of categories, wherein the first match corresponds to the first category in the plurality of categories, and the second match corresponds to the second category in the plurality of categories.
[0030] Based on this indication information, network devices can identify packets with larger or smaller statistical values from the received network traffic. This allows them to adopt appropriate processing strategies for network flows with larger or smaller statistical values, thereby saving device resources while ensuring the processing performance of high-speed links.
[0031] Since cluster labels can be manually selected, it is beneficial to monitor the clustering process and to optimize the generated instruction information by adjusting the cluster labels.
[0032] In one possible implementation, the information of the first network traffic includes information about each packet in the first network traffic, such as the packet header of each packet in the first network traffic.
[0033] In one possible implementation, the information of the first network traffic includes the flow identifier and measurement information of each network flow within the first network traffic. The measurement information is used to determine the value of the statistical information of the corresponding network flow. Assuming that the statistical value of the network flow is the flow rate of the network flow, and the measurement information of the network flow includes the number of packets within the network flow and the start and end times of the network flow being received, the third device can determine the flow rate of the network flow based on the measurement information of the network flow.
[0034] In one possible implementation, the statistical value in each of the multiple samples is positively correlated with the value of the flow size of the corresponding network flow. In this way, the indication information is helpful for identifying packets of larger or smaller flow size from the network traffic, thereby facilitating the application of appropriate processing strategies for network flows of larger or smaller flow size.
[0035] In one possible implementation, the statistical value in each of the multiple samples is positively correlated with the flow rate of the corresponding network flow. In this way, the indication information is helpful for identifying packets with higher or lower flow rates from the network traffic, thereby facilitating the application of appropriate processing strategies for network flows with higher or lower flow rates.
[0036] In one possible implementation, the statistical value in each of the multiple samples is positively correlated with the value of the flow size of the corresponding network flow and also positively correlated with the value of the flow rate of the corresponding network flow. In this way, the indication information is helpful for identifying packets from network traffic whose flow rate and flow size are both large or small, thereby facilitating the application of appropriate processing strategies for network flows with both large and small flow rates and flow sizes.
[0037] In one possible implementation, each of the multiple matches corresponds to one of the multiple clusters, each match is determined based on the clustering label of its corresponding cluster, and the category corresponding to each match is determined based on the statistical value of the samples within its corresponding cluster. Optionally, after obtaining the multiple clusters, the third device can identify at least two types of clusters (referred to as the cluster corresponding to the first category and the cluster corresponding to the second category): the statistical value of the samples in the cluster corresponding to the first category is greater than the statistical value of the samples in the cluster corresponding to the second category. Furthermore, the third device can determine the match corresponding to the first category based on the clustering label of the cluster corresponding to the first category, and indicate in the indication information that the match corresponds to the first category. The third device can also determine the match corresponding to the second category based on the clustering label of the cluster corresponding to the second category, and indicate in the indication information that the match corresponds to the second category.
[0038] Network flows with the same flow identifier generally have similar statistical values. Clustering analysis of the flow identifiers and statistical values of network flows in historical network traffic helps identify commonalities (e.g., identical cluster labels) among packets with relatively large or small statistical values. The indication information generated from these cluster labels is helpful in predicting the magnitude of the statistical value of the network flow to which a received packet belongs. Based on this indication information, network devices can identify packets with relatively large or small statistical values from received network traffic, thus enabling them to adopt appropriate processing strategies for these flows and conserving device resources while maintaining high-speed link performance.
[0039] In one possible implementation, the statistical value of samples within the cluster corresponding to the first category is less than a threshold, while the statistical value of samples within the cluster corresponding to the second category is greater than the threshold. Compared to sorting the statistical values of samples within each cluster to determine the clusters corresponding to the first and second categories, determining the clusters corresponding to the first and second categories by comparing the statistical values of samples within each cluster with the threshold helps reduce the number of comparisons, improves the efficiency of generating indication information, and solves computational resource issues.
[0040] In one possible implementation, the threshold is determined based on statistical values (e.g., the number of packets within a flow) of each network flow in the second network traffic. Compared to determining the threshold based on user experience, determining the threshold based on statistical values of each network flow in the network traffic is beneficial for classifying network flows in the first network traffic according to user needs. For example, if the user determines a threshold (i.e., the number of packets within a flow) to distinguish between elephant flows and mouse flows based on the number of packets within each network flow in the second network traffic, then the indication information generated by the third device based on this threshold is helpful in identifying packets in elephant flows and packets in mouse flows from the network traffic respectively.
[0041] In one possible implementation, the threshold is determined based on a threshold function. This threshold function takes a threshold variable as its independent variable, and the threshold is a value within the domain of the threshold variable. The threshold function is the difference between a function proportional to the number of small flows and a function proportional to the statistical values of small flows. Both the value of the function proportional to the number of small flows and the value of the function proportional to the size of small flows are related to the value of the threshold variable; therefore, the value of the threshold function is related to the value of the threshold variable. Optionally, network flows in the second network traffic that are smaller than the value of the threshold variable are called small flows. The ratio of the number of small flows in the second network traffic to the total number of network flows in the second network traffic is defined as the function proportional to the number of small flows. The ratio between the sum of the statistical values of small flows in the second network traffic and the sum of the statistical values of all network flows in the second network traffic is defined as the function proportional to the statistical values of small flows. In one possible implementation, among multiple candidate values of the threshold variable, the threshold function is maximized when the threshold variable takes the value of the threshold.
[0042] If the first message matches the first matching item, then the flow identifier in the first message satisfies the flow identification rule of the message indicated by the first matching item. For example, the first matching item includes a specific flow identifier, such as: 10.171.111.124 (source IP address), 445 (source port), 10.136.82.178 (destination IP address), 6 (transport layer protocol code), 9610 (destination port). The flow identification rule of the message indicated by the first matching item is: the five-tuple of the message is the same as the flow identifier in the first matching item. It should be noted that the embodiments of this application are not limited; if the flow identifier in the first message satisfies the flow identification rule of the message indicated by the first matching item, then the first message must match the first matching item. For example, in one possible implementation, the control information also indicates the matching order of the first matching item and the second matching item with the message. Assuming that the first matching item comes before the second matching item, even if the first message satisfies the message flow identification rules indicated by the first matching item and the second matching item at the same time, since the first matching item comes before the second matching item, the third device determines that the first message matches the first matching item based on the control information, and will not determine that the first message matches the second matching item.
[0043] In one possible implementation, the indication information also indicates the order among multiple matching items. If the target message matches the target matching item among the multiple matching items, then the target matching item is the first matching item that the target message matches from among the multiple matching items according to the order. It should be noted that the target message mentioned in this application refers to any message that can match any one of the multiple matching items, and the target matching item is the matching item that the target message matches among the multiple matching items. The word "target" is only for convenience of illustration, and this application does not limit the word "target" to having other special meanings.
[0044] Optionally, if the target packet matches the target match among multiple matches, then the target packet satisfies the target match, and the target packet does not satisfy any of the multiple matches that precede the target match in the order listed. By indicating the order among multiple matches in the indication information, it is beneficial to prioritize the matches with higher identification accuracy, thereby improving the accuracy of predicting the magnitude of the statistical value of the network flow to which the packet belongs based on the indication information.
[0045] In one possible implementation, a third device clusters multiple samples based on a first clustering label to obtain a first cluster set. The third device then clusters multiple samples based on a second clustering label to obtain a second cluster set. The multiple clusters obtained by the third device include clusters from the first cluster set and clusters from the second cluster set. In another possible implementation, the number of information units corresponding to the first clustering label is equal to the number of information units corresponding to the second clustering label. However, there is one information unit in the first clustering label whose information type differs from the information type described by any information unit in the second clustering label.
[0046] In one possible implementation, if the number of samples in a cluster (called the first cluster) in the first cluster set is greater than the number of samples in a cluster (called the second cluster) in the second cluster set, then the matching items corresponding to the first cluster among the multiple matching items take precedence over the matching items corresponding to the second cluster among the multiple matching items.
[0047] The third device performs multiple clustering operations on multiple samples based on clustering labels that describe different information types, which helps to discover more clusters, and thus helps to discover more clusters with larger or smaller statistical values. The indication information generated based on more clusters helps to identify more messages.
[0048] In one possible implementation, the third device clusters multiple samples according to a third cluster label to obtain a third cluster set. The third device then clusters multiple samples according to a fourth cluster label to obtain a fourth cluster set. The multiple clusters obtained by the third device include clusters from the third cluster set and clusters from the fourth cluster set. In another possible implementation, the number of information units corresponding to the third cluster label is greater than the number of information units corresponding to the fourth cluster label.
[0049] In one possible implementation, the matches corresponding to the third cluster among the multiple matches are prioritized over the matches corresponding to the fourth cluster among the multiple matches.
[0050] The third device performs multiple clustering operations on multiple samples based on clustering labels with different numbers of information units, which helps to discover more clusters, thereby helping to discover more clusters with larger or smaller statistical values. The indication information generated based on more clusters helps to identify more messages.
[0051] In one possible implementation, cluster labels correspond to at least two information units, which helps improve the accuracy of identifying indication information.
[0052] In one possible implementation, at least two information units include an information unit for describing the protocol number.
[0053] In one possible implementation, the first device involved in the first aspect or any possible implementation of the first aspect may also perform the method described in the second aspect or any possible implementation of the second aspect.
[0054] In one possible implementation, after the third device generates indication information based on samples within each of the multiple clusters, the third device can generate control information based on the indication information. The control information indicates that a first matching item corresponds to a first measurement strategy and a second matching item corresponds to a second measurement strategy. The first and second measurement strategies are used to perform network flow-based measurements.
[0055] In one possible implementation, the third device can receive a message, which is either a first message or a second message. The first message matches a first matching item, and the second message matches a second matching item. Then, based on the message being the first message, the third device can execute a first measurement strategy on the first message according to control information; based on the message being the second message, the third device can execute a second measurement strategy on the second message according to indication information. Optionally, the third device generates control information based on the indication information, and then executes the first or second measurement strategy on the received message according to the control information. This method can be understood by referring to the first aspect and its various possible implementations above, and will not be repeated here. After the third device generates the indication information, it can identify the magnitude of the statistical value of the network flow to which the received message belongs, thereby facilitating the execution of a suitable measurement strategy (e.g., the first or second measurement strategy) for the message, thus saving device resources while ensuring the processing performance of the high-speed link.
[0056] In one possible implementation, after the third device generates indication information based on samples within each of multiple clusters, the third device sends a message to the first device. This message, generated based on the indication information, instructs the first device to execute a first measurement strategy on a received third packet or a second measurement strategy on a received fourth packet. The third packet matches the first matching item, and the fourth packet matches the second matching item. The first and second measurement strategies are used for network flow-based measurements, respectively. After generating the indication information, the third device generates a message based on that information and sends it to the first device, as in the first aspect and its various possible implementations. This facilitates instructing the first device to identify the magnitude of the statistical value of the network flow to which the received packet belongs, thereby enabling the first device to execute a measurement strategy suitable for the packet (e.g., the first or second measurement strategy). This, in turn, helps conserve the first device's resources while ensuring the processing performance of the high-speed link.
[0057] Thirdly, embodiments of this application provide a processing apparatus, which can be the first device mentioned in the first aspect, a device within the first device, or a device compatible with the first device. In one design, the information processing apparatus may include a module corresponding to executing the method / operation / step / action described in the first aspect or any possible implementation of the first aspect. This module can be a hardware circuit, a software module, or a module implemented by combining hardware circuitry with software. For example, the information processing mask includes a storage unit and a processing unit. The storage unit stores program code, and the processing unit executes the program code in the storage unit to implement the method described in the first aspect or any possible implementation of the first aspect.
[0058] In one possible implementation, the processing device includes a communication module and a measurement module. The communication module receives a message, which is either a first message or a second message. The first message and the second message are respectively matched against a first matching item and a second matching item in control information. The control information instructs the execution of a first measurement strategy on the first message and also instructs the execution of a second measurement strategy on the second message. The first matching item and the second matching item respectively indicate flow identification rules for the message. The first measurement strategy or the second measurement strategy is used to perform network flow-based measurements. The measurement module executes the first measurement strategy or the second measurement strategy on the message according to the control information.
[0059] In one possible implementation, the first measurement strategy and the second measurement strategy utilize different device resources.
[0060] In one possible implementation, a first measurement strategy is used to perform a first sampling process on the message, and a second measurement strategy is used to perform a second sampling process on the message. The first sampling process and the second sampling process record the number of samples in different storage addresses or the same sampling counter.
[0061] In one possible implementation, the statistical value of the network flow to which the first message belongs is less than the statistical value of the network flow to which the second message belongs, and the statistical value is positively correlated with the flow size and / or positively correlated with the flow rate.
[0062] In one possible implementation, a first sampling process is used to sample the message according to a first sampling rate, a second sampling process is used to sample the message according to a second sampling rate, and the first sampling rate is less than the second sampling rate.
[0063] In one possible implementation, the first measurement strategy is further used to instruct the first device to notify the second device to record statistical information of the network flow to which the first message belongs, wherein the first message is the message sampled by the first sampling process.
[0064] In one possible implementation, the first measurement strategy is also used to send message digest information to the second device, the message digest information indicating the network flow to which the first message belongs and the time when the first message was received.
[0065] In one possible implementation, the processing device is located in the forwarding plane device, and the second device is the control plane device.
[0066] In one possible implementation, the first device includes a first storage component and a second storage component, wherein the memory access performance of the second storage component is superior to that of the first storage component. The first measurement strategy is further configured to use the first storage component to record statistical information about the network flow to which the packets sampled in the first sampling process belong. The second measurement strategy is further configured to use the second storage component to record statistical information about the network flow to which the packets sampled in the second sampling process belong.
[0067] In one possible implementation, the control information is determined based on statistical values of multiple network flows.
[0068] In one possible implementation, multiple network flows are received by the first device before it receives the message.
[0069] For the beneficial effects in this regard, please refer to the relevant introduction in the first aspect mentioned above; details will not be repeated here.
[0070] Fourthly, embodiments of this application provide a processing apparatus, which can be the third device mentioned in the second aspect, a device within the third device, or a device compatible with the third device. In one design, the information processing apparatus may include a module corresponding to executing the method / operation / step / action described in the second aspect or any possible implementation of the second aspect. This module may be a hardware circuit, a software module, or a module implemented by combining hardware circuitry with software. For example, the information processing mask includes a storage unit and a processing unit. The storage unit stores program code, and the processing unit executes the program code in the storage unit to implement the method described in the second aspect or any possible implementation of the second aspect.
[0071] In one possible implementation, the processing apparatus includes a determination module, a clustering module, and a generation module. The first device determines multiple samples based on information from network traffic (referred to as first network traffic). Each sample corresponds to a network flow within the first network traffic, and each sample includes a flow identifier and a statistical value for the corresponding network flow. The clustering module clusters the multiple samples to obtain multiple clusters. Each cluster corresponds to a cluster label, which is the same information unit of the flow identifier in all samples within the corresponding cluster. The generation module generates indication information based on the samples within each cluster. The indication information includes multiple matching items and the category corresponding to each matching item in multiple categories. Each matching item indicates the flow identification rule of the packet. The indication information indicates that the statistical value of the network flow to which the packet matching the first matching item belongs is less than the statistical value of the network flow to which the packet matching the second matching item belongs. The first matching item is the matching item corresponding to the first category among the multiple matching items, and the second matching item is the matching item corresponding to the second category among the multiple matching items. The first category and the second category are different categories among the multiple categories.
[0072] In one possible implementation, each match corresponds to one of multiple clusters, each match is determined based on the clustering label of the corresponding cluster, and the category corresponding to each match is determined based on the statistical values in the samples within the corresponding cluster.
[0073] In one possible implementation, the statistical value of the samples within the cluster corresponding to the first category is less than a threshold, and the statistical value of the samples within the cluster corresponding to the second category is greater than a threshold.
[0074] In one possible implementation, the indication information also indicates the order among multiple matching items. If the target message matches the target matching item among the multiple matching items, then the target message satisfies the target matching item, and the target message does not satisfy any of the multiple matching items that are preceding the target matching item in the order of priority.
[0075] In one possible implementation, the first cluster and the second cluster in the multiple clusters correspond to the first cluster label and the second cluster label, respectively. The number of information units corresponding to the first cluster label is equal to the number of information units corresponding to the second cluster label. The number of samples in the first cluster is greater than the number of samples in the second cluster. Among the multiple matching items, the matching items corresponding to the first cluster are prioritized before the matching items corresponding to the second cluster.
[0076] In one possible implementation, the third and fourth clusters in the multiple clusters correspond to the third cluster label and the fourth cluster label, respectively. The number of information units corresponding to the third cluster label is greater than the number of information units corresponding to the fourth cluster label. Among the multiple matching items, the matching items corresponding to the third cluster are prioritized over the matching items corresponding to the fourth cluster.
[0077] In one possible implementation, the statistical value in each sample is positively correlated with the value of the flow size of the corresponding network flow, and / or positively correlated with the value of the flow rate of the corresponding network flow.
[0078] In one possible implementation, the cluster label corresponds to at least two information units.
[0079] In one possible implementation, at least two information units include an information unit for describing the protocol number.
[0080] In one possible implementation, the processing device further includes a communication module for generating control information based on the indication information after the generation module generates indication information based on samples in each of the multiple clusters. The control information indicates that a first matching item corresponds to a first measurement strategy and a second matching item corresponds to a second measurement strategy. The first and second measurement strategies are used to perform network flow-based measurements.
[0081] In one possible implementation, the third device can receive a message, which is either a first message or a second message, wherein the first message matches a first matching item and the second message matches a second matching item. The processing apparatus further includes a measurement module configured to execute a first measurement strategy on the first message based on indication information, or to execute a second measurement strategy on the second message based on indication information.
[0082] In one possible implementation, the processing device further includes a communication module, which is used to send a message to a second device after the generation module generates indication information based on samples in each of the multiple clusters. The message is generated based on the indication information and is used to instruct the second device to perform a first measurement strategy on a received third message or a second measurement strategy on a received fourth message. The third message matches the first matching item, and the fourth message matches the second matching item. The first measurement strategy and the second measurement strategy are respectively used to perform network flow-based measurements.
[0083] For the beneficial effects in this regard, please refer to the relevant introduction in the second aspect above, which will not be repeated here.
[0084] Fifthly, embodiments of this application provide a processing apparatus including a processor and a memory, the processor and the memory being coupled together, the memory being used to store program code, and when the processor executes the program code stored in the memory, it is able to execute the method described in the first aspect or any possible implementation of the first aspect or the second aspect or any possible implementation of the second aspect.
[0085] In one possible implementation, these instructions are stored in memory external to the processing device. When these instructions are decoded and executed by the processor of the processing device, some or all of the contents of the instructions are temporarily stored in the memory internal to the processing device. Optionally, some of the contents of these instructions are stored in memory external to the processing device, while other portions of the instructions are stored in memory internal to the processing device.
[0086] Based on the fifth aspect, in one possible design, the processing device may also include a communication interface through which the processor can, for example, receive messages or send message summary information to a second device.
[0087] A sixth aspect of this application provides a chip system including a processor and an interface circuit. The processor is coupled to a memory via the interface circuit. The processor executes program code in the memory to perform the methods described in the first aspect or any possible implementation thereof, or the second aspect or any possible implementation thereof. The chip system may be composed of chips or may include chips and other discrete devices.
[0088] A seventh aspect of this application provides a computer-readable storage medium storing instructions (or computer-readable instructions, computer program instructions, functional programs, or program code) that, when executed on a computer device, cause the computer device to perform the methods described in the embodiments of this application, which are capable of performing the first aspect or any possible implementation of the first aspect, or the second aspect or any possible implementation of the second aspect.
[0089] The eighth aspect of this application provides a computer program product containing instructions (or computer-readable instructions or computer program instructions or functional programs or program code) that, when executed by a computer device, implement the method described in the embodiments of this application, which can perform the first aspect or any possible implementation of the first aspect or the second aspect or any possible implementation of the second aspect.
[0090] Since the devices provided in this application can be used to execute the methods of the corresponding embodiments described above, the technical effects that can be obtained by the device embodiments of this application can be referred to the technical effects obtained by the corresponding method embodiments described above, and will not be repeated here.
[0091] As can be seen from the above technical solutions, the embodiments of this application have the following advantages:
[0092] In this embodiment, the third device can identify multiple samples, then cluster these samples according to clustering labels, and generate indication information based on the samples within each of the resulting clusters. This indication information indicates that a first matching item corresponds to a first category, a second matching item corresponds to a second category, and indicates that the statistical value of the network flow to which the packet matching the first matching item belongs is less than the statistical value of the network flow to which the packet matching the second matching item belongs. Therefore, this indication information helps the network device identify packets with larger or smaller statistical values from the received network traffic, thereby enabling the network device to adopt appropriate processing strategies for network flows with larger or smaller statistical values, thus saving device resources while ensuring the processing performance of high-speed links.
[0093] The specification of this application document serves to support the technical solutions described in the claims of this application document. The third device mentioned in the specification corresponds to the first device mentioned in the claims of this application document, and the first device mentioned in the specification corresponds to the second device mentioned in the claims of this application document. The method performed by the first device in the claims of this application document can be understood by referring to the corresponding method performed by the third device in the specification of this application document, and the second device involved in the claims of this application document can be understood by referring to the first device mentioned in the specification of this application document. Attached Figure Description
[0094] Figure 1 This application illustrates one possible network environment related to an embodiment of the present application;
[0095] Figure 2 This illustrates one possible structure for a network device;
[0096] Figure 3 This application illustrates a possible flow of a method for processing network data according to an embodiment of the present application;
[0097] Figure 4A This illustrates a sampling process in an existing network measurement workflow;
[0098] Figure 4B and Figure 4C Each of these illustrates a possible sampling process in the network measurement workflow of this application;
[0099] Figure 4D This illustration shows a possible schematic diagram of the measurement results generated by an embodiment of this application;
[0100] Figure 5 This application illustrates another possible flow of the network data processing method according to an embodiment of the present application;
[0101] Figure 6A , Figure 6B , Figure 6C , Figure 6D and Figure 6E They are shown respectively Figure 5 The step 502 shown may result in multiple clusters;
[0102] Figure 7A , Figure 7B and Figure 7C The following are examples of clustering parameters that may be used in embodiments of this application;
[0103] Figure 8 F is shown f (T) and F p (T) Possible function graphs;
[0104] Figure 9 A schematic diagram of a possible structure of the processing apparatus according to an embodiment of this application is shown;
[0105] Figure 10 This paper illustrates another possible structural diagram of the processing apparatus according to an embodiment of the present application;
[0106] Figure 11 This illustration shows another possible structural diagram of the processing apparatus according to an embodiment of this application. Detailed Implementation
[0107] The following explanations of some terms used in this application are provided to facilitate understanding by those skilled in the art.
[0108] Network traffic refers to all messages that arrive at a specific observation point within a certain time period.
[0109] A network flow refers to one or more packets in network traffic that share the same flow identifier.
[0110] A flow identifier, or flow tag, is one or more information elements in a message that define the network flow to which the message belongs. These elements describe one or more of the following information types: source IP address, destination IP address, source port number, network layer protocol type, IP service type, or router / switch interface number. These elements may all be in the message header, all in the message payload, or partially in the header and partially in the payload. The information elements describing the source IP address, source port, destination IP address, transport layer protocol, and destination port are generally called the message's 5-tuple, which is most commonly used as the flow identifier to define the network flow to which the message belongs.
[0111] Stream size refers to the number of packets and / or bytes within a stream. The number of packets refers to the total number of packets in the network stream, and the number of bytes refers to the total number of bytes in the network stream.
[0112] Flow rate refers to the average arrival rate of a network flow. For example, the flow rate can be calculated based on the start and end times of the network flow being received and the number of packets (or bytes) within the flow.
[0113] Network flow statistics refer to information obtained by statistically analyzing the packets in a network flow. For example, network flow statistics generally include the number of packets and / or bytes in the flow. In addition, network flow statistics may also include the start and end times of the network flow arrival.
[0114] Network flow-based measurement is a commonly used technique in network measurement. By measuring network traffic in a flow-based manner, multiple packets within the network traffic can be aggregated into multiple network flows, and statistical information of each network flow can be recorded, thereby achieving the purpose of measuring network traffic. The measurement results of network flow-based measurement generally include statistical information of each network flow within the network traffic.
[0115] Network flow statistics are obtained by recording information about each packet in the network flow. Therefore, measuring network traffic based on network flow can be understood as measuring each packet in the network flow based on network flow.
[0116] The following is an example of a network environment in which network traffic can be measured.
[0117] Figure 1 This is a schematic diagram of a network environment. (Reference) Figure 1 The network environment includes source device 11, receiving device 12, receiving device 13 and network device 14, and source device 11 and receiving device 12 are connected via network device 14, and source device 11 and receiving device 13 are connected via network device 14.
[0118] In this configuration, source device 11 generates network traffic, which then reaches network device 14. Network device 14 performs flow-based measurements on the network traffic and forwards the network flows within it. For example, refer to... Figure 1 Network device 14 forwards network flow 1 from the network traffic to receiving device 12, and forwards network flow 2 from the network traffic to receiving device 13. Receiving devices 12 and 13 are used to receive the network flows forwarded by network device 14.
[0119] Figure 1 The source device 11 in the diagram is a server. Figure 1The receiving devices 12 and 13 in the illustrations are based on user equipment (desktop computers). These illustrations are for illustrative purposes only and are not limiting. Source device 11, receiving device 12, and receiving device 13 can be other types of devices. For example, source device 11 can be a user equipment, and receiving device 12 can be a server. In some scenarios, source device 11 can also be used to receive network traffic, and receiving device 12 or receiving device 13 can also be used to generate and send network traffic.
[0120] Figure 1 The network device 14 in the illustration is a switch, but this illustration is for illustrative purposes only and not a limitation. The network device 14 can be other types of devices. For example, the network device 14 can be a physical entity in the network, such as a router, switch, or firewall; the network device 14 can also be a virtual network device running within a physical entity, such as a virtual router or virtual switch; the network device 14 can also be a mainframe computer, a server with computing capabilities, or other computing devices. This application does not limit the specific form of the first device, as long as the first device can execute the method provided in this application.
[0121] The following is combined Figure 2 Example of network device 14 is given.
[0122] Figure 2 This is a schematic diagram of the structure of network device 14. (Reference) Figure 2 Network device 14 includes a control plane device 141 (or control plane) and a forwarding plane device 142 (or forwarding plane or data plane). The control plane device 141 is generally used to transmit instructions and calculate entries, such as protocol message forwarding, protocol entry calculation and maintenance. The forwarding plane device 142 is generally used to receive, decapsulate, encapsulate and forward data packets. After performing protocol interaction and route calculation, the control plane device 141 generates several entries and sends them to the forwarding plane device 142 to guide the forwarding plane device 142 in forwarding packets.
[0123] Optionally, the control plane device 141 and the forwarding plane device 142 can be physically separated. For example, the central processing unit (CPU) on the main control board of the network device is not responsible for packet forwarding but focuses on system control, while the service board focuses on packet forwarding. Alternatively, the control plane device 141 and the forwarding plane device 142 can also be logically separated. That is, after the network device starts up, the system allocates CPU and memory resources to different processes. The process corresponding to the control plane device 141 is responsible for learning routes, and the process corresponding to the forwarding plane device 142 is responsible for packet forwarding.
[0124] Figure 1 and Figure 2 The network device 14 shown can not only forward received packets but also perform network flow-based measurements on them. However, as network link speeds continue to increase, the device resources (e.g., processing power and storage resources) consumed by network device 14 in measuring network traffic are constantly increasing, affecting basic forwarding performance. To ensure forwarding performance, in the prior art, network device 14 uses resource-saving methods such as packet sampling to measure network traffic. While these methods help reduce the device resources consumed by network device 14 during measurement, they can easily reduce the measurement performance (e.g., accuracy and / or efficiency) of some network flows. For example, a high sampling rate can easily lead to missed sampling of mouse flows, threatening network security. Furthermore, using storage resources with low memory access performance to measure high-flow-rate network flows can easily lead to packet loss, reducing measurement and forwarding performance.
[0125] To address the aforementioned technical problems, embodiments of this application provide a method for processing network data. This method is used to implement differentiated measurement strategies for different types of network flows in network traffic. This facilitates the adoption of suitable measurement strategies for each network flow, thereby reducing the resource occupancy rate of the measurement process while ensuring the accuracy of the measurement results for each network flow.
[0126] The following is a detailed description of the network data processing method provided in the embodiments of this application.
[0127] refer to Figure 3 This application provides a method for processing network data. In one possible implementation, the method includes steps 301 and 302.
[0128] Step 301. The first device receives the message;
[0129] Step 302. The first device executes the first measurement strategy or the second measurement strategy on the message according to the control information.
[0130] Steps 301 and 302 will be described below.
[0131] The first device can be a physical entity in a network, such as a router, switch, or firewall; it can also be a virtual network device running within a physical entity, such as a virtual router or virtual switch; or it can be a mainframe computer, a server with computing capabilities, or other computing devices. This application does not limit the specific form of the first device, as long as it can execute the method provided in this application. For example, the first device is... Figure 1 or Figure 2The network device 14 shown is, or is Figure 2 The forwarding plane device 142 shown.
[0132] The first device is capable of acquiring control information. This control information includes at least two matching items (fields), which are referred to as the first matching item and the second matching item, respectively.
[0133] This matching term indicates the flow identification rules for the message. For any one of multiple matching terms (called the target matching term), in one possible implementation, the target matching term indicates the conditions that the hash value of the flow identifier in the message must satisfy. For example, if the hash value of the flow identifier in the message is the same as the target matching term, then the message is determined to satisfy the flow identification rules indicated by the target matching term. Alternatively, in one possible implementation, the target matching term is the flow identifier in the matched message, such as a 5-tuple (10.136.82.184, 8080, 10.136.82.180, 9610, 6), or a partial field of a 5-tuple (10.136.82, 8080, 10.136.82, 9610, 6). If the flow identifier in the message is the same as the target matching term, then the message is determined to satisfy the flow identification rules indicated by the target matching term.
[0134] In this embodiment, if a message matches a first matching item, then the message satisfies the flow identification rule indicated by the first matching item. However, this embodiment does not limit that a message satisfying the flow identification rule indicated by the first matching item must necessarily match that matching item. For example, the control information may also indicate the order among multiple matching items. If a message satisfies matching item 1 and matching item 2 among multiple matching items, and the control information indicates that matching item 1 precedes matching item 2, then the message matches matching item 1 but not matching item 2.
[0135] Furthermore, the control information instructs the execution of a first measurement policy for packets matching a first match and a second measurement policy for packets matching a second match. Both the first and second measurement policies are used for network flow-based measurements. Optionally, the control information may indicate at least two mappings, one of which (called the first mapping) represents the mapping between the first match and the first measurement policy, and the other (called the second mapping) represents the mapping between the second match and the second measurement policy. The first device can execute the first measurement policy on packets matching the first match based on the first mapping, and execute the second measurement policy on packets matching the second match based on the second mapping. In one possible implementation, the control information may be stored, for example, in the form of Table 1. For example, the control information may be an access control list (ACL) used to control the measurement process of network traffic.
[0136] Table 1
[0137] First match First Measurement Strategy Second match Second measurement strategy
[0138] In one possible implementation, the control information uses a first storage address and a second storage address to represent a first measurement strategy and a second measurement strategy, respectively. The first and second storage addresses store codes for executing the first and second measurement strategies, respectively. After receiving a message matching the first matching item, the first device can execute the first measurement strategy on the message by looking up the code in the first storage address and executing the corresponding code. Similarly, after receiving a message matching the second matching item, the first device can execute the second measurement strategy on the message by looking up the code in the second storage address and executing the corresponding code.
[0139] The method by which the first device acquires control information will be further described later, and will not be elaborated here.
[0140] During operation, the first device can receive a message that matches either the first matching item or the second matching item mentioned above. Upon receiving the message, the first device can execute either the first measurement strategy or the second measurement strategy based on the control information. For example, based on the message matching the first matching item, the first device executes the first measurement strategy based on the control information. Similarly, based on the message matching the second matching item, the first device executes the second measurement strategy based on the control information.
[0141] Alternatively, during operation, the first device can receive a message matching a first matching item (referred to as a first message) and a message matching a second matching item (referred to as a second message). Furthermore, upon receiving the first message, the first device can execute a first measurement strategy on the first message. And upon receiving the second message, the first device can execute a second measurement strategy on the second message.
[0142] Because the matching entries in the control information indicate the flow identification rules for packets, all packets in the same network flow satisfy the same matching entries in the control information. Therefore, Figure 3 The corresponding implementation method can perform differentiated measurement strategies for different network flows in the network traffic. This is beneficial for adopting a suitable measurement strategy for each network flow, thereby reducing the resource consumption of the measurement process while ensuring the accuracy of the measurement results of each network flow.
[0143] In this embodiment of the application, the first measurement strategy and the second measurement strategy are different. The possible differences between the two are described below from different perspectives.
[0144] Perspective 1: In one possible implementation, the difference between the two lies in the different equipment resources utilized.
[0145] In the process of measuring the received messages, the first device needs to utilize its own device resources. These resources include, for example, storage resources (e.g., storage units in a storage component), computing resources, and network resources. The first measurement strategy differs from the second measurement strategy. In one possible implementation, this difference lies in the different device resources utilized by the first and second measurement strategies. Table 2 provides examples of several possible differences.
[0146] Table 2
[0147]
[0148] Based on Table 2, the equipment resources utilized by the first measurement strategy and the equipment resources utilized by the second measurement strategy are different. In one possible implementation, this difference can be reflected in any one or more of the differences in Table 2.
[0149] The differences between the first and second measurement strategies are illustrated below with examples from Table 2.
[0150] For ease of description, the storage resource utilized by the first measurement strategy is referred to as the first storage resource, the storage resource utilized by the second measurement strategy is referred to as the second storage resource, the storage component containing the first storage resource is referred to as the first storage component, and the storage component containing the second storage resource is referred to as the second storage component. Referring to Table 2, optionally, the first storage resource and the second storage resource are different. In one possible implementation, this difference may manifest as different memory access performance between the first storage component and the second storage component. Specifically, in one possible implementation, the first storage component and the second storage component are of different types. For example, the first storage component is static random-access memory (SRAM), and the second storage component is reduced-latency dynamic random-access memory (RLDRAM). Different types of storage components may have different memory access performance. Alternatively, in one possible implementation, the first storage component and the second storage component are located at different positions relative to the chip. For example, the first storage component is located on-chip within the chip containing the processor, while the second storage component is located off-chip within the chip containing the processor. Thus, even if the first and second storage components are of the same type, their memory access performance differs due to their different locations relative to the chip.
[0151] Referring to Table 2, optionally, the first storage resource and the second storage resource are different. In one possible implementation, this difference is reflected in the fact that even if the first storage component and the second storage component are the same, the addresses of the first storage resource and the second storage resource in the storage component are different.
[0152] For ease of description, the computing resources utilized by the first measurement strategy are referred to as the first computing resources, and the computing resources utilized by the second measurement strategy are referred to as the second computing resources. Referring to Table 2, optionally, the first and second computing resources may differ. In one possible implementation, this difference manifests as a difference in the computational complexity involved in the first and second measurement strategies. For example, the first and second computing resources may correspond to different computation types and / or different number of computations. Computation types may include arithmetic operations, read operations, write operations, send operations, and receive operations. Different computation types generally have different complexities. For the same computation type, the more computations performed, the higher the complexity. Different computational complexities generally result in different processor utilization rates.
[0153] For ease of description, the network resources utilized by the first measurement strategy are referred to as the first network resources, and the network resources utilized by the second measurement strategy are referred to as the second network resources. Referring to Table 2, optionally, the first network resources and the second network resources are different. In one possible implementation, this difference is reflected in the different network bandwidths utilized by the first and second measurement strategies.
[0154] Angle 2: In one possible implementation, the difference between the two lies in the different measurement processes.
[0155] The measurement process of a packet by the first device generally includes a sampling process and a process of recording statistical information of the network flow to which the packet belongs (referred to as the recording process). For ease of description, the sampling process corresponding to the first measurement strategy is called the first sampling process, the sampling process corresponding to the second measurement strategy is called the second sampling process, the recording process corresponding to the first measurement strategy is called the first recording process, and the recording process corresponding to the second measurement strategy is called the second recording process. Table 3 illustrates several possible differences.
[0156] Table 3
[0157] different same same different different different
[0158] The first measurement strategy differs from the second measurement strategy. In one possible implementation, this difference can be reflected in the difference between the first sampling process and the second sampling process. Alternatively, the difference between the two sampling processes can be reflected in the difference in the equipment resources used in the two sampling processes, or in the difference in the sampling parameters used in the two sampling processes.
[0159] Alternatively, the first measurement strategy may differ from the second measurement strategy. In one possible implementation, this difference may be reflected in the difference between the first recording process and the second recording process. Optionally, the difference between the two recording processes may be reflected in the difference in the equipment resources utilized by the two recording processes.
[0160] Alternatively, the first measurement strategy may differ from the second measurement strategy. In one possible implementation, this difference may be reflected both in the first and second sampling processes being different, and in the first and second recording processes being different.
[0161] The previous text introduced the possible differences between the first and second measurement strategies from different perspectives. The following text uses examples from perspectives 1 and 2 to illustrate the differences between the first and second measurement strategies.
[0162] Example 1: The first sampling process and the second sampling process utilize storage resources at different addresses in the storage unit.
[0163] In existing network measurement technologies, network devices perform uniform sampling on received packets. This uniform sampling is achieved by using storage resources at the same address in the storage unit for sampling. For example, they may use storage resources at the same address or the same sampling counter to record the number of samples, and then sample the received packets based on the recorded number of samples.
[0164] To facilitate comparison between the embodiments of this application and the prior art, the following provides an example of an existing network measurement process.
[0165] Continue to refer to Figure 1 The network environment shown is assumed to be... Figure 1 The network traffic shown (filled with black dots) contains packets, such as Figure 4A As shown, the packets arrive in the following order from first to last: P1, P2, P3, ..., and P9, with arrival times of t1, t2, t3, ..., and t9 respectively. It is assumed that the flow identifiers in P1, P3, P4, P5, P6, P7, and P9 are all flow identifier 1, corresponding to... Figure 1 The network flow shown is 1. Assume that the flow identifiers for P2 and P8 are both flow identifier 2, corresponding to... Figure 1 Network flow 2 is shown.
[0166] A sampling rate of X:1 means that sampling is performed once every X-1 packets, or in other words, one packet is sampled for every X packets received. In existing network measurement sampling processes, it is assumed that network device 14 is configured with a sampling rate of 3:1, and network device 14 utilizes storage resources at the same address (e.g., the same sampling counter) to record the number of samples. Figure 4A As shown, network device 14 utilizes the storage resources corresponding to address 1 in the storage component or the sampling counter (in Figure 4A The number of samples is recorded using the box indicated by address 1. Therefore, when P9 arrives at network device 14, as... Figure 4A As shown, the number of samples corresponding to address 1 is 9. Assuming the network device starts sampling when the number of samples is 1, since the number of samples corresponding to packets P1, P4, and P7 are 1, 4, and 7 respectively, therefore, as... Figure 4A The packets indicated by the arrows in the diagram include samples P1, P4, and P7 collected by network device 14. After collecting the packets, network device 14 performs a recording process on the collected packets. The following section combines... Figure 4A This describes the recording process.
[0167] like Figure 4AAs shown in table a, which is pointed to by P1, when network device 14 collects P1, it creates a new record 1 corresponding to network flow 1 based on the information of P1. Record 1 represents the following in sequence: the flow identifier of network flow 1 is flow identifier 1, the number of packets in network flow 1 is 1, the arrival time of the first packet in network flow 1 is t1, and the arrival time of the last packet in network flow 1 is t1.
[0168] like Figure 4A As shown in table a pointed to by P4, when network device 14 collects P4, it updates record 1 corresponding to network flow 1 based on the information in P4. By comparing table a pointed to by P1 and P4, it is easy to see that network device 14 updates the number of packets in network flow 1 to 2 and updates the arrival time of the last packet in network flow 1 to t4.
[0169] like Figure 4A As shown in table a pointed to by P7, when network device 14 collects P7, it updates record 1 corresponding to network flow 1 based on the information in P7. By comparing table a pointed to by P4 and P7, it is easy to see that network device 14 updates the number of packets in network flow 1 to 3 and updates the arrival time of the last packet in network flow 1 to t7.
[0170] Because network device 14 did not sample P8 and P9, therefore, Figure 4A Table a, pointed to by P7, shows the measurement results of network traffic (filled with black dots) for network device 14.
[0171] pass Figure 4A As can be seen from the network measurement process shown, since network device 14 did not sample P2 or P8, the measurement results obtained by network device 14 cannot reflect the statistical information of network flow 2, the accuracy of the measurement results is low, and it is not conducive to discovering threats to network security.
[0172] In the embodiments of this application, the first sampling process and the second sampling process sample the received messages independently, which is reflected in the use of storage resources at different addresses in the storage component for sampling. For example, storage resources at different addresses (e.g., different sampling counters) are used to record the number of samples, and then the received messages are sampled independently based on the recorded number of samples.
[0173] The following example illustrates the network measurement process corresponding to the embodiments of this application, so as to help understand the differences between the embodiments of this application and the aforementioned prior art.
[0174] Continue to refer to Figure 1 The network environment shown is assumed to be... Figure 1 The network traffic shown (filled with black dots) contains packets such as Figure 4B As shown, with Figure 4AThe situation is the same as shown, that is, according to the order of arrival from first to last: P1, P2, P3, ..., and P9, and the arrival times of each message are t1, t2, t3, ..., and t9 respectively. Here, it is assumed that the flow identifiers in P1, P3, P4, P5, P6, P7, and P9 are all flow identifier 1, corresponding to... Figure 1 The network flow shown is 1. Assume that the flow identifiers for P2 and P8 are both flow identifier 2, corresponding to... Figure 1 Network flow 2 is shown.
[0175] During the sampling process of network measurement in this embodiment of the application, and Figure 4A Similarly, assuming network device 14 is configured with a sampling rate of 3:1 during the sampling process, and... Figure 4A Unlike other examples, network device 14 utilizes storage resources (e.g., a sample counter) at different addresses to record the number of samples. Figure 4B As shown, network device 14 utilizes the storage resources corresponding to address 1 in the storage component (in Figure 4B The number of samples of the packet corresponding to stream identifier 1 is recorded using the box indicated by address 1, and the storage resource corresponding to address 2 is used (in...). Figure 4B The number of samples of the packet corresponding to flow identifier 2 is recorded using the box indicated by address 2. This embodiment does not limit the storage resources corresponding to address 1 and address 2 to be in the same storage component. When P9 arrives at network device 14, as... Figure 4B As shown, the number of samples corresponding to address 1 is 7, and the number of samples corresponding to address 2 is 2. Assuming the network device starts sampling when the number of samples is 1, since the number of samples for packets P1, P5, and P9 corresponding to address 1 are 1, 4, and 7 respectively, and the number of samples for packet 2 corresponding to address 2 is 1, the samples collected by network device 14 include P1, P2, P5, and P9. After collecting the packets, network device 14 performs a recording process on the collected packets. The following section combines... Figure 4B Let's continue with the recording process.
[0176] like Figure 4B As shown in table b pointed to by P1, when network device 14 collects P1, it creates a new record 1 corresponding to network flow 1 based on the information of P1. Record 1 represents the following in sequence: the flow identifier of network flow 1 is flow identifier 1, the number of packets in network flow 1 is 1, the arrival time of the first packet in network flow 1 is t1, and the arrival time of the last packet in network flow 1 is t1.
[0177] like Figure 4BAs shown in table b pointed to by P2, when network device 14 collects P2, it creates a new record 2 corresponding to network flow 2 based on the information of P2. Record 2 represents the following in sequence: the flow identifier of network flow 2 is flow identifier 2, the number of packets in network flow 2 is 1, the arrival time of the first packet in network flow 2 is t2, and the arrival time of the last packet in network flow 2 is t2.
[0178] like Figure 4B As shown in table b pointed to by P5, when network device 14 collects P5, it updates record 1 corresponding to network flow 1 based on the information in P5. By comparing table b pointed to by P1 and P5, it is easy to see that network device 14 updates the number of packets in network flow 1 to 2 and updates the arrival time of the last packet in network flow 1 to t5.
[0179] like Figure 4B As shown in table b pointed to by P9, when network device 14 collects P9, it updates record 1 corresponding to network flow 1 based on the information in P9. By comparing table b pointed to by P5 and P9, it is easy to see that network device 14 updates the number of packets in network flow 1 to 3 and updates the arrival time of the last packet in network flow 1 to t9.
[0180] Since P9 is the last packet collected by network device 14 from the network traffic (filled with black dots), therefore, Figure 4B Table b, pointed to by P9, shows the measurement results of network traffic (filled with black dots) for network device 14.
[0181] pass Figure 4B As can be seen from the network measurement process shown, since network device 14 samples the packets of network flow 1 and network flow 2, the measurement results obtained by network device 14 can reflect the statistical information of each network flow in the network traffic (black dot filling). Compared with the existing technology, it is beneficial to reduce the probability of missing network flow samples, improve the accuracy of measurement results, and facilitate the timely detection of threats to network security.
[0182] In this embodiment, the first device performs a differentiated sampling process on the received packets based on matching items in the control information, which includes a first matching item and a second matching item. The following assumptions are made regarding the network flows corresponding to the first and second matching items: the statistical value of the network flow to which the packet matching the first matching item belongs is less than the statistical value of the network flow to which the packet matching the second matching item belongs.
[0183] For ease of description, in the embodiments of this application, 's' represents a statistical value, 'n' represents the flow size, and 'r' represents the flow rate. In one possible implementation, the 's' of the network flow is determined based on the 'n' of the network flow. Optionally, the 's' of the network flow is positively correlated with the 'n' of the network flow, for example, 's = k1 * n', where k1 is a determined value. For example, the network flow to which the packet matching the first match belongs is the 'mouse' flow, and the network flow to which the packet matching the second match belongs is the 'elephant' flow. Alternatively, in one possible implementation, the 's' of the network flow is determined based on the 'r' of the network flow. Optionally, the 's' of the network flow is positively correlated with the 'r' of the network flow, for example, 's = k2 * r', where k2 is a determined value. Alternatively, in one possible implementation, the network flow s is determined based on the network flow n and r. Optionally, the network flow s is positively correlated with the network flow n, and the network flow s is positively correlated with the network flow r, for example, s = k1*n + k2*r, where k1 and k2 are determined values. It should be noted that the above formula is only an example, and the embodiments of this application do not limit the specific method of determining the statistical value of the network flow.
[0184] Based on this assumption, in one possible implementation, the first sampling process samples the packets according to a first sampling rate, and the second sampling process samples the packets according to a second sampling rate, where the first sampling rate is less than the second sampling rate. Below, we take the network flow matching the first match as a "mouse flow" and the network flow matching the second match as an "elephant flow" as examples, combined with... Figure 4C This implementation method will be illustrated with an example.
[0185] Continue to refer to Figure 1 The network environment shown is assumed to be... Figure 1 The network traffic shown (filled with black dots) contains packets such as Figure 4C As shown, with Figure 4B The situation is the same as shown, so it will not be repeated here. Assume that packets P2 and P8 match the first match, and network flow 1 is the mouse flow; packets P1, P3, P4, P5, P6, P7, and P9 match the second match, and network flow 2 is the elephant flow. Figure 4C During the sampling process of the network measurement shown, with Figure 4B Unlike other methods, suppose the sampling rate of the first sampling process of network device 14 is 1:1, and the sampling rate of the second sampling process is 3:1. The sampling rate of the first sampling process is less than the sampling rate of the second sampling process. Figure 4B The same example is found in... Figure 4C In the corresponding example, network device 14 utilizes storage resources at different addresses to execute the first sampling process and the second sampling process respectively. For example... Figure 4C As shown, network device 14 utilizes the storage resources corresponding to address 1 in the storage component (in Figure 4CThe number of samples of the packet corresponding to stream identifier 1 is recorded using the box indicated by address 1, and the storage resource corresponding to address 2 is used (in...). Figure 4C The number of samples of the message corresponding to flow identifier 2 is recorded using the box indicated by address 2. This embodiment of the application does not limit the storage resources corresponding to address 1 and address 2 to be in the same storage component.
[0186] Assuming the network device starts sampling when the number of samples is 1, and Figure 4B The corresponding examples are different, in Figure 4C In the corresponding example, the samples collected by network device 14, in addition to P1, P2, P5, and P9, also include P8. Accordingly, refer to... Figure 4C After network device 14 collects data from P8, such as Figure 4C As shown in table c pointed to by P8, when network device 14 collects P8, it updates record 2 corresponding to network flow 2 based on the information in P8. By comparing table c pointed to by P2 and P8, it is easy to see that network device 14 updates the number of packets in network flow 2 to 2 and updates the arrival time of the last packet in network flow 2 to t8.
[0187] Figure 4C The corresponding example is with Figure 4B For the same parts, please refer to the previous text. Figure 4B The corresponding description, for example, for Figure 4C The description of table b corresponding to P1 can be found in the previous text. Figure 4B The description of the corresponding table b will not be repeated here.
[0188] Since P9 is the last packet collected by network device 14 from the network traffic (filled with black dots), therefore, Figure 4C Table c, pointed to by P9, contains the measurement results of network traffic (filled with black dots) for network device 14.
[0189] By comparison Figure 4B and Figure 4C The network measurement process shown is not difficult to see. Figure 4C The sampling process shown is beneficial for more comprehensive collection of packets in network flows with small statistical values (such as mouse flows), which helps improve the measurement accuracy of such network flows, and thus improves the measurement accuracy of network traffic (black dot filling).
[0190] In one possible implementation, the first device can send the measurement results of the packets to other devices, such as an analysis center. Since the measurement results record the statistical information of the sampled packets, and the first device uses different sampling rates in the first and second sampling processes, optionally, the measurement results sent by the first device may include not only the statistical information of the network flow, but also the sampling rate corresponding to the network flow, so that, for example, the analysis center can determine the statistical information of all packets received by the first device based on the sampling rate.
[0191] For example, the first device can be in accordance with Figure 4D The measurement results are generated in the format shown, and then sent to other devices, such as an analysis center.
[0192] refer to Figure 4D In the table shown, rows 1 through 6 are referred to as a set, which corresponds to the template representation area. In row 1, "set ID = 1" indicates the set identifier (ID) is 1, and "set length = 24 Bytes" indicates the set length is 24 bytes. In row 2, "template ID = 256" indicates the template ID is 256. In row 2, "number of fields = 4" indicates the number of defined information units (numberof fields) is 4. In row 3, "destinationIPv4Address" indicates the first information unit (field 1) is of type destination IP address, and "field length = 4 Bytes" indicates the size of the first information unit (fieldlength) is 4 bytes. In row 4, "destinationPort" indicates the second information unit (field 2) is of type destination port, and "field length = 2 Bytes" indicates the size of the second information unit (field length) is 2 bytes. In line 5, "Sampling Interval" indicates that the type of the third information unit (field 3) is sampling rate, and "field length = 2 Bytes" indicates that the size (field length) of the third information unit is 2 bytes. In line 6, "packetcount" indicates that the type of the fourth information unit (field 4) is packet count, and "fieldlength = 4 Bytes" indicates that the size (field length) of the fourth information unit is 4 bytes.
[0193] Continue to refer to Figure 4DIn the table shown, rows 7 to 13 are referred to as a set, which corresponds to the data area. In row 7, "set ID = 2" indicates that the set identifier is 2, and "set length = 28Bytes" indicates that the set length is 28 bytes. Rows 8 to 10 represent the destination IP address, destination port, sampling rate, and number of packets for record 1, respectively, while rows 11 to 13 represent the destination IP address, destination port, sampling rate, and number of packets for record 2, respectively.
[0194] Continuing with the aforementioned assumptions about the network flow corresponding to the first matching item and the network flow corresponding to the second matching item, the differentiated recording process provided by the embodiments of this application will be illustrated below through Examples 2 and 3.
[0195] Example 2: The memory access performance of the storage components used by the first and second recording procedures is different.
[0196] Suppose a first device includes multiple storage components with different memory access performances, and the memory access performance of a second storage component among these components is superior to that of a first storage resource among the multiple storage components. In one possible implementation, the first device uses the first storage component to perform a first recording process and the second storage component to perform a second recording process. This helps to conserve the first device's valuable storage resources while ensuring the efficiency of network flow measurement, thereby improving the forwarding performance of the first device.
[0197] Referring to the preceding description of Table 2, for example, the first storage component can be an off-chip storage component, and the second storage component can be an on-chip storage component. Combined with... Figure 4B or Figure 4C In a corresponding example, in one possible implementation, network device 14 may store record 1 in a second storage component (e.g., on-chip storage component) and network device 14 may store record 2 in a first storage component (e.g., off-chip storage component).
[0198] Example 3: The first and second recording procedures utilize different computing and network resources.
[0199] In one possible implementation, when the first device performs a first recording procedure on a message that matches a first match (i.e., the first message), it sends an instruction message to the second device, which instructs the second device to record statistical information about the network flow to which the first message belongs. The computational resources utilized by the first recording procedure include, for example, the sending operation, and the network resources utilized by the first recording procedure include, for example, the network bandwidth between the first and second devices.
[0200] In one possible implementation, when the first device performs a second recording procedure on a message that matches a second match (i.e., the second message), it locally records statistical information about the network flow to which the second message belongs. The computational resources utilized by the second recording procedure include, for example, statistical operations on the message. Figure 4B Taking the recording process shown as an example, the statistical operation involves reading the contents of table b (e.g., the number of messages), calculating the updated number of messages, and writing the calculation result into table b, etc.
[0201] By comparing the first and second recording processes described above, it can be seen that the first recording process primarily utilizes computing resources corresponding to the sending operation, while the second recording process primarily utilizes computing resources corresponding to statistical operations (as described above). Therefore, the first recording process generally utilizes fewer computing resources than the second recording process. Furthermore, the first recording process utilizes network resources, including those used for sending instruction information, while the second recording process can be executed locally and thus does not consume network resources. Therefore, the first recording process generally utilizes more network resources than the second recording process. It is evident that, compared to the first device recording the statistical information of the network flow to which the first packet belongs locally, the first device's instruction to the second device to record the statistical information of the network flow to which the first packet belongs saves computing resources. Moreover, compared to the first device instructing the second device to record the statistical information of the network flow to which the second packet belongs, since the statistical value of the network flow to which the first packet belongs is smaller than that of the network flow to which the second packet belongs, the first device's instruction to the second device to record the statistical information of the network flow to which the first packet belongs saves network resources. In summary, the first device executes the first recording process and the second recording process on the first message and the second message respectively, which helps to save computing resources while consuming less network resources.
[0202] Optionally, the indication information is message digest information, which can indicate the network flow to which the first message belongs and the time when the first message was received, etc. The first device instructs the second device to record statistical information of the network flow to which the first message belongs through the message digest information.
[0203] In one possible implementation, the first device and the second device are respectively a forwarding plane device and a control plane device. For example, the first device and the second device are respectively... Figure 2 The forwarding plane device 142 and control plane device 141 are shown. In this embodiment, by transferring part of the recording process of the forwarding plane device to the control plane device, the forwarding performance of the forwarding plane device is better ensured.
[0204] In one possible implementation, after receiving the first packet, the forwarding plane device 142 can send the packet header of the first packet to the control plane device 141. After receiving the packet header, the control plane device 141 records statistical information about the network flow to which the first packet belongs. This application embodiment does not limit the method by which the control plane device 141 records the statistical information about the network flow to which the first packet belongs. For example, optionally, after receiving the control information, the control plane device 141 can, based on the control information,... Figure 4B The statistical information of the network flow to which the packet belongs is created or updated in Table b shown. Alternatively, for example, after receiving the control information, the control plane device 141 may not record the statistical information of the network flow to which the packet belongs, but instead accumulate the headers of multiple packets and then calculate the statistical information of the network flow to which each packet belongs based on the headers of the multiple packets.
[0205] The above have introduced the possible differences between the first and second measurement strategies from different perspectives (angle 1 and angle 2) and different examples (example 1, example 2 and example 3). The following are a few additional points that are not limited.
[0206] This application does not limit the type or number of differences between the first measurement strategy and the second measurement strategy. In one possible implementation, the difference between the first measurement strategy and the second measurement strategy may include only one type of difference. For example, there may be a difference between the first measurement strategy and the second measurement strategy as described in angle 1, or a difference as described in angle 2, or a difference as described in example 1, or a difference as described in example 2, or a difference as described in example 3. Alternatively, in one possible implementation, the difference between the first measurement strategy and the second measurement strategy may include at least two types of differences. For example, there may be a difference between the first measurement strategy and the second measurement strategy as described in angles 1 and 2, or a difference as described in any two of examples 1 to 3.
[0207] This application does not limit the number of matches or measurement strategies indicated by the control information. In one possible implementation, the control information may also include more matches; for example, the control information may further indicate the execution of a first measurement strategy or a second measurement strategy for messages matching a third match. In another possible implementation, the control information may not only indicate more matches but also more measurement strategies. For example, the control information may further indicate the execution of a third measurement strategy for messages matching a fourth match. The difference between the third measurement strategy and the first measurement strategy can be understood with reference to the difference between the first and second measurement strategies described above, and similarly, the difference between the third measurement strategy and the second measurement strategy can be understood with reference to the difference between the first and second measurement strategies described above, and will not be repeated here.
[0208] This application does not limit the measurement of a message to being completed by executing a first measurement strategy or a second measurement strategy. In one possible implementation, during the execution of the first or second measurement strategy, the first device can sequentially perform the sampling and recording processes described above on the message. Alternatively, in one possible implementation, the first device performs the same sampling strategy (e.g., corresponding to codes in the same storage address) on the received first and second messages, then performs the first measurement strategy on the first message, such as performing the first recording process described above, and performs the second measurement strategy on the second message, such as performing the second recording process described above. Alternatively, in one possible implementation, the first device performs the first measurement strategy on the received first message, such as performing the first sampling process described above, and performs the second measurement strategy on the received second message, such as performing the second sampling process described above. The first device also performs the same recording strategy (e.g., corresponding to codes in the same storage address) on both the first and second messages.
[0209] In this embodiment, the first device can execute a first measurement strategy on a received first message and a second measurement strategy on a received second message based on control information. The differences between the first and second measurement strategies have been described above; the control information and its acquisition method are described below.
[0210] In one possible implementation, the matching terms in the control information can indicate the relative magnitude of statistical values. In another possible implementation, the statistical value of the network flow to which the packet matching the first matching term belongs is less than the statistical value of the network flow to which the packet matching the second matching term belongs. In yet another possible implementation, the control information is determined based on the statistical values of multiple network flows.
[0211] The following describes several possible ways for the first device to obtain this control information.
[0212] Control information acquisition method 1: The first device generates control information.
[0213] Taking the first device as Figure 2 Taking network device 14 as an example, after generating control information, control plane device 141 sends the control information to forwarding plane device 142. Based on this control information, forwarding plane device 142 can perform, for example... Figure 3 The corresponding method.
[0214] In one possible implementation, the first device generates the control information based on user configuration. Alternatively, in another possible implementation, the first device generates the control information based on instruction information.
[0215] The following is a description of this instruction.
[0216] The indication information includes multiple matching items and the category corresponding to each matching item in multiple categories. Each matching item in the multiple matching items indicates the flow identification rule of the packet. The indication information indicates that the statistical value of the network flow to which the first packet belongs is greater than the statistical value of the network flow to which the second packet belongs. The first packet matches the first matching item in the multiple matching items, and the second packet matches the second matching item in the multiple matching items. The first matching item and the second matching item correspond to the first category and the second category in the multiple categories, respectively.
[0217] In one possible implementation, the indication information can be stored in, for example, the form of Table 4, where "0" represents the first category and "1" represents the second category. Based on this indication information, if a packet (referred to as packet 1) matches the first match and other packets (referred to as packet 2) match the second match, then it can be determined that the statistical value of the network flow to which packet 1 belongs is less than the statistical value of the network flow to which packet 2 belongs.
[0218] Table 4
[0219] First match 0 Second match 1
[0220] In one possible implementation, the multiple categories in the indication information include multiple categories of a first type and multiple categories of a second type. Optionally, the first type and the second type correspond to statistical values of different types. For example, the first type corresponds to the value of the number of packets within a flow, and the multiple categories of the first type are used to indicate the magnitude relationship between the values of the number of packets within a flow. For example, the second type corresponds to the value of the flow rate, and the multiple categories of the second type are used to indicate the magnitude relationship between the values of the flow rate. This indication information not only indicates the correspondence between the matching item and the category of the first type, but also indicates the relationship between the matching item and the category of the second type. The first category and the second category mentioned in the embodiments of this application can refer to different categories within the first type or different categories within the second type.
[0221] In one possible implementation, the indication information can be stored, for example, in the form of Table 5. In the indication information shown in Table 5, the column corresponding to category a represents the size relationship between the number of packets within a flow; specifically, "0" represents a smaller number of packets within a flow than "1". Furthermore, in the indication information shown in Table 5, the column corresponding to category b represents the size relationship between flow rates; specifically, "0" represents a smaller flow rate than "1". If a packet matches the first match, it can be considered that the flow size and flow rate of the network flow to which the packet belongs are both small; for example, the network flow to which the packet belongs can be called a short-connection, low-speed flow. If a packet matches the second match, it can be considered that the flow size of the network flow to which the packet belongs is large, but the flow rate is small; for example, the network flow to which the packet belongs can be called a long-connection, low-speed flow. If a packet matches the third match, it can be considered that the flow size of the network flow to which the packet belongs is small, but the flow rate is large; for example, the network flow to which the packet belongs can be called a short-connection, high-speed flow. If a message matches the fourth matching item, it can be considered that the network flow to which the message belongs has a large flow size and flow rate. For example, the network flow to which the message belongs can be called a long-connection high-speed flow.
[0222] Table 5
[0223] First match 0 0 Second match 1 0 Third match 0 1 Fourth matching item 1 1
[0224] In this indication, the statistical value corresponding to the first category of network flow is less than the statistical value corresponding to the second category of network flow. In one possible implementation, the statistical value corresponding to the first category of network flow is less than a threshold, and the statistical value corresponding to the second category of network flow is greater than the threshold. In another possible implementation, this threshold is used to distinguish between "elephant flows" and "mouse flows" in network traffic. Based on this indication, it can be determined that packets matching the first match belong to "mouse flows," and packets matching the second match belong to "elephant flows."
[0225] The following describes several possible ways for the first device to obtain this instruction information.
[0226] Method 1 for obtaining instruction information: The first device receives instruction information from the third device.
[0227] In this embodiment of the application, optionally, the network environment may further include, for example, the network environment of this embodiment of the application. Figure 1 The third device 15 is shown. Figure 1The third device 15 in this embodiment is illustrated using a server as an example; this illustration is for illustrative purposes only and not as a limitation. The third device can be a physical entity in the network, such as a router, switch, or firewall; it can also be a virtual network device running within a physical entity, such as a virtual router or virtual switch; or it can be a mainframe computer, a server with computing capabilities, or other computing devices. This application does not limit the specific form of the third device, as long as it can perform the methods provided in this application. For example, the third device 15 could also be, for example... Figure 1 or Figure 2 The network device 14 shown can be, for example, Figure 2 The control surface device 141 shown.
[0228] In one possible implementation, the indication information is generated by a third device. In another possible implementation, the third device generates the indication information based on user configuration. Alternatively, in a possible implementation, the third device can generate the indication information based on statistical values of multiple network flows. The following is a combination of... Figure 5 This paper introduces one possible method for a third device to generate this instruction information.
[0229] refer to Figure 5 This application also provides a method for processing network data. In one possible implementation, the method includes steps 501 to 503.
[0230] Step 501. The third device identifies multiple samples;
[0231] The third device can identify multiple samples, each corresponding to a network flow, and each sample includes the flow identifier and statistics of the corresponding network flow.
[0232] In one possible implementation, the third device determines multiple samples based on information from the first network traffic, which can be used to determine the flow identifier and statistics of the network flows in the first network traffic.
[0233] In one possible implementation, the statistical value of the network flow can be determined based on the flow size and / or flow rate. For an understanding of the statistical value of the network flow, please refer to the relevant content in the aforementioned assumptions made about the network flow corresponding to the first match and the network flow corresponding to the second match; it will not be repeated here. In one possible implementation, the statistical value s of the network flow is positively correlated with the flow size n, for example, s = k1 * n, where k1 is a fixed value. Optionally, the flow size can refer to the number of packets within the flow.
[0234] In one possible implementation, the information of the first network traffic includes information about each packet within the first network traffic. Optionally, the packet information may include the packet itself. Alternatively, the packet information may include partial information from the packet, such as the packet header or the flow identifier within the packet.
[0235] Alternatively, in one possible implementation, the information of the first network traffic includes information about each network flow within the first network traffic. Optionally, the information of a network flow includes the flow identifier and statistical information of that network flow. The statistical information of the network flow is used to determine the statistical value of that network flow. For example, if the statistical information of a network flow is the value n of the number of packets within that flow, the third device can determine the statistical value s of the network flow as k1*n based on the value n of the number of packets within that flow.
[0236] Based on the information from the first network traffic, the third device can identify multiple samples, each corresponding to a network flow within the first network traffic. Each sample includes a flow identifier and a statistical value for the corresponding network flow. For example, assuming the flow identifier is a 5-tuple and the statistical value is the number of packets within the flow, the multiple samples include: sample {(a1, b1, c1, d1, e1), s1}, sample {(a1, b1, c1, d1, e2), s2}, sample {(a2, b1, c1, d1, e1), s3}, sample {(a2, b1, c1, d1, e2), s4}, and sample {(a2, b1, c1, d1, e3), s5}. In this embodiment, for ease of description, 'a' represents the information type corresponding to the source IP address, and 'a1' and 'a2' represent different information units describing the source IP address, such as 10.171.111.124 and 10.136.82.184, respectively. For ease of description, in this embodiment, 'b' represents the information type corresponding to the source port, and 'b1' and 'b2' represent different information units describing the source port, such as 445 and 8080, respectively. For ease of description, in this embodiment, 'c' represents the information type corresponding to the destination IP address, and 'c1' and 'c2' represent different information units describing the destination IP address, such as 10.136.82.178 and 10.247.68.243, respectively. For ease of description, in this embodiment, 'd' represents the information type corresponding to the transport layer protocol, and 'd1' and 'd2' represent different information units describing the transport layer protocol, respectively. For ease of description, in the embodiments of this application, e represents the information type corresponding to the destination port, and e1 and e2 represent different information units describing the destination port, such as 9610 and 58645 respectively.
[0237] Step 502. The third device clusters multiple samples to obtain multiple clusters;
[0238] After identifying multiple samples, the third device can cluster these samples to obtain multiple clusters. Continuing with the example of the multiple samples listed in step 501, these multiple clusters can be as follows: Figure 6A As shown. Alternatively, the multiple clusters can be as follows: Figure 6B As shown. Alternatively, the multiple clusters can be as follows: Figure 6C As shown. Alternatively, the multiple clusters can be as follows: Figure 6D As shown, or, the multiple clusters can be as follows: Figure 6E As shown. Figures 6A to 6E In any of the graphs, samples within the same circle belong to the same cluster.
[0239] refer to Figures 6A to 6E In any of the accompanying figures, each of the plurality of clusters comprises one or more samples, and all samples within each cluster share the same information unit in their flow identifiers. For ease of description, for any one of the plurality of clusters (referred to as the target cluster), the same information unit in the flow identifiers of all samples within the target cluster is referred to as the cluster label of the target cluster. It should be noted that there may be multiple identical information units in the flow identifiers of all samples within the target cluster; the cluster label of the target cluster refers to all identical information units among the flow identifiers of all samples within the target cluster. Figures 6A to 6E In the diagram, each circle is identified as the cluster label of the cluster it represents.
[0240] refer to Figures 6A to 6E In any of the attached figures, in one possible implementation, any two different clusters among the multiple clusters will have different cluster labels.
[0241] refer to Figures 6A to 6E In any of the attached figures, in one possible implementation, the number of information units corresponding to the clustering label of the target cluster is less than the number of information units corresponding to the flow identifier. Figures 6A to 6E Taking the multiple clusters shown as an example, the flow identifier corresponds to 5 information units, and the information units corresponding to the clustering label of each cluster are all less than 5.
[0242] refer to Figures 6A to 6E In any of the accompanying figures, in one possible implementation, there exist two clusters among the plurality of clusters whose cluster labels correspond to the same number of information units as the described information type. For example, see reference... Figure 6A The multiple clusters shown, clusters (a1, b1, c1, d1) and (a2, b1, c1, d1) each have 4 information units corresponding to their cluster labels. The information types described by the 4 information units corresponding to the cluster labels of clusters (a1, b1, c1, d1) and (a2, b1, c1, d1) are source IP address, source port, destination IP address and transport layer protocol information units, respectively.
[0243] refer to Figures 6A to 6E In any of the accompanying figures, in one possible implementation, the clustering label of the target cluster includes at least an information unit for describing the transport layer protocol.
[0244] refer to Figure 6D or Figure 6E In one possible implementation, there exist two clusters among these clusters whose cluster labels correspond to the same number of information units, but the information types described by at least one information unit corresponding to the cluster label are different. For example, refer to... Figure 6D The clusters shown, clusters (a1, b1, c1, d1) and (b1, c1, d1, e1) each have 4 information units corresponding to their cluster labels. The information types described by the 4 information units corresponding to the cluster labels of cluster (a1, b1, c1, d1) are source IP address, source port, destination IP address and transport layer protocol information units, respectively. The information types described by the 4 information units corresponding to the cluster labels of cluster (b1, c1, d1, e1) are source port, destination IP address, destination port and transport layer protocol information units, respectively.
[0245] refer to Figure 6E In one possible implementation, there exist two clusters among these multiple clusters whose cluster labels correspond to different numbers of information units. For example, see reference... Figure 6C or Figure 6E The number of information units corresponding to the cluster labels of the multiple clusters shown are 4 and 3 for cluster (a1, b1, c1, d1) and cluster (b1, c1, d1), respectively.
[0246] The multiple clusters obtained in step 502 are used to generate indication information. The possible implementations provided above are beneficial for enriching the types of clustering labels for multiple clusters, thereby facilitating the generation of more clusters based on these multiple samples, and further providing more basis for the indication information, thus improving the accuracy of the indication information.
[0247] Figures 6A to 6E Several possible clustering results are illustrated. The multiple clusters obtained by clustering the multiple samples listed in step 501 are not limited to the clustering results described above, and can also be other clustering results. For example, there is a cluster among these multiple clusters whose cluster label corresponds to 2 information units.
[0248] The above describes the process of clustering multiple clusters. The following example illustrates the process of clustering multiple samples.
[0249] In one possible implementation, after obtaining multiple samples, the third device can select at least two information types from the information types described by the stream identifiers of the samples as clustering parameters. Then, the third device can cluster the multiple samples according to the selected clustering parameters.
[0250] Assuming the stream identifiers of multiple samples describe information types including, for example... Figure 7A Given boxes A1 containing a, b, c, d, and e, if four information types are selected as clustering parameters from the information types described by the flow identifier, and assuming that the transport layer protocol is a mandatory information type, then the selected clustering parameters can be: Figure 7A The information types listed in any one of the boxes A2 to A5.
[0251] Assuming the stream identifiers of multiple samples describe information types including, for example... Figure 7B For boxes B1 showing a, b, c, d, and e, if three information types are selected as clustering parameters from the information types described by the flow identifier, and assuming that the transport layer protocol is a mandatory information type, then the selected clustering parameters can be: Figure 7B The information types listed in any one of the boxes B2 to B7.
[0252] Assuming the stream identifiers of multiple samples describe information types including, for example... Figure 7C For boxes C1 showing a, b, c, d, and e, if two information types described by the flow identifier are selected as clustering parameters, and assuming that the transport layer protocol is a mandatory information type, then the selected clustering parameters can be: Figure 7C The information types listed in any one of the boxes C2 to C5.
[0253] The above describes the clustering parameters that may be used in the clustering process. The following describes a clustering method for multiple samples based on the clustering parameters.
[0254] In one possible implementation, the third device can cluster multiple samples based on a clustering parameter. For example, the third device can select... Figure 7A The information types listed in box A2 shown are used as clustering parameters. Based on these clustering parameters, multiple samples exemplified in step 501 are clustered, and the resulting clusters can be, for example... Figure 6A As shown. For example, a third device can be selected. Figure 7A The information types listed in box A5 are used as clustering parameters. Based on these parameters, multiple samples from step 501 are clustered, resulting in multiple clusters, for example... Figure 6B As shown.
[0255] Alternatively, in one possible implementation, the third device can cluster multiple samples according to multiple clustering parameters, and the multiple clusters obtained in step 502 include clusters obtained by clustering multiple samples according to each clustering parameter. In one possible implementation, refer to Figure 5 Step 502 includes steps 5021 and 5022.
[0256] Step 5021. The third device clusters multiple samples according to the first clustering parameters to obtain the first cluster set;
[0257] Step 5022. The third device clusters multiple samples according to the second clustering parameters to obtain a second cluster set;
[0258] The first cluster set includes one or at least two clusters, the second cluster set includes one or at least two clusters, and the multiple clusters obtained in step 502 include the first cluster set and the second cluster set. The information types corresponding to the first clustering parameter and the second clustering parameter are different.
[0259] In one possible implementation, the number of information types corresponding to the first clustering parameter and the second clustering parameter is the same, and at least one information type corresponding to the first clustering parameter does not appear in the information type corresponding to the second clustering parameter. Continuing with the example of the third device clustering multiple samples in step 501, assuming the third device selects respectively... Figure 7A The information types listed in boxes A2 and A5 are used as the first and second clustering parameters. The third device performs step 5021 on multiple samples to obtain, for example... Figure 6A The first cluster set shown, the third device performs step 5022 on multiple samples to obtain, for example... Figure 6B The second cluster set is shown. After obtaining the first and second cluster sets, the third device can obtain, for example... Figure 6D The multiple clusters shown.
[0260] In one possible implementation, the number of information types corresponding to the first clustering parameter and the second clustering parameter are different, and the number of information types corresponding to the first clustering parameter is greater than the number of information types corresponding to the second clustering parameter. Continuing with the example of the third device clustering multiple samples in step 501, assuming the third device selects... Figure 7A Box A2 and shown Figure 7B The information types listed in box B4 are used as the first and second clustering parameters. The third device performs steps 5021 and 5022 on multiple samples to obtain, for example... Figure 6C The multiple clusters shown.
[0261] If clustering multiple samples results in two clusters, where one cluster (called the large cluster) contains more samples than the other cluster (called the small cluster), and the number of information units corresponding to the cluster labels of the large cluster and the small cluster are equal, and both the large and small clusters include a subset of samples (called intermediate samples), then in one possible implementation, these intermediate samples can be removed from the small cluster. This could potentially remove all samples from the small cluster, preventing it from being included in the multiple clusters and thus reducing the number of clusters and consequently the number of matching terms in the indication information.
[0262] refer to Figure 5 In one possible implementation, step 502 further includes step 5023.
[0263] Step 5023. The third device clusters multiple samples according to the third clustering parameters to obtain a third cluster set;
[0264] The number of information types corresponding to the first clustering parameter and the second clustering parameter is the same; at least one information type corresponding to the first clustering parameter does not appear in the information type corresponding to the second clustering parameter; the number of information types corresponding to the first clustering parameter and the third clustering parameter is different; and the number of information types corresponding to the first clustering parameter is greater than the number of information types corresponding to the third clustering parameter. Continuing with the example of the third device clustering multiple samples in step 501, assuming the third device selects... Figure 7A Boxes A2 and A5 shown Figure 7B The information types listed in box B4 are used as the first clustering parameter, the second clustering parameter, and the third clustering parameter. The third device performs steps 5021, 5022, and 5023 on multiple samples to obtain, for example... Figure 6E The multiple clusters shown.
[0265] The embodiments of this application do not limit the order of any two steps in steps 5021 to 5023.
[0266] The multiple clusters obtained in step 502 are used to generate indication information. This embodiment of the application uses multiple clustering parameters to cluster multiple samples, which helps to obtain more clusters, thereby providing more evidence for the indication information and improving its accuracy. Furthermore, since the clustering parameters are user-selectable, the clustering process provided in this embodiment is visualized, facilitating understanding and control of the clustering process, and thus enabling optimization of the accuracy of the obtained indication information.
[0267] Through step 502, the third device obtains multiple clusters. The following section, in conjunction with step 503, describes how the third device generates indication information based on these multiple clusters.
[0268] Step 503. The third device generates indication information based on the samples in each of the multiple clusters.
[0269] This instruction can be understood by referring to the previous introduction on instruction information, and will not be repeated here.
[0270] In one possible implementation, the third device determines the matching items in the indication information based on samples within each of the plurality of clusters. Optionally, each matching item in the indication information corresponds to one of the plurality of clusters, and each matching item is determined based on the clustering label of the corresponding cluster. For any matching item in the indication information (referred to as a target matching item), its corresponding cluster is called the target cluster, and the target matching item is determined based on the clustering label of the target cluster. Optionally, the target matching item is the clustering label of the target cluster, or a hash value of the clustering label of the target cluster. In one possible implementation, different matching items in the indication information correspond to different clusters in the plurality of clusters.
[0271] In one possible implementation, the third device determines the category of each matching item in the indication information across multiple categories based on samples within each of the multiple clusters. Optionally, the category corresponding to each matching item is determined based on statistical values from samples within the corresponding cluster.
[0272] Since the category corresponding to each matching item in the indication information is determined based on the statistical values in the samples within the corresponding cluster, the method for determining the category corresponding to the matching item is described below.
[0273] In one possible implementation, the statistical value of the samples within the cluster corresponding to the first category is less than a threshold, and the statistical value of the samples within the cluster corresponding to the second category is greater than a threshold.
[0274] For ease of description, for any one of multiple clusters (referred to as the target cluster), samples in the target cluster whose statistical values are less than the threshold are called small flow samples. Furthermore, for ease of description, the ratio of the number of small flow samples in the target cluster to the total number of samples in the target cluster is called the small flow number ratio of the target cluster, and the ratio between the sum of the statistical values of the small flow samples in the target cluster and the sum of the statistical values of all samples in the target cluster is called the small flow statistical value ratio of the target cluster. In one possible implementation, this application embodiment does not limit the statistical values of all samples in the cluster corresponding to the first category to be less than the threshold. For example, optionally, the small flow number ratio of the cluster corresponding to the first category is greater than a first ratio, and the small flow statistical value ratio of the cluster corresponding to the first category is greater than a second ratio. Taking the statistical value as the number of packets in the flow as an example, optionally, both the first ratio and the second ratio are 90%.
[0275] For ease of description, for any one of multiple clusters (referred to as the target cluster), samples in the target cluster whose statistical values are greater than the threshold are called high-flow samples. Furthermore, for ease of description, the ratio of the number of high-flow samples in the target cluster to the total number of samples in the target cluster is called the high-flow number ratio of the target cluster, and the ratio between the sum of the statistical values of the high-flow samples in the target cluster and the sum of the statistical values of all samples in the target cluster is called the high-flow statistical value ratio of the target cluster. In one possible implementation, this application embodiment does not limit the statistical values of all samples in the cluster corresponding to the second category to be greater than the threshold. For example, optionally, the high-flow number ratio of the cluster corresponding to the second category is greater than the third ratio, and the high-flow statistical value ratio of the cluster corresponding to the second category is greater than the fourth ratio. Taking the statistical value as the number of packets in the flow as an example, optionally, the third ratio and the fourth ratio are 10% and 90%, respectively.
[0276] The conditions that the clusters corresponding to the first category and the clusters corresponding to the second category must meet have been described above. In one possible implementation, this application embodiment does not limit each of the multiple clusters to correspond to a matching item in the indication information. For example, if a certain cluster among the multiple clusters does not meet the conditions described above, then the third device may not generate a matching item in the indication information based on that cluster.
[0277] Alternatively, in one possible implementation, for a cluster among multiple clusters, if the cluster satisfies the conditions corresponding to the first category described above, but the total number of samples within the cluster does not exceed a first number, then the third device may not generate a matching item in the indication information based on that cluster. Alternatively, in one possible implementation, for a cluster among multiple clusters, if the cluster satisfies the conditions corresponding to the second category described above, but the total number of samples within the cluster does not exceed a second number, then the third device may not generate a matching item in the indication information based on that cluster. Taking the statistical value as the number of packets within a flow as an example, optionally, the first number is greater than the second number; for example, the first number is 20, and the second number is 1.
[0278] As mentioned above, in one possible implementation, the statistical value of the samples within the cluster corresponding to the first category is less than a threshold, while the statistical value of the samples within the cluster corresponding to the second category is greater than the threshold. In one possible implementation, this threshold is manually configured by the user based on experience. Alternatively, in another possible implementation, the threshold is determined based on information from the second network traffic.
[0279] Based on the fact that the threshold is determined according to information from the second network traffic, in one possible implementation, the threshold is determined by a fourth device, and the third device can obtain the threshold from the fourth device. The fourth device is a computer device, such as a hardware-based device or a software-based device (e.g., a cloud server). Alternatively, in one possible implementation, the threshold is determined by the third device. In one possible implementation, the second network traffic is the aforementioned first network traffic, or in one possible implementation, the information of the second network traffic was received by the third device prior to the information of the first network traffic. In one possible implementation, some or all of the network flow in the second network traffic is received by the first device. For example, after receiving a message, the first device can send the message header to the third device, or send the measurement results of the message to the third device.
[0280] The following describes one possible method for determining this threshold. Optionally, the threshold is determined based on a threshold function. This threshold function takes a threshold variable as its independent variable, and the threshold is a value within the domain of that threshold variable. First, we introduce one possible form of this threshold function.
[0281] The threshold function represents the difference between the small flow number ratio function and the small flow statistical value ratio function. Both the value of the small flow number ratio function and the small flow size ratio function are related to the value of the threshold variable; therefore, the value of the threshold function is related to the value of the threshold variable. Optionally, network flows in the second network traffic that are smaller than the value of the threshold variable are called small flows. The ratio of the number of small flows in the second network traffic to the total number of network flows in the second network traffic is defined as the small flow number ratio function. The ratio between the sum of the statistical values of small flows in the second network traffic and the sum of the statistical values of all network flows in the second network traffic is defined as the small flow statistical value ratio function. In one possible implementation, among multiple candidate values of the threshold variable, the threshold function is maximized when the threshold variable takes the specified threshold value.
[0282] In one possible implementation, the threshold function, the stream number ratio function, and the stream statistic ratio function are shown in Equations 1), 2), and 3), respectively.
[0283] Formula 1): F(T) = F f (T)-F p (T),
[0284] Formula 2): F f (T)=N f (s <T) / N f (total),
[0285] Formula 3): F p (T)=Np (s<T) / N p (total),
[0286] Formula 4): T0=argmax T (F f (T)-F p (T)).
[0287] Wherein, F(T) represents the threshold function, F f (T) represents the proportion function of the number of mice flows, F p (T) represents the proportion function of statistical values of mice flows, T represents the threshold variable, N f (s<T) represents the number of network flows with statistical values less than T in the second network traffic, N f (total) represents the total number of network flows in the second network traffic, N p (s<T) represents the sum of in-flow statistical values of network flows with statistical values less than T in the second network traffic, N p (total) represents the sum of in-flow statistical values of all network flows in the second network traffic, T0 represents the threshold involved in the embodiment of the present application, argmax T (F f (T)-F p (T)) is the value of the threshold variable T that enables F f (T)-F p (T) to obtain the maximum value. For ease of understanding, Figure 8 the function graphs of Ff(T) and Fp(T) are exemplarily shown by curve 1 and curve 2 respectively.
[0288] In real network traffic, there exists the "Pareto principle". The number of elephant flows in the network accounts for a very low proportion of the number of all network flows in the network traffic, but the sum of the number of in-flow packets or the number of in-flow bytes of elephant flows accounts for a very high proportion of the sum of the number of in-flow packets or the number of in-flow bytes of all flows in the network traffic. On the contrary, the number of mice flows in the network accounts for a very high proportion of the number of all network flows in the network traffic, but the sum of the number of in-flow packets or the number of in-flow bytes of mice flows accounts for a very low proportion of the sum of the number of in-flow packets or the number of in-flow bytes of all flows in the network traffic. If the value of the number of in-flow packets or the number of in-flow bytes is taken as the statistical value, using the threshold obtained by the above threshold determination method (e.g., formulas 1) to 4)) to determine the category corresponding to the matching item will facilitate accurate identification of elephant flows and mice flows in network traffic.
[0289] The above example illustrates how a third device determines each match and its corresponding category in the indication information based on samples within each of multiple clusters. In one possible implementation, the categories determined in step 503 may include multiple types of categories. Assuming that after step 502, the third device obtains... Figure 6E Based on the multiple clusters shown, after step 503, the third device obtains indication information, such as that shown in Table 6. As shown in Table 6, each match corresponds to a category in category a and a category in category b. Assuming... Figure 6E The clusters (b1, c1, d1, e2) shown do not meet the conditions required for the clusters corresponding to the first or second categories (see related content above). Therefore, Table 6 does not show the matching items corresponding to the clusters (b1, c1, d1, e2).
[0290] Table 6
[0291]
[0292] Referring back to the description of Table 5, if a packet matches the matching item (a2, b1, c1, d1), then the packet is a short-connection, low-speed flow. If a packet matches the matching item (a1, b1, c1, d1), then the packet is a long-connection, low-speed flow. If a packet matches the matching item (b1, c1, d1, e1), then the packet is a short-connection, high-speed flow. If a packet matches the matching item (b1, c1, d1), then the packet is a long-connection, high-speed flow.
[0293] In one possible implementation, the indication information also indicates the order of the matching items. This order instructs the message to be compared with the matching items in the indication information sequentially to determine whether the message meets the flow identification rules indicated by the currently matched matching item. When a message meets the flow identification rules indicated by the currently matched matching item, it can be determined that the message matches that matching item, and the message is no longer compared with other subsequent matching items. In other words, any two matching items in the indication information have a definite order. For any matching item (called the target matching item), for any message (called the target message) that matches the target matching item, the target matching item corresponding to the target message is the first matching item that the target message matches from the multiple matching items according to the order. Optionally, the target message meets the target matching item, but the target message does not meet any of the matching items that precede the target matching item.
[0294] In one possible implementation, the indication information indicates the order of the matching items in the table. For example, referring to Table 6, the four matching items are, in order from first to last, matching item (a2, b1, c1, d1), matching item (a1, b1, c1, d1), matching item (b1, c1, d1, e1), and matching item (b1, c1, d1). Assuming that a message simultaneously satisfies the flow identification rules indicated by matching item (a1, b1, c1, d1) and matching item (b1, c1, d1, e1), according to the order indicated in Table 6, it can be determined that the message matches matching item (a1, b1, c1, d1).
[0295] The above describes how the matching items in the indication information can exist in a definite order. In one possible implementation, the more samples within the cluster corresponding to a matching item, the more advantageous it is to be compared with the message first. Alternatively, suppose there are two clusters among the multiple clusters obtained in step 502, where the number of information units corresponding to the cluster label (called the first cluster label) of one cluster (called the second cluster label) is greater than the number of information units corresponding to the cluster label (called the second cluster label) of the other cluster (called the first cluster label), and the number of samples within the first cluster is greater than the number of samples within the second cluster. In this case, the matching item corresponding to the first cluster takes precedence over the matching item corresponding to the second cluster in the indication information. The last column of Table 6 identifies the number of samples within the cluster corresponding to each matching item. Taking Table 6 as an example, since the number of samples within the cluster corresponding to the matching item (a2, b1, c1, d1) is greater than the number of samples within the cluster corresponding to the matching item (a1, b1, c1, d1), this indication information indicates that the matching item (a2, b1, c1, d1) precedes the matching item (a1, b1, c1, d1). Taking Table 6 as an example, since the number of samples within the cluster corresponding to the matching item (a1, b1, c1, d1) is greater than the number of samples within the cluster corresponding to the matching item (b1, c1, d1, e1), this indication information indicates that the matching item (a1, b1, c1, d1) precedes the matching item (b1, c1, d1, e1). The last column of Table 6 is used to facilitate understanding of one possible basis for determining the order between matching items. This application embodiment does not limit the indication information generated by the third device to necessarily including the number of samples within the cluster corresponding to the matching item.
[0296] In one possible implementation, the more information units a matching item corresponds to, the more advantageous it is for that matching item to be compared with the message first. Alternatively, suppose that among the multiple clusters obtained in step 502, there are two clusters where the number of information units corresponding to the cluster label (called the third cluster label) of one cluster (called the third cluster label) is greater than the number of information units corresponding to the cluster label (called the fourth cluster label) of the other cluster. In this indication information, the matching item corresponding to the third cluster precedes the matching item corresponding to the fourth cluster. Taking Table 6 as an example, since the number of information units corresponding to the matching items (a2, b1, c1, d1), (a1, b1, c1, d1), and (b1, c1, d1, e1) is greater than the number of information units corresponding to the matching item (b1, c1, d1), this indication information indicates that the matching items (a2, b1, c1, d1), (a1, b1, c1, d1), and (b1, c1, d1, e1) all precede the matching item (b1, c1, d1).
[0297] The above process Figure 5 In a corresponding embodiment, the third device can generate instruction information. After generating the instruction information, in one possible implementation, the third device can send a message to the first device based on the instruction information. This message is used to instruct the first device to perform, for example... Figure 3 The corresponding method. Optionally, the message carries the indication information. Taking the first device as an example. Figure 2 Taking the network device 14 shown as an example, after receiving instruction information from the third device, the control plane device 141 generates control information based on the instruction information, and then sends the control information to the forwarding plane device 142. Based on this control information, the forwarding plane device 142 can perform, for example... Figure 3 The corresponding method.
[0298] In one possible implementation, in this embodiment of the application, the third device can execute any of the methods executed by the first device as described above (e.g., before the indication information acquisition method 1). For example, after generating the indication information, the third device can execute... Figure 3 The corresponding steps are 301 and 302.
[0299] Figure 5 In the corresponding embodiment, the third device can generate the instruction information. It should be noted that the embodiments of this application do not limit the third device to instruct the first device to perform actions based on the instruction information after generating it. Figure 3 The corresponding embodiments do not limit the execution of the instruction information generated by the third device. Figure 3 Corresponding implementation examples.
[0300] The above exemplifies a method for a third device to acquire instruction information. In instruction information acquisition method 1, the first device can receive instruction information from the third device. In one possible implementation, after receiving the instruction information, the first device can generate control information based on the instruction information.
[0301] For example, in conjunction with the descriptions of Examples 1 to 3 above, for the matching item corresponding to the first category in the indication information, its corresponding measurement strategy is set to perform the sampling process at a smaller sampling rate, or, its corresponding measurement strategy is set to perform the recording process using a storage component with lower memory access performance, or, its corresponding measurement strategy is set to instruct the second device to perform the recording process, or, its corresponding measurement strategy is set to perform the sampling process at a smaller sampling rate and, and, its corresponding measurement strategy is set to perform the recording process using a storage component with lower memory access performance, or, its corresponding measurement strategy is set to perform the sampling process at a smaller sampling rate and, and, its corresponding measurement strategy is set to instruct the second device to perform the recording process.
[0302] For example, in conjunction with the descriptions of Examples 1 to 3 above, for the matching item corresponding to the second category in the indication information, its corresponding measurement strategy is set to perform the sampling process at a larger sampling rate, or its corresponding measurement strategy is set to perform the recording process using a storage component with high memory access performance, or its corresponding measurement strategy is set to perform the sampling process at a larger sampling rate and, its corresponding measurement strategy is set to perform the recording process using a storage component with high memory access performance.
[0303] Assuming the instruction information received by the first device is as shown in Table 6 above, in one possible implementation, the first device can generate control information as shown in Table 7 based on the instruction information shown in Table 6. In Table 7, N is an integer greater than 1. To facilitate understanding of the differences between measurement strategies corresponding to different matching items, Table 7 uses sampling ratio, "second device," "on-chip," and "off-chip" to describe the characteristics of the measurement strategy. It should be noted that the embodiments of this application do not limit the control information to describe the measurement strategy in the form of Table 7. For example, referring to the relevant description of Table 1 above, in one possible implementation, the control information uses a first storage address and a second storage address to represent the first measurement strategy and the second measurement strategy, respectively, and the first storage address and the second storage address store the code used to execute the first measurement strategy and the second measurement strategy, respectively.
[0304] Table 7
[0305]
[0306] Alternatively, in one possible implementation, the category in the indication information can indicate a measurement strategy; for example, the first category corresponds to a first measurement strategy, and the second category corresponds to a second measurement strategy. In this case, the indication information generated by the third device can be, for example, the control information shown in Table 7. If the indication information can serve as control information to instruct the first device to perform differentiated measurements on the message, then the first device, after obtaining the indication information from the third device, does not need to generate control information based on that indication information.
[0307] Method 2 for obtaining instruction information: The first device generates the instruction information.
[0308] In one possible implementation, the first device generates the indication information based on user configuration. Alternatively, in another possible implementation, the first device can generate the indication information based on statistical values of multiple network flows. Optionally, the method by which the first device generates the indication information can refer to the method for generating indication information by the third device described in indication information acquisition method 1; for example, the first device can execute... Figure 5 The steps performed by the third device in the corresponding embodiment will not be repeated here.
[0309] Taking the first device as Figure 2 Taking network device 14 as an example, after the control plane device 141 generates the instruction information, it generates control information based on the instruction information, and then sends the control information to the forwarding plane device 142. Based on the control information, the forwarding plane device 142 can perform, for example... Figure 3 The corresponding method.
[0310] After the first device generates the indication information, it can generate control information based on the indication information, as described above. Alternatively, in one possible implementation, the category in the indication information indicates the measurement strategy; for example, the first category corresponds to the first measurement strategy, and the second category corresponds to the second measurement strategy. In this case, the indication information can be, for example, the control information shown in Table 3. If the indication information can serve as control information, instructing the first device to perform differential measurements on the message, then the first device does not need to generate control information based on the indication information after generating it.
[0311] The above describes the relevant content of control information acquisition method 1. The following provides another control information acquisition method.
[0312] Control information acquisition method 2: The first device receives control information from the fifth device.
[0313] In one possible implementation, the control information is generated by a fifth device. The method by which the fifth device generates the control information can be referenced from the method described in Control Information Acquisition Method 1 above, where the first device generates the control information; this will not be repeated here.
[0314] In one possible implementation, the fifth device sends a message to the first device carrying the control information, the message being used to instruct the first device to perform, for example... Figure 3 The corresponding method. In one possible implementation, the control information is generated based on the instruction information. The method by which the fifth device obtains the instruction information can be understood by referring to Instruction Information Acquisition Method 1 or Instruction Information Acquisition Method 2 described above, and will not be repeated here. In one possible implementation, the fifth device generates the instruction information, generates the control information based on the instruction information, and then sends the message to the first device. The method by which the fifth device generates the instruction information can be referred to the method by which the third device generates the instruction information in Instruction Information Acquisition Method 1 described above, and will not be repeated here.
[0315] In this embodiment of the application, the fifth device may optionally be a computer device, which may be a hardware-based device or a software-based device (such as a cloud server), as long as the computer device is capable of generating the control information.
[0316] Taking the first device as Figure 2 The forwarding plane device 142 shown is the fifth device. Figure 2 Taking the control plane device 141 as an example, after generating control information, the control plane device 141 sends the control information to the forwarding plane device 142. Based on the control information, the forwarding plane device 142 can perform, for example... Figure 3 The corresponding method.
[0317] The methods of the embodiments of this application have been described above. The apparatus provided by the embodiments of this application will be described below.
[0318] This application provides a processing device. Figure 9 This is a schematic diagram of a possible structure of the processing device 9 according to an embodiment of this application. (See attached diagram.) Figure 9 The processing device 9 includes a processor 901 and a memory 902.
[0319] The processor 901 can be one or more CPUs, which can be a single-core CPU or a multi-core CPU.
[0320] The memory 902 includes, but is not limited to, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), flash memory, or optical memory. The memory 902 stores computer-readable programs or instructions. Optionally, the memory 902 can be non-volatile memory or volatile memory.
[0321] Optionally, the processing device 9 further includes a communication interface 903, which can be a wired interface, such as a Fiber Distributed Data Interface (FDDI) or a Gigabit Ethernet (GE) interface; or it can be a wireless interface. The communication interface 903 is used to receive network data from an internal network and / or an external network, or to send network data to an internal network and / or an external network.
[0322] Optionally, the processing device 9 also includes a bus 904, through which the processor 901 and memory 902 are typically interconnected, or in other ways.
[0323] Processor 901 reads and executes program instructions stored in memory 902 to cause processing device 9 to perform the method executed by the first device or the method executed by the third device in the above method embodiments. For example, processor 901 reads and executes program instructions stored in memory 902 to cause processing device 9 to perform the above method. Figure 3 The steps in the illustrated embodiments, or Figure 5 The steps in the illustrated embodiment are as follows. For further details regarding the processor 901 reading and executing the stored program instructions to cause the processing device 9 to perform the above steps, please refer to the corresponding descriptions in the preceding method embodiments, which will not be repeated here.
[0324] Optionally, these instructions are stored in external memory of the computer device. When these instructions are decoded and executed by the processor 901 of the processing device 9, part or all of the contents of the instructions are temporarily stored in the internal memory 902 of the processing device 9. Optionally, part of the contents of these instructions are stored in external memory of the processing device 9, and other parts of the contents of these instructions are stored in the internal memory 902 of the processing device 9.
[0325] This application also provides a chip system including a processor and an interface circuit. The processor is used to couple with a memory through the interface circuit, and the processor is used to run program code stored in the memory, thereby implementing the method provided in any of the above method embodiments of this application.
[0326] In one example, the processor can execute instructions stored in memory to cause the chip system to perform any of the method embodiments described above. Optionally, the memory can be a storage unit within the chip system, such as a register, cache, etc., or the memory can be a memory located outside the chip system within a computer device, such as read-only memory (ROM) or other types of static storage components capable of storing static information and instructions, random access memory (RAM), etc. Optionally, the memory can be non-volatile memory or volatile memory. Optionally, the processor can be a general-purpose central processing unit (CPU), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits for controlling the execution of programs in any of the method embodiments described above.
[0327] This application provides a processing apparatus. The processing apparatus may be a first device mentioned in the method of this application, a device within the first device, or a device compatible with the first device. In one possible implementation, the processing apparatus may include performing... Figure 3 The module of the method shown can be a hardware circuit, a software module, or a module implemented by combining hardware circuit and software.
[0328] Figure 10 This is a schematic diagram of a possible structure of the processing device 10 according to an embodiment of this application. (See attached diagram.) Figure 10 The processing device 10 includes a communication module 1001 and a measurement module 1002. The communication module 1001 receives a message, which is either a first message or a second message. The first message and the second message are respectively matched against a first matching item and a second matching item in control information. The control information instructs the execution of a first measurement strategy on the first message and also instructs the execution of a second measurement strategy on the second message. The first matching item and the second matching item respectively indicate the flow identification rules of the message. The first measurement strategy or the second measurement strategy is used to perform network flow-based measurements. The measurement module 1002 executes the first measurement strategy or the second measurement strategy on the message according to the control information.
[0329] Specifically, for example, the communication module 1001 is used to perform... Figure 3In step 301 of the illustrated embodiment, the measurement module 1002 is used to perform... Figure 3 Step 302 in the illustrated embodiment.
[0330] Alternatively, in one possible implementation, the processing device 10 further includes a determining module 1003, a clustering module 1004, and a generating module 1005. The determining module 1003 is used to perform... Figure 5 Step 501 in the illustrated embodiment. The clustering module 1004 is used to perform... Figure 5 Steps 502, 5021, 5022, or 5023 in the illustrated embodiment. The generation module 1005 is used to execute... Figure 5 Step 503 in the illustrated embodiment. The control information can be determined based on the instruction information generated by the generation module 1005.
[0331] Figure 10 The module division of the processing device 10 described is merely illustrative and represents only a logical functional division. In actual implementation, there may be other division methods. For example, multiple modules or components may be combined or integrated into another system, or some features may be ignored or not executed. The functional modules in the processing device may be integrated into one processing module, or each module may exist physically separately, or two or more modules may be integrated into one module.
[0332] Figure 10 The modules can be implemented either in hardware or as software functional units. For example, when implemented in software, Figure 10 The various modules in can be generated by Figure 9 The processor 901 reads the program code stored in the memory 902 and generates software function modules to implement it. Figure 10 The modules in can also be made by Figure 9 Different hardware components are implemented separately. For example, the communication module 1001 can be implemented using the communication interface 903, and the measurement module 1002 is implemented using... Figure 9 The processing resources of the processor 901 (such as other cores in a multi-core processor) can be used, or programmable devices such as field-programmable gate arrays (FPGAs) or coprocessors can be employed to complete the task. Clearly, the above functional modules can also be implemented using a combination of software and hardware.
[0333] This application also provides a processing apparatus. This processing apparatus may be a third device mentioned in the method of this application, or a device within a third device, or a device compatible with a third device. In one possible implementation, the processing apparatus may include executing... Figure 5The module of the method shown can be a hardware circuit, a software module, or a module implemented by combining hardware circuit and software.
[0334] Figure 11 This is a schematic diagram of another possible structure of the processing device 11 according to an embodiment of this application. (See attached diagram) Figure 11 The processing device 11 includes a determining module 1101, a clustering module 1102, and a generating module 1103. The determining module 1101 determines multiple samples, each corresponding to a network flow, and each sample includes a flow identifier and statistical value for the corresponding network flow. The clustering module 1102 clusters the multiple samples to obtain multiple clusters, each cluster corresponding to a cluster label, which is the same information unit of the flow identifier in all samples within the corresponding cluster. The generating module 1103 generates indication information based on the samples within each cluster. The indication information includes multiple matching items and the category corresponding to each matching item. Each matching item indicates the flow identification rule of the packet. The indication information indicates that the statistical value of the network flow to which the packet matching the first matching item belongs is less than the statistical value of the network flow to which the packet matching the second matching item belongs. The first matching item is the matching item corresponding to the first category among the multiple matching items, and the second matching item is the matching item corresponding to the second category among the multiple matching items.
[0335] Specifically, for example, the determining module 1101 is used to perform... Figure 5 Step 501 in the illustrated embodiment. The clustering module 1102 is used to perform... Figure 5 Steps 502, 5021, 5022, or 5023 in the illustrated embodiment. The generation module 1103 is used to execute... Figure 5 Step 503 in the illustrated embodiment.
[0336] In one possible implementation, the processing device 11 further includes a communication module 1104 and a measurement module 1105. For example, the communication module 1104 is used to perform... Figure 3 In step 301 of the illustrated embodiment, the measurement module 1105 is used to perform... Figure 3 Step 302 in the illustrated embodiment.
[0337] Figure 11 The module division of the processing device 11 described is merely illustrative and represents only a logical functional division. In actual implementation, there may be other division methods. For example, multiple modules or components may be combined or integrated into another system, or some features may be ignored or not executed. The functional modules in the processing device may be integrated into one processing module, or each module may exist physically separately, or two or more modules may be integrated into one module.
[0338] Figure 11 The modules can be implemented either in hardware or as software functional units. For example, when implemented in software, Figure 11 The various modules in can be generated by Figure 9 The processor 901 reads the program code stored in the memory 902 and generates software function modules to implement it. Figure 11 The modules in can also be made by Figure 9 The different hardware components can be implemented separately, or the above functional modules can be implemented using a combination of software and hardware.
[0339] The coupling in the embodiments of this application is an indirect coupling or communication connection between devices, units, or modules, which can be electrical, mechanical, or other forms, and is used for information interaction between devices, units, or modules.
[0340] Those skilled in the art will understand that when various aspects or possible implementations of the embodiments of this application are implemented using software, all or part of the aforementioned aspects or possible implementations can be implemented in the form of a computer program product. A computer program product refers to instructions (or computer-readable instructions, computer program instructions, functional programs, or program code) stored in a computer-readable medium. When these instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of this application are generated.
[0341] Computer-readable media can be computer-readable signal media or computer-readable storage media. Computer-readable storage media include, but are not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices, or apparatuses, or any suitable combination thereof. Examples of computer-readable storage media include random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM), or compact disc read-only memory (CD-ROM).
[0342] The terms "first," "second," "third," "fourth," etc., used in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such terms are interchangeable where appropriate; this is merely a way of distinguishing objects with the same attributes in the embodiments of this application. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion, so that a process, method, system, product, or apparatus that comprises a series of units is not necessarily limited to those units, but may include other units not explicitly listed or inherent to those processes, methods, products, or apparatuses. The term "multiple" appearing in the embodiments of this application refers to two or more. It should be understood that the term "and / or" in this document is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A alone, A and B simultaneously, and B alone. Additionally, the character " / " in this document generally indicates that the preceding and following related objects are in an "or" relationship.
[0343] It should be understood that in the various embodiments of this application, the order of the above-mentioned processes does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.
[0344] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.
[0345] The above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application.
Claims
1. A method for processing network data, characterized in that, include: The first device identifies multiple samples, each of which corresponds to a network flow, and each sample includes a flow identifier and statistical value of the corresponding network flow; The first device clusters the multiple samples to obtain multiple clusters. Each cluster corresponds to a clustering label, and the clustering label is the same information unit of the flow identifier in all samples within the corresponding cluster. The first device generates indication information based on samples within each of the plurality of clusters. The indication information includes a plurality of matching items and a category corresponding to each of the plurality of matching items. Each of the plurality of matching items indicates the flow identification rules of the packet. The indication information indicates that the statistical value of the network flow to which the packet matching the first matching item belongs is less than the statistical value of the network flow to which the packet matching the second matching item belongs. The first matching item is the matching item corresponding to the first category among the plurality of matching items, and the second matching item is the matching item corresponding to the second category among the plurality of matching items. The indication information also indicates the order among the plurality of matching items. The target matching item for the target packet is the first matching item that the target packet matches from the plurality of matching items according to the order.
2. The method according to claim 1, characterized in that, Each of the plurality of matching terms corresponds to one of the plurality of clusters, each matching term is determined based on the clustering label of the corresponding cluster, and the category corresponding to each matching term is determined based on the statistical value in the samples within the corresponding cluster.
3. The method according to claim 2, characterized in that, The statistical value of the samples within the cluster corresponding to the first category is less than the threshold, and the statistical value of the samples within the cluster corresponding to the second category is greater than the threshold.
4. The method according to claim 1, characterized in that, The first cluster and the second cluster in the plurality of clusters correspond to the first cluster label and the second cluster label, respectively. The number of information units corresponding to the first cluster label is equal to the number of information units corresponding to the second cluster label. The number of samples in the first cluster is greater than the number of samples in the second cluster. Among the plurality of matching items, the matching items corresponding to the first cluster are in the order stated above, which precede the matching items corresponding to the second cluster.
5. The method according to any one of claims 1-4, characterized in that, The third and fourth clusters in the plurality of clusters correspond to the third cluster label and the fourth cluster label, respectively. The number of information units corresponding to the third cluster label is greater than the number of information units corresponding to the fourth cluster label. The matching items corresponding to the third cluster in the plurality of matching items are in the order mentioned above, before the matching items corresponding to the fourth cluster in the plurality of matching items.
6. The method according to any one of claims 1 to 4, characterized in that, The statistical values in each sample are positively correlated with the flow size of the corresponding network flow and / or with the flow rate of the corresponding network flow.
7. The method according to any one of claims 1 to 4, characterized in that, The clustering label corresponds to at least two information units.
8. The method according to claim 7, characterized in that, The at least two information units include an information unit for describing the protocol number.
9. The method according to any one of claims 1 to 4, characterized in that, After the first device generates indication information based on samples within each of the plurality of clusters, the method further includes: The first device generates control information based on the indication information. The control information indicates that the first matching item corresponds to a first measurement strategy, and the second matching item corresponds to a second measurement strategy. The first measurement strategy and the second measurement strategy are used to perform network flow-based measurements.
10. The method according to claim 9, characterized in that, The method further includes: The first device receives a message, which is either a first message or a second message, wherein the first message matches the first matching item, and the second message matches the second matching item; The first device executes the first measurement strategy or the second measurement strategy on the message according to the control information.
11. The method according to any one of claims 1 to 4, characterized in that, After the first device generates indication information based on samples within each of the plurality of clusters, the method further includes: The first device sends a message to the second device, the message being generated based on the indication information. The message is used to instruct the second device to execute a first measurement strategy on a received third message or to execute a second measurement strategy on a received fourth message. The third message matches the first matching item, and the fourth message matches the second matching item. The first measurement strategy and the second measurement strategy are respectively used for network flow-based measurement.
12. A processing apparatus, characterized in that, include: The determination module is used to determine multiple samples, each of which corresponds to a network flow, and each sample includes a flow identifier and statistical value of the corresponding network flow; The clustering module is used to cluster the multiple samples to obtain multiple clusters. Each cluster corresponds to a clustering label, and the clustering label is the same information unit of the flow identifier in all samples within the corresponding cluster. A generation module is used to generate indication information based on samples within each of the plurality of clusters. The indication information includes a plurality of matching items and a category corresponding to each of the plurality of matching items. Each of the plurality of matching items indicates the flow identification rules of the packet. The indication information indicates that the statistical value of the network flow to which the packet matching the first matching item belongs is less than the statistical value of the network flow to which the packet matching the second matching item belongs. The first matching item is the matching item corresponding to the first category among the plurality of matching items, and the second matching item is the matching item corresponding to the second category among the plurality of matching items. The indication information also indicates the order among the plurality of matching items, and the target matching item for the target packet is the first matching item that the target packet matches from the plurality of matching items according to the order.
13. A processing apparatus, characterized in that, It includes a processor and a memory, the memory being coupled to the processor, the memory being used to store program code; the processor being used to invoke the program code to execute the method of any one of claims 1 to 11.
14. A chip system, characterized in that, The device includes a processor and an interface circuit, the processor being coupled to a memory via the interface circuit, the processor being used to execute program code in the memory to implement the method as described in any one of claims 1 to 11.
Citation Information
Patent Citations
Adaptive fair sampling method based on reduction of number of flows
CN106789444A
Network data flow detection method and device
CN110213227A