Traffic processing method and device, electronic equipment, computer readable storage medium and computer program product
By extracting features and clustering network traffic data to generate feature labels, the problem that network traffic scheduling methods cannot flexibly respond to changes in traffic status is solved, thus achieving efficient utilization of network resources and cost reduction.
Patent Information
- Application Number
- CN202410624362.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-05-15
- Publication Date
- 2025-11-18
AI Technical Summary
In existing technologies, network traffic scheduling methods cannot flexibly respond to changes in traffic status, resulting in low efficiency in network resource utilization.
By performing feature extraction, clustering, and label determination on traffic data sets, feature labels are generated, enabling flexible scheduling of traffic data.
It improved the utilization rate of network resources, reduced network transmission costs, and ensured the stability and flexibility of application services.
Smart Images

Figure CN120980035A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the computer technology field, and particularly relates to a traffic processing method and device, electronic equipment, computer readable storage medium and computer program product. BACKGROUND
[0002] In a network system, especially in the cloud computing and big data era, a large amount of traffic data is transmitted every moment, and therefore, the traffic data needs to be scheduled to ensure efficient use of network resources. In the related art, a fixed method is usually used to classify and schedule the traffic data, and the scheduling manner of the traffic data is determined according to historical experience data or the experience of an operation and maintenance personnel, without considering the traffic state change, and the network traffic cannot be flexibly scheduled. SUMMARY
[0003] The embodiments of the present application provide a traffic processing method and device, electronic equipment, computer readable storage medium and computer program product, which can flexibly schedule the traffic data and improve network performance and reduce network cost.
[0004] The technical scheme of the embodiments of the present application is implemented as follows:
[0005] The embodiments of the present application provide a traffic processing method, which comprises the following steps.
[0006] Obtaining a plurality of traffic data sets collected from a network system, performing feature extraction on each traffic data set to obtain feature data corresponding to each traffic data set, wherein the traffic data set comprises a plurality of traffic data under the same source.
[0007] Performing clustering processing on the plurality of traffic data sets based on the feature data corresponding to each traffic data set to obtain a clustering result.
[0008] Determining a feature label of each traffic data set based on the clustering result.
[0009] Performing scheduling processing on each traffic data set based on the feature label of each traffic data set.
[0010] The embodiments of the present application provide a traffic processing device, which comprises the following.
[0011] An obtaining module is configured to obtain a plurality of traffic data sets collected from a network system, perform feature extraction on each traffic data set to obtain feature data corresponding to each traffic data set, wherein the traffic data set comprises a plurality of traffic data under the same source.
[0012] The clustering processing module is configured to perform clustering processing on the plurality of traffic data sets based on the feature data corresponding to each traffic data set, to obtain a clustering result.
[0013] The determining module is configured to determine a feature label of each traffic data set based on the clustering result.
[0014] The scheduling processing module is configured to perform scheduling processing on each traffic data set based on the feature label of each traffic data set.
[0015] An electronic device is provided in an embodiment of the present application, and the electronic device comprises:
[0016] A memory is configured to store computer executable instructions.
[0017] A processor is configured to execute the computer executable instructions stored in the memory, to implement the method provided in the embodiments of the present application.
[0018] A computer readable storage medium is provided in an embodiment of the present application, and the computer readable storage medium stores a computer program or computer executable instructions, and is configured to be executed by a processor to implement the traffic processing method provided in the embodiments of the present application.
[0019] A computer program product is provided in an embodiment of the present application, and the computer program product comprises a computer program or computer executable instructions, and the computer program or computer executable instructions are executed by a processor to implement the traffic processing method provided in the embodiments of the present application.
[0020] The embodiments of the present application have the following beneficial effects:
[0021] In the embodiments of the present application, a plurality of traffic data sets collected from a network system are obtained, feature extraction is performed on each traffic data set to obtain corresponding feature data, clustering processing is performed on the plurality of traffic data sets based on the feature data to obtain a clustering result, a feature label of each traffic data set is determined based on the clustering result, and finally scheduling processing is performed on the traffic data set based on the feature label. In this way, the traffic data set can be analyzed in real time by combining data collection, feature extraction, data clustering and other technologies, so as to adapt to the changes of the traffic data set in the network system and improve the accuracy of classification and identification of the traffic data set. Moreover, the corresponding feature label is generated by performing clustering processing on the traffic data set, the traffic data can be scheduled and processed according to the features possessed by different traffic data sets, so as to realize flexible and accurate scheduling processing of the traffic data set, improve the network resource utilization rate and reduce the traffic data transmission cost. BRIEF DESCRIPTION OF DRAWINGS
[0022] Figure 1 is a structural schematic diagram of the architecture of the traffic processing system 100 provided in the embodiments of the present application.
[0023] Figure 2 is a structural schematic diagram of a server 200 provided by an embodiment of the present application;
[0024] Figure 3A is an implementation flowchart of a traffic processing method provided by an embodiment of the present application;
[0025] Figure 3B is an implementation flowchart of a feature extraction method provided by an embodiment of the present application;
[0026] Figure 3C is an implementation flowchart of a smoothing processing method provided by an embodiment of the present application;
[0027] Figure 3D is an implementation flowchart of a clustering processing method provided by an embodiment of the present application;
[0028] Figure 3E is an implementation flowchart of a model optimization method provided by an embodiment of the present application;
[0029] Figure 4A is a distribution diagram of a daytime traffic data set provided by an embodiment of the present application;
[0030] Figure 4B is a distribution diagram of a nighttime traffic data set provided by an embodiment of the present application;
[0031] Figure 4C is a distribution diagram of a regular traffic data set provided by an embodiment of the present application;
[0032] Figure 4D is a distribution diagram of a burst traffic data set provided by an embodiment of the present application. DETAILED DESCRIPTION
[0033] In order to make the objectives, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings, and the described embodiments should not be regarded as limiting the present application, and all other embodiments obtained by those of ordinary skill in the art without making any creative effort fall within the scope of protection of the present application.
[0034] In the following description, “some embodiments” are described, which describe a subset of all possible embodiments, but it can be understood that “some embodiments” can be the same subset or different subsets of all possible embodiments, and can be combined with each other without conflict.
[0035] In the following description, the terms "first", "second", "third" are merely used to distinguish similar objects, and do not represent a specific order or sequence of the objects. Understandably, the "first", "second", "third" can be interchanged in a specific order or sequence as allowed, so that the embodiments described herein can be implemented in an order other than that illustrated or described herein.
[0036] In the embodiments of the present application, the term "module" or "unit" refers to a computer program or a part of a computer program with a predetermined function, and works together with other related parts to achieve a predetermined target, and can be implemented entirely or partially by using software, hardware (such as a processing circuit or a memory) or a combination thereof. Similarly, one processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be a part of an integral module or unit that includes the functions of the module or unit.
[0037] Unless otherwise defined, all technical and scientific terms used in the embodiments of the present application have the same meanings as those commonly understood by one skilled in the art. The terms used in the embodiments of the present application are only for the purpose of describing the embodiments of the present application and are not intended to limit the present application.
[0038] The relevant data collection process in the embodiments of the present application should strictly comply with the requirements of relevant national laws and regulations, obtain the informed consent or separate consent of the personal information subject, and within the scope of authorization of laws and regulations and the personal information subject, carry out subsequent data use and processing.
[0039] Before further detailing the embodiments of the present application, the terms and phrases involved in the embodiments of the present application are explained, and the terms and phrases involved in the embodiments of the present application are applicable to the following explanations.
[0040] 1) sFlow flow analysis: a packet sampling-based flow statistics technology that can monitor all types of network traffic in real time and continuously. sFlow uses random sampling to count network traffic, so it has less impact on device performance and resource consumption, and is suitable for large-scale network environments.
[0041] 2) Network traffic flow analysis (NetFlow): a flow-based flow statistics technology that classifies packets passing through a router to form flows, and then collects and collects these flows. NetFlow can provide accurate flow statistics data, but since it needs to count all data flows, it has a large resource consumption on the device.
[0042] 3) Internet Protocol Flow Information Export (IPFIX): Provides a standardized way to export information about network traffic, including the source IP address, destination IP address, port number, protocol type, traffic size, etc. IPFIX is an open standard version of NetFlow, inheriting the characteristics of NetFlow and adding some new features, such as supporting more information elements, supporting variable length fields, etc. IPFIX, like NetFlow, is also a flow-based statistics, which consumes a lot of resources of the device.
[0043] 4) Exponential Weighted Moving Average (EWMA): A statistical method used to smooth time series data to identify trends and periodic changes in the data. EWMA is a weighted moving average, where each data point is weighted exponentially decreasing according to its distance from the current point.
[0044] 5) Constrained K-means: An improved K-means clustering algorithm, the main idea is to add constraints in the traditional K-means algorithm, in order to obtain the clustering results that meet the actual application requirements.
[0045] 6) Network system: A complex system composed of various hardware, software, protocols and personnel, used to realize network communication, resource sharing and data exchange functions. The main purpose of the network system is to connect various devices, applications and services to form a unified network environment, to provide efficient, reliable and secure communication services. The main components of the network system include: network devices, network protocols, servers, network security, users and application programs.
[0046] 7) Traffic data: The amount of data passing through a network interface or device within a certain period of time. Traffic data is usually used to monitor and manage network usage, which can help users understand network usage and optimize the allocation of network resources.
[0047] 8) Traffic data set: Including multiple same type traffic data under the same source, which can help network administrators identify which applications or services consume the most bandwidth in the network, and then take corresponding measures to optimize network performance, improve user experience, or manage traffic data to ensure the availability of critical applications.
[0048] In a network system, traffic data is transmitted every moment, involving various applications and services, such as online video, social media, Internet of Things, and the like. Different types of traffic data have different characteristics, and scheduling traffic data according to business requirements can improve network performance. In the related art, a fixed method is usually used to classify and schedule traffic data, and the way of scheduling traffic data is determined according to historical experience data or the experience of operation and maintenance personnel, without considering the change of traffic state, and the network traffic cannot be flexibly scheduled and processed.
[0049] Embodiments of the present application provide a traffic processing method and device, electronic equipment, computer readable storage medium and computer program product, which can flexibly schedule traffic data, improve network performance and reduce network cost. The following describes an exemplary application of the electronic equipment provided by the embodiments of the present application. The device provided by the embodiments of the present application can be implemented as a notebook computer, a tablet computer, a desktop computer, a set-top box, a smart phone, a smart speaker, a smart watch, a smart television, a vehicle-mounted terminal, and the like. The device can also be implemented as a server. The following describes an exemplary application when the device is implemented as a server.
[0050] Referring to Figure 1 , Figure 1 is an architecture schematic diagram of a traffic processing system 100 provided by the embodiments of the present application. The terminal 400 includes a graphical interface 410, and the terminal 400 is connected to the server 200 through the network 300. The network 300 can be a wide area network or a local area network, or a combination of the two.
[0051] Massive traffics are transmitted in terminal 400 and server 200 every moment, which are related to various applications and services, such as online video, social media, Internet of Things, and the like. Terminal 400 is configured to display pages of various applications and services on graphical interface 410. Server 200 is configured to obtain a plurality of traffic data sets collected from a network system (the network system at least includes terminal 400 and server 200), and perform feature extraction on each traffic data set to obtain corresponding feature data. Then, server 200 performs clustering processing on the plurality of traffic data sets based on the feature data to obtain clustering results, and determines a feature label of each traffic data set according to the clustering results, which is used for scheduling processing of the traffic data set. For example, in a scenario where terminal 400 includes an office application and a game application, server 200 collects, in real time, a traffic data set corresponding to the office application and a traffic data set corresponding to the game application from the network system, and performs feature extraction on the two traffic data sets to obtain corresponding feature data. Then, server 200 performs clustering processing on the two traffic data sets according to the feature data, and obtains that the traffic data set corresponding to the office application is a clustering cluster (distribution time is of a daytime type) and the traffic data set corresponding to the game application is another clustering cluster (distribution time is of a night type). Then, server 200 determines that the feature label of the traffic data set corresponding to the office application is “daytime type” and the feature label of the traffic data set corresponding to the game application is “night type” according to the clustering results. Since the two traffic data sets belong to different time periods, in order to save network system resources, the traffic data set corresponding to the office application and the traffic data set corresponding to the game application can be combined and scheduled to the same scheduling link, such as scheduling link A. In this way, when a user uses the office application in the daytime, the traffic data set related to the office application can be transmitted through scheduling link A, so that the user can normally use the related service of the office application; when the user uses the game application at night, the traffic data set related to the game application can be transmitted through scheduling link A, so that the user can normally use the related service of the game application, thereby realizing flexible and accurate scheduling processing of the traffic data set, and providing stable application-related services for the user in the case of saving network resources.
[0052] Referring to Figure 2 , Figure 2 is a structural schematic diagram of server 200 provided by an embodiment of the present application, Figure 2 Server 200 shown in FIG. 2 includes at least one processor 210, a memory 230, and at least one network interface 220. Various components in server 200 are coupled together through a bus system 240. It can be understood that the bus system 240 is used to realize the connection and communication between the components. In addition to a data bus, the bus system 240 also includes a power bus, a control bus, and a status signal bus. However, in order to clearly illustrate, all kinds of buses are marked as bus system 240 in Figure 2 .
[0053] The processor 210 can be an integrated circuit chip that has a processing capability of signals, such as a general purpose processor, a digital signal processor (DSP), or other programmable logic device, discrete gate or transistor logic device, discrete hardware component, etc., where the general purpose processor can be a microprocessor or any conventional processor.
[0054] The memory 230 can be removable, non-removable, or a combination thereof. Exemplary hardware devices include solid-state memory, hard drives, optical drives, etc. The memory 230 optionally includes one or more storage devices remotely located from the processor 210.
[0055] The memory 230 includes volatile memory or non-volatile memory, and can also include both volatile and non-volatile memory. The non-volatile memory can be read only memory (ROM), and the volatile memory can be random access memory (RAM). The memory 230 described in the embodiments of the present application is intended to include any suitable type of memory.
[0056] In some embodiments, the memory 230 is capable of storing data to support various operations, examples of which include programs, modules, and data structures or subsets or supersets thereof, which are exemplarily illustrated below.
[0057] The operating system 231 includes system programs for processing various basic system services and performing hardware-related tasks, such as a framework layer, a core library layer, a driver layer, etc., for implementing various basic services and processing hardware-based tasks.
[0058] The network communication module 232 is used to communicate with other electronic devices via one or more (wired or wireless) network interfaces 220, and exemplary network interfaces 220 include Bluetooth, wireless fidelity (WiFi), universal serial bus (USB), etc.
[0059] In some embodiments, the apparatus provided by the embodiments of the present application can be realized in software, Figure 2 The traffic processing apparatus 233 stored in the memory 230 is shown, which can be software in the form of programs and plug-ins, etc., including the following software modules: an acquisition module 2331, a clustering processing module 2332, a determination module 2333, and a scheduling processing module 2334. These modules are logical, and thus can be combined or further split according to the implemented functions. The functions of each module will be described below.
[0060] In some embodiments, the apparatus provided by the embodiments of the present application can be implemented in a hardware manner. For example, the apparatus provided by the embodiments of the present application can be a processor in the form of a hardware decoding processor, which is programmed to perform the traffic processing method provided by the embodiments of the present application. For example, the processor in the form of a hardware decoding processor can be implemented by using one or more application specific integrated circuits (ASICs), digital signal processors (DSPs), programmable logic devices (PLDs), complex programmable logic devices (CPLDs), field programmable gate arrays (FPGAs), or other electronic elements.
[0061] The traffic processing method provided by the embodiments of the present application will be described in combination with an exemplary application and implementation of the server provided by the embodiments of the present application.
[0062] In the following, the traffic processing method provided by the embodiments of the present application will be described. As described above, the electronic device implementing the traffic processing method provided by the embodiments of the present application can be a server. Therefore, the execution subject of each step will not be repeatedly described hereinafter.
[0063] It should be noted that, in the following examples of traffic processing, network traffic is taken as an example for description. Those skilled in the art can apply the traffic processing method provided by the embodiments of the present application to the processing of other types of traffic according to the understanding of the following description.
[0064] Referring to Figure 3A , Figure 3A is a flowchart of the traffic processing method provided by the embodiments of the present application. The steps shown in Figure 3A will be described in combination.
[0065] In step 101, a plurality of traffic data sets collected from a network system are obtained, and feature extraction is performed on each traffic data set to obtain feature data corresponding to each traffic data set.
[0066] Here, the network system is the overall structure and organization of the interconnected network composed of network devices, software and protocols, etc., including various components connected and interacting on the physical and logical levels to achieve data transmission, communication and resource sharing, etc. A large amount of traffic data is transmitted in the network system, which refers to the amount of data transmitted in the network system, usually used to measure the usage, performance and load of the network, which can include various indicators and statistical information such as bandwidth utilization, packet quantity, data transmission rate, transmission protocol distribution, etc. The traffic data set includes multiple same type traffic data collected at different times from the same source, for example, the same type of traffic data of the same application collected at different times can be used as a traffic data set. Among them, the way to collect the traffic data set is to collect the traffic data through a dedicated network protocol (such as sFlow traffic analysis) on network devices such as routers, switches, etc. and report it to the cloud data system. Each traffic data in the traffic data set includes: packet header information (including source address, destination address, source port, destination port, protocol type, etc.), data flow information (including the number of data packets, the number of bytes, the start time and end time of the flow, etc.), device information (including the IP address of the device, the port number of the device, etc.). Feature data is obtained by feature extraction on the traffic data set, which can include packet quantity, byte quantity, flow start time and end time, source address, destination address, source port, destination port, protocol type, etc. information, which is used to help describe the basic information of the traffic data set. Among them, in order to effectively smooth the feature data, the feature data is calculated by Exponential Weighted Moving Average (EWMA), which is a sequence composed of multiple EWMA values, which can represent the trend of the characteristics of the traffic data over time.
[0067] In some embodiments, referring to Figure 3B The "feature extraction is performed on each traffic data set to obtain the feature data corresponding to each traffic data set" in step 101 can be implemented by steps 1011 to 1013, which are described in detail as follows:
[0068] In step 1011, for each traffic data set, the pre-processing is performed on each traffic data in the traffic data set to obtain the pre-processed traffic data.
[0069] Here, since multiple traffic data sets are collected from the network system, each traffic data in each traffic data set needs to be preprocessed. The preprocessing includes data cleaning (for removing invalid or abnormal traffic data), data conversion (converting the original traffic data into a format that can be used for feature extraction), and the like. The preprocessed traffic data can improve the accuracy of feature extraction.
[0070] In step 1012, the preprocessed traffic data is subjected to feature extraction to obtain initial feature values.
[0071] Here, the specified features in the traffic data can be extracted by a feature extraction tool (such as a traffic data analysis tool, a traffic data processing tool, etc.) to obtain initial feature values. The initial feature values include initial feature values of each traffic data in the traffic data set, and the initial feature values are not subjected to smoothing processing and can include the number of traffic data, the number of bytes, the start time and end time of the flow, the source address, the destination address, the source port, the destination port, the protocol type, etc., which can help describe the basic information of the traffic data.
[0072] In step 1013, each initial feature value is subjected to smoothing processing to obtain feature data corresponding to the traffic data set.
[0073] Here, since the traffic data has certain sporadic characteristics, in order to effectively smooth the traffic data, eliminate short-term fluctuations, and retain long-term trends, the initial feature values need to be subjected to smoothing processing. For example, each initial feature value is processed by applying an exponential weighted moving average algorithm, and the obtained feature data is a sequence of multiple EWMA values, which can represent the trend of the traffic data features over time.
[0074] In the embodiments of the present application, for each traffic data set, each traffic data in the traffic data set is preprocessed to obtain preprocessed traffic data, the preprocessed traffic data is subjected to feature extraction to obtain initial feature values, and each initial feature value is subjected to smoothing processing to obtain feature data corresponding to the traffic data set. In this way, the key information is filtered out from the traffic data set by feature extraction, which can simplify the traffic data set and make it easier to understand and analyze the traffic data set. Moreover, since the feature data is obtained after smoothing processing, the long-term trend of the traffic data set is retained, and the accuracy of the feature data is improved.
[0075] In some embodiments, referring to Figure 3C The above step 1013 can be implemented by steps 10131 to 10134, which are described in detail as follows.
[0076] In step 10131, a smoothing factor and the collection time of the traffic data corresponding to each initial feature value are obtained.
[0077] Here, first determine a smoothing factor a(alpha), the value range is [0, 1] wherein, the smoothing factor is closer to 1, it is higher to the flow data of the collection time nearest to give weight;The smoothing factor is closer to 0, it is higher to the flow data of the collection time farthest to give weight.Collection time is the specific time of each flow data collection in each flow data set, for example, time stamp, it can be in the form of year, month and day, or specific time point form, such as 2024.07.06, 18:44 etc.
[0078] In step 10132, based on each collection time, determine the time point information corresponding to each initial feature value.
[0079] Here, for the convenience of calculation and statistics, according to the collection time, determine the corresponding time point information.Time point information can be a specific numerical value determined according to the collection time, and the collection time is converted into a numerical value, which needs a common reference point or reference time.For example, each collection time is 2024-05-1218:44, 18:2024-05-1250, 2024-05-1218:47, and 2024-05-1218:40 is taken as the reference time, then the time point information corresponding to 2024-05-1218:44 can be determined as 4, the time point information corresponding to 2024-05-1218:50 is 10, and the time point information corresponding to 2024-05-1218:45 is 7.
[0080] In step 10133, for each time point information, based on the smoothing factor and the initial feature value corresponding to the time point information, determine the smoothing feature value corresponding to the time point information.
[0081] Here, for each time point information corresponding to each initial feature value, the smoothing feature value corresponding to the time point information can be calculated by formula 1, formula 1 is as follows:
[0082] EWMA(t)=α×X(t)+(1-α)×EWMA(t-1)(1)
[0083] Wherein, t indicates the time point information corresponding to the initial feature value, X(t) indicates the initial feature value at time point information t, EWMA(t) indicates the smoothing feature value at time point information t.Exemplarily, for the first time point information, EWMA(1) can be initialized as X(1).Traverse each time point information, use the above formula (1) to calculate the EWMA value of each time point information as the corresponding smoothing feature value.
[0084] In step 10134, the smoothing feature value corresponding to each time point information is determined as the feature data corresponding to the flow data set.
[0085] Here, the smooth characteristic value sequence can be obtained according to the smooth characteristic values corresponding to the time point information, and the smooth characteristic value sequence is determined as the characteristic data corresponding to the traffic data set, which is used to represent the change trend of the characteristic data of each traffic data in the traffic data set over time.
[0086] In the embodiment of the present application, the smoothing factor and the collection time of the traffic data corresponding to each initial characteristic value are obtained, the time point information corresponding to each initial characteristic value is determined based on the collection time, for each time point information, the smooth characteristic value corresponding to the time point information is determined based on the smoothing factor and the initial characteristic value corresponding to the time point information, and the smooth characteristic value corresponding to each time point information is determined as the characteristic data corresponding to the traffic data set. In this way, the fluctuation of the traffic data in the short term is effectively eliminated by smoothing the characteristic data, the long-term development trend is retained, and the accuracy of the characteristic data is improved.
[0087] In step 102, the multiple traffic data sets are clustered based on the characteristic data corresponding to each traffic data set, and a clustering result is obtained.
[0088] Here, after the feature extraction of the traffic data set, the clustering model can be used to classify the traffic data set through the clustering algorithm, and the traffic data sets with similar characteristics can be gathered together. The clustering result is used to represent that the traffic data sets of the same type are gathered into a clustering cluster, and the traffic data sets of different types are divided into different clustering clusters. For example, the traffic data sets are A, B, C and D, the traffic distribution time of the traffic data sets A and B is concentrated in the evening, and the traffic distribution time of the traffic data sets C and D is concentrated in the daytime, and the clustering result is that the traffic data sets A and B are a clustering cluster, and the traffic data sets C and D are another clustering cluster.
[0089] In some embodiments, referring to Figure 3D The above step 102 can be implemented by steps 1021 to 1023, which are specifically described as follows:
[0090] In step 1021, at least one target clustering model is obtained.
[0091] Here, the target clustering model is obtained by pre-training. By obtaining at least one initial clustering model and the training data corresponding to each initial clustering model, for each initial clustering model, the initial clustering model is trained based on the training data, and the target clustering model is obtained.
[0092] The initial clustering model refers to a model of an initial clustering algorithm, such as a K-means clustering model. Different clustering algorithms can cluster traffic data sets from different perspectives, and thus multiple types of target clustering models can be constructed, such as a target clustering model for clustering according to traffic distribution time, a target clustering model for clustering according to traffic size, and a target clustering model for clustering according to traffic security level. Each type of target clustering model has corresponding training data, which includes multiple sample traffic data sets and sample clustering results of the multiple sample traffic data sets. The training data can be obtained by manual screening.
[0093] Specifically, in the initial clustering model training process, each sample traffic data set in the training data is input into the initial clustering model to enable the initial clustering model to perform clustering processing based on the sample traffic data set, obtain a clustering result, and then determine a predicted loss value according to a preset loss function and a corresponding sample clustering result in the training data. The loss value is back-propagated to the initial clustering model, and the parameters of the initial clustering model are adjusted by a gradient descent algorithm based on the loss value. The above steps are repeated for other sample traffic data sets in the training data until a target clustering model is obtained when a training end condition is reached. The training end condition can be that the loss value is lower than a loss threshold or that the difference between the current loss value and the previous adjacent loss value is less than a preset difference value.
[0094] In step 1022, for each target clustering model, multiple traffic data sets are clustered based on the feature data corresponding to each traffic data set, to obtain an initial clustering result corresponding to the target clustering model.
[0095] Here, the initial clustering result corresponds to the target clustering model and is used to indicate that traffic data sets of the same type are clustered into one clustering cluster and traffic data sets of different types are divided into different clustering clusters. For example, if the target clustering model is used to cluster multiple traffic data sets according to the distribution time of the traffic data, the corresponding initial clustering result is related to the distribution time of the traffic data, and traffic data sets A and B can be one clustering cluster (the distribution time belongs to the daytime type), and traffic data sets C and D can be another clustering cluster (the distribution time belongs to the nighttime type).
[0096] In step 1023, the initial clustering result corresponding to each target clustering model is determined as the clustering result.
[0097] Here, the clustering result includes initial clustering results corresponding to multiple target clustering models. For example, the target clustering models include a target clustering model that clusters multiple traffic data sets according to distribution time of the traffic data, and a target clustering model that clusters multiple traffic data sets according to size of the traffic data. Two initial clustering results of the traffic data sets are as follows: traffic data sets A and B form one cluster (distribution time belongs to daytime type), and traffic data sets C and D form another cluster (distribution time belongs to nighttime type); traffic data sets A and D form one cluster (size of the traffic data belongs to large type), and traffic data sets B and C form another cluster (size of the traffic data belongs to small type). The clustering result is as follows: traffic data sets A and B form one cluster (daytime type), traffic data sets C and D form one cluster (nighttime type), traffic data sets A and D form one cluster (large type), and traffic data sets B and C form one cluster (small type).
[0098] In the embodiment of the present application, at least one target clustering model is obtained, for each target clustering model, multiple traffic data sets are clustered based on feature data corresponding to each traffic data set, to obtain an initial clustering result corresponding to the target clustering model, and then the initial clustering result corresponding to each target clustering model is determined as the clustering result. In this way, the same traffic data set can be divided into different clusters from different clustering perspectives by multiple target clustering models, different clustering results are obtained, comprehensive analysis of the traffic data sets is performed, and the clustering accuracy of the traffic data sets is improved.
[0099] In some embodiments, referring to Figure 3E To ensure the clustering accuracy of the target clustering model, the target clustering model can be optimized by steps 201 to 205, which will be described in detail below.
[0100] In step 201, for each target clustering model, at least one historical clustering result corresponding to the target clustering model is obtained when a model optimization opportunity is reached.
[0101] Here, the model optimization opportunity is pre-set, for example, the target clustering model is optimized every pre-set time (such as 5 minutes), or the target clustering model is optimized when the historical clustering result reaches a pre-set number (such as 10). The historical clustering result is a clustering result generated before the model optimization opportunity is reached, which corresponds to the clustering type of the target clustering model. For example, if the target clustering model can classify traffic data sets according to the size of the traffic data, then the clustering result obtained by clustering the traffic data sets according to the size of the traffic data before the model optimization opportunity is reached is obtained as the historical clustering result corresponding to the target clustering model.
[0102] In step 202, the verification processing is performed on each historical clustering result to obtain a verification result corresponding to each historical clustering result.
[0103] Here, the verification processing is used to determine the clustering accuracy of each historical clustering result, and the verification result includes verification pass and verification fail. For example, the historical clustering result is that traffic data sets A and B are a group, and traffic data sets C and D are a group. If it is determined through verification that traffic data sets A and B are data of one type and traffic data sets C and D are data of another type, a verification result of verification pass is obtained; otherwise, a verification result of verification fail is obtained.
[0104] In step 203, the historical clustering result with the verification result of verification pass is determined as a target historical clustering result.
[0105] Here, the target historical clustering result is used to optimize the target clustering model, and the verification result of verification pass indicates that the corresponding historical clustering result is accurate. The historical clustering result with the verification result of verification pass is determined as the target historical clustering result, which is used to improve the clustering accuracy of the target clustering model.
[0106] In step 204, the training data is updated based on the target historical clustering result to obtain new training data.
[0107] Here, the target historical clustering result and the target traffic data set corresponding to the target historical clustering result can be added to the training data to obtain the new training data, or the target historical clustering result and the target traffic data set can be used to replace the original data in the training data to obtain the new training data.
[0108] In step 205, the target clustering model is optimized based on the new training data to obtain an optimized target clustering model.
[0109] Here, taking an example in which the new training data includes the target historical clustering result and the target traffic data set, each target traffic data set in the new training data is input into the target clustering model, so that the target clustering model performs clustering processing based on the target traffic data set to obtain a clustering result. Then, a predicted loss value is determined according to a preset loss function and the target historical clustering result corresponding to the new training data, and the loss value is back propagated to the target clustering model. Based on the loss value, the parameters of the target clustering model are adjusted by a gradient descent algorithm. The above steps are repeated for other target traffic data sets in the new training data until a training end condition is reached to obtain the optimized target clustering model. The training end condition can be that the loss value is lower than a loss threshold or the difference between the current loss value and the previous adjacent loss value is less than a preset difference value.
[0110] In addition, the accuracy of the optimized target clustering model can be evaluated by other evaluation indexes, such as recall (an important index in binary classification model evaluation, which measures the ability of the model to capture all actual positive samples), F1 score (an index for measuring the accuracy of a binary classification model, which takes into account the accuracy and recall of the classification model), and the like. If the index does not meet the basic requirements, the target clustering model needs to be adjusted. In addition, in the actual use process, the traffic data set needs to undergo multiple clustering calculations, and therefore the performance of the algorithm is also a factor to be considered. The target clustering model can be measured by indexes such as within-cluster sum of squares (WCSS), silhouette coefficient, and running time. If the performance does not meet the requirements, the algorithm parameters of the target clustering model are adjusted, for example, the initialization of the centroid is optimized, the preprocessing of the outliers is optimized, and the multi-dimensional feature value standardization processing is performed.
[0111] Correspondingly, the step 102 can also be implemented by the following process: obtaining at least one optimized target clustering model, for each optimized target clustering model, performing clustering processing on the plurality of traffic data sets based on the feature data corresponding to each traffic data set to obtain an initial clustering result corresponding to the optimized target clustering model, and determining the initial clustering result corresponding to each optimized target clustering model as the clustering result. In this way, the accuracy of the initial clustering result can be improved by performing clustering processing on the plurality of traffic data sets by using the optimized target clustering model.
[0112] In the embodiments of the present application, for each target clustering model, when the model optimization opportunity is reached, at least one historical clustering result corresponding to the target clustering model is obtained, each historical clustering result is verified to obtain a verification result corresponding to each historical clustering result, the historical clustering result with a verification result of verification passed is determined as a target historical clustering result, and the target clustering model is optimized based on new training data to obtain an optimized target clustering model. In this way, the clustering accuracy of the target clustering model can be improved by optimizing the target clustering model based on the target historical clustering result that has been generated and verified. In addition, the training data can be updated in real time to realize iterative optimization of the model and continuously improve the clustering accuracy. Therefore, the accuracy of clustering the plurality of traffic data sets can be continuously improved by using the optimized target clustering model to perform clustering processing on the plurality of traffic data sets, which facilitates the analysis and scheduling processing of the traffic data set.
[0113] In step 103, the feature label of each traffic data set is determined based on the clustering result.
[0114] Here, the feature label is a label generated according to the clustering result, which is a label list including a plurality of initial feature labels, and is used to represent the characteristics of the traffic data set, such as daytime type, burst, large, etc., and the feature label can be directly marked on the traffic data set.
[0115] Specifically, the preset label setting rule can be obtained, and for each initial clustering result in the clustering result, the initial feature label corresponding to the initial clustering result is determined based on the label setting rule, and the initial feature label corresponding to each initial clustering result is determined as the feature label. Wherein, the preset label setting rule is set in advance, after the target clustering model clustering processing is completed to obtain the initial clustering result, the corresponding initial feature label can be directly determined by obtaining the preset label setting rule. When training the target clustering model, the clustering number K of the target clustering model and the feature label corresponding to each type of initial clustering result are determined, such as K=4, and the initial feature labels corresponding to the four types of initial clustering results are night type, daytime type, regular type and burst type. In this way, the corresponding initial feature label can be directly generated after the target clustering model clustering processing is completed to obtain the initial clustering result.
[0116] Alternatively, the label setting information is obtained, and for each initial clustering result in the clustering result, the initial feature label corresponding to the initial clustering result is determined based on the label setting information; and the initial feature label corresponding to each initial clustering result is determined as the feature label. Here, after the target clustering model clustering processing is completed to obtain the initial clustering result, the initial clustering result can be output to the user for display, and then the user can manually determine the type of the traffic data set corresponding to each initial clustering result according to the experience value, and input the corresponding label setting information. By obtaining the label setting information input by the user, the initial feature label corresponding to each initial clustering result is determined, and the feature label is further obtained.
[0117] In the embodiments of the present application, the preset label setting rule can be obtained, and for each initial clustering result in the clustering result, the initial feature label corresponding to the initial clustering result is determined based on the label setting rule. Alternatively, the label setting information is obtained, and the initial feature label corresponding to the initial clustering result is determined based on the label setting information; and the initial feature label corresponding to each initial clustering result. Then the initial feature label corresponding to each initial clustering result is determined as the feature label. In this way, not only the initial feature label can be directly generated by the preset label setting rule, but also the initial feature label can be generated by the manually input label setting information, and then the feature label is generated according to the initial feature label, which improves the accuracy of the feature label and facilitates the analysis and scheduling processing of the traffic data set.
[0118] In step 104, each traffic data set is scheduled based on the feature label of each traffic data set.
[0119] Here, the traffic data sets with different usage periods (e.g., a daytime traffic data set and a nighttime traffic data set) can be combined and scheduled on the same link to improve the utilization of the link. Alternatively, the traffic data set with low requirements for network quality can be scheduled on a low-priority link to reduce costs. In addition, if a traffic data set has an abnormal feature and may have a security risk, a copy of the traffic data set can be sent to a security detection server for security detection. When the detection triggers an alarm, it indicates that the traffic data set is abnormal. The abnormal traffic data set is shielded to ensure the security of the network system.
[0120] In some embodiments, the above step 104 can be implemented by the following process: obtaining a traffic data scheduling rule, for each feature label of a traffic data set, determining a preset feature label in the traffic data scheduling rule that is the same as the feature label of the traffic data set as a target feature label, determining a preset scheduling link in the traffic data scheduling rule that corresponds to the target feature label as a target scheduling link, and scheduling the traffic data set corresponding to the feature label of the traffic data set to the target scheduling link.
[0121] Here, the traffic data scheduling rule is used to indicate that the traffic data set corresponding to the preset feature label is scheduled to the corresponding preset scheduling link. For example, the feature labels of two traffic data sets are "daytime" and "nighttime", respectively. The traffic data scheduling rule sets that the traffic data sets of "daytime" and "nighttime" are combined and scheduled to a preset scheduling link Y. Then, the preset feature labels of "daytime" and "nighttime" are determined as target feature labels, and the preset scheduling link Y is determined as a target scheduling link. Therefore, the traffic data corresponding to "daytime" and "nighttime" is scheduled to the target scheduling link Y.
[0122] Among them, the setting information for the traffic data scheduling rule can be obtained first, and then the multiple preset feature labels and the preset scheduling link corresponding to each preset feature label are determined based on the setting information. For example, the setting information is that the traffic data set with low requirements is scheduled to scheduling link C. Then, the preset feature label is "low requirements", and the corresponding preset scheduling link is "scheduling link C". Then, the traffic data scheduling rule is determined based on each preset feature label and the preset scheduling link corresponding to the preset feature label.
[0123] In the embodiment of the present application, the scheduling rule for the traffic data is determined according to the setting information of the scheduling rule, and then the scheduling rule is obtained when the traffic data set is scheduled. For the feature label of each traffic data set, the preset feature label corresponding to the target feature label in the traffic data scheduling rule is determined as the target feature label, and the preset scheduling link corresponding to the target feature label in the traffic data scheduling rule is determined as the target scheduling link. Then the traffic data corresponding to the feature label of the traffic data set is scheduled to the target scheduling link. In this way, the corresponding scheduling link can be set in advance according to the characteristics of different traffic data sets, so as to realize flexible scheduling of traffic data.
[0124] In the embodiment of the present application, a plurality of traffic data sets collected from a network system are obtained, the feature data corresponding to each traffic data set is obtained by feature extraction, and the clustering result of the plurality of traffic data sets is obtained by clustering based on the feature data. The feature label of each traffic data set is determined based on the clustering result, and finally the traffic data set is scheduled based on the feature label. In this way, the combination of data collection, feature extraction, data clustering and other technologies can analyze the traffic data set in real time, so as to adapt to the change of the traffic data set in the network system and improve the accuracy of classification and identification of the traffic data set. Moreover, by clustering the traffic data set to generate the corresponding feature label, the traffic data can be scheduled according to the characteristics of different traffic data sets, so as to realize flexible and accurate scheduling of the traffic data set, improve the utilization of network resources and reduce the cost of traffic data transmission.
[0125] In the following, an exemplary application of the embodiment of the present application in an actual application scenario will be described.
[0126] The embodiment of the present application provides an adaptive network traffic data identification and classification system, which can provide real-time network traffic classification and identification capability through data collection, data analysis, clustering and other methods, so that the system can automatically schedule the network traffic flexibly, improve the network performance and reduce the cost, so as to effectively manage the network resources, improve the network performance, prevent network attacks and meet different business needs. The embodiment of the present application will use big data technology to process the collected traffic data (including source and destination IP, network rate, traffic size, time distribution and other indicators), and through self-learning and iterative optimization, the accuracy of classification and identification is continuously improved.
[0127] First, a traffic data set needs to be collected. The collection method is to collect traffic data through a dedicated network protocol on network devices such as routers, switches, etc., and report it to the cloud data system. Among them, the common traffic collection protocols are as follows: sFlow traffic analysis, network traffic flow (NetFlow), and network traffic flow guide (Internet Protocol Flow Information Export, IPFIX). Since the embodiments of the present application consider application in large network systems, the sFlow protocol is preferred. Because the sFlow protocol has good real-time performance, it samples each data packet, so it can reflect the network status in real time; at the same time, sFlow uses random sampling, which consumes less device resources and is more suitable for large-scale network environments; in addition, sFlow is an open standard supported by many network device manufacturers, so it can be used on devices from different manufacturers. The collected traffic data includes: packet header information (including source address, destination address, source port, destination port, protocol type, etc.), data flow information (including the number of data packets, the number of bytes, the start time and end time of the flow, etc.), device information (including the IP address of the device, the port number of the device, etc.).
[0128] Then, feature extraction is performed on the traffic data in the traffic data set. First, the traffic data is preprocessed, including data cleaning (to remove invalid or abnormal traffic data), data conversion (to convert the original traffic data into a format that can be used for feature extraction), and other operations. Then, useful initial feature data is extracted from the preprocessed traffic data. These initial feature data may include the number of data packets, the number of bytes, the start time and end time of the flow, the source address, the destination address, the source port, the destination port, the protocol type, etc. These initial feature data can help describe the basic information of network traffic data.
[0129] Among them, since a particular network traffic data has certain sporadic characteristics, in order to effectively smooth the data, eliminate short-term fluctuations, and retain long-term development trends, the Exponential Weighted Moving Average (EWMA) algorithm is applied to process the feature values of the traffic data. The specific process is as follows: determine a smoothing factor α (alpha), whose value range is 0 to 1. The closer α is to 1, the higher the weight assigned to the most recent data points; the closer α is to 0, the higher the weight assigned to historical data points. For each original feature data, such as the number of data packets, the smoothed feature value (EWMA value) can be calculated using the following smoothing formula (1):
[0130] EWMA(t) = α × X(t) + (1 - α) × EWMA(t - 1) (1)
[0131] wherein t represents time point information corresponding to the initial feature value, X(t) represents the initial feature value at the time point information t, and EWMA(t) represents the smoothed feature value at the time point information t. For example, for the first time point information, EWMA(1) can be initialized as X(1). By traversing each time point information, the EWMA value of each time point information is calculated as the corresponding smoothed feature value using the above formula (1), and the obtained sequence of EWMA values (smoothed feature values corresponding to each time point information) is taken as the feature data, which can represent the change trend of the feature data over time.
[0132] Then, after feature extraction is performed on the traffic data sets, a clustering algorithm can be used to perform clustering processing on the traffic data sets, so as to realize classification and gather traffic data sets with similar features together, thereby helping to identify traffic data sets of different types. There are many clustering algorithms, and here a constrained K-means clustering algorithm is used. The constrained K-means clustering algorithm adds some constraints to the traditional K-means algorithm to solve some limitations of the K-means algorithm. In addition, since the traffic data has certain regularity, the traffic data sets can be classified according to empirical values in advance.
[0133] The traffic data sets are clustered according to the time of traffic data set distribution. Some traffic data in some traffic data sets are concentrated in the early morning, as shown in FIG. 6, wherein the abscissa represents time point information, and the ordinate represents traffic data. The traffic data is concentrated between time point information 130 (an approximate value) and 300 (corresponding to the early morning period). Figure 4A Some traffic data in some traffic data sets are concentrated in the evening, as shown in FIG. 7, wherein the abscissa represents time point information, and the ordinate represents traffic data. The traffic data is concentrated between time point information 0 and 210 (an approximate value) (corresponding to the evening period). Figure 4B Some traffic data in some traffic data sets fluctuate stably, as shown in FIG. 8, wherein the abscissa represents time point information, and the ordinate represents traffic data. The traffic data fluctuates stably between time point information 0 and 300 without sudden changes. Figure 4C Some traffic data in some traffic data sets are all sudden, as shown in FIG. 9, wherein the abscissa represents time point information, and the ordinate represents traffic data. The traffic data fluctuates suddenly and continuously between time point information 0 and 300. Figure 4D
[0134] For a target clustering model, the number of clusters K and the feature tags of each type need to be determined first, such as K = 4 and four types of feature tags: night type, day type, regular type, and burst type. During the training process of the target clustering model, a batch of training data needs to be prepared for each type of feature tag. The first time these historical experience data are generated, manual screening is required. In the subsequent iterative upgrade model, the clustering results that have been calculated and verified by the previous clustering algorithm can be used as new training data. In order to improve the accuracy of the target clustering model, more training data should be used to optimize the model. The optimized target clustering model will learn an initial clustering result that meets the constraint conditions.
[0135] After a set of multiple traffic data sets is calculated by a clustering algorithm, the clustering result of the set of traffic data sets will be obtained. According to the clustering result, the initial feature tags of the traffic data sets are determined. Different clustering algorithms can generate different initial feature tags for traffic data sets from different perspectives. Therefore, multiple types of target clustering models can be constructed to cluster traffic data sets and obtain a list of multiple initial tags as feature tags.
[0136] For a target clustering model, the accuracy of the target clustering model needs to be evaluated by evaluation indicators. The recall rate and F1 score can be used to measure the accuracy. If the indicators do not meet the basic requirements, the target clustering model needs to be adjusted. In addition, the traffic data sets need to be calculated by multiple clustering algorithms, so the performance of the algorithm is also a factor to be considered. The performance of the target clustering model can be measured by calculating the within-cluster sum of squares (WCSS), silhouette coefficient, and running time. If the performance does not meet the requirements, the algorithm parameters need to be adjusted. The methods that can be used include optimizing the initialization of the centroid, optimizing the preprocessing of outliers, and multi-dimensional feature value standardization processing.
[0137] Finally, the traffic data sets are flexibly scheduled. Through the processing of the above manners, the feature labels of the traffic data sets can be obtained, and the traffic data sets are optimally scheduled according to the feature labels. For example, traffic data sets with different usage periods (for example, traffic data sets of daytime type and traffic data sets of nighttime type) are combined and scheduled on the same scheduling link to improve the utilization rate of the scheduling link. Or, traffic data sets with low network quality requirements usually have large latency, and users have low quality requirements for the traffic data sets, so the traffic data sets can be scheduled on a low-priority scheduling link to reduce costs. In addition, if the traffic data sets have abnormal features, indicating possible security risks, the traffic data sets can be copied to a security detection server. When the detection triggers an alarm, it indicates that the traffic data sets are abnormal traffic sets, and the abnormal traffic sets are shielded to protect network security.
[0138] The embodiment of the present application proposes a network scheduling system, which can flexibly schedule real-time network traffic data sets to reduce costs and increase efficiency of the network system. The target clustering model in the embodiment of the present application can iteratively adapt to the changing data traffic sets according to the feature information of the traffic data sets, improve the accuracy and efficiency of clustering processing of the traffic data sets, help flexibly schedule the traffic data, and improve network performance and reduce network costs.
[0139] The following continues to illustrate an exemplary structure of the traffic processing device 233 implemented as a software module provided by the embodiment of the present application. In some embodiments, as shown in Figure 2 The software module in the traffic processing device 233 stored in the memory 230 can include:
[0140] The acquisition module 2331 is configured to acquire a plurality of traffic data sets collected from a network system, extract features of each traffic data set, and obtain feature data corresponding to each traffic data set. The traffic data set includes a plurality of traffic data from the same source.
[0141] The clustering processing module 2332 is configured to cluster the plurality of traffic data sets based on the feature data corresponding to each traffic data set to obtain a clustering result.
[0142] The determination module 2333 is configured to determine a feature label of each traffic data set based on the clustering result.
[0143] The scheduling processing module 2334 is configured to schedule each traffic data set based on the feature label of each traffic data set.
[0144] In some embodiments, the obtaining module 2331 is further configured to, for each of the traffic data sets, pre-process each traffic data in the traffic data set to obtain pre-processed traffic data, and perform feature extraction on the pre-processed traffic data to obtain initial feature values, and perform smoothing processing on each initial feature value to obtain feature data corresponding to the traffic data set.
[0145] In some embodiments, the obtaining module 2331 is further configured to obtain a smoothing factor and collection times of traffic data corresponding to each initial feature value, determine time point information corresponding to each initial feature value based on each collection time, for each time point information, determine a smoothing feature value corresponding to the time point information based on the smoothing factor and the initial feature value corresponding to the time point information, and determine the smoothing feature value corresponding to each time point information as the feature data corresponding to the traffic data set.
[0146] In some embodiments, the clustering processing module 2332 is further configured to obtain at least one target clustering model, for each target clustering model, perform clustering processing on the plurality of traffic data sets based on the feature data corresponding to each traffic data set to obtain an initial clustering result corresponding to the target clustering model, and determine the initial clustering result corresponding to each target clustering model as the clustering result.
[0147] In some embodiments, the obtaining module 2331 is further configured to obtain at least one initial clustering model and training data corresponding to each initial clustering model, for each initial clustering model, train the initial clustering model based on the training data to obtain a target clustering model.
[0148] In some embodiments, the obtaining module 2331 is further configured to, for each target clustering model, obtain at least one historical clustering result corresponding to the target clustering model in a case where a model optimization opportunity is reached, perform verification processing on each historical clustering result to obtain a verification result corresponding to each historical clustering result, determine a historical clustering result with a verification result of verification passed as a target historical clustering result, update the training data based on the target historical clustering result to obtain new training data, optimize the target clustering model based on the new training data to obtain an optimized target clustering model.
[0149] In some embodiments, the clustering processing module 2332 is further configured to obtain at least one of the optimized target clustering model; for each optimized target clustering model, perform clustering processing on the plurality of traffic data sets based on the feature data corresponding to each traffic data set, to obtain an initial clustering result corresponding to the optimized target clustering model; and determine the initial clustering result corresponding to each optimized target clustering model as the clustering result.
[0150] In some embodiments, the determining module 2333 is further configured to obtain a preset label setting rule, determine, for each initial clustering result in the clustering result, an initial feature label corresponding to the initial clustering result based on the label setting rule; or obtain label setting information, determine, for each initial clustering result in the clustering result, an initial feature label corresponding to the initial clustering result based on the label setting information; and determine the initial feature label corresponding to each initial clustering result as the feature label.
[0151] In some embodiments, the scheduling processing module 2334 is further configured to obtain a traffic data scheduling rule; for the feature label of each traffic data set, determine a preset feature label in the traffic data scheduling rule that is the same as the feature label of the traffic data set as a target feature label; determine a preset scheduling link in the traffic data scheduling rule that corresponds to the target feature label as a target scheduling link; and schedule the traffic data corresponding to the feature label of the traffic data set to the target scheduling link.
[0152] In some embodiments, the determining module 2333 is further configured to obtain setting information for a traffic data scheduling rule; determine a plurality of preset feature labels and a preset scheduling link corresponding to each preset feature label based on the setting information; and determine the traffic data scheduling rule based on each preset feature label and the preset scheduling link corresponding to the preset feature label.
[0153] Embodiments of the present application provide a computer program product, which includes a computer program or computer executable instructions stored in a computer readable storage medium. A processor of an electronic device reads the computer executable instructions from the computer readable storage medium, and the processor executes the computer executable instructions to cause the electronic device to perform the traffic processing method provided in the embodiments of the present application.
[0154] Embodiments of the present application provide a computer readable storage medium, which stores computer executable instructions or computer programs. When the computer executable instructions or computer programs are executed by a processor, the processor will execute the traffic processing method provided in the embodiments of the present application, for example, as described above. Figure 3AA traffic processing method is shown.
[0155] In some embodiments, the computer-readable storage media can be a RAM, a ROM, a flash memory, a magnetic surface memory, an optical disk, or a CD-ROM memory, etc. It can also be various devices including one or any combination of the above-mentioned memories.
[0156] In some embodiments, the computer-executable instructions can be in the form of programs, software, software modules, scripts or code, written in any form of programming language, including compiled or interpreted languages, or declarative or procedural languages, and can be deployed in any form, including as a stand-alone program or as a module, component, subroutine, or other unit suitable for use in a computing environment.
[0157] As an example, computer-executable instructions can, but need not, reside in a file system's files, can be stored in a portion of a file that holds other programs or data, e.g., one or more scripts stored in a markup language document, in a single file dedicated to the program in question, or in multiple coordinated files, e.g., files that store one or more modules, sub programs, or portions of code.
[0158] As an example, computer-executable instructions can be deployed to be executed on one electronic device or on multiple electronic devices that are located at one site or distributed across multiple sites and interconnected by a communication network.
[0159] To sum up, through the embodiments of the present application, a plurality of traffic data sets collected from a network system are acquired, feature extraction is performed on each traffic data set to obtain corresponding feature data, clustering processing is performed on the plurality of traffic data sets based on the feature data to obtain a clustering result, a feature label of each traffic data set is determined based on the clustering result, and finally, scheduling processing is performed on the traffic data sets based on the feature label. In this way, by combining the technologies of data collection, feature extraction, data clustering, etc., the traffic data sets can be analyzed in real time, so as to adapt to the changes of the traffic data sets in the network system and improve the accuracy of classification and identification of the traffic data sets. Moreover, by performing clustering processing on the traffic data sets to generate corresponding feature labels, the traffic data can be scheduled and processed in a targeted manner according to the features possessed by different traffic data sets, so as to realize flexible and accurate scheduling processing of the traffic data sets, improve the utilization rate of network resources, and reduce the cost of traffic data transmission.
[0160] The above merely provides an example of the present application, but is not intended to limit the protection scope of the present application. Any modification, equivalent replacement, and improvement made within the spirit and scope of the present application shall be included in the protection scope of the present application.
Claims
1. A flow processing method, characterized in that, The method includes: Multiple traffic data sets collected from the network system are acquired, and feature extraction is performed on each traffic data set to obtain feature data corresponding to each traffic data set; wherein, the traffic data set includes multiple traffic data from the same source; Based on the feature data corresponding to each traffic data set, the multiple traffic data sets are clustered to obtain clustering results; Based on the clustering results, the feature labels of each of the traffic data sets are determined; Each traffic data set is scheduled based on its feature label.
2. The method according to claim 1, characterized in that, The step of extracting features from each traffic data set to obtain feature data corresponding to each traffic data set includes: For each of the aforementioned traffic data sets, each traffic data in the traffic data set is preprocessed to obtain preprocessed traffic data. Feature extraction is performed on the preprocessed traffic data to obtain initial feature values; The initial feature values are smoothed to obtain the feature data corresponding to the traffic data set.
3. The method according to claim 2, characterized in that, The smoothing process for each initial feature value to obtain the feature data corresponding to the traffic data set includes: Obtain the smoothing factor and the acquisition time of the traffic data corresponding to each initial feature value; The time point information corresponding to each initial feature value is determined based on each acquisition time. For each of the aforementioned time point information, a smoothing feature value corresponding to the time point information is determined based on the smoothing factor and the initial feature value corresponding to the time point information. The smoothed feature values corresponding to the information at each time point are determined as the feature data corresponding to the traffic data set.
4. The method according to any one of claims 1 to 3, characterized in that, The clustering process, which involves performing clustering on the multiple traffic data sets based on the feature data corresponding to each of the traffic data sets, to obtain clustering results, includes: Obtain at least one target clustering model; For each target clustering model, clustering is performed on the multiple traffic data sets based on the feature data corresponding to each traffic data set to obtain the initial clustering result corresponding to the target clustering model; The initial clustering results corresponding to each of the target clustering models are determined as the clustering results.
5. The method according to claim 4, characterized in that, The method further includes: Obtain at least one initial clustering model and the training data corresponding to each initial clustering model; For each initial clustering model, the initial clustering model is trained based on the training data to obtain the target clustering model.
6. The method according to claim 5, characterized in that, The method further includes: For each target clustering model, when the model optimization opportunity is reached, at least one historical clustering result corresponding to the target clustering model is obtained; Each historical clustering result is validated to obtain the validation result corresponding to each historical clustering result; The historical clustering results that pass the verification are identified as the target historical clustering results; The training data is updated based on the target historical clustering results to obtain new training data; The target clustering model is optimized based on the new training data to obtain the optimized target clustering model. The clustering process, which involves performing clustering on the multiple traffic data sets based on the feature data corresponding to each of the traffic data sets, to obtain clustering results, includes: Obtain at least one of the optimized target clustering models; For each optimized target clustering model, clustering is performed on the multiple traffic data sets based on the feature data corresponding to each traffic data set to obtain the initial clustering result corresponding to the optimized target clustering model. The initial clustering results corresponding to each of the optimized target clustering models are determined as the clustering results.
7. The method according to any one of claims 1 to 3, characterized in that, The step of determining the feature label for each traffic data set based on the clustering results includes: Obtain preset label setting rules, and for each initial clustering result in the clustering results, determine the initial feature label corresponding to the initial clustering result based on the label setting rules; Alternatively, obtain label setting information, and for each initial clustering result in the clustering results, determine the initial feature label corresponding to the initial clustering result based on the label setting information; The initial feature labels corresponding to each of the initial clustering results are determined as the feature labels.
8. The method according to any one of claims 1 to 3, characterized in that, The scheduling process for each traffic data based on the feature tags of each traffic data set includes: Obtain traffic data scheduling rules; For each of the traffic data sets, a preset feature label that is the same as the feature label of the traffic data set in the traffic data scheduling rule is determined as the target feature label; The preset scheduling link corresponding to the target feature label in the traffic data scheduling rule is determined as the target scheduling link; The traffic data set corresponding to the feature label of the traffic data set is scheduled to the target scheduling link.
9. The method according to claim 8, characterized in that, The method further includes: Obtain the settings information for traffic data scheduling rules; Based on the setting information, multiple preset feature tags and a preset scheduling link corresponding to each preset feature tag are determined. The traffic data scheduling rules are determined based on each preset feature label and the preset scheduling link corresponding to the preset feature label.
10. A flow processing device, characterized in that, The device includes: The acquisition module is used to acquire multiple traffic data sets collected from the network system, extract features from each traffic data set, and obtain feature data corresponding to each traffic data set; wherein, the traffic data set includes multiple traffic data from the same source; The clustering processing module is used to perform clustering processing on the multiple traffic data sets based on the feature data corresponding to each traffic data set, and obtain the clustering result; The determination module is used to determine the feature label of each of the traffic data sets based on the clustering results; The scheduling processing module is used to perform scheduling processing on each of the traffic data sets based on the feature labels of each traffic data set.
11. An electronic device, characterized in that, The electronic device includes: Memory is used to store executable instructions for a computer; A processor, when executing computer-executable instructions stored in the memory, implements the method according to any one of claims 1 to 9.
12. A computer-readable storage medium storing computer-executable instructions or a computer program, characterized in that, When the computer-executable instructions or computer program are executed by a processor, they implement the method described in any one of claims 1 to 9.
13. A computer program product comprising computer-executable instructions or a computer program, characterized in that, When the computer-executable instructions or computer program are executed by a processor, they implement the method according to any one of claims 1 to 9.