Programmable network flow analysis model training method and device, equipment and medium

By constructing an unsupervised training method based on spatiotemporal similarity and topological direction characteristics on programmable network element devices, the problem of insufficient training accuracy of network traffic analysis model under resource constraints is solved, and efficient and accurate traffic analysis is achieved.

CN120378178APending Publication Date: 2025-07-25INFORMATION & COMM BRANCH OF STATE GRID JIANGSU ELECTRIC POWER +2
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510576701.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-06
Publication Date
2025-07-25

AI Technical Summary

Technical Problem

On programmable network element devices with limited resources, how to efficiently and accurately realize traffic analysis, while fully considering the physical location information of network traffic data, the existing methods have problems such as insufficient model training accuracy and neglecting physical location information.

Method used

By obtaining the network traffic data set of network nodes, determining the same network address, determining the spatiotemporal similarity characteristics based on spatial similarity and temporal similarity, determining the topological direction characteristics based on the protocol distribution vector and Wasserstein distance, unsupervised training is performed to obtain the position characteristics of network nodes, and a programmable network traffic analysis model is constructed.

Benefits of technology

It improves the accuracy and flexibility of model training, reduces data redundancy and dependence on labeled data, improves the utilization rate of computing and storage resources, and expands the adaptability of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120378178A_ABST
    Figure CN120378178A_ABST
Patent Text Reader

Abstract

The invention discloses a programmable network traffic analysis model training method and device, equipment and a medium, and the method comprises the steps: determining the same network address of each network traffic data set of each network node, and determining target network traffic data in the corresponding network traffic data set according to the same network address; determining space-time similarity characteristics of the network nodes based on the space similarity and the time similarity between the target network flow data of the different network nodes; determining topological direction characteristics of each network node based on the protocol distribution vectors of different network nodes and the Warisstein distance between the network flow data sets; and determining position features of the corresponding network nodes according to the space-time similarity features and the topological direction features, and performing unsupervised training on the target model based on the position features of the network nodes and the network traffic data set to obtain a programmable network traffic analysis model. According to the invention, by considering the position features among the network nodes, the training precision of the model training method is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of data processing, and in particular, to a method, apparatus, device, and medium for training a programmable network traffic analysis model. Background Art

[0002] With the rapid development of Internet technology, network security issues have become increasingly prominent, and the application of programmable network element devices in network security services has gradually become a research hotspot. These devices can support complex network security tasks, such as encrypted traffic classification and abnormal traffic detection, which are crucial for maintaining the stability and security of the network environment. However, due to the limited on-chip resources of programmable network element devices, deploying efficient machine learning or deep learning models on resource-constrained programmable network elements still faces many challenges.

[0003] Traditional methods usually rely on complex model structures and a large amount of computing resources to achieve high-precision traffic analysis. For example, some methods describe traffic behavior by constructing a knowledge graph, which improves the detection efficiency and accuracy to a certain extent, but faces the problem of high knowledge graph construction costs. Some other methods use the characteristics of programmable switches for two-stage encrypted traffic classification, which improves the classification efficiency and accuracy, but may introduce delays. There are also some methods that use knowledge distillation technology to improve the detection efficiency and accuracy of network traffic. However, the feature extraction ability of knowledge distillation technology is limited, and there are still performance bottlenecks. More critically, traditional network traffic analysis models often ignore the physical location information of network traffic data during the training process, which affects the final accuracy of model training to a certain extent.

[0004] Therefore, how to efficiently and accurately perform traffic analysis on resource-constrained programmable network element devices while fully considering the physical location information of network traffic data has become an urgent problem to be solved in the field of data processing. Summary of the Invention

[0005] The present invention provides a method, apparatus, device, and medium for training a programmable network traffic analysis model, and the present invention improves the training accuracy of the method for training a programmable network traffic analysis model.

[0006] On the one hand, an embodiment of the present invention provides a method for training a programmable network traffic analysis model, including:

[0007] Obtain the network traffic data sets of each network node, determine the same network addresses among the network traffic data sets, and determine the target network traffic data in the corresponding network traffic data sets according to the same network addresses;

[0008] Determine the spatio-temporal similarity features of the network nodes based on the spatial similarity and temporal similarity between the target network traffic data of different network nodes;

[0009] Determine the protocol distribution vector of the corresponding network node according to the traffic protocol and traffic direction in the network traffic dataset, and determine the topological direction characteristics of each network node based on the Wasserstein distance between the protocol distribution vectors of different network nodes and the network traffic dataset;

[0010] Determine the location characteristics of the corresponding network node according to the spatio-temporal similarity characteristics and topological direction characteristics, and perform unsupervised training on the target model based on the location characteristics of each network node and the network traffic dataset to obtain a programmable network traffic analysis model.

[0011] On the other hand, an embodiment of the present invention provides a model training device, including:

[0012] A data determination module, configured to obtain the network traffic datasets of each network node, determine the same network addresses between the network traffic datasets, and determine the target network traffic data in the corresponding network traffic dataset according to the same network address;

[0013] A spatio-temporal feature determination module, configured to determine the spatio-temporal similarity characteristics of the network node based on the spatial similarity and temporal similarity between the target network traffic data of different network nodes;

[0014] A direction feature determination module, configured to determine the distribution vector of the corresponding network node according to the traffic protocol and traffic direction in the network traffic dataset, and determine the topological direction characteristics of each network node based on the Wasserstein distance between the distribution vectors of different network nodes and the network traffic dataset;

[0015] A model acquisition module, configured to determine the location characteristics of the corresponding network node according to the spatio-temporal similarity characteristics and topological direction characteristics, and perform unsupervised training on the target model based on the location characteristics of each network node and the network traffic dataset to obtain a programmable network traffic analysis model.

[0016] On the other hand, an embodiment of the present invention provides a device, including:

[0017] At least one processor;

[0018] And a memory communicatively connected to the at least one processor;

[0019] Wherein, the memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute the programmable network traffic analysis model training method of any embodiment of the present invention.

[0020] On the other hand, an embodiment of the present invention provides a computer-readable medium, including: computer instructions for causing a processor to execute the programmable network traffic analysis model training method according to any embodiment of the present invention.

[0021] In an embodiment of the present invention, a series of network traffic data of each network node can be obtained to form a network traffic data set. Network nodes can be paired two by two as network node pairs. Repeated IP addresses can be determined in the two network traffic data sets of a network node pair. The IP addresses that appear repeatedly in the two network traffic data sets can be used as the same network addresses. Based on the same network addresses, traffic data subsets that interact with the same network addresses can be filtered out from the two network traffic data sets. The filtered traffic data subsets can be used as target network traffic data. Based on the obtained target network traffic data, a spatial similarity used to measure the similarity degree of the content of the target network traffic data can be obtained, and a temporal similarity used to measure the similarity degree of the change trend of the target network traffic data can be obtained. Based on the determined spatial similarity and temporal similarity, a spatio-temporal similarity feature can be jointly determined. In the network traffic data sets corresponding to each network node, the uplink protocol ratio of the number of each type of uplink traffic protocol to the total number of traffic protocols, and the downlink protocol ratio of the number of each type of downlink traffic protocol to the total number of traffic protocols can be determined. The uplink protocol ratios and the downlink protocol ratios can be used as vector elements to construct a protocol distribution vector. The Euclidean distance between the protocol distribution vectors of any two network nodes can be determined according to the Euclidean norm. The Wasserstein distance and the maximum Wasserstein distance between network traffic data sets can be statistically calculated. The ratio of the product of the Euclidean distance and the Wasserstein distance to the maximum Wasserstein distance can be used as a topological direction feature. A target model for analyzing network traffic can be obtained. Based on the obtained spatio-temporal similarity feature and topological direction feature, a position feature used to describe the position of a network node in the network can be determined. Positive and negative samples for training a network traffic analysis model can be constructed based on the network traffic data set. A training objective function of the network traffic analysis model can be constructed based on the positive samples and negative samples. The target model can be trained based on the above position feature and training objective function to obtain a programmable network traffic analysis model.In the embodiments of the present invention, by accurately determining the target network traffic data with the same network address, the process of network traffic analysis can focus on the key network traffic data, reduce data redundancy and unnecessary data processing overhead, and improve the utilization rate of computing and storage resources for the programmable network traffic analysis model training method; this programmable network traffic analysis model training method can infer the location characteristics between network nodes without relying on physical location tags, enabling the training of the network traffic analysis model to be not limited by the acquisition of physical location tags, expanding the adaptability of the programmable network traffic analysis model training method, and enhancing the flexibility of the programmable network traffic analysis model training method; the location characteristics of network nodes can comprehensively reflect the spatio-temporal characteristics and topological structure information of network nodes. By training the network traffic analysis model based on the location characteristics, a more comprehensive feature representation can be provided for the training of the network traffic analysis model, improving the accuracy of the programmable network traffic analysis model training method; through unsupervised training based on the location characteristics of each network node and the network traffic data set, the dependence on a large amount of labeled data can be avoided, reducing the training cost.

[0022] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present invention, nor is it used to limit the scope of the present invention. Other features of the present invention will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0024] Figure 1 is a flowchart of a programmable network traffic analysis model training method provided in Embodiment 1 of the present invention;

[0025] Figure 2 is a flowchart of another programmable network traffic analysis model training method provided in Embodiment 2 of the present invention;

[0026] Figure 3 is a flowchart of a training method for implementing a programmable network traffic analysis model through implicit location modeling provided in Embodiment 3 of the present invention;

[0027] Figure 4 is a flowchart of another programmable network traffic analysis model training method provided in Embodiment 3 of the present invention;

[0028] Figure 5It is a schematic structural diagram of a programmable network traffic analysis model training device provided in Embodiment 4 of the present invention;

[0029] Figure 6 It is a block diagram of a device for executing the programmable network traffic analysis model training method provided in Embodiment 5 of the present invention. Detailed implementation manners

[0030] In order to enable those skilled in the art to better understand the solution of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative work shall fall within the protection scope of the present invention.

[0031] It should be noted that the terms "first", "second", etc. in the specification and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily need to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present invention described here can be implemented in an order different from those illustrated or described here. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device including a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0032] Embodiment 1

[0033] Figure 1 This is a flowchart of a programmable network traffic analysis model training method provided in Embodiment 1 of the present invention. The embodiments of the present invention are applicable to the situation of realizing efficient and accurate traffic analysis on resource-constrained programmable network element devices. This method can be executed by a model training device, which can be implemented in the form of hardware and / or software, and the model training can be configured in the device. As Figure 1 shown, the method includes:

[0034] S110. Obtain the network traffic data sets of each network node, determine the same network addresses among the network traffic data sets, and determine the target network traffic data in the corresponding network traffic data sets according to the same network addresses.

[0035] Among them, the network traffic dataset can be understood as a collection of a series of network traffic data. For example, network traffic data can include data such as IP addresses, protocol types, or port information. It can be understood that the network traffic dataset is the original traffic dataset obtained from each network node, and each original traffic dataset can include network traffic data interacting with different network addresses.

[0036] The same network address can be understood as the IP address that repeatedly appears in each network traffic dataset. The same network address can be used to count the intersection of each network traffic dataset. For example, the same network address can be the same source IP address and / or the same destination IP address.

[0037] The target network traffic data can be understood as a collection of a series of network traffic data associated with the same network address, and can be used to determine spatio-temporal similarity. It can be understood that the target network traffic data is the network traffic data in which any two network nodes interact with the same network address, and is a subset of traffic data filtered from the network traffic dataset based on the same network address.

[0038] Specifically, a series of network traffic data of each network node can be obtained to form a network traffic dataset. The network nodes can be paired two by two as network node pairs. The IP addresses that repeatedly appear can be determined in the two network traffic datasets of the network node pair. The IP addresses that repeatedly appear in the two network traffic datasets can be used as the same network address. The subset of traffic data interacting with the same network address can be filtered from the two network traffic datasets based on the same network address, and the filtered subset of traffic data can be used as the target network traffic data.

[0039] For example, the method of obtaining the network traffic dataset of each network node can include: reading the network traffic data inside each network node through the device interface connecting each network node, or a traffic collection tool can be deployed in each network node, and the network traffic data inside each network node can be obtained by accessing the traffic collection tool. The traffic collection tool can include tools such as Tcpdump or Wireshark.

[0040] For example, the method of determining the IP addresses that repeatedly appear in the two network traffic datasets of the network node pair can include: filtering the network traffic data based on a Bloom filter to determine the IP addresses that repeatedly appear, or directly extracting the IP addresses in the network traffic dataset for comparison and acquisition.

[0041] For example, the arrangement structure of each network node can include: a distributed structure, a hierarchical structure, or a fully connected structure, etc.

[0042] S120. Determine the spatio-temporal similarity features of network nodes based on the spatial similarity and temporal similarity between the target network traffic data of different network nodes.

[0043] Among them, the spatial similarity can be understood as a quantization index, which can be used to measure the similarity degree of the content of the target network traffic data. It can be understood that the greater the spatial similarity, the higher the similarity degree of the content of the target network traffic data, and the closer the physical positions of the network nodes corresponding to the target network traffic data.

[0044] The temporal similarity can be understood as another quantization index, which can be used to measure the similarity degree of the change trend of the target network traffic data. It can be understood that the greater the temporal similarity, the higher the similarity degree of the change trend of the target network traffic data, and the closer the physical positions of the network nodes corresponding to the target network traffic data.

[0045] The spatio-temporal similarity features can be understood as a comprehensive evaluation criterion, which can be used to determine the position features of each network node.

[0046] Specifically, obtain the target network traffic data of each network node. Based on the obtained target network traffic data, the spatial similarity used to measure the similarity degree of the content of the target network traffic data can be obtained, the temporal similarity used to measure the similarity degree of the change trend of the target network traffic data can be obtained, and the spatio-temporal similarity features can be jointly determined based on the determined spatial similarity and temporal similarity.

[0047] Exemplarily, the steps to determine the spatial similarity between the target network traffic data of different network nodes may include:

[0048] S1201. Pair each network node with another as a network node pair;

[0049] S1202. Determine the total number of data fields of all non-repeating data fields in the two target network traffic data of the network node pair;

[0050] S1203. Determine the total number of identical data fields with the same data fields in the two target network traffic data of the network node pair;

[0051] S1204. Take the ratio of the total number of identical data fields to the total number of data fields as the spatial similarity between the two target network traffic data of the network node pair.

[0052] The steps to determine the temporal similarity between the target network traffic data of different network nodes may include:

[0053] S1205. Sort the target network traffic data of each network node according to its respective timestamp to obtain the network traffic time series of each network node;

[0054] S1206. Use the dynamic time warping distance between the network traffic time series determined by the dynamic programming algorithm as the time similarity of each network node.

[0055] The steps to determine the spatio-temporal characteristics between the target network traffic data of different network nodes may include:

[0056] S1207. Use the product of the spatial similarity and the time similarity of each network node as the spatio-temporal similarity feature, or use the weighted product of the spatial similarity and the time similarity of each network node as the spatio-temporal similarity feature.

[0057] S130. Determine the protocol distribution vector of the corresponding network node according to the traffic protocol and traffic direction in the network traffic data set, and determine the topological direction feature of each network node based on the Wasserstein distance between the protocol distribution vectors of different network nodes and the network traffic data set.

[0058] Among them, the traffic protocol refers to the communication protocol used by each network node during the data packet interaction process. For example, the traffic protocol may include: Transmission Control Protocol, Internet Protocol, and / or User Datagram Protocol, etc.

[0059] The traffic direction refers to the transmission direction of the data packet in the network traffic data. For example, the traffic direction may include: uplink traffic and downlink traffic. It can be understood that the uplink traffic is the data traffic that the network node is responsible for sending, and the downlink traffic is the data traffic that the network node is responsible for receiving.

[0060] The protocol distribution vector can be understood as a proportion vector of various traffic protocols, and the protocol distribution vector can be used to describe the distribution of each traffic protocol in the network traffic data set.

[0061] The Wasserstein distance can be understood as a mathematical tool for quantifying the difference between network traffic data sets between network nodes, and can be used to reveal the topological relationship between network nodes.

[0062] The topological direction feature is similar to the spatio-temporal similarity feature, and can also be understood as an evaluation criterion, and can also be used to determine the position feature of each network node.

[0063] Specifically, the uplink protocol ratio of the number of various uplink traffic protocols to the total number of traffic protocols and the downlink protocol ratio of the number of various downlink traffic protocols to the total number of traffic protocols can be determined in the network traffic datasets corresponding to each network node. The uplink protocol ratio and the downlink protocol ratio can be used as vector elements to construct a protocol distribution vector. The Euclidean distance between the protocol distribution vectors of any two network nodes can be determined according to the Euclidean norm. The Wasserstein distance and the maximum Wasserstein distance between network traffic datasets can be statistically calculated. The ratio of the product of the Euclidean distance and the Wasserstein distance to the maximum Wasserstein distance can be used as the topological direction feature.

[0064] Exemplarily, determining the protocol distribution vector of the corresponding network node according to the traffic protocols and traffic directions in the network traffic dataset may include the following steps:

[0065] S1301. Determine the uplink protocol ratio of the number of various uplink traffic protocols to the total number of traffic protocols in the two target network traffic data of the network node pair;

[0066] S1302. Determine the downlink protocol ratio of the number of various downlink traffic protocols to the total number of traffic protocols in the two target network traffic data of the network node pair;

[0067] S1303. Use the uplink protocol ratio and the downlink protocol ratio as vector elements to construct the protocol distribution vector of the corresponding network node.

[0068] Determining the topological direction feature of each network node based on the Wasserstein distance between the protocol distribution vectors of different network nodes and the network traffic dataset may include the following steps:

[0069] S1304. Determine the Wasserstein distance and the maximum Wasserstein distance between the network traffic datasets of any two network nodes;

[0070] S1305. Determine the Euclidean distance between the protocol distribution vectors of any two network nodes, and use the ratio of the product of the Euclidean distance and the Wasserstein distance to the maximum Wasserstein distance as the topological direction feature.

[0071] S140. Determine the position feature of the corresponding network node according to the spatio-temporal similarity feature and the topological direction feature, and perform unsupervised training on the target model based on the position feature of each network node and the network traffic dataset to obtain a programmable network traffic analysis model.

[0072] Among them, the position feature can be understood as a comprehensive feature used to describe the position of a network node in the network. It can be understood that the position feature of each network node takes into account both the physical position relationship of the network node and the similarity in traffic behavior, which can help improve the accuracy and robustness of the network traffic analysis model.

[0073] The target model can be understood as a model for network traffic analysis. For example, the target model can include: a deep learning model, a machine learning model, etc. The uses of the target model can include traffic classification, traffic anomaly detection, etc.

[0074] Specifically, a target model for analyzing network traffic can be obtained, the spatio-temporal similarity features and topological direction features of each network node can be obtained, the position features for describing the position of the network node in the network can be determined based on the obtained spatio-temporal similarity features and topological direction features, the positive samples and negative samples for training the network traffic analysis model can be constructed based on the network traffic dataset, the training objective function of the network traffic analysis model can be constructed based on the positive samples and negative samples, and the target model can be trained based on the above position features and training objective function to obtain a programmable network traffic analysis model.

[0075] For example, the steps of determining the position features of the corresponding network nodes according to the spatio-temporal similarity features and topological direction features can include:

[0076] S1401. Select the larger feature among the spatio-temporal similarity feature and the topological direction feature as the normalization factor;

[0077] S1402. Use the ratio of the product of the spatio-temporal similarity feature and the topological direction feature to the normalization factor as the traffic interaction intensity of the corresponding network node;

[0078] S1403. Use each network node as a vertex, determine the associated edges between the vertices according to the traffic interaction intensity, and construct an implicit position graph of the network nodes based on the vertices and the associated edges;

[0079] S1404. Determine the node degree of the vertex in the implicit position graph, and construct a structured feature vector of the corresponding associated edge with the node degree and the traffic interaction intensity as vector elements;

[0080] S1405. Perform an exponential operation on the structured feature vector processed by the LeakyRelu activation function to obtain an exponentialized feature vector;

[0081] S1406. Determine all the structured feature vectors of all the associated edges connected to the vertex and the corresponding all exponentialized feature vectors, and use the ratio of the exponentialized feature vector to all the exponentialized feature vectors as the attention score of the vertex;

[0082] S1407. Use the Relu activation function to map the attention score to the position feature.

[0083] For example, the unsupervised training of the target model based on the position features of each network node and the network traffic dataset to obtain a programmable network traffic analysis model can include the following steps:

[0084] S1408. Encode all data fields included in each network traffic dataset to obtain traffic features. The data fields at least include: source IP address, destination IP address, port information, protocol type, and timestamp. The encoding formats for encoding the data fields at least include: hash encoding and One-Hot encoding;

[0085] S1409. Perform an outer product operation on the location features and the traffic features to obtain the traffic location cross features corresponding to each network node;

[0086] S140a. Construct a fusion feature vector with the location features, the traffic features, and the traffic location cross features as vector elements;

[0087] S140b. Perform unsupervised training on the deep learning architecture based on the fusion feature vector to obtain a programmable network traffic analysis model.

[0088] In an embodiment of the present invention, a series of network traffic data of each network node can be obtained to form a network traffic data set. Network nodes can be paired two by two as network node pairs. Repeated IP addresses can be determined in the two network traffic data sets of a network node pair. The IP addresses that appear repeatedly in the two network traffic data sets can be used as the same network addresses. Based on the same network addresses, subsets of traffic data that interact with the same network addresses can be filtered out from the two network traffic data sets. The filtered subsets of traffic data can be used as target network traffic data. Based on the obtained target network traffic data, a spatial similarity used to measure the similarity degree of the content of the target network traffic data can be obtained, and a temporal similarity used to measure the similarity degree of the change trend of the target network traffic data can be obtained. Based on the determined spatial similarity and temporal similarity, spatio-temporal similarity features can be jointly determined. In the network traffic data sets corresponding to each network node, the uplink protocol ratio of the number of each type of uplink traffic protocol to the total number of traffic protocols, and the downlink protocol ratio of the number of each type of downlink traffic protocol to the total number of traffic protocols can be determined. The uplink protocol ratios and the downlink protocol ratios can be used as vector elements to construct a protocol distribution vector. The Euclidean distance between the protocol distribution vectors of any two network nodes can be determined according to the Euclidean norm. The Wasserstein distance and the maximum Wasserstein distance between network traffic data sets can be statistically calculated. The ratio of the product of the Euclidean distance and the Wasserstein distance to the maximum Wasserstein distance can be used as a topological direction feature. A target model for analyzing network traffic can be obtained. Based on the obtained spatio-temporal similarity features and topological direction features, position features used to describe the positions of network nodes in the network can be determined. Positive samples and negative samples for training a network traffic analysis model can be constructed based on the network traffic data set. Based on the positive samples and negative samples, a training objective function of the network traffic analysis model can be constructed. Based on the above position features and the training objective function, the target model can be trained to obtain a programmable network traffic analysis model.

[0089] In the embodiments of the present invention, by accurately determining the target network traffic data with the same network address, the process of network traffic analysis can focus on the key network traffic data, reducing data redundancy and unnecessary data processing overhead, and improving the utilization rate of computing and storage resources for the programmable network traffic analysis model training method; this programmable network traffic analysis model training method can infer the location characteristics between network nodes without relying on physical location tags, enabling the training of the network traffic analysis model to be unrestricted by the acquisition of physical location tags, expanding the adaptability of the programmable network traffic analysis model training method, and enhancing the flexibility of the programmable network traffic analysis model training method; the location characteristics of network nodes can comprehensively reflect the spatio-temporal characteristics and topological structure information of network nodes. By training the network traffic analysis model based on the location characteristics, a more comprehensive feature representation can be provided for the training of the network traffic analysis model, improving the accuracy of the programmable network traffic analysis model training method; through unsupervised training based on the location characteristics of each network node and the network traffic data set, the dependence on a large number of labeled data can be avoided, reducing the training cost.

[0090] Embodiment 2

[0091] Figure 2 As shown in the flowchart of another programmable network traffic analysis model training method provided by Embodiment 2 of the present invention, this embodiment of the present invention is a refinement of the above embodiment. Specifically, it refines the specific steps of how to determine the spatio-temporal similarity features of network nodes based on the spatial similarity and temporal similarity between the target network traffic data of different network nodes, refines the specific steps of how to determine the protocol distribution vector of the corresponding network nodes according to the traffic protocol and traffic direction in the network traffic data set, refines the specific steps of how to determine the topological direction features of each network node based on the Wasserstein distance between the protocol distribution vectors of different network nodes and the network traffic data set, refines the specific steps of how to determine the location features of the corresponding network nodes according to the spatio-temporal similarity features and topological direction features, and refines the specific steps of how to train and obtain the programmable network traffic analysis model.

[0092] As Figure 2 shown, another programmable network traffic analysis model training method may include the following steps:

[0093] S201. Obtain the network traffic data sets of each network node, determine the same network address between the network traffic data sets, and determine the target network traffic data in the corresponding network traffic data set according to the same network address.

[0094] Exemplarily, the network traffic data set D i of network node i and the network traffic data set D j of network node j can be obtained. Based on D i and Dj For the same network address among them, the target network traffic data of network node i can be determined as S i , and the target network traffic data of network node j is S j .

[0095] S202. Pair up each network node to form network node pairs.

[0096] Specifically, any two network nodes can be randomly selected from all the network nodes for pairing to form network node pairs.

[0097] It can be understood that the set of network node pairs formed by network node pairs can also include the network node pair formed by network node i and itself.

[0098] S203. Determine the total number of field types of all non-repeated data fields within the two target network traffic data of the network node pair.

[0099] Among them, a data field can be understood as the basic unit constituting a data set. By way of example, data fields can include fields such as IP address, protocol type, or port information, etc.

[0100] The total number of field types refers to the number of all non-repeated data fields in the two target network traffic data, and can be used to determine the spatial similarity between the two target network traffic data.

[0101] Specifically, the two target network traffic data of the network node pair can be obtained, the two obtained target network traffic data can be combined to form a target network traffic data set of the network node pair, the non-repeated data fields can be counted in the target network traffic data set, and the first number of non-repeated data fields can be recorded, and this first number can be used as the total number of field types of the network node pair.

[0102] By way of example, the representation method of the total number of field types of all non-repeated data fields within each target network traffic data can include: |S i ∪S j |, where |S i ∪S j | represents the scale of the union of the target network traffic data S i and the target network traffic data S j .

[0103] S204. Determine the total number of identical field types of the same data fields within the two target network traffic data of the network node pair.

[0104] Among them, the total number of identical field types refers to the number of all identical data fields in the two target network traffic data, and can be used to determine the spatial similarity between the two target network traffic data.

[0105] Specifically, two target network traffic data of a network node pair can be obtained, and the two obtained target network traffic data can be combined to form a target network traffic data set of a network node pair. The same data fields can be counted in the target network traffic data set, and the second quantity of the same data fields can be recorded. The second quantity can be used as the total number of the same fields of the network node pair.

[0106] Exemplarily, the representation of the total number of the same fields of the same data fields in each target network traffic data may include: |S i ∩S j |, where |S i ∩S j | represents the scale of the intersection of the target network traffic data S i and the target network traffic data S j .

[0107] S205. Use the ratio of the total number of the same fields to the total number of fields as the spatial similarity between the two target network traffic data of the network node pair.

[0108] Specifically, the total number of fields and the total number of the same fields of all non-repeated data fields can be determined in the target network traffic data of the network node pair. The total number of the same fields can be compared with the total number of fields to obtain the ratio between the total number of the same fields and the total number of fields. The ratio can be used as the spatial similarity between the two target network traffic data.

[0109] Exemplarily, the spatial similarity between the two target network traffic data can be obtained through .

[0110] S206. Sort the target network traffic data of each network node according to its respective timestamp to obtain the network traffic time series of each network node.

[0111] Among them, the network traffic time series can be understood as a data series, which is obtained by sorting the target network traffic data of each network node according to its respective timestamp. The network traffic time series can reflect the change of network traffic data over time.

[0112] Specifically, the timestamp for data interaction of the target network traffic data can be obtained in the network traffic data of each network node. Based on the obtained timestamp, the data included in the target network traffic data can be sorted to form a data series, and the data series can be used as the network traffic time series of each network node.

[0113] Exemplarily, the network traffic time series of network node i and network node j can be respectively expressed as T i (S i ∩Sj ) and T j (S i ∩S j )。

[0114] S207. Use the dynamic time warping distance between the network traffic time series determined by the dynamic programming algorithm as the time similarity of each network node.

[0115] Among them, the dynamic programming algorithm is an algorithm for measuring the similarity between time series. The dynamic programming algorithm can determine the dynamic time warping distance between two time series by finding the path of the minimum cumulative distance between the two time series.

[0116] The dynamic time warping distance can be understood as another quantization index. The smaller the dynamic time warping distance, the higher the time similarity degree between the target network traffic data.

[0117] Specifically, the dynamic programming algorithm and the network traffic time series of each network node can be obtained. The dynamic programming algorithm obtained can be used to process the network traffic time series to obtain the distances between the network traffic time series. The cumulative distances between the network traffic time series can be counted. The minimum cumulative distance can be used as the dynamic time warping distance between two network traffic time series, and this dynamic time warping distance can be used as the time similarity of each network node.

[0118] For example, the dynamic time warping distance between the network traffic time series corresponding to network node i and network node j can be obtained through DTW(T i (S i ∩S j ),T j (S i ∩S j ).

[0119] S208. Use the product of the spatial similarity and the time similarity of each network node as the spatio-temporal similarity feature.

[0120] Specifically, the spatial similarity and the time similarity of each network node can be obtained. The spatial similarity and the time similarity of each network node obtained can be multiplied to obtain the multiplication result, and this multiplication result can be used as the spatio-temporal similarity feature of each network node.

[0121] For example, the spatio-temporal similarity feature of network node i can be obtained through .

[0122] S209. Determine the uplink protocol ratio of the number of various uplink traffic protocols in the two target network traffic data of the network node pair to the total number of all traffic protocols.

[0123] Among them, the proportion of uplink protocols can be understood as another quantization metric, which refers to the proportion of the number of various uplink traffic protocols to the number of all traffic protocols. The proportion of the number of various uplink traffic protocols to the number of all traffic protocols can be used as vector elements to construct a protocol distribution vector.

[0124] Specifically, two target network traffic data of a network node pair can be obtained. The two obtained target network traffic data can be combined to form a target network traffic data set of a network node pair. The number of various uplink traffic protocols and the number of all traffic protocols can be counted in this target network traffic data set. The number of various uplink traffic protocols can be compared with the number of all traffic protocols respectively to obtain the proportion of the number of various uplink traffic protocols to the number of all traffic protocols. This proportion can be used as the uplink protocol proportion corresponding to various uplink traffic protocols. It can be understood that this uplink protocol proportion can be used as a vector element to construct a protocol distribution vector.

[0125] S210. Determine the downlink protocol proportion of the number of various downlink traffic protocols to the number of all traffic protocols in the two target network traffic data of the network node pair.

[0126] Among them, the downlink protocol proportion can be understood as another quantization metric, which refers to the proportion of the number of various downlink traffic protocols to the number of all traffic protocols. The proportion of the number of various downlink traffic protocols to the number of all traffic protocols can be used as vector elements to construct a protocol distribution vector.

[0127] Specifically, two target network traffic data of a network node pair can be obtained. The two obtained target network traffic data can be combined to form a target network traffic data set of a network node pair. The number of various downlink traffic protocols and the number of all traffic protocols can be counted in this target network traffic data set. The number of various downlink traffic protocols can be compared with the number of all traffic protocols respectively to obtain the proportion of the number of various downlink traffic protocols to the number of all traffic protocols. This proportion can be used as the downlink protocol proportion corresponding to various downlink traffic protocols. It can be understood that this downlink protocol proportion can be used as a vector element to construct a protocol distribution vector.

[0128] S211. Use each uplink protocol proportion and each downlink protocol proportion as vector elements to construct a protocol distribution vector for the corresponding network node.

[0129] Specifically, the uplink protocol proportion corresponding to various uplink traffic protocols and the downlink protocol proportion corresponding to various downlink traffic protocols can be obtained. The obtained uplink protocol proportions and downlink protocol proportions can be used as vector elements, and the vector elements can be arranged to form a protocol distribution vector.

[0130] Exemplarily, the expression of the protocol distribution vector can be Among them, represents the ratio of the number of the first type of uplink traffic protocols in network node i to the total number of traffic protocols. represents the ratio of the number of the first type of downlink traffic protocols in network node i to the total number of traffic protocols.

[0131] S212. Determine the Wasserstein distance and the maximum Wasserstein distance between the network traffic data sets of any two network nodes.

[0132] Specifically, the network traffic data sets of any two network nodes can be obtained, and the Wasserstein distance between the network traffic data sets corresponding to any two network nodes can be determined through the Wasserstein distance algorithm, and the maximum Wasserstein distance can be determined from the determined Wasserstein distances.

[0133] Exemplarily, the network traffic data sets D i and D j corresponding to network node i and network node j, respectively, the Wasserstein distance between them can be expressed as Wasserstein(D i , D j ), and the maximum Wasserstein distance between D i and D j can be expressed as max k,l Wasserstein(D k , D l ), where k is not equal to l.

[0134] S213. Determine the Euclidean distance between the protocol distribution vectors of any two network nodes, and use the ratio of the product of the Euclidean distance and the Wasserstein distance to the maximum Wasserstein distance as the topological direction feature.

[0135] Specifically, the Euclidean distance between the protocol distribution vectors of any two network nodes can be determined according to the Euclidean norm. The product of the Euclidean distance and the Wasserstein distance can be calculated, and the ratio of the product of the Euclidean distance and the Wasserstein distance to the maximum Wasserstein distance can be obtained by comparison, and this ratio can be used as the topological direction feature of each network node.

[0136] Exemplarily, the Euclidean distance between the protocol distribution vectors of any two network nodes can be obtained according to the formula ||d i -d j ||2, and the topological direction feature of each network node can be obtained according to .

[0137] S214. Select the larger feature between the spatio-temporal similarity feature and the topological direction feature as the normalization factor.

[0138] Among them, the normalization factor can be understood as a scaling coefficient, which can be used to adjust the numerical range of the traffic interaction intensity.

[0139] Specifically, the spatio-temporal similarity feature and the topological direction feature can be obtained, the obtained spatio-temporal similarity feature and topological direction feature can be compared, and the larger feature among the spatio-temporal similarity feature and the topological direction feature can be selected as the normalization factor for adjusting the numerical range of the traffic interaction intensity.

[0140] Exemplarily, the normalization factor can be expressed as

[0141] S215. Use the ratio of the product of the spatio-temporal similarity feature and the topological direction feature to the normalization factor as the traffic interaction intensity of the corresponding network node.

[0142] Among them, the traffic interaction intensity can be understood as a numerical index, which can be used to quantify the closeness of the traffic interaction between two network nodes.

[0143] Specifically, the normalization factor can be obtained, the spatio-temporal similarity feature and the topological direction feature can be multiplied to obtain the product result of the spatio-temporal similarity feature and the topological direction feature, the ratio of the product result to the normalization factor can be obtained by comparing the product result with the normalization factor, and this ratio can be used as the traffic interaction intensity of the network node.

[0144] Exemplarily, the traffic interaction intensity of each network node can be expressed as

[0145] S216. Use each network node as a vertex, determine the associated edges between the vertices according to the traffic interaction intensity, and construct an implicit position graph of the network nodes based on the vertices and the associated edges.

[0146] Among them, the implicit position graph can be understood as a graph structure, which can be used to represent the relative position relationship between each network node. It can be understood that in the implicit position graph, the positions of the network nodes are implicitly represented by their associated edges with other nodes.

[0147] Specifically, each network node can be used as a vertex, the traffic interaction intensity between each network node can be used as the associated edge between the vertices, and an implicit position graph representing the topological relationship between each network node can be constructed based on the vertices and the associated edges between the vertices.

[0148] Exemplarily, a preset threshold can be obtained, the traffic interaction intensity can be compared with the above preset threshold, only the strong associated edges with the traffic interaction intensity greater than the preset threshold can be retained in the implicit position graph, and only the structured feature vectors of the strong associated edges can be determined.

[0149] S217. Determine the node degree of vertices in the implicit position graph, and construct a structured feature vector corresponding to the associated edge with the node degree and the traffic interaction intensity as vector elements.

[0150] Among them, the node degree refers to the number of associated edges directly connected to each vertex in the implicit position graph. For example, the node degree can include 2 or 3.

[0151] The structured feature vector can be understood as a vector used to represent the edge features in the implicit position graph. For example, the structured feature vector includes at least the node degree vector element and the traffic interaction intensity vector element.

[0152] Specifically, the traffic interaction intensity of each vertex can be obtained, the number of associated edges directly connected to each vertex can be determined in the implicit position graph, this number can be used as the node degree of each vertex, the obtained traffic interaction intensity and node degree can be used as vector elements, and the traffic interaction intensity vector element and the node degree vector element can be arranged to form a structured feature vector for representing the edge features in the implicit position graph.

[0153] For example, the structured feature vector φ of the associated edge ij can also be expressed as φ ij = [T ij , log(d i + 1), log(d j + 1)] T .

[0154] S218. Perform an exponential operation on the structured feature vector processed by the LeakyRelu activation function to obtain an exponentialized feature vector.

[0155] Among them, the exponentialized feature vector can be understood as a feature vector after an exponential operation, which can be used to determine the attention score of each vertex.

[0156] Specifically, the structured feature vector can be used as the input of the LeakyRelu activation function. The LeakyRelu activation function can process each vector element in the structured feature vector according to the definition of the LeakyRelu activation function. The feature vector processed by the LeakyRelu activation function can be used as an exponential operation, and the vector elements in this feature vector can be exponentiated to obtain the output vector of the exponential operation. This output vector can be used as the exponentialized feature vector.

[0157] For example, the exponentialized feature vector can be expressed as exp(LeakyReLU(a T φ ij )), where a is a learnable parameter.

[0158] S219: Determine all the indexed feature vectors corresponding to all the associated edges connecting each vertex, and use the ratio of the indexed feature vector to all the indexed feature vectors as the attention score of the vertex.

[0159] Among them, the total indexed feature vector can be understood as a new feature vector obtained by adding the indexed feature vectors, and the total indexed feature vector can be used to determine the attention score of each vertex.

[0160] The attention score can be understood as a quantitative value that can be used to measure the importance of the associated edge.

[0161] Specifically, the exponential feature vector corresponding to each associated edge directly connected to each vertex can be obtained, the exponential feature vector corresponding to each associated edge can be added to obtain all the exponential feature vectors, the exponential feature vector corresponding to each vertex can be compared with all the exponential feature vectors to obtain a ratio, which can be used as the attention score of the vertex.

[0162] For example, all exponential eigenvectors can be expressed as The attention score can be expressed as

[0163] S220, using the Relu activation function to map the attention score to a position feature.

[0164] Specifically, the attention score can be used as the input of the Relu activation function, the attention score can be processed according to the definition of the Relu activation function, and the output of the Relu activation function can be used as the position feature.

[0165] For example, after obtaining the attention score, the following steps may be included: the above attention score may be processed using the Elu activation function to obtain an implicit position feature z can be transformed by Relu activation function i Represented as position feature v i =ReLU(Wz i +b), where w∈R is a learnable scalar parameter, W∈R is a position weight matrix, and b is a learnable parameter. The learnable scalar parameter, the position weight matrix, and the learnable parameter can be automatically optimized during the model training process.

[0166] S221. Encode all data fields contained in each network traffic data set to obtain traffic characteristics.

[0167] The data field includes at least: source IP address, destination IP address, port information, protocol type and timestamp, and the encoding format for encoding the data field includes at least: hash encoding and One-Hot encoding.

[0168] Among them, the traffic feature can be understood as a numerical feature and can be used to construct a fused feature vector.

[0169] Specifically, all data fields can be obtained from the network traffic datasets corresponding to each network node, all the obtained data fields can be encoded, vectors with fixed lengths corresponding to the respective data fields can be obtained, and the obtained vectors with fixed lengths can be used as vector elements to construct the traffic feature.

[0170] S222. Perform an outer product operation on the location feature and the traffic feature to obtain the traffic-location cross feature corresponding to each network node.

[0171] Among them, the traffic-location cross feature is obtained by performing an outer product operation on the location feature and the traffic feature and can be used to construct a fused feature vector.

[0172] Specifically, the location feature and the traffic feature can be obtained, the location feature and the traffic feature can be subjected to an outer product operation, a new feature can be obtained by processing the elements of the location feature and the elements of the traffic feature according to the definition of the outer product operation, and the new feature can be used as the traffic-location cross feature.

[0173] S223. Use the location feature, the traffic feature, and the traffic-location cross feature as vector elements to construct a fused feature vector.

[0174] Among them, the fused feature vector can be understood as a combined vector and can be used to obtain a network traffic analysis model. For example, the vector elements included in the fused feature vector at least include: location feature vector elements, traffic feature vector elements, and traffic-location cross feature vector elements.

[0175] Specifically, the location feature, the traffic feature, and the traffic-location cross feature can be obtained, and the vector formed by arranging the location feature vector elements, the traffic feature vector elements, and the traffic-location cross feature vector elements can be used as the fused feature vector.

[0176] S224. Perform unsupervised training on the deep learning architecture based on the fused feature vector to obtain a programmable network traffic analysis model.

[0177] Specifically, the fused feature vector and the deep learning architecture can be obtained, the fused feature vector can be input into the deep learning architecture, the deep learning architecture can be trained using the fused feature vector, and the trained deep learning architecture can be used as the programmable network traffic analysis model.

[0178] In an embodiment of the present invention, a network traffic dataset of each network node can be obtained, the same network addresses among the network traffic datasets can be determined, and target network traffic data can be determined within the corresponding network traffic dataset according to the same network address. Any two network nodes can be randomly selected from all the network nodes for matching to form a network node pair. The two target network traffic data of the network node pair can be obtained, and the two obtained target network traffic data can be combined to form a target network traffic dataset of the network node pair. The non-repeated data fields can be counted in the target network traffic dataset, and the first quantity of the non-repeated data fields can be recorded. The first quantity can be used as the total number of fields of the network node pair. The same data fields can be counted in the target network traffic dataset, and the second quantity of the same data fields can be recorded. The second quantity can be used as the total number of same fields of the network node pair. The total number of same fields can be compared with the total number of fields to obtain the ratio between the total number of same fields and the total number of fields. The ratio can be used as the spatial similarity between the two target network traffic data. The timestamps of the data interaction of the target network traffic data can be obtained from the network traffic data of each network node. Based on the obtained timestamps, the data included in the target network traffic data can be sorted to form a data sequence. The data sequence can be used as the network traffic time sequence of each network node. The dynamic time warping algorithm and the network traffic time sequences of each network node can be obtained. The obtained dynamic time warping algorithm can be used to process the network traffic time sequences to obtain the distances between the network traffic time sequences. The cumulative distances between the network traffic time sequences can be counted. The minimum cumulative distance can be used as the dynamic time warping distance between the two network traffic time sequences. The dynamic time warping distance can be used as the time similarity of each network node. The spatial similarity and the time similarity of each network node obtained can be multiplied to obtain a multiplication result. The multiplication result can be used as the spatio-temporal similarity feature of each network node. The quantities of various types of upstream traffic protocols and the quantity of all traffic protocols can be counted in the target network traffic dataset. The quantities of various types of upstream traffic protocols can be respectively compared with the quantity of all traffic protocols to obtain the proportion of the quantity of each type of upstream traffic protocol in the quantity of all traffic protocols. The proportion can be used as the upstream protocol proportion corresponding to each type of upstream traffic protocol. The quantities of various types of downstream traffic protocols and the quantity of all traffic protocols can be counted in the target network traffic dataset. The quantities of various types of downstream traffic protocols can be respectively compared with the quantity of all traffic protocols to obtain the proportion of the quantity of each type of downstream traffic protocol in the quantity of all traffic protocols. The proportion can be used as the downstream protocol proportion corresponding to each type of downstream traffic protocol. The quantities of various types of downstream traffic protocols and the quantity of all traffic protocols can be counted in the target network traffic dataset. The quantities of various types of downstream traffic protocols can be respectively compared with the quantity of all traffic protocols to obtain the proportion of the quantity of each type of downstream traffic protocol in the quantity of all traffic protocols. The proportion can be used as the downstream protocol proportion corresponding to each type of downstream traffic protocol.A network traffic dataset of any two network nodes can be obtained. The Wasserstein distance between the network traffic datasets corresponding to any two network nodes can be determined by the Wasserstein distance algorithm, and the maximum Wasserstein distance can be determined from the determined Wasserstein distances. The Euclidean distance between the protocol distribution vectors of any two network nodes can be determined according to the Euclidean norm. The Euclidean distance can be multiplied by the Wasserstein distance. The product of the multiplication of the Euclidean distance and the Wasserstein distance can be compared with the maximum Wasserstein distance to obtain the ratio of the product to the maximum Wasserstein distance. This ratio can be used as the topological direction feature of each network node. The obtained spatio-temporal similarity feature and topological direction feature can be compared, and the larger feature of the spatio-temporal similarity feature and topological direction feature can be selected as the normalization factor for adjusting the numerical range of the traffic interaction intensity. The spatio-temporal similarity feature and topological direction feature can be multiplied to obtain the product of the spatio-temporal similarity feature and topological direction feature. The product can be compared with the normalization factor to obtain the ratio of the product to the normalization factor. This ratio can be used as the traffic interaction intensity of the network node. Each network node can be used as a vertex, and the traffic interaction intensity between each network node can be used as the associated edge between each vertex. An implicit position map representing the topological relationship between each network node can be constructed based on the vertices and the associated edges between the vertices. The number of associated edges directly connected to each vertex can be determined in the implicit position map, and this number can be used as the node degree of each vertex. The obtained traffic interaction intensity and node degree can be used as vector elements. The traffic interaction intensity vector elements and node degree vector elements can be arranged to form a structured feature vector representing the edge features in the implicit position map. The structured feature vector can be used as the input of the LeakyRelu activation function. The LeakyRelu activation function can process each vector element in the structured feature vector according to the definition of the LeakyRelu activation function. The feature vector processed by the LeakyRelu activation function can be used for exponential operation. The vector elements in the feature vector can be subjected to exponential operation to obtain the output vector of the exponential operation. This output vector can be used as the exponentialized feature vector. The exponentialized feature vector corresponding to each associated edge directly connected to each vertex can be obtained. The exponentialized feature vectors corresponding to each associated edge can be added to obtain all the exponentialized feature vectors. The exponentialized feature vectors corresponding to each vertex and all the exponentialized feature vectors can be compared to obtain a ratio. This ratio can be used as the attention score of the vertex. The attention score can be used as the input of the Relu activation function. The attention score can be processed according to the definition of the Relu activation function. The output of the Relu activation function can be used as the position feature. All data fields can be obtained from the network traffic datasets corresponding to each network node. The obtained all data fields can be encoded to obtain vectors with fixed lengths corresponding to each data field. The obtained vectors with fixed lengths can be used as vector elements to construct the traffic feature.The outer product operation can be performed on the location feature and the traffic feature. According to the definition of the outer product operation, the elements of the location feature and the traffic feature can be processed to obtain a new feature. This new feature can be used as the traffic-location cross feature. The vector formed by arranging the elements of the location feature vector, the elements of the traffic feature vector, and the elements of the traffic-location cross feature vector can be used as the fusion feature vector. A deep learning architecture can be obtained, and this fusion feature vector can be input into the deep learning architecture to train the deep learning architecture using this fusion feature vector. The trained deep learning architecture can be used as a programmable network traffic analysis model. In the embodiments of the present invention, by quantifying the total number of fields and the total number of identical fields in the target network traffic data, the content similarity between network traffic data can be evaluated more precisely, avoiding the operation of comparing one by one and improving the data processing efficiency; by using the dynamic time warping algorithm to calculate the time similarity between network traffic data, considering the time series characteristics of data interaction, the accuracy of traffic analysis is improved; through the product operation, the spatial similarity and time similarity in two dimensions are fused into one feature, simplifying the subsequent processing flow and thus improving the efficiency of network traffic analysis; by introducing the topological direction feature combining the Wasserstein distance and the Euclidean distance, the model can obtain richer feature information during the training process, which can improve the stability of model training; by constructing an implicit location map, the location relationship between each network node is implicitly represented by the traffic interaction intensity between each network node, so that the training method of the programmable network traffic analysis model can consider the location information of each network node during the training process of the model without relying on physical location tags, improving the accuracy of the training method of the programmable network traffic analysis model; through the calculation of attention scores, the model can focus on important network nodes, improving the accuracy of traffic analysis; using the deep learning architecture for model training improves the generalization ability and accuracy of the model.,

[0179] Embodiment III

[0180] Based on the above embodiments, the embodiments of the present invention provide a training method for a programmable network traffic analysis model through implicit location modeling, as Figure 3 shown. This method may include the following processes:

[0181] The network traffic data sets of each network node can be collected. Among them, the original data fields in each network traffic data set can be as shown in Table 1:

[0182] Table 1 Data Fields and Meanings of Data Fields

[0183]

[0184] The data fields in each network traffic dataset can be preprocessed, that is, encoded to obtain the traffic characteristics of each network node. As shown in Table 2, the data fields IP address and IP port can be mapped to fixed-length vectors through hash encoding, the data field protocol type can be one-hot encoded to obtain another fixed-length vector, the traffic size and timestamp can be directly used as numerical features, and the above fixed-length vectors and numerical features can be used as vector elements to construct the traffic characteristics u of each network node ij 。

[0185] Table 2 Data fields and corresponding encoding formats

[0186]

[0187] The spatio-temporal similarity weights between the network traffic datasets corresponding to any two network nodes can be calculated Among them, S i and S j are the interaction datasets with the same IP address of the network traffic dataset i corresponding to network node i and the network traffic dataset j corresponding to network node j respectively. |S i ∩S j | is the total number of identical fields of the same data fields in the interaction datasets corresponding to network node i and network node j respectively. |S i ∪S j | is the total number of fields of all non-repeating data fields in the interaction datasets corresponding to network node i and network node j respectively. T i (S i ∩S j ) is the time series obtained by sorting the interaction dataset according to the timestamp of the data field. DTW(.) is the Dynamic Time Warping algorithm, which is used to calculate the dynamic time warping distance between two time series

[0188] The topological direction distance between the network traffic datasets corresponding to any two network nodes can be calculated Among them, d i is the protocol distribution vector of network node i, D iDenote the network traffic dataset of network node \(i\). \(\|\cdot\|_2\) is used to calculate the Euclidean distance between two protocol distribution vectors, and Wasserstein\((\cdot,\cdot)\) represents calculating the Wasserstein distance between two network traffic datasets. Among them, the method for obtaining the protocol distribution vector is as follows: In the two target network traffic data of the network node pair, count the uplink protocol ratio of the number of each type of uplink traffic protocol to the total number of all traffic protocols and the downlink protocol ratio of the number of each type of downlink traffic protocol to the total number of all traffic protocols. The uplink protocol ratios and the downlink protocol ratios can be used as vector elements to construct the protocol distribution vector \(d\) corresponding to the network node i :

[0189] The traffic interaction intensity \(T\) between network node \(i\) and network node \(j\) can be calculated through where \(Z\) is a normalization factor ij Each network node can be regarded as a vertex. According to the traffic interaction intensity, determine the existence of associated edges between the vertices, and construct the implicit position graph of the network node based on the vertices and the associated edges By setting a threshold \(\tau\), weak associated edges in the implicit position graph can be filtered, and only the edges with \(T ij >\(\tau\) are retained For the implicit position graph \(G f In it, for the edge The structured feature vector \(\varphi i,j of edge \(e ij =[T ij ,\log(c + 1),\log(c + 1)] T can be extracted, where \(c\) is the degree of network node \(i\). A learnable parameter is defined to calculate the attention score of edge \(e i,j where LeakyReLU is an activation function. A learnable scalar parameter \(w\in\mathbb{R}\) is defined. According to the Elu activation function, \(w\) and \(\alpha ij are processed to generate the position feature embedding vector The position feature \(v i can be mapped according to the Relu activation function for \(z i = ReLU(Wz i + b), where is the position weight matrix, \(d u is the traffic feature dimension, \(d z is the implicit position feature dimension, and \(b\) is a learnable parameter

[0190] The traffic feature \(u ij can be combined with the position feature \(v i ​​Mapped to the same dimension: u' ij = W u u ij + b u , v' i = W v v i + b v , where b u , b v ∈ R d is a learnable parameter. Then, explicitly construct the flow-location cross-feature matrix through the outer product operation: It can be flattened into a vector C ij ∈ R d×d and compressed through a fully connected layer: c' ij = ReLU(W c c ij + b c ). The flow features, location features, and cross-features can be concatenated into a fusion vector w ij = [u' ij ; v' i ; c' ij , Finally, high-dimensional features can be extracted through a multi-layer perceptron layer: h ij = MLP(w ij ).

[0191] To improve the robustness and spatial consistency of feature representation, a contrastive learning mechanism can be designed. First, construct positive samples: Apply random masking or temporal perturbation to the flow data d ij to generate enhanced samples Negative samples: Randomly select flow data from other nodes Then, perform a location-weighted similarity measurement and define a location-aware similarity function: sim loc (h ij , h kj ) = sim(h ij , h kj ) · exp(-γ|z i - z k | 2 ), where γ is the control parameter for the decay rate of the location weight, and |z i - z k | 2 is the Euclidean distance between the implicit location embeddings of node i and node k. The normalized temperature-scaled contrastive loss can be calculated: where τ is the temperature parameter. The training objective of the network traffic analysis model can be to minimize the composite loss function: where is the task-related loss (such as categorical cross-entropy), is the implicit position-related loss, and λ is the balance coefficient.

[0192] The global neural network consists of a shared feature layer and a task-specific output layer, which realizes distributed feature aggregation and knowledge transfer. It can be understood that the global neural network can be used as a network traffic analysis model. In shared feature extraction, the encoded features h of all nodes ij Input to the shared MLP layer: h shared = MLP shared (h ij ), and then output according to the specific task. For the traffic classification problem: Node i uses the softmax output layer: For the anomaly detection problem: Node i uses the sigmoid output layer:

[0193] The goal of the network traffic analysis model is to optimize the task-related loss and contrastive learning loss of all nodes respectively through a multi-task learning architecture, and adopt an alternating optimization mechanism with two ways of local parameter update and global parameter update: Local parameter update: Fix the global network g and update the encoders of each node where η is the learning rate and θ is the parameter of the encoder f i The parameter of. Global parameter update: Fix the local encoder and update the parameters of the global network to minimize the task loss where φ is the parameter of the global network g.

[0194] Based on the above embodiments, the embodiment of the present invention provides another flowchart of a programmable network traffic analysis model training method, as Figure 4 shown. The method includes collecting network traffic data sets of N network nodes (DataSet). Each node can independently model local traffic data and can enhance the sensitivity to local traffic patterns by fusing implicit position embedding features; An implicit position graph (Graph) can be constructed based on the N network traffic data sets, and node implicit position information can be introduced into the N encoders (Encoder) corresponding to the N network nodes based on the implicit position graph to construct a contrastive loss function, which can improve the discriminability and spatial consistency of feature representation. By sharing the global neural network and the task-specific output layer, the balance between local feature learning and global knowledge transfer is achieved.

[0195] Embodiment Four

[0196] Figure 5 A programmable network traffic analysis model training device provided in Embodiment Five of the present invention, as Figure 5 shown. The device includes:

[0197] A data determination module 310, configured to obtain network traffic data sets of each network node, determine the same network addresses among the network traffic data sets, and determine target network traffic data within the corresponding network traffic data sets according to the same network addresses;

[0198] A spatio-temporal feature determination module 320, configured to determine spatio-temporal similarity features of network nodes based on spatial similarity and temporal similarity between target network traffic data of different network nodes;

[0199] A direction feature determination module 330, configured to determine a distribution vector of a corresponding network node according to traffic protocols and traffic directions in a network traffic data set, and determine topological direction features of each network node based on the Wasserstein distance between distribution vectors of different network nodes and between network traffic data sets;

[0200] A model acquisition module 340, configured to determine position features of corresponding network nodes according to spatio-temporal similarity features and topological direction features, and perform unsupervised training on a target model based on the position features of each network node and network traffic data sets to obtain a programmable network traffic analysis model.

[0201] In the embodiments of the present invention, a series of network traffic data of each network node can be obtained to form a network traffic data set. Network nodes can be paired up as network node pairs. Repeated IP addresses can be determined in the two network traffic data sets of a network node pair. The IP addresses that appear repeatedly in the two network traffic data sets can be used as the same network addresses. Based on the same network addresses, subsets of traffic data that interact with the same network addresses can be filtered out from the two network traffic data sets. The filtered subsets of traffic data can be used as target network traffic data. Based on the obtained target network traffic data, a spatial similarity used to measure the similarity degree of the content of the target network traffic data can be obtained, and a temporal similarity used to measure the similarity degree of the change trend of the target network traffic data can be obtained. The spatio-temporal similarity features can be jointly determined based on the determined spatial similarity and temporal similarity. In the network traffic data sets corresponding to each network node, the uplink protocol ratio of the number of each type of uplink traffic protocol to the total number of traffic protocols, and the downlink protocol ratio of the number of each type of downlink traffic protocol to the total number of traffic protocols can be determined. The uplink protocol ratios and the downlink protocol ratios can be used as vector elements to construct a protocol distribution vector. The Euclidean distance between the protocol distribution vectors of any two network nodes can be determined according to the Euclidean norm. The Wasserstein distance and the maximum Wasserstein distance between network traffic data sets can be statistically calculated. The ratio of the product of the Euclidean distance and the Wasserstein distance to the maximum Wasserstein distance can be used as a topological direction feature. A target model for analyzing network traffic can be obtained. Based on the obtained spatio-temporal similarity features and topological direction features, position features used to describe the positions of network nodes in the network can be determined. Positive and negative samples for training a network traffic analysis model can be constructed based on the network traffic data set. A training objective function of the network traffic analysis model can be constructed based on the positive samples and the negative samples. The target model can be trained based on the above position features and the training objective function to obtain a programmable network traffic analysis model.In the embodiment of the present invention, by accurately determining the target network traffic data with the same network address, the process of network traffic analysis can focus on the key network traffic data, reducing data redundancy and unnecessary data processing overhead, and improving the utilization rate of computing and storage resources for the programmable network traffic analysis model training method; this programmable network traffic analysis model training method can infer the location characteristics between network nodes without relying on physical location tags, enabling the training of the network traffic analysis model not to be limited by the acquisition of physical location tags, expanding the adaptability of the programmable network traffic analysis model training method, and enhancing the flexibility of the programmable network traffic analysis model training method; the location characteristics of network nodes can comprehensively reflect the spatio-temporal characteristics and topological structure information of network nodes. By training the network traffic analysis model based on the location characteristics, a more comprehensive feature representation can be provided for the training of the network traffic analysis model, improving the accuracy of the programmable network traffic analysis model training method; through unsupervised training based on the location characteristics of each network node and the network traffic data set, the dependence on a large amount of labeled data can be avoided, reducing the training cost.

[0202] Based on the above embodiments, in the embodiment of the present invention, the spatio-temporal feature determination module 320 further includes: a node pair matching unit for pairwise matching each network node as a network node pair;

[0203] a field determination unit for determining the total number of fields of all non-repeating data fields in the two target network traffic data of the network node pair;

[0204] a same field determination unit for determining the total number of same data fields in the two target network traffic data of the network node pair;

[0205] a spatial similarity determination unit for using the ratio of the total number of same fields to the total number of fields as the spatial similarity between the two target network traffic data of the network node pair;

[0206] a time series determination unit for sorting the target network traffic data of each network node according to its respective timestamp to obtain the network traffic time series of each network node;

[0207] a time similarity determination unit for using the dynamic time warping distance determined by the dynamic programming algorithm between the network traffic time series as the time similarity of each network node;

[0208] a similar feature determination unit for using the product of the spatial similarity and the time similarity of each network node as the spatio-temporal similar feature.

[0209] The direction feature determination module 330 further includes: an uplink protocol ratio determination unit for determining the uplink protocol ratio of the number of various uplink traffic protocols in the two target network traffic data of the network node pair to the number of all traffic protocols;

[0210] A downlink protocol ratio determination unit for determining the downlink protocol ratio of the number of various downlink traffic protocols in the two target network traffic data of the network node pair to the number of all traffic protocols;

[0211] A protocol distribution vector construction unit for constructing a protocol distribution vector of the corresponding network node by using each uplink protocol ratio and each downlink protocol ratio as vector elements;

[0212] A Wasserstein distance determination unit for determining the Wasserstein distance and the maximum Wasserstein distance between the network traffic data sets of any two network nodes;

[0213] A topological direction feature determination unit for determining the Euclidean distance between the protocol distribution vectors of any two network nodes, and taking the ratio of the product of the Euclidean distance and the Wasserstein distance to the maximum Wasserstein distance as the topological direction feature.

[0214] The model acquisition module 340 further includes: a normalization factor selection unit for selecting the larger feature among the spatio-temporal similarity feature and the topological direction feature as the normalization factor;

[0215] A traffic interaction intensity determination unit for taking the ratio of the product of the spatio-temporal similarity feature and the topological direction feature to the normalization factor as the traffic interaction intensity of the corresponding network node;

[0216] An implicit position graph construction unit for taking each network node as a vertex, determining that there are associated edges between the vertices according to the traffic interaction intensity, and constructing an implicit position graph of the network node based on the vertices and the associated edges;

[0217] A structured feature vector construction unit for determining the node degree of the vertex in the implicit position graph, and constructing a structured feature vector of the corresponding associated edge by using the node degree and the traffic interaction intensity as vector elements;

[0218] An exponentiated feature vector determination unit for performing an exponential operation on the structured feature vector processed by the LeakyRelu activation function to obtain an exponentiated feature vector;

[0219] A position embedding feature vector determination unit for determining all the exponentiated feature vectors corresponding to the associated edges connected by each vertex, and taking the ratio of the exponentiated feature vector to all the exponentiated feature vectors as the attention score of the vertex;

[0220] A position feature mapping unit for mapping the attention score to a position feature by using the Relu activation function;

[0221] A traffic feature encoding unit, configured to encode all data fields included in each network traffic dataset to obtain traffic features. The data fields at least include: source IP address, destination IP address, port information, protocol type, and timestamp. The encoding formats for encoding the data fields at least include: hash encoding and One-Hot encoding;

[0222] An intersection feature determination unit, configured to perform an outer product operation on the location feature and the traffic feature to obtain traffic location intersection features corresponding to each network node;

[0223] A fusion feature construction unit, configured to construct a fusion feature vector with the location feature, the traffic feature, and the traffic location intersection feature as vector elements;

[0224] A model acquisition unit, configured to perform unsupervised training on a deep learning architecture based on the fusion feature vector to obtain a network traffic analysis model.

[0225] The data communication device provided by the embodiments of the present invention can execute the data communication method provided by any embodiment of the present invention, and has corresponding functional modules and beneficial effects for executing the method.

[0226] Embodiment Five

[0227] The embodiments of the present invention provide a device for executing a programmable network traffic analysis model training method, a computer-readable medium, and a computer program product.

[0228] Figure 6 There is shown a device that can be used to execute a programmable network traffic analysis model training method. The device is intended to represent various forms of digital computers, such as, a laptop computer, a desktop computer, a workbench, a personal digital assistant, a server, a blade server, a mainframe computer, and other suitable computers. The device can also represent various forms of mobile devices, such as, a personal digital processor, a cellular phone, a smart phone, a wearable device (such as a helmet, glasses, a watch, etc.) and other similar computing devices. The components shown in the embodiments of the present invention, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the embodiments of the present invention described herein and / or claimed.

[0229] Such as Figure 6As shown, the device includes at least one processor 11 and a memory communicatively connected to the at least one processor 11, such as a Read-Only Memory (ROM) 12, a Random Access Memory (RAM) 13, etc. The memory stores a computer program executable by the at least one processor. The processor 11 can perform various appropriate actions and processes according to the computer program stored in the ROM 12 or the computer program loaded from the storage unit 18 into the RAM 13. In the RAM 13, various programs and data required for device operation can also be stored. The processor 11, the ROM 12, and the RAM 13 are connected to each other via a bus 14. An Input / Output (I / O) interface 15 is also connected to the bus 14.

[0230] Multiple components in the device are connected to the I / O interface 15, including: an input unit 16, such as a keyboard, a mouse, etc.; an output unit 17, such as various types of displays, speakers, etc.; a storage unit 18, such as a disk, an optical disc, etc.; and a communication unit 19, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 19 allows the device to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0231] The processor 11 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the processor 11 include, but are not limited to, a central processing unit, a graphics processing unit, various dedicated artificial intelligence computing chips, various processors running machine learning model algorithms, a digital signal processor, and any appropriate processor, controller, microcontroller, etc. The processor 11 executes the various methods and processes described above, such as the programmable network traffic analysis model training method.

[0232] In some embodiments, the programmable network traffic analysis model training method can be implemented as a computer program tangibly embodied in a computer-readable medium, such as the storage unit 18. In some embodiments, part or all of the computer program can be loaded and / or installed onto the device via the ROM 12 and / or the communication unit 19. When the computer program is loaded into the RAM 13 and executed by the processor 11, one or more steps of the programmable network traffic analysis model training method can be performed. Alternatively, in other embodiments, the processor 11 can be configured for the programmable network traffic analysis model training method by any other appropriate means (e.g., by means of firmware).

[0233] In various embodiments of the systems and techniques described above in the embodiments of the present invention, they can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays, application specific integrated circuits, application specific standard products, systems on a chip, programmable logic devices loaded, computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include: being implemented in one or more computer programs, which can be executed and / or interpreted on a programmable system including at least one programmable processor. The programmable processor can be a dedicated or general programmable processor, and can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit the data and instructions to the storage system, the at least one input device, and the at least one output device.

[0234] The computer programs for implementing the methods of the embodiments of the present invention can be written in any combination of one or more programming languages. These computer programs can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing devices, such that when the computer programs are executed by the processor, the functions / operations specified in the flowchart and / or block diagram are implemented. The computer programs can be executed entirely on the machine, partially on the machine, executed partially on the machine as an independent software package and partially on a remote machine, or executed entirely on a remote machine or server.

[0235] In the context of the embodiments of the present invention, a computer-readable medium can be a tangible medium that can contain or store a computer program for use by or in connection with an instruction execution system, apparatus, or device. The computer-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. Alternatively, the computer-readable medium can be a machine-readable signal medium. More specific examples of the machine-readable medium would include electrical connections based on one or more wires, portable computer disks, hard disks, RAM, ROM, erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disc read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0236] To provide interaction with a user, the systems and techniques described herein can be implemented on a device having: a display device (e.g., a cathode ray tube or a liquid crystal display monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) through which the user can provide input to the device. Other kinds of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including: acoustic input, voice input, or tactile input).

[0237] The systems and techniques described herein can be implemented in a computing system including backend components (e.g., as a data server), or a computing system including middleware components (e.g., an application server), or a computing system including frontend components (e.g., a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or a computing system including any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected to each other by digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include: local area networks, wide area networks, blockchain networks, and the Internet.

[0238] A computing system can include a client and a server. The client and the server are generally far from each other and typically interact through a communication network. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or a cloud host, which is a host product in the cloud computing service system and solves the defects of difficult management and weak business scalability existing in traditional physical hosts and virtual private server services.

[0239] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps recited in the present invention can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solution of the present invention can be achieved, and no limitation is imposed herein.

[0240] The above specific embodiments do not constitute a limitation on the protection scope of the embodiments of the present invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. A method for training a programmable network traffic analysis model, characterized in that, The method includes: S110. Obtain the network traffic data sets of each network node, determine the same network addresses among the network traffic data sets, and determine the target network traffic data within the corresponding network traffic data sets according to the same network addresses; S120. Determine the spatio-temporal similarity features of the network nodes based on the spatial similarity and temporal similarity between the target network traffic data of different network nodes; S130. Determine the protocol distribution vector corresponding to the network node according to the traffic protocol and traffic direction in the network traffic data set, and determine the topological direction features of each network node based on the Wasserstein distance between the protocol distribution vectors of different network nodes and the network traffic data sets; S140. Determine the position features corresponding to the network nodes according to the spatio-temporal similarity features and the topological direction features, and perform unsupervised training on the target model based on the position features of each network node and the network traffic data set to obtain a programmable network traffic analysis model.

2. The method according to claim 1, wherein In step S120, the determining the spatio-temporal similarity features of the network nodes based on the spatial similarity and temporal similarity between the target network traffic data of different network nodes includes: Pairwise match each network node as a network node pair; Determine the total number of data fields of all non-repeated data fields in the two target network traffic data of the network node pair; Determine the total number of identical data fields with the same data fields in the two target network traffic data of the network node pair; Take the ratio of the total number of identical data fields to the total number of data fields as the spatial similarity between the two target network traffic data of the network node pair; Sort the target network traffic data of each network node according to its respective timestamp to obtain the network traffic time series of each network node; Take the dynamic time warping distance between the network traffic time series determined by using the dynamic programming algorithm as the temporal similarity of each network node.

3. According to the method described in any one of claims 1 or 2, characterized in that, In the step S120, the determining the spatio-temporal similarity features of the network nodes based on the spatial similarity and temporal similarity between the target network traffic data of different network nodes includes: Take the product of the spatial similarity and the temporal similarity of each network node as the spatio-temporal similarity feature.

4. The method according to claim 2, wherein In step S130, the determining the protocol distribution vector corresponding to the network node according to the traffic protocol and traffic direction in the network traffic data set includes: Determine the uplink protocol ratio of the number of various uplink traffic protocols in the two target network traffic data of the network node pair to the number of all traffic protocols; Determine the downlink protocol ratio of the number of various downlink traffic protocols in the two target network traffic data of the network node pair to the number of all traffic protocols; Take each uplink protocol ratio and each downlink protocol ratio as vector elements to construct the protocol distribution vector corresponding to the network node.

5. The method according to claim 1, characterized in that, In step S130, determining the topological direction features of each network node based on the Wasserstein distance between the protocol distribution vectors of different network nodes and the network traffic data set includes: Determining the Wasserstein distance between the network traffic data sets of any two network nodes and the maximum Wasserstein distance; Determining the Euclidean distance between the protocol distribution vectors of any two network nodes, and taking the ratio of the product of the Euclidean distance and the Wasserstein distance to the maximum Wasserstein distance as the topological direction feature.

6. The method according to claim 1, characterized in that In step S140, determining the position features corresponding to the network nodes according to the spatio-temporal similarity features and the topological direction features includes: Selecting the larger feature among the spatio-temporal similarity features and the topological direction features as the normalization factor; Taking the ratio of the product of the spatio-temporal similarity features and the topological direction features to the normalization factor as the traffic interaction intensity corresponding to the network node; Taking each network node as a vertex, determining the existence of associated edges between the vertices according to the traffic interaction intensity, and constructing an implicit position graph of the network nodes based on the vertices and the associated edges; Determining the node degree of the vertex in the implicit position graph, and constructing a structured feature vector corresponding to the associated edge with the node degree and the traffic interaction intensity as vector elements; Performing an exponential operation on the structured feature vector processed by the LeakyRelu activation function to obtain an exponentialized feature vector; Determining all the exponentialized feature vectors corresponding to the associated edges connected by each vertex, and taking the ratio of the exponentialized feature vector to all the exponentialized feature vectors as the attention score of the vertex; Mapping the attention score to the position feature by using the Relu activation function.

7. The method according to claim 1, wherein In step S140, the unsupervised training of the target model based on the position features of each network node and the network traffic data set to obtain a programmable network traffic analysis model includes: Encoding all the data fields included in each network traffic data set to obtain the traffic features. The data fields at least include: source IP address, destination IP address, port information, protocol type, and timestamp. The encoding formats for encoding the data fields at least include: hash encoding and One-Hot encoding; Performing an outer product operation on the position feature and the traffic feature to obtain the traffic position cross feature corresponding to each network node; Constructing a fusion feature vector with the position feature, the traffic feature, and the traffic position cross feature as vector elements; Performing unsupervised training on the deep learning architecture based on the fusion feature vector to obtain a programmable network traffic analysis model.

8. A model training device, characterized in that, The device includes: A data determination module, configured to obtain the network traffic data sets of each network node, determine the same network addresses between the network traffic data sets, and determine the target network traffic data in the corresponding network traffic data sets according to the same network addresses; A spatio-temporal feature determination module, configured to determine the spatio-temporal similarity features of the network nodes based on the spatial similarity and temporal similarity between the target network traffic data of different network nodes; A direction feature determination module, configured to determine a protocol distribution vector corresponding to the network nodes according to the traffic protocols and traffic directions in the network traffic data set, and determine the topological direction features of each network node based on the Wasserstein distance between the protocol distribution vectors of different network nodes and the network traffic data set; A model acquisition module, configured to determine the position features corresponding to the network nodes according to the spatio-temporal similarity features and the topological direction features, and perform unsupervised training on a target model based on the position features of each network node and the network traffic data set to obtain a programmable network traffic analysis model.

9. A device, characterized in that, The device includes: At least one processor; and A memory communicatively connected to the at least one processor; wherein, The memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor to enable the at least one processor to execute the programmable network traffic analysis model training method according to any one of claims 1-7.

10. A computer-readable medium, characterized in that, The computer-readable medium storage includes: Computer instructions for causing a processor to implement the programmable network traffic analysis model training method according to any one of claims 1-7 when executed.