Network data classification method based on artificial intelligence

By constructing a structured protocol field graph and a multi-layer graph convolutional network, and combining protocol attention and behavioral residual sparse regularization, the problems of low accuracy and insufficient interpretability of traditional methods in highly encrypted network environments are solved, and high-precision and interpretable network threat detection is achieved.

CN120951035AActive Publication Date: 2025-11-14广州小白信息技术有限公司 +1
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
CN202511043712.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-28
Publication Date
2025-11-14
Estimated Expiration
2045-07-28

AI Technical Summary

Technical Problem

In the face of network traffic environments with high encryption, high concurrency, and multiple protocols, existing technologies suffer from low accuracy of traditional classification methods, difficulty in revealing the global intent of threat detection, and lack of interpretability. Machine learning solutions also lack host-time dimension behavior aggregation and context alignment.

Method used

An AI-based network data classification method is adopted. By constructing a structured protocol field graph, it is aggregated using a multi-layer graph convolutional network to generate connection semantic embedding vectors. Protocol attention masks and behavior residual sparse regularization are introduced to generate host behavior tensors and protocol attention score matrices. Finally, the attention-guided classifier outputs the classification score.

Benefits of technology

It achieves high-precision network threat detection in encrypted and protocol variation environments, provides interpretable classification results and direct evidence chains, and is suitable for high-encryption traffic and multi-protocol mixed environments, improving the comprehensiveness and reliability of detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120951035A_ABST
    Figure CN120951035A_ABST
Patent Text Reader

Abstract

The invention provides a network data classification method based on artificial intelligence, and the method comprises the steps: obtaining original connection data, and constructing a structured protocol field diagram; aggregating the structured protocol field graph through a multilayer graph convolutional network to generate a connection semantic embedding vector; performing context protocol semantic embedding enhancement on the connection semantic embedding vector to obtain an enhanced connection semantic vector; aggregating the connection semantic vector according to a host and a time window, and introducing a protocol attention mask and behavior residual sparse regularization to obtain and generate a host behavior tensor and a protocol attention score matrix; generating a behavior semantic vector through an aggregation tensor according to the host behavior tensor, and introducing a connection semantic distribution constraint to generate a host behavior semantic vector; and based on the semantic vector of the host behavior, outputting a classification score of the host behavior through a classifier based on attention guidance. According to the method, the comprehensiveness and credibility of network threat detection and service flow insight can be remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of network data classification, and in particular relates to a network data classification method based on artificial intelligence. Background Technology

[0002] With the continuous rise in the proportion of encrypted communication and the large-scale deployment of new multiplexing protocols such as HTTP / 3-QUIC, network traffic has exhibited complex characteristics of high encryption, high concurrency, and multiple protocol overlays. Traditional classification methods that rely solely on ports, keywords, or shallow statistical information experience a sharp drop in accuracy when faced with encrypted payloads and dynamic field reuse, while deep packet inspection is difficult to deploy comprehensively due to privacy compliance and computational overhead. Furthermore, modern attacks and business behaviors often span multiple connections: from DNS queries and TLS connection establishment to data push, forming a complete chain process, modeling only isolated connections cannot reveal their global intent; this leads to both errors and omissions in threat detection, and even normal business operations are easily misjudged. While some existing machine learning-based solutions introduce neural network feature extraction, they mostly remain at the single-connection level, lacking host-time dimension behavior aggregation and context alignment; moreover, model outputs are often presented as labels or scores, lacking readable reasoning paths, making it difficult to meet the hard requirements for interpretability in scenarios such as enterprise auditing and compliance forensics.

[0003] Therefore, network data classification urgently needs a new technical framework that can robustly extract semantics in encrypted and protocol variation environments, characterize behavioral sequences across connections, and provide structured interpretations of the results. Summary of the Invention

[0004] The purpose of this invention is to propose a network data classification method based on artificial intelligence to solve the above-mentioned problems.

[0005] To achieve the above objectives, this invention provides a network data classification method based on artificial intelligence, the method comprising:

[0006] Obtain the raw connection data and construct a structured protocol field graph;

[0007] The structured protocol field graph is aggregated using a multi-layer graph convolutional network to generate connection semantic embedding vectors.

[0008] The connection semantic embedding vector is enhanced by context protocol semantic embedding to obtain an enhanced connection semantic vector.

[0009] The connection semantic vectors are aggregated by host and time window, and a protocol attention mask and behavior residual sparse regularization are introduced to obtain the generated host behavior tensor and protocol attention score matrix.

[0010] Based on the host behavior tensor, a behavior semantic vector is generated by aggregating the tensor, and a connection semantic distribution constraint is introduced to generate the host behavior semantic vector.

[0011] Based on the host behavior semantic vector, an attention-guided classifier is used to output the classification score of the host behavior.

[0012] Furthermore, each node in the structured protocol field graph corresponds to a protocol field, and the node attributes are the numerical / category / text vector representation of the field and the field type label; the edges between nodes represent the logical or protocol dependencies between fields.

[0013] Furthermore, the step of aggregating the structured protocol field graph through a multi-layer graph convolutional network to generate connection semantic embedding vectors specifically includes:

[0014] Suppose that the structured protocol field graph has N field nodes, and the initial feature dimension of each node is d0, forming an initial feature matrix;

[0015] Based on the initial feature matrix, the neighbor nodes are aggregated using the adjacency matrix in each layer through graph convolution operations to update the node embedding;

[0016] The connection embeddings are concatenated using max pooling and average pooling to generate a unified connection embedding, resulting in a connection semantic embedding vector.

[0017] Furthermore, the step of enhancing the connection semantic embedding vector by performing context protocol semantic embedding to obtain an enhanced connection semantic vector specifically includes:

[0018] The connection semantic embedding vector is multiplebroadcast to each node and concatenated with the original node features to form a context-aware input;

[0019] Based on the context-aware input, the following convolutional embedding process is designed to update node representations:

[0020]

[0021] in, This is the normalized adjacency matrix; σ represents the convolution weights of the trainable graph; σ is the non-linear activation function; φ(E) conn ) is to embed the connection semantics into the vector E conn The control terms mapped to the node dimension adopt a shared perceptual vector approach; λ is a regularization control coefficient that determines the degree of influence of contextual intent on node semantic updates;

[0022] After the node representations are updated, an attention-based weighted aggregation method is used to merge the updated node representations into an enhanced connection semantic vector.

[0023] Furthermore, the enhanced connection semantic vectors exhibit separability for connections with the same field structure but different behavioral intentions. At the same time, they can automatically amplify the semantics of important fields during model training and adapt to cross-protocol and multi-mode network connection classification scenarios.

[0024] Further, the steps of aggregating the connection semantic vectors by host and time window, and introducing protocol attention mask and behavior residual sparse regularization to obtain the host behavior tensor and protocol attention score matrix specifically include:

[0025] Suppose there are H hosts in the entire network. First, based on the source IP address in the connection metadata, all connections are categorized by host, forming host sets host1, host2, ..., host H Each host corresponds to several connections;

[0026] Within the sliding window, the enhanced connection semantic vector Z for each host is collected sequentially. conn The behavior sequence is obtained by sorting the time sequence in ascending order. If the sequence length is less than the maximum step size L, it is padded with zero vectors; if it exceeds the maximum step size L, only the nearest L sequences are retained.

[0027] Introduce protocol attention masks to enhance key information; introduce behavioral residual sparse regularization to highlight transitions.

[0028] A weighted sequence is generated and filled into a three-dimensional tensor to generate a host behavior tensor. At the same time, the protocol attention score matrix is ​​recorded.

[0029] Furthermore, the protocol attention mask introduces protocol-aware attention α. (i,k) Each connection is weighted, with the weight jointly determined by the connection semantics and protocol type, specifically as follows:

[0030]

[0031] in, It is a learnable embedding of the protocol type (dimension p, implemented by lookup table), w is a parameter vector; σ uses Sigmoid to ensure α∈(0,1); among them, high-risk protocols or abnormal semantic vectors automatically obtain higher weights;

[0032] The behavioral residual sparse regularization encourages the model to focus on behavioral transitions by applying L1 regularization to the vector differences of consecutive connections within the same subject during the training phase. Specifically:

[0033]

[0034] in, is the sparse regularization term for behavioral residuals, where i and j are hosts i and j; Let j be the host behavior tensor of host j and j-1; where, if the host makes repeated indiscriminate requests within a period of time, the adjacent difference approaches zero, and the behavior residual sparse regularization prompts the model to reduce its gradient contribution, thereby reducing redundancy; when a scan → login or DNS → TLS jump occurs, the difference increases.

[0035] Furthermore, the step of generating a host behavior semantic vector by aggregating the host behavior tensor and introducing connection semantic distribution constraints specifically includes:

[0036] The host behavior tensor is weighted and aggregated based on the protocol attention score matrix to obtain the attention-guided behavior semantic vector for each host.

[0037] A connection semantic distribution constraint is introduced onto the behavior semantic vector and then concatenated to generate a host behavior semantic vector; wherein the connection semantic distribution constraint is modified by designing a regular vector based on the connection semantic distribution constraint, and the regular vector is used to characterize the semantic variance in the connection sequence.

[0038] Furthermore, the step of outputting a classification score of the host behavior based on the host behavior semantic vector and by using an attention-guided classifier specifically includes:

[0039] The host behavior semantic vector is fed into a set of parallel feedforward encoding modules for feature transformation to obtain the host behavior semantic representation.

[0040] Design an attention weighting mechanism among hosts to enable the attention-guided classifier to consider the local host group behavior trends during classification and output the classification score of host behavior.

[0041] Furthermore, the inter-host attention weighting mechanism is calculated based on a cross-host architecture attention correction term; the cross-host architecture attention correction term is calculated as follows:

[0042] Average pooling is performed on the host behavior semantic representations of all hosts to obtain the host group behavior center;

[0043] Calculate the cosine difference between the current host behavior semantic representation and the host group row-centered value as a cross-host architecture attention correction term:

[0044]

[0045] in, λ is an adjustable coefficient, ranging from [0,1].

[0046] The beneficial technical effects of the present invention are at least as follows:

[0047] The intelligent network data classification method provided by this invention is based on a three-level progressive architecture of "connection semantics—behavioral sequence—host profiling". First, a protocol field graph is constructed at the connection level, and context-injected graph neural networks are used to dynamically reconstruct the semantic embeddings of fields, solving the feature mismatch problem caused by protocol diversity and field drift. Then, multiple connections are aggregated into a time-series tensor according to host and time window, introducing protocol attention and residual sparsity regularization to highlight key connections with attack or anomaly characteristics, completing a structured expression of complex behavioral chains. Next, host behavioral semantic vectors are obtained through weighted aggregation and variance characterization, and cross-host comparative attention is added to the classification network to achieve simultaneous identification of isolated anomalies and group anomalies. Finally, the system outputs traceable classification results and key connection indexes, providing a direct chain of evidence for security operations and compliance audits. The overall solution balances high accuracy, real-time performance, and interpretability, and is particularly suitable for production environments with high encryption traffic, high protocol mixing, and a large number of concurrent hosts, significantly improving the comprehensiveness and reliability of network threat detection and business traffic insights. Attached Figure Description

[0048] The present invention will be further described with reference to the accompanying drawings, but the embodiments in the drawings do not constitute any limitation on the present invention. For those skilled in the art, other drawings can be obtained based on the following drawings without creative effort.

[0049] Figure 1 This is a flowchart of a network data classification method based on artificial intelligence according to the present invention. Detailed Implementation

[0050] Embodiments of the present invention are described in detail below. Examples of these embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present invention, and should not be construed as limiting the present invention.

[0051] like Figure 1 As shown in the figure, an embodiment of the present invention provides a network data classification method based on artificial intelligence, the method comprising the following steps S1-S5:

[0052] S1. Obtain the original connection data and construct a structured protocol field graph; aggregate the structured protocol field graph through a multi-layer graph convolutional network to generate connection semantic embedding vectors.

[0053] Specifically, the goal of this step is to obtain raw connection data from a real network environment, extract the structured semantic vector of each connection through protocol field graph modeling and graph neural network feature learning, and provide reproducible basic input for subsequent multi-connection behavior modeling.

[0054] In actual deployment, connection data is collected as follows: Data collection modules are deployed on egress firewalls, core switches, or mirror probes, and network traffic is captured in real time using tools such as NetFlow, sFlow, or dedicated packet capture tools (such as TShark and nDPI). For each complete connection, a five-tuple of information (source IP, destination IP, source port, destination port, and protocol type) and its protocol payload are recorded. For example, a typical HTTPS connection record includes the following fields: source IP is 192.168.1.8, destination IP is 203.0.113.10, protocol is TCP, source port is 443, destination port is 53128, and the payload includes the TLSClientHello message field, encryption algorithm family, domain name SNI, etc. After all fields are processed by a protocol parser (such as the Wireshark protocol stack or nDPI library), a uniformly formatted data structure is output, and the field type is automatically marked as numeric, categorical, or text.

[0055] Furthermore, in the field preprocessing stage:

[0056] Numeric fields (such as port numbers and sequence numbers) retain their original values ​​directly.

[0057] Categorical fields (such as protocol type, TCPFlag) use One-Hot or categorical index encoding;

[0058] Text fields (such as HTTP paths, SNIs) are embedded as fixed-length vectors using encoding models based on Word2Vec or BERT;

[0059] Ultimately, each field is represented by a unified vector, which facilitates subsequent calculations.

[0060] The processed field results are constructed into a structured protocol field graph. Each node in this graph corresponds to a protocol field, and the node attributes are the field's numerical / category / text vector representation and field type label; the edges between nodes represent the logical or protocol dependencies between fields, for example:

[0061] TCPFlags are strongly correlated with sequence numbers in terms of timing.

[0062] The HTTP method field and the request path field have a contextual dependency;

[0063] The fields in SNI and ClientHello are permanently bound.

[0064] The edge weights consist of two parts: first, the structural dependencies parsed from the protocol specification (such as mandatory field combinations); and second, the statistical co-occurrence frequency of fields in millions of connection data (such as the probability of two fields appearing simultaneously in the same connection). After standard normalization, weighted directed edges are generated. The final constructed protocol field graph is used to represent the contextual structure and semantic coupling between fields in each connection.

[0065] Furthermore, the graph structure is input into a multi-layer graph convolutional network (GCN). Assume the graph has N field nodes, each node having an initial feature dimension of d0, forming an initial feature matrix. By using graph convolution operations, neighboring nodes are aggregated using the adjacency matrix in each layer, and the node embeddings are updated. The specific update formula is as follows:

[0066]

[0067] Among them, H (l) W is the feature matrix of the nodes in the l-th layer. (l) This is the trainable weight matrix for this layer; Let A be the adjacency matrix with self-loops, and I be the adjacency matrix of the protocol field graph. yes The degree matrix; σ is the activation function, such as ReLU.

[0068] Taking a graph convolutional layer with L=2 layers as an example, the output is a context-aware embedding vector for each node. These vectors fuse semantic information from the field itself and its neighboring fields. To aggregate node-level information into a unified representation at the connection level, connection embeddings are generated using a concatenation strategy combining max pooling and average pooling.

[0069]

[0070] in, This is the final embedding of the i-th node in layer L. The max operation takes the maximum response of each feature dimension, average pooling captures the overall distribution, and the result is obtained after concatenation. It is the semantic representation vector of the connection.

[0071] This step ultimately outputs two variables: the structured protocol field graph G. conn and connect semantic embedding vector E conn G conn It is a graph structure input used in subsequent behavioral modeling steps; E conn It is a vector form that connects semantic features and is one of the core inputs for the entire system to achieve intelligent network classification.

[0072] S2. Perform context protocol semantic embedding enhancement on the connection semantic embedding vector to obtain an enhanced connection semantic vector.

[0073] Specifically, the goal of this step is to further introduce contextual dependencies, scenario constraints, and network behavior intentions, while preserving the structural information of the protocol fields, thereby improving the connection semantic vector E generated in the previous step. conn Enhanced reconstruction is performed to ultimately generate a connection semantic representation Z with global context awareness. conn This indicates that it will serve as the direct input for subsequent behavior recognition and connection classification, and is a core step in the entire scheme that connects the preceding and following steps.

[0074] In real-world network data scenarios, issues arise such as "same field structure but different behaviors." For example, two HTTPS connections may have the same protocol field graph structure, but one accesses an enterprise OA system while the other accesses an intranet penetration tool (such as frp). From the field graph perspective, there is no significant difference, and traditional field topology-based GCNs struggle to effectively distinguish between such connections with "same structure but different intentions." Therefore, this paper designs a dual-attention graph neural network mechanism (Context-Guided GNN) that integrates contextual semantics and connection structure to enhance the connection representation's ability to perceive behavioral intentions and protocol usage.

[0075] The input for this step is: G conn The protocol field graph generated in the previous step includes node fields and their edge structure; Connection semantic embedding vectors generated by multi-layer GCN aggregation.

[0076] Furthermore, to enable each node to embed more comprehensive connection context information, this invention introduces a context injection mechanism. Specifically, this involves embedding E... conn The data is repeatedly broadcast to each node and concatenated with the original node features to form the context-aware input X. ctx :

[0077] X ctx =[X||Repeat(E) conn ,N)] (3)

[0078] in, The original features of each field node, where N is the number of field nodes. To create a context representation for each node by copying the connection vector N times, the following steps are performed: After concatenation... This means that each field simultaneously possesses its own structural characteristics and global context information.

[0079] Based on this context-enhanced feature, the convolutional embedding process is designed as shown in the figure below:

[0080]

[0081] in: This is the normalized adjacency matrix; σ represents the convolution weights of the trainable graph; σ is the non-linear activation function; φ(E) conn ) for E conn The control terms mapped to the node dimension adopt a shared perceptual vector approach; λ is a regularization control coefficient that determines the degree of influence of contextual intent on node semantic updates.

[0082] The innovation lies in the addition of φ(E) to the formula. conn This term is equivalent to introducing a "connection intent vector control channel" during the node update process. It is used to adjust whether the node update conforms to the semantic positioning of the connection in the global network behavior. For example, in connections with strong attack intent, this channel will push certain sensitive fields (such as request paths and DNS query names) closer to the abnormal semantic space, thereby improving the recognition ability of subsequent classification models.

[0083] Furthermore, after the node representations are updated, an attention-based weighted aggregation method is used to merge the node representations into the final connection semantic embedding:

[0084]

[0085] Among them, h i For the updated representation of the i-th field node, α i For E-based conn and h i Relevance scoring function This attention machine All parameters are trainable and used to calculate the relevance of field nodes to global semantics. This approach strengthens the representation weight of important fields in the join representation, allowing context-sensitive fields such as "HTTP method," "SNI field," and "DNS query name" to occupy a higher proportion in the final representation, thus enhancing discriminative capabilities.

[0086] Final output To enhance the semantic vector of connections, it possesses the following characteristics: connections with the same field structure but different behavioral intentions have the following Z-axis. conn It exhibits separability; it can automatically amplify the semantics of important fields during model training; and it can adapt to cross-protocol and multi-mode network connection classification scenarios, such as the coexistence of plaintext and encryption protocols, covert communication based on DNS tunnels, and high-variable fields in P2P networks.

[0087] In summary, this step introduces global semantics and behavioral intent into the protocol field structure by using context-aware graph convolution, connection intent regularization, and attention aggregation mechanism. This significantly improves the fine-grainedness and generalization ability of network connection semantic modeling, and is the core module of this invention for dealing with the diversity of network protocols and the variability of attacks.

[0088] S3. Aggregate the connection semantic vectors by host and time window, and introduce protocol attention mask and behavior residual sparse regularization to obtain the generated host behavior tensor and protocol attention score matrix.

[0089] Specifically, this step involves the already obtained connection semantic representation Z. conn Based on this, host-layer behavior sequences are constructed and converted into high-dimensional temporal tensors T. seq This allows the model to observe "the collaborative behavior of multiple connections at the time axis and protocol level." This representation inherits the contextual semantics provided in step two and adds cross-connection temporal, protocol type, and behavioral differences information, which is crucial for identifying scenarios such as lateral movement, distributed scanning, and DNS tunneling. It serves as a bridge for subsequent classifiers to understand "behavioral intent."

[0090] enter: The contextual semantic vector for each connection.

[0091] in,<src_ip,dst_ip,timestamp,proto> : Connection metadata, where proto is the protocol type number (TCP=1, UDP=2, QUIC=3...).

[0092] Furthermore, assuming there are H hosts in the entire network, firstly, based on the source IP address in the connection metadata, all connections are categorized by host, forming host sets host1, host2, ..., host H Each host corresponds to several connections. Then, within a sliding window (default 300s, step size 60s), the connection semantic vector Z of each host is collected sequentially. conn The behavior sequence is obtained by sorting the time sequence in ascending order. If the sequence length is less than the maximum step size L (empirical value 64), it is padded with zero vectors; if it exceeds the limit, only the nearest L sequences are retained to ensure that the tensor size is fixed and the real-time performance is controllable.

[0093] ① Protocol Attention Mask (Enhancing Key Information): To prevent massive normal ACKs and Keep-alive from preempting sequence capacity, protocol-aware attention α is introduced. (i,k) Each connection is weighted, with the weight jointly determined by the connection semantics and protocol type:

[0094]

[0095] It is a learnable embedding of the protocol type (dimension p, implemented by lookup table), w is a parameter vector; σ uses Sigmoid to ensure α∈(0,1). High-risk protocols (such as QUIC-0RTT, unknown port UDP) or anomalous semantic vectors automatically receive higher weights. This construction... It both reduces redundancy and highlights potential attack connections.

[0096] ② Behavioral Residual Sparse Regularization (Highlighting Jumps): During the training phase, L1 regularization is applied to the differences in vectors connected to the same subject, encouraging the model to focus on behavioral inflection points.

[0097]

[0098] If a host makes repeated indiscriminate requests over a period of time, the difference between adjacent values ​​approaches zero. The regularization term prompts the model to reduce its gradient contribution, thereby reducing redundancy. When there is a transition from scanning to login or from DNS to TLS, the difference increases, and the model is able to capture the real behavioral transition.

[0099] Fill the weighted sequence into the three-dimensional tensor: Tensor size (H,L,d1). Simultaneously record the protocol attention score matrix Γ[i,j]=α. (i,j) This is for subsequent interpretable analysis.

[0100] Output The host behavior tensor can be directly input into Transformer or TCN for behavior classification. The protocol attention score matrix is ​​used only for reasoning interpretability or visualization.

[0101] By leveraging protocol attention masks and residual sparsity regularization, this step extends the semantics of a single connection into a structural tensor that expresses "behavioral rhythm," enabling the system to maintain high sensitivity to abnormal intents even in encrypted traffic, mixed protocols, and high-frequency short connection environments, providing high-quality temporal-semantic integrated input for subsequent classification decisions.

[0102] S4. Generate a behavior semantic vector by aggregating the host behavior tensor and introducing connection semantic distribution constraints to generate the host behavior semantic vector.

[0103] Specifically, the goal of this step is to: extract the host behavior tensor from the previous step. With connection attention matrix Based on this, a structured behavioral semantic modeling framework Φ is constructed. sem It is used to characterize the multi-connection behavior pattern of a host within a certain time window.

[0104] First, this invention is based on Γ on the host behavior tensor Tseq We perform weighted aggregation to obtain the attention-guided behavioral semantic vector μ for each host. i :

[0105]

[0106] Where i represents the host ID, and j represents the j-th connection of that host. This step aggregates the L connection vectors into a d1-dimensional vector, representing the host's "compressed semantic behavior" within that time window. The attention weights are derived from step three and have been obtained through adaptive learning of the time-protocol-semantic structure.

[0107] To improve the discriminative power of this modeling vector, this invention uses μ i Based on this, a connection semantic distribution constraint is introduced, and a regular vector δ is designed. i This is used to characterize the semantic variance in the connected sequences and is embedded into the model as an independent feature:

[0108]

[0109] This vector measures the degree of change in the semantics of internal connections within a host, and can significantly distinguish between aggressive hosts (with drastic changes in behavior patterns) and stable hosts (with strong behavioral aggregation). This is achieved by simultaneously considering μ. i and δ i This invention defines host behavior modeling features. for:

[0110]

[0111] Here, || represents a vector concatenation operation, resulting in a host-level semantic representation of dimension 2d1. All vectors have been normalized.

[0112] Output The host behavior semantic vector is the result of splicing together the behavior models of all hosts.

[0113] S5. Based on the host behavior semantic vector, output the classification score of the host behavior by using an attention-guided classifier.

[0114] Specifically, this step receives the host behavior semantic vector constructed in the previous step. By constructing a lightweight attention-guided classifier, host behavior can be discriminated and classified.

[0115] Furthermore, firstly, the input behavioral semantic vector is fed into a set of parallel feedforward encoding modules for feature transformation. This module consists of two linear layers (input 2d1-dimensional, output d...). h (Dimension), with a ReLU activation function in the middle:

[0116]

[0117] Here, This represents the weights and biases of the fully connected layer. Dimension d h The value is typically set to 128 or 256, depending on the business deployment. This module's function is to extract deeper behavioral difference features, improving the model's non-linear expressive power; i Host behavior semantic representation.

[0118] To introduce an information-sharing mechanism across hosts, this invention designs an inter-host attention weighting mechanism, allowing the model to consider the behavioral trends of local host groups during classification. The formula is as follows:

[0119]

[0120] in, Let C be the classification result vector for host i (where C is the total number of behavior categories). α i It is a cross-host architecture attention correction term, and its calculation method is as follows:

[0121] Representation z for all hosts j Perform average pooling to obtain the host group behavior center;

[0122] Calculate the cosine difference between host i and the center as a correction term:

[0123]

[0124] in λ is an adjustable coefficient, ranging from [0,1].

[0125] The mechanism aims to enhance the model's ability to identify isolated or clustered abnormal behaviors by utilizing comparative information between host behaviors. It is particularly suitable for detection tasks in high-concurrency host environments such as enterprise networks and campus networks.

[0126] Final output The classification score vector for each host can be used for subsequent attack detection, auditing alerts, and alert tracing.

[0127] It should be noted that the workflow described above is merely illustrative and does not limit the scope of protection of this invention. In practical applications, those skilled in the art can select some or all of the workflow to achieve the purpose of this embodiment according to actual needs, and no restrictions are imposed here.

[0128] In addition, for technical details not described in detail in this embodiment, please refer to the parameter operation method provided in any embodiment of the present invention, which will not be repeated here.

[0129] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or system that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or system. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or system that includes that element.

[0130] The sequence numbers of the above embodiments of the present invention are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments.

[0131] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as read-only memory / random access memory, magnetic disk, optical disk) and includes several instructions to cause a terminal device (which may be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in the various embodiments of the present invention.

[0132] The above are merely preferred embodiments of the present invention and do not limit the scope of the patent. Any equivalent structural or procedural transformations made based on the description and drawings of the present invention, or direct or indirect applications in other related technical fields, are similarly included within the scope of patent protection of the present invention.

Claims

1. A network data classification method based on artificial intelligence, characterized in that, The method includes: Obtain the raw connection data and construct a structured protocol field graph; The structured protocol field graph is aggregated using a multi-layer graph convolutional network to generate connection semantic embedding vectors. The connection semantic embedding vector is enhanced by context protocol semantic embedding to obtain an enhanced connection semantic vector; The connection semantic vectors are aggregated by host and time window, and a protocol attention mask and behavior residual sparse regularization are introduced to obtain the generated host behavior tensor and protocol attention score matrix. Based on the host behavior tensor, a behavior semantic vector is generated by aggregating the tensor, and a connection semantic distribution constraint is introduced to generate the host behavior semantic vector. Based on the host behavior semantic vector, an attention-guided classifier is used to output the classification score of the host behavior.

2. The network data classification method based on artificial intelligence according to claim 1, characterized in that, Each node in the structured protocol field graph corresponds to a protocol field. The node attributes are the numerical / category / text vector representation of the field and the field type label. The edges between nodes represent the logical or protocol dependencies between fields.

3. The network data classification method based on artificial intelligence according to claim 1, characterized in that, The step of aggregating the structured protocol field graph using a multi-layer graph convolutional network to generate connection semantic embedding vectors specifically includes: Suppose that the structured protocol field graph has N field nodes, and the initial feature dimension of each node is d0, forming an initial feature matrix; Based on the initial feature matrix, the neighbor nodes are aggregated using the adjacency matrix in each layer through graph convolution operations to update the node embedding; The connection embeddings are concatenated using max pooling and average pooling to generate a unified connection embedding, resulting in a connection semantic embedding vector.

4. The network data classification method based on artificial intelligence according to claim 1, characterized in that, The step of performing context protocol semantic embedding enhancement on the connection semantic embedding vector to obtain the enhanced connection semantic vector specifically includes: The connection semantic embedding vector is multiplebroadcast to each node and concatenated with the original node features to form a context-aware input; Based on the context-aware input, the following convolutional embedding process is designed to update node representations: in, This is the normalized adjacency matrix; σ represents the convolution weights of the trainable graph; σ is the non-linear activation function; φ(E) conn ) is to embed the connection semantics into the vector E conn The control terms mapped to the node dimension adopt a shared perceptual vector approach; λ is a regularization control coefficient that determines the degree of influence of contextual intent on node semantic updates; After the node representations are updated, an attention-based weighted aggregation method is used to merge the updated node representations into an enhanced connection semantic vector.

5. The network data classification method based on artificial intelligence according to claim 1, characterized in that, The enhanced connection semantic vector exhibits separability for connections with the same field structure but different behavioral intentions. It can also automatically amplify the semantics of important fields during model training and adapt to cross-protocol and multi-modal network connection classification scenarios.

6. The network data classification method based on artificial intelligence according to claim 1, characterized in that, The steps of aggregating the connection semantic vectors by host and time window, and introducing protocol attention masks and behavioral residual sparse regularization to obtain the host behavior tensor and protocol attention score matrix specifically include: Suppose there are H hosts in the entire network. First, based on the source IP address in the connection metadata, all connections are categorized by host, forming host sets host1, host2, ..., host H Each host corresponds to several connections; Within the sliding window, the enhanced connection semantic vector Z for each host is collected sequentially. conn The behavior sequence is obtained by sorting the time sequence in ascending order. If the sequence length is less than the maximum step size L, it is padded with zero vectors; if it exceeds the maximum step size L, only the nearest L sequences are retained. Introduce protocol attention masks to enhance key information; introduce behavioral residual sparse regularization to highlight transitions. A weighted sequence is generated and filled into a three-dimensional tensor to generate a host behavior tensor. At the same time, the protocol attention score matrix is ​​recorded.

7. The network data classification method based on artificial intelligence according to claim 6, characterized in that, The protocol attention mask introduces protocol-aware attention α. (i,k) Each connection is weighted, with the weight jointly determined by the connection semantics and protocol type, specifically as follows: in, It is a learnable embedding of the protocol type (dimension p, implemented by lookup table), w is a parameter vector; σ uses Sigmoid to ensure α∈(0,1); among them, high-risk protocols or abnormal semantic vectors automatically obtain higher weights; The behavioral residual sparse regularization encourages the model to focus on behavioral transitions by applying L1 regularization to the vector differences of consecutive connections within the same subject during the training phase. Specifically: in, is the sparse regularization term for behavioral residuals, where i and j are hosts i and j; Let j be the host behavior tensor of host j and j-1; where, if the host makes repeated indiscriminate requests within a period of time, the adjacent difference approaches zero, and the behavior residual sparse regularization prompts the model to reduce its gradient contribution, thereby reducing redundancy; when a scan → login or DNS → TLS jump occurs, the difference increases.

8. The network data classification method based on artificial intelligence according to claim 1, characterized in that, The step of generating a host behavior semantic vector by aggregating the host behavior tensor and introducing connection semantic distribution constraints specifically includes: The host behavior tensor is weighted and aggregated based on the protocol attention score matrix to obtain the attention-guided behavior semantic vector for each host. A connection semantic distribution constraint is introduced into the behavioral semantic vector and concatenated to generate a host behavioral semantic vector; wherein the connection semantic distribution constraint is modified by designing a regular vector based on the connection semantic distribution constraint, and the regular vector is used to characterize the semantic variance in the connection sequence.

9. The network data classification method based on artificial intelligence according to claim 1, characterized in that, The step of outputting a classification score for host behavior based on the host behavior semantic vector and through an attention-guided classifier specifically includes: The host behavior semantic vector is fed into a set of parallel feedforward encoding modules for feature transformation to obtain the host behavior semantic representation. Design an attention weighting mechanism among hosts to enable the attention-guided classifier to consider the local host group behavior trends during classification and output the classification score of host behavior.

10. The network data classification method based on artificial intelligence according to claim 9, characterized in that, The inter-host attention weighting mechanism is calculated based on a cross-host architecture attention correction term; the cross-host architecture attention correction term is calculated as follows: Average pooling is performed on the host behavior semantic representations of all hosts to obtain the host group behavior center; Calculate the cosine difference between the current host behavior semantic representation and the host group row-centered value as a cross-host architecture attention correction term: in, λ is an adjustable coefficient, ranging from [0,1].

Citation Information

Patent Citations

  • Android malicious software detection method based on graph convolutional neural network

    CN117909975A

  • Encrypted traffic classification method and system

    CN118864921A

  • Network anomaly detection method, computer equipment and computer program product

    CN120110764A

  • Network protocol reverse analysis method based on deep learning and graph neural network

    CN120321322A

  • Network attack dynamic detection and security protection method and system based on artificial intelligence

    CN120342748A