Network traffic anomaly monitoring and early warning method and system based on deep learning

By combining the long-sequence Transformer variant model and capsule neural network, the sparse attention mechanism and dynamic routing mechanism are adopted, and the problem of identifying complex attack behaviors in complex network traffic environments is solved, achieving efficient and accurate anomaly detection and early warning.

CN120050195AActive Publication Date: 2025-05-27NANJING HUIZHOU TONGQIAO INFORMATION TECH CO LTD

Patent Information

Application Number
CN202510200675.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-24
Publication Date
2025-05-27
Estimated Expiration
2045-02-24

AI Technical Summary

Technical Problem

It is difficult for the prior art to effectively identify new and complex attack behaviors in complex high-dimensional network traffic environments, especially in multi-protocol, multi-stage and long-span attack scenarios. Traditional methods lack the ability to identify unknown threats or covert attacks.

Method used

The network traffic anomaly monitoring and early warning method based on deep learning is adopted, and the timing dependence and multi-level structural characteristics in network traffic are captured through the combination of the long-sequence Transformer variant model and the capsule neural network, and the detection efficiency and explanatory nature are improved through the sparse attention mechanism and dynamic routing mechanism.

Benefits of technology

Efficiently and accurately identify complex or latent attack behaviors in large-scale or multi-protocol scenarios, reduce false alarm rates and improve detection accuracy, and achieve flexible alarm output through the combination of early warning modules and external threat intelligence databases, meeting the real-time or quasi-real-time protection needs of multiple scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120050195A_ABST
    Figure CN120050195A_ABST
Patent Text Reader

Abstract

The invention provides a network traffic anomaly monitoring and early warning method and system based on deep learning, and relates to the field of network security. The method comprises the following steps: collecting target network traffic and preprocessing the target network traffic to generate serialized traffic data; inputting the traffic data into a long sequence Transform variant model with a sparse attention mechanism, representing the time sequence dependence and global context features of the traffic data, and generating a first feature representation; inputting the first feature representation into a capsule neural network, and obtaining a multi-level depth feature representation through a dynamic routing mechanism of a capsule unit; and performing anomaly judgment on the multi-level depth feature representation, generating an anomaly judgment result, and generating early warning information based on the anomaly judgment result. The method gives consideration to calculation efficiency and global dependence in a long sequence environment, can improve the recognition accuracy of complex attacks, and is suitable for real-time protection of large-scale network traffic.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of network security technology, and more specifically, to a method and system for monitoring and warning network traffic anomalies based on deep learning. Background Art

[0002] As the scale and complexity of networks continue to expand, traditional anomaly detection methods based on rules or simple statistical models are gradually unable to adapt to complex high-dimensional traffic environments, especially in the face of multi-protocol, multi-stage and long-span attack scenarios. Such methods are insufficient in identifying unknown threats or hidden attacks. In order to cope with new and complex attack forms, deep learning technology has received widespread attention in the field of network security, but how to balance global dependency and efficiency in ultra-long sequence traffic data and retain good interpretability in the detection results is still the main challenge facing existing technologies. Summary of the invention

[0003] In response to the deficiencies in the prior art, the present application provides a network traffic anomaly monitoring and early warning method and system based on deep learning.

[0004] In the first aspect, the present application provides a network traffic anomaly monitoring and early warning method based on deep learning, including:

[0005] Collect target network traffic and preprocess it to generate serialized traffic data;

[0006] Inputting the traffic data into a long sequence Transformer variant model with a sparse attention mechanism, characterizing the temporal dependency and global context features of the traffic data, and generating a first feature representation;

[0007] Inputting the first feature representation into a capsule neural network, and obtaining a multi-level deep feature representation through a dynamic routing mechanism of capsule units;

[0008] Anomaly determination is performed on the multi-level deep feature representation to generate anomaly determination results, and warning information is generated based on the anomaly determination results; wherein the anomaly determination results include: anomaly scores and category labels.

[0009] As an optional implementation, the generating of serialized traffic data includes:

[0010] Based on a preset stream data scheme, periodically deriving connection summary information as the stream data;

[0011] De-duplication, merging and correlation analysis are performed on the stream data to obtain key fields;

[0012] The key fields include: source IP, destination IP, port number, protocol type and timestamp;

[0013] The key fields are organized according to time sequence and connection relationship to generate serialized traffic data.

[0014] As an optional implementation manner, generating the first feature representation includes:

[0015] Dividing the serialized traffic data into a number of blocks of equal size;

[0016] Perform local attention calculation on sequence elements in each block to obtain local attention scores;

[0017] Randomly select some elements as global connection tokens to establish global attention associations between adjacent blocks and obtain global attention scores;

[0018] An attention matrix is ​​constructed based on the local attention score and the global attention score, and the traffic data is multi-layer encoded to generate the first feature representation.

[0019] As an optional implementation, the obtaining of multi-level deep feature representation through the dynamic routing mechanism of capsule units includes:

[0020] Setting a plurality of protocol capsules at a lower layer of the capsule network for generating initial activation vectors related to different protocols based on the first feature representation;

[0021] A plurality of attack phase capsules are arranged at the high layer of the capsule network, and an update vector outputted by the protocol capsule after iteration is received during the dynamic routing process; wherein the update vector is an activation vector generated by the protocol capsule according to the initial activation vector and high layer feedback in the dynamic routing iteration, and is used to transmit protocol-related features to the attack phase capsule;

[0022] The attack stage capsules are used to characterize different abnormal stage features respectively; the abnormal stage features include: scanning, vulnerability exploitation, lateral movement and data transmission;

[0023] Multi-level deep feature representations are generated at the output of high-level capsules.

[0024] As an optional implementation manner, generating warning information includes:

[0025] Based on normal traffic data, build a reference traffic sample library;

[0026] Determine the similarity between the current multi-level deep feature representation and the reference traffic sample library;

[0027] In response to the similarity being lower than a preset threshold, performing label determination in combination with an external threat intelligence library;

[0028] The anomaly score and category label are output according to the judgment result, and different levels of warning information are generated according to the preset multi-level thresholds.

[0029] As an optional implementation, the normal traffic is traffic that has been confirmed by security audit to have no attack behavior and is collected within a pre-set no-attack time period.

[0030] As an optional implementation, before randomly selecting some elements as global connection tokens, the method further includes:

[0031] Based on at least one of the port number, the protocol type and the keyword, the serialized traffic data is screened to determine a high-risk Token, and the high-risk Token is included in a candidate set of global connection tokens.

[0032] As an optional implementation, the multi-layer encoding of the traffic data further includes:

[0033] Set up sparse attention mechanisms for different Transformer encoder layers; where:

[0034] Local window attention is used at the first level;

[0035] Use random global attention in the middle layers;

[0036] The global connection scope is divided based on the protocol type or time dimension at least one layer.

[0037] As an optional implementation, protocol capsules and attack phase capsules are simultaneously set at the high level of the capsule network, and a multi-path cross-coupling relationship is established through a dynamic routing algorithm, so that the update vector output by the protocol capsule and the initial activation vector of the attack phase capsule are transmitted to each other.

[0038] In the second aspect, the present application provides a network traffic anomaly monitoring and early warning system based on deep learning, including:

[0039] A collection unit, used to collect target network traffic and perform preprocessing to generate serialized traffic data;

[0040] A first processing unit, configured to input the traffic data into a long sequence Transformer variant model with a sparse attention mechanism, characterize the temporal dependency and global context features of the traffic data, and generate a first feature representation;

[0041] A second processing unit, configured to input the first feature representation into a capsule neural network, and obtain a multi-level deep feature representation through a dynamic routing mechanism of capsule units;

[0042] A generating unit is used to perform anomaly determination on the multi-level deep feature representation, generate anomaly determination results, and generate warning information based on the anomaly determination results; wherein the anomaly determination results include: anomaly scores and category labels.

[0043] Compared with the prior art, this application combines the long-sequence Transformer and the capsule network to capture the long-term dependencies and multi-level structural features in network traffic, so that complex or latent attack behaviors can be efficiently and accurately identified in large-scale or multi-protocol scenarios. The sparse attention mechanism can retain global key information while reducing computation and memory usage; the dynamic routing of the capsule network improves the characterization and interpretability of protocol details and attack stage behaviors, thereby achieving lower false alarm rates and higher detection accuracy in various network environments. Through the combination of the early warning module and the external threat intelligence library, the present invention can flexibly output alarms at different risk levels to meet the real-time or quasi-real-time protection needs of multiple scenarios. BRIEF DESCRIPTION OF THE DRAWINGS

[0044] Figure 1 A flowchart of a method for monitoring and warning network traffic anomalies based on deep learning provided in an embodiment of the present application;

[0045] Figure 2 A flowchart of a method for generating early warning information provided in an embodiment of the present application;

[0046] Figure 3 A schematic diagram of a deep learning-based network traffic anomaly monitoring and early warning system provided in an embodiment of the present application.

[0047] Reference numerals: 10, acquisition unit; 20, first processing unit; 30, second processing unit; 40, generation unit. DETAILED DESCRIPTION

[0048] The technical solutions in the embodiments of the present invention will be described clearly and completely below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, rather than all the embodiments.

[0049] See also Figure 1 As shown, it is a flowchart of a method for monitoring and warning network traffic anomalies based on deep learning provided by an embodiment of the present application, and the method includes steps S101 to S104, wherein:

[0050] S101: Collect target network traffic and perform preprocessing to generate serialized traffic data;

[0051] S102: Inputting the traffic data into a long sequence Transformer variant model with a sparse attention mechanism, characterizing the temporal dependency and global context features of the traffic data, and generating a first feature representation;

[0052] S103: Input the first feature representation into a capsule neural network, and obtain a multi-level deep feature representation through a dynamic routing mechanism of capsule units;

[0053] S104: Perform anomaly determination on the multi-level deep feature representation, generate anomaly determination results, and generate warning information based on the anomaly determination results; wherein the anomaly determination results include: anomaly scores and category labels.

[0054] In a specific implementation, in S101, the data traffic in the target network environment is collected, and the collected original data packets are deduplicated and preliminarily screened, retaining key fields such as source IP, destination IP, port number, protocol type, timestamp, packet size, etc. By integrating the data packets in the same session or connection in time sequence, multiple packets can be serialized into an ordered and relatively complete traffic data. If necessary, the traffic characteristics can also be standardized or normalized to ensure a balanced distribution of feature values ​​and reduce the impact of extreme values ​​on subsequent model training. If labeled data is used in the model training stage, the data can be annotated at this time; if it is an unsupervised or self-supervised detection method, subsequent analysis can be performed without additional label information.

[0055] In S102, the serialized traffic data obtained above is input into a long-sequence Transformer variant model with a sparse attention mechanism to characterize the temporal dependencies and global context features contained therein at a deeper level. The variant model can process ultra-long sequences by combining local windows with global attention, thereby effectively reducing the computational complexity while retaining long-distance dependency information. In actual implementation, the input sequence is usually vectorized and embedded and position encoding is superimposed so that the Transformer can distinguish packets or fragments of different orders in the traffic. After multiple layers of sparse attention calculations, the model will produce a high-dimensional representation vector that captures the correlation between network traffic at multiple levels such as time, protocol, and context. This high-dimensional representation vector is used as the first feature representation described in the present invention.

[0056] After the first feature representation is obtained, enter S103, input the first feature representation into the capsule neural network, and obtain multi-level deep feature representation through the dynamic routing mechanism of the capsule unit. The capsule network is composed of a series of "capsules" in the form of vectors, and each capsule can represent the existence and importance of a certain protocol mode or potential abnormal mode. The low-level capsules are first activated according to the input features, and then the activation vectors with similar or related properties are iteratively aggregated into high-level capsules through the dynamic routing algorithm. In this way, the key contextual information that the attack chain may span multiple protocols and multiple periods of time can be retained, so that the high-level capsule activation vectors finally output can more comprehensively characterize the complex network traffic characteristics. If targeting specific protocols (such as HTTP, DNS, etc.) or different attack process stages (such as scanning, vulnerability exploitation, lateral movement), the capsule network structure and routing strategy can also be appropriately adjusted to improve the ability to capture the corresponding abnormal patterns.

[0057] Finally, in S104, the above-mentioned multi-level deep feature representation is subjected to anomaly determination, and corresponding warning information is generated based on the determination result. Specifically, one or more classifiers (such as Softmax, Sigmoid, SVM, etc.) can be used to determine the high-dimensional vector output by the capsule network, and output anomaly scores and category labels. When the anomaly score is higher than a certain set threshold or matches a known attack type, the system immediately generates an early warning signal, and informs the security administrator to conduct inspection and defense in a variety of ways (such as logging, visual interface, API callback, etc.). For different threat levels, multiple threshold limits can be set at this stage to divide the risk level, thereby achieving differentiated early warning strategies. Based on these output results, network security personnel can also track the source IP, destination IP, protocol and port information, and timestamp records, and include suspicious traffic in key monitoring or automatic blocking. If the model needs to be continuously updated to deal with emerging new threats, new labeled data can be regularly collected according to business needs or operation and maintenance cycles to fine-tune the parameters of Transformer and capsule networks.

[0058] Through the above steps, the present invention realizes comprehensive in-depth analysis on the time series processing, protocol association and anomaly determination of high-dimensional network traffic data. Compared with the traditional method of relying on rules or statistical thresholds, the sparse attention mechanism used in the present invention can effectively capture long-term time series dependency information at a lower computational cost, and the multi-level vector representation method of the capsule network can further analyze the complex attack clues hidden in the traffic. In actual deployment, the model depth or capsule network structure can also be flexibly adjusted according to equipment resources and network scale to balance detection performance and system load to meet network security requirements in multiple scenarios. The anomaly determination results and early warning information generated by the above process can help enterprises or operators respond to potential threats in a timely manner and reduce losses.

[0059] In this way, this application forms a complete deep learning anomaly detection and alarm mechanism around the four main stages of network traffic collection and preprocessing, deep learning model (Transformer variant and capsule network) training and reasoning, and anomaly judgment and warning output. Based on the understanding of the core idea of ​​the present invention, those skilled in the art can make appropriate modifications or replacements to the implementation details without departing from the principles of the present invention, which should all fall within the scope of protection of the present invention.

[0060] For example, in a large enterprise, there are usually multiple subnets and departments, and the network size can range from hundreds to thousands of terminals. In order to meet the timely detection of internal threats, this application can be implemented according to the following process:

[0061] Enterprises can deploy traffic collection modules at each core switch or firewall to continuously collect raw data packets entering or leaving the LAN. Agents can also be set up at the front end of key servers to obtain application layer sessions. After deduplication and blacklist filtering of the collected raw data packets, the source IP, destination IP, port number, protocol type and other fields are parsed and integrated into serialized traffic data in chronological order or by session identifier. If labeled data is used for training, the security team can mark known attack traffic; if unsupervised learning is adopted, subsequent modeling can be done without labeling.

[0062] Common connection sessions in enterprise networks span from seconds to minutes, and some complex applications (such as big data transmission or video conferencing) may last longer. The present invention uses a long-sequence Transformer variant model with a sparse attention mechanism to vectorize and position-encode traffic data in multiple time windows to generate a first feature representation that can capture long-range dependencies. In this way, potential scanning, password blasting or hidden connection behaviors can be mined over a longer time series range.

[0063] In order to further identify common behaviors such as lateral movement, privilege escalation, and database detection within the enterprise network, after obtaining the first feature representation, it is input into the capsule network. Through the dynamic routing algorithm, the low-level capsules aggregate the activation vectors of the same source protocol, and the high-level capsules correspond to abnormal patterns at different levels or different links, thereby obtaining a multi-level deep feature representation. In this way, when the attack path spans multiple internal subnets or protocols, clues can also be shown in the high-level capsule output.

[0064] Enterprises can use a set of classifiers (such as Softmax) to score abnormalities in high-level vectors of capsule networks, and immediately generate warnings when they exceed a given threshold or identify specific attack labels. According to actual needs, different threshold levels such as low risk, medium risk, and high risk can also be set to apply more rapid security policies (such as isolation or blocking) to suspicious hosts or connections that pose serious threats. Warning information is usually visualized through the enterprise SOC (Security Operation Center) platform to help security personnel verify and trace the source in a timely manner.

[0065] Among them, the long sequence Transformer of this application can better capture abnormal correlations across time periods or subnets, while the capsule network performs hierarchical analysis at the protocol layer and behavior stage level to help identify complex intranet penetration or data theft attacks.

[0066] For example, this application can also be applied to a larger-scale operator backbone network or cloud data center environment, where the traffic scale, protocol types, and number of user connections are significantly increased, placing higher requirements on the processing efficiency and scalability of the model. The implementation process is as follows:

[0067] Operators or cloud platforms can export key fields such as source / destination IP, port, protocol, number of packets, flow size, etc. at fixed intervals or flow thresholds at backbone routers and virtual switches (vSwitches). Some data centers can also use NetFlow, sFlow, or mirror port technology to collect summary information, and then deduplicate, merge, and organize it in chronological order to form long serialized flow samples. If it is necessary to compare the overall flow levels of different time periods horizontally, preliminary aggregation or splitting can also be performed at this stage to adapt to the requirements of the Transformer variant for sequence length.

[0068] When faced with extremely long or dense connection data, the algorithm complexity can be significantly reduced by pre-dividing the sequence into blocks and then applying sparse attention in each block. Large-scale data centers often have thousands or tens of thousands of concurrent connections. Characterizing global dependencies through long-sequence Transformer variants is conducive to capturing certain large-scale flooding attacks such as DDoS attacks, botnets, and abnormal bandwidth usage. The first feature representation of the model output can be combined with additional information such as time period features to generate a high-dimensional vector.

[0069] To cope with multiple protocols, multiple tenants, and multiple service types, the capsule network can aggregate different traffic features based on protocol features or tenant identifiers at the lower layer, and identify potential abnormal behaviors with distributed attack features or across VLANs in a dynamic routing manner at the higher layer. The multi-level deep feature representation formed in this way can also retain the potential abnormal coupling between different services in a complex connection environment.

[0070] When a suspicious traffic vector is detected that exceeds the threshold or matches a typical attack category (such as DDoS surge, cross-tenant network scanning), an alarm is immediately issued in the operator's NOC (Network Operation Center) or cloud security console. The border firewall or traffic cleaning system can also be linked to limit the speed or drop packets of suspicious connections. For cloud tenants, the alarm information can be pushed to their dedicated management interface, and the tenants decide on subsequent strategies. The deep anomaly detection mechanism provided by the present invention still maintains a high degree of real-time and accuracy in large-scale data scenarios, and is suitable for the complex requirements of high concurrent access to the backbone network and multi-tenant isolation of the cloud platform.

[0071] Among them, this application reduces the processing cost of massive traffic through sparse attention Transformer, and uses capsule network to perform layered feature representation of multi-protocol and multi-tenant traffic, effectively improving the detection capability of sudden and covert attacks in a high-throughput environment.

[0072] As an optional implementation manner, the generating serialized traffic data includes:

[0073] Based on a preset stream data scheme, periodically deriving connection summary information as the stream data;

[0074] De-duplication, merging and correlation analysis are performed on the stream data to obtain key fields;

[0075] The key fields include: source IP, destination IP, port number, protocol type and timestamp;

[0076] The key fields are organized according to time sequence and connection relationship to generate serialized traffic data.

[0077] In specific implementation, NetFlow, sFlow and other flow data solutions can be enabled in advance on network routers, switches or edge devices, and connection summary information can be periodically exported at fixed time intervals or based on traffic triggering. The summary information usually includes records such as source IP, destination IP, source port, destination port, protocol type, number of traffic bytes, number of packets, and flow start and end time.

[0078] In some scenarios, additional fields may be added according to operational or security requirements, such as TCP flag statistics, session duration, application layer protocol identifier, etc. The exported connection profile information is regarded as the initial form of "the flow data".

[0079] It should be noted that in order to ensure the accuracy of subsequent processing, these stream data need to be deduplicated and merged after export, duplicate connection entries need to be removed, and records with the same quintuple and in the same time window need to be merged into one to avoid information redundancy.

[0080] After deduplication and merging, the obtained flow data is subjected to correlation analysis to connect scattered records and obtain key fields for subsequent modeling. Here, key fields usually include source IP, destination IP, port number, protocol type, and timestamp, and may also contain additional identification information (such as connection duration, application layer characteristics) to distinguish traffic patterns at a higher dimension. By utilizing quintuples and derived timestamps, the originally scattered traffic overview can be reconstructed into relatively continuous connection fragments. When traffic records meet the same connection conditions (for example, the source IP, destination IP, port, and protocol are exactly the same) and are adjacent or overlapping in time, they can be connected sequentially to form a more complete traffic data sequence.

[0081] On this basis, for each reconstructed or merged traffic record, the key fields are organized according to the collection order or connection relationship. For example, several records with similar time and the same five-tuple identifier are merged into the same sequence to maintain context continuity on the timeline.

[0082] In addition, if a longer time span is required for analysis in the system, the connection profile information of several minutes, hours or even longer periods can be integrated into a longer serialized record according to business needs. Through the above steps, serialized flow data with time sequence characteristics can be gradually obtained from NetFlow, sFlow and other flow data.

[0083] It is understandable that the export cycle, five-tuple matching strategy or duration threshold can be flexibly set according to the network scale and actual application scenarios to strike a balance between ensuring data integrity and reducing storage overhead. Through this collection and preprocessing method based on the preset flow data solution, it can not only adapt to the massive connection data in large-scale network environments such as operators and cloud data centers, but also play a role in monitoring the dynamic changes of traffic in the enterprise intranet, thereby meeting the needs for efficient serialized traffic data generation in different environments.

[0084] In this way, the high bandwidth occupancy and huge storage pressure caused by collecting and storing all original data packets in a large-scale network environment can be effectively alleviated, while retaining the core connection information to support subsequent long-sequence Transformer and capsule network modeling. Compared with the traditional method of collecting full packets, the periodic export of connection summary information can reduce the redundancy of traffic records while maintaining high accuracy, and improve adaptability to large-scale traffic scenarios. This collection and preprocessing method based on streaming data solutions is not only suitable for massive traffic environments such as operators and cloud data centers, but also meets the needs of timely monitoring of dynamic changes in traffic in corporate intranets.

[0085] As an optional implementation manner, generating the first feature representation includes:

[0086] Dividing the serialized traffic data into a number of blocks of equal size;

[0087] Perform local attention calculation on sequence elements in each block to obtain local attention scores;

[0088] Randomly select some elements as global connection tokens to establish global attention associations between adjacent blocks and obtain global attention scores;

[0089] An attention matrix is ​​constructed based on the local attention score and the global attention score, and the traffic data is multi-layer encoded to generate the first feature representation.

[0090] In order to balance the high computational overhead brought by global attention and the ability to capture remote dependency information in long sequence scenarios, this application introduces a sparse attention strategy of "blocking + random global connection" when performing attention modeling on the serialized traffic data.

[0091] The traditional approach of directly applying global attention to the entire long sequence often requires huge storage and computing resources, especially when the network traffic data may reach thousands or even tens of thousands of words in length, which is more likely to cause memory overflow or a significant decrease in processing efficiency.

[0092] This optional implementation first divides the serialized traffic data into several blocks of equal size, then performs local attention calculations within the blocks, and randomly selects some elements as global connection tokens to establish attention associations across blocks. This effectively reduces the size of the attention matrix while still retaining cross-block remote dependency information, taking into account both computational efficiency and global information acquisition.

[0093] In the specific implementation, the serialized traffic data is first divided into several continuous blocks according to the preset block size parameters. Each block can contain several network traffic records or data slices, which is usually determined by the user based on the model scale and hardware resources.

[0094] Subsequently, a local attention mechanism is used within each block to calculate the query, key, and value and perform dot product operations on the elements at each position in the block (such as the traffic feature vector) to obtain the corresponding local attention score. Since the attention range is limited to the block at this time, the computational overhead is significantly lower than directly performing global attention on all sequence elements. By aggregating features with strong temporal or contextual associations within the block, local dependency information can be extracted at a lower cost.

[0095] In order to further take into account long-distance dependencies, it is necessary to establish global attention associations between adjacent blocks. Specifically, some elements can be randomly selected from the multiple blocks divided as global connection tokens, so that the selected elements can be dot-product matched across blocks in subsequent attention calculations, thereby capturing global information across block boundaries.

[0096] If there are many blocks, the remote dependencies can be gradually transferred to the internal elements of the corresponding blocks in multiple iterations through randomly assigned global connection tokens; for unselected elements, local attention calculations are still performed only within the range of the block in which they are located. In this way, a balance can be achieved between overall memory capacity and computational complexity: the attention matrix will not be expanded to an unmanageable scale when the sequence is extremely long, and sufficient cross-block information can be retained with the help of some global connection tokens.

[0097] After completing the above local and global attention score calculations, the local attention scores are combined with the global attention scores into a sparse attention matrix to perform multi-layer encoding on the entire sequence.

[0098] Among them, each layer of encoder will perform weighted summation on the input vector based on the sparse attention matrix and perform nonlinear transformation through the feedforward network, and then pass the result to the next layer.

[0099] If you need to focus on different feature subspaces or temporal relationships multiple times, you can stack multiple attention heads and reuse or change the attention strategy in different encoding layers. After multi-layer encoding is completed, the output set of vectors is the first feature representation, which retains the comprehensive information about network traffic between local windows and remote dependencies.

[0100] Through this optional implementation, it is possible to take into account both local pattern capture and long-range dependencies across blocks when processing ultra-long sequences, which not only reduces the burden of global attention on computing and memory resources, but also maintains the ability to identify key long-range interactions, providing a more complete sequence representation for subsequent capsule networks or other downstream modules.

[0101] For example, in Internet streaming services (such as video on demand and live streaming), there are often long-lasting and large-volume session connections between users and servers. Traditional global attention schemes will bring high computational overhead or memory usage when processing hundreds or thousands of streaming media traffic data. In order to balance global information capture and computational efficiency, the following approach can be adopted:

[0102] Operators or cloud platforms first collect and pre-process streaming sessions between users and servers at the edge of the network or near content distribution nodes (CDNs), and construct relevant fields (such as timestamps, packet sizes, protocol types, etc.) into serialized traffic data. Since it may involve viewing / pushing behaviors of several minutes or even hours, the sequence length increases significantly.

[0103] Each long sequence is evenly divided into several blocks according to a predetermined block size (e.g., 128, 256, 512, etc.). Local attention calculation is performed for several records in each block (e.g., time-close packets or summary entries). Within a block, the correlation between elements is usually high (e.g., the same small segment of video data or adjacent time slices), so local attention can fully exploit local dependency information while avoiding the high load of global calculation.

[0104] Among the multiple blocks obtained after segmentation, in order to capture long-distance dependencies and cross-period anomalies (such as sudden bandwidth jitter, malicious injection, etc.), some elements are randomly selected as global connection tokens, allowing these tokens to participate in cross-block attention calculations.

[0105] For example, 5% to 10% of the elements are retained in each block to establish global associations in adjacent blocks. In this way, once some features in block A appear suspicious, blocks B and C can gradually perceive the information after multiple layers of sparse attention stacking.

[0106] The local attention score is combined with the global attention score to form a relatively sparse attention matrix, which is used to perform weighted aggregation and nonlinear transformation on the feature vector of each block. Since the number of cross-block elements is relatively limited, the matrix dimension and the amount of computation can be kept within a controllable range. After stacking multiple layers, the generated first feature representation can take into account both short-range and long-range dependencies, and effectively identify possible abnormal patterns in the video stream (such as abnormally high bitrates, forged data blocks, etc.).

[0107] This first feature representation can be combined with subsequent capsule networks or other detection modules to identify potential attacks and anomalies during the transmission process on the user side or server side, such as large-scale brushing, forged requests, video tampering, abnormally high traffic usage, etc. Once the system detects an anomaly, it can issue an early warning and notify the operation and maintenance personnel to block, limit the speed or adjust the strategy in a timely manner.

[0108] As an optional implementation, the obtaining of multi-level deep feature representation through the dynamic routing mechanism of capsule units includes:

[0109] Setting a plurality of protocol capsules at a lower layer of the capsule network for generating initial activation vectors related to different protocols based on the first feature representation;

[0110] A plurality of attack phase capsules are arranged at the high layer of the capsule network, and an update vector outputted by the protocol capsule after iteration is received during the dynamic routing process; wherein the update vector is an activation vector generated by the protocol capsule according to the initial activation vector and high layer feedback in the dynamic routing iteration, and is used to transmit protocol-related features to the attack phase capsule;

[0111] The attack stage capsules are used to characterize different abnormal stage features respectively; the abnormal stage features include: scanning, vulnerability exploitation, lateral movement and data transmission;

[0112] Multi-level deep feature representations are generated at the output of high-level capsules.

[0113] In the specific implementation process, the first feature representation output by the previous step (such as a long sequence Transformer variant) can be mapped to the low-level protocol capsule of the capsule network according to the protocol category or feature block. Specifically, several protocol capsules can be pre-defined, such as TCP capsules, UDP capsules, HTTP capsules, DNS capsules, etc. Each protocol capsule has a set of learnable weight matrices and vectorized representations for generating initial activation vectors related to its corresponding protocol.

[0114] For example, a vector segment that is determined to be a TCP feature in the first feature representation will be sent to the TCP capsule, and a segment marked as an HTTP feature will be sent to the HTTP capsule. In this way, different protocol capsules can more specifically aggregate the context information of the protocol at the low level and generate the corresponding preliminary activation output.

[0115] Subsequently, the output vector of the low-level protocol capsule is iteratively passed to the high-level attack phase capsule through a dynamic routing algorithm. The high-level capsule can be designed to establish corresponding capsule groups for different attack phases (scanning, vulnerability exploitation, lateral movement, and data outflow).

[0116] For example, capsules in the scanning phase may be sensitive to the characteristics of network port detection or a surge in connection requests, while capsules in the vulnerability exploitation phase focus on identifying features such as abnormal request loads and suspicious command injections. The dynamic routing process may include multiple iterations, and each iteration updates the coupling coefficient based on the similarity between the prediction vector output by each protocol capsule and the activation vector of the high-level attack phase capsule. In each iteration of dynamic routing, the protocol capsule will modify or recalculate its initial activation vector based on the feedback information and similarity calculation results of the high-level attack phase capsule, thereby generating an updated vector, and pass the updated vector to the high-level attack phase capsule for the next iteration.

[0117] Among them, if the output of the protocol capsule has a high degree of match with the characteristics of a capsule in a certain attack stage, the coupling coefficient will increase with iteration, so that more information will be gathered in the corresponding high-level capsule; otherwise, it will gradually decay. After each coupling coefficient update, the protocol capsule will weight or recalculate its output vector accordingly to form an updated vector. This routing mechanism can "route" the truly relevant protocol activation vector to the appropriate attack stage capsule, thereby obtaining a more accurate anomaly representation.

[0118] After multiple iterations, each high-level attack phase capsule will output its refined activation vector, which can be regarded as the degree of representation of the attack phase in the current traffic data. If there is a strong activation for one of the phases such as scanning, vulnerability exploitation, lateral movement, or data outflow, it means that there is a high probability of an abnormal pattern corresponding to the attack phase in the input data of this protocol or multi-protocol combination. By combining the activation vectors of all high-level capsules, a multi-level deep feature representation can be obtained at the output of the capsule network. These vectors also contain an aggregated understanding of different protocol levels and different attack phases.

[0119] During the detection or downstream classification process, if the activation of one of the paths increases significantly, it can prompt the security operation and maintenance personnel which attack stage needs the most attention, and further cooperate with the system log, threat intelligence or security policy to block and alarm. Through the mapping of protocol capsules and attack stage capsules, the present invention can take into account the differences in protocols and the diversity of attack chains under the same model framework, and improve the accuracy and interpretability of identifying complex network attacks.

[0120] For example, in the scenario of monitoring large-scale enterprise intranet traffic, traffic collection and first feature representation can be obtained at multiple core switches or firewalls, and the first feature representation can be input into the capsule network. To achieve multi-level deep feature representation, the capsule network includes bottom-level protocol capsules and high-level attack stage capsules, and completes the deep analysis of protocol features and abnormal stages through dynamic routing iteration.

[0121] The system first maps different protocol elements (such as TCP, UDP, HTTP, DNS, etc.) in the serialized traffic data to corresponding protocol capsules. Each protocol capsule analyzes the first feature representation of the input based on its own internal learnable weights to generate an initial activation vector that represents the protocol context information. For example, the TCP capsule extracts feature activations related to session control and connection status; the HTTP capsule aggregates information such as abnormal patterns in the HTTP header or payload.

[0122] After generating the initial activation vector, the underlying protocol capsule interacts with the high-level attack phase capsule (scanning, vulnerability exploitation, lateral movement, data outflow) through multiple rounds of iterations. In each round of iteration, the high-level attack phase capsule will focus on higher risk or more matching protocol features based on similarity or coupling coefficient and pass feedback to the underlying capsule.

[0123] After receiving the feedback information, the bottom-level protocol capsule will modify or recalculate the original initial activation vector to generate a new update vector. This update vector will focus more on the feature elements in the protocol capsule that are highly similar to the specific attack stage, and then pass it to the high-level attack stage capsule again in the next round of iteration.

[0124] Through such iterations, the bottom-level output is adjusted each time according to the feedback from the upper level, so that truly suspicious or abnormal elements with obvious protocol characteristics can be more prominently reflected in the update vector. With the continuous transmission of the bottom-level protocol capsule update vector, the characteristics of the capsules in the high-level attack stage, such as scanning, vulnerability exploitation, lateral movement, and data transmission, will be gradually strengthened or suppressed.

[0125] For example, if the update vectors generated by the protocol capsule aggregate a large number of characteristics of a port detection or a surge in connection requests in a short period of time, the high-level scanning phase capsule will increase its activation intensity; if suspicious command injection or SQL keywords are detected, the vulnerability exploitation phase capsule may also be triggered and enhanced.

[0126] When the number of iterations reaches the preset threshold or converges, each attack phase capsule outputs the final activation vector, and the high-level capsule network integrates these vectors into a multi-level deep feature representation. In this representation, not only is the refinement of the protocol layer context by each protocol capsule retained, but it can also highlight which attack phase capsule is significantly activated in the current traffic combination. As a result, the system can accurately indicate potential abnormal protocol paths and attack chain stages.

[0127] See also Figure 2 , Figure 2 A flowchart of a method for generating early warning information provided in an embodiment of the present application, as an optional implementation, the generating early warning information includes steps S201 to S204, wherein:

[0128] S201: Building a reference traffic sample library based on normal traffic data;

[0129] S202: Determine the similarity between the current multi-level deep feature representation and the reference traffic sample library;

[0130] S203: In response to the similarity being lower than a preset threshold, label determination is performed in combination with an external threat intelligence library;

[0131] S204: Outputting the anomaly score and category label according to the determination result, and generating different levels of warning information according to preset multi-level thresholds.

[0132] In specific implementation, a section of traffic data that has been confirmed by security audits to contain no attack behavior or rare anomalies can be first regarded as "normal traffic", and its key features (such as the first feature representation or some vector representations of the capsule network output) can be collected to build a reference traffic sample library. The sample library can be generated offline or in the initial deployment stage, or it can be dynamically updated in an incremental manner to continuously enrich the coverage of normal traffic patterns in long-term operation.

[0133] To improve efficiency, the traffic vectors included in the sample library can be clustered or hashed to facilitate fast retrieval and similarity calculation.

[0134] After obtaining the multi-level deep feature representation, the similarity between the current traffic representation and the reference traffic sample library is calculated. Common measurement methods include Euclidean distance, cosine similarity, Mahalanobis distance, etc., and the indicator that can best distinguish normal distribution from abnormal distribution can also be selected in combination with actual business scenarios or experimental results.

[0135] If the similarity is higher than the preset threshold, it is usually considered to be close to the normal mode, and no alarm is triggered or only recorded; conversely, when the similarity is lower than the preset threshold, it means that the current traffic is significantly different from the existing normal distribution. At this time, it can be further combined with the external threat intelligence library for label judgment.

[0136] In this judgment phase, more specific attack information can be annotated for suspected abnormal traffic based on the matching results of external threat intelligence databases (such as known malicious IP blacklists, suspicious domain name C2 lists, common attack feature signatures, etc.). If a high-risk label is hit, such as a blacklisted IP or a widely monitored ransomware C2 node, the anomaly score can be moderately increased or directly classified as a high-risk category; if no external intelligence is matched, but there is still a large deviation from the normal distribution, it can be temporarily classified as an unknown anomaly or gray list, pending subsequent audit confirmation.

[0137] According to the final judgment result, the corresponding anomaly score and category label can be output, and different levels of warning information can be generated according to the preset multi-level thresholds.

[0138] For example, when the anomaly score is between the medium and low risk thresholds, it is only recorded in the security log and a prompt is sent to the administrator; when the anomaly score exceeds the high risk threshold or matches a known attack tag, an emergency alarm is issued and automatic blocking, isolation or speed limit policies are triggered. Administrators can also view relevant IP, port, attack stage information and traffic feature vectors on the security platform for further tracing analysis.

[0139] Through this optional method of building a reference library based on normal traffic data and assisting in judgment with the help of an external threat intelligence library, potential threats can still be identified through similarity measurement when there are only a small number of labeled abnormal samples or a lack of complete attack signatures. At the same time, the intelligence library can be used to improve the accurate recognition rate of known malicious behaviors, providing a more flexible and effective means for multi-level security monitoring and response.

[0140] In existing anomaly determination schemes based on reference libraries, it is usually necessary to collect and label a certain amount of "normal traffic" data in advance to support subsequent similarity or comparative learning methods. However, if the "normal traffic" itself contains hidden malicious behavior, or its collection source lacks strict auditing, once the contaminated data is included in the reference library, subsequent anomaly detection will be difficult to maintain sufficient accuracy, and may even result in high false positives or omissions. Some traditional schemes rely only on simple traffic screening or short-term monitoring to assert that traffic is "attack-free", but it is difficult to fully guarantee the purity of the reference library in large-scale networks or complex multi-protocol environments.

[0141] As an optional implementation, the normal traffic is traffic that has been confirmed by security audit to have no attack behavior and is collected within a pre-set no-attack time period.

[0142] During the specific implementation process, one or more security audits can be conducted on the target network environment (such as the corporate intranet, cloud data center, or operator backbone network) to check and confirm that there are no obvious signs of attacks or abnormal activities during a specific time period. For example, during regular working days or nighttime hours, traffic can be monitored in multiple dimensions through a variety of security tools (such as firewall logs, IDS / IPS detection, endpoint security detection, etc.) to confirm that no attack behaviors such as large-scale scanning, abnormal transmission, or suspicious command injection have been detected. When the audit results do not find known threats or high-risk events, and the environment has not received external threat intelligence alerts recently, the traffic during this period can be regarded as a "no-attack period."

[0143] After determining the attack-free period, the traffic in the network is collected. The specific operations can be consistent with the previous steps (such as preprocessing, de-redundancy, feature extraction, etc.), and the final traffic data or feature vector is marked to indicate that it comes from "normal traffic" confirmed by security audits.

[0144] To ensure reliability, multiple sampling or cross-validation methods can also be used to collect traffic within the period to ensure statistical coverage. If a few suspicious connections or accesses are found, operation and maintenance personnel can exclude them during the audit phase to avoid mixing potential attacks into normal traffic data.

[0145] By collecting data during the attack-free period set in this way and combining multiple security audit methods to confirm that there is no attack behavior, a batch of high-confidence normal traffic records can be obtained relatively accurately. When building a reference traffic sample library later, this part of the data can be directly used as a benchmark for normal traffic characteristics to compare and match suspicious characteristics detected during subsequent runtime. With this kind of attack-free traffic with audit endorsement and clear time periods, the accuracy of similarity determination and anomaly score calculation can be effectively improved, and a purer normal data baseline is provided for security analysis, thereby reducing false positives and enhancing the model's ability to identify truly abnormal behaviors.

[0146] In existing long sequence processing solutions, global connection tokens are usually selected in a completely random manner to capture long-distance dependencies across block boundaries or local windows. However, such practices may ignore prior knowledge about port numbers, protocol types, keywords, etc. in the security field, resulting in insufficient attention to key high-risk elements. Especially in network security scenarios, if the randomly selected global token happens to be irrelevant to potential attack features, it may lead to insufficient mining of remote dependencies, thereby missing some hidden abnormal behaviors.

[0147] As an optional implementation, before randomly selecting some elements as global connection tokens, the method further includes:

[0148] Based on at least one of the port number, the protocol type and the keyword, the serialized traffic data is screened to determine a high-risk Token, and the high-risk Token is included in a candidate set of global connection tokens.

[0149] In a specific implementation, a port, protocol or keyword level scan may be first performed on the serialized traffic data.

[0150] For example, you can define a "high-risk port list", such as common blasting or Trojan ports (22, 3389, 445, etc.), or identify certain suspicious request headers or payload keywords (such as "cmd", "shell", "sql", etc.) in specific protocol types (such as HTTP / HTTPS, DNS, FTP, etc.) according to security policies. Mark or score the traffic data tokens that meet the above conditions, and place these "high-risk tokens" in a candidate set.

[0151] Subsequently, when it is necessary to randomly select some elements as global connection tokens, a certain proportion of tokens can be preferentially or forcibly extracted from the candidate set, or directly incorporated into the global connection token set, so that these tokens with potential abnormal signs can be associated with attention across blocks.

[0152] In this way, not only the universality of random global connection tokens is preserved, but also tokens with significant risk characteristics in serialized traffic data are not ignored. When long-distance dependencies are subsequently calculated between blocks, these high-risk tokens can generate stronger cross-block attention links, thereby amplifying possible abnormal feature signals, which is helpful for subsequent feature extraction or abnormality determination.

[0153] This method effectively improves the sensitivity to potential attack clues under the long sequence Transformer variant framework of the present invention, and has significant value in identifying port blasting, malicious protocol abuse or keyword injection attacks. Those skilled in the art can flexibly specify high-risk ports, protocols or keyword lists according to different network environments and security requirements, and combine the configuration parameters of the random strategy (such as the proportion extracted from the candidate set) to balance the attention to high-risk features and the balance of random distribution.

[0154] In existing long sequence processing methods, the same sparse attention configuration is often used uniformly in all Transformer encoding layers (such as all local windows or all random global). Although this can simplify the implementation, it is difficult to take into account the differentiated needs of different layers for local feature aggregation and long-distance dependency extraction. Especially in multi-protocol and multi-time span network traffic, different layers may have different preferences for short-range context or cross-time association.

[0155] Exemplarily, the random strategy can be set in combination with the overall distribution of the candidate set of high-risk tokens and ordinary tokens. First, define one or more ratio thresholds to determine the proportion of tokens drawn from the high-risk candidate set in each block or in the global scope of the sequence. For example, you can specify in the system parameters that "guarantee that high-risk candidate tokens account for 10% to 30% of the global connection tokens", and the rest are randomly generated by ordinary tokens.

[0156] When starting to select global connection tokens, you can randomly select from the candidate set and the common token set according to the set ratio. If you need more flexibility, you can also use a two-step extraction: first extract the candidate set with a higher priority (such as a larger sampling probability), and then supplement the common token set in a completely random manner, so as to maintain a stable number of global connection tokens.

[0157] In addition, within the candidate set, different weights can be assigned according to the risk level of the token (secondary screening). If certain ports, protocols or keywords are assessed as extremely high risk in the security policy, they can be preferentially or required to be included in the global connection token, while high-risk tokens with relatively low risks can be extracted with a lower probability.

[0158] As an optional implementation, the multi-layer encoding of the traffic data further includes:

[0159] Set up sparse attention mechanisms for different Transformer encoder layers; where:

[0160] Local window attention is used at the first level;

[0161] Use random global attention in the middle layers;

[0162] The global connection scope is divided based on the protocol type or time dimension at least one layer.

[0163] In the specific implementation, local window attention is used in the first layer to model the close dependencies between adjacent or time-continuous elements in the traffic data. In this layer, only the relationship between local adjacent fields (such as a small time window) is concerned to effectively aggregate basic features and reduce the computational burden; in the middle layer, it switches to random global attention, so that distant elements in the sequence can also establish associations with a certain probability, thereby obtaining remote context information at the block or window boundary of the network traffic. Through this random global mechanism, dependencies across multiple time periods or protocol fragments can be gradually propagated to the vector representation of the middle layer.

[0164] In addition, at least one layer (which may be placed in a deeper layer or in parallel with the middle layer) further divides the global connection scope based on the protocol type or the time dimension.

[0165] For example, a part of the sequence elements can be divided into the same protocol subspace for global association, or all elements in the same period can be linked across blocks according to time slices. In this way, when identifying protocol switching or consistency across time periods, the model can focus more on aggregating relevant features without being disturbed by other irrelevant subsequences.

[0166] This differentiated sparse attention configuration inherits the advantages of local windows and random global attention, and can further refine the global attention scope through protocol or time division at a specific level, thereby improving the ability to capture multi-protocol and large-span anomalies.

[0167] After multiple layers of encoding are completed, the model can fuse the local features aggregated by the first layer, the long-range dependencies captured by the middle layer of random global attention, and the cross-domain information based on protocol / time division to form the final first feature representation.

[0168] This scheme shows the advantages of paying equal attention to local details and large-scale associations in network traffic scenarios by setting sparse attention in a hierarchical and strategic manner. For large-scale or highly heterogeneous network traffic analysis tasks, this optional implementation can effectively reduce the huge overhead caused by the large-scale deployment of global attention in all layers, while retaining the fine modeling of long-distance dependencies or specific protocol-related aggregations in key layers, thereby achieving better recognition of complex sequence anomalies.

[0169] In existing capsule network designs, protocol capsules and attack phase capsules are often viewed as two layers or two unidirectional mapping relationships: the update vector is first output from the protocol capsule, and then received by the attack phase capsule to generate the final activation result. Although such a unidirectional structure can partially reflect multi-level information, it is difficult to make targeted corrections to the underlying protocol activation under the feedback of the attack phase capsule. Once new key information is discovered at a high level (such as the strength of attack features at a specific stage), it cannot be passed back to the protocol capsule for secondary routing, resulting in insufficient recognition of multi-protocol, multi-stage complex attacks.

[0170] As an optional implementation, the present application also includes:

[0171] The protocol capsule and the attack phase capsule are set at the same time in the high layer of the capsule network, and a multi-path cross-coupling relationship is established through a dynamic routing algorithm, so that the update vector output by the protocol capsule and the initial activation vector of the attack phase capsule are transmitted to each other.

[0172] In the specific implementation, in the high-level capsule network, each protocol capsule is first retained to aggregate the feature vector outputs of different protocols; at the same time, several attack stage capsules (scanning, vulnerability exploitation, lateral movement, data transmission, etc.) are set.

[0173] In the dynamic routing process, the update vector output by the protocol capsule will be passed to the corresponding attack phase capsule to measure the similarity and update the coupling coefficient; at the same time, the initial activation vector of each attack phase capsule can also reversely affect the activation weight of the protocol capsule. Through this multi-way cross-coupling mechanism, if a capsule in an attack phase shows a strong response to a certain protocol feature, it is possible to improve the coupling coefficient of the corresponding protocol capsule in the next iteration, and vice versa.

[0174] In this way, a closed loop of two-way learning and information updating is formed between the protocol capsule and the attack stage capsule, which enhances the ability to capture the association between protocols and stage progression in complex attack chains.

[0175] After completing multiple rounds of cross-coupling iterations, both the protocol capsules and the attack phase capsules will output more accurate activation vectors with a higher degree of matching, which can be further aggregated or projected into the multi-level deep feature representation described in the present invention for anomaly identification and subsequent early warning. For those advanced persistent threats (APTs) or combined attack scenarios across multiple protocols and stages, this two-way information transmission mechanism can effectively integrate suspicious vectors under each protocol, and compete for activation or reinforce each other in capsules at different attack stages, thereby helping security analysts quickly locate multi-protocol, multi-node attack chains. The number and coupling strategy parameters of protocol capsules and attack phase capsules can also be flexibly adjusted according to actual needs during deployment to balance recognition accuracy and computational overhead.

[0176] For example, in an industrial control system (ICS) or factory automation environment, different devices often use multiple network protocols (such as Modbus, OPC UA, S7, etc.) to exchange data. If an attacker wants to successfully penetrate and manipulate the production process, they usually need to go through network scanning, vulnerability exploitation, privilege escalation, and the final data transmission or destruction phase in sequence, forming a multi-stage attack chain. For security operations personnel, it is crucial to detect these multi-protocol, multi-stage attacks in a timely manner.

[0177] The industrial control traffic within the factory can first pass through the front-end deep learning model (such as Transformer) to convert the time-series data into the first feature representation. Then set up several "protocol capsules" in the capsule network high layer, such as "Modbus capsule", "OPC UA capsule", "S7 capsule" and "TCP / IP auxiliary capsule", etc., to aggregate the activation vectors for each specific protocol. If some traffic fragments are identified as consistent with the Modbus communication mode, they are preferentially routed to the Modbus capsule; if OPC UA or S7 related traffic appears, they enter the corresponding capsules respectively, so that the features of different protocols can be aggregated and weighted activated separately.

[0178] At the same time, multiple attack stage capsules are retained at the high level, such as "scanning stage capsule", "vulnerability exploitation stage capsule", "lateral movement stage capsule" and "data destruction / exfiltration stage capsule". Each attack stage capsule is configured with weights and activation thresholds to capture the typical characteristics of the corresponding stage behavior. For example, the "scanning stage capsule" will be more sensitive to traffic that quickly traverses IP or ports, and the "vulnerability exploitation stage capsule" may be more sensitive to suspicious instruction injection or function code anomalies.

[0179] When the system performs dynamic routing, the update vector output by the low-level protocol capsule is transmitted to the initial activation vector of the attack phase capsule. Once the attack phase capsule detects typical "vulnerability exploitation" or "lateral movement" features, the corresponding activation vector may reversely affect the protocol capsule in the iteration, increasing their attention to the protocol traffic; if a protocol capsule detects a low similarity with a specific attack phase feature, it will gradually weaken the coupling coefficient to the high-level capsule in the next round of routing. Through multiple iterations of convergence, this two-way feedback can help the system achieve more refined feature aggregation and judgment when identifying attacks across multiple protocols and multiple stages.

[0180] Finally, at the high-level output end, both the protocol capsule and the attack phase capsule will generate their own activation vectors; the system can further fuse these activation vectors into a multi-level deep feature representation for the next step of abnormality determination or early warning. If a protocol channel in the factory network suddenly carries lateral mobile scanning behavior, or a specific OPC UA session triggers vulnerability exploitation signs multiple times, the above-mentioned bidirectional routing mechanism will significantly improve the efficiency of capturing these anomalies and show a strong activation level on the attack phase capsule. Based on this, security personnel can accurately track the attacker's specific action path in the ICS network and take timely protective measures.

[0181] Through this multi-way cross-coupling relationship, the industrial protocols and phased attack modes within the factory can be deeply characterized and bidirectionally exchanged in the same capsule network structure. When attacks are lurking or moving laterally at different times and with different protocols, the system can also dynamically update capsule activation to reflect new suspicious points, thereby greatly improving the visualization and rapid detection capabilities of attacks in complex industrial control network scenarios.

[0182] Based on the same inventive concept, the embodiments of the present disclosure also provide a deep learning-based network traffic anomaly monitoring and early warning system corresponding to the deep learning-based network traffic anomaly monitoring and early warning method. Since the principle of the system in the embodiments of the present disclosure to solve the problem is similar to the above-mentioned deep learning-based network traffic anomaly monitoring and early warning method in the embodiments of the present disclosure, the implementation of the system can refer to the implementation of the method, and the repeated parts will not be repeated.

[0183] Reference Figure 3 FIG. 1 is a schematic diagram of a network traffic anomaly monitoring and early warning system based on deep learning provided in an embodiment of the present application, wherein the system includes:

[0184] The collection unit 10 is used to collect target network traffic and perform preprocessing to generate serialized traffic data;

[0185] The first processing unit 20 is configured to input the traffic data into a variant of the long sequence Transformer model with a sparse attention mechanism, characterize the temporal dependence and global context features of the traffic data, and generate a first feature representation;

[0186] The second processing unit 30 is configured to input the first feature representation into a capsule neural network, and obtain a multi-level deep feature representation through the dynamic routing mechanism of the capsule unit;

[0187] The generating unit 40 is configured to perform anomaly determination on the multi-level deep feature representation, generate an anomaly determination result, and generate a warning message based on the anomaly determination result; wherein, the anomaly determination result includes: an anomaly score and a class label.

[0188] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed in the present invention can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present invention.

Claims

1. A network traffic anomaly monitoring and early warning method based on deep learning, characterized in that: include: Collect target network traffic and preprocess it to generate serialized traffic data; Inputting the traffic data into a long sequence Transformer variant model with a sparse attention mechanism, characterizing the temporal dependency and global context features of the traffic data, and generating a first feature representation; Inputting the first feature representation into a capsule neural network, and obtaining a multi-level deep feature representation through a dynamic routing mechanism of capsule units; Anomaly determination is performed on the multi-level deep feature representation to generate anomaly determination results, and warning information is generated based on the anomaly determination results; wherein the anomaly determination results include: anomaly scores and category labels.

2. The method according to claim 1, characterized in that The generating of serialized traffic data comprises: Based on a preset stream data scheme, periodically deriving connection summary information as the stream data; De-duplication, merging and correlation analysis are performed on the stream data to obtain key fields; The key fields include: source IP, destination IP, port number, protocol type and timestamp; The key fields are organized according to time sequence and connection relationship to generate serialized traffic data.

3. The method according to claim 1, characterized in that Generating the first feature representation comprises: Dividing the serialized traffic data into a number of blocks of equal size; Perform local attention calculation on sequence elements in each block to obtain local attention scores; Randomly select some elements as global connection tokens to establish global attention associations between adjacent blocks and obtain global attention scores; An attention matrix is ​​constructed based on the local attention score and the global attention score, and the traffic data is multi-layer encoded to generate the first feature representation.

4. The method according to claim 3, characterized in that: The method of obtaining a multi-level deep feature representation through a dynamic routing mechanism of capsule units includes: Setting a plurality of protocol capsules at a lower layer of the capsule network for generating initial activation vectors related to different protocols based on the first feature representation; A plurality of attack phase capsules are arranged at the high layer of the capsule network, and an update vector outputted by the protocol capsule after iteration is received during the dynamic routing process; wherein the update vector is an activation vector generated by the protocol capsule according to the initial activation vector and high layer feedback in the dynamic routing iteration, and is used to transmit protocol-related features to the attack phase capsule; The attack stage capsules are used to characterize different abnormal stage features respectively; the abnormal stage features include: scanning, vulnerability exploitation, lateral movement and data transmission; Multi-level deep feature representations are generated at the output of high-level capsules.

5. The method according to claim 4, characterized in that The generating of early warning information includes: Based on normal traffic data, build a reference traffic sample library; Determine the similarity between the current multi-level deep feature representation and the reference traffic sample library; In response to the similarity being lower than a preset threshold, performing label determination in combination with an external threat intelligence library; The anomaly score and category label are output according to the judgment result, and different levels of warning information are generated according to the preset multi-level thresholds.

6. The method according to claim 5, characterized in that The normal traffic refers to the traffic that has been confirmed by security audit to have no attack behavior and is collected within a pre-set no-attack time period.

7. The method according to claim 3, characterized in that Before randomly selecting some elements as global connection tokens, it also includes: Based on at least one of the port number, the protocol type and the keyword, the serialized traffic data is screened to determine a high-risk Token, and the high-risk Token is included in a candidate set of global connection tokens.

8. The method according to claim 3, characterized in that The multi-layer encoding of the traffic data further comprises: Set up sparse attention mechanisms for different Transformer encoder layers; where: Local window attention is used at the first level; Use random global attention in the middle layers; The global connection scope is divided based on the protocol type or time dimension at least one layer.

9. The method according to claim 4, characterized in that Also includes: The protocol capsule and the attack phase capsule are set at the same time in the high layer of the capsule network, and a multi-path cross-coupling relationship is established through a dynamic routing algorithm, so that the update vector output by the protocol capsule and the initial activation vector of the attack phase capsule are transmitted to each other.

10. A network traffic anomaly monitoring and early warning system based on deep learning, characterized in that: include: A collection unit, used to collect target network traffic and perform preprocessing to generate serialized traffic data; A first processing unit, configured to input the traffic data into a long sequence Transformer variant model with a sparse attention mechanism, characterize the temporal dependency and global context features of the traffic data, and generate a first feature representation; A second processing unit, configured to input the first feature representation into a capsule neural network, and obtain a multi-level deep feature representation through a dynamic routing mechanism of capsule units; A generating unit is used to perform anomaly determination on the multi-level deep feature representation, generate anomaly determination results, and generate warning information based on the anomaly determination results; wherein the anomaly determination results include: anomaly scores and category labels.

Citation Information

Patent Citations

  • Method for predicting RBP binding site of lncRNA through attention mechanism

    CN112270955A

  • Rotary machinery fault intelligent diagnosis method based on semantics and capsule network

    CN117113214A

  • Remote control Trojan flow detection method based on fusion sequence

    CN117176382A

  • Abnormal traffic identification method based on capsule neural network, medium and equipment

    CN117857213A

  • Safety alarm false alarm identification method based on deep learning and text analysis

    CN119341792A

Cited By

  • Network traffic anomaly classification method and system based on artificial intelligence

    CN120512317A

  • Hydropower station dam safety monitoring data acquisition and transmission system

    CN120975335A

  • Network intrusion detection method, system and equipment based on adaptive entropy sampling and Transform

    CN122348863A