Network traffic deep packet inspection multi-task classification method and system

By identifying and analyzing network traffic data streams, generating topology maps, fitting transmission curves, and constructing a deep packet inspection framework, this approach solves the problems of detection accuracy and real-time performance in complex network environments, and achieves efficient classification and collaborative optimization of multi-task traffic.

CN122053437APending Publication Date: 2026-05-15SHENZHEN RUIWANG YUNLIAN TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610255542.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-03-04
Publication Date
2026-05-15

AI Technical Summary

Technical Problem

Existing deep packet inspection methods struggle to achieve high accuracy, real-time performance, and multi-task collaboration in modern network environments characterized by high dynamism, encryption, and mixed protocols. This results in high false positive rates, large response delays, and an inability to effectively address complex attack behaviors.

Method used

By acquiring traffic data streams in the target network environment, identifying traffic protocols and payload content, generating traffic topology maps, parsing topology data sequences, fitting traffic transmission curves, identifying abnormal traffic peaks and core transmission hotspots, calculating traffic throughput thresholds, constructing a deep packet inspection framework, monitoring classification paths and calculating traffic coordination indices, and reconstructing data coordination tasks to achieve multi-task classification.

Benefits of technology

It improves the efficiency of multi-task traffic detection and classification in complex network environments, ensures the accuracy and real-time performance of the detection process, enhances resilience in dealing with complex traffic scenarios, achieves dynamic balance between resources and tasks, and improves detection efficiency and collaborative optimization of network transmission.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122053437A_ABST
    Figure CN122053437A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of data communication, and discloses a network flow deep packet inspection multi-task classification method and system, and the method comprises the steps: firstly obtaining a target network flow data flow, recognizing the corresponding protocol and load content of the target network flow data flow, and generating a flow topological graph related to the target network flow data flow and the load content; collecting a topological data sequence in the atlas, analyzing a corresponding data transmission type, analyzing a type guiding mode, and fitting a flow transmission curve; identifying an abnormal traffic peak value and a core transmission hotspot in the curve, calculating a traffic throughput threshold and analyzing a network cache state; then, a deep packet detection framework is constructed based on the cache state, a processing delay value and a classification misjudgment rate of a real-time classification path in the framework are collected, and a traffic coordination index is calculated; and finally, reconstructing a data coordination task according to the index, identifying a task classification protocol, and generating a multi-task-oriented classification result of the flow data flow. According to the invention, the detection and classification efficiency of multi-task traffic in a complex network environment can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application relates to a network flow deep packet inspection multi-task classification method and system, and belongs to the technical field of data communication. BACKGROUND

[0002] Network flow deep packet inspection (DPI) is a key technology in network security and flow management, which is used for identifying and analyzing the content and protocol type of network data packets to realize flow classification, threat detection and policy control.

[0003] At present, common deep packet inspection methods mostly rely on fixed rule matching, port identification or single feature analysis, and combine machine learning models for flow classification. However, these methods often cannot adapt to the modern network environment with high dynamics, encryption and mixed protocols, and have problems such as low detection accuracy, poor real-time performance and insufficient multi-task coordination capability, resulting in high misjudgment rate and large response delay, and thus cannot effectively cope with sudden traffic and complex attack behaviors. Therefore, a network flow deep packet inspection multi-task classification method is needed to improve the detection and classification efficiency of multi-task flow in complex network environment. SUMMARY

[0004] The application provides a network flow deep packet inspection multi-task classification method and system, which aims to improve the detection and classification efficiency of multi-task flow in complex network environment.

[0005] To achieve the above purpose, the application provides a network flow deep packet inspection multi-task classification method, which comprises: Obtaining flow data stream in a target network environment, identifying the flow protocol and load content corresponding to the flow data stream, and generating a flow topology graph associated with the flow protocol and the load content; Collecting topology data sequences in the flow topology graph, analyzing the data transmission type corresponding to the topology data sequences, analyzing the type-oriented mode corresponding to the data transmission type, and fitting the flow transmission curve under the type-oriented mode; Identifying abnormal flow peaks and core transmission hotspots in the flow transmission curve, calculating the flow throughput threshold corresponding to the flow data stream based on the abnormal flow peaks and the core transmission hotspots, and analyzing the network cache state of the flow throughput threshold in the target network environment; Based on the network cache state, constructing a deep packet inspection framework corresponding to the target network environment, monitoring real-time classification paths in the deep packet inspection framework, collecting processing delay values and classification misjudgment rates in real-time classification paths, and calculating a flow coordination index corresponding to the flow data stream based on the processing delay values and the classification misjudgment rates; Based on the traffic coordination index, reconstruct the data coordination task corresponding to the traffic data stream, identify the task classification protocol in the data coordination task, and generate the task classification result of the traffic data stream facing multiple tasks based on the task classification protocol.

[0006] Optionally, the generating the traffic topology atlas associated with the traffic protocol and the load content comprises: Parsing the protocol header information in the traffic protocol; Identifying the protocol attribute in the protocol header information; Analyzing the content identifier corresponding to the content feature in the load content; Associating the protocol-content mapping relationship corresponding to the protocol attribute and the content identifier; Based on the protocol-content mapping relationship, generating the traffic topology atlas associated with the traffic protocol and the load content.

[0007] Optionally, the associating the protocol-content mapping relationship corresponding to the protocol attribute and the content identifier comprises: Identifying the key protocol field in the protocol attribute; Extracting the feature value sequence corresponding to the content identifier; According to the predefined rule, the key protocol field and the feature value sequence are matched to obtain the matching field sequence; Analyzing the mapping confidence corresponding to the matching field sequence; Based on the mapping confidence, the protocol-content mapping relationship corresponding to the protocol attribute and the content identifier is generated.

[0008] Optionally, the constructing the deep packet detection framework corresponding to the target network environment based on the network cache state comprises: Analyzing the cache utilization rate corresponding to the network cache state; Based on the cache utilization rate, determine the traffic detection level in the target network environment; Based on the traffic detection level, identify the environment deep packet in the target network environment; Query the performance detection data corresponding to the environment deep packet; Based on the performance detection data, construct the deep packet detection framework corresponding to the target network environment

[0009] Optionally, the identifying the environment deep packet in the target network environment based on the traffic detection level comprises: Parsing the level detection baseline corresponding to the traffic detection level; Deeply associate the level detection baseline with real-time traffic states in the target network environment to obtain a candidate deep packet; Generate a deep packet identification list corresponding to the candidate deep packet; Extract list protocol factors in the deep packet identification list; Based on the list protocol factor, identify the environment deep packet in the target network environment.

[0010] Optionally, the calculation of the traffic throughput threshold corresponding to the traffic data stream based on the abnormal traffic peak and the core transmission hotspot includes: Analyze the amplitude characteristics and occurrence frequency corresponding to the abnormal traffic peak; Determine the traffic upper limit value corresponding to the amplitude characteristics and the occurrence frequency; Extract the sustained transmission rate of the core transmission hotspot; Based on the traffic upper limit value and the sustained transmission rate, determine the throughput load value corresponding to the traffic data stream; Based on the throughput load value, calculate the traffic throughput threshold corresponding to the traffic data stream by the following formula: ; Wherein, The traffic throughput threshold corresponding to the traffic data stream is represented by The network bandwidth capacity is represented by The throughput load value is represented by The average duration of the core transmission hotspot is represented by The occurrence frequency of the abnormal traffic peak is represented by

[0011] Optionally, the calculation of the traffic coordination index corresponding to the traffic data stream based on the processing delay value and the classification misjudgment rate includes: Analyze the synergistic change characteristics associated with the processing delay value and the classification misjudgment rate; Extract the abnormal synergistic mark in the synergistic change characteristics; Based on the abnormal synergistic mark, map the classification coordination link in the deep packet detection framework; Statistical delay fluctuation value and misjudgment fluctuation value in the classification coordination link; Based on the delay fluctuation value and the misjudgment fluctuation value, calculate the traffic coordination index corresponding to the traffic data stream by the following formula: ; Wherein, The traffic coordination index corresponding to the traffic data stream is represented by The total number of the classification coordination link is represented by This represents the quantity index corresponding to the classification coordination step. Indicates the first The delay fluctuation value corresponding to each classification coordination link Indicates the reference delay value. Indicates the first The error fluctuation value corresponding to each classification coordination link This indicates the reference misjudgment rate.

[0012] Optionally, fitting the traffic transmission curve under the type-oriented mode includes: Extract the guidance transmission timing from the guidance mode of the aforementioned type; Based on the aforementioned guidance transmission timing, the traffic sampling frequency corresponding to the type of guidance mode is set; Based on the traffic sampling frequency, the real-time traffic sequence in the guided transmission time sequence is collected; Extract traffic data points from the real-time traffic sequence; Based on the traffic data points, fit the traffic transmission curve under the type-oriented mode.

[0013] Optionally, the step of reconstructing the data coordination task corresponding to the traffic data stream based on the traffic coordination index includes: Analyze the coordination equilibrium threshold corresponding to the traffic coordination index; Based on the coordination and balancing threshold, obtain the data coordination task corresponding to the traffic data stream; Query the task coordination rules corresponding to the data coordination task; Generate the task sequence data in the task coordination rules; Based on the task sequence data, reconstruct the data coordination task corresponding to the traffic data stream.

[0014] To address the aforementioned problems, the present invention also provides a network traffic deep packet inspection multi-task classification system, the system comprising: The topology graph construction module is used to acquire traffic data streams in the target network environment, identify the traffic protocols and load content corresponding to the traffic data streams, and generate a traffic topology graph associated with the traffic protocols and load content. The curve fitting module is used to collect topology data sequences in the traffic topology map, parse the data transmission types corresponding to the topology data sequences, analyze the type-oriented mode corresponding to the data transmission types, and fit the traffic transmission curve under the type-oriented mode. The status resolution module is used to identify abnormal traffic peaks and core transmission hotspots in the traffic transmission curve, calculate the traffic throughput threshold corresponding to the traffic data stream based on the abnormal traffic peaks and the core transmission hotspots, and resolve the network cache status of the traffic throughput threshold in the target network environment. The index calculation module is used to construct a deep packet inspection framework corresponding to the target network environment based on the network cache state, monitor the real-time classification path in the deep packet inspection framework, collect the processing latency value and classification misclassification rate in the real-time classification path, and calculate the traffic coordination index corresponding to the traffic data stream based on the processing latency value and the classification misclassification rate. The result generation module is used to reconstruct the data coordination task corresponding to the traffic data stream based on the traffic coordination index, identify the task classification protocol in the data coordination task, and generate the task classification result of the traffic data stream when it is oriented towards multiple tasks based on the task classification protocol.

[0015] Compared to the problems described in the background technology, this invention, by acquiring traffic data streams in the target network environment, provides raw data support for accurately identifying traffic protocols and analyzing load content. Simultaneously, it allows for a direct understanding of the basic operational status of network traffic, ensuring the accuracy and effectiveness of the subsequent detection and classification process. By collecting topology data sequences from the traffic topology map and analyzing the corresponding data transmission types, this invention transforms the scattered protocol-load association information and device transmission relationships in the map into an ordered and analyzable dataset. This provides complete data support for the subsequent accurate analysis of data transmission types, avoiding analytical biases caused by data fragmentation and improving the efficiency and accuracy of subsequent analysis stages. Furthermore, by identifying abnormal traffic peaks and core transmission hotspots in the traffic transmission curve, this invention can promptly capture traffic anomalies deviating from normal transmission trends, enabling rapid location of potential network threats. This invention provides direct evidence for anomalies such as network attacks and link congestion, avoiding response delays caused by concealed anomalies. This ensures that the subsequent deep packet inspection framework is more aligned with actual network operation, improving overall detection efficiency. Furthermore, based on the network cache state, this invention constructs a deep packet inspection framework corresponding to the target network environment. It can dynamically allocate detection resources based on cache capacity saturation and load pressure, avoiding ineffective loss of detection capabilities when the cache is overloaded. This enhances resilience in handling complex traffic scenarios and achieves synergistic optimization of detection efficiency and network transmission. Finally, based on the traffic coordination index, this invention reconstructs the data coordination tasks corresponding to the traffic data stream. Based on the latency and misjudgment coordination characteristics revealed by the index, it can dynamically adjust the task allocation strategy in each classification coordination stage, allowing high-efficiency stages to bear more load and low-efficiency stages to be optimized first, achieving a dynamic balance between resources and tasks and enhancing the framework's resilience in handling complex traffic. Therefore, the network traffic deep packet inspection multi-task classification method and system provided by this invention can improve the detection and classification efficiency of multi-task traffic in complex network environments. Attached Figure Description

[0016] Figure 1 This is a flowchart illustrating a multi-task classification method for deep packet inspection of network traffic according to an embodiment of the present invention. Figure 2 This is a schematic diagram of traffic coordination logic in a multi-task classification method for deep packet inspection of network traffic according to an embodiment of the present invention; Figure 3 This is a schematic diagram of a module for implementing a network traffic deep packet inspection multi-task classification system according to an embodiment of the present invention.

[0017] The objectives, features, and advantages of this invention will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation

[0018] It should be understood that the specific embodiments described herein are merely illustrative of the invention and are not intended to limit the invention.

[0019] This application provides a multi-task classification method for deep packet inspection of network traffic. The executing entity of this method includes, but is not limited to, at least one of the following electronic devices that can be configured to execute the method provided in this application: a server, a terminal, etc. In other words, the multi-task classification method for deep packet inspection of network traffic can be executed by software or hardware installed on a terminal device or a server device. The server includes, but is not limited to, a single server, a server cluster, a cloud server, or a cloud server cluster.

[0020] Reference Figure 1 The diagram shown is a flowchart illustrating a multi-task classification method for deep packet inspection of network traffic according to an embodiment of the present invention. In this embodiment, the multi-task classification method for deep packet inspection of network traffic includes: S1. Obtain traffic data streams in the target network environment, identify the traffic protocols and load content corresponding to the traffic data streams, and generate a traffic topology map associated with the traffic protocols and load content.

[0021] This invention acquires traffic data streams from the target network environment, providing raw data support for accurate identification of traffic protocols and parsing of load content; at the same time, it allows for a direct understanding of the basic operational status of network traffic, ensuring the accuracy and effectiveness of the entire subsequent detection and classification process.

[0022] The target network environment refers to a specific network space requiring traffic detection and management, encompassing elements such as the network's physical architecture, equipment composition, protocol specifications, and operational scenarios. This includes closed environments like enterprise LANs and data center networks, as well as open scenarios like metropolitan area networks (MANs) and wide area networks (WANs). For example, an enterprise's office network might consist of 50 employee terminals, 8 switches, 2 core routers, and 1 firewall, running protocols such as TCP / IP, HTTP, and FTP, and handling daily office communication and file transfers. This specific network space constitutes the target network environment. The traffic data flow refers to the data generated within the target network environment through network devices and links. A network traffic data stream is a series of consecutive data packets transmitted over a network. These packets contain information such as source address, destination address, protocol type, and payload content, reflecting the dynamic transmission process of data from the sender to the receiver. For example, when a user browses a webpage in a corporate network, 15 HTTP data packets are generated per second, containing webpage request instructions, image data, etc. These packets are transmitted from the user terminal through switches and routers to the target server. This sequence of consecutively transmitted data packets constitutes the traffic data stream. Optionally, the acquisition of the traffic data stream in the target network environment can be achieved through network traffic capture methods, such as using Wireshark software to monitor network data packets to obtain the traffic data stream.

[0023] Furthermore, by identifying the traffic protocol and load content corresponding to the traffic data stream, the present invention can provide accurate feature basis for subsequent steps, avoid analysis deviations caused by missing key information, and clearly distinguish the traffic attributes of different services, reducing misjudgments caused by unclear protocol or load information, so as to ensure the effectiveness of the overall detection process.

[0024] The traffic protocol refers to the rules and standards followed when devices and terminals transmit data in a network. It includes core elements such as data format definitions, interaction process specifications, and error handling mechanisms. It ensures that different devices can understand and correctly parse the transmitted information and is the foundation for orderly data transmission. For example, when a user accesses a news website through a browser, the terminal and the website server communicate using the HTTP / 1.1 protocol. The header of the data packet will clearly indicate the request instruction "GET / news / 202405.html HTTP / 1.1", specify the communication port as 80, and carry fields such as Host and Connection. The server will return a 200 OK response code according to the protocol to ensure accurate transmission of web page data. The payload content refers to the actual data in the network data packet that truly carries the business requirements, excluding the header control information. It is the core part that reflects the business attributes of the traffic and is directly related to the specific purpose of data transmission. It is also the key object for analyzing the essence of traffic in deep packet inspection. For example, when an employee sends a 1.5MB project report via corporate email, the header of the email data packet will contain control information such as source / destination IP, port number, and email protocol identifier, while the payload is the binary data of the report document. This data will be split into multiple data packets for transmission, with each data packet's payload being approximately 1024 bytes. It is this payload content that determines that the traffic belongs to the "email file transfer" type. Optionally, identifying the traffic protocol corresponding to the traffic data stream can be achieved through deep packet inspection methods, such as using the Zeek network analysis framework to parse the data packet header information to obtain the traffic protocol; identifying the payload content corresponding to the traffic data stream can be achieved through payload extraction methods, such as using the Scapy Python library to decode application layer data to obtain the payload content.

[0025] Furthermore, by generating a traffic topology map that associates the traffic protocol with the load content, this invention can integrate the scattered protocol rules and load data characteristics into an intuitive structured view, breaking the isolation between the two information and providing a clear global data association foundation for subsequent analysis of data transmission types and type-oriented patterns.

[0026] The traffic topology map refers to a visualized or structured network traffic association model built based on protocol-content mapping relationships. With nodes (devices, protocol types, load types) and edges (transmission relationships, association rules) as core elements, it intuitively presents the source of traffic, transmission path, and the correspondence between protocols and loads. It is a key carrier for globally understanding traffic characteristics. For example, the map will set nodes such as "client (TCP source port 12345)," "core switch," and "web server (HTTP destination port 80)," and the edges between nodes will be labeled "transmitting HTML content (identifier a1b2), transmission rate 1Mbps." At the same time, different colors will be used to distinguish "web page traffic" and "file transfer traffic," clearly showing the association logic between protocols, loads, and devices.

[0027] As an embodiment of the present invention, generating a traffic topology map associated with the traffic protocol and the load content includes: parsing the protocol header information in the traffic protocol; identifying the protocol attributes in the protocol header information; analyzing the content identifiers corresponding to the content features in the load content; associating the protocol attributes with the protocol-content mapping relationship corresponding to the content identifiers; and generating a traffic topology map associated with the traffic protocol and the load content based on the protocol-content mapping relationship.

[0028] The protocol header information refers to the structured information in the network data packet header used to control data transmission. It contains key parameters required for data parsing, routing, and interaction rules. It is the core basis for devices to identify protocol types and process data packets. It does not contain actual business data and only undertakes transmission control functions. For example, the protocol header information of a TCP protocol data packet will contain parameters such as a 16-bit source port (e.g., 12345), a 16-bit destination port (e.g., 80), a 32-bit sequence number (e.g., 1234567890), and 6-bit flag bits (e.g., the SYN connection request flag). By parsing this information, the receiving end can determine the source, destination, and processing method of the data packet. The protocol attributes refer to the key parameters extracted from the protocol header information that reflect the essential characteristics and operating rules of the protocol, covering protocol type, transmission direction, etc. Port range, connection mechanism, and error handling methods are the core bridges for distinguishing different protocols and associating payload content. For example, the protocol attributes of the HTTP protocol include: application layer protocol type, one-way / two-way transmission direction from client to server, default destination port 80 (HTTP) or 443 (HTTPS), TCP-based connection-oriented mechanism, and 30-second connection timeout threshold. These attributes clearly indicate that the payload content corresponding to this protocol is mostly web page data, providing a basis for subsequent correlation analysis. The content characteristics refer to the unique identifiers in the payload content that reflect business attributes and data types, covering data format, encoding method, key fields, length range, business logic tags, etc., which are the core basis for judging the purpose of payload data and distinguishing different business traffic. For example, the content characteristics of a video stream payload include: H.The payload includes a .264 video encoding format, a frame length of 1024-2048 bytes, a timestamp interval of 40ms (corresponding to 25 frames / second), and a media type field containing "video / mp4". These characteristics directly reflect that the payload belongs to video transmission services, clearly distinguishing it from text, images, and other payloads. The content identifier is a specific identifier used to uniquely mark the characteristics of the payload content. It can be presented in the form of hash value, type encoding, key field summary, etc., and can quickly associate the same or similar payload content. It is the "link" for establishing the correspondence between the protocol and the payload. For example, the content identifier of an HTML webpage payload can be set as the SHA-256 hash value of the key code in its tags (such as a1b2c3d4e5f6a7b8c9d0e1f2a3b4c5d6e7f8a9b0c1d2e3f4a5b6c7d8e9f0a1b2); the content identifier of a JPEG image payload can directly use the JPEG format start code. The identifier 0xFFD8 allows for quick identification of the load as an image type. The protocol-content mapping relationship refers to a structured correspondence formed by formally associating key protocol fields with feature value sequences based on mapping confidence (usually requiring a preset threshold, such as 85%). This clearly defines that "a specific combination of protocol features corresponds to a specific load content type," serving as the core link integrating protocol and load information. For example, when the mapping confidence is 92%, a mapping relationship of "HTTP (GET method, port 80) → HTML webpage load" is generated based on the key protocol fields "HTTP + GET + 80" and the feature value sequence "hash + UTF-8 encoding." This relationship can be directly used to subsequently construct a traffic topology map, ensuring accurate association between protocols and loads. When the protocol attributes are "FTP protocol, control port 21," the corresponding content identifier is "file MD5, file type encoding," forming a mapping relationship of "FTP (port 21) → file identifier → file transfer content."

[0029] Furthermore, parsing the protocol header information in the traffic protocol can be achieved through packet structure parsing methods, such as using Wireshark's protocol decoder to extract protocol fields at each layer to obtain the protocol header information; identifying the protocol attributes in the protocol header information can be achieved through protocol feature extraction methods, such as using a NetFlow collector to analyze source and destination address and port combinations to obtain protocol attributes; analyzing the content identifiers corresponding to content features in the payload content can be achieved through pattern matching algorithms, such as applying the Snort rule engine to detect specific payload feature codes to obtain content identifiers; associating the protocol attributes with the protocol-content mapping relationship corresponding to the content identifiers can be achieved through association rule mining methods, such as using the Apriori algorithm to discover the symbiotic pattern between protocol types and content features to obtain the protocol-content mapping relationship; generating a traffic topology map associated with the traffic protocol and the payload content can be achieved through graph construction algorithms, such as using the Gephi tool to visualize the association network between protocol and content entities to obtain the traffic topology map.

[0030] As another embodiment of the present invention, the association of the protocol attribute and the protocol-content mapping relationship corresponding to the content identifier includes: identifying key protocol fields in the protocol attribute; extracting the feature value sequence corresponding to the content identifier; matching the key protocol fields with the feature value sequence according to predefined rules to obtain a matching field sequence; analyzing the mapping confidence corresponding to the matching field sequence; and generating the protocol-content mapping relationship between the protocol attribute and the content identifier based on the mapping confidence.

[0031] The key protocol fields refer to the core information units extracted from the protocol attributes that determine the protocol's functional positioning and data transmission logic. They are directly related to protocol type identification and payload content matching, and typically include irreplaceable fields such as protocol type identifier, port number, request method, and version number. For example, in the HTTP protocol attributes, key protocol fields include request method (such as GET, POST), destination port (80 or 443), and protocol version (HTTP / 1).(1 or HTTP / 2); where the "GET" field indicates that the protocol is used to retrieve resources, and the "80" port specifies its default communication port. These fields together determine that the protocol must match web page-type payload content; the feature value sequence refers to an ordered set of multiple feature values ​​or identifiers extracted from the content identifier that reflect the core attributes of the payload content. Each feature value corresponds to a unique attribute of the payload. After being arranged in an ordered manner, the payload type can be accurately characterized, providing a specific basis for associating protocol fields. For example, the feature value sequence corresponding to the content identifier of a JPEG image payload is: start code (0xFFD8), resolution identifier (1920×1080), data compression ratio (1:8), color space code (0x03 corresponding to RG). B); The sequence is arranged in the order of "format-resolution-compression ratio-color", which can clearly distinguish images from text, video and other payloads; predefined rules refer to the logical criteria set in advance based on network protocol standards and payload business characteristics to match key protocol fields and feature value sequences, clarifying the corresponding conditions and matching priorities of the two, and ensuring the standardization and accuracy of the association process. For example, the predefined rule for web traffic is: "If the key protocol field contains 'HTTP protocol + GET method + port 80', then the feature value sequence must match at least one of 'tag hash (such as a1b2)' or 'UTF-8 encoding identifier (0xEFBBBF)'"; the rule for file transfer is: "FTP protocol + The port 21 field needs to match the 'file type encoding (e.g., 0x504446 corresponds to PDF)' in the feature value sequence. The matching field sequence refers to the ordered set of corresponding fields and feature values ​​obtained after matching the key protocol field and the feature value sequence according to predefined rules. Each element consists of "key protocol field + successfully matched feature value", intuitively presenting the association result between the two. For example, when the key protocol field is "HTTP+GET+80", and the feature value sequence is "a1b2 (hash), 512 bytes (data length), 0xEFBBBF (UTF-8)", after matching according to the rules, the matching field sequence is "HTTP→a1b2, GET→0xEFBBBF, 8". "0 → 512 bytes"; this sequence records each pair of matches, providing the basis for subsequent calculation of mapping confidence. The mapping confidence refers to a quantitative indicator measuring the reliability of the association between "key protocol fields and feature values" in the matching field sequence, calculated by statistically analyzing the number of successful matches and the matching accuracy. The value typically ranges from 0 to 100%, with higher values ​​indicating a more reliable mapping relationship. For example, if a matching field sequence contains 3 pairs of matches, 2 of which fully conform to predefined rules (e.g., HTTP → hash, 80 → webpage data length), and 1 pair partially conforms (GET → UTF-8 encoding), the calculated mapping confidence is 92%. If the confidence is below 80%, the matching process needs to be re-checked to ensure the association result is valid.

[0032] Furthermore, the identification of key protocol fields in the protocol attributes can be achieved through protocol field location methods, such as using a YAML syntax parser to extract required fields from the protocol definition to obtain key protocol fields; the extraction of the feature value sequence corresponding to the content identifier can be achieved through a feature hashing algorithm, such as using the SimHash algorithm to generate a binary fingerprint sequence of content features to obtain a feature value sequence; the matching of the key protocol fields with the feature value sequence can be achieved through a sequence alignment algorithm, such as applying the Smith-Waterman dynamic programming algorithm to calculate the similarity between the fields and the feature sequence to obtain a matching field sequence; the analysis of the mapping confidence corresponding to the matching field sequence can be achieved through a probabilistic statistical algorithm, such as using a Bayesian inference model to calculate the posterior probability value of sequence matching to obtain the mapping confidence; the generation of the protocol-content mapping relationship between the protocol attributes and the content identifier can be achieved through a relational modeling method, such as using a Protobuf protocol compiler to generate a structured mapping definition file to obtain the protocol-content mapping relationship.

[0033] S2. Collect the topology data sequence in the traffic topology map, parse the data transmission type corresponding to the topology data sequence, analyze the type-oriented mode corresponding to the data transmission type, and fit the traffic transmission curve under the type-oriented mode.

[0034] This invention collects topology data sequences from the traffic topology map and analyzes the data transmission types corresponding to these sequences. This transforms the scattered protocol-load association information, device transmission relationships, and other structured data in the map into an ordered and analyzable dataset, providing complete data support for subsequent accurate analysis of data transmission types. This avoids analytical biases caused by data fragmentation and improves the efficiency and accuracy of subsequent analysis steps.

[0035] The topology data sequence refers to an ordered, structured data set containing key information about network nodes and transmission relationships, collected from the traffic topology map according to time or transmission logic. It encompasses core elements such as timestamps, node identifiers (e.g., device IP, port), transmission parameters (rate, data volume), and protocol types. It serves as a carrier for transforming the visualized associations of the map into computable and analyzable data, providing a foundation for analyzing data transmission types. For example, a 10-second topology data sequence of a company network, ordered by timestamps (16:30:00, 16:30:02…16:…), could be used to analyze the network topology data. The data is sorted in a 30:10 ratio. Each data point includes the source IP (192.168.2.35), destination IP (203.0.113.8), protocol (HTTPS), transmission rate (0.8Mbps, 0.9Mbps…1.0Mbps), and single frame data size (1024 bytes, 1200 bytes…980 bytes), completely recording the transmission dynamics within this period. The data transmission type refers to the traffic category based on the characteristics of the topology data sequence (protocol type, transmission rate, connection method, and service data volume), reflecting the data... The technical characteristics and practical applications of transmission are the core identifiers that distinguish different business traffic (such as office communication, media transmission, and file interaction), directly determining the direction of subsequent type-oriented pattern analysis. For example, when parsing a certain topology data sequence, if its protocol is TCP, destination port 443, average transmission rate 2.5Mbps, long connection (lasting 120 seconds), single session data volume 150MB, and load characteristics match video encoding identifiers, the data transmission type corresponding to this sequence is determined to be "HTTPS video streaming transmission"; another sequence has a protocol of UDP, port 53, rate 0.1Mbps, and short connection (single 100ms), then it is determined to be "DNS domain name resolution transmission". Optionally, the collection of topology data sequences in the traffic topology map can be achieved through graph traversal algorithms, such as using a breadth-first search algorithm to traverse the relationship data of network nodes and edges to obtain the topology data sequence; the parsing of the data transmission type corresponding to the topology data sequence can be achieved through pattern classification methods, such as using a decision tree algorithm to classify data packets based on their size and frequency characteristics to obtain the data transmission type.

[0036] Furthermore, by analyzing the type-oriented patterns corresponding to the data transmission types, this invention can extract differentiated transmission patterns (such as rate fluctuation range, connection duration, data interaction frequency, etc.) from the characteristics of different transmission types, providing targeted pattern basis for subsequent fitting of traffic transmission curves, avoiding deviation of curve fitting from actual transmission scenarios due to lack of typified standards, and improving the targeting and efficiency of the overall detection process.

[0037] The type-oriented pattern refers to a framework of rules extracted from the topology data sequence for a specific data transmission type. This framework includes key characteristic parameters and constraints of transmission behavior, clarifying the normal range of this type of traffic in dimensions such as rate, connection duration, interaction frequency, and data volume fluctuations. It serves as a benchmark for judging whether transmission behavior conforms to its business attributes. For example, the type-oriented pattern for "DNS domain name resolution transmission type" specifies: transmission rate 0.05-0.2Mbps, single connection duration 50-200ms, interaction frequency 10-30 times per minute, and single data volume 100-500 bytes. If a DNS traffic connection duration lasts for 500ms or the interaction frequency reaches 60 times / minute, it is determined to deviate from this pattern and may be abnormal. Optionally, the analysis of the type-oriented pattern corresponding to the data transmission type can be achieved through clustering analysis methods, such as using the K-means algorithm to classify patterns based on transmission rate and data packet size characteristics, thereby obtaining the type-oriented pattern.

[0038] This invention, by fitting the traffic transmission curve under the type-oriented mode, can transform the abstract transmission patterns (such as rate fluctuations and data volume changes) under this mode into an intuitive dynamic trend model, making it easier to capture the characteristics of traffic changes in the time dimension. This provides quantifiable analytical basis for accurately identifying abnormal traffic peaks and locating core transmission hotspots, ensuring the accuracy of subsequent network cache status analysis.

[0039] The traffic transmission curve refers to a continuous curve generated by mathematical fitting (such as polynomial fitting or moving average fitting) based on multiple traffic data points, reflecting the change of traffic indicators over time under a certain type of guidance mode. It intuitively presents the dynamic trend of traffic (rising, falling, fluctuating) and is a visualization tool for identifying abnormal peaks and core hotspots. For example, after fitting 60 data points (rate 1.8-2.3Mbps) of "HTTPS video stream", the curve, with time as the horizontal axis and rate as the vertical axis, shows a trend of "small fluctuations and rise followed by stabilization". The peak is at 16:30:30 (2.3Mbps) and the trough is at 16:30:15 (1.8Mbps), clearly showing the time dimension characteristics of this type of traffic.

[0040] As an embodiment of the present invention, fitting the traffic transmission curve under the type-oriented mode includes: extracting the guidance transmission timing sequence in the type-oriented mode; setting the traffic sampling frequency corresponding to the type-oriented mode based on the guidance transmission timing sequence; collecting the real-time traffic sequence in the guidance transmission timing sequence according to the traffic sampling frequency; extracting traffic data points in the real-time traffic sequence; and fitting the traffic transmission curve under the type-oriented mode based on the traffic data points.

[0041] The guided transmission timing refers to the sequence of transmission characteristic data extracted from the type-guided pattern and arranged in chronological order. It records the core transmission parameters (such as rate, data volume, and connection status) of this type of traffic at different time points under normal scenarios. It serves as the time reference for setting the sampling frequency and collecting real-time traffic. For example, the guided transmission timing for "HTTPS video streaming" records the average rate (1.8Mbps, 2.2Mbps, 2.1Mbps...2.0Mbps) and the data volume per minute (13.5MB, 16.5MB, 15.8MB...15MB) every minute from 0 to 10 minutes, clearly presenting this type of traffic. The underlying transmission trend over time provides a timing reference for subsequent sampling. The traffic sampling frequency refers to the time interval for collecting real-time traffic data, set based on the time density of the guided transmission timing and the degree of fluctuation in transmission characteristics. This determines the granularity of the real-time traffic sequence—transmission types with large fluctuations require higher frequencies to capture details, while those with small fluctuations can reduce frequency balancing accuracy and resource consumption. For example, the rate fluctuation in the guided timing of "HTTPS video streaming transmission" is ±0.4Mbps, requiring a sampling frequency of 1 second / time; the rate fluctuation of "DNS resolution transmission" is only ±0.05Mbps, requiring only 5 seconds / time. This frequency ensures that the sampled data reflects the true trend while avoiding wasted computational resources due to oversampling. The real-time traffic sequence refers to a set of real-time traffic data of the corresponding type and guidance mode collected from the target network environment at a set traffic sampling frequency. It is arranged in order according to the sampling timestamp and contains the core transmission parameters at each sampling moment. It is the direct source for extracting traffic data points. For example, collecting real-time data of "HTTPS video streaming" at a frequency of 1 second / time results in the following sequence: timestamp 16:30:01 (rate 2.1Mbps), 16:30:02 (rate 1.9Mbps), 16:30:03 (rate 2.3Mbps)...16:30:60 (rate 2.0Mbps), a total of 60 sampling items, completely recording the real-time data within one minute. Transmission status; the traffic data point refers to a specific traffic information unit corresponding to a single sampling moment extracted from the real-time traffic sequence, including the sampling timestamp and the core traffic indicators at that moment (such as instantaneous rate, single frame data size, and number of connections). It is the basic "coordinate point" that constitutes the traffic transmission curve and directly reflects the traffic status at a certain moment. For example, a data point extracted from the real-time sequence of "HTTPS video stream" is: timestamp 16:30:05, instantaneous rate 2.2Mbps, single frame data size 1800 bytes; another data point is 16:30:10, instantaneous rate 1.8Mbps, single frame data size 1500 bytes. These discrete points together provide the original data support for curve fitting.

[0042] Furthermore, the extraction of the guided transmission time sequence in the type-oriented mode can be achieved through time series segmentation methods, such as using a sliding window algorithm to segment continuous traffic data at fixed time intervals to obtain the guided transmission time sequence; the setting of the traffic sampling frequency corresponding to the type-oriented mode can be achieved by applying the Nyquist sampling theorem, such as using the Shannon sampling formula to calculate twice the highest frequency of the mode to obtain the traffic sampling frequency; the acquisition of the real-time traffic sequence in the guided transmission time sequence can be achieved through traffic mirroring technology, such as using the SPAN port mirroring function to capture network device port traffic to obtain the real-time traffic sequence; the extraction of traffic data points in the real-time traffic sequence can be achieved through peak detection algorithms, such as applying the local maximum detection method to identify key data points in the traffic sequence to obtain traffic data points; the fitting of the traffic transmission curve in the type-oriented mode can be achieved through regression analysis methods, such as using a multinomial regression algorithm to fit the data points to generate a smooth transmission curve to obtain the traffic transmission curve.

[0043] S3. Identify abnormal traffic peaks and core transmission hotspots in the traffic transmission curve, calculate the traffic throughput threshold corresponding to the traffic data stream based on the abnormal traffic peaks and the core transmission hotspots, and analyze the network cache status of the traffic throughput threshold in the target network environment.

[0044] This invention identifies abnormal traffic peaks and core transmission hotspots in the traffic transmission curve, enabling timely capture of traffic anomalies that deviate from the normal transmission trend. This provides direct evidence for quickly locating potential network attacks, link congestion, and other abnormal issues, avoiding response delays caused by concealed anomalies. It ensures that the subsequent construction of the deep packet inspection framework is more in line with the actual network operation, thereby improving overall detection efficiency.

[0045] The abnormal traffic peak refers to a sudden surge in traffic value in the traffic transmission curve that deviates from the normal traffic fluctuation range of this type of guidance mode. It typically needs to meet the criteria of "exceeding the normal peak threshold by a certain percentage (e.g., more than 30%)" or "traffic increase exceeding 50% in a short period." It is often caused by sudden business requests, network attacks, or equipment failures, and is a key indicator for identifying network anomalies. For example, the traffic transmission curve for "HTTPS video streaming" shows a normal peak value stable at 2.3 Mbps. At a certain moment (16:45:20), the traffic suddenly surges to 4.5 Mbps, exceeding the normal peak value by 95.6%, and then quickly drops back after 15 seconds. This sudden surge in traffic is an abnormal traffic peak, possibly caused by a malicious DDoS attack. The core transmission hotspot refers to a specific transmission node (e.g., router, switch), link, or time period where traffic is highly concentrated under the traffic transmission curve and network topology correlation. Its characteristic is that its traffic share is significantly higher than other areas (usually accounting for the same percentage). Traffic of this type (accounting for more than 25% of the total traffic) and continuously operating under high load is a key area affecting network resource consumption and transmission efficiency. For example, in an enterprise network, the core switch (IP: 192.168.1.1) connecting the headquarters and R&D department accounts for 35% of the total daily traffic of this type between 8:00 and 10:00 AM, with a stable transmission rate of 1.2-1.5 Gbps, far exceeding other switches (average 0.3-0.5 Gbps). This switch and the corresponding time period constitute the core transmission hotspot. Optionally, the identification of abnormal traffic peaks in the traffic transmission curve can be achieved through anomaly detection algorithms, such as using the Z-score statistical algorithm to identify data points that deviate from three standard deviations from the mean, thereby obtaining abnormal traffic peaks. The identification of core transmission hotspots in the traffic transmission curve can be achieved through density clustering methods, such as using the DBSCAN algorithm to discover high-density continuous fluctuation areas, thereby obtaining core transmission hotspots.

[0046] Based on the abnormal traffic peak and the core transmission hotspot, this invention calculates the traffic throughput threshold corresponding to the traffic data stream, which can break through the limitations of static setting of the threshold, fully integrate the actual characteristics of extreme network load (abnormal peak) and high frequency and high load (core hotspot), more accurately reflect the real carrying capacity of the target network, avoid misjudgment caused by the disconnect between the threshold and the actual load, and reduce the processing delay caused by the load exceeding the threshold.

[0047] The traffic throughput threshold refers to a critical value calculated based on the throughput load value, combined with the resource redundancy requirements of network devices (usually reserving 5%-10% buffer space) and the rate fluctuation range of core transmission hotspots. It is used to define whether the traffic data flow exceeds the network's stable carrying capacity. For example, given a throughput load value of approximately 1215.48 Mbps, a core transmission hotspot rate fluctuation of ±0.15 Gbps (150 Mbps), and a device redundancy coefficient of 1.05, the traffic throughput threshold is calculated using the formula "(throughput load value + rate fluctuation limit) × redundancy coefficient" as (1215.48 + 150) × 1.05 ≈ 1432.25 Mbps. When the actual traffic exceeds this value, the system triggers a resource expansion or traffic throttling warning.

[0048] As an embodiment of the present invention, the step of calculating the traffic throughput threshold corresponding to the traffic data stream based on the abnormal traffic peak and the core transmission hotspot includes: analyzing the amplitude characteristics and occurrence frequency corresponding to the abnormal traffic peak; determining the upper limit value of the traffic corresponding to the amplitude characteristics and occurrence frequency; extracting the continuous transmission rate of the core transmission hotspot; determining the throughput load value corresponding to the traffic data stream based on the upper limit value of the traffic and the continuous transmission rate; and calculating the traffic throughput threshold corresponding to the traffic data stream based on the throughput load value.

[0049] The amplitude characteristic refers to the degree of deviation of the abnormal traffic peak from the normal traffic baseline value under this type of guidance mode. It is usually quantified by "the difference between the peak value and the normal average value" or "the proportion of the peak value exceeding the normal range," reflecting the intensity of the abnormal traffic and serving as a core indicator for judging the degree of impact of the anomaly. For example, the normal average traffic for "HTTPS video streaming transmission" is 2.0 Mbps, and the normal peak limit is 2.3 Mbps. If an abnormal traffic peak reaches 4.5 Mbps, its amplitude characteristic can be described as "a difference of 2.5 Mbps from the normal average value" and "exceeding the normal peak limit by 95.6%." These data directly reflect the intensity of the abnormal peak value, providing a basis for subsequent... This provides a basis for further determining the traffic limit. The frequency of occurrence refers to the number of times or probability that abnormal traffic peaks occur within a preset time period (e.g., 1 hour, 1 day), reflecting the frequency of abnormal traffic occurrences. It needs to be combined with time span statistics to avoid excessive influence of single, occasional anomalies on threshold calculation. For example, if the abnormal traffic peaks of "HTTPS video streaming" occur 3 times within 1 day (24 hours), at 9:15, 14:30, and 19:45 respectively, with an interval of 5-6 hours between adjacent occurrences, the frequency of occurrence can be expressed as "occurring 3 times in 1 day, with a frequency of 0.125 times / hour". This data can help determine whether the anomaly belongs to a high-frequency risk, and thus adjust the traffic limit. The safety factor is limited; the upper limit of traffic refers to the maximum traffic volume that can be tolerated based on the amplitude characteristics and frequency of abnormal traffic peaks, combined with the hardware capacity of network devices (such as routers and switches). It is a critical value to avoid equipment overload or link congestion, and needs to balance "anti-abnormal capability" and "resource utilization". For example, if the abnormal peak amplitude difference is known to be 2.5Mbps, the frequency of occurrence is 0.125 times / hour, and the maximum hardware capacity of the network device is 5.0Mbps, the traffic upper limit is set to 4.8Mbps after comprehensive calculation—which covers 95.6% of the abnormal amplitude (4.5Mbps < 4.8Mbps) and reserves 0.2Mbps. A buffer space of bps is provided to prevent hardware overload. The sustained transmission rate refers to the stable transmission rate (not instantaneous fluctuation value) maintained by the core transmission hotspot during its high-load period. It is usually expressed as the "average rate within the period" or "rate fluctuation range," reflecting the normalized load intensity of the core hotspot. This is key to judging whether network resources are under long-term strain. For example, the enterprise core switch (192.168.1.1), as the core hotspot for "office document transmission," has a stable transmission rate between 1.2-1.5Gbps during the period from 8:00 to 10:00 AM, with an average rate of 1.35Gbps calculated. This 1.35Gbps and the "fluctuation ±0."The range of "15Gbps" represents the sustained transmission rate of the core hotspot. The throughput load value is a quantified value reflecting the actual load on the traffic data stream, calculated by weighting the overall traffic limit and the sustained transmission rate of the core hotspot. The sustained rate of the core hotspot has a higher weight (due to its long-term resource occupation) and serves as the direct basis for subsequent calculations of the traffic throughput threshold. For example, if the traffic limit is 4.8Mbps (corresponding to abnormal peak scenarios) and the sustained transmission rate of the core hotspot is 1.35Gbps (equivalent to 1350Mbps, corresponding to normal high load), weighting it with "90% for sustained rate and 10% for limit" yields a throughput load value of (1350×0.9+4.8×0.1)≈1215.48Mbps. This value directly reflects the actual load intensity of traffic on the network.

[0050] Furthermore, the analysis of the amplitude characteristics corresponding to the abnormal traffic peaks can be achieved through extreme value analysis methods, such as using box plot statistical methods to calculate the interquartile range and boundary values ​​of the peak data to obtain the amplitude characteristics; the analysis of the occurrence frequency corresponding to the abnormal traffic peaks can be achieved through frequency statistics algorithms, such as using a sliding window counter to count the number of peak triggers per unit time to obtain the occurrence frequency; the determination of the upper limit value of the traffic corresponding to the amplitude characteristics and the occurrence frequency can be achieved through threshold calculation methods, such as applying the POT over-threshold model to fit the peak distribution to determine the safety boundary, thereby obtaining the upper limit value of the traffic; the extraction of the continuous transmission rate of the core transmission hotspot can be achieved through a moving average algorithm, such as using the EMA exponential moving average method to calculate the stable transmission rate of the hotspot area, thereby obtaining the continuous transmission rate; the determination of the throughput load value corresponding to the traffic data stream can be achieved through integral summation methods, such as using the trapezoidal integral method to calculate the area under the traffic curve per unit time to obtain the throughput load value; the calculation of the traffic throughput threshold corresponding to the traffic data stream can be achieved through the following formula.

[0051] In another embodiment of the present invention, the traffic throughput threshold corresponding to the traffic data stream is calculated based on the throughput load value using the following formula: ; in, This represents the throughput threshold (unit: bps) corresponding to the traffic data stream. This indicates network bandwidth capacity (unit: bps). This represents the throughput load value (unit: bps). This indicates the average duration (in seconds) of the core transmission hotspot. This indicates the frequency of abnormal traffic peaks (unit: 1 / second).

[0052] Specifically, the traffic throughput threshold represents the critical throughput value for determining "normal / overload" for a specific traffic data stream in the current network environment, considering factors such as bandwidth, actual load, duration of core hotspots, and abnormal peak frequency. For example, when network bandwidth B = 1000Mbps (1Gbps), throughput load L = 800Mbps, average duration of core hotspots D = 1800 seconds (30 minutes), and abnormal peak frequency F = 6.94 × When the frequency is 1 / second (6 times a day), Ty dynamically outputs a threshold that adapts to the actual network carrying capacity by integrating these factors through a formula, avoiding misjudgments based on static thresholds. The network bandwidth capacity can represent the theoretical maximum transmission capacity of a network link / device (in bps), which is the basic benchmark of the formula. It reflects the network's "inherent carrying capacity limit." The larger B is in the formula, the higher the benchmark value in the denominator, and theoretically, the higher the upper limit of Ty's calculation. For example, for home broadband, B = 100Mbps (10 8 bps), Enterprise backbone network B=10Gbps (10 10 The former (B) is smaller, so under the same load, the calculation result of Ty will be more stringent (the proportion of B in the denominator has a greater impact); the throughput load value can represent the actual carrying value (unit bps) of the combined abnormal traffic peak and core transmission hotspot load, representing the normal pressure of traffic on the network. In the formula, the ratio of L to B (L / B) reflects the "proportion of load to bandwidth". The larger the ratio, the larger the denominator, and the smaller Ty (high load compression threshold). For example, if B=1000Mbps and L=900Mbps (load proportion of 90%), then L / B=0.9, which will significantly increase the denominator and reduce Ty, reflecting the constraint of high load on the threshold; the average duration Time can represent the average duration (in seconds) that a core transmission hotspot remains under high load, reflecting the "temporal intensity" of resource consumption by the hotspot. In the formula, D is multiplied by the abnormal frequency F (D·F), reflecting the combined effect of "time × frequency": the longer D is (e.g., a core switch under high load for 3600 seconds = 1 hour), the larger D·F is if F is fixed, the larger the denominator, and the smaller Ty is (a longer hotspot duration will lower the threshold). The occurrence frequency can represent the number of times abnormal traffic peaks occur per second (in 1 / second), reflecting the "temporal density" of the anomaly. In the formula, F and D are linked; the higher F is (e.g., 5 occurrences per hour, F = 5 / 3600 ≈ 1.39 × 1000 / 10000), the higher the frequency. (1 / second). The larger the product of D and F, the larger the denominator and the smaller Ty (the more frequent the anomalies, the stricter the threshold). For example, F = 1 × 1 / second (8.64 times per day) is equal to F = 5 × 1 / second (4.32 times per day) will increase D·F, further reducing Ty.

[0053] Furthermore, by analyzing the network cache status of the traffic throughput threshold in the target network environment, this invention aims to accurately determine the cache overload, saturation, or redundancy status by correlating the threshold with real-time traffic, providing a quantitative basis for dynamic resource allocation, enabling the cache capacity to continuously adapt to network carrying requirements, and improving transmission efficiency.

[0054] The network cache status refers to the operational status of the network cache system at a specific moment, considering its capacity, real-time storage volume, traffic throughput load, and data hit efficiency. It reflects the cache's ability to handle and respond to traffic requests, encompassing dimensions such as capacity saturation, load pressure, and data hit rate. For example, a video platform has an edge cache capacity of 1000GB, currently storing 880GB (88% saturation). Under the load corresponding to the traffic throughput threshold, the cache hit rate drops from 92% to 75%, and there are 15 cache misses per second. At this time, the status is "high saturation + low hit rate, requiring optimization of the cache strategy or expansion." Optionally, the analysis of the network cache status in the target network environment based on the traffic throughput threshold can be achieved through cache hit rate analysis methods, such as using the Redis INFO command to obtain cache key space hit rate and memory usage indicators, thereby obtaining the network cache status.

[0055] S4. Based on the network cache state, construct a deep packet inspection framework corresponding to the target network environment, monitor the real-time classification path in the deep packet inspection framework, and collect the processing latency value and classification misclassification rate in the real-time classification path. Based on the processing latency value and the classification misclassification rate, calculate the traffic coordination index corresponding to the traffic data stream.

[0056] Based on the network cache state, this invention constructs a deep packet inspection framework corresponding to the target network environment. It can dynamically allocate detection resources according to the cache capacity saturation, load pressure, etc., avoids the ineffective loss of detection capability when the cache is overloaded, enhances the resilience to cope with complex traffic scenarios, and achieves synergistic optimization of detection efficiency and network transmission.

[0057] The deep packet inspection framework refers to a detection system that integrates cache status analysis, traffic level classification, deep packet identification, and performance data application. It features dynamic rule updates and intelligent resource scheduling, and can adapt to the dynamic changes of the target network. For example, the framework adjusts the size of the detection thread pool based on cache utilization (expanding by 30% under high load) and optimizes detection rules based on the performance data of deep packets in the environment (such as skipping known harmless video stream header fields), ultimately forming a closed-loop detection system of "resources allocated on demand + precise rule adaptation".

[0058] As an embodiment of the present invention, the step of constructing a deep packet inspection framework corresponding to the target network environment based on the network cache state includes: analyzing the cache utilization rate corresponding to the network cache state; determining the traffic detection level in the target network environment based on the cache utilization rate; identifying environmental deep packets in the target network environment based on the traffic detection level; querying the performance detection data corresponding to the environmental deep packets; and constructing a deep packet inspection framework corresponding to the target network environment based on the performance detection data.

[0059] The cache utilization rate refers to the proportion of network cache capacity used to the total capacity. It also needs to be considered in conjunction with cache hit efficiency, read / write frequency, and other dimensions to comprehensively reflect the cache's load pressure and resource utilization efficiency. For example, a certain edge cache has a total capacity of 500GB, currently storing 420GB (84% capacity), but the cache hit rate for popular videos is only 65% ​​(far lower than the expected 90%). This indicates that although the capacity utilization rate is high, the uneven distribution of hot and cold data leads to inefficiency, requiring optimization of cache content through testing. The traffic detection level refers to the detection priority (e.g., high, medium, low) of target network traffic based on cache utilization rate, network load, and other statuses, determining the allocation strategy for detection resources (computing power, bandwidth). For example, when the cache utilization rate exceeds 85% (high load), video streams and core business traffic are set to high-level detection (prioritizing 70% of detection resources); utilization rates of 50%-85% are set to medium-level (balanced allocation); and below 50% are set to low-level (basic detection only), achieving on-demand resource deployment. The environment depth packet refers to the layer based on the list protocol factor, combined with the target... The deep data packets that are ultimately determined from the candidate deep packets and require key detection are based on the actual business scenarios of the target network environment (such as core business and high-load cache-related traffic). These packets possess the dual attributes of "protocol feature matching + strong business relevance". For example, considering a scenario where "video conferencing is the core business" and "the cache video fragment hit rate is low", data packets with the protocol factor of "TLS1.3 + port 443 + video fragment request" are selected from the candidate HTTPS packets. These packets are directly related to core business and cache optimization requirements, and are thus the environment deep packets. The performance detection data refers to the quantitative indicators of the environment deep packets during transmission and processing, covering dimensions such as network layer (latency, throughput, packet loss rate) and detection layer (single packet processing time, detection module concurrency capability), providing a basis for framework optimization. For example, a video stream deep packet has a network latency of 45ms, a throughput of 600Mbps, a single packet parsing time of 1.2ms for the detection module, and a concurrency processing limit of 8000 packets / second. These data reflect the packet transmission performance and detection difficulty, guiding framework resource scheduling.

[0060] Furthermore, the analysis of the cache utilization rate corresponding to the network cache status can be achieved through resource monitoring methods, such as using the Prometheus monitoring system to collect cache memory usage and exchange frequency indicators to obtain the cache utilization rate; the determination of the traffic detection level in the target network environment can be achieved through a hierarchical evaluation model, such as using a fuzzy comprehensive evaluation method to combine traffic scale and risk factors to divide the detection level to obtain the traffic detection level; the identification of the environmental deep packets in the target network environment can be achieved through deep packet inspection technology, such as using the nDPI open-source library to identify the seven-layer protocol and extract metadata to obtain the environmental deep packets; the query of the performance detection data corresponding to the environmental deep packets can be achieved through performance profiling tools, such as using the Perf tool to analyze the CPU cycles and cache hit rate during deep packet processing to obtain the performance detection data; the construction of the deep packet inspection framework corresponding to the target network environment can be achieved through a modular design method, such as using the DPDK data plane development kit to build a high-performance packet processing pipeline to obtain the deep packet inspection framework.

[0061] As another embodiment of the present invention, the step of identifying environmental deep packets in the target network environment based on the traffic detection level includes: parsing the level detection baseline corresponding to the traffic detection level; performing deep correlation between the level detection baseline and the real-time traffic status in the target network environment to obtain candidate deep packets; generating a deep packet identifier list corresponding to the candidate deep packets; extracting the list protocol factors from the deep packet identifier list; and identifying environmental deep packets in the target network environment based on the list protocol factors.

[0062] The aforementioned level detection baseline refers to a standard framework set for the corresponding traffic detection level (high / medium / low), which includes detection dimensions and quantitative thresholds. It clarifies the key traffic characteristics (such as protocol type, packet size, and interaction frequency) that need to be focused on at that level, serving as a "reference benchmark" for screening deep packet inspections. For example, the baseline for "high-level detection" is set as follows: protocol type limited to HTTPS / QUIC, single packet size 1024-2048 bytes, and interaction frequency ≥ 50 times per second; the baseline for "low-level detection" is relaxed to: no protocol restrictions, packet size 512-4096 bytes, and frequency ≥ 10 times per second, matching different detection requirements through differentiated thresholds. The aforementioned real-time traffic status refers to the real-time traffic status in the target network environment at a given time. The dynamic feature set of the traffic being transmitted at any given moment, encompassing key parameters such as traffic protocol type, instantaneous transmission rate, single packet size, source / destination IP port, and data interaction frequency, serves as a "real-time data source" correlated with the level detection baseline. For example, the real-time traffic status of a company's network at 14:30 might be: HTTPS protocol traffic accounting for 65%, instantaneous rate 800Mbps, average single packet size 1500 bytes, and interaction frequency between 192.168.2.35 (client) and 203.0.113.8 (video server) 62 times / second. This data is directly used for comparison with the baseline. The candidate deep packet refers to the filter after deeply correlating the level detection baseline with the real-time traffic status. The data packets that meet the baseline characteristic thresholds are the preliminary screening results of "environmental deep packets," only satisfying the "basic matching conditions" and not yet further verified by protocol factors. For example, based on the "advanced detection baseline" (HTTPS, 1024-2048 bytes, ≥50 times / second), HTTPS data packets of 1200-1800 bytes transmitted between 192.168.2.35 and 203.0.113.8 are filtered from real-time traffic, totaling 58 packets / second. These data packets that meet the baseline are candidate deep packets. The deep packet identifier list refers to a structured list containing unique identification information generated for all candidate deep packets. Each record corresponds to one candidate packet, covering the packet's... The unique ID, protocol type, source / destination port, and summary value (such as hash value) of key payload features facilitate subsequent extraction of protocol factors and tracing. For example, a record in the list might have the following characteristics: ID (DP20240529143001), protocol (HTTPS), source port (54321), destination port (443), and payload hash (a1b2c3d4e5f6). This list allows for quick location and extraction of the core information of each candidate packet. The protocol factors in the list refer to the set of key features extracted from the deep packet identifier list that reflect the protocol attributes of candidate deep packets, covering protocol type, port number, request method (such as HTTP GET / POST), and protocol version (such as TLS1).3) These factors are the core criteria for distinguishing whether candidate packets belong to the "key environmental focus objects." For example, the protocol factors extracted from the above list include: protocol type (HTTPS), destination port (443), TLS version (1.3), and associated application layer request (video segment acquisition). These factors can clearly define the association between candidate packets and "video streaming transmission," providing support for the final identification.

[0063] Furthermore, the parsing of the level detection baseline corresponding to the traffic detection level can be achieved through baseline modeling methods, such as generating the minimum-maximum threshold range for each level using historical traffic data statistical analysis, thereby obtaining the level detection baseline; the deep correlation between the level detection baseline and the real-time traffic status in the target network environment can be achieved through correlation analysis algorithms, such as calculating the correlation strength between real-time traffic indicators and baseline features using Pearson correlation coefficient, thereby obtaining candidate deep packets; the generation of the deep packet identifier list corresponding to the candidate deep packets can be achieved through identifier generation methods, such as generating unique feature identifiers for candidate deep packets using the SHA-256 hash algorithm, thereby obtaining the deep packet identifier list; the extraction of list protocol factors from the deep packet identifier list can be achieved through factor analysis methods, such as extracting the most influential protocol feature dimensions from the identifier list using principal component analysis (PCA), thereby obtaining list protocol factors; the identification of environmental deep packets in the target network environment can be achieved through pattern recognition technology, such as matching the feature patterns of real-time traffic packets and protocol factors using convolutional neural networks (CNN), thereby obtaining environmental deep packets.

[0064] This invention monitors the real-time classification path in the deep packet inspection framework and collects the processing delay value and classification misclassification rate in the real-time classification path. It can dynamically track the execution logic of traffic classification within the framework, promptly detect abnormalities such as path deviation and branch congestion, avoid some traffic being missed or misdetected due to path failure, ensure the continuity and integrity of the detection process, and reduce the interference of misjudgments on normal network transmission.

[0065] The real-time classification path refers to the dynamically generated execution path in the deep packet inspection framework when incoming data packets are classified according to preset rules (such as protocol type and business characteristics). This path includes the complete chain of detection nodes (such as protocol parsing modules and feature matching modules), branch judgment logic (such as whether the cache is hit), and the final classification result. For example, after an HTTPS video packet enters the framework, the path is: "Ingress traffic filtering → TLS protocol parsing node (extract certificate information) → Application layer payload feature matching node (identify video segment identifier) ​​→ Core business classification branch → Marked as 'video stream'". This path records the packet's processing trajectory in real time, facilitating anomaly tracking. The processing latency value refers to the total time it takes for a data packet to complete classification processing (outputting the classification result) from entering the deep packet inspection framework, usually measured in milliseconds (ms). It includes the processing time of each detection node and the transmission time between nodes, reflecting the framework's response efficiency to real-time traffic. For example, in a deep packet inspection framework for an enterprise network, the average processing latency of video stream data packets under normal load is 8ms (TLS parsing 3ms + feature matching 4ms + transmission 1ms). When the stream... When the peak bandwidth reaches 1000Mbps, the latency increases to 22ms (mainly due to a 15ms increase in queuing time for the feature matching module). This value directly reflects the framework's real-time processing capability. The classification misclassification rate refers to the proportion of misclassified data packets to the total number of classified data packets in the deep packet inspection framework, usually expressed as a percentage (%). It includes two categories: "misclassifying normal packets as abnormal" and "misclassifying service type" (such as misclassifying video streams as file transfers). It is a core indicator for measuring classification accuracy. For example, if the framework classifies 10,000 data packets in one hour, of which 50 are normal... Video packets were misclassified as "abnormal attack packets," and 30 file transfer packets were incorrectly classified as "web traffic," for a total of 80 misclassifications. The classification misclassification rate is 80 / 10000 = 0.8%. The lower this value, the more accurate the classification rules. Optionally, the monitoring of the real-time classification path in the deep packet inspection framework can be achieved through path tracing technology, such as using an eBPF kernel probe to capture the processing path of data packets in the classification engine in real time, thereby obtaining the real-time classification path. The collection of processing latency values ​​in the real-time classification path can be achieved through high-precision timing methods, such as using an Intel TSC timestamp counter to measure the dwell time of data packets in each processing module, thereby obtaining the processing latency value. The collection of the classification misclassification rate in the real-time classification path can be achieved through error statistics methods, such as applying a confusion matrix to compare the difference ratio between the actual classification result and the expected result, thereby obtaining the classification misclassification rate.

[0066] Furthermore, based on the processing latency value and the classification misclassification rate, the present invention calculates the traffic coordination index corresponding to the traffic data stream, which can integrate the latency index reflecting the framework response efficiency with the misclassification rate index reflecting the classification accuracy to form a comprehensive quantitative standard for measuring the coordination level of "efficiency-accuracy" in traffic processing. This avoids one-sided evaluation caused by relying solely on latency or misclassification rate, and achieves an overall balance in traffic processing performance.

[0067] The flow coordination index refers to a quantitative indicator obtained by normalizing the delay fluctuation and misjudgment fluctuation values ​​of the comprehensive classification coordination links after considering the reference delay value and the reference misjudgment rate. It reflects the degree of coordination and stability between processing delay and classification misjudgment rate. The smaller the value, the better the balance between the two. For example, a framework contains two coordination links: link 1 has a delay fluctuation of 10ms and a misjudgment fluctuation of 2.2%; link 2 has a delay fluctuation of 8ms and a misjudgment fluctuation of 1.8%. Assuming a reference delay of 20ms and a reference misjudgment rate of 3%, the normalized fluctuations of the two links are calculated to be 0.71 and 0.63, respectively. The index takes an average of 0.67, indicating good coordination. If the index rises to 1.2, it indicates that the fluctuation is unbalanced and needs optimization.

[0068] As an embodiment of the present invention, the step of calculating the traffic coordination index corresponding to the traffic data stream based on the processing latency value and the classification misclassification rate includes: analyzing the cooperative change features associated with the processing latency value and the classification misclassification rate; extracting abnormal cooperative markers from the cooperative change features; mapping the classification coordination link in the deep packet inspection framework based on the abnormal cooperative markers; statistically analyzing the latency fluctuation value and misclassification fluctuation value in the classification coordination link; and calculating the traffic coordination index corresponding to the traffic data stream based on the latency fluctuation value and the misclassification fluctuation value.

[0069] The aforementioned coordinated change characteristic refers to the correlation between processing latency and classification misclassification rate as network load and framework resource allocation change. This can manifest as "simultaneous increase / decrease," "inverse fluctuation," or "one remaining stable while the other changes," reflecting their interrelationship in traffic processing. For example, when network load increases from 500Mbps to 1000Mbps, processing latency increases from 8ms to 22ms (an increase of 175%), and classification misclassification rate increases from 0.8% to 2.5% (an increase of 212.5%), exhibiting a coordinated characteristic of "latency and misclassification rate increasing synchronously with increased load." However, after optimizing the feature matching rules, latency decreases to 15ms (a decrease of 31.8%), and the misclassification rate decreases to 1.1% (a decrease of 56%). If the processing latency and classification misclassification rate deviate from the preset normal collaboration mode (e.g., "latency increases, misclassification rate increases slightly"), then the collaborative characteristic of "both decrease simultaneously during rule optimization" will be exhibited. The abnormal collaboration marker refers to the identifier generated by the system to mark the abnormal association when the changes in processing latency and classification misclassification rate deviate from the preset normal collaboration mode (e.g., "latency increases, misclassification rate increases slightly"). It usually corresponds to situations that do not conform to normal logic, such as "latency drops sharply but misclassification rate rises sharply" or "latency rises sharply but misclassification rate remains unchanged". For example, the preset normal collaboration mode is "for every 10ms increase in latency, the misclassification rate increases by 0.5%-0.8%". If the latency decreases from 15ms to 10ms in a certain period (a decrease of 33.3%), but the misclassification rate rises sharply from 1.1% to 3.2% (an increase of 190.9%), completely deviating from the normal range, the system will generate a "latency-misclassification rate increase". The "Reverse Anomaly" flag indicates a rule conflict or resource scheduling anomaly. The classification coordination stage refers to a core functional module within the deep packet inspection framework specifically designed to balance latency and classification misclassification rate. It encompasses sub-stages such as "feature matching rule optimization," "dynamic scheduling of detection resources," and "data packet priority sorting," and is a crucial part determining the synergistic relationship between latency and misclassification rate. Anomaly coordination flags typically map to specific sub-modules within this stage. For example, when the anomaly coordination flag points to "sudden latency decreases but sudden increase in misclassification rate," mapping reveals that the anomaly originates from the "feature matching rule simplification" sub-stage—three key feature detections were removed to reduce latency, resulting in a decrease in latency but a surge in the misclassification rate. The "stage" refers to the corresponding classification and coordination stage. The delay fluctuation value refers to the variation in processing delay value around the average delay of that stage within the classification and coordination stage. It is usually expressed as the "difference between the maximum and minimum delay" or the "delay standard deviation," reflecting the stability of the delay in that stage. It is an important fluctuation parameter for calculating the traffic coordination index. For example, for the "feature matching rule simplification" classification and coordination stage, the processing delay data for this stage over one hour is statistically analyzed: average delay 12ms, maximum delay 18ms, minimum delay 8ms, with a delay fluctuation value of 18ms - 8ms = 10ms. If the standard deviation is calculated, the delay data distribution is 8ms, 10ms, 12ms, 15ms, and 18ms, with a standard deviation of approximately 3.Both methods can quantify the latency fluctuation of this stage, using 74ms. The misjudgment fluctuation value refers to the variation in the classification misjudgment rate around the average misjudgment rate of the classification coordination stage. It is usually expressed as the "difference between the maximum and minimum misjudgment rates" or the "standard deviation of the misjudgment rate," reflecting the stability of the misjudgment rate in this stage. Together with the latency fluctuation value, it forms the basis for calculating the traffic coordination index. For example, for the "feature matching rule simplification" stage, statistical analysis of the misjudgment rate data over one hour shows an average misjudgment rate of 2.1%, a maximum misjudgment rate of 3.2%, and a minimum misjudgment rate of 1.0%, with a misjudgment fluctuation value of 3.2% - 1.0% = 2.2%. If the standard deviation is calculated, the misjudgment rate data distribution is 1.0%, 1.5%, 2.1%, 2.8%, and 3.2%, with a standard deviation of approximately 0.83%. This value clearly indicates the intensity of the misjudgment rate fluctuation in this stage.

[0070] Furthermore, the analysis of the co-variance characteristics between the processing latency and the classification misclassification rate can be achieved through covariance analysis, such as using the Pearson correlation coefficient to calculate the synchronous change trend of latency and misclassification rate, thereby obtaining co-variance characteristics; the extraction of abnormal co-variance markers from the co-variance characteristics can be achieved through abnormal pattern detection, such as using the isolated forest algorithm to identify outlier data points in the co-variance characteristics, thereby obtaining abnormal co-variance markers; the mapping of the classification coordination link in the deep packet inspection framework can be achieved through dependency graph construction, such as using the Jaeger distributed tracing system to draw the call relationship graph of each classification module, thereby obtaining the classification coordination link; the statistical analysis of latency fluctuation values ​​in the classification coordination link can be achieved through standard deviation calculation, such as using the sliding window statistical method to calculate the standard deviation of latency time of each link, thereby obtaining latency fluctuation values; the statistical analysis of misclassification fluctuation values ​​in the classification coordination link can be achieved through coefficient of variation analysis, such as using the CV coefficient to calculate the relative dispersion of misclassification rate of each link, thereby obtaining misclassification fluctuation values; the calculation of the traffic coordination index corresponding to the traffic data stream can be achieved through the following formula.

[0071] In another embodiment of the present invention, the traffic coordination index corresponding to the traffic data stream is calculated based on the delay fluctuation value and the misjudgment fluctuation value using the following formula: ; in, This represents the traffic coordination index corresponding to the traffic data stream. This indicates the total number of the classification and coordination links. This represents the quantity index corresponding to the classification coordination step. Indicates the first Delay fluctuation values ​​(unit: seconds) corresponding to each classification coordination link. Indicates the reference delay value (unit: seconds). Indicates the first The error fluctuation value corresponding to each classification coordination link This indicates the reference misjudgment rate.

[0072] In detail, the flow coordination index can represent the delay fluctuation of the K classification coordination links ( ) and misjudgment fluctuations ( ), based on reference values ​​( , After normalization, the average "comprehensive fluctuation deviation" is calculated. In the formula, the deviation is first calculated for each stage j. Then sum them up and divide by K. The smaller the value, the better the synergy between delay and misjudgment. For example, K=2, step 1: =0.01s (10ms) =0.015 (1.5%); Step 2: =0.008s (8ms) =0.012 (1.2%); Let =0.02s (20ms) =0.03 (3%), then step 1 = ≈0.707, and similarly for stage 2, C=(0.707+0.707) / 2≈0.707, reflecting the overall coordination deviation; the delay fluctuation value can represent the fluctuation amplitude (e.g., standard deviation, unit: seconds) of the processing delay in the j-th classification coordination stage, characterizing the stability of the delay in that stage. In the formula, it is compared with the reference delay. In comparison, this measures the degree to which actual delay fluctuations deviate from the "normal baseline." For example, if the delay sequence of a certain stage is [0.01s, 0.012s, 0.008s] (corresponding to 10ms, 12ms, and 8ms), the standard deviation is calculated. ≈0.0016s (approximately 1.6ms), which is the delay fluctuation value of this stage, used for subsequent normalization calculations; the reference delay value can represent a predefined "normal delay fluctuation benchmark for classification and coordination stages" (unit: seconds), as a normalization factor, allowing for horizontal comparison of delay fluctuations in different stages. In the formula, it will... Converting to a "relative fluctuation ratio" eliminates the impact of absolute numerical differences; for example, setting based on historical best data. =0.005s (5ms), if a certain link =0.003s (3ms), then =0.6, indicating that the delay fluctuation in this stage is only 60% of the reference value, showing good stability; the misclassification fluctuation value can represent the fluctuation range (e.g., standard deviation) of the classification misclassification rate in the j-th classification coordination stage, reflecting the stability of the misclassification rate in this stage. In the formula, it is related to the reference misclassification rate. In combination, the "relative deviation" of quantified misjudgment fluctuations is calculated. For example, if the misjudgment rate sequence for a certain stage is [1.2%, 1.5%, 1.0%] (converted to 0.012, 0.015, 0.01), the standard deviation is calculated. ≈0.0021 (2.1%), which is the misjudgment fluctuation value of this stage, used for subsequent normalization; the reference misjudgment rate can represent the pre-set "normal misjudgment fluctuation benchmark of the classification and coordination stage", as a normalization reference for misjudgment fluctuation. In the formula, it will... Converting to a relative proportion, we can standardize the measurement of misjudgment fluctuations across different scenarios. For example, in industry standards, the normal misjudgment fluctuation for deep packet inspection should not exceed 3% (assuming...). =0.03), if a certain link =0.015 (1.5%), then =0.5, indicating that the misjudgment fluctuation is only 50% of the reference value, and the stability is relatively good.

[0073] Specifically, for a more intuitive understanding of the execution logic and data flow relationships in deep packet inspection and traffic coordination processes, please refer to [link to relevant documentation]. Figure 2 ,Should Figure 2 As a schematic diagram of traffic coordination logic, it clearly presents the complete link from resource initialization to coordination index output: The input layer focuses on basic resources and data: through "initialization and resource loading," the load analysis model and state classification rules are loaded, and "data input" imports consultation load values, forming the basis for analysis; the processing layer relies on the "core analysis engine" to link topology analysis (rendering network structure, traversing node paths), state and performance analysis (analyzing classification indicators, detecting bottlenecks), and platform load analysis (assessing carrying pressure) in a step-by-step logical linkage, transforming raw data into the basis for building the detection framework; the output layer takes "deep packet detection and traffic coordination" as its core, and through "building the detection framework → monitoring paths and collecting indicators → calculating the coordination index," it realizes the quantitative optimization of traffic processing efficiency. It should be noted that the flowchart is an abstract distillation of the "caching-detection-coordination" logic. In reality, the parameter coupling of dynamic topology tracking and index calculation is more complex. This architecture is only used to show the core logic to help understand the systematic thinking.

[0074] S5. Based on the traffic coordination index, reconstruct the data coordination task corresponding to the traffic data stream, identify the task classification protocol in the data coordination task, and generate the task classification result of the traffic data stream when it is oriented towards multiple tasks based on the task classification protocol.

[0075] Based on the traffic coordination index, this invention reconstructs the data coordination tasks corresponding to the traffic data stream. According to the latency and misjudgment coordination characteristics revealed by the index, the allocation strategy of tasks in each category of coordination stage can be dynamically adjusted, allowing efficient stages to bear more load and inefficient stages to be optimized first, thereby achieving a dynamic balance between resources and tasks and enhancing the framework's resilience in dealing with complex traffic.

[0076] The data coordination task refers to a new set of tasks that are optimized and adjusted based on task sequence data. By correcting inefficient links and balancing resource allocation, the tasks are made more suitable for the current state corresponding to the traffic coordination index, thereby improving the "efficiency-precision" coordination level. For example, in the original task, "all video packets are processed by the master node" resulted in C=0.92 (exceeding the threshold). After reconstruction, the task is "the master node processes 50% (delay-sensitive packets) + the backup node processes 50% (non-sensitive packets)", and the secondary feature detection is simplified, so that C is reduced to 0.75 (within the regression threshold), thereby achieving dynamic adaptation between the task and the coordination state.

[0077] As an embodiment of the present invention, the step of reconstructing the data coordination task corresponding to the traffic data stream based on the traffic coordination index includes: analyzing the coordination balance threshold corresponding to the traffic coordination index; obtaining the data coordination task corresponding to the traffic data stream based on the coordination balance threshold; querying the task coordination rule corresponding to the data coordination task; generating task sequence data in the task coordination rule; and reconstructing the data coordination task corresponding to the traffic data stream based on the task sequence data.

[0078] The coordination and balancing threshold, defined based on the traffic coordination index (C), is a critical range used to determine whether a data coordination task is in a "balanced state of efficiency and accuracy." It typically includes an upper and lower limit; exceeding this range triggers task refactoring. For example, based on historical best data, the coordination and balancing threshold can be set to 0.6-0.8 (a reasonable range for the traffic coordination index C): when C = 0.5 (below the lower limit), it indicates that latency and misjudgment fluctuations are too small, suggesting potential resource redundancy; when C = 0.9 (above the upper limit), it indicates excessive fluctuations, requiring urgent task refactoring to reduce coordination deviations. This threshold serves as a reference for task adjustment. Provide clear triggering criteria; the data coordination task refers to a series of operations performed during the transmission and processing of traffic data streams to achieve the collaborative goal of "low latency and low false positives". It covers traffic scheduling, resource allocation, rule adaptation, etc., and is directly related to the operating efficiency of the deep packet inspection framework. For example, the coordination tasks for video stream data streams include: "allocating 70% of high-definition video packets to detection nodes with a latency of ≤10ms", "automatically updating the feature matching library when the false positive rate exceeds 1.5%", and "dynamically adjusting the load ratio of each node every 5 minutes". These tasks ensure the balance of traffic processing through collaborative actions. The task coordination rules refer to the binding clauses that regulate the execution logic of data coordination tasks, clarifying task triggering conditions, execution priorities, resource allocation ratios, etc., and serve as the basis for generating task sequences, ensuring that tasks proceed in an orderly manner according to preset logic. For example, a rule might stipulate: "When the coordination balance threshold is exceeded (C > 0.8), the 'reduce feature matching complexity' task (priority 1) is executed first, followed by the 'transfer 20% load to the backup node' task (priority 2), and the load transfer amount each time does not exceed 30% of the current node's load." These rules clearly define the order and limitations of task execution; the number of task sequences... Data refers to a structured list of tasks generated according to task coordination rules and arranged in execution order. Each record contains task ID, execution time window, required resources, and associated parameters. It serves as the direct operational basis for reconstructing data coordination tasks. For example, the sequence data generated based on the above rules is: Task 1 (ID: T001, execution window: 0-30s, resources: 20% computing power of feature matching module, parameters: retain 3 core features) → Task 2 (ID: T002, execution window: 30-60s, resources: 5GB memory of spare node, parameters: transfer load 15%), clearly defining the execution chain of tasks.

[0079] Furthermore, the analysis of the coordination equilibrium threshold corresponding to the traffic coordination index can be achieved through a dynamic threshold calculation method, such as using an exponentially weighted moving average method to dynamically adjust the threshold boundary based on historical coordination indices to obtain the coordination equilibrium threshold; the acquisition of the data coordination task corresponding to the traffic data stream can be achieved through a task extraction method, such as using a rule-based matching algorithm to extract the task items to be coordinated from traffic metadata to obtain the data coordination task; the query of the task coordination rule corresponding to the data coordination task can be achieved through a rule engine query method, such as using the Drools rule engine to retrieve the coordination strategy set corresponding to the task type to obtain the task coordination rule; the generation of task sequence data in the task coordination rule can be achieved through a topological sorting algorithm, such as using the Kahn algorithm to generate an ordered execution sequence based on task dependencies to obtain the task sequence data; the reconstruction of the data coordination task corresponding to the traffic data stream can be achieved through a task optimization method, such as using a genetic algorithm to perform multi-objective recombination optimization on the coordination task to obtain the data coordination task.

[0080] This invention identifies the task classification protocols in the data coordination task, accurately anchors the protocol attributes of the task, provides a decision basis for task scheduling based on the protocol dimension, avoids processing disorder caused by cross-protocol mismatch, promotes deep collaboration between task and protocol features, and optimizes the operational efficiency of the overall data coordination system.

[0081] The task classification protocol refers to the mapping system between the traffic protocol type and task processing rules in data coordination tasks. It includes protocol identifiers (such as HTTPS, QUIC), adapted task scheduling strategies (such as resource allocation weights), and protocol-specific processing logic (such as pre-decryption of encryption protocols). For example, the classification protocol of a video stream data coordination task is identified as HTTPS (TLS 1.3 version). The adaptation rule requires "prioritizing the allocation of 70% of TLS decryption computing power to this protocol task" and "extracting video segment identifiers after decryption before performing load scheduling". By binding protocol features with task rules, it ensures that task processing conforms to protocol transmission characteristics (such as encryption and high load), improving the protocol adaptation accuracy of data coordination. Optionally, the identification of the task classification protocol in the data coordination task can be achieved through protocol clustering analysis methods, such as using the K-means algorithm to automatically divide the protocol categories according to the task feature vector, thereby obtaining the task classification protocol.

[0082] Furthermore, based on the task classification protocol, the present invention generates task classification results for the traffic data stream when it is oriented towards multiple tasks. This allows the data stream to accurately match the corresponding processing tasks according to the protocol attributes, avoiding task mismatch and processing disorder in multi-task scenarios, ensuring that each task proceeds in an orderly manner according to the protocol adaptation logic, while enhancing the dynamic adaptation capability of the multi-task system to data streams of different protocols and optimizing the overall task processing efficiency.

[0083] The task classification result refers to the structured output generated after matching and configuring the traffic data stream in a multi-task scenario based on the task classification protocol. It includes the task ID associated with the data stream, the protocol matching identifier, the resource allocation ratio, the execution priority, and the specific processing rules, which clarify the specific processing path and parameters of the data stream in the multi-task system. For example, the classification result of a 100Mbps HTTPS video stream (protocol TLS1.3) of an enterprise during a certain period is: the assigned task ID (Tsk-V001, video stream-specific coordination task), resource allocation (50% TLS decryption computing power + 30% core transmission bandwidth), execution priority (P1, the highest level), and processing rules (completing fragment decryption first and then executing load scheduling). This result directly guides the data stream to execute efficiently in multi-task according to the protocol adaptation logic. Optionally, the task classification result of the traffic data stream when facing multiple tasks can be generated by a multi-label classification method, such as using a BERT neural network model combined with traffic features to perform multi-task parallel classification, thereby obtaining the task classification result.

[0084] Compared to the problems described in the background technology, this invention, by acquiring traffic data streams in the target network environment, provides raw data support for accurately identifying traffic protocols and analyzing load content. Simultaneously, it allows for a direct understanding of the basic operational status of network traffic, ensuring the accuracy and effectiveness of the subsequent detection and classification process. By collecting topology data sequences from the traffic topology map and analyzing the corresponding data transmission types, this invention transforms the scattered protocol-load association information and device transmission relationships in the map into an ordered and analyzable dataset. This provides complete data support for the subsequent accurate analysis of data transmission types, avoiding analytical biases caused by data fragmentation and improving the efficiency and accuracy of subsequent analysis stages. Furthermore, by identifying abnormal traffic peaks and core transmission hotspots in the traffic transmission curve, this invention can promptly capture traffic anomalies deviating from normal transmission trends, enabling rapid location of potential network threats. This invention provides direct evidence for anomalies such as network attacks and link congestion, avoiding response delays caused by hidden anomalies. This ensures that the subsequent deep packet inspection framework is more aligned with actual network operation, improving overall detection efficiency. Furthermore, based on the network cache state, this invention constructs a deep packet inspection framework corresponding to the target network environment. It can dynamically allocate detection resources based on cache capacity saturation and load pressure, avoiding ineffective loss of detection capabilities when the cache is overloaded, enhancing resilience in complex traffic scenarios, and achieving synergistic optimization of detection efficiency and network transmission. Finally, based on the traffic coordination index, this invention reconstructs the data coordination tasks corresponding to the traffic data stream. Based on the latency and misjudgment coordination characteristics revealed by the index, it can dynamically adjust the task allocation strategy in each classification coordination stage, allowing high-efficiency stages to bear more load and low-efficiency stages to be optimized first, achieving a dynamic balance between resources and tasks, and enhancing the framework's resilience in handling complex traffic. Therefore, the network traffic deep packet inspection multi-task classification method and system provided by this invention can improve the detection and classification efficiency of multi-task traffic in complex network environments. like Figure 3 The diagram shown is a functional block diagram of a network traffic deep packet inspection multi-task classification system according to the present invention.

[0085] The deep packet inspection (DBI) multi-task classification system 200 for network traffic described in this invention can be installed in an electronic device. Depending on the functions implemented, the DBI multi-task classification system may include a graph construction module 201, a curve fitting module 202, a state analysis module 203, an exponent calculation module 204, and a result generation module 205. The modules described in this invention can also be referred to as units, which are a series of computer program segments that can be executed by the processor of an electronic device and perform a fixed function, and are stored in the memory of the electronic device.

[0086] In this embodiment of the invention, the functions of each module / unit are as follows: The topology construction module 201 is used to acquire traffic data streams in the target network environment, identify the traffic protocols and load content corresponding to the traffic data streams, and generate a traffic topology map associated with the traffic protocols and load content. The curve fitting module 202 is used to collect topology data sequences in the traffic topology map, parse the data transmission type corresponding to the topology data sequence, analyze the type-oriented mode corresponding to the data transmission type, and fit the traffic transmission curve under the type-oriented mode. The state parsing module 203 is used to identify abnormal traffic peaks and core transmission hotspots in the traffic transmission curve, calculate the traffic throughput threshold corresponding to the traffic data stream based on the abnormal traffic peaks and the core transmission hotspots, and parse the network cache status of the traffic throughput threshold in the target network environment. The index calculation module 204 is used to construct a deep packet inspection framework corresponding to the target network environment based on the network cache state, monitor the real-time classification path in the deep packet inspection framework, collect the processing delay value and classification misclassification rate in the real-time classification path, and calculate the traffic coordination index corresponding to the traffic data stream based on the processing delay value and the classification misclassification rate. The result generation module 205 is used to reconstruct the data coordination task corresponding to the traffic data stream based on the traffic coordination index, identify the task classification protocol in the data coordination task, and generate the task classification result of the traffic data stream when it is oriented towards multiple tasks based on the task classification protocol.

[0087] In detail, the modules in the network traffic deep packet inspection multi-task classification system 200 described in this embodiment of the invention employ the same methods as described above. Figure 1 This method employs the same technical means as the deep packet inspection multi-task classification method for network traffic described above, and can produce the same technical effect, so it will not be elaborated here.

[0088] It will be apparent to those skilled in the art that the present invention is not limited to the details of the exemplary embodiments described above, and that the present invention can be implemented in other specific forms without departing from the spirit or essential characteristics of the present invention.

[0089] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. In the above multiple embodiments, each embodiment can be combined with each other or independent. Deleting any one of them will not affect the technical implementation of other embodiments. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention.

Claims

1. A multi-task classification method for deep packet inspection of network traffic, characterized in that, The method includes: Acquire traffic data streams in the target network environment, identify the traffic protocols and load content corresponding to the traffic data streams, and generate a traffic topology map associated with the traffic protocols and load content; Collect topology data sequences from the traffic topology map, parse the data transmission types corresponding to the topology data sequences, analyze the type-oriented patterns corresponding to the data transmission types, and fit the traffic transmission curves under the type-oriented patterns; Identify abnormal traffic peaks and core transmission hotspots in the traffic transmission curve, calculate the traffic throughput threshold corresponding to the traffic data stream based on the abnormal traffic peaks and the core transmission hotspots, and analyze the network cache status of the traffic throughput threshold in the target network environment. Based on the network cache state, a deep packet inspection framework corresponding to the target network environment is constructed. The real-time classification path in the deep packet inspection framework is monitored, and the processing latency value and classification misclassification rate in the real-time classification path are collected. Based on the processing latency value and the classification misclassification rate, the traffic coordination index corresponding to the traffic data stream is calculated. Based on the traffic coordination index, the data coordination task corresponding to the traffic data stream is reconstructed, the task classification protocol in the data coordination task is identified, and based on the task classification protocol, the task classification result of the traffic data stream when it is oriented towards multiple tasks is generated.

2. The multi-task classification method for deep packet inspection of network traffic as described in claim 1, characterized in that, The generation of the traffic topology map associated with the traffic protocol and the load content includes: Parse the protocol header information in the traffic protocol; Identify the protocol attributes in the protocol header information; Analyze the content identifiers corresponding to the content features in the load content; Associat the protocol attributes with the protocol-content mapping relationship corresponding to the content identifier; Based on the protocol-content mapping relationship, a traffic topology map is generated that associates the traffic protocol with the load content.

3. The network traffic deep packet inspection multi-task classification method as described in claim 2, characterized in that, The association of the protocol attribute with the protocol-content mapping relationship corresponding to the content identifier includes: Identify the key protocol fields in the protocol attributes; Extract the feature value sequence corresponding to the content identifier; According to predefined rules, the key protocol fields are matched with the feature value sequence to obtain the matching field sequence; Analyze the mapping confidence corresponding to the matching field sequence; Based on the mapping confidence, a protocol-content mapping relationship is generated between the protocol attributes and the content identifiers.

4. The multi-task classification method for deep packet inspection of network traffic as described in claim 1, characterized in that, The step of constructing a deep packet inspection framework corresponding to the target network environment based on the network cache state includes: Analyze the cache utilization rate corresponding to the network cache status; Based on the cache utilization rate, the traffic detection level in the target network environment is determined; Based on the traffic detection level, environmental deep packets in the target network environment are identified; Query the performance test data corresponding to the environment depth package; Based on the performance testing data, a deep packet inspection framework corresponding to the target network environment is constructed.

5. The network traffic deep packet inspection multi-task classification method as described in claim 4, characterized in that, The step of identifying environmental deep packets in the target network environment based on the traffic detection level includes: Analyze the level detection baseline corresponding to the traffic detection level; The level detection baseline is deeply correlated with the real-time traffic status in the target network environment to obtain candidate depth packets; Generate a list of depth packet identifiers corresponding to the candidate depth packets; Extract the list protocol factors from the deep packet identifier list; Based on the list of protocol factors, the environment depth packets in the target network environment are identified.

6. The multi-task classification method for deep packet inspection of network traffic as described in claim 1, characterized in that, The step of calculating the traffic throughput threshold corresponding to the traffic data stream based on the abnormal traffic peak and the core transmission hotspot includes: Analyze the amplitude characteristics and frequency of occurrence of the abnormal traffic peaks; Determine the upper limit value of the flow rate corresponding to the amplitude feature and the frequency of occurrence; Extract the continuous transmission rate of the core transmission hotspot; Based on the upper limit of traffic and the continuous transmission rate, the throughput load value corresponding to the traffic data stream is determined; Based on the throughput load value, the traffic throughput threshold corresponding to the traffic data stream is calculated using the following formula: ; in, This indicates the traffic throughput threshold corresponding to the traffic data stream. Indicates network bandwidth capacity. This represents the throughput load value. This indicates the average duration of the core transmission hotspot. This indicates the frequency of abnormal traffic peaks.

7. The multi-task classification method for deep packet inspection of network traffic as described in claim 1, characterized in that, The step of calculating the traffic coordination index corresponding to the traffic data stream based on the processing latency value and the classification misclassification rate includes: Analyze the co-variation characteristics of the processing delay value and the classification misclassification rate; Extract anomalous cooperative markers from the cooperative change features; Based on the aforementioned anomaly co-labeling, the classification coordination step in the deep packet detection framework is mapped; Statistically analyze the delay fluctuation value and misjudgment fluctuation value in the classification and coordination process; Based on the latency fluctuation value and the misjudgment fluctuation value, the traffic coordination index corresponding to the traffic data stream is calculated using the following formula: ; in, This represents the traffic coordination index corresponding to the traffic data stream. This indicates the total number of the classification and coordination links. This represents the quantity index corresponding to the classification coordination step. Indicates the first The delay fluctuation value corresponding to each classification coordination link Indicates the reference delay value. Indicates the first The error fluctuation value corresponding to each classification coordination link This indicates the reference misjudgment rate.

8. The multi-task classification method for deep packet inspection of network traffic as described in claim 1, characterized in that, The fitting of the traffic transmission curve under the type-oriented mode includes: Extract the guidance transmission timing from the guidance mode of the aforementioned type; Based on the aforementioned guidance transmission timing, the traffic sampling frequency corresponding to the type of guidance mode is set; Based on the traffic sampling frequency, the real-time traffic sequence in the guided transmission time sequence is collected; Extract traffic data points from the real-time traffic sequence; Based on the traffic data points, fit the traffic transmission curve under the type-oriented mode.

9. The multi-task classification method for deep packet inspection of network traffic as described in claim 1, characterized in that, The data coordination task, which reconstructs the data flow corresponding to the traffic data stream based on the traffic coordination index, includes: Analyze the coordination equilibrium threshold corresponding to the traffic coordination index; Based on the coordination and balancing threshold, obtain the data coordination task corresponding to the traffic data stream; Query the task coordination rules corresponding to the data coordination task; Generate the task sequence data in the task coordination rules; Based on the task sequence data, reconstruct the data coordination task corresponding to the traffic data stream.

10. A multi-task classification system for deep packet inspection of network traffic, characterized in that, The system includes: The topology graph construction module is used to acquire traffic data streams in the target network environment, identify the traffic protocols and load content corresponding to the traffic data streams, and generate a traffic topology graph associated with the traffic protocols and load content. The curve fitting module is used to collect topology data sequences in the traffic topology map, parse the data transmission types corresponding to the topology data sequences, analyze the type-oriented mode corresponding to the data transmission types, and fit the traffic transmission curve under the type-oriented mode. The status resolution module is used to identify abnormal traffic peaks and core transmission hotspots in the traffic transmission curve, calculate the traffic throughput threshold corresponding to the traffic data stream based on the abnormal traffic peaks and the core transmission hotspots, and resolve the network cache status of the traffic throughput threshold in the target network environment. An index calculation module is used to construct a deep packet inspection framework corresponding to the target network environment based on the network cache state, monitor the real-time classification path in the deep packet inspection framework, and collect the processing latency value and classification misclassification rate in the real-time classification path. Based on the processing latency value and the classification misclassification rate, it calculates the traffic coordination index corresponding to the traffic data stream. A result generation module is used to reconstruct the data coordination task corresponding to the traffic data stream based on the traffic coordination index, identify the task classification protocol in the data coordination task, and generate the task classification result of the traffic data stream when it is oriented towards multiple tasks based on the task classification protocol.