Network traffic analysis and management system based on artificial intelligence
The generation of high-quality feature vectors through the protocol semantic deep deconstruction engine and the self-evolution knowledge base, solving the problem of insufficient understanding of protocol semantics in network traffic analysis, and improving the analysis performance and system adaptability of artificial intelligence models.
Patent Information
- Application Number
- CN202510878675.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-27
- Publication Date
- 2025-08-22
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
In the network traffic analysis and management, it is difficult to automatically and intelligently understand and extract the context semantics of protocol fields and encrypted traffic metadata information, resulting in limited ability to generate feature vectors, affecting the generalization ability and analysis accuracy of artificial intelligence models.
Using the protocol semantics deep deconstruction engine, context-aware dynamic feature generator, closed-loop verification optimization module and self-evolution knowledge base, high-quality and high-distinction feature vectors are generated through incremental protocol syntax trees, lightweight domain knowledge graphs and deformable convolution kernels, and feature quality is ensured through the dual closed-loop verification mechanism.
It realizes intelligent analysis of unknown protocols and encrypted traffic, improves the generalization ability and detection accuracy of artificial intelligence models in complex network environments, reduces operation and maintenance costs, and ensures system adaptability and online service continuity.
Smart Images

Figure CN120528845A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of network traffic analysis and management, and in particular to an artificial intelligence-based network traffic analysis and management system. Background Art
[0002] In the field of network traffic analysis and management, especially when it comes to the application of artificial intelligence (AI), existing systems based on electronic digital data processing face a significant bottleneck. Raw network traffic data is inherently high-dimensional, sparse, and highly heterogeneous, containing a wealth of protocol semantics and encrypted traffic metadata. Current technical solutions for feature engineering suffer from two main limitations: First, they rely heavily on domain experts to manually design feature extraction rules and templates based on their experience. This approach is not only inefficient and labor-intensive, but also lacks adaptability to the constant emergence of new protocols, new application models, or unknown attack behaviors, resulting in unstable feature quality and difficulty in effectively capturing deep, dynamically changing contextual meaning. Second, they employ general automated feature extraction algorithms, such as basic numerical transformations or standard deep learning models, to directly process raw data. While these methods reduce reliance on manual effort, they often lack an understanding of the complex logical relationships between protocol fields, the true semantics of field values in specific contexts, and the patterns hidden in encrypted traffic metadata. They tend to perform shallow processing and fail to intelligently parse and utilize the key semantic information contained in the data. This results in the generated feature vectors having limited expressive power and low discriminability, failing to fully reflect the essential characteristics of network traffic. Ultimately, this feature engineering flaw severely restricts the generalization ability, analytical accuracy, and adaptability of subsequent AI models to complex and changing network environments, becoming a core obstacle to improving the overall performance of the system. Therefore, the current technical problem that needs to be solved is how to automatically and intelligently understand and extract the contextual semantics of protocol fields and encrypted traffic metadata information in raw network traffic data to generate high-quality, highly discriminative, and generalizable feature vectors, thereby overcoming the over-reliance on expert experience and improving the analytical performance of AI models. Summary of the Invention
[0003] To achieve the above objectives, the present invention is implemented through the following technical solutions: an artificial intelligence-based network traffic analysis and management system, including a network traffic data acquisition module, a protocol semantic deep deconstruction engine, a context-aware dynamic feature generator, a closed-loop verification and optimization module, a self-evolving knowledge base, and an AI analysis engine: The protocol semantics deep deconstruction engine is connected to the network traffic data acquisition module to analyze the protocol syntax structure, field semantics and context association of the traffic data; The context-aware dynamic feature generator is connected to the protocol semantics deep deconstruction engine to generate a feature vector of fused semantics; The closed-loop verification and optimization module bidirectionally connects the context-aware dynamic feature generator and the self-evolving knowledge base to perform double verification and iterative optimization on the feature vector; The AI analysis engine is connected to the closed-loop verification optimization module and uses the optimized feature vector to perform traffic analysis tasks.
[0004] Preferably, the protocol semantic deep deconstruction engine includes an incremental protocol syntax tree unit, a lightweight domain knowledge graph unit and a temporal attention analysis unit; The incremental protocol syntax tree unit dynamically constructs a syntax tree of an unknown protocol; The lightweight domain knowledge graph unit maps field values to predefined semantic nodes and identifies logical dependencies between nodes; The temporal attention analysis unit captures dynamic combination patterns of multiple field values within the same data stream.
[0005] Preferably, the context-aware dynamic feature generator comprises a structured feature generation path, an unstructured feature generation path and a gated fusion unit; The structured feature generation path converts the semantic deconstruction result into a semantic vector with a weight label; The unstructured feature generation path processes the original message byte stream through a deformable convolutional semantic encoder; The gated fusion unit dynamically weights and fuses the dual-path outputs.
[0006] Preferably, the deformable convolutional semantic encoder includes three layers of convolution kernels: The first layer of fixed convolution kernel extracts the fixed pattern of protocol header; The second layer of deformable convolution kernel dynamically adjusts the kernel shape according to the grammatical structure output by the incremental protocol syntax tree; The third layer of spatiotemporal convolution kernel integrates cross-message temporal features.
[0007] Preferably, the closed-loop verification optimization module includes a semantically driven adversarial generative network screen and a feature utility evaluation screen; The semantic-driven adversarial generative network screen generates simulated traffic by reconstructing feature vectors, determines semantic consistency based on the protocol syntax tree and the knowledge graph, and screens out features with reconstruction errors exceeding a threshold. The feature utility evaluation screen applies semantic label mask perturbations to feature vectors and determines feature discrimination through the stability of lightweight proxy model output.
[0008] Preferably, the semantically driven generative adversarial network sieve comprises a generator and a discriminator; The generator deconstructs the feature vector into simulated traffic data; The discriminator calls the semantic rules in the lightweight domain knowledge graph unit to perform consistency verification.
[0009] Preferably, the feature utility evaluation screen performs the following operations: Randomly mask the semantic labels of the feature vector; Input the masked vector to the lightweight proxy model and calculate the output standard deviation; If the standard deviation exceeds the preset tolerance threshold, feature regeneration is triggered.
[0010] Preferably, the self-evolving knowledge base is updated in the following manner: When the incremental protocol syntax tree unit detects an unrecognized protocol pattern, it initiates offline clustering analysis to generate new semantic nodes; After the new semantic nodes are double-verified by the semantic-driven adversarial generative network screen and the feature utility evaluation screen, the knowledge graph nodes and association rules are automatically expanded to lightweight domain knowledge graph units; For encrypted traffic metadata, perform protocol behavior timing chain transformation and risk mapping processing.
[0011] Preferably, the processing method for encrypted traffic metadata is: converting the certificate chain length, cipher suite selection and session resumption frequency in the TLS handshake sequence into a protocol behavior timing chain; mapping the timing chain to risk semantic labels through lightweight domain knowledge graph units.
[0012] Preferably, the AI analysis engine includes a pre-trained deep learning model, and the input of the deep learning model is the feature vector output by the closed-loop verification optimization module.
[0013] The present invention provides a network traffic analysis and management system based on artificial intelligence. It has the following beneficial effects: This AI-based network traffic analysis and management system achieves intelligent parsing of unknown protocols and encrypted traffic through the dynamic syntax tree expansion and knowledge graph conflict resolution mechanism of the protocol semantic deep deconstruction engine, avoiding manually predefined rules; combined with the dynamic adaptation of field boundaries of deformable convolution kernels and gated fusion weight distribution, it generates feature vectors with both semantic accuracy and pattern integrity, solving the bottleneck of traditional feature engineering relying on expert experience; the dual closed-loop verification mechanism ensures adaptive optimization of feature quality through three-level semantic reconstruction inspection and weighted graded perturbation evaluation, thereby improving the generalization ability and detection accuracy of AI models in complex network environments.
[0014] This AI-based network traffic analysis and management system features a self-evolving knowledge base based on cluster analysis and a dual-authentication access mechanism, enabling automatic digestion of new protocols and attack patterns and generation of risk labels, enabling the system to continuously adapt to changes in the network environment. The intelligent analysis engine's multi-task collaboration mechanism, combined with a confidence decay strategy and cross-task knowledge transfer, reduces false alarm rates while building a closed-loop defense enhancement system. Dynamic model incremental training and local feature review mechanisms ensure self-optimization while ensuring online service continuity. This complete closed-loop process encompasses "protocol analysis - feature generation - verification and evolution - analysis and decision-making," reducing operational and maintenance costs while improving network traffic analysis efficiency and security management. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] Figure 1 It is a schematic diagram of the overall framework of the present invention; Figure 2 It is a schematic diagram of the process of the present invention. DETAILED DESCRIPTION
[0016] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0017] See also Figure 1 and Figure 2 The present invention provides a technical solution: an artificial intelligence-based network traffic analysis and management system, including a network traffic data acquisition module, a protocol semantic deep deconstruction engine, a context-aware dynamic feature generator, a closed-loop verification and optimization module, a self-evolving knowledge base, and an AI analysis engine: The protocol semantics deep deconstruction engine is connected to the network traffic data collection module to analyze the protocol syntax structure, field semantics and context association of traffic data; The context-aware dynamic feature generator connects the protocol semantic deep deconstruction engine to generate feature vectors with fused semantics; The closed-loop verification and optimization module bidirectionally connects the context-aware dynamic feature generator and the self-evolving knowledge base to perform dual verification and iterative optimization on the feature vector; The AI analysis engine is connected to the closed-loop verification optimization module and uses the optimized feature vectors to perform traffic analysis tasks.
[0018] It should be further explained that during the specific implementation process, the network traffic data collection module captures raw traffic data packets in real time and transmits them to the protocol semantic deep deconstruction engine. This engine first dynamically parses the packet structure through an incremental protocol syntax tree: when an unknown protocol field is identified, the syntax tree branches are automatically expanded to construct a hierarchical relationship between fields. At the same time, a lightweight domain knowledge graph is called to map field values to predefined semantic nodes, such as mapping the HTTP User-Agent value to "mobile iOS device", and identifying logical dependencies between nodes, such as "login request must include credential fields". The temporal attention analysis unit is further used to capture the dynamic combination patterns of multiple fields in the same data stream, such as associating "HTTP method = POST", "URL path = / api / login", and "response code = 401" to generate the contextual semantics of "authentication failed".
[0019] After receiving the deconstruction results, the context-aware dynamic feature generator starts a two-path processing: Structured path: Convert semantic node combinations into weighted label vectors, such as [Device type: mobile, Behavior: Authentication failed, Weight: 0.92]. The weight is dynamically calculated based on the node association strength in the knowledge graph. Unstructured path: The original byte stream is processed through a deformable convolutional semantic encoder, namely: the first layer of fixed convolution kernels extracts the protocol header flag features; the second layer of convolution kernels dynamically deforms according to the field boundaries output by the syntax tree to adapt to variable-length fields; the third layer of spatiotemporal convolution fuses the time series features of adjacent messages.
[0020] The dual-path outputs are dynamically weighted by the gated fusion unit to generate the final feature vector, where the structured semantic weight accounts for 70% and the unstructured spatiotemporal features account for 30%.
[0021] The closed-loop validation optimization module performs a double check on the feature vector: First-level verification: The semantically driven generative adversarial network (SdGAN) inputs the feature vector into the generator and deconstructs it into simulated traffic data. The discriminator compares the logical consistency of the fields in the simulated traffic with those in the real traffic based on the protocol syntax tree rules and knowledge graph semantics. For example, it checks whether the "authentication failure" field in the simulated traffic contains the required response code field. If the reconstruction error exceeds a threshold, it is marked as a low-confidence feature. Second test: The feature utility evaluation screen randomly masks the key semantic labels of the feature vector, such as blocking the "device type" label, and inputs it into the lightweight proxy model to perform the classification task. If the standard deviation of the masked model output exceeds the tolerance threshold, such as the fluctuation of the prediction result is greater than 10%, it is judged as a low-discrimination feature.
[0022] Features that fail any test trigger regeneration, and features that pass the test are fed into the AI analysis engine.
[0023] The self-evolving knowledge base is updated as follows: When the incremental protocol syntax tree detects unrecognized protocol patterns, such as new IoT protocols, offline clustering analysis is initiated to generate candidate semantic nodes. These candidate nodes are then double-checked by SdGAN semantic consistency verification and feature utility evaluation before being expanded into the knowledge graph. For encrypted traffic, the system converts metadata such as certificate chain length and cipher suite selection in the TLS handshake sequence into a time-series chain, which is then mapped to risk labels in the knowledge graph, such as "short certificate chain + unconventional cipher suite → high risk."
[0024] The AI analysis engine receives the optimized feature vectors and performs intrusion detection or protocol classification. If traffic analysis results show persistent anomalies, they are automatically fed back to the closed-loop verification module to initiate feature review.
[0025] The protocol semantic deep deconstruction engine includes an incremental protocol syntax tree unit, a lightweight domain knowledge graph unit, and a temporal attention analysis unit; The incremental protocol syntax tree unit dynamically constructs the syntax tree of the unknown protocol; The lightweight domain knowledge graph unit maps field values to predefined semantic nodes and identifies logical dependencies between nodes; The temporal attention analysis unit captures the dynamic combination patterns of multiple field values within the same data stream.
[0026] It should be further explained that, in the specific implementation process, when network traffic data is input into the protocol semantic deep deconstruction engine, the incremental protocol syntax tree unit first performs dynamic parsing, including: if the data packet matches the known protocol HTTP, directly call the pre-stored syntax tree to extract fields, request line, header, and body; if an unknown protocol flag is detected, such as the new IoT protocol signature code, the branch expansion mechanism is activated; The process of starting the branch extension mechanism includes: inferring the hierarchical structure based on the relationship between field length and position, such as identifying the field with the first byte of 0x10 as the control header; when similar field sequences appear repeatedly in the same data stream, automatically creating subtree nodes and marking them as variable-length field groups; skipping content parsing for encrypted traffic, but recording metadata timing patterns, such as TLS handshake packet intervals.
[0027] The lightweight domain knowledge graph unit is activated synchronously, including semantic mapping of parsed fields and cross-message correlation of data streams by the temporal attention analysis unit. Semantic mapping of parsed fields includes: directly matching basic fields to predefined nodes, such as IP address → geographic location; and splitting composite fields to associate multiple nodes, such as "User-Agent: Mozilla / 5.0 (Android)" → operating system = Android, device type = mobile. When field values conflict, such as when the same IP address is mapped to both the "server" and "client" nodes, logic dependency verification is initiated: verifying the constraints between nodes, such as "client" cannot be "server" at the same time; If the constraint is violated, the label is overwritten based on the timing context priority. For example, if the IP is the request initiator in subsequent traffic, it is marked as "client".
[0028] The temporal attention analysis unit's cross-message correlation of data streams involves key field focus and dynamic combination detection. Key field focus involves calculating the frequency and entropy of fields and screening high-value fields, such as frequently occurring "cookies" or encrypted payloads with high entropy. Dynamic combination detection includes: when a preset combination pattern is detected, such as "HTTP method = GET" + "URL contains / admin" + "response code = 403", a strong correlation semantic tag is generated; For undefined abnormal combinations, such as "DNS query domain name length is greater than 255 bytes", it is marked as pending verification mode and the knowledge graph update process is triggered.
[0029] The three-unit architecture that deeply couples protocol analysis, semantic mapping, and context association forms an inseparable technical closed loop, solving the problem of insufficient understanding of protocol semantics.
[0030] The context-aware dynamic feature generator includes a structured feature generation path, an unstructured feature generation path, and a gated fusion unit; The structured feature generation path converts the semantic deconstruction results into semantic vectors with weighted labels; The unstructured feature generation path processes the original message byte stream through a deformable convolutional semantic encoder; The gated fusion unit dynamically weights and fuses the dual-path outputs.
[0031] It should be further explained that, during implementation, the structured semantic labels output by the protocol semantic deep deconstruction engine are synchronously fed into the context-aware dynamic feature generator along with the unstructured raw byte stream. The structured feature generation path performs semantic vector conversion, including extracting the association strength values of semantic nodes from the knowledge graph. For example, the association strength between "login failure" and "response code 401" is 0.95. This is combined with the contextual confidence values output by the temporal attention unit (e.g., if three consecutive failures occur, the confidence is increased by 0.2), and the label weights are dynamically calculated. When multiple sets of semantic labels exist for the same traffic flow, such as "device type = mobile" and "geographic location = overseas," the relevant labels are merged based on the node distance in the knowledge graph. If the distance is less than a threshold, the labels are merged into "overseas mobile terminal high-risk access," and a compact semantic vector is subsequently generated.
[0032] The unstructured feature generation path starts the deformable convolutional semantic encoder to process the raw byte stream, which includes: The first layer of fixed convolution kernel slides and scans the packet header to extract protocol flag features, which include TCP port number and HTTP method type. The second layer of deformable convolution kernels dynamically deforms according to the field boundary information output by the incremental protocol syntax tree: if the field length exceeds the preset value, such as a URL length greater than 200 characters, the convolution kernel width is automatically expanded to cover the variable-length content; when the syntax tree marks a field as an encrypted segment, such as TLSApplicationData, it switches to sparse sampling mode and skips deep parsing; The third layer of spatiotemporal convolution kernel associates adjacent messages. The process includes the following: For request-response traffic, such as HTTP, capture the correlation pattern of request and response packets along the timeline; For continuous streaming traffic, such as video streaming, statistical features are aggregated according to fixed time windows.
[0033] After receiving the dual-path output, the gated fusion unit starts dynamic weighting. The process includes the following: Calculate the average weighted confidence of the structured semantic vector and compare it with the entropy value of the unstructured features. If the confidence is greater than the entropy value, the weight of the structured path is increased. If an unknown protocol or encrypted traffic is detected, the proportion of unstructured paths is automatically increased. For conflicting features, such as the structured path tag "normal access" and the unstructured path detection abnormal byte pattern, the process of initiating secondary verification includes: isolating the conflicting feature fragment, triggering the local re-analysis of the closed-loop verification module, and overwriting the original weight according to the verification result.
[0034] The three-component architecture of structured semantic compression, dynamic adaptation of deformable convolution, and intelligent fusion of gated features addresses the issue of limited feature expression capabilities. Compared to purely structured rule-based or deep learning-based approaches, the dual-path collaborative mechanism ensures semantic interpretability while improving adaptability to complex traffic.
[0035] The deformable convolutional semantic encoder consists of three layers of convolutional kernels: The first layer of fixed convolution kernel extracts the fixed pattern of protocol header; The second layer of deformable convolution kernel dynamically adjusts the kernel shape according to the grammatical structure output by the incremental protocol syntax tree; The third layer of spatiotemporal convolution kernel integrates cross-message temporal features.
[0036] It should be further explained that in the specific implementation process, after the original message byte stream is input into the deformable convolutional semantic encoder, the first layer of fixed convolution kernel slides along the message header and matches key flag bits using a preset protocol template. The process includes: when an HTTP message is detected, the fixed convolution kernel is positioned at the beginning of the request line and extracts the method type and target URL features, where the method type includes GET and POST; Identify the source / destination port number combination for TCP traffic. If the port belongs to a predefined high-risk service, automatically enhance the feature extraction depth. If the message header contains an encryption identifier, skip the payload parsing but record the encryption protocol type feature.
[0037] The second layer of deformable convolution kernel receives the field boundary coordinates output by the incremental protocol syntax tree. The process includes: For fixed-length fields, such as IP addresses with a fixed length of 4 bytes, a standard square convolution kernel is used for scanning; For variable-length fields, such as HTTP URLs, the following approach is used: the actual length is calculated based on the start and end positions of the field marked in the syntax tree. If the length exceeds the default size of the convolution kernel, such as greater than 128 bytes, the kernel width is expanded along the length direction to the actual size of the field. After expansion, the convolution kernel adopts a sparse sampling strategy to enhance feature capture at key delimiters, where key delimiters include question marks and equal signs. For encrypted fields, the following method is used: switch to metadata extraction mode and scan only external features. The external features include packet length and time interval, where the packet length is the byte length of the data packet; When the number of consecutive encrypted packets exceeds the threshold, a time series statistical feature is generated, which includes the average packet length and transmission rate.
[0038] The process of performing cross-message association in the third-layer spatiotemporal convolution kernel is as follows: For request-response interactions, we align request and response packets along the time axis and use a three-dimensional convolution kernel to capture feature variation patterns, such as the correlation between request packet size and response latency. For continuous streaming transmission, the data stream is divided into fixed time windows and statistical features are aggregated within the window, including packet size variance and number of transmission interruptions. If an abnormal pattern is detected, such as a high-frequency burst of small packets, a flow behavior anomaly label is generated.
[0039] The three-layer convolutional architecture, featuring dynamic kernel deformation, intelligent spatiotemporal correlation, and multi-layer fault tolerance, addresses the inability of general algorithms to understand protocol structures. Compared to fixed-size CNNs or independent time series models, this achieves a deep coupling of protocol structure and feature extraction, improving feature generation quality.
[0040] The closed-loop verification and optimization module includes a semantically driven adversarial generative network screen and a feature utility evaluation screen; The semantically driven adversarial generative network screen generates simulated traffic by reconstructing feature vectors, determines semantic consistency based on the protocol syntax tree and knowledge graph, and screens out features with reconstruction errors exceeding a threshold. The feature utility evaluation screen applies semantic label mask perturbations to the feature vectors and determines the feature discrimination through the stability of the lightweight proxy model output.
[0041] It should be further explained that in the specific implementation process, after the context-aware dynamic feature generator outputs a feature vector, the semantic-driven adversarial network screen starts the verification process, including: the generator receives the feature vector, deconstructs the simulated traffic based on the protocol syntax tree rules and the semantic relationship of the knowledge graph, and partially restores the structured features to protocol fields, such as converting [Device: Mobile] to generate the HTTP User-Agent header; unstructured features are converted into raw byte approximations through deconvolution; The discriminator compares the semantic consistency of simulated traffic and real traffic. The process is as follows: first, the integrity of key fields is checked. For example, "Login Failure" in simulated traffic must include the response code field; then the field logical dependencies are verified; if a field is detected to be missing or inconsistent, the reconstruction error is marked as exceeding the threshold, triggering feature regeneration.
[0042] The process of parallel execution of model feedback verification for the feature utility evaluation screen is as follows: First, the semantic labels of the feature vector are randomly masked and disturbed. High-weight labels are blocked probabilistically, such as blocking the "high-risk operation" label with a 20% probability. Low-weight labels are masked in batches, such as blocking all labels with a weight less than 0.3 at once. The perturbed vector is then fed into a lightweight proxy model, i.e., a simplified AI analysis engine. The process is as follows: The same traffic analysis task is performed, and the standard deviation of the inference results is recorded for N consecutive times. If the standard deviation exceeds the tolerance threshold, such as the result fluctuates wildly between "normal" and "suspicious", it is determined that the feature discrimination is insufficient: the knowledge graph node weight adjustment is triggered, or the local reconstruction process of the feature generator is started.
[0043] Among them, the dual-screen collaborative mechanism includes the following: when the first screen reports an error, the second screen suspends execution and prioritizes semantic integrity; if only the second screen reports an error, the generator intervenes to enhance feature significance, such as strengthening the weight of abnormal labels; for encrypted traffic features, the semantic consistency check is skipped and only model utility verification is performed.
[0044] The dual-verification architecture, which includes semantic reconstruction testing, intelligent perturbation assessment, and dynamic collaborative fault tolerance, addresses the pain point of uncontrollable feature quality. Compared to single statistical testing or manual review, it achieves automated, multi-dimensional, and adaptive feature quality assurance.
[0045] The semantically driven adversarial generative network sieve consists of a generator and a discriminator; The generator deconstructs the feature vector into simulated traffic data; The discriminator calls the semantic rules in the lightweight domain knowledge graph unit for consistency verification.
[0046] It should be further explained that in the specific implementation process, when the feature vector is input into the semantic-driven adversarial network sieve, the generator starts the deconstruction process as follows: For the structured features, semantic labels and weights are extracted from the feature vectors, and the original fields are restored based on the field structure rules stored in the protocol syntax tree unit, including: Map the "Device Type: Mobile" tag back to the User-Agent field value of the HTTP protocol. If the tags have a composite relationship, such as "Overseas Mobile High-Risk Access," they are split into independent fields based on the relevance of the knowledge graph nodes. For features that lack required fields, such as HTTPS traffic lacking certificate chain information, default values are automatically filled in and marked as simulated fields. For the unstructured feature part, the original byte stream is reconstructed through the deconvolution operation, including: The third layer of spatiotemporal convolutional features is decoded into message time series; The second layer of deformable convolutional features restores field-level byte segments; The first layer fixes the convolution feature reconstruction protocol header flag; If a feature fuzzy area is detected, such as a low-confidence feature in an encrypted segment, an approximate value is generated by interpolation of adjacent messages.
[0047] After receiving the simulated traffic output by the generator, the discriminator performs three-level consistency verification. The process is as follows: For syntax layer verification: call incremental protocol syntax tree unit rules, namely: check field position compliance, such as the TCP source port must be located in the 0-2 bytes of the header; verify variable-length field size limits, such as HTTP URL length ≤ 2048 bytes; if violations are found, such as the simulated DNS query packet length exceeding 512 bytes, mark the syntax error; For semantic layer verification: lightweight domain knowledge graph unit rules are called, namely: matching field values with semantic nodes, such as port 443 must be associated with "HTTPS service"; verifying logical dependency chains, such as "authentication response" must exist after "login request"; if a contradiction is detected, such as port 80 mapping "HTTP service" but the content contains SQL injection features, a semantic conflict warning is triggered; For context verification: Use the temporal attention analysis unit to compare with real traffic, that is, verify the rationality of cross-message response delay, such as DNS response time less than 3 seconds; detect the difference in session continuity between simulated traffic and real traffic, such as the lack of FIN packets in simulated TCP flows.
[0048] When any level fails to verify, a reconstruction error report is generated and transmitted to the feature generator. The process is as follows: Syntax errors trigger re-parsing of the original traffic; Semantic conflicts trigger a review of knowledge graph rules; Contextual anomalies force updates of temporal attention models.
[0049] Through a generator-discriminator architecture, rule-driven deconstruction, a three-level verification system, and targeted error feedback, this approach addresses the disconnect between features and protocol logic. Compared to general-purpose GANs used in generative verification schemes, this approach achieves precise simulation and deep verification under the constraints of protocol rules, improving the detection rate of reconstruction errors.
[0050] The Feature Utility Evaluation Screen performs the following operations: Randomly mask the semantic labels of the feature vector; Input the masked vector to the lightweight proxy model and calculate the output standard deviation; If the standard deviation exceeds the preset tolerance threshold, feature regeneration is triggered.
[0051] It should be further explained that in the specific implementation process, when the feature vector enters the feature utility evaluation screen, the system performs intelligent perturbation injection, and the process is as follows: For the semantic tag weight analysis phase: extract the weight values of all semantic tags in the feature vector and group them by weight interval, that is: High-weight labels (>0.8) are randomly masked using a single label to generate a few interference versions. For example, the "high-risk operation" label is randomly masked twice in five tests. Medium-weighted tags (0.3-0.8) implement masking of associated tag groups, such as simultaneously masking the frequently associated tags "mobile" and "overseas access"; For low-weight labels (<0.3), batch masking is performed, i.e., masking all underweighted secondary labels at once; For the perturbation execution stage: noise injection is applied to the unstructured feature part, that is: Gaussian noise is added to spatiotemporal convolution features; Deformable convolution features randomly set local areas to zero; The perturbed feature vector is input into the lightweight proxy model, which performs the same analysis task and records the continuous output. The process is as follows: Output stability assessment, including: for classification tasks, calculating the standard deviation of the category distribution of N perturbations; for regression tasks, calculating the fluctuation range of the predicted value; if the standard deviation exceeds the tolerance threshold, such as the category distribution jumping between "normal" and "suspicious", it is determined that the feature discrimination is insufficient; The triggering process of the hierarchical response mechanism includes: If only a single perturbation fails, the corresponding label in the knowledge graph is marked as "weak robustness"; subsequent feature generation reduces the weight of this label; If the masking of the associated label group fails, check the label logic dependency, such as whether there is a false association between "mobile terminal" and "overseas"; adjust the connection strength of the knowledge graph node; If batch masking causes the model to crash, the unstructured path reinforcement mode of the feature generator is activated; supplementary spatiotemporal features are generated to replace the failed semantic labels.
[0052] Through intelligent disturbance injection, hierarchical response mechanism, and cross-module collaborative evaluation process, the problem of uncontrollable feature discrimination is solved, and a dynamic, directional, and adaptive feature optimization closed loop is achieved.
[0053] The self-evolving knowledge base is updated in the following ways: When the incremental protocol syntax tree unit detects an unrecognized protocol pattern, it initiates offline clustering analysis to generate new semantic nodes; After the new semantic nodes are double-verified by the semantic-driven adversarial generative network screen and the feature utility evaluation screen, the knowledge graph nodes and association rules are automatically expanded to lightweight domain knowledge graph units; For encrypted traffic metadata, perform protocol behavior timing chain transformation and risk mapping processing.
[0054] It should be further explained that, in the specific implementation process, when the incremental protocol syntax tree unit detects an unrecognized protocol pattern, an offline clustering analysis process is initiated, which includes traffic segment capture, multi-dimensional clustering, and candidate node generation. Among them, traffic segment capture includes: intercepting a continuous message sequence containing an unknown protocol pattern and retaining the complete session context; multi-dimensional clustering includes: clustering based on message length distribution to identify similar structure groups, separating encrypted and non-encrypted groups through byte entropy analysis, and distinguishing interactive and streaming transmissions by time interval pattern grouping; candidate node generation includes: for strong cluster groups with intra-group similarity greater than a threshold, automatically extracting common features as new semantic nodes, such as "Zigbee control frame: function code = 0x03"; weak cluster groups generate draft nodes to be reviewed and mark the confidence level; The process of inputting candidate nodes into the dual verification screen includes semantic reconstruction verification and model utility verification. The semantic reconstruction verification includes: adding the candidate node to the temporary area of the knowledge graph, using the SdGAN generator to simulate traffic based on the node, and the discriminator to verify the protocol logic consistency of the simulated traffic and the actual unidentified traffic; Model effectiveness verification includes: converting candidate nodes into feature labels and injecting them into test vectors, using a lightweight proxy model to perform classification tasks, and evaluating output stability. If the new labels significantly improve classification accuracy and the fluctuations are controllable, the effectiveness is determined to be met. After passing the double verification, the knowledge graph is expanded, which includes node formalization and association rule establishment. Node formalization includes: writing strong clustering nodes directly into the protocol syntax tree unit; adding an "observation period" mark to the nodes to be reviewed, and requiring historical traffic backtesting. Association rule building includes: analyzing the co-occurrence frequency of new nodes and existing nodes, such as new protocols are often associated with specific IP segments; when the co-occurrence frequency exceeds a threshold, automatically creating a logical dependency edge, such as "Zigbee control frame → Industrial IoT device"; For encrypted traffic metadata, the system performs protocol behavior time chain transformation, including metadata extraction, time chain encoding, and risk label mapping. Metadata extraction includes: capturing certificate chain length, session resumption flag, and non-standard cipher suites from the TLS handshake sequence; recording the time interval and retry count during the key exchange phase; Temporal chain coding includes: encoding discrete metadata into a state sequence in the order of interaction, such as "short certificate → unconventional cipher suite → session resumption failure"; Risk label mapping includes: matching similar patterns in the historical attack database, such as associating "short certificate + unconventional cipher suite" with phishing attacks; when the matching confidence is greater than the threshold, creating a "high-risk encrypted session" node in the knowledge graph and associating the feature label.
[0055] Through an evolutionary process involving intelligent clustering and grading, dual-authentication access, and encrypted metadata conversion, the system overcomes its inability to adapt to new network protocols. Compared to static rule base solutions, this system achieves autonomous evolution of the knowledge system, improving its generalization capabilities.
[0056] The processing method for encrypted traffic metadata is to convert the certificate chain length, cipher suite selection, and session resumption frequency in the TLS handshake sequence into a protocol behavior time chain; and map the time chain to risk semantic labels through lightweight domain knowledge graph units. It should be further explained that in the specific implementation process, the system performs a refined protocol behavior time chain conversion for encrypted traffic metadata. The process is as follows: Metadata capture phase: locates key interaction nodes in the TLS handshake sequence and extracts non-content parameters based on the handshake phase sequence, including certificate chain features, cipher suite selection, and session behavior parameters. Certificate chain features include capturing the length of the leaf certificate body field, the number of intermediate certificates, and the root certificate authority identifier. Cipher suite selection includes recording the server's preferred suite type and unconventional suite flags. Session behavior parameters include extracting the number of session resumption attempts, session ticket length, and key exchange phase retransmission interval. During the time chain construction phase, discrete metadata is encoded into a state transition sequence in the order of interaction. This includes: generating a "Certificate Risk Level Increased" state node for certificate chain anomalies, such as a short certificate chain plus a self-signed root certificate; adding a "Non-standard Encryption Configuration" state node when an unconventional cipher suite is detected; and creating an "Abnormal Session Reset" state transition edge if the session fails to recover and then immediately reconnects. Risk mapping phase: Matching the timing chain with the knowledge graph risk pattern library, including: Exact matching mode: If the timing chain completely matches the historical attack characteristics, such as "short certificate → non-standard suite → session recovery failure" associated with malware communication, it is directly associated with the "high-risk encrypted session" label; Similarity matching mode: When a partial match occurs, such as when only a "short certificate + non-standard suite" is detected, the pattern similarity is calculated. If the similarity exceeds the first-level threshold, a temporary label of "medium-risk encrypted session" is generated; if the similarity is below the threshold but above the baseline, it is marked as "encryption behavior to be observed"; Zero-day pattern processing: For new time series chains without historical matching, this includes: extracting statistical anomaly features; launching a lightweight agent model for real-time behavior analysis; and creating candidate risk nodes and triggering a double verification process if the analysis results show malicious features, such as high-frequency heartbeat packets. The dynamic confidence management mechanism includes: Formal risk labels such as "high-risk encrypted session" are directly injected into the feature generator to generate corresponding labels in the structured path; Temporary labels must undergo traffic retrospective verification, including continuous monitoring of subsequent encrypted traffic behavior, such as whether it triggers data transmission; downgrading the confidence level if no malicious behavior is detected within 24 hours; and converting the label to a formal label if the associated historical attack pattern library updates the matching record. The behavior to be observed starts a dedicated monitoring channel, improves the spatiotemporal convolution sampling rate of the traffic, creates a "new encryption mode" observation node in the knowledge graph, and upgrades to a candidate node after the same mode is triggered three times cumulatively.
[0057] Through the encryption processing architecture of time chain transformation, dynamic risk mapping, and zero-day response mechanism, the defect of invisible encrypted traffic semantics is solved, and a deep understanding of encryption behavior and self-evolutionary perception of threats are achieved.
[0058] The AI analysis engine includes a pre-trained deep learning model, and the input of the deep learning model is the feature vector output by the closed-loop verification optimization module. It should be further explained that in the specific implementation process, after the optimized feature vector is input into the pre-trained deep learning model of the AI analysis engine, multi-task collaborative analysis is performed, including intrusion detection tasks and protocol classification tasks; among them, for intrusion detection tasks: the model extracts high-risk semantic labels in the feature vector as the first-level judgment basis; combined with the abnormal byte pattern in the unstructured feature, a high-risk alarm is generated when both are triggered at the same time; if only a single indicator is abnormal, the confidence decay mechanism is activated; In this regard, if only a single indicator is abnormal, the process of initiating the confidence decay mechanism includes: marking it as a low-confidence alarm when it occurs for the first time; upgrading it to an actionable alarm if it occurs three times in a row; For protocol classification tasks: the device type label output by the structured path directly limits the protocol range; the spatiotemporal characteristics of the unstructured path are matched against the protocol fingerprint library; when a new protocol label conflicts with an existing classification, the conflicting traffic segments are isolated and the protocol deconstruction engine is called to perform deep packet inspection arbitration; The dynamic model enhancement mechanism is activated when analyzing anomalies. The activation process includes local feature review and online fine-tuning of model parameters. The local feature review includes the following situations: When intrusion detection continuously generates false positives: extract the feature vector of the false positive traffic; request the closed-loop verification module to re-perform the semantic consistency check; if the check finds that the feature is distorted, trigger the reconstruction of the feature generator; When protocol classification remains ambiguous: increase the number of channels in the spatiotemporal convolution kernel; increase the weight of the unstructured path to a dominant position in the gated fusion unit; Online fine-tuning of model parameters involves collecting data on deviations between analysis results and actual traffic. When the accumulated deviation exceeds a threshold in a specific scenario, such as when the misclassification rate of IoT traffic increases, the model's core parameters are frozen, and only the last layer of classifiers is decoupled for incremental training. Training data comes from recent snapshots of similar traffic characteristics. The cross-task knowledge transfer mechanism includes the following: When the intrusion detection model identifies a new attack pattern, it extracts the spatiotemporal characteristic patterns of the attack traffic, such as abnormal heartbeat packet intervals, and converts them into auxiliary features for the protocol classification task, marking them as "potentially malicious protocols." When the same features reappear in the protocol classification, the threat level is automatically increased. When protocol classification discovers an unknown service port: the communication behavior characteristics of the port are captured, such as the number of connections within a fixed time window, and injected into the feature input layer of the intrusion detection model. If the behavior matches the historical attack pattern, a collaborative alarm is generated.
[0059] Through multi-task collaboration, dynamic enhancement, and knowledge transfer, the rigidity and lag of analytical models are resolved. Compared to static analytical models, closed-loop self-optimization and multi-dimensional defense collaboration are achieved.
[0060] It should be further explained that during implementation, the network traffic data collection module continuously captures raw traffic packets and transmits them to the protocol semantic deep deconstruction engine. This engine first dynamically parses the packet structure using an incremental protocol syntax tree. When unknown protocol features are identified, it automatically expands the syntax tree branches. It infers the hierarchical structure based on the relationship between field length and position. Subtree nodes are created for recurring sequences of similar fields and marked as variable-length field groups. Encrypted traffic skips content parsing but records metadata temporal patterns. The lightweight domain knowledge graph unit performs semantic mapping simultaneously, directly matching basic fields to predefined nodes and splitting composite fields into multi-node associations. If field value mapping conflicts occur, such as when an address is marked as both a server and a client, the labels are overwritten based on temporal context priority. For example, subsequent traffic in which the address appears as the request initiator is identified as the client. The temporal attention analysis unit screens high-value fields, such as frequently occurring authentication identifiers. It detects predefined combination patterns, such as generating an unauthorized access semantic label after multiple failed logins. Unpredefined anomalous combinations, such as overly long domain name queries, are marked as pending verification patterns and trigger a knowledge base update.
[0061] After receiving the decomposition results, the context-aware dynamic feature generator initiates dual-path processing. The structured path merges semantic nodes into compact labels based on knowledge graph relevance, such as "high-risk domestic mobile access." Weights are dynamically calculated based on node relevance strength and contextual confidence. The unstructured path processes the raw byte stream using a deformable convolutional semantic encoder. The first layer uses fixed convolution kernels to scan protocol header flags. The second layer dynamically adjusts the kernel shape based on syntax tree field boundaries, for example, expanding the kernel width to accommodate overly long address fields and switching to sparse sampling for encrypted fields to extract metadata. The third layer uses spatiotemporal convolution kernels to correlate adjacent packets, aligning features along the time axis for request-response traffic and aggregating statistical features over a fixed window for streaming traffic. The dual-path outputs are dynamically weighted by a gated fusion unit. When the confidence of the structured label is significantly higher than the entropy of the unstructured feature, the weight of the structured label is increased, while the weight of the unstructured feature is increased for encrypted or unknown protocols. If the dual-path outputs conflict, such as when the structured tag is normal and the unstructured detects anomalous byte patterns, the conflicting segment is isolated and a closed-loop verification module is requested to perform a partial reparse.
[0062] The closed-loop verification and optimization module performs a double check on feature vectors. A semantically driven generative adversarial network screen first converts structured labels into simulated fields based on the protocol syntax tree rules. Unstructured features are reconstructed into byte streams through deconvolution, and default values are filled in when mandatory fields are missing. The discriminator performs three levels of validation: the syntax layer checks field position compliance and size constraints; the semantic layer matches knowledge graph nodes and verifies logical dependency chains, such as the requirement for an authentication response after a login request; and the context layer uses temporal attention units to compare session continuity with real traffic. Failures at any level generate a reconstruction error report. Syntax errors trigger re-parsing of the original traffic, and semantic conflicts trigger a review of knowledge base rules. The feature utility evaluation screen injects perturbations in parallel, processing them hierarchically based on label weight: high-weight labels are randomly blocked individually, medium-weight labels are blocked in groups, low-weight labels are masked in batches, and adaptive noise is added to unstructured features. The perturbed vector is input into a lightweight proxy model, which calculates the standard deviation of multiple outputs. If the classification results fluctuate significantly between normal and suspicious, the discrimination is considered insufficient. When only a single perturbation fails, the knowledge graph nodes are marked as weakly robust, the node connection strength is adjusted when the association group fails, and the unstructured paths are strengthened when batch masking causes the model to collapse.
[0063] When detecting an unrecognized protocol, the self-evolving knowledge base initiates offline clustering analysis. Traffic segments are intercepted and grouped by packet length distribution, byte entropy, and time interval patterns. Strong clusters extract common features as candidates for new semantic nodes. Candidate nodes undergo two verification steps. Semantic reconstruction verification adds them to a temporary knowledge graph and simulates traffic for protocol logical consistency. Model effectiveness verification is then converted into label testing to improve classification accuracy and controllable fluctuations. After passing verification, strong cluster nodes are written into the protocol syntax tree. Nodes awaiting review are marked with an observation period and tested with historical traffic backtesting. For encrypted traffic, certificate chain features, cipher suite types, and session behavior parameters in the handshake sequence are converted into a time-series state chain. Similar patterns are matched against a historical attack database. Exact matches are associated with high-risk encrypted session labels. Partial matches generate temporary risk labels. No matches but statistical anomalies trigger real-time behavioral analysis. Temporary labels are subject to subsequent traffic monitoring. If no malicious activity is detected within 24 hours, the confidence level is downgraded. After three cumulative triggers of the new pattern, they are upgraded to candidate nodes.
[0064] The optimized feature vector is input into the pre-trained model of the analysis engine to perform multi-task analysis. The intrusion detection task combines high-risk semantic labels with unstructured abnormal byte patterns. When the two coexist, a high-risk alarm is generated. The first abnormality of a single indicator is marked as a low-confidence alarm, and three consecutive occurrences are upgraded to an actionable alarm. The protocol classification task limits the scope with the device type label, matches the spatiotemporal features to the protocol fingerprint library, isolates the traffic and requests deep packet inspection arbitration when the new protocol label conflicts with the historical rules. Dynamic enhancement is initiated when analyzing anomalies. When there are continuous false alarms, feature vectors are extracted to request closed-loop verification and review. Reconstruction is triggered when features are distorted. When deviations accumulate in specific scenarios, the last layer of classifiers is decoupled for incremental training. Cross-task knowledge transfer extracts spatiotemporal features when identifying new attacks and converts them into auxiliary labels for protocol classification. When the protocol classification finds unknown ports, behavioral features are captured and injected into the intrusion detection model. The migration process must be double-verified and confirmed, and a lightweight proxy model stress test is performed after addition.
[0065] A network traffic analysis and management method based on artificial intelligence includes the following steps: Step S1: The network traffic data acquisition module captures the original data packets and transmits them to the protocol semantic deep deconstruction engine.
[0066] Step S2: The incremental protocol syntax tree parses the known protocol structure, automatically expands branches when detecting unknown protocols, infers the hierarchy based on the relationship between field length and position, creates variable-length subtrees for repeated field sequences, skips content parsing for encrypted traffic, and records metadata timing patterns.
[0067] Step S3: The knowledge graph maps field values, basic fields match predefined nodes, composite fields are split into multi-node associations, and labels are overwritten with temporal context when fields conflict; the temporal attention unit captures field combination patterns, preset combinations generate semantic labels, and abnormal combinations trigger knowledge base updates.
[0068] Step S4: For structured paths, labels are merged according to node association, and weights are calculated based on association strength and confidence. For unstructured paths, the deformable convolution kernel adapts to the syntax tree field boundaries: the kernel width is expanded to cover variable-length fields, the encrypted segments switch to sparse sampling, and spatiotemporal convolution associates adjacent message features. Gated fusion is then performed, and dual-path weights are dynamically allocated according to protocol type. In case of conflict, segments are isolated and verification arbitration is requested.
[0069] Step S5: Perform double closed-loop verification, including semantic reconstruction verification and three-level discriminator verification. During the semantic reconstruction verification, the generator restores simulated traffic according to the syntax tree rules. During the three-level discriminator verification, the syntax layer checks field compliance, the semantic layer verifies the logical dependency chain, and the context layer compares session continuity. In case of failure, re-parsing or knowledge review is triggered in a targeted manner. Then, the feature utility is evaluated. On the one hand, the perturbations are graded according to the label weights, with high-weight random single points blocked, medium-weight associated groups blocked, and low-weight batches blocked. On the other hand, if the output fluctuation of the proxy model exceeds the threshold, the weakly robust nodes are marked and the unstructured paths are strengthened.
[0070] Step S6: Knowledge self-evolution: Unidentified protocols are clustered to generate candidate nodes; Double verification access: semantic reconstruction ensures logical consistency of the protocol, and model utility verification improves classification accuracy; Conversion of encrypted traffic time series chains: Match historical attack patterns to associated risk labels. Initiate behavioral analysis when there is no match but statistical anomalies occur. Temporary labels become official after traffic backtracking verification.
[0071] Step S7: Intelligent analysis and decision making: Intrusion Detection Combining Semantic Tags and Byte Patterns: The coexistence of two indicators generates a high-risk alarm, a single anomaly triggers a confidence attenuation mechanism, and deep packet inspection arbitrates when there is a conflict in protocol classification; When the model deviation exceeds the threshold, the local features are reviewed and the last layer is incrementally trained.
[0072] Step S8: Cross-task knowledge transfer: New attack features are converted into protocol classification labels, and unknown port behavior features are injected into the intrusion detection model. The migration process is verified by the agent model stress test.
[0073] Step S9: Result output and response: The AI analysis engine outputs an alarm or classification result and performs blocking or traffic control.
[0074] Step S10: Abnormal feedback loop: Continuous analysis of result deviations triggers dynamic optimization of steps S5 / S7.
[0075] Through the dynamic syntax tree expansion and knowledge graph conflict resolution mechanism of the protocol semantic deep deconstruction engine, intelligent parsing of unknown protocols and encrypted traffic is achieved, avoiding manually predefined rules; combined with the dynamic adaptation of field boundaries of deformable convolution kernels and gated fusion weight distribution, feature vectors with both semantic accuracy and pattern integrity are generated, solving the bottleneck of traditional feature engineering relying on expert experience; the dual closed-loop verification mechanism ensures adaptive optimization of feature quality through three-level semantic reconstruction inspection and weighted graded perturbation evaluation, thereby improving the generalization ability and detection accuracy of artificial intelligence models in complex network environments.
[0076] The self-evolving knowledge base, based on cluster analysis and a dual-verification access mechanism, automatically digests new protocols and attack patterns and generates risk labels, enabling the system to continuously adapt to changes in the network environment. The intelligent analysis engine's multi-task collaboration mechanism, combined with a confidence decay strategy and cross-task knowledge transfer, reduces false alarm rates while building a closed-loop defense enhancement system. Dynamic model incremental training and local feature review mechanisms ensure self-optimization while ensuring the continuity of the system's online services. This complete closed-loop process of "protocol analysis - feature generation - verification and evolution - analysis and decision-making" reduces operational and maintenance costs while improving network traffic analysis efficiency and security management.
[0077] It should be noted that, in this document, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or device. In the absence of further limitations, an element defined by the phrase "comprising a..." does not exclude the presence of additional identical elements in the process, method, article, or device comprising the element.
[0078] While embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions, and variations may be made to these embodiments without departing from the principles and spirit of the invention, and that the scope of the invention is defined by the appended claims and their equivalents.
Claims
1. An artificial intelligence-based network traffic analysis and management system, comprising a network traffic data acquisition module, a protocol semantic deep deconstruction engine, a context-aware dynamic feature generator, a closed-loop verification and optimization module, a self-evolving knowledge base, and an AI analysis engine, characterized by: The protocol semantics deep deconstruction engine is connected to the network traffic data acquisition module to analyze the protocol syntax structure, field semantics and context association of the traffic data; The context-aware dynamic feature generator is connected to the protocol semantics deep deconstruction engine to generate a feature vector of fused semantics; The closed-loop verification and optimization module bidirectionally connects the context-aware dynamic feature generator and the self-evolving knowledge base to perform double verification and iterative optimization on the feature vector; The AI analysis engine is connected to the closed-loop verification optimization module and uses the optimized feature vector to perform traffic analysis tasks.
2. The artificial intelligence-based network traffic analysis and management system according to claim 1, characterized in that: The protocol semantic deep deconstruction engine includes an incremental protocol syntax tree unit, a lightweight domain knowledge graph unit, and a temporal attention analysis unit; The incremental protocol syntax tree unit dynamically constructs a syntax tree of an unknown protocol; The lightweight domain knowledge graph unit maps field values to predefined semantic nodes and identifies logical dependencies between nodes; The temporal attention analysis unit captures dynamic combination patterns of multiple field values within the same data stream.
3. The artificial intelligence-based network traffic analysis and management system according to claim 2, characterized in that: The context-aware dynamic feature generator includes a structured feature generation path, an unstructured feature generation path and a gated fusion unit; The structured feature generation path converts the semantic deconstruction result into a semantic vector with a weight label; The unstructured feature generation path processes the original message byte stream through a deformable convolutional semantic encoder; The gated fusion unit dynamically weights and fuses the dual-path outputs.
4. The artificial intelligence-based network traffic analysis and management system according to claim 3, characterized in that: The deformable convolutional semantic encoder includes three layers of convolution kernels: The first layer of fixed convolution kernel extracts the fixed pattern of protocol header; The second layer of deformable convolution kernel dynamically adjusts the kernel shape according to the grammatical structure output by the incremental protocol syntax tree; The third layer of spatiotemporal convolution kernel integrates cross-message temporal features.
5. The network traffic analysis and management system based on artificial intelligence according to claim 1, characterized in that: The closed-loop verification optimization module includes a semantically driven adversarial generative network screen and a feature utility evaluation screen; The semantic-driven adversarial generative network screen generates simulated traffic by reconstructing feature vectors, determines semantic consistency based on the protocol syntax tree and the knowledge graph, and screens out features with reconstruction errors exceeding a threshold. The feature utility evaluation screen applies semantic label mask perturbations to feature vectors and determines feature discrimination through the stability of lightweight proxy model output.
6. The artificial intelligence-based network traffic analysis and management system according to claim 5, characterized in that: The semantically driven adversarial generative network sieve includes a generator and a discriminator; The generator deconstructs the feature vector into simulated traffic data; The discriminator calls the semantic rules in the lightweight domain knowledge graph unit to perform consistency verification.
7. The network traffic analysis and management system based on artificial intelligence according to claim 5, characterized in that: The Feature Utility Evaluation Screen performs the following operations: Randomly mask the semantic labels of the feature vector; Input the masked vector to the lightweight proxy model and calculate the output standard deviation; If the standard deviation exceeds the preset tolerance threshold, feature regeneration is triggered.
8. The network traffic analysis and management system based on artificial intelligence according to claim 1, characterized in that: The self-evolving knowledge base is updated in the following ways: When the incremental protocol syntax tree unit detects an unrecognized protocol pattern, it initiates offline clustering analysis to generate new semantic nodes; After the new semantic nodes are double-verified by the semantic-driven adversarial generative network screen and the feature utility evaluation screen, the knowledge graph nodes and association rules are automatically expanded to lightweight domain knowledge graph units; For encrypted traffic metadata, perform protocol behavior timing chain transformation and risk mapping processing.
9. The artificial intelligence-based network traffic analysis and management system according to claim 8, characterized in that: The method for processing encrypted traffic metadata is as follows: convert the certificate chain length, cipher suite selection and session resumption frequency in the TLS handshake sequence into a protocol behavior timing chain; map the timing chain to risk semantic labels through lightweight domain knowledge graph units.
10. The network traffic analysis and management system based on artificial intelligence according to claim 1, characterized in that: The AI analysis engine includes a pre-trained deep learning model, the input of which is the feature vector output by the closed-loop verification optimization module.
Citation Information
Cited By
Network flow restoring and monitoring method
CN120825342A