Dynamic network flow control method based on large model training

By deploying data processing servers and probes on the intranet, and combining K-Means clustering and silhouette coefficient methods, abnormal traffic is dynamically identified and intercepted. This solves the limitations of static defense strategies and the efficiency bottleneck of manual annotation, enabling flexible and real-time intranet traffic monitoring, and reducing attack success rate and detection latency.

CN120979713APending Publication Date: 2025-11-18GUANGXI PUBLIC INFORMATION IND CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202511098288.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-06
Publication Date
2025-11-18

AI Technical Summary

Technical Problem

In existing technologies, static defense strategies cannot dynamically adapt to changes in internal network traffic, allowing attackers to bypass external defenses through lateral movement. Static firewalls and ACL policies are difficult to detect abnormal internal traffic. Manually labeling datasets is time-consuming and costly. The complexity of system integration leads to a high success rate of social engineering attacks. Traditional detection methods have long latency and a high false negative rate.

Method used

A dynamic network traffic control method based on large model training is adopted. By deploying data processing servers and data probes in the intranet, traffic data is captured and parsed in real time. K-Means clustering and average profile coefficient method are used to identify outliers, and firewall policies are dynamically adjusted to intercept abnormal traffic to avoid affecting normal business.

Benefits of technology

It achieves dynamic and controllable traffic control, blocking only malicious traffic without affecting normal business operations, enhancing model sensitivity and adaptive defense capabilities, and reducing detection latency and false negative rate.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120979713A_ABST
    Figure CN120979713A_ABST
Patent Text Reader

Abstract

The invention discloses a dynamic network flow control method based on large model training, which comprises the following steps of: capturing and processing intranet flow data through a data probe and a data processing server, storing the data into a flow database through a data server according to a data capturing result, forming a data set, and analyzing the characteristics of the flow through a data mining model; the method comprises the following steps: training a data processing server, calculating whether the traffic is an outlier based on an average contour coefficient method according to the trained model and the captured traffic, and finally judging whether the traffic is the outlier and whether the traffic is intercepted according to the judgment of the data processing server, thereby ensuring that the conventional traffic is normally passed and the unconventional traffic is intercepted, and preventing the server from being damaged. The method has the advantages of dynamic controllability, no influence on normal business and capability of only intercepting malicious traffic; compared with existing hard configuration through access control, the method has the advantage of being more flexible.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of network security technology, and specifically relates to a dynamic network traffic control method based on large model training. Background Technology

[0002] With the proliferation of digital assets, industrial automation systems, and IoT devices, enterprise intranet traffic has become a crucial component of cybersecurity. Intranet traffic encompasses all communication data between internal enterprise networks, including access sources, target IPs, access times, and request frequencies. Analyzing this data allows for real-time monitoring of network security posture and timely detection of abnormal behavior.

[0003] However, once attackers successfully breach the internal network and gain control of a "zombie" (a device on the internal network), they often choose to further disrupt enterprise operations through lateral movement. Since many enterprise internal systems are deployed on the internal network and generate normal internal network traffic during operation, blocking access to certain devices via ACLs would affect normal operations. Therefore, enterprise internal networks generally do not have access control policies. Consequently, attacking legitimate devices and gaining control of "zombies" through social engineering techniques such as phishing emails and sending text messages from fake base stations has become a common attack method with a high success rate.

[0004] The main problems with the existing technology are as follows:

[0005] 1. Limitations of Static Defense: Traditional firewalls and ACLs rely on fixed rules and cannot dynamically adapt to changes in internal network traffic. After attackers bypass external defenses through lateral movement, static policies struggle to detect abnormal internal traffic. According to a 2023 Gartner report, 70% of internal network attacks are carried out through lateral movement, with traditional detection methods averaging over 48 hours of latency. Internal network devices often lack strict access controls to avoid impacting normal business operations, leading to a higher success rate for social engineering attacks (such as phishing emails).

[0006] 2. Dependence on Manual Labor and Efficiency Bottlenecks: Existing machine learning systems, such as patent CN119583359A, require manual annotation of datasets, which is time-consuming and costly (annotating 1TB of traffic data requires 200 hours with an error rate of 15%). Manually generated rules are insufficient to cover new types of attacks such as zero-day vulnerabilities, resulting in a false negative rate exceeding 30%.

[0007] 3. System integration complexity: For example, the large-scale model solution proposed in patent CN118316736A requires API integration with multiple types of security devices (such as firewalls and IDS), with a deployment cycle of 3-6 months and high maintenance costs.

[0008] Lateral penetration refers to the process by which an attacker, after successfully compromising a system or device within a network, uses that system's privileges to further penetrate other systems or devices within the network. This typically occurs after the attacker has breached the network's outer defenses (such as firewalls and intrusion detection systems) but has not yet gained complete control of the entire network.

[0009] K-Means clustering algorithm: A classic unsupervised learning algorithm used to divide a dataset into K clusters, such that data points within each cluster are as similar as possible, while data points between different clusters are as different as possible. The goal of the K-Means algorithm is to discover structure in the data by minimizing the squared error within each cluster. It has the advantages of being simple, easy to understand and implement, suitable for large-scale datasets, and having a relatively fast convergence speed. Summary of the Invention

[0010] To address the problems of existing technologies, this invention provides a dynamic network traffic control method based on large model training, which has the advantages of being dynamically controllable, not affecting normal business operations, only blocking malicious traffic, and being flexible.

[0011] The technical solution of the present invention is as follows:

[0012] A dynamic network flow control method based on large model training includes the following steps:

[0013] S1. Deploy a data processing server, data probe, and traffic database in the intranet; the data probe is located in the intranet device and is equipped with a protocol parsing engine based on deep packet inspection technology;

[0014] S2. The data probe captures intranet traffic data in real time, deeply analyzes network packets, and automatically extracts rich traffic semantic information. Simultaneously, the data probe also acquires basic data, including source IP, source port, destination IP, destination port, protocol type, connection time, and packet size. During data capture, the data probe automatically performs preliminary cleaning, classification, and labeling of the data. Data cleaning removes useless or missing data, ensuring the quality of data uploaded to the data processing server. The entire process requires no manual intervention. Furthermore, the data probe performs preliminary classification and labeling of the captured data to enable the data processing server to perform subsequent processing more efficiently.

[0015] After receiving the data transmitted by the data probe, the data processing server first standardizes the data to unify the data format, which facilitates subsequent analysis and modeling. Then, the data processing server automatically constructs behavioral feature vectors using time series analysis technology.

[0016] S3. The data processing server stores the captured data into the traffic database and forms a dataset. Then, it analyzes the characteristics of the traffic through a data mining model.

[0017] S4. When the data probe detects abnormal traffic, it uploads it to the data processing server. The data processing server calculates whether it is an outlier based on the trained model and the current traffic flow using the average profile coefficient method, as follows:

[0018]

[0019] For an uploaded traffic data point x, calculate the average distance within the cluster a(x) and the average distance to the nearest cluster b(x), and then calculate the silhouette coefficient s(x). Points with a silhouette coefficient close to -1 can be considered outliers.

[0020] S5. If the data processing server determines that the traffic is an outlier, it sends a command to the data probe. The data probe then adjusts its local firewall configuration to block the traffic and prevent it from accessing the outlier. If the traffic is not an outlier, it simply enters the data into the database as learning data and does not block the traffic. Through these methods, regular traffic is ensured to pass normally, while non-regular traffic is blocked to prevent damage to the server.

[0021] To further illustrate the present invention, step S3 is performed according to the following steps:

[0022] S31. The data processing server preprocesses the captured data, which includes the following:

[0023] Source IP: The IP address from which the connection was initiated, used to confirm the source of the access;

[0024] Source port: The port from which the connection was initiated, used to confirm the source of the access;

[0025] Destination IP: Connect to the destination IP address to confirm the service being accessed;

[0026] Destination Port: Connect to the destination port to confirm the service being accessed;

[0027] Protocol type: The protocol type of the connection, used to help determine the service being accessed;

[0028] Connection time: The time it takes to establish a connection, used to determine whether the connection period is within a normal timeframe;

[0029] Packet size: The size of packets transmitted during the connection, used to help determine whether the connection is normal;

[0030] The data processing server cleans, transforms, and normalizes the data, removes useless or missing data, handles missing values, and then stores the processed data into the traffic database.

[0031] S32. Normalize the data. Due to the large range of packet sizes, data normalization is necessary in big data analysis to avoid gradient explosion or gradient vanishing. Data normalization is performed using the following formula:

[0032]

[0033] Where: x is the original data point, min(x) and max(x) are the minimum and maximum values ​​of the data respectively, and x′ is the normalized data; after normalization, all values ​​of the packet size are converted to the range [0,1] and stored in the database;

[0034] S33. Select features for K-Means clustering. Different features are selected for K-Means clustering depending on the scenario. The features include:

[0035] Source IP: For servers that provide services to the outside world, data that is accessed between internal servers will be removed before clustering to ensure that only connections initiated from outside are filtered.

[0036] Source Port: If it is a server providing services to the outside world, this item is excluded because the source port is random; if it is an internal communication case, this item will be selected as a feature because the communication port between servers is fixed.

[0037] Destination IP: The server used to confirm access;

[0038] Destination port: The service port used to confirm the access, which can determine the specific system being accessed;

[0039] Protocol type: Used to filter abnormal protocol access;

[0040] Connection time: Used to filter access during abnormal periods;

[0041] Packet size: Select this as a feature for servers that do not upload files to determine whether they have been subjected to malicious file uploads; for servers that do upload files, this item can also be selected as a feature to filter abnormal data packets.

[0042] S34. Determine the K value based on the elbow coefficient of the dataset; determine the appropriate K value by calculating the sum of squared errors of the clustering results under different K values;

[0043] S35. Enter the obtained K value and perform cluster analysis using the K-Means algorithm.

[0044] To further illustrate the present invention, the step of determining a suitable K value in step S34 is as follows:

[0045] S341. For different K values, run the K-Means algorithm and calculate the total sum of squared errors for each K value;

[0046] S342. Plot the K value against the corresponding total sum of squared errors to obtain a curve.

[0047] S343. Observe the trend of the total sum of squared errors as K changes in the graph. As the value of K increases, the total sum of squared errors gradually decreases.

[0048] S344. Find the "elbow" point where the rate of decrease of the total sum of squared errors slows down. This point usually corresponds to the optimal K value for the number of clusters, because after this point, increasing the K value has little effect on improving SSE.

[0049] To further illustrate the present invention, step S35 uses the K-Means algorithm for cluster analysis, and the steps are as follows:

[0050] S351. Initialize K cluster centers, either by randomly selecting them or by using the K-Means++ initialization method, to avoid getting trapped in local optima;

[0051] S352. Assign each data point, i.e., each network traffic sample, to the nearest cluster center;

[0052] S353, Update Cluster Center: Update the cluster center based on the mean of all data points within the cluster;

[0053] S354. Repeat the above two steps until the cluster center no longer changes or the maximum number of iterations is reached;

[0054] The goal of the K-Means algorithm is to minimize the sum of squared Euclidean distances between each data point in a cluster and the cluster center, expressed by the following objective function:

[0055]

[0056] Where: K is the number of clusters; Ci is the i-th cluster; μi is the center of the i-th cluster; xj is the data point; |xj-μi| is the distance between the data point and the cluster center.

[0057] The beneficial effects of this invention are:

[0058] This method captures and analyzes intranet traffic that is significantly abnormal from normal traffic. If the traffic is confirmed to be suspicious, it will be blocked. This method has the advantages of being dynamically controllable, not affecting normal business operations, and only blocking malicious traffic. Compared with the existing hard configuration of access control, it has the advantage of being more flexible. At the same time, this method enhances the sensitivity of the model, realizes the "detection-blocking-model update" closed loop, and outlier data is fed back to the training set in real time, forming an adaptive defense capability. Attached Figure Description

[0059] Figure 1 This is a flowchart illustrating the method of the present invention. Detailed Implementation

[0060] The present invention will be further described below with reference to specific embodiments.

[0061] Example:

[0062] A dynamic network flow control method based on large model training includes the following steps:

[0063] S1. Deploy a data processing server, data probe, and traffic database in the intranet; the data probe is located in the intranet device and is equipped with a protocol parsing engine based on deep packet inspection technology;

[0064] S2. The data probe captures intranet traffic data in real time, deeply analyzes network packets, and automatically extracts rich traffic semantic information, such as HTTP methods and DNS query types. Simultaneously, the data probe also acquires basic data, including source IP, source port, destination IP, destination port, protocol type, connection time, and packet size. During data capture, the data probe automatically performs preliminary cleaning, classification, and labeling of the data. Cleaning removes useless or missing data, ensuring the quality of data uploaded to the data processing server. The entire process requires no manual intervention. Furthermore, the data probe performs preliminary classification and labeling of the captured data to enable the data processing server to perform subsequent processing more efficiently.

[0065] After receiving the data transmitted by the data probe, the data processing server first standardizes the data to unify the data format, which facilitates subsequent analysis and modeling. Then, the data processing server uses time series analysis technology to automatically construct behavioral feature vectors, such as constructing access frequency time series patterns based on connection time, mining traffic behavior features from multiple dimensions, and further extracting key data to provide strong support for subsequent traffic analysis based on data mining models.

[0066] S3. The data processing server stores the captured data into the traffic database and forms a dataset. Then, it analyzes the characteristics of the traffic through a data mining model.

[0067] S4. When the data probe detects abnormal traffic, it uploads it to the data processing server. The data processing server calculates whether it is an outlier based on the trained model and the current traffic flow using the average profile coefficient method, as follows:

[0068]

[0069] For an uploaded traffic data point x, calculate the average distance within the cluster a(x) and the average distance to the nearest cluster b(x), and then calculate the silhouette coefficient s(x). Points with a silhouette coefficient close to -1 can be considered outliers.

[0070] S5. If the data processing server determines that the traffic is an outlier, it sends a command to the data probe. The data probe then adjusts its local firewall configuration to block the traffic and prevent it from accessing the outlier. If the traffic is not an outlier, it simply enters the data into the database as learning data and does not block the traffic. Through these methods, regular traffic is ensured to pass normally, while non-regular traffic is blocked to prevent damage to the server.

[0071] To further illustrate the present invention, step S3 is performed according to the following steps:

[0072] S31. The data processing server preprocesses the captured data, which includes the following:

[0073] Source IP: The IP address from which the connection was initiated, used to confirm the source of the access;

[0074] Source port: The port from which the connection was initiated, used to confirm the source of the access;

[0075] Destination IP: Connect to the destination IP address to confirm the service being accessed;

[0076] Destination Port: Connect to the destination port to confirm the service being accessed;

[0077] Protocol type: The protocol type of the connection, used to help determine the service being accessed;

[0078] Connection time: The time it takes to establish a connection, used to determine whether the connection period is within a normal timeframe;

[0079] Packet size: The size of packets transmitted during the connection, used to help determine whether the connection is normal;

[0080] The data processing server cleans, transforms, and normalizes the data, removes useless or missing data, handles missing values, and then stores the processed data into the traffic database.

[0081] S32. Normalize the data. Due to the large range of packet sizes, data normalization is necessary in big data analysis to avoid gradient explosion or gradient vanishing. Data normalization is performed using the following formula:

[0082]

[0083] Where: x is the original data point, min(x) and max(x) are the minimum and maximum values ​​of the data respectively, and x′ is the normalized data; after normalization, all values ​​of the packet size are converted to the range [0,1] and stored in the database;

[0084] S33. Select features for K-Means clustering. Different features are selected for K-Means clustering depending on the scenario. The features include:

[0085] Source IP: For servers that provide services to the outside world, data that is accessed between internal servers will be removed before clustering to ensure that only connections initiated from outside are filtered.

[0086] Source Port: If it is a server providing services to the outside world, this item is excluded because the source port is random; if it is an internal communication case, this item will be selected as a feature because the communication port between servers is fixed.

[0087] Destination IP: The server used to confirm access;

[0088] Destination port: The service port used to confirm the access, which can determine the specific system being accessed;

[0089] Protocol type: Used to filter abnormal protocol access;

[0090] Connection time: Used to filter access during abnormal periods;

[0091] Packet size: Select this as a feature for servers that do not upload files to determine whether they have been subjected to malicious file uploads; for servers that do upload files, this item can also be selected as a feature to filter abnormal data packets.

[0092] S34. Determine the K value based on the elbow coefficient of the dataset; determine the appropriate K value by calculating the sum of squared errors of the clustering results under different K values;

[0093] S35. Enter the obtained K value and perform cluster analysis using the K-Means algorithm.

[0094] To further illustrate the present invention, the step of determining a suitable K value in step S34 is as follows:

[0095] S341. For different K values, run the K-Means algorithm and calculate the total sum of squared errors for each K value;

[0096] S342. Plot the K value against the corresponding total sum of squared errors to obtain a curve.

[0097] S343. Observe the trend of the total sum of squared errors as K changes in the graph. As the value of K increases, the total sum of squared errors gradually decreases.

[0098] S344. Find the "elbow" point where the rate of decrease of the total sum of squared errors slows down. This point usually corresponds to the optimal K value for the number of clusters, because after this point, increasing the K value has little effect on improving SSE.

[0099] To further illustrate the present invention, step S35 uses the K-Means algorithm for cluster analysis, and the steps are as follows:

[0100] S351. Initialize K cluster centers, either by randomly selecting them or by using the K-Means++ initialization method, to avoid getting trapped in local optima;

[0101] S352. Assign each data point, i.e., each network traffic sample, to the nearest cluster center;

[0102] S353, Update Cluster Center: Update the cluster center based on the mean of all data points within the cluster;

[0103] S354. Repeat the above two steps until the cluster center no longer changes or the maximum number of iterations is reached;

[0104] The goal of the K-Means algorithm is to minimize the sum of squared Euclidean distances between each data point in a cluster and the cluster center, expressed by the following objective function:

[0105]

[0106] Where: K is the number of clusters; Ci is the i-th cluster; μi is the center of the i-th cluster; xj is the data point; |xj-μi| is the distance between the data point and the cluster center.

Claims

1. A dynamic network flow control method based on large model training, characterized in that: Includes the following steps: S1. Deploy a data processing server, a data probe, and a traffic database in the intranet; the data probe is located in the intranet device and is equipped with a protocol parsing engine based on deep packet inspection technology; S2. The data probe captures intranet traffic data in real time, deeply analyzes network packets, and automatically extracts rich traffic semantic information. At the same time, the data probe also obtains basic data, including source IP, source port, destination IP, destination port, protocol type, connection time, and packet size. During the data capture process, the data probe automatically performs preliminary cleaning, classification, and labeling of the data. After receiving the data transmitted by the data probe, the data processing server first standardizes the data to unify the data format, which facilitates subsequent analysis and modeling. Then, the data processing server automatically constructs behavioral feature vectors using time series analysis technology. S3. The data processing server stores the captured data into the traffic database and forms a dataset. Then, it analyzes the characteristics of the traffic through a data mining model. S4. When the data probe detects abnormal traffic, it uploads it to the data processing server. The data processing server calculates whether it is an outlier based on the trained model and the current traffic flow using the average profile coefficient method, as follows: For an uploaded traffic data point x, calculate the average distance within the cluster a(x) and the average distance to the nearest cluster b(x), and then calculate the silhouette coefficient s(x). Points with a silhouette coefficient close to -1 can be considered outliers. S5. If the data processing server determines that it is an outlier, it sends an instruction to the data probe. The data probe then adjusts its local firewall configuration to block the traffic and prevent it from accessing the network further. If it is not an outlier, the data will only be entered into the database as learning data, and the traffic will not be blocked.

2. The dynamic network traffic control method based on large model training according to claim 1, characterized in that: Step S3 is performed according to the following steps: S31. The data processing server preprocesses the captured data, which includes the following: Source IP: The IP address from which the connection was initiated, used to confirm the source of the access; Source port: The port from which the connection was initiated, used to confirm the source of the access; Destination IP: Connect to the destination IP address to confirm the service being accessed; Destination Port: Connect to the destination port to confirm the service being accessed; Protocol type: The protocol type of the connection, used to help determine the service being accessed; Connection time: The time it takes to establish a connection, used to determine whether the connection period is within a normal timeframe; Packet size: The size of packets transmitted during the connection, used to help determine whether the connection is normal; The data processing server cleans, transforms, and normalizes the data, removes useless or missing data, handles missing values, and then stores the processed data into the traffic database. S32. Normalize the data. Due to the large range of packet sizes, data normalization is necessary in big data analysis to avoid gradient explosion or gradient vanishing. Data normalization is performed using the following formula: Where: x is the original data point, min(x) and max(x) are the minimum and maximum values ​​of the data respectively, and x′ is the normalized data; after normalization, all values ​​of the packet size are converted to the range [0,1] and stored in the database; S33. Select features for K-Means clustering. Different features are selected for K-Means clustering depending on the scenario. The features include: Source IP: For servers that provide services to the outside world, data that is accessed between internal servers will be removed before clustering to ensure that only connections initiated from outside are filtered. Source Port: If it is a server providing services to the outside world, this item is excluded because the source port is random; if it is an internal communication case, this item will be selected as a feature because the communication port between servers is fixed. Destination IP: The server used to confirm access; Destination port: The service port used to confirm the access, which can determine the specific system being accessed; Protocol type: Used to filter abnormal protocol access; Connection time: Used to filter access during abnormal periods; Packet size: Select this as a feature for servers that do not upload files to determine whether they have been subjected to malicious file uploads; for servers that do upload files, this item can also be selected as a feature to filter abnormal data packets. S34. Determine the K value based on the elbow coefficient of the dataset; determine the appropriate K value by calculating the sum of squared errors of the clustering results under different K values; S35. Enter the obtained K value and perform cluster analysis using the K-Means algorithm.

3. The dynamic network traffic control method based on large model training according to claim 2, characterized in that: The steps for determining a suitable K value in step S34 are as follows: S341. For different K values, run the K-Means algorithm and calculate the total sum of squared errors for each K value; S342. Plot the K value against the corresponding total sum of squared errors to obtain a curve. S343. Observe the trend of the total sum of squared errors as K changes in the graph. As the value of K increases, the total sum of squared errors gradually decreases. S344. Find the "elbow" point where the rate of decrease of the total sum of squared errors slows down. This point usually corresponds to the optimal K value for the number of clusters, because after this point, increasing the K value has little effect on improving SSE.

4. The dynamic network traffic control method based on large model training according to claim 2, characterized in that: Step S35 uses the K-Means algorithm for cluster analysis, and the steps are as follows: S351. Initialize K cluster centers, either by randomly selecting them or by using the K-Means++ initialization method, to avoid getting trapped in local optima; S352. Assign each data point, i.e., each network traffic sample, to the nearest cluster center; S353, Update Cluster Center: Update the cluster center based on the mean of all data points within the cluster; S354. Repeat the above two steps until the cluster center no longer changes or the maximum number of iterations is reached; The goal of the K-Means algorithm is to minimize the sum of squared Euclidean distances between each data point in a cluster and the cluster center, expressed by the following objective function: Where: K is the number of clusters; C i It is the i-th cluster; μ i It is the center of the i-th cluster; x j These are data points; ||x j -μ i ||2 is the distance between the data point and the cluster center.

Citation Information

Patent Citations

  • Network traffic data analysis method based on traffic large model

    CN119583359A