Flow filtering method and device for extracting quintuple flow characteristics
By obtaining the five-tuple feature sequence of historical network communication data, constructing timing-related features and using large language models for training, the problem of low accuracy of abnormal traffic filtering in the existing technology is solved, and efficient abnormal identification and filtering of network communication behavior is achieved.
Patent Information
- Application Number
- CN202510506649.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-22
- Publication Date
- 2025-08-05
AI Technical Summary
The prior art is difficult to effectively capture the dynamic correlation of network communication behaviors in multiple time windows, resulting in low accuracy of abnormal traffic filtering. Especially when facing intelligent and concealed network attacks, traditional security protection mechanisms based on rule matching and feature detection are difficult to deal with complex and changeable network threats.
By obtaining the five-tuple feature sequence of historical network communication data samples, the joint probability distribution between the five-tuple features is extracted in multiple time windows, the timing correlation features are constructed, and the large language model is used for training to generate communication prompt words to realize abnormal analysis and filtering of real-time network communication data.
It improves the accuracy of abnormal traffic filtering, improves the context semantic understanding of communication behavior, and can effectively identify potential abnormal behaviors and perform precise filtering.
Smart Images

Figure CN120433972A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of traffic filtering, and specifically to a traffic filtering method and device for extracting quintuple traffic features. Background Art
[0002] With the widespread adoption and application of information networks, network security issues are becoming increasingly prominent, especially in key sectors such as government, finance, e-commerce, and industrial control. The integrity and security of network communication data are directly related to the stable operation of business systems and the protection of data assets. As network attack methods become more intelligent, covert, and automated, traditional security protection mechanisms based on rule matching and signature detection are no longer able to effectively address complex and ever-changing network threats. Therefore, improving the ability to understand network communication behavior and promptly identify potential abnormal traffic has become a key topic in network security research.
[0003] In the existing technology, network traffic is usually analyzed and modeled by extracting the five-tuple features in network communication - namely, source IP address, destination IP address, source port number, destination port number and transmission protocol - and some machine learning or deep learning algorithms are combined to achieve classification and identification of communication behavior.
[0004] However, these methods primarily focus on modeling static features, making it difficult to fully reflect the dynamic evolution of communication behavior over multiple time periods. In practice, some attack behaviors, such as distributed scanning, low-frequency probing, and lateral movement, exhibit significant temporal characteristics and contextual dependencies. Their communication behavior is often not noticeable over a short period of time. Consequently, existing analysis methods are often unable to effectively capture these behavioral correlations across time windows, impacting the accuracy of anomaly traffic filtering. Summary of the Invention
[0005] The present application provides a traffic filtering method and device for extracting quintuple traffic features, which are used to improve the accuracy of abnormal traffic filtering.
[0006] In a first aspect of the present application, a traffic filtering method for extracting quintuple traffic features is provided, which is applied to a server. The method includes: obtaining a quintuple feature sequence of historical network communication data samples, the quintuple feature sequence including a source IP address, a destination IP address, a source port number, a destination port number, and a transmission protocol; extracting a joint probability distribution between different quintuple features in the quintuple feature sequence within multiple time windows to obtain a time series correlation feature; constructing a communication prompt word based on the time series correlation feature; inputting the communication prompt word into a preset large language model for training to obtain a target large language model; performing an anomaly analysis on real-time network communication data using the target large language model to obtain an analysis result; determining an anomaly type based on the analysis result, determining a traffic filtering strategy based on the anomaly type and a preset strategy library, and performing traffic filtering using the traffic filtering strategy.
[0007] Optionally, within multiple time windows, before extracting the joint probability distribution between different quintuple features in the quintuple feature sequence and obtaining the time series correlation feature, the method includes: analyzing the correlation frequency between the source IP address and the destination IP address in the quintuple feature sequence to obtain the correlation distribution feature; counting the combination patterns between the target features in the quintuple feature sequence to obtain the combination distribution feature, the target features including the source port number, the destination port number, and the transmission protocol; and determining multiple time windows based on the correlation distribution feature and the combination distribution feature.
[0008] Optionally, multiple time windows are determined based on the associated distribution characteristics and the combined distribution characteristics, specifically including: determining multiple first time windows based on the associated distribution characteristics; determining multiple second time windows based on the combined distribution characteristics; calculating the overlap between the target first time window and the target second time window, where the target first time window is any first time window and the target second time window is any second time window; combining the target first time window and the target second time window whose overlap is greater than a preset threshold to obtain a time window pair; extracting the overlapping time period in each time window pair to obtain a time window.
[0009] Optionally, within multiple time windows, the joint probability distribution between different quintuple features in the quintuple feature sequence is extracted to obtain time series correlation features, specifically including: within each time window, counting the occurrence frequency of each quintuple feature and the co-occurrence frequency between different quintuple features according to the joint probability distribution, and constructing a quintuple feature co-occurrence matrix; determining the time series correlation features based on the quintuple feature co-occurrence matrix.
[0010] Optionally, the time series correlation features are determined based on the quintuple feature co-occurrence matrix, specifically including: constructing a probability transmission chain based on the quintuple feature co-occurrence matrix, calculating the correlation probability between different quintuple features; determining the joint probability distribution based on the correlation probability; and analyzing the changing trend of the joint probability distribution within multiple time windows to generate time series correlation features.
[0011] Optionally, communication prompt words are constructed based on the time series correlation features, specifically including: extracting the time series distribution characteristics of the joint probability distribution in the time series correlation features and the probability change characteristics of the probability transmission chain; identifying the time series dependency between different five-tuple features based on the time series distribution characteristics and the probability change characteristics; and constructing communication prompt words based on the time series dependency.
[0012] Optionally, the communication prompt words are input into a preset large language model for training to obtain a target large language model, specifically including: constructing a training sample set based on the communication prompt words; inputting the training sample set into the preset large language model to obtain training parameters; verifying the accuracy of the preset large language model in recognizing the temporal features of the five-tuple feature sequence based on the training parameters, and ending the training when the accuracy of the temporal feature recognition meets the preset requirements.
[0013] In a second aspect of the present application, a traffic filtering system for extracting quintuple traffic features is provided, comprising: an acquisition module for acquiring a quintuple feature sequence of historical network communication data samples; an extraction module for extracting a joint probability distribution between different quintuple features in the quintuple feature sequence within multiple time windows to obtain time series correlation features; a construction module for constructing communication prompt words based on the time series correlation features; a training module for inputting the communication prompt words into a preset large language model for training to obtain a target large language model; an analysis module for performing anomaly analysis on real-time network communication data using the target large language model to obtain analysis results; and a filtering module for determining anomaly types based on the analysis results, determining a traffic filtering strategy based on the anomaly type and a preset strategy library, and performing traffic filtering using the traffic filtering strategy.
[0014] In the third aspect of the present application, an electronic device is provided, including a processor, a memory, a user interface and a network interface, the memory is used to store instructions, the user interface and the network interface are both used to communicate with other devices, and the processor is used to execute the instructions stored in the memory so that the electronic device performs any of the methods described above.
[0015] In a fourth aspect of the present application, a computer-readable storage medium is provided, wherein the computer-readable storage medium stores instructions, and when the instructions are executed, any one of the above methods is executed.
[0016] In summary, one or more technical solutions provided by this application have at least the following technical effects or advantages: 1. This application introduces historical network communication data samples as a reference basis to obtain a five-tuple feature sequence to provide statistical support for understanding real-time communication behavior; extracts the joint probability distribution between the five-tuple features in multiple time windows, models the temporal correlation characteristics of communication behavior, and improves the ability to characterize the potential correlation between different communication modes; constructs communication prompt words based on the extracted temporal correlation characteristics, and inputs the communication prompt words into a preset large language model for training, so that the target large language model has the ability to express and understand the semantics of communication behavior; with the help of the target large language model, the real-time network communication data is analyzed for anomalies, which can realize the contextual semantic understanding and accurate identification of abnormal behavior, and then determine the anomaly type based on the analysis results, and match the corresponding traffic filtering strategy according to the preset strategy library, and finally realize the effective filtering of abnormal traffic. Therefore, this application improves the accuracy of abnormal traffic filtering through the collaboration of historical data modeling, temporal correlation mining, language model training and strategy matching.
[0017] 2. Before extracting the joint probability distribution between different quintuple features in the quintuple feature sequence to form the temporal correlation feature, the correlation frequency between the source IP address and the destination IP address in the quintuple feature sequence is analyzed to obtain the correlation distribution feature, and the combination pattern between the source port number, the destination port number and the transmission protocol type is counted to obtain the combined distribution feature, which can extract representative statistical features from dimensions such as the structural distribution and protocol combination of the communication behavior; further, multiple first time windows are determined according to the correlation distribution feature, and multiple second time windows are determined according to the combined distribution feature, and by calculating the overlap between any first time window and any second time window, time window pairs with an overlap greater than a preset threshold are screened out, and the overlapping time period is extracted as the final time window, while retaining the key communication behavior features, effectively avoiding the interference of redundant time periods, thereby ensuring that the subsequently extracted temporal correlation features have higher timeliness and representativeness, and improving the accuracy and context stability of abnormal behavior modeling.
[0018] 3. Within multiple time windows, the occurrence frequency of each quintuple feature and the co-occurrence frequency between different quintuple features are counted according to the joint probability distribution to construct a quintuple feature co-occurrence matrix, which can effectively reflect the synchronization relationship between various features in network communication behavior; on this basis, a probability transmission chain is constructed based on the quintuple feature co-occurrence matrix, and the association probability between quintuple features is calculated to form a joint probability distribution; by analyzing the changing trend of the joint probability distribution within multiple time windows, a time series correlation feature is generated that can reflect the evolution law of feature association over time, thereby improving the modeling ability of dynamic communication behavior patterns and providing more accurate and predictive input features for subsequent anomaly analysis. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] Figure 1 This is a system architecture diagram involved in a traffic filtering method for extracting quintuple traffic features or a traffic filtering system for extracting quintuple traffic features in an embodiment of the present application; Figure 2 This is a flow chart of a traffic filtering method for extracting quintuple traffic features in an embodiment of the present application; Figure 3 This is another flow chart of a traffic filtering method for extracting quintuple traffic features in an embodiment of the present application; Figure 4 Schematic diagram of the structure of a traffic filtering system for extracting quintuple traffic features in an embodiment of the present application; Figure 5 It is a structural diagram of an electronic device in an embodiment of the present application.
[0020] Explanation of the accompanying drawings: 401, acquisition module; 402, extraction module; 403, construction module; 404, training module; 405, analysis module; 406, filtering module; 407, determination module; 501, processor; 502, communication bus; 503, user interface; 504, network interface; 505, memory. DETAILED DESCRIPTION
[0021] In order to enable those skilled in the art to better understand the technical solutions in this specification, the technical solutions in the embodiments of this specification will be clearly and completely described below in conjunction with the drawings in the embodiments of this specification. Obviously, the described embodiments are only part of the embodiments of this application, not all of the embodiments.
[0022] In the description of the embodiments of this application, words such as "for example" or "for instance" are used to indicate examples, illustrations, or explanations. Any embodiment or design described as "for example" or "for instance" in the embodiments of this application should not be construed as being preferred or advantageous over other embodiments or designs. Rather, the use of words such as "for example" or "for instance" is intended to present the relevant concepts in a concrete manner.
[0023] In the description of the embodiments of the present application, the term "multiple" means two or more. For example, multiple systems refer to two or more systems, and multiple screen terminals refer to two or more screen terminals. In addition, the terms "first" and "second" are used for descriptive purposes only and are not to be understood as indicating or implying relative importance or implicitly indicating the indicated technical features. Thus, the features defined as "first" and "second" may explicitly or implicitly include one or more of the features. The terms "including", "comprising", "having" and their variations all mean "including but not limited to", unless otherwise specifically emphasized.
[0024] Figure 1 An exemplary system architecture 100 is shown to which an embodiment of a traffic filtering method for extracting quintuple traffic features or a traffic filtering system for extracting quintuple traffic features of the present application can be applied.
[0025] like Figure 1 As shown, system architecture 100 may include terminal devices 101, 102, 103, a network 104, and a server 105. Network 104 is a medium for providing communication links between terminal devices 101, 102, 103 and server 105. Network 104 may include various connection types, such as wired or wireless communication links or fiber optic cables.
[0026] Users can use terminal devices 101, 102, and 103 to interact with server 105 via network 104 to receive or send messages, etc. Various communication client applications can be installed on terminal devices 101, 102, and 103, such as model training applications, video recognition applications, web browser applications, social platform software, etc.
[0027] Terminal devices 101, 102, 103 can be hardware or software. When terminal devices 101, 102, 103 are hardware, they can be various electronic devices with display screens, including but not limited to smart phones, tablet computers, e-book readers, MP3 (Moving Picture Experts Group Audio Layer III, Moving Picture Experts Compression Standard Audio Layer 3) players, MP4 (Moving Picture Experts Group Audio Layer IV, Moving Picture Experts Compression Standard Audio Layer 4) players, laptop computers and desktop computers, etc. When terminal devices 101, 102, 103 are software, they can be installed in the electronic devices listed above. It can be implemented as multiple software or software modules (for example, multiple software or software modules used to provide distributed services), or it can be implemented as a single software or software module. No specific limitation is made here.
[0028] When terminals 101, 102, and 103 are hardware, they may also be equipped with a video capture device. The video capture device may be any device capable of capturing video, such as a camera, a sensor, and the like. Users can use the video capture device on terminals 101, 102, and 103 to capture video.
[0029] The server 105 may be a server that provides various services, such as a background server that processes data displayed on the terminal devices 101, 102, and 103. The background server may analyze and process the received data, and may feed back the processing results (such as recognition results) to the terminal device.
[0030] It should be noted that the server can be either hardware or software. When the server is hardware, it can be implemented as a distributed server cluster consisting of multiple servers, or as a single server. When the server is software, it can be implemented as multiple software or software modules (e.g., multiple software or software modules used to provide distributed services), or as a single software or software module. No specific limitations are given here.
[0031] It should be understood that Figure 1 The number of terminal devices, networks, and servers in the above description is merely illustrative. Any number of terminal devices, networks, and servers may be used as needed. In particular, if target data does not need to be acquired remotely, the above system architecture may not include a network, but may instead include only terminal devices or servers.
[0032] Figure 2 This is a flow chart of a traffic filtering method for extracting five-tuple traffic features in an embodiment of the present application.
[0033] See also Figure 2 In an embodiment of the present application, a traffic filtering method for extracting quintuple traffic features is applied to a server, and the method includes: S201, obtaining a five-tuple feature sequence of a historical network communication data sample; The system first needs to extract representative structured features from historical network communication behaviors. Therefore, the core goal of step S201 is to obtain a five-tuple feature sequence of historical network communication data samples, providing raw data support for the subsequent extraction of temporal correlation features and construction of communication prompt words.
[0034] To achieve the above goals, the system collects historical network communication data samples from multiple sources, including but not limited to the following three categories: Actual communication log data within an enterprise or organization's internal network environment, such as traffic collection devices deployed on edge routers, core switches, intrusion detection systems (IDS), or traffic mirroring probes, capture and store large amounts of network data packets generated during daily business operations over a long period of time. Public network traffic datasets, such as standard traffic sample sets published by research institutions or network security laboratories. These datasets often contain labeled normal and abnormal traffic data and have good versatility and verifiability. Network communication data generated in a simulation or test environment, on a test platform that simulates a real network environment, generates controllable and repeatable historical communication data samples through preset communication tasks or attack behaviors (such as scanning, malicious login, etc.), which are used to enhance the model's ability to recognize specific abnormal patterns.
[0035] The system uniformly cleans and analyzes the communication data collected from the above sources. During the data analysis process, the system extracts a five-tuple feature from each network connection or data packet record: source IP (Internet Protocol) address, destination IP address, source port number, destination port number, and protocol type. As the fundamental identifier of network communication behavior, the five-tuple feature uniquely characterizes the end-to-end communication process and is a key indicator for identifying communication relationships and behavioral patterns.
[0036] The system further organizes the quintuple features in chronological order, forming a continuous, timestamped sequence of quintuple features. To ensure the temporal and contextual integrity of data analysis, the system uses a window sliding mechanism to aggregate the quintuple sequences within a set time window and sort them by communication start time, forming a structured time series dataset.
[0037] To improve data processing efficiency and subsequent access capabilities, the system encodes the quintuple feature sequence in a structured format (such as JSON or Parquet) and stores it in a high-performance time series database or distributed file system (such as a time series database (TSDB)). Furthermore, to support high-dimensional feature queries and model training, the system also tags each quintuple feature, including auxiliary fields such as communication direction, number of packets, number of bytes, and connection duration.
[0038] S202. Extracting the joint probability distribution between different quintuple features in the quintuple feature sequence within multiple time windows to obtain time series correlation features; Network communication data naturally has significant temporal characteristics, and there are obvious differences in communication behavior patterns in different time periods. Reasonable time window division can not only avoid the averaging of feature information due to the window being too large, but also prevent the failure to reflect the complete behavior pattern due to the window being too small. Therefore, before step S202, a dual verification mechanism of IP-related distribution characteristics and combined distribution characteristics is used to ensure that the division of time windows can simultaneously meet the behavioral characteristic requirements of the network layer and the transport layer, and based on the overlap screening, the representativeness of the selected time window in different dimensions is guaranteed, which provides a reasonable time scale for the subsequent extraction of the joint probability distribution between the five-tuple features.
[0039] For details, please refer to Figure 3 , Figure 3 This is another flow chart of a traffic filtering method for extracting quintuple traffic features in an embodiment of the present application.
[0040] S301, analyzing the association frequency between the source IP address and the destination IP address in the five-tuple feature sequence to obtain the association distribution feature; In order to identify the behavioral structural characteristics between the communication entities in the network communication and extract the distribution information that can quantify the communication relationship, the correlation frequency between the source IP address and the destination IP address in the five-tuple feature sequence is analyzed in step S301 to obtain the correlation distribution characteristics, and the correlation distribution characteristics used to represent the communication relationship structure are extracted. In the specific implementation process, by traversing the historical five-tuple feature sequence, the source IP address and the destination IP address in each five-tuple record are extracted, and statistics are performed in the form of communication pairs. Based on the statistical results, a communication frequency matrix of the source IP and the destination IP is constructed. In the communication frequency matrix, the rows represent the source IP address, the columns represent the destination IP address, and the values in the communication frequency matrix represent the number of communications between the corresponding source and destination in the entire historical data sequence.
[0041] After obtaining a complete correlation frequency matrix, we further normalize the frequency matrix to eliminate bias caused by differences in activity between different source IP addresses. This normalization process involves normalizing each row, converting the communication frequency between each source IP address and different destination IP addresses into relative proportions, forming a normalized probability distribution. After normalization, each row forms a correlation distribution vector, which reflects the communication preferences of the corresponding source IP address within a certain timeframe and its distribution characteristics within the destination address space.
[0042] To incorporate time sensitivity, the aforementioned association distribution vectors are time-sliced using a sliding time window mechanism. Trends in the communication distribution between source and destination IP addresses are extracted over different time periods, ultimately generating association distribution features with a temporal structure. These association distribution features characterize the evolution of network communication relationships and serve as the basis for determining the first time window in subsequent steps. They also provide behavioral structure support for constructing communication prompts based on semantic modeling.
[0043] S302, counting the combination patterns between target features in the quintuple feature sequence to obtain a combination distribution feature; To further explore the structural patterns of network communication behavior at the protocol and port levels, in step S302, statistical analysis is performed on the combination patterns of target features in the quintuple feature sequence to extract combined distribution features that characterize the distribution of communication service features. Target features include source port number, destination port number, and transport protocol type. The combination of target features reflects the call structure and traffic characteristics of application-layer services in network communications.
[0044] During implementation, the five-tuple feature sequence is parsed item by item, extracting the target feature combination of the source port number, destination port number, and transport protocol from each record. To enhance the temporal sensitivity of the combination pattern, a time sliding window mechanism is introduced, dividing the five-tuple sequence into several adjacent or overlapping time periods along the time axis. The frequency of the target feature combination is then independently counted within each time period.
[0045] The time sliding window is configured based on the overall time span of historical communication data and the density of changes in communication behavior. A fixed sliding window width, such as 5 minutes, 10 minutes, or 1 hour, is typically set, adjusting based on the network's communication activity. A sliding step size is also specified to determine the degree of overlap between adjacent windows. The window width and step size can be preset empirically, automatically optimized, or dynamically adjusted based on historical behavior fluctuations to ensure a balance between statistical stability and change detection sensitivity.
[0046] Within each sliding window, the frequency of all target feature combinations that occur is calculated to generate a combined frequency vector. To eliminate the impact of differences in communication activity within different windows, each combined frequency vector is further normalized to form a combined probability distribution vector. The combined distribution vectors within all windows are arranged in chronological order, forming a combined distribution feature with time series properties.
[0047] The resulting combined distribution features, specifically expressed as multi-dimensional indicators such as the high-frequency service port combinations, protocol usage bias, and combined structural complexity that occur within a given time period, reflect the structural characteristics of communication behavior at the application level. These combined distribution features can be used to distinguish between normal business patterns and abnormal behavior, such as identifying atypical communication patterns such as port scanning, protocol mixing, and service changes. These features also serve as the structural basis for determining the second time window in step S304. Furthermore, these features provide protocol-level context for the semantic modeling of subsequent communication prompts.
[0048] S303, determining a plurality of first time windows according to the correlation distribution; After completing the extraction of the associated distribution features, in order to further analyze the changing trend of the communication behavior pattern of the source-destination IP in the time dimension, in step S303, multiple first time windows are divided according to the associated distribution features. Specifically, the system uses a sliding window method to perform time series analysis on the normalized source-destination IP association frequency, and identifies the time period when the communication behavior changes significantly by calculating the rate of change, variance or mutation point of the communication density within the window. Based on each change inflection point, the system divides the time axis into multiple continuous or overlapping first time windows to capture the time segments when the strength of the source-destination IP relationship in the network changes significantly. This window division method based on the evolution of communication relationships helps to focus on time periods with behavioral value in subsequent analysis and improve the context relevance of model training.
[0049] S304, determining a plurality of second time windows according to the combined distribution characteristics; Unlike the first time window which focuses on the evolution of the relationship between communication entities, step S304 determines multiple second time windows based on the aforementioned combination distribution characteristics to capture the time evolution characteristics of the protocol and port combination pattern. The system also uses a sliding analysis window to process the time series of the source port number, destination port number and transmission protocol type combination frequency, and determines the structural change period of the communication behavior at the protocol level by detecting the frequent change points, sudden increase points or continuous high-frequency state of the combination pattern. Each second time window corresponds to a stable or abnormal combination pattern interval, which is used to locate the time point when potential service anomalies, port scanning or protocol deception occur. By introducing the time division method of the behavioral semantic dimension, it complements the first time window and provides a multi-dimensional time reference for the subsequent construction of a time series cross-analysis model.
[0050] S305. Calculate the overlap between a target first time window and a target second time window, where the target first time window is any first time window and the target second time window is any second time window; On the basis of obtaining multiple first time windows and second time windows, the overlap between any first time window and any second time window is further calculated in step S305 to evaluate the behavioral co-occurrence of the two types of time periods in the time dimension. Specifically, the system takes each first time window as the target first time window and each second time window as the target second time window, and calculates the ratio of their time overlapping intervals to joint intervals by comparing their start and end times, that is, the overlap. The calculation of the overlap is used to reflect whether the changes in the communication relationship and the changes in the service combination pattern tend to be consistent, and has a strong behavioral coupling significance. By calculating the overlap of all time window pairs, time segments with strong correlations can be identified, providing a basis for the subsequent construction of a joint temporal feature space.
[0051] S306: combining the target first time window and the target second time window with an overlap greater than a preset threshold to obtain a time window pair; Based on the aforementioned overlap calculation results, in step S306, time window pairs with an overlap greater than a preset threshold are further screened and combined into valid time window pairs as analysis objects. The preset threshold can be dynamically set based on historical data experience, system sensitivity, or security policy requirements. The system sorts all calculated time window pairs by overlap and selects window pairs with an overlap exceeding the threshold as time periods where communication behavior and combination patterns are highly coupled. During these time periods, network communication entity relationships and service structures change simultaneously, which is very likely to correspond to key events such as network anomalies, attack behaviors, or policy changes.
[0052] S307: Extract the overlapping time periods in each time window pair to obtain a time window.
[0053] After determining the time window pairs with high behavioral coupling, the overlapping time periods in each time window pair are further extracted in step S307, and used as the final time window for subsequent analysis. The system traverses all valid time window pairs, calculates the time intersection interval of each pair of windows, and uses the intersection as the time period representing the synchronous occurrence of changes in the source-destination communication relationship and changes in the service combination pattern. These time windows not only have strong semantic consistency, but also significantly improve the accuracy and interpretability of the subsequent constructed temporal correlation features. The extracted time windows serve as high-information-density analysis slices that can be used to drive communication prompt word generation, large language model training, and real-time anomaly analysis, thereby enhancing the responsiveness of the entire system in handling complex and dynamic network security incidents.
[0054] After the time window is determined, in step S202, within multiple time windows, the joint probability distribution between different quintuple features in the quintuple feature sequence is extracted to obtain the time series correlation feature, which may include the following steps: In each time window, the occurrence frequency of each quintuple feature and the co-occurrence frequency between different quintuple features are counted according to the joint probability distribution to construct a quintuple feature co-occurrence matrix; the temporal correlation features are determined based on the quintuple feature co-occurrence matrix.
[0055] In one embodiment, in order to further characterize the joint relationship structure of network communication behavior at the fine-grained feature level and improve the contextual expression ability of subsequent semantic modeling and behavior recognition, after completing the division of multiple time windows, the system counts the occurrence frequency of each quintuple feature and the co-occurrence frequency between different quintuple features in each time window according to the joint probability distribution, and constructs a quintuple feature co-occurrence matrix based on this.
[0056] During implementation, the system first traverses each five-tuple communication record within a defined time window to obtain a joint probability distribution. For each five-tuple record, the system extracts its five basic feature dimensions—source IP address, destination IP address, source port number, destination port number, and transport protocol type—and uses these features as the basic analysis units. During the statistical process, the system not only records the number of independent occurrences of each five-tuple feature within the current time window, but also further counts the frequency of co-occurrence of any two different five-tuple features within the same five-tuple record. For example, when a record contains source IP 192.168.1.10, destination IP 192.168.1.20, source port 52345, destination port 80, and protocol TCP (Transmission Control Protocol), the system will increase the independent counts of these features and increase the co-occurrence counts of the following feature pairs by one: <source IP, destination IP>, <source IP, source port>, <source IP, protocol>, <destination IP, destination port>, <destination port, protocol>, etc. By counting all records, a five-tuple feature co-occurrence matrix of dimension N×N is ultimately formed, where N is the total number of features involved (including various discrete values such as IP addresses, port numbers, and protocols). Each element in the matrix represents the number of times the corresponding two feature values co-occur within the current time window.
[0057] The purpose of constructing a quintuple feature co-occurrence matrix is to discover the inherent correlations between different feature dimensions in network communication behavior and reveal coupling patterns between features. For example, frequently co-occurring <destination port, communication protocol> pairs may correspond to a specific application service, while frequently co-occurring <source IP, source port> pairs may reflect the port allocation mechanism of the client application. Frequently changing or suddenly appearing new feature combinations may indicate anomalous behavior such as port scanning or spoofed communication. Further operations such as spectral clustering, graph modeling, or embedding learning on this co-occurrence matrix can extract semantically consistent feature substructures and construct high-dimensional feature representations for cue word generation or model input. Furthermore, the quintuple feature co-occurrence matrix exhibits good temporal scalability. Because the quintuple feature co-occurrence matrix construction process is based on time window partitioning, the system can construct and compare multiple co-occurrence matrices across different time windows, thereby analyzing the temporal evolution of feature relationship structures and identifying potential behavioral mutation points or anomalous patterns.
[0058] In a preferred embodiment, after obtaining the quintuple feature co-occurrence matrix, determining the time series correlation features based on the quintuple feature co-occurrence matrix may include the following steps: constructing a probability transfer chain based on the quintuple feature co-occurrence matrix, and calculating the correlation probability between different quintuple features; determining the joint probability distribution based on the correlation probability; and analyzing the changing trend of the joint probability distribution within multiple time windows to generate time series correlation features.
[0059] In order to further explore the structural probabilistic relationship between quintuple features from the quintuple feature co-occurrence matrix and improve the system's expressive ability in multi-dimensional feature joint modeling, the system constructs a probability transfer chain based on the quintuple feature co-occurrence matrix in each time window to obtain the association probability between different quintuple features. Specifically, the system regards the quintuple feature co-occurrence matrix as an adjacency matrix of a weighted undirected graph, where the nodes represent different quintuple feature values (such as a source IP, destination port or protocol type), and the edge weights represent the frequency of the two features co-occurring in the same quintuple record. Based on the weighted undirected graph, the system normalizes the outgoing edge weight of each node to obtain the conditional probability distribution from the feature to other co-occurring features, that is, to construct the first-order probability transfer vector of each node. Furthermore, the system can recursively construct high-order transmission chains through matrix multiplication. For example, it can calculate the probability that a certain feature is indirectly associated with other features through two-hop, three-hop, etc. paths, forming a multi-order probability transmission chain structure, and finally obtaining a probability transmission chain. The probability transmission chain can not only quantify the direct co-occurrence relationship, but also reveal the potential indirect coupling path between features, thereby more comprehensively characterizing the network structured association between the five-tuple features.
[0060] After obtaining the first-order and higher-order association probabilities between different quintuple features, the system further determines the joint probability distribution based on the association probabilities, which is used to express the probability structure of the combined appearance of multiple quintuple features within the current time window. In specific implementation, the system uses all edge conditional probabilities obtained in the probability transfer chain as the basis, combined with the boundary conditions of feature co-occurrence frequency, and adopts maximum entropy modeling or Bayesian joint inference methods to construct a joint probability model. Taking <destination port, protocol, source IP> as an example, the joint probability can be approximately calculated as follows: P(destination port, protocol, source IP)≈P(destination port | protocol)·P(protocol | source IP)·P(source IP).
[0061] In this modeling process, conditional probabilities are extracted from the probability transfer chain, and marginal probabilities are estimated from feature frequencies, thus forming a complete joint probability distribution model. This joint probability distribution not only reflects the co-occurrence patterns of multiple quintuple features in communication behavior but also quantitatively assesses the probability of occurrence of specific feature combinations, providing refined structural information for subsequent anomaly detection and semantic modeling.
[0062] After constructing the joint probability distribution over multiple time windows, the system further analyzes the structural evolution trend of these joint probability distribution sequences on the time axis, thereby generating time series correlation features that reflect the dynamic changes in the joint relationship of features. Specifically, the system represents the joint probability distribution structure within each time window as a high-dimensional vector or tensor, and can use time series modeling methods (such as sliding window difference analysis, principal component change tracking, or distribution distance measurement) to quantify the degree of change in the joint distribution between adjacent windows. For example, the system can calculate the KL divergence (Kullback-Leibler Divergence), JS divergence (Jensen-Shannon Divergence), or cosine similarity between the joint distributions of two adjacent time windows to measure the stability or mutation of the feature relationship structure. When the system detects that the joint distribution changes dramatically within a certain time period, or a new high-probability feature combination appears, it can mark the time period as a potential behavioral turning point or anomaly trigger point. Ultimately, the system encodes the joint distribution change trend over the entire time series into a temporal correlation feature, which can be used to drive multiple core tasks such as communication prompt word generation, abnormal behavior recognition, or historical behavior modeling, significantly enhancing the system's temporal understanding of complex communication behaviors and the accuracy of semantic expression.
[0063] S203, constructing a communication prompt word according to the temporal correlation feature; Extract the time series distribution characteristics of the joint probability distribution in the time series correlation features and the probability change characteristics of the probability transmission chain; identify the time series dependency between different five-tuple features based on the time series distribution characteristics and probability change characteristics; and construct communication prompt words based on the time series dependency.
[0064] In a preferred embodiment, to map the five-tuple feature sequences in historical network communication data samples into semantically expressive language model input, thereby enabling the large language model to understand and learn network behavior patterns, the system constructs communication prompts based on temporal correlation features in step S203 to guide language model learning. By extracting behavioral dependency features with temporal structure from joint probability distributions and probability transfer chains and encoding them in natural language or language-like structures, a prompt sequence with contextual relevance, semantic consistency, and behavioral interpretability is constructed, serving as an important input for subsequent training of the target large language model.
[0065] During the specific implementation process, the system first extracts the time series distribution characteristics of the joint probability distribution contained in the time series correlation characteristics, that is, the change trend of the joint probability of different five-tuple feature combinations (such as <source IP, destination port, protocol>) over time in multiple time windows. The change trend of the joint probability over time can be modeled by constructing a probability trajectory curve to identify the active period, mutation time point or trend transfer pattern of the feature combination. At the same time, the system also extracts the probability change characteristics of the probability transmission chain, that is, the dynamic evolution of the conditional probability path between each five-tuple feature in the time window dimension, such as whether the probability transfer from a source IP to a destination port shows an enhancement, attenuation or reversal trend. The probability change information reflects the potential dependencies and behavior stage characteristics of communication behaviors at different semantic levels.
[0066] After obtaining the above-mentioned time series distribution characteristics and probability change characteristics, the system further identifies the temporal dependency relationship between different five-tuple features. Specifically, by analyzing feature pairs with significant linkage changes, combining the temporal order of their appearance with the strength of probability coupling, the system constructs a dependency path diagram between multi-dimensional features, expressed as "the probability of source IP-A accessing destination port D under protocol P increases significantly during the T1-T2 period, followed by an increase in the frequency of associated communications with destination IP-B", thereby deducing the temporal behavior chain of "source IP-A → protocol P → destination port D → destination IP-B". This dependency relationship not only retains the structural information between the original features, but also introduces a temporal context, making the expression of communication behavior causal and phased.
[0067] Based on the identified timing dependencies, the system further maps the timing dependencies into communication prompt words in a format that can be understood by the language model. The construction form of the prompt words can adopt templated natural language sentences, nested structural expressions or vectorized encoding methods. For example, the system can convert the above dependencies into the following prompt words: "In the time period T1-T2, the source IP-A initiates communication with protocol P, frequently accesses port D, and then the destination IP-B becomes the main response target", or use a structured representation such as: "[T1-T2]A→P→D→B". This prompt word not only retains the multi-dimensional information of the original five-tuple features, but also gives it context continuity and semantic abstraction capabilities through linguistic organization, which facilitates the language model to perform semantic modeling, pattern learning and abnormal reasoning.
[0068] To enhance the target large language model's ability to generalize and identify diverse threat behaviors in complex network environments, the system introduces a universal semantic representation mechanism across attack types during the construction of communication prompt words. By semantically abstracting abnormal communication behaviors under different network protocol types and converting them into prompt word representations with unified behavioral intent, the target large language model is guided to learn universal semantic feature representations across protocols and attack types during training. This enables the target large language model to identify abnormal communication behaviors with heterogeneous protocols but similar behaviors during subsequent reasoning.
[0069] Traditional security detection methods typically rely on static attack signature matching or protocol-specific rules, making it difficult to capture the common behavioral characteristics of different attack types, resulting in blind spots in identification when facing attack scenarios with heterogeneous protocols and diverse behaviors. However, this implementation, by constructing a unified semantic representation space, enables the system to identify behavioral consistency between, for example, abnormal HTTP (HyperText Transfer Protocol) header field operations and abnormal RDP (Remote Desktop Protocol) connection modes. This breaks down protocol boundaries at the semantic level and significantly improves the model's ability to understand the nature of attacks.
[0070] During the specific implementation process, when constructing communication prompt words, the system does not simply input the five-tuple features as isolated fields, but instead extracts behavioral intentions, structural anomalies, and temporal features in combination with context to generate prompt word fragments with semantic abstraction capabilities. For example, for abnormal HTTP requests, the prompt words not only include "the source IP sends a large number of HTTP requests to the target server in a short period of time," but also further abstract semantic elements such as "abnormal arrangement of request header fields" and "includes rare User-Agent information." For abnormal RDP connection behavior, the prompt words may include features such as the source host attempting to establish multiple remote desktop connections during non-working hours. Although the above prompt words come from different protocols, they can all be summarized at the semantic level as unified behavioral intentions such as abnormal access behavior and connection patterns that differ significantly from normal ones, thereby achieving semantic normalization and generalized modeling across protocols.
[0071] S204: input the communication prompt word into a preset large language model for training to obtain a target large language model; A training sample set is constructed based on communication prompt words; the training sample set is input into a preset large language model to obtain training parameters; the accuracy of the preset large language model in recognizing the temporal features of the five-tuple feature sequence is verified based on the training parameters, and the training is terminated when the accuracy of the temporal feature recognition meets the preset requirements.
[0072] In a preferred embodiment, to enable a pre-set large language model to accurately model and identify the temporal features of quintuple feature sequences in network communication data, the system constructs a training sample set based on the previously constructed communication prompt words and uses this training sample set as input to train the pre-set large language model, thereby obtaining a target large language model capable of performing communication behavior semantic modeling and anomaly identification tasks. The core goal of the training phase is to map complex, multi-dimensional communication behaviors into sequence inputs with contextual expression capabilities through linguistic prompt words, enabling the language model to learn the potential dependencies between quintuple features, temporal evolution patterns, and behavioral semantic features.
[0073] During the specific implementation process, communication prompt words are first constructed based on the extracted temporal correlation features, and these prompt words are paired with the corresponding original five-tuple feature sequences to form a training sample set. Each training sample consists of an input part and an output part, where the input part is the communication prompt word, and the output part is the five-tuple feature sequence or its structured representation (such as vector encoding, behavior label, dependency graph structure, etc.) within the corresponding time window. For example, the input prompt word is: "In the time period T3-T4, the source IP 192.168.1.10 frequently accesses port 443 using the TCP protocol and establishes communication with the destination IP 192.168.1.20", and the corresponding output sequence is multiple five-tuple records, such as <192.168.1.10, 192.168.1.20, *, 443, TCP>, where * represents any source port. Through the construction of a large number of sample pairs, the training sample set comprehensively covers different communication behavior patterns and their temporal feature expressions.
[0074] Subsequently, the training sample set is input into the preset large language model for training. During the training process, the system adopts a supervised learning approach, with communication prompt words as the input sequence of the language model and a five-tuple feature sequence as the output sequence. The language model is trained to restore or predict the corresponding communication behavior sequence by encoding the semantic information in the prompt words. During the training process, the system uses cross-entropy loss functions, attention mechanism optimization methods and other technologies to iteratively update the language model parameters, improving the model's ability to model the temporal logic, feature combination dependencies and behavioral semantics in the prompt information. After multiple rounds of training, the model parameters gradually converged, and the language model was able to accurately capture the five-tuple behavior structure implied by the prompt words, achieving an effective mapping from language expression to behavioral patterns.
[0075] To ensure that training results meet practical application requirements, the system validates the model based on the current training parameters after each training cycle, evaluating the language model's accuracy in recognizing temporal features in the quintuple feature sequence. This accuracy is calculated by comparing the model's predictions with the actual quintuple sequence and statistically analyzing the degree of match across dimensions such as time order, feature combination, and behavioral phase. If the current temporal feature recognition accuracy meets preset requirements, these requirements are derived from statistical analysis of a large amount of real-world network communication data and attack samples, combined with security experts' analysis of common attack behavior phases, quintuple sequence structures, and their evolution patterns, and further referenced by the system's performance requirements in typical application scenarios (such as APT attack detection and lateral movement identification). For example, threshold standards are established, such as a sequence structure recovery accuracy of no less than 95% and a stage behavior consistency score of no less than 90%. These preset standards reflect the minimum operational requirements for model recognition capabilities and ensure that the system training process has clear convergence targets and quantifiable performance evaluation criteria. If the current temporal feature recognition accuracy meets the preset requirements, the system determines that the model has demonstrated good behavioral modeling capabilities, terminates the training process, and outputs the target large language model. If the accuracy rate does not meet the preset standard, the system will continue to iterate training until the termination condition is met.
[0076] To further enhance the target large language model's responsiveness and generalization capabilities to new types of abnormal communication behaviors during actual deployment, after completing the target large language model training based on communication prompt words, the system introduces a small sample prompting learning (Few-shot Prompting) mechanism to achieve immediate adaptation to emerging threat patterns. Traditional security models built based on large-scale training samples often suffer from long model update cycles and poor adaptability when faced with sudden, mutated, or unknown attack behaviors. However, through the Few-shot Prompting technology in this embodiment, the target large language model can complete the modeling and identification of new threat characteristics without retraining, relying only on a very small number of representative abnormal communication prompt words, thereby achieving high-efficiency and low-latency threat adaptation capabilities.
[0077] During implementation, when a new communication behavior is detected that isn't matched by the current policy, or when an administrator manually labels a newly discovered abnormal communication sample, the system structures the five-tuple feature sequence corresponding to the new communication behavior and generates a new prompt word segment based on the same construction rules as the original communication prompt word. This prompt word typically includes a contextual description of the abnormal behavior, a typical five-tuple combination, information about the time window in which it occurred, and possible attack intent keywords. For example, for a new type of abnormal DNS tunneling behavior, the prompt word might be constructed as: "Source IP 10.1.1.5 sends a high frequency of DNS requests to multiple external IP addresses within 30 seconds. The request packet header length is abnormal, and it is suspected that a data channel has been established."
[0078] The system then feeds this cue word, along with a small number of known anomaly cue word samples, into the target large language model. Through contextual learning, the model is guided to understand the semantic connections and behavioral structural similarities between this new behavior and existing anomaly patterns. Because the target large language model has already learned the joint probability distribution and temporal correlation features between a large number of quintuple features during training, it possesses strong semantic generalization and pattern transfer capabilities. Given a small number of new cue word samples, it can automatically construct a semantic representation of the new behavior and incorporate it into the existing anomaly recognition framework.
[0079] Through this small-sample cue learning mechanism, the system can rapidly adapt and identify new communication threats without retraining or fine-tuning the parameters of the target large language model. During subsequent real-time anomaly analysis, if the system receives a real-time five-tuple feature sequence with a similar structure to the cue word, the model can directly identify it as abnormal behavior based on the learned semantic pattern and output the corresponding analysis results, triggering the subsequent anomaly type determination and strategy generation process.
[0080] S205, performing anomaly analysis on the real-time network communication data using the target large language model to obtain analysis results; In a preferred embodiment of the present invention, to achieve intelligent identification and anomaly detection of complex and dynamic network communication behavior, in step S205, a target large language model is used to perform anomaly analysis on real-time network communication data to obtain analysis results. Using the target large language model trained on communication prompt words, the real-time 5-tuple feature sequence is semantically understood and compared with temporal patterns, thereby identifying deviations in the communication structure and temporal evolution of abnormal behavior.
[0081] During implementation, the system continuously collects real-time network communication data and parses each communication record into standardized five-tuple features, including fields such as source IP, destination IP, source port, destination port, and protocol. The system organizes these five-tuple records within consecutive time periods into feature sequences. Using a time window partitioning method consistent with the historical modeling process, the system sequentially organizes the real-time data into multiple short-term sequence blocks, forming a real-time five-tuple feature sequence.
[0082] Subsequently, the system calls the target large language model and inputs the real-time five-tuple feature sequence into the model for reasoning and analysis. During the training phase, the target large language model has learned a large number of joint probability structures, conditional dependency paths, and temporal evolution laws between five-tuple features through communication prompt words, and has the ability to perform context modeling, semantic understanding, and structural recognition of communication behaviors. During the reasoning process, the model compares and analyzes the current input sequence with the normal communication behavior patterns it has learned internally, and automatically identifies whether there are probability anomalies, path anomalies, or temporal structure anomalies in the feature combination. For example, if the model recognizes that a source IP accesses multiple unauthorized ports in a short period of time, or that a protocol frequently appears in an atypical time period, the model will determine that the behavior deviates from the normal pattern based on its semantic knowledge and output an anomaly mark.
[0083] To enhance the interpretability of anomaly analysis, the target large language model also outputs corresponding semantic annotations, such as "The source IP communication path does not match the normal pattern in the trained model" or "The current behavior combination is rare in the historical distribution," thereby enhancing the actionability and credibility of the analysis results. The model outputs analysis results including anomaly detection labels, anomaly scores, and key triggering features, providing foundational support for subsequent anomaly type determination and strategy selection.
[0084] S206: Determine the abnormality type based on the analysis result, determine the traffic filtering strategy according to the abnormality type and the preset strategy library, and perform traffic filtering using the traffic filtering strategy.
[0085] To respond to and intercept anomalous network behavior, in step S206, the system determines the anomaly type based on the analysis results output by the target large language model. Based on this, it combines the pre-set policy library to determine the corresponding traffic filtering strategy, ultimately executing targeted traffic filtering operations. This step is designed to further introduce a dynamic decision-making mechanism after anomaly detection, enabling the system to adopt differentiated handling strategies based on different types of anomalous behavior, thereby improving the intelligence and adaptability of network protection.
[0086] During implementation, the system first receives the analysis results output by the target large language model in step S205. These results contain information such as anomaly tags, anomaly scores, feature trigger paths, and semantic annotations for real-time network communication data. The system then performs a structured analysis of the analysis results, extracting key fields that reflect the characteristics of abnormal behavior, such as the anomaly triggering quintuple combination, the time period of abnormal behavior, the communication direction, the protocol type, and the access frequency. Combined with the anomaly annotations provided by the large language model, the system performs semantic understanding and preliminarily identifies the manifestation characteristics and potential intentions of the abnormal behavior.
[0087] After identifying relevant abnormal behavior characteristics, the system compares the current behavior characteristics with the anomaly type rule base through a preset anomaly type matching mechanism. The anomaly type rule base predefines various network anomaly types and their behavioral characteristic patterns, such as "port scanning behavior," "lateral movement behavior," "abnormal data leakage," and "denial of service attacks." Each type is configured with corresponding judgment criteria and semantic descriptions. Based on the current abnormal behavior's performance in terms of feature combination, frequency evolution, and target concentration, the system performs similarity matching or semantic inference with the type definitions in the rule base to determine the anomaly type to which the abnormal behavior belongs.
[0088] After identifying the anomaly type, the system invokes the traffic filtering policy corresponding to the anomaly type from the preset policy library based on the identification results. This preset policy library is constructed based on a large amount of historical network security event data, industry security standards, and security expert knowledge. By categorizing and analyzing typical abnormal behavior patterns (such as port scanning, DDoS (Distributed Denial of Service) attacks, and data leaks), it formulates corresponding filtering rules and response policies based on behavioral characteristics and protection requirements. These policies are then stored in the system in a structured format. This allows for rapid matching and generation of corresponding traffic filtering policies after the anomaly type is determined, ensuring the system has automated and precise protection capabilities against various network threats. Traffic filtering policies include rule configurations across multiple dimensions, including but not limited to source / destination IP blocking, port blocking, protocol restrictions, session termination, and access frequency threshold control. Based on the specific parameters of the current abnormal behavior, the system generates a set of filtering rules that conform to the policy template and sends these rules to network traffic control modules (such as firewalls, intrusion prevention systems, or border gateway devices) for real-time execution.
[0089] For example, when the system identifies that a source IP address accesses port 22 of a large number of destination IPs in a short period of time, and the behavior pattern highly matches "lateral movement behavior", the system will generate the following filtering policy based on the corresponding policy in the policy library: "Block all access requests from the source IP to port 22 of the internal network segment in the next 30 minutes" and immediately apply the policy to the network edge device, thereby achieving immediate interception of the abnormal communication behavior.
[0090] See also Figure 4 , is a structural diagram of a traffic filtering system for extracting quintuple traffic features provided in an embodiment of the present application. A traffic filtering system 400 for extracting quintuple traffic features specifically includes: The acquisition module 401 is used to obtain a five-tuple feature sequence of historical network communication data samples, where the five-tuple feature sequence includes a source IP address, a destination IP address, a source port number, a destination port number, and a transmission protocol; the extraction module 402 is used to extract the joint probability distribution between different five-tuple features in the five-tuple feature sequence within multiple time windows to obtain time series correlation features; the construction module 403 is used to construct communication prompt words based on the time series correlation features; the training module 404 is used to input the communication prompt words into a preset large language model for training to obtain a target large language model; the analysis module 405 is used to perform anomaly analysis on real-time network communication data using the target large language model to obtain analysis results; the filtering module 406 is used to determine the anomaly type based on the analysis results, determine the traffic filtering strategy based on the anomaly type and the preset strategy library, and perform traffic filtering using the traffic filtering strategy.
[0091] Optionally, the extraction module 402 is specifically used to: within each time window, count the occurrence frequency of each quintuple feature and the co-occurrence frequency between different quintuple features according to the joint probability distribution, and construct a quintuple feature co-occurrence matrix; determine the temporal correlation features according to the quintuple feature co-occurrence matrix.
[0092] Optionally, the extraction module 402 is further specifically used to: construct a probability transfer chain based on the co-occurrence matrix of the five-tuple features, calculate the association probability between different five-tuple features; determine the joint probability distribution based on the association probability; analyze the changing trend of the joint probability distribution in multiple time windows, and generate time series association features.
[0093] Optionally, construction module 403 is specifically used to: extract the time series distribution characteristics of the joint probability distribution in the time series correlation characteristics and the probability change characteristics of the probability transmission chain; identify the time series dependency between different quintuple features based on the time series distribution characteristics and the probability change characteristics; and construct communication prompt words based on the time series dependency.
[0094] Optionally, the training module 404 is specifically used to: construct a training sample set based on the communication prompt words; input the training sample set into a preset large language model to obtain training parameters; verify the accuracy of the preset large language model in recognizing the temporal features of the five-tuple feature sequence according to the training parameters, and end the training when the accuracy of the temporal feature recognition meets the preset requirements.
[0095] Optionally, the system also includes a determination module 407, which is specifically used to: analyze the association frequency between the source IP address and the destination IP address in the five-tuple feature sequence to obtain the association distribution characteristics; count the combination patterns between the target features in the five-tuple feature sequence to obtain the combination distribution characteristics, the target features include the source port number, the destination port number, and the transmission protocol; determine multiple time windows based on the association distribution characteristics and the combination distribution characteristics.
[0096] Optionally, the determination module 407 is further specifically used to: determine multiple first time windows based on associated distribution characteristics; determine multiple second time windows based on combined distribution characteristics; calculate the overlap between the target first time window and the target second time window, where the target first time window is any first time window and the target second time window is any second time window; combine the target first time window and the target second time window whose overlap is greater than a preset threshold to obtain a time window pair; extract the overlapping time period in each time window pair to obtain a time window.
[0097] It should be noted that the above embodiments provide devices that implement their functions using only the division of the above functional modules as examples. In actual applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. In addition, the device and method embodiments provided in the above embodiments are based on the same concept. The specific implementation process is detailed in the method embodiment and will not be repeated here.
[0098] This embodiment also discloses an electronic device, referring to Figure 5The electronic device may include: at least one processor 501, at least one communication bus 502, a user interface 503, a network interface 504, and at least one memory 505. The communication bus 502 is used to implement connection and communication between these components. The user interface 503 may include a display screen (Display) and a camera (Camera), and optionally the user interface 503 may also include a standard wired interface and a wireless interface. The network interface 504 may optionally include a standard wired interface and a wireless interface (such as a WI-FI interface). The processor 501 may include one or more processing cores. The processor 501 uses various interfaces and lines to connect various parts within the entire server, and executes various functions of the server and processes data by running or executing instructions, programs, code sets or instruction sets stored in the memory 505, and calling data stored in the memory 505. Optionally, the processor 501 can be implemented in at least one hardware form of digital signal processing (DSP), field-programmable gate array (FPGA), and programmable logic array (PLA). Processor 501 may integrate one or a combination of a central processing unit (CPU), a graphics processing unit (GPU), and a modem. The CPU primarily handles the operating system, user interface, and application programs; the GPU is responsible for rendering and drawing content displayed on the display; and the modem handles wireless communications. It is understood that the modem may not be integrated into processor 501 and may be implemented as a separate chip.
[0099] Among them, the memory 505 may include a random access memory (RAM) or a read-only memory (Read-Only Memory). Optionally, the memory 505 includes a non-transitory computer-readable storage medium. The memory 505 can be used to store instructions, programs, codes, code sets or instruction sets. The memory 505 may include a program storage area and a data storage area, wherein the program storage area may store instructions for implementing an operating system, instructions for at least one function (such as a touch function, a sound playback function, an image playback function, etc.), instructions for implementing the above-mentioned various method embodiments, etc.; the data storage area may store data involved in the above-mentioned various method embodiments, etc. The memory 505 may also be optionally at least one storage device located away from the aforementioned processor 501. As Figure 5 As shown, the memory 505 as a computer storage medium may include an operating system, a network communication module, a user interface module, and an application program of a traffic filtering method for extracting quintuple traffic features.
[0100] exist Figure 5 In the electronic device shown, the user interface 503 is mainly used to provide an input interface for the user and obtain data input by the user; and the processor 501 can be used to call an application program stored in the memory 505 for a traffic filtering method for extracting five-tuple traffic features. When executed by one or more processors 501, the electronic device executes one or more methods as in the above-mentioned embodiments.
[0101] It should be noted that for the aforementioned method embodiments, for simplicity of description, they are all expressed as a series of action combinations, but those skilled in the art should be aware that this application is not limited by the order of the actions described, because according to this application, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily required for this application.
[0102] In the above embodiments, the description of each embodiment has its own emphasis. For parts that are not described in detail in a certain embodiment, please refer to the relevant description of other embodiments. In the several embodiments provided in this application, it should be understood that the disclosed device can be implemented in other ways. For example, the device embodiments described above are only schematic. For example, the division of units is only a logical function division. There may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some service interfaces, and the indirect coupling or communication connection of devices or units can be electrical or other forms.
[0103] The units described as separate components may or may not be physically separate, and the components displayed as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the scheme of this embodiment. In addition, the functional units in the various embodiments of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated units may be implemented in the form of hardware or in the form of software functional units.
[0104] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable memory. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product. The computer software product is stored in a memory 505 and includes several instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the various embodiments of the present application. The aforementioned memory 505 includes various media that can store program codes, such as a USB flash drive, a mobile hard disk, a magnetic disk, or an optical disk.
[0105] The above is merely an exemplary embodiment of the present disclosure and is not intended to limit the scope of the present disclosure. That is, any equivalent changes and modifications made in accordance with the teachings of the present disclosure are still within the scope of the present disclosure. After considering the disclosure of the specification, those skilled in the art will easily think of other embodiments of the present disclosure. This application is intended to cover any variations, uses, or adaptive changes of the present disclosure, which follow the general principles of the present disclosure and include common knowledge or customary technical means in the art that are not described in the present disclosure.
Claims
1. A traffic filtering method for extracting quintuple traffic features, characterized in that: Applied to a server, the method includes: Obtain a five-tuple feature sequence of a historical network communication data sample, wherein the five-tuple feature sequence includes a source IP address, a destination IP address, a source port number, a destination port number, and a transmission protocol; Extracting the joint probability distribution between different quintuple features in the quintuple feature sequence within multiple time windows to obtain time series correlation features; Constructing a communication prompt word according to the temporal correlation feature; Inputting the communication prompt words into a preset large language model for training to obtain a target large language model; Performing anomaly analysis on real-time network communication data using the target large language model to obtain analysis results; The abnormality type is determined based on the analysis result, a traffic filtering strategy is determined according to the abnormality type and a preset strategy library, and traffic filtering is performed using the traffic filtering strategy.
2. The method according to claim 1, characterized in that Before extracting the joint probability distribution between different quintuple features in the quintuple feature sequence within multiple time windows to obtain the time series correlation features, the method includes: Analyzing the association frequency between the source IP address and the destination IP address in the five-tuple feature sequence to obtain an association distribution feature; Counting the combination patterns between target features in the quintuple feature sequence to obtain a combination distribution feature, where the target features include the source port number, the destination port number, and the transmission protocol; A plurality of time windows are determined according to the associated distribution feature and the combined distribution feature.
3. The method according to claim 2, characterized in that The determining of the plurality of time windows according to the associated distribution feature and the combined distribution feature specifically includes: determining a plurality of first time windows according to the associated distribution characteristics; determining a plurality of second time windows according to the combined distribution characteristics; Calculating an overlap between a target first time window and a target second time window, where the target first time window is any of the first time windows and the target second time window is any of the second time windows; Combining the target first time window and the target second time window whose overlap is greater than a preset threshold to obtain a time window pair; The overlapping time period in each of the time window pairs is extracted to obtain the time window.
4. The method according to claim 1, wherein Extracting the joint probability distribution between different quintuple features in the quintuple feature sequence within multiple time windows to obtain time series correlation features specifically includes: In each of the time windows, the occurrence frequency of each of the five-tuple features and the co-occurrence frequency between different five-tuple features are counted according to the joint probability distribution to construct a five-tuple feature co-occurrence matrix; The temporal correlation feature is determined according to the quintuple feature co-occurrence matrix.
5. The method according to claim 4, characterized in that The determining the temporal correlation feature according to the quintuple feature co-occurrence matrix specifically includes: Constructing a probability transfer chain based on the co-occurrence matrix of the five-tuple features, and calculating the association probability between different five-tuple features; determining the joint probability distribution according to the association probability; Within the multiple time windows, the changing trend of the joint probability distribution is analyzed to generate the time series correlation feature.
6. The method according to claim 5, characterized in that The step of constructing a communication prompt word according to the temporal correlation feature specifically includes: Extracting the time series distribution characteristics of the joint probability distribution and the probability change characteristics of the probability transfer chain from the time series correlation characteristics; Identifying temporal dependencies between different quintuple features based on the time series distribution characteristics and the probability change characteristics; The communication prompt words are constructed according to the temporal dependency relationship.
7. The method according to claim 1, characterized in that The inputting the communication prompt word into a preset large language model for training to obtain a target large language model specifically includes: Constructing a training sample set based on the communication prompt words; Inputting the training sample set into the preset large language model to obtain training parameters; The accuracy of the temporal feature recognition of the quintuple feature sequence by the preset large language model is verified according to the training parameters, and the training is terminated when the accuracy of the temporal feature recognition meets the preset requirements.
8. A traffic filtering system for extracting quintuple traffic features, characterized in that: include: An acquisition module is used to obtain a five-tuple feature sequence of a historical network communication data sample, wherein the five-tuple feature sequence includes a source IP address, a destination IP address, a source port number, a destination port number, and a transmission protocol; An extraction module, configured to extract, within a plurality of time windows, a joint probability distribution between different quintuple features in the quintuple feature sequence to obtain a time series correlation feature; A construction module, configured to construct a communication prompt word according to the temporal correlation feature; A training module, configured to input the communication prompt words into a preset large language model for training to obtain a target large language model; An analysis module, configured to perform an anomaly analysis on the real-time network communication data using the target large language model to obtain an analysis result; The filtering module is used to determine the abnormality type based on the analysis result, determine the traffic filtering strategy according to the abnormality type and a preset strategy library, and perform traffic filtering according to the traffic filtering strategy.
9. A traffic filtering device for extracting quintuple traffic features, characterized in that: include: one or more processors and memory; The memory is coupled to the one or more processors, and the memory is used to store computer program code, wherein the computer program code includes computer instructions, and the one or more processors call the computer instructions to enable the traffic filtering device for extracting quintuple traffic features based on a large language model to perform the method as described in any one of claims 1 to 7.
10. A computer-readable storage medium comprising instructions, characterized in that: When the instruction is executed on a traffic filtering device that extracts quintuple traffic features based on a large language model, the traffic filtering device that extracts quintuple traffic features based on a large language model executes the method according to any one of claims 1 to 7.
Citation Information
Cited By
Library entity dependency conflict identification method for Python environment
CN120631736A
Method and system for dynamically optimizing firewall access control strategy based on cooperation of flow analysis and large language model
CN121000489A