Network security intrusion detection method and system based on large model

By building access behavior benchmark comparison tables and large models to predict potential attack nodes, the problem of traditional intrusion detection methods insufficient identification capabilities in complex network environments is solved, and efficient detection and early warning of unknown and collaborative attacks is achieved.

CN120378227AActive Publication Date: 2025-07-25JIANGSU GUOBAO INFORMATION SYST EVALUATION CENT CO LTD

Patent Information

Application Number
CN202510865080.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-26
Publication Date
2025-07-25
Estimated Expiration
2045-06-26

AI Technical Summary

Technical Problem

Traditional intrusion detection methods lack recognition capabilities when dealing with unknown or mutant attacks, and lack effective means of identifying the coordinated behavior patterns between multiple access nodes, resulting in low detection accuracy, lagging response and easy missed reporting.

Method used

By constructing an access behavior benchmark comparison table, identify access deviation nodes and synchronous behavior nodes, and predict potential attack nodes and their attack probability in combination with large models, and formulate node detection strategies for network security intrusion detection.

Benefits of technology

It significantly improves the detection capabilities of unknown attacks and collaborative attacks, realizes intelligent identification and accurate early warning of network intrusion behavior, and improves the accuracy, real-time and active defense capabilities of intrusion detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120378227A_ABST
    Figure CN120378227A_ABST
Patent Text Reader

Abstract

The invention provides a network security intrusion detection method and system based on a large model, and relates to the technical field of network security, and the method comprises the steps: monitoring and obtaining a plurality of access behavior data sequences in a current time period, carrying out node difference comparison through employing an access behavior reference comparison table, and constructing a deviation node distribution sequence; performing node similarity comparison on the plurality of access behavior data sequences, performing same-frequency node identification, and constructing a same-frequency node distribution sequence; in combination with the large model, potential attack node distribution and potential attack probability distribution are obtained through prediction according to the deviation node distribution sequence and the same-frequency node distribution sequence, and a node detection strategy is made for detection. The objective of the invention is to solve the problems of low detection accuracy, response lag, easy missing report and the like of an intrusion detection system in a complex network environment due to the lack of effective identification of a cooperative behavior mode among a plurality of access nodes in a traditional detection method. And the accuracy, the real-time performance and the active defense capability of intrusion detection can be remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of network security technology, and in particular, to a network security intrusion detection method and system based on a large model. Background Art

[0002] With the development of information technology and the complexity of enterprise network environments, network attack means have been continuously evolving, showing new characteristics such as strong concealment, frequent mutation, and collaborative attacks.

[0003] Traditional intrusion detection methods mainly rely on signature matching, rule bases, or simple statistical models. Although they have certain effects in detecting known attacks, their ability to identify unknown attacks, zero-day attacks, and variant attacks is limited; in addition, as attack behaviors tend to be distributed and collaborative, multiple access nodes may exhibit highly similar abnormal behaviors in a short period of time, forming collaborative attack patterns such as lateral movement and botnets. However, existing detection methods often lack a modeling and identification mechanism for such collaborative behaviors and are difficult to perceive potential threats in a timely manner. Summary of the Invention

[0004] The purpose of the present invention is to provide a network security intrusion detection method and system based on a large model to solve the problems that traditional intrusion detection methods have insufficient identification capabilities when dealing with unknown or variant attacks and lack effective identification means for collaborative behavior patterns between multiple access nodes, resulting in low detection accuracy, response lag, and easy false negatives in the intrusion detection system in a complex network environment, including: In a first aspect, the present invention provides a network security intrusion detection method based on a large model, including: clustering the access behaviors of several access nodes according to the network logs of the enterprise internal network in a preset historical time zone to construct an access behavior benchmark comparison table; monitoring and obtaining several access behavior data sequences of the several access nodes in the current period, using the access behavior benchmark comparison table to compare the node differences of the several access behavior data sequences to construct a deviation node distribution sequence; comparing the node similarities of the several access behavior data sequences and identifying co-frequency nodes to construct a co-frequency node distribution sequence; combining a large model, predicting and obtaining a potential attack node distribution and a potential attack probability distribution according to the deviation node distribution sequence and the co-frequency node distribution sequence, and formulating a node detection strategy for network security intrusion detection.

[0005] Preferably, the network security intrusion detection method based on a large model further includes: randomly selecting a first access node, screening the network logs according to the first access node to determine a first access behavior dataset; dividing the first access behavior dataset based on a preset time interval to obtain multiple access behavior datasets; performing access behavior clustering according to the multiple access behavior datasets to determine multiple high-frequency access behavior characteristics, where the access behavior characteristics include access frequency, single-time data traffic, and single-time duration; establishing a first access behavior benchmark for the first access node according to the multiple access time periods and multiple high-frequency access behavior characteristics, and sequentially analyzing and determining several access behavior benchmarks to construct an access behavior benchmark comparison table.

[0006] Preferably, the network security intrusion detection method based on a large model further includes: randomly selecting a first access behavior dataset, performing clustering using the K-means algorithm to obtain multiple first clustering clusters and multiple first central feature values; counting the data volumes of the multiple first clustering clusters, and selecting the first clustering clusters with the largest data volumes at a preset ratio as high-frequency clustering clusters to determine multiple high-frequency clustering clusters, and mapping to obtain multiple high-frequency central feature values; calculating the mean of the multiple high-frequency central feature values to obtain a first high-frequency access behavior characteristic, and sequentially analyzing to obtain multiple high-frequency access behavior characteristics.

[0007] Preferably, the network security intrusion detection method based on a large model further includes: using the access behavior benchmark comparison table to match and obtain several adapted access behavior benchmarks for several access nodes in the current time period; based on the several adapted access behavior benchmarks, sequentially comparing the several access behavior data at the same time point in the several access behavior data sequences to determine multiple deviation node distributions, and constructing a deviation node distribution sequence.

[0008] Preferably, the network security intrusion detection method based on a large model further includes: respectively calculating the deviation amplitudes between the adapted access behavior benchmarks and the corresponding access behavior data to obtain an access frequency deviation ratio, a single-time data traffic deviation ratio, and a single-time duration deviation ratio, and performing weighted calculation to obtain a node deviation value, and sequentially analyzing to obtain several deviation values for several access nodes; obtaining the node position distributions of several access nodes, and using the several deviation values to perform mapping identification on the node position distributions to obtain a deviation node distribution.

[0009] Preferably, the network security intrusion detection method based on a large model further includes: performing similarity clustering on several access behavior data at the same time point in the several access behavior data sequences according to a predetermined similarity threshold, and performing co-frequency node identification on the access nodes corresponding to the access behavior data that meet the predetermined similarity threshold to determine a co-frequency node set sequence; respectively identifying the node position distributions of several access nodes according to the co-frequency node set sequence to obtain a co-frequency node distribution sequence.

[0010] Preferably, the network security intrusion detection method based on a large model further includes: screening sample data according to the network logs of similar enterprises to obtain a sample deviation node distribution sequence set and a sample co-frequency node distribution sequence set, and analyzing to obtain a high-frequency attack node distribution and a high-frequency attack probability distribution under different sample deviation node distribution sequences and sample co-frequency node distribution sequences, obtaining a sample attack node distribution set and a sample attack probability distribution set; using the sample deviation node distribution sequence set and the sample co-frequency node distribution sequence set as inputs, using the sample attack node distribution set and the sample attack probability distribution set as supervision, training a large model to construct a network attack recognition plug-in; using the network attack recognition plug-in to predict and obtain a potential attack node distribution and a potential attack probability distribution according to the deviation node distribution sequence and the co-frequency node distribution sequence.

[0011] Preferably, the network security intrusion detection method based on a large model further includes: performing risk identification on potential attack nodes that meet a preset probability threshold to determine a potential risk node distribution; performing a concentration analysis of the position distribution of the potential risk node distribution to determine a risk node distribution concentration; if the risk node distribution concentration is less than or equal to a preset concentration index, formulating a node detection strategy as local detection, where local detection is to perform network security intrusion detection on risk nodes in descending order of potential attack probability.

[0012] Preferably, the network security intrusion detection method based on a large model further includes: if the risk node distribution concentration is greater than a preset concentration index, formulating a node detection strategy as global detection, where global detection is to perform network security intrusion detection on all nodes in descending order of potential attack probability.

[0013] Second aspect, the present invention further provides a network security intrusion detection system based on a large model, which is used to execute a network security intrusion detection method based on a large model as described in the first aspect, including: an access behavior clustering module, which is used to cluster the access behaviors of a number of access nodes according to the network logs of the enterprise internal network in a preset historical time zone, and construct an access behavior benchmark comparison table; a deviation node distribution sequence construction module, which is used to monitor and obtain a number of access behavior data sequences of the number of access nodes in the current period, and use the access behavior benchmark comparison table to perform node difference comparison on the number of access behavior data sequences to construct a deviation node distribution sequence; a same-frequency node distribution sequence construction module, which is used to perform node similarity comparison on the number of access behavior data sequences and perform same-frequency node identification to construct a same-frequency node distribution sequence; a node detection strategy formulation module, which is used to combine a large model, predict and obtain a potential attack node distribution and a potential attack probability distribution according to the deviation node distribution sequence and the same-frequency node distribution sequence, and formulate a node detection strategy for network security intrusion detection.

[0014] The embodiments of the present invention have the following advantages: By clustering the access behaviors of a number of access nodes according to the network logs of the enterprise internal network in a preset historical time zone, an access behavior benchmark comparison table is constructed; then, a number of access behavior data sequences of the number of access nodes in the current period are monitored and obtained, and the access behavior benchmark comparison table is used to perform node difference comparison on the number of access behavior data sequences to construct a deviation node distribution sequence; on the other hand, node similarity comparison is performed on the number of access behavior data sequences, and same-frequency node identification is performed to construct a same-frequency node distribution sequence; then, a large model is combined, and a potential attack node distribution and a potential attack probability distribution are predicted and obtained according to the deviation node distribution sequence and the same-frequency node distribution sequence; finally, a node detection strategy is formulated according to the potential attack node distribution and the potential attack probability distribution for network security intrusion detection. That is to say, by constructing an access behavior benchmark comparison table, identifying access deviation nodes and same-frequency collaborative behavior nodes, and combining a large model to predict potential attack nodes and their attack probabilities, the detection ability for unknown attacks and collaborative attacks is effectively improved, the intelligent identification and accurate early warning of network intrusion behaviors are realized, and the accuracy, real-time performance and active defense ability of intrusion detection are significantly improved. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] Figure 1 is a step flow chart of a network security intrusion detection method based on a large model of the present invention; Figure 2 is a schematic structural diagram of a network security intrusion detection system based on a large model of the present invention.

[0016] Description of the reference numerals: Access behavior clustering module 11, deviated node distribution sequence construction module 12, same-frequency node distribution sequence construction module 13, node detection strategy formulation module 14. Detailed implementation

[0017] By providing a network security intrusion detection method and system based on a large model, the present invention solves the problems of insufficient recognition ability of traditional intrusion detection methods in dealing with unknown or mutated attacks, and the lack of effective means to identify the collaborative behavior patterns between multiple access nodes, resulting in low detection accuracy, response lag, and easy missed reporting in the intrusion detection system in a complex network environment. By constructing a reference table for access behavior benchmarks, identifying access deviated nodes and same-frequency collaborative behavior nodes, and combining a large model to predict potential attack nodes and their attack probabilities, the detection ability for unknown attacks and collaborative attacks is effectively improved, intelligent recognition and accurate early warning of network intrusion behaviors are realized, and the accuracy, real-time performance, and active defense ability of intrusion detection are significantly improved.

[0018] Next, the technical solutions in the present invention will be clearly and completely described with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments of the present invention. It should be understood that the present invention is not limited by the example embodiments described herein. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the scope of protection of the present invention. Additionally, it should be noted that for the sake of description, only the parts related to the present invention are shown in the accompanying drawings rather than all of them.

[0019] Example 1, please refer to the attached Figure 1 , the present invention provides a network security intrusion detection method based on a large model, which is applied to a network security intrusion detection system based on a large model, and specifically includes the following steps: S10: According to the network logs of the enterprise internal network in a preset historical time zone, perform access behavior clustering on a number of access nodes to construct a reference table for access behavior benchmarks.

[0020] Further, step S10 of the present invention further includes: S11: Randomly select a first access node, screen the network logs according to the first access node to determine a first access behavior dataset; S12: Divide a plurality of access time periods based on a preset time interval, and divide the first access behavior dataset to obtain a plurality of access behavior datasets.

[0021] Specifically, first, within a preset historical time interval (such as the most recent 30 days), obtain the network log data of the enterprise intranet. The preset historical time zone refers to a historical time range set by the system for obtaining network log data within a certain period in the past. This time zone can be configured according to actual needs, such as being set to the most recent 7 days, 30 days, or 90 days, etc. In this embodiment, "the most recent 30 days" is used as an example to collect complete access behavior patterns through a long enough time window for subsequent modeling and analysis. Network logs refer to various access records, connection requests, data transmission, and other behavior information recorded by enterprise internal network devices (such as switches, servers, etc.). The log content usually includes fields such as timestamp, source IP address, destination IP address, port, protocol type, data traffic, and session duration. The enterprise intranet includes multiple access nodes. An access node refers to a terminal device or server with an independent network identifier, such as an employee computer, printer, business system host, etc. Each node leaves access behavior records in the network, serving as the basic unit for behavior modeling. Then, randomly select any one of the multiple access nodes as the first access node, and use the unique identifier (such as the IP address) of the first access node as a screening condition to filter the obtained network logs, only retaining all access records generated by this node within the preset historical time zone, forming a first access behavior dataset for subsequent behavior modeling and analysis.

[0022] Then, divide each day based on a preset time interval. That is, in order to analyze the distribution law of access behavior in the time dimension more precisely, the entire historical time zone is further divided into multiple consecutive and equal-length time periods. This time interval can be custom-set according to actual business characteristics, such as 30 minutes, 1 hour, 2 hours, 4 hours, etc. The purpose is to more accurately depict the differences in access behavior of nodes in different time periods. For example, divide a day (24 hours) into 24 access time periods with each 1 hour as a unit time interval. Further, based on the above division results, classify the access behavior data of the first access node within the entire historical time zone into the corresponding access time periods according to its timestamp attribute. After the above time division operation, the original first access behavior dataset will be split into several subsets, and each subset corresponds to the access behavior within a time period, obtaining multiple access behavior datasets.

[0023] S13: Perform access behavior clustering based on the multiple access behavior datasets to determine multiple high-frequency access behavior characteristics, where the access behavior characteristics include access frequency, single data traffic, and single duration.

[0024] Furthermore, step S13 of the present invention further includes: S131: Randomly select the first access behavior dataset, perform clustering using the K-means algorithm to obtain multiple first clustering clusters and multiple first central feature values; S132: Count the data volumes of the multiple first clustering clusters, and select the first clustering clusters with the largest data volumes in a preset proportion as high-frequency clustering clusters, determine multiple high-frequency clustering clusters, and map to obtain multiple high-frequency central feature values; S133: Calculate the mean value of the multiple high-frequency central feature values to obtain the first high-frequency access behavior feature, and sequentially analyze to obtain multiple high-frequency access behavior features.

[0025] Specifically, first, randomly select any dataset from multiple access behavior datasets as the first access behavior dataset; then, use the K-means algorithm to cluster the first access behavior dataset. First, randomly initialize several clustering center points, and then assign each data sample to the clustering cluster to which the nearest center point belongs; then, recalculate the mean value of all samples in each clustering cluster as the new center point, and repeat the above steps until the clustering result converges or reaches the set number of iterations, thereby obtaining multiple access behavior clustering clusters and their corresponding central feature values. Among them, each behavior sample is represented by a set of access behavior feature vectors (such as access frequency, single-data traffic, single-duration), and the goal of clustering is to minimize the distance between samples within the same cluster and the center point, resulting in several clustering clusters with similar behavior characteristics.

[0026] Next, count the data volumes of the multiple first clustering clusters. According to the preset proportion (for example, select the clustering clusters with the top 20% or 30% of the data volumes), filter out the clustering clusters with larger data volumes and label them as high-frequency clustering clusters; obtain the corresponding center point feature values of these high-frequency clustering clusters, called high-frequency central feature values, which represent the access patterns that most frequently appear in the historical behavior of the node. Then, calculate the mean value of the multiple high-frequency central feature values according to the dimensions respectively to obtain the first high-frequency access behavior feature, that is, the typical access behavior portrait reflecting the first access node in this time period. Repeat this analysis process to sequentially extract multiple high-frequency access behavior features in multiple time periods. Among them, the access behavior features include access frequency, single-data traffic, and single-duration. The access frequency refers to the number of connection or access requests initiated in a certain time period, the single-data traffic refers to the data volume transmitted in each access, and the single-duration refers to the time length of each connection maintenance. By using K-means clustering for access behavior pattern recognition, it can effectively strip noise data, highlight high-frequency and stable behavior features, and then establish an accurate access behavior portrait for each node.

[0027] S14: Establish the first access behavior benchmark of the first access node based on the multiple access time periods and multiple high-frequency access behavior features, and sequentially analyze to determine several access behavior benchmarks to construct an access behavior benchmark comparison table.

[0028] Specifically, first, a first access behavior benchmark for the first access node is established according to the multiple access time periods and multiple high-frequency access behavior characteristics. The access behavior benchmark refers to modeling the historical access behavior characteristics of a node in different time periods to form the normal behavior pattern of the node, which is used to determine whether the current access behavior deviates from the normal pattern in the subsequent process and is an important basis for judging abnormalities in intrusion detection. Using the same method for obtaining the first access behavior benchmark, the access behavior benchmarks of several access nodes are sequentially analyzed and obtained, and an access behavior benchmark comparison table is constructed according to the combination of several access behavior benchmarks.

[0029] S20: Monitor and obtain several access behavior data sequences of the several access nodes in the current time period, and use the access behavior benchmark comparison table to perform node difference comparison on the several access behavior data sequences to construct a deviation node distribution sequence.

[0030] Furthermore, step S20 of the present invention further includes: S21: Use the access behavior benchmark comparison table to match and obtain several adapted access behavior benchmarks of the several access nodes in the current time period.

[0031] Specifically, monitor and obtain several access behavior data sequences of the several access nodes in the current time period. The current time period refers to the real-time time window when the intrusion detection task is in progress, usually consistent with the time interval used for the historical behavior benchmark (for example: the current 1 hour, the current half hour, etc.), which ensures the comparability between the current access behavior and the historical benchmark; in the current time period, perform network behavior monitoring on all or part of the key monitored access nodes in the enterprise internal network, collect the access request logs sent by each node, generate structured behavior data records, and obtain several access behavior data sequences.

[0032] Then, use the access behavior benchmark comparison table to match and obtain several adapted access behavior benchmarks of the several access nodes in the current time period, that is, select a set of historical behavior characteristics that most conforms to the current time period and the current node behavior from the access behavior benchmark comparison table as the comparison template for the current behavior.

[0033] S22: Based on the several adapted access behavior benchmarks, sequentially perform node difference comparison on the several access behavior data at the same time point in the several access behavior data sequences to determine multiple deviation node distributions and construct a deviation node distribution sequence.

[0034] Furthermore, step S22 of the present invention further includes: S221: Calculate the deviation amplitude between the adapted access behavior benchmark and the corresponding access behavior data respectively, obtain the access frequency deviation ratio, the single - data - flow deviation ratio, and the single - duration deviation ratio, perform weighted calculation to obtain the node deviation value, and analyze in sequence to obtain the deviation values of several access nodes; S222: Obtain the node position distribution of several access nodes, use the several deviation values to perform mapping identification on the node position distribution, and obtain the deviation node distribution.

[0035] Specifically, first, for each access node within the current time period, perform deviation analysis based on its current access behavior data and its corresponding adapted access behavior benchmark. Calculate the deviation amplitude between the adapted access behavior benchmark and the corresponding access behavior data respectively. The deviation amplitude refers to the ratio of the difference between the two to the benchmark value. Obtain the access frequency deviation ratio, the single - data - flow deviation ratio, and the single - duration deviation ratio. For example, assume the current access frequency is 5 times per minute and the benchmark frequency is 4 times per minute, then the frequency deviation amplitude is (5 - 4) / 4, which is equal to 25%. Further configure the weight ratios of access frequency, single - data - flow, and single - duration, which can be set according to experience or optimized through training. Perform weighted calculation on the access frequency deviation ratio, the single - data - flow deviation ratio, and the single - duration deviation ratio to obtain the node deviation value, and analyze in sequence to obtain the deviation values of several access nodes.

[0036] Then obtain the node position distribution of several access nodes. The topological position or logical coordinates of each access node within the enterprise intranet can be determined according to methods such as IP address segments, subnet locations, device numbers, physical locations, or logical service groupings. Then use the several deviation values to perform mapping identification on the node position distribution, obtain the deviation node distribution, and analyze in sequence to obtain multiple deviation node distributions, constructing a deviation node distribution sequence.

[0037] S30: Perform node similarity comparison on the several access behavior data sequences, and perform same - frequency node identification to construct a same - frequency node distribution sequence.

[0038] Furthermore, step S30 of the present invention further includes: S31: According to a predetermined similarity threshold, perform similarity clustering on the several access behavior data at the same time point in the several access behavior data sequences, and perform same - frequency node identification on the access nodes corresponding to the access behavior data that meet the predetermined similarity threshold to determine a same - frequency node set sequence; S32: According to the same - frequency node set sequence, respectively perform identification on the node position distributions of several access nodes to obtain a same - frequency node distribution sequence.

[0039] Specifically, first, at the same time point, the behavior data of all access nodes is regarded as a feature vector. Pairwise comparison is performed among all nodes using similarity metrics (such as cosine similarity, Euclidean distance, etc.), and a group of nodes that meet a predetermined similarity threshold (which can be set according to the actual scenario, such as 70%) are classified into the same co-frequency node set, indicating that they exhibit consistent or collaborative access behavior characteristics at this time point. Co-frequency node identification is carried out, and the above operations are repeated for each time point to obtain a sequence of co-frequency node sets evolving over time. Then, according to the sequence of co-frequency node sets, the node position distributions of several access nodes are respectively identified. Each time point corresponds to a co-frequency node distribution structure diagram, obtaining a co-frequency node distribution sequence, where the co-frequency node distribution sequence reflects the evolution law of potential attack behaviors in the cyber space or logically. By constructing the co-frequency node distribution sequence, abnormal collaborative behavior characteristics among multiple access nodes (such as group attacks, botnet synchronization behaviors, etc.) can be effectively identified, providing structured input data for subsequent prediction of collaborative attack paths and identification of potential attack nodes through large models.

[0040] S40: Combine with a large model, predict and obtain the potential attack node distribution and potential attack probability distribution according to the deviated node distribution sequence and the co-frequency node distribution sequence, and formulate a node detection strategy for network security intrusion detection.

[0041] Furthermore, step S40 of the present invention further includes: S41: According to the network logs of similar enterprises, screen sample data to obtain a sample deviated node distribution sequence set and a sample co-frequency node distribution sequence set, and analyze to obtain the high-frequency attack node distribution and high-frequency attack probability distribution under different sample deviated node distribution sequences and sample co-frequency node distribution sequences, obtaining a sample attack node distribution set and a sample attack probability distribution set; S42: Use the sample deviated node distribution sequence set and the sample co-frequency node distribution sequence set as inputs, and use the sample attack node distribution set and the sample attack probability distribution set as supervision to train a large model to construct a network attack recognition plug-in; S43: Use the network attack recognition plug-in to predict and obtain the potential attack node distribution and potential attack probability distribution according to the deviated node distribution sequence and the co-frequency node distribution sequence.

[0042] Specifically, first, collect the historical network log data of enterprises with similar business scenarios, network structures, or asset scales. Each log sample contains access behavior data, node access time series, labeled attack events, or alarm records, which serve as the network logs of similar enterprises. Then, based on the network logs of similar enterprises, conduct sample data screening to obtain the sample deviated node distribution sequence set and the sample co-frequency node distribution sequence set; further analyze to obtain the high-frequency attack node distribution and high-frequency attack probability distribution under different sample deviated node distribution sequences and sample co-frequency node distribution sequences, and obtain the sample attack node distribution set and the sample attack probability distribution set. Among them, the sample attack node distribution refers to the true distribution of attack nodes in each sample, and the sample attack probability refers to the recorded or estimated value of the attack probability of nodes at different time periods.

[0043] Next, use the sample deviated node distribution sequence set and the sample co-frequency node distribution sequence set as inputs, and use the sample attack node distribution set and the sample attack probability distribution set as supervision to train the large model. First, use the sample deviated node distribution sequence and the co-frequency node distribution sequence as model inputs, and the model performs forward propagation to output the prediction of the attack node distribution and attack probability; then, calculate the error between the prediction result and the true label through a loss function (such as cross-entropy, mean squared error, etc.); then use the backpropagation algorithm and an optimizer (such as Adam) to adjust the model parameters to reduce the prediction error; finally, periodically evaluate the model performance on the validation set and adjust the training strategy (such as learning rate, regularization) to prevent overfitting. After the training is completed, save the model parameters with the best performance, package them as a network attack recognition plug-in module, and integrate the trained recognition plug-in into the network security monitoring platform to achieve real-time prediction of potential attack nodes and their attack probabilities.

[0044] Finally, use the trained network attack recognition plug-in to input the deviated node distribution sequence and the co-frequency node distribution sequence monitored in the current network environment. Based on the attack pattern learning ability built into the large model, the plug-in predicts the spatial distribution of potential attack nodes and the corresponding attack probability distribution, so as to achieve real-time identification and early warning of network intrusion behavior.

[0045] Furthermore, step S40 of the present invention further includes: S44: Perform risk identification on potential attack nodes that meet the preset probability threshold to determine the potential risk node distribution; S45: Conduct a concentration analysis of the position distribution of the potential risk node distribution to determine the risk node distribution concentration; S46: If the risk node distribution concentration is less than or equal to the preset concentration index, formulate the node detection strategy as local detection, where local detection is to perform network security intrusion detection on risk nodes in descending order of potential attack probability.

[0046] Specifically, first, according to the attack probability of each node predicted by the network attack recognition plugin, it is compared with a preset probability threshold (which can be set according to the actual scenario, such as 50%); for nodes with an attack probability higher than this threshold, the system marks them as "potential risk nodes", and these marked potential risk nodes form a distribution of potential risk nodes in terms of network topology or physical location, reflecting the set of nodes that may be attacked currently and their distribution status. Then, a concentration analysis is performed on the spatial distribution of the above-mentioned potential risk nodes, that is, to evaluate whether these risk nodes are aggregated or relatively dispersed in certain regions of the network. The analysis indicators may include the physical distance between nodes, the network topology distance, or statistical distribution concentration indicators (such as clustering coefficient, density, etc.). By calculation, a numerical value of the risk node distribution concentration is obtained, which quantifies the "aggregation degree" of the risk nodes. Further, the risk node distribution concentration is compared with the preset concentration index threshold. If the risk node distribution concentration is less than or equal to this threshold, it indicates that the risk nodes are relatively dispersed and not highly concentrated. Based on this, a local detection strategy is formulated, that is, targeted detection is only carried out on high-risk nodes. The specific approach is to sort the potential risk nodes in descending order according to the potential attack probability, and give priority to detecting the node with the highest attack probability, and gradually conduct in-depth network security intrusion detection. This strategy can concentrate limited security resources, improve the detection efficiency and response speed, avoid performing undifferentiated detection on the entire network, and save computing and labor costs.

[0047] Furthermore, step S40 of the present invention further includes: S47: If the risk node distribution concentration is greater than the preset concentration index, formulate the node detection strategy as global detection, where global detection is to perform network security intrusion detection on all nodes in descending order according to the potential attack probability.

[0048] Specifically, when the risk node distribution concentration is greater than the preset concentration index threshold, it indicates that the potential risk nodes show an obvious aggregation or concentration trend in the network. Due to the high concentration of risk nodes, there may be a possibility of cross-node propagation or coordinated attacks of attack risks. Simple local detection may not be able to fully cover potential threats. Therefore, a global detection strategy is adopted to comprehensively scan and detect all access nodes in the network. Global detection is sorted from high to low according to the potential attack probability, and network security intrusion detection is performed on all nodes one by one to ensure the identification of potential attack behaviors without omission. By using the global detection strategy to check all nodes one by one, it is possible to effectively prevent cluster attacks and security threats that spread rapidly, and enhance the comprehensiveness and depth of network security defense.

[0049] In summary, a network security intrusion detection method based on a large model provided by the present invention has the following technical effects: By clustering the access behaviors of a number of access nodes based on the network logs of the enterprise intranet in a preset historical time zone, an access behavior benchmark comparison table is constructed; then, a number of access behavior data sequences of the number of access nodes in the current period are monitored and obtained, and the access behavior benchmark comparison table is used to perform node difference comparison on the number of access behavior data sequences to construct a deviation node distribution sequence; on the other hand, node similarity comparison is performed on the number of access behavior data sequences, and co-frequency node identification is performed to construct a co-frequency node distribution sequence; then, in combination with a large model, based on the deviation node distribution sequence and the co-frequency node distribution sequence, a potential attack node distribution and a potential attack probability distribution are predicted and obtained; finally, a node detection strategy is formulated based on the potential attack node distribution and the potential attack probability distribution for network security intrusion detection. That is to say, by constructing an access behavior benchmark comparison table, access deviation nodes and co-frequency collaborative behavior nodes are identified, and in combination with a large model, potential attack nodes and their attack probabilities are predicted, effectively improving the detection ability for unknown attacks and collaborative attacks, realizing the intelligent identification and precise early warning of network intrusion behaviors, and significantly improving the accuracy, real-time performance, and active defense ability of intrusion detection.

[0050] Embodiment 2. Based on the same inventive concept as the network security intrusion detection method based on a large model in the foregoing embodiment, the present invention also provides a network security intrusion detection system based on a large model. Please refer to the attached Figure 2 , including: an access behavior clustering module 11, configured to cluster the access behaviors of a number of access nodes according to the network logs of the enterprise intranet in a preset historical time zone, and construct an access behavior benchmark comparison table; a deviation node distribution sequence construction module 12, configured to monitor and obtain a number of access behavior data sequences of the number of access nodes in the current period, and use the access behavior benchmark comparison table to perform node difference comparison on the number of access behavior data sequences to construct a deviation node distribution sequence; a co-frequency node distribution sequence construction module 13, configured to perform node similarity comparison on the number of access behavior data sequences, and perform co-frequency node identification to construct a co-frequency node distribution sequence; a node detection strategy formulation module 14, configured to combine a large model, and based on the deviation node distribution sequence and the co-frequency node distribution sequence, predict and obtain a potential attack node distribution and a potential attack probability distribution, and formulate a node detection strategy for network security intrusion detection.

[0051] Further, the large model-based network security intrusion detection system is also used for: randomly selecting a first access node, screening the network logs according to the first access node to determine a first access behavior dataset; dividing the first access behavior dataset based on a preset time interval to obtain multiple access behavior datasets; performing access behavior clustering according to the multiple access behavior datasets to determine multiple high-frequency access behavior characteristics, where the access behavior characteristics include access frequency, single data traffic, and single duration; establishing a first access behavior benchmark for the first access node according to the multiple access time periods and multiple high-frequency access behavior characteristics, and sequentially analyzing and determining several access behavior benchmarks to construct an access behavior benchmark comparison table.

[0052] Further, the large model-based network security intrusion detection system is also used for: randomly selecting a first access behavior dataset, performing clustering using the K-means algorithm to obtain multiple first clustering clusters and multiple first central feature values; counting the data volume of the multiple first clustering clusters, and selecting the first clustering clusters with the largest data volume at a preset ratio as high-frequency clustering clusters to determine multiple high-frequency clustering clusters, and mapping to obtain multiple high-frequency central feature values; calculating the mean value of the multiple high-frequency central feature values to obtain a first high-frequency access behavior characteristic, and sequentially analyzing to obtain multiple high-frequency access behavior characteristics.

[0053] Further, the large model-based network security intrusion detection system is also used for: using the access behavior benchmark comparison table to match and obtain several adapted access behavior benchmarks for several access nodes in the current time period; based on the several adapted access behavior benchmarks, sequentially comparing the several access behavior data at the same time point in the several access behavior data sequences to determine multiple deviation node distributions, and constructing a deviation node distribution sequence.

[0054] Further, the large model-based network security intrusion detection system is also used for: respectively calculating the deviation amplitude between the adapted access behavior benchmark and the corresponding access behavior data to obtain an access frequency deviation ratio, a single data traffic deviation ratio, and a single duration deviation ratio, and performing weighted calculation to obtain a node deviation value, and sequentially analyzing to obtain several deviation values for several access nodes; obtaining the node position distribution of several access nodes, and using the several deviation values to perform mapping and identification on the node position distribution to obtain a deviation node distribution.

[0055] Furthermore, the large model-based network security intrusion detection system is also used for: performing similarity clustering on several access behavior data at the same time point in the several access behavior data sequences according to a predetermined similarity threshold, and performing co-frequency node identification on the access nodes corresponding to the access behavior data that meet the predetermined similarity threshold to determine a co-frequency node set sequence; respectively identifying the node position distributions of several access nodes according to the co-frequency node set sequence to obtain a co-frequency node distribution sequence.

[0056] Furthermore, the large model-based network security intrusion detection system is also used for: screening sample data according to the network logs of similar enterprises, obtaining a sample deviation node distribution sequence set and a sample co-frequency node distribution sequence set, and analyzing to obtain a high-frequency attack node distribution and a high-frequency attack probability distribution under different sample deviation node distribution sequences and sample co-frequency node distribution sequences, obtaining a sample attack node distribution set and a sample attack probability distribution set; using the sample deviation node distribution sequence set and the sample co-frequency node distribution sequence set as inputs, using the sample attack node distribution set and the sample attack probability distribution set as supervision, training a large model to construct a network attack recognition plug-in; using the network attack recognition plug-in to predict and obtain a potential attack node distribution and a potential attack probability distribution according to the deviation node distribution sequence and the co-frequency node distribution sequence.

[0057] Furthermore, the large model-based network security intrusion detection system is also used for: performing risk identification on potential attack nodes that meet a preset probability threshold to determine a potential risk node distribution; performing a concentration analysis of the position distribution of the potential risk node distribution to determine a risk node distribution concentration; if the risk node distribution concentration is less than or equal to a preset concentration index, formulating a node detection strategy as local detection, where local detection is to perform network security intrusion detection on risk nodes in descending order of potential attack probability.

[0058] Furthermore, the large model-based network security intrusion detection system is also used for: if the risk node distribution concentration is greater than a preset concentration index, formulating a node detection strategy as global detection, where global detection is to perform network security intrusion detection on all nodes in descending order of potential attack probability.

[0059] In the present specification, the various embodiments are described in a progressive manner. The key point of each embodiment is to illustrate the differences from other embodiments. The method and specific examples of a large model-based network security intrusion detection method in the foregoing Embodiment 1 are equally applicable to a large model-based network security intrusion detection system in this embodiment. Through the foregoing detailed description of a large model-based network security intrusion detection method, those skilled in the art can clearly understand the large model-based network security intrusion detection system in this embodiment. Therefore, for the sake of brevity of the specification, it will not be elaborated herein. For the system disclosed in the embodiment, since it corresponds to the method disclosed in the embodiment, the description is relatively simple. For related parts, reference may be made to the description in the method part.

[0060] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present invention. Various modifications to these embodiments will be obvious to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to the embodiments shown herein, but will be accorded the widest scope consistent with the principles and novel features disclosed herein.

[0061] Obviously, those skilled in the art can make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the present invention and its equivalent technologies, the present invention is also intended to include these changes and modifications.

Claims

1. A network security intrusion detection method based on a large model, characterized in that, The method includes: Performing access behavior clustering on a number of access nodes according to the network logs of the enterprise intranet in a preset historical time zone, and constructing an access behavior benchmark comparison table; Monitoring and obtaining a number of access behavior data sequences of the number of access nodes in the current period, using the access behavior benchmark comparison table to perform node difference comparison on the number of access behavior data sequences, and constructing a deviation node distribution sequence; Performing node similarity comparison on the number of access behavior data sequences, and performing co-frequency node identification to construct a co-frequency node distribution sequence; Combined with a large model, predicting and obtaining a potential attack node distribution and a potential attack probability distribution according to the deviation node distribution sequence and the co-frequency node distribution sequence, and formulating a node detection strategy for network security intrusion detection.

2. The network security intrusion detection method based on a large model according to claim 1, characterized in that Performing access behavior clustering on a number of access nodes according to the network logs of the enterprise intranet in a preset historical time zone, and constructing an access behavior benchmark comparison table, including: Randomly selecting a first access node, screening the network logs according to the first access node, and determining a first access behavior data set; Dividing a plurality of access time periods based on a preset time interval, and dividing the first access behavior data set to obtain a plurality of access behavior data sets; Performing access behavior clustering according to the plurality of access behavior data sets, and determining a plurality of high-frequency access behavior characteristics, where the access behavior characteristics include access frequency, single data traffic, and single duration; Establishing a first access behavior benchmark of the first access node according to the plurality of access time periods and the plurality of high-frequency access behavior characteristics, and sequentially analyzing and determining a number of access behavior benchmarks to construct an access behavior benchmark comparison table.

3. The network security intrusion detection method based on a large model according to claim 2, characterized in that, Performing access behavior clustering according to the plurality of access behavior data sets, and determining a plurality of high-frequency access behavior characteristics, including: Randomly selecting a first access behavior data set, performing clustering using the K-means algorithm to obtain a plurality of first clustering clusters and a plurality of first central feature values; Counting the data volumes of the plurality of first clustering clusters, and selecting the first clustering clusters with the largest data volume in a preset proportion as high-frequency clustering clusters, determining a plurality of high-frequency clustering clusters, and mapping to obtain a plurality of high-frequency central feature values; Calculating the mean value of the plurality of high-frequency central feature values to obtain a first high-frequency access behavior characteristic, and sequentially analyzing to obtain a plurality of high-frequency access behavior characteristics.

4. The network security intrusion detection method based on a large model according to claim 1, wherein, Using the access behavior benchmark comparison table to perform node difference comparison on the number of access behavior data sequences, and constructing a deviation node distribution sequence, including: Using the access behavior benchmark comparison table to match and obtain a number of adapted access behavior benchmarks of the number of access nodes in the current period; Based on the number of adapted access behavior benchmarks, sequentially performing node difference comparison on the number of access behavior data at the same time point in the number of access behavior data sequences, determining a number of deviation node distributions, and constructing a deviation node distribution sequence.

5. The network security intrusion detection method based on a large model according to claim 4, characterized in that Calculate the deviation amplitude between the adapted access behavior benchmark and the corresponding access behavior data respectively, obtain the access frequency deviation ratio, the single - data - flow deviation ratio, and the single - duration deviation ratio, and perform weighted calculation to obtain the node deviation value. Analyze and obtain the deviation values of several access nodes in sequence. Obtain the node position distribution of several access nodes, and use the several deviation values to map and identify the node position distribution to obtain the deviation node distribution.

6. The network security intrusion detection method based on a large model according to claim 1, characterized in that, Perform node similarity comparison on the several access behavior data sequences and perform same - frequency node identification to construct a same - frequency node distribution sequence, including: According to a predetermined similarity threshold, perform similarity clustering on several access behavior data at the same time point in the several access behavior data sequences, and perform same - frequency node identification on the access nodes corresponding to the access behavior data that meet the predetermined similarity threshold to determine the same - frequency node set sequence. According to the same - frequency node set sequence, respectively identify the node position distributions of several access nodes to obtain the same - frequency node distribution sequence.

7. The network security intrusion detection method based on a large model according to claim 1, characterized in that, Combined with the large model, predict and obtain the potential attack node distribution and the potential attack probability distribution according to the deviation node distribution sequence and the same - frequency node distribution sequence, including: According to the network logs of similar enterprises, perform sample data screening to obtain a sample deviation node distribution sequence set and a sample same - frequency node distribution sequence set, and analyze to obtain the high - frequency attack node distribution and the high - frequency attack probability distribution under different sample deviation node distribution sequences and sample same - frequency node distribution sequences, and obtain the sample attack node distribution set and the sample attack probability distribution set. Use the sample deviation node distribution sequence set and the sample same - frequency node distribution sequence set as inputs, and use the sample attack node distribution set and the sample attack probability distribution set as supervision to train the large model and construct a network attack recognition plugin. Use the network attack recognition plugin to predict and obtain the potential attack node distribution and the potential attack probability distribution according to the deviation node distribution sequence and the same - frequency node distribution sequence.

8. The network security intrusion detection method based on a large model according to claim 1, wherein The formulating of the node detection strategy includes: Perform risk identification on potential attack nodes that meet the preset probability threshold to determine the potential risk node distribution. Perform an analysis of the concentration of the position distribution of the potential risk node distribution to determine the risk node distribution concentration. If the risk node distribution concentration is less than or equal to the preset concentration index, formulate the node detection strategy as local detection, where local detection is to perform network security intrusion detection on risk nodes in descending order of potential attack probability.

9. The network security intrusion detection method based on a large model according to claim 8, wherein If the risk node distribution concentration is greater than the preset concentration index, formulate the node detection strategy as global detection, where global detection is to perform network security intrusion detection on all nodes in descending order of potential attack probability.

10. A network security intrusion detection system based on a large model, characterized in that, Steps for implementing the network security intrusion detection method based on a large model according to any one of claims 1 to 9, including: An access behavior clustering module, configured to perform access behavior clustering on several access nodes according to the network logs of the enterprise internal network in a preset historical time zone, and construct an access behavior benchmark comparison table. The deviation node distribution sequence construction module is used to monitor and obtain a plurality of access behavior data sequences of the plurality of access nodes during the current period, and use the access behavior benchmark comparison table to perform node difference comparison on the plurality of access behavior data sequences to construct a deviation node distribution sequence; The co-frequency node distribution sequence construction module is used to perform node similarity comparison on the plurality of access behavior data sequences, and perform co-frequency node identification to construct a co-frequency node distribution sequence; The node detection strategy formulation module is used to combine with a large model, predict and obtain the potential attack node distribution and potential attack probability distribution according to the deviation node distribution sequence and the co-frequency node distribution sequence, and formulate a node detection strategy for network security intrusion detection.

Citation Information

Patent Citations

  • Gateway access abnormity monitoring and early warning method and system

    CN119363481A

  • Large-scale network security defense system based on collaborative intrusion detection

    CN119544381A

  • Software supply chain threat monitoring method and device based on large model

    CN119557881A

  • Large model intrusion detection method and device based on user behavior data

    CN119675900A

  • Anomalous behavior detection in processor based systems

    US20200293657A1

Cited By

  • Intrusion detection system based on big data analysis

    CN120811756A