Network vulnerability identification method, system and equipment based on data analysis

By using high-dimensional vectors that integrate temporal and spatial features to identify network vulnerabilities, and combining graph neural networks and extreme gradient boosting models, the problem of low identification accuracy in traditional methods is solved, achieving efficient and accurate identification of new attacks and unknown vulnerabilities.

CN120880747APending Publication Date: 2025-10-31ELECTRIC POWER RESEARCH INSTITUTE OF STATE GRID NINGXIA ELECTRIC POWER COMPANY
View PDF 0 Cites 4 Cited by

Patent Information

Application Number
CN202511087718.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-05
Publication Date
2025-10-31

AI Technical Summary

Technical Problem

Traditional network vulnerability identification methods rely on feature libraries of known vulnerabilities, which makes it difficult to effectively identify new attacks and unknown vulnerabilities, resulting in low identification accuracy and frequent false negatives.

Method used

By extracting the temporal and spatial features of network data and fusing them into a high-dimensional feature vector, and combining graph neural networks and extreme gradient boosting models for vulnerability identification, vulnerabilities are processed in a hierarchical manner and their authenticity is verified and correlation is analyzed to ensure accurate vulnerability identification.

Benefits of technology

It improves the accuracy and efficiency of network vulnerability identification, effectively identifies new attacks and unknown vulnerabilities, reduces the false negative rate, and ensures network security.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120880747A_ABST
    Figure CN120880747A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of information network security, in particular to a network vulnerability identification method, system and equipment based on data analysis, and the method comprises the steps: extracting time sequence features and spatial features of network data, and fusing the time sequence features and the spatial features to obtain a high-dimensional feature vector; determining the type and confidence of the network vulnerability according to the high-dimensional feature vector; determining the hierarchy to which the network vulnerability belongs based on the type and confidence of the network vulnerability; performing authenticity verification and correlation analysis on the network vulnerabilities of the first level and the second level to determine the authenticity and attack chain of the network vulnerabilities; and identifying the network vulnerability of the third level again, and marking the type to which the network vulnerability of the third level belongs according to an identification result. Therefore, the identification precision of the network vulnerability is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of information network security technology, specifically to a network vulnerability identification method, system, and device based on data analysis. Background Technology

[0002] With the rapid iteration of internet technology, the deep integration of technologies such as cloud computing and the Internet of Things has driven the exponential expansion of network scale, making network architecture increasingly complex. However, behind this rapid development, the forms of cyberattacks are also constantly evolving, from early simple virus infections to today's diversified and organized new threats such as APT attacks, ransomware, and supply chain attacks.

[0003] Against this backdrop, traditional network vulnerability identification methods rely on rule-based feature matching or signature scanning techniques. However, these techniques heavily depend on known vulnerability signature databases and lack effective identification capabilities for novel attacks and unknown vulnerabilities. Attackers can bypass signature detection with minor modifications to their attack code, leading to frequent false negatives in the detection system. Furthermore, unidentified vulnerabilities are permanently discarded, causing similar attacks to recur. Therefore, the accuracy of network vulnerability identification using these methods is relatively low. Summary of the Invention

[0004] To address the technical problem of low accuracy in identifying network vulnerabilities, the present invention aims to provide a network vulnerability identification method, system, and device based on data analysis. The specific technical solution adopted is as follows:

[0005] In a first aspect, embodiments of the present invention disclose a network vulnerability identification method based on data analysis, comprising: extracting and fusing temporal and spatial features of network data to obtain a high-dimensional feature vector; determining the type and confidence level of the network vulnerability based on the high-dimensional feature vector; determining the level to which the network vulnerability belongs based on the type and confidence level of the network vulnerability, the level including a first level, a second level, and a third level, wherein the first and second levels indicate that the network vulnerability is identified, and the third level indicates that the network vulnerability is unidentified, the confidence level of the first level is greater than the confidence level of the second level, and the confidence level of the second level is greater than the confidence level of the third level; performing authenticity verification and correlation analysis on the network vulnerabilities of the first and second levels to determine the authenticity and attack chain of the network vulnerabilities; and identifying the network vulnerabilities of the third level again, and labeling the type to which the network vulnerabilities of the third level belong based on the identification results.

[0006] Secondly, embodiments of the present invention disclose a network vulnerability identification system based on data analysis, comprising: an extraction module for extracting and fusing temporal and spatial features of network data to obtain a high-dimensional feature vector; a determination module for determining the type and confidence level of a network vulnerability based on the high-dimensional feature vector; the determination module is further configured to determine the level to which a network vulnerability belongs based on its type and confidence level, the levels including a first level, a second level, and a third level, wherein the first and second levels indicate that the network vulnerability is identified, and the third level indicates that the network vulnerability is unidentified, the confidence level of the first level is greater than that of the second level, and the confidence level of the second level is greater than that of the third level; an analysis module for performing authenticity verification and correlation analysis on the network vulnerabilities of the first and second levels to determine the authenticity and attack chain of the network vulnerabilities; and an identification module for re-identifying the network vulnerabilities of the third level and labeling the type to which the network vulnerabilities of the third level belong based on the identification results.

[0007] Thirdly, embodiments of the present invention disclose an electronic device, including: a processor and a memory; wherein the memory is used to store a computer program that can run on the processor; the processor is used to execute the program stored in the memory to implement the steps of the network vulnerability identification method based on data analysis mentioned in the first aspect.

[0008] The technical solution disclosed in this invention extracts and fuses temporal and spatial features of network data to obtain a high-dimensional feature vector, which can comprehensively capture key information in the network data. This fusion avoids the limitations of a single feature dimension, providing a more comprehensive basis for subsequent vulnerability identification and helping to improve the accuracy of vulnerability identification. The type and confidence level of the network vulnerability are determined based on the high-dimensional feature vector, achieving a precise preliminary judgment of the vulnerability. The high-dimensional feature vector contains a large amount of vulnerability-related information, enabling effective identification of new attacks and unknown vulnerabilities, improving the accuracy of identifying various types of network vulnerabilities. The confidence level classifies various vulnerabilities. The first and second levels clearly define vulnerabilities in an identified state, with the first level having a higher confidence level than the second, helping staff prioritize high-confidence vulnerabilities, rationally allocate resources, and improve the efficiency of vulnerability handling. The third level defines unidentified vulnerabilities, avoiding the risk of omission and ensuring that all possible vulnerabilities are addressed. Subsequently, the authenticity verification and correlation analysis of the network vulnerabilities in the first and second levels further ensure the reliability of vulnerability information and uncover attack chains between vulnerabilities. Authenticity verification can eliminate false positives and further improve the accuracy of vulnerability identification; attack chains help determine the overall network security posture, enabling targeted defensive measures to prevent the spread of attacks. Re-identifying and labeling third-level network vulnerabilities improves the completeness and accuracy of vulnerability identification. This avoids situations where vulnerabilities are overlooked due to a single misidentification; by re-identifying, unidentified vulnerabilities are converted to identified status as much as possible, preventing the recurrence of similar attacks and reducing the risk of network attacks. Attached Figure Description

[0009] Figure 1 This is a flowchart illustrating a network vulnerability identification method based on data analysis, provided as an embodiment of the present invention.

[0010] Figure 2 This is a schematic diagram of a network vulnerability identification system based on data analysis, provided as an embodiment of the present invention.

[0011] Figure 3 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present invention. Detailed Implementation

[0012] To further illustrate the technical means and effects adopted by the present invention to achieve its intended purpose, the following, in conjunction with the accompanying drawings and preferred embodiments, details the specific implementation, structure, features, and effects of a network vulnerability identification method, system, and device based on data analysis proposed according to the present invention. In the following description, different "one embodiment" or "another embodiment" do not necessarily refer to the same embodiment. Furthermore, specific features, structures, or characteristics in one or more embodiments can be combined in any suitable form.

[0013] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains. The specific solutions of a data analysis-based network vulnerability identification method, system, and device provided by this invention are described below in conjunction with the accompanying drawings.

[0014] like Figure 1 As shown, Figure 1 This invention provides a flowchart illustrating a network vulnerability identification method based on data analysis, which includes the following steps:

[0015] Step S101: Extract the temporal and spatial features of the network data and fuse them to obtain a high-dimensional feature vector.

[0016] Specifically, in this embodiment of the invention, network data is first collected, including but not limited to network traffic data, system logs, and asset configuration information. Network traffic data includes, but is not limited to, PCAP format and real-time packet capture data; system logs include, but are not limited to, Syslog / JSON, login records, and abnormal behavior; asset configuration information includes, but is not limited to, Extensible Markup Language (XML) / databases, such as Internet Protocol (IP), ports, and service types. Then, the network data is preprocessed, including alignment and standardization through dynamic time window slicing, cleaning outlier data, filling missing values, improving data quality, and obtaining structured data. Standardization refers to converting data of different formats into a unified Apache Avro format to ensure consistency in subsequent module processing. Cleaning outlier data can be done using an improved Z-score, and filling missing values ​​can be done using the mean of neighboring time windows. Alignment can be done by unifying timestamps, outputting structured data.

[0017] Furthermore, when capturing network traffic data in real time and synchronizing logs and asset information, the length of the data acquisition window can be dynamically adjusted by time window slicing based on network traffic bursts. Lightweight probes (CPU usage <5%) are deployed to capture network traffic data in real time, and logs and asset information are synchronized via an Application Programming Interface (API). This is achieved through W... t ={x i |t-δ≤t i Time window slicing is performed within the range ≤t} to avoid data redundancy. Where W... t Let x be the data set in the current time data acquisition window, x be the i-th data point, t be the current time, and δ be the length of the current time data acquisition window. i Let be the timestamp of the i-th data point. The time window slicing method is as follows: data points with timestamps in the range [t-δ, t] are grouped into the same window. Redundancy is avoided by dynamically adjusting δ (based on traffic bursts); for example, δ is increased during traffic bursts and decreased during stable periods. More specifically, the length of the data acquisition window is adjusted according to traffic bursts using the following formula:

[0018]

[0019] Where, δ t δ0 is the data acquisition window length (in seconds) for the current time, calculated in real time; δ0 is the baseline window length (default 5 seconds), which is a system preset parameter. For bursty traffic gradients, where, The burst factor B can be obtained from the traffic flow at past scheduled times (e.g., 1 hour in the past). t The rate of change is expressed as, σ 流量 The standard deviation of traffic data for past scheduled times, μ 流量 B is the average of traffic flow data over a predetermined period of time. avg This is the historical average traffic burst coefficient (average traffic burst coefficient over the past 1 hour). Therefore, dynamically adjusting the length of the data collection window based on traffic bursts avoids data redundancy and improves the quality of the collected data.

[0020] Furthermore, when cleaning outlier data, the degree of deviation of data points can be determined by calculating the median and median absolute deviation of the data distribution, and outliers can be removed based on the degree of deviation. The specific process is as follows: Calculate the median... median The median is the median value of a dataset. Compared to the mean (which is easily affected by extreme values), it more stably represents the "typical value" of the dataset. For example, in Hypertext Transfer Protocol (HTTP) request response time data, if there are a few exceptionally long responses (such as 900ms), the mean will be inflated, but the median still reflects the level of most normal requests. Calculate the median for each data point. The absolute deviation is used to obtain the median absolute deviation (MAD). For example, for a response time dataset of 120ms, 118ms, 115ms, 900ms, and 122ms, the median is 118ms, and the absolute deviations at each point are 2, 0, 3, 782, 42, 0, 3, 782, 4, with a MAD of 3ms. Compared to the standard deviation (which is sensitive to outliers), MAD is more robust to extreme values.

[0021] Furthermore, the degree of deviation for each data point is calculated using the following formula:

[0022]

[0023] In the above formula, Z' i Represents the i-th data point x i The degree of deviation from the median, Z' is the median, and MAD is the median absolute deviation. The coefficient 0.6745 is used to calibrate MAD to the standard deviation units of a standard normal distribution. i If the value is greater than 3.5, it is considered an outlier and the data point is removed.

[0024] Thus, by cleaning outliers, noise interference can be removed from the data source. After cleaning, regular attacks can be identified more accurately. Data cleaning can improve the robustness of the model, making subsequent feature extraction and model training more focused on real threats. It can also improve the quality of subsequent data fusion, thereby ensuring data consistency, significantly reducing false alarms and improving detection efficiency.

[0025] Furthermore, embodiments of the present invention also extract and fuse the temporal and spatial features of network data. As an optional embodiment of the present invention, extracting and fusing the temporal and spatial features of network data to obtain a high-dimensional feature vector includes: extracting the temporal features of network data using a sliding window Long Short-Term Memory (LSTM) encoder; determining the spatial features of network data based on the PageRank value of the network node where the network data resides; determining the attention weights of the network data based on the temporal and spatial features; and fusing the attention weights, temporal features, and spatial features to obtain a high-dimensional feature vector.

[0026] Specifically, this embodiment of the invention uses a sliding window LSTM encoder to capture the timing pattern of network traffic data in network data, specifically using the following formula:

[0027] h t =LSTM(x t-k:t W h )

[0028] Q i =W Q ·h t

[0029] In the above formula, Q i Let h be the temporal characteristics of the network traffic data at time i; t The hidden state of the LSTM at time step t captures the temporal dependencies; W Q for h t Weight matrix mapped to the query space (dimension: 128×64); x t-k:t To output a time-series data window from time step tk to t (e.g., traffic data over the past 10 seconds), W h The weight parameters of the LSTM model are obtained through training with historical data.

[0030] For example, suppose the input data acquisition window is an HTTP request frequency of 10 time steps: 10, 12, 15, 20, 18, 22, 25, 30, 28, 35 (unit: times / second). After LSTM processing, the output h t = [0.7, -0.3, ..., 0.5] (a 128-dimensional vector), through linear transformation W Q , get Q i =

[0031] [0.8, 0.2, ..., 0.6] (64 dimensions).

[0032] Thus, by calculating the time series feature Q i It can capture traffic fluctuations, thereby capturing dynamic changes in attack behavior. For example, a sudden surge in HTTP requests may indicate a DDoS attack, and the memory cells of LSTM can identify periodic attack patterns (such as timed scans).

[0033] Furthermore, embodiments of the present invention use the following formula to calculate the spatial characteristics of network asset data in network data:

[0034]

[0035] In the above formula, K jLet PR(u) represent the spatial characteristics of the j-th network asset data, where PR(u) is the PageRank value of node u, reflecting its importance in the topology, N is the total number of asset nodes (e.g., 200 servers), d is the damping coefficient, simulating the probability of a user randomly switching, with a default value of d = 0.85, PR(v) is the PageRank value of node v, L(v) is the number of outgoing links of node v (e.g., the number of external connections of a server), and B... u W is the set of all nodes pointing to node u; K The weight matrix (dimension: 72×64) maps PageRank values ​​to the key space.

[0036] For example, in the PageRank calculation of the core database server (node ​​A), N = 200, d = 0.85, assuming 10 application servers (nodes B1-B10) point to A, and each node has L(v) = 2, if each node B has a PageRank of 2... i If ) = 0.1, then Through W K , obtain K j = [0.6, -0.1, ..., 0.4] (64 dimensions).

[0037] Thus, by calculating spatial features K j The purpose and role of assessing the importance of network assets include quantifying the importance of network assets (such as servers and terminal devices) in the topology, focusing on the security risks of critical nodes, identifying critical nodes (such as core servers and devices with a high number of connections), and identifying nodes that, once compromised, may cause a wider range of security risks. Nodes with high PageRank values ​​(such as core databases) are prioritized for monitoring. Attack path prediction can also be performed: attackers may move laterally through high PR nodes, which can assist in vulnerability correlation analysis.

[0038] Furthermore, assets can be prioritized, with assets having higher PageRank values ​​receiving greater weight in vulnerability detection and receiving priority in deep scanning. Attack path prediction: Attackers tend to move laterally through high-importance nodes, and PageRank values ​​can help predict potential attack paths. Relationship with vulnerability identification: Asset importance scores (PageRank values) serve as feature inputs to the vulnerability identification engine, enabling the model to focus more on the abnormal behavior of critical assets and improve the detection rate of high-risk vulnerabilities.

[0039] Furthermore, embodiments of the present invention integrate temporal and spatial features to achieve deep fusion of network traffic temporal features and asset topology spatial features, specifically using the following formula:

[0040]

[0041] In the above formula, Aij Representing the temporal features Q i With spatial features K j The attention weights after fusion are more critical for vulnerability detection if the weights are higher. i is the index of the temporal feature (corresponding to the temporal feature extracted by LSTM), j is the index of the spatial feature (corresponding to the spatial feature extracted by PageRank), and d is the feature dimension, which is 64 by default. Calculate the time series feature Q i With spatial features K j Similarity, scaling factor To prevent the gradient from vanishing due to an excessively large inner product, ∑ j The upper bound is the index of all spatial features (i.e., traversing all spatial features). exp() represents the normalization function, used for normalization.

[0042] For example, suppose Q i = [0.8, 0.2], K j = [0.6, 0.4], then Output timing feature Q i With spatial features K j The fused high-dimensional feature vector is V = A ij ·(Q i +K j ).

[0043] Thus, by calculating A ij Dynamic feature weighting is performed, with high attention weights (such as A). ij =0.8) indicates that the combination of temporal and spatial features is more critical for vulnerability detection. Therefore, the high attention weights are allocated to improve the F1-score of the fused features, thereby improving detection accuracy. In addition, it can adapt to different network architectures (such as cloud environments and OT networks), enhancing generalization ability. Temporal features (LSTM output) reflect changes in behavioral patterns, while spatial features (PageRank value) reflect topological importance. After fusion, the vulnerability features are comprehensively described, the complementarity is enhanced, key features are selected, and the dimensionality is compressed from the original 1000+ to ≤200, achieving dimensionality reduction, purification and optimization.

[0044] Step S102: Determine the type and confidence level of the network vulnerability based on the high-dimensional feature vector.

[0045] Specifically, embodiments of the present invention output the type and confidence level of a network vulnerability through a hybrid model. As an optional embodiment of the present invention, determining the type and confidence level of a network vulnerability based on a high-dimensional feature vector includes: inputting the network asset association graph corresponding to the high-dimensional feature vector into a first vulnerability identification model to obtain a first output; inputting the high-dimensional feature vector into a second vulnerability identification model to obtain a second output; weighting the first and second outputs to obtain a score for each type of network vulnerability; selecting the type corresponding to the maximum score as the type of network vulnerability; mapping the scores of each type of network vulnerability to a probability distribution; and selecting the network vulnerability with the highest probability as the confidence level of the type.

[0046] Specifically, the hybrid model in this embodiment of the invention can be composed of a first vulnerability identification model and a second vulnerability identification model. The first vulnerability identification model can be a Graph Neural Network (GNN) model and an Extreme Gradient Boosting (XGBoost) model. GNN analyzes the network asset association graph (capturing topological features), and XGBoost processes statistical features (such as traffic frequency and log anomaly count). The two are weighted and fused to output the vulnerability type (such as SQL injection, XSS, etc.). Finally, the type corresponding to the highest probability is taken as the type of network vulnerability and the confidence level is determined. The network vulnerability data with high confidence, low confidence, and unidentified vulnerabilities are then identified.

[0047] Furthermore, the system can identify the following vulnerability types: SQL injection (SQLi), cross-site scripting (XSS), buffer overflow, privilege escalation, remote code execution (RCE), zero-day vulnerabilities (unknown attack patterns but unusual characteristics), and other (cannot be clearly categorized but pose a threat). The detection targets for vulnerability identification include network traffic data, system logs, and asset configuration data. Detection accuracy is improved by weighted fusion of the outputs of GNN and XGBoost.

[0048] For the GNN model, high-dimensional feature vectors are input into the GNN model, and the formula GNN(G) = GraphConv(h) is used. (l) W G ) Analyze the network asset association graph (node ​​= asset, edge = interaction relationship) to capture topological features; where G is the asset association graph (node ​​= asset, edge = interaction relationship), h(l) is the hidden state of the l-th layer of GNN (capturing topological features), W GThis is the weight matrix of the GNN. Then, using the formula... f k Decision trees process network traffic data, such as traffic frequency and log anomaly counts, where F is the feature vector (e.g., structured features of traffic and logs), f k Let K be the k-th decision tree, and K be the total number of decision trees.

[0049] Finally, the outputs of the two models are combined using the following formula:

[0050] y=α·GNN(G)+(1-α)·XGBoost(F)

[0051] In the above formula, y is the predicted result of the vulnerability type after fusion calculation, which represents the score of each type of vulnerability (such as SQL injection, buffer overflow). G is the asset association graph (nodes are assets, edges are interaction relationships), F is the feature vector, and α is the weight coefficient. Through cross-validation optimization, α = 0.6. The final classification takes the vulnerability type corresponding to the maximum score.

[0052] Furthermore, the confidence level is calculated using the following formula:

[0053] θ c =max(p1, p2, ..., p i )

[0054] In the above formula, θ c θ represents the confidence level, where p1 is the predicted probability of the hybrid model for the i-th type of vulnerability (obtained by normalizing the model output score using Softmax), and θ c ∈[0,1], the confidence threshold is used for the classification result (high / low confidence).

[0055] Furthermore, embodiments of the present invention can also calculate the loss value using the following formula:

[0056]

[0057] In the above formula, The value represents the loss, used to measure the difference between the vulnerability classification prediction result and the actual result, where p t Let γ be the predicted probability of the model for the true class, and let γ be an adjustment factor to reduce the weight of easily classified samples; γ = 2. The upper and lower bounds of the summation are calculated by iterating through all samples (x, y), i.e., summing the losses for each sample in the training set. Model training is performed through loss calculation to improve the accuracy of confidence judgments and address the class imbalance problem.

[0058] Furthermore, in the confidence calculation: the hybrid model outputs GNN to extract asset association features and XGBoost to analyze statistical features, and the two are weighted (α = 0.6) to generate a comprehensive score; Softmax normalization maps the score to a probability distribution, and takes the maximum value as the confidence score (θc = max(p1, p2, ..., pn)). In vulnerability type judgment: rule matching: known vulnerabilities (such as CVE numbers) are directly matched to the rule base (such as Snort rules). AI inference: for novel vulnerabilities (such as zero-day attacks), they are clustered based on feature similarity to be classified into the closest vulnerability type.

[0059] Step S103: Determine the level to which the network vulnerability belongs based on its type and confidence level.

[0060] The hierarchy includes a first level, a second level, and a third level. The first and second levels indicate that the network vulnerability is identified, while the third level indicates that the network vulnerability is not identified. The confidence level of the first level is greater than that of the second level, and the confidence level of the second level is greater than that of the third level.

[0061] Specifically, after calculating the confidence levels for various types of network vulnerabilities, θ is... c Network vulnerabilities with a value ≥0.85 are classified as Level 1, and vulnerabilities with a value ≤0.4 are classified as Level 2. c Network vulnerabilities with a vulnerability value <0.85 were classified as Level 2, and θ was... c Network vulnerabilities with a vulnerability size of <0.4 are classified as Level 3.

[0062] Furthermore, after determining the hierarchy, as an optional embodiment of the present invention, before performing authenticity verification and correlation analysis on the first-level and second-level network vulnerabilities to determine the authenticity and attack chain of the network vulnerabilities, the method further includes: scoring the network vulnerabilities at the first level using a general vulnerability scoring system to obtain a score, the score indicating the severity of the network vulnerability; determining the vulnerability risk of the first-level network vulnerabilities based on the asset value, environmental threat coefficient, patch status, traffic burst coefficient, and severity of the actual network state corresponding to the first-level network vulnerabilities; determining the risk level of the first-level network vulnerabilities based on the vulnerability risk; implementing a handling strategy corresponding to the risk level for the network vulnerabilities based on the risk level, and, when the risk level is high, performing authenticity verification and correlation analysis on the first-level and second-level network vulnerabilities to determine the authenticity and attack chain of the network vulnerabilities includes: performing authenticity verification and correlation analysis on high-risk level network vulnerabilities in the first-level network vulnerabilities and second-level network vulnerabilities to determine the authenticity and attack chain of the network vulnerabilities, where the risk level of the high-risk level is greater than a preset level.

[0063] Specifically, this embodiment of the invention performs risk assessment on first-level network vulnerabilities, dynamically quantifies the actual threat of vulnerabilities by combining static vulnerability scores with real-time network status, and implements tiered responses. Specifically, this embodiment utilizes the Common Vulnerability Scoring System (CVSS) for scoring, which measures the severity of vulnerabilities (such as attack complexity and scope of impact); vulnerability types include SQL injection and remote code execution (RCE), with different types of vulnerabilities posing different potential harms. Then, the asset value A is determined based on the contextual information of the network environment. v Pre-set scores based on business importance (e.g., core database = 10, test server = 2); Environmental threat coefficient E c Based on real-time threat data (such as the number of intrusion attempts and scanning activity frequency in the past hour); Patch status T p : Vulnerability patch release time (in hours), reflecting the exposure risk caused by patch delays; Traffic burst factor B t Current network traffic fluctuations are used to detect potential attacks.

[0064] Finally, risk quantification is performed. First, a multi-dimensional weighted fusion is conducted, and the CVSS score provides a basic threat level, but this needs to be adjusted based on environmental factors. For example, a vulnerability might have a CVSS of 7.5 (high risk), but if the asset value is low (A... v =2) and no recent attacks (E) c =1), the actual risk may only be medium; then nonlinear relationship modeling: using the formula Capture the synergy between asset value and environmental threats. For example, high-value assets (A v =9) Encountering frequent attacks (E c When T = 4), the risk increases exponentially. p +1) This reduces the linear impact of patch time, better reflecting real-world security operations (e.g., risk increases rapidly in the first 72 hours after a patch delay, then slows down). Finally, the vulnerability risk is calculated using the following formula:

[0065]

[0066] Here, Risk represents the vulnerability risk, CVSS is the general vulnerability scoring system score (0-10), obtained from a vulnerability database, A v The asset value is scored (0-10), preset by the user (e.g., core servers = 10), E c The environmental threat factor (0-5) is calculated based on the number of real-time intrusion attempts, T. p The patch release time (in hours) is obtained from the vendor's announcement. t This represents the real-time traffic burst coefficient.

[0067] After identifying the vulnerability risks, the results are categorized into 5 levels, each corresponding to a different handling strategy:

[0068] Risk ≥ 8.0, Risk Level 5 (Emergency), immediately isolate assets, automatically deploy patches or virtual patches, and provide multi-channel alerts (SMS, email, dashboard).

[0069] Risk: 6.0~7.9, Risk level 4 (high risk), rate limiting (bandwidth / connection count reduction), generate repair ticket (processed within 24 hours).

[0070] Risk: 3.0~5.9, risk level 3 (medium risk), log entries are recorded and marked "requires attention", and a risk report is generated weekly.

[0071] Risk < 3.0, risk level 1-2 (low risk), only log entries are recorded, and periodic sampling reviews are conducted (e.g., 5% per month).

[0072] Based on Risk results, implement precise tiered response and optimized resource allocation:

[0073] Precise Tiered Response: Transforming vulnerability threats from abstract scores (such as CVSS) into actionable risk levels (1-5) to guide differentiated handling. Example: A certain SQL injection vulnerability (CVSS=8.0) in the core database (A v A burst of flow (B) was detected on (=10). t =0.6), Risk=9.2, Level 5, triggering immediate isolation; the same vulnerability exists on edge device (A v =2) On, Risk=3.1, Level 2, log only.

[0074] Resource optimization: Avoid over-responding to low-risk vulnerabilities (such as mistakenly isolating normal business operations) to save security team processing time; concentrate resources on handling high-risk vulnerabilities (such as zero-day attacks) to shorten the average remediation time.

[0075] Thus, by integrating vulnerability attributes, asset value, environmental threats, and real-time status, risk quantification and tiered response are achieved, overcoming the shortcomings of traditional assessment methods that are static and detached from the actual environment, thus providing accuracy. Furthermore, it avoids over-responding to low-risk vulnerabilities while ensuring rapid and adaptable handling of high-risk vulnerabilities. Real-time adjustment of risk levels addresses the dynamic changes in cyberattacks, and automation, along with linkage with the response system, shortens threat exposure time.

[0076] Furthermore, in this embodiment of the invention, a high-confidence vulnerability (θ) c ≥0.85): Output to the dynamic risk assessment module; secondary verification is triggered only when the risk assessment is high risk (≥4 level) (e.g., vulnerability associated with core assets); Example: If SQL injection vulnerability (θ)c If the risk level is assessed as 5 (=0.92), then the feasibility of its attack path needs to be verified a second time.

[0077] Low confidence vulnerability (0.4≤θ) c <0.85): Directly triggers secondary verification without waiting for risk assessment. The verification results are divided into two categories: confirmed vulnerability: update the vulnerability database and reassess the risk; false positive: add to the training set to optimize the model.

[0078] Unidentified Vulnerability (θ) c <0.4): Stored independently in the unidentified vulnerability handling module, retaining the original data and characteristics.

[0079] The default level can be 4.

[0080] Step S104: Perform authenticity verification and correlation analysis on the first-level and second-level network vulnerabilities to determine the authenticity of the network vulnerabilities and the attack chain.

[0081] Specifically, this embodiment of the invention re-verifies network vulnerabilities at the first and second levels to determine their authenticity and attack chains. If a vulnerability is identified, it is confirmed, and the vulnerability database is updated. If a false positive is identified, the sample is added to the training set of the hybrid model, and the hybrid model is optimized after the false positive sample is added to the training set. Causal reasoning is used to verify whether the vulnerability can be actually exploited (e.g., excluding false attacks blocked by firewalls) and to confirm whether the vulnerability is associated with other identified vulnerabilities in an attack chain. Authenticity verification excludes false associations through causal reasoning, while association analysis determines whether an attack chain exists (e.g., "SQL injection → privilege escalation → lateral movement"), achieved by analyzing the dependencies between vulnerabilities (whether the exploitation of a previous vulnerability creates conditions for the exploitation of a subsequent vulnerability). A genuine vulnerability is defined as one that satisfies a causal reasoning probability ≥ 0.7 and is verifiable as exploitable through adversarial testing.

[0082] Furthermore, as an optional embodiment of the present invention, the authenticity verification and correlation analysis of the first-level and second-level network vulnerabilities to determine the authenticity of the network vulnerabilities and the attack chain include: intervening in the high-dimensional feature vectors of the first-level and second-level network vulnerabilities to obtain values, and calculating the conditional probability distribution of successful attacks on the first-level and second-level network vulnerabilities after the intervention of the high-dimensional feature vector values; determining the marginal probability of the network vulnerabilities based on the confounding variables of the network in which the network vulnerabilities are located; determining the true probability of successful attacks on the first-level and second-level network vulnerabilities after intervention based on the conditional probability distribution and the marginal probability, so as to eliminate false correlations between the network vulnerabilities and the confounding variables; determining the authenticity of the network vulnerabilities based on the true probability, and determining the attack chain based on the dependencies between the network vulnerabilities.

[0083] Specifically, this invention employs intervention experiments to remove environmental interference, causal reasoning to eliminate false associations, and adversarial sample testing to verify the authenticity of the vulnerability and the attack chain. Specifically, the vulnerability's characteristic X and successful attack Y may be indirectly related through confounding variables Z (such as network configuration) rather than a direct causal relationship, requiring association determination. This invention eliminates false associations through intervention experiments. For example, the attack is reproduced after forcibly disabling the firewall to verify the vulnerability's existence and eliminate environmental interference (such as false vulnerabilities caused by firewall misinterpretations). The true probability of the network vulnerability is calculated using the following formula:

[0084]

[0085] In the above formula, P(Y|do(X=x)) represents the probability P(Y|do(X=x)) of Y taking a specific value when X is intervened to take the value x. This is used for causal inference to eliminate spurious associations. do(X=x) represents the intervention operation on variable X; Z represents confounding variables (such as network topology, asset value); P(Y|X=x, Z=z) represents the conditional probability distribution; P(Z=z) represents the marginal probability of confounding variable Z taking the value z. By controlling confounding variables and removing environmental interference, this function calculates the "true probability of Y after intervention X" and eliminates spurious associations (such as false vulnerabilities caused by firewall interception).

[0086] After calculating the true probability, if P(Y|do(X=x))≥0.7, it is determined to be a real vulnerability, the vulnerability database is updated, and risk assessment and dynamic response (such as asset isolation) are triggered; if 0.3≤P(Y|do(X=x))<0.7, further analysis is required, combined with adversarial testing for verification, and if it is still uncertain, manual review is required.

[0087] If P(Y|do(X=x))<0.3, it is judged as a false association, marked as a false alarm, added to the training set to optimize the model, and the feature engineering rules are adjusted (such as ignoring specific firewall interference).

[0088] For example, suppose the original detection probability of a vulnerability is P(Y|X=x)=0.7, but there is a confounding factor Z in the environment (such as firewall blocking). Intervention is performed: forcibly disable the firewall (do(X=x)), severing the association between X and Z. The probability is calculated as P(Y|do(X=x)). If in the actual test P(Y|do(X=x))=0.05, which is much lower than the original value of 0.7, it means that the vulnerability was misjudged (actually blocked by the firewall); if the vulnerability can still be triggered after disabling the firewall (do(firewall=off)), it is determined to be a real vulnerability.

[0089] Furthermore, in adversarial example testing, WGAN-GAN can be used to generate adversarial traffic (such as modifying payload offsets or SQL injection statement structures) to test model robustness. Environment simulation: Attack scenarios are reproduced in a sandbox to verify vulnerability exploitability and confirm whether the vulnerability can be exploited by actual attackers, rather than being a theoretical risk. After verification, for confirmed real vulnerabilities, the vulnerability database is updated, alerts are triggered, and they are included in risk assessments. For false positives: they are added to the training set to optimize the hybrid model, and the reasons for false positives (such as interference from specific firewall rules) are recorded to reduce subsequent false positive rates. By re-identifying first- and second-level network vulnerabilities and removing environmental interference through intervention experiments, real vulnerabilities can be accurately identified.

[0090] Step S105: Re-identify the third-level network vulnerabilities and label the type of the third-level network vulnerabilities based on the identification results.

[0091] Specifically, for unidentified Level 3 network vulnerabilities, this embodiment of the invention can independently store, actively learn and sample, and incrementally train them before re-identifying the vulnerabilities. This embodiment of the invention generates a dataset of unidentified Level 3 network vulnerabilities labeled as Suspected Unclassified Vulnerabilities (SUVs), and selects high-value samples to update the parameters of the hybrid model.

[0092] Furthermore, as an optional embodiment of the present invention, before re-identifying the third-level network vulnerabilities and labeling the type of the third-level network vulnerabilities according to the identification results, the method further includes: determining the retention period for the third-level network vulnerabilities based on a preset minimum retention period and the number of third-level network vulnerabilities; and structuring the network data corresponding to the third-level network vulnerabilities and retaining the retention period in an independent database.

[0093] Specifically, in this embodiment of the invention, the independent database can be MySQL or MongoDB, and stored in shards according to timestamps to maintain data correlation. The retention days can be calculated using the following formula:

[0094]

[0095] In the above formula, T 保留 N represents the data retention period (in days), where d is the number of days, 30d is the minimum retention period, and N is the minimum retention period. suv This represents the total number of unidentified vulnerabilities in the current SUV. Example: If there are currently 100 unidentified vulnerabilities (N... suv =100), then the retention time is

[0096] In this way, by storing data independently and dynamically adjusting it periodically, we can prevent the direct discarding of potentially threatening data, ensure subsequent detection opportunities, and avoid false negatives. Data is retained for model iteration and optimization to improve future detection capabilities and support incremental learning.

[0097] Furthermore, as an optional embodiment of the present invention, the third-level network vulnerability is identified again, and the type of the third-level network vulnerability is labeled according to the identification result, including: inputting the high-dimensional feature vector of the third-level network vulnerability into multiple detection models, and using multiple detection models to output the prediction result of the type of the third-level network vulnerability; determining the mean difference of the high-dimensional feature vector of the third-level network vulnerability based on the difference between the prediction results of each pair of detection models; selecting the high-dimensional feature vector corresponding to the maximum value of the mean difference as the sample of the real vulnerability, and labeling the type of the sample of the real vulnerability.

[0098] Specifically, in this embodiment of the invention, for low-confidence samples, the difference in prediction results from different models is calculated, and the sample with the most information is selected to achieve the screening of high-value samples. Specifically, the mean difference of the high-dimensional feature vectors of third-level network vulnerabilities is calculated using the following formula:

[0099]

[0100] In the above formula, MMD(x) represents the mean difference of the high-dimensional feature vector of the x-th network vulnerability in the input, K is the number of all detection models in the model set, and P i (y|x) represents the prediction result of the i-th detection model for the input x, P j (y|x) is the prediction result of the j-th detection model for the input sample x. The sample with the highest MMD value (largest difference), i.e., the sample with the most information, is selected as the high-value sample of the real vulnerability. The square of the L2 norm, This represents the quantified value of the "difference" between the prediction results of the two detection models, i and j.

[0101] Furthermore, after screening, samples are manually labeled, and security experts confirm whether they are real vulnerabilities or false positives. Labeled samples are added to the training set for incremental training of the hybrid model, optimizing its parameters. Example: If the predicted probability distributions of five detection models for a sample differ significantly (e.g., model A identifies it as SQL injection, while model B identifies it as normal), then samples with higher MMD values ​​are prioritized for labeling. Using MMD to filter samples improves labeling efficiency: it reduces manual labeling workload, focuses on key samples, prioritizes labeling samples where the model is uncertain, and accelerates model optimization.

[0102] Furthermore, as an optional embodiment of the present invention, after selecting the high-dimensional feature vector corresponding to the maximum value in the mean difference as a sample of real vulnerability and labeling the type of the sample of real vulnerability, the method further includes: using the labeled sample of real vulnerability as a training sample of vulnerability identification model to train the first vulnerability identification model and the second vulnerability identification model.

[0103] Specifically, the importance of each parameter in the hybrid model is calculated using real vulnerability samples, and constraints are updated to limit the variation of important old parameters, thus balancing the detection capabilities of new and old vulnerabilities. This embodiment of the invention calculates the Fisher information content of each parameter based on historical data. When training on new data, penalties are applied to important parameters to limit their deviation from their original values, constraining updates and preventing the forgetting of historical knowledge. This ensures that the model does not lose its ability to detect old vulnerabilities when learning new ones. After the update, the optimized parameters are deployed to the vulnerability identification engine, historical data is re-scanned, and the new model is used to re-detect the unidentified vulnerability database, outputting the updated vulnerability identification model M. t+1 .

[0104] Furthermore, when the model learns new vulnerability knowledge, it protects old knowledge from being overwritten by using the formula:

[0105]

[0106] In the above formula, This represents the conventional loss for the new sample. The prediction error of the hybrid model for high-value samples selected by active learning is used to measure the error of the model, and it is the core optimization objective of incremental training. Where (x, y∈D) 新 ) represents traversing the samples in the new dataset; M(x) is the predicted output of the mixture model for sample x; y is the true label of the sample (manually labeled, such as "SQL injection vulnerability", "false positive", "0-day vulnerability", etc.); The loss function measures the difference between the predicted output M(x) and the true label y. Specifically, cross-entropy loss or mean squared error can be used.

[0107] The prediction result M(x) of the vulnerability identification model (hybrid model: GNN + ensemble learning) for the input feature x includes: the probability distribution of the vulnerability type (e.g., "SQL injection" probability 0.8, "XSS" probability 0.2); and the predicted probability of whether it is a vulnerability (e.g., "is a vulnerability" probability 0.9).

[0108] Specifically through the formula

[0109]

[0110] In the above formula, This represents the total loss, which is used for incremental training with elastic weights, where λ is the elastic constraint coefficient, balancing the learning weights of new and old knowledge. j Let θ be the Fisher information matrix, which represents the Fisher information content of the j-th parameter, measuring the parameter's importance (calculated using historical data). j θ represents the current model parameters. j,旧 Let j be the j-th parameter of the historical model.

[0111] In incremental learning, dynamic adaptation is achieved by adjusting the λ control system's emphasis on new and old knowledge. The larger λ is, the more conservative the model (more reliant on old knowledge), and the smaller λ is, the more aggressive the model (learns new features faster).

[0112] During constraint updates, parameter importance filtering is performed: the F-values ​​of all parameters are calculated. j Set a threshold (e.g., F) j >0.1) Select key parameters; impose gradient update constraints. During training, add penalty terms to the gradients of key parameters to ensure their update direction follows the optimal path of the previous task as closely as possible. The calculation formula is as follows:

[0113]

[0114] In the above formula, The updated parameter value; The parameter values ​​before the update; θ j,旧 The parameter values ​​are historical reference values; η is the learning rate, which controls the update step size. For the loss of the new sample, the parameter θ j The gradient of F; j λ is the Fisher information content of the j-th parameter; λ is the regularization coefficient; 2λF j (θ j -θ j,旧 ) is a penalty term (limiting the range of variation of important parameters); if θ j Trying to move away from the old value θ j,旧 The penalty item will pull it back.

[0115] Furthermore, through the Fisher information matrix F j Parameter θ j The importance of old tasks is calculated as the expected value of the second derivative of the parameters under historical data. The update range of important parameters is fixed by elastic weights. The elastic constraints protect old knowledge and achieve a balance between the detection capabilities of new and old vulnerabilities.

[0116] The technical solution disclosed in this invention extracts and fuses temporal and spatial features of network data to obtain a high-dimensional feature vector, which can comprehensively capture key information in the network data. This fusion avoids the limitations of a single feature dimension, providing a more comprehensive basis for subsequent vulnerability identification and helping to improve the accuracy of vulnerability identification. The type and confidence level of the network vulnerability are determined based on the high-dimensional feature vector, achieving a precise preliminary judgment of the vulnerability. The high-dimensional feature vector contains a large amount of vulnerability-related information, enabling effective identification of new attacks and unknown vulnerabilities, improving the accuracy of identifying various types of network vulnerabilities. The confidence level classifies various vulnerabilities. The first and second levels clearly define vulnerabilities in an identified state, with the first level having a higher confidence level than the second, helping staff prioritize high-confidence vulnerabilities, rationally allocate resources, and improve the efficiency of vulnerability handling. The third level defines unidentified vulnerabilities, avoiding the risk of omission and ensuring that all possible vulnerabilities are addressed. Subsequently, the authenticity verification and correlation analysis of the network vulnerabilities in the first and second levels further ensure the reliability of vulnerability information and uncover attack chains between vulnerabilities. Authenticity verification can eliminate false positives and further improve the accuracy of vulnerability identification; attack chains help determine the overall network security posture, enabling targeted defensive measures to prevent the spread of attacks. Re-identifying and labeling third-level network vulnerabilities improves the completeness and accuracy of vulnerability identification. This avoids situations where vulnerabilities are overlooked due to a single misidentification; by re-identifying, unidentified vulnerabilities are converted to identified status as much as possible, preventing the recurrence of similar attacks and reducing the risk of network attacks.

[0117] For example, this embodiment of the invention uses a financial enterprise cloud platform that processes an average of 10Gbps of network traffic, 2TB of system logs, and has assets of 200+ servers. Its detection requirements include real-time identification of vulnerabilities such as SQL injection and XSS vulnerabilities, dynamic risk assessment, and rapid response.

[0118] First, during the dynamic window adjustment, δ0 = 5 seconds. (Current traffic burst gradient, calculate B) t = 0.8 (rate of change), B avg =0.3 (historical average suddenness coefficient), calculate Output a standardized data stream (data within a 6.67-second time window is packaged in Avro format);

[0119] Then, the data is standardized. Example data points: response times of 10 HTTP requests: 120ms, 118ms, 115ms, 900ms, 122ms, 119ms, 1100ms, 117ms, 121ms, 119ms. Sorted by: 1100, 900, 122, 121, 120, 119, 119, 118, 117, 115. The median is then calculated. The absolute deviations of each point from the median are: 980.5, 780.5, 2.5, 1.5, 0.5, 0.5, 0.5, 1.5, 2.5, 4.5. The median deviation (MAD) is calculated as 1.5 ms. For outlier points x... i =Calculation is performed within 1100ms. Remove outlier data and output: cleaned data (remove 1100ms and 900ms, keep the remaining 8 points).

[0120] Secondly, when extracting time-series features from the cleaned data (HTTP request response time, log status codes, asset topology), h is calculated. t =LSTM(x t-k:t W h ), (k=10, W h The weights obtained during training are used to output a 128-dimensional temporal feature vector. During spatial feature extraction, the asset topology map is used, with the core database server (node ​​A) connecting to 10 application servers (nodes B1-B10). Each B node has a PR value of 0.1, and the number of outgoing chains L = 2. The calculation... Temporal features Q i =0.8, spatial feature K J =0.6, perform feature fusion, calculate Output the fused 200-dimensional feature vector (128 temporal + 72 spatial).

[0121] Next, the vulnerabilities are identified. The GNN branch takes the asset association graph (node ​​= server, edge = interaction relationship) as input and outputs graph features G = [0.7, 0.3, ..., 0.5] (64 dimensions). The XGBoost branch takes the statistical features F = [120ms, 200 logins, ...] (136 dimensions) as input and outputs the weighted result f of the tree model. k (F) = 0.4 (assuming single-tree output), weighted output, calculate y = α·GNN(G) + (1-α)·XGBoost(F) = 0.6 × 0.7 + 0.4 × 0.4 = 0.58, Softmax output probabilities: 0.58 (SQLi), 0.2 (XSS), ..., 0.02 (others), calculate confidence θ. c =0.58, classified as a low-confidence vulnerability (SUV);

[0122] Conduct a risk assessment for high-confidence network vulnerabilities: Input vulnerability information (SQLi, CVSS=7.5), asset value A v =8, Environmental threat E c =3, patch time T p =48h, traffic burst B t =4, calculate risk The output shows a risk level of 5 (urgent), triggering asset isolation and automatic repair.

[0123] Next, verify the low-confidence SQL injection vulnerability by forcibly modifying the firewall configuration through causal reasoning, check if the vulnerability still exists, generate mutated SQL statements through adversarial testing, check if the model still identifies it as a vulnerability, and if the vulnerability still exists, confirm it as a real vulnerability and update the vulnerability database.

[0124] Finally, unidentified vulnerabilities were addressed by storing low-confidence data in MySQL for 30 days, actively learning and sampling, and calculating...

[0125] Samples with MMD > 0.3 were manually labeled and incremental training was performed to calculate...

[0126] Update model parameters to prevent historical knowledge from being forgotten.

[0127] Based on the same inventive concept, embodiments of the present invention also provide a network vulnerability identification system based on data analysis, such as... Figure 2 As shown, Figure 2This is a schematic diagram of a network vulnerability identification system based on data analysis provided in an embodiment of the present invention. The network vulnerability identification system based on data analysis includes: an extraction module 201, used to extract and fuse the temporal and spatial features of network data to obtain a high-dimensional feature vector; a determination module 202, used to determine the type and confidence level of the network vulnerability based on the high-dimensional feature vector; the determination module 202 is further used to determine the level to which the network vulnerability belongs based on the type and confidence level of the network vulnerability, the level including a first level, a second level and a third level, the first level and the second level indicating that the network vulnerability is identified, the third level indicating that the network vulnerability is unidentified, the confidence level of the first level is greater than the confidence level of the second level, and the confidence level of the second level is greater than the confidence level of the third level; an analysis module 203, used to perform authenticity verification and correlation analysis on the network vulnerabilities of the first level and the second level to determine the authenticity and attack chain of the network vulnerability; and an identification module 204, used to re-identify the network vulnerability of the third level and label the type to which the network vulnerability of the third level belongs based on the identification result.

[0128] It should be noted that the network vulnerability identification system based on data analysis provided in this embodiment of the invention and the network vulnerability identification method based on data analysis in the above embodiments are based on the same application concept. The same or similar aspects of the specific implementation of the two embodiments can be referred to each other and have the same or similar beneficial effects. The repeated parts will not be described again.

[0129] Corresponding to the method provided in the above embodiments, and based on the same technical concept, this embodiment of the invention also provides an electronic device for executing the above-described network vulnerability identification method based on data analysis. Figure 3 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present invention, such as... Figure 3 As shown. Electronic devices can vary considerably due to differences in configuration or performance, and may include one or more processors 301 and memories 302. The memory 302 stores computer programs that can run on the processor 301, and the processor 301 executes the programs stored in the memory 302 to achieve the above. Figure 1 The various steps in the method embodiment are described. The memory 302 can be temporary or persistent storage. The application stored in the memory 302 may include one or more modules (not shown), each module may include a series of computer-executable instructions for the electronic device.

[0130] Furthermore, the processor 301 may be configured to communicate with the memory 302 and execute a series of computer-executable instructions stored in the memory 302 on the electronic device. The electronic device may also include one or more power supplies 303, one or more wired or wireless network interfaces 304, one or more input / output interfaces 305, and one or more keyboards 306.

[0131] Specifically, in this embodiment, the electronic device includes a processor, a communication interface, a memory, and a communication bus; wherein, the processor, the communication interface, and the memory communicate with each other via the bus; the memory is used to store computer programs; and the processor is used to execute the programs stored in the memory to achieve the above. Figure 1 The various steps in the method embodiments are the same as those in the above method embodiments, and have the same beneficial effects. To avoid repetition, the embodiments of the present invention will not be described again here.

[0132] Furthermore, based on the same inventive concept, embodiments of the present invention provide a computer storage medium on which a computer program is stored, and when the computer program is executed by a processor, it implements the steps of the network vulnerability identification method based on data analysis as described in the above method implementation.

[0133] It should be noted that the order of the above embodiments of the present invention is merely for descriptive purposes and does not represent the superiority or inferiority of the embodiments. The processes depicted in the accompanying drawings do not necessarily require a specific or sequential order to achieve the desired result. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0134] The various embodiments in this specification are described in a progressive manner. The same or similar parts between the various embodiments can be referred to each other. Each embodiment focuses on describing the differences from other embodiments.

Claims

1. A network vulnerability identification method based on data analysis, characterized in that, The data analysis-based network vulnerability identification method includes: Extract and fuse the temporal and spatial features of the network data to obtain a high-dimensional feature vector; The type and confidence level of the network vulnerability are determined based on the high-dimensional feature vector. The network vulnerability is determined to be at a certain level based on its type and confidence level. The levels include a first level, a second level, and a third level. The first level and the second level indicate that the network vulnerability is identified, and the third level indicates that the network vulnerability is unidentified. The confidence level of the first level is greater than that of the second level, and the confidence level of the second level is greater than that of the third level. The authenticity of the network vulnerabilities at the first and second levels is verified and correlation analysis is performed to determine the authenticity of the network vulnerabilities and the attack chain. The network vulnerabilities at the third level are identified again, and their types are labeled based on the identification results.

2. The network vulnerability identification method based on data analysis according to claim 1, characterized in that, Before performing authenticity verification and correlation analysis on the first-level and second-level network vulnerabilities to determine the authenticity and attack chain of the network vulnerabilities, the method further includes: A general vulnerability scoring system is used to score network vulnerabilities at the first level, and a score is obtained, which indicates the severity of the network vulnerability. The vulnerability risk of the first-level network vulnerability is determined based on the asset value, environmental threat coefficient, patch status, traffic burst coefficient, and severity of the actual network state corresponding to the first-level network vulnerability. The risk level of the first-level network vulnerability is determined based on the described vulnerability risk. Based on the risk level, a corresponding handling strategy is implemented for the network vulnerability. When the risk level is high, the verification and correlation analysis of the authenticity and attack chain of the first and second level network vulnerabilities are performed, including: The authenticity and correlation analysis of high-risk network vulnerabilities in the first level of network vulnerabilities and network vulnerabilities in the second level are performed to determine the authenticity and attack chain of the network vulnerabilities, wherein the risk level of the high-risk level is greater than the preset level.

3. The network vulnerability identification method based on data analysis according to claim 1, characterized in that, Before re-identifying the third-level network vulnerability and labeling its type based on the identification results, the method further includes: The retention period for the third-level network vulnerabilities is determined based on the preset minimum retention period and the number of third-level network vulnerabilities. The network data corresponding to the third-level network vulnerability is structured and then stored in an independent database for the specified number of retention days.

4. The network vulnerability identification method based on data analysis according to claim 1, characterized in that, The step of determining the type and confidence level of a network vulnerability based on the high-dimensional feature vector includes: The network asset association graph of the network corresponding to the high-dimensional feature vector is input into the first vulnerability identification model to obtain the first output, and the high-dimensional feature vector is input into the second vulnerability identification model to obtain the second output. The first output and the second output are weighted to obtain the score of each type of network vulnerability. The type corresponding to the maximum score is selected as the type of the network vulnerability, and the scores of the network vulnerability belonging to each type are mapped to a probability distribution. The maximum probability is selected as the confidence level of the network vulnerability of that type.

5. The network vulnerability identification method based on data analysis according to claim 1, characterized in that, The process of verifying the authenticity and performing correlation analysis on the network vulnerabilities at the first and second levels to determine the authenticity and attack chain of the network vulnerabilities includes: Intervene and extract values ​​from the high-dimensional feature vectors of the first-level and second-level network vulnerabilities, and calculate the conditional probability distribution of successful attacks on the first-level and second-level network vulnerabilities after the intervention and extraction of the high-dimensional feature vectors. The marginal probability of the network vulnerability is determined based on the confounding variables of the network in which the network vulnerability is located. Based on the conditional probability distribution and the marginal probability, determine the true probability of successful attack when attacking the network vulnerability after intervention, for the first-level and second-level network vulnerabilities, in order to eliminate the false association between the network vulnerability and the confounding variable. The authenticity of the network vulnerability is determined based on the true probability, and the attack chain is determined based on the dependency relationship between the network vulnerabilities.

6. The network vulnerability identification method based on data analysis according to claim 1, characterized in that, The process of re-identifying the third-level network vulnerabilities and labeling their type based on the identification results includes: The high-dimensional feature vector of the third-level network vulnerability is input into multiple detection models, and the prediction results of the type of the third-level network vulnerability are output using the multiple detection models. The mean difference of the high-dimensional feature vectors of the third-level network vulnerability is determined based on the difference between the prediction results of the pairwise detection models. The high-dimensional feature vector corresponding to the maximum value among the mean differences is selected as a sample of real vulnerabilities, and the type of the sample of real vulnerabilities is labeled.

7. The network vulnerability identification method based on data analysis according to claim 6, characterized in that, After selecting the high-dimensional feature vector corresponding to the maximum value among the mean differences as the sample of the real vulnerability, and labeling the type of the sample of the real vulnerability, the method further includes: The labeled real vulnerability samples are used as training samples for the vulnerability identification model to train the first vulnerability identification model and the second vulnerability identification model.

8. The network vulnerability identification method based on data analysis according to claim 1, characterized in that, The process of extracting and fusing the temporal and spatial features of the network data to obtain a high-dimensional feature vector includes: The temporal features of the network data were extracted using a sliding window long short-term memory network encoder. The spatial characteristics of the network data are determined based on the PageRank value of the network node where the network data is located; The attention weights of the network data are determined based on the temporal and spatial characteristics. The attention weights, temporal features, and spatial features are fused to obtain the high-dimensional feature vector.

9. A network vulnerability identification system based on data analysis, characterized in that, include: The extraction module is used to extract and fuse the temporal and spatial features of network data to obtain a high-dimensional feature vector. The determination module is used to determine the type and confidence level of the network vulnerability based on the high-dimensional feature vector; The determining module is further configured to determine the level to which the network vulnerability belongs based on the type of the network vulnerability and its confidence level. The level includes a first level, a second level, and a third level. The first level and the second level indicate that the network vulnerability is identified, and the third level indicates that the network vulnerability is unidentified. The confidence level of the first level is greater than the confidence level of the second level, and the confidence level of the second level is greater than the confidence level of the third level. The analysis module is used to perform authenticity verification and correlation analysis on the network vulnerabilities of the first level and the second level, so as to determine the authenticity of the network vulnerabilities and the attack chain; The identification module is used to re-identify the third-level network vulnerabilities and label the type of the third-level network vulnerabilities based on the identification results.

10. An electronic device, characterized in that, include: Processor and memory; wherein the memory is used to store computer programs that can run on the processor; A processor for executing a program stored in memory to implement the steps of the network vulnerability identification method based on data analysis as described in any one of claims 1-8.

Citation Information

Cited By

  • AI agent early warning method and device oriented to 0day vulnerability

    CN121125351A

  • An ai agent early warning method and device for 0day vulnerabilities

    CN121125351B

  • Intelligent contract vulnerability detection and auditing method and system based on large language model

    CN121302379A

  • Intelligent contract vulnerability detection and auditing method and system based on large language model

    CN121302379B